Using Tools & Agents
What this means
Measures how well the model handles multi-step conversations and correctly operates external tools and systems to complete a task.
Use case
Power citizen-service agents that work through requests end-to-end, like retrieving a record, looking up eligibility criteria, or filling in a form.
Interactive chart. Arrow keys left and right move between benchmarks, up and down between models. Enter highlights a model, Escape clears the selection.
| Benchmark | Kolibri | Qwen3.6 35B-A3B | Mistral Small 4 119B-A6B | Nemotron 3 Super 120B-A12B |
|---|---|---|---|---|
| Tau2-Bench (Telecom) | 94.7% | 99.1% | 41.5% | 68.1% |
| Tau2-Bench (Retail) | 69.9% | 71.6% | 62.9% | 67.5% |
| Tau2-Bench (Airline) | 76.7% | 70.7% | 40% | 72.7% |
| Tau3-Bench (Banking) | 38.1% | 10.6% | 5.7% | 15.5% |
| BFCL v4 (overall) | 61.4% | 67.2% | 58% | 61% |