Skip to main content
ModelScale

Agentic benchmark

τ³-bench results leaderboard

τ³-Bench Tool-Agent-User Evaluation. Every model the catalog carries a published τ³-bench results value for, ranked by that value.

CategoryAgentic
MeasurePublished domain or average success results
TasksCorrected customer-service tasks plus knowledge and voice evaluation modes
DifficultyLong-horizon, multimodal, and knowledge-aware tool use

τ³-bench is the current evolution of Sierra's tool-agent-user framework, adding corrected task releases and newer knowledge and voice evaluation modes alongside airline, retail, and telecom.

τ³-bench results ranking

18 models with a published τ³-bench results value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published τ³-Bench Tool-Agent-User Evaluation value
RankModelProviderPublished domain or average success results
1Mercury 2.5Inception96.0
2Mistral Medium 3.5 128BMistral91.4
3MiMo-V2.5-ProXiaomi72.9
4Nemotron 3 UltraNVIDIA70.9
5Qwen3.6 PlusAlibaba70.7
6GLM-5.1Z.AI70.6
7Claude Opus 4.5Anthropic70.2
8Qwen3.5 397BAlibaba68.4
9Qwen3.6-35B-A3BAlibaba67.2
10PAPokee-Isaac 28BPokee AI66.2
11Kimi K2.5Moonshot AI65.7
12GLM-5Z.AI65.6
13Granite 4.2 30BIBM62.0
14Granite 4.2 8BIBM58.1
15Granite 4.2 3BIBM45.8
16Atria Dawn PreviewShanghai Artificial Intelligence Laboratory41.2
17Nemotron 3.5 Lightning 30B A3B NVFP4NVIDIA9.5
18LFM2.5-2.6BLiquidAI5.7

Evidence key: Observed

Rows are ordered by the value Official τ³-bench repository and release notes published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard