Skip to main content
ModelScale

Agentic benchmark

τ²-bench results leaderboard

τ²-Bench Tool-Agent-User Evaluation. Every model the catalog carries a published τ²-bench results value for, ranked by that value.

CategoryAgentic
MeasurePublished domain success or pass^k results
TasksAirline, retail, and telecom customer-service task sets
DifficultyDual-control customer-service workflows

This route is a sourced ledger for published τ²-bench results. Most current rows come from Artificial Analysis's telecom implementation, while named provider rows can use telecom, airline, retail, or aggregate setups.

τ²-bench results ranking

133 models with a published τ²-bench results value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published τ²-Bench Tool-Agent-User Evaluation value
RankModelProviderPublished domain success or pass^k results
1GLM-5.2Z.AI99.1
2GPT-5.4OpenAI98.9
3Claude Fable 5Anthropic98.5
3GLM-5-TurboZ.AI98.5
3GLM-5V-TurboZ.AI98.5
3Step 3.7 FlashStepFun98.5
7GLM-5Z.AI98.2
8GPT-5.5OpenAI98
9GLM-5.1Z.AI97.7
9Grok 4.3xAI97.7
9Qwen3.6 PlusAlibaba97.7
12DeepSeek V4 Pro 0813DeepSeek96.2
13GLM-4.7Z.AI95.9
13Kimi K2.5Moonshot AI95.9
13Kimi K2.5 (Reasoning)Moonshot AI95.9
13Kimi K2.6Moonshot AI95.9
13Qwen 3.6 Max (preview)Alibaba95.9
18Gemini 3.1 ProGoogle95.6
19Gemini 3.5 FlashGoogle95.3
19Qwen3.6-35B-A3BAlibaba95.3
21MiMo-V2-ProXiaomi95
22Qwen3.7 MaxAlibaba94.7
23Claude Opus 4.8Anthropic94.4
24MiMo-V2.5-ProXiaomi94.2
24Mistral Medium 3.5 128BMistral94.2
24Qwen3.6-27BAlibaba94.2
27Qwen3.5-27BAlibaba93.9
28Qwen3.5-122B-A10BAlibaba93.6
29GPT-5.4 miniOpenAI93.4
30Grok 4.1 Fast (Reasoning)xAI93.3
31Qwen3.7 PlusAlibaba93
32GPT-5.4 nanoOpenAI92.5
33Claude Opus 4.6 (Adaptive)Anthropic92.1
33GPT-5.2-CodexOpenAI92.1
35Muse SparkMeta91.5
36MiMo-V2-OmniXiaomi91.2
37Kimi K2.7 CodeMoonshot AI90.1
37Trinity-Large-PreviewArcee AI90.1
37Trinity-Large-ThinkingArcee AI90.1
40Claude Opus 4.5 ThinkingAnthropic89.5
41Qwen3.5-35B-A3BAlibaba89.2
42MiniMax M3MiniMax88.9
43Claude Opus 4.7 (Adaptive)Anthropic88.6
44LFM2.5-8B-A1BLiquidAI88.07
45Gemini 3 ProGoogle87.1
46GPT-5 (medium)OpenAI86.5
47Claude Opus 4.5Anthropic86.3
47GPT-5.6 TerraOpenAI86.3
47Solar Pro 3Upstage86.3
50GPT-5.3 CodexOpenAI86
50Ling 2.6 FlashInclusionAI86
52GPT-5.6 SolOpenAI85.1
53Command A+Cohere85
54Claude Opus 4.6Anthropic84.8
54GPT-5 (high)OpenAI84.8
54GPT-5.2OpenAI84.8
54MiniMax M2.7MiniMax84.8
58MiMo-V2-FlashXiaomi83.9
58Qwen3.5 397BAlibaba83.9
58Qwen3.5 397B (Reasoning)Alibaba83.9
61Nemotron 3 UltraNVIDIA83.3
62GPT-5.1-CodexOpenAI83
62GPT-5.1-Codex-MaxOpenAI83
64GPT-5.1OpenAI81.9
65o3OpenAI80.7
66LLaDA2.2-flashInclusionAI80.33
67PMTernary Bonsai 2 27BPrism ML80.22
68Claude Sonnet 4.6Anthropic79.5
69DeepSeek V3.2DeepSeek78.9
70Agents-A1-4BInternScience78.2
71GLM-4.6Z.AI76.9
72Grok Code Fast 1xAI75.7
73Grok 4xAI74.9
74K-ExaoneLG AI Research74.3
74Qwen3 MaxAlibaba74.3
76Claude Opus 4.7Anthropic74
77Claude 4.1 Opus ThinkingAnthropic71.4
78Nemotron 3 Super 100BNVIDIA67.8
79GPT-OSS 120BOpenAI65.8
79Grok 4 Fast (Reasoning)xAI65.8
81Grok 4.1 FastxAI63.7
82o1OpenAI62.6
83Kimi K2Moonshot AI61.1
84GPT-OSS 20BOpenAI60.2
85Gemma 4 31BGoogle59.9
86LLaDA2.2-miniInclusionAI57.5
87Gemini 2.5 ProGoogle54.1
88GPT-4.1 miniOpenAI52.9
89Claude 4 SonnetAnthropic52.3
90GPT-4.1OpenAI47.1
91SASarvam 105BSarvam46.8
92GLM-4.5-AirZ.AI46.5
93Nemotron 3 Nano Omni 30B A3BNVIDIA45.3
94Gemma 4 26B A4BGoogle43.6
95Gemini 3 FlashGoogle43.3
96Mistral Small 4Mistral41.2
96Mistral Small 4 (Reasoning)Mistral41.2
98Nemotron 3 Nano 30BNVIDIA40.9
99DeepSeek V3.1 (Reasoning)DeepSeek37.4
99North Mini CodeCohere37.4
101DeepSeek-R1DeepSeek36.5
102Gemma 4 12BGoogle36.3
103DeepSeek V3.1DeepSeek34.8
104SASarvam 30BSarvam34.5
105Solar Pro 2Upstage31.9
106Mistral Large 2Mistral30.7
107o3-miniOpenAI28.7
108FAUltravox v0.6 Llama 3.3 70BFixie AI26.6
109GPT-4oOpenAI25.1
110Mistral Large 3Mistral24.6
111Mistral Medium 3Mistral24.3
112DeepSeek V3DeepSeek22.8
112Granite-4.0-1BIBM22.8
114Qwen3-Omni-30B-A3B-ThinkingAlibaba21.3
115Claude 3 HaikuAnthropic21.1
116Gemma 4 E2BGoogle20.8
116Gemma 4 E4BGoogle20.8
118Exaone 4.0 1.2BLG AI Research20.5
119Granite-4.0-H-1BIBM19.6
120Llama 3.1 405BMeta19
121Llama 4 MaverickMeta17.8
122GPT-4.1 nanoOpenAI17.3
123Qwen3-Omni-30B-A3B-InstructAlibaba16.4
124Llama 4 ScoutMeta15.5
125Gemini 2.5 FlashGoogle14.9
126Granite-4.0-H-350MIBM14.6
127Nova ProAmazon14
128Granite-4.0-350MIBM13.2
129Nemotron Ultra 253BNVIDIA11.4
130Gemma 3 27BGoogle10.5
131LFM2.5-VL-1.6B-ExtractLiquidAI8.5
132Exaone 4.0 32BLG AI Research4.1
133Phi-4Microsoft0

Evidence key: ObservedLast good

Rows are ordered by the value τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard