Skip to main content
ModelScale

Agentic benchmark

Toolathlon leaderboard

Every model the catalog carries a published Toolathlon value for, ranked by that value.

CategoryAgentic
MeasureInteractive tool-calling evaluation
TasksMulti-tool workflows
DifficultyAdvanced tool use

A tool-use benchmark focused on selecting, sequencing, and completing tasks with external tools.

Toolathlon ranking

22 models with a published Toolathlon value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Toolathlon value
RankModelProviderInteractive tool-calling evaluation
1Muse Spark 1.1Meta75.6
2Claude Opus 4.8Anthropic59.9
3GPT-5.6 SolOpenAI58
4Gemini 3.5 FlashGoogle56.5
5GPT-5.5OpenAI55.6
6GPT-5.4OpenAI54.6
7GPT-5.6 LunaOpenAI53.4
8GPT-5.6 TerraOpenAI53.1
9DeepSeek V4 Pro 0813DeepSeek51.8
10Kimi K2.6Moonshot AI50
11Step 3.7 FlashStepFun49.5
12GLM-5.2Z.AI48.2
13DeepSeek V4 Flash 0731DeepSeek47.8
14MiniMax M2.7MiniMax46.3
15Claude Opus 4.5Anthropic43.5
16GPT-5.4 miniOpenAI42.9
17Qwen3.6 PlusAlibaba39.8
18GLM-5Z.AI38
19Qwen3.5 397BAlibaba36.3
20GPT-5.4 nanoOpenAI35.5
21Kimi K2.5Moonshot AI27.8
22Qwen3.6-35B-A3BAlibaba26.9

Evidence key: Observed

Rows are ordered by the value Introducing GPT-5.4 mini and nano published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard