Skip to main content
ModelScale

Agentic benchmark

Toolathlon-Verified leaderboard

Every model the catalog carries a published Toolathlon-Verified value for, ranked by that value.

CategoryAgentic
MeasureInteractive tool-use score
TasksVerified multi-tool workflows
DifficultyAdvanced tool use

A verified tool-use benchmark variant for completing multi-step workflows with external tools.

Toolathlon-Verified ranking

16 models with a published Toolathlon-Verified value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Toolathlon-Verified value
RankModelProviderInteractive tool-use score
1Claude Opus 5Anthropic80.6
2GLM-5.3-FlashZ.AI78.4
3Claude Fable 5.1Anthropic77.8
4DeepSeek V4 Pro 0813DeepSeek74.1
4Hy4 previewTencent74.1
6Qwen3.8-Flash-NextAlibaba73.5
7Kimi K3Moonshot AI73.2
8GLM-5.3Z.AI73.0
9Qwen3.8 MaxAlibaba72.5
10OAOrnith-1.5-397BOrnith AI71.2
11DeepSeek V4 Flash 0731DeepSeek70.3
12DSdots3-note PreviewDots Studio55.6
13TMInkling-SmallThinking Machines Lab54.4
14Laguna S 2.1Poolside49.7
15OAOrnith-1.5-35B-A3BOrnith AI48.7
16OAOrnith-1.5-9BOrnith AI41.2

Evidence key: Observed

Rows are ordered by the value Kimi K3: Open Frontier Intelligence published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard