Skip to main content
ModelScale

Agentic benchmark

Claw-Eval leaderboard

Every model the catalog carries a published Claw-Eval value for, ranked by that value.

CategoryAgentic
MeasureEnd-to-end autonomous-agent evaluation with Pass^3 scoring
Tasks300 tasks, 2,159 rubrics
DifficultyReal-world general, multi-turn, and native multimodal agent execution

A transparent real-world autonomous-agent benchmark with 300 human-verified tasks, 2,159 rubric items, and Pass^3 scoring across general, multi-turn, and native multimodal agent tasks.

Claw-Eval ranking

39 models with a published Claw-Eval value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Claw-Eval value
RankModelProviderEnd-to-end autonomous-agent evaluation with Pass^3 scoring
1OAOrnith-1.5-397BOrnith AI81.4
2K-EXAONE 2.0LG AI Research77.7
3DAOrnith-1.0-397BDeepReinforce AI77.1
4MiniMax M3MiniMax74.5
5DSdots3-note PreviewDots Studio73.4
6OAOrnith-1.5-35B-A3BOrnith AI72.5
7Qwen3.6-27BAlibaba72.4
8Claude Opus 4.6Anthropic70.4
9DAOrnith-1.0-35BDeepReinforce AI69.8
10Qwen3.6-35B-A3BAlibaba68.7
11Claude Sonnet 4.6Anthropic67.8
12Step 3.7 FlashStepFun67.1
13OAOrnith-1.5-9BOrnith AI66.5
14Qwen3.7 MaxAlibaba65.2
15LLaDA2.2-flashInclusionAI64.2
16MiMo-V2.5-ProXiaomi63.8
16Muse SparkMeta63.8
18DAOrnith-1.0-9BDeepReinforce AI63.1
19LFM2.5-2.6BLiquidAI62.9
20Qwen3.7 PlusAlibaba62.7
21GLM-5.1Z.AI62.3
21Kimi K2.6Moonshot AI62.3
21MiMo-V2.5Xiaomi62.3
24GPT-5.4OpenAI60.3
25Claude Opus 4.5Anthropic59.6
26Qwen3.6 PlusAlibaba58.8
27Gemini 3.1 ProGoogle57.8
27MiMo-V2-ProXiaomi57.8
29GLM-5Z.AI57.7
30LLaDA2.2-miniInclusionAI57.2
31Qwen3.5 397BAlibaba56.8
32GLM-5-TurboZ.AI55.8
33GLM-5V-TurboZ.AI53.8
34Kimi K2.5Moonshot AI52.3
35Gemini 3 FlashGoogle49.2
36MiniMax M2.7MiniMax48.7
37MiMo-V2-OmniXiaomi45.2
38DeepSeek V3.2DeepSeek40.2
39Nemotron 3 Super 100BNVIDIA5.5

Evidence key: Observed

Rows are ordered by the value Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard