Skip to main content
ModelScale

Agentic benchmark

Terminal-Bench 2.0 leaderboard

Every model the catalog carries a published Terminal-Bench 2.0 value for, ranked by that value.

CategoryAgentic
MeasureInteractive CLI agent evaluation
TasksTerminal-based software tasks
DifficultyProfessional software engineering
Published byTerminal-Bench 2.0

A benchmark for agentic software engineering tasks executed in real terminal environments. Models must inspect files, run commands, edit code, and recover from errors over multi-step workflows.

Terminal-Bench 2.0 ranking

67 models with a published Terminal-Bench 2.0 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Terminal-Bench 2.0 value
RankModelProviderInteractive CLI agent evaluation
1GPT-5.6 SolOpenAI91.9
2Kimi K3Moonshot AI88.3
3Claude Mythos 5Anthropic88
4GPT-5.6 TerraOpenAI87.4
5GPT-5.6 LunaOpenAI84.7
6Claude Fable 5Anthropic84.3
7Grok 4.5xAI83.3
8SASakana Fugu-UltraSakana AI82.1
9GPT-5.5OpenAI82
10SWE-1.7Cognition81.5
11GLM-5.2Z.AI81
12Claude Sonnet 5Anthropic80.4
13SASakana FuguSakana AI80.2
14Muse Spark 1.1Meta80
15DAOrnith-1.0-397BDeepReinforce AI77.5
16GPT-5.3 CodexOpenAI77.3
17Gemini 3.5 FlashGoogle76.2
18GPT-5.4OpenAI75.1
19Claude Opus 4.8Anthropic74.6
20Qwen3.7 PlusAlibaba70.3
21Laguna S 2.1Poolside70.2
22Qwen3.7 MaxAlibaba69.7
23Claude Opus 4.7 (Adaptive)Anthropic69.4
24Composer 2.5Cursor69.3
25MiMo-V2.5-ProXiaomi68.4
26DeepSeek V4 Pro 0813DeepSeek67.9
27Kimi K2.6Moonshot AI66.7
28MiniMax M3MiniMax66
29MiMo-V2.5Xiaomi65.8
30Claude Opus 4.6Anthropic65.4
30Qwen 3.6 Max (preview)Alibaba65.4
32TMInkling-SmallThinking Machines Lab64.7
33DAOrnith-1.0-35BDeepReinforce AI64.2
34TMInklingThinking Machines Lab63.8
35GLM-5.1Z.AI63.5
36Composer 2Cursor61.7
37Qwen3.6 PlusAlibaba61.6
38GPT-5.4 miniOpenAI60
39Step 3.7 FlashStepFun59.5
40Claude Opus 4.5Anthropic59.3
40Qwen3.6-27BAlibaba59.3
42Claude Sonnet 4.6Anthropic59.1
43Muse SparkMeta59
44MiniMax M2.7MiniMax57
45DeepSeek V4 Flash 0731DeepSeek56.9
46Nemotron 3 UltraNVIDIA56.4
47GLM-5Z.AI56.2
48Hy3 PreviewTencent54.4
49Gemini 3.5 Flash-LiteGoogle54
50Qwen3.5 397BAlibaba52.5
51Qwen3.6-35B-A3BAlibaba51.5
52Kimi K2.5Moonshot AI50.8
52Kimi K2.5 (Reasoning)Moonshot AI50.8
54Claude Sonnet 4.5Anthropic50
55Qwen3.5-122B-A10BAlibaba49.4
56Grok 4.20xAI47.1
57GPT-5.4 nanoOpenAI46.3
58MAI-Thinking-1Microsoft46
59Laguna M.1Poolside45.8
60DAOrnith-1.0-9BDeepReinforce AI43.1
61Qwen3.5-27BAlibaba41.6
62GLM-4.7Z.AI41
63Qwen3.5-35B-A3BAlibaba40.5
64Laguna XS 2.1Poolside37.5
65Laguna XS.2Poolside35.7
66LongCat-Flash-Lite-SparseMeituan33.7
67Nemotron 3.5 Lightning 30B A3B NVFP4NVIDIA23.46

Evidence key: Observed

Rows are ordered by the value Terminal-Bench 2.0 published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard