Skip to main content
ModelScale

Coding benchmark

Terminal-Bench 2.0 leaderboard

Every model the catalog carries a published Terminal-Bench 2.0 value for, ranked by that value.

CategoryCoding
MeasureInteractive CLI agent evaluation
TasksTerminal-based software tasks
DifficultyProfessional software engineering
Published byTerminal-Bench 2.0

A benchmark for agentic software engineering tasks executed in real terminal environments. DeepSeek reports it in the agentic section, while BenchLM also mirrors it in coding for models that publish it as a developer-task signal.

Terminal-Bench 2.0 ranking

45 models with a published Terminal-Bench 2.0 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Terminal-Bench 2.0 value
RankModelProviderInteractive CLI agent evaluation
1GPT-5.6 SolOpenAI91.9
2Claude Mythos 5Anthropic88.0
3GPT-5.6 TerraOpenAI87.4
4GPT-5.6 LunaOpenAI84.7
5Claude Fable 5Anthropic84.3
6Grok 4.5xAI83.3
7SASakana Fugu-UltraSakana AI82.1
8GPT-5.5OpenAI82.0
9SWE-1.7Cognition81.5
10GLM-5.2Z.AI81.0
11Claude Sonnet 5Anthropic80.4
12SASakana FuguSakana AI80.2
13Muse Spark 1.1Meta80.0
14DAOrnith-1.0-397BDeepReinforce AI77.5
15Gemini 3.5 FlashGoogle76.2
16Claude Opus 4.8Anthropic74.6
17Qwen3.7 PlusAlibaba70.3
18Laguna S 2.1Poolside70.2
19Qwen3.7 MaxAlibaba69.7
20Claude Opus 4.7 (Adaptive)Anthropic69.4
21Composer 2.5Cursor69.3
22MiMo-V2.5-ProXiaomi68.4
23DeepSeek V4 Pro 0813DeepSeek67.9
24Kimi K2.6Moonshot AI66.7
25MiniMax M3MiniMax66.0
26MiMo-V2.5Xiaomi65.8
27Qwen 3.6 Max (preview)Alibaba65.4
28TMInkling-SmallThinking Machines Lab64.7
29DAOrnith-1.0-35BDeepReinforce AI64.2
30TMInklingThinking Machines Lab63.8
31Composer 2Cursor61.7
32Step 3.7 FlashStepFun59.5
33Qwen3.6-27BAlibaba59.3
34DeepSeek V4 Flash 0731DeepSeek56.9
35Nemotron 3 UltraNVIDIA56.4
36Hy3 PreviewTencent54.4
37Gemini 3.5 Flash-LiteGoogle54.0
38Qwen3.6-35B-A3BAlibaba51.5
39MAI-Thinking-1Microsoft46.0
40Laguna M.1Poolside45.8
41DAOrnith-1.0-9BDeepReinforce AI43.1
42Laguna XS 2.1Poolside37.5
43Laguna XS.2Poolside35.7
44LongCat-Flash-Lite-SparseMeituan33.7
45Nemotron 3.5 Lightning 30B A3B NVFP4NVIDIA23.5

Evidence key: Observed

Rows are ordered by the value Terminal-Bench 2.0 published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Coding capability leaderboard