Coding benchmark
Terminal-Bench 2.0 leaderboard
Every model the catalog carries a published Terminal-Bench 2.0 value for, ranked by that value.
A benchmark for agentic software engineering tasks executed in real terminal environments. DeepSeek reports it in the agentic section, while BenchLM also mirrors it in coding for models that publish it as a developer-task signal.
Terminal-Bench 2.0 ranking
45 models with a published Terminal-Bench 2.0 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Interactive CLI agent evaluation |
|---|---|---|---|
| 1 | OpenAI | 91.9 | |
| 2 | Anthropic | 88.0 | |
| 3 | OpenAI | 87.4 | |
| 4 | OpenAI | 84.7 | |
| 5 | Anthropic | 84.3 | |
| 6 | xAI | 83.3 | |
| 7 | SASakana Fugu-Ultra | Sakana AI | 82.1 |
| 8 | OpenAI | 82.0 | |
| 9 | Cognition | 81.5 | |
| 10 | Z.AI | 81.0 | |
| 11 | Anthropic | 80.4 | |
| 12 | SASakana Fugu | Sakana AI | 80.2 |
| 13 | Meta | 80.0 | |
| 14 | DAOrnith-1.0-397B | DeepReinforce AI | 77.5 |
| 15 | 76.2 | ||
| 16 | Anthropic | 74.6 | |
| 17 | Alibaba | 70.3 | |
| 18 | Poolside | 70.2 | |
| 19 | Alibaba | 69.7 | |
| 20 | Anthropic | 69.4 | |
| 21 | Cursor | 69.3 | |
| 22 | Xiaomi | 68.4 | |
| 23 | DeepSeek | 67.9 | |
| 24 | Moonshot AI | 66.7 | |
| 25 | MiniMax | 66.0 | |
| 26 | Xiaomi | 65.8 | |
| 27 | Alibaba | 65.4 | |
| 28 | TMInkling-Small | Thinking Machines Lab | 64.7 |
| 29 | DAOrnith-1.0-35B | DeepReinforce AI | 64.2 |
| 30 | TMInkling | Thinking Machines Lab | 63.8 |
| 31 | Cursor | 61.7 | |
| 32 | StepFun | 59.5 | |
| 33 | Alibaba | 59.3 | |
| 34 | DeepSeek | 56.9 | |
| 35 | NVIDIA | 56.4 | |
| 36 | Tencent | 54.4 | |
| 37 | 54.0 | ||
| 38 | Alibaba | 51.5 | |
| 39 | Microsoft | 46.0 | |
| 40 | Poolside | 45.8 | |
| 41 | DAOrnith-1.0-9B | DeepReinforce AI | 43.1 |
| 42 | Poolside | 37.5 | |
| 43 | Poolside | 35.7 | |
| 44 | Meituan | 33.7 | |
| 45 | NVIDIA | 23.5 |
Evidence key: Observed
Rows are ordered by the value Terminal-Bench 2.0 published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.