Agentic benchmark
Terminal-Bench 2.0 leaderboard
Every model the catalog carries a published Terminal-Bench 2.0 value for, ranked by that value.
A benchmark for agentic software engineering tasks executed in real terminal environments. Models must inspect files, run commands, edit code, and recover from errors over multi-step workflows.
Terminal-Bench 2.0 ranking
67 models with a published Terminal-Bench 2.0 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Interactive CLI agent evaluation |
|---|---|---|---|
| 1 | OpenAI | 91.9 | |
| 2 | Moonshot AI | 88.3 | |
| 3 | Anthropic | 88 | |
| 4 | OpenAI | 87.4 | |
| 5 | OpenAI | 84.7 | |
| 6 | Anthropic | 84.3 | |
| 7 | xAI | 83.3 | |
| 8 | SASakana Fugu-Ultra | Sakana AI | 82.1 |
| 9 | OpenAI | 82 | |
| 10 | Cognition | 81.5 | |
| 11 | Z.AI | 81 | |
| 12 | Anthropic | 80.4 | |
| 13 | SASakana Fugu | Sakana AI | 80.2 |
| 14 | Meta | 80 | |
| 15 | DAOrnith-1.0-397B | DeepReinforce AI | 77.5 |
| 16 | OpenAI | 77.3 | |
| 17 | 76.2 | ||
| 18 | OpenAI | 75.1 | |
| 19 | Anthropic | 74.6 | |
| 20 | Alibaba | 70.3 | |
| 21 | Poolside | 70.2 | |
| 22 | Alibaba | 69.7 | |
| 23 | Anthropic | 69.4 | |
| 24 | Cursor | 69.3 | |
| 25 | Xiaomi | 68.4 | |
| 26 | DeepSeek | 67.9 | |
| 27 | Moonshot AI | 66.7 | |
| 28 | MiniMax | 66 | |
| 29 | Xiaomi | 65.8 | |
| 30 | Anthropic | 65.4 | |
| 30 | Alibaba | 65.4 | |
| 32 | TMInkling-Small | Thinking Machines Lab | 64.7 |
| 33 | DAOrnith-1.0-35B | DeepReinforce AI | 64.2 |
| 34 | TMInkling | Thinking Machines Lab | 63.8 |
| 35 | Z.AI | 63.5 | |
| 36 | Cursor | 61.7 | |
| 37 | Alibaba | 61.6 | |
| 38 | OpenAI | 60 | |
| 39 | StepFun | 59.5 | |
| 40 | Anthropic | 59.3 | |
| 40 | Alibaba | 59.3 | |
| 42 | Anthropic | 59.1 | |
| 43 | Meta | 59 | |
| 44 | MiniMax | 57 | |
| 45 | DeepSeek | 56.9 | |
| 46 | NVIDIA | 56.4 | |
| 47 | Z.AI | 56.2 | |
| 48 | Tencent | 54.4 | |
| 49 | 54 | ||
| 50 | Alibaba | 52.5 | |
| 51 | Alibaba | 51.5 | |
| 52 | Moonshot AI | 50.8 | |
| 52 | Moonshot AI | 50.8 | |
| 54 | Anthropic | 50 | |
| 55 | Alibaba | 49.4 | |
| 56 | xAI | 47.1 | |
| 57 | OpenAI | 46.3 | |
| 58 | Microsoft | 46 | |
| 59 | Poolside | 45.8 | |
| 60 | DAOrnith-1.0-9B | DeepReinforce AI | 43.1 |
| 61 | Alibaba | 41.6 | |
| 62 | Z.AI | 41 | |
| 63 | Alibaba | 40.5 | |
| 64 | Poolside | 37.5 | |
| 65 | Poolside | 35.7 | |
| 66 | Meituan | 33.7 | |
| 67 | NVIDIA | 23.46 |
Evidence key: Observed
Rows are ordered by the value Terminal-Bench 2.0 published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.