Agentic benchmark
Terminal-Bench 2.1 (Vals) leaderboard
Terminal-Bench 2.1, Vals AI run. Every model the catalog carries a published Terminal-Bench 2.1 (Vals) value for, ranked by that value.
Vals AI’s independent run of the Terminal-Bench 2.1 terminal-task suite with published easy, medium, and hard splits.
Terminal-Bench 2.1 (Vals) ranking
53 models with a published Terminal-Bench 2.1 (Vals) value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Task success rate |
|---|---|---|---|
| 1 | OpenAI | 87.3 | |
| 2 | OpenAI | 85.8 | |
| 3 | Anthropic | 85.0 | |
| 4 | Anthropic | 84.6 | |
| 5 | 81.3 | ||
| 6 | Moonshot AI | 80.9 | |
| 7 | Anthropic | 80.5 | |
| 8 | OpenAI | 79.0 | |
| 9 | xAI | 78.3 | |
| 10 | 77.5 | ||
| 10 | OpenAI | 77.5 | |
| 12 | OpenAI | 76.4 | |
| 13 | Anthropic | 74.5 | |
| 14 | 74.2 | ||
| 15 | 73.8 | ||
| 16 | Anthropic | 71.9 | |
| 17 | Z.AI | 71.5 | |
| 18 | 70.8 | ||
| 19 | Meta | 69.7 | |
| 20 | Meta | 69.3 | |
| 21 | Anthropic | 68.5 | |
| 22 | Z.AI | 67.8 | |
| 22 | xAI | 67.8 | |
| 24 | Alibaba | 67.4 | |
| 25 | DeepSeek | 67.0 | |
| 26 | Z.AI | 62.9 | |
| 27 | Alibaba | 61.0 | |
| 28 | Xiaomi | 60.7 | |
| 29 | Alibaba | 58.4 | |
| 30 | Anthropic | 57.3 | |
| 30 | Xiaomi | 57.3 | |
| 32 | Z.AI | 56.9 | |
| 33 | TMInkling-Small | Thinking Machines Lab | 55.1 |
| 34 | DeepSeek | 54.7 | |
| 34 | OpenAI | 54.7 | |
| 36 | 53.9 | ||
| 37 | Moonshot AI | 53.6 | |
| 37 | MiniMax | 53.6 | |
| 39 | Alibaba | 53.2 | |
| 40 | Alibaba | 52.8 | |
| 41 | NVIDIA | 50.9 | |
| 42 | 50.2 | ||
| 42 | InclusionAI | 50.2 | |
| 44 | MiniMax | 48.7 | |
| 45 | TMInkling | Thinking Machines Lab | 47.6 |
| 46 | xAI | 44.2 | |
| 47 | Anthropic | 43.8 | |
| 48 | xAI | 41.9 | |
| 49 | OpenAI | 41.6 | |
| 50 | Mistral | 39.0 | |
| 51 | 34.1 | ||
| 51 | Poolside | 34.1 | |
| 53 | Poolside | 25.8 |
Evidence key: Observed
Rows are ordered by the value Vals AI Terminal-Bench 2.1, Vals AI run leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.