Skip to main content
ModelScale

Agentic benchmark

Terminal-Bench 2.1 (Vals) leaderboard

Terminal-Bench 2.1, Vals AI run. Every model the catalog carries a published Terminal-Bench 2.1 (Vals) value for, ranked by that value.

CategoryAgentic
MeasureTask success rate
TasksDifficult terminal tasks
DifficultyFrontier agentic

Vals AI’s independent run of the Terminal-Bench 2.1 terminal-task suite with published easy, medium, and hard splits.

Terminal-Bench 2.1 (Vals) ranking

53 models with a published Terminal-Bench 2.1 (Vals) value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Terminal-Bench 2.1, Vals AI run value
RankModelProviderTask success rate
1GPT-6 AstraOpenAI87.3
2GPT-5.6 SolOpenAI85.8
3Claude Fable 5.1Anthropic85.0
4Claude Opus 5Anthropic84.6
5Gemini 3.8 FlashGoogle81.3
6Kimi K3Moonshot AI80.9
7Claude Fable 5Anthropic80.5
8GPT-5.6 LunaOpenAI79.0
9Grok 4.6xAI78.3
10Gemini 3.7 FlashGoogle77.5
10GPT-5.6 TerraOpenAI77.5
12GPT-5.5OpenAI76.4
13Claude Sonnet 5Anthropic74.5
14Gemini 3.5 FlashGoogle74.2
15Gemini 3.6 FlashGoogle73.8
16Claude Opus 4.8Anthropic71.9
17GLM-5.3Z.AI71.5
18Gemini 3.1 ProGoogle70.8
19Muse Spark 1.2Meta69.7
20Muse Spark 1.1Meta69.3
21Claude Opus 4.7Anthropic68.5
22GLM-5.2Z.AI67.8
22Grok 4.5xAI67.8
24Qwen3.8 MaxAlibaba67.4
25DeepSeek V4 Flash 0731DeepSeek67.0
26GLM-5.3-FlashZ.AI62.9
27Qwen3.7 MaxAlibaba61.0
28MiMo-V2.5Xiaomi60.7
29Qwen3.8-27BAlibaba58.4
30Claude Sonnet 4.6Anthropic57.3
30MiMo-V2.5-ProXiaomi57.3
32GLM-5.1Z.AI56.9
33TMInkling-SmallThinking Machines Lab55.1
34DeepSeek V4 Pro 0813DeepSeek54.7
34GPT-5.4 miniOpenAI54.7
36Gemini 3 FlashGoogle53.9
37Kimi K2.6Moonshot AI53.6
37MiniMax M3MiniMax53.6
39Qwen3.6 PlusAlibaba53.2
40Qwen3.7 PlusAlibaba52.8
41Nemotron 3 UltraNVIDIA50.9
42Gemini 3.5 Flash-LiteGoogle50.2
42Ling 3.0 FlashInclusionAI50.2
44MiniMax M2.7MiniMax48.7
45TMInklingThinking Machines Lab47.6
46Grok 4.20xAI44.2
47Claude Haiku 4.5Anthropic43.8
48Grok 4.3xAI41.9
49GPT-5.4 nanoOpenAI41.6
50Mistral Medium 3.5 128BMistral39.0
51Gemini 3.1 Flash-LiteGoogle34.1
51Laguna M.1Poolside34.1
53Laguna XS.2Poolside25.8

Evidence key: Observed

Rows are ordered by the value Vals AI Terminal-Bench 2.1, Vals AI run leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard