Skip to main content
ModelScale

Agentic benchmark

AA Terminal-Bench 4.0 leaderboard

Artificial Analysis Terminal-Bench v4.0. Every model the catalog carries a published AA Terminal-Bench 4.0 value for, ranked by that value.

CategoryAgentic
MeasureTask success rate
TasksTerminal-based agent tasks
DifficultyAgentic software engineering

An independently evaluated Terminal-Bench v4.0 result from Artificial Analysis, one of the ten components of its Intelligence Index v4.3.

AA Terminal-Bench 4.0 ranking

15 models with a published AA Terminal-Bench 4.0 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Artificial Analysis Terminal-Bench v4.0 value
RankModelProviderTask success rate
1GPT-6 AstraOpenAI59.1
2Claude Fable 5.1Anthropic52.0
3Claude Opus 5Anthropic49.0
4Claude Fable 5Anthropic42.4
5GLM-5.3Z.AI41.9
6GPT-5.6 SolOpenAI39.9
7GPT-5.6 TerraOpenAI35.4
8Muse Spark 1.3Meta33.3
9GLM-5.3-FlashZ.AI32.8
10DeepSeek V4.1 FlashDeepSeek26.8
11Grok 4.6xAI21.2
12Gemini 3.8 FlashGoogle19.7
13DeepSeek V4 Pro 0813DeepSeek14.1
14Kimi K3Moonshot AI12.6
15GPT-5.6 LunaOpenAI11.6

Evidence key: Last good

Rows are ordered by the value Artificial Analysis Terminal-Bench v4.0 Benchmark Leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard