Skip to main content
ModelScale

Agentic benchmark

Terminal-Bench 3.0 leaderboard

Every model the catalog carries a published Terminal-Bench 3.0 value for, ranked by that value.

CategoryAgentic
MeasureTask completion rate
Tasks74 professional computer-work tasks across 7 domains
DifficultyFrontier autonomous knowledge work
Published byTerminal-Bench 3.0

A continuously maintained benchmark for difficult computer work, including coding, deep learning, finance, engineering, math, and science tasks.

Terminal-Bench 3.0 ranking

11 models with a published Terminal-Bench 3.0 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Terminal-Bench 3.0 value
RankModelProviderTask completion rate
1Claude Opus 5Anthropic42.7
2GPT-5.6 SolOpenAI34.6
3Claude Fable 5Anthropic34.0
4Grok 4.6xAI26.5
5Claude Opus 4.8Anthropic21.1
6GPT-5.6 TerraOpenAI20.8
7Grok 4.5xAI15.7
8Gemini 3.7 FlashGoogle14.9
9Claude Sonnet 5Anthropic14.6
10GPT-5.6 LunaOpenAI14.3
11GLM-5.2Z.AI4.6

Evidence key: Observed

Rows are ordered by the value Terminal-Bench 3.0 published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard