Skip to main content
ModelScale

Agentic benchmark

Terminal-Bench 4.0 leaderboard

Every model the catalog carries a published Terminal-Bench 4.0 value for, ranked by that value.

CategoryAgentic
MeasureTask completion rate across 5 trials per task
Tasks66 professional computer-work tasks
DifficultyFrontier autonomous knowledge work
Published byTerminal-Bench 4.0

The current Terminal-Bench release measures difficult computer work after recalibrating task resources, fixing unstable tasks, and removing tasks that no longer separate frontier systems.

Terminal-Bench 4.0 ranking

7 models with a published Terminal-Bench 4.0 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Terminal-Bench 4.0 value
RankModelProviderTask completion rate across 5 trials per task
1Claude Mythos 5.1Anthropic60.90
2GPT-6 AstraOpenAI57.90
3Claude Fable 5.1Anthropic55.80
4UNPareto 26.9Unbiased51.00
5DeepSeek V4.1 FlashDeepSeek31.20
6SWE-2Cognition27.30
7Gemini 3.8 FlashGoogle19.10

Evidence key: Observed

Rows are ordered by the value Terminal-Bench 4.0 published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard