Skip to main content
ModelScale

Coding benchmark

Terminal-Bench Hard leaderboard

Every model the catalog carries a published Terminal-Bench Hard value for, ranked by that value.

CategoryCoding
MeasureTask success rate
TasksAgentic coding and terminal tasks
DifficultyProfessional software engineering

A display-only Artificial Analysis coding metric for agentic coding and terminal use on a harder Terminal-Bench slice.

No model in the catalog has a published Terminal-Bench Hard score.

The benchmark is defined by Artificial Analysis model benchmarks, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

All leaderboards · Coding capability leaderboard