Coding benchmark
Terminal-Bench Hard leaderboard
Every model the catalog carries a published Terminal-Bench Hard value for, ranked by that value.
CategoryCoding
MeasureTask success rate
TasksAgentic coding and terminal tasks
DifficultyProfessional software engineering
Published byArtificial Analysis model benchmarks
A display-only Artificial Analysis coding metric for agentic coding and terminal use on a harder Terminal-Bench slice.
No model in the catalog has a published Terminal-Bench Hard score.
The benchmark is defined by Artificial Analysis model benchmarks, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.