Coding benchmark
AA Terminal-Bench 2.1 leaderboard
Artificial Analysis Terminal-Bench v2.1. Every model the catalog carries a published AA Terminal-Bench 2.1 value for, ranked by that value.
CategoryCoding
MeasureTask success rate
TasksTerminal-based agent tasks
DifficultyAgentic software engineering
An independently evaluated Terminal-Bench v2.1 result from Artificial Analysis.
No model in the catalog has a published AA Terminal-Bench 2.1 score.
The benchmark is defined by Artificial Analysis Terminal-Bench v2.1 Benchmark Leaderboard, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.