Skip to main content
ModelScale

Coding benchmark

AA Terminal-Bench 2.1 leaderboard

Artificial Analysis Terminal-Bench v2.1. Every model the catalog carries a published AA Terminal-Bench 2.1 value for, ranked by that value.

CategoryCoding
MeasureTask success rate
TasksTerminal-based agent tasks
DifficultyAgentic software engineering

An independently evaluated Terminal-Bench v2.1 result from Artificial Analysis.

No model in the catalog has a published AA Terminal-Bench 2.1 score.

The benchmark is defined by Artificial Analysis Terminal-Bench v2.1 Benchmark Leaderboard, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

All leaderboards · Coding capability leaderboard