Skip to main content
ModelScale

external benchmark

Terminal-Bench 4.0 (Vals) leaderboard

Vals Terminal-Bench 4.0. Every model the catalog carries a published Terminal-Bench 4.0 (Vals) value for, ranked by that value.

Categoryexternal
MeasureAccuracy score
TasksFrontier-difficulty terminal tasks across seven domains
DifficultyFrontier terminal-agent execution

Vals AI’s independent run of the Terminal-Bench 4.0 frontier terminal suite, with per-domain splits.

No model in the catalog has a published Terminal-Bench 4.0 (Vals) score.

The benchmark is defined by Vals Terminal-Bench 4.0, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

All leaderboards