Coding benchmark
CursorBench leaderboard
Every model the catalog carries a published CursorBench value for, ranked by that value.
CategoryCoding
MeasureCursor agent-loop evaluation
TasksLong-horizon multi-file agentic coding tasks
DifficultyProfessional agentic software engineering
Published byCursorBench 4.0
Cursor's first-party benchmark for ambiguous, multi-file coding-agent tasks drawn from real Cursor sessions, currently on the 4.0 task set.
No model in the catalog has a published CursorBench score.
The benchmark is defined by CursorBench 4.0, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.