Skip to main content
ModelScale

Coding benchmark

CursorBench leaderboard

Every model the catalog carries a published CursorBench value for, ranked by that value.

CategoryCoding
MeasureCursor agent-loop evaluation
TasksLong-horizon multi-file agentic coding tasks
DifficultyProfessional agentic software engineering
Published byCursorBench 4.0

Cursor's first-party benchmark for ambiguous, multi-file coding-agent tasks drawn from real Cursor sessions, currently on the 4.0 task set.

No model in the catalog has a published CursorBench score.

The benchmark is defined by CursorBench 4.0, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

All leaderboards · Coding capability leaderboard