Coding benchmark
IDE-Bench leaderboard
Every model the catalog carries a published IDE-Bench value for, ranked by that value.
CategoryCoding
MeasureAutonomous IDE-agent task completion (pass@1)
Tasks80 tasks across 8 repositories
DifficultyEnd-to-end software engineering
Published byIDE-Bench
An 80-task software-engineering benchmark across eight repositories that tests whether autonomous IDE agents can explore, edit, run, and verify code changes end to end.
No model in the catalog has a published IDE-Bench score.
The benchmark is defined by IDE-Bench, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.