external benchmark
DeepSWE leaderboard
Every model the catalog carries a published DeepSWE value for, ranked by that value.
Categoryexternal
MeasurePass@1 with confidence interval, cost, time, and token metadata
Tasks113 software engineering tasks across 91 repositories and 5 languages
DifficultyLong-horizon software engineering
Published byDeepSWE benchmark blog
A long-horizon software engineering benchmark from Datacurve for measuring frontier coding agents on original tasks drawn from active open-source repositories.
No model in the catalog has a published DeepSWE score.
The benchmark is defined by DeepSWE benchmark blog, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.