external benchmark
SWE-Marathon leaderboard
Every model the catalog carries a published SWE-Marathon value for, ranked by that value.
Categoryexternal
MeasureTask resolution and trajectory review
Tasks20 multi-hour software engineering tasks
DifficultyUltra-long-horizon software engineering
A long-horizon software engineering benchmark from Abundant AI with multi-hour tasks spanning library reproductions, full-stack product clones, and ML engineering.
No model in the catalog has a published SWE-Marathon score.
The benchmark is defined by SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.