Agentic benchmark
SWE-Atlas Refactoring leaderboard
Every model the catalog carries a published SWE-Atlas Refactoring value for, ranked by that value.
CategoryAgentic
MeasureRefactoring score with confidence intervals
TasksSWE-Atlas refactoring tasks
DifficultyReal-world software-engineering agent tasks
Published bySWE-Atlas
A Scale SWE-Atlas software-engineering agent benchmark focused on refactoring tasks.
No model in the catalog has a published SWE-Atlas Refactoring score.
The benchmark is defined by SWE-Atlas, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.