Skip to main content
ModelScale

Agentic benchmark

SWE Refactor Bench leaderboard

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?. Every model the catalog carries a published SWE Refactor Bench value for, ranked by that value.

CategoryAgentic
MeasureMigration audit, frozen behavioral checks, and agentic verification
Tasks20 whole-repository stack migrations
Difficulty6- to 30-hour autonomous repository migrations

Tests whether coding agents can complete long-horizon, whole-repository stack migrations while preserving the original program's behavior.

No model in the catalog has a published SWE Refactor Bench score.

The benchmark is defined by SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

All leaderboards · Agentic capability leaderboard