Agentic benchmark
OpenHands Index leaderboard
Every model the catalog carries a published OpenHands Index value for, ranked by that value.
CategoryAgentic
MeasureMacro-average across five coding-agent categories
TasksSWE-bench Verified, SWE-bench Multimodal, Commit0, SWT-bench Verified, and GAIA
DifficultyReal-world software engineering agent tasks
Published byOpenHands Index methodology
A holistic coding-agent benchmark that evaluates AI agents across issue resolution, frontend work, greenfield development, testing, and information gathering.
No model in the catalog has a published OpenHands Index score.
The benchmark is defined by OpenHands Index methodology, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.