Skip to main content
ModelScale

Agentic benchmark

OpenHands Index leaderboard

Every model the catalog carries a published OpenHands Index value for, ranked by that value.

CategoryAgentic
MeasureMacro-average across five coding-agent categories
TasksSWE-bench Verified, SWE-bench Multimodal, Commit0, SWT-bench Verified, and GAIA
DifficultyReal-world software engineering agent tasks

A holistic coding-agent benchmark that evaluates AI agents across issue resolution, frontend work, greenfield development, testing, and information gathering.

No model in the catalog has a published OpenHands Index score.

The benchmark is defined by OpenHands Index methodology, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

All leaderboards · Agentic capability leaderboard