Skip to main content
ModelScale

Coding benchmark

VulcanBench CII v1 leaderboard

VulcanBench Coding Intelligence Index v1. Every model the catalog carries a published VulcanBench CII v1 value for, ranked by that value.

CategoryCoding
MeasurePass@1 with vendor coding-agent harnesses
Tasks38 validated post-cutoff repository tasks
DifficultyMid-band frontier software engineering

A post-cutoff software-engineering benchmark with hidden functional tests and regression guards, reported for vendor coding-agent harnesses.

VulcanBench CII v1 ranking

3 models with a published VulcanBench CII v1 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published VulcanBench Coding Intelligence Index v1 value
RankModelProviderPass@1 with vendor coding-agent harnesses
1Claude Opus 5Anthropic96.4
2Claude Sonnet 5Anthropic89.2
3GPT-5.6 SolOpenAI86.5

Evidence key: Observed

Rows are ordered by the value CII v1 frontier results published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Coding capability leaderboard