Skip to main content
ModelScale

Coding benchmark

ProgramBench (episode 1) leaderboard

ProgramBench hidden-test pass rate after episode 1. Every model the catalog carries a published ProgramBench (episode 1) value for, ranked by that value.

CategoryCoding
MeasureHidden-test pass rate after episode 1
Tasks166 golden program-reconstruction tasks
DifficultyLong-context clean-room software engineering

Program-reconstruction hidden-test pass rate after the first of five sequential long-context episodes.

ProgramBench (episode 1) ranking

1 model with a published ProgramBench (episode 1) value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published ProgramBench hidden-test pass rate after episode 1 value
RankModelProviderHidden-test pass rate after episode 1
1Claude Opus 5Anthropic83.0

Evidence key: Observed

Rows are ordered by the value ProgramBench: Can language models rebuild programs from scratch? published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Coding capability leaderboard