Coding benchmark
ProgramBench (episode 1) leaderboard
ProgramBench hidden-test pass rate after episode 1. Every model the catalog carries a published ProgramBench (episode 1) value for, ranked by that value.
Program-reconstruction hidden-test pass rate after the first of five sequential long-context episodes.
ProgramBench (episode 1) ranking
1 model with a published ProgramBench (episode 1) value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Hidden-test pass rate after episode 1 |
|---|---|---|---|
| 1 | Anthropic | 83.0 |
Evidence key: Observed
Rows are ordered by the value ProgramBench: Can language models rebuild programs from scratch? published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.