Skip to main content
ModelScale

Coding benchmark

ProgramBench leaderboard

ProgramBench: Can Language Models Rebuild Programs From Scratch?. Every model the catalog carries a published ProgramBench value for, ranked by that value.

CategoryCoding
MeasureCleanroom executable reimplementation
Tasks200 program reconstruction tasks
DifficultyFull-repository software architecture

A cleanroom software-engineering benchmark where agents receive only a compiled executable and documentation, then must architect and implement a complete codebase that reproduces the original program's behavior.

ProgramBench ranking

8 models with a published ProgramBench value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published ProgramBench: Can Language Models Rebuild Programs From Scratch? value
RankModelProviderCleanroom executable reimplementation
1Claude Opus 5Anthropic93.0
2Claude Fable 5.1Anthropic87.6
3Kimi K3Moonshot AI77.8
4GLM-5.2Z.AI63.7
5Kimi K2.7 CodeMoonshot AI53.6
6DeepSeek V4.1 FlashDeepSeek20.3
7GLM-5.3Z.AI19.0
8Hy4 previewTencent17.5

Evidence key: Observed

Rows are ordered by the value ProgramBench: Can Language Models Rebuild Programs From Scratch? published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Coding capability leaderboard