Skip to main content
ModelScale

Coding benchmark

HumanEval leaderboard

Evaluating Large Language Models Trained on Code. Every model the catalog carries a published HumanEval value for, ranked by that value.

CategoryCoding
MeasurePython function generation
Tasks164 problems
DifficultyIntroductory to intermediate programming

A set of 164 handwritten Python function-generation problems. HumanEval is useful as a historical floor check, but BenchLM's current exact-source table is too small to support a broad frontier-coding verdict.

HumanEval ranking

2 models with a published HumanEval value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Evaluating Large Language Models Trained on Code value
RankModelProviderPython function generation
1BTBTL-3Bad Theory Labs95.12
2SPSoofi S 30B-A3BSoofi Project73.8

Evidence key: Observed

Rows are ordered by the value Evaluating Large Language Models Trained on Code published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Coding capability leaderboard