Coding benchmark
HumanEval leaderboard
Evaluating Large Language Models Trained on Code. Every model the catalog carries a published HumanEval value for, ranked by that value.
A set of 164 handwritten Python function-generation problems. HumanEval is useful as a historical floor check, but BenchLM's current exact-source table is too small to support a broad frontier-coding verdict.
HumanEval ranking
2 models with a published HumanEval value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Python function generation |
|---|---|---|---|
| 1 | BTBTL-3 | Bad Theory Labs | 95.12 |
| 2 | SPSoofi S 30B-A3B | Soofi Project | 73.8 |
Evidence key: Observed
Rows are ordered by the value Evaluating Large Language Models Trained on Code published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.