Knowledge benchmark
MMLU-Redux leaderboard
Every model the catalog carries a published MMLU-Redux value for, ranked by that value.
A harder refresh of MMLU intended to keep broad knowledge evaluation useful after the original benchmark became too easy for frontier models.
MMLU-Redux ranking
11 models with a published MMLU-Redux value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Multiple choice questions |
|---|---|---|---|
| 1 | Anthropic | 96.6 | |
| 2 | Alibaba | 95 | |
| 3 | Alibaba | 94.9 | |
| 4 | Alibaba | 94.5 | |
| 4 | Alibaba | 94.5 | |
| 6 | Alibaba | 93.5 | |
| 7 | PMTernary Bonsai 2 27B | Prism ML | 89.09 |
| 8 | JEMellum2-12B-A2.5B-Thinking | JetBrains | 86.2 |
| 9 | OPMiniCPM5-2B | OpenBMB | 84.7 |
| 10 | JEMellum2-12B-A2.5B-Instruct | JetBrains | 78.1 |
| 11 | OPMiniCPM5-1B | OpenBMB | 70.06 |
Evidence key: Observed
Rows are ordered by the value Qwen3.6 launch benchmarks published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.