Skip to main content
ModelScale

Knowledge benchmark

MMLU-Redux leaderboard

Every model the catalog carries a published MMLU-Redux value for, ranked by that value.

CategoryKnowledge
MeasureMultiple choice questions
TasksBroad academic QA
DifficultyAdvanced general knowledge

A harder refresh of MMLU intended to keep broad knowledge evaluation useful after the original benchmark became too easy for frontier models.

MMLU-Redux ranking

11 models with a published MMLU-Redux value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published MMLU-Redux value
RankModelProviderMultiple choice questions
1Claude Opus 4.5Anthropic96.6
2Qwen3.7 MaxAlibaba95
3Qwen3.5 397BAlibaba94.9
4Qwen3.6 PlusAlibaba94.5
4Qwen3.7 PlusAlibaba94.5
6Qwen3.6-27BAlibaba93.5
7PMTernary Bonsai 2 27BPrism ML89.09
8JEMellum2-12B-A2.5B-ThinkingJetBrains86.2
9OPMiniCPM5-2BOpenBMB84.7
10JEMellum2-12B-A2.5B-InstructJetBrains78.1
11OPMiniCPM5-1BOpenBMB70.06

Evidence key: Observed

Rows are ordered by the value Qwen3.6 launch benchmarks published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Knowledge capability leaderboard