Skip to main content
ModelScale

Knowledge benchmark

MMLU-Pro leaderboard

Massive Multitask Language Understanding Professional. Every model the catalog carries a published MMLU-Pro value for, ranked by that value.

CategoryKnowledge
Measure10-way multiple choice
TasksMultiple subjects
DifficultyProfessional level

An enhanced version of MMLU with 10 answer choices instead of 4, featuring more reasoning-focused questions that better differentiate frontier models.

MMLU-Pro ranking

46 models with a published MMLU-Pro value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Massive Multitask Language Understanding Professional value
RankModelProvider10-way multiple choice
1Qwen3.7 MaxAlibaba89.6
2Claude Opus 4.5Anthropic89.5
3Qwen3.6 PlusAlibaba88.5
3Qwen3.7 PlusAlibaba88.5
5Qwen3.5 397BAlibaba87.8
6DeepSeek V4 Pro 0813DeepSeek87.5
7Kimi K2.5Moonshot AI87.1
7Kimi K2.5 (Reasoning)Moonshot AI87.1
9Nemotron 3 UltraNVIDIA86.8
10Qwen3.5-122B-A10BAlibaba86.7
11Solar Pro 4Upstage86.3
12DeepSeek V4 Flash 0731DeepSeek86.2
12Qwen3.6-27BAlibaba86.2
12Solar Open 2Upstage86.2
15Qwen3.5-27BAlibaba86.1
16GLM-5Z.AI85.7
17Qwen3.5-35B-A3BAlibaba85.3
18Gemma 4 31BGoogle85.2
18Qwen3.6-35B-A3BAlibaba85.2
20MAI-Thinking-1Microsoft85
21MiMo-V2-FlashXiaomi84.9
22GLM-4.7Z.AI84.3
23K-EXAONE 2.0LG AI Research83.5
24Qwen3 235B 2507Alibaba83
25Gemma 4 26B A4BGoogle82.6
26Claude Opus 4.6Anthropic82
27Exaone 4.0 32BLG AI Research81.8
28Nemotron 3.5 Lightning 30B A3B NVFP4NVIDIA81.62
29LongCat-Flash-Lite-SparseMeituan79.24
30Claude Sonnet 4.6Anthropic79.2
31Granite 4.2 30BIBM77.6
32Nemotron 3 Nano Omni 30B A3BNVIDIA77.3
33Gemma 4 12BGoogle77.2
34CECeleris-1Celeris75.9
34DeepSeek V3DeepSeek75.9
36ZYZAYA1-8BZyphra74.2
37Granite 4.2 8BIBM74.04
38OPMiniCPM5-2BOpenBMB70.8
39Gemma 4 E4BGoogle69.4
40ZYZAYA1-74B-PreviewZyphra68.1
41Granite 4.2 3BIBM67.84
42Gemma 4 E2BGoogle60
43SPSoofi S 30B-A3BSoofi Project51.4
44OPMiniCPM5-1BOpenBMB48.85
45LFM2.5-230MLiquidAI20.25
46LFM2.5-VL-450MLiquidAI19.32

Evidence key: Observed

Rows are ordered by the value MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Knowledge capability leaderboard