Skip to main content
ModelScale

Knowledge benchmark

MMLU-Pro (Vals) leaderboard

MMLU-Pro, Vals AI run. Every model the catalog carries a published MMLU-Pro (Vals) value for, ranked by that value.

CategoryKnowledge
MeasureAccuracy
TasksAcademic multiple-choice questions
DifficultyBroad academic knowledge

Vals AI’s independent run of MMLU-Pro across fourteen academic subjects.

MMLU-Pro (Vals) ranking

51 models with a published MMLU-Pro (Vals) value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published MMLU-Pro, Vals AI run value
RankModelProviderAccuracy
1Claude Fable 5.1Anthropic92.4
2Claude Opus 5Anthropic91.6
3Claude Fable 5Anthropic91.5
4Gemini 3.1 ProGoogle91.0
5Gemini 3.8 FlashGoogle90.2
6Gemini 3.7 FlashGoogle90.1
7Claude Opus 4.7Anthropic89.9
8Claude Opus 4.8Anthropic89.6
9Gemini 3.5 FlashGoogle89.5
10Grok 4.6xAI89.4
11Gemini 3.6 FlashGoogle89.3
11Qwen3.7 MaxAlibaba89.3
13Grok 4.5xAI89.2
14GPT-5.6 SolOpenAI89.1
15Muse Spark 1.1Meta88.7
16Gemini 3 FlashGoogle88.6
16Qwen3.8 MaxAlibaba88.6
18Muse Spark 1.2Meta88.3
19GPT-5.5OpenAI88.1
20Kimi K3Moonshot AI88.0
21Qwen3.6 PlusAlibaba87.7
22Kimi K2.6Moonshot AI87.6
23Claude Sonnet 5Anthropic87.5
24Claude Sonnet 4.6Anthropic87.3
25DeepSeek V4 Pro 0813DeepSeek87.0
26GLM-5.1Z.AI86.9
27GLM-5.3Z.AI86.8
28GLM-5.2Z.AI86.7
28GPT-5.6 TerraOpenAI86.7
30Grok 4.20xAI86.3
30TMInklingThinking Machines Lab86.3
32DeepSeek V4 Flash 0731DeepSeek86.2
32Gemini 3.1 Flash-LiteGoogle86.2
34GLM-5.3-FlashZ.AI86.1
35GPT-5.6 LunaOpenAI86.0
36Gemini 3.5 Flash-LiteGoogle85.8
36Grok 4.3xAI85.8
36Nemotron 3 UltraNVIDIA85.8
39TMInkling-SmallThinking Machines Lab85.6
40GPT-5.4 miniOpenAI84.6
40MiMo-V2.5-ProXiaomi84.6
42Qwen3.8-27BAlibaba84.3
43MiniMax M3MiniMax84.2
44MiMo-V2.5Xiaomi82.9
45Ling 3.0 FlashInclusionAI82.0
46MiniMax M2.7MiniMax80.4
47Claude Haiku 4.5Anthropic78.7
48GPT-5.4 nanoOpenAI77.2
49Mistral Medium 3.5 128BMistral75.3
50Laguna XS.2Poolside69.1
51Laguna M.1Poolside68.8

Evidence key: Observed

Rows are ordered by the value Vals AI MMLU-Pro, Vals AI run leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Knowledge capability leaderboard