Skip to main content
ModelScale

Multimodal & Grounded benchmark

MMMU-Pro leaderboard

Massive Multi-discipline Multimodal Understanding Pro. Every model the catalog carries a published MMMU-Pro value for, ranked by that value.

CategoryMultimodal & Grounded
MeasureImage + text question answering
TasksMultimodal academic reasoning
DifficultyFrontier multimodal

A harder multimodal benchmark for frontier models that combines text with images, diagrams, charts, and academic visual reasoning tasks.

MMMU-Pro ranking

41 models with a published MMMU-Pro value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Massive Multi-discipline Multimodal Understanding Pro value
RankModelProviderImage + text question answering
1Gemini 3.1 ProGoogle83.9
2Gemini 3.5 FlashGoogle83.6
3GPT-5.6 SolOpenAI83
4Qwen3.8 MaxAlibaba82.3
5Kimi K3Moonshot AI81.6
5Seed 2.1 ProByteDance81.6
7GPT-5.4OpenAI81.2
7GPT-5.5OpenAI81.2
9Gemini 3 ProGoogle81
10GPT-5.6 TerraOpenAI80.7
11Muse SparkMeta80.4
12Seed 2.1 TurboByteDance80.1
13GPT-5.2OpenAI79.5
14Kimi K2.6Moonshot AI79.4
15DSdots3-note PreviewDots Studio79.1
16Qwen3.5 397BAlibaba79
16Qwen3.7 PlusAlibaba79
18Qwen3.6 PlusAlibaba78.8
19Kimi K2.5Moonshot AI78.5
19Kimi K2.5 (Reasoning)Moonshot AI78.5
21GPT-5.6 LunaOpenAI78.4
22Grok 4.3xAI78.1
22MiniMax M3MiniMax78.1
24UNPareto 26.9Unbiased78
25MiMo-V2.5Xiaomi77.9
26Claude Opus 4.6Anthropic77.3
27Gemma 4 31BGoogle76.9
28GPT-5.4 miniOpenAI76.6
29Qwen3.6-27BAlibaba75.8
30Qwen3.6-35B-A3BAlibaba75.3
31Grok 4.20xAI75.2
32TMInkling-SmallThinking Machines Lab74
32Muse Glimmer 30BMeta74
34Gemma 4 26B A4BGoogle73.8
35TMInklingThinking Machines Lab73.5
36INInterfaze BetaInterfaze71.1
37Claude Opus 4.5Anthropic70.6
38Gemma 4 12BGoogle69.1
39GPT-5.4 nanoOpenAI66.1
40Command A+Cohere63
41LFM2.5-VL-3BLiquidAI30.5

Evidence key: Observed

Rows are ordered by the value MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards