Skip to main content
ModelScale

Multimodal & Grounded benchmark

AA-MMMU-Pro leaderboard

Artificial Analysis MMMU-Pro. Every model the catalog carries a published AA-MMMU-Pro value for, ranked by that value.

CategoryMultimodal & Grounded
MeasureImage + text question answering
TasksMultimodal academic reasoning
DifficultyFrontier multimodal

A display-only Artificial Analysis MMMU-Pro score.

AA-MMMU-Pro ranking

97 models with a published AA-MMMU-Pro value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Artificial Analysis MMMU-Pro value
RankModelProviderImage + text question answering
1GPT-6 AstraOpenAI86.9
2Gemini 3.8 FlashGoogle85.6
3Gemini 3.7 FlashGoogle85.5
4Claude Opus 5Anthropic84.7
5Gemini 3.5 FlashGoogle84.3
6GPT-5.6 SolOpenAI83.4
7Gemini 3.6 FlashGoogle83.2
8Qwen3.8 Max PreviewAlibaba82.8
9Gemini 3.1 ProGoogle82.4
10GPT-5.6 TerraOpenAI80.7
11Kimi K3Moonshot AI80.5
11Muse SparkMeta80.5
11Qwen3.7 PlusAlibaba80.5
14Grok 4.5xAI80.4
15Gemini 3 ProGoogle80.2
16GPT-5.5OpenAI79.9
17Qwen3.8-Flash-NextAlibaba79.8
18Kimi K2.6Moonshot AI79.4
19APApodex 1.1Apodex79.2
19APApodex 1.1 MiniApodex79.2
21Gemini 3.5 Flash-LiteGoogle79.0
21Ling 3.0 Flash VLInclusionAI79.0
23Claude Opus 4.7 (Adaptive)Anthropic78.8
24Gemini 3 FlashGoogle78.6
24GPT-5.6 LunaOpenAI78.6
24MiniMax M3MiniMax78.6
27GPT-5.3 CodexOpenAI78.5
28GPT-5.4OpenAI78.4
29Grok 4.3xAI78.1
30Qwen3.6 PlusAlibaba78.0
31Claude Sonnet 5Anthropic77.3
32DeepSeek V4.1 FlashDeepSeek77.0
33Claude Opus 4.7Anthropic76.4
34GPT-5.2-CodexOpenAI76.3
34Qwen3.8-27BAlibaba76.3
36GPT-5.1OpenAI75.5
37Claude Opus 4.6 (Adaptive)Anthropic75.4
37Kimi K2.5Moonshot AI75.4
37Kimi K2.5 (Reasoning)Moonshot AI75.4
40Step 3.7 FlashStepFun75.3
41Qwen3.5-122B-A10BAlibaba75.0
41Qwen3.5-27BAlibaba75.0
41Qwen3.6-35B-A3BAlibaba75.0
44Gemini 2.5 ProGoogle74.9
45Qwen3.6-27BAlibaba74.6
46GPT-5 (medium)OpenAI74.3
46Muse Glimmer 30BMeta74.3
48GPT-5 (high)OpenAI74.2
49Claude Opus 4.5 ThinkingAnthropic74.0
49TMInkling-SmallThinking Machines Lab74.0
51TMInklingThinking Machines Lab73.5
52Gemma 4 31BGoogle73.4
53GPT-5.4 miniOpenAI73.3
54GLM-5V-TurboZ.AI72.8
55Qwen3.5-35B-A3BAlibaba72.7
56Claude Opus 4.6Anthropic72.5
56GPT-5.1-CodexOpenAI72.5
56GPT-5.1-Codex-MaxOpenAI72.5
59Claude Opus 4.5Anthropic71.2
60Claude Sonnet 4.6Anthropic70.6
61o3OpenAI70.1
62MiMo-V2-OmniXiaomi69.9
63Gemma 4 12BGoogle69.7
64Gemma 4 26B A4BGoogle69.2
65Grok 4xAI68.8
66Claude 4.1 Opus ThinkingAnthropic67.9
67Gemini 2.5 FlashGoogle65.5
68GPT-5.4 nanoOpenAI65.4
69Mistral Medium 3.5 128BMistral64.9
70Grok 4.1 Fast (Reasoning)xAI63.3
71Command A+Cohere63.2
72Claude 4 SonnetAnthropic62.4
73Llama 4 MaverickMeta62.1
74Grok 4 Fast (Reasoning)xAI61.8
75GPT-4.1OpenAI61.2
76Qwen3-Omni-30B-A3B-ThinkingAlibaba60.2
77GPT-4.1 miniOpenAI58.7
78Mistral Small 4Mistral56.8
78Mistral Small 4 (Reasoning)Mistral56.8
80Mistral Large 3Mistral55.7
81Qwen3-Omni-30B-A3B-InstructAlibaba55.5
82Gemini 1.5 ProGoogle55.0
83Nemotron 3 Nano Omni 30B A3BNVIDIA53.2
84Mistral Medium 3Mistral53.0
85Llama 4 ScoutMeta52.9
86Qwen3.5 397BAlibaba52.7
86Qwen3.5 397B (Reasoning)Alibaba52.7
88Gemma 4 E4BGoogle51.4
89Grok 4.1 FastxAI48.4
90Gemma 3 27BGoogle48.0
91Gemma 4 E2BGoogle44.6
92Nova ProAmazon44.3
93GPT-4o miniOpenAI41.5
94GPT-4.1 nanoOpenAI40.1
95Claude 3 HaikuAnthropic30.8
96LFM2.5-VL-1.6B-ExtractLiquidAI26.5
97Phi-4 Multimodal InstructMicrosoft14.5

Evidence key: Observed

Rows are ordered by the value Artificial Analysis MMMU-Pro Benchmark Leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards