Skip to main content
ModelScale

Multimodal & Grounded benchmark

MathVision leaderboard

Every model the catalog carries a published MathVision value for, ranked by that value.

CategoryMultimodal & Grounded
MeasureImage + math reasoning
TasksVisually grounded math problems
DifficultyAdvanced multimodal mathematics

A visual mathematics benchmark that tests whether a model can solve math problems grounded in diagrams, equations, figures, and other visual inputs.

MathVision ranking

19 models with a published MathVision value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published MathVision value
RankModelProviderImage + math reasoning
1Qwen3.8 MaxAlibaba95.2
2Kimi K3Moonshot AI94.3
3Seed 2.1 ProByteDance92.6
4Qwen3.8-Omni-FlashAlibaba91.8
5Qwen3.8-Flash-NextAlibaba90.6
6Qwen3.7 PlusAlibaba90.3
7Seed 2.1 TurboByteDance90.1
8Qwen3.8-27BAlibaba90.0
9Qwen3.5 397BAlibaba88.6
10Qwen3.6 PlusAlibaba88.0
11DSdots3-note PreviewDots Studio87.7
12Kimi K2.6Moonshot AI87.4
13Gemini 3 ProGoogle86.6
14Qwen3.5-122B-A10BAlibaba86.2
15Qwen3.5-27BAlibaba86.0
16Qwen3.5-35B-A3BAlibaba83.9
17GPT-5.2OpenAI83.0
18Gemma 4 12BGoogle79.7
19Claude Opus 4.5Anthropic74.3

Evidence key: Observed

Rows are ordered by the value Qwen3.6 launch benchmarks published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards