Multimodal & Grounded benchmark
MathVision leaderboard
Every model the catalog carries a published MathVision value for, ranked by that value.
A visual mathematics benchmark that tests whether a model can solve math problems grounded in diagrams, equations, figures, and other visual inputs.
MathVision ranking
19 models with a published MathVision value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Image + math reasoning |
|---|---|---|---|
| 1 | Alibaba | 95.2 | |
| 2 | Moonshot AI | 94.3 | |
| 3 | ByteDance | 92.6 | |
| 4 | Alibaba | 91.8 | |
| 5 | Alibaba | 90.6 | |
| 6 | Alibaba | 90.3 | |
| 7 | ByteDance | 90.1 | |
| 8 | Alibaba | 90.0 | |
| 9 | Alibaba | 88.6 | |
| 10 | Alibaba | 88.0 | |
| 11 | DSdots3-note Preview | Dots Studio | 87.7 |
| 12 | Moonshot AI | 87.4 | |
| 13 | 86.6 | ||
| 14 | Alibaba | 86.2 | |
| 15 | Alibaba | 86.0 | |
| 16 | Alibaba | 83.9 | |
| 17 | OpenAI | 83.0 | |
| 18 | 79.7 | ||
| 19 | Anthropic | 74.3 |
Evidence key: Observed
Rows are ordered by the value Qwen3.6 launch benchmarks published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.