Skip to main content
ModelScale

Multimodal & Grounded benchmark

CharXiv leaderboard

CharXiv Reasoning. Every model the catalog carries a published CharXiv value for, ranked by that value.

CategoryMultimodal & Grounded
MeasureChart understanding and reasoning
TasksScientific chart reasoning
DifficultyScientific visualization reasoning

A scientific chart reasoning benchmark that tests whether models can understand, interpret, and reason about complex scientific visualizations including plots, diagrams, and data charts.

CharXiv ranking

39 models with a published CharXiv value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published CharXiv Reasoning value
RankModelProviderChart understanding and reasoning
1Claude Mythos 5Anthropic93.5
1Qwen3.8 MaxAlibaba93.5
3Qwen3.8-Omni-FlashAlibaba91.4
4Kimi K3Moonshot AI91.3
5Claude Opus 4.7 (Adaptive)Anthropic91
6Qwen3.8-Flash-NextAlibaba90.6
7Qwen3.8-27BAlibaba90.2
8Claude Opus 4.8Anthropic89.9
9GLM-5.3-FlashZ.AI89.4
10Gemini 3.7 FlashGoogle88.7
11Muse Spark 1.1Meta88.4
12Claude Sonnet 5Anthropic88.3
13SASakana Fugu-UltraSakana AI86.6
14Muse SparkMeta86.4
15Qwen3.7 PlusAlibaba85.9
16Seed 2.1 ProByteDance85.4
17SASakana FuguSakana AI85.1
18Gemini 3.5 FlashGoogle84.2
19GPT-5.4OpenAI82.8
20Seed 2.1 TurboByteDance82.5
21GPT-5.2OpenAI82.1
22TMInklingThinking Machines Lab82
23Qwen3.6 PlusAlibaba81.5
24Gemini 3 ProGoogle81.4
25TMInkling-SmallThinking Machines Lab81.3
26MiMo-V2.5Xiaomi81
27Qwen3.5 397BAlibaba80.8
28Kimi K2.6Moonshot AI80.4
29Gemini 3.1 ProGoogle80.2
30Muse Glimmer 30BMeta78.8
31Qwen3.6-27BAlibaba78.4
32Qwen3.6-35B-A3BAlibaba78
33Claude Sonnet 4.6Anthropic77.4
34Qwen3.5-122B-A10BAlibaba77.2
35Nemotron 3 Nano Omni 30B A3BNVIDIA76.25
36Gemini 3.1 Flash-LiteGoogle73.2
37Claude Opus 4.5Anthropic68.5
38Grok 4.20xAI60.9
39Command A+Cohere52.7

Evidence key: Observed

Rows are ordered by the value CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards