Skip to main content
ModelScale

Multimodal & Grounded benchmark

CharXiv w/o tools leaderboard

CharXiv Reasoning without tools. Every model the catalog carries a published CharXiv w/o tools value for, ranked by that value.

CategoryMultimodal & Grounded
MeasureChart understanding without tools
TasksScientific chart reasoning (tool-free)
DifficultyScientific visualization reasoning

Tool-free variant of CharXiv that isolates raw visual reasoning ability without code execution or tool augmentation.

CharXiv w/o tools ranking

14 models with a published CharXiv w/o tools value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published CharXiv Reasoning without tools value
RankModelProviderChart understanding without tools
1Claude Mythos 5Anthropic88.9
2Qwen3.8 MaxAlibaba88.4
3Gemini 3.8 FlashGoogle86.2
4Kimi K3Moonshot AI84.8
5Qwen3.8-Flash-NextAlibaba84.6
6Gemini 3.7 FlashGoogle84.5
7Qwen3.8-27BAlibaba83.7
8Qwen3.8-Omni-FlashAlibaba83.5
9DSdots3-note PreviewDots Studio83.1
10Claude Opus 4.7 (Adaptive)Anthropic82.1
11Claude Opus 4.8Anthropic80.5
12TMInklingThinking Machines Lab78.1
13TMInkling-SmallThinking Machines Lab77.4
14Claude Sonnet 5Anthropic77

Evidence key: Observed

Rows are ordered by the value CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards