Multimodal & Grounded benchmark
CharXiv leaderboard
CharXiv Reasoning. Every model the catalog carries a published CharXiv value for, ranked by that value.
A scientific chart reasoning benchmark that tests whether models can understand, interpret, and reason about complex scientific visualizations including plots, diagrams, and data charts.
CharXiv ranking
39 models with a published CharXiv value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Chart understanding and reasoning |
|---|---|---|---|
| 1 | Anthropic | 93.5 | |
| 1 | Alibaba | 93.5 | |
| 3 | Alibaba | 91.4 | |
| 4 | Moonshot AI | 91.3 | |
| 5 | Anthropic | 91 | |
| 6 | Alibaba | 90.6 | |
| 7 | Alibaba | 90.2 | |
| 8 | Anthropic | 89.9 | |
| 9 | Z.AI | 89.4 | |
| 10 | 88.7 | ||
| 11 | Meta | 88.4 | |
| 12 | Anthropic | 88.3 | |
| 13 | SASakana Fugu-Ultra | Sakana AI | 86.6 |
| 14 | Meta | 86.4 | |
| 15 | Alibaba | 85.9 | |
| 16 | ByteDance | 85.4 | |
| 17 | SASakana Fugu | Sakana AI | 85.1 |
| 18 | 84.2 | ||
| 19 | OpenAI | 82.8 | |
| 20 | ByteDance | 82.5 | |
| 21 | OpenAI | 82.1 | |
| 22 | TMInkling | Thinking Machines Lab | 82 |
| 23 | Alibaba | 81.5 | |
| 24 | 81.4 | ||
| 25 | TMInkling-Small | Thinking Machines Lab | 81.3 |
| 26 | Xiaomi | 81 | |
| 27 | Alibaba | 80.8 | |
| 28 | Moonshot AI | 80.4 | |
| 29 | 80.2 | ||
| 30 | Meta | 78.8 | |
| 31 | Alibaba | 78.4 | |
| 32 | Alibaba | 78 | |
| 33 | Anthropic | 77.4 | |
| 34 | Alibaba | 77.2 | |
| 35 | NVIDIA | 76.25 | |
| 36 | 73.2 | ||
| 37 | Anthropic | 68.5 | |
| 38 | xAI | 60.9 | |
| 39 | Cohere | 52.7 |
Evidence key: Observed
Rows are ordered by the value CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.