Multimodal & Grounded benchmark
SimpleVQA leaderboard
Every model the catalog carries a published SimpleVQA value for, ranked by that value.
A visual question answering benchmark focused on straightforward image-grounded understanding.
SimpleVQA ranking
11 models with a published SimpleVQA value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Image-grounded question answering |
|---|---|---|---|
| 1 | Alibaba | 81.7 | |
| 2 | StepFun | 79.2 | |
| 3 | Alibaba | 75.0 | |
| 4 | DSdots3-note Preview | Dots Studio | 72.5 |
| 5 | 72.4 | ||
| 6 | Meta | 71.3 | |
| 7 | OpenAI | 61.1 | |
| 8 | Alibaba | 58.9 | |
| 9 | xAI | 57.4 | |
| 10 | Alibaba | 56.1 | |
| 11 | LiquidAI | 35.4 |
Evidence key: Observed
Rows are ordered by the value GLM-5V-Turbo published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.