Skip to main content
ModelScale

Multimodal & Grounded benchmark

SimpleVQA leaderboard

Every model the catalog carries a published SimpleVQA value for, ranked by that value.

CategoryMultimodal & Grounded
MeasureImage-grounded question answering
TasksVisual QA tasks
DifficultyGeneral visual understanding
Published byGLM-5V-Turbo

A visual question answering benchmark focused on straightforward image-grounded understanding.

SimpleVQA ranking

11 models with a published SimpleVQA value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published SimpleVQA value
RankModelProviderImage-grounded question answering
1Qwen3.7 PlusAlibaba81.7
2Step 3.7 FlashStepFun79.2
3Qwen3.8 MaxAlibaba75.0
4DSdots3-note PreviewDots Studio72.5
5Gemini 3.1 ProGoogle72.4
6Muse SparkMeta71.3
7GPT-5.4OpenAI61.1
8Qwen3.6-35B-A3BAlibaba58.9
9Grok 4.20xAI57.4
10Qwen3.6-27BAlibaba56.1
11LFM2.5-VL-3BLiquidAI35.4

Evidence key: Observed

Rows are ordered by the value GLM-5V-Turbo published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards