Skip to main content
ModelScale

Multimodal & Grounded benchmark

V* leaderboard

Every model the catalog carries a published V* value for, ranked by that value.

CategoryMultimodal & Grounded
MeasureVision-centric reasoning benchmark
TasksFrontier multimodal reasoning tasks
DifficultyFrontier multimodal
Published byGLM-5V-Turbo

A vision-centric benchmark for high-level multimodal reasoning and perception quality.

V* ranking

11 models with a published V* value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published V* value
RankModelProviderVision-centric reasoning benchmark
1Kimi K2.6Moonshot AI96.9
1Qwen3.6 PlusAlibaba96.9
3Qwen3.5 397BAlibaba95.8
4Step 3.7 FlashStepFun95.3
5Qwen3.6-27BAlibaba94.7
6Qwen3.5-27BAlibaba93.7
7Qwen3.5-122B-A10BAlibaba93.2
8Qwen3.5-35B-A3BAlibaba92.7
9Gemini 3 ProGoogle88.0
10GPT-5.2OpenAI75.9
11Claude Opus 4.5Anthropic67.0

Evidence key: Observed

Rows are ordered by the value GLM-5V-Turbo published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards