Skip to main content
ModelScale

Multimodal & Grounded benchmark

BabyVision leaderboard

Every model the catalog carries a published BabyVision value for, ranked by that value.

CategoryMultimodal & Grounded
MeasureMultimodal visual reasoning
TasksVisual perception tasks
DifficultyFine-grained visual perception

A multimodal benchmark for fine-grained visual perception and grounded reasoning tasks.

BabyVision ranking

7 models with a published BabyVision value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published BabyVision value
RankModelProviderMultimodal visual reasoning
1Qwen3.8 MaxAlibaba82.0
2Muse Spark 1.1Meta76.3
3Seed 2.1 ProByteDance73.7
4Qwen3.8-27BAlibaba65.7
5Seed 2.1 TurboByteDance62.9
6GLM-5.3-FlashZ.AI53.4
7DSdots3-note PreviewDots Studio50.0

Evidence key: Observed

Rows are ordered by the value Muse Spark 1.1 Evaluation Report published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards