Skip to main content
ModelScale

Multimodal & Grounded benchmark

RefCOCO (avg) leaderboard

RefCOCO average. Every model the catalog carries a published RefCOCO (avg) value for, ranked by that value.

CategoryMultimodal & Grounded
MeasureGrounded visual localization
TasksReferring-expression grounding
DifficultyFine-grained visual grounding

A referring-expression grounding benchmark averaged across RefCOCO variants to test whether a model can localize described objects correctly.

RefCOCO (avg) ranking

6 models with a published RefCOCO (avg) value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published RefCOCO average value
RankModelProviderGrounded visual localization
1Qwen3.6-27BAlibaba92.5
2Qwen3.6-35B-A3BAlibaba92.0
3Nemotron 3 Nano Omni 30B A3BNVIDIA90.5
4LFM2.5-VL-3BLiquidAI87.9
5ZYZAYA1-VL-8BZyphra84.3
6INInterfaze BetaInterfaze82.1

Evidence key: Observed

Rows are ordered by the value RefCOCO referring expression datasets published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards