Skip to main content
ModelScale

Multimodal & Grounded benchmark

ChatCVQA leaderboard

Every model the catalog carries a published ChatCVQA value for, ranked by that value.

CategoryMultimodal & Grounded
MeasureMulti-turn image-grounded QA
TasksConversational visual QA
DifficultyConversational multimodal reasoning

A conversational visual QA benchmark that tests multi-turn grounded answering over images and documents.

No model in the catalog has a published ChatCVQA score.

The benchmark is defined by Qwen3.6 launch benchmarks, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

All leaderboards