Multimodal & Grounded benchmark
MMVU leaderboard
Multimodal Multi-disciplinary Video Understanding. Every model the catalog carries a published MMVU value for, ranked by that value.
A benchmark for evaluating multimodal models on video understanding tasks across multiple disciplines, emphasizing temporal reasoning and comprehension over video content.
MMVU ranking
7 models with a published MMVU value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Video reasoning benchmark |
|---|---|---|---|
| 1 | Alibaba | 82.4 | |
| 2 | Z.AI | 80.5 | |
| 3 | Moonshot AI | 80.4 | |
| 4 | DSdots3-note Preview | Dots Studio | 79.9 |
| 5 | Alibaba | 74.7 | |
| 6 | Alibaba | 73.3 | |
| 7 | Alibaba | 72.3 |
Evidence key: Observed
Rows are ordered by the value Kimi K2.5 benchmark release surface published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.