Skip to main content
ModelScale

Multimodal & Grounded benchmark

OmniDocBench 1.5 leaderboard

Every model the catalog carries a published OmniDocBench 1.5 value for, ranked by that value.

CategoryMultimodal & Grounded
MeasureDocument understanding benchmark
TasksDocument understanding tasks
DifficultyGrounded document reasoning

A document understanding benchmark used in frontier-model comparison tables to measure extraction and grounded reasoning quality on complex documents.

OmniDocBench 1.5 ranking

6 models with a published OmniDocBench 1.5 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published OmniDocBench 1.5 value
RankModelProviderDocument understanding benchmark
1Qwen3.8 MaxAlibaba92.1
2MiniMax M3MiniMax91.6
3Qwen3.7 PlusAlibaba91.4
4Qwen3.8-27BAlibaba91.1
5Qwen3.6-35B-A3BAlibaba89.9
6Muse Glimmer 30BMeta75.8

Evidence key: Observed

Rows are ordered by the value Introducing GPT-5.4 mini and nano published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards