Skip to main content
ModelScale

Multilingual benchmark

MMLU-ProX leaderboard

Every model the catalog carries a published MMLU-ProX value for, ranked by that value.

CategoryMultilingual
MeasureMultilingual multiple choice
TasksMultilingual professional QA
DifficultyProfessional multilingual

A multilingual extension of professional-level academic evaluation across many languages.

MMLU-ProX ranking

12 models with a published MMLU-ProX value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published MMLU-ProX value
RankModelProviderMultilingual multiple choice
1Qwen3.7 MaxAlibaba87
2Claude Opus 4.5Anthropic85.7
3Qwen3.7 PlusAlibaba85.4
4Qwen3.5 397BAlibaba84.7
4Qwen3.6 PlusAlibaba84.7
6GLM-5Z.AI83.1
7Nemotron 3 UltraNVIDIA83
8Kimi K2.5Moonshot AI82.3
9Qwen3.5-122B-A10BAlibaba82.2
9Qwen3.5-27BAlibaba82.2
11Qwen3.5-35B-A3BAlibaba81
12Qwen3 235B 2507Alibaba79.4

Evidence key: Observed

Rows are ordered by the value MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards