Skip to main content
ModelScale

Knowledge benchmark

MMLU leaderboard

Massive Multitask Language Understanding. Every model the catalog carries a published MMLU value for, ranked by that value.

CategoryKnowledge
MeasureMultiple choice questions
Tasks57 subjects
DifficultyElementary to professional level

A comprehensive multiple-choice question answering test covering 57 tasks including elementary mathematics, US history, computer science, law, and more. Tests knowledge across diverse academic subjects from high school to professional level.

MMLU ranking

7 models with a published MMLU value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Massive Multitask Language Understanding value
RankModelProviderMultiple choice questions
1o1OpenAI91.8
2GPT-4.1OpenAI90.2
3GPT-4.1 miniOpenAI87.5
4Trinity-Large-PreviewArcee AI87.2
5o3-miniOpenAI86.9
6LongCat-Flash-Lite-SparseMeituan85.31
7GPT-4.1 nanoOpenAI80.1

Evidence key: Observed

Rows are ordered by the value Measuring Massive Multitask Language Understanding published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Knowledge capability leaderboard