Knowledge benchmark
MMLU leaderboard
Massive Multitask Language Understanding. Every model the catalog carries a published MMLU value for, ranked by that value.
A comprehensive multiple-choice question answering test covering 57 tasks including elementary mathematics, US history, computer science, law, and more. Tests knowledge across diverse academic subjects from high school to professional level.
MMLU ranking
7 models with a published MMLU value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Multiple choice questions |
|---|---|---|---|
| 1 | OpenAI | 91.8 | |
| 2 | OpenAI | 90.2 | |
| 3 | OpenAI | 87.5 | |
| 4 | Arcee AI | 87.2 | |
| 5 | OpenAI | 86.9 | |
| 6 | Meituan | 85.31 | |
| 7 | OpenAI | 80.1 |
Evidence key: Observed
Rows are ordered by the value Measuring Massive Multitask Language Understanding published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.