Knowledge benchmark
HLE leaderboard
Humanity's Last Exam. Every model the catalog carries a published HLE value for, ranked by that value.
An expert-authored benchmark designed to probe frontier knowledge and reasoning. BenchLM keeps protocol differences visible because tool-assisted and closed-book HLE runs answer different questions.
HLE ranking
58 models with a published HLE value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Open-ended and multiple choice |
|---|---|---|---|
| 1 | Anthropic | 65 | |
| 2 | Anthropic | 64.7 | |
| 3 | Anthropic | 64.5 | |
| 4 | Meta | 62.1 | |
| 5 | OpenAI | 58.7 | |
| 6 | Anthropic | 57.9 | |
| 7 | Anthropic | 57.4 | |
| 8 | OpenAI | 57.2 | |
| 9 | APApodex 1.1 | Apodex | 56.1 |
| 10 | Moonshot AI | 56 | |
| 11 | Tencent | 55.4 | |
| 12 | Anthropic | 54.7 | |
| 12 | Z.AI | 54.7 | |
| 14 | Anthropic | 53 | |
| 15 | DSdots3-note Preview | Dots Studio | 52.6 |
| 16 | Z.AI | 52.3 | |
| 17 | OpenAI | 52.2 | |
| 18 | OpenAI | 52.1 | |
| 19 | Z.AI | 50.4 | |
| 19 | Meta | 50.4 | |
| 21 | Anthropic | 49 | |
| 22 | Xiaomi | 48 | |
| 23 | TMInkling-Small | Thinking Machines Lab | 47.8 |
| 24 | InternScience | 47.6 | |
| 25 | TMInkling | Thinking Machines Lab | 46 |
| 26 | OAOrnith-1.5-397B | Ornith AI | 44.6 |
| 27 | Alibaba | 43.6 | |
| 28 | DeepSeek | 42.7 | |
| 29 | OpenAI | 41.5 | |
| 30 | Alibaba | 41.4 | |
| 31 | 40.2 | ||
| 32 | OpenAI | 37.7 | |
| 33 | DeepSeek | 36.8 | |
| 34 | Alibaba | 36.5 | |
| 35 | Alibaba | 35.9 | |
| 36 | xAI | 35 | |
| 37 | DeepSeek | 34.8 | |
| 38 | Moonshot AI | 34.7 | |
| 38 | Alibaba | 34.7 | |
| 40 | Anthropic | 30.8 | |
| 40 | Alibaba | 30.8 | |
| 42 | Moonshot AI | 30.1 | |
| 43 | Alibaba | 28.8 | |
| 44 | Alibaba | 28.7 | |
| 45 | STA.X K2 | SK Telecom | 27.8 |
| 46 | NVIDIA | 26.7 | |
| 47 | 26.5 | ||
| 48 | OAOrnith-1.5-35B-A3B | Ornith AI | 25.6 |
| 49 | Tencent | 25.5 | |
| 50 | Z.AI | 24.8 | |
| 51 | Alibaba | 24 | |
| 52 | InclusionAI | 22.7 | |
| 53 | Alibaba | 21.4 | |
| 54 | OAOrnith-1.5-9B | Ornith AI | 20.2 |
| 55 | 18.8 | ||
| 56 | LG AI Research | 18.3 | |
| 57 | 17.2 | ||
| 58 | OPMiniCPM5-2B | OpenBMB | 8.9 |
Evidence key: Observed
Rows are ordered by the value Humanity's Last Exam published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.