Skip to main content
ModelScale

Knowledge benchmark

HLE leaderboard

Humanity's Last Exam. Every model the catalog carries a published HLE value for, ranked by that value.

CategoryKnowledge
MeasureOpen-ended and multiple choice
TasksExpert-level questions
DifficultyFrontier expert level

An expert-authored benchmark designed to probe frontier knowledge and reasoning. BenchLM keeps protocol differences visible because tool-assisted and closed-book HLE runs answer different questions.

HLE ranking

58 models with a published HLE value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Humanity's Last Exam value
RankModelProviderOpen-ended and multiple choice
1Claude Fable 5.1Anthropic65
2Claude Opus 5Anthropic64.7
3Claude Mythos 5Anthropic64.5
4Muse Spark 1.1Meta62.1
5GPT-5.4 ProOpenAI58.7
6Claude Opus 4.8Anthropic57.9
7Claude Sonnet 5Anthropic57.4
8GPT-5.5 ProOpenAI57.2
9APApodex 1.1Apodex56.1
10Kimi K3Moonshot AI56
11Hy4 previewTencent55.4
12Claude Opus 4.7 (Adaptive)Anthropic54.7
12GLM-5.2Z.AI54.7
14Claude Opus 4.6Anthropic53
15DSdots3-note PreviewDots Studio52.6
16GLM-5.1Z.AI52.3
17GPT-5.5OpenAI52.2
18GPT-5.4OpenAI52.1
19GLM-5Z.AI50.4
19Muse SparkMeta50.4
21Claude Sonnet 4.6Anthropic49
22MiMo-V2.5-ProXiaomi48
23TMInkling-SmallThinking Machines Lab47.8
24Agents-A1InternScience47.6
25TMInklingThinking Machines Lab46
26OAOrnith-1.5-397BOrnith AI44.6
27Qwen3.8 MaxAlibaba43.6
28DeepSeek V4 Pro 0813DeepSeek42.7
29GPT-5.4 miniOpenAI41.5
30Qwen3.7 MaxAlibaba41.4
31Gemini 3.5 FlashGoogle40.2
32GPT-5.4 nanoOpenAI37.7
33DeepSeek V4.1 FlashDeepSeek36.8
34Qwen3.8-Omni-FlashAlibaba36.5
35Qwen3.8-Flash-NextAlibaba35.9
36Grok 4.3xAI35
37DeepSeek V4 Flash 0731DeepSeek34.8
38Kimi K2.6Moonshot AI34.7
38Qwen3.7 PlusAlibaba34.7
40Claude Opus 4.5Anthropic30.8
40Qwen3.8-27BAlibaba30.8
42Kimi K2.5Moonshot AI30.1
43Qwen3.6 PlusAlibaba28.8
44Qwen3.5 397BAlibaba28.7
45STA.X K2SK Telecom27.8
46Nemotron 3 UltraNVIDIA26.7
47Gemma 4 31BGoogle26.5
48OAOrnith-1.5-35B-A3BOrnith AI25.6
49Hy3 PreviewTencent25.5
50GLM-4.7Z.AI24.8
51Qwen3.6-27BAlibaba24
52Ling 3.0 FlashInclusionAI22.7
53Qwen3.6-35B-A3BAlibaba21.4
54OAOrnith-1.5-9BOrnith AI20.2
55Gemini 2.5 ProGoogle18.8
56K-EXAONE 2.0LG AI Research18.3
57Gemma 4 26B A4BGoogle17.2
58OPMiniCPM5-2BOpenBMB8.9

Evidence key: Observed

Rows are ordered by the value Humanity's Last Exam published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Knowledge capability leaderboard