Skip to main content
ModelScale

Knowledge benchmark

AA-HLE leaderboard

Artificial Analysis Humanity's Last Exam. Every model the catalog carries a published AA-HLE value for, ranked by that value.

CategoryKnowledge
MeasureAccuracy
TasksExpert-level questions
DifficultyFrontier expert reasoning

A display-only Artificial Analysis Humanity's Last Exam score.

AA-HLE ranking

181 models with a published AA-HLE value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Artificial Analysis Humanity's Last Exam value
RankModelProviderAccuracy
1Claude Fable 5.1Anthropic59.1
2Claude Fable 5Anthropic55.5
3Claude Opus 5Anthropic54.9
4GPT-6 AstraOpenAI54.7
5GPT-5.6 SolOpenAI49.5
6Claude Opus 4.8Anthropic48.7
6Muse Spark 1.3Meta48.7
8Gemini 3.7 FlashGoogle47.9
9Gemini 3.8 FlashGoogle47.8
10Gemini 3.1 ProGoogle47.0
11Kimi K3Moonshot AI46.9
12Muse Spark 1.1Meta46.2
13GPT-5.5OpenAI45.8
14Muse Spark 1.2Meta45.5
15GPT-5.4OpenAI43.7
16Qwen3.8 Max PreviewAlibaba43.1
17GPT-5.6 TerraOpenAI42.9
17Grok 4.6xAI42.9
19Gemini 3.5 FlashGoogle42.7
19Grok 4.5xAI42.7
21GPT-5.3 CodexOpenAI42.5
22Claude Opus 4.7 (Adaptive)Anthropic42.3
22GLM-5.3Z.AI42.3
24Claude Sonnet 5Anthropic41.3
25GLM-5.2Z.AI41.1
26DeepSeek V4 Pro 0813DeepSeek41.0
27Gemini 3.6 FlashGoogle40.8
28Muse SparkMeta40.7
29Qwen3.7 MaxAlibaba40.5
30Claude Opus 4.6 (Adaptive)Anthropic39.9
30GLM-5.3-FlashZ.AI39.9
32Gemini 3 ProGoogle39.7
33GPT-5.6 LunaOpenAI39.5
34DeepSeek V4.1 FlashDeepSeek39.2
35MiniMax M3MiniMax39.0
36DeepSeek V4 Flash 0731DeepSeek38.6
37Qwen3.8-Flash-NextAlibaba38.0
38GPT-5.2OpenAI37.7
39Kimi K2.6Moonshot AI37.5
40Grok 4.3xAI37.2
41GPT-5.2-CodexOpenAI35.7
41MiMo-V2.5-ProXiaomi35.7
43Qwen3.7 PlusAlibaba35.6
44Kimi K2.7 CodeMoonshot AI35.0
45APApodex 1.1Apodex34.1
45APApodex 1.1 MiniApodex34.1
47Qwen3.8-27BAlibaba33.9
48Hy3Tencent33.5
48Hy3 PreviewTencent33.5
50Claude Opus 4.7Anthropic33.3
50TMInkling-SmallThinking Machines Lab33.3
52TMInklingThinking Machines Lab31.9
53Qwen 3.6 Max (preview)Alibaba30.8
54Kimi K2.5Moonshot AI30.7
54Kimi K2.5 (Reasoning)Moonshot AI30.7
56MiMo-V2-ProXiaomi30.4
57Claude Opus 4.5 ThinkingAnthropic30.1
57GLM-5.1Z.AI30.1
59STA.X K2SK Telecom29.6
59MiniMax M2.7MiniMax29.6
61GLM-5Z.AI29.3
62Solar Pro 4Upstage29.2
63GPT-5 (high)OpenAI28.5
63GPT-5.1OpenAI28.5
65Nemotron 3 UltraNVIDIA28.4
66GPT-5.4 nanoOpenAI28.3
67GPT-5.4 miniOpenAI28.1
68GLM-5-TurboZ.AI27.8
68Qwen3.6 PlusAlibaba27.8
70GLM-4.7Z.AI27.4
71Grok 4xAI26.7
72GPT-5.1-CodexOpenAI25.7
72GPT-5.1-Codex-MaxOpenAI25.7
74GPT-5 (medium)OpenAI25.4
75Qwen3.5-122B-A10BAlibaba25.2
76Qwen3.5-27BAlibaba23.9
77Ling 3.0 FlashInclusionAI23.7
77Ling 3.0 Flash FP8InclusionAI23.7
79Gemma 4 31BGoogle23.6
80Qwen3.6-27BAlibaba23.1
81Gemini 2.5 ProGoogle22.5
82Qwen3.6-35B-A3BAlibaba22.2
83MiMo-V2-OmniXiaomi22.1
84Ling 3.0 Flash VLInclusionAI22.0
84Muse Glimmer 30BMeta22.0
86Step 3.7 FlashStepFun21.4
87Qwen3.5-35B-A3BAlibaba21.0
88Nemotron 3 Super 100BNVIDIA20.8
89o3OpenAI20.1
90Qwen3.5 397BAlibaba19.8
90Qwen3.5 397B (Reasoning)Alibaba19.8
92GPT-OSS 120BOpenAI19.6
93Gemma 4 26B A4BGoogle19.3
93Grok 4.1 Fast (Reasoning)xAI19.3
95Claude Opus 4.6Anthropic19.1
95Grok 4 Fast (Reasoning)xAI19.1
97Gemini 3.5 Flash-LiteGoogle18.8
98MCQuasar 438BMultiverse Computing18.7
99K-EXAONE 2.0LG AI Research18.6
100GLM-5V-TurboZ.AI17.1
101DeepSeek-R1DeepSeek15.8
101Trinity-Large-PreviewArcee AI15.8
101Trinity-Large-ThinkingArcee AI15.8
104Gemma 4 12BGoogle15.7
105Gemini 3 FlashGoogle15.0
106DeepSeek V3.1 (Reasoning)DeepSeek14.3
107K-ExaoneLG AI Research13.9
108Mistral Medium 3.5 128BMistral13.8
109Claude Sonnet 4.6Anthropic13.3
110Claude Opus 4.5Anthropic13.2
111Claude 4.1 Opus ThinkingAnthropic12.5
112Command A+Cohere12.0
113Qwen3 MaxAlibaba11.9
114Nemotron 3 Nano 30BNVIDIA11.4
115DeepSeek V3.2DeepSeek11.2
115Granite 4.2 30BIBM11.2
117North Mini CodeCohere11.1
118GPT-OSS 20BOpenAI11.0
118SASarvam 105BSarvam11.0
120Nemotron 3.5 Lightning 30B A3B NVFP4NVIDIA10.6
121Solar Pro 3Upstage10.3
122Mistral Small 4Mistral9.9
122Mistral Small 4 (Reasoning)Mistral9.9
124Granite 4.2 8BIBM9.7
125Ling 3.0 TinyInclusionAI9.3
126OPMiniCPM5-2BOpenBMB8.9
127MiMo-V2-FlashXiaomi8.6
128Grok Code Fast 1xAI8.0
129o3-miniOpenAI7.9
130Qwen3-Omni-30B-A3B-ThinkingAlibaba7.5
130SASarvam 30BSarvam7.5
132Kimi K2Moonshot AI7.4
132Nemotron Ultra 253BNVIDIA7.4
134GLM-4.5-AirZ.AI7.0
134o1OpenAI7.0
136LFM2.5-8B-A1BLiquidAI6.9
137CECeleris-1Celeris6.8
138DeepSeek V3.1DeepSeek6.7
139Granite 4.2 3BIBM6.6
140Granite-4.0-H-350MIBM6.4
141Ling 2.6 FlashInclusionAI6.3
142LFM2.5-2.6BLiquidAI6.2
143Exaone 4.0 1.2BLG AI Research5.7
144GLM-4.6Z.AI5.5
144Granite-4.0-350MIBM5.5
146Grok 4.1 FastxAI5.1
146LFM2.5-VL-1.6B-ExtractLiquidAI5.1
148Exaone 4.0 32BLG AI Research5.0
148GPT-4.1 miniOpenAI5.0
148Granite-4.0-H-1BIBM5.0
148Phi-4 Multimodal InstructMicrosoft5.0
152Llama 4 MaverickMeta4.9
153Gemma 4 E2BGoogle4.8
153Granite-4.0-1BIBM4.8
153Nemotron 3 Nano Omni 30B A3BNVIDIA4.8
156Gemini 2.5 FlashGoogle4.7
157DeepSeek R1 Distill Qwen 32BDeepSeek4.6
157Gemini 1.5 ProGoogle4.6
157Qwen3-Omni-30B-A3B-InstructAlibaba4.6
160Gemma 3 27BGoogle4.4
161Claude 4 SonnetAnthropic4.3
162Gemini 1.0 ProGoogle4.2
162GPT-4.1OpenAI4.2
162GPT-4o miniOpenAI4.2
162Mistral Large 3Mistral4.2
166Claude 3 HaikuAnthropic4.1
166Mistral Medium 3Mistral4.1
168Llama 3.1 405BMeta4.0
169Gemma 4 E4BGoogle3.8
169GPT-4.1 nanoOpenAI3.8
169Llama 4 ScoutMeta3.8
169Phi-4Microsoft3.8
173Solar Pro 2Upstage3.7
174FAUltravox v0.6 Llama 3.3 70BFixie AI3.6
175Qwen2.5 Coder 32B InstructAlibaba3.5
176Mistral Large 2Mistral3.3
177Nova ProAmazon3.2
178GPT-4 TurboOpenAI3.1
179DeepSeek V3DeepSeek2.9
180Claude 3 OpusAnthropic2.8
181GPT-4oOpenAI2.4

Evidence key: ObservedLast good

Rows are ordered by the value Artificial Analysis Humanity's Last Exam Benchmark Leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Knowledge capability leaderboard