Skip to main content
ModelScale

Knowledge benchmark

HLE w/o tools leaderboard

Humanity's Last Exam without tools. Every model the catalog carries a published HLE w/o tools value for, ranked by that value.

CategoryKnowledge
MeasureTool-free expert QA
TasksExpert-level questions
DifficultyFrontier expert level

Tool-free variant of Humanity's Last Exam that isolates a model's raw frontier reasoning.

HLE w/o tools ranking

38 models with a published HLE w/o tools value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Humanity's Last Exam without tools value
RankModelProviderTool-free expert QA
1Claude Fable 5.1Anthropic60.9
2Claude Mythos 5Anthropic59
3Claude Opus 5Anthropic56.3
4Muse Spark 1.1Meta52.2
5SASakana Fugu-UltraSakana AI50
6Claude Opus 4.8Anthropic49.8
7UNPareto 26.9Unbiased49
8SASakana FuguSakana AI47.2
9Claude Opus 4.7 (Adaptive)Anthropic46.9
10Gemini 3.1 ProGoogle45.4
11OAOrnith-1.5-397BOrnith AI44.6
12Qwen3.8 MaxAlibaba43.6
13Kimi K3Moonshot AI43.5
14Hy4 previewTencent43.4
15Claude Sonnet 5Anthropic43.2
16GPT-5.5 ProOpenAI43.1
17Muse SparkMeta42.8
18GPT-5.4 ProOpenAI42.7
19GPT-5.5OpenAI41.4
20GLM-5.2Z.AI40.5
21Claude Opus 4.6Anthropic40
22GPT-5.4OpenAI39.8
23Qwen3.8-Flash-NextAlibaba35.9
24MiMo-V2.5-ProXiaomi34
25Grok 4.20xAI31.6
25TMInkling-SmallThinking Machines Lab31.6
27Qwen3.8-27BAlibaba30.8
28TMInklingThinking Machines Lab30
29Solar Open 2Upstage28.8
30GPT-5.4 miniOpenAI28.2
31Nemotron 3 UltraNVIDIA26.7
32OAOrnith-1.5-35B-A3BOrnith AI25.6
33GPT-5.4 nanoOpenAI24.3
34OAOrnith-1.5-9BOrnith AI20.2
35Gemma 4 31BGoogle19.5
36Nemotron 3.5 Lightning 30B A3B NVFP4NVIDIA10.47
37Gemma 4 26B A4BGoogle8.7
38Gemma 4 12BGoogle5.2

Evidence key: Observed

Rows are ordered by the value Introducing GPT-5.4 mini and nano published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Knowledge capability leaderboard