Knowledge benchmark
HLE w/o tools leaderboard
Humanity's Last Exam without tools. Every model the catalog carries a published HLE w/o tools value for, ranked by that value.
Tool-free variant of Humanity's Last Exam that isolates a model's raw frontier reasoning.
HLE w/o tools ranking
38 models with a published HLE w/o tools value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Tool-free expert QA |
|---|---|---|---|
| 1 | Anthropic | 60.9 | |
| 2 | Anthropic | 59 | |
| 3 | Anthropic | 56.3 | |
| 4 | Meta | 52.2 | |
| 5 | SASakana Fugu-Ultra | Sakana AI | 50 |
| 6 | Anthropic | 49.8 | |
| 7 | UNPareto 26.9 | Unbiased | 49 |
| 8 | SASakana Fugu | Sakana AI | 47.2 |
| 9 | Anthropic | 46.9 | |
| 10 | 45.4 | ||
| 11 | OAOrnith-1.5-397B | Ornith AI | 44.6 |
| 12 | Alibaba | 43.6 | |
| 13 | Moonshot AI | 43.5 | |
| 14 | Tencent | 43.4 | |
| 15 | Anthropic | 43.2 | |
| 16 | OpenAI | 43.1 | |
| 17 | Meta | 42.8 | |
| 18 | OpenAI | 42.7 | |
| 19 | OpenAI | 41.4 | |
| 20 | Z.AI | 40.5 | |
| 21 | Anthropic | 40 | |
| 22 | OpenAI | 39.8 | |
| 23 | Alibaba | 35.9 | |
| 24 | Xiaomi | 34 | |
| 25 | xAI | 31.6 | |
| 25 | TMInkling-Small | Thinking Machines Lab | 31.6 |
| 27 | Alibaba | 30.8 | |
| 28 | TMInkling | Thinking Machines Lab | 30 |
| 29 | Upstage | 28.8 | |
| 30 | OpenAI | 28.2 | |
| 31 | NVIDIA | 26.7 | |
| 32 | OAOrnith-1.5-35B-A3B | Ornith AI | 25.6 |
| 33 | OpenAI | 24.3 | |
| 34 | OAOrnith-1.5-9B | Ornith AI | 20.2 |
| 35 | 19.5 | ||
| 36 | NVIDIA | 10.47 | |
| 37 | 8.7 | ||
| 38 | 5.2 |
Evidence key: Observed
Rows are ordered by the value Introducing GPT-5.4 mini and nano published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.