Skip to main content
ModelScale

Agentic benchmark

HLE w/ tools leaderboard

Humanity's Last Exam with tools. Every model the catalog carries a published HLE w/ tools value for, ranked by that value.

CategoryAgentic
MeasurePass@1
TasksExpert questions with tool use
DifficultyFrontier tool-augmented reasoning

Tool-augmented Humanity's Last Exam scores reported in DeepSeek-V4 thinking-mode evaluations.

HLE w/ tools ranking

19 models with a published HLE w/ tools value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Humanity's Last Exam with tools value
RankModelProviderPass@1
1Claude Opus 5Anthropic64.7
2DeepSeek V4.1 FlashDeepSeek63.9
3GLM-5.3Z.AI62.5
4DeepSeek V4 Pro 0813DeepSeek60.0
5Claude Sonnet 5Anthropic57.4
6GPT-6 AstraOpenAI57.2
7Qwen3.8 MaxAlibaba56.2
8APApodex 1.1Apodex56.1
8OAOrnith-1.5-397BOrnith AI56.1
10Hy4 previewTencent55.4
11GLM-5.3-FlashZ.AI55.3
12Qwen3.7 MaxAlibaba53.5
13DSdots3-note PreviewDots Studio52.6
14Agents-A1InternScience47.6
15Step 3.7 FlashStepFun47.2
16DeepSeek V4 Flash 0731DeepSeek45.1
17Nemotron 3 UltraNVIDIA37.4
18OAOrnith-1.5-35B-A3BOrnith AI33.4
19OAOrnith-1.5-9BOrnith AI30.5

Evidence key: Observed

Rows are ordered by the value DeepSeek-V4 Technical Report published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard