Skip to main content
ModelScale

Knowledge benchmark

HLE-Verified leaderboard

Every model the catalog carries a published HLE-Verified value for, ranked by that value.

CategoryKnowledge
MeasureFull-set accuracy
Tasks1,811 verified or revised expert questions
DifficultyFrontier multidisciplinary expert reasoning

A verified and revised Humanity's Last Exam set that removes uncertain items and repairs fixable questions before evaluating frontier expert reasoning.

HLE-Verified ranking

6 models with a published HLE-Verified value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published HLE-Verified value
RankModelProviderFull-set accuracy
1Gemini 3.8 FlashGoogle54.9
2GPT-5.6 SolOpenAI54.5
3Claude Opus 5Anthropic54.4
4Gemini 3.7 FlashGoogle53.6
5GPT-5.6 TerraOpenAI51.1
6Claude Sonnet 5Anthropic31.0

Evidence key: Observed

Rows are ordered by the value HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Knowledge capability leaderboard