Skip to main content
ModelScale

Knowledge benchmark

SuperGPQA leaderboard

SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines. Every model the catalog carries a published SuperGPQA value for, ranked by that value.

CategoryKnowledge
MeasureMultiple choice questions
Tasks285 disciplines
DifficultyGraduate level

An expanded version of GPQA that evaluates graduate-level knowledge and reasoning capabilities across 285 disciplines, providing comprehensive coverage of academic domains.

SuperGPQA ranking

20 models with a published SuperGPQA value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines value
RankModelProviderMultiple choice questions
1Claude Opus 4.6Anthropic95
1Claude Sonnet 4.6Anthropic95
3Qwen 3.6 Max (preview)Alibaba73.9
4Qwen3.7 MaxAlibaba73.6
5Qwen3.6 PlusAlibaba71.6
6Qwen3.7 PlusAlibaba71.4
7Seed 2.1 ProByteDance70.8
8Claude Opus 4.5Anthropic70.6
9Qwen3.5 397BAlibaba70.4
10Kimi K2.5Moonshot AI69.2
11Seed 2.1 TurboByteDance67.4
12Qwen3.5-122B-A10BAlibaba67.1
13GLM-5Z.AI66.8
14Qwen3.6-27BAlibaba66
15Qwen3.5-27BAlibaba65.6
16Qwen3.6-35B-A3BAlibaba64.7
17Qwen3.5-35B-A3BAlibaba63.4
18Qwen3 235B 2507Alibaba62.6
19OPMiniCPM5-2BOpenBMB40.8
20OPMiniCPM5-1BOpenBMB23.14

Evidence key: Observed

Rows are ordered by the value SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Knowledge capability leaderboard