Skip to main content
ModelScale

Knowledge benchmark

GPQA leaderboard

Graduate-Level Google-Proof Q&A. Every model the catalog carries a published GPQA value for, ranked by that value.

CategoryKnowledge
MeasureMultiple choice questions
Tasks448 questions
DifficultyGraduate level

A challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. Designed to be difficult even for skilled non-experts with access to Google.

GPQA ranking

84 models with a published GPQA value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Graduate-Level Google-Proof Q&A value
RankModelProviderMultiple choice questions
1GPT-6 AstraOpenAI96
2SASakana FuguSakana AI95.5
2SASakana Fugu-UltraSakana AI95.5
4GPT-5.6 SolOpenAI94.6
5Claude Opus 4.7 (Adaptive)Anthropic94.2
6Claude Mythos 5Anthropic94.1
7Claude Opus 4.8Anthropic93.6
7GPT-5.5OpenAI93.6
9Kimi K3Moonshot AI93.5
10GPT-5.6 TerraOpenAI92.9
11GPT-5.4OpenAI92.8
11OAOrnith-1.5-397BOrnith AI92.8
13Qwen3.8 MaxAlibaba92.6
14GPT-5.2OpenAI92.4
14Qwen3.7 MaxAlibaba92.4
16GPT-5.6 LunaOpenAI92.3
16Hy4 previewTencent92.3
18Gemini 3.5 FlashGoogle92.2
19Qwen3.8-Flash-NextAlibaba91.7
20Claude Opus 4.6Anthropic91.3
21GLM-5.2Z.AI91.2
22Qwen3.8-Omni-FlashAlibaba91
23DeepSeek V4.1 FlashDeepSeek90.9
24Kimi K2.6Moonshot AI90.5
25Qwen3.6 PlusAlibaba90.4
26Qwen3.7 PlusAlibaba90.3
27DeepSeek V4 Pro 0813DeepSeek90.1
27Grok 4.3xAI90.1
29Claude Sonnet 4.6Anthropic89.9
29INInterfaze BetaInterfaze89.9
31TMInkling-SmallThinking Machines Lab89.5
32OAOrnith-1.5-35B-A3BOrnith AI89.2
32Qwen3.8-27BAlibaba89.2
34Qwen3.5 397BAlibaba88.4
35DeepSeek V4 Flash 0731DeepSeek88.1
36GPT-5.4 miniOpenAI88
37TMInklingThinking Machines Lab87.9
38Qwen3.6-27BAlibaba87.8
39Kimi K2.5Moonshot AI87.6
39Kimi K2.5 (Reasoning)Moonshot AI87.6
41Hy3 PreviewTencent87.2
42Claude Opus 4.5Anthropic87
42Nemotron 3 UltraNVIDIA87
44Qwen3.5-122B-A10BAlibaba86.6
45OAOrnith-1.5-9BOrnith AI86.4
46GLM-5Z.AI86
46Qwen3.6-35B-A3BAlibaba86
48PMTernary Bonsai 2 27BPrism ML85.76
49GLM-4.7Z.AI85.7
50Qwen3.5-27BAlibaba85.5
51Ling 3.0 FlashInclusionAI84.97
52Gemma 4 31BGoogle84.3
53MAI-Thinking-1Microsoft84.2
53Qwen3.5-35B-A3BAlibaba84.2
55Ling 3.0 Flash FP8InclusionAI84
56MiMo-V2-FlashXiaomi83.7
57Claude Sonnet 4.5Anthropic83.4
58Gemini 2.5 ProGoogle83
59GPT-5.4 nanoOpenAI82.8
60o1-proOpenAI79
61Gemma 4 12BGoogle78.8
62Qwen3 235B 2507Alibaba77.5
63o3-miniOpenAI77.2
64o1OpenAI75.7
65Nemotron 3.5 Lightning 30B A3B NVFP4NVIDIA75.57
66Nemotron 3 Nano Omni 30B A3BNVIDIA72.2
67ZYZAYA1-8BZyphra71
68Granite 4.2 30BIBM66.41
69GPT-4.1OpenAI66.3
70GPT-4.1 miniOpenAI64.2
71Granite 4.2 8BIBM64.14
72Claude 3.5 SonnetAnthropic59.4
73DeepSeek V3DeepSeek59.1
74Ling 2.6 FlashInclusionAI59
75Gemma 4 E4BGoogle58.6
76JEMellum2-12B-A2.5B-ThinkingJetBrains57.6
77ZYZAYA1-74B-PreviewZyphra57.3
78Granite 4.2 3BIBM54.8
79GPT-4.1 nanoOpenAI50.3
80Gemma 4 E2BGoogle43.4
80SPSoofi S 30B-A3BSoofi Project43.4
82JEMellum2-12B-A2.5B-InstructJetBrains40.9
83LFM2.5-VL-450MLiquidAI25.66
84LFM2.5-230MLiquidAI25.41

Evidence key: Observed

Rows are ordered by the value GPQA: A Graduate-Level Google-Proof Q&A Benchmark published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Knowledge capability leaderboard