Skip to main content
ModelScale

Knowledge benchmark

GPQA-D leaderboard

GPQA Diamond. Every model the catalog carries a published GPQA-D value for, ranked by that value.

CategoryKnowledge
MeasureMultiple choice questions
TasksGraduate-level science questions
DifficultyGraduate level

A display-only GPQA Diamond reference from provider comparison charts.

GPQA-D ranking

64 models with a published GPQA-D value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published GPQA Diamond value
RankModelProviderMultiple choice questions
1GPT-6 AstraOpenAI96.0
2SASakana FuguSakana AI95.5
2SASakana Fugu-UltraSakana AI95.5
4GPT-5.6 SolOpenAI94.6
5Gemini 3.1 ProGoogle94.3
6Claude Opus 4.7 (Adaptive)Anthropic94.2
7Claude Opus 4.8Anthropic93.6
7GPT-5.5OpenAI93.6
9Kimi K3Moonshot AI93.5
10GPT-5.6 TerraOpenAI92.9
11GPT-5.4OpenAI92.8
11OAOrnith-1.5-397BOrnith AI92.8
13Gemini 3.5 FlashGoogle92.7
14Qwen3.8 MaxAlibaba92.6
15Qwen3.7 MaxAlibaba92.4
16GPT-5.6 LunaOpenAI92.3
16Hy4 previewTencent92.3
18Qwen3.8-Flash-NextAlibaba91.7
19GLM-5.2Z.AI91.2
20Qwen3.8-Omni-FlashAlibaba91.0
21DeepSeek V4.1 FlashDeepSeek90.9
22Kimi K2.6Moonshot AI90.5
23Qwen3.7 PlusAlibaba90.3
24DeepSeek V4 Pro 0813DeepSeek90.1
25INInterfaze BetaInterfaze89.9
26TMInkling-SmallThinking Machines Lab89.5
26Muse SparkMeta89.5
28Claude Opus 4.6Anthropic89.2
28OAOrnith-1.5-35B-A3BOrnith AI89.2
28Qwen3.8-27BAlibaba89.2
31Solar Pro 4Upstage89.0
32Grok 4.20xAI88.5
33DeepSeek V4 Flash 0731DeepSeek88.1
34TMInklingThinking Machines Lab87.9
35Kimi K2.5Moonshot AI87.6
36Hy3 PreviewTencent87.2
37MiniMax M2.7MiniMax87.0
37Nemotron 3 UltraNVIDIA87.0
39OAOrnith-1.5-9BOrnith AI86.4
40Solar Open 2Upstage86.3
41GLM-5.1Z.AI86.2
42GLM-5Z.AI86.0
43PMTernary Bonsai 2 27BPrism ML85.8
44STA.X K2SK Telecom85.6
45Ling 3.0 FlashInclusionAI85.0
46MAI-Thinking-1Microsoft84.2
47Ling 3.0 Flash FP8InclusionAI84.0
48K-EXAONE 2.0LG AI Research82.2
49Mercury 2.5Inception79.0
50Gemma 4 12BGoogle78.8
51Trinity-Large-ThinkingArcee AI76.3
52Nemotron 3.5 Lightning 30B A3B NVFP4NVIDIA75.6
53Nemotron 3 Nano Omni 30B A3BNVIDIA72.2
54ZYZAYA1-8BZyphra71.0
55OPMiniCPM5-2BOpenBMB70.2
56LongCat-Flash-Lite-SparseMeituan69.5
57Trinity-Large-PreviewArcee AI63.3
58JEMellum2-12B-A2.5B-ThinkingJetBrains57.6
59ZYZAYA1-74B-PreviewZyphra57.3
60LLaDA2.2-miniInclusionAI44.4
61SPSoofi S 30B-A3BSoofi Project43.4
62JEMellum2-12B-A2.5B-InstructJetBrains40.9
63OPMiniCPM5-1BOpenBMB26.3
64LFM2.5-230MLiquidAI25.4

Evidence key: Observed

Rows are ordered by the value Trinity-Large-Thinking: Scaling an Open Source Frontier Agent published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Knowledge capability leaderboard