Skip to main content
ModelScale

Knowledge benchmark

GPQA Diamond (Vals) leaderboard

GPQA Diamond, Vals AI run. Every model the catalog carries a published GPQA Diamond (Vals) value for, ranked by that value.

CategoryKnowledge
MeasureAccuracy
TasksGraduate-level science questions
DifficultyExpert reasoning

Vals AI’s independent run of the GPQA Diamond graduate-level science questions.

GPQA Diamond (Vals) ranking

50 models with a published GPQA Diamond (Vals) value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published GPQA Diamond, Vals AI run value
RankModelProviderAccuracy
1Gemini 3.1 ProGoogle95.5
2GPT-5.6 SolOpenAI95.2
3Grok 4.6xAI94.7
4Gemini 3.8 FlashGoogle94.4
5Gemini 3.7 FlashGoogle93.9
6Qwen3.8 MaxAlibaba93.7
7Claude Fable 5.1Anthropic93.4
7Claude Opus 5Anthropic93.4
7Gemini 3.6 FlashGoogle93.4
10Claude Fable 5Anthropic93.2
10GPT-5.5OpenAI93.2
12Grok 4.5xAI92.9
12Kimi K3Moonshot AI92.9
14Gemini 3.5 FlashGoogle92.7
14MiniMax M3MiniMax92.7
16Claude Opus 4.8Anthropic92.4
16DeepSeek V4 Pro 0813DeepSeek92.4
18GPT-5.6 LunaOpenAI91.7
19Grok 4.3xAI91.4
20Muse Spark 1.1Meta91.2
21GPT-5.6 TerraOpenAI90.9
22Claude Opus 4.7Anthropic90.2
22Qwen3.7 MaxAlibaba90.2
24DeepSeek V4 Flash 0731DeepSeek89.9
25Kimi K2.6Moonshot AI89.1
26Claude Sonnet 5Anthropic88.9
26Qwen3.8-27BAlibaba88.9
28Grok 4.20xAI88.6
29GLM-5.3Z.AI88.1
30Gemini 3 FlashGoogle87.9
31Qwen3.6 PlusAlibaba87.4
32TMInklingThinking Machines Lab87.1
33MiniMax M2.7MiniMax86.6
34GLM-5.3-FlashZ.AI86.4
35Nemotron 3 UltraNVIDIA86.1
36Claude Sonnet 4.6Anthropic85.6
36GLM-5.2Z.AI85.6
38Ling 3.0 FlashInclusionAI84.8
39GLM-5.1Z.AI84.5
40Gemini 3.5 Flash-LiteGoogle83.8
41TMInkling-SmallThinking Machines Lab83.6
42GPT-5.4 miniOpenAI83.1
43MiMo-V2.5-ProXiaomi82.6
44MiMo-V2.5Xiaomi81.6
45Gemini 3.1 Flash-LiteGoogle81.1
46GPT-5.4 nanoOpenAI77.5
47Claude Haiku 4.5Anthropic72.2
48Laguna XS.2Poolside55.1
49Mistral Medium 3.5 128BMistral34.8
50Laguna M.1Poolside27.0

Evidence key: Observed

Rows are ordered by the value Vals AI GPQA Diamond, Vals AI run leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Knowledge capability leaderboard