Knowledge benchmark
GPQA-D leaderboard
GPQA Diamond. Every model the catalog carries a published GPQA-D value for, ranked by that value.
A display-only GPQA Diamond reference from provider comparison charts.
GPQA-D ranking
64 models with a published GPQA-D value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Multiple choice questions |
|---|---|---|---|
| 1 | OpenAI | 96.0 | |
| 2 | SASakana Fugu | Sakana AI | 95.5 |
| 2 | SASakana Fugu-Ultra | Sakana AI | 95.5 |
| 4 | OpenAI | 94.6 | |
| 5 | 94.3 | ||
| 6 | Anthropic | 94.2 | |
| 7 | Anthropic | 93.6 | |
| 7 | OpenAI | 93.6 | |
| 9 | Moonshot AI | 93.5 | |
| 10 | OpenAI | 92.9 | |
| 11 | OpenAI | 92.8 | |
| 11 | OAOrnith-1.5-397B | Ornith AI | 92.8 |
| 13 | 92.7 | ||
| 14 | Alibaba | 92.6 | |
| 15 | Alibaba | 92.4 | |
| 16 | OpenAI | 92.3 | |
| 16 | Tencent | 92.3 | |
| 18 | Alibaba | 91.7 | |
| 19 | Z.AI | 91.2 | |
| 20 | Alibaba | 91.0 | |
| 21 | DeepSeek | 90.9 | |
| 22 | Moonshot AI | 90.5 | |
| 23 | Alibaba | 90.3 | |
| 24 | DeepSeek | 90.1 | |
| 25 | INInterfaze Beta | Interfaze | 89.9 |
| 26 | TMInkling-Small | Thinking Machines Lab | 89.5 |
| 26 | Meta | 89.5 | |
| 28 | Anthropic | 89.2 | |
| 28 | OAOrnith-1.5-35B-A3B | Ornith AI | 89.2 |
| 28 | Alibaba | 89.2 | |
| 31 | Upstage | 89.0 | |
| 32 | xAI | 88.5 | |
| 33 | DeepSeek | 88.1 | |
| 34 | TMInkling | Thinking Machines Lab | 87.9 |
| 35 | Moonshot AI | 87.6 | |
| 36 | Tencent | 87.2 | |
| 37 | MiniMax | 87.0 | |
| 37 | NVIDIA | 87.0 | |
| 39 | OAOrnith-1.5-9B | Ornith AI | 86.4 |
| 40 | Upstage | 86.3 | |
| 41 | Z.AI | 86.2 | |
| 42 | Z.AI | 86.0 | |
| 43 | PMTernary Bonsai 2 27B | Prism ML | 85.8 |
| 44 | STA.X K2 | SK Telecom | 85.6 |
| 45 | InclusionAI | 85.0 | |
| 46 | Microsoft | 84.2 | |
| 47 | InclusionAI | 84.0 | |
| 48 | LG AI Research | 82.2 | |
| 49 | Inception | 79.0 | |
| 50 | 78.8 | ||
| 51 | Arcee AI | 76.3 | |
| 52 | NVIDIA | 75.6 | |
| 53 | NVIDIA | 72.2 | |
| 54 | ZYZAYA1-8B | Zyphra | 71.0 |
| 55 | OPMiniCPM5-2B | OpenBMB | 70.2 |
| 56 | Meituan | 69.5 | |
| 57 | Arcee AI | 63.3 | |
| 58 | JEMellum2-12B-A2.5B-Thinking | JetBrains | 57.6 |
| 59 | ZYZAYA1-74B-Preview | Zyphra | 57.3 |
| 60 | InclusionAI | 44.4 | |
| 61 | SPSoofi S 30B-A3B | Soofi Project | 43.4 |
| 62 | JEMellum2-12B-A2.5B-Instruct | JetBrains | 40.9 |
| 63 | OPMiniCPM5-1B | OpenBMB | 26.3 |
| 64 | LiquidAI | 25.4 |
Evidence key: Observed
Rows are ordered by the value Trinity-Large-Thinking: Scaling an Open Source Frontier Agent published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.