Knowledge benchmark
GPQA leaderboard
Graduate-Level Google-Proof Q&A. Every model the catalog carries a published GPQA value for, ranked by that value.
A challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. Designed to be difficult even for skilled non-experts with access to Google.
GPQA ranking
84 models with a published GPQA value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Multiple choice questions |
|---|---|---|---|
| 1 | OpenAI | 96 | |
| 2 | SASakana Fugu | Sakana AI | 95.5 |
| 2 | SASakana Fugu-Ultra | Sakana AI | 95.5 |
| 4 | OpenAI | 94.6 | |
| 5 | Anthropic | 94.2 | |
| 6 | Anthropic | 94.1 | |
| 7 | Anthropic | 93.6 | |
| 7 | OpenAI | 93.6 | |
| 9 | Moonshot AI | 93.5 | |
| 10 | OpenAI | 92.9 | |
| 11 | OpenAI | 92.8 | |
| 11 | OAOrnith-1.5-397B | Ornith AI | 92.8 |
| 13 | Alibaba | 92.6 | |
| 14 | OpenAI | 92.4 | |
| 14 | Alibaba | 92.4 | |
| 16 | OpenAI | 92.3 | |
| 16 | Tencent | 92.3 | |
| 18 | 92.2 | ||
| 19 | Alibaba | 91.7 | |
| 20 | Anthropic | 91.3 | |
| 21 | Z.AI | 91.2 | |
| 22 | Alibaba | 91 | |
| 23 | DeepSeek | 90.9 | |
| 24 | Moonshot AI | 90.5 | |
| 25 | Alibaba | 90.4 | |
| 26 | Alibaba | 90.3 | |
| 27 | DeepSeek | 90.1 | |
| 27 | xAI | 90.1 | |
| 29 | Anthropic | 89.9 | |
| 29 | INInterfaze Beta | Interfaze | 89.9 |
| 31 | TMInkling-Small | Thinking Machines Lab | 89.5 |
| 32 | OAOrnith-1.5-35B-A3B | Ornith AI | 89.2 |
| 32 | Alibaba | 89.2 | |
| 34 | Alibaba | 88.4 | |
| 35 | DeepSeek | 88.1 | |
| 36 | OpenAI | 88 | |
| 37 | TMInkling | Thinking Machines Lab | 87.9 |
| 38 | Alibaba | 87.8 | |
| 39 | Moonshot AI | 87.6 | |
| 39 | Moonshot AI | 87.6 | |
| 41 | Tencent | 87.2 | |
| 42 | Anthropic | 87 | |
| 42 | NVIDIA | 87 | |
| 44 | Alibaba | 86.6 | |
| 45 | OAOrnith-1.5-9B | Ornith AI | 86.4 |
| 46 | Z.AI | 86 | |
| 46 | Alibaba | 86 | |
| 48 | PMTernary Bonsai 2 27B | Prism ML | 85.76 |
| 49 | Z.AI | 85.7 | |
| 50 | Alibaba | 85.5 | |
| 51 | InclusionAI | 84.97 | |
| 52 | 84.3 | ||
| 53 | Microsoft | 84.2 | |
| 53 | Alibaba | 84.2 | |
| 55 | InclusionAI | 84 | |
| 56 | Xiaomi | 83.7 | |
| 57 | Anthropic | 83.4 | |
| 58 | 83 | ||
| 59 | OpenAI | 82.8 | |
| 60 | OpenAI | 79 | |
| 61 | 78.8 | ||
| 62 | Alibaba | 77.5 | |
| 63 | OpenAI | 77.2 | |
| 64 | OpenAI | 75.7 | |
| 65 | NVIDIA | 75.57 | |
| 66 | NVIDIA | 72.2 | |
| 67 | ZYZAYA1-8B | Zyphra | 71 |
| 68 | IBM | 66.41 | |
| 69 | OpenAI | 66.3 | |
| 70 | OpenAI | 64.2 | |
| 71 | IBM | 64.14 | |
| 72 | Anthropic | 59.4 | |
| 73 | DeepSeek | 59.1 | |
| 74 | InclusionAI | 59 | |
| 75 | 58.6 | ||
| 76 | JEMellum2-12B-A2.5B-Thinking | JetBrains | 57.6 |
| 77 | ZYZAYA1-74B-Preview | Zyphra | 57.3 |
| 78 | IBM | 54.8 | |
| 79 | OpenAI | 50.3 | |
| 80 | 43.4 | ||
| 80 | SPSoofi S 30B-A3B | Soofi Project | 43.4 |
| 82 | JEMellum2-12B-A2.5B-Instruct | JetBrains | 40.9 |
| 83 | LiquidAI | 25.66 | |
| 84 | LiquidAI | 25.41 |
Evidence key: Observed
Rows are ordered by the value GPQA: A Graduate-Level Google-Proof Q&A Benchmark published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.