Reasoning benchmark
ARC-AGI-3 leaderboard
Abstraction and Reasoning Corpus for AGI v3. Every model the catalog carries a published ARC-AGI-3 value for, ranked by that value.
An interactive successor to ARC-AGI-2 that evaluates whether an AI agent can learn unfamiliar task mechanics through action and feedback.
ARC-AGI-3 ranking
12 models with a published ARC-AGI-3 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Agentic task completion under a capped evaluation budget |
|---|---|---|---|
| 1 | OpenAI | 62.71 | |
| 2 | Anthropic | 30.16 | |
| 3 | OpenAI | 7.78 | |
| 4 | Anthropic | 1.52 | |
| 5 | OpenAI | 0.8 | |
| 6 | OpenAI | 0.43 | |
| 7 | 0.42 | ||
| 8 | xAI | 0.3 | |
| 9 | OpenAI | 0.21 | |
| 10 | Anthropic | 0.18 | |
| 10 | OpenAI | 0.18 | |
| 12 | xAI | 0.09 |
Evidence key: Observed
Rows are ordered by the value ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.