Skip to main content
ModelScale

Reasoning benchmark

ARC-AGI-2 leaderboard

Abstraction and Reasoning Corpus for AGI v2. Every model the catalog carries a published ARC-AGI-2 value for, ranked by that value.

CategoryReasoning
MeasureGrid transformation puzzles with novel rules
TasksVisual pattern completion and abstract reasoning
DifficultyExpert-level — hardest public reasoning benchmark

A benchmark measuring fluid intelligence and novel abstract reasoning through visual grid puzzles. Models must identify patterns in input-output pairs and generate the correct output for unseen inputs. Considered the hardest public reasoning benchmark — average individual human performance is 66%.

ARC-AGI-2 ranking

22 models with a published ARC-AGI-2 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Abstraction and Reasoning Corpus for AGI v2 value
RankModelProviderGrid transformation puzzles with novel rules
1GPT-6 AstraOpenAI95
2GPT-5.6 SolOpenAI92.5
3Claude Opus 5Anthropic90.4
4Claude Fable 5.1Anthropic90
5GPT-5.5OpenAI85
6GPT-5.6 TerraOpenAI83.9
7GPT-5.4 ProOpenAI83.3
8DSdots3-note PreviewDots Studio81.4
9Gemini 3.1 ProGoogle77.08
10Claude Opus 4.7 (Adaptive)Anthropic75.8
11GPT-5.4OpenAI73.95
12Gemini 3.5 FlashGoogle72.1
13Claude Opus 4.8Anthropic72.08
14GPT-5.6 LunaOpenAI59.54
15Grok 4.20xAI53.3
16GPT-5.2OpenAI52.9
17Grok 4.5xAI52.64
18Gemini 3 Pro Deep ThinkGoogle45.1
19Muse SparkMeta42.5
20TMInkling-SmallThinking Machines Lab40.1
21Gemini 3 ProGoogle31.1
22Claude Sonnet 4.5Anthropic13.6

Evidence key: Observed

Rows are ordered by the value ARC-AGI-2: A Harder General Intelligence Benchmark published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Reasoning capability leaderboard