Skip to main content
ModelScale

Agentic benchmark

GDP.pdf leaderboard

Artificial Analysis GDP.pdf. Every model the catalog carries a published GDP.pdf value for, ranked by that value.

CategoryAgentic
MeasureTask success rate
TasksProfessional document-production tasks
DifficultyProfessional knowledge work

Artificial Analysis' document-output benchmark: professional tasks whose deliverable is a produced PDF, one of the ten components of its Intelligence Index v4.3.

GDP.pdf ranking

14 models with a published GDP.pdf value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Artificial Analysis GDP.pdf value
RankModelProviderTask success rate
1GPT-6 AstraOpenAI31.0
2GPT-5.6 SolOpenAI27.2
3Muse Spark 1.3Meta26.6
4Claude Fable 5.1Anthropic26.2
5Claude Fable 5Anthropic24.0
5GPT-5.6 LunaOpenAI24.0
5GPT-5.6 TerraOpenAI24.0
8Kimi K3Moonshot AI22.0
9Claude Opus 5Anthropic21.6
10Gemini 3.8 FlashGoogle21.0
11Grok 4.6xAI17.0
12Qwen3.8-27BAlibaba16.4
13GLM-5.3-FlashZ.AI15.4
14Gemini 3.5 Flash-LiteGoogle13.6

Evidence key: Observed

Rows are ordered by the value GDP.pdf Benchmark Leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard