Skip to main content
ModelScale

Agentic benchmark

AA-AnalystAgent leaderboard

Artificial Analysis AnalystAgent. Every model the catalog carries a published AA-AnalystAgent value for, ranked by that value.

CategoryAgentic
MeasureTask success rate
TasksSpreadsheet and document analysis questions
DifficultyBusiness and data analysis

Artificial Analysis' data-analysis benchmark, testing agents on spreadsheet and document work to answer the quantitative questions a business or data analyst faces day to day.

AA-AnalystAgent ranking

16 models with a published AA-AnalystAgent value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Artificial Analysis AnalystAgent value
RankModelProviderTask success rate
1Gemini 3.7 FlashGoogle60.0
2Claude Fable 5.1Anthropic57.5
3Claude Opus 5Anthropic53.8
4GPT-6 AstraOpenAI51.2
5GPT-5.5OpenAI50.0
6Claude Fable 5Anthropic48.8
7GPT-5.6 SolOpenAI47.5
8Claude Sonnet 5Anthropic46.3
9Claude Opus 4.8Anthropic45.0
9Gemini 3.5 FlashGoogle45.0
11Grok 4.6xAI41.3
12Kimi K3Moonshot AI38.8
13TMInklingThinking Machines Lab23.8
14Mistral Medium 3.5 128BMistral12.5
15MiniMax M3MiniMax10.0
16Nemotron 3 UltraNVIDIA6.3

Evidence key: Observed

Rows are ordered by the value AA-AnalystAgent Benchmark Leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard