Agentic benchmark
AA-AnalystAgent leaderboard
Artificial Analysis AnalystAgent. Every model the catalog carries a published AA-AnalystAgent value for, ranked by that value.
Artificial Analysis' data-analysis benchmark, testing agents on spreadsheet and document work to answer the quantitative questions a business or data analyst faces day to day.
AA-AnalystAgent ranking
16 models with a published AA-AnalystAgent value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Task success rate |
|---|---|---|---|
| 1 | 60.0 | ||
| 2 | Anthropic | 57.5 | |
| 3 | Anthropic | 53.8 | |
| 4 | OpenAI | 51.2 | |
| 5 | OpenAI | 50.0 | |
| 6 | Anthropic | 48.8 | |
| 7 | OpenAI | 47.5 | |
| 8 | Anthropic | 46.3 | |
| 9 | Anthropic | 45.0 | |
| 9 | 45.0 | ||
| 11 | xAI | 41.3 | |
| 12 | Moonshot AI | 38.8 | |
| 13 | TMInkling | Thinking Machines Lab | 23.8 |
| 14 | Mistral | 12.5 | |
| 15 | MiniMax | 10.0 | |
| 16 | NVIDIA | 6.3 |
Evidence key: Observed
Rows are ordered by the value AA-AnalystAgent Benchmark Leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.