Skip to main content
ModelScale

Agentic benchmark

AA AutomationBench leaderboard

Artificial Analysis AutomationBench. Every model the catalog carries a published AA AutomationBench value for, ranked by that value.

CategoryAgentic
MeasureTask success rate
TasksBusiness-process automation tasks
DifficultyAgentic automation

An independently evaluated automation benchmark from Artificial Analysis.

AA AutomationBench ranking

14 models with a published AA AutomationBench value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Artificial Analysis AutomationBench value
RankModelProviderTask success rate
1DeepSeek V4.1 FlashDeepSeek68.9
2GPT-6 AstraOpenAI68.5
3Grok 4.6xAI66.7
4GLM-5.3Z.AI62.2
5GLM-5.3-FlashZ.AI60.4
6GPT-5.6 SolOpenAI60.1
7Gemini 3.8 FlashGoogle59.9
8GPT-5.6 TerraOpenAI59.6
9Claude Fable 5.1Anthropic59.4
10Kimi K3Moonshot AI58.3
11Muse Spark 1.3Meta57.9
12DeepSeek V4 Pro 0813DeepSeek56.7
13Claude Opus 5Anthropic56.6
14Claude Fable 5Anthropic54.1

Evidence key: Observed

Rows are ordered by the value Artificial Analysis AutomationBench Benchmark Leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard