Skip to main content
ModelScale

Agentic benchmark

AA ITBench leaderboard

Artificial Analysis ITBench-AA. Every model the catalog carries a published AA ITBench value for, ranked by that value.

CategoryAgentic
MeasureTask success rate
TasksIT incident-response tasks
DifficultyEnterprise IT operations

An independently evaluated IT-operations benchmark from Artificial Analysis.

AA ITBench ranking

10 models with a published AA ITBench value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Artificial Analysis ITBench-AA value
RankModelProviderTask success rate
1GPT-5.6 SolOpenAI56.2
2GPT-5.6 TerraOpenAI51.0
3Kimi K3Moonshot AI47.7
4Claude Opus 4.7 (Adaptive)Anthropic46.7
5GPT-5.5OpenAI45.8
6GLM-5.2Z.AI42.7
7Qwen3.7 MaxAlibaba42.5
8Gemini 3.5 FlashGoogle40.3
8GPT-5.6 LunaOpenAI40.3
10GPT-OSS 120BOpenAI5.6

Evidence key: Observed

Rows are ordered by the value Artificial Analysis ITBench-AA Benchmark Leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard