Skip to main content
ModelScale

Agentic benchmark

AA EnterpriseOps-Gym leaderboard

Artificial Analysis EnterpriseOps-Gym. Every model the catalog carries a published AA EnterpriseOps-Gym value for, ranked by that value.

CategoryAgentic
MeasureTask success rate
TasksEnterprise operations workflows
DifficultyEnterprise agent operations

An independently evaluated enterprise-operations benchmark from Artificial Analysis.

AA EnterpriseOps-Gym ranking

17 models with a published AA EnterpriseOps-Gym value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Artificial Analysis EnterpriseOps-Gym value
RankModelProviderTask success rate
1Claude Fable 5Anthropic51.1
2Gemini 3.5 FlashGoogle50.1
3DeepSeek V4 Pro 0813DeepSeek49.6
4Grok 4.6xAI48.3
5Claude Opus 5Anthropic47.5
6Muse Spark 1.2Meta47.3
7GPT-5.5OpenAI46.6
8Kimi K3Moonshot AI45.3
9Qwen3.8-27BAlibaba44.2
10GPT-5.6 SolOpenAI42.9
11Gemini 3.5 Flash-LiteGoogle42.3
12GPT-5.6 LunaOpenAI40.8
13GPT-5.6 TerraOpenAI38.5
14TMInklingThinking Machines Lab38.0
15GLM-5.3Z.AI36.4
16Muse Glimmer 30BMeta34.7
17Mistral Medium 3.5 128BMistral33.7

Evidence key: ObservedLast good

Rows are ordered by the value Artificial Analysis EnterpriseOps-Gym Benchmark Leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard