Skip to main content
ModelScale

Agentic benchmark

APEX-Agents-AA leaderboard

Every model the catalog carries a published APEX-Agents-AA value for, ranked by that value.

CategoryAgentic
MeasurePass@1
Tasks452 professional-services agent tasks
DifficultyLong-horizon workplace agent tasks

Artificial Analysis' implementation of the APEX-Agents benchmark for long-horizon professional-services agent tasks.

APEX-Agents-AA ranking

26 models with a published APEX-Agents-AA value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published APEX-Agents-AA value
RankModelProviderPass@1
1Gemini 3.5 FlashGoogle47.1
2Kimi K3Moonshot AI41.3
3GPT-5.6 TerraOpenAI38.9
4GPT-5.5OpenAI37.7
5GPT-5.6 LunaOpenAI35.8
6GLM-5.2Z.AI33.7
7GPT-5.4OpenAI33.3
8Claude Opus 4.6 (Adaptive)Anthropic33.0
9Gemini 3.1 ProGoogle32.0
10APApodex 1.1Apodex31.2
10APApodex 1.1 MiniApodex31.2
12Kimi K2.6Moonshot AI28.5
13GPT-5.4 miniOpenAI28.2
14GPT-5.4 nanoOpenAI24.9
15DeepSeek V4 Pro 0813DeepSeek24.3
16Qwen3.7 PlusAlibaba22.4
17Grok 4.3xAI17.0
18Step 3.7 FlashStepFun14.8
19GLM-5Z.AI14.5
20Kimi K2.5Moonshot AI11.5
20Kimi K2.5 (Reasoning)Moonshot AI11.5
22MiniMax M2.7MiniMax10.6
23GPT-OSS 120BOpenAI3.1
24MiMo-V2.5-ProXiaomi2.4
25Nemotron 3 Super 100BNVIDIA1.8
26GPT-OSS 20BOpenAI0.7

Evidence key: Observed

Rows are ordered by the value APEX-Agents-AA Benchmark Leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard