Skip to main content
ModelScale

Agentic benchmark

JobBench leaderboard

Every model the catalog carries a published JobBench value for, ranked by that value.

CategoryAgentic
MeasureAgentic workplace deliverables
Tasks130 tasks across 35 occupations
DifficultyProfessional multi-source workflows

An occupational agent benchmark for professional workflows that workers say they most want delegated to AI.

JobBench ranking

27 models with a published JobBench value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published JobBench value
RankModelProviderAgentic workplace deliverables
1Muse Spark 1.3Meta64.9
2Hy4 previewTencent61.7
3Qwen3.8-Flash-NextAlibaba55.7
4Muse Spark 1.1Meta54.7
5Qwen3.8 MaxAlibaba53.4
6Kimi K3Moonshot AI52.9
7Atria Dawn PreviewShanghai Artificial Intelligence Laboratory50.3
8Claude Opus 4.7 (Adaptive)Anthropic45.9
9GPT-5.5OpenAI42.7
10GPT-5.4OpenAI38.9
11Claude Sonnet 4.6Anthropic36.9
12Claude Opus 4.6Anthropic36.7
13GPT-5.2OpenAI34.3
14GPT-5.3 CodexOpenAI33.7
15Qwen3.8-27BAlibaba33.4
16Claude Opus 4.5Anthropic32.3
17Claude Sonnet 4.5Anthropic27.7
18GPT-5.1-CodexOpenAI26.2
19GPT-5.2-CodexOpenAI26.0
20Claude 4.1 OpusAnthropic21.9
21Qwen3.5 PlusAlibaba18.5
22Claude 4 SonnetAnthropic18.4
23Claude Haiku 4.5Anthropic16.0
24Gemini 3 FlashGoogle11.4
24Gemini 3 ProGoogle11.4
26Kimi K2.5Moonshot AI8.7
27GPT-5 (high)OpenAI8.5

Evidence key: Observed

Rows are ordered by the value JobBench: Aligning Agent Work With Human Will published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard