Skip to main content
ModelScale

Agentic benchmark

ApprenticeBench leaderboard

ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable job. Every model the catalog carries a published ApprenticeBench value for, ranked by that value.

CategoryAgentic
MeasureCumulative success rate over 100 bills
Tasks100 vendor bills processed in sequence inside a simulated construction company
DifficultyLong-horizon computer use with offline and online continual learning

Tests whether a computer-use agent can learn a real accounts-payable job on the job, processing 100 vendor bills in a company ERP system with only the handbook, historical records, and mentor feedback a new hire would get.

ApprenticeBench ranking

18 models with a published ApprenticeBench value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable job value
RankModelProviderCumulative success rate over 100 bills
1Claude Fable 5.1Anthropic72
2GPT-6 AstraOpenAI68
3Claude Opus 5Anthropic36
4Claude Fable 5Anthropic34
5GPT-5.6 SolOpenAI26
6Gemini 3.8 FlashGoogle24
7GPT-5.5OpenAI20
8Muse Spark 1.3Meta19
9Kimi K3Moonshot AI18
10Claude Sonnet 5Anthropic16
10Gemini 3.7 FlashGoogle16
10GPT-5.6 TerraOpenAI16
13Grok 4.6xAI13
14GPT-5.4OpenAI11
15Claude Opus 4.7Anthropic7
15GPT-5.6 LunaOpenAI7
17Claude Opus 4.6Anthropic5
18Claude Sonnet 4.6Anthropic2

Evidence key: Observed

Rows are ordered by the value ApprenticeBench: a step change in AI's job readiness published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard