Skip to main content
ModelScale

Agentic benchmark

OSWorld 2.0 leaderboard

Every model the catalog carries a published OSWorld 2.0 value for, ranked by that value.

CategoryAgentic
MeasureInteractive computer-use evaluation
Tasks108 long-horizon computer-use workflows
DifficultyLong-horizon professional workflows

A long-horizon computer-use benchmark covering realistic workflows across everyday and professional desktop tasks.

OSWorld 2.0 ranking

20 models with a published OSWorld 2.0 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published OSWorld 2.0 value
RankModelProviderInteractive computer-use evaluation
1GPT-6 AstraOpenAI72.6
2Claude Opus 5Anthropic70.6
3Muse Spark 1.3Meta66.9
4GPT-5.6 SolOpenAI62.6
5Gemini 3.8 FlashGoogle59.0
6GPT-5.6 TerraOpenAI50.2
7Gemini 3.7 FlashGoogle47.9
8GPT-5.6 LunaOpenAI45.6
9Claude Fable 5.1Anthropic41.7
10Claude Opus 4.8Anthropic20.6
11Qwen3.8 MaxAlibaba19.4
11Qwen3.8-Flash-NextAlibaba19.4
13Claude Opus 4.7 (Adaptive)Anthropic18.2
14Muse Spark 1.1Meta14.2
15Claude Opus 4.7Anthropic13.9
16GPT-5.5OpenAI13.0
17Claude Sonnet 4.6Anthropic8.3
18Kimi K2.6Moonshot AI4.6
18MiniMax M3MiniMax4.6
20Qwen3.7 PlusAlibaba2.8

Evidence key: Observed

Rows are ordered by the value OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard