Skip to main content
ModelScale

Agentic benchmark

OSWorld-Verified leaderboard

Every model the catalog carries a published OSWorld-Verified value for, ranked by that value.

CategoryAgentic
MeasureExecution-based interactive task success
Tasks369 real-world computer tasks (361 when eight Google Drive tasks are excluded)
DifficultyMulti-step desktop and cross-application workflows
Published byOSWorld

OSWorld-Verified is the July 2025 repaired release of OSWorld's real-computer evaluation. It measures whether a model-agent system can finish desktop and web tasks from configured starting states, with success checked by execution-based evaluators.

OSWorld-Verified ranking

32 models with a published OSWorld-Verified value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published OSWorld-Verified value
RankModelProviderExecution-based interactive task success
1Qwen3.8 MaxAlibaba86.1
2Claude Fable 5Anthropic85
2Claude Mythos 5Anthropic85
4Qwen3.8-27BAlibaba84.3
5Claude Opus 4.8Anthropic83.4
6Gemini 3.6 FlashGoogle83
7HCHolo3-35B-A3BH Company82.56
8Claude Sonnet 5Anthropic81.2
9Muse Spark 1.1Meta80.8
10HCHolo3-122B-A10BH Company78.85
11GPT-5.5OpenAI78.7
12Gemini 3.5 FlashGoogle78.4
13Claude Opus 4.7 (Adaptive)Anthropic78
14UI-Mate-27BTencent77
15GPT-5.4OpenAI75
16Gemini 3.5 Flash-LiteGoogle74
17Qwen3.7 PlusAlibaba73.3
18Kimi K2.6Moonshot AI73.1
19Claude Opus 4.6Anthropic72.7
20Claude Sonnet 4.6Anthropic72.1
20GPT-5.4 miniOpenAI72.1
22MiniMax M3MiniMax70.06
23Claude Opus 4.5Anthropic66.3
24UI-Mate-9BTencent66.2
25Muse Glimmer 30BMeta65.9
26GPT-5.3 CodexOpenAI64.7
27Claude Sonnet 4.5Anthropic61.4
28Qwen3.5-122B-A10BAlibaba58
29Qwen3.5-27BAlibaba56.2
30Qwen3.5-35B-A3BAlibaba54.5
31GPT-5.2OpenAI47.3
32GPT-5.4 nanoOpenAI39

Evidence key: Observed

Rows are ordered by the value OSWorld published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard