Skip to main content
ModelScale

Agentic benchmark

AndroidWorld leaderboard

Every model the catalog carries a published AndroidWorld value for, ranked by that value.

CategoryAgentic
MeasureInteractive mobile-agent evaluation
TasksAndroid app workflows
DifficultyComplex mobile task completion
Published byGLM-5V-Turbo

A mobile GUI agent benchmark for completing Android app workflows and on-device tasks.

AndroidWorld ranking

9 models with a published AndroidWorld value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published AndroidWorld value
RankModelProviderInteractive mobile-agent evaluation
1Qwen3.8-Omni-FlashAlibaba87.1
2Qwen3.8 MaxAlibaba85.3
3Qwen3.8-Flash-NextAlibaba84.5
4Qwen3.8-27BAlibaba81.9
5Qwen3.7 PlusAlibaba81.0
6HCHolo3.1-35B-A3BH Company79.3
7HCHolo3.1-4BH Company71.0
7HCHolo3.1-9BH Company71.0
9Qwen3.6-27BAlibaba70.3

Evidence key: Observed

Rows are ordered by the value GLM-5V-Turbo published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard