Skip to main content
ModelScale

Multimodal & Grounded benchmark

ScreenSpot Pro leaderboard

Every model the catalog carries a published ScreenSpot Pro value for, ranked by that value.

CategoryMultimodal & Grounded
MeasureStatic interface element localization
Tasks1,581 grounding instructions
DifficultyProfessional GUI grounding

A GUI-grounding benchmark for 1,581 instructions in full-screen, high-resolution professional interfaces. It tests where a target is, not whether an agent can finish the surrounding workflow.

ScreenSpot Pro ranking

18 models with a published ScreenSpot Pro value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published ScreenSpot Pro value
RankModelProviderStatic interface element localization
1GPT-6 AstraOpenAI92.7
2Claude Opus 4.8Anthropic87.9
3GPT-5.4OpenAI85.4
4Qwen3.8 MaxAlibaba84.5
5Gemini 3.1 ProGoogle84.4
6Muse SparkMeta84.1
7Claude Opus 4.6Anthropic83.1
8Qwen3.7 PlusAlibaba79.0
9Muse Glimmer 30BMeta75.4
10Gemini 3 ProGoogle72.7
11HCHolo2-235B-A22BH Company70.6
12Qwen3.6 PlusAlibaba68.2
13HCHolo2-30B-A3BH Company66.1
14Qwen3.5 397BAlibaba65.6
15HCHolo2-8BH Company58.9
16Nemotron 3 Nano Omni 30B A3BNVIDIA57.8
17HCHolo2-4BH Company57.2
18Claude Opus 4.5Anthropic45.7

Evidence key: Observed

Rows are ordered by the value ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards