Skip to main content
ModelScale

Multimodal & Grounded benchmark

TIR-Bench leaderboard

Every model the catalog carries a published TIR-Bench value for, ranked by that value.

CategoryMultimodal & Grounded
MeasureScreenshot-grounded task reasoning
TasksVisual agent and interface reasoning
DifficultyComputer-use visual reasoning

A visual agent benchmark for interface reasoning and task execution over screenshots or software surfaces.

No model in the catalog has a published TIR-Bench score.

The benchmark is defined by Qwen3.6 launch benchmarks, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

All leaderboards