Multimodal & Grounded benchmark
TIR-Bench leaderboard
Every model the catalog carries a published TIR-Bench value for, ranked by that value.
CategoryMultimodal & Grounded
MeasureScreenshot-grounded task reasoning
TasksVisual agent and interface reasoning
DifficultyComputer-use visual reasoning
Published byQwen3.6 launch benchmarks
A visual agent benchmark for interface reasoning and task execution over screenshots or software surfaces.
No model in the catalog has a published TIR-Bench score.
The benchmark is defined by Qwen3.6 launch benchmarks, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.