Skip to main content
ModelScale

external benchmark

CUA-bench leaderboard

Vals CUA-bench. Every model the catalog carries a published CUA-bench value for, ranked by that value.

Categoryexternal
MeasureOverall score with game-level task splits
TasksSix commercial video-game control tasks, including held-out games
DifficultyKeyboard-and-mouse computer use in real-time games
Published byCUA-bench

A Vals AI computer-use benchmark in which agents play six commercial video games with a keyboard and mouse.

No model in the catalog has a published CUA-bench score.

The benchmark is defined by CUA-bench, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

About CUA-bench

What the benchmark measures and how its value is published, from the source’s own record.

What is CUA-bench?

CUA-bench (Vals CUA-bench) is an external benchmark published by CUA-bench. A Vals AI computer-use benchmark in which agents play six commercial video games with a keyboard and mouse. Its evaluation set is described by the source as six commercial video-game control tasks, including held-out games. The source rates its difficulty as keyboard-and-mouse computer use in real-time games.

How is CUA-bench scored?

Each model's result is published by CUA-bench as overall score with game-level task splits. ModelScale prints that value to 2 decimal places, exactly as published, and never converts or rescales it. This page ranks the 0 models that carry a published value, highest value first; a model the source never scored is left out rather than ranked at zero. The source does not state whether a higher value is the better result, so neither does this page.

All leaderboards