# CUA-bench leaderboard

external benchmark

> Every figure below is reproduced as its upstream source published it: nothing is modelled, estimated, interpolated or converted. `Unavailable` means no source published the value — it is never a zero. Each value carries its evidence state and the date it was observed.

Page: https://modelscale.dev/leaderboards/benchmarks/external-vals-cua-bench

Vals CUA-bench. Every model the catalog carries a published CUA-bench value for, ranked by that value.

## Benchmark record

| Field | Value |
| --- | --- |
| Category | external |
| Measure | Overall score with game-level task splits |
| Tasks | Six commercial video-game control tasks, including held-out games |
| Difficulty | Keyboard-and-mouse computer use in real-time games |
| Published by | [CUA-bench](https://www.vals.ai/benchmarks/cua_bench) |

A Vals AI computer-use benchmark in which agents play six commercial video games with a keyboard and mouse.

## CUA-bench ranking

No model in the catalog has a published CUA-bench score. The benchmark is defined by CUA-bench, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

Rows are ordered by the value the source published, highest first. The source does not state whether a higher value is the better result, so this document does not either: for a benchmark that measures a rate of failure, read the table from the bottom.
