Quickstart
Run a Cua Bench registry taskset on a local desktop sandbox with one command, read the results, then run it with an agent.
Run a Cua Bench registry taskset on a local desktop sandbox with one command, read the results, then run it with an agent.
Pick a taskset from the Cua Bench registry, run it on a local desktop sandbox, and read the score. No API key until the agent step.
cua runtime doctor.uv tool install cua-benchList the tasksets, then run one task of cua-bench-basic with its built-in
oracle:
cb dataset list
cb run cua-bench-basic --task-filter click-button --max-variants 1cb fetches the taskset at its pinned version, starts a Linux desktop
sandbox, opens the task's app as a window on it, runs the oracle, scores the
result (reward=1) and releases the sandbox. A failing oracle means the task
or your setup is broken, not the agent.
cb run list
cb run info <run-id>Each variant's result.json, run.log and trajectory.json land in
~/.local/share/cua-bench/runs/<run-id>/, next to a summary.json.
cb run cua-bench-basic --agent cua-agent --model anthropic/claude-sonnet-4-20250514 -j 4
cb run cua-bench-basic --agent cua-agent --model anthropic/claude-sonnet-4-20250514 -j 16 --on cloudThe first runs every task on this machine (--on local is the default); the
second runs them in the Cua cloud after cb login.
cb interact sets up a task on a desktop you can see, waits while you solve
it, then evaluates it and releases the sandbox:
cb interact click-button --dataset cua-bench-basicRun results stay in ~/.local/share/cua-bench/runs (a warning past 10 GiB,
CUA_BENCH_WARN_SIZE). Keep fewer with --keep-runs N, --max-age DAYS or
--max-results-size SIZE on cb run, or:
cb prune --runs --keep 10