Cua Bench
Run the Cua Bench registry tasksets against any computer-use agent on local or cloud desktop sandboxes, or write your own verifiable tasks.
Run the Cua Bench registry tasksets against any computer-use agent on local or cloud desktop sandboxes, or write your own verifiable tasks.
Cua Bench (cb, MIT) runs verifiable computer-use tasks on real desktop
sandboxes and scores any agent on them. Start from a taskset in the
Cua Bench registry: one command runs it
on this machine or in the cloud.
cb run cua-bench-basic --agent cua-agent --model anthropic/claude-sonnet-4-20250514A task goes from the registry through variations, agent adapters, the CLI runner, and self-hosted execution.
dataset -> task -> variant -> session -> trajectory -> evaluator -> rewardA task module (main.py) has four functions:
| Function | Does |
|---|---|
@tasks_config | Returns the variants: same objective, different prompt, data, theme or OS. |
@setup_task | Prepares the variant's environment and opens the apps. |
@solve_task | Optional oracle: one known way to succeed. A failing oracle means a broken task, not a bad agent. |
@evaluate_task | Inspects the final state and returns rewards. It scores an oracle, a human or an agent the same way. |
Tasks run in a real sandbox from the cua SDK: on this machine (cb run --on local, the default) or in the Cua cloud (--on cloud).