cb run
Run a task or dataset with an agent or the oracle, run one interactively, and follow runs.
Run a task or dataset with an agent or the oracle, run one interactively, and follow runs.
| Command | Description |
|---|---|
cb run | Run a task or dataset, and manage runs. |
cb run task | Run one task variant. |
cb run dataset | Run every task and variant of a dataset in parallel. |
cb run list | List all runs with status. |
cb run info | Show detailed info about a run. |
cb run watch | Watch a run in real-time with live updates. |
cb run stop | Stop a run and release its sandboxes. |
cb run logs | View combined logs from a run or session. |
cb interact | Run a task's setup with its desktop visible, then evaluate. |
Every command also accepts the global options.
cb run#Run a task or dataset, and manage runs.
Run a task (a directory with main.py) or a dataset, with an agent or the task's oracle. cb run <path> is shorthand for cb run task|dataset <path>.
cb run [COMMAND]Examples
# Run a task with its oracle solution
cb run ./tasks/hello_file_env
# Run a dataset with an agent, four variants at once
cb run cua-bench-basic --agent cua-agent --model anthropic/claude-sonnet-4-20250514 -j 4
# List runs
cb run listcb run task#Run one task variant.
cb run task [OPTIONS] <task_path>| Argument | Type | Default | Description |
|---|---|---|---|
<task_path> | string | required | Path to task directory (containing main.py). |
| Flag | Short | Type | Default | Description |
|---|---|---|---|---|
--variant-id | integer | 0 | Task variant index (default: 0). | |
--agent | string | Agent to use for evaluation (e.g., cua-agent). | ||
--agent-import-path | string | Import path for custom agent (e.g., "path.to.agent:MyCustomAgent"). | ||
--model | string | Model to use with the agent (e.g., anthropic/claude-sonnet-4-20250514). | ||
--max-steps | integer | 100 | Maximum number of steps for agent execution (default: 100). | |
--on | local | cloud | e2b | daytona | modal | cloudflare | vercel | morph | runloop | fly | blaxel | codesandbox | northflank | aws | gcp | Where sandboxes run: local (this machine), cloud, your own cloud account (aws, gcp, modal; connect it first with cua cloud connect <provider>), or a contrib provider (e2b, daytona; needs a cua SDK built with contrib and the provider's key). Default: CUA_DEFAULT_ON, else cua config set default.on, else local. | ||
--kind | auto | container | vm | What kind of sandbox: container or vm. Default auto: Linux as a container unless the task or image says VM; Windows, macOS and Android are VM-only. Default: CUA_DEFAULT_KIND, else cua config set default.kind. | ||
--runtime | auto | gvisor | runc | qemu | lume | kubevirt | Which engine: gvisor or runc (local containers), qemu or lume (local VMs), gvisor or kubevirt (cloud containers, VMs). Implies its kind. Default auto: the SDK picks (CUA_DEFAULT_RUNTIME, else cua config set default.runtime). | ||
--image | string | Image for every task (overrides the task's setup_config.image): an OS alias (linux, windows, macos:tahoe), ghcr.io/trycua/<os>, any registry ref or a digest, or pool:<name> for an existing Fleet pool. Env: CUA_BENCH_IMAGE. | ||
--cpu | integer | vCPUs per sandbox. | ||
--memory | string | Memory per sandbox (e.g. 4096, 4G). | ||
--claim-ttl | string | cloud: how long a claim outlives a crashed run (default 15m; renewed while running). | ||
--output-dir | string | Output directory for session results. | ||
--attempts | integer | 1 | Run every variant this many times (fresh sandbox each); summary.json reports pass@k (default: 1). | |
--retries | integer | 0 | Retry a variant that failed before evaluation (sandbox start, connection) up to this many times, with backoff (default: 0). | |
--keep-runs | integer | After the run, delete all but the N newest runs in the runs directory (default: keep everything; env CUA_BENCH_KEEP_RUNS). | ||
--max-age | number | After the run, delete runs older than DAYS (env CUA_BENCH_MAX_AGE_DAYS). | ||
--max-results-size | string | After the run, delete the oldest runs until the runs directory fits in SIZE, e.g. 20G (env CUA_BENCH_MAX_RESULTS_SIZE). | ||
--oracle | boolean | false | Run the oracle solution (the default when no agent is given). | |
--noop | boolean | false | Set up and evaluate with no actions (the null baseline for eval parity). | |
--warm | boolean | false | cloud: keep one replica ready when the managed pool is first created. | |
--detach | -d | boolean | false | Run in the background and return (follow with cb run watch <id>). |
--dry-run | boolean | false | Print each variant's image, kind, variant and backend, then exit without starting a sandbox. |
Examples
# Run variant 0 with the oracle solution
cb run task ./tasks/hello_file_env
# Run variant 2 with an agent in the cloud
cb run task ./tasks/hello_file_env --variant-id 2 --agent cua-agent --model anthropic/claude-sonnet-4-20250514 --on cloud
# Show what would run without starting a sandbox
cb run ./tasks/hello_file_env --dry-runcb run dataset#Run every task and variant of a dataset in parallel.
cb run dataset [OPTIONS] <dataset_path>| Argument | Type | Default | Description |
|---|---|---|---|
<dataset_path> | string | required | Path to dataset directory, or dataset name from registry. |
| Flag | Short | Type | Default | Description |
|---|---|---|---|---|
--max-parallel | -j | integer | 4 | Variants running at once; in the cloud also the managed pool's max size (default: 4). |
--max-variants | integer | Maximum number of variants to run per task (default: all). | ||
--task-filter | string | Filter tasks by name pattern (glob). | ||
--agent | string | Agent to use for evaluation (e.g., cua-agent). | ||
--agent-import-path | string | Import path for custom agent (e.g., "path.to.agent:MyCustomAgent"). | ||
--model | string | Model to use with the agent (e.g., anthropic/claude-sonnet-4-20250514). | ||
--max-steps | integer | 100 | Maximum number of steps for agent execution (default: 100). | |
--on | local | cloud | e2b | daytona | modal | cloudflare | vercel | morph | runloop | fly | blaxel | codesandbox | northflank | aws | gcp | Where sandboxes run: local (this machine), cloud, your own cloud account (aws, gcp, modal; connect it first with cua cloud connect <provider>), or a contrib provider (e2b, daytona; needs a cua SDK built with contrib and the provider's key). Default: CUA_DEFAULT_ON, else cua config set default.on, else local. | ||
--kind | auto | container | vm | What kind of sandbox: container or vm. Default auto: Linux as a container unless the task or image says VM; Windows, macOS and Android are VM-only. Default: CUA_DEFAULT_KIND, else cua config set default.kind. | ||
--runtime | auto | gvisor | runc | qemu | lume | kubevirt | Which engine: gvisor or runc (local containers), qemu or lume (local VMs), gvisor or kubevirt (cloud containers, VMs). Implies its kind. Default auto: the SDK picks (CUA_DEFAULT_RUNTIME, else cua config set default.runtime). | ||
--image | string | Image for every task (overrides the task's setup_config.image): an OS alias (linux, windows, macos:tahoe), ghcr.io/trycua/<os>, any registry ref or a digest, or pool:<name> for an existing Fleet pool. Env: CUA_BENCH_IMAGE. | ||
--cpu | integer | vCPUs per sandbox. | ||
--memory | string | Memory per sandbox (e.g. 4096, 4G). | ||
--claim-ttl | string | cloud: how long a claim outlives a crashed run (default 15m; renewed while running). | ||
--output-dir | string | Output directory for session results. | ||
--attempts | integer | 1 | Run every variant this many times (fresh sandbox each); summary.json reports pass@k (default: 1). | |
--retries | integer | 0 | Retry a variant that failed before evaluation (sandbox start, connection) up to this many times, with backoff (default: 0). | |
--keep-runs | integer | After the run, delete all but the N newest runs in the runs directory (default: keep everything; env CUA_BENCH_KEEP_RUNS). | ||
--max-age | number | After the run, delete runs older than DAYS (env CUA_BENCH_MAX_AGE_DAYS). | ||
--max-results-size | string | After the run, delete the oldest runs until the runs directory fits in SIZE, e.g. 20G (env CUA_BENCH_MAX_RESULTS_SIZE). | ||
--oracle | boolean | false | Run the oracle solution (the default when no agent is given). | |
--noop | boolean | false | Set up and evaluate with no actions (the null baseline for eval parity). | |
--warm | boolean | false | cloud: keep one replica ready when the managed pool is first created. | |
--detach | -d | boolean | false | Run in the background and return (follow with cb run watch <id>). |
--dry-run | boolean | false | Print each variant's image, kind, variant and backend, then exit without starting a sandbox. |
Examples
# Run a registry dataset with its oracle solutions
cb run dataset cua-bench-basic
# Run one variant of each matching task, eight at once, in the background
cb run ./tasks --task-filter 'click*' --max-variants 1 -j 8 --detach
# Run every variant three times and report pass@k
cb run dataset cua-bench-basic --agent cua-agent --attempts 3cb run list#List all runs with status.
cb run list [OPTIONS]| Flag | Short | Type | Default | Description |
|---|---|---|---|---|
--verbose | -v | boolean | false | Show verbose debugging information. |
Examples
# List runs
cb run list
# Include debugging details
cb run list --verbosecb run info#Show detailed info about a run.
cb run info <run_id>| Argument | Type | Default | Description |
|---|---|---|---|
<run_id> | string | required | Run ID to show info for. |
Examples
# Show one run
cb run info 30c12572cb run watch#Watch a run in real-time with live updates.
cb run watch <run_id>| Argument | Type | Default | Description |
|---|---|---|---|
<run_id> | string | required | Run ID to watch. |
Examples
# Follow a detached run
cb run watch 30c12572cb run stop#Stop a run and release its sandboxes.
cb run stop <run_id>| Argument | Type | Default | Description |
|---|---|---|---|
<run_id> | string | required | Run ID to stop. |
Examples
# Stop a run
cb run stop 30c12572cb run logs#View combined logs from a run or session.
cb run logs [OPTIONS] <identifier>| Argument | Type | Default | Description |
|---|---|---|---|
<identifier> | string | required | Run ID or Session ID to view logs for. |
| Flag | Type | Default | Description |
|---|---|---|---|
--tail | integer | Show only the last N lines. |
Examples
# Print a run's logs
cb run logs 30c12572
# Print the last 50 lines
cb run logs 30c12572 --tail 50cb interact#Run a task's setup with its desktop visible, then evaluate.
cb interact [OPTIONS] <env_path>| Argument | Type | Default | Description |
|---|---|---|---|
<env_path> | string | required | Path to the task directory, or a task name with --dataset or --dataset-path. |
| Flag | Type | Default | Description |
|---|---|---|---|
--variant-id | integer | 0 | Task variant index (default: 0). |
--dataset | string | Registry dataset to resolve the task from (CUA_REGISTRY_HOME). | |
--dataset-path | string | Path to dataset directory containing multiple tasks. | |
--max-steps | integer | Maximum number of env.step() calls before stopping. | |
--screenshot | string | Save a screenshot to this file. | |
--trace-out | string | Record a trace and save it as a dataset to this path on exit. | |
--on | local | cloud | e2b | daytona | modal | cloudflare | vercel | morph | runloop | fly | blaxel | codesandbox | northflank | aws | gcp | Where sandboxes run: local (this machine), cloud, your own cloud account (aws, gcp, modal; connect it first with cua cloud connect <provider>), or a contrib provider (e2b, daytona; needs a cua SDK built with contrib and the provider's key). Default: CUA_DEFAULT_ON, else cua config set default.on, else local. | |
--kind | auto | container | vm | What kind of sandbox: container or vm. Default auto: Linux as a container unless the task or image says VM; Windows, macOS and Android are VM-only. Default: CUA_DEFAULT_KIND, else cua config set default.kind. | |
--runtime | auto | gvisor | runc | qemu | lume | kubevirt | Which engine: gvisor or runc (local containers), qemu or lume (local VMs), gvisor or kubevirt (cloud containers, VMs). Implies its kind. Default auto: the SDK picks (CUA_DEFAULT_RUNTIME, else cua config set default.runtime). | |
--image | string | Image for every task (overrides the task's setup_config.image): an OS alias (linux, windows, macos:tahoe), ghcr.io/trycua/<os>, any registry ref or a digest, or pool:<name> for an existing Fleet pool. Env: CUA_BENCH_IMAGE. | |
--cpu | integer | vCPUs per sandbox. | |
--memory | string | Memory per sandbox (e.g. 4096, 4G). | |
--claim-ttl | string | cloud: how long a claim outlives a crashed run (default 15m; renewed while running). | |
--oracle | boolean | false | Run the solution after setup. |
--view | boolean | false | Open the trace viewer when done. |
--no-wait | boolean | false | Skip the interactive prompt (useful for SSH/CI testing). |
--no-browser | boolean | false | Print the Display: URL without opening it (also CUA_BENCH_NO_BROWSER=1). |
--warm | boolean | false | cloud: keep one replica ready when the managed pool is first created. |
Examples
# Open a task and its display, wait for Enter, then evaluate
cb interact ./tasks/hello_file_env
# Run variant 1 with its oracle and save a screenshot
cb interact ./tasks/hello_file_env --variant-id 1 --oracle --screenshot shot.png
# Resolve a task from a registry dataset
cb interact click-button --dataset cua-bench-basic| Code | Meaning |
|---|---|
0 | Success: every variant completed. |
1 | Failure: a variant failed or was cancelled, or the command could not run. |
2 | Usage error: an unknown command, flag or value. |
130 | Interrupted (Ctrl-C); running variants are cancelled. |