Task definition
The task module contract: cb.Task, setup_config, the lifecycle decorators, the session API and actions.
The task module contract: cb.Task, setup_config, the lifecycle decorators, the session API and actions.
A task is a directory with a main.py. cb imports it, calls the @cb.tasks_config function to list the variants, then for each variant starts a sandbox from its computer and runs setup, the agent (or the @cb.solve_task oracle) and evaluation. Each lifecycle function receives the variant's cb.Task and the session on its sandbox.
cb.Task#One task variant, as returned by the @cb.tasks_config function.
| Field | Type | Default | Description |
|---|---|---|---|
description | str | required | The instruction the agent receives. |
task_id | str | null | None | A stable identifier for the variant (optional; the index is used otherwise). |
metadata | dict | null | None | Variant-specific values your setup, solve and evaluate functions read. |
computer | dict | null | None | The sandbox: {"provider": "native", "setup_config": {...}} (see DesktopSetupConfig). native (the default) is the only provider. |
setup_config#A task's computer["setup_config"]: the sandbox the task runs in. Every key is optional. Command-line flags (--image, --kind, --cpu, --memory) override the task's values; --runtime (the engine) is a command-line choice only.
| Key | Type | Description |
|---|---|---|
os_type | string | Guest OS: linux (the default; ubuntu is an alias), windows (win11, win10 and older names are aliases), macos or android. macOS and Android run only with --on local. |
width | int | Screen width in pixels. |
height | int | Screen height in pixels. |
image | str | Registry image or OS alias (linux, windows, macos:tahoe), or pool:<name> for an existing Fleet pool. --image wins, then this, then CUA_BENCH_IMAGE, then the canonical image of os_type. |
kind | str | Preferred kind: container (Linux only) or vm. |
kinds | list[str] | Kinds the task supports (a requirement, unlike kind): for example ["vm"] for a task that needs its own kernel. |
requires | list[str] | What the target must provide: kvm, env:<NAME> (a set environment variable), openai (env:OPENAI_API_KEY) or hf-gated (env:HF_TOKEN). |
server_port | int | A port the image serves itself, used as the readiness probe. |
memory | str | VM memory, for example "8GB". |
cpu | str | VM CPUs, for example "4". |
background | str | Ignored (kept so older tasks still load). |
wallpaper | str | Ignored (kept so older tasks still load). |
installed_apps | list[str] | Ignored (kept so older tasks still load). |
storage | str | Deprecated and ignored. |
provider_type | str | Deprecated: "cloud" runs on Fleet; use --on cloud instead. |
| Decorator | Description |
|---|---|
@cb.tasks_config | Marks the function that lists the task's variants. The function takes no arguments and returns a list of cb.Task; the CLI indexes them from 0 (--variant-id). Use as @cb.tasks_config or with a split: @cb.tasks_config("train") (default train). |
@cb.setup_task | Marks the function that prepares a variant's sandbox. async def setup(task_cfg: cb.Task, session: cb.DesktopSession) runs after the sandbox starts and before the agent or the oracle. Use as @cb.setup_task or with a split: @cb.setup_task("train"). |
@cb.solve_task | Marks the oracle: the function that solves a variant (optional). async def solve(task_cfg: cb.Task, session: cb.DesktopSession) runs instead of an agent with --oracle or when no agent is given. Use as @cb.solve_task or with a split: @cb.solve_task("train"). |
@cb.evaluate_task | Marks the function that scores a variant. async def evaluate(task_cfg: cb.Task, session: cb.DesktopSession) returns the evaluation, commonly list[float]; the reward is the mean of its numbers. Use as @cb.evaluate_task or with a split: @cb.evaluate_task("train"). |
Unified desktop session using the cua-sandbox SDK.
| Method | Returns | Description |
|---|---|---|
computer | Any | The cua-computer style Computer of this session (0.2.x). |
sandbox | Any | The connected cua_sandbox.Sandbox (None for computer-server sessions). |
await step(action: Action) | None | Execute an action (alias for execute_action, for env.step() compatibility). |
await start(config: Optional[DesktopSetupConfig]=None, headless: Optional[bool]=None) | None | Start the session and connect to the environment. |
interface | Any | The cua-computer 0.5 interface surface of this session. |
await serve_static(url_path: str, local_path: str) | None | Serve static files - not applicable for remote environments. |
await launch_window(url: Optional[str]=None, *, html: Optional[str]=None, folder: Optional[str]=None, title: str='Window', x: Optional[int]=None, y: Optional[int]=None, width: int=600, height: int=400, icon: Optional[str]=None, use_inner_size: bool=False, title_bar_style: str='default') | int | str | Launch a window in the remote environment using bench_ui (pywebview). |
await get_element_rect(pid: int | str, selector: str, *, space: Literal['window', 'screen']='window', timeout: float=0.5) | dict[str, Any] | None | Get element rect by CSS selector using bench_ui. |
await execute_javascript(pid: int | str, javascript: str) | Any | Execute JavaScript in a pywebview window using bench_ui. |
await execute_action(action: Action) | None | Execute an action on the remote desktop using the SDK. |
await screenshot() | bytes | Capture screenshot from remote environment. |
await get_snapshot() | Snapshot | Get snapshot of desktop state with active window info. |
await close() | None | Close the session and cleanup resources. |
await close_all_windows() | None | Close all windows - best effort. |
page | Any | Return underlying page object - not applicable for remote. |
vnc_url | str | Return the VNC URL for accessing the environment. |
apps | AppsProxy | Access registered apps via session.apps.{app_name}. |
await click_element(pid: int | str, selector: str) | None | Find element by CSS selector and click its center. |
await right_click_element(pid: int | str, selector: str) | None | Find element by CSS selector and right-click its center. |
await get_accessibility_tree() | Dict[str, Any] | Get the accessibility tree if supported ({} when not). |
await shell_command(command: str, *, check: bool=True, timeout: Optional[float]=None) | Dict[str, Any] | Execute a shell command. |
await read_file(path: str) | str | Read a text file from the environment. |
await write_file(path: str, content: str) | None | Write a text file to the environment. |
await read_bytes(path: str) | bytes | Read a file as bytes from the environment. |
await write_bytes(path: str, data: bytes) | None | Write bytes to a file in the environment. |
await file_exists(path: str) | bool | Whether path is an existing file (not a directory), as in 0.2.x. |
await directory_exists(path: str) | bool | Check if a directory exists in the environment. |
await list_dir(path: str) | list[str] | List contents of a directory in the environment. |
await run_command(command: str, *, check: bool=True, timeout: Optional[float]=None) | Dict[str, Any] | Execute a shell command (alias for shell_command). |
await launch_application(app_name: str) | None | Launch an application by name. |
await check_status() | bool | Check if the environment is responsive. |
await wait_until_ready(timeout: int=60, poll_interval: float=2.0) | bool | Wait until the environment is ready. |
await click(x: int, y: int) | None | Click at coordinates. |
await right_click(x: int, y: int) | None | Right-click at coordinates. |
await double_click(x: int, y: int) | None | Double-click at coordinates. |
await type(text: str) | None | Type text. |
await key(key: str) | None | Press a key. |
await hotkey(keys: list[str]) | None | Press a key combination. |
await scroll(direction: str='down', amount: int=300) | None | Scroll the screen. |
await move_to(x: int, y: int) | None | Move cursor to coordinates. |
await drag(from_x: int, from_y: int, to_x: int, to_y: int) | None | Drag from one position to another. |
os_type | str | Return the OS type for this session. |
await install_app(app_name: str, *, with_shortcut: bool=True, **kwargs) | None | Install a registered app on the native desktop environment. |
await launch_app(app_name: str, **kwargs) | None | Launch a registered app on the native desktop environment. |
Input actions for session.execute_action (from cua_bench).
| Action | Fields |
|---|---|
ClickAction | x: int, y: int |
RightClickAction | x: int, y: int |
DoubleClickAction | x: int, y: int |
MiddleClickAction | x: int, y: int |
DragAction | from_x: int, from_y: int, to_x: int, to_y: int, duration: float = 1.0 |
MoveToAction | x: int, y: int, duration: float = 0.0 |
ScrollAction | direction: string = 'up', amount: int = 100 |
TypeAction | text: str |
KeyAction | key: str |
HotkeyAction | keys: list[str] |
DoneAction | none |
WaitAction | seconds: float = 1.0 |
"""A tiny native task for smoke-testing a target: write a word to a file.
Runs on any Linux image with cua-spacesd (shell and files), for example:
cb run example_tasks/hello_file_env # local gVisor container
cb run example_tasks/hello_file_env --on cloud --image <registry ref>
"""
import cua_bench as cb
WORDS = ["hello", "bench"]
@cb.tasks_config(split="train")
def load():
return [
cb.Task(
description=f"Write the word '{word}' into /tmp/cb_answer.txt.",
metadata={"word": word},
computer={"provider": "native", "setup_config": {"os_type": "linux"}},
)
for word in WORDS
]
@cb.setup_task(split="train")
async def setup(task_cfg, session: cb.DesktopSession):
await session.run_command("rm -f /tmp/cb_answer.txt", check=False)
await session.write_file("/tmp/cb_goal.txt", task_cfg.metadata["word"])
@cb.solve_task(split="train")
async def solve(task_cfg, session: cb.DesktopSession):
await session.run_command("cp /tmp/cb_goal.txt /tmp/cb_answer.txt")
@cb.evaluate_task(split="train")
async def evaluate(task_cfg, session: cb.DesktopSession) -> list[float]:
if not await session.file_exists("/tmp/cb_answer.txt"):
return [0.0]
answer = (await session.read_file("/tmp/cb_answer.txt")).strip()
return [1.0 if answer == task_cfg.metadata["word"] else 0.0]