Results
The result.json and summary.json a run writes: every key, its type and meaning.
The result.json and summary.json a run writes: every key, its type and meaning.
A run writes one directory per variant and a summary.json next to them:
<output-dir>/
summary.json
<task>_v<variant>/ (with _a<attempt> for --attempts repeats)
result.json
run.log
trajectory.json (ATIF-v1.8, screenshots under imgs/)
task_<variant>_trace/ (a Hugging Face dataset of the trace events)Both files carry schema_version; new keys are additive, and a rename or removal bumps the version.
result.json#One variant's outcome: the cua-bench fields of its result.json.
| Key | Type | Description |
|---|---|---|
schema_version | int | Layout version; a rename or removal bumps it, new keys do not. |
session_id | str | The session id (also the variant's log and trace key). |
task | str | Task directory name. |
variant | int | Variant index (--variant-id). |
status | str | completed, failed or cancelled. |
reward | float | null | The mean of the evaluation's numbers (true counts as 1), or null when evaluation did not run. |
evaluation | any | What @cb.evaluate_task returned (commonly a list of floats). |
error | str | null | The error message when the variant failed. |
duration_s | float | Wall time of the variant in seconds. |
on | str | Where it ran: local or cloud. |
backend | str | What ran it, <location>-<engine> (local-gvisor, local-qemu, cloud-gvisor, cloud-kubevirt, cloud-pool:<name>). |
image | str | null | The image as the task or --image named it. |
pool | str | null | The Fleet pool that served the claim (cloud runs). |
output_dir | str | The variant's output directory. |
extra | dict | Extra values the task or agent reported. |
kind | str | null | container or vm. |
runtime | str | null | The engine (gvisor, runc, qemu, lume, kubevirt) when one was chosen; None: the SDK picked. |
image_ref | str | null | The resolved image reference. |
image_variant | str | null | The image variant that ran: rootfs, containerdisk or lume. |
image_digest | str | null | The pinned repo@sha256:... when the SDK reports it. |
arch | str | null | Guest architecture, when known. |
attempt | int | The --attempts repeat index (0 for the first). |
retries | int | How many infrastructure retries (--retries) the variant took. |
sandbox | str | null | The sandbox's qualified ref (local:<name>, cloud:<name>), when known. |
id | str | A random id for this trial. |
task_name | str | The task directory name. |
trial_name | str | <task>_v<variant>, with _a<attempt> for repeats. |
trial_uri | str | null | file:// URI of the variant's output directory. |
source | str | Always cua-bench. |
agent_info | dict | The agent label, version and model ({name, provider}). |
verifier_result | dict | null | {"rewards": {"reward": <float>, ...}} (numeric evaluation keys included), or null. |
exception_info | dict | null | Exception type, message, traceback and time when the variant raised, else null. |
started_at | str | null | When the trial started. |
finished_at | str | null | When the trial finished. |
environment_setup | Span | null | Sandbox start and task setup. |
agent_setup | null | Always null (agents run in the cb process). |
agent_execution | Span | null | The agent's (or the oracle's) run. |
verifier | Span | null | The evaluation. |
Phase spans (environment_setup, agent_execution, verifier) are started_at and finished_at timestamps (ISO 8601, UTC).
summary.json#summary.json: the run's totals, written next to each variant's directory.
| Key | Type | Description |
|---|---|---|
schema_version | int | Layout version; a rename or removal bumps it, new keys do not. |
total | int | Variants run. |
completed | int | Variants that completed. |
failed | int | Variants that failed or were cancelled. |
avg_reward | float | null | Mean reward over variants with a reward, or null. |
results | list[SummaryResult] | One row per variant (SummaryResult). |
n_total_trials | int | The number of trials (Harbor), equal to total. |
stats | SummaryStats | Harbor-compatible statistics (SummaryStats). |
pass_at_k | dict | null | {"<k>": <pass@k>} with --attempts above 1 (binary rewards, unbiased estimator), else null. |
targets | list[SummaryTarget] | Where the variants ran (SummaryTarget). |
results[]#One row of summary.json results: a variant's outcome.
| Key | Type | Description |
|---|---|---|
results[].task | str | Task directory name. |
results[].variant | int | Variant index (--variant-id). |
results[].status | str | completed, failed or cancelled. |
results[].reward | float | null | The mean of the evaluation's numbers (true counts as 1), or null when evaluation did not run. |
results[].pool | str | null | The Fleet pool that served the claim (cloud runs). |
results[].duration_s | float | Wall time of the variant in seconds. |
results[].error | str | null | The error message when the variant failed. |
results[].kind | str | null | container or vm. |
results[].runtime | str | null | The engine (gvisor, runc, qemu, lume, kubevirt) when one was chosen, else null (the SDK picked). |
results[].image_variant | str | null | The image variant that ran: rootfs, containerdisk or lume. |
results[].image_digest | str | null | The pinned repo@sha256:... when the SDK reports it. |
results[].attempt | int | The --attempts repeat index (0 for the first). |
results[].retries | int | How many infrastructure retries (--retries) the variant took. |
stats#| Key | Type | Description |
|---|---|---|
stats.n_completed | int | Variants that completed. |
stats.n_errored | int | Variants with no reward that did not complete (the run broke). |
stats.n_retries | int | Infrastructure retries across the run. |
targets[]#Where variants ran: one row per (on, kind, runtime, image variant, image).
| Key | Type | Description |
|---|---|---|
targets[].on | str | Where it ran: local or cloud. |
targets[].kind | str | null | container or vm. |
targets[].runtime | str | null | The engine (gvisor, runc, qemu, lume, kubevirt) when one was chosen; None: the SDK picked. |
targets[].image_variant | str | null | The image variant that ran: rootfs, containerdisk or lume. |
targets[].image | str | null | The pinned digest when known, else the image reference. |
targets[].count | int | Variants that ran there. |