cua-bench is an RL / eval environment for GUI Computer-Use Agents on ERPNext v15. It ships as Harbor tasks: each task packages a pre-seeded ERPNext stack, an in-container cua-bench agent-api (FastAPI + Playwright + Chromium), and a grader that scores the agent against real accounting state via the Frappe REST API.
- What an agent sees: PNG screenshots of a Chromium window running ERPNext (default 1280×800).
- What an agent does:
click / type / key / scroll / drag / wait— in cua-bench's canonical schema or the lab's native format (anthropic,openai,google). - How it's graded: state-based checkpoints with weighted partial credit, GL cross-validation, and episode-scoped creation timestamps to resist common reward hacks; the open attack surface is documented in docs/reward-hacking.md.
This project was previously developed under the name veraf. Some committed historical artifacts (the sweep summary in
jobs/) predate the rename. The git history was squashed from a private repository at that re-home, so the current commits date from the public re-home rather than reflecting the original development timeline.
cua-bench/
├── docker/
│ ├── agent.Dockerfile # "main" image: Playwright + agent-api + graders
│ ├── erpnext-seeded.Dockerfile # ERPNext backend with seed baked in
│ ├── mariadb-seeded.Dockerfile # MariaDB with golden snapshot baked in
│ ├── canary.Dockerfile # Minimal smoke image
│ └── cua-bench-dind.Dockerfile # DinD snapshot template (bakes GHCR creds for Daytona)
├── cua_bench/ # FastAPI agent-api (runs inside env container on :5000)
│ └── {server,environment,executor,reset,screenshot,…}.py
├── cua_bench_harbor/ # Harbor-native multi-vendor CUA runner
│ └── src/cua_bench_harbor/{agent,runner,clients,translators,trajectory}.py
├── cua-patterns/ # Cross-env spec-gaming detection predicate library
│ └── src/cua_patterns/{predicates,judges,*_adapter}.py
├── graders/erpnext/ # Grader defs (JSON + Python) baked into the agent image
├── tasks/<id>/ # Harbor task packages
│ ├── task.toml # compute / image / healthcheck
│ ├── instruction.md # what the agent is told
│ ├── environment/docker-compose.yaml
│ ├── solution/solve.sh # oracle (upper-bound check)
│ └── tests/test.sh # verifier → reward
├── db-snapshot/golden.sql # committed; restored per-episode by /reset
├── scripts/build_images.sh # builds all prebuilt images
└── docs/reward-hacking.md
cua-bench runs on Daytona cloud sandboxes. Each trial spins up its own isolated Docker-in-Docker sandbox with the ERPNext stack — no local Docker Desktop, no port conflicts, trivial horizontal scaling (we routinely run 30+ trials in parallel).
# 1. Configure
cp .env.example .env
# Edit .env — you need at minimum:
# DAYTONA_API_KEY=dtn_... (https://app.daytona.io/dashboard)
# DAYTONA_API_URL=https://app.daytona.io/api
# one of ANTHROPIC_API_KEY / OPENAI_API_KEY / GEMINI_API_KEY
uv venv && uv pip install -e cua_bench_harbor
# 2. Run one task on Daytona
EXECUTOR=daytona bash scripts/run_sweep.sh anthropic 1 po-001
# 3. Inspect what happened
python scripts/inspect_trajectory.py jobs/anthropic/po-001/r1
# 4. Full failure-rate sweep: 19 tasks × 2 vendors (anthropic + openai) × N rollouts
EXECUTOR=daytona bash scripts/run_sweep.sh all 3
python scripts/build_sweep_table.py > docs/failure-rate.mdThe task packages reference prebuilt images at
ghcr.io/vaiyr/cua-bench-*. Those images are built locally withscripts/build_images.shor published at release, so a fresh clone will not pull them until you build them or a tagged release is cut.
Read on for details, troubleshooting, and the full set of tasks.
- A Daytona account with API key (sign up)
- Python 3.11+ and
uv harbor≥ 0.3- At least one of:
ANTHROPIC_API_KEY,OPENAI_API_KEY,GEMINI_API_KEY - Optional (for ad-hoc runs only, not swept):
QWEN_API_KEY(Alibaba Dashscope),HAI_API_KEY(Hcompany Holo)
Network policy on Daytona: by default Daytona restricts sandbox egress
in a way that blocks some benchmarks. Apply coupon HARBOR_NETWORK on your
Daytona account to remove the restriction before running.
cua-bench uses a custom DinD snapshot (cua-bench-dind) that bakes in GHCR pull
credentials, so the in-sandbox Docker daemon can pull private images from
ghcr.io/vaiyr/cua-bench-*. This is per-org, one-time.
- Register GHCR with Daytona (Dashboard → Registries → Add Registry, or
POST /api/docker-registrywith your GitHub PAT —write:packagesandread:packagesscopes). - Build + push the
cua-bench-dindimage with your GHCR creds baked in (seedocker/cua-bench-dind.Dockerfilefor the template), or adapt the snippet in the Daytona Dashboard Snapshots UI. The snapshot namecua-bench-dindis whatscripts/run_sweep.shpasses via--ek dind_snapshot=cua-bench-dind. - Flag: 10 GB per-sandbox disk is the default org cap. The stack fits,
but barely. If you see
Disk request 20GB exceeds maximum allowed, emailsupport@daytona.ioto raise the quota.
Three rate limits to plan around when sweeping at scale. Defaults in
scripts/run_sweep.sh are tuned for the most restrictive case; bump them
once you've confirmed your account/setup tolerates it.
| limit | scope | default cap | symptom when hit |
|---|---|---|---|
MAX_JOBS=5 |
Anthropic concurrent sandboxes | safe | none — well below all caps |
OPENAI_MAX_JOBS=6 |
OpenAI concurrent sandboxes | OpenAI Tier 3 (2M TPM) | llm_failure mid-trial |
| Daytona vCPU | concurrent sandboxes total | 25 (default org plan) | sandbox queueing |
| Docker Hub anonymous | 100 pulls / 6hr / IP | per Daytona egress IP | RuntimeError: toomanyrequests, trial exits with null reward |
Docker Hub rate limit was the dominant failure mode in the 2026-04-16
sweep — at 13 concurrent sandboxes, anonymous pulls of redis:6.2-alpine
exhausted the limit and ~12% of trials returned null rewards. Two
mitigations, in order of preference:
- Mirror the offending image. The compose at
tasks/po-001/environment/docker-compose.yamlnow referencesmirror.gcr.io/library/redis:6.2-alpineinstead ofredis:6.2-alpine. Google's mirror is free, no auth, no rate limit. If you add new tasks that pull docker.io images, mirror them the same way. - Authenticate Docker Hub inside the DinD snapshot (boosts to 200
pulls/6hr — temporary workaround, doesn't scale to large sweeps). Bake
docker logininto thecua-bench-dindsnapshot rebuild rather than per-sweep, so creds don't leak into harbor logs and login latency doesn't multiply across sandboxes.
Tuning for higher OpenAI tiers. TPM caps roughly: Tier 2 = 1M, Tier 3
= 2M ($100 spent), Tier 4 = 4M ($250), Tier 5 = 40M ($1k). Default is
OPENAI_MAX_JOBS=6 (sized for Tier 3); bump higher in proportion to TPM
headroom once you cross.
Pre-caching transitive images in cua-bench-dind (the proper long-term
fix) requires rebuilding the snapshot via the Daytona dashboard. Spin up
a sandbox from the current cua-bench-dind, run docker pull for every
image referenced in any task compose, then save as cua-bench-dind:v2. Update
DIND_SNAPSHOT=cua-bench-dind:v2 in scripts/run_sweep.sh env or .env.
git clone <repo> && cd cua-bench
cp .env.example .env # then edit: DAYTONA_API_KEY=..., ANTHROPIC_API_KEY=...cd cua_bench_harbor
uv venv && uv pip install -e ".[dev]"
.venv/bin/pytest # sanitycua-bench's prebuilt images live on GitHub Container Registry at
ghcr.io/vaiyr/cua-bench-*. Daytona sandboxes pull them via a custom
cua-bench-dind snapshot that bakes in GHCR credentials. You only need to
rebuild if cua_bench/, graders/, docker/, or db-snapshot/golden.sql
changes.
erpnext-seeded needs sites-data/ (the initialized Frappe site dir). If
missing, extract it from a running backend:
docker cp cua-bench-backend-1:/home/frappe/frappe-bench/sites/. sites-data/Then rebuild + push all four images (forces linux/amd64 for Daytona):
PUSH=1 CUA_BENCH_TAG=v2 bash scripts/build_images.shFast path when only cua_bench/ or graders/ changed (rebuilds just the
agent image):
PUSH=1 CUA_BENCH_TAG=v2 bash scripts/rebuild_agent.shAfter pushing a new tag, update the docker_image field in
tasks/*/task.toml and the image: fields in
tasks/po-001/environment/docker-compose.yaml to the new tag.
All run paths default to Daytona when EXECUTOR=daytona is set. The
single-task wrapper and the sweep both pass the Daytona-specific flags
(-e daytona --ek dind_snapshot=cua-bench-dind --override-storage-mb 10000)
for you.
# Sweep-of-one is the cleanest single-task entry point:
EXECUTOR=daytona bash scripts/run_sweep.sh anthropic 1 po-001
EXECUTOR=daytona bash scripts/run_sweep.sh google 1 pe-001The raw harbor run form is also supported — useful for one-off debugging:
# Smoke test: canary task (no ERPNext, ~30s)
harbor run \
--env-file .env \
-p tasks/cua-bench-canary \
--agent-import-path cua_bench_harbor.agent:CuaBenchComputerUseAgent \
-e daytona --ek dind_snapshot=cua-bench-dind --override-storage-mb 10000 \
--ak vendor=anthropic --ak max_steps=20 --no-delete
# Real task: create + submit a Purchase Order (~5–7 min end-to-end)
harbor run \
--env-file .env \
-p tasks/po-001 \
--agent-import-path cua_bench_harbor.agent:CuaBenchComputerUseAgent \
--model anthropic/claude-opus-4-7 \
-e daytona --ek dind_snapshot=cua-bench-dind --override-storage-mb 10000 \
--ak vendor=anthropic --ak max_steps=50 --no-delete
# Same task with GPT-5.4 (Responses API + built-in computer tool)
harbor run \
--env-file .env \
-p tasks/po-001 \
--agent-import-path cua_bench_harbor.agent:CuaBenchComputerUseAgent \
--model openai/gpt-5.4 \
-e daytona --ek dind_snapshot=cua-bench-dind --override-storage-mb 10000 \
--ak vendor=openai --ak max_steps=50 --no-delete
# Same task with Gemini 3 flash preview
harbor run \
--env-file .env \
-p tasks/po-001 \
--agent-import-path cua_bench_harbor.agent:CuaBenchComputerUseAgent \
--model google/gemini-3-flash-preview \
-e daytona --ek dind_snapshot=cua-bench-dind --override-storage-mb 10000 \
--ak vendor=google --ak max_steps=50 --no-deleteQwen3-VL and Hcompany Holo3 are wired as additional vendors but excluded
from run_sweep.sh. Both currently show capability-bound failure on
po-001 — the models localize the Items table row at y≈830, which is
off-canvas (viewport is 800 tall), and don't self-correct even with textual
or visual feedback signals. The traces are clean and useful to inspect, but
sweeping them across 19 tasks × N rollouts burns budget without changing the
failure-rate story. Re-enable in the sweep once a stronger variant (e.g.
qwen3-vl-235b-a22b-thinking) is worth benchmarking — see the comment in
scripts/run_sweep.sh.
# Qwen3-VL-235B-A22B-Instruct via Alibaba Dashscope (OpenAI-compat, native
# tool_calls). Requires QWEN_API_KEY.
bash scripts/run_task.sh qwen po-001
# Holo3-122B-A10B via Hcompany hub (OpenAI-compat, structured JSON output +
# thinking mode — does NOT use native tool_calls). Requires HAI_API_KEY.
bash scripts/run_task.sh holo po-001
# Override model within a vendor:
QWEN_MODEL=qwen/qwen3-vl-235b-a22b-thinking bash scripts/run_task.sh qwen po-001
HOLO_MODEL=holo/holo3-35b-a3b bash scripts/run_task.sh holo po-001Gotchas if you change these clients:
- Holo coord convention is prompt-dependent. With an explicit "viewport is WxH" hint in the system prompt (which we send), Holo emits absolute pixel coords. Without it, it defaults to a normalized 0–999 grid. Don't remove the hint unless you also add denormalization.
- Holo thinking mode eats tokens. We set
max_tokens=4000because reasoning_content can be ~1–2k tokens before the JSON action is emitted. With the default 800, responses come backfinish=length, content=None. - Holo rejects tiny screenshots. The Hcompany API returns a 500 on images smaller than ~a few hundred pixels per side. Our 1280×800 task screenshots are fine; unit tests that build 10×10 probe images are not.
- Holo accumulates context aggressively. We trim the in-flight message
history to the last 2 screenshots (
_trim_screenshots) — without this the API 400s around step 6.
## Multi-vendor failure-rate sweep
`scripts/run_sweep.sh` runs every task N times per vendor and writes a
JSONL summary to `jobs/sweep-summary.jsonl`. `scripts/build_sweep_table.py`
renders it as a markdown table with Wilson 95% CIs.
```bash
# Smoke first: one easy task × each vendor × 1 rollout.
EXECUTOR=daytona bash scripts/run_sweep.sh all 1 po-001
# Pilot: all 19 tasks × 2 vendors (anthropic + openai) × 1 rollout (~15-30 min at 30-way, ~$30).
EXECUTOR=daytona bash scripts/run_sweep.sh all 1
# Full sweep: n=3 × 19 tasks × 2 vendors (~1h at 30-way, ~$75–$150).
EXECUTOR=daytona bash scripts/run_sweep.sh all 3
# Render the table.
.venv/bin/python scripts/build_sweep_table.py > docs/failure-rate.md
Expected cost: ~$300–800 for the full sweep depending on task lengths
and cache-hit rates. Each client logs cached-token counts — verify your
cache is warming via the ATIF trajectories before letting the long tail
run. Cap per-rollout spend with --ak max_cost_usd=....
Model availability caveat: gpt-5.4 and
gemini-3-flash-preview are both relatively new/preview.
If your account doesn't have access, the first smoke run will fail with
a 404 from the vendor API — swap to gpt-4o or gemini-2.5-flash
respectively (neither has the same computer-use fidelity, but smoke still
validates plumbing).
Oracle check (bypasses the UI, hits Frappe REST directly — upper bound):
harbor run -p tasks/po-001 --agent oracle \
-e daytona --ek dind_snapshot=cua-bench-dind --override-storage-mb 10000 \
--env-file .env --no-deleteThe verifier (tasks/<id>/tests/test.sh) calls POST /evaluate
inside the env container and writes the grader score to
/logs/verifier/reward.txt.
# Per-step actions, repeat-coord smell, cache-hit %, total cost.
python scripts/inspect_trajectory.py jobs/runs/google-po-001-*
# Rendered videos and per-step PNGs:
open jobs/runs/<run>/<task>__*/artifacts/videos/*.webm
ls jobs/runs/<run>/<task>__*/agent/screenshots/
# Inner loop when editing cua_bench/ (executor, server, reset, etc.):
PUSH=1 CUA_BENCH_TAG=v2 bash scripts/rebuild_agent.sh # rebuild + push to GHCR
EXECUTOR=daytona bash scripts/run_sweep.sh openai 1 po-001inspect_trajectory.py is the fastest way to spot model-quality vs
plumbing issues: a "stuck-coord" cluster + repeated typed text usually
means the model isn't learning from the screenshot; high env_errors
means a translator/executor bug.
tasks/ holds 25 packages: 23 ERPNext accounting tasks plus 2 canary smoke
images (canary, cua-bench-canary). The failure-rate sweep in
docs/failure-rate.md covers the 19 tasks below
(difficulty is the raw config rating, 1 easiest to 7 hardest). That file has
the per-task pass rates, per-checkpoint decomposition, and cost per rollout.
| ID | Title | Difficulty |
|---|---|---|
po-001 |
Create and submit a Purchase Order | 1 |
po-002 |
Create and submit a two-line Purchase Order | 1 |
je-001 |
Record an expense accrual Journal Entry | 2 |
pe-001 |
Record payment against an outstanding Purchase Invoice | 2 |
sinv-from-dn-001 |
Create a Sales Invoice from a Delivery Note | 2 |
pi-3way-001 |
Three-way match with quantity discrepancy (PO to PR to PI) | 3 |
sinv-multi-001 |
Sales Invoice with multiple lines and a discount | 3 |
je-reversal-001 |
Accrual Journal Entry plus next-period reversal | 3 |
je-period-001 |
Accrue an expense in the correct period (cutoff judgment) | 3 |
pi-price-var-001 |
Purchase Invoice with price variance against the receipt | 4 |
tb-invest-001 |
Investigate and correct a Trial Balance misclassification | 4 |
pi-tolerance-fail-001 |
Resolve a Purchase Invoice blocked by over-billing tolerance | 4 |
po-reject-chain-001 |
Reject a PO with a downstream PI, trace the chain, resubmit | 5 |
bank-recon-001 |
Reconcile the Feb 2025 bank statement against the GL | 5 |
ar-aging-invest-001 |
Correct a misapplied customer payment hidden by AR aging | 5 |
stock-recon-001 |
Reconcile the Feb 2025 physical stock count against the system | 5 |
pi-misapp-001 |
Correct a Purchase Invoice misapplied to the wrong Receipt | 5 |
gl-subledger-recon-001 |
Reconcile AR subledger vs GL Debtors, fix a misapplied payment | 6 |
ap-subledger-recon-001 |
Reconcile AP subledger vs GL Creditors, fix a misapplied payment | 6 |
Four further task packages ship with graders and oracles but were not part of
this sweep: si-credit-block-001 (4), cogs-reclass-001 (5),
revenue-cutoff-001 (5), and vendor-statement-recon-001 (7).
Expected: frontier models fail a large share of the harder tasks within the step limit. A score < 1.0 is valid output, not a bug.
The graders above resist reward hacking through task design: state-based
checks, GL cross-validation, and episode-scoped creation timestamps.
cua-patterns/ is the complementary transcript-level layer. It is a
cross-environment library of spec-gaming detection predicates that read an
agent's trajectory and flag behaviors such as grader-state tampering,
verification bypass, answer lookup through side channels, instrument
manipulation, and fabricated completion. Each predicate is grounded in at
least two published cross-environment incidents (Baker, METR, Apollo,
Palisade, Petri, and published lab system cards), and the package ships with
adapters for ERPNext, OSWorld, and Petri transcripts. It is the more heavily
tested half of this repo (~100 unit tests). See
cua-patterns/README.md.
POST /reset {"task_id":"po-001"}
POST /step {"action":"click","x":..,"y":..}
POST /step/anthropic {"action":"left_click","coordinate":[x,y]}
POST /step/openai {"type":"click","x":..,"y":..}
POST /step/google {"name":"click_at","args":{"x":..,"y":..}}
# qwen and holo both translate to the canonical /step via cua_bench_harbor,
# using the Qwen `computer_use` action schema.
POST /evaluate {"task_id":"po-001"} → {score, checkpoints}
GET /health → {status, checks}
GET /tasks
The server listens on :5000 inside an env container.
# One venv, all three packages editable + dev tools (pytest, ruff)
uv venv
uv pip install -e ".[dev]" -e cua_bench_harbor -e cua-patterns
# Grader / agent-api unit tests (no Docker)
.venv/bin/pytest -v
# Harbor runner tests
.venv/bin/pytest cua_bench_harbor/tests
# Cross-env spec-gaming predicate tests
.venv/bin/pytest cua-patterns
# Lint / typecheck
.venv/bin/ruff check cua_bench/ tests/ graders/
.venv/bin/ty check cua_bench/When you change graders/, cua_bench/, or the golden snapshot, rebuild
cua-bench:erpnext-agent (and cua-bench:*-seeded if seeds changed).
cua-bench:erpnext-agent not found— runbash scripts/build_images.sh.- Task healthcheck times out — first run of
tasks/po-001takes 5–10 min while thesitesvolume initializes; subsequent runs reuse it. - Grader returns 0 — check
/logs/verifier/evaluate.jsoninside the run's log dir for the raw checkpoint breakdown. no matching manifest for linux/arm64/v8— Apple Silicon: the agent image is Playwright-based (multi-arch).