Skip to content
vaiyrPublic

About

Computer-use agent benchmark on ERP accounting

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

cua-bench

cua-bench is an RL / eval environment for GUI Computer-Use Agents on ERPNext v15. It ships as Harbor tasks: each task packages a pre-seeded ERPNext stack, an in-container cua-bench agent-api (FastAPI + Playwright + Chromium), and a grader that scores the agent against real accounting state via the Frappe REST API.

  • What an agent sees: PNG screenshots of a Chromium window running ERPNext (default 1280×800).
  • What an agent does: click / type / key / scroll / drag / wait — in cua-bench's canonical schema or the lab's native format (anthropic, openai, google).
  • How it's graded: state-based checkpoints with weighted partial credit, GL cross-validation, and episode-scoped creation timestamps to resist common reward hacks; the open attack surface is documented in docs/reward-hacking.md.

This project was previously developed under the name veraf. Some committed historical artifacts (the sweep summary in jobs/) predate the rename. The git history was squashed from a private repository at that re-home, so the current commits date from the public re-home rather than reflecting the original development timeline.

Repo layout

cua-bench/
├── docker/
│   ├── agent.Dockerfile            # "main" image: Playwright + agent-api + graders
│   ├── erpnext-seeded.Dockerfile   # ERPNext backend with seed baked in
│   ├── mariadb-seeded.Dockerfile   # MariaDB with golden snapshot baked in
│   ├── canary.Dockerfile           # Minimal smoke image
│   └── cua-bench-dind.Dockerfile   # DinD snapshot template (bakes GHCR creds for Daytona)
├── cua_bench/                          # FastAPI agent-api (runs inside env container on :5000)
│   └── {server,environment,executor,reset,screenshot,…}.py
├── cua_bench_harbor/                   # Harbor-native multi-vendor CUA runner
│   └── src/cua_bench_harbor/{agent,runner,clients,translators,trajectory}.py
├── cua-patterns/             # Cross-env spec-gaming detection predicate library
│   └── src/cua_patterns/{predicates,judges,*_adapter}.py
├── graders/erpnext/                # Grader defs (JSON + Python) baked into the agent image
├── tasks/<id>/                     # Harbor task packages
│   ├── task.toml                   # compute / image / healthcheck
│   ├── instruction.md              # what the agent is told
│   ├── environment/docker-compose.yaml
│   ├── solution/solve.sh           # oracle (upper-bound check)
│   └── tests/test.sh               # verifier → reward
├── db-snapshot/golden.sql          # committed; restored per-episode by /reset
├── scripts/build_images.sh         # builds all prebuilt images
└── docs/reward-hacking.md

Quickstart

cua-bench runs on Daytona cloud sandboxes. Each trial spins up its own isolated Docker-in-Docker sandbox with the ERPNext stack — no local Docker Desktop, no port conflicts, trivial horizontal scaling (we routinely run 30+ trials in parallel).

# 1. Configure
cp .env.example .env
# Edit .env — you need at minimum:
#   DAYTONA_API_KEY=dtn_...                 (https://app.daytona.io/dashboard)
#   DAYTONA_API_URL=https://app.daytona.io/api
#   one of ANTHROPIC_API_KEY / OPENAI_API_KEY / GEMINI_API_KEY
uv venv && uv pip install -e cua_bench_harbor

# 2. Run one task on Daytona
EXECUTOR=daytona bash scripts/run_sweep.sh anthropic 1 po-001

# 3. Inspect what happened
python scripts/inspect_trajectory.py jobs/anthropic/po-001/r1

# 4. Full failure-rate sweep: 19 tasks × 2 vendors (anthropic + openai) × N rollouts
EXECUTOR=daytona bash scripts/run_sweep.sh all 3
python scripts/build_sweep_table.py > docs/failure-rate.md

The task packages reference prebuilt images at ghcr.io/vaiyr/cua-bench-*. Those images are built locally with scripts/build_images.sh or published at release, so a fresh clone will not pull them until you build them or a tagged release is cut.

Read on for details, troubleshooting, and the full set of tasks.

Prerequisites

  • A Daytona account with API key (sign up)
  • Python 3.11+ and uv
  • harbor ≥ 0.3
  • At least one of: ANTHROPIC_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY
  • Optional (for ad-hoc runs only, not swept): QWEN_API_KEY (Alibaba Dashscope), HAI_API_KEY (Hcompany Holo)

Network policy on Daytona: by default Daytona restricts sandbox egress in a way that blocks some benchmarks. Apply coupon HARBOR_NETWORK on your Daytona account to remove the restriction before running.

One-time Daytona setup (per Daytona org)

cua-bench uses a custom DinD snapshot (cua-bench-dind) that bakes in GHCR pull credentials, so the in-sandbox Docker daemon can pull private images from ghcr.io/vaiyr/cua-bench-*. This is per-org, one-time.

  1. Register GHCR with Daytona (Dashboard → Registries → Add Registry, or POST /api/docker-registry with your GitHub PAT — write:packages and read:packages scopes).
  2. Build + push the cua-bench-dind image with your GHCR creds baked in (see docker/cua-bench-dind.Dockerfile for the template), or adapt the snippet in the Daytona Dashboard Snapshots UI. The snapshot name cua-bench-dind is what scripts/run_sweep.sh passes via --ek dind_snapshot=cua-bench-dind.
  3. Flag: 10 GB per-sandbox disk is the default org cap. The stack fits, but barely. If you see Disk request 20GB exceeds maximum allowed, email support@daytona.io to raise the quota.

Sweep concurrency and rate limits

Three rate limits to plan around when sweeping at scale. Defaults in scripts/run_sweep.sh are tuned for the most restrictive case; bump them once you've confirmed your account/setup tolerates it.

limit scope default cap symptom when hit
MAX_JOBS=5 Anthropic concurrent sandboxes safe none — well below all caps
OPENAI_MAX_JOBS=6 OpenAI concurrent sandboxes OpenAI Tier 3 (2M TPM) llm_failure mid-trial
Daytona vCPU concurrent sandboxes total 25 (default org plan) sandbox queueing
Docker Hub anonymous 100 pulls / 6hr / IP per Daytona egress IP RuntimeError: toomanyrequests, trial exits with null reward

Docker Hub rate limit was the dominant failure mode in the 2026-04-16 sweep — at 13 concurrent sandboxes, anonymous pulls of redis:6.2-alpine exhausted the limit and ~12% of trials returned null rewards. Two mitigations, in order of preference:

  1. Mirror the offending image. The compose at tasks/po-001/environment/docker-compose.yaml now references mirror.gcr.io/library/redis:6.2-alpine instead of redis:6.2-alpine. Google's mirror is free, no auth, no rate limit. If you add new tasks that pull docker.io images, mirror them the same way.
  2. Authenticate Docker Hub inside the DinD snapshot (boosts to 200 pulls/6hr — temporary workaround, doesn't scale to large sweeps). Bake docker login into the cua-bench-dind snapshot rebuild rather than per-sweep, so creds don't leak into harbor logs and login latency doesn't multiply across sandboxes.

Tuning for higher OpenAI tiers. TPM caps roughly: Tier 2 = 1M, Tier 3 = 2M ($100 spent), Tier 4 = 4M ($250), Tier 5 = 40M ($1k). Default is OPENAI_MAX_JOBS=6 (sized for Tier 3); bump higher in proportion to TPM headroom once you cross.

Pre-caching transitive images in cua-bench-dind (the proper long-term fix) requires rebuilding the snapshot via the Daytona dashboard. Spin up a sandbox from the current cua-bench-dind, run docker pull for every image referenced in any task compose, then save as cua-bench-dind:v2. Update DIND_SNAPSHOT=cua-bench-dind:v2 in scripts/run_sweep.sh env or .env.

Setup

1. Clone and configure

git clone <repo> && cd cua-bench
cp .env.example .env        # then edit: DAYTONA_API_KEY=..., ANTHROPIC_API_KEY=...

2. Install the Harbor runner

cd cua_bench_harbor
uv venv && uv pip install -e ".[dev]"
.venv/bin/pytest          # sanity

3. (One-time per image change) Rebuild and push prebuilt images to GHCR

cua-bench's prebuilt images live on GitHub Container Registry at ghcr.io/vaiyr/cua-bench-*. Daytona sandboxes pull them via a custom cua-bench-dind snapshot that bakes in GHCR credentials. You only need to rebuild if cua_bench/, graders/, docker/, or db-snapshot/golden.sql changes.

erpnext-seeded needs sites-data/ (the initialized Frappe site dir). If missing, extract it from a running backend:

docker cp cua-bench-backend-1:/home/frappe/frappe-bench/sites/. sites-data/

Then rebuild + push all four images (forces linux/amd64 for Daytona):

PUSH=1 CUA_BENCH_TAG=v2 bash scripts/build_images.sh

Fast path when only cua_bench/ or graders/ changed (rebuilds just the agent image):

PUSH=1 CUA_BENCH_TAG=v2 bash scripts/rebuild_agent.sh

After pushing a new tag, update the docker_image field in tasks/*/task.toml and the image: fields in tasks/po-001/environment/docker-compose.yaml to the new tag.

Running a task

All run paths default to Daytona when EXECUTOR=daytona is set. The single-task wrapper and the sweep both pass the Daytona-specific flags (-e daytona --ek dind_snapshot=cua-bench-dind --override-storage-mb 10000) for you.

# Sweep-of-one is the cleanest single-task entry point:
EXECUTOR=daytona bash scripts/run_sweep.sh anthropic 1 po-001
EXECUTOR=daytona bash scripts/run_sweep.sh google 1 pe-001

The raw harbor run form is also supported — useful for one-off debugging:

# Smoke test: canary task (no ERPNext, ~30s)
harbor run \
  --env-file .env \
  -p tasks/cua-bench-canary \
  --agent-import-path cua_bench_harbor.agent:CuaBenchComputerUseAgent \
  -e daytona --ek dind_snapshot=cua-bench-dind --override-storage-mb 10000 \
  --ak vendor=anthropic --ak max_steps=20 --no-delete

# Real task: create + submit a Purchase Order (~5–7 min end-to-end)
harbor run \
  --env-file .env \
  -p tasks/po-001 \
  --agent-import-path cua_bench_harbor.agent:CuaBenchComputerUseAgent \
  --model anthropic/claude-opus-4-7 \
  -e daytona --ek dind_snapshot=cua-bench-dind --override-storage-mb 10000 \
  --ak vendor=anthropic --ak max_steps=50 --no-delete

# Same task with GPT-5.4 (Responses API + built-in computer tool)
harbor run \
  --env-file .env \
  -p tasks/po-001 \
  --agent-import-path cua_bench_harbor.agent:CuaBenchComputerUseAgent \
  --model openai/gpt-5.4 \
  -e daytona --ek dind_snapshot=cua-bench-dind --override-storage-mb 10000 \
  --ak vendor=openai --ak max_steps=50 --no-delete

# Same task with Gemini 3 flash preview
harbor run \
  --env-file .env \
  -p tasks/po-001 \
  --agent-import-path cua_bench_harbor.agent:CuaBenchComputerUseAgent \
  --model google/gemini-3-flash-preview \
  -e daytona --ek dind_snapshot=cua-bench-dind --override-storage-mb 10000 \
  --ak vendor=google --ak max_steps=50 --no-delete

Qwen and Holo (ad-hoc only)

Qwen3-VL and Hcompany Holo3 are wired as additional vendors but excluded from run_sweep.sh. Both currently show capability-bound failure on po-001 — the models localize the Items table row at y≈830, which is off-canvas (viewport is 800 tall), and don't self-correct even with textual or visual feedback signals. The traces are clean and useful to inspect, but sweeping them across 19 tasks × N rollouts burns budget without changing the failure-rate story. Re-enable in the sweep once a stronger variant (e.g. qwen3-vl-235b-a22b-thinking) is worth benchmarking — see the comment in scripts/run_sweep.sh.

# Qwen3-VL-235B-A22B-Instruct via Alibaba Dashscope (OpenAI-compat, native
# tool_calls). Requires QWEN_API_KEY.
bash scripts/run_task.sh qwen po-001

# Holo3-122B-A10B via Hcompany hub (OpenAI-compat, structured JSON output +
# thinking mode — does NOT use native tool_calls). Requires HAI_API_KEY.
bash scripts/run_task.sh holo po-001

# Override model within a vendor:
QWEN_MODEL=qwen/qwen3-vl-235b-a22b-thinking bash scripts/run_task.sh qwen po-001
HOLO_MODEL=holo/holo3-35b-a3b                bash scripts/run_task.sh holo po-001

Gotchas if you change these clients:

  • Holo coord convention is prompt-dependent. With an explicit "viewport is WxH" hint in the system prompt (which we send), Holo emits absolute pixel coords. Without it, it defaults to a normalized 0–999 grid. Don't remove the hint unless you also add denormalization.
  • Holo thinking mode eats tokens. We set max_tokens=4000 because reasoning_content can be ~1–2k tokens before the JSON action is emitted. With the default 800, responses come back finish=length, content=None.
  • Holo rejects tiny screenshots. The Hcompany API returns a 500 on images smaller than ~a few hundred pixels per side. Our 1280×800 task screenshots are fine; unit tests that build 10×10 probe images are not.
  • Holo accumulates context aggressively. We trim the in-flight message history to the last 2 screenshots (_trim_screenshots) — without this the API 400s around step 6.

## Multi-vendor failure-rate sweep

`scripts/run_sweep.sh` runs every task N times per vendor and writes a
JSONL summary to `jobs/sweep-summary.jsonl`. `scripts/build_sweep_table.py`
renders it as a markdown table with Wilson 95% CIs.

```bash
# Smoke first: one easy task × each vendor × 1 rollout.
EXECUTOR=daytona bash scripts/run_sweep.sh all 1 po-001

# Pilot: all 19 tasks × 2 vendors (anthropic + openai) × 1 rollout (~15-30 min at 30-way, ~$30).
EXECUTOR=daytona bash scripts/run_sweep.sh all 1

# Full sweep: n=3 × 19 tasks × 2 vendors (~1h at 30-way, ~$75–$150).
EXECUTOR=daytona bash scripts/run_sweep.sh all 3

# Render the table.
.venv/bin/python scripts/build_sweep_table.py > docs/failure-rate.md

Expected cost: ~$300–800 for the full sweep depending on task lengths and cache-hit rates. Each client logs cached-token counts — verify your cache is warming via the ATIF trajectories before letting the long tail run. Cap per-rollout spend with --ak max_cost_usd=....

Model availability caveat: gpt-5.4 and gemini-3-flash-preview are both relatively new/preview. If your account doesn't have access, the first smoke run will fail with a 404 from the vendor API — swap to gpt-4o or gemini-2.5-flash respectively (neither has the same computer-use fidelity, but smoke still validates plumbing).

Oracle check (bypasses the UI, hits Frappe REST directly — upper bound):

harbor run -p tasks/po-001 --agent oracle \
  -e daytona --ek dind_snapshot=cua-bench-dind --override-storage-mb 10000 \
  --env-file .env --no-delete

The verifier (tasks/<id>/tests/test.sh) calls POST /evaluate inside the env container and writes the grader score to /logs/verifier/reward.txt.

Debugging a failed run

# Per-step actions, repeat-coord smell, cache-hit %, total cost.
python scripts/inspect_trajectory.py jobs/runs/google-po-001-*

# Rendered videos and per-step PNGs:
open jobs/runs/<run>/<task>__*/artifacts/videos/*.webm
ls   jobs/runs/<run>/<task>__*/agent/screenshots/

# Inner loop when editing cua_bench/ (executor, server, reset, etc.):
PUSH=1 CUA_BENCH_TAG=v2 bash scripts/rebuild_agent.sh   # rebuild + push to GHCR
EXECUTOR=daytona bash scripts/run_sweep.sh openai 1 po-001

inspect_trajectory.py is the fastest way to spot model-quality vs plumbing issues: a "stuck-coord" cluster + repeated typed text usually means the model isn't learning from the screenshot; high env_errors means a translator/executor bug.

Tasks

tasks/ holds 25 packages: 23 ERPNext accounting tasks plus 2 canary smoke images (canary, cua-bench-canary). The failure-rate sweep in docs/failure-rate.md covers the 19 tasks below (difficulty is the raw config rating, 1 easiest to 7 hardest). That file has the per-task pass rates, per-checkpoint decomposition, and cost per rollout.

ID Title Difficulty
po-001 Create and submit a Purchase Order 1
po-002 Create and submit a two-line Purchase Order 1
je-001 Record an expense accrual Journal Entry 2
pe-001 Record payment against an outstanding Purchase Invoice 2
sinv-from-dn-001 Create a Sales Invoice from a Delivery Note 2
pi-3way-001 Three-way match with quantity discrepancy (PO to PR to PI) 3
sinv-multi-001 Sales Invoice with multiple lines and a discount 3
je-reversal-001 Accrual Journal Entry plus next-period reversal 3
je-period-001 Accrue an expense in the correct period (cutoff judgment) 3
pi-price-var-001 Purchase Invoice with price variance against the receipt 4
tb-invest-001 Investigate and correct a Trial Balance misclassification 4
pi-tolerance-fail-001 Resolve a Purchase Invoice blocked by over-billing tolerance 4
po-reject-chain-001 Reject a PO with a downstream PI, trace the chain, resubmit 5
bank-recon-001 Reconcile the Feb 2025 bank statement against the GL 5
ar-aging-invest-001 Correct a misapplied customer payment hidden by AR aging 5
stock-recon-001 Reconcile the Feb 2025 physical stock count against the system 5
pi-misapp-001 Correct a Purchase Invoice misapplied to the wrong Receipt 5
gl-subledger-recon-001 Reconcile AR subledger vs GL Debtors, fix a misapplied payment 6
ap-subledger-recon-001 Reconcile AP subledger vs GL Creditors, fix a misapplied payment 6

Four further task packages ship with graders and oracles but were not part of this sweep: si-credit-block-001 (4), cogs-reclass-001 (5), revenue-cutoff-001 (5), and vendor-statement-recon-001 (7).

Expected: frontier models fail a large share of the harder tasks within the step limit. A score < 1.0 is valid output, not a bug.

Reward-hacking detection: cua-patterns/

The graders above resist reward hacking through task design: state-based checks, GL cross-validation, and episode-scoped creation timestamps. cua-patterns/ is the complementary transcript-level layer. It is a cross-environment library of spec-gaming detection predicates that read an agent's trajectory and flag behaviors such as grader-state tampering, verification bypass, answer lookup through side channels, instrument manipulation, and fabricated completion. Each predicate is grounded in at least two published cross-environment incidents (Baker, METR, Apollo, Palisade, Petri, and published lab system cards), and the package ships with adapters for ERPNext, OSWorld, and Petri transcripts. It is the more heavily tested half of this repo (~100 unit tests). See cua-patterns/README.md.

Agent API (inside the env container)

POST /reset           {"task_id":"po-001"}
POST /step            {"action":"click","x":..,"y":..}
POST /step/anthropic  {"action":"left_click","coordinate":[x,y]}
POST /step/openai     {"type":"click","x":..,"y":..}
POST /step/google     {"name":"click_at","args":{"x":..,"y":..}}
# qwen and holo both translate to the canonical /step via cua_bench_harbor,
# using the Qwen `computer_use` action schema.
POST /evaluate        {"task_id":"po-001"}  → {score, checkpoints}
GET  /health          → {status, checks}
GET  /tasks

The server listens on :5000 inside an env container.

Development

# One venv, all three packages editable + dev tools (pytest, ruff)
uv venv
uv pip install -e ".[dev]" -e cua_bench_harbor -e cua-patterns

# Grader / agent-api unit tests (no Docker)
.venv/bin/pytest -v

# Harbor runner tests
.venv/bin/pytest cua_bench_harbor/tests

# Cross-env spec-gaming predicate tests
.venv/bin/pytest cua-patterns

# Lint / typecheck
.venv/bin/ruff check cua_bench/ tests/ graders/
.venv/bin/ty check cua_bench/

When you change graders/, cua_bench/, or the golden snapshot, rebuild cua-bench:erpnext-agent (and cua-bench:*-seeded if seeds changed).

Troubleshooting

  • cua-bench:erpnext-agent not found — run bash scripts/build_images.sh.
  • Task healthcheck times out — first run of tasks/po-001 takes 5–10 min while the sites volume initializes; subsequent runs reuse it.
  • Grader returns 0 — check /logs/verifier/evaluate.json inside the run's log dir for the raw checkpoint breakdown.
  • no matching manifest for linux/arm64/v8 — Apple Silicon: the agent image is Playwright-based (multi-arch).

About

Computer-use agent benchmark on ERP accounting

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages