- Status: done
- Date: 2026-08-11
- Specs touched: docs/specs/UPDATES.md (# 3, no text change — this realizes it)
Fifth implementation slice of the update work designed in #369, after control-plane-image-ledger.md gave the box a declaration to write. Closes #380.
New package internal/hostagent/cpupdate — the stream-B apply and rollback that UPDATES.md # 3 specifies and # 8 says is shared by both profiles ("the expensive half of the update machinery is shared, and we only build it once"). It is trigger-free by design: nothing in it knows whether the target came from a signed release manifest, from the cloud, or from an admin typing two refs. A caller hands it a pair; it applies it or puts the box back.
The transaction:
- Pull every moved ref. Nothing on the box has changed yet, so a pull failure costs nothing.
- If the brain moved: stop it, then snapshot its SQLite. In that order, because the brain opens the database with
journal_mode=WALand a copy taken under a live writer restores into a corrupt database — a safety net that only fails when used. - Write the declaration — ledger and staged compose — before starting anything. The # 8.3 handoff.
- Recreate only what moved, the brain from
brainlaunch's own run spec. - Health-check what moved — the brain on
/healthz, the UI on/, bounded at 60s. - On any failure from step 2 on: revert both refs, restore the snapshot, put the containers back, and report which step failed.
Both refs revert even when only one moved. A coordinated ship is tested as a pair, so a box left running half of one is in a combination nobody has ever run.
The brain is recreated from brainlaunch.RunSpecFor, not from a second docker run builder. An updated brain has to be identical to a first-boot one except for the image. A second builder would drift the first time a mount or env var is added, the box would boot fine, and the divergence would appear only after an update — the worst possible time to find out.
Any HTTP response counts as "serving", including 500. The question is "did the container come up and bind its port", not "is every route behaving". A brain answering 500 somewhere is still a running brain, and reverting the control plane over it would take a working box backwards. /healthz is 200-as-soon-as-serving by construction (brain-healthz-probe.md).
A test caught a real ordering subtlety in this code. The first version of the ordering assertion counted any container call, and it failed: the brain is removed (step 2) before the declaration is written (step 3). That turned out to be correct rather than a bug — the removal is needed for a consistent snapshot, and a crash in that window leaves the ledger naming the old pair, so the next boot brings the old brain back. The claim UPDATES.md makes is about starting a container on an image the declaration does not name, so the assertion now watches starts. The distinction is written down because it is exactly the kind of thing a later change could quietly break.
UPDATES.md# 3 step 3 — pull, snapshot, recreate-only-what-changed, 60s health wait; step 4 — revert both on failure of either; the 7-day retention and GC.UPDATES.md# 8.3 / # 8.4 step 3 — one actor, declaration-before-recreate, and the same transaction under both triggers.CLAUDE.md# Go code discipline — consumer-sideDockerandProberinterfaces,slogwithimage/step/err.
Eight findings on the PR, all fixed in it. Two were correctness bugs worth naming:
- The recreate skipped
brainlaunch's protocol-major lockstep check. A brain image declaring a major this host-agent does not speak would have started, answered/healthz(the brain serves HTTP before it needs host-agent), and committed — then failed at the next reboot, whenbrainlaunchreads that ref out of the ledger, applies the guard, and refuses. The box would have had no brain and no failed update to point at. The check now runs before the start, and a mismatch is arecreatefailure, so the revert fires. - The revert ran on the caller's context. A cancelled or expired context — a job past its deadline, a client that hung up, host-agent shutting down — would have made every
dockercall in the rollback fail instantly, leaving the box on the new brain with a ledger naming the old pair. The revert now runs oncontext.WithoutCancelwith its own 5-minute budget: a cancelled apply is exactly when the rollback matters most.
Also fixed: the snapshot dir is cleared before reuse (a retried update could otherwise carry a stale -wal from the earlier attempt into the new snapshot); GC no longer deletes the snapshot the retained rollback target needs when both generations name the same brain; waitHealthy probes only what moved, so a pre-existing UI outage cannot revert a good brain-only update; two FailureMode values were pointing at the wrong step; BrainCfg.Image's doc contradicted the code in a way that would have produced an unrollbackable ledger; and the snapshot now fsyncs its directory, not just its files.
The first cancellation test was vacuous and a mutation check caught it: the Docker fake ignored the context entirely, so it kept "running docker" after cancellation and the test passed with the fix removed. The fake now fails on a dead context the way exec.CommandContext does. Both Block fixes are mutation-checked.
- Nothing calls
Applyyet. The trigger (#381) is the next slice: this ships the transaction and no way to start one. - Proven against a fake Docker only. Fifteen transaction tests plus four snapshot tests cover both-changed, UI-only, brain-only, declaration ordering, declaration restore, health-failure revert with a simulated schema migration, pull failure, no-op, GC in and out of the window, and the no-ledger derivation — but no real daemon, no real registry, no real brain restart. That is #382, and it is the gap that matters most here.
- The health probe reaches containers by IP (
docker inspect), because the control-plane containers publish no ports. That works from the host, where host-agent runs, and is not a path anything else uses today. brainPortanduiPortare constants. The brain's port iscmd/brain'sMALMO_LISTENdefault; a box that overrode it would be probed on the wrong port. Nothing overrides it today, and threading the value through would mean host-agent reading the brain's environment.- No retry, and no "three consecutive failures then pin" (# 3 step 5). One attempt, one revert. The counter belongs with the trigger that decides when to try again.
- GC deletes images and the snapshot, but never the ledger's own history, since there is only ever one previous generation (# 3: "single-generation rollback").
- A revert that itself fails is reported, not repaired.
Result.RevertErris the operator's signal; there is no third generation to fall back to and inventing one would be guessing.
- #381 — the trigger. A minimal
system-updatejob on the host socket plus the admin endpoint that starts it, taking explicit refs. - #382 — the QEMU proof. A real update and a real failed-update-then-revert on a booted box, against a registry inside the guest.
- The box → cloud target and report (
UPDATES.md# 8.1, # 8.4 step 5), still blocked on the box↔cloud credential (NEXT.mdTier 1).