The growth rule in docs/PARITY.md / CLAUDE.md says to "periodically mine the reference's own e2e fixtures and real-world configs for scenarios the suite does not have yet" — this issue makes that actionable for the richest single vein: the reference CLI's own test suite.
The idea
@devcontainers/cli ships hundreds of e2e test configs (src/test/** in devcontainers/cli — per-scenario devcontainer.jsons, Dockerfiles, compose files, feature fixtures). Each one encodes a behavior upstream cared enough to pin. Running BOTH CLIs over them differentially asks, for free, "does deacon handle everything upstream's own tests exercise?" — coverage chosen by the reference's authors, not ours, which is exactly the blind-spot antidote.
Method (data-only — this is the suite's whole design)
- Harvest from the tag matching the oracle pin (v0.87.0, so fixtures and binary agree), not main.
- Triage the harvest: drop configs exercising out-of-scope surfaces (feature authoring/test/publish), dedupe against existing
parity/fixtures/ coverage (many shapes are already covered — the win is the tail), group the rest by area.
- For each survivor: vendor the fixture into
parity/fixtures/fx-upstream-*/ (MIT-licensed; keep upstream attribution in a fixture README line), author a live-differential case — a pure data edit, no new Rust.
- Run the batch. Every divergence follows the standard triage from
docs/PARITY.md: deacon wrong → issue + red case; reference deviates from spec → recorded conformant row; spec-silent deliberate → maintainer ruling before any tolerance.
- Land in batches by area, each batch's SPEC_STATUS rows in the same commit. Expect the first batch to be the most informative and the cheapest.
Guardrails
- Data and fixtures only; if a fixture needs a runner capability the model lacks, skip it and list it in the report — no machinery growth to chase a fixture.
- Environment-pinned fixtures (uid assumptions, huge images, arch-specific) get the same treatment the suite already documents: first run is a measurement, and
scripts/parity/prepull-fixture-images.sh must discover anything image-hungry.
- Volume control: this is a periodic mining pass, not a bulk import — a batch that adds 500 cases nobody triaged is worse than 30 that were each read.
Second vein, same method, separate pass: real-world devcontainer.json files from popular public repos.
The growth rule in
docs/PARITY.md/ CLAUDE.md says to "periodically mine the reference's own e2e fixtures and real-world configs for scenarios the suite does not have yet" — this issue makes that actionable for the richest single vein: the reference CLI's own test suite.The idea
@devcontainers/cliships hundreds of e2e test configs (src/test/**in devcontainers/cli — per-scenariodevcontainer.jsons, Dockerfiles, compose files, feature fixtures). Each one encodes a behavior upstream cared enough to pin. Running BOTH CLIs over them differentially asks, for free, "does deacon handle everything upstream's own tests exercise?" — coverage chosen by the reference's authors, not ours, which is exactly the blind-spot antidote.Method (data-only — this is the suite's whole design)
parity/fixtures/coverage (many shapes are already covered — the win is the tail), group the rest by area.parity/fixtures/fx-upstream-*/(MIT-licensed; keep upstream attribution in a fixture README line), author alive-differentialcase — a pure data edit, no new Rust.docs/PARITY.md: deacon wrong → issue + red case; reference deviates from spec → recorded conformant row; spec-silent deliberate → maintainer ruling before any tolerance.Guardrails
scripts/parity/prepull-fixture-images.shmust discover anything image-hungry.Second vein, same method, separate pass: real-world
devcontainer.jsonfiles from popular public repos.