Skip to content

chore(ci): Podman lane intermittently fails an exec with conmon bytes "" — two hypotheses refuted #723

Description

@pofallon

Symptom

smoke_compose_override_command::test_compose_override_command_lifecycle_runs fails intermittently on the Podman lane:

Container command failed with exit code 255: Error: container create failed
(no logs from conmon): conmon bytes "": readObjectStart: expect { or n, but found ^@

podman is failing to start an exec session's conmon. It surfaces as a bare deacon up failed with postCreateCommand … exit code 255, which reads like a lifecycle defect and is not one. The env probe fails the same way ~0.6s earlier, so two consecutive podman exec calls failed.

History

Run Result
#715 run 1 different failure — a genuine test defect (hardcoded docker), fixed
#715 run 2 conmon → passed on re-run
#717 conmon
#720 passed
#722 conmon — after the concurrency change below

Never reproduced locally: the test passes 5/5 serially and the full suite passes 433/433 against real rootless podman 4.9.3 in the dev container.

Refuted: contention

smoke-lite was halved from 4 to 2 (#718) on the theory that concurrent rootless-podman compose projects were starving conmon — plausible, since #715 moved those projects from docker to podman for the first time. The failure recurred unchanged at 2. Reverted; it cost ~14% on every smoke lane and bought nothing.

Worth noting for whoever picks this up: at max-threads = 2 the pair running concurrently was exactly the two sibling tests in that binary, one of which passed. So "two concurrent podman compose projects" is not sufficient to cause it either.

Refuted: container not running

get_primary_container_id (crates/core/src/compose.rs:1723) matches a service by name with no state check, while its sibling handle_port_events filters on s.state == "running". A not-yet-running or already-exited container looked like a strong candidate.

Measured against rootless podman 4.9.3:

  • podman compose up -d returns with the container already running; polling compose ps immediately never showed an intermediate state.
  • exec into an exited compose container gives can only create exec sessions on running containers: container state improper — a clean, different error.

So this is not it. (The missing state check is still an asymmetry worth tidying, but it is not this bug and should not be "fixed" as though it were.)

What is known

  • It is podman failing to create the exec session, not deacon misusing it.
  • readObjectStart: expect { or n, but found ^@ means podman read NUL from conmon's sync pipe — conmon died or wrote nothing.
  • Typical causes for that shape are resource ceilings (pids, fds, memory / OOM-kill of conmon), not container state.
  • No OOM or resource message appears in the job log, but nothing currently looks.

Next step is evidence, not a third hypothesis

A failure-only diagnostics step is added to the Podman job: memory, load, ulimit -a, pids.max, conmon presence and version, podman info, container states, and dmesg tail. Everything is || true and if: failure(), so it reports and never masks the real failure.

  • Wait for the next occurrence and read the diagnostics.
  • If it is a resource ceiling, raise it or reduce whole-suite concurrency — not smoke-lite alone, which has been tried.
  • Only then consider a change to deacon.

Deliberately not doing: quarantining the test, adding a retry, or reducing concurrency further. Each would hide the signal, and two of them would hide it permanently.

Activity

  1. pofallon commented on Aug 29, 2026

    @pofallon
    ContributorAuthor

    The diagnostics fired, and the first thing they show is that the lane runs a different podman than anyone thought

    From the failing Podman job on #725 (run 33214...):

    Setting up podman (4.9.3+ds1-1ubuntu0.2) ...     <- what apt installed, into /usr/bin
    Client:       Podman Engine
    Version:      5.8.4                              <- what actually ran
    

    The ubuntu-24.04 runner image already ships podman 5.8.4 under /usr/local/bin, which precedes /usr/bin on PATH. So sudo apt-get install -y podman in the setup step does not provision the podman the lane uses — it installs an unused second copy, and pulls in a second conmon alongside it:

    Path Version
    command -v conmon /usr/bin/conmon 2.1.10 (apt)
    what podman resolves /usr/local/lib/podman/conmon 2.2.1 (bundled)

    podman info confirms the rest of the 5.x stack is in play: buildahVersion 1.43.2, crun 1.28 from /usr/local/bin, netavark, pasta.

    Why this matters for this issue

    It explains "never reproduces locally." Every local reproduction attempt used podman 4.9.3 — the version apt reports and the version in the dev container. CI fails on 5.8.4. Different major version, different conmon, different network backend (netavark/pasta vs slirp4netns). The two refuted hypotheses were tested against the wrong substrate.

    I am deliberately not claiming the version skew is the cause. Two hypotheses have already been wrong here, and "there are two installs" is not the same fact as "that is why conmon fails". What it does is make a real reproduction possible for the first time.

    Other diagnostics, for the record

    • Memory was not tight: 14.5 GB available of 16 GB, load 2.48 on 4 CPUs.
    • pids.max: n/a, max user processes 63795, open files 65536 — no ceiling in sight.
    • No OOM in dmesg.
    • Leftover containers were accumulating (14 listed, 3 still Up), which is worth a separate look but is not obviously this.

    Next

    • Reproduce against podman 5.8.4, not 4.9.3.
    • Decide what the setup step should do: drop the apt install (the runner already has 5.8.4) or pin deliberately to a chosen version. Either way it should stop claiming to install the podman under test when it does not.
  2. pofallon commented on Aug 29, 2026

    @pofallon
    ContributorAuthor

    The lane now runs a coherent podman stack, and the flake survived it

    #726 removes the apt-get install -y podman that was feeding a 4.9.3 configuration to the 5.8.4 engine the lane actually runs. The setup step passed, so the correction is validated on its own terms:

    === podman as provisioned by the runner image ===
    /usr/local/bin/podman
    Version:      5.8.4
    === /etc/containers ===
    containers.conf  policy.json  registries.conf  seccomp.json  storage.conf   <- all image-provided
    === rootless helpers ===
      pasta          /usr/local/bin/pasta
      slirp4netns    <not on PATH>
      crun           /usr/local/bin/crun
      runc           /usr/local/bin/runc
    

    Nothing was depending on the apt package: /etc/containers is shipped complete by the image (dated Jun 26 / Aug 23, i.e. from the image build, not from apt), and pasta — podman 5's default rootlessNetworkCmd — is already at /usr/local/bin. The apt install was purely overwriting image config with an older release's.

    The diagnostics are also no longer self-contradictory:

    === conmon (the one PODMAN resolves, not the one on $PATH) ===
      podman resolves: /usr/local/lib/podman/conmon
    conmon version 2.2.1
      none on PATH (expected: apt's podman is no longer installed)
    

    What that rules out

    The version/config mismatch is not the cause of this flake. Same test, same error, on a stack with no mixed versions:

    FAIL [ 11.605s] (149/436) deacon::smoke_compose_override_command test_compose_override_command_lifecycle_runs
    Summary [132.961s] 436 tests run: 435 passed, 1 failed, 3 skipped
    
    Container command failed with exit code 255: Error: container create failed
    (no logs from conmon): conmon bytes "": readObjectStart: expect { or n, but found ...
    Error: Lifecycle error: Lifecycle command failed (source: config) in phase postCreate
    with exit code 255: touch /tmp/deacon-lifecycle-marker
    

    So the hypothesis list now stands at three refuted: daemon contention (#718, reverted in #724), container-not-running (an exec into an exited container gives can only create exec sessions on running containers, a different error), and now conmon/podman version skew.

    What is still standing

    The failure is podman failing to spin up the conmon for an exec session — postCreate — against a compose-created container, i.e. one created through podman's API socket by docker-compose rather than by podman run. Every observed instance has been that same test. Worth probing next, in rough order of cheapness:

    • whether the exec is racing the container's own startup in a way that podman ps state does not expose (distinct from the refuted not-running hypothesis)
    • whether an API-socket-created container's exec path differs from a CLI-created one
    • rootless resource ceilings in the user namespace at that moment (the earlier diagnostics ruled out host memory, pids and fd ceilings, and OOM)

    Observed rate on that one test across recent runs: roughly 6 fail / 3 pass. It is currently the sole blocker on both #725 and #726, neither of which touches the code path it exercises.

  3. pofallon commented on Aug 29, 2026

    @pofallon
    ContributorAuthor

    Reopened — closed by accident, not by resolution.

    The merge commit for #726 contained the sentence This does NOT fix <this issue's number>, written specifically to avoid overclaiming. GitHub's closing keywords fire on the verb next to the number regardless of the surrounding prose: negated, quoted or merely described, the pattern fires. This repository's HANDOFF.md documents that trap, and I walked into it anyway.

    Nothing is fixed. State is unchanged from the previous comment: three hypotheses refuted (daemon contention, container-not-running, podman/conmon version skew), the failure still intermittent at roughly half of runs, still confined to test_compose_override_command_lifecycle_runs.

    Observed again today across four Podman-lane runs on #725/#726: fail, pass, fail, pass.

  4. pofallon commented on Aug 29, 2026

    @pofallon
    ContributorAuthor

    A measurement pass, not a fourth hypothesis

    Nothing here changes the state of the bug. What it does is put numbers under "never reproduces locally", correct two things the error message implies but does not mean, and add the instrumentation that would make the next occurrence self-describing. PR #728.

    Local non-reproduction, quantified

    Substrate: podman 4.9.3, conmon 2.1.10, runc 1.5.1, netavark, cgroups v2 + cgroupfs, vfs, in the dev container.

    Shape Runs Failures
    test_compose_override_command_lifecycle_runs alone, serial 20 0
    Same test, 4 concurrent copies (6 rounds) 24 0
    The whole Podman lane, mvp-integration -E 'not binary(=parity_docker)', 436 tests 3 0

    The earlier note said "passes 5/5 serially" — that was n=5. This is n=44 on the test plus three whole suites, including the compose-through-the-API-socket path the issue singled out. Every run went through podman compose against the API socket, so "compose-created container + exec" alone is not sufficient on 4.9.3.

    Two corrections to what the error implies

    Both read out of the reference's own source at v5.8.4.

    1. container create failed (no logs from conmon) is a shared message, not a claim about container creation. It is produced inside readConmonPipeData (libpod/oci_conmon_common.go:1443), and both the container-create path and the exec path call it. So the wording is not evidence against the exec framing — the exec framing is right.

    2. conmon is spawned successfully and then dies silently. In ExecContainer (libpod/oci_conmon_exec_common.go:78-82), execCmd.Wait() runs before readConmonPipeData, and a failure to spawn conmon returns cannot run conmon: …. We get the later error, so:

    • conmon's parent process exited 0 — this is not a fork/exec failure;
    • the sync pipe then returned EOF with zero bytes (conmon bytes "");
    • conmon wrote nothing to stderr either, since podman's stderr carried only the one Error: line.

    Whatever kills it happens after fork and before it writes sync data.

    A trap for whoever instruments this next

    Do not run the lane with podman --log-level=debug. It is the instinctive move and it destroys the best evidence:

    var ociLog string
    if logrus.GetLevel() != logrus.DebugLevel && r.supportsJSON {
        ociLog = c.execOCILog(sessionID)
    }

    — libpod/oci_conmon_exec_common.go:50-53, v5.8.4

    ociLog is the file readConmonPipeData reads to replace the generic message with the OCI runtime's real error, and podman only writes it when the level is not debug. That we saw the generic message means the log existed and was empty or unparseable — i.e. crun never wrote an error either, which is consistent with conmon dying before it ever ran crun.

    Refuted cheaply

    The test harness points TMPDIR at a per-test-process temp dir (support::deacon_command), which every podman child inherits, and that looked like a candidate for divergent runtime paths between the API service and the CLI. Measured: TMPDIR=… podman info reports the same RunRoot and GraphRoot. Not it.

    What #728 adds

    Failure-time state capture. The three compose-override tests now call support::runtime_state_dump(workspace) on the up-failure path before deacon_down. The job-level if: failure() diagnostics can never answer the question that matters — was the container running when the exec was attempted — because the test removes it first. Watched to fail with postCreateCommand set to exit 3:

    --- every container the runtime can see ---
    deacon_tmpdvb56d_b8fa1f86_c467676a-app-1	Up 4 seconds	f0372a834c09
    
    --- inspect f0372a834c09 ---
    status=running running=true exit=0 oom=false pid=2967310 err=
    

    A dispatch-only probe (.github/workflows/podman-flake-probe.yml) that asks the runner one question and returns a number. Phase A: the failing test alone, N times. Phase B: the whole lane, M times. The rates discriminate:

    • A fails, B fails → the shape reproduces on its own; iterate on phase A.
    • A passes, B fails → the suite context is load-bearing, and phase A is the control that proves it is not the test.
    • Both pass → that runner is not one that fails; re-dispatch before concluding anything.

    workflow_dispatch only fires for workflows on the default branch, so the probe has to land before it can be run.

    Next

  5. pofallon commented on Aug 29, 2026

    @pofallon
    ContributorAuthor

    The new capture fired on its first CI failure, and it retires a fourth hypothesis

    #728's failure-time dump landed on the very next Podman-lane run (job 99121571737). This is the first time anyone has seen the container's state at the moment of the failing exec rather than after teardown.

    --- inspect abb67e046fd9 ---
    status=running running=true exit=0 oom=false pid=41057 err=
    --- logs abb67e046fd9 (tail) ---
      (no output)
    

    The container was running, healthy, un-OOM-killed, with a live PID, at the moment podman exec failed to start conmon. That retires the last surviving variant of the container-state theory — not just "already exited", which was refuted earlier with a different error message, but also "exiting right now, in a window podman ps does not expose". It was not in a window. It was up.

    What the same dump shows that is new

    The dump lists every container the runtime can see, and the picture is not the one any local reproduction has had:

    deacon_tmpxatdnj_c2cdbc8d_1f306951-app-1   Up 3 seconds     abb67e046fd9   <- ours
    deacon_tmpdwzmsy_31c8432b_49d897f7-app-1   Up 4 seconds     b968379844a2
    deacon-d6d37754                            Up 3 seconds     d5cba866a342
    deacon-ef7bca29                            Up 2 seconds     010b18fb0c91
    deacon_tmpuombg8_af7f2c16_4bdad49c-app-1   Exited (0) 5 seconds ago
    deacon_tmpuombg8_af7f2c16_4bdad49c-db-1    Exited (0) 5 seconds ago
    deacon_tmpa1wrum_2633c22c_86bceb2f-app-1   Exited (0) 6 seconds ago
    deacon_project_99107214_de39c4b7-app-1     Exited (0) 10 seconds ago
    

    28 containers total, and seven of them were created or destroyed inside the same few seconds as the failing exec. Our own container was 3 seconds old. The timeline from the job log:

    Time Event
    15:38:04.135 env probe exec fails — conmon bytes ""
    15:38:05.344 second probe fails the same way
    15:38:05.465 postCreate exec fails, up exits 255
    15:38:07.54 dump captured; container still Up, running=true

    This is not the contention hypothesis that was refuted in #718. That one halved smoke-lite from 4 to 2, which bounds concurrency within one nextest group. The churn visible here comes from the whole 436-test suite across every group — Compose projects, single containers, testcontainers — all creating and tearing down against one rootless podman at once. Lowering smoke-lite never touched it, which is exactly why lowering it changed nothing.

    It also explains the local non-reproduction reported above better than the version gap does. Locally the test ran alone (20x) or as 4 copies of itself (24x); neither produces this. The three local full-suite runs do produce it and still passed — so if concurrent container churn is the mechanism, podman 4.9.3 tolerates a level that 5.8.4 does not, or the runner's 4 CPUs matter. Both remain open.

    Where that leaves the prediction

    The probe in #728 has two phases, and this evidence makes a falsifiable prediction about them:

    • Phase A (the test alone, N times) should pass.
    • Phase B (the whole lane, M times) should fail at roughly the observed rate.

    If that is what comes back, the question stops being "what is wrong with this test" and becomes "what does rootless podman 5.8.4 do to an exec session when N containers are being created around it" — which is answerable by bisecting the concurrency, not by staring at the one victim.

    If phase A fails too, the churn is a red herring and I am wrong about this.

    Status

    Still open, still unfixed, three refuted hypotheses now four. The exec's conmon dies after a successful fork, against a container that is verifiably running, on a daemon that is servicing several other container creations at that instant.

  6. pofallon commented on Aug 29, 2026

    @pofallon
    ContributorAuthor

    Second occurrence, and the neighbourhood is identical

    #728's Podman lane failed again on the immediate re-run (job 99122605950) — same test, same signature, 435 passed, 1 failed. Two independent captures now exist, and putting them side by side says more than either alone.

    Container state at the failing exec, both times:

    run 1:  status=running running=true exit=0 oom=false pid=41057 err=
    run 2:  status=running running=true exit=0 oom=false pid=42038 err=
    

    The co-scheduled neighbourhood is the same in both, down to the config hashes.

    run 1 run 2
    ours (…_1f306951-app-1) Up 3s Up 4s
    a sibling compose project …_49d897f7-app-1 Up 4s Up 5s
    a two-container project …_4bdad49c-app-1 + -db-1 Exited 5s ago Exited 8s ago
    …_86bceb2f-app-1 Exited 6s ago Exited 8s ago
    deacon_project_…_de39c4b7-app-1 Exited 10s ago Exited 10s ago
    bare single containers created within 1–6s 2 3

    Those hashes are config hashes, so they name the same tests each time. nextest schedules the suite in a stable order, so this is not coincidence: our test's exec lands in the same window of concurrent container churn on every run. What varies between a pass and a fail is only whether the timing inside that window lines up — which is exactly the shape of a ~50% flake that is nonetheless completely deterministic in its setup.

    Two things follow:

    1. It is not "some load somewhere". The neighbours are nameable. A sibling Compose project comes up roughly one second before ours in both failures, and a two-service project tore down a few seconds earlier. Those are the first concrete candidates for what to bisect.
    2. It is not the concurrency knob that was already tried. fix(ci): halve smoke-lite concurrency, which #715 made too high for podman #718 halved smoke-lite, which bounds parallelism inside one nextest group. Every neighbour listed above is in a different group. That is why lowering it changed nothing, and why lowering it further would also change nothing.

    Also visible, from the runner's own orphan-process reaping at job end:

    Terminate orphan process: pid (2461)  (podman pause)
    Terminate orphan process: pid (16794) (podman pause)
    Terminate orphan process: pid (16980) (podman pause)
    Terminate orphan process: pid (18104) (podman pause)
    

    Four rootless podman pause processes. The pause process is what holds the rootless user namespace open, and the normal steady state is one. Four is not obviously the cause of anything, but it is a fact about this lane that nobody has looked at, and "which user namespace is a given exec actually joining" is a reasonable question to have an answer for before ruling it out.

    Standing prediction, unchanged

    The probe in #728 should return phase A passing, phase B failing. Both captures are consistent with that and neither proves it.

  7. pofallon commented on Aug 29, 2026

    @pofallon
    ContributorAuthor

    The reference waits for a start event before it execs. deacon does not.

    Third failure in a row on #728, same fingerprint (…_1f306951-app-1 Up 3s, sibling …_49d897f7-app-1 Up 4s, the two-service …_4bdad49c project torn down 6s earlier, running=true oom=false). Three-for-three on the co-scheduling pattern.

    That made the timing question sharp enough to go and ask what the reference does between "compose up returned" and "run a command in the container". It does something deacon does not do at all.

    From the pinned oracle's own compiled source (@devcontainers/cli@0.87.0, dist/spec-node/devContainersSpecCLI.js), the compose path — function at index 1574585, jV:

    let p, D = new Promise((J,b) => p = b),
        {started: m} = await qQ(A, {[QG]: i, [uG]: B.service}, D, I.output, …);
    //  ^ subscription opens BEFORE compose up, filtered by project + service labels
    …
    let k = ["--project-name", i, ...l];
    k.push("up","-d");
    …
    A.isTTY ? await QQ({...e, output:h}, ...k) : await Go({...e, output:h}, ...k);
    …
    return await m, { containerId: await lG(A, i, B.service) };
    //     ^^^^^^^ blocks until the start event has been SEEN

    and qQ itself:

    async function qQ(A, e, t, i, r) {
      let n = await Od(A, {event: ["start"]});          // docker/podman events --event start
      return { started: new Promise((o, s) => { …
          (B.status || B.Status || B.Action) === "start" && await C9(A, B, e) && (n.terminate(), o())
      })};
    }

    C9 confirms the event's container carries the expected labels. The single-container path does the identical thing ({started:R} = await qQ(A, a9(r), k, …) … await R) around docker run.

    So the reference gates on the container runtime's own start event before it will touch the container.

    deacon's compose path gates on nothing of the kind. resolve_primary_container_id_with_retry (crates/deacon/src/commands/up/compose.rs:44) retries compose ps until an ID appears, and get_primary_container_id matches the service by name with no state check — a point already noted on this issue and set aside as "an asymmetry, not this bug". With the reference's behavior in hand it reads differently: it is not that deacon checks state loosely, it is that deacon has no readiness gate at all, where the reference has an event-driven one. grep -rn 'docker events' across deacon finds exactly one mention, in docker.rs, about the destroy path (#688), where polling was chosen deliberately as an equivalent guarantee. Nothing equivalent exists for start.

    Why this fits every observation

    • The container reads running=true in inspect while the exec fails — podman sets that state early; the start event is what says the runtime finished the job.
    • The exec fails ~1 second after compose up returns, and both captured failures show the container 3–4 seconds old.
    • It gets worse under concurrent container churn, which is exactly when the gap between "state says running" and "start actually completed" widens.
    • It does not reproduce on an idle machine, where that gap is too small to land in.
    • smoke-lite never affected it, because the churn comes from every nextest group at once.

    What I am and am not claiming

    Measured and certain: deacon diverges from the reference here. The reference opens an event subscription before starting the container, filters it by the container's labels, and blocks on it; deacon returns as soon as an ID exists. That is a parity gap on the up path regardless of what causes this flake, and there is no SPEC_STATUS.md row for it.

    Hypothesis, unproven: that this gap is why the exec races conmon. It is consistent with everything above and it is the first candidate that is independently a defect, but "the reference synchronizes more" is not the same fact as "this is the cause". Implementing the gate and watching the rate is weak evidence on its own when the base rate is roughly one in two — it would need several runs, on both a fixed and an unfixed build.

    Suggested next step

    Treat the parity gap as its own issue and fix it on parity grounds. If the flake goes away, that is a bonus and a strong hint; if it does not, deacon still stops skipping a synchronization step the reference performs, and this issue has lost its most plausible hypothesis honestly.

  8. pofallon commented on Aug 29, 2026

    @pofallon
    ContributorAuthor

    Probe result: the suite context is load-bearing, and the test is not the problem

    Run 33261826509 on 1a0670c, substrate printed by the probe itself: podman 5.8.4, overlay, netavark, cgroupfs, 4 CPUs — i.e. the exact stack the lane fails on, not the 4.9.3 one local reproduction had.

    PHASE A: 0 failed / 30 runs
    conmon signature seen in: 0 iterations
    
    PHASE B: 3 failed / 3 suite runs
    conmon signature seen in: 3 suite runs
    

    Phase A ran test_compose_override_command_lifecycle_runs alone, thirty times, on the failing substrate. Zero failures, zero conmon signatures.

    Phase B ran the whole lane three times on that same runner, in the same job, minutes apart. All three failed, all three with the conmon signature.

    That is the discriminator this issue has been missing. It rules out the version gap as a sufficient explanation — 5.8.4 runs this test thirty times without complaint — and it rules out the test itself. What fails is the test inside the suite; what passes is the same test on the same podman with nothing else running.

    Combined with the three failure captures from #728 — container running=true oom=false with a live pid, and an identical co-scheduled neighbourhood by config hash every time — the shape of the bug is now specific: an exec races something that only goes wrong while other containers are being created and destroyed around it.

    Consequence for the hypothesis list

    Every previously refuted hypothesis stays refuted. The live one is #729: podman sets a container's state to running early, the reference CLI waits for the runtime's own start event before it will touch the container, and deacon waits for nothing. Under churn the gap between those two moments is where this lands.

    Still a hypothesis. The probe does not test it — it establishes the precondition (suite-level concurrency is required) that makes it plausible, nothing more.

    A note on the probe itself

    Its per-failure detail block came out empty: grep -E '^\s+(FAIL|TRY|Summary)' does not match nextest's output because the ANSI colour codes sit before the leading whitespace. The counts and the conmon bytes signature grep are unaffected and are what the phases are for, but anyone re-dispatching it for detail should strip escapes first.

    Suggested reading of "3 of 3"

    Do not read it as a rate. The lane's observed rate across real PRs is 5 red of 9 (fail, pass, fail, pass on #725/#726; fail, fail, fail, pass on #728). Three consecutive phase-B failures in one job on one runner is consistent with that and with something about a warm runner making it likelier, and the probe was not designed to separate those.

  9. pofallon commented on Aug 29, 2026

    @pofallon
    ContributorAuthor

    The #729 experiment has a first result, and it goes against the hypothesis

    Run 33268702458, Podman job 99144406001, on a branch based on f09c5f2 — i.e. with the start-event gate from #731 in.

    436 tests run: 435 passed, 1 failed. The one failure is test_compose_override_command_lifecycle_runs, with the usual signature:

    Error: container create failed (no logs from conmon): conmon bytes "":
    readObjectStart: expect { or n, but found , error found in #0 byte of ...||...
    

    and the #728 capture again showing a live container at the moment of the failing exec:

    === runtime state at failure ===
    status=running running=true exit=0 oom=false pid=41177 err=
    

    Why this is evidence rather than just another occurrence

    The gate is designed to close exactly the window this failure sits in: the failing exec is the environment probe, which now runs strictly after up's compose path has blocked on the runtime's own start event.

    And the gate held. It degrades loudly by design — a timeout or an unavailable events stream logs at WARN. The captured output contains two WARN lines (both the env-probe failure itself), and no gate warning of either kind. Same process, same stream. So this is not a case of the gate silently not applying; it applied, and the flake happened anyway.

    That does not make podman-reports-running-early wrong as a description, but it does mean it is not sufficient to explain this failure. A fifth hypothesis is owed.

    What this does not establish

    One occurrence is not a rate. The lane's own historical rate is 5 red of 9, so nothing here says the gate changed the frequency in either direction — only that the failure survives a correctly-functioning gate at least once. The prior three hypotheses (daemon contention, container-not-running, podman/conmon version skew) remain refuted; this adds a fourth to that list rather than replacing it.

    Suggested next step

    The podman-flake-probe.yml rig already discriminates "the test alone" from "the test inside the suite", and its earlier run showed PHASE A: 0/30 versus PHASE B: 3/3 — the suite around it is load-bearing. With the start-event race now ruled out, the remaining candidates point at what the suite does to the runtime rather than at what the test does: exec/conmon fork pressure from co-scheduled projects, and the identical co-scheduled neighbourhood already documented (nextest's order is stable, so the setup is deterministic and only the timing varies).

    Worth fixing the known defect in the probe's reporting before the next run: its per-failure block greps '^\s+(FAIL|TRY|Summary)', which never matches because nextest's ANSI codes precede the leading whitespace.

  10. added a commit that references this issue on Aug 29, 2026
  11. pofallon commented on Aug 29, 2026

    @pofallon
    ContributorAuthor

    Second occurrence with the gate in — identical, so n=2

    Run 33269557759, job 99146932731, same branch, one commit later (a docs-only change).

    Identical on every axis to the occurrence above:

    run 33268702458 run 33269557759
    failing test test_compose_override_command_lifecycle_runs same
    signature conmon bytes "" same
    container at failure status=running running=true exit=0 oom=false pid=41177 … pid=43262
    gate warning present no no

    Two consecutive reds on the same branch, both with the start-event gate in place and demonstrably holding. Taken with the earlier probe result — PHASE A: 0 failed / 30 runs for the test alone versus PHASE B: 3 failed / 3 runs for the whole suite — the picture is consistent and narrow:

    • the test on its own does not fail, at any repetition count tried;
    • the test inside the suite fails often;
    • closing the start-event race did not change that.

    So whatever is left is a property of the suite around the test, not of the test and not of when the container is considered started. The co-scheduled neighbourhood is already known to be deterministic (nextest's order is stable; the same three sibling projects appear at the same relative offsets across failures), which means the remaining variable is timing under that fixed load — pointing at exec/conmon fork pressure rather than at anything deacon sequences.

    No claim here about the rate. Two reds in a row is entirely consistent with the historical 5-of-9 and says nothing new about frequency; what it adds is that the gate is now ruled out with n=2 rather than n=1.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions