Repository navigation
chore(ci): Podman lane intermittently fails an exec with conmon bytes "" — two hypotheses refuted #723
Description
Activity
The diagnostics fired, and the first thing they show is that the lane runs a different podman than anyone thought
From the failing Podman job on #725 (run 33214...):
Setting up podman (4.9.3+ds1-1ubuntu0.2) ... <- what apt installed, into /usr/bin Client: Podman Engine Version: 5.8.4 <- what actually ranThe
ubuntu-24.04runner image already ships podman 5.8.4 under/usr/local/bin, which precedes/usr/binon PATH. Sosudo apt-get install -y podmanin the setup step does not provision the podman the lane uses — it installs an unused second copy, and pulls in a second conmon alongside it:Path Version command -v conmon/usr/bin/conmon2.1.10 (apt) what podman resolves /usr/local/lib/podman/conmon2.2.1 (bundled) podman infoconfirms the rest of the 5.x stack is in play:buildahVersion 1.43.2,crun 1.28from/usr/local/bin,netavark,pasta.Why this matters for this issue
It explains "never reproduces locally." Every local reproduction attempt used podman 4.9.3 — the version apt reports and the version in the dev container. CI fails on 5.8.4. Different major version, different conmon, different network backend (netavark/pasta vs slirp4netns). The two refuted hypotheses were tested against the wrong substrate.
I am deliberately not claiming the version skew is the cause. Two hypotheses have already been wrong here, and "there are two installs" is not the same fact as "that is why conmon fails". What it does is make a real reproduction possible for the first time.
Other diagnostics, for the record
- Memory was not tight: 14.5 GB available of 16 GB, load 2.48 on 4 CPUs.
pids.max: n/a,max user processes 63795,open files 65536— no ceiling in sight.- No OOM in
dmesg. - Leftover containers were accumulating (14 listed, 3 still
Up), which is worth a separate look but is not obviously this.
Next
- Reproduce against podman 5.8.4, not 4.9.3.
- Decide what the setup step should do: drop the apt install (the runner already has 5.8.4) or pin deliberately to a chosen version. Either way it should stop claiming to install the podman under test when it does not.
The lane now runs a coherent podman stack, and the flake survived it
#726 removes the
apt-get install -y podmanthat was feeding a 4.9.3 configuration to the 5.8.4 engine the lane actually runs. The setup step passed, so the correction is validated on its own terms:=== podman as provisioned by the runner image === /usr/local/bin/podman Version: 5.8.4 === /etc/containers === containers.conf policy.json registries.conf seccomp.json storage.conf <- all image-provided === rootless helpers === pasta /usr/local/bin/pasta slirp4netns <not on PATH> crun /usr/local/bin/crun runc /usr/local/bin/runcNothing was depending on the apt package:
/etc/containersis shipped complete by the image (dated Jun 26 / Aug 23, i.e. from the image build, not from apt), andpasta— podman 5's defaultrootlessNetworkCmd— is already at/usr/local/bin. The apt install was purely overwriting image config with an older release's.The diagnostics are also no longer self-contradictory:
=== conmon (the one PODMAN resolves, not the one on $PATH) === podman resolves: /usr/local/lib/podman/conmon conmon version 2.2.1 none on PATH (expected: apt's podman is no longer installed)What that rules out
The version/config mismatch is not the cause of this flake. Same test, same error, on a stack with no mixed versions:
FAIL [ 11.605s] (149/436) deacon::smoke_compose_override_command test_compose_override_command_lifecycle_runs Summary [132.961s] 436 tests run: 435 passed, 1 failed, 3 skippedContainer command failed with exit code 255: Error: container create failed (no logs from conmon): conmon bytes "": readObjectStart: expect { or n, but found ... Error: Lifecycle error: Lifecycle command failed (source: config) in phase postCreate with exit code 255: touch /tmp/deacon-lifecycle-markerSo the hypothesis list now stands at three refuted: daemon contention (#718, reverted in #724), container-not-running (an exec into an exited container gives
can only create exec sessions on running containers, a different error), and now conmon/podman version skew.What is still standing
The failure is podman failing to spin up the conmon for an exec session —
postCreate— against a compose-created container, i.e. one created through podman's API socket by docker-compose rather than bypodman run. Every observed instance has been that same test. Worth probing next, in rough order of cheapness:- whether the exec is racing the container's own startup in a way that
podman psstate does not expose (distinct from the refuted not-running hypothesis) - whether an API-socket-created container's exec path differs from a CLI-created one
- rootless resource ceilings in the user namespace at that moment (the earlier diagnostics ruled out host memory, pids and fd ceilings, and OOM)
Observed rate on that one test across recent runs: roughly 6 fail / 3 pass. It is currently the sole blocker on both #725 and #726, neither of which touches the code path it exercises.
- whether the exec is racing the container's own startup in a way that
- added a commit that references this issue
on Aug 29, 2026 Reopened — closed by accident, not by resolution.
The merge commit for #726 contained the sentence
This does NOT fix <this issue's number>, written specifically to avoid overclaiming. GitHub's closing keywords fire on the verb next to the number regardless of the surrounding prose: negated, quoted or merely described, the pattern fires. This repository's HANDOFF.md documents that trap, and I walked into it anyway.Nothing is fixed. State is unchanged from the previous comment: three hypotheses refuted (daemon contention, container-not-running, podman/conmon version skew), the failure still intermittent at roughly half of runs, still confined to
test_compose_override_command_lifecycle_runs.Observed again today across four Podman-lane runs on #725/#726: fail, pass, fail, pass.
A measurement pass, not a fourth hypothesis
Nothing here changes the state of the bug. What it does is put numbers under "never reproduces locally", correct two things the error message implies but does not mean, and add the instrumentation that would make the next occurrence self-describing. PR #728.
Local non-reproduction, quantified
Substrate: podman 4.9.3, conmon 2.1.10, runc 1.5.1, netavark, cgroups v2 +
cgroupfs,vfs, in the dev container.Shape Runs Failures test_compose_override_command_lifecycle_runsalone, serial20 0 Same test, 4 concurrent copies (6 rounds) 24 0 The whole Podman lane, mvp-integration -E 'not binary(=parity_docker)', 436 tests3 0 The earlier note said "passes 5/5 serially" — that was n=5. This is n=44 on the test plus three whole suites, including the compose-through-the-API-socket path the issue singled out. Every run went through
podman composeagainst the API socket, so "compose-created container + exec" alone is not sufficient on 4.9.3.Two corrections to what the error implies
Both read out of the reference's own source at
v5.8.4.1.
container create failed (no logs from conmon)is a shared message, not a claim about container creation. It is produced insidereadConmonPipeData(libpod/oci_conmon_common.go:1443), and both the container-create path and the exec path call it. So the wording is not evidence against the exec framing — the exec framing is right.2. conmon is spawned successfully and then dies silently. In
ExecContainer(libpod/oci_conmon_exec_common.go:78-82),execCmd.Wait()runs beforereadConmonPipeData, and a failure to spawn conmon returnscannot run conmon: …. We get the later error, so:- conmon's parent process exited 0 — this is not a fork/exec failure;
- the sync pipe then returned EOF with zero bytes (
conmon bytes ""); - conmon wrote nothing to stderr either, since podman's stderr carried only the one
Error:line.
Whatever kills it happens after fork and before it writes sync data.
A trap for whoever instruments this next
Do not run the lane with
podman --log-level=debug. It is the instinctive move and it destroys the best evidence:var ociLog string if logrus.GetLevel() != logrus.DebugLevel && r.supportsJSON { ociLog = c.execOCILog(sessionID) }
—
libpod/oci_conmon_exec_common.go:50-53, v5.8.4ociLogis the filereadConmonPipeDatareads to replace the generic message with the OCI runtime's real error, and podman only writes it when the level is not debug. That we saw the generic message means the log existed and was empty or unparseable — i.e. crun never wrote an error either, which is consistent with conmon dying before it ever ran crun.Refuted cheaply
The test harness points
TMPDIRat a per-test-process temp dir (support::deacon_command), which everypodmanchild inherits, and that looked like a candidate for divergent runtime paths between the API service and the CLI. Measured:TMPDIR=… podman inforeports the sameRunRootandGraphRoot. Not it.What #728 adds
Failure-time state capture. The three compose-override tests now call
support::runtime_state_dump(workspace)on the up-failure path beforedeacon_down. The job-levelif: failure()diagnostics can never answer the question that matters — was the container running when the exec was attempted — because the test removes it first. Watched to fail withpostCreateCommandset toexit 3:--- every container the runtime can see --- deacon_tmpdvb56d_b8fa1f86_c467676a-app-1 Up 4 seconds f0372a834c09 --- inspect f0372a834c09 --- status=running running=true exit=0 oom=false pid=2967310 err=A dispatch-only probe (
.github/workflows/podman-flake-probe.yml) that asks the runner one question and returns a number. Phase A: the failing test alone, N times. Phase B: the whole lane, M times. The rates discriminate:- A fails, B fails → the shape reproduces on its own; iterate on phase A.
- A passes, B fails → the suite context is load-bearing, and phase A is the control that proves it is not the test.
- Both pass → that runner is not one that fails; re-dispatch before concluding anything.
workflow_dispatchonly fires for workflows on the default branch, so the probe has to land before it can be run.Next
- Land chore(ci): make the Podman conmon flake carry its own evidence #728, dispatch the probe, and record which of the three outcomes came back.
- Whichever it is, the next question is narrower than the one this issue currently asks.
The new capture fired on its first CI failure, and it retires a fourth hypothesis
#728's failure-time dump landed on the very next Podman-lane run (job 99121571737). This is the first time anyone has seen the container's state at the moment of the failing exec rather than after teardown.
--- inspect abb67e046fd9 --- status=running running=true exit=0 oom=false pid=41057 err= --- logs abb67e046fd9 (tail) --- (no output)The container was running, healthy, un-OOM-killed, with a live PID, at the moment
podman execfailed to start conmon. That retires the last surviving variant of the container-state theory — not just "already exited", which was refuted earlier with a different error message, but also "exiting right now, in a windowpodman psdoes not expose". It was not in a window. It was up.What the same dump shows that is new
The dump lists every container the runtime can see, and the picture is not the one any local reproduction has had:
deacon_tmpxatdnj_c2cdbc8d_1f306951-app-1 Up 3 seconds abb67e046fd9 <- ours deacon_tmpdwzmsy_31c8432b_49d897f7-app-1 Up 4 seconds b968379844a2 deacon-d6d37754 Up 3 seconds d5cba866a342 deacon-ef7bca29 Up 2 seconds 010b18fb0c91 deacon_tmpuombg8_af7f2c16_4bdad49c-app-1 Exited (0) 5 seconds ago deacon_tmpuombg8_af7f2c16_4bdad49c-db-1 Exited (0) 5 seconds ago deacon_tmpa1wrum_2633c22c_86bceb2f-app-1 Exited (0) 6 seconds ago deacon_project_99107214_de39c4b7-app-1 Exited (0) 10 seconds ago28 containers total, and seven of them were created or destroyed inside the same few seconds as the failing exec. Our own container was 3 seconds old. The timeline from the job log:
Time Event 15:38:04.135 env probe exec fails — conmon bytes ""15:38:05.344 second probe fails the same way 15:38:05.465 postCreate exec fails, upexits 25515:38:07.54 dump captured; container still Up,running=trueThis is not the contention hypothesis that was refuted in #718. That one halved
smoke-litefrom 4 to 2, which bounds concurrency within one nextest group. The churn visible here comes from the whole 436-test suite across every group — Compose projects, single containers, testcontainers — all creating and tearing down against one rootless podman at once. Loweringsmoke-litenever touched it, which is exactly why lowering it changed nothing.It also explains the local non-reproduction reported above better than the version gap does. Locally the test ran alone (20x) or as 4 copies of itself (24x); neither produces this. The three local full-suite runs do produce it and still passed — so if concurrent container churn is the mechanism, podman 4.9.3 tolerates a level that 5.8.4 does not, or the runner's 4 CPUs matter. Both remain open.
Where that leaves the prediction
The probe in #728 has two phases, and this evidence makes a falsifiable prediction about them:
- Phase A (the test alone, N times) should pass.
- Phase B (the whole lane, M times) should fail at roughly the observed rate.
If that is what comes back, the question stops being "what is wrong with this test" and becomes "what does rootless podman 5.8.4 do to an exec session when N containers are being created around it" — which is answerable by bisecting the concurrency, not by staring at the one victim.
If phase A fails too, the churn is a red herring and I am wrong about this.
Status
Still open, still unfixed, three refuted hypotheses now four. The exec's conmon dies after a successful fork, against a container that is verifiably running, on a daemon that is servicing several other container creations at that instant.
Second occurrence, and the neighbourhood is identical
#728's Podman lane failed again on the immediate re-run (job 99122605950) — same test, same signature,
435 passed, 1 failed. Two independent captures now exist, and putting them side by side says more than either alone.Container state at the failing exec, both times:
run 1: status=running running=true exit=0 oom=false pid=41057 err= run 2: status=running running=true exit=0 oom=false pid=42038 err=The co-scheduled neighbourhood is the same in both, down to the config hashes.
run 1 run 2 ours ( …_1f306951-app-1)Up 3s Up 4s a sibling compose project …_49d897f7-app-1Up 4s Up 5s a two-container project …_4bdad49c-app-1+-db-1Exited 5s ago Exited 8s ago …_86bceb2f-app-1Exited 6s ago Exited 8s ago deacon_project_…_de39c4b7-app-1Exited 10s ago Exited 10s ago bare single containers created within 1–6s 2 3 Those hashes are config hashes, so they name the same tests each time. nextest schedules the suite in a stable order, so this is not coincidence: our test's exec lands in the same window of concurrent container churn on every run. What varies between a pass and a fail is only whether the timing inside that window lines up — which is exactly the shape of a ~50% flake that is nonetheless completely deterministic in its setup.
Two things follow:
- It is not "some load somewhere". The neighbours are nameable. A sibling Compose project comes up roughly one second before ours in both failures, and a two-service project tore down a few seconds earlier. Those are the first concrete candidates for what to bisect.
- It is not the concurrency knob that was already tried. fix(ci): halve smoke-lite concurrency, which #715 made too high for podman #718 halved
smoke-lite, which bounds parallelism inside one nextest group. Every neighbour listed above is in a different group. That is why lowering it changed nothing, and why lowering it further would also change nothing.
Also visible, from the runner's own orphan-process reaping at job end:
Terminate orphan process: pid (2461) (podman pause) Terminate orphan process: pid (16794) (podman pause) Terminate orphan process: pid (16980) (podman pause) Terminate orphan process: pid (18104) (podman pause)Four rootless
podman pauseprocesses. The pause process is what holds the rootless user namespace open, and the normal steady state is one. Four is not obviously the cause of anything, but it is a fact about this lane that nobody has looked at, and "which user namespace is a given exec actually joining" is a reasonable question to have an answer for before ruling it out.Standing prediction, unchanged
The probe in #728 should return phase A passing, phase B failing. Both captures are consistent with that and neither proves it.
The reference waits for a
startevent before it execs. deacon does not.Third failure in a row on #728, same fingerprint (
…_1f306951-app-1Up 3s, sibling…_49d897f7-app-1Up 4s, the two-service…_4bdad49cproject torn down 6s earlier,running=true oom=false). Three-for-three on the co-scheduling pattern.That made the timing question sharp enough to go and ask what the reference does between "compose up returned" and "run a command in the container". It does something deacon does not do at all.
From the pinned oracle's own compiled source (
@devcontainers/cli@0.87.0,dist/spec-node/devContainersSpecCLI.js), the compose path — function at index 1574585,jV:let p, D = new Promise((J,b) => p = b), {started: m} = await qQ(A, {[QG]: i, [uG]: B.service}, D, I.output, …); // ^ subscription opens BEFORE compose up, filtered by project + service labels … let k = ["--project-name", i, ...l]; k.push("up","-d"); … A.isTTY ? await QQ({...e, output:h}, ...k) : await Go({...e, output:h}, ...k); … return await m, { containerId: await lG(A, i, B.service) }; // ^^^^^^^ blocks until the start event has been SEEN
and
qQitself:async function qQ(A, e, t, i, r) { let n = await Od(A, {event: ["start"]}); // docker/podman events --event start return { started: new Promise((o, s) => { … (B.status || B.Status || B.Action) === "start" && await C9(A, B, e) && (n.terminate(), o()) })}; }
C9confirms the event's container carries the expected labels. The single-container path does the identical thing ({started:R} = await qQ(A, a9(r), k, …)…await R) arounddocker run.So the reference gates on the container runtime's own
startevent before it will touch the container.deacon's compose path gates on nothing of the kind.
resolve_primary_container_id_with_retry(crates/deacon/src/commands/up/compose.rs:44) retriescompose psuntil an ID appears, andget_primary_container_idmatches the service by name with no state check — a point already noted on this issue and set aside as "an asymmetry, not this bug". With the reference's behavior in hand it reads differently: it is not that deacon checks state loosely, it is that deacon has no readiness gate at all, where the reference has an event-driven one.grep -rn 'docker events'across deacon finds exactly one mention, indocker.rs, about the destroy path (#688), where polling was chosen deliberately as an equivalent guarantee. Nothing equivalent exists for start.Why this fits every observation
- The container reads
running=trueininspectwhile the exec fails — podman sets that state early; thestartevent is what says the runtime finished the job. - The exec fails ~1 second after compose up returns, and both captured failures show the container 3–4 seconds old.
- It gets worse under concurrent container churn, which is exactly when the gap between "state says running" and "start actually completed" widens.
- It does not reproduce on an idle machine, where that gap is too small to land in.
smoke-litenever affected it, because the churn comes from every nextest group at once.
What I am and am not claiming
Measured and certain: deacon diverges from the reference here. The reference opens an event subscription before starting the container, filters it by the container's labels, and blocks on it; deacon returns as soon as an ID exists. That is a parity gap on the
uppath regardless of what causes this flake, and there is noSPEC_STATUS.mdrow for it.Hypothesis, unproven: that this gap is why the exec races conmon. It is consistent with everything above and it is the first candidate that is independently a defect, but "the reference synchronizes more" is not the same fact as "this is the cause". Implementing the gate and watching the rate is weak evidence on its own when the base rate is roughly one in two — it would need several runs, on both a fixed and an unfixed build.
Suggested next step
Treat the parity gap as its own issue and fix it on parity grounds. If the flake goes away, that is a bonus and a strong hint; if it does not, deacon still stops skipping a synchronization step the reference performs, and this issue has lost its most plausible hypothesis honestly.
- The container reads
- added a commit that references this issue
on Aug 29, 2026 Probe result: the suite context is load-bearing, and the test is not the problem
Run 33261826509 on
1a0670c, substrate printed by the probe itself: podman 5.8.4,overlay,netavark,cgroupfs, 4 CPUs — i.e. the exact stack the lane fails on, not the 4.9.3 one local reproduction had.PHASE A: 0 failed / 30 runs conmon signature seen in: 0 iterations PHASE B: 3 failed / 3 suite runs conmon signature seen in: 3 suite runsPhase A ran
test_compose_override_command_lifecycle_runsalone, thirty times, on the failing substrate. Zero failures, zero conmon signatures.Phase B ran the whole lane three times on that same runner, in the same job, minutes apart. All three failed, all three with the conmon signature.
That is the discriminator this issue has been missing. It rules out the version gap as a sufficient explanation — 5.8.4 runs this test thirty times without complaint — and it rules out the test itself. What fails is the test inside the suite; what passes is the same test on the same podman with nothing else running.
Combined with the three failure captures from #728 — container
running=true oom=falsewith a live pid, and an identical co-scheduled neighbourhood by config hash every time — the shape of the bug is now specific: an exec races something that only goes wrong while other containers are being created and destroyed around it.Consequence for the hypothesis list
Every previously refuted hypothesis stays refuted. The live one is #729: podman sets a container's state to
runningearly, the reference CLI waits for the runtime's ownstartevent before it will touch the container, and deacon waits for nothing. Under churn the gap between those two moments is where this lands.Still a hypothesis. The probe does not test it — it establishes the precondition (suite-level concurrency is required) that makes it plausible, nothing more.
A note on the probe itself
Its per-failure detail block came out empty:
grep -E '^\s+(FAIL|TRY|Summary)'does not match nextest's output because the ANSI colour codes sit before the leading whitespace. The counts and theconmon bytessignature grep are unaffected and are what the phases are for, but anyone re-dispatching it for detail should strip escapes first.Suggested reading of "3 of 3"
Do not read it as a rate. The lane's observed rate across real PRs is 5 red of 9 (
fail, pass, fail, passon #725/#726;fail, fail, fail, passon #728). Three consecutive phase-B failures in one job on one runner is consistent with that and with something about a warm runner making it likelier, and the probe was not designed to separate those.The #729 experiment has a first result, and it goes against the hypothesis
Run 33268702458, Podman job
99144406001, on a branch based onf09c5f2— i.e. with the start-event gate from #731 in.436 tests run: 435 passed, 1 failed. The one failure istest_compose_override_command_lifecycle_runs, with the usual signature:Error: container create failed (no logs from conmon): conmon bytes "": readObjectStart: expect { or n, but found , error found in #0 byte of ...||...and the #728 capture again showing a live container at the moment of the failing exec:
=== runtime state at failure === status=running running=true exit=0 oom=false pid=41177 err=Why this is evidence rather than just another occurrence
The gate is designed to close exactly the window this failure sits in: the failing exec is the environment probe, which now runs strictly after
up's compose path has blocked on the runtime's ownstartevent.And the gate held. It degrades loudly by design — a timeout or an unavailable events stream logs at WARN. The captured output contains two WARN lines (both the env-probe failure itself), and no gate warning of either kind. Same process, same stream. So this is not a case of the gate silently not applying; it applied, and the flake happened anyway.
That does not make podman-reports-running-early wrong as a description, but it does mean it is not sufficient to explain this failure. A fifth hypothesis is owed.
What this does not establish
One occurrence is not a rate. The lane's own historical rate is 5 red of 9, so nothing here says the gate changed the frequency in either direction — only that the failure survives a correctly-functioning gate at least once. The prior three hypotheses (daemon contention, container-not-running, podman/conmon version skew) remain refuted; this adds a fourth to that list rather than replacing it.
Suggested next step
The
podman-flake-probe.ymlrig already discriminates "the test alone" from "the test inside the suite", and its earlier run showedPHASE A: 0/30versusPHASE B: 3/3— the suite around it is load-bearing. With the start-event race now ruled out, the remaining candidates point at what the suite does to the runtime rather than at what the test does: exec/conmon fork pressure from co-scheduled projects, and the identical co-scheduled neighbourhood already documented (nextest's order is stable, so the setup is deterministic and only the timing varies).Worth fixing the known defect in the probe's reporting before the next run: its per-failure block greps
'^\s+(FAIL|TRY|Summary)', which never matches because nextest's ANSI codes precede the leading whitespace.- added a commit that references this issue
on Aug 29, 2026 Second occurrence with the gate in — identical, so n=2
Run 33269557759, job
99146932731, same branch, one commit later (a docs-only change).Identical on every axis to the occurrence above:
run 33268702458 run 33269557759 failing test test_compose_override_command_lifecycle_runssame signature conmon bytes ""same container at failure status=running running=true exit=0 oom=false pid=41177… pid=43262gate warning present no no Two consecutive reds on the same branch, both with the start-event gate in place and demonstrably holding. Taken with the earlier probe result —
PHASE A: 0 failed / 30 runsfor the test alone versusPHASE B: 3 failed / 3 runsfor the whole suite — the picture is consistent and narrow:- the test on its own does not fail, at any repetition count tried;
- the test inside the suite fails often;
- closing the start-event race did not change that.
So whatever is left is a property of the suite around the test, not of the test and not of when the container is considered started. The co-scheduled neighbourhood is already known to be deterministic (nextest's order is stable; the same three sibling projects appear at the same relative offsets across failures), which means the remaining variable is timing under that fixed load — pointing at exec/conmon fork pressure rather than at anything deacon sequences.
No claim here about the rate. Two reds in a row is entirely consistent with the historical 5-of-9 and says nothing new about frequency; what it adds is that the gate is now ruled out with n=2 rather than n=1.
- added a commit that references this issue
on Aug 30, 2026
Symptom
smoke_compose_override_command::test_compose_override_command_lifecycle_runsfails intermittently on the Podman lane:podman is failing to start an exec session's conmon. It surfaces as a bare
deacon up failedwithpostCreateCommand … exit code 255, which reads like a lifecycle defect and is not one. The env probe fails the same way ~0.6s earlier, so two consecutivepodman execcalls failed.History
docker), fixedconmon→ passed on re-runconmonconmon— after the concurrency change belowNever reproduced locally: the test passes 5/5 serially and the full suite passes 433/433 against real rootless podman 4.9.3 in the dev container.
Refuted: contention
smoke-litewas halved from 4 to 2 (#718) on the theory that concurrent rootless-podman compose projects were starving conmon — plausible, since #715 moved those projects from docker to podman for the first time. The failure recurred unchanged at 2. Reverted; it cost ~14% on every smoke lane and bought nothing.Worth noting for whoever picks this up: at
max-threads = 2the pair running concurrently was exactly the two sibling tests in that binary, one of which passed. So "two concurrent podman compose projects" is not sufficient to cause it either.Refuted: container not running
get_primary_container_id(crates/core/src/compose.rs:1723) matches a service by name with no state check, while its siblinghandle_port_eventsfilters ons.state == "running". A not-yet-running or already-exited container looked like a strong candidate.Measured against rootless podman 4.9.3:
podman compose up -dreturns with the container alreadyrunning; pollingcompose psimmediately never showed an intermediate state.can only create exec sessions on running containers: container state improper— a clean, different error.So this is not it. (The missing state check is still an asymmetry worth tidying, but it is not this bug and should not be "fixed" as though it were.)
What is known
readObjectStart: expect { or n, but found ^@means podman read NUL from conmon's sync pipe — conmon died or wrote nothing.Next step is evidence, not a third hypothesis
A failure-only diagnostics step is added to the Podman job: memory, load,
ulimit -a,pids.max, conmon presence and version,podman info, container states, anddmesgtail. Everything is|| trueandif: failure(), so it reports and never masks the real failure.smoke-litealone, which has been tried.Deliberately not doing: quarantining the test, adding a retry, or reducing concurrency further. Each would hide the signal, and two of them would hide it permanently.