Conversation
…ximizes mega-evm branch coverage Adds an offline tool that answers "which mainnet blocks, replayed, exercise every mega-evm branch we have ever seen taken?" — the fixture set a stateless validator wants for regression coverage, without replaying 21M blocks each time. `backfill` replays blocks under LLVM branch instrumentation: resident worker subprocesses reset counters, replay one block, and capture a profraw; a judge dedups the resulting per-block bitmaps into "patterns" in a redb store, keeping the lightest block of each pattern as its representative. `set-cover` runs a greedy cover with antichain pruning and redundancy elimination over those patterns, `report` renders an llvm-cov summary for the selected set, `inspect` prints store statistics, and `merge` folds per-machine shard stores from a distributed scan into one (remapping each shard's private dense indices through machine-stable counter ids). Counter ids belong to one instrumented build, so `binary_id` (mega-evm rev + toolchain) namespaces every store and the write paths refuse a mismatch. That would strand a scan at each mega-evm bump, since a full sweep costs weeks — so `inspect --dump-pool` exports the antichain's representatives as a block list and `backfill --blocks-file` replays one, carrying the block numbers (which survive) rather than the bitmaps (which do not). `inspect` deliberately skips the binary-id check, so a pool can still be extracted long after the bump. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude review status
🛠️ Review did not finish Attempted This round did not publish: MODEL_ACTION_FAILED in phase review_retry. Anything listed below is from the last round that did. Re-run the workflow or push a new commit to try again. |
|
Label check: this PR currently has no labels applied. Given the title ( Suggest adding |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 76c04a183c
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
The workspace pins clap with `default-features = false` and only
`derive`/`env`/`std`, which drops clap's `help`, `usage` and `error-context`
features. `--help` is one of those features, not something the derive
provides, so every binary answered `coverage-replayer --help`,
`stateless-validator --help` and `debug-trace-server --help` with
error: unexpected argument found
— and the error did not even name the argument, because without
`error-context` and `usage` clap's messages carry no offender and no usage
line. Anyone deploying these binaries had no way to discover a flag from the
binary itself.
Opting the three features back in restores `--help` on the root and on every
subcommand, and makes an unknown flag report itself with a usage line. Two
tests pin the behaviour by error kind (`DisplayHelp` / `UnknownArgument` plus
the rendered text), so trimming the features again fails the suite instead of
silently removing `--help`; both were verified to fail with the features
removed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: fda8347481
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
clippy::items_after_test_module — the help regression tests landed above `fn main`. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… split Two costs that only show at full-history scale (21M blocks, 1.5M patterns), both measured on the consolidated mainnet scan: `merge` wrote ~35 GB to produce a 3.6 GB store. `write_table` drained a `HashMap` straight into 100k-row transactions, i.e. random-order insertion: every batch dirtied pages all over the B-tree, so each commit rewrote a slice of the whole tree. Sorting the rows by key first makes every batch land on the right edge, and a commit costs what the batch holds. `select_cover`'s antichain prune tested each pattern against every earlier one — quadratic in the pattern count, ~10^12 iterations for 1.5M patterns, and it runs inside `inspect` as well as `set-cover`. `split_antichain` scans only kept patterns (the dominated 97% never dominate anything a kept pattern does not) and looks candidates up through an inverted index from counter to the kept patterns containing it: a superset must contain the candidate's rarest counter, so the shortest posting list bounds the search, and a counter no kept pattern has proves the candidate maximal outright. The quadratic scan stays in the tests as the oracle for a differential test over randomized stores shaped like real ones (hubs, derived subsets, equal-bitmap twins, an empty pattern); the test was verified to fail on a mutant that checks only the first posting-list entry. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 4b1c3a9cc7
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
`inspect` always ran the real set-cover algorithm to print the antichain count and the selection preview. With the indexed antichain split that is no longer a twenty-minute step, but on a full-history store it is still about half of the command — the rest is a single pass over the tables. `--no-cover-preview` skips it. The preview stays on by default: it is the only way to see "N blocks cover X/Y" without running `set-cover`, which deletes dominated profiles, and on anything short of a full-history store it is effectively free. The flag conflicts with `--dump-pool`, since the pool is made of the antichain that pass computes — rejected by clap up front rather than producing a run that silently writes nothing. Also covers `inspect::run` end to end over a real redb store, which the direct `write_pool` tests did not reach. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 84604b1641
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
…eview
The tool took "which physical counters are non-zero" as a block's coverage
bitmap, on the assumption that with `-Z coverage-options=branch` both arms of
a branch are plain counters. They are not: rustc minimizes physical counters,
so an `if`/`else` gets two (entry, then-arm) and the else-arm exists only as
the expression `entry - then`. Reproduced on nightly-2026-02-03 — a then-only
run reads `[1,1]`, an else-only run `[1,0]` — and no `-Z coverage-options`
value changes it. Over physical counters the else-only block is `{entry}`, a
strict subset of `{entry, then}`: the judge never archived its profile,
set-cover pruned it as dominated, and the cover lost a branch arm the scan had
covered. On the consolidated v1.6.1 scan the 94-block cover reported Branches
59.35% where the union of the archived profiles reached at least 60.02%.
A counter is now an evaluated item: every block's profile goes through
`llvm-cov export --format=text --skip-functions`, scoped to the source dirs,
and the items are its region entries and branch arms with a non-zero count,
keyed by source span and OR-ed across instantiations — the arithmetic `report`
runs. ~0.6 s per block. The scope replaces the symbol-substring filter
(`--source-dir`, default the mega-evm checkout the binary was built against),
which also stops dependency generics instantiated with a mega-evm type from
entering the universe; an unscoped export of this binary segfaults llvm-cov,
and a scope that matches nothing fails the block instead of recording an empty
bitmap. Stores carry a universe stamp; a legacy store is recognized by its
symbol-filter key and labelled `physical-counters/v0`, so it is never mixed
with current ids yet still inspects and merges with its own kind. The
provenance fields of `CounterInfo` are renamed, not re-encoded: schema v1
stores still decode.
Validated on 296 blocks replayed under mega-evm v1.7.0: `report` over the 14
selected blocks is identical, file by file, to `report` over the union of all
174 archived profiles.
Review items (chatgpt-codex-connector on #222), each verified against the code:
- report: refuse a manifest from another `binary_id` — llvm-cov does not fail
on a mismatch, it drops the functions and reports them uncovered.
- backfill: fail unless every selected block was judged; a panicked fetch task
or a dead manager dropped its block and the run still exited 0.
- backfill: load the chain spec in the dispatcher first. A worker that cannot
start looks like one that crashed mid-block, and blocks are retried forever,
so a mistyped `--genesis-file` wedged the run in a respawn loop.
- set-cover: load patterns only, point-look-up the selected hashes; deleting
dominated patterns' profiles is opt-in (`--prune-profiles`) and happens
after the manifest is written.
- inspect: fold the BLOCKS table in a streaming pass instead of materializing
it (7.2 GB RSS on a 21M-row store); siblings by a second bounded pass.
- merge: mark the output incomplete until the last batch, and refuse such a
store on every open; fold a height two shards share when the replays agree
(counting its pattern hit once), stop when they disagree, and let a clean
record win over a quarantined one.
- docs: the cover is greedy and redundancy-eliminated, not minimum-cardinality;
scans are for final blocks; the ValidatorDB caveat is scoped to the
validator.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 6a109fc1d2
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
Replay is fetch-bound: on the pool re-sweep the worker pool could take roughly twenty times what eight concurrent fetches delivered, because a deep-history block costs far longer to download than to execute. The spool backlog stays bounded whatever the value — a full dispatch queue blocks the fetch loop — so the higher default costs nothing but idle workers. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The checked `get_block` recovers the signer of each transaction. On mainnet's stress-test blocks — tens of thousands of transactions each — that, not the download, was what the fetch stage spent its time on: a profile of the dispatcher showed over 90% of its CPU in secp256k1 arithmetic, with the concurrent fetches contending for the same coverage counters, and the sweep advanced by one block in nine minutes. Fetch unchecked, as the trace server does. Nothing is lost: a wrong sender or transaction set cannot reproduce the header's gas, receipts root and logs bloom, which the worker compares after replaying, and a divergence stops the run. The header hash is still checked here — it is one keccak, and it is what the manifest publishes and the witness is addressed by. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…trument only what is measured Scope. Coverage was scoped to the mega-evm sources alone. But mega-evm shapes execution *through* revm — its host and handler are type arguments of revm's generic interpreter and handler — so which EVM paths mainnet exercises is a fact about revm's source as much as mega-evm's. (The first, symbol-substring universe had included revm by accident: a mangled name carries its generic arguments and its instantiating crate, so 37% of that universe was dependency code — revm's interpreter, but also k256, generic-array and typenum.) The default scope is now the mega-evm checkout plus the crates listed in `measured-crates.txt`: revm-interpreter, revm-handler, revm-context, revm-context-interface and op-revm. build.rs resolves their locked versions into the default `--source-dir`s, and the same file drives the rustc wrapper, so what is instrumented and what is measured cannot drift. Item ids now start with the source dir's own name — two scoped crates both have a `src/lib.rs` — and the universe stamp moves to v2 so ids from the two derivations never share a store. Instrumentation. The build line used RUSTFLAGS, which instruments every crate. Every basic block of an instrumented crate bumps a process-global counter; in tight loops that dwarfs the work itself (k256 field arithmetic, per-byte bincode encoding), and threads running the same code contend for the same counter cache lines. `cov-rustc-wrapper.sh` instruments only the measured code and this workspace — the workspace because most of the measured code is generic and a generic function is compiled, counters included, in the crate that instantiates it; nothing outside the workspace depends on mega-evm, so nothing else can. Host artifacts are skipped (instrumented build scripts drop default_*.profraw files into the source tree), which is what the explicit `--target` in the build line is for. Verified on mainnet blocks: replaying the antichain of a finished store under full and under selective instrumentation gives the same universe and a byte-identical per-file report, at roughly a tenth of the time. `report` documents one llvm-cov convention that shows up with generic code in scope: it summarizes a generic function by its best single instantiation, so its region total can trail a report over the whole scan by a region in a const-generic family even though every source region is covered. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b1ce000c91
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
Four holes, all of them opened or widened by taking revm's execution engine into the scope. Review threads on #222. `binary_id` fingerprinted mega-evm's revision and the toolchain. The measured dependencies are now measured code, not merely linked code, so a revm bump with mega-evm unchanged rewrote part of the coverage map while leaving the id identical — and `report` validates a manifest by that id alone, so it would have accepted profiles whose functions llvm-cov then drops, reporting the difference as uncovered. The locked versions of the measured crates join the hash (sorted in build.rs, so reordering measured-crates.txt is not a new build). The universe stamp embedded absolute source paths, so two shards of one distributed scan stamped differently under different home directories and `merge`, which compares stamps byte for byte, refused them — while their item ids, which strip the prefix, were identical and mergeable. Stamps are now built from root labels, like the ids. A root that does not match the coverage map is answered by llvm-cov with a warning on stderr and a success exit, and the check here was that the scope as a whole matched something: a valid mega-evm root beside a stale revm one measured no revm at all while the store's stamp went on claiming that scope, which would invalidate a scan silently and late. Every root must now contribute a file. That is well defined per block because llvm-cov lists the files of the coverage map under the scope, not the files this block executed — verified against real per-block profiles, where each of the six roots appears on every block. `report` grew the same check against its table. A root is identified by its final path component, in both the ids and the stamp, so two roots sharing one would have had their coverage merged under a stamp still claiming two; nested roots would have made "which root owns this file" depend on the order the roots were listed in, which the sorted stamp cannot see. Both are now refused when the scope is resolved, which is also what lets the ids take the first matching root rather than the most specific. Roots are matched as spelled, never canonicalized: llvm-cov matches the absolute paths baked in at build time, so a canonical form could disagree with what llvm-cov itself matches. A path spelled a way that does not prefix the coverage map is caught by the per-root check instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b0dafdb9c5
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
…ed for Two ways `report` could run cleanly and still be wrong. Review threads on #222. A cover only promises the scope it was computed over, but a manifest carried just its `binary_id`. The archived profiles hold counters for every instrumented crate, so reporting a cover over a wider scope than its scan's — or over a root swapped for another — rendered a clean table that read whatever the cover never had to reach as uncovered code. The manifest now records the store's universe stamp, read from the store rather than re-derived from flags so it cannot drift from the scan that produced the profiles, and `report` refuses any other scope. Manifests written before the field existed still report on the build check alone, with a warning that the scope went unchecked. `report` checked each root by looking for its label anywhere in the table, because llvm-cov strips the common prefix from the paths it prints. That is a substring test, and a stale root named `src` passed it on every other root's `/src/` paths. It now reads a `--summary-only` export of the merged profile, whose filenames are absolute, and checks roots through the same prefix-matching function the per-block export uses — one definition for both. Verified on the consolidated scan's store: re-running set-cover keeps the same 86 blocks and binary_id and records the universe; `report` over the scan's scope reproduces the same totals; a narrower scope is refused; and a scan scoped to `mega-evm + revm-handler/src`, reported with a stale `/…/src` in place of the real one — same universe stamp, so only the root check stands in the way — is refused by name where the substring test had let it through. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: a32b7868a6
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
…ling, and harden the scan - Workers build the run-once precompile tables (mega-evm `rex`/`mini_rex`, op-revm `isthmus`/`granite`/`fjord`) before capturing anything. Left to the blocks they were credited to whichever block a worker replayed first, and to the first block of each later table, so ~2.5% of blocks recorded different coverage from run to run. The item universe moves to v3, so stores, shards and manifests from before are refused. - `report` re-derives the covered items from the merged profiles and refuses a count that differs from the manifest's: `binary_id` does not see the workspace crates whose metadata the measured instances are named by. - `set-cover --incumbent-manifest` matches patterns, not blocks: a pattern's representative moves to lighter blocks between runs. - Workers are spawned from, and profiles evaluated against, the running image (`/proc/<pid>/exe`), so a rebuild mid-scan cannot mix coverage maps or wedge respawns; a `%` in `--data-dir` is refused up front. - The default scope and the LLVM tools come from the cargo home and the sysroot the binary was built with; a crate found under two registry indexes is refused instead of picked by listing order. - The worker keeps stdout for its frames (library prints go to stderr), so frames are parsed strictly; item provenance travels in the frame, only for ids new to the worker, instead of a fsynced per-block sidecar file. - Contract files verified once are not re-read for every block, and the file work runs off the async runtime. - `report` works in a private temp dir and no longer scaffolds a data dir. - The wrapper reads measured-crates.txt the way build.rs does; docs corrected (R2 retention, report scope, clap `error-context` statements elsewhere in the workspace). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…the store, judge and worker plumbing Reuse: - R2 witnesses go through the workspace's R2WitnessTransport, fetcher and light decoder (one attempt per round: the fetch loop's retry stays the policy); `r2.rs`, its RedactedSecret copy and the chrono/reqwest deps go. - `LlvmArgs` is the one definition of the scope and tool flags, flattened into backfill, report and the worker; extraction is a method of the resolved value. - `Manifest::read` / `ManifestBlock::pattern_key` replace three parsers. Simplification: - Store reads are counters / patterns / blocks(range) / block_records / load; one encode/decode helper; the universe stamp is required (the legacy physical-counter label could no longer be reached), and readers that interpret dense indices (set-cover, merge) open read-only through one build check. - `PatternRecord::first_seen` / `absorb` are the one pattern fold the judge and merge share; merge keeps a single id map. - Redundancy elimination is one forward pass; the judge's universe bitmap and the response's `ok` flag were derivable and are gone. - The worker takes `--data-dir`; `SpoolEntry::open` is the one check a spool entry passes, for the worker and for a resumed run alike (the fetch now checks the block number itself); `DataDir` owns the per-block scratch name and clears what a killed run left, replacing an age-based sweep. - `report` works in a system temp dir; comments that restated one policy in several places now state it once. Efficiency: - The judge checks domination only against patterns that arrived undominated (transitivity makes that complete), not against every pattern. - Judged blocks are flushed to disk every 64 commits and on every exit path; spool entries and contract codes are written without fsync (both are re-checked before use); blocks are serialized off the async runtime; ids resolve in one Fx-hashed pass; a resumed run keeps only block statuses. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Codecov Report✅ All modified and coverable lines are covered by tests. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 1ded72fcc3
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
A profile names each instance of the measured generics by a symbol that carries the cargo metadata of the crate instantiating it, and that metadata moves with any dependency or version change. binary_id did not see it: after this branch's own dependency changes, a cover taken by the previous build evaluated to 4,124 of its 13,006 items under the new binary. `report` refused it, but backfill would still have resumed such a store and mixed the two builds' profiles in its archive. build.rs now fingerprints Cargo.lock (FNV-1a, stable across machines, so shards of one scan still agree) into binary_id, so backfill, set-cover and merge refuse a store whose profiles this build cannot read. `report`'s re-derivation stays as the check for what the id cannot see (features, compiler flags). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…arget is the switch As in the validator and the trace server: the `--r2-*` flags go through the shared `validate_r2_flags` rules (a half-configured target or a blank value is refused by name), a configured target makes R2 the first witness path, and a block R2 cannot serve within three attempts and a minute goes to the witness RPC — a second path to the same bucket. `--witness-endpoint` is therefore always required; the placeholder witness endpoints are gone. S3 target only. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b546acfe66
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| } else { | ||
| explicit.to_vec() |
There was a problem hiding this comment.
Normalize source-root paths before matching exports
When an operator passes a relative --source-dir such as . or ../mega-evm, this accepts the path, but llvm-cov exports the absolute filenames embedded in the coverage map. ensure_every_root_matched then compares those absolute filenames against this relative root with Path::starts_with, finds no match, and makes both backfill and report fail despite the root naming the correct directory. The same failure can affect the default scope if the build used a relative CARGO_HOME; convert roots to absolute paths at resolution time (without resolving symlinks) or reject relative roots explicitly.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Relative --source-dir never matches. resolve_source_dirs doesn't absolutize explicit roots, but llvm-cov exports absolute paths, so --source-dir . fails every block. Running the roots through std::path::absolute at resolution time should fix it.
There was a problem hiding this comment.
Fixed in 9578055. resolve_source_dirs now runs every root through std::path::absolute (bin/coverage-replayer/src/llvm.rs:243), so --source-dir . and --source-dir mega-evm match. It also fixes the label: . used to stamp the universe as ".".
absolute alone does not cover ../mega-evm. On POSIX it keeps .., because it does not resolve symlinks, and roots must match as spelled. A root that still contains .. is therefore refused by name (llvm.rs:245) instead of failing later as an unmatched root. Test: resolve_makes_relative_roots_absolute (llvm.rs:687).
A relative CARGO_HOME is not covered. Cargo resolves it against the directory it was invoked from, which build.rs cannot see.
…n its flushing, tidy the R2 path - The worker protocol names a block by number alone: both ends derive its spool entry and profiles from the data dir (`DataDir::block_profdata`), and new items travel as `CoveredItem` itself instead of a string-typed copy. - `DataDir` owns the archived-profile format (write and read); `llvm-cov report` is a method of `Llvm` like the other tool calls. - The store decides which commits flush (every 64th, plus `flush` on exit); callers no longer pass a durability flag. Metadata keys are named once. - Dense indices are the counter count; the judge's archive and "undominated" branches are one; merge's re-key warning lives in its insert branch. - R2: each of the three attempts gets its share of the minute, so a stalled GET no longer uses up the rest; the no-op metrics are `impl R2Metrics for ()` in stateless-common; the command-line-secret warning only this binary had is gone (the flag's doc says to prefer the env var, as in the other binaries). - `block_json` goes through serde's byte-array hooks: the same bytes, written and read in one call instead of one per byte. - Removed a block-number check the RPC client already makes, clones the fetch task did not need, and a test asserting a deleted flag stays deleted; the block statuses are dropped once the todo list exists; stale text fixed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
…onger needs - `merge` and the store's bulk-write path: the workflow is a candidate-pool re-sweep on one machine; pools from several stores still union as block lists. - The R2 witness route: witnesses come from `--witness-endpoint`, the gateway in front of the same bucket. - `inspect`'s rarity and growth statistics and `--pool-siblings`. - `set-cover --incumbent-manifest` and `--prune-profiles`; gain ties go to the higher block number. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Comment lines 1,097 → 568. Each comment now states what would break, once, at the item that owns it; history, measurements and restatements of the code are gone (the rationale lives in the PR description). No code line changed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The pruning flag is gone; dominated patterns are simply never cover candidates. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
| } | ||
|
|
||
| #[derive(Subcommand, Debug)] | ||
| enum Cmd { |
There was a problem hiding this comment.
[Minor] PR description promises features this delta removes (merge, --pool-siblings, coverage-replayer R2)
This repo uses squash-and-merge and AGENTS.md states the PR description becomes the squash commit message. git log will describe features coverage-replayer --help no longer offers, and readers grepping history for merge/--pool-siblings/R2 in this tool find promises the code does not keep. The 84 blocks cover 13,006/13,006 items claim was produced with --pool-siblings per the current text, so reproducing it after this delta needs an updated method note as well.
Suggested fix: Update the PR description before squash-merge: drop the merge bullet from ## How it works, drop the --pool-siblings N sentence from ## Carrying a scan across a mega-evm bump, and rewrite the ## Notes / Witness sources paragraph so it no longer describes a coverage-replayer R2 route (or notes that R2 was removed and this binary now uses --witness-endpoint alone). Also revisit the merge row of the Performance at full-history scale table or move it under a Prior scan heading so it reads as history.
Restoring clap's `help` feature made `--help` render every env-backed flag as `[env: NAME=<current value>]`, so on a host or CI job with the R2 credentials in its environment, `stateless-validator --help` and `debug-trace-server --help` printed them in the clear. `RedactedSecret` does not help there: clap prints the raw env string, not the parsed value's `Debug`. The three flags each binary already redacts (the Access client id and secret, and the S3 secret access key) now set `hide_env_values`, so the help still names the variable but not its value. A test per binary sets all three, renders the long help, and checks that the names are there and the values are not. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
llvm-cov lists the absolute source paths baked into the coverage map, and a root is matched against them with `Path::starts_with`, so a relative `--source-dir` such as `.` or `mega-evm` matched nothing and failed every block in `backfill` and `report` although it named the right directory. It also gave `.` the label `"."`, so the universe stamp depended on how the root was spelled. Roots now go through `std::path::absolute` when they are resolved. It does not resolve symlinks, which keeps roots matching as spelled, and for the same reason it keeps `..`; no listed path contains `..`, so a root that still does is refused by name rather than failing later as an unmatched root. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 9578055b61
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| .filter(|l| { | ||
| l.starts_with("release:") || l.starts_with("host:") || l.starts_with("LLVM version:") | ||
| }) | ||
| .collect::<Vec<_>>() | ||
| .join(";"); |
There was a problem hiding this comment.
Include the rustc commit in the coverage namespace
When scans are built with two nightly snapshots that share release:, host, and LLVM version, this filter emits the same COVERAGE_RUSTC_VERSION even though rustc -vV also exposes distinct commit-hash/commit-date values and compiler changes can alter MIR coverage mapping. Unlike the earlier optimization-profile concern, this is a compiler-version change. Resuming into the same data directory then retains old Ok bitmaps and profiles alongside counters from the new compiler; include the commit hash (or the complete -vV output) in the fingerprint so this starts a new coverage namespace.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in a7fe2a1. The toolchain fingerprint now keeps commit-hash: alongside release:, host: and LLVM version: (bin/coverage-replayer/build.rs:37), so two nightlies of the same release cycle no longer share a namespace. A store written by an earlier build is refused by this one.
| let codes = client | ||
| .get_codes(&missing, true) | ||
| .await | ||
| .map_err(|e| eyre::eyre!("fetch {} bytecodes: {e}", missing.len()))?; |
There was a problem hiding this comment.
Bound bytecode fetch concurrency
With the default 32 fetch tasks, an uncached block that references N distinct contract hashes reaches this call with all N hashes at once. RpcClient::new leaves the data-request semaphore unlimited, and get_codes starts every hash through try_join_all, so this can burst to 32×N concurrent eth_getCodeByHash calls against one endpoint. Gate or batch these code fetches (and expose a data concurrency limit) so a high-fan-out block cannot trigger gateway rate limits and stall the scan in its infinite retry loop.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in 42e49d4. backfill now builds its client with data_max_concurrent_requests, set by a new --data-max-concurrent-requests flag (env COVERAGE_REPLAYER_DATA_MAX_CONCURRENT_REQUESTS, default 64, at least 1): bin/coverage-replayer/src/backfill.rs:96, :259. The cap covers block and bytecode requests together, so the get_codes fan-out at backfill.rs:502 queues on the same semaphore instead of bursting. Test: data_requests_are_capped_by_default (backfill.rs:1035).
| Self::Range(r) => store.blocks(r.clone(), |n, record| { | ||
| found.insert(n, record.status); | ||
| })?, |
There was a problem hiding this comment.
Stream range statuses instead of materializing the block table
On a resumed full-history range, this path inserts every BLOCKS row into a HashMap solely to decide whether it is Ok; it then separately materializes the remaining block numbers in todo. With tens of millions of completed rows, the status map alone consumes hundreds of megabytes before any worker starts and can prevent the documented resumable scan from running. Stream the range records while deriving pending work, or use bounded/batched status lookups rather than retaining every status.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in 06b5732. Selection::pending (bin/coverage-replayer/src/backfill.rs:166) walks a range against the store's rows in block order and yields the pending blocks and the retry count directly, so no per-row status map is built; only the pending list itself is held. Block lists are still looked up point by point, since they are pool-sized. Test: pending_skips_only_clean_replays (backfill.rs:1044), which fails on each of three mutants of the walk.
CodeQL's cleartext-logging query reads a binding named `secrets` as sensitive data and flagged the `assert!` messages that name each variable. The table holds variable names and dummy values, and the messages print only the names; calling it `vars` says that and keeps the query quiet. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The toolchain part of `binary_id` kept only the `release:`, `host:` and `LLVM version:` lines of `rustc -vV`. Every nightly of a release cycle shares `release:` (1.95.0-nightly for all of them) and often the LLVM version, so two different compilers produced the same namespace, and a store could be resumed by a build whose compiler maps coverage differently. `commit-hash:` now joins the fingerprint. A store written by an earlier build of this branch is refused by this one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`backfill` built its RPC client with the default config, which leaves the data-request semaphore unlimited, and a block fetches every bytecode it lacks in one `try_join_all`. With 32 block fetches in flight, a stretch of blocks that each reference many new contracts could put 32 x N `eth_getCodeByHash` calls on one endpoint at once, enough to hit its rate limit and stall the scan in its retry loop. `--data-max-concurrent-requests` (env COVERAGE_REPLAYER_DATA_MAX_CONCURRENT_REQUESTS, default 64, at least 1) now caps block and bytecode requests together. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Resuming a range inserted every stored row of it into a HashMap only to tell clean replays from the rest, then built the pending list from the selection in a second pass. On a resumed full-history range that is tens of millions of rows, hundreds of megabytes before any worker starts. `Selection::pending` walks the range against the store's rows in block order and yields the pending blocks, in order, with the count that have a record to retry; a block list is still looked up point by point, since lists are pool-sized. `Selection::iter`, whose only caller was the old filter, is gone. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Summary
Adds
coverage-replayer, an offline tool that answers "which mainnet blocks, replayed, exercise every execution path — in mega-evm and in the revm execution engine it drives — that we have ever seen taken?" — the fixture set a stateless validator wants for regression coverage, without replaying the whole chain each time.Every mainnet block from 1 to 27,814,972 has been replayed, all matching their headers, and 84 blocks reproduce all 13,419 coverage items those replays observed under mega-evm v1.7.0 — 61.8% of the branch arms and 63.6% of the regions in the measured code; the rest are paths mainnet has never taken.
Supersedes #153, which was opened before several large refactors landed on
mainand whose diff had become dominated by contentmainalready carries. This branch ismainplus the tool only, and the port to mega-evm v1.7.0 needed a single argument dropped from tworeplay_blockcall sites (the EIP-3155 trace writer removed in #195, which this tool always passed asNone).How it works
backfillreplays blocks under LLVM branch instrumentation. Resident worker subprocesses reset counters, replay one block, and capture a profraw — having first built the per-hardfork precompile tables, which mega-evm and op-revm construct once per process and would otherwise be credited to whichever block a worker happened to replay first; a judge dedups the resulting per-block bitmaps into patterns in a redb store, keeping the lightest block of each pattern as its representative (bin/coverage-replayer/src/backfill.rs,worker.rs). Nothing block-sized is retained — the only per-pattern artifact kept is a small sparse profdata forreport.A bit is an evaluated coverage item, not a physical counter. rustc minimizes physical counters: an
if/elsegets two (entry, then-arm) and the else-arm exists only as the expressionentry - then. Over physical counters an else-only block reads{entry}— a strict subset of a then-only block's{entry, then}— so it looks dominated, is never archived, gets pruned, and the cover silently loses a branch arm the scan had covered. No-Z coverage-optionsvalue turns the minimization off. So each block's profile goes throughllvm-cov export --format=text --skip-functions, scoped to the source dirs, and the items are its region entries and branch arms with a non-zero count, keyed by source span (llvm.rs) — the same arithmeticreportruns. It costs ~0.6 s per block. An unscoped export of this binary segfaults llvm-cov, so the scope is not optional.set-cover— greedy cover over the patterns with antichain pruning and a redundancy-elimination pass. The completeness contract is that the selected set always covers the full observed universe; no selected block can be dropped, but the set is not guaranteed minimum-cardinality (setcover.rs:1). Gain ties go to the higher block number, so a store always yields the same selection.report— llvm-cov summary for the selected set.inspect— read-only store statistics. The one subcommand that skips the binary-id namespace check, so a store from another build can be analyzed.Kept out of this PR to keep it reviewable, and recoverable from the branch history when needed: a
mergesubcommand for combining per-machine shard stores (the workflow is now a pool re-sweep on one machine; pools from several stores still union as plain block lists), reading witnesses straight from R2 (the witness RPC is the gateway in front of the same bucket),inspect's rarity and growth statistics and--pool-siblings, andset-cover's--incumbent-manifestand--prune-profiles.Carrying a scan across a mega-evm bump
Counter ids belong to one instrumented build and one item definition, so
binary_id(mega-evm rev, toolchain, measured crates and the lockfile) and a universe stamp (item definition + source scope) namespace every store and the write paths refuse a mismatch (store.rs:1). That would strand a scan at every mega-evm bump, since a full sweep costs weeks of machine time. What survives a bump is the block numbers, so the tool carries those over instead of the bitmaps:The pool is the antichain's representatives, not the previous minimal set: a cover is minimal only for the universe that produced it and carries no slack once a new build splits patterns the old one merged. Everything the antichain leaves out was strictly dominated — its coverage a subset of a kept block's — which is the closest available stand-in for "adds nothing".
This branch's store carries its own pool for the next bump: 35,678 blocks, the antichain of its 226,158 patterns (30,123 of them from blocks up to 20,980,000, 5,555 from the extension past it), and the next bump re-sweeps only the pool. Kept next to it are the block numbers of all 226,158 pattern representatives: bitmaps, profiles and stores can all be rebuilt by replaying that list, the list itself only by replaying the whole range again.
Results: every block from 1 to 27,814,972
Every mainnet block from 1 to 27,814,972 has been replayed and matched its header, in three passes:
divergent=0 error=0(63 of the sampled blocks fall inside that stretch).06b5732; itsbinary_idalso keys the rustc commit, so earlier stores do not open under it) re-swept those 77,915 blocks into a fresh store — the same 84 blocks in the cover, with an identical per-file report over them — then replayed every block from 20,980,001 to 27,814,972, the finalized head when the run started: 6,912,887 blocks, 762.9M transactions,divergent=0 error=0.mega-evm v1.7.0 (this branch's build)
84 blocks cover 13,419/13,419 items. Coverage of that set, by measured crate:
The table is llvm-cov's count; the tool's 13,419 items tally slightly differently from the table's 12,307 regions and 1,252 branch arms (13,559 together): an item is keyed by source position, so regions that start at one position are one item and a generic function's instantiations are OR-ed, while llvm-cov counts regions per function and summarizes a generic by its best single instantiation.
The extension past 20,980,000 added 413 items (13,006 → 13,419), 392 of them from blocks 24,829,789–24,850,253, and lifted branch-arm coverage from 58.4% to 61.8% (mega-evm alone from 56.1% to 61.5%). The cover stayed at 84 blocks, 35 of them replaced, 27 by blocks from the extension.
In mega-evm the largest gap is
sandbox/, with 37 of its 212 branch arms covered (17.5%): mainnet barely exercises it.reportover the selected blocks matchedreportover the union of every archived profile (35,188) exactly on lines, functions and branches, and on regions for 124 of 126 files. The other two differ by one region each, both in const-generic families (push::<N>in revm, one in mega-evm's instructions): llvm-cov summarizes a generic function by its best single instantiation, and comparing region by region across instantiations the selected set misses none.OnceBox::get_or_init. Those closures used to be credited to the first block each worker replayed, and to its first block of each later table, so about 2.5% of blocks recorded different coverage from run to run and re-deriving the cover gave a different selection each time. Workers now build every table before capturing anything (worker.rs), and the item universe moved to v3, so stores, shards and manifests from before are refused. Verified: the same 401 blocks replayed twice by 30 workers in fetch-completion order give identical patterns block for block; a REX-era block replayed alone and after a MINI_REX block gives the same pattern, where the previous build differed by 55 items; and the two definitions do not mix —backfillwill not resume a v2 store andreportwill not read a v2 manifest. Against the previous run of the same 52,340 blocks, the per-file report changes in the twoprecompiles.rsfiles only — 54 regions of table construction (13 in mega-evm, 41 in op-revm), now reported uncovered, matching the 54 items the universe lost — and the cover is 84 blocks instead of 86 (71 shared) with branch coverage identical. Three further full re-sweeps selected the same 84 patterns — the last with every witness fetched through the RPC gateway, down to the same blocks in the same order; which block stands for a pattern can still differ between runs, since a pattern's representative is its fastest-replaying block (one of the 84 did once, for an equivalent block with the identical bitmap). Adding the 25,638 blocks of 6,624,362–6,649,999 left the selection unchanged, down to the block numbers.sandbox::execution: keyless deploy nonce read failed error=Metadata not in witnessfor a handful of stress-range blocks whose replays nevertheless matched their headers — worth a look on the mega-evm side.Per-file coverage of the 84 selected blocks —
llvm-cov reportover their profiles, 126 files (paths relative to each source root:mega-evm/is the v1.7.0 checkout's crate, the revm crates are at their locked versions)The 84 selected blocks — sorted by block number: block number, items it newly covered when picked, items it covers on its own, pattern key, block hash
The full-history scan (mega-evm v1.6.1, physical-counter universe)
These numbers come from the scan that ran before the item definition above existed; its stores are legacy stores (recognized as such and never mixed with current ones). They stand as a scan of mainnet, with one known bias stated below.
Four shards, combined with a
mergesubcommand since removed (mega-evm v1.6.1, nightly-2026-02-03):ok=20954362 divergent=0 error=0— every block the four shards were assigned replayed cleanly against its header (gas / receipts root / logs bloom). Their ranges left one stretch, 6,624,362–6,649,999, to none of them; this branch's build replayed it instead (above).sandbox/*alone accounts for 168 of the 426 missed arms.evm/execution.rs,evm/precompiles.rs,evm/instructions.rs). That is a lower bound: the archive was filtered by the same domination. This is the finding that led to the item definition above.Performance at full-history scale
One cost only shows at this scale — the antichain prune, which both commands below run; it was found on the consolidation run above and fixed here, then re-measured on the same store (old-pin build, same machine):
select_coverinsideset-coverinspect --dump-poolThe antichain prune tested each pattern against every earlier one (~10¹² iterations for 1.5M patterns).
split_antichain(setcover.rs) scans only kept patterns and finds candidates through an inverted index from counter to the kept patterns containing it: a superset must contain the candidate's rarest counter, and a counter no kept pattern has proves the candidate maximal outright.Equivalence was checked three ways: the manifest from the new code is identical to the quadratic run's (same 94 blocks, order and gains), the exported pool is line-for-line identical, and a differential test keeps the quadratic scan as the oracle over randomized stores (verified to fail on a mutant).
Replay throughput. Two costs, neither of them the network, kept the stress-test range at a crawl (one block in nine minutes at worst):
RUSTFLAGSinstrumented every crate, and tiny hot functions (k256, per-byte bincode, the interpreter loop) paid a shared-counter increment per basic block, with threads contending for the same cache lines.cov-rustc-wrapper.shinstruments the measured crates and the workspace only. The workspace is required — most measured code is generic and is compiled where it is instantiated — and nothing outside the workspace depends on mega-evm. Replaying a finished store's antichain both ways gives the same universe and a byte-identical report, in about a tenth of the time. The registry crates that depend on the measured revm crates (alloy-evm, alloy-op-evm, revm, revm-inspector, four reth crates) stay uninstrumented on evidence: a build instrumenting them too, replaying the same 401 blocks, gave the same 11,706-item universe and an identical per-file report for all 127 measured files, at ~7% more worker time per block.--fetch-concurrencynow defaults to 32: the dispatch queue bounds the spool backlog, so the only cost of a higher value was idle workers. The densest stretch, 6,624,362–6,649,999 at ~21,000 transactions a block (~540M in all), runs at ~4.8 blocks/s on 32 workers — CPU-bound on the replay itself, with fetched blocks queued ahead of every worker — so its 25,638 blocks take about 90 minutes.inspect --no-cover-previewskips the set-cover pass, which after the above is about half ofinspecton a full-history store; it stays on by default because it is the only non-destructive way to see the cover.Testing
cargo test --workspacegreen. Unit tests cover the item extraction against committedllvm-cov exportfixtures (the then-only / else-only pair that physical counters got wrong), the judge's dedup/promotion paths, the set-cover algorithm (domination strictness, gain tie-breaks, redundancy elimination), the report's re-derivation check, scope detection under several registry indexes, tool lookup order, the judge's domination check through the undominated set, the pattern-key probing walk, store namespacing, spool checksums, and the block-list/pool formats.bin/coverage-replayer/tests/replay_fixtures.rsreplays every mainnet fixture through the exact path the worker uses and checks the header sanity triple — it runs uninstrumented, so it guards the replay glue in normal CI.tests/worker_protocol.rsdrives the real binary as a worker and checks that a request gets exactly one frame on stdout. Two properties are verified only by instrumented runs, since an uninstrumented test sees no counters and no stray prints: the warm-up (the replays above) and the worker's exclusive stdout.Notes
[profile.coverage]build, made throughRUSTC_WRAPPER=bin/coverage-replayer/cov-rustc-wrapper.sh;Cargo.tomlcarries the build line and the reason-C link-dead-codemust not be added.measured-crates.txtis the single list both the wrapper and the default--source-dirscope are derived from.build.rs, so neithersudonorCARGO_HOMEmoves it), a crate present under two registry indexes is refused rather than picked by listing order, and the LLVM tools default to the ones shipped with the building toolchain. A scope that matches nothing fails the block instead of recording an empty bitmap. An explicit--source-diris made absolute (symlinks left unresolved) and one still containing..is refused, since the paths llvm-cov lists are absolute and never contain...reportscopes llvm-cov to the measured sources: reporting over the full covmap crashes llvm-cov on an instantiation-group bug.--helpfix, workspace-wide. clap was pinned withdefault-features = falseand onlyderive/env/std, which drops itshelp,usageanderror-contextfeatures — so--helpanswerederror: unexpected argument foundon all three binaries, includingstateless-validatoranddebug-trace-server, and errors named neither the offending flag nor the usage. This PR opts those three features back in (Cargo.toml) and pins the behaviour with tests by error kind. Withhelpon, clap renders an env-backed flag as[env: NAME=<current value>], so the three R2 secrets each binary already redacts inDebug(the Access client id and secret, the S3 secret access key) sethide_env_values:--helpnames the variable and not its value, which a test per binary checks with all three set.reportagainst its table) — reproduced first: a stale revm root beside a valid mega-evm one dropped the file list from 126 to 47 while the store's stamp still claimed the full scope.binary_idhashes the measured crates' locked versions, since a revm bump with mega-evm unchanged rewrites part of the coverage map andreportvalidates a manifest by that id alone. The universe stamp is built from root labels rather than absolute paths, so the same scope stamps identically whatever home it was resolved under. Duplicate and nested roots are refused when the scope is resolved, which is what makes a label a sound identity and lets ids take the first matching root. The manifest records the universe it was computed over, andreportrefuses any other scope: the archived profiles carry counters for every instrumented crate, so a wider scope would otherwise report cleanly and read the gap as uncovered code.reportre-derives the covered items from the merged profiles through the same extraction the per-block path uses — which also checks every root — and refuses a count that differs from the manifest's. The measured generics are instantiated in the workspace, and their profile names carry the workspace crates' cargo metadata, so a version bump or a dependency change renames them. That was observed, not hypothetical: after this PR's own dependency changes, the previous build's cover evaluated to 4,124 of its 13,006 items under the new binary, andreportrefused it.binary_idtherefore fingerprints the lockfile too, so such a build refuses the store up front instead of resuming it and mixing two builds' profiles in its archive; the re-derivation stays as the check for what the id cannot see (features, compiler flags).ifinside amacro_rules!body (verified on this toolchain — noBranchline at the call site, and no expansion branch records in a 400-profile export of the measured binary). revm's instruction macros are the main place this matters./proc/<pid>/exe), so rebuilding the binary mid-scan can neither mix coverage maps nor wedge respawns on a deleted executable. A worker keeps stdout for its frames alone — library prints go to stderr — so frames are parsed strictly instead of salvaged. Item provenance travels in the frame, only for ids new to that worker, replacing a per-block, fsynced sidecar file. Contract files verified once are not re-read and re-hashed for every block, and that file work runs off the async runtime. A%in--data-dir(a filename pattern to the profile runtime) is refused up front, andreportworks in a private temp dir.error-contextelsewhere in the workspace. Enabling it made five comments and three AGENTS.md sentences in the validator, the trace server andstateless-commonfalse ("clap names no argument"); they are corrected. The R2 rules still run after parsing, for the reasons that remain: one verdict in one wording for both binaries, and a blank env line diagnosed as blank rather than as a phantom conflict.LlvmArgs, shared bybackfill,reportand the worker, and a spool entry passes one check,SpoolEntry::open, for the worker and a resumed run alike.-9, the resumed run ended on the identical pattern table). Spool entries and contract codes skip the fsync, since both are re-checked before use.🤖 Generated with Claude Code