Skip to content

Fail-open resolver errors are indistinguishable from "already latest" after the fact #35

Description

@algomaster99

Summary

When the PreToolUse hook's resolver check fails to flag an outdated pin, there's currently no way to tell, after the fact, whether that's because the pin genuinely was already latest, or because the resolver silently errored and yul failed open. Both look identical from the outside: exit 0, no visible output.

What's happening today

main.go's runHook does write resolver errors to stderr before failing open:

mismatches, err := checker.Check(before, after)
if err != nil {
    fmt.Fprintf(os.Stderr, "hook: %v\n", err)
    os.Exit(0) // fail open: a resolver/network error shouldn't block the write
}

if len(mismatches) == 0 {
    os.Exit(0)
}

But that's the only place this information exists. It's not written anywhere persistent (no log file — runHook and runScan both have zero logging infrastructure beyond ad hoc fmt.Fprintf(os.Stderr, ...) immediately before an os.Exit(0)/os.Exit(2)). And since Claude Code only surfaces a PreToolUse hook's stderr to the model when the hook blocks (non-zero exit), a fail-open resolver error is invisible both to the user in a live session and to anyone auditing a completed run's transcript (e.g. benchmark/run_case.sh's transcript.jsonl, which is what the linked PR comment's analysis was built from) — the error was written to a stderr that nothing downstream captured.

Why this matters beyond the benchmark

This is the same blind spot noted in #34 as an open question ("is this a resolver/PURL problem or something else") — without persistent logging there's no way to build confidence about how often or why yul fails open in real usage, only that it sometimes does. It also means a user who hits this in a live session has no way to know their outdated pin slipped through silently rather than being genuinely fine.

Possible directions (not prescribing one)

  • A structured, persistent log (e.g. under ~/.cache/yul/, alongside the existing scan cache) that every fail-open path appends to, distinguishing "no mismatches found" from "resolver error, failed open" from "manifest/parse error, failed open."
  • A hookSpecificOutput.additionalContext on the success path too (not just blocking), when the resolver errored — though that risks being noisy on every keystroke if resolver hiccups are frequent.
  • At minimum, giving checker.Check's error path and the "genuinely 0 mismatches" path distinguishable exit codes or stderr markers, so a wrapper script (or this benchmark's own tooling) could tell them apart without needing new logging infra.

Relationship to #34

Filing separately since #34 is about a specific, reproducible bug (spring-core/jetty-server never resolving, even outside the benchmark's exact scenario). This issue is about the general lack of observability into any fail-open path — pandas/scipy might turn out to share #34's root cause, or might not; there's currently no way to know without the kind of manual, one-off reproduction the PR comment already did by hand.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions