Skip to content

#426 - L1 taxonomy: classify cancellations and cascade-only TCPStore logs - #439

Open
continue-revolution wants to merge 11 commits into
mainfrom
chengsongz/taxonomy-category-fixes
Open

continue-revolution wants to merge 11 commits into
mainfrom
chengsongz/taxonomy-category-fixes

Conversation

@continue-revolution

Copy link
Copy Markdown
Contributor

Five category descriptions in l1/categories.json. No decision values change; no code changes.

Each was found by reading a production misclassification back to the description that caused it.

What changed

cat problem fix
12 680 chars, longest of 38 (median 142), and the only one in prohibitive voice. No clause for "the initiating rank's output is absent", so the model declined instead of classifying. Say what to do when TCPStore is the only surface: pick it, with reduced confidence. Keep the guard that a visible upstream fault wins.
1 Read, in full, "no application fault" — false on its face, since a kill leaves plenty of application-shaped noise. Give it both surfaces a cancellation actually leaves: the bare slurmstepd marker, and progress-then-teardown. Name the cat 2 discriminator.
28 Names "TypeError/AttributeError on None", which matches interpreter-teardown noise verbatim. Exclude anything raised during shutdown: a deterministic code fault recurs on unchanged retry, so it fails on the first attempt, not after 3390 iterations.
34 Pinned rank 608 and the GPTDataset index-load path. Described as a class; discriminating signatures kept.
35 Pinned ckpt 27600. Same. (Its row indices were already removed on main by 3e96f3a, which also covered cat 29.)

Why it mattered

Hand-cancelled jobs were being recommended STOP. Their logs hold 751 SIGTERMs, 847 Exception ignored and healthy iterations to the last line before the signal; the grounded primary failure was Python's own weakref finalizer losing a module global during teardown — which Python labels Exception ignored precisely because it did not affect execution.

Separately, cascade-only TCPStore logs returned category_id: 0 ("none matched") at confidence 0, so the alert told the reader nothing about a job that had cleanly checkpointed and then lost its store.

Validation

  • 26 live Nemotron logs, re-analyzed on an isolated instance: 5 changes, all accounted for. Three 0 → 12 (intended), one 0 → 1 (a correction — that log has 6180 SIGTERMs and 3 TCPStore lines), one 26 ↔ 32 shown to be model variance by repeat trials on identical input, both STOP.
  • 13 labeled dataset cases covering every edited category, with positive and negative controls: 0 regressions, 1 improvement. Config errors still STOP, node failures still cat 2, labeled TCPStore cases still cat 12.
  • Seven incidents inspected individually. Five cancellations corrected STOP → RESTART. 4188567 correctly left at STOP/cat 17 — its kill followed a real MissingEpochIndex dataset fault (7104 occurrences, zero completed iterations), and the broadened cat 1 did not claim it.

Note for reviewers

An earlier revision of cat 1 described only the progress-then-teardown shape and broke the canonical case — two labeled cases whose entire log is *** STEP <id> ON <node> CANCELLED AT <time> *** drifted to cat 2. That is the same over-fitting this PR removes from cat 12, and it was caught only because the control set covered every edited category; both categories are RESTART, so headline accuracy was unchanged. Worth keeping that shape of control for future taxonomy edits.

Two follow-ups not addressed here: cats 26 and 32 are not reliably distinguishable (demonstrated twice), and cats 1 and 3 both have a claim on a log containing only a cancellation marker.

🤖 Generated with Claude Code

continue-revolution and others added 3 commits October 8, 2026 14:57
The description was the longest of the 38 at 680 chars against a median of
142, and the only one written in prohibitive voice. Most of it warned the
model away from the category, with no clause for the case where the
initiating rank's output is absent from the log entirely, so the model
resolved that ambiguity by declining: 3 of 28 live analyses returned
category 0 at confidence 0 with a TCPStore observed failure, while category
12 was only ever picked when a primary failure had been identified.

Say what to do when TCPStore is the only surface present - pick it, with
reduced confidence - and keep the guard that a visible upstream fault wins.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both were written from the seed incident that produced them: cat 34 pinned a
specific rank (608) and the GPTDataset index-load path, cat 35 a specific
checkpoint (27600). Those name nothing a model can find in a fresh log.

Keep every signature that actually discriminates - the wild-vs-near-null
faulting address, the node-local blast radius - and drop the identifiers that
only described one log. Category 29's row indices were already removed on main
by 3e96f3a, so it is left as it stands.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Hand-cancelled jobs 4285548 and 4285549 both came back STOP under category 28,
on the strength of "TypeError: 'NoneType' object is not callable". Their logs
hold 751 SIGTERMs, 847 "Exception ignored" and healthy iterations to the last
line before the signal; the grounded primary failure is Python's own weakref
finalizer losing a module global during interpreter teardown. Python labels
those Exception ignored precisely because they did not affect execution.

Category 1 was the right answer and read, in full, "no application fault" -
false on its face here, since a kill produces plenty of application-shaped
noise. Give it the two surfaces a cancellation actually leaves: the bare
slurmstepd marker, which is the whole content of some logs, and the
progress-then-teardown pattern. Name the cat 2 discriminator so a node-failure
reason still lands there.

Exclude teardown from 28 on a general principle: a deterministic code fault
recurs on unchanged retry, so it fails on the first attempt, not after 3390
iterations.

Verified on five cancellations (4285548, 4285549, 4240576, 4189651, 4255225),
all STOP -> RESTART, and on 4188567, whose kill followed a real dataset fault
and correctly stayed STOP at category 17.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@greptile-apps

greptile-apps Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[Low impact] Updates taxonomy descriptions for log classification.

The PR should not merge until category 28 stops excluding repeatable code errors merely because they occur after successful training.

Findings

  1. P1 Late code errors lose STOP ▶

Summary

Updates descriptions for scheduler cancellations, downstream termination signals, TCPStore failures, and native crashes. Removes decision wording from several descriptions without changing their stored decisions.

  • Cancellation descriptions now cover both bare scheduler markers and progress followed by signal-driven teardown.
  • TCPStore descriptions distinguish setup failures from downstream errors after training progress.
  • continue-revolution explicitly deferred the overlap between categories 26 and 32, and between categories 1 and 3 for cancellation-only logs.

Reviews (9) · Last reviewed commit: "Merge branch 'main' into chengsongz/taxo..." · Reviewed by Greptile

"id": 28,
"name": "Launch config or workload code deterministic error (TypeError/AttributeError)",
"description": "Deterministic runtime error inside already-loaded workload code that will recur on unchanged retry: TypeError/AttributeError on None value threaded through model code (e.g. float(None) in moe_layer postprocess), missing MASTER_ADDR, argparse error, or world-size mismatch. Excludes ModuleNotFoundError / ImportError where the failing import path resolves under Lustre or a mounted network filesystem - those are cat 16 (filesystem visibility), not a workload code fault, even if no adjacent FileNotFoundError is present on the same line.",
"description": "Deterministic runtime error inside already-loaded workload code that will recur on unchanged retry: TypeError/AttributeError on None value threaded through model code (e.g. float(None) in moe_layer postprocess), missing MASTER_ADDR, argparse error, or world-size mismatch. Excludes ModuleNotFoundError / ImportError where the failing import path resolves under Lustre or a mounted network filesystem - those are cat 16 (filesystem visibility), not a workload code fault, even if no adjacent FileNotFoundError is present on the same line. Also excludes anything raised during interpreter shutdown - inside \"Exception ignored\" / atexit callbacks, or appearing only after sustained successful iterations. A deterministic code fault recurs on an unchanged retry and so fails early and every time; one that appears only after thousands of iterations, at teardown, is the wake of an external kill (cat 1).",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Late code errors lose STOP

Category 28 now excludes errors “appearing only after sustained successful iterations” and says deterministic faults fail early. A reproducible TypeError or AttributeError can first occur in a later evaluation step or a data-dependent model branch, without any shutdown or external kill. The new description excludes that error from its correct STOP category.

If the category-driven restart override is enabled and the model selects category 1 instead, the job can restart only to hit the same error. Limit the exclusion to errors shown to come from interpreter shutdown; successful earlier iterations do not rule out a deterministic workload bug.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

continue-revolution and others added 2 commits October 8, 2026 15:26
Cat 12 was claiming logs where training had already run. The labeled cases it
is meant for - b1-048/049/050 - all have zero completed iterations and fail
during rendezvous; the production complaints all progressed first (iteration
48310, 6810, 12) and then saw TCPStore resets as survivors noticed a peer
vanish. In 4303520 that peer was a worker an operator killed by hand, so the
alert read "Distributed rendezvous / TCPStore socket timeout" for what was
really someone running kill -9.

Completed iterations settle it: a rendezvous failure cannot happen after
iteration 48310, because rendezvous already succeeded. Gate cat 12 on no
progress, and give cat 3 the general mid-training case - faulthandler dump,
store resets, destroy_process_group-not-called at abrupt exit, atexit noise -
with the cross-references that keep cat 1 (explicit scheduler marker) and cat
12 (never started) distinct.

Both are RESTART, so no verdict changes; the explanation stops asserting a
rendezvous fault that did not occur. Reported by the attribution channel
thread, where the NVRx view was that "no primary failure" is the correct
statement since the applog never sees the SIGKILL.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two cases from the attribution channel exposed descriptions that list symptoms
without the condition that decides them.

4225797 stopped at iteration 4967 of 601014 with exit_interval unset, and came
back cat 4 "clean planned exit" at confidence 96 - matched on a final
checkpoint save, a training-energy summary and teardown warnings, all of which
a graceful shutdown on signal also produces. Require positive evidence the run
reached its configured end; last iteration against the schedule total settles
it.

4224356 carries 6878 SIGTERMs and went to cat 3, whose whole premise is that
the cause is not in the log. The cat 1 cross-reference was a trailing sentence
against a description that advertises atexit noise as its own signature. Make
the absence of any termination signal a requirement of cat 3 rather than a
footnote.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Comment thread src/nvidia_resiliency_ext/attribution/restart_agent/l1/categories.json Outdated
continue-revolution and others added 4 commits October 9, 2026 15:08
…ause"

This reverts commit 0a94ce1.

Neither gate worked as prompt text. 4225797 stayed on cat 4 at confidence 96
despite a requirement that the run reach its configured end, and 4224356 moved
off cat 3 only to land on cat 6 (EXCLUDED) rather than cat 1. A trailing
condition does not override a leading symptom list.

The cost was real: cat 3 at least yields RESTART, which is the correct action
for both. Trading a slightly-wrong category for a category with no useful
action is the wrong trade. These two want a deterministic gate - last iteration
against the schedule total, presence of SIGTERM - computed in evidence rather
than asserted in a description. Left for later.

The cat 3/12 split from 8590c43 stays; it is what fixes six of these eight.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Nine descriptions asserted the action as well as the evidence - "so product
policy calls this a STOP", "external guidance is restart", "so this stays
retryable" (the last one mine). The decision field already carries that, so
the description was a second, unenforced copy that can drift from it, and it
mixes two jobs: L1 classifies evidence, L4 decides what to do about it.

Every claim about the nature of the failure is kept, since that is the
evidence a classifier should weigh - an unchanged workload hitting the same
allocation ceiling, a CUDA OOM reproducing at fixed config, a single log being
unable to separate a library race from memory corruption. Only the verdict
drawn from it is removed.

Cats 7 and 8 still say a fault may "clear on an in-place restart"; that
describes transience rather than prescribing an action, so it stays.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Stripping these alongside the policy verdicts was too blunt. "Restart from
last good checkpoint is plausibly safe" is not a verdict duplicating the
decision field; it says a known-good recovery point exists and resuming from
it will not carry the corruption forward - a property of the failure. It also
carries an operational distinction the replacement lost: roll back rather than
resume in place.

The line is between asserting the action ("product policy calls this a STOP")
and describing what recovery looks like. Only the former is removed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
4234179 c4 and 4307920 c2 both came back category 0 on a trace-export
connection failure. Neither is a telemetry failure. 4234179 c4 runs normally
for 84k lines to iteration 45250 of 615915, then the exporter starts failing,
then worker processes are reaped leaving "resource_tracker: 27 leaked
semaphore objects", then the log stops. The exporter losing its collector is
an early symptom of the same shutdown, and the only error-shaped text in the
file, so it got reported as the finding.

Cat 3 already owns this story but listed none of these surfaces - it names
faulthandler dumps, TCPStore resets, destroy_process_group and atexit, and
neither log has any of them. Add the two that do appear, and say plainly that
exporter errors are residue rather than fault.

Described by behaviour rather than by endpoint: naming the collector's port
would re-introduce the seed-incident specificity this branch removes from
cats 29, 34 and 35.

No new category: the failure is an external termination with the cause
off-log, which is what cat 3 is.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@continue-revolution
continue-revolution force-pushed the chengsongz/taxonomy-category-fixes branch from 37a32a7 to eb7be87 Compare October 10, 2026 01:03
Comment thread src/nvidia_resiliency_ext/attribution/restart_agent/l1/categories.json Outdated
…rfaces"

This reverts commit eb7be87.

It did not work. 4234179 c4 and 4307920 c2 stayed on category 0 with the new
surfaces named. The only thing it moved was 4225797, from cat 4 to cat 0 - a
case the change does not describe - which is prompt coupling, not a fix.

It also took cat 3 to 1125 characters against a 142 median, making it the
longest in the taxonomy. An over-long description written around specific
incidents is what made cat 12 misbehave in the first place.

Third attempt at steering these by wording, after the cat 3 and cat 4 gates.
They want a deterministic signal - exporter-only error surface, last iteration
against the schedule total - computed in evidence rather than asserted in a
description.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@continue-revolution

continue-revolution commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor Author

Generated by human --

This PR is mainly to fix:

  • Patterned TypeError failure caused by slurm cancellation. Previously go to launch config error category (c28) and STOP, now go to slurm cancellation category (c1) / off-site termination category (c3) and RESTART
  • Patterned TCPStore failure caused by outside termination. Previously go to no match category (c0), now go to off-site termination category (c3), action unaffected
  • Remove any actions / policy in category description, and remove any data index leak from category description

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant