Repository navigation
#426 - L1 taxonomy: classify cancellations and cascade-only TCPStore logs - #439
continue-revolution wants to merge 11 commits into
Conversation
The description was the longest of the 38 at 680 chars against a median of 142, and the only one written in prohibitive voice. Most of it warned the model away from the category, with no clause for the case where the initiating rank's output is absent from the log entirely, so the model resolved that ambiguity by declining: 3 of 28 live analyses returned category 0 at confidence 0 with a TCPStore observed failure, while category 12 was only ever picked when a primary failure had been identified. Say what to do when TCPStore is the only surface present - pick it, with reduced confidence - and keep the guard that a visible upstream fault wins. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both were written from the seed incident that produced them: cat 34 pinned a specific rank (608) and the GPTDataset index-load path, cat 35 a specific checkpoint (27600). Those name nothing a model can find in a fresh log. Keep every signature that actually discriminates - the wild-vs-near-null faulting address, the node-local blast radius - and drop the identifiers that only described one log. Category 29's row indices were already removed on main by 3e96f3a, so it is left as it stands. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Hand-cancelled jobs 4285548 and 4285549 both came back STOP under category 28, on the strength of "TypeError: 'NoneType' object is not callable". Their logs hold 751 SIGTERMs, 847 "Exception ignored" and healthy iterations to the last line before the signal; the grounded primary failure is Python's own weakref finalizer losing a module global during interpreter teardown. Python labels those Exception ignored precisely because they did not affect execution. Category 1 was the right answer and read, in full, "no application fault" - false on its face here, since a kill produces plenty of application-shaped noise. Give it the two surfaces a cancellation actually leaves: the bare slurmstepd marker, which is the whole content of some logs, and the progress-then-teardown pattern. Name the cat 2 discriminator so a node-failure reason still lands there. Exclude teardown from 28 on a general principle: a deterministic code fault recurs on unchanged retry, so it fails on the first attempt, not after 3390 iterations. Verified on five cancellations (4285548, 4285549, 4240576, 4189651, 4255225), all STOP -> RESTART, and on 4188567, whose kill followed a real dataset fault and correctly stayed STOP at category 17. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
| "id": 28, | ||
| "name": "Launch config or workload code deterministic error (TypeError/AttributeError)", | ||
| "description": "Deterministic runtime error inside already-loaded workload code that will recur on unchanged retry: TypeError/AttributeError on None value threaded through model code (e.g. float(None) in moe_layer postprocess), missing MASTER_ADDR, argparse error, or world-size mismatch. Excludes ModuleNotFoundError / ImportError where the failing import path resolves under Lustre or a mounted network filesystem - those are cat 16 (filesystem visibility), not a workload code fault, even if no adjacent FileNotFoundError is present on the same line.", | ||
| "description": "Deterministic runtime error inside already-loaded workload code that will recur on unchanged retry: TypeError/AttributeError on None value threaded through model code (e.g. float(None) in moe_layer postprocess), missing MASTER_ADDR, argparse error, or world-size mismatch. Excludes ModuleNotFoundError / ImportError where the failing import path resolves under Lustre or a mounted network filesystem - those are cat 16 (filesystem visibility), not a workload code fault, even if no adjacent FileNotFoundError is present on the same line. Also excludes anything raised during interpreter shutdown - inside \"Exception ignored\" / atexit callbacks, or appearing only after sustained successful iterations. A deterministic code fault recurs on an unchanged retry and so fails early and every time; one that appears only after thousands of iterations, at teardown, is the wake of an external kill (cat 1).", |
There was a problem hiding this comment.
Category 28 now excludes errors “appearing only after sustained successful iterations” and says deterministic faults fail early. A reproducible TypeError or AttributeError can first occur in a later evaluation step or a data-dependent model branch, without any shutdown or external kill. The new description excludes that error from its correct STOP category.
If the category-driven restart override is enabled and the model selects category 1 instead, the job can restart only to hit the same error. Limit the exclusion to errors shown to come from interpreter shutdown; successful earlier iterations do not rule out a deterministic workload bug.
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
Cat 12 was claiming logs where training had already run. The labeled cases it is meant for - b1-048/049/050 - all have zero completed iterations and fail during rendezvous; the production complaints all progressed first (iteration 48310, 6810, 12) and then saw TCPStore resets as survivors noticed a peer vanish. In 4303520 that peer was a worker an operator killed by hand, so the alert read "Distributed rendezvous / TCPStore socket timeout" for what was really someone running kill -9. Completed iterations settle it: a rendezvous failure cannot happen after iteration 48310, because rendezvous already succeeded. Gate cat 12 on no progress, and give cat 3 the general mid-training case - faulthandler dump, store resets, destroy_process_group-not-called at abrupt exit, atexit noise - with the cross-references that keep cat 1 (explicit scheduler marker) and cat 12 (never started) distinct. Both are RESTART, so no verdict changes; the explanation stops asserting a rendezvous fault that did not occur. Reported by the attribution channel thread, where the NVRx view was that "no primary failure" is the correct statement since the applog never sees the SIGKILL. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two cases from the attribution channel exposed descriptions that list symptoms without the condition that decides them. 4225797 stopped at iteration 4967 of 601014 with exit_interval unset, and came back cat 4 "clean planned exit" at confidence 96 - matched on a final checkpoint save, a training-energy summary and teardown warnings, all of which a graceful shutdown on signal also produces. Require positive evidence the run reached its configured end; last iteration against the schedule total settles it. 4224356 carries 6878 SIGTERMs and went to cat 3, whose whole premise is that the cause is not in the log. The cat 1 cross-reference was a trailing sentence against a description that advertises atexit noise as its own signature. Make the absence of any termination signal a requirement of cat 3 rather than a footnote. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ause" This reverts commit 0a94ce1. Neither gate worked as prompt text. 4225797 stayed on cat 4 at confidence 96 despite a requirement that the run reach its configured end, and 4224356 moved off cat 3 only to land on cat 6 (EXCLUDED) rather than cat 1. A trailing condition does not override a leading symptom list. The cost was real: cat 3 at least yields RESTART, which is the correct action for both. Trading a slightly-wrong category for a category with no useful action is the wrong trade. These two want a deterministic gate - last iteration against the schedule total, presence of SIGTERM - computed in evidence rather than asserted in a description. Left for later. The cat 3/12 split from 8590c43 stays; it is what fixes six of these eight. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Nine descriptions asserted the action as well as the evidence - "so product policy calls this a STOP", "external guidance is restart", "so this stays retryable" (the last one mine). The decision field already carries that, so the description was a second, unenforced copy that can drift from it, and it mixes two jobs: L1 classifies evidence, L4 decides what to do about it. Every claim about the nature of the failure is kept, since that is the evidence a classifier should weigh - an unchanged workload hitting the same allocation ceiling, a CUDA OOM reproducing at fixed config, a single log being unable to separate a library race from memory corruption. Only the verdict drawn from it is removed. Cats 7 and 8 still say a fault may "clear on an in-place restart"; that describes transience rather than prescribing an action, so it stays. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Stripping these alongside the policy verdicts was too blunt. "Restart from
last good checkpoint is plausibly safe" is not a verdict duplicating the
decision field; it says a known-good recovery point exists and resuming from
it will not carry the corruption forward - a property of the failure. It also
carries an operational distinction the replacement lost: roll back rather than
resume in place.
The line is between asserting the action ("product policy calls this a STOP")
and describing what recovery looks like. Only the former is removed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
4234179 c4 and 4307920 c2 both came back category 0 on a trace-export connection failure. Neither is a telemetry failure. 4234179 c4 runs normally for 84k lines to iteration 45250 of 615915, then the exporter starts failing, then worker processes are reaped leaving "resource_tracker: 27 leaked semaphore objects", then the log stops. The exporter losing its collector is an early symptom of the same shutdown, and the only error-shaped text in the file, so it got reported as the finding. Cat 3 already owns this story but listed none of these surfaces - it names faulthandler dumps, TCPStore resets, destroy_process_group and atexit, and neither log has any of them. Add the two that do appear, and say plainly that exporter errors are residue rather than fault. Described by behaviour rather than by endpoint: naming the collector's port would re-introduce the seed-incident specificity this branch removes from cats 29, 34 and 35. No new category: the failure is an external termination with the cause off-log, which is what cat 3 is. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
37a32a7 to
eb7be87
Compare
…rfaces" This reverts commit eb7be87. It did not work. 4234179 c4 and 4307920 c2 stayed on category 0 with the new surfaces named. The only thing it moved was 4225797, from cat 4 to cat 0 - a case the change does not describe - which is prompt coupling, not a fix. It also took cat 3 to 1125 characters against a 142 median, making it the longest in the taxonomy. An over-long description written around specific incidents is what made cat 12 misbehave in the first place. Third attempt at steering these by wording, after the cat 3 and cat 4 gates. They want a deterministic signal - exporter-only error surface, last iteration against the schedule total - computed in evidence rather than asserted in a description. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
Generated by human -- This PR is mainly to fix:
|
Five category descriptions in
l1/categories.json. No decision values change; no code changes.Each was found by reading a production misclassification back to the description that caused it.
What changed
rank 608and the GPTDataset index-load path.ckpt 27600.Why it mattered
Hand-cancelled jobs were being recommended STOP. Their logs hold 751 SIGTERMs, 847
Exception ignoredand healthy iterations to the last line before the signal; the grounded primary failure was Python's own weakref finalizer losing a module global during teardown — which Python labelsException ignoredprecisely because it did not affect execution.Separately, cascade-only TCPStore logs returned
category_id: 0("none matched") at confidence 0, so the alert told the reader nothing about a job that had cleanly checkpointed and then lost its store.Validation
0 → 12(intended), one0 → 1(a correction — that log has 6180 SIGTERMs and 3 TCPStore lines), one26 ↔ 32shown to be model variance by repeat trials on identical input, both STOP.MissingEpochIndexdataset fault (7104 occurrences, zero completed iterations), and the broadened cat 1 did not claim it.Note for reviewers
An earlier revision of cat 1 described only the progress-then-teardown shape and broke the canonical case — two labeled cases whose entire log is
*** STEP <id> ON <node> CANCELLED AT <time> ***drifted to cat 2. That is the same over-fitting this PR removes from cat 12, and it was caught only because the control set covered every edited category; both categories are RESTART, so headline accuracy was unchanged. Worth keeping that shape of control for future taxonomy edits.Two follow-ups not addressed here: cats 26 and 32 are not reliably distinguishable (demonstrated twice), and cats 1 and 3 both have a claim on a log containing only a cancellation marker.
🤖 Generated with Claude Code