You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Encountered an IMA in DeepEPv2 combine running on H200 connected via Infiniband on Coreweave.
The GIST is that DeepEP v2's EP barrier uses a deprecated API (ncclGin_SignalInc) issued on a single GIN context, while data Puts use the other contexts. Currently A "done" signal can overtake in-flight data, in which case the receiver can read stale buffer rows. The suggested fix is to migrate the barrier to ncclGinBarrier with ncclGinAllContexts / ncclGinFenceLevel::Put, which was added to NCCL to address this failure mode.
Below is generated by GLM 5.3 after a long investigation.
Summary
gin_barrier_wo_local_sync signals completion on context 0 while dispatch data
puts ride contexts 1..N. On the GDAKI transport, flush() completion means the
data WQEs have been ACKed ("issued toward the receiver's GPU memory"), not that
the payload is visible to the receiver's GPU reads, and there is no IB/PCIe ordering
between a context-0 signal and context-X puts (different QPs, possibly different
HCAs). The tiny 8-byte signal can therefore be performed at the receiver while bulk
posted writes are still in flight. The receiver's dispatch kernel exits, PDL launches
the copy epilogue, and the epilogue reads stale raw-buffer content — the previous
dispatch's rows.
The result is silent model-input corruption:
phantom routes to experts with count 0 (stale rows),
occasionally out-of-bounds slot indices, which surface as the combine_impl illegal-memory-access crashes seen in production nightlies.
We reproduced and root-caused this on pristine DeepEP with zero kernel
instrumentation (host-side audit only). All of DeepEP's count computation, prefix
construction, notify exchange, decoding, and route generation are correct on both
sides — the fault is purely the barrier's visibility guarantee.
Verified still present in V2.5 (93eb6eb, "DeepEP V2.5 (#763)"): see
"V2.5 status" below.
Expert 2 received 7 rows (rank 6's token-1 row missing)
Expert 7 received 1 row (phantom, counted 0)
Expert 20 received 4, but one via a stale row
The phantom landed on slot 8, double-assigned with expert 20's
exclusive-prefix cursor; one counted row (recv_x[7]) was never written
Decisive evidence that the epilogue read stale raw-buffer content: rank 6's two
received rows both carry token-0 row headers
(src_token_global = 12288 = 6·2048 + 0) in different lanes — one fresh (the
token-0 route) and one stale (rank 6's previous dispatch's token-0 route). A
second token-0 row is impossible for the current dispatch; rank 6's fresh token-1
put (expert 226, which would have filled expert 2's 8th slot) was not yet visible.
An earlier capture also showed partial-word visibility within a single row
(lane 5 fresh, lane 4 stale) — i.e., individual cache lines of one put commit out
of order, consistent with a multi-line posted-write burst distributed across GPU
L2 slices.
Root cause
Barrier structure (deep_ep/comm/barrier.cuh in V2.5; deep_ep/common/comm.cuh
at d4f41e4e), as invoked by the final dispatch barrier
(impls/ep/dispatch.cuh, kDispatchTag1, kFlushStores=true, kSyncAtStart=true):
All warps finish issuing data puts (contexts 1..N via get_qp_mode, kQPStartIdx = 1 when notify warps are present) and grid-sync.
Flush phase: warps sweep all contexts with flush(ncclCoopWarp()).
Grid sync.
SM 0 signals each peer on context 0 with ncclGin_SignalInc.
Receivers wait_signal on the context-0 signals table, the kernel exits,
PDL launches the copy epilogue, which reads the raw receive buffer.
The barrier's bookkeeping is sound (all puts are doorbelled before any flush warp
reads sq_rsvd_index). The problem is what flush completion proves. From the
NCCL 2.30.x gin.h contract:
flush(): "Flush does not guarantee that data has settled in remote memory."
Strong signals cover same-context puts only.
ncclGin_SignalInc (which the barrier uses) is deprecated, and on GDAKI isStrong is not read by the backend at all — strong and weak signals are
byte-for-byte identical; the only ordering a signal ever gets is same-QP WQE
order.
And at the transport level, GDAKI data puts are plain RDMA_WRITE WQEs while any
signal is an 8-byte ATOMIC_FA. For RC writes into GPUDirect memory, the
responder HCA ACKs once it has issued the PCIe posted write toward the GPU —
posted writes have no completion signal, so:
CQE ⇒ "write issued toward GPU memory", not "visible to GPU reads".
The signal atomic must be performed at the target before completing; the data
burst of posted writes needs no such settlement.
Signal (context 0) and data (contexts 1..N) are different QPs, plausibly
different HCAs → no IB- or PCIe-level ordering between them.
So "flush completes → grid sync → signal posted" only proves the data left the
senders' NICs. The signal can be performed at the receiver while the data is
still in flight, and the epilogue reads the previous dispatch's rows still
standing at those addresses — which is exactly what makes this dangerous: the
stale content is valid-looking routing data.
The window is normally tiny relative to kernel-launch/PDL overhead, which is why
this only fires rarely and preferentially on small decode/warmup batches (minimal
data, short epilogue racing immediately behind the barrier, tight dispatch cadence
reusing the raw buffer with the previous layer's rows resident). Every observed
fault was a 1–2-token-per-rank batch.
NCCL already fixed this exact bug in its own barrier
The strongest evidence is on the record from NVIDIA themselves:
NCCL issue #2302 ("[Question/RFE]
Cross-context ordering before a single GIN signal", July 2026): a user asked
exactly the question DeepEP's barrier depends on — "Is there an NCCL GIN
operation equivalent to nvshmem_fence() for establishing ordering across
multiple GIN contexts before sending a single signal to one peer?" — and
NVIDIA's answer was unambiguous:
"There are no ordering guarantees across queue pairs. For most GIN backends,
each context has one queue pair per peer. Hence, there is no ordering across
contexts."
"Is there an existing one-sided GIN primitive that orders puts across multiple
contexts before one subsequent signal to the same peer? — No, there is not."
NVIDIA's recommended pattern for this use case was one signal per context with
fan-in at the target — i.e., precisely what DeepEP's barrier fails to do, and
precisely what ncclGinBarrier with ncclGinAllContexts now implements.
This confirms DeepEP's barrier pattern (puts on contexts 1..N, single
context-0 signal) has no ordering guarantee at the transport level.
The v2.30.7-1 release notes
list, under "GIN Enhancements": "Adds explicit signal semantics with Strong and
Weak signals" and "Adds proper ncclGinFenceLevel semantics for barriers" — the
same release line as our build. The same notes also add "logic to gin.flush to
ensure all prior gets are visible" — notably gets only, reinforcing that
flush has never settled remote puts.
"Signal on each context, not just context 0: signals and puts on different QPs
are not ordered at the receiving NIC, so a ctx-0 signal could overtake an
in-flight ctx-X put. Each context has its own signal memory and shadow slot for
the same signal id, so no extra slot allocation is needed."
The fence-all-contexts switch — gin_barrier.h#L43-L46
(ncclGinAllContexts), and the visibility contract DeepEP needs — gin_barrier.h#L26-L31
(ncclGinFenceLevel::Put: "After the barrier returns, puts issued by other team
members targeting the calling rank prior to the barrier are visible in the
calling rank's memory")
The deprecation of what DeepEP's barrier uses — gin.h#L87
("Deprecated: use ncclGin_StrongSignalInc or ncclGin_WeakSignalInc explicitly")
— and the flush disclaimer the barrier rests on — gin.h#L310-L312
("Flush does not guarantee that data has settled in remote memory")
NCCL 2.30.x ships ncclGinBarrierSession / ncclGinBarrier with the ncclGinAllContexts fence-all-contexts mode, whose implementation comments on
precisely this failure (quoted above). Notably NCCL's fixed barrier still uses SignalInc — the fix is purely signal on every context, i.e., same-QP WQE
ordering is the operative mechanism, not signal "strength".
This is independent confirmation of the root cause, on the NCCL side, in the same
2.30.x line DeepEP already requires.
V2.5 status (verified from source, 93eb6eb)
The race is not fixed in V2.5:
deep_ep/comm/barrier.cuh:137 — gin_barrier_wo_local_sync still constructs a
single context-0 handle and signals every peer on it with ncclGin_SignalInc.
Data puts still ride contexts ≥ 1; notify warps still share QP 0 with the
barrier signal.
The only barrier changes since d4f41e4e are: (a) migration of the wait loop
to the official gin.waitSignal API with timeout + shadow counters (resolves
an old TODO(NCCL) — diagnostics/robustness only; waitSignal waits for the
signal value, it does not add data-settlement semantics), and (b) fix: add system-scope release before the GIN barrier when scaleup spans NVLink and RDMA #715's fence_acq_rel_sys() before world-team barriers, which fixes a different
race (NVLink writes vs GIN signal ordering for mixed-fabric scaleup) and cannot
order remote NIC-posted writes.
ncclGinBarrier / ncclGinBarrierSession / ncclGinFenceLevel / per-context
signals appear nowhere in the tree. Gen-2 signal types
(ncclGin_StrongVASignalAdd etc.) are used only in the new bucket/PP
collectives, not the EP dispatch/combine barriers.
The copy epilogue (impls/ep/dispatch_copy_epilogue.cuh) still relies solely
on cudaGridDependencySynchronize() and performs no generation/staleness
validation of row headers.
Suggested fixes (in order of preference)
Port NCCL's per-context signal pattern into gin_barrier_wo_local_sync
(self-contained, no host changes): after flush + grid sync, signal every peer
on every context and wait on the per-(context, signal-id) slots — each
context has its own signals table and shadow slot, so no extra allocation is
needed. A signal posted on context c is SQ-FIFO-ordered behind all WQEs
previously reserved on context c, covering the data puts. Cost: ginContextCount × kNumRanks 8-byte FA signals per barrier. Use the runtime ginContextCount (not per-call-site kNumQPs) so every barrier invocation
signals and waits on the same set of contexts, keeping shadow counters
consistent across call sites.
Migrate to ncclGinBarrier with ncclGinAllContexts — maintained fix
with documented semantics; needs host-side barrier-handle setup
(ncclGinBarrierCreateRequirement) and restructuring of the SM-0 /
scaleup-scaleout-parallel barrier drivers.
Defense-in-depth: generation/sequence counters in row headers, validated by
the copy epilogue — fail fast, not retry. Retry is not implementable (once
the barrier passes, senders have moved on), but this converts silent
corruption into a clean trap and depends on no transport semantics at all.
Given that the same symptom class (receiver reads before data settles) shows
up on other transports (see related issues below), this may be worth having
regardless.
Switching to ncclGin_StrongVASignalInc would not fix this: strong signals
cover same-context puts only, and GDAKI ignores isStrong regardless.
Config: DeepSeek-V4-Flash, DP=8 direct mode (hybrid_mode=0), --enforce-eager,
fp8 KV cache, 3× H200 nodes (any 4+2+2 split; not node-specific).
Method: host-side sitecustomize.py monkeypatch (zero kernel instrumentation)
recording per-dispatch host counts, device cursors, and recv_src_metadata
per rank, with a fault trigger on received-vs-counted lane contradiction.
Hit rate: once during warmup of a single pristine run; six guarded captures and
one stock-image warmup crash in earlier runs. Rare but reproducible within
tens of minutes of server start under warmup churn.
We have a full artifact set (per-dispatch audit ring buffers, the faulting
dispatch's row-by-row decode with the impossible token-0 header, CUDA cores,
NCCL probe, environment captures) and a standalone report with the complete
evidence chain — happy to attach or link on request.
This report was prepared with AI assistance; all evidence was produced by the
runs and artifacts referenced above and has been reviewed end-to-end.
Encountered an IMA in DeepEPv2 combine running on H200 connected via Infiniband on Coreweave.
The GIST is that DeepEP v2's EP barrier uses a deprecated API (
ncclGin_SignalInc) issued on a single GIN context, while data Puts use the other contexts. Currently A "done" signal can overtake in-flight data, in which case the receiver can read stale buffer rows. The suggested fix is to migrate the barrier toncclGinBarrierwithncclGinAllContexts/ncclGinFenceLevel::Put, which was added to NCCL to address this failure mode.Below is generated by GLM 5.3 after a long investigation.
Summary
gin_barrier_wo_local_syncsignals completion on context 0 while dispatch dataputs ride contexts 1..N. On the GDAKI transport,
flush()completion means thedata WQEs have been ACKed ("issued toward the receiver's GPU memory"), not that
the payload is visible to the receiver's GPU reads, and there is no IB/PCIe ordering
between a context-0 signal and context-X puts (different QPs, possibly different
HCAs). The tiny 8-byte signal can therefore be performed at the receiver while bulk
posted writes are still in flight. The receiver's dispatch kernel exits, PDL launches
the copy epilogue, and the epilogue reads stale raw-buffer content — the previous
dispatch's rows.
The result is silent model-input corruption:
recv_xrows,combine_implillegal-memory-access crashes seen in production nightlies.We reproduced and root-caused this on pristine DeepEP with zero kernel
instrumentation (host-side audit only). All of DeepEP's count computation, prefix
construction, notify exchange, decoding, and route generation are correct on both
sides — the fault is purely the barrier's visibility guarantee.
Verified still present in V2.5 (
93eb6eb, "DeepEP V2.5 (#763)"): see"V2.5 status" below.
Environment
d4f41e4e(2.0.0) — audited build; V2.593eb6eb— re-verified from sourceGIN_IB_GDAKItransport)36768d1b, DeepSeek-V4-Flash, fp8 KV cache, direct mode (hybrid_mode=0),--enforce-eager(every dispatchdo_expand=True,do_cpu_sync=True)GPU Direct enabled), DP=8 arranged 4+2+2
decode batches
The fault (pristine build, host-side audit only)
Rank 7, dispatch sequence 56, warmup, 2 tokens/rank. Host-side counts consistent
across all ranks (send == recv on every sequence-matched dispatch):
Device state after the dispatch barrier:
exclusive-prefix cursor; one counted row (
recv_x[7]) was never writtenDecisive evidence that the epilogue read stale raw-buffer content: rank 6's two
received rows both carry token-0 row headers
(
src_token_global = 12288 = 6·2048 + 0) in different lanes — one fresh (thetoken-0 route) and one stale (rank 6's previous dispatch's token-0 route). A
second token-0 row is impossible for the current dispatch; rank 6's fresh token-1
put (expert 226, which would have filled expert 2's 8th slot) was not yet visible.
An earlier capture also showed partial-word visibility within a single row
(lane 5 fresh, lane 4 stale) — i.e., individual cache lines of one put commit out
of order, consistent with a multi-line posted-write burst distributed across GPU
L2 slices.
Root cause
Barrier structure (
deep_ep/comm/barrier.cuhin V2.5;deep_ep/common/comm.cuhat
d4f41e4e), as invoked by the final dispatch barrier(
impls/ep/dispatch.cuh,kDispatchTag1, kFlushStores=true, kSyncAtStart=true):get_qp_mode,kQPStartIdx = 1when notify warps are present) and grid-sync.flush(ncclCoopWarp()).ncclGin_SignalInc.wait_signalon the context-0 signals table, the kernel exits,PDL launches the copy epilogue, which reads the raw receive buffer.
The barrier's bookkeeping is sound (all puts are doorbelled before any flush warp
reads
sq_rsvd_index). The problem is what flush completion proves. From theNCCL 2.30.x
gin.hcontract:flush(): "Flush does not guarantee that data has settled in remote memory."ncclGin_SignalInc(which the barrier uses) is deprecated, and on GDAKIisStrongis not read by the backend at all — strong and weak signals arebyte-for-byte identical; the only ordering a signal ever gets is same-QP WQE
order.
And at the transport level, GDAKI data puts are plain
RDMA_WRITEWQEs while anysignal is an 8-byte
ATOMIC_FA. For RC writes into GPUDirect memory, theresponder HCA ACKs once it has issued the PCIe posted write toward the GPU —
posted writes have no completion signal, so:
burst of posted writes needs no such settlement.
different HCAs → no IB- or PCIe-level ordering between them.
So "flush completes → grid sync → signal posted" only proves the data left the
senders' NICs. The signal can be performed at the receiver while the data is
still in flight, and the epilogue reads the previous dispatch's rows still
standing at those addresses — which is exactly what makes this dangerous: the
stale content is valid-looking routing data.
The window is normally tiny relative to kernel-launch/PDL overhead, which is why
this only fires rarely and preferentially on small decode/warmup batches (minimal
data, short epilogue racing immediately behind the barrier, tight dispatch cadence
reusing the raw buffer with the previous layer's rows resident). Every observed
fault was a 1–2-token-per-rank batch.
NCCL already fixed this exact bug in its own barrier
The strongest evidence is on the record from NVIDIA themselves:
NCCL issue #2302 ("[Question/RFE]
Cross-context ordering before a single GIN signal", July 2026): a user asked
exactly the question DeepEP's barrier depends on — "Is there an NCCL GIN
operation equivalent to
nvshmem_fence()for establishing ordering acrossmultiple GIN contexts before sending a single signal to one peer?" — and
NVIDIA's answer was unambiguous:
NVIDIA's recommended pattern for this use case was one signal per context with
fan-in at the target — i.e., precisely what DeepEP's barrier fails to do, and
precisely what
ncclGinBarrierwithncclGinAllContextsnow implements.This confirms DeepEP's barrier pattern (puts on contexts 1..N, single
context-0 signal) has no ordering guarantee at the transport level.
The v2.30.7-1 release notes
list, under "GIN Enhancements": "Adds explicit signal semantics with Strong and
Weak signals" and "Adds proper
ncclGinFenceLevelsemantics for barriers" — thesame release line as our build. The same notes also add "logic to
gin.flushtoensure all prior gets are visible" — notably gets only, reinforcing that
flush has never settled remote puts.
Source permalinks at the
v2.30.7-1tag:impl/gin_barrier__funcs.h#L129-L135:gin_barrier.h#L43-L46(
ncclGinAllContexts), and the visibility contract DeepEP needs —gin_barrier.h#L26-L31(
ncclGinFenceLevel::Put: "After the barrier returns, puts issued by other teammembers targeting the calling rank prior to the barrier are visible in the
calling rank's memory")
gin.h#L87("Deprecated: use ncclGin_StrongSignalInc or ncclGin_WeakSignalInc explicitly")
— and the flush disclaimer the barrier rests on —
gin.h#L310-L312("Flush does not guarantee that data has settled in remote memory")
NCCL 2.30.x ships
ncclGinBarrierSession/ncclGinBarrierwith thencclGinAllContextsfence-all-contexts mode, whose implementation comments onprecisely this failure (quoted above). Notably NCCL's fixed barrier still uses
SignalInc— the fix is purely signal on every context, i.e., same-QP WQEordering is the operative mechanism, not signal "strength".
This is independent confirmation of the root cause, on the NCCL side, in the same
2.30.x line DeepEP already requires.
V2.5 status (verified from source,
93eb6eb)The race is not fixed in V2.5:
deep_ep/comm/barrier.cuh:137—gin_barrier_wo_local_syncstill constructs asingle context-0 handle and signals every peer on it with
ncclGin_SignalInc.barrier signal.
d4f41e4eare: (a) migration of the wait loopto the official
gin.waitSignalAPI with timeout + shadow counters (resolvesan old
TODO(NCCL)— diagnostics/robustness only;waitSignalwaits for thesignal value, it does not add data-settlement semantics), and (b) fix: add system-scope release before the GIN barrier when scaleup spans NVLink and RDMA #715's
fence_acq_rel_sys()before world-team barriers, which fixes a differentrace (NVLink writes vs GIN signal ordering for mixed-fabric scaleup) and cannot
order remote NIC-posted writes.
ncclGinBarrier/ncclGinBarrierSession/ncclGinFenceLevel/ per-contextsignals appear nowhere in the tree. Gen-2 signal types
(
ncclGin_StrongVASignalAddetc.) are used only in the new bucket/PPcollectives, not the EP dispatch/combine barriers.
impls/ep/dispatch_copy_epilogue.cuh) still relies solelyon
cudaGridDependencySynchronize()and performs no generation/stalenessvalidation of row headers.
Suggested fixes (in order of preference)
gin_barrier_wo_local_sync(self-contained, no host changes): after flush + grid sync, signal every peer
on every context and wait on the per-(context, signal-id) slots — each
context has its own signals table and shadow slot, so no extra allocation is
needed. A signal posted on context
cis SQ-FIFO-ordered behind all WQEspreviously reserved on context
c, covering the data puts. Cost:ginContextCount × kNumRanks8-byte FA signals per barrier. Use the runtimeginContextCount(not per-call-sitekNumQPs) so every barrier invocationsignals and waits on the same set of contexts, keeping shadow counters
consistent across call sites.
ncclGinBarrierwithncclGinAllContexts— maintained fixwith documented semantics; needs host-side barrier-handle setup
(
ncclGinBarrierCreateRequirement) and restructuring of the SM-0 /scaleup-scaleout-parallel barrier drivers.
the copy epilogue — fail fast, not retry. Retry is not implementable (once
the barrier passes, senders have moved on), but this converts silent
corruption into a clean trap and depends on no transport semantics at all.
Given that the same symptom class (receiver reads before data settles) shows
up on other transports (see related issues below), this may be worth having
regardless.
Switching to
ncclGin_StrongVASignalIncwould not fix this: strong signalscover same-context puts only, and GDAKI ignores
isStrongregardless.Related issues
Possibly the same symptom class (stale data → OOB slot indices) on a
different transport; still open.
fixed by Add fence.proxy.async.shared::cta between mbarrier wait and TMA load. #642, which is unrelated to this race.
direction of this race).
Reproduction
hybrid_mode=0),--enforce-eager,fp8 KV cache, 3× H200 nodes (any 4+2+2 split; not node-specific).
sitecustomize.pymonkeypatch (zero kernel instrumentation)recording per-dispatch host counts, device cursors, and
recv_src_metadataper rank, with a fault trigger on received-vs-counted lane contradiction.
one stock-image warmup crash in earlier runs. Rare but reproducible within
tens of minutes of server start under warmup churn.
We have a full artifact set (per-dispatch audit ring buffers, the faulting
dispatch's row-by-row decode with the impossible token-0 header, CUDA cores,
NCCL probe, environment captures) and a standalone report with the complete
evidence chain — happy to attach or link on request.
This report was prepared with AI assistance; all evidence was produced by the
runs and artifacts referenced above and has been reviewed end-to-end.