Skip to content

[Bug] Data-visibility race in V2 low-latency dispatch over NCCL GIN (GDAKI) #772

Description

@tlrmchlsmth

Encountered an IMA in DeepEPv2 combine running on H200 connected via Infiniband on Coreweave.

The GIST is that DeepEP v2's EP barrier uses a deprecated API (ncclGin_SignalInc) issued on a single GIN context, while data Puts use the other contexts. Currently A "done" signal can overtake in-flight data, in which case the receiver can read stale buffer rows. The suggested fix is to migrate the barrier to ncclGinBarrier with ncclGinAllContexts / ncclGinFenceLevel::Put, which was added to NCCL to address this failure mode.

Below is generated by GLM 5.3 after a long investigation.

Summary

gin_barrier_wo_local_sync signals completion on context 0 while dispatch data
puts ride contexts 1..N. On the GDAKI transport, flush() completion means the
data WQEs have been ACKed ("issued toward the receiver's GPU memory"), not that
the payload is visible to the receiver's GPU reads, and there is no IB/PCIe ordering
between a context-0 signal and context-X puts (different QPs, possibly different
HCAs). The tiny 8-byte signal can therefore be performed at the receiver while bulk
posted writes are still in flight. The receiver's dispatch kernel exits, PDL launches
the copy epilogue, and the epilogue reads stale raw-buffer content — the previous
dispatch's rows
.

The result is silent model-input corruption:

  • phantom routes to experts with count 0 (stale rows),
  • missing routes for experts with nonzero counts,
  • double-assigned slots (two experts' cursors collide),
  • unwritten recv_x rows,
  • occasionally out-of-bounds slot indices, which surface as the
    combine_impl illegal-memory-access crashes seen in production nightlies.

We reproduced and root-caused this on pristine DeepEP with zero kernel
instrumentation
(host-side audit only). All of DeepEP's count computation, prefix
construction, notify exchange, decoding, and route generation are correct on both
sides — the fault is purely the barrier's visibility guarantee.

Verified still present in V2.5 (93eb6eb, "DeepEP V2.5 (#763)"): see
"V2.5 status" below.

Environment

  • DeepEP d4f41e4e (2.0.0) — audited build; V2.5 93eb6eb — re-verified from source
  • NCCL 2.30.7 (internal build, GIN_IB_GDAKI transport)
  • vLLM 36768d1b, DeepSeek-V4-Flash, fp8 KV cache, direct mode (hybrid_mode=0),
    --enforce-eager (every dispatch do_expand=True, do_cpu_sync=True)
  • 3× H200 nodes (XE9680-class, 8× GPU-direct-RDMA IB HCAs visible, ibp0–ibp3
    GPU Direct enabled), DP=8 arranged 4+2+2
  • Fault observed during warmup (~3.5 min after server start), 1–2-token-per-rank
    decode batches

The fault (pristine build, host-side audit only)

Rank 7, dispatch sequence 56, warmup, 2 tokens/rank. Host-side counts consistent
across all ranks (send == recv on every sequence-matched dispatch):

  • Expert 2 (global 226): count 8 — every rank's token 1, lane 3
  • Expert 20 (global 244): count 4 — ranks 0/3/4/6's token 0, lane 5
  • Expert 7 (global 231): count 0

Device state after the dispatch barrier:

  • Expert 2 received 7 rows (rank 6's token-1 row missing)
  • Expert 7 received 1 row (phantom, counted 0)
  • Expert 20 received 4, but one via a stale row
  • The phantom landed on slot 8, double-assigned with expert 20's
    exclusive-prefix cursor; one counted row (recv_x[7]) was never written

Decisive evidence that the epilogue read stale raw-buffer content: rank 6's two
received rows both carry token-0 row headers
(src_token_global = 12288 = 6·2048 + 0) in different lanes — one fresh (the
token-0 route) and one stale (rank 6's previous dispatch's token-0 route). A
second token-0 row is impossible for the current dispatch; rank 6's fresh token-1
put (expert 226, which would have filled expert 2's 8th slot) was not yet visible.

An earlier capture also showed partial-word visibility within a single row
(lane 5 fresh, lane 4 stale) — i.e., individual cache lines of one put commit out
of order, consistent with a multi-line posted-write burst distributed across GPU
L2 slices.

Root cause

Barrier structure (deep_ep/comm/barrier.cuh in V2.5; deep_ep/common/comm.cuh
at d4f41e4e), as invoked by the final dispatch barrier
(impls/ep/dispatch.cuh, kDispatchTag1, kFlushStores=true, kSyncAtStart=true):

  1. All warps finish issuing data puts (contexts 1..N via get_qp_mode,
    kQPStartIdx = 1 when notify warps are present) and grid-sync.
  2. Flush phase: warps sweep all contexts with flush(ncclCoopWarp()).
  3. Grid sync.
  4. SM 0 signals each peer on context 0 with ncclGin_SignalInc.
  5. Receivers wait_signal on the context-0 signals table, the kernel exits,
    PDL launches the copy epilogue, which reads the raw receive buffer.

The barrier's bookkeeping is sound (all puts are doorbelled before any flush warp
reads sq_rsvd_index). The problem is what flush completion proves. From the
NCCL 2.30.x gin.h contract:

  • flush(): "Flush does not guarantee that data has settled in remote memory."
  • Strong signals cover same-context puts only.
  • ncclGin_SignalInc (which the barrier uses) is deprecated, and on GDAKI
    isStrong is not read by the backend at all — strong and weak signals are
    byte-for-byte identical; the only ordering a signal ever gets is same-QP WQE
    order.

And at the transport level, GDAKI data puts are plain RDMA_WRITE WQEs while any
signal is an 8-byte ATOMIC_FA. For RC writes into GPUDirect memory, the
responder HCA ACKs once it has issued the PCIe posted write toward the GPU —
posted writes have no completion signal, so:

  • CQE ⇒ "write issued toward GPU memory", not "visible to GPU reads".
  • The signal atomic must be performed at the target before completing; the data
    burst of posted writes needs no such settlement.
  • Signal (context 0) and data (contexts 1..N) are different QPs, plausibly
    different HCAs → no IB- or PCIe-level ordering between them.

So "flush completes → grid sync → signal posted" only proves the data left the
senders' NICs
. The signal can be performed at the receiver while the data is
still in flight, and the epilogue reads the previous dispatch's rows still
standing at those addresses — which is exactly what makes this dangerous: the
stale content is valid-looking routing data.

The window is normally tiny relative to kernel-launch/PDL overhead, which is why
this only fires rarely and preferentially on small decode/warmup batches (minimal
data, short epilogue racing immediately behind the barrier, tight dispatch cadence
reusing the raw buffer with the previous layer's rows resident). Every observed
fault was a 1–2-token-per-rank batch.

NCCL already fixed this exact bug in its own barrier

The strongest evidence is on the record from NVIDIA themselves:

  1. NCCL issue #2302 ("[Question/RFE]
    Cross-context ordering before a single GIN signal", July 2026): a user asked
    exactly the question DeepEP's barrier depends on — "Is there an NCCL GIN
    operation equivalent to nvshmem_fence() for establishing ordering across
    multiple GIN contexts before sending a single signal to one peer?"
    — and
    NVIDIA's answer was unambiguous:

    "There are no ordering guarantees across queue pairs. For most GIN backends,
    each context has one queue pair per peer. Hence, there is no ordering across
    contexts."

    "Is there an existing one-sided GIN primitive that orders puts across multiple
    contexts before one subsequent signal to the same peer? — No, there is not."

    NVIDIA's recommended pattern for this use case was one signal per context with
    fan-in at the target
    — i.e., precisely what DeepEP's barrier fails to do, and
    precisely what ncclGinBarrier with ncclGinAllContexts now implements.
    This confirms DeepEP's barrier pattern (puts on contexts 1..N, single
    context-0 signal) has no ordering guarantee at the transport level.

  2. The v2.30.7-1 release notes
    list, under "GIN Enhancements": "Adds explicit signal semantics with Strong and
    Weak signals" and "Adds proper ncclGinFenceLevel semantics for barriers" — the
    same release line as our build. The same notes also add "logic to gin.flush to
    ensure all prior gets are visible" — notably gets only, reinforcing that
    flush has never settled remote puts.

  3. Source permalinks at the v2.30.7-1 tag:

    • The multi-context signal fix and its rationale —
      impl/gin_barrier__funcs.h#L129-L135:

      "Signal on each context, not just context 0: signals and puts on different QPs
      are not ordered at the receiving NIC, so a ctx-0 signal could overtake an
      in-flight ctx-X put. Each context has its own signal memory and shadow slot for
      the same signal id, so no extra slot allocation is needed."

    • The fence-all-contexts switch —
      gin_barrier.h#L43-L46
      (ncclGinAllContexts), and the visibility contract DeepEP needs —
      gin_barrier.h#L26-L31
      (ncclGinFenceLevel::Put: "After the barrier returns, puts issued by other team
      members targeting the calling rank prior to the barrier are visible in the
      calling rank's memory")
    • The deprecation of what DeepEP's barrier uses —
      gin.h#L87
      ("Deprecated: use ncclGin_StrongSignalInc or ncclGin_WeakSignalInc explicitly")
      — and the flush disclaimer the barrier rests on —
      gin.h#L310-L312
      ("Flush does not guarantee that data has settled in remote memory")

NCCL 2.30.x ships ncclGinBarrierSession / ncclGinBarrier with the
ncclGinAllContexts fence-all-contexts mode, whose implementation comments on
precisely this failure (quoted above). Notably NCCL's fixed barrier still uses
SignalInc — the fix is purely signal on every context, i.e., same-QP WQE
ordering is the operative mechanism, not signal "strength".

This is independent confirmation of the root cause, on the NCCL side, in the same
2.30.x line DeepEP already requires.

V2.5 status (verified from source, 93eb6eb)

The race is not fixed in V2.5:

  • deep_ep/comm/barrier.cuh:137 — gin_barrier_wo_local_sync still constructs a
    single context-0 handle and signals every peer on it with ncclGin_SignalInc.
  • Data puts still ride contexts ≥ 1; notify warps still share QP 0 with the
    barrier signal.
  • The only barrier changes since d4f41e4e are: (a) migration of the wait loop
    to the official gin.waitSignal API with timeout + shadow counters (resolves
    an old TODO(NCCL) — diagnostics/robustness only; waitSignal waits for the
    signal value, it does not add data-settlement semantics), and (b) fix: add system-scope release before the GIN barrier when scaleup spans NVLink and RDMA #715's
    fence_acq_rel_sys() before world-team barriers, which fixes a different
    race (NVLink writes vs GIN signal ordering for mixed-fabric scaleup) and cannot
    order remote NIC-posted writes.
  • ncclGinBarrier / ncclGinBarrierSession / ncclGinFenceLevel / per-context
    signals appear nowhere in the tree. Gen-2 signal types
    (ncclGin_StrongVASignalAdd etc.) are used only in the new bucket/PP
    collectives, not the EP dispatch/combine barriers.
  • The copy epilogue (impls/ep/dispatch_copy_epilogue.cuh) still relies solely
    on cudaGridDependencySynchronize() and performs no generation/staleness
    validation of row headers.

Suggested fixes (in order of preference)

  1. Port NCCL's per-context signal pattern into gin_barrier_wo_local_sync
    (self-contained, no host changes): after flush + grid sync, signal every peer
    on every context and wait on the per-(context, signal-id) slots — each
    context has its own signals table and shadow slot, so no extra allocation is
    needed. A signal posted on context c is SQ-FIFO-ordered behind all WQEs
    previously reserved on context c, covering the data puts. Cost:
    ginContextCount × kNumRanks 8-byte FA signals per barrier. Use the runtime
    ginContextCount (not per-call-site kNumQPs) so every barrier invocation
    signals and waits on the same set of contexts, keeping shadow counters
    consistent across call sites.
  2. Migrate to ncclGinBarrier with ncclGinAllContexts — maintained fix
    with documented semantics; needs host-side barrier-handle setup
    (ncclGinBarrierCreateRequirement) and restructuring of the SM-0 /
    scaleup-scaleout-parallel barrier drivers.
  3. Defense-in-depth: generation/sequence counters in row headers, validated by
    the copy epilogue — fail fast, not retry.
    Retry is not implementable (once
    the barrier passes, senders have moved on), but this converts silent
    corruption into a clean trap and depends on no transport semantics at all.
    Given that the same symptom class (receiver reads before data settles) shows
    up on other transports (see related issues below), this may be worth having
    regardless.

Switching to ncclGin_StrongVASignalInc would not fix this: strong signals
cover same-context puts only, and GDAKI ignores isStrong regardless.

Related issues

Reproduction

  • Config: DeepSeek-V4-Flash, DP=8 direct mode (hybrid_mode=0), --enforce-eager,
    fp8 KV cache, 3× H200 nodes (any 4+2+2 split; not node-specific).
  • Method: host-side sitecustomize.py monkeypatch (zero kernel instrumentation)
    recording per-dispatch host counts, device cursors, and recv_src_metadata
    per rank, with a fault trigger on received-vs-counted lane contradiction.
  • Hit rate: once during warmup of a single pristine run; six guarded captures and
    one stock-image warmup crash in earlier runs. Rare but reproducible within
    tens of minutes of server start under warmup churn.

We have a full artifact set (per-dispatch audit ring buffers, the faulting
dispatch's row-by-row decode with the impossible token-0 header, CUDA cores,
NCCL probe, environment captures) and a standalone report with the complete
evidence chain — happy to attach or link on request.


This report was prepared with AI assistance; all evidence was produced by the
runs and artifacts referenced above and has been reviewed end-to-end.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions