Add DeepEP v2 flex dispatcher backend - #5153
Autumn1998 wants to merge 17 commits into
Conversation
|
This PR has been automatically converted to draft because all PRs must start as drafts. When you are ready for review, click Ready for Review to begin the review process. This will:
See the contribution guide for more details. |
|
/ok to test 646cceb |
|
/ok to test c0a73a0 |
zhongbozhu
left a comment
There was a problem hiding this comment.
Please add test coverage to this file: https://github.com/NVIDIA/Megatron-LM/blob/main/tests/unit_tests/transformer/moe/test_grouped_tensor_dispatcher_numerics.py
| HAVE_DEEP_EP = False | ||
|
|
||
| try: | ||
| from deep_ep import ElasticBuffer |
There was a problem hiding this comment.
Is there anyway we can leave some instructions for people who want to use deepepv2 about how to update their container and its potential impact like HybridEP might need to be uninstalled?
There was a problem hiding this comment.
Added the installation instructions and the note about potentially uninstalling HybridEP here in the missing-DeepEP-v2 error message
(
Megatron-LM/megatron/core/transformer/moe/token_dispatcher.py
Lines 1511 to 1518 in 4e37802
guihong-nv
left a comment
There was a problem hiding this comment.
The NCCL warmup and input-lifetime fixes address my earlier comments. I left one new comment about a race in backward. The uneven-token-count issue in the existing thread still needs to be handled. I couldn't run the multi-GPU tests here.
| previous_event = ( | ||
| ctx.buffer.capture() if ctx.async_finish and ctx.allocate_on_comm_stream else None | ||
| ) | ||
| grad_x, _, _, _, event = ctx.buffer.dispatch( | ||
| grad_output.contiguous(), |
There was a problem hiding this comment.
Can we make grad_output contiguous before capturing previous_event, in both backward methods? With the default async settings, a non-contiguous gradient queues its copy after the event, but DeepEP only waits for that event before reading the copied tensor. The communication stream can then read unfinished data and produce wrong gradients. Moving the input conversions before capture would close this race.
Signed-off-by: tongliu <tongliu@nvidia.com>
Signed-off-by: tongliu <tongliu@nvidia.com>
Signed-off-by: tongliu <tongliu@nvidia.com>
64b8723 to
4e37802
Compare
|
could support deepep v2.5:deepseek-ai/DeepEP#763 |
|
/claude review |
| raise ValueError( | ||
| f"Invalid backend: {self.config.moe_flex_dispatcher_backend}" | ||
| "Please set --moe-flex-dispatcher-backend to deepep, hybridep, or ncclep" | ||
| "Please set --moe-flex-dispatcher-backend to deepep, deepepv2, hybridep, or ncclep" |
There was a problem hiding this comment.
Hi @Autumn1998 , do we have perf comparison between deepepv2 and our ncclep on any popular models?
There was a problem hiding this comment.
It would be nice if we have some shared public comparison across these different dispatcher-backend. So that it's easier for us to make a choice between them
|
/ok to test 4e37802 |
Light review —
|
| changed source | manifest entry | required test |
|---|---|---|
megatron/core/transformer/moe/fused_a2a.py |
moe_fused_a2a (kind="external-lib") |
tests/unit_tests/determinism/kernels/test_moe_kernels.py |
megatron/core/transformer/moe/token_dispatcher.py |
moe_token_dispatchers (kind="torch-op") |
tests/unit_tests/determinism/kernels/test_moe_kernels.py |
test_moe_kernels.py is not in this PR's diff, and the PR carries only complexity: medium / Approved — no determinism-exempt. So check() emits a violation per source ("kernel … changed but none of its determinism tests did").
This is also substantively warranted, not just a gate: ElasticBuffer.dispatch / .combine is a brand-new external-library reduction over the EP group, exactly the \b\w*buffer\.(low_latency_)?(dispatch|combine)\s*\( pattern the registry tracks. The existing flex-deepep-ep2 cell covers v1 only.
Adding a sibling cell in test_moe_kernels.py::test_moe_layer_replays, immediately after the existing flex-deepep-ep2 entry in its @pytest.mark.parametrize("dispatcher,ep,extra", [...]) list:
from megatron.core.transformer.moe.fused_a2a import HAVE_DEEP_EP, HAVE_DEEP_EP_V2
pytest.param(
"flex",
2,
{"moe_flex_dispatcher_backend": "deepepv2"},
id="flex-deepepv2-ep2",
marks=pytest.mark.skipif(not HAVE_DEEP_EP_V2, reason="DeepEP v2 not installed"),
),and refreshing the two manifest notes so they no longer read v1-only:
KernelEntry(
name="moe_token_dispatchers",
sources=(
"megatron/core/transformer/moe/token_dispatcher.py",
"megatron/core/transformer/moe/moe_layer.py",
),
tests=(K + "test_moe_kernels.py",),
kind="torch-op",
notes="MoELayer replay through the allgather / alltoall dispatchers (EP=1, EP=2) and "
"flex+DeepEP v1/v2 when available.",
),
KernelEntry(
name="moe_fused_a2a",
sources=("megatron/core/transformer/moe/fused_a2a.py",),
tests=(K + "test_moe_kernels.py",),
kind="external-lib",
notes="DeepEP v1 and v2 flex dispatcher cells (skipped without deep_ep / >=2 GPUs). "
"HybridEP and NCCL-EP backends are not yet replayed bit-exactly here.",
),The cell skips cleanly wherever deep_ep.ElasticBuffer is absent, matching how the v1 cell already behaves — so it satisfies the gate without making CI depend on a v2 container.
2. Stale error text now reachable via deepepv2
megatron/core/transformer/transformer_config.py:2055 — the predicate above it was widened to in ("deepep", "deepepv2"), but the message still names only deepep, so a deepepv2 user gets an error that does not mention their backend:
raise ValueError(
"Flex token dispatcher with deepep/deepepv2 backend does not support "
"moe_pad_expert_input_to_capacity"
)Not blocking
tests/unit_tests/transformer/moe/test_token_dispatcher_capacity.py still parametrizes ["deepep", "hybridep"]. Extending it is optional here — moe_expert_capacity_factor reaches _DeepepV2Manager.setup_metadata through the inherited v1 implementation, so the masking path is shared code rather than anything new in this PR.
|
/claude fix |
|
❌ Claude fix stopped because a workflow step failed. Inspect the run. |
|
/claude fix |
|
❌ Claude fix stopped because a workflow step failed. Inspect the run. |
What does this PR do ?
The DeepEP V2 support as a backend of flex dispatcher.
PR on dev: #4793
Issue tracking
For PRs from open-source community contributors:
Linked issue:
Contribution process
Pre-checks
Code review
Feel free to message or comment @NVIDIA/mcore-oncall to help accelerate your merge into main. The less complex your PR is, the faster it will be approved and merged!
All PRs start as draft. If you open a non-draft PR, it will be automatically converted to draft.
Step 1: Mark PR as "Ready for Review"
.github/CODEOWNERS.Final Review might get declined if these requirements are not fulfilled.
Step 2: Final Review
For PRs that change
megatron/core, once all expert reviewers have approved, theFinal Reviewlabel is applied automatically and final reviewers are assigned.For PRs outside
megatron/core, this step is skipped.Step 3: Approved
Once all required reviewers have approved, the
Approvedlabel is applied automatically.Merge
Any member of mcore-engineers will be able to merge your PR.