Skip to content

fix(moe): align HybridEP capacity growth with kernel chunk sizes - #7913

Draft
jalbericiola wants to merge 1 commit into
NVIDIA:mainfrom
jalbericiola:shared-prefix/hybridep-capacity-20261006
Draft

jalbericiola wants to merge 1 commit into
NVIDIA:mainfrom
jalbericiola:shared-prefix/hybridep-capacity-20261006

Conversation

@jalbericiola

Copy link
Copy Markdown
Contributor

Ragged batches can grow HybridEP's token reservation to a capacity that violates the installed dispatch/combine kernels' chunk alignment. Derive a common alignment from their configured chunk sizes and apply it to initial and grown reservations. Release unused Torch allocator cache before raw CUDA allocation growth so free cached memory is available to the external allocator.

The change affects reserved capacity, not actual token counts, routing probabilities or losses. It applies without shared-prefix execution. Add tests for initial/grown capacity, no-shrink behavior and invalid chunk sizes. Real dispatch/combine backward across a capacity boundary remains a required GPU qualification gate.

An issue is optional for this bounded bug fix. The GPU regression must verify the supported DeepEP version before merge.

Validation and review status

7 focused CPU helper checks pass for chunk alignment, invalid settings and no-shrink/growth behavior. CUDA cache release is mocked; this does not qualify real dispatch/combine or allocator behavior.

The adapted patch applies to the pinned current-main base and reproduces candidate file hashes exactly. This remains a draft pending the stated runtime/CI gates.

Final local checks

The native tools/autoformat.sh workflow, Black, isort, pylint, Ruff and git diff --check pass against the final scoped files. The shallow-history merge base was verified before accepting these results. Mypy remains advisory and is not clean; dependency/type findings are recorded separately. Real GPU allocator and dispatch/combine qualification remains pending.

Diff size

Incremental contribution: 74 added lines, including 38 added test lines (51.35%); 0 test lines removed. These are physical diff lines in test directories, including comments and blank lines, not test coverage.

Commits are cryptographically signed and locally verified. GitHub signing-key registration is pending, so GitHub may label the signature unverified.

Signed-off-by: Jorge Albericio <jalbericiola@nvidia.com>
@copy-pr-bot

copy-pr-bot Bot commented Oct 6, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant