Skip to content

fix(moe): align HybridEP capacity growth with kernel chunk sizes - #7913

Draft
jalbericiola wants to merge 2 commits into
NVIDIA:mainfrom
jalbericiola:shared-prefix/hybridep-capacity-20261006
Draft

jalbericiola wants to merge 2 commits into
NVIDIA:mainfrom
jalbericiola:shared-prefix/hybridep-capacity-20261006

Conversation

@jalbericiola

@jalbericiola jalbericiola commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Ragged batches can grow HybridEP's token reservation to a capacity that violates the installed dispatch/combine kernels' chunk alignment. Derive a common alignment from their configured chunk sizes and apply it to initial and grown reservations. Release unused Torch allocator cache before raw CUDA allocation growth so free cached memory is available to the external allocator.

The change affects reserved capacity, not actual token counts, routing probabilities or losses. It applies without shared-prefix execution. The tests cover alignment, growth without shrinking, invalid chunk sizes and the cache-release call.

Validation and review status

Merged main at cff9cf4e0935311ca12faa84820f2d329c17271a without conflicts. The PR remains scoped to the capacity helper changes and their tests.

  • Seven focused CPU helper checks pass. CUDA cache release is mocked.
  • Both scoped Python files compile.
  • The native tools/autoformat.sh workflow passes its Black, isort, pylint and Ruff gates; git diff --check passes.
  • Mypy remains advisory and reports dependency/type findings in the local lint environment.

This remains a draft pending real GPU allocator and dispatch/combine backward qualification across a capacity boundary with the supported DeepEP version.

Signed-off-by: Jorge Albericio <jalbericiola@nvidia.com>
@copy-pr-bot

copy-pr-bot Bot commented Oct 6, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

Signed-off-by: Jorge Albericio <jalbericiola@nvidia.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant