Skip to content

feat(sweeper): add heterogeneous P/D search [AIC-1772] - #22

Closed
jasonqinzhou wants to merge 6 commits into
ai-dynamo:mainfrom
jasonqinzhou:jasonzho/aic-1772-sweeper-heterogeneous-pd
Closed

jasonqinzhou wants to merge 6 commits into
ai-dynamo:mainfrom
jasonqinzhou:jasonzho/aic-1772-sweeper-heterogeneous-pd

Conversation

@jasonqinzhou

@jasonqinzhou jasonqinzhou commented Aug 19, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • add independent prefill/decode model, system, backend/version, estimator/data, engine-control, scheduler, quantization, and KV-feasibility resolution with documented inheritance
  • search backend pairs and role-specific parallel shapes under one GPU budget, retaining exact role estimator, engine-request, and candidate provenance
  • support mixed vLLM/SGLang disaggregated replay, role-correct lowering, GPU accounting, legacy-calibrated rate matching, and categorized role failures
  • fail closed for unsupported heterogeneous backend pairs, mismatched role identities/model architectures, invalid finite correction values, and adapters without explicit heterogeneous-P/D capability opt-in
  • forward role systems paths and estimator database/transfer/FPM controls losslessly into the runner and Rust/Python engine compile path
  • preserve the homogeneous input and replay path unchanged when no role override is configured

Stack dependencies

Parity remediation

  • require explicit prefill/decode backend-pair capability and retain deterministic considered/accepted/pruned diagnostics even when another pair remains viable
  • keep RoleIdentity identical to the resolved per-role EstimatorSpec, including inherited backend-version fallback; reject incompatible model architectures before KV handoff
  • forward configured role/shared systems_paths, performance database mode, transfer policy, FPM selection, and exact data roots through engine lowering
  • prevent shared chunked-prefill controls from leaking into the decode request/provenance
  • make all five disaggregated legacy correction knobs executable through a canonical replay correction spec and finite timing-scale validation
  • reject standalone or effective per-role rate arithmetic overflow and validate every canonical role_rates value before serialization
  • retain ordered counts plus pair/role/category pruning details on terminal NoViableParallelConfig.enumeration_reports when every heterogeneous pair is pruned, including explicit pinned domains

Exact-head validation (555f5695)

  • pinned/enumerated all-pairs-pruned integration and search-space suites: 27 passed
  • full root Python/Sweeper suite: 381 passed, 3 skipped
  • exact CI Ruff scope: passed
  • git diff --check: passed
  • generator bridge diff audit: no files touched

The preceding remediation head also passed 456 workspace Rust tests. Earlier proportional legacy Python validation completed 3,578 passed, 10 skipped, 0 failures before a bounded stop; 66 selected slow build-tail cases were not completed locally.

Residual boundaries

  • TensorRT-LLM disaggregated execution remains unsupported and is rejected by role/pair capability.
  • decode-side chunked prefill remains disabled because it is a prefill-role scheduler control.
  • optional adapters must explicitly advertise each supported heterogeneous backend pair after auditing role-specific candidate materialization.
  • the 66 slow legacy build-tail cases remain for hosted CI; no failure was observed in the completed 3,588 outcomes.

Linear: AIC-1772

Signed-off-by: Jason Zhou (Engrg-Hardware 1) <jasonzho@nvidia.com>
Signed-off-by: Jason Zhou (Engrg-Hardware 1) <jasonzho@nvidia.com>
Signed-off-by: Jason Zhou (Engrg-Hardware 1) <jasonzho@nvidia.com>
@linear-code

linear-code Bot commented Aug 19, 2026 •

Copy link
Copy Markdown

Signed-off-by: Jason Zhou (Engrg-Hardware 1) <jasonzho@nvidia.com>
Signed-off-by: Jason Zhou (Engrg-Hardware 1) <jasonzho@nvidia.com>
Signed-off-by: Jason Zhou (Engrg-Hardware 1) <jasonzho@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant