Repository navigation
data: collect DeepSeek V4.1 H100/H200/B200/GB200 operator profiles - #305
Conversation
Signed-off-by: Harry Lee <harrli@nvidia.com>
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Important Review skippedReview was skipped as selected files did not have any reviewable changes. ⛔ Files ignored due to path filters (16)
⚙️ Run configurationConfiguration used: Repository: ai-dynamo/aisimulate/.coderabbit.yaml Review profile: ASSERTIVE Plan: Enterprise Run ID: ⛔ Files ignored due to path filters (16)
You can disable this status message by setting the Use the checkbox below for a quick retry:
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: ai-dynamo/aisimulate/.coderabbit.yaml Review profile: ASSERTIVE Plan: Enterprise Run ID: ⛔ Files ignored due to path filters (16)
📒 Files selected for processing (6)
🔗 Linked repositories identifiedCodeRabbit considers these linked repositories for cross-repo context during reviews:
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review. 📜 Recent review details⏰ Context from checks skipped due to timeout. (21)
🧰 Additional context used📓 Path-based instructions (7)Preserve the Rust single oracle: Python may describe operations, load raw data, orchestrate, and present results, but must not compute per-op performance values.⚙️ CodeRabbit configuration file Files:
Enforce the single-oracle and golden-diff rules in python/aisimulate/.claude/rules/rust-core/parity.md.⚙️ CodeRabbit configuration file Files:
Verify upstream provenance, license text, collection and overlay hashes, local modification notices, and generated derivative records remain complete and internally consistent.⚙️ CodeRabbit configuration file Files:
Read REVIEW.md before commenting.⚙️ CodeRabbit configuration file Files:
Source excerpt: Do NOT add Python-side interpolation, roofline/SOL formulas, empirical-utilization estimates, or per-call table lookups anywhere under `python/aisimulate/src/aisimulate_core/sdk/` (banned def shapes: the `_query_*` and `_loo...📄 CodeRabbit inference engine (python/aisimulate/.claude/rules/rust-core/parity.md) Files:
Before making any change under: `python/aisimulate/src/aisimulate/generator/**` MUST read: `python/aisimulate/.claude/rules/generator-development.md` Before making any change under `python/aisimulate/collector/**` MUST read: `python/aisimul...📄 CodeRabbit inference engine (AGENTS.md) Files:
Source excerpt: Only workflows under the repository-root `.github/workflows/` run for this repository.📄 CodeRabbit inference engine (REVIEW.md) Files:
🔀 Multi-repo context ai-dynamo/aiconfigurator, ai-dynamo/dynamoLinked repositories findingsai-dynamo/aiconfigurator
ai-dynamo/dynamo
🔇 Additional comments (1)
📝 SummaryRisk: HighHuman attention should focus on:
Changed behavior and contracts
Evidence supplied
Evidence still missing
Technical qualityThe changes add explicit checks for execution identity, provenance, schema compatibility, and aggregation inputs. Incorrect admission or aggregation could still produce measurements that do not represent the intended native execution path. Merge readinessMerge readiness is not established by the supplied evidence. Verify the reported CI and validation results, and inspect the generated dataset and provenance artifacts. The reported CI status is evidence to verify, not review approval. WalkthroughThe change adds Humming MXFP4/BF16 support, Hopper-specific DeepSeek V4.1 scoring, native SGLang attention and isolated collectors, bounded evidence aggregation, and runtime and collection metadata for multiple systems. ChangesDeepSeek V4.1 performance modeling
Priority: ➖ Normal Estimated code review effort: 5 (Critical) | ~90 minutes Merge Risk: ⚪ Minimal · up to No actionable merge-blocking issue is established for these changes; the PR is mergeable after normal checks. 🚥 Pre-merge checks | ✅ 8✅ Passed checks (8 passed)
Comment |
Signed-off-by: Harry Lee <harrli@nvidia.com>
Signed-off-by: Harry Lee <harrli@nvidia.com>
Signed-off-by: Harry Lee <harrli@nvidia.com>
Signed-off-by: Harry Lee <harrli@nvidia.com>
Signed-off-by: Harry Lee <harrli@nvidia.com>
Signed-off-by: Harry Lee <harrli@nvidia.com>
Signed-off-by: Harry Lee <harrli@nvidia.com>
Signed-off-by: Harry Lee <harrli@nvidia.com>
Signed-off-by: Harry Lee <harrli@nvidia.com>
Signed-off-by: Harry Lee <harrli@nvidia.com>
Signed-off-by: Harry Lee <harrli@nvidia.com>
Signed-off-by: Harry Lee <harrli@nvidia.com>
Signed-off-by: Harry Lee <harrli@nvidia.com>
Signed-off-by: Harry Lee <harrli@nvidia.com>
What changes
Add measured DeepSeek-V4.1-Flash
fullanddecoder_boundedoperator profiles for H100, H200, B200 and GB200 at TP2 and TP4. The eight cells contain 10,120 physical keys in 32 parquet/provenance files (206,830 bytes). Existing tables are preserved; raw campaign archives remain outside GitHub.Native Hopper CUTLASS W4A16, explicit Humming W4A16 and Blackwell TRTLLM W4A8 retain distinct execution identities. Humming adds a Python/Rust quantization value without changing existing Rust bincode discriminants. Strict schema-v2 readers accept the provenance fields already emitted by the collector while rejecting unknown fields and invalid types.
The isolated producers measure native linear, Engram, mHC, MoE, GEMM, NCCL and attention using declared reproducible synthetic parameters and actual tokenizer/native KV semantics. GB200 TP4 attention comes from the whole pinned checkpoint; the other attention cells use the isolated native module. All eight cells use the same isolated primitive method. Duplicate non-TP mHC/router keys retain TP2 deterministically after a shape/dtype/source/scope audit; selection never depends on latency, and TP4 raw observations are retained.
Coverage and independent accuracy
All eight primitive and formal attention collections completed. The 1,200 exact primitive queries pass through the database and Rust SILICON interface. Each full cell passes 183 strict predictions: 145 calibration geometries and 38 independent held-out geometries, 1,464 full predictions. Each bounded cell additionally passes 154 calibration and 38 held-out geometries, for 3,000 strict predictions total. All original full predictions and rows are unchanged. This is the declared case grid, not arbitrary batch/length coverage.
Independent whole-checkpoint held-out validation uses 38 geometries, five repetitions and four ranks per cell:
Predictions and table hashes were frozen before reading each system's truth. Text and tokenized corpora are separate, and calibration/held-out geometry overlap is zero. Ground truth is the median of five per-repetition rank maxima, using native synchronized wall time including prepare/forward/sample. No accuracy threshold is asserted, and the larger H200 decode error is retained without fitting to held-out data. These are OP/SILICON results, not FPM results.
decoder_boundedadds 260 non-overlapping attention keys per TP cell. Eight physical-key collision groups were reviewed against actual native layouts, integer mappings and source contracts. The explicit reducer keeps every logical owner, takes each owner's median of per-repetition rank maxima, then averages owners equally; it does not assert equal inputs or latency. Original full rows take fixed precedence at shared keys. The default collector still rejects collisions, and the joint exporter requires separately admitted rows. H200 bounded independent whole-model heldout4449635completed on the verified node 333, and all 760 retained rank/repetition observations reproduce the reported metrics. B200 full/bounded heldout4448904completed on node 018 with exit 0. All 168 archived files (8,248,737 bytes) were verified against the original remote SHA inventory; each profile retains 38 geometries × five repetitions × four ranks. Original native admission and a separate raw-rank arithmetic review reproduce the metrics above. Every B200 case underpredicts actual time; decode errors of 25.51% and 26.97% are retained without fitting. Requested isolated operator data and independent whole-model validation for all feasible TP4 cells are complete. TP2 and H100 TP4 whole-model accuracy remain outside the feasible residency scope.Whole-model TP2 on all four GPUs and H100 TP4 exceed HBM under the requested checkpoint/placement. Isolated operator coverage and latency predictions do not establish whole-model fit or accuracy for those cells.
Runtime and use
Select SGLang
dev-1aa0e962b206102b7c439a4a0c4981cfec6e87bc, FP8-block GEMM and FP8 FMHA. Hopper TP2 usesw4a16_mxfp4_cutlass; Hopper TP4 usesw4a16_mxfp4_humming; Blackwell usesw4a8_mxfp4_mxfp8_trtllm. MoE TP equals model TP, with EP1.The measured NCCL provider is 2.30.7, distinct from Torch's 2.29.7 build macros. Same-rank ARM getter/library receipts and explicitly labeled AMD post-run library-identity evidence support that identity. Use the README's existing
systems_pathsoverlay to select 2.30.7 without changing system defaults. Each collection event records its actual SquashFS identity; the source audit records the installed vision-only guard and does not describe the image as an unmodified upstream checkout.Review and validation
Start with
python/aisimulate/collector/sglang/README.dsv41.md, the three native producers (dsv41_isolated_runner.py,dsv41_attention_runner.py,dsv41_native_runner.py), their tests and the schema-v2 reader insrc/aisimulate_core/sdk/perf_database.py.-I, it reproduces all 3,000 predictions and the three README CLI examples. Wheel SHA256:d11fc2adf6cb65210b341e1ce88860957672523fe82fd189489173b636127a40.5cfe2827: Fast CI and Full CI both pass. The previous8744b810Fast/Full CI pass is historical evidence. Registered backend facts distinguish Humming and conservatively unresolved composite modules; table timings are unchanged by registration.The source-audited Hopper index-score fallback now uses native BF16 tensor-core arithmetic and a two-byte query operand for the SGLang FP8/BF16 layout on SM90. Packed FP4 index-K storage remains packed; SM100 and other layouts keep their original requirements. This also permits H200 FPM SOL transfer without inventing an FP4 throughput field. The 51 architecture and 121 FPM tests pass (one existing ignored test). Source attribution and byte-identical packaged notices are retained.
Companion FPM PR: #304. Builds on merged #160/#233 and the shared V4.1 architecture contract.