Skip to content

Load MIMO-trained VLM checkpoints into LLaVAModel for dynamic inference - #7839

Merged
ericharper merged 11 commits into
NVIDIA:mainfrom
mathemakitten:helenn-mimo-vlm-inference
Oct 6, 2026
Merged

ericharper merged 11 commits into
NVIDIA:mainfrom
mathemakitten:helenn-mimo-vlm-inference

Conversation

@mathemakitten

Copy link
Copy Markdown
Contributor
  • I, the PR author, have personally reviewed every line of this PR.

What does this PR do?

Use tools/run_dynamic_text_generation_server.py to serve VLMs trained with MIMO, so image + text evals run on Megatron's own inference engine.

MIMO checkpoints nest each module under its own prefix (language_model.module.module., modality_submodules..module.module.encoders.., ...input_projections.0.*). Instead of building a MIMO model for inference, the server builds a LLaVAModel from the same parts and remaps its sharded state-dict keys onto the checkpoint's names.

New flags:

--vision-num-layers, --vision-hidden-size, --vision-ffn-hidden-size, --vision-num-attention-heads, --vision-num-query-groups, --vision-kv-channels (override the registry sizes of --vision-model-type)- --vision-projection-activation {gelu,fast_gelu} - --image-token-id

Needs #7799, #7752, #7772 to merge.

Issue tracking

For PRs from open-source community contributors:

  • New features: a linked issue is required. Please open a feature request and reference it here before submitting the PR.
  • Small updates (bug fixes, minor improvements): a linked issue is recommended and will accelerate the PR review process.

Linked issue:

Contribution process

Pre-checks

  • I have added relevant unit tests
  • I have added relevant functional tests
  • If this PR adds or changes a GPU kernel (Triton, jit_fuser/torch.compile, CUDA extension, TE or external-library dispatch, or a scatter/index accumulation), I have added or updated its bit-exact determinism test and registered it in tests/unit_tests/determinism/kernels/manifest.py (guide)
  • I have added proper typing to my code Typing guidelines
  • I have added relevant documentation
  • I have run the autoformatter.sh on my PR

Code review

Feel free to message or comment @NVIDIA/mcore-oncall to help accelerate your merge into main. The less complex your PR is, the faster it will be approved and merged!

All PRs start as draft. If you open a non-draft PR, it will be automatically converted to draft.

Step 1: Mark PR as "Ready for Review"

  1. When your PR is ready, click Ready for Review.
  2. An oncall reviewer is auto-assigned and expert reviewers are notified based on your changes.
    • Some PRs may jump straight to step 2. This is determined by .github/CODEOWNERS.

⚠️ Only mark as ready once merge-conflicts are resolved and the CI is passing.
Final Review might get declined if these requirements are not fulfilled.

Step 2: Final Review

For PRs that change megatron/core, once all expert reviewers have approved, the Final Review label is applied automatically and final reviewers are assigned.

For PRs outside megatron/core, this step is skipped.

Step 3: Approved

Once all required reviewers have approved, the Approved label is applied automatically.

Merge

Any member of mcore-engineers will be able to merge your PR.

@copy-pr-bot

copy-pr-bot Bot commented Oct 2, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

Signed-off-by: Helen Ngo <helenn@nvidia.com>
Signed-off-by: Helen Ngo <helenn@nvidia.com>
Signed-off-by: Helen Ngo <helenn@nvidia.com>
Signed-off-by: Helen Ngo <helenn@nvidia.com>
Signed-off-by: Helen Ngo <helenn@nvidia.com>
Signed-off-by: Helen Ngo <helenn@nvidia.com>
Signed-off-by: Helen Ngo <helenn@nvidia.com>
@mathemakitten
mathemakitten force-pushed the helenn-mimo-vlm-inference branch from 096db4f to 8b3b8fd Compare October 5, 2026 18:55
Signed-off-by: Helen Ngo <helenn@nvidia.com>
@mathemakitten
mathemakitten marked this pull request as ready for review October 5, 2026 19:01
@mathemakitten
mathemakitten requested review from a team as code owners October 5, 2026 19:01
@wdykas

wdykas commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

In case it helps we do have mimo to llava remapping stuff here https://github.com/NVIDIA/Megatron-LM/pull/7367/changes#diff-585c880dcd23609d8158a7da63e68841621f9630ec33e8614fd22742b1707571R106

Signed-off-by: Helen Ngo <helenn@nvidia.com>
@svcnvidia-nemo-ci svcnvidia-nemo-ci added the Final Review PR is in the "final review" stage label Oct 6, 2026
@svcnvidia-nemo-ci svcnvidia-nemo-ci added Approved All necessary approvals have been made and removed Final Review PR is in the "final review" stage labels Oct 6, 2026
@ericharper
ericharper enabled auto-merge October 6, 2026 17:03
@ericharper
ericharper added this pull request to the merge queue Oct 6, 2026
@nemo-automation-bot

Copy link
Copy Markdown

🔄 Merge queue validation started!

You can track the progress here: https://github.com/NVIDIA/Megatron-LM/actions/runs/37515456785

@nemo-automation-bot

Copy link
Copy Markdown

🔄 Merge queue validation started!

You can track the progress here: https://github.com/NVIDIA/Megatron-LM/actions/runs/37521041312

Merged via the queue into NVIDIA:main with commit d3a9e24 Oct 6, 2026
117 of 121 checks passed

This branch was successfully deployed

2 active deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Approved All necessary approvals have been made complexity: high

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants