Skip to content

feat(sft): add Nemotron task encoders and cookers - #4108

Open
rohitrango wants to merge 31 commits into
mainfrom
rohit/sft_v2_nemotron_cooker
Open

rohitrango wants to merge 31 commits into
mainfrom
rohit/sft_v2_nemotron_cooker

Conversation

@rohitrango

Copy link
Copy Markdown
Contributor

What does this PR do ?

Adds Nemotron-specific cookers and multimodal task encoding to the Energon SFT v2 loader on top of the PackedTensor optimization branch.

Key changes:

  • Add Nemotron text, image, video, and audio cookers.
  • Add Nemotron tokenization, visual processing, and multimodal task encoding.
  • Add the skip-chat-template path for preformatted Nemotron data.
  • Register the new components and expose their loader configuration.
  • Add focused unit coverage for the cookers, encoders, loader wiring, and message handling.

Issues

None.

Usage

Select the nemotron_multimodal task encoder and the appropriate registered Nemotron cooker in an Energon SFT v2 loader configuration.

Before your PR is "Ready for review"

  • Make sure you read and followed Contributor guidelines
  • Did you write any new necessary tests?
  • Did you run the unit tests and functional tests locally?
  • Did you add or update any necessary documentation?

Additional Information

This is a stacked PR based on rohit/packedtensor_opt. Local tests were not rerun during the final rebase and push. No uv commands were run.

Stack and previous PR

This PR is stacked on #4107. Previous PR: rohitrango#8.

@copy-pr-bot

copy-pr-bot Bot commented Sep 12, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@rohitrango
rohitrango added this pull request to stack #4109 September 12, 2026 00:15
@rohitrango
rohitrango marked this pull request as ready for review September 12, 2026 00:15
@rohitrango
rohitrango requested review from a team as code owners September 12, 2026 00:15
@copy-pr-bot

copy-pr-bot Bot commented Sep 12, 2026

Copy link
Copy Markdown

Auto-sync is disabled for ready for review pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@github-actions

Copy link
Copy Markdown

❌ Submodule Fast-Forward Check Failed

Check based on commit: f2c6c94 (PR #4108 from rohit/sft_v2_nemotron_cooker)

❌ Submodules that need attention:

Megatron-Bridge: ❌ Commits have DIVERGED from a common ancestor
TARGET (rohit/packedtensor_opt branch): https://github.com/NVIDIA-NeMo/Megatron-Bridge/commits/5ed97996cc2b422904d18179375b6d7366915097/
CURRENT (PR #4108 from rohit/sft_v2_nemotron_cooker): https://github.com/NVIDIA-NeMo/Megatron-Bridge/commits/3961f399ef181bca689de8e984110b85c4df00fe/

Please ensure all submodule commits are fast-forwards of the rohit/packedtensor_opt branch before merging.

@rohitrango
rohitrango force-pushed the rohit/sft_v2_nemotron_cooker branch from f2c6c94 to 124ca91 Compare September 12, 2026 00:19
@github-actions

Copy link
Copy Markdown

❌ Submodule Fast-Forward Check Failed

Check based on commit: 124ca91 (PR #4108 from rohit/sft_v2_nemotron_cooker)

❌ Submodules that need attention:

Megatron-Bridge: ❌ Commits have DIVERGED from a common ancestor
TARGET (main branch): https://github.com/NVIDIA-NeMo/Megatron-Bridge/commits/5ed97996cc2b422904d18179375b6d7366915097/
CURRENT (PR #4108 from rohit/sft_v2_nemotron_cooker): https://github.com/NVIDIA-NeMo/Megatron-Bridge/commits/3961f399ef181bca689de8e984110b85c4df00fe/

Please ensure all submodule commits are fast-forwards of the main branch before merging.

1 similar comment
@github-actions

Copy link
Copy Markdown

❌ Submodule Fast-Forward Check Failed

Check based on commit: 124ca91 (PR #4108 from rohit/sft_v2_nemotron_cooker)

❌ Submodules that need attention:

Megatron-Bridge: ❌ Commits have DIVERGED from a common ancestor
TARGET (main branch): https://github.com/NVIDIA-NeMo/Megatron-Bridge/commits/5ed97996cc2b422904d18179375b6d7366915097/
CURRENT (PR #4108 from rohit/sft_v2_nemotron_cooker): https://github.com/NVIDIA-NeMo/Megatron-Bridge/commits/3961f399ef181bca689de8e984110b85c4df00fe/

Please ensure all submodule commits are fast-forwards of the main branch before merging.

rohitrango and others added 11 commits September 14, 2026 10:13
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: Rohit Jena <rohitrango@users.noreply.github.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
@rohitrango
rohitrango force-pushed the rohit/sft_v2_nemotron_cooker branch from 124ca91 to d044f40 Compare September 14, 2026 17:14
@github-actions

Copy link
Copy Markdown

❌ Submodule Fast-Forward Check Failed

Check based on commit: d044f40 (PR #4108 from rohit/sft_v2_nemotron_cooker)

❌ Submodules that need attention:

Megatron-Bridge: ❌ Commits have DIVERGED from a common ancestor
TARGET (main branch): https://github.com/NVIDIA-NeMo/Megatron-Bridge/commits/5ed97996cc2b422904d18179375b6d7366915097/
CURRENT (PR #4108 from rohit/sft_v2_nemotron_cooker): https://github.com/NVIDIA-NeMo/Megatron-Bridge/commits/3961f399ef181bca689de8e984110b85c4df00fe/

Please ensure all submodule commits are fast-forwards of the main branch before merging.

@github-actions

Copy link
Copy Markdown

❌ Submodule Fast-Forward Check Failed

Check based on commit: 66b1237 (PR #4108 from rohit/sft_v2_nemotron_cooker)

❌ Submodules that need attention:

Megatron-Bridge: ❌ Commits have DIVERGED from a common ancestor
TARGET (main branch): https://github.com/NVIDIA-NeMo/Megatron-Bridge/commits/5ed97996cc2b422904d18179375b6d7366915097/
CURRENT (PR #4108 from rohit/sft_v2_nemotron_cooker): https://github.com/NVIDIA-NeMo/Megatron-Bridge/commits/3961f399ef181bca689de8e984110b85c4df00fe/

Please ensure all submodule commits are fast-forwards of the main branch before merging.

rohitrango and others added 16 commits September 15, 2026 10:55
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Keep one-field-at-a-time staging while moving PackedTensor segments to the NCCL device before concatenation. This avoids the large CPU concatenation and subsequent flat host-to-device copy in worker fetch.

Signed-off-by: Rohit Jena <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: Rohit Jena <rohit.rango@gmail.com>
Signed-off-by: Rohit Jena <rohit.rango@gmail.com>
Preserve precomputed Nemotron token masks during physical pack preparation and keep loss-mask mode metadata out of the local data plane. Disable Energon's fatal sample tolerance so malformed source rows are skipped by the configured handler.

Signed-off-by: Rohit Jena <rohit.rango@gmail.com>
Move packing ownership into the Energon task encoder so the upstream cooker path can reproduce the v1 loader parameters. Normalize hybrid-model MoE metrics by the actual MoE layer count and add regression coverage and pipeclean configs.

Signed-off-by: Rohit Jena <rohit.rango@gmail.com>
Signed-off-by: Rohit Jena <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
With calculate_per_token_loss=True the mcore router pre-multiplies its aux
loss by the microbatch's padded token count (num_local_tokens *
tp_cp_group.size() in MoETopKRouter.attach_and_log_load_balancing_loss), on
the assumption that the same count is the gradient denominator. NeMo-RL
normalizes gradients by the supervised token count instead, so on a heavily
masked recipe the aux gradient is inflated by padded/supervised. The 67B
Super VLM SFT recipe trains ~22% of its packed tokens, which gave the aux
loss ~4.5x its configured moe_aux_loss_coeff and drove
moe/seq_load_balancing_loss well below the Megatron-LM reference run
(~0.92 vs ~0.96 at the same step) while lm loss stayed comparable.

Megatron-LM runs --no-calculate-per-token-loss and never enters that path,
so its aux loss carries exactly the configured coefficient.

_compute_moe_grad_scale now returns local_valid_toks / padded_toks, which
cancels the router's padded count and leaves sum_mb(valid_toks_mb *
d(aux_mb)); the later 1/global_valid_toks rescale turns that into the
supervised-token-weighted mean that calculate_per_token_loss=True is
defined to produce. global_valid_toks is now optional: the synchronous path
supplies it because its gradients are never rescaled afterwards, while the
split path omits it because _finish_train_step_body already applies 1/N to
every gradient.

The split path sets the scale per train_microbatch call and clears it
straight after the forward-backward, alongside the existing MTP clears in
begin_train_step, the error paths and abort_train_step, so no step-local
callable survives into a serialized config.

process_global_batch also returns local_valid_toks so the synchronous path
can reach this rank's own count.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
@rohitrango
rohitrango force-pushed the rohit/sft_v2_nemotron_cooker branch from 66b1237 to 5dd87fb Compare September 15, 2026 17:57
@github-actions

Copy link
Copy Markdown

❌ Submodule Fast-Forward Check Failed

Check based on commit: 5dd87fb (PR #4108 from rohit/sft_v2_nemotron_cooker)

❌ Submodules that need attention:

Megatron-Bridge: ❌ Commits have DIVERGED from a common ancestor
TARGET (main branch): https://github.com/NVIDIA-NeMo/Megatron-Bridge/commits/5ed97996cc2b422904d18179375b6d7366915097/
CURRENT (PR #4108 from rohit/sft_v2_nemotron_cooker): https://github.com/NVIDIA-NeMo/Megatron-Bridge/commits/3961f399ef181bca689de8e984110b85c4df00fe/

Please ensure all submodule commits are fast-forwards of the main branch before merging.

1 similar comment
@github-actions

Copy link
Copy Markdown

❌ Submodule Fast-Forward Check Failed

Check based on commit: 5dd87fb (PR #4108 from rohit/sft_v2_nemotron_cooker)

❌ Submodules that need attention:

Megatron-Bridge: ❌ Commits have DIVERGED from a common ancestor
TARGET (main branch): https://github.com/NVIDIA-NeMo/Megatron-Bridge/commits/5ed97996cc2b422904d18179375b6d7366915097/
CURRENT (PR #4108 from rohit/sft_v2_nemotron_cooker): https://github.com/NVIDIA-NeMo/Megatron-Bridge/commits/3961f399ef181bca689de8e984110b85c4df00fe/

Please ensure all submodule commits are fast-forwards of the main branch before merging.

Signed-off-by: rohitrango <rohit.rango@gmail.com>
Reverts 331cb14.

Signed-off-by: rohitrango <rohit.rango@gmail.com>
@rohitrango
rohitrango requested a review from a team as a code owner September 19, 2026 18:06
Base automatically changed from rohit/packedtensor_opt to main September 19, 2026 21:02
Signed-off-by: rohitrango <rohit.rango@gmail.com>
@github-actions github-actions Bot added the Documentation Improvements or additions to documentation label Sep 26, 2026
@rohitrango

Copy link
Copy Markdown
Contributor Author

/ok to test 15d258c

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants