feat(sft): add Nemotron task encoders and cookers - #4108
rohitrango wants to merge 31 commits into
Conversation
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
|
Auto-sync is disabled for ready for review pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
❌ Submodule Fast-Forward Check FailedCheck based on commit: f2c6c94 (PR #4108 from ❌ Submodules that need attention:Megatron-Bridge: ❌ Commits have DIVERGED from a common ancestor Please ensure all submodule commits are fast-forwards of the rohit/packedtensor_opt branch before merging. |
f2c6c94 to
124ca91
Compare
❌ Submodule Fast-Forward Check FailedCheck based on commit: 124ca91 (PR #4108 from ❌ Submodules that need attention:Megatron-Bridge: ❌ Commits have DIVERGED from a common ancestor Please ensure all submodule commits are fast-forwards of the main branch before merging. |
1 similar comment
❌ Submodule Fast-Forward Check FailedCheck based on commit: 124ca91 (PR #4108 from ❌ Submodules that need attention:Megatron-Bridge: ❌ Commits have DIVERGED from a common ancestor Please ensure all submodule commits are fast-forwards of the main branch before merging. |
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: Rohit Jena <rohitrango@users.noreply.github.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
124ca91 to
d044f40
Compare
❌ Submodule Fast-Forward Check FailedCheck based on commit: d044f40 (PR #4108 from ❌ Submodules that need attention:Megatron-Bridge: ❌ Commits have DIVERGED from a common ancestor Please ensure all submodule commits are fast-forwards of the main branch before merging. |
❌ Submodule Fast-Forward Check FailedCheck based on commit: 66b1237 (PR #4108 from ❌ Submodules that need attention:Megatron-Bridge: ❌ Commits have DIVERGED from a common ancestor Please ensure all submodule commits are fast-forwards of the main branch before merging. |
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Keep one-field-at-a-time staging while moving PackedTensor segments to the NCCL device before concatenation. This avoids the large CPU concatenation and subsequent flat host-to-device copy in worker fetch. Signed-off-by: Rohit Jena <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: Rohit Jena <rohit.rango@gmail.com>
Signed-off-by: Rohit Jena <rohit.rango@gmail.com>
Preserve precomputed Nemotron token masks during physical pack preparation and keep loss-mask mode metadata out of the local data plane. Disable Energon's fatal sample tolerance so malformed source rows are skipped by the configured handler. Signed-off-by: Rohit Jena <rohit.rango@gmail.com>
Move packing ownership into the Energon task encoder so the upstream cooker path can reproduce the v1 loader parameters. Normalize hybrid-model MoE metrics by the actual MoE layer count and add regression coverage and pipeclean configs. Signed-off-by: Rohit Jena <rohit.rango@gmail.com>
Signed-off-by: Rohit Jena <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
With calculate_per_token_loss=True the mcore router pre-multiplies its aux loss by the microbatch's padded token count (num_local_tokens * tp_cp_group.size() in MoETopKRouter.attach_and_log_load_balancing_loss), on the assumption that the same count is the gradient denominator. NeMo-RL normalizes gradients by the supervised token count instead, so on a heavily masked recipe the aux gradient is inflated by padded/supervised. The 67B Super VLM SFT recipe trains ~22% of its packed tokens, which gave the aux loss ~4.5x its configured moe_aux_loss_coeff and drove moe/seq_load_balancing_loss well below the Megatron-LM reference run (~0.92 vs ~0.96 at the same step) while lm loss stayed comparable. Megatron-LM runs --no-calculate-per-token-loss and never enters that path, so its aux loss carries exactly the configured coefficient. _compute_moe_grad_scale now returns local_valid_toks / padded_toks, which cancels the router's padded count and leaves sum_mb(valid_toks_mb * d(aux_mb)); the later 1/global_valid_toks rescale turns that into the supervised-token-weighted mean that calculate_per_token_loss=True is defined to produce. global_valid_toks is now optional: the synchronous path supplies it because its gradients are never rescaled afterwards, while the split path omits it because _finish_train_step_body already applies 1/N to every gradient. The split path sets the scale per train_microbatch call and clears it straight after the forward-backward, alongside the existing MTP clears in begin_train_step, the error paths and abort_train_step, so no step-local callable survives into a serialized config. process_global_batch also returns local_valid_toks so the synchronous path can reach this rank's own count. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: rohitrango <rohit.rango@gmail.com>
66b1237 to
5dd87fb
Compare
❌ Submodule Fast-Forward Check FailedCheck based on commit: 5dd87fb (PR #4108 from ❌ Submodules that need attention:Megatron-Bridge: ❌ Commits have DIVERGED from a common ancestor Please ensure all submodule commits are fast-forwards of the main branch before merging. |
1 similar comment
❌ Submodule Fast-Forward Check FailedCheck based on commit: 5dd87fb (PR #4108 from ❌ Submodules that need attention:Megatron-Bridge: ❌ Commits have DIVERGED from a common ancestor Please ensure all submodule commits are fast-forwards of the main branch before merging. |
Signed-off-by: rohitrango <rohit.rango@gmail.com>
Reverts 331cb14. Signed-off-by: rohitrango <rohit.rango@gmail.com>
Signed-off-by: rohitrango <rohit.rango@gmail.com>
|
/ok to test 15d258c |
What does this PR do ?
Adds Nemotron-specific cookers and multimodal task encoding to the Energon SFT v2 loader on top of the PackedTensor optimization branch.
Key changes:
Issues
None.
Usage
Select the nemotron_multimodal task encoder and the appropriate registered Nemotron cooker in an Energon SFT v2 loader configuration.
Before your PR is "Ready for review"
Additional Information
This is a stacked PR based on rohit/packedtensor_opt. Local tests were not rerun during the final rebase and push. No uv commands were run.
Stack and previous PR
This PR is stacked on #4107. Previous PR: rohitrango#8.