Skip to content

Pull requests: NVIDIA/Megatron-LM

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

Require an explicit process-group collection in LanguageModule
#6303 opened Aug 5, 2026 by Connor-XY Contributor Draft
5 of 6 tasks
Support direct hybrid MTP layer specs
#6298 opened Aug 5, 2026 by Phlip79 Member Draft
Add direct hybrid architecture descriptors
#6295 opened Aug 5, 2026 by Phlip79 Member Draft
Remove global process-group reads from megatron/core (207 -> 155)
#6293 opened Aug 5, 2026 by Connor-XY Contributor Draft
5 of 6 tasks
Support MXFP8 parameter gather in MIMO training
#6284 opened Aug 5, 2026 by yashaswikarnati Contributor Draft
3 of 6 tasks
save tokenizer assets [not ready for review]
#6280 opened Aug 5, 2026 by dimapihtar Contributor Draft
1 of 6 tasks
Fix latent-MoE expert fc2 initialization community-request
#6278 opened Aug 5, 2026 by bo-ke Loading…
6 tasks
Print params L2 norm before the start of training complexity: low
#6276 opened Aug 5, 2026 by gautham-kollu Contributor Loading…
6 tasks
TE Layernorm dtype guard with fp32 residuals
#6272 opened Aug 5, 2026 by mkhona-nvidia Contributor Draft
1 of 6 tasks
Fix fused MLA and MTP checkpointing for Megatron-FSDP. complexity: medium
#6271 opened Aug 4, 2026 by rapatel Contributor Loading…
6 tasks done
ProTip! no:milestone will show everything without a milestone.