Skip to content

feat(rollout): replay MInf routing indices - #4307

Open
lauradang wants to merge 65 commits into
NVIDIA-NeMo:mainfrom
lauradang:laurad/minf-routing-indices-replay
Open

lauradang wants to merge 65 commits into
NVIDIA-NeMo:mainfrom
lauradang:laurad/minf-routing-indices-replay

Conversation

@lauradang

@lauradang lauradang commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • enable router replay with Megatron generation on the Gym token-capture path
  • join MInf's native [T-1, L, K] route payload with Gym's CaptureAdmission in the serving worker
  • validate and delta-align routes before synchronously staging the canonical call row in TransferQueue
  • return only ng_commit_coords with the HTTP completion, then commit a token-free CallRecord to Gym's lineage ledger
  • verify and linearize staged calls in the CPU reassembler before publishing [B, S, L, K] routes to the canonical training batch

Builds on #4129 and NVIDIA/Megatron-LM#7015 (both merged). Reviewed on the fork as lauradang#2.

Flow

flowchart LR
    subgraph Gym["Gym model server"]
        Admit["resolve parent<br/>create CaptureAdmission"]
        Ledger["CaptureLedger<br/>token-free CallRecord"]
    end

    subgraph Serve["MInf serving worker"]
        Forward["prepare admitted request<br/>MoE forward records top-k IDs"]
        Native["OffloadedRequestPayload<br/>tokens + logprobs<br/>routes: [T-1, L, K]"]
        Stage["TQMegatronTokenStager<br/>validate T-1 and (L, K)<br/>append -1<br/>slice from prev_len"]
    end

    subgraph TQ["TransferQueue"]
        Staging["canonical staging row<br/>key: rollout_id/model_call_id<br/>routes: [delta_len, L, K]"]
        Training["canonical training batch<br/>routed_experts: [B, S, L, K]"]
    end

    Reply["HTTP response<br/>content + ng_commit_coords<br/>no token or route payload"]

    subgraph Finalizer["NeMo-RL CPU reassembler"]
        Verify["fetch staged rows<br/>verify digests and lineage<br/>linearize call chain"]
        Assemble["assemble call route spans<br/>pad the rollout group"]
    end

    Trainer["Megatron trainer<br/>force expert IDs<br/>compute current router scores"]

    Admit -->|"request + offload_params.ng_capture"| Forward
    Forward -. "preparer reads staging_chain prefix" .- Staging
    Forward --> Native --> Stage
    Admit -. "prev_len + lineage" .-> Stage
    Stage -->|"synchronous durable write"| Staging
    Stage -->|"staged, or capture_failed coords"| Reply
    Reply -->|"validate coordinates; strip transport fields"| Ledger
    Ledger -->|"RolloutReceipt manifest"| Verify
    Staging --> Verify --> Assemble --> Training --> Trainer
Loading

The conversion is owned by NeMo-RL through an adapter installed into the MInf
engine, but it runs in the serving worker. Gym forwards the lineage admission,
including prev_len, in offload_params, so routes are already canonical and
delta-aligned when they enter TransferQueue. Before admission, the prompt
preparer in the same worker splices in the prior turns' tokens from the staged
rows named by the admission's staging_chain. Neither the route payload nor
token arrays return in the HTTP response. If staging fails, the reply still
carries ng_commit_coords, marked capture_failed, so Gym records the call as
poisoned.

Test plan

  • Ruff check and format check pass for all touched Python files
  • On macOS against this branch merged with main: tests/unit/single_controller/test_setup.py (141 passed), tests/unit/algorithms/test_grpo.py (210 passed), tests/unit/data_plane/test_rollout_reassembler.py (14 passed; the MInf stager test needs megatron.core, which only the Linux CI environment has)
  • tests/unit/models/megatron/test_router_replay.py, tests/unit/data_plane/test_tq_token_sink.py, and tests/unit/models/generation/test_megatron_token_capture_hosting.py need Megatron/Linux and are covered in CI

🤖 Generated with Claude Code

lauradang and others added 30 commits September 14, 2026 13:09
Signed-off-by: Laura Dang <laurad@nvidia.com>
…very.sh

Signed-off-by: Laura Dang <lauradang.2000@gmail.com>
Signed-off-by: Laura Dang <lauradang.2000@gmail.com>
Signed-off-by: Laura Dang <lauradang.2000@gmail.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
… cache

Extract the vLLM worker's _fetch_chain_prefix and _resolve_admission_prefix
into tq_token_sink.py as ChainPrefixCache and resolve_admission_prefix, and
make the worker methods one-line delegates. TQMegatronPromptPreparer now
resolves staging chains through the same pair, so both backends share one
cached TQ read (256 entries keyed by the chain's last staging key, deepest
cached key bounds the fetch to the uncached suffix).

The preparer reads the splice boundary from the request-metadata keys the
Megatron chat endpoint writes (prefix_splice_suffix_token_ids,
prefix_splice_boundary_token_id). The key strings are spelled out here rather
than imported so this module stays importable in the vLLM worker and
finalizer environments; a test asserts parity with Megatron's constants when
Megatron is importable.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
- Declare the nemo_gym extra on MegatronPolicyWorker in actor_environments.py
  (the venv source of truth) instead of swapping ACTOR_ENVIRONMENT_REGISTRY at
  runtime, which the prebuilt container venv ignored; drop PY_EXECUTABLES.MCORE_GYM
  and the ModuleNotFoundError string-match remediation.
- Gate backend=megatron token capture at setup on the MInf capture hook
  protocols (RequestPayloadStager / RequestPromptPreparer) from Megatron-LM
  PR #7015, failing with a NotImplementedError that names the dependency while
  the Megatron-Bridge pin predates it.
- Point the Gym submodule at Gym PR NVIDIA-NeMo#2823 (fc08bf19), which is reachable from
  NVIDIA-NeMo/Gym, fast-forwards from Gym main, and nests ng_capture under
  request_metadata as the Megatron endpoint requires.
- Skip test_prefix_splice_keys_match_megatron_constants when the pinned
  megatron-core lacks the constants, and importorskip megatron.core in the
  Megatron hosting test, so the Nemo_Gym shard skips instead of aborting.
- Restore the ++ Hydra override for checkpointing.save_data_plane in the
  streaming recovery script (the key is absent from the Gym config chain).
- Document the Megatron-LM #7015 dependency in the design doc.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
- Add design-docs/token-capture-ledger.md to the docs toctree and drop
  links to guides that do not exist yet (Sphinx treats both as errors).
- Read vllm_cfg / mcore_generation_config through a dict cast in the
  token-capture validation so pyrefly does not reject the TypedDict keys.
- Apply ruff formatting to megatron_worker.py and test_checkpointing.py.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…anup

Stamp a Megatron request that spans a refit with its admission (oldest)
policy epoch instead of failing capture and masking the whole rollout.
This matches vLLM, which freezes the version at begin_call, and the
finalizer's min-over-calls group tag; spans are counted
(epoch_span_count) and logged at WARNING.

Also fold in the mechanical review items: drop the undefined
rollout_max_attempts_to_avoid_lp_nan knob, remove the dead KeyError
branch in receipt parsing, route fetch_for_finalization through
TQStagingStore, collapse the duplicate prefix predicate, call
set_generation_epoch directly, fix import order, revert diff churn, fix
the FIFO docstring and the manifest control route in the design doc,
and remove the unused logical_request_id test parameter.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…dapter

TQMegatronTokenStager now hands the offloaded payload to
RolloutTokenCapture.complete_call_from_response with Gym's
MegatronCaptureAdapter instead of extracting fields by hand. Malformed
payloads poison the call with capture_failed coordinates, which Gym
records as worker_capture_failed (matching vLLM) rather than
worker_response_missing_commit_coordinates.

Requires the Gym-side adapter from NVIDIA-NeMo/Gym#2823.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
- add parent_chain_hash to the inline-prefix token_in admission
- drop logical_request_id from _manifest_record (no such CallRecord field)

Signed-off-by: Laura Dang <laurad@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- reword _assemble_receipt docstring so the named failure reasons are
  examples, not an exhaustive list; add unresolved_parent and the
  capture_failed fallback for reason-less rows
- terminal_selection is left unchanged: the pinned Gym RolloutReceipt
  (fc08bf19) types it Literal["declared","response_id","content",
  "heuristic"] with no None, so the invalid_manifest_row path cannot
  report "no heuristic ran" without a Gym-side schema change

Signed-off-by: Laura Dang <laurad@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…rics code

Signed-off-by: Laura Dang <laurad@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…wiring

Signed-off-by: Laura Dang <laurad@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…fload_params

Megatron Inference renamed the opaque per-request dict it forwards to the
payload stager and prompt preparer (NVIDIA/Megatron-LM PR #7015); follow it
on the TQMegatronPromptPreparer / TQMegatronTokenStager side.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…_prefix_tokens

The Megatron chat endpoint now ships the chat-template render through the
last assistant message (template_prefix_token_ids) and the EOS id in
offload_params instead of a precomputed boundary (NVIDIA/Megatron-LM
PR #7015). TQMegatronPromptPreparer feeds those, plus the prefix it resolved
from the staging chain, to the same replace_prefix_tokens the vLLM worker
uses, so one splicing algorithm serves both backends.

replace_prefix_tokens gains a keyword-only eos_token_id override so callers
that hold only token ids can use it without a tokenizer; the boundary
failure message skips the detokenized reprs in that case. Existing callers
are unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- bump Gym to 37dc751f (RolloutReceipt.terminal_selection is now Optional)
- _assemble_receipt no longer pre-labels parse failures as heuristic
- finalizer metric loop skips the None member of the Optional annotation

Signed-off-by: Laura Dang <laurad@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Keep the Gym submodule at 37dc751f (Gym NVIDIA-NeMo#2823 head); main's pin 267305e2 is an ancestor.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
lauradang and others added 20 commits September 22, 2026 22:23
Co-authored-by: Terry Kong <terrycurtiskong@gmail.com>
Signed-off-by: Laura Dang <lauradang.2000@gmail.com>
Co-authored-by: Terry Kong <terrycurtiskong@gmail.com>
Signed-off-by: Laura Dang <lauradang.2000@gmail.com>
Co-authored-by: Terry Kong <terrycurtiskong@gmail.com>
Signed-off-by: Laura Dang <lauradang.2000@gmail.com>
Co-authored-by: Terry Kong <terrycurtiskong@gmail.com>
Signed-off-by: Laura Dang <lauradang.2000@gmail.com>
…e hosting

- vLLM worker: pass the resolved prefix straight to Gym's begin_call and
  drop the duplicated staging_chain / prev_len checks; tests match Gym's
  error text and read the prefix off the ActiveCall.
- Delete the driver-side #7015 hook gate (_require_minf_capture_hooks) and
  its tests; pin the requirement as an mcore-lane unit test that fails,
  rather than skips, if a Megatron-Bridge bump drops the engine hooks.
- Import the prefix-splice field names from megatron-core instead of
  keeping copies in tq_token_sink.
- Remove the unused LOGGER / logging import in megatron_generation.
- Collect test_interfaces.py in the three vLLM L0 lanes.
- Add a backend parity test that drives both the vLLM and the Megatron
  capture glue through Gym's worked_example rollout (steady and mid-request
  refit) and requires byte-identical staged rows and commit coordinates.
- Sibling-recovery functional test gates phase 2 on finalize/invalid_row_rate
  and finalize/capture_poisoned_rollouts staying at zero.
- New SingleController nightly sibling of the Megatron-inference async Gym
  recipe with token_capture.enabled=true and the same finalizer gates.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
The InferenceClient's ZMQ socket is read by its listener task on the
inference loop thread, and ZMQ sockets are not thread safe. Marshal
set_generation_epoch onto that loop via run_coroutine_threadsafe, the
same way _sleep()/_wake() already do.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
- Replace the stale router-replay rejection test with an accept-path test
  that pins the narrowed setup guard end to end
- Drop dead router_replay lines from the mocked gym setup scenario
- Document _router_replay_enabled in the MegatronGenerationMixin contract
- Add a stager -> TQ -> reassembler test for MInf [T-1, L, K] routes fed
  as numpy int16, asserting the published tensor and sentinel fraction
- Pin the require_routed_experts raise reason via caplog
- Parametrize the stager passthrough test over both flag values
- Remove the dtype check duplicated by encode_routed_experts
- Point the router-replay guide at the SingleController entrypoint and
  fix the MTP wording

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
_assemble_receipt validated the whole manifest in one list comprehension
and shipped every raw row regardless. With one malformed row the
finalizer's RolloutReceipt.model_validate failed on that same row and
rejected the receipt as invalid_receipt with no staging keys, so the
good rows' staged TQ entries leaked until the restart sweep and the
terminal_selection=None path was unreachable.

Validate row by row, mask the rollout as invalid_manifest_row, and ship
only the rows that parsed. The finalizer now rejects as
rollout_failed:invalid_manifest_row with the good rows' staging keys.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
TQMegatronPromptPreparer and TQMegatronTokenStager lived in
nemo_rl/data_plane/tq_token_sink.py and pulled in
nemo_rl.models.generation.openai_server_utils at module scope. That made
every importer of tq_token_sink, including the finalizer actor, execute
nemo_rl/models/generation/__init__.py and load the vLLM and TRT-LLM
config modules.

Move both classes to nemo_rl/models/generation/megatron/token_capture.py,
mirroring how the vLLM backend keeps its capture glue in
vllm_worker_async.py and imports only TQTokenSink/TQTokenSource from
data_plane. Drop the local MegatronPayloadStageResult copy in favor of
Megatron-LM's RequestPayloadStageResult, imported inside the method like
RequestPromptPreparationResult already is. ChainPrefixCache and
resolve_admission_prefix stay in tq_token_sink since both backends use
them. Importing rollout_reassembler no longer loads
nemo_rl.models.generation.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Resolve the uv.lock conflict by taking main's lockfile and relocking with
the Dockerfile-pinned uv (0.11.28). The remaining lock deltas against main
(hydra-core floor, flashinfer 0.6.18.post1 on the flashinfer index, and
its cutlass-dsl transitive) come from this branch's Megatron-Bridge bump.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…registry

- validate_router_replay_config rejects Megatron generation with inference
  pipeline_model_parallel_size != 1 (resolved through
  merged_inference_megatron_cfg, so inherited training PP is caught) and
  with mcore_generation_config.async_sched_mode=async, both of which
  Megatron-Core cannot serve with routing replay
- Add reset_global_router_replay_instances_for_model and call it in
  _initialize_inference_engine before DynamicInferenceEngine is built:
  setup_reference_model_state clears RouterReplay's process-wide registry
  (or a reshard leaves both models' routers in it), so colocated MInf
  otherwise records no routes and every capture fails
- TQMegatronTokenStager takes the model-owned (L, K) and rejects a
  mismatched route payload at the first request instead of a rollout later
- Document the PP=1 / legacy-scheduler requirements in the router replay
  guide; add unit tests for the gates, the registry reset, and the dims check

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…ices-replay

Port the MInf route staging (delta-aligned [T-1, L, K] routes, strict
require_routed_experts, expected (L, K) check, fail_call poisoning) onto
the relocated nemo_rl/models/generation/megatron/token_capture.py and pick
up base's Megatron-Bridge 4d472695d / Automodel pins and lockfile.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…ttern

Nemotron-H style hybrid stacks (e.g. Nano 3.5: 52 layers, 23 MoE) place MoE
layers by the 'E' symbol in hybrid_layer_pattern; moe_layer_freq stays at
its default of 1. _global_moe_layer_numbers therefore predicted an MoE
layer at every position, so router_replay_dimensions reported (52, K) and
MInf's [T, 23, K] payload was rejected as a layout mismatch (the vLLM
full-layer layout happened to match total_layers). Read the pattern when
present, ignoring pipeline separators and MTP depths.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Replace the remaining bare "ng_capture" / "ng_commit_coords" literals in
TQMegatronPromptPreparer.prepare_prompt and TQMegatronTokenStager with
NG_CAPTURE_FIELD / NG_COMMIT_COORDS_FIELD exported by
nemo_gym.token_id_capture, so a Gym rename fails at import time instead
of silently declining every request.

Addresses NVIDIA-NeMo#4129 (comment)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
CI's lint job fails when a file type-checks clean but is absent from
pyrefly.toml's project-includes; the new Megatron capture module was
missing, which failed Lint check and skipped every downstream test job
on the previous PR head.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Conflict resolutions:
- Megatron-Bridge submodule: take main's 1f8873bb (NVIDIA-NeMo#4139); it already contains
  the PR's 4d472695 gather_output fix.
- community_import.py / test_community_import.py: take main, which removed the
  _prefer_nvrx_for_dist_ckpt_save shim the PR had extended with a signature check.
- megatron_policy_worker.py: take main's unconditional FileSystemWriterAsync
  import and direct cleanup_tensor_caches() call.
- L1_Functional_Tests_SingleController: main split the script into _1/_2/_3;
  the PR's Megatron sibling-recovery variant now lives in _3 next to its vLLM twin.
- pyproject.toml: keep the PR's comment on the flashinfer index URL spelling.
- uv.lock: relocked with uv 0.11.28 (Dockerfile pin); result matches main.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…dget

- test_backend_capture_glue_reproduces_the_gym_worked_example compared coords
  against every CallRecord field except mode/response_id, but CallRecord also
  carries attribution-only fields (admitted_at, output_fingerprint,
  continuation_fingerprint, fingerprint_version) that CommitCoords never has,
  so the lookup raised KeyError on whichever one the set yielded first.
  Compare on the fields the two models share.
- The new token-capture nightly cost 48 GPU-hours and pushed the suite from
  4698 to 4746, over the 4720 cap asserted by
  test_nightly_compute_stays_below_4720_hours. Cut it to 3 steps / 82 min
  (21 GPU-hours) so the suite lands at 4719.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…-replay

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
PR NVIDIA-NeMo#4129 landed on main as squash e073dab, whose tree matches the
628b7bd head merged in the parent commit. Conflicts were resolved by
three-way merging against e073dab so the result is main plus only the
routing-indices-replay changes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…ackend module

The stager moved to nemo_rl.models.generation.megatron.token_capture in
4652e32, but the reassembler test still imported it from
nemo_rl.data_plane.tq_token_sink, so the module failed to collect.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
@copy-pr-bot

copy-pr-bot Bot commented Sep 28, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@github-actions github-actions Bot added the Documentation Improvements or additions to documentation label Sep 28, 2026
@lauradang
lauradang marked this pull request as ready for review September 28, 2026 23:49

@terrykong terrykong left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, Laura. This round addresses the fork review (lauradang#2): 17 of its 21 threads are fixed, including the colocated MInf + KL > 0 stall, and 3 more no longer apply. The last one is the vLLM arange(K) question for @zyzhou5, re-asked inline. The Gym capture API calls are correct against Gym source, and keeping the MInf route alignment separate from vLLM's is the right call. All changed tests pass on a CPU venv and sit in CI lanes that run them.

CI: no unit or functional CI has run yet (copy-pr-bot has not vetted the PR). ruff check and ruff format pass on all changed files.

Most important, in order:

  1. Colocated MInf leaves the policy routers recording: the reference pass can crash when prev-logprobs are skipped.
  2. Hybrid layer count change: vLLM with deferred routes quietly turns router replay off for hybrid MoE models.
  3. async_sched_mode check: an unset key passes, and MCore defaults to async.

Not raised on purpose: router replay on the PPO and distillation entry points fails late for both vLLM and MInf, because only GRPO and SingleController call configure_vllm_for_router_replay (grpo.py, setup.py). That is not new in this PR; it may be worth a tracking issue.

Visual explainer: https://terrykong.github.io/gh-pages-poc/terryk/pr-4307-minf-vs-vllm-routes.html

Generated by Claude Code

Comment thread nemo_rl/models/generation/megatron/megatron_worker.py
Comment thread nemo_rl/models/megatron/router_replay.py
Comment thread nemo_rl/models/megatron/router_replay.py Outdated
Comment thread tests/unit/models/megatron/test_router_replay.py
Comment thread tests/unit/models/generation/test_megatron_token_capture_hosting.py
Comment on lines +132 to +133
its current router select experts. This deliberately differs from vLLM, which
uses an in-range placeholder for its terminal row. All other positions normally

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No action needed from @lauradang in this PR — tracked in #4355, assigned to @zyzhou5.

TL;DR — vLLM fills the last token's route row with experts 0…K-1, and on the SingleController path that row becomes a prompt token of the next turn, so the trainer replays made-up experts there.

@zyzhou5, bringing this over from the fork review (lauradang#2). This line says MInf and vLLM fill the last row differently. MInf uses -1, which means "let the trainer's router pick".

  • vLLM builds that row from arange(K) in pad_and_align_routed_expert_indices.
  • The SingleController vLLM stager keeps the row through experts[prev_len:] (vllm_worker_async.py).
  • In a multi-turn rollout, the parent call's last sampled token is a prompt token of the next turn, so the trainer replays experts 0…K-1 for it instead of using its own router.

Follow-up

Tracked in #4355, which covers the whole fix: pad_and_align_routed_expert_indices writes R3_MISSING_ROUTE_SENTINEL (-1) into the last real token's row, batch-padding rows keep arange(K), plus a unit test. Once that lands, this doc line no longer needs to say MInf "deliberately differs from vLLM".

lauradang and others added 3 commits September 29, 2026 23:05
… layouts

- finish_generation clears the RouterReplay action and static buffers in
  the colocated branch, so a following non-replaying forward (the
  reference pass) cannot record into MInf's fixed-size route buffer.
- Route fragments carry a RouteLayout: the executor accepts either the
  MoE-only layer count (MInf) or the total layer count (vLLM) and keeps
  only the MoE columns, so hybrid models on the deferred vLLM path no
  longer fall back to empty routes.
- Megatron generation with router replay requires an explicit
  async_sched_mode of "legacy"; an unset key is rejected too because
  MCore defaults to async scheduling.
- Tests: the registry reset skips MTP routers, and the worker repoints
  the router registry at the served model before building the engine.

Signed-off-by: Laura Dang <laurad@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…d SC Gym test

The SingleController colocated reshard Gym functional test now runs with
router replay and token capture on, so L1 exercises MInf routing-index
capture, canonical staging, finalizer route assembly, and trainer replay
without adding a test. No small pretrained MoE exists, so the served
model is a tiny random-init Qwen3 MoE built in-script on Qwen3-0.6B's
tokenizer. The reward gate is replaced by full routed-experts row
coverage and zero poisoned rollouts; the engine/trainer parity gates
stay.

Signed-off-by: Laura Dang <laurad@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants