Conversation
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
This was referenced Oct 1, 2026
Draft
Draft
ananthsub
force-pushed
the
ananthsub/partial-ckpt-model-worker-cuts
branch
from
October 1, 2026 19:38
55f0706 to
f99bcb5
Compare
ananthsub
force-pushed
the
ananthsub/partial-ckpt-model-worker-cuts
branch
from
October 1, 2026 21:53
f99bcb5 to
3409be3
Compare
This was referenced Oct 1, 2026
ananthsub
force-pushed
the
ananthsub/partial-ckpt-model-worker-cuts
branch
from
October 1, 2026 22:05
3409be3 to
78bf698
Compare
ananthsub
force-pushed
the
ananthsub/partial-ckpt-model-worker-cuts
branch
2 times, most recently
from
October 2, 2026 19:43
4a6c0f8 to
8c661c4
Compare
ananthsub
force-pushed
the
ananthsub/partial-ckpt-model-worker-cuts
branch
from
October 2, 2026 21:17
8c661c4 to
b6a8734
Compare
Ported from the prefix-recovery work (Gym #3412) onto the re-cut policy model participant. An inference worker that serves generation cuts needs three things from RolloutTokenCapture: - begin_call(generation_cut=, generation_cut_staging_keys=) validates the durable prefix a replacement call continues against its admission: the staging keys, source call, digest, lineage, generated-token count, and that the prefix's policy version is not newer than the current one. - build_prefix_record snapshots a live call without completing it, so a cut can stage the prefix while the request keeps decoding. For a continued call it produces the old prefix plus the new tail, stamped with the oldest contributing policy version. - build_generation_chunk_record stages only newly generated tokens, for workers that flush a long generation in chunks. complete_call now builds its record through build_prefix_record and claims completion only after the record is valid. The wire contract (GenerationCutContinuation on CaptureAdmission) was already in place. Signed-off-by: Ananth Subramaniam <ansubramania@nvidia.com>
pthombre
added this pull request to stack #3962
October 2, 2026 22:50
ananthsub
force-pushed
the
ananthsub/partial-ckpt-model-worker-cuts
branch
from
October 2, 2026 23:32
b6a8734 to
a432062
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed and why
An inference worker that serves generation cuts needs three helpers from
RolloutTokenCapture. This PR adds them:begin_call(generation_cut=..., generation_cut_staging_keys=...)validates the durable prefix a replacement call continues against its admission: the staging keys, source call, digest, lineage, generated-token count, and that the prefix's policy version is not newer than the current one.build_prefix_recordsnapshots a live call without completing it, so a cut can stage the prefix while the request keeps decoding. For a continued call it produces the old prefix plus the new tail, stamped with the oldest contributing policy version.build_generation_chunk_recordstages only newly generated tokens, for workers that flush a long generation in chunks.complete_callnow builds its record throughbuild_prefix_recordand claims completion only after the record is valid.Why: the policy model PR sends cut requests and attaches restored cuts, but a real worker (for example NeMo RL's vLLM worker) could not serve them, because these helpers existed only on the old prefix-recovery branch.
How it works
Where this PR sits in the overall flow
The highlighted part is what this PR adds.
flowchart LR C["Controller<br/>NeMo RL, or rollout collection"] CO["Coordination<br/>prepare, commit, restore, resume, retire"] K["Control plane on every server<br/>phases, lease, storage,<br/>retire: stop, then free"] subgraph G["One participant per Gym server"] E["Environment server<br/>episode steps"] M["Policy model<br/>held responses, generation cuts"] A["Agent<br/>sessions parked at boundaries"] R["Resources server<br/>session state"] end W["Inference worker<br/>stages cut prefixes"] D[("Checkpoint directory<br/>records, then manifest")] L[("Capture ledger<br/>retire and delete from #3938, #3939")] C --> CO --> K K --> E & M & A & R M --> W M --> L G --> D classDef this fill:#fde68a,stroke:#b45309,stroke-width:2px,color:#1f2937 class W thisOne checkpoint, a crash, and the restore, end to end:
This PR
What an inference worker does with these helpers when the policy model asks for a cut, and when a replacement call continues it.
Where this sits in the stack
This is one PR in a stack of draft PRs that re-cut partial-rollout checkpointing onto environment servers. Each PR's base is the branch of the PR before it, so each diff shows only that PR's commits.
The stack is based on the token-capture cleanup PRs #3938 (capture ledger
retireanddelete) and #3939 (complete-recordretireanddelete): #3882's base is #3939's branch. The checkpoint stack uses those operations to free the ledgers of retired attempts and to clear a dead execution's capture files before a restore. The striped lock files of #3937 are independent of the stack.ananthsub/partial-ckpt-core): feat(checkpoint): add the participant control plane, episode steps, and coordinationananthsub/partial-ckpt-policy-model): feat(checkpoint): make policy model servers checkpoint participants with generation cutsananthsub/partial-ckpt-model-worker-cuts): feat(token-capture): add worker staging helpers for generation cuts (this PR)ananthsub/partial-ckpt-environment): feat(checkpoint): continue environment server episodes from their boundariesananthsub/partial-ckpt-resources): feat(checkpoint): add the resources server participant with asynchronous session hooksananthsub/partial-ckpt-verifier-declarations): feat(checkpoint): declare five training verifiers stateless with replayable verificationananthsub/partial-ckpt-agent): feat(checkpoint): add the agent session participant and Simple Agent continuationananthsub/partial-ckpt-e2e): test(checkpoint): add a process-level end-to-end suite driven by coordinationananthsub/partial-ckpt-rollout-collection): feat(checkpoint): checkpoint evaluation runs from rollout collectionananthsub/partial-ckpt-telemetry): feat(checkpoint): spans and metrics for partial-rollout checkpointsananthsub/partial-ckpt-multi-worker): feat(checkpoint): resources, agent, and environment servers with several workersPartial-rollout checkpointing lets a training controller, such as NeMo RL, checkpoint Gym while rollouts are in flight and, after a crash, continue those rollouts from their last safe point instead of starting them over. Checkpointing is off by default; with the
checkpoint:block unset, no server installs any checkpoint routes or behavior.Relationship to the old stack (#2939 to #2946)
None of #2939 to #2946 had these helpers. They are ported from #3412 (prefix token-level recovery), which was built on the old stack; the wire contract is unchanged.
Issue
No tracking issue exists for this re-cut. The design and the mapping from the old stack are described in this PR series, and the old stack's PRs (#2939 to #2946) carry the original discussion.
Validation
Run on this branch, on top of #3939's branch:
RAY_TMPDIR=/tmp .venv/bin/python -m pytest -q -p no:cacheprovider tests/unit_tests/test_checkpoint_*.py: 86 passed.RAY_TMPDIR=/tmp .venv/bin/python -m pytest -q -p no:cacheprovider tests/unit_tests/test_token_capture_*.py tests/unit_tests/test_token_id_capture.py tests/unit_tests/test_base_responses_api_model.py: 545 passed.tests/unit_tests, eight processes, withouttest_opensandbox_compose.py, which needs the optionalopensandboxpackage): 6,507 passed, 2 failed. The two failures, a sandbox retry test and a Slurm script test, fail the same way onmainin this development environment.pre-commit run --files <files changed by this PR>: all hooks passed, and no hook modified a file.Signed-off-byline.Rollout evidence
Compatibility
Additive keyword arguments and new methods on
RolloutTokenCapture; existing callers are unaffected.complete_callproduces the same records as before.