Conversation
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
This was referenced Oct 1, 2026
Draft
Draft
ananthsub
force-pushed
the
ananthsub/partial-ckpt-e2e
branch
from
October 1, 2026 19:38
5f8b5af to
87e1f0c
Compare
This was referenced Oct 1, 2026
This was referenced Oct 1, 2026
ananthsub
force-pushed
the
ananthsub/partial-ckpt-e2e
branch
from
October 1, 2026 21:53
87e1f0c to
207f588
Compare
This was referenced Oct 1, 2026
ananthsub
force-pushed
the
ananthsub/partial-ckpt-e2e
branch
from
October 1, 2026 22:05
207f588 to
afee845
Compare
ananthsub
force-pushed
the
ananthsub/partial-ckpt-e2e
branch
2 times, most recently
from
October 2, 2026 19:43
6a3c667 to
f3ade85
Compare
ananthsub
force-pushed
the
ananthsub/partial-ckpt-e2e
branch
from
October 2, 2026 21:17
f3ade85 to
ca27fd2
Compare
pthombre
added this pull request to stack #3962
October 2, 2026 22:50
Real server processes against a fake inference backend and token store that survive a Gym crash. The fake worker stages generation cuts with the real staging helpers and continues them through begin_call, so the cut record Gym restores is validated exactly as a real worker validates it. Each scenario checkpoints mid-episode through the coordination functions, usually kills every Gym process, restores, and runs the replacement attempt, asserting on the exact model calls: - native and legacy /run episodes continue from their last boundary; - restored resources state (counter) is required for the right reward, with a negative control that skips the resources restore; - token lineage continues across the crash; - an in-flight generation is cut and its prefix continued; - a checkpoint without a crash parks and then releases the episode; - verification in flight is replayed or waited for, per checkpoint_verify; - an episode that finishes during prepare is neither lost nor redone; - a second crash, after a checkpoint taken before the replacement attempt starts, still continues the episode. An opt-in real-model test (NEMO_GYM_CHECKPOINT_VLLM_URL) checkpoints 16 concurrent rollouts against vLLM and checks the controller contract: every rollout either replied before the checkpoint or is exported, none replies while Gym is prepared, and every continued rollout completes after a crash and restore. The native, double-crash, token-lineage, generation-cut, and park scenarios run with the policy model in one process and with two uvicorn workers. A simulated crash kills each server's whole process group, as a node failure takes a server's workers with it. An opt-in scale scenario (NEMO_GYM_CHECKPOINT_SCALE=<rollouts>) checkpoints thousands of in-flight rollouts with token capture and generation cuts, crashes, restores, and checks the controller contract: every rollout either finished before the checkpoint or was exported, and every exported one completes after the restore. It reports per-stage prepare, commit, and restore times. Skipped unless NEMO_GYM_CHECKPOINT_E2E=1; about 3 minutes in total. Signed-off-by: Ananth Subramaniam <ansubramania@nvidia.com>
…de model calls The replacement attempt commits model calls to its capture ledger, Gym crashes before the next checkpoint, and the same checkpoint is restored again. The second run of the attempt continues from the checkpoint's boundary, and none of the dead run's calls are in its lineage. Signed-off-by: Ananth Subramaniam <ansubramania@nvidia.com>
…ng behind A native episode is retired mid-flight with token capture on. Afterwards no server is stopping or tracking anything for it, the model makes no further call for it, and its capture ledger is retired. Signed-off-by: Ananth Subramaniam <ansubramania@nvidia.com>
ananthsub
force-pushed
the
ananthsub/partial-ckpt-e2e
branch
from
October 2, 2026 23:32
ca27fd2 to
f98be7a
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed and why
A process-level end-to-end suite for partial-rollout checkpointing, in
tests/e2e/checkpoint. It starts real Gym server processes (environment server, Simple Agent, resources server, policy model server) against a fake inference backend and token store that survive a Gym crash. The fake worker stages generation cuts with the real staging helpers and continues them throughbegin_call, so the cut record Gym restores is validated exactly as a real worker validates it.Each scenario checkpoints mid-episode through the coordination functions, usually kills every Gym process group, restores into fresh processes, and runs the replacement attempt, asserting on the exact model calls made:
/runepisodes continue from their last boundary;checkpoint_verify;Retire no longer leaves lasting fences, so a stale
/runfor a replaced attempt is not refused withstale_attemptany more. It runs a wasted episode that cannot touch the replacement (a different capture key and different sessions), and its capture rows are discarded because the restore retired that attempt's ledger. Idempotent/run, separate work, will refuse it.The native, double-crash, token-lineage, generation-cut, and park scenarios run with the policy model in one process and with two uvicorn workers.
Two opt-in scenarios are skipped by default:
NEMO_GYM_CHECKPOINT_VLLM_URL) that checkpoints 16 concurrent rollouts against vLLM and checks the controller contract: every rollout either replied before the checkpoint or is exported, none replies while Gym is prepared, and every continued rollout completes after a crash and restore;NEMO_GYM_CHECKPOINT_SCALE=<rollouts>) that checkpoints thousands of in-flight rollouts with token capture and generation cuts, crashes, restores, checks the same contract, and reports per-stage prepare, commit, and restore times.The whole suite is skipped unless
NEMO_GYM_CHECKPOINT_E2E=1is set, so it does not run in the default unit test job. It takes about 6 minutes on this branch.How it works
Where this PR sits in the overall flow
The highlighted part is what this PR adds.
flowchart LR C["Controller<br/>NeMo RL, or rollout collection"] CO["Coordination<br/>prepare, commit, restore, resume, retire"] K["Control plane on every server<br/>phases, lease, storage,<br/>retire: stop, then free"] subgraph G["One participant per Gym server"] E["Environment server<br/>episode steps"] M["Policy model<br/>held responses, generation cuts"] A["Agent<br/>sessions parked at boundaries"] R["Resources server<br/>session state"] end W["Inference worker<br/>stages cut prefixes"] D[("Checkpoint directory<br/>records, then manifest")] L[("Capture ledger<br/>retire and delete from #3938, #3939")] C --> CO --> K K --> E & M & A & R M --> W M --> L G --> D T["Process-level e2e suite<br/>plays the controller"] T --> CO classDef this fill:#fde68a,stroke:#b45309,stroke-width:2px,color:#1f2937 class T thisOne checkpoint, a crash, and the restore, end to end:
This PR
How the suite runs real server processes and crashes them.
Where this sits in the stack
This is one PR in a stack of draft PRs that re-cut partial-rollout checkpointing onto environment servers. Each PR's base is the branch of the PR before it, so each diff shows only that PR's commits.
The stack is based on the token-capture cleanup PRs #3938 (capture ledger
retireanddelete) and #3939 (complete-recordretireanddelete): #3882's base is #3939's branch. The checkpoint stack uses those operations to free the ledgers of retired attempts and to clear a dead execution's capture files before a restore. The striped lock files of #3937 are independent of the stack.ananthsub/partial-ckpt-core): feat(checkpoint): add the participant control plane, episode steps, and coordinationananthsub/partial-ckpt-policy-model): feat(checkpoint): make policy model servers checkpoint participants with generation cutsananthsub/partial-ckpt-model-worker-cuts): feat(token-capture): add worker staging helpers for generation cutsananthsub/partial-ckpt-environment): feat(checkpoint): continue environment server episodes from their boundariesananthsub/partial-ckpt-resources): feat(checkpoint): add the resources server participant with asynchronous session hooksananthsub/partial-ckpt-verifier-declarations): feat(checkpoint): declare five training verifiers stateless with replayable verificationananthsub/partial-ckpt-agent): feat(checkpoint): add the agent session participant and Simple Agent continuationananthsub/partial-ckpt-e2e): test(checkpoint): add a process-level end-to-end suite driven by coordination (this PR)ananthsub/partial-ckpt-rollout-collection): feat(checkpoint): checkpoint evaluation runs from rollout collectionananthsub/partial-ckpt-telemetry): feat(checkpoint): spans and metrics for partial-rollout checkpointsananthsub/partial-ckpt-multi-worker): feat(checkpoint): resources, agent, and environment servers with several workersPartial-rollout checkpointing lets a training controller, such as NeMo RL, checkpoint Gym while rollouts are in flight and, after a crash, continue those rollouts from their last safe point instead of starting them over. Checkpointing is off by default; with the
checkpoint:block unset, no server installs any checkpoint routes or behavior.Relationship to the old stack (#2939 to #2946)
Issue
No tracking issue exists for this re-cut. The design and the mapping from the old stack are described in this PR series, and the old stack's PRs (#2939 to #2946) carry the original discussion.
Validation
Run on this branch, on top of #3939's branch:
NEMO_GYM_CHECKPOINT_E2E=1 RAY_TMPDIR=/tmp .venv/bin/python -m pytest -q -p no:cacheprovider tests/e2e/checkpoint: all 23 collected tests ran, 20 passed and 3 skipped (the opt-in real-model and scale tests), in 5 minutes 31 seconds.capture ledger for again-1-a1 already holds rows from another execution.RAY_TMPDIR=/tmp .venv/bin/python -m pytest -q -p no:cacheprovider tests/unit_tests/test_checkpoint_*.py: 136 passed.tests/unit_tests, eight processes, withouttest_opensandbox_compose.py, which needs the optionalopensandboxpackage): 6,557 passed, 2 failed. The two failures, a sandbox retry test and a Slurm script test, fail the same way onmainin this development environment.pre-commit run --files <files changed by this PR>: all hooks passed, and no hook modified a file.Signed-off-byline.Rollout evidence
Process-level crash and restore: the end-to-end suite above, with a fake inference backend: all scenarios pass.
Scale, on one workstation with every Gym server and the fake inference backend (native
single_agent_turnweather rollouts, token capture and generation cuts on, each rollout's second model call held so most rollouts are in flight at the checkpoint), run withNEMO_GYM_CHECKPOINT_E2E=1 NEMO_GYM_CHECKPOINT_SCALE=<N> pytest tests/e2e/checkpoint -k scale -s:Every rollout either finished before the checkpoint or was exported, and every exported rollout completed after the restore. These numbers come from the lane branches this stack was cut from; the scale test was not rerun on this stack.
Real model: the opt-in vLLM test has not been run on this stack yet: pending. Not covered yet: a real vLLM worker serving generation cuts, and several hosts.
Compatibility
Tests only. The suite is skipped unless
NEMO_GYM_CHECKPOINT_E2E=1is set.