Conversation
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
This was referenced Oct 1, 2026
Draft
ananthsub
force-pushed
the
ananthsub/partial-ckpt-verifier-declarations
branch
from
October 1, 2026 19:38
77bfad7 to
f91e01f
Compare
ananthsub
force-pushed
the
ananthsub/partial-ckpt-verifier-declarations
branch
from
October 1, 2026 21:53
f91e01f to
39f45c2
Compare
This was referenced Oct 1, 2026
ananthsub
force-pushed
the
ananthsub/partial-ckpt-verifier-declarations
branch
from
October 1, 2026 22:05
39f45c2 to
2ef25f4
Compare
ananthsub
force-pushed
the
ananthsub/partial-ckpt-verifier-declarations
branch
from
October 2, 2026 13:22
2ef25f4 to
82f0814
Compare
ananthsub
force-pushed
the
ananthsub/partial-ckpt-verifier-declarations
branch
2 times, most recently
from
October 2, 2026 21:17
31ca329 to
ae94453
Compare
…e verification A resources server that declares nothing is restart-only, so every checkpoint has to retire the rollouts that use it. These five servers keep no session state, and their verification can run again after a crash, so they declare checkpoint_mode = "stateless" and checkpoint_verify = "replay": - code_gen, competitive_coding_challenges, equivalence_llm_judge, and math_with_judge verify each request on its own. - genrm_compare keys its cohorts by prompt, not by session. Its verification must replay: a member waits for its siblings, so waiting on it would deadlock a checkpoint while siblings are still generating. After a crash, the members that had not recorded their reward re-verify and rebuild the cohort. A crash in the moment between a cohort's result and every member recording it leaves the rest waiting for siblings that will not re-verify; they fail after cohort_collection_timeout_s and are retried from input. The declarations are in code, not config: they describe what each server can do. A test reads them from source, because each server's dependencies live in its own environment. Signed-off-by: Ananth Subramaniam <ansubramania@nvidia.com>
pthombre
added this pull request to stack #3962
October 2, 2026 22:50
ananthsub
force-pushed
the
ananthsub/partial-ckpt-verifier-declarations
branch
from
October 2, 2026 23:32
ae94453 to
5f0c222
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed and why
A resources server that declares nothing is restart-only, so every checkpoint has to retire the rollouts that use it. These five training verifiers keep no session state and their verification can run again after a crash, so they declare
checkpoint_mode = "stateless"andcheckpoint_verify = "replay":code_gen,competitive_coding_challenges,equivalence_llm_judge, andmath_with_judgeverify each request on its own.genrm_comparekeys its cohorts by prompt, not by session. Its verification must replay: a member waits for its siblings, so waiting on it would deadlock a checkpoint while siblings are still generating. After a crash, members that had not recorded their reward re-verify and rebuild the cohort. A crash between a cohort's result and every member recording it leaves the remaining members waiting; they fail aftercohort_collection_timeout_sand are retried from input.The declarations are class attributes, not config, because they describe what each server's code can do. The test reads them from source, because each server's dependencies live in its own environment.
How it works
Where this PR sits in the overall flow
The highlighted part is what this PR adds.
flowchart LR C["Controller<br/>NeMo RL, or rollout collection"] CO["Coordination<br/>prepare, commit, restore, resume, retire"] K["Control plane on every server<br/>phases, lease, storage,<br/>retire: stop, then free"] subgraph G["One participant per Gym server"] E["Environment server<br/>episode steps"] M["Policy model<br/>held responses, generation cuts"] A["Agent<br/>sessions parked at boundaries"] R["Resources server<br/>session state"] end W["Inference worker<br/>stages cut prefixes"] D[("Checkpoint directory<br/>records, then manifest")] L[("Capture ledger<br/>retire and delete from #3938, #3939")] C --> CO --> K K --> E & M & A & R M --> W M --> L G --> D V["Five training verifiers<br/>stateless, replayable verify"] R --- V classDef this fill:#fde68a,stroke:#b45309,stroke-width:2px,color:#1f2937 class V thisOne checkpoint, a crash, and the restore, end to end:
This PR
These five servers declare that verification can run again after a crash, so a checkpoint never waits for a slow judge.
Where this sits in the stack
This is one PR in a stack of draft PRs that re-cut partial-rollout checkpointing onto environment servers. Each PR's base is the branch of the PR before it, so each diff shows only that PR's commits.
The stack is based on the token-capture cleanup PRs #3938 (capture ledger
retireanddelete) and #3939 (complete-recordretireanddelete): #3882's base is #3939's branch. The checkpoint stack uses those operations to free the ledgers of retired attempts and to clear a dead execution's capture files before a restore. The striped lock files of #3937 are independent of the stack.ananthsub/partial-ckpt-core): feat(checkpoint): add the participant control plane, episode steps, and coordinationananthsub/partial-ckpt-policy-model): feat(checkpoint): make policy model servers checkpoint participants with generation cutsananthsub/partial-ckpt-model-worker-cuts): feat(token-capture): add worker staging helpers for generation cutsananthsub/partial-ckpt-environment): feat(checkpoint): continue environment server episodes from their boundariesananthsub/partial-ckpt-resources): feat(checkpoint): add the resources server participant with asynchronous session hooksananthsub/partial-ckpt-verifier-declarations): feat(checkpoint): declare five training verifiers stateless with replayable verification (this PR)ananthsub/partial-ckpt-agent): feat(checkpoint): add the agent session participant and Simple Agent continuationananthsub/partial-ckpt-e2e): test(checkpoint): add a process-level end-to-end suite driven by coordinationananthsub/partial-ckpt-rollout-collection): feat(checkpoint): checkpoint evaluation runs from rollout collectionananthsub/partial-ckpt-telemetry): feat(checkpoint): spans and metrics for partial-rollout checkpointsananthsub/partial-ckpt-multi-worker): feat(checkpoint): resources, agent, and environment servers with several workersPartial-rollout checkpointing lets a training controller, such as NeMo RL, checkpoint Gym while rollouts are in flight and, after a crash, continue those rollouts from their last safe point instead of starting them over. Checkpointing is off by default; with the
checkpoint:block unset, no server installs any checkpoint routes or behavior.Relationship to the old stack (#2939 to #2946)
None of #2939 to #2946 declared these servers. The declarations come from #3349 (durable turn-level recovery), where they were config on the agent; here they are class attributes on the resources servers.
Issue
No tracking issue exists for this re-cut. The design and the mapping from the old stack are described in this PR series, and the old stack's PRs (#2939 to #2946) carry the original discussion.
Validation
Run on this branch, on top of #3939's branch:
RAY_TMPDIR=/tmp .venv/bin/python -m pytest -q -p no:cacheprovider tests/unit_tests/test_checkpoint_*.py: 122 passed.RAY_TMPDIR=/tmp .venv/bin/python -m pytest -q -p no:cacheprovider resources_servers/code_gen/tests: 62 passed.RAY_TMPDIR=/tmp .venv/bin/python -m pytest -q -p no:cacheprovider resources_servers/competitive_coding_challenges/tests: 33 passed.RAY_TMPDIR=/tmp .venv/bin/python -m pytest -q -p no:cacheprovider resources_servers/equivalence_llm_judge/tests: 56 passed.RAY_TMPDIR=/tmp .venv/bin/python -m pytest -q -p no:cacheprovider resources_servers/genrm_compare/tests: 194 passed.RAY_TMPDIR=/tmp .venv/bin/python -m pytest -q -p no:cacheprovider resources_servers/math_with_judge/tests: not run in this environment. It needs the server'smath-verify==0.8.0requirement, which this environment does not install. It passed (22) on an earlier version of this branch with that requirement installed.tests/unit_tests, eight processes, withouttest_opensandbox_compose.py, which needs the optionalopensandboxpackage): 6,543 passed, 2 failed. The two failures, a sandbox retry test and a Slurm script test, fail the same way onmainin this development environment.pre-commit run --files <files changed by this PR>: all hooks passed, and no hook modified a file.Signed-off-byline.Rollout evidence
Compatibility
checkpoint:block setsenabled: true. With it on, these five servers no longer block checkpoints, and their/verifymay run again after a crash.