Conversation
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
This was referenced Oct 1, 2026
Draft
Draft
ananthsub
force-pushed
the
ananthsub/partial-ckpt-telemetry
branch
from
October 1, 2026 22:05
fd8308c to
9775d38
Compare
ananthsub
force-pushed
the
ananthsub/partial-ckpt-rollout-collection
branch
from
October 1, 2026 22:05
324d650 to
29a2feb
Compare
ananthsub
force-pushed
the
ananthsub/partial-ckpt-rollout-collection
branch
from
October 2, 2026 13:22
29a2feb to
d1bac74
Compare
ananthsub
force-pushed
the
ananthsub/partial-ckpt-telemetry
branch
2 times, most recently
from
October 2, 2026 19:43
dc49274 to
f9affd7
Compare
ananthsub
force-pushed
the
ananthsub/partial-ckpt-rollout-collection
branch
from
October 2, 2026 19:43
d1bac74 to
37f360b
Compare
Checkpoints emitted nothing: a prepare that stalled, a lease that expired and silently aborted a checkpoint, or a restore that took minutes left no trace. Spans, in a new opt-in `checkpoint` span group (`default,checkpoint`): - gym.checkpoint.<operation> on every participant for prepare, commit, restore, resume, and retire, with the participant kind, checkpoint ID, outcome, and record counts; - child spans for the steps that grow with the live set: wait_ready, export, write, read, install, and the generation-cut round; - gym.checkpoint.coordinate.<operation> in the controller, with one child per prepare stage, and spans for rollout collection's own checkpoint and restore. Control calls carry the trace context, so one checkpoint is one trace across Gym's processes. Metrics, recorded whenever telemetry exports, on Gym-owned instruments like the sandbox ones: gym.checkpoint.operation_duration_ms (by operation, kind, and outcome), gym.checkpoint.records_total and bytes_total (commit and restore), and gym.checkpoint.events_total for lease_expired, prepare_not_ready, refused (by error code), and generation_cut (by disposition). Signed-off-by: Ananth Subramaniam <ansubramania@nvidia.com>
ananthsub
force-pushed
the
ananthsub/partial-ckpt-telemetry
branch
from
October 2, 2026 21:17
f9affd7 to
f8f2b11
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed and why
Checkpoints emitted nothing. That left three failure modes invisible:
This PR adds spans and metrics through the existing nemo-lens telemetry, on the same pattern as the sandbox metrics.
checkpointspan group. It is opt-in withdefault,checkpoint, because thedefaultpreset is deliberately coarse.gym.checkpoint.<operation>forprepare,commit,restore,resumeandretire. Each carries the checkpoint ID, participant kind, server instance, outcome and record count.wait_ready,export,write,read,install, and the policy model'sgeneration_cutround, which also records each cut's disposition.gym.checkpoint.coordinate.<operation>, with one child per prepare stage, plusgym.checkpoint.collection.checkpointandgym.checkpoint.collection.restorefor rollout collection.gym.checkpoint.operation_duration_ms, by operation, participant kind and outcome;gym.checkpoint.records_totalandgym.checkpoint.bytes_total, for commit and restore;gym.checkpoint.events_total, which counts:lease_expired: a participant resumed on its own, which aborts the checkpoint;prepare_not_ready: a prepare reached its deadline with blockers;refused: a request refused for a checkpoint, by error code;generation_cut: each cut, by disposition.keyandtoken, which the span helpers redact.The spans were used to profile commit and restore at 16,000 rollouts. The fixes that profile led to are in the core, policy model, environment, resources, and agent PRs below (cheaper record payloads and ledger imports).
How it works
Where this PR sits in the overall flow
The highlighted part is what this PR adds.
flowchart LR C["Controller<br/>NeMo RL, or rollout collection"] CO["Coordination<br/>prepare, commit, restore, resume, retire"] K["Control plane on every server<br/>phases, fencing, lease, storage"] subgraph G["One participant per Gym server"] E["Environment server<br/>episode steps"] M["Policy model<br/>held responses, generation cuts"] A["Agent<br/>sessions parked at boundaries"] R["Resources server<br/>session state"] end W["Inference worker<br/>stages cut prefixes"] D[("Checkpoint directory<br/>records, then manifest")] C --> CO --> K K --> E & M & A & R M --> W G --> D classDef this fill:#fde68a,stroke:#b45309,stroke-width:2px,color:#1f2937 class CO,K thisOne checkpoint, a crash, and the restore, end to end:
This PR
The spans of one checkpoint. Control calls carry the trace context, so the controller's spans and every participant's spans form one trace across Gym's processes.
Where this sits in the stack
This is one PR in a stack of draft PRs for partial-rollout checkpointing. Each PR's base is the branch of the PR before it, so each diff shows only that PR's commits.
The stack is based on the token-capture cleanup PRs #3938 (capture ledger
retireanddelete) and #3939 (complete-recordretireanddelete): #3882's base is #3939's branch. The checkpoint stack uses those operations to free the ledgers of retired attempts and to clear a dead execution's capture files before a restore. The striped lock files of #3937 are independent of the stack.ananthsub/partial-ckpt-core): feat(checkpoint): add the participant control plane, episode steps, and coordinationananthsub/partial-ckpt-policy-model): feat(checkpoint): make policy model servers checkpoint participants with generation cutsananthsub/partial-ckpt-model-worker-cuts): feat(token-capture): add worker staging helpers for generation cutsananthsub/partial-ckpt-environment): feat(checkpoint): continue environment server episodes from their boundariesananthsub/partial-ckpt-resources): feat(checkpoint): add the resources server participant with asynchronous session hooksananthsub/partial-ckpt-verifier-declarations): feat(checkpoint): declare five training verifiers stateless with replayable verificationananthsub/partial-ckpt-agent): feat(checkpoint): add the agent session participant and Simple Agent continuationananthsub/partial-ckpt-e2e): test(checkpoint): add a process-level end-to-end suite driven by coordinationananthsub/partial-ckpt-rollout-collection): feat(checkpoint): checkpoint evaluation runs from rollout collectionananthsub/partial-ckpt-telemetry): feat(checkpoint): spans and metrics for partial-rollout checkpoints (this PR)ananthsub/partial-ckpt-multi-worker): feat(checkpoint): resources, agent, and environment servers with several workersServer ports that build on the stack: #3894 (Workplace Assistant), #3895 (Gymnasium), #3896 (Blackjack), #3897 (indirect prompt injection), #3898 (proof refinement). Session routing for several workers is #3905, against main.
Issue
No tracking issue exists. Observability was a planned follow-up for the stack.
Validation
pytest tests/unit_tests/telemetry tests/unit_tests/test_checkpoint_*.py: 328 passed.tests/unit_tests/telemetry/test_checkpoint_telemetry.pydrives a real resources participant through its control routes and asserts against a real in-memory span exporter and metric reader:outcome=error;checkpoint; thedefaultpreset is unchanged.pre-commit run --files <changed files>: passed.Rollout evidence
The 16,000-rollout scale test was profiled with the console exporter. The per-participant breakdown is in the next PR. Real-model runs are pending, as for the rest of the stack.
Compatibility
checkpoint, opt-in. Existing presets are unchanged.gym.checkpoint.*.