Skip to content

feat(checkpoint): spans and metrics for partial-rollout checkpoints - #3903

Draft
ananthsub wants to merge 1 commit into
ananthsub/partial-ckpt-rollout-collectionfrom
ananthsub/partial-ckpt-telemetry
Draft

ananthsub wants to merge 1 commit into
ananthsub/partial-ckpt-rollout-collectionfrom
ananthsub/partial-ckpt-telemetry

Conversation

@ananthsub

@ananthsub ananthsub commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

What changed and why

Checkpoints emitted nothing. That left three failure modes invisible:

  • a prepare that stalled on one server;
  • a lease that expired and silently aborted a checkpoint;
  • a restore that took minutes.

This PR adds spans and metrics through the existing nemo-lens telemetry, on the same pattern as the sandbox metrics.

  • Spans are in a new checkpoint span group. It is opt-in with default,checkpoint, because the default preset is deliberately coarse.
    • On every participant: gym.checkpoint.<operation> for prepare, commit, restore, resume and retire. Each carries the checkpoint ID, participant kind, server instance, outcome and record count.
    • Children for the steps that grow with the live set: wait_ready, export, write, read, install, and the policy model's generation_cut round, which also records each cut's disposition.
    • In the controller: gym.checkpoint.coordinate.<operation>, with one child per prepare stage, plus gym.checkpoint.collection.checkpoint and gym.checkpoint.collection.restore for rollout collection.
    • One trace per checkpoint: control calls go through Gym's traced HTTP client, so a checkpoint is one trace across every Gym process.
  • Metrics are recorded whenever telemetry exports, independent of span groups:
    • gym.checkpoint.operation_duration_ms, by operation, participant kind and outcome;
    • gym.checkpoint.records_total and gym.checkpoint.bytes_total, for commit and restore;
    • gym.checkpoint.events_total, which counts:
      • lease_expired: a participant resumed on its own, which aborts the checkpoint;
      • prepare_not_ready: a prepare reached its deadline with blockers;
      • refused: a request refused for a checkpoint, by error code;
      • generation_cut: each cut, by disposition.
  • Cost when telemetry is off: none. Every span sits behind the span-group check, and every metric recorder is a no-op.
  • Attribute names avoid key and token, which the span helpers redact.

The spans were used to profile commit and restore at 16,000 rollouts. The fixes that profile led to are in the core, policy model, environment, resources, and agent PRs below (cheaper record payloads and ledger imports).

How it works

Where this PR sits in the overall flow

The highlighted part is what this PR adds.

flowchart LR
  C["Controller<br/>NeMo RL, or rollout collection"]
  CO["Coordination<br/>prepare, commit, restore, resume, retire"]
  K["Control plane on every server<br/>phases, fencing, lease, storage"]
  subgraph G["One participant per Gym server"]
    E["Environment server<br/>episode steps"]
    M["Policy model<br/>held responses, generation cuts"]
    A["Agent<br/>sessions parked at boundaries"]
    R["Resources server<br/>session state"]
  end
  W["Inference worker<br/>stages cut prefixes"]
  D[("Checkpoint directory<br/>records, then manifest")]
  C --> CO --> K
  K --> E & M & A & R
  M --> W
  G --> D
  classDef this fill:#fde68a,stroke:#b45309,stroke-width:2px,color:#1f2937
  class CO,K this
Loading

One checkpoint, a crash, and the restore, end to end:

sequenceDiagram
  participant C as Controller
  participant G as Gym participants
  participant D as Checkpoint directory
  C->>G: prepare, in order environment, model, agent, resources
  Note over G: admission closes, in-flight work parks at a boundary,<br/>undelivered model responses are held
  G-->>C: prepared, or blockers at the deadline
  C->>G: commit with the episodes the controller continues
  G->>D: each participant writes its records, then its manifest
  C->>C: publish the checkpoint with the controller's own state
  C->>G: resume, in order resources, agent, model, environment
  Note over C,G: crash - every Gym process dies
  C->>G: restore the checkpoint in fresh processes, all or nothing
  D-->>G: records installed under attempt + 1
  C->>G: resume
  C->>G: /run as attempt + 1 continues each episode from its boundary
  G-->>C: a late call from attempt 0 gets 409 stale_attempt
Loading

This PR

The spans of one checkpoint. Control calls carry the trace context, so the controller's spans and every participant's spans form one trace across Gym's processes.

flowchart LR
  P["coordinate.prepare"] --> PS["one span per stage<br/>environment, model, agent, resources"]
  PS --> SP["prepare on each participant"] --> WR["wait_ready"]
  SP --> GC["generation_cut<br/>policy model only"]
  C["coordinate.commit"] --> SC["commit on each participant"] --> EX["export"]
  SC --> WRT["write"]
  R["coordinate.restore"] --> SR["restore on each participant"] --> RD["read"]
  SR --> IN["install"]
Loading

Where this sits in the stack

This is one PR in a stack of draft PRs for partial-rollout checkpointing. Each PR's base is the branch of the PR before it, so each diff shows only that PR's commits.

The stack is based on the token-capture cleanup PRs #3938 (capture ledger retire and delete) and #3939 (complete-record retire and delete): #3882's base is #3939's branch. The checkpoint stack uses those operations to free the ledgers of retired attempts and to clear a dead execution's capture files before a restore. The striped lock files of #3937 are independent of the stack.

  1. feat(checkpoint): add the participant control plane, episode steps, and coordination #3882 (ananthsub/partial-ckpt-core): feat(checkpoint): add the participant control plane, episode steps, and coordination
  2. feat(checkpoint): make policy model servers checkpoint participants with generation cuts #3883 (ananthsub/partial-ckpt-policy-model): feat(checkpoint): make policy model servers checkpoint participants with generation cuts
  3. feat(token-capture): add worker staging helpers for generation cuts #3884 (ananthsub/partial-ckpt-model-worker-cuts): feat(token-capture): add worker staging helpers for generation cuts
  4. feat(checkpoint): continue environment server episodes from their boundaries #3885 (ananthsub/partial-ckpt-environment): feat(checkpoint): continue environment server episodes from their boundaries
  5. feat(checkpoint): add the resources server participant with asynchronous session hooks #3886 (ananthsub/partial-ckpt-resources): feat(checkpoint): add the resources server participant with asynchronous session hooks
  6. feat(checkpoint): declare five training verifiers stateless with replayable verification #3887 (ananthsub/partial-ckpt-verifier-declarations): feat(checkpoint): declare five training verifiers stateless with replayable verification
  7. feat(checkpoint): add the agent session participant and Simple Agent continuation #3888 (ananthsub/partial-ckpt-agent): feat(checkpoint): add the agent session participant and Simple Agent continuation
  8. test(checkpoint): add a process-level end-to-end suite driven by coordination #3889 (ananthsub/partial-ckpt-e2e): test(checkpoint): add a process-level end-to-end suite driven by coordination
  9. feat(checkpoint): checkpoint evaluation runs from rollout collection #3893 (ananthsub/partial-ckpt-rollout-collection): feat(checkpoint): checkpoint evaluation runs from rollout collection
  10. feat(checkpoint): spans and metrics for partial-rollout checkpoints #3903 (ananthsub/partial-ckpt-telemetry): feat(checkpoint): spans and metrics for partial-rollout checkpoints (this PR)
  11. feat(checkpoint): resources, agent, and environment servers with several workers #3909 (ananthsub/partial-ckpt-multi-worker): feat(checkpoint): resources, agent, and environment servers with several workers

Server ports that build on the stack: #3894 (Workplace Assistant), #3895 (Gymnasium), #3896 (Blackjack), #3897 (indirect prompt injection), #3898 (proof refinement). Session routing for several workers is #3905, against main.

Issue

No tracking issue exists. Observability was a planned follow-up for the stack.

Validation

  • pytest tests/unit_tests/telemetry tests/unit_tests/test_checkpoint_*.py: 328 passed.
  • tests/unit_tests/telemetry/test_checkpoint_telemetry.py drives a real resources participant through its control routes and asserts against a real in-memory span exporter and metric reader:
    • every operation span, with export and write nested under commit and read and install under restore;
    • durations, records and bytes;
    • a refusal counted by its error code, and a failed operation recorded with outcome=error;
    • a lease expiry counted;
    • one coordination span per prepare stage.
  • The span-group tests now include checkpoint; the default preset is unchanged.
  • The process-level e2e suite passes on this branch.
  • pre-commit run --files <changed files>: passed.

Rollout evidence

The 16,000-rollout scale test was profiled with the console exporter. The per-participant breakdown is in the next PR. Real-model runs are pending, as for the rest of the stack.

Compatibility

  • New span group: checkpoint, opt-in. Existing presets are unchanged.
  • New metric instruments under gym.checkpoint.*.
  • Nothing changes when telemetry is off.

@ananthsub ananthsub added feature New capabilities, enhancements, or enablement work area:core Shared APIs, servers, telemetry, health, and registries labels Oct 1, 2026
@copy-pr-bot

copy-pr-bot Bot commented Oct 1, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@ananthsub
ananthsub force-pushed the ananthsub/partial-ckpt-telemetry branch from fd8308c to 9775d38 Compare October 1, 2026 22:05
@ananthsub
ananthsub force-pushed the ananthsub/partial-ckpt-rollout-collection branch from 324d650 to 29a2feb Compare October 1, 2026 22:05
@ananthsub
ananthsub force-pushed the ananthsub/partial-ckpt-rollout-collection branch from 29a2feb to d1bac74 Compare October 2, 2026 13:22
@ananthsub
ananthsub force-pushed the ananthsub/partial-ckpt-telemetry branch 2 times, most recently from dc49274 to f9affd7 Compare October 2, 2026 19:43
@ananthsub
ananthsub force-pushed the ananthsub/partial-ckpt-rollout-collection branch from d1bac74 to 37f360b Compare October 2, 2026 19:43
Checkpoints emitted nothing: a prepare that stalled, a lease that expired and
silently aborted a checkpoint, or a restore that took minutes left no trace.

Spans, in a new opt-in `checkpoint` span group (`default,checkpoint`):
- gym.checkpoint.<operation> on every participant for prepare, commit,
  restore, resume, and retire, with the participant kind, checkpoint ID,
  outcome, and record counts;
- child spans for the steps that grow with the live set: wait_ready, export,
  write, read, install, and the generation-cut round;
- gym.checkpoint.coordinate.<operation> in the controller, with one child per
  prepare stage, and spans for rollout collection's own checkpoint and
  restore. Control calls carry the trace context, so one checkpoint is one
  trace across Gym's processes.

Metrics, recorded whenever telemetry exports, on Gym-owned instruments like
the sandbox ones: gym.checkpoint.operation_duration_ms (by operation, kind,
and outcome), gym.checkpoint.records_total and bytes_total (commit and
restore), and gym.checkpoint.events_total for lease_expired,
prepare_not_ready, refused (by error code), and generation_cut (by
disposition).

Signed-off-by: Ananth Subramaniam <ansubramania@nvidia.com>
@ananthsub
ananthsub force-pushed the ananthsub/partial-ckpt-telemetry branch from f9affd7 to f8f2b11 Compare October 2, 2026 21:17

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:core Shared APIs, servers, telemetry, health, and registries feature New capabilities, enhancements, or enablement work

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant