Skip to content

feat(checkpoint): add the resources server participant with asynchronous session hooks - #3886

Draft
ananthsub wants to merge 6 commits into
ananthsub/partial-ckpt-environmentfrom
ananthsub/partial-ckpt-resources
Draft

ananthsub wants to merge 6 commits into
ananthsub/partial-ckpt-environmentfrom
ananthsub/partial-ckpt-resources

Conversation

@ananthsub

@ananthsub ananthsub commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

What changed and why

Every resources server takes part in checkpoints once checkpointing is on, in one of three modes, declared with the checkpoint_mode class attribute:

  • restart_only (the default): live sessions block a checkpoint until their rollouts are retired, so a server without support fails closed.
  • stateless: tool results depend only on the request, so nothing is exported. The weather example server uses it.
  • exported: the server implements the session hooks. The counter example server (example_session_state_mgmt) uses it.

The session hooks are coroutines, and export takes every session at once: export_session_states(session_ids), restore_session_states(states), and retire_session_state(session_id). A server whose session state lives outside the process, such as a sandbox per session, can snapshot all of it concurrently within the commit. A session missing from the export result is treated as already dropped (for example, after a failed verification) and is no longer tracked, instead of failing the commit.

The session lifecycle comes from /seed_session, /close_session, and /verify, keyed by the rollout attempt, so no request body is inspected. Servers that create sessions elsewhere (Gymnasium-style servers) call checkpoint_session_started and checkpoint_session_ended.

The checkpoint_verify class attribute declares whether /verify may run again after a crash (replay) or must finish before a checkpoint (wait, the default). The server reports it in the x-ng-checkpoint-verify header on every /seed_session reply and in its checkpoint status. Replay-safe requests are admitted even after admission closes, because the checkpoint does not wait for the episode steps that send them.

Tools called over MCP (expose_tools_over_mcp: true). A CLI agent harness calls tools over MCP with a signed session token and no session cookie. Before this change, such a server failed at startup with checkpointing on: MCP auto-exposure refuses middleware that direct MCP dispatch would skip, and the checkpoint middleware was one of them.

  • The checkpoint middleware works on the whole MCP request, so MCP auto-exposure accepts it. Admission and in-flight counting apply to the one POST that carries each tool call.
  • The middleware reads the session from the MCP token, so a retire drains and refuses MCP calls like any other request on the session. The token key is deterministic, so a harness keeps its token across a restore.
  • A tool call made while a checkpoint is open waits for resume instead of getting a 409. The harness is third-party code that would pass the refusal to the model as a tool error. A waiting call is not in flight, so it does not hold up prepare.
  • Checkpoint control routes are no longer harvested as MCP tools.

Ending sessions.

  • /close_session is never refused. Closing is how a server releases a session's state. A close that arrives while a checkpoint is open, from an episode that ended during it, waits for resume as an MCP call does; a refusal would lose the release, since the environment server closes once.
  • A retire refuses the attempt's sessions while it stops them, waits until no request is in flight on them (over HTTP or MCP), then releases them through retire_session_state for exported servers. Nothing about a retired session remains afterwards.
  • Restored sessions no request has used yet are reported as pending, so a commit whose scope no longer continues their episode retires them.

How it works

Where this PR sits in the overall flow

The highlighted part is what this PR adds.

flowchart LR
  C["Controller<br/>NeMo RL, or rollout collection"]
  CO["Coordination<br/>prepare, commit, restore, resume, retire"]
  K["Control plane on every server<br/>phases, lease, storage,<br/>retire: stop, then free"]
  subgraph G["One participant per Gym server"]
    E["Environment server<br/>episode steps"]
    M["Policy model<br/>held responses, generation cuts"]
    A["Agent<br/>sessions parked at boundaries"]
    R["Resources server<br/>session state"]
  end
  W["Inference worker<br/>stages cut prefixes"]
  D[("Checkpoint directory<br/>records, then manifest")]
  L[("Capture ledger<br/>retire and delete from #3938, #3939")]
  C --> CO --> K
  K --> E & M & A & R
  M --> W
  M --> L
  G --> D
  classDef this fill:#fde68a,stroke:#b45309,stroke-width:2px,color:#1f2937
  class R this
Loading

One checkpoint, a crash, and the restore, end to end:

sequenceDiagram
  participant C as Controller
  participant G as Gym participants
  participant D as Checkpoint directory
  C->>G: prepare, in order environment, model, agent, resources
  Note over G: admission closes, in-flight work parks at a boundary,<br/>undelivered model responses are held
  G-->>C: prepared, or blockers at the deadline
  C->>G: commit with the episodes the controller continues
  G->>D: each participant writes its records, then its manifest
  Note over G: restored state the commit's scope leaves out is released
  C->>C: publish the checkpoint with the controller's own state
  C->>G: resume, in order resources, agent, model, environment
  Note over C,G: crash - every Gym process dies
  C->>G: restore the checkpoint in fresh processes, all or nothing
  D-->>G: records installed under attempt + 1, attempt N's capture ledger retired
  C->>G: resume
  C->>G: /run as attempt + 1 continues each episode from its boundary
  Note over C,G: dropping an episode, only while no checkpoint is open
  C->>G: retire - environment, then agent, then model and resources
  Note over G: each server stops the attempt's work, waits, frees its state, then replies
Loading

This PR

How a resources server takes part, chosen by a class attribute.

flowchart TD
  M{"checkpoint_mode"} -->|stateless| S["Nothing to export.<br/>Never blocks a checkpoint."]
  M -->|exported| X["export_session_states at commit<br/>restore_session_states after a crash<br/>retire_session_state on retire"]
  M -->|"restart_only, the default"| O["A live session blocks prepare<br/>until its rollout is retired"]
Loading

A session's lifecycle, as the participant learns it from the routes that already define it.

sequenceDiagram
  participant E as Environment server or agent
  participant R as Resources server
  E->>R: /seed_session for rollout r
  R-->>E: session cookie, plus x-ng-checkpoint-verify - wait or replay
  E->>R: tool calls - refused while a checkpoint is open, except over MCP, which waits
  E->>R: /verify, then /close_session from final cleanup, which waits for resume if a checkpoint is open
Loading

Retire drains a session before releasing it.

sequenceDiagram
  participant C as Controller
  participant R as Resources server
  C->>R: retire(r, N)
  Note over R: refuse the attempt's sessions, except /close_session
  Note over R: wait until no request is in flight on them, over HTTP or MCP
  Note over R: retire_session_state for exported servers, stop tracking
  R-->>C: stopped and freed
Loading

Where this sits in the stack

This is one PR in a stack of draft PRs that re-cut partial-rollout checkpointing onto environment servers. Each PR's base is the branch of the PR before it, so each diff shows only that PR's commits.

The stack is based on the token-capture cleanup PRs #3938 (capture ledger retire and delete) and #3939 (complete-record retire and delete): #3882's base is #3939's branch. The checkpoint stack uses those operations to free the ledgers of retired attempts and to clear a dead execution's capture files before a restore. The striped lock files of #3937 are independent of the stack.

  1. feat(checkpoint): add the participant control plane, episode steps, and coordination #3882 (ananthsub/partial-ckpt-core): feat(checkpoint): add the participant control plane, episode steps, and coordination
  2. feat(checkpoint): make policy model servers checkpoint participants with generation cuts #3883 (ananthsub/partial-ckpt-policy-model): feat(checkpoint): make policy model servers checkpoint participants with generation cuts
  3. feat(token-capture): add worker staging helpers for generation cuts #3884 (ananthsub/partial-ckpt-model-worker-cuts): feat(token-capture): add worker staging helpers for generation cuts
  4. feat(checkpoint): continue environment server episodes from their boundaries #3885 (ananthsub/partial-ckpt-environment): feat(checkpoint): continue environment server episodes from their boundaries
  5. feat(checkpoint): add the resources server participant with asynchronous session hooks #3886 (ananthsub/partial-ckpt-resources): feat(checkpoint): add the resources server participant with asynchronous session hooks (this PR)
  6. feat(checkpoint): declare five training verifiers stateless with replayable verification #3887 (ananthsub/partial-ckpt-verifier-declarations): feat(checkpoint): declare five training verifiers stateless with replayable verification
  7. feat(checkpoint): add the agent session participant and Simple Agent continuation #3888 (ananthsub/partial-ckpt-agent): feat(checkpoint): add the agent session participant and Simple Agent continuation
  8. test(checkpoint): add a process-level end-to-end suite driven by coordination #3889 (ananthsub/partial-ckpt-e2e): test(checkpoint): add a process-level end-to-end suite driven by coordination
  9. feat(checkpoint): checkpoint evaluation runs from rollout collection #3893 (ananthsub/partial-ckpt-rollout-collection): feat(checkpoint): checkpoint evaluation runs from rollout collection
  10. feat(checkpoint): spans and metrics for partial-rollout checkpoints #3903 (ananthsub/partial-ckpt-telemetry): feat(checkpoint): spans and metrics for partial-rollout checkpoints
  11. feat(checkpoint): resources, agent, and environment servers with several workers #3909 (ananthsub/partial-ckpt-multi-worker): feat(checkpoint): resources, agent, and environment servers with several workers

Partial-rollout checkpointing lets a training controller, such as NeMo RL, checkpoint Gym while rollouts are in flight and, after a crash, continue those rollouts from their last safe point instead of starting them over. Checkpointing is off by default; with the checkpoint: block unset, no server installs any checkpoint routes or behavior.

Relationship to the old stack (#2939 to #2946)

  • Supersedes feat(checkpoint): add resources state participant #2945 (resources state participant). Route classification, mutation receipts, revision reconciliation, and the terminal context variable are gone; the prepare order (agents park before resources servers close) makes them unnecessary. MCP sessions are not supported yet.
  • Takes the counter adapter from feat(checkpoint): adapt initial stateful environments #2946. The Blackjack and Workplace Assistant adapters are not carried: Blackjack moves to a separate Gymnasium PR, and Workplace Assistant can be ported onto these hooks later.

Issue

No tracking issue exists for this re-cut. The design and the mapping from the old stack are described in this PR series, and the old stack's PRs (#2939 to #2946) carry the original discussion.

Validation

Run on this branch, on top of #3939's branch:

  • RAY_TMPDIR=/tmp .venv/bin/python -m pytest -q -p no:cacheprovider tests/unit_tests/test_checkpoint_*.py: 117 passed, including, each failing without its change:
    • a /close_session during a checkpoint waits for resume, then ends the session;
    • a retire waits for the session's request in flight, and refuses a new one meanwhile;
    • a commit whose scope leaves out a restored session releases it through the server's hook.
  • RAY_TMPDIR=/tmp .venv/bin/python -m pytest -q -p no:cacheprovider tests/unit_tests/test_mcp_auto_exposure.py tests/unit_tests/test_checkpoint_resources_mcp.py: 51 passed. The MCP checkpoint tests go through the real control routes and MCP JSON-RPC:
    • a tool call during a checkpoint waits and runs after resume, and the checkpoint holds the state from before it;
    • a tool call in flight holds up prepare;
    • a retire waits for a tool call in flight over MCP;
    • a restored server accepts the old token;
    • control routes are not tools.
  • RAY_TMPDIR=/tmp .venv/bin/python -m pytest -q -p no:cacheprovider resources_servers/example_session_state_mgmt/tests: 1 passed.
  • RAY_TMPDIR=/tmp .venv/bin/python -m pytest -q -p no:cacheprovider resources_servers/example_single_tool_call/tests: 2 passed.
  • The full core unit suite (tests/unit_tests, eight processes, without test_opensandbox_compose.py, which needs the optional opensandbox package): 6,538 passed, 2 failed. The two failures, a sandbox retry test and a Slurm script test, fail the same way on main in this development environment.
  • pre-commit run --files <files changed by this PR>: all hooks passed, and no hook modified a file.
  • Every commit carries a Signed-off-by line.

Rollout evidence

  • Process-level evidence is in the top PR: restored counter state is required for the right reward, with a negative control that skips the resources restore and gets reward 0.0.
  • Real-model rollout evidence: pending.

Compatibility

  • New resources server hooks and class attributes: checkpoint_mode, checkpoint_verify, export_session_states, restore_session_states, retire_session_state, checkpoint_session_started, checkpoint_session_ended. All are new; existing servers need no change and default to restart_only with wait verification.
  • /seed_session replies carry a new x-ng-checkpoint-verify header when checkpointing is on.
  • Nothing changes unless the global checkpoint: block sets enabled: true.

@ananthsub ananthsub added feature New capabilities, enhancements, or enablement work area:env-infra Shared environment framework, lifecycle, registry, scaffolding, and validation labels Oct 1, 2026
@copy-pr-bot

copy-pr-bot Bot commented Oct 1, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@ananthsub
ananthsub requested a review from zyzhou5 October 1, 2026 19:04
@ananthsub
ananthsub force-pushed the ananthsub/partial-ckpt-resources branch from aa41669 to defe5d3 Compare October 1, 2026 19:38
@ananthsub
ananthsub force-pushed the ananthsub/partial-ckpt-resources branch from defe5d3 to f481624 Compare October 1, 2026 21:53
@ananthsub
ananthsub force-pushed the ananthsub/partial-ckpt-resources branch from f481624 to 4e9f81d Compare October 1, 2026 22:05
@ananthsub
ananthsub force-pushed the ananthsub/partial-ckpt-resources branch from 4e9f81d to b72fe50 Compare October 2, 2026 13:22
@ananthsub
ananthsub force-pushed the ananthsub/partial-ckpt-resources branch from b72fe50 to 7f1cfb6 Compare October 2, 2026 19:43
@ananthsub
ananthsub force-pushed the ananthsub/partial-ckpt-resources branch from 7f1cfb6 to 412e854 Compare October 2, 2026 21:17
Every resources server takes part in checkpoints once checkpointing is
on, in one of three modes:

- restart_only (the default): live sessions block a checkpoint until
  their rollouts are retired, so an unsupported server fails closed.
- stateless: tool results depend only on the request, so nothing is
  exported. The weather example server uses it.
- exported: the server implements export_session_state,
  restore_session_states, and retire_session_state. The counter example
  server uses it. export_session_state raises KeyError for a session the
  server already dropped (for example after a failed verification); the
  participant stops tracking it instead of failing the commit.

The session lifecycle comes from /seed_session, /close_session, and
/verify, keyed by the rollout attempt from the rollout context, so no
request body is inspected. Servers whose protocol creates sessions
elsewhere, such as Gymnasium-style servers, call
checkpoint_session_started and checkpoint_session_ended.

The checkpoint_verify class attribute declares whether /verify is safe
to run again after a crash ("replay") or must finish before a
checkpoint ("wait", the default). It is a property of the code, so it is
not configurable. The server reports it in an x-ng-checkpoint-verify
header on every /seed_session reply, which is how callers learn it
without another request, and in its checkpoint status.

Replay-safe requests (/seed_session, and /verify when declared replay)
are admitted even after admission closes. The checkpoint does not wait for
the episode steps that send them, so refusing one would fail its episode.
A session seeded while closed is not part of that checkpoint: it neither
blocks it nor is exported by it, and becomes an ordinary session on
resume.

Signed-off-by: Ananth Subramaniam <ansubramania@nvidia.com>
Resources session hooks become coroutines, and export takes every session at
once: export_session_states(session_ids) returns the state of each session the
server still holds. A server whose session state lives outside the process,
such as a sandbox per session, can then snapshot it concurrently within the
commit. A session left out of the result replaces the KeyError signal for a
session the server already dropped.

Signed-off-by: Ananth Subramaniam <ansubramania@nvidia.com>
…over MCP

A resources server with expose_tools_over_mcp and checkpointing failed at
startup: MCP auto-exposure refused the checkpoint middleware, because direct
MCP dispatch skips middleware.

- The checkpoint middleware declares that it applies to the MCP request as a
  whole, and MCP auto-exposure accepts it. Admission, in-flight counting, and
  attempt fencing work on the one POST that carries each tool call.
- An MCP request names its session with the signed token, not the cookie,
  so the middleware reads the session from the token and fences retired
  sessions. The token key is deterministic, so a harness keeps its token
  across a restore.
- While a checkpoint is open, an MCP call waits for resume instead of being
  refused. MCP clients are third-party agent harnesses that would hand the
  refusal to the model as a tool error. A waiting call is not in flight, so
  it does not hold up prepare.
- Checkpoint control routes are no longer harvested as MCP tools.

Signed-off-by: Ananth Subramaniam <ansubramania@nvidia.com>
…for resume

While a checkpoint was open, the resources participant refused every request
except seeds and replayable verifications, including /close_session. An
episode that ends while a checkpoint is open, for example one a controller
retired or one whose replay step failed, closes its resources session once
from its final cleanup. The refusal was only logged, so the server's release
of that session was lost.

A /close_session now waits for resume, as an MCP tool call already does. A
waiting close is not in flight, so it does not hold up prepare, and it never
changes state a commit is exporting.

Signed-off-by: Ananth Subramaniam <ansubramania@nvidia.com>
- A /close_session for a retired attempt's session is no longer refused as
  stale. Closing is how the server releases that session's state, and a
  restart_only server's retire runs no hook, so its state leaked.
- Restored sessions no request has used yet are reported as pending, so a
  commit that no longer continues their episode retires them through the
  server's retire hook instead of exporting them at every checkpoint.

Signed-off-by: Ananth Subramaniam <ansubramania@nvidia.com>
… requests in flight

A retire now refuses the attempt's sessions while it stops them, waits until every request in flight on them, tool calls over MCP included, has finished, then releases them. The set of sessions retired over the life of the process is gone: nothing remains of a session once its retire replies.

Signed-off-by: Ananth Subramaniam <ansubramania@nvidia.com>
@pthombre
pthombre added this pull request to stack #3962 October 2, 2026 22:50
@ananthsub
ananthsub force-pushed the ananthsub/partial-ckpt-resources branch from 412e854 to fb28578 Compare October 2, 2026 23:32

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:env-infra Shared environment framework, lifecycle, registry, scaffolding, and validation feature New capabilities, enhancements, or enablement work

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant