Skip to content

feat(checkpoint): resources, agent, and environment servers with several workers - #3909

Draft
ananthsub wants to merge 8 commits into
ananthsub/partial-ckpt-telemetryfrom
ananthsub/partial-ckpt-multi-worker
Draft

ananthsub wants to merge 8 commits into
ananthsub/partial-ckpt-telemetryfrom
ananthsub/partial-ckpt-multi-worker

Conversation

@ananthsub

@ananthsub ananthsub commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

What changed and why

The bottom two commits are #3905 (session routing). This PR carries them until #3905 merges. Review only the commits after them:

  1. refactor(checkpoint): a worker coordinator for any participant kind
  2. feat(checkpoint): resources, agent, and environment servers with several workers
  3. test(checkpoint): run the e2e scenarios with two workers per server
  4. feat(checkpoint): restore sessions onto a different number of workers

Until now, resources, agent, and environment servers refused num_workers > 1 when checkpointing was on. uvicorn's workers share one port, so a control call reaches one arbitrary worker. If each worker ran its own participant, a checkpoint would close, export, and restore only that worker's share of the episodes and sessions. Production servers run several workers (swe_rebench runs 8, several sandboxed agents run 4), so this guard kept them out of partial-rollout checkpoints.

This PR makes each of these servers one participant across its workers:

  • A generic worker coordinator. The policy model server already runs as one participant across workers. The parts of it that do not depend on the policy model move to nemo_gym/_checkpoint/workers.py:

    • the framed Unix-socket channel;
    • worker registration;
    • closing and reopening every worker;
    • sequenced readiness reports;
    • the blockers for unreported and lost workers;
    • a worker that loses the coordinator stops itself.

    The policy model keeps only its own state, and its behavior and tests are unchanged.

  • Prepare closes every worker. It is ready once every worker reports ready. While a checkpoint is open, a worker re-reports its readiness on a short poll, so readiness that changes without a notification still reaches the coordinator. One example is a resources session that ends.

  • Commit asks every worker for its records. The coordinator writes the participant's one records file and manifest, so the store format and the controller's view do not change.

    • Resources and agent session records carry their routing owner: the worker ID in the session's cookie or MCP token.
    • The manifest lists every owner a cookie may name at commit time.
  • Restore of sessions that span requests (resources sessions and native agent sessions):

    • All sessions of one pre-crash owner are installed on the same live worker, and owners are spread evenly.
    • Every worker's router gets an alias table from each old owner to its new worker. A request whose cookie or MCP token names a pre-crash worker reaches the worker that now holds its session. Without the alias, routing would answer 410.
    • Owners whose sessions exported nothing are aliased too. One example is a stateless resources server, whose old cookies must still reach a live worker.
    • Sessions checkpointed by a server with one worker name no owner, because routing was off. Each one is installed on a live worker on its own, and these sessions are spread evenly too. Every router gets a second table, from session ID to worker, for these sessions only. When a cookie or MCP token names no owner, the router looks up its session ID in this table before it handles the request locally. The router already decodes the session ID from both.
      • The worker a session is placed on becomes its owner from then on. Its next response stamps that owner in the cookie.
      • A later checkpoint lists the placed sessions in its manifest. If the server crashes again before a placed session's cookie is stamped, the next restore keeps a table entry for it and points the entry at the session's new worker.
      • The table is replaced at each restore and holds at most one entry per restored session. An entry is not dropped when its session ends, because the session ends on one worker while the table is on every worker. That costs about 100 bytes per restored session until the next restore.
      • Agent session records now also carry the cookie's session ID, read from the request that opened the session.
    • So the worker count can change between a checkpoint and a restore: 1 to N, N to M with both above one (old owners are aliased onto however many workers there are), and N to 1. With one worker there is no router, and the single process holds every session.
    • A worker that joins later, or restarts, gets both current tables when it registers.
  • Restore of state that lives inside one request (environment episodes and legacy agent /run episodes):

    • The coordinator holds the records.
    • The worker that receives the replacement attempt's /run claims its record, and exactly one claim succeeds.
    • A claim is answered before any later message to that worker. A claimed episode is therefore either exported by its worker or waited for by the next checkpoint, never missed.
    • Claims are refused while a checkpoint is open, with the same code a new episode gets, so the caller retries the episode the same way.
    • A record nobody has claimed is exported again by the next checkpoint.
  • Retire reaches every worker. Each worker refuses the attempt only while its own retire stops it, and the coordinator drops unclaimed records of retired attempts. No lasting fence is copied to workers or broadcast on resume.

  • Restored state no longer continued is released on every worker. The coordinator gathers the restored records no worker has claimed and the restored sessions each worker has not used yet, so a commit whose scope leaves them out retires them.

  • Fail closed. Fewer registered workers than num_workers blocks prepare. So does a worker lost while a checkpoint is open, until the controller resumes.

Two smaller changes support this:

  • A legacy /run calls the agent's own /v1/responses for its turn loop. That call now names its worker in a new x-ng-session-owner header, so the turn loop's activation runs on the worker that tracks the episode. The router honors the header only for a request without a session cookie or MCP token.
  • Simple Agent no longer refuses sessions with several workers, now that routing sends a session's calls to its worker. Simple Agent and the weather resources server gain the module-level app that multi-worker uvicorn imports.

How it works

Where this PR sits in the overall flow

The highlighted part is what this PR adds.

flowchart LR
  C["Controller<br/>NeMo RL, or rollout collection"]
  CO["Coordination<br/>prepare, commit, restore, resume, retire"]
  K["Control plane on every server<br/>phases, lease, storage,<br/>retire: stop, then free"]
  subgraph G["One participant per Gym server"]
    E["Environment server<br/>episode steps"]
    M["Policy model<br/>held responses, generation cuts"]
    A["Agent<br/>sessions parked at boundaries"]
    R["Resources server<br/>session state"]
  end
  W["Inference worker<br/>stages cut prefixes"]
  D[("Checkpoint directory<br/>records, then manifest")]
  L[("Capture ledger<br/>retire and delete from #3938, #3939")]
  C --> CO --> K
  K --> E & M & A & R
  M --> W
  M --> L
  G --> D
  classDef this fill:#fde68a,stroke:#b45309,stroke-width:2px,color:#1f2937
  class E,A,R this
Loading

One checkpoint, a crash, and the restore, end to end:

sequenceDiagram
  participant C as Controller
  participant G as Gym participants
  participant D as Checkpoint directory
  C->>G: prepare, in order environment, model, agent, resources
  Note over G: admission closes, in-flight work parks at a boundary,<br/>undelivered model responses are held
  G-->>C: prepared, or blockers at the deadline
  C->>G: commit with the episodes the controller continues
  G->>D: each participant writes its records, then its manifest
  Note over G: restored state the commit's scope leaves out is released
  C->>C: publish the checkpoint with the controller's own state
  C->>G: resume, in order resources, agent, model, environment
  Note over C,G: crash - every Gym process dies
  C->>G: restore the checkpoint in fresh processes, all or nothing
  D-->>G: records installed under attempt + 1, attempt N's capture ledger retired
  C->>G: resume
  C->>G: /run as attempt + 1 continues each episode from its boundary
  Note over C,G: dropping an episode, only while no checkpoint is open
  C->>G: retire - environment, then agent, then model and resources
  Note over G: each server stops the attempt's work, waits, frees its state, then replies
Loading

This PR

One server with several uvicorn workers is still one participant. Its coordinator, in the main process, drives each worker's own participant; the router sends a session's requests to the worker that holds it.

flowchart LR
  CO["Coordination"] -->|"control call on the shared port"| W1["worker 1<br/>participant and router"]
  CO -->|"control call on the shared port"| W2["worker 2<br/>participant and router"]
  W1 <-->|"Unix socket"| C["Coordinator in the main process<br/>one participant: phases, records,<br/>owner aliases, restored episodes"]
  W2 <-->|"Unix socket"| C
  W1 <-->|"forward to the session's owner"| W2
Loading

Restoring sessions into new workers. A cookie or MCP token names the worker that held the session before the crash; the alias table sends it to the worker that holds it now.

sequenceDiagram
  participant C as Coordinator
  participant W1 as Worker 1
  participant W2 as Worker 2
  participant X as Caller
  C->>C: read the records, place each old owner on one live worker
  C->>W2: install the sessions of old owner A
  C->>W1: alias table - A is now worker 2
  C->>W2: alias table - A is now worker 2
  X->>W1: request with a cookie naming owner A
  W1->>W2: forward
  W2-->>X: reply from the restored session
Loading

Whole-episode state, such as an environment episode or a legacy /run, waits in the coordinator. Exactly one worker claims it, whichever receives the replacement attempt's /run.

sequenceDiagram
  participant X as Controller
  participant W as Worker that receives the call
  participant C as Coordinator
  X->>W: /run for attempt + 1
  W->>C: claim the episode's restored record
  C-->>W: the record, or nothing if another worker claimed it
  W->>W: continue the episode from its boundary
  Note over C: unclaimed records stay while the controller's commit scope lists them,<br/>and are released when it no longer does
Loading

Where this sits in the stack

This is one PR in a stack of draft PRs for partial-rollout checkpointing. Each PR's base is the branch of the PR before it, so each diff shows only that PR's commits.

The stack is based on the token-capture cleanup PRs #3938 (capture ledger retire and delete) and #3939 (complete-record retire and delete): #3882's base is #3939's branch. The checkpoint stack uses those operations to free the ledgers of retired attempts and to clear a dead execution's capture files before a restore. The striped lock files of #3937 are independent of the stack.

  1. feat(checkpoint): add the participant control plane, episode steps, and coordination #3882 (ananthsub/partial-ckpt-core): feat(checkpoint): add the participant control plane, episode steps, and coordination
  2. feat(checkpoint): make policy model servers checkpoint participants with generation cuts #3883 (ananthsub/partial-ckpt-policy-model): feat(checkpoint): make policy model servers checkpoint participants with generation cuts
  3. feat(token-capture): add worker staging helpers for generation cuts #3884 (ananthsub/partial-ckpt-model-worker-cuts): feat(token-capture): add worker staging helpers for generation cuts
  4. feat(checkpoint): continue environment server episodes from their boundaries #3885 (ananthsub/partial-ckpt-environment): feat(checkpoint): continue environment server episodes from their boundaries
  5. feat(checkpoint): add the resources server participant with asynchronous session hooks #3886 (ananthsub/partial-ckpt-resources): feat(checkpoint): add the resources server participant with asynchronous session hooks
  6. feat(checkpoint): declare five training verifiers stateless with replayable verification #3887 (ananthsub/partial-ckpt-verifier-declarations): feat(checkpoint): declare five training verifiers stateless with replayable verification
  7. feat(checkpoint): add the agent session participant and Simple Agent continuation #3888 (ananthsub/partial-ckpt-agent): feat(checkpoint): add the agent session participant and Simple Agent continuation
  8. test(checkpoint): add a process-level end-to-end suite driven by coordination #3889 (ananthsub/partial-ckpt-e2e): test(checkpoint): add a process-level end-to-end suite driven by coordination
  9. feat(checkpoint): checkpoint evaluation runs from rollout collection #3893 (ananthsub/partial-ckpt-rollout-collection): feat(checkpoint): checkpoint evaluation runs from rollout collection
  10. feat(checkpoint): spans and metrics for partial-rollout checkpoints #3903 (ananthsub/partial-ckpt-telemetry): feat(checkpoint): spans and metrics for partial-rollout checkpoints
  11. feat(checkpoint): resources, agent, and environment servers with several workers #3909 (ananthsub/partial-ckpt-multi-worker): feat(checkpoint): resources, agent, and environment servers with several workers (this PR)

Server ports that build on the stack: #3894 (Workplace Assistant), #3895 (Gymnasium), #3896 (Blackjack), #3897 (indirect prompt injection), #3898 (proof refinement). Session routing for several workers is #3905, against main.

Relationship to the old stack

The old stack (#2939 to #2946) required one worker for every checkpointing server except the policy model. The policy model already had multi-worker support in #3883, as a coordinator in the main process with worker gates. This PR generalizes that coordinator rather than adding a second mechanism. It replaces the three num_workers=1 guards the re-cut stack kept for resources, agent, and environment servers.

Issue

No tracking issue exists. Multi-worker support for these servers was a stack requirement. Session routing, which it depends on, is #3905.

Validation

Run on this branch, on top of #3939's branch:

  • RAY_TMPDIR=/tmp .venv/bin/python -m pytest -q -p no:cacheprovider tests/unit_tests/test_checkpoint_*.py tests/unit_tests/test_session_routing.py: 185 passed. test_checkpoint_participant_workers.py runs each scenario with a real coordinator socket and two worker links:
    • prepare closes every worker and waits for the episodes of all of them;
    • commit collects every worker's records;
    • readiness that changes without a notification still reaches the coordinator;
    • a lost worker blocks prepare and commit until resume, and fewer workers than configured block prepare;
    • a restored environment episode and a restored legacy agent /run are each claimed by exactly one worker;
    • an unclaimed restored episode is exported by the next checkpoint while the scope continues it, and released, on every worker, when it does not;
    • a restored agent session is released on the worker that holds it when the scope no longer continues it;
    • claims are refused while a checkpoint is open;
    • restored agent sessions are placed one owner per worker, every router gets the alias table, and a late worker gets it too;
    • owners without records are aliased;
    • a session placed without an owner stays routable through a later checkpoint and restore;
    • a retire leaves no fence on any worker, including one that joins later;
    • two real counter resources servers with session routing serve a restored session from its old cookie on either worker;
    • sessions checkpointed by a single-process counter server are restored onto two routed counter workers.
  • RAY_TMPDIR=/tmp .venv/bin/python -m pytest -q -p no:cacheprovider responses_api_agents/simple_agent/tests: 30 passed. environment_servers/single_agent_turn/tests: 14 passed, after removing the test of the one-worker refusal this PR removes.
  • NEMO_GYM_CHECKPOINT_E2E=1 RAY_TMPDIR=/tmp .venv/bin/python -m pytest -q -p no:cacheprovider tests/e2e/checkpoint tests/e2e/session_routing, one test at a time: 43 passed, 9 skipped (the opt-in real-model and scale tests), in about 17 minutes. The scenarios run with the environment, agent and resources servers at one and two workers, and the policy model at one and two, including:
    • native and legacy crash and continue;
    • a second crash before the replacement starts, and a restore again after the replacement made calls;
    • 16 concurrent rollouts at 1 to 1, 2 to 2, 1 to 2, 2 to 1 and 2 to 4 workers;
    • MCP tool calls that wait out a checkpoint on whichever worker owns the session, then keep their sessions across a restore onto three workers;
    • a retired episode leaves nothing behind.
  • The full core unit suite (tests/unit_tests, eight processes, without test_opensandbox_compose.py, which needs the optional opensandbox package): 6,722 passed, 2 failed. The two failures, a sandbox retry test and a Slurm script test, fail the same way on main in this development environment.
  • pre-commit run --files <files changed by this PR>: all hooks passed, and no hook modified a file.
  • Every commit carries a Signed-off-by line.

Rollout evidence

  • Real-model rollout evidence: pending, as for the rest of the stack. The opt-in real vLLM test was not run.

  • Scale, on one workstation with every Gym server and the fake inference backend. The run used native single_agent_turn weather rollouts, with token capture and generation cuts on. Each rollout's second model call was held, so about half the rollouts were in flight at the checkpoint. The test checkpoints, kills every Gym process group, restarts, restores, and runs the replacement attempts. Every server ran with two workers:

    Rollouts Policy workers Environment, agent, and resources workers In flight at checkpoint (exported) Prepare (s) Commit (s) Restore (s) Replacement failures
    16,000 2 2 9,350 2.73 (model 2.62) 1.32 1.47 0
    • Records: 8.4 MB for the environment, 18.2 MB for the agent, and 25.2 MB for the model.
    • 8,002 generation cuts, and all 8,002 cut calls were continued.
    • Every rollout either finished before the checkpoint or was exported, and every exported rollout scored 1.0 after the restore.

Compatibility

  • Resources, agent, and environment servers now accept num_workers > 1 with checkpointing on. Single-worker servers behave as before. The visible differences are in the records:

    • a new owner field, null with one worker, in resources and agent session records;
    • a new session_id field in agent session records;
    • session_owners and placed_sessions in a multi-worker participant's manifest.
  • The worker count may change between a checkpoint and a restore, in either direction: from one worker to several, from several to one, or between different counts.

  • With several workers, a restored environment or legacy episode whose replacement /run arrives while a checkpoint is open is refused with checkpoint_parked, as a new episode is, and the caller retries it. With one worker, it is admitted and parks at its first boundary.

  • With several workers, commit and restore carry each worker's records over the coordinator's socket. The frame limit is now 1 GiB, up from the policy model's 64 MiB, which only rejects a corrupt length prefix. Servers with very large sessions put all their records through the main process. Per-worker record files are the alternative if that becomes a bottleneck.

  • New header x-ng-session-owner: on a server with several workers, a request without a session cookie or MCP token that sets it is forwarded to that worker.

@ananthsub ananthsub added feature New capabilities, enhancements, or enablement work area:core Shared APIs, servers, telemetry, health, and registries labels Oct 1, 2026
@copy-pr-bot

copy-pr-bot Bot commented Oct 1, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@ananthsub
ananthsub force-pushed the ananthsub/partial-ckpt-telemetry branch from 9775d38 to dc49274 Compare October 2, 2026 13:22
@ananthsub
ananthsub force-pushed the ananthsub/partial-ckpt-multi-worker branch from d64723f to a92580a Compare October 2, 2026 13:22
@ananthsub
ananthsub force-pushed the ananthsub/partial-ckpt-telemetry branch from dc49274 to f9affd7 Compare October 2, 2026 19:43
@ananthsub
ananthsub force-pushed the ananthsub/partial-ckpt-multi-worker branch from a92580a to 6647f1c Compare October 2, 2026 19:43
@ananthsub
ananthsub force-pushed the ananthsub/partial-ckpt-telemetry branch from f9affd7 to f8f2b11 Compare October 2, 2026 21:17
@ananthsub
ananthsub force-pushed the ananthsub/partial-ckpt-multi-worker branch from 6647f1c to ec6b311 Compare October 2, 2026 21:17
uvicorn's workers share one listening socket, so each request goes to
whichever worker accepts the connection, while session state lives in the
memory of the worker that created the session. With the stateful counter
resources server at 4 workers, 200 concurrent episodes through /run scored
1.0 on only 7; swe_rebench (8 workers) can lose a session's sandbox at
/verify and score 0.

With num_workers > 1, resources and agent servers now:

- stamp each new session with the ID of the worker that created it, inside
  the signed session cookie;
- serve the same app on a private Unix socket per worker, in a directory
  the main process creates under /tmp;
- forward a request for a session another worker owns to that worker's
  private socket, unchanged, and stream the reply back unchanged. A
  forwarded request is always handled locally;
- answer 410 when the owner's socket is gone, instead of running against
  empty state.

The router sits outside SessionMiddleware and decodes the cookie with the
same secret, so a forwarded reply carries only the owner's Set-Cookie.
Single-worker servers and model servers are unchanged.

The counter resources server gains the module-level app that multi-worker
uvicorn imports, and an opt-in e2e test (NEMO_GYM_MULTIWORKER_E2E=1) runs
it with 4 workers behind the legacy /run relay and Simple Agent.

Signed-off-by: Ananth Subramaniam <ansubramania@nvidia.com>
CLI harnesses in sandboxes reach a resources server's tools over /mcp with
the signed session token minted at /seed_session, and send no Gym session
cookie. With several workers, those calls ran on whichever worker accepted
them: 4 workers, 200 concurrent counter sequences over MCP, 65 of 200
correct on main and 72 of 200 with cookie routing alone.

With routing active, the MCP token now also carries the owning worker's ID.
The router reads the owner from the token on the MCP path, and on any
request without a session cookie, verifying it with the same serializer
and salt. A token without an owner (minted by a single-worker server or
before this change) or with a bad signature is handled locally, where the
MCP endpoint's own check applies as before. Single-worker tokens are
unchanged. With this change, 200 of 200 are correct.

Signed-off-by: Ananth Subramaniam <ansubramania@nvidia.com>
The policy model server already runs as one checkpoint participant across
several uvicorn workers: a coordinator in the main process owns the
controller, and each worker links to it over a Unix socket. Environment,
agent, and resources servers need the same machinery.

Move the parts that do not depend on the policy model into
nemo_gym/_checkpoint/workers.py:
- the framed channel between the coordinator and its workers;
- CoordinatedParticipant: worker registration, closing and reopening every
  worker, sequenced readiness reports, fence broadcast, and the blockers
  for unreported or lost workers;
- WorkerCoordinator, which serves it on a background-thread event loop;
- WorkerLink: forwarding control calls, joining while a checkpoint is open,
  reporting, and stopping the worker when the coordinator is gone.

The policy model keeps only its own state: the gate, the restored
generation cuts and their claims, and the ledger export and import. Its
behavior and its tests are unchanged.

Signed-off-by: Ananth Subramaniam <ansubramania@nvidia.com>
…ral workers

These servers refused num_workers > 1 when checkpointing was on: each
worker would have closed, exported, and restored only its own share of
the episodes and sessions. With session routing in place, they now run
as one participant across workers, on the worker coordinator the policy
model server uses.

Each worker runs the participant a single process runs, linked to the
coordinator in the main process:

- Prepare closes every worker and is ready once every worker reports
  ready. A worker re-reports while a checkpoint is open, so readiness
  that changes without a notification, such as a session ending, still
  reaches the coordinator.
- Commit asks every worker for its records and writes the participant's
  one records file and manifest. Resources and agent session records
  carry their routing owner, the worker ID in the session's cookie or
  MCP token. The manifest lists every owner a cookie may name.
- Restore installs resources sessions and native agent sessions on live
  workers, one worker per pre-crash owner, spread evenly. Every worker's
  router gets the alias table from old owner to new worker, also for
  owners that exported nothing, such as a stateless resources server's.
  A request whose cookie or token names a pre-crash worker reaches the
  worker that holds its session instead of getting 410. A worker that
  joins later gets the table when it registers.
- Environment episodes and legacy agent /run episodes stay with the
  coordinator. The worker that receives the replacement /run claims the
  record, and exactly one claim succeeds. A claim is answered before
  any later message to that worker, so a claimed episode is always
  either exported by its worker or waited for by the next checkpoint.
  Claims are refused while a checkpoint is open, as new episodes are.
  Unclaimed records are exported again by the next checkpoint.
- Retire and attempt fences reach every worker; the coordinator drops
  unclaimed records of retired attempts.
- An unregistered or lost worker blocks prepare and commit, and a
  worker that loses the coordinator stops itself.

A legacy /run's call to the agent's own /v1/responses names its worker
in the new x-ng-session-owner header, so the turn loop's activation
runs where the episode is tracked. Simple Agent no longer refuses
sessions with several workers, and it and the weather resources server
gain the module-level app that multi-worker uvicorn imports.

Signed-off-by: Ananth Subramaniam <ansubramania@nvidia.com>
The environment, agent, and resources servers in the e2e deployment take
a server_workers count, and the crash and continue scenarios (native,
legacy, resources state with its negative control, a second crash before
the replacement, and a checkpoint without a crash) run with one worker
and with two. The scale test takes the same parameter.

A new scenario starts 16 rollouts at once, native over agent sessions and
legacy over counter resources sessions, checkpoints them mid-episode,
kills every Gym process, restores, and runs the replacements. With two
workers the checkpoint's sessions come from both workers; every
replacement scores 1.0, and the backend sees only the calls each rollout
still needed.

Signed-off-by: Ananth Subramaniam <ansubramania@nvidia.com>
A checkpoint taken with one worker could not restore resources sessions
or native agent sessions into servers with several workers. Routing was
off at commit, so their records and cookies name no owner, and the
restore failed. Users will change the worker count between a checkpoint
and a restore, so it must work in any direction.

The coordinator now places each session record without an owner on a
live worker of its own, spread evenly with the owners' groups, and
installs it there; that worker owns it from then on. Every worker's
router gets a second table, alongside the owner aliases, from restored
session ID to worker, for these sessions only. When a cookie or MCP
token names no owner, the router looks the decoded session ID up in it
before handling the request locally. Workers that register later get it
too.

The manifest lists the placed sessions, so a session restored again
before its cookie was stamped with its new owner keeps a table entry,
pointed at its owner's new worker. The table is replaced at each restore
and holds at most one entry per restored session; an entry outlives its
session until then, since a session ends on one worker and the table is
on all of them.

Agent session records now carry the cookie's session ID, which the
session middleware exposes to code without the request at hand.

The e2e scenario with 16 concurrent rollouts now restarts the servers
with a different worker count: 1 to 2, 2 to 1, and 2 to 4, as well as
unchanged. Before this change, both 1 to 2 cases failed at restore.

Signed-off-by: Ananth Subramaniam <ansubramania@nvidia.com>
…kers

A CLI agent harness reaches the resources server only through MCP tokens. Its calls during an open checkpoint wait on the worker that owns the session, and each harness keeps its token when Gym restarts with a different worker count.

Signed-off-by: Ananth Subramaniam <ansubramania@nvidia.com>
…eral workers

The coordinator now reports as pending both the restored records no worker has claimed and the restored sessions each worker has not used yet, so a commit that no longer continues their episode retires them on every worker.

Signed-off-by: Ananth Subramaniam <ansubramania@nvidia.com>
@ananthsub
ananthsub force-pushed the ananthsub/partial-ckpt-telemetry branch from f8f2b11 to 14f23fa Compare October 2, 2026 23:32
@ananthsub
ananthsub force-pushed the ananthsub/partial-ckpt-multi-worker branch from ec6b311 to 51dc9c8 Compare October 2, 2026 23:32

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:core Shared APIs, servers, telemetry, health, and registries feature New capabilities, enhancements, or enablement work

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant