Skip to content

feat(checkpoint): add durable turn-level rollout recovery - #3349

Open
macandro96 wants to merge 39 commits into
ananthsub/checkpoint-environment-adaptersfrom
amahishi/gym-turn-level-recovery
Open

macandro96 wants to merge 39 commits into
ananthsub/checkpoint-environment-adaptersfrom
amahishi/gym-turn-level-recovery

Conversation

@macandro96

@macandro96 macandro96 commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor

What does this PR do?

Summary

Adds durable turn-level rollout recovery to NeMo Gym on top of #2946.

This change allows an accepted multi-turn rollout to stop at a committed agent/environment boundary, persist coordinated Gym state, and continue under a replacement rollout attempt after process or job restart.

What this adds above #2946

#2946 introduces checkpoint adapters for the initial stateful environments.

This PR integrates those adapters into the complete Gym rollout lifecycle:

  • durably acknowledge completed agent executions
  • preserve agent continuations and model-call lineage across attempts
  • restore conversation cookies and resource-state revisions
  • make resource mutations idempotent across recovery
  • coordinate model, agent, and resources admission during checkpointing
  • classify environments as export/restore, restart-only, or stateless
  • restart only executions that depend on non-restorable resources
  • use indexed checkpoint artifacts instead of directory scans
  • archive model lineage and agent state in bounded, checksummed shards

Recovery flow

  1. The checkpoint coordinator closes model admission.
  2. Accepted model calls finish and their capture lineage becomes durable.
  3. The agent parks at its latest committed turn boundary.
  4. Resource servers export the corresponding environment state and revision.
  5. Model, agent, and resource manifests are committed under one checkpoint.
  6. After restart, source attempt N remains fenced.
  7. The checkpoint is installed for replacement attempt N+1 while admission
    remains paused.
  8. The participants resume together and the rollout continues from the saved
    turn without replaying completed resource mutations.
sequenceDiagram
    participant C as Checkpoint coordinator
    participant M as Model server
    participant A as Agent server
    participant R as Resource servers
    participant S as Checkpoint store

    Note over C,S: Checkpoint

    C->>M: Close admission and prepare
    Note over M: Accepted calls finish<br/>No mid-generation recovery
    M-->>C: Model-lineage manifest

    C->>A: Prepare checkpoint
    A-->>C: Parked turn and continuation manifest

    C->>R: Export environment state
    R-->>C: State and revision<br/>or restart-only classification

    C->>S: Commit coordinated checkpoint
    S-->>C: Checkpoint committed

    Note over C,S: Restore after restart

    C->>S: Load committed checkpoint
    S-->>C: Return participant manifests

    C->>M: Restore lineage and fence attempt N
    C->>A: Install continuation as attempt N+1
    C->>R: Restore state or restart dependent execution

    C->>M: Reopen admission
    C->>A: Resume execution
    C->>R: Resume resources
Loading

Correctness guarantees

  • completed /run results remain available until acknowledged
  • source attempts cannot mutate restored successor state
  • resource mutations are replay-safe and revision checked
  • the first restored model call links to the saved parent exactly once
  • failures during prepare or restore leave participants paused or abortable
  • environments without export/restore support restart only their dependent unfinished executions

Scalability

Checkpoint artifacts use:

  • explicit continuation and storage-reference indexes
  • scoped artifact retrieval
  • bounded model-lineage tar shards
  • bounded agent-state tar shards
  • per-file size and SHA-256 validation

This avoids scanning or materializing one loose checkpoint file per active
rollout during restore.

Testing

  • unit coverage for model, agent and resources checkpoint participants
  • recovery across replacement attempts
  • completed-result acknowledgement and replay
  • exactly-once resource mutation behavior
  • restart-only resource fallback
  • SimpleAgent turn recovery
  • deterministic Workplace Assistant state recovery
  • archive integrity, corruption and unsafe-member validation

Scope

Included:

  • turn-boundary recovery
  • agent conversation/continuation recovery
  • supported environment-state recovery
  • restart fallback for unsupported stateful environments

Not included:

  • token-prefix recovery inside an unfinished model call
  • vLLM KV-cache persistence
  • TransferQueue checkpoint orchestration
  • sandbox process or in-memory snapshotting

Stack

Checklist

  • I have read the contributing guidelines.
  • The change is focused; unrelated "drive-by" edits are tracked as separate issues/PRs.
  • Tests added or updated and pass locally, or N/A for docs-only / non-code changes (so CI unit/server checks pass when applicable).
  • Pre-commit checks pass locally (pre-commit run --all-files) (so CI lint/format/copyright pass).
  • All commits have DCO sign-off (git commit -s) (so the DCO check passes).

@copy-pr-bot

copy-pr-bot Bot commented Sep 13, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@yaoyu-33 yaoyu-33 added area:core Shared APIs, servers, telemetry, health, and registries feature New capabilities, enhancements, or enablement work needs-review PR is ready for code review and waiting on a reviewer labels Sep 13, 2026
@macandro96
macandro96 force-pushed the amahishi/gym-turn-level-recovery branch from 6473f49 to 6bba7b8 Compare September 14, 2026 01:40
@macandro96
macandro96 marked this pull request as draft September 14, 2026 03:47
@macandro96 macandro96 changed the title feat(checkpoint): Gym turn level recovery feat(checkpoint): add durable turn-level rollout recovery Sep 14, 2026
@macandro96
macandro96 force-pushed the amahishi/gym-turn-level-recovery branch from dbb8b76 to 80bd2a9 Compare September 15, 2026 16:33
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
@macandro96
macandro96 force-pushed the amahishi/gym-turn-level-recovery branch from 5b0b444 to f4fcf8c Compare September 23, 2026 00:36
@macandro96
macandro96 changed the base branch from main to ananthsub/checkpoint-environment-adapters September 23, 2026 00:39
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
@macandro96
macandro96 marked this pull request as ready for review September 29, 2026 21:52
@yaoyu-33 yaoyu-33 added the complexity:high Cross-component or shared-contract change with a large review surface label Sep 29, 2026
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:core Shared APIs, servers, telemetry, health, and registries complexity:high Cross-component or shared-contract change with a large review surface feature New capabilities, enhancements, or enablement work needs-review PR is ready for code review and waiting on a reviewer sla:triage-overdue Review assignment is over the one-business-day SLA

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants