Skip to content

feat(checkpoint): add generation-prefix recovery - #4319

Open
macandro96 wants to merge 5 commits into
amahishi/gym-turn-recovery-orchestrationfrom
amahishi/prefix-recovery-core
Open

macandro96 wants to merge 5 commits into
amahishi/gym-turn-recovery-orchestrationfrom
amahishi/prefix-recovery-core

Conversation

@macandro96

@macandro96 macandro96 commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

This PR adds generation-prefix recovery on top of the turn-level Gym recovery contract in #4266.

When a periodic rollout snapshot occurs while vLLM is still decoding, the active output prefix is staged in TransferQueue (TQ) and referenced from Gym's durable model lineage. After a restart, a replacement model call reconstructs that prefix from TQ and generates only the remaining suffix instead of restarting the entire model call.

This does not checkpoint the vLLM KV cache. Recovery reconstructs the model input from durable token IDs.

Checkpoint flow

flowchart TD
    A[Periodic rollout snapshot requested] --> B[Single Controller closes rollout admission]
    B --> C[Fence terminal token-capture writes<br/>vLLM continues decoding]
    C --> D[Gym prepares active agents and model calls]
    D --> E[Each vLLM worker freezes/swaps<br/>the current capture buffer]
    E --> F[Stage active token prefixes in TQ]
    F --> G[Return token-free TQ coordinates<br/>to the Gym model ledger]
    G --> H[Gym commits agent/resource state<br/>and model lineage]
    H --> I[Single Controller validates<br/>Gym references against TQ]
    I --> J[Save TQ + controller recovery state]
    J --> K[Atomically publish rollout snapshot]
    K --> L[Resume Gym and release generation fence]
Loading

Recovery flow

flowchart TD
    A[Select latest compatible rollout snapshot] --> B[Restore TQ and controller state]
    B --> C[Restore Gym participants and lineage]
    C --> D[Redispatch unfinished rollout attempt]
    D --> E[Gym attaches saved prefix coordinates<br/>to the replacement model call]
    E --> F[vLLM fetches and validates<br/>the saved token chunks from TQ]
    F --> G[Rebuild engine input from the prefix]
    G --> H[Generate only the remaining suffix]
    H --> I[Stage one canonical terminal call<br/>and continue normal finalization]
Loading

Implementation

  • Adds rollout_checkpointing.gym.mode: prefix_recovery.
  • Tracks active async-vLLM request capture state and safely freezes/swaps its current buffer at a checkpoint boundary.
  • Supports optional periodic chunk staging through generation_chunk_flush_tokens; the checkpoint cut flushes any remaining unstaged tokens.
  • Stores token/logprob payloads in TQ and only durable coordinates plus lineage metadata in Gym.
  • Extends Gym checkpoint validation so a snapshot is published only when every referenced generation-cut row exists in the same TQ checkpoint.
  • Restores staged chunks into replacement calls while preserving the effective output budget, policy-version checks, masks, logprobs, and terminal metadata needed by finalization.
  • Cleans obsolete generation-cut rows after a canonical terminal row has been staged.

Scope and non-goals

  • Requires token capture and async vLLM generation.
  • Router replay is rejected for this mode in this PR because per-token routed-expert traces are not yet restored.
  • This PR contains the core single-Gym-instance prefix path. Sharded Gym and replica ownership are handled in downstream work.
  • KV-cache, process-memory, and sandbox-memory snapshots are not included.

Stack

Coverage added

  • Unit coverage for worker capture-buffer cuts, TQ staging, Gym checkpoint records, restore wiring, and Single Controller checkpoint ordering.
  • Functional recovery coverage for simple-agent and Workplace rollouts.
  • The new functional tests are registered in the Single Controller L1 suite.

Validation

  • ruff check passed on touched Python files.
  • ruff format --check passed on the final stacked tree.
  • Python compilation passed for the checkpoint orchestration modules.
  • Cluster unit and functional recovery suites still need to run on the published stack.

Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>
@copy-pr-bot

copy-pr-bot Bot commented Sep 29, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@github-actions

Copy link
Copy Markdown

❌ Submodule Fast-Forward Check Failed

Check based on commit: c31dc5c (PR #4319 from amahishi/prefix-recovery-core)

❌ Submodules that need attention:

Gym: ❌ Commits have DIVERGED from a common ancestor
TARGET (amahishi/gym-turn-recovery-orchestration branch): https://github.com/NVIDIA-NeMo/Gym/commits/d1abebd0f08f2aba9cac0acf5d3fa82d2058d154/
CURRENT (PR #4319 from amahishi/prefix-recovery-core): https://github.com/NVIDIA-NeMo/Gym/commits/d474ef6b26124db28ff2d9a3617ae5d8d04e323f/

Please ensure all submodule commits are fast-forwards of the amahishi/gym-turn-recovery-orchestration branch before merging.

@macandro96
macandro96 added this pull request to stack #4267 September 29, 2026 04:25
@macandro96
macandro96 marked this pull request as ready for review September 29, 2026 04:28
Signed-off-by: Anish Mahishi <amahishi@nvidia.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant