Skip to content

test: exercise Gym checkpoint recovery with CPU mocks - #4306

Draft
terrykong wants to merge 4 commits into
amahishi/gym-turn-recovery-orchestrationfrom
terryk/rl-1581-cpu-mocks
Draft

terrykong wants to merge 4 commits into
amahishi/gym-turn-recovery-orchestrationfrom
terryk/rl-1581-cpu-mocks

Conversation

@terrykong

Copy link
Copy Markdown
Collaborator

What does this PR do?

Adds replaceable CPU policy, generation, and refit components under tests/mock_stack/ to exercise the real Single Controller, Gym, token capture, row assembly, and SimpleStorage recovery path without GPUs. Production recovery code is unchanged.

The test runs four prompt groups with three siblings each, saves at steps 2 and 4, removes step 4, starts a fresh Ray runtime, and restores step 2 through normal discovery. It checks training order, reuse of one completed sibling, continuation of five partial siblings, exactly 20 new model calls, exact saved token/logprob prefixes, and restored policy weights.

Issues

Depends on #4266 and targets its branch. Draft for viewing the additional test diff.

Usage

uv run pytest tests/mock_stack -q

See tests/mock_stack/README.md for the workload, component factories, and environment requirements.

Validation

  • Full local CPU suite: 24 passed in 257.42 seconds.
  • After review added explicit checks that the first five resumed calls use the saved step-2 weights; the checkpoint test passed again in 196.46 seconds.
  • Ruff 0.9.9 and whitespace checks passed.
  • Tested on macOS with Python 3.13.14 and torch 2.10.0. The pinned Linux/torch 2.11 environment remains unvalidated.

Before ready for review

  • Add component and controller integration tests.
  • Document the scenario and replacement interfaces.
  • Validate in the pinned Linux environment.

Additional information

Rollout timing uses elapsed sleeps, with no event gates. The assertions fail when the expected checkpoint cut is missed. The first version uses deliberate checkpoints and normal shutdown; signal/crash injection and mixed real/GPU components are outside this test's scope.

Per-call timing records are currently returned in memory, not saved as an observed-timeline artifact. Existing text logs alone cannot reconstruct the complete rollout timeline.

@copy-pr-bot

copy-pr-bot Bot commented Sep 28, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@macandro96
macandro96 force-pushed the amahishi/gym-turn-recovery-orchestration branch from 4db8f66 to 7ebf4aa Compare September 29, 2026 03:42
Signed-off-by: Terry Kong <terryk@nvidia.com>
Signed-off-by: Terry Kong <terryk@nvidia.com>
Signed-off-by: Terry Kong <terryk@nvidia.com>
Signed-off-by: Terry Kong <terryk@nvidia.com>
@terrykong
terrykong force-pushed the terryk/rl-1581-cpu-mocks branch from d61ee61 to 1e7c0cb Compare September 29, 2026 20:15

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant