Skip to content

Add Slurm run and DSE results to experiment output - #1040

Draft
podkidyshev wants to merge 1 commit into
ipod/unified-outputfrom
ipod/unified-slurm
Draft

podkidyshev wants to merge 1 commit into
ipod/unified-outputfrom
ipod/unified-slurm

Conversation

@podkidyshev

Copy link
Copy Markdown
Contributor

Summary

  • Populate experiment.json with Slurm run status, UTC accounting timestamps, duration, and canonical workload metrics before iteration or DSE state advances.
  • Fill case metrics from the sole successful normal run or the successful DSE trial with the highest configured reward. Preserve every run's original measurements and record the DSE search space, best step, and best configuration.
  • Preserve completed results after failure or cancellation, finalize unresolved records as unknown, and keep extraction/write failures nonfatal. Workload success continues to use was_run_successful.

Test Plan

Affected tests on macOS / Python 3.14.3:

uv run --locked --extra dev pytest tests/systems/slurm/test_runner.py tests/systems/slurm/test_system.py tests/test_base_runner.py tests/test_single_sbatch_runner.py tests/systems/standalone/test_runner.py tests/test_output.py tests/test_handlers.py tests/test_acceptance.py tests/workloads/nccl_test tests/workloads/nixl_bench
262 passed in 3.94s

All pre-commit checks passed for the 12 changed files, including Pyright, Ruff, Vulture, import-linter, and Taplo. Existing output/status assertions were extended; three focused test functions cover Slurm capture, missing metadata/metric extraction failure, and DSE winner selection.

Earlier live Slurm smoke tests used one node with eight H100 GPUs: normal NCCL, two-step NCCL and NIXL DSE, NCCL single-sbatch, and Slurm dry-run. Workloads completed successfully; dry-run generated scripts without a real submission. Artifacts were retrieved and analyzed locally. The final case-metric changes were replayed locally against those artifacts: normal NCCL has 36 case measurements, and the selected DSE trials have 36 NCCL / 12 NIXL measurements, with all per-run values and UTC timing preserved.

Additional Notes

Stacked on #1030 (ipod/unified-output). Single-sbatch per-run output, live running-state updates, iteration aggregation, and DSE with iterations remain separate work. Single-sbatch retains scenario-level output; dry-run has no actual run records.

Signed-off-by: Ivan Podkidyshev <ipodkidyshev@nvidia.com>
@coderabbitai

coderabbitai Bot commented Sep 18, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Comment @coderabbitai help to get the list of available commands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant