Add Slurm run and DSE results to experiment output - #1040
Draft
podkidyshev wants to merge 1 commit into
Draft
podkidyshev wants to merge 1 commit into
podkidyshev wants to merge 1 commit into
Conversation
Signed-off-by: Ivan Podkidyshev <ipodkidyshev@nvidia.com>
Contributor
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: trueComment |
podkidyshev
added this pull request to stack #1042
September 18, 2026 18:12
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
experiment.jsonwith Slurm run status, UTC accounting timestamps, duration, and canonical workload metrics before iteration or DSE state advances.was_run_successful.Test Plan
Affected tests on macOS / Python 3.14.3:
All pre-commit checks passed for the 12 changed files, including Pyright, Ruff, Vulture, import-linter, and Taplo. Existing output/status assertions were extended; three focused test functions cover Slurm capture, missing metadata/metric extraction failure, and DSE winner selection.
Earlier live Slurm smoke tests used one node with eight H100 GPUs: normal NCCL, two-step NCCL and NIXL DSE, NCCL single-sbatch, and Slurm dry-run. Workloads completed successfully; dry-run generated scripts without a real submission. Artifacts were retrieved and analyzed locally. The final case-metric changes were replayed locally against those artifacts: normal NCCL has 36 case measurements, and the selected DSE trials have 36 NCCL / 12 NIXL measurements, with all per-run values and UTC timing preserved.
Additional Notes
Stacked on #1030 (
ipod/unified-output). Single-sbatch per-run output, live running-state updates, iteration aggregation, and DSE with iterations remain separate work. Single-sbatch retains scenario-level output; dry-run has no actual run records.