fix: abort process groups on rank-local training failure - #781
Merged
Conversation
5 tasks
maocheng23
changed the base branch from
maocheng/colocate-1-capture-rows
to
main
September 1, 2026 23:30
This was referenced Sep 1, 2026
maocheng23
force-pushed
the
maocheng/colocate-2-teardown-abort
branch
from
September 1, 2026 23:30
b8f2ba5 to
6a5cccd
Compare
destroy_process_group is collective; when one rank dies inside the run while its peers are blocked in a CUDA/FSDP collective, teardown hangs in NCCL communicator destruction, the originating traceback never surfaces, and torchrun cannot reap the job. On the exceptional path the CLI now calls destroy_distributed(abort=True), which uses the non-collective ProcessGroup.abort() on every distinct cached group (including the default group) so elastic sees the real exception. The success path keeps the existing collective destroy. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Expose capture_rows(input_ids) on OfflineSGLangCaptureBackend and OfflineSGLangCapture so callers can capture variable-length rows in one packed prefill without building padded tensors. capture_eagle3 delegates to it and keeps its exact request construction and outputs. DSpark capture-layer setup now falls back from the native set_dspark_layers_to_capture hook to the dense DFlash hook (with a log line naming the resolved hook) so dense targets such as Qwen3 can serve DSpark capture on stock SGLang 0.5.14 models. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Every trainer rank constructed its own tracker, so W&B/MLflow runs were duplicated world-size times and TensorBoard ranks wrote the same directory. An external tracker is now created on global rank zero only; every rank keeps the console logger, so console output is unchanged.
maocheng23
force-pushed
the
maocheng/colocate-2-teardown-abort
branch
from
September 2, 2026 00:24
6a5cccd to
bb22955
Compare
maocheng23
marked this pull request as ready for review
September 2, 2026 00:25
…rows refactor: extract packed capture_rows from the offline SGLang backend
…acker fix: create the external metrics tracker only on trainer rank zero
jiapingW
approved these changes
Sep 2, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Stack 1/5 replacing draft #766 (moved to the front of the stack: it is independent of colocated training and useful on its own). When one rank dies inside a run (a data or capture error) while its peers are blocked in a CUDA/FSDP collective, the collective
destroy_process_groupin the CLI'sfinallyhangs in NCCL communicator destruction: the originating traceback never surfaces and torchrun cannot reap the job. Colocated online training makes rank-local capture failures a realistic event, but this fix applies to every distributed topology.Modifications
destroy_distributed(abort=False): on the exceptional path, call the non-collectiveProcessGroup.abort()on every distinct cached group, including the default group. Aborting does not unblock the peers by itself; torchrun/elastic terminates them once this rank exits non-zero with its real traceback (the docstring now says so). The success path keeps the existing collective destroy; torch builds withoutProcessGroup.abortfall back to destroy._trainmarks the exceptional path with afailedflag and passesabort=failedto teardown.Related Issues
Splits #766. Stack: #1 (this, on
main) ← #782 rank0-tracker ← #780 capture-rows ← #783 colocated-core ← #784 hybrid-shard.Accuracy Test
test_failure_teardown_aborts_each_distinct_process_groupverifies abort is called once per distinct group and collective destroy is not used on the failure path. The existing teardown test keeps covering the success path.main2fc99307(sglang 0.5.18). CPU suite: failure set identical tomain.Checklist
black --checkandisort --check-only).