Cancel a PR's orphaned workflow runs when the PR is closed - #415
Merged
Merged
Conversation
GitHub does not cancel queued or in-progress runs when a PR is closed, and `main.yml`'s concurrency group only supersedes a run when a *newer* run appears on the same head ref -- closing a PR is not a new run, so it never fires. The orphaned runs then sit in the single self-hosted GPU queue and block every later PR behind them. Observed on 2026-08-03: PR #413 was open for 2m45s, and its run outlived it by ~84 minutes, holding the head of the GPU queue in front of #414 until the runs were cancelled by hand. Notes on the implementation: - `pull_request_target`, not `pull_request`, so the token is writable for fork PRs too. Safe only because nothing here checks out or executes PR code. - The head ref reaches jq via `$ENV` and gh via `-f`, never a shell word or a raw URL -- on a fork PR it is attacker-controlled. - It skips if any *other* open PR still points at the same head repo + ref, so retargeting a stack branch cannot cancel a live PR's runs. - Bounded escalation to `force-cancel`: a long-running self-hosted job can ignore the graceful cancel, which is what kept the GPU queue blocked above. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QipFjBYxb5aPnmZE76cxhi
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #415 +/- ##
==========================================
- Coverage 98.30% 90.39% -7.92%
==========================================
Files 42 170 +128
Lines 2415 15153 +12738
==========================================
+ Hits 2374 13697 +11323
- Misses 41 1456 +1415 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
`closed` also fires on merge, so the first version cancelled a merged PR's runs
too. HMG's call is to let those finish, so the job is now guarded by
if: github.event.pull_request.merged == false
The cost is documented at the guard rather than left to be rediscovered: a
`benchmark-pr` run sits in concurrency group `gpu-benchmarks-<head_ref>` while
the post-merge `benchmark-main` run sits in `gpu-benchmarks-main`, so merging
does not supersede it. Combined with `timeout-minutes: 14400` (10 days) and the
"wait for GPU to be free" spin loop, a benchmark run started on a PR can hold
the benchmark runner well past the merge and needs cancelling by hand.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QipFjBYxb5aPnmZE76cxhi
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes the gap that blocked the GPU queue for 84 minutes today.
The problem
GitHub does not cancel a PR's queued or in-progress runs when the PR is closed.
main.yml's concurrency guard cannot cover this:That only supersedes a run when a newer run appears on the same head ref. Closing a PR
is not a new run, so it never fires.
With a single self-hosted GPU runner, the orphans become head-of-line blockers.
Observed 2026-08-03. PR #413 (
grid-search-ez→reachability) was open from 11:21:03to 11:23:48 — 2m45s. Its run
30809182760outlived it by ~84 minutes, sitting in the GPUqueue in front of #414 until the runs were cancelled by hand at 12:45. #390's run also had
to be force-cancelled: its 32-bit GPU job ignored the ordinary cancel endpoint entirely.
The fix
A reaper on
pull_request_target: [closed]that cancels runs for that head ref.Five things worth reviewing:
closedalso fires on merge, so the job isguarded by
if: github.event.pull_request.merged == false. Merged PRs' runs aredeliberately left to finish.
pull_request_target, notpull_request— the token is read-only for fork PRs onpull_request, so the reaper would silently no-op on exactly the PRs least likely to becleaned up by hand. This is only safe because the job never checks out or executes PR
code; it reads the head ref and calls the REST API. Do not add a checkout step.
jq via
$ENVand gh via-f(which also URL-encodes it) — never a shell word or a rawquery string. Branch names may legally contain
&and#.the obvious case here. The job bails if any other open PR still points at the same head
repo + ref, so closing one cannot cancel a live one's runs.
force-cancel. 120s grace, then force. Without this theworkflow would not have fixed today's incident: the graceful cancel was ignored.
Known cost of exempting merged PRs
This is documented at the guard so it isn't rediscovered the hard way.
benchmark-prrunsin concurrency group
gpu-benchmarks-<head_ref>, while the post-mergebenchmark-mainrunuses
gpu-benchmarks-main— different groups, so merging does not supersede the PR'sbenchmark run. With
timeout-minutes: 14400(10 days) and awhile nvidia-smi … sleep 60wait-for-GPU loop, a benchmark run started on a PR can hold the benchmark runner well past
the merge. Cancel it by hand if that happens.
Verification
Ran against live API state before committing:
-f branch="feat/dcegm"filters correctly (slash in ref)$ENVworks in gh's jq (gojq)reachability(#414 open)still_open=1→ skipsgrid-search-ez30809182760queued→ inside the filter, would have been reapedset -esurvives[ … ] && continuebash -non the extractedrun:blockyamllint,check-github-workflows, fullprekNot exercised end-to-end on GitHub —
pull_request_targetruns the version of the workflowon the base branch, so it does nothing until this lands on
main. The first real testis the next PR closed without merging.