Repository navigation
Evict failing merge-queue entries ~35 min sooner - #1146
Merged
Mikola Lysenko (mikolalysenko) merged 2 commits intoOct 8, 2026
Merged
Conversation
ci-ok needs every CI job, so a merge-queue entry whose CI fails in minute 2 (a semantic conflict that stops compilation, say) only reports failure after the slowest Gradle e2e leg ends ~45 minutes later. Until then the entry, and every entry stacked on it, holds the queue and keeps burning macOS and Windows runners on a run that can no longer pass. Add a merge_group-only watcher that polls the entry's CI run and cancels it on the first failed or timed-out job. No CI job sets continue-on-error, so such a run can't pass. ci-ok is if: always(), so it still runs on the cancelled run and fails in seconds; the queue then evicts the entry and rebuilds the ones behind it. A test pins both ci.yml properties the watcher relies on. The watcher is a separate workflow, not a required check, and never fails its own job. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CtHdJ7b6ncYXfs8LSJMsPy
Collaborator
Author
|
bugbot run Generated by Claude Code |
Tanmay Singla (Tanmay182003)
approved these changes
Oct 8, 2026
The watch job holds an actions:write token. Checking out the merge-group head meant a queued PR could edit scripts/merge-queue-fail-fast.py and run arbitrary Python with that token before it ever lands. Check the script out at merge_group.base_sha instead so changing it requires the change to have already merged, and skip cleanly while the script is not on base yet. Co-Authored-By: Claude <noreply@anthropic.com>
Collaborator
Author
|
bugbot run Generated by Claude Code |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 66aa864. Configure here.
Collaborator
Author
|
Burn-down agent: labeled Ready for review at 66aa864.
Generated by Claude Code |
Collaborator
Author
|
Tanmay Singla (@Tanmay182003) — final reviewer: not enqueueing yet. One non-merge commit landed after your approval at 5bf4852:
CI is green on 66aa864 (ci-ok, clippy). Could you take another look at the workflow diff since your approval? Once you re-approve, it can go straight to the merge queue. Generated by Claude Code |
Mikola Lysenko (mikolalysenko)
deleted the
ci-janitor/merge-queue-fail-fast
branch
October 8, 2026 21:05
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
ci-ok(the merge queue's required check)needs:every CI job, so a queue entry whose CI is already doomed only reports failure after the slowest leg ends. Today's queue shows the cost:hosted::engine::rewrite, and Fix open bun issues (#992, #861, #784, #764, #735, #635, #599, #578, #497, #443, #371) #1009 adds a test that still passes it (error[E0061]).coverage,test (macos/windows)andtest-releasewere all red by 15:20Z.ci-okstill sat pending behind the Gradle e2e legs for another ~40 min.ci-okonly reported at 15:51Z, ~35 min after its first failure. The queue can't evict and rebuild until then, and each of these runs (~190 jobs, macOS and Windows included) keeps burning runners.Last 60 merge-group CI runs (2026-10-07 22:07Z to 2026-10-08 15:33Z): 12 failed runs got a `ci-ok` verdict, and 5 more were still pending at the time of writing. In those 12, the first failed job finished a median of ~2 min into the run: `clippy` at 0.6–1.7 min, or the compile error in `coverage`/`test`/`test-release` at ~1.7 min. `ci-ok` then took another 31–43 min (434 min total, ~36 min per run) to report:
These runs also ran ~490 macOS job-minutes, counting every macOS job that finished after the run's first failure.
Root cause
A job with
needs:can't start until all of its needs finish, and the Gradle e2e legs take 28–35 min. Nothing in the run stops the other jobs once a failure makesci-okcertain to fail.Fix
A new
merge_group-only workflow,merge-queue-fail-fast.yml, runsscripts/merge-queue-fail-fast.py. The script polls the entry's CI run every 30 s and cancels it as soon as any job concludesfailureortimed_out.continue-on-error, so a failed job already meansci-okfails. Cancelling only changes when it fails.ci-okisif: always(), so it runs on a cancelled run and fails within seconds. This repo's own cancelled PR runs show it: 37800029383, 37799319868 and 37782495575 each got aci-okfailure 3–7 s after the cancel.scripts/tests/test_merge_queue_fail_fast.pypins both of those ci.yml properties. It fails if anyone adds a job-levelcontinue-on-erroror dropsif: always()fromci-ok.cancelledjobs don't trigger a cancel, because someone (the queue, a person, Cut merge-group CI from ~46 to ~20 min: shard Gradle e2e and test legs, skip test-release in queue, cancel orphaned runs #1133's orphan janitor) is already cancelling.No
ci.ymlchange, so this doesn't conflict with #1133 or #1139 and keeps every required-check name. All tests still run exactly where they did before. The only difference is that a doomed merge-group run stops early.Proof
coverage,test (macos-latest),test (windows-latest)andtest-release.python3 -B -m unittest discover -s scripts/tests: 259 tests OK, 5 of them new (this suite runs inlint-ecosystems).actionlintclean;zizmorreports no findings.Expected effect: a failing entry leaves the queue ~35 min sooner (about 7 queue-hours saved in the window above), and the entries behind it rebuild that much sooner. Each failed group run stops using macOS and Windows runners minutes after its first failure instead of running all ~190 jobs. The cost is one ubuntu job per queue entry that polls for the length of the run.
Note: like any merge_group workflow, this takes effect only after it lands on main.
🤖 Generated with Claude Code
https://claude.ai/code/session_01CtHdJ7b6ncYXfs8LSJMsPy
Note
Medium Risk
Introduces automated workflow-run cancellation with
actions:write, gated to merge queue and base-branch script; mis-detection could cancel runs that could still pass if CI invariants change without updating the contract tests.Overview
Adds a merge-queue-only workflow that watches each entry’s
ci.ymlrun and cancels it as soon as any job finishes withfailureortimed_out, so the requiredci-okgate (which runsif: always()) fails in seconds instead of waiting on slow e2e legs.The watcher checks out
base_shaonly (sparsescripts/merge-queue-fail-fast.py) so queued PR code never runs with anactions:writetoken; if the script isn’t on main yet, the job exits cleanly. The Python poller finds the merge-group run by head SHA, posts cancel via the Actions API, logs to the step summary, and never fails the workflow on API errors.scripts/tests/test_merge_queue_fail_fast.pycoversfatal_jobslogic and assertsci.ymlkeeps no job-levelcontinue-on-errorandci-okkeepsif: always(), so early cancel stays safe.Reviewed by Cursor Bugbot for commit 66aa864. Configure here.
Generated by Claude Code