Skip to content

Evict failing merge-queue entries ~35 min sooner - #1146

Merged
Mikola Lysenko (mikolalysenko) merged 2 commits into
mainfrom
ci-janitor/merge-queue-fail-fast
Oct 8, 2026
Merged

Mikola Lysenko (mikolalysenko) merged 2 commits into
mainfrom
ci-janitor/merge-queue-fail-fast

Conversation

@mikolalysenko

@mikolalysenko Mikola Lysenko (mikolalysenko) commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

ci-ok (the merge queue's required check) needs: every CI job, so a queue entry whose CI is already doomed only reports failure after the slowest leg ends. Today's queue shows the cost:

Last 60 merge-group CI runs (2026-10-07 22:07Z to 2026-10-08 15:33Z): 12 failed runs got a `ci-ok` verdict, and 5 more were still pending at the time of writing. In those 12, the first failed job finished a median of ~2 min into the run: `clippy` at 0.6–1.7 min, or the compile error in `coverage`/`test`/`test-release` at ~1.7 min. `ci-ok` then took another 31–43 min (434 min total, ~36 min per run) to report:

run entry first failure (min) verdict (min) wasted (min) failed jobs
37741022380 #1057 0.6 44.0 43.4 clippy
37738827925 #1050 1.7 42.5 40.9 clippy, e2e
37739124591 #1029 1.7 38.7 37.0 clippy
37730827414 #724 2.1 44.5 42.4 coverage, test, test-release
37730570198 #1038 1.7 41.9 40.2 coverage, test, test-release
37729900304 #1083 1.7 40.0 38.2 coverage, e2e, test, test-release
37730608490 #1035 1.6 36.2 34.6 coverage, test, test-release
37713148083 #1046 11.2 53.8 42.6 coverage, test, test-release
37713149707 #1003 14.8 52.0 37.2 coverage, test, test-release
37759566368 #768 11.2 42.6 31.4 e2e
37726258998 #724 18.0 49.0 31.0 e2e
37799137861 #1121 3.1 18.2 15.1 coverage, test, test-release

These runs also ran ~490 macOS job-minutes, counting every macOS job that finished after the run's first failure.

Root cause

A job with needs: can't start until all of its needs finish, and the Gradle e2e legs take 28–35 min. Nothing in the run stops the other jobs once a failure makes ci-ok certain to fail.

Fix

A new merge_group-only workflow, merge-queue-fail-fast.yml, runs scripts/merge-queue-fail-fast.py. The script polls the entry's CI run every 30 s and cancels it as soon as any job concludes failure or timed_out.

  • Safe: no CI job sets continue-on-error, so a failed job already means ci-ok fails. Cancelling only changes when it fails.
  • The verdict still arrives: ci-ok is if: always(), so it runs on a cancelled run and fails within seconds. This repo's own cancelled PR runs show it: 37800029383, 37799319868 and 37782495575 each got a ci-ok failure 3–7 s after the cancel.
  • Guarded: scripts/tests/test_merge_queue_fail_fast.py pins both of those ci.yml properties. It fails if anyone adds a job-level continue-on-error or drops if: always() from ci-ok.
  • Never in the way: it isn't a required check and never fails its job. An API error is a warning, and the run then finishes as before.
  • cancelled jobs don't trigger a cancel, because someone (the queue, a person, Cut merge-group CI from ~46 to ~20 min: shard Gradle e2e and test legs, skip test-release in queue, cancel orphaned runs #1133's orphan janitor) is already cancelling.

No ci.yml change, so this doesn't conflict with #1133 or #1139 and keeps every required-check name. All tests still run exactly where they did before. The only difference is that a doomed merge-group run stops early.

Proof

  • Dry run against the live doomed run (POST stubbed out): it found 37799050567 by the group's head SHA and would have cancelled it for coverage, test (macos-latest), test (windows-latest) and test-release.
  • python3 -B -m unittest discover -s scripts/tests: 259 tests OK, 5 of them new (this suite runs in lint-ecosystems).
  • actionlint clean; zizmor reports no findings.

Expected effect: a failing entry leaves the queue ~35 min sooner (about 7 queue-hours saved in the window above), and the entries behind it rebuild that much sooner. Each failed group run stops using macOS and Windows runners minutes after its first failure instead of running all ~190 jobs. The cost is one ubuntu job per queue entry that polls for the length of the run.

Note: like any merge_group workflow, this takes effect only after it lands on main.

🤖 Generated with Claude Code

https://claude.ai/code/session_01CtHdJ7b6ncYXfs8LSJMsPy


Note

Medium Risk
Introduces automated workflow-run cancellation with actions:write, gated to merge queue and base-branch script; mis-detection could cancel runs that could still pass if CI invariants change without updating the contract tests.

Overview
Adds a merge-queue-only workflow that watches each entry’s ci.yml run and cancels it as soon as any job finishes with failure or timed_out, so the required ci-ok gate (which runs if: always()) fails in seconds instead of waiting on slow e2e legs.

The watcher checks out base_sha only (sparse scripts/merge-queue-fail-fast.py) so queued PR code never runs with an actions:write token; if the script isn’t on main yet, the job exits cleanly. The Python poller finds the merge-group run by head SHA, posts cancel via the Actions API, logs to the step summary, and never fails the workflow on API errors.

scripts/tests/test_merge_queue_fail_fast.py covers fatal_jobs logic and asserts ci.yml keeps no job-level continue-on-error and ci-ok keeps if: always(), so early cancel stays safe.

Reviewed by Cursor Bugbot for commit 66aa864. Configure here.


Generated by Claude Code

ci-ok needs every CI job, so a merge-queue entry whose CI fails
in minute 2 (a semantic conflict that stops compilation, say) only
reports failure after the slowest Gradle e2e leg ends ~45 minutes
later. Until then the entry, and every entry stacked on it, holds
the queue and keeps burning macOS and Windows runners on a run that
can no longer pass.

Add a merge_group-only watcher that polls the entry's CI run and
cancels it on the first failed or timed-out job. No CI job sets
continue-on-error, so such a run can't pass. ci-ok is if: always(),
so it still runs on the cancelled run and fails in seconds; the
queue then evicts the entry and rebuilds the ones behind it. A test
pins both ci.yml properties the watcher relies on.

The watcher is a separate workflow, not a required check, and never
fails its own job.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CtHdJ7b6ncYXfs8LSJMsPy
@mikolalysenko Mikola Lysenko (mikolalysenko) added the ci-janitor Opened by the CI janitor routine (flakes, redundant tests, CI perf) label Oct 8, 2026
@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

bugbot run


Generated by Claude Code

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread .github/workflows/merge-queue-fail-fast.yml Outdated
The watch job holds an actions:write token. Checking out the merge-group
head meant a queued PR could edit scripts/merge-queue-fail-fast.py and run
arbitrary Python with that token before it ever lands. Check the script out
at merge_group.base_sha instead so changing it requires the change to have
already merged, and skip cleanly while the script is not on base yet.

Co-Authored-By: Claude <noreply@anthropic.com>
@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

bugbot run


Generated by Claude Code

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 66aa864. Configure here.

@mikolalysenko Mikola Lysenko (mikolalysenko) added the Ready for review Agent-verified: mergeable, CI green, Bugbot clean — awaiting human review label Oct 8, 2026
@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

Burn-down agent: labeled Ready for review at 66aa864.

  • CI: 189/189 non-skipped checks green (ci-ok success).
  • Bugbot: the MEDIUM security finding (the actions:write watch job ran the queued PR's copy of scripts/merge-queue-fail-fast.py) is fixed in 66aa864. The canceller is now checked out at merge_group.base_sha, and the step skips while the script is not on base yet. Bugbot re-reviewed 66aa864 with no findings.
  • Reviewer note: on this PR's own queue entry the watcher is a no-op, because base has no script yet. It starts working from the next queue entry after this lands.

Generated by Claude Code

@mikolalysenko
Mikola Lysenko (mikolalysenko) added this pull request to the merge queue Oct 8, 2026
@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

Tanmay Singla (@Tanmay182003) — final reviewer: not enqueueing yet. One non-merge commit landed after your approval at 5bf4852:

  • 66aa864 Run the merge-queue canceller from the base branch, not the queue head — fixes Bugbot's MEDIUM security finding: the actions:write watch job in .github/workflows/merge-queue-fail-fast.yml now checks out github.event.merge_group.base_sha before running scripts/merge-queue-fail-fast.py, so a queued PR can't run its own copy of the script with that token.

CI is green on 66aa864 (ci-ok, clippy). Could you take another look at the workflow diff since your approval? Once you re-approve, it can go straight to the merge queue.


Generated by Claude Code

Merged via the queue into main with commit 34a152f Oct 8, 2026
196 checks passed
@mikolalysenko
Mikola Lysenko (mikolalysenko) deleted the ci-janitor/merge-queue-fail-fast branch October 8, 2026 21:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci-janitor Opened by the CI janitor routine (flakes, redundant tests, CI perf) Ready for review Agent-verified: mergeable, CI green, Bugbot clean — awaiting human review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants