Skip to content

CI perf: CI on push to main re-tests the exact SHA the merge queue just passed (~16,000 job-min/day, ~1,800 macOS) #1170

Description

Measurement

Window: 2026-10-08 02:16Z → 19:36Z (0.72 days). That is after #1093 moved macOS off the PR path and the merge queue became the landing path. Source: actions/runs and runs/<id>/jobs for all CI runs.

  • 30 of 30 CI push runs on main since 2026-10-07 21:00Z have a head_sha that equals the head SHA of a merge_group CI run that already concluded success. The merge queue fast-forwards main to the commit it tested, so the push run re-tests an identical tree.

  • Push runs per day: 35 (25 runs in the window). Each run has 197 jobs and costs:

    OS job-min/run job-min/day
    Linux 443 15,335
    macOS 62 2,146
    Windows 63 2,167

    For comparison, push runs before 10-08 02:16Z were 33 jobs and 58 Linux min.

  • 3 of the push runs failed on SHAs whose merge_group run was green: 37725703000 (merge_group 37718964794), 37747365139 (merge_group 37743052751) and 37796729798 (merge_group 37790956990). Those are flake-only reds on main.

  • Example pairs (push run ↔ merge_group run on the same SHA):

  • Linux runners are saturated in busy hours. Linux job queue wait had p50 7.3 / 10.3 / 6.3 min in the 02Z / 03Z / 04Z hours and 5.5 min at 11Z, with about 3,000 jobs started per hour. Every push run adds about 197 jobs into the same pool as the merge_group runs that make up the queue's critical path. Its macOS jobs (e2e-macos ×23, test macOS, yarn/cargo macOS) also compete for the ~20-slot macOS pool.

Where the time goes

Mean per push run (post-cut):

Job Linux min/run
e2e 256 (117 jobs)
coverage-docker 68
test-release 26
e2e-full 22 (30 jobs)
coverage 16
yarn-classic 14
yarn-berry 12

Non-Linux per push run: e2e-macos 28 macOS min (23 jobs), test windows 27, test macOS 22, e2e windows 21.5.

Only the full tier (e2e-full, cargo-vex-matrix-full, yarn-berry-full, about 25 min/run in total) does not also run in merge_group.

Root cause

ci.yml gates the macOS legs (e2e-macos, yarn-berry-e2e-macos, cargo-vex-matrix-macos) and the pull_request-tier jobs on github.event_name != 'pull_request' or draft != true. So push runs the whole merge_group set plus the full tier. With the merge queue in place, push is a byte-identical re-run of the queue's last verification.

Proposed fix

In .github/workflows/ci.yml:

  1. On push to main, skip the test-execution jobs that merge_group already ran on the same SHA: e2e, e2e-macos, coverage-docker, docker-base, coverage-merge, yarn-classic-matrix, yarn-berry-e2e(-macos), cargo-vex-matrix(-macos) and hosted-e2e.

    • Add github.event_name != 'push' to their if:.
    • Optionally add a workflow_dispatch input to force them.
  2. Keep the jobs that save caches on main. These are the save-if: github.ref == 'refs/heads/main' users: clippy, node-addon, e2e-build, test, test-release, coverage, cargo-old-toolchains and the yarn jobs. Two ways to do that:

    • Turn the heavy ones into compile-only on push. For test, run only the Build step plus cargo test --no-run, so target/ is warm for PR caches.
    • Or add one cache-warm job per OS/profile that builds and saves without running tests.

    Main remains the only cache writer, as Decouple CI platform builds and share compatible Rust caches #1143 intends.

  3. Keep the full tier (e2e-full, cargo-vex-matrix-full, yarn-berry-full) on push. It's the only per-merge coverage they get. If the full tier also moves to the nightly or a 4-hourly cron, push CI shrinks to cache warming only.

  4. ci-ok treats skipped as success, so no required-check change is needed.

Expected saving

Assumes 35 pushes/day. Skipping e2e, coverage-docker, docker-base, the yarn/cargo matrices and the test step of test/test-release removes about:

OS job-min/run job-min/day
Linux ~350 ~12,200
macOS ~52 ~1,800
Windows ~48 ~1,700

Jobs started on Linux drop by about 140 per merge, about 4,900/day, which directly relieves the busy-hour Linux queue wait the merge_group runs sit in. It also removes 3 flake-only red main runs per ~25 pushes.

Coverage and risk

  • Every skipped job already ran and passed in merge_group on the identical SHA. The merge queue's ALLGREEN grouping and the required ci-ok guarantee that.
  • Direct pushes to main that bypass the queue (admin) would lose CI. Mitigation: only skip when the push SHA has a successful merge_group CI run. Check it in a 10-second gate job with gh api repos/{repo}/actions/runs?head_sha=<sha>&event=merge_group, and otherwise run everything.
  • Nightly (schedule) still runs the full matrix including e2e-docker.
  • Required check names ci-ok and clippy are unchanged.
  • Cache freshness for PRs is preserved by keeping the cache-saving compile jobs.

Effort

M. This touches the if: conditions of about 12 jobs, adds the gate job, and needs the CI harness tests (scripts/test_ci_*.py) that pin job conditions to be updated.

ROI

ROI = weighted saving × confidence / effort.

  • Weighted saving (in thousands): (Linux 12.2 + 3×macOS 1.8 + 2×Windows 1.7) = 21.0.
  • Confidence: 0.8. Effort: M = 2.
  • ROI = 8.4

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions