You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
CI perf: compatibility workflows (sbt, PDM, vlt, Bun, Composer, +5) — full matrices on ~80% of PR pushes, 0 failures (~80,000 Linux + 26,000 Windows job-min/day) #1198
Window: 2026-10-08 00:17 → 2026-10-09 00:17 UTC. Daily run counts are estimated from six 4-hour slices of the Actions runs list (4,850 runs in the window). Per-run job-minutes come from the job data of 2–3 sampled PR runs per workflow, so treat each workflow's figure as ±50%.
For comparison, the whole CI workflow on pull_request costs about 89,000 Linux + 16,000 Windows job-min/day.
Signal: 0 failures. The runs list has 438 completed PR runs of these 10 workflows in the window: 371 success, 33 cancelled (superseded), 63 skipped, 0 failure. None of these workflows is a required check (ruleset 24668460 requires only ci-ok and clippy), so a red result would not block a merge anyway.
Push to main already runs the full matrices: about 40 runs/day each. That is where the real breaks surface. In the same window, the push runs that failed were Composer 3/13, Pipenv 2/13, Bun 1/10 and sbt 1/13. Every one of those commits had a green PR run.
The trigger rate is near total: these workflows start on 59–85% of PR pushes. That is ~276 sbt runs/day against ~325 CI PR runs/day. Their pull_request.paths cover the shared engine (crates/socket-patch-core/src/{formats,crawlers,hosted,patch,manifest,api,utils,vendor,vex}/**, crates/socket-patch-cli/src/**, Cargo.toml/Cargo.lock; sbt-compatibility.yml alone lists 11 such globs), and almost every PR touches one of them.
Where the time goes
sbt: an image build job (~12 min). Then 26 Docker cells, at 2–10 min each, across 6 sbt lines × {agent, hosted, vendored} plus Mill and scala-cli. Then 2 Windows native cells (11.4 + 7.3 min).
vlt and Bun: half the minutes are on Windows (37–40 job-min per run).
Each workflow's per-ecosystem blocking slice already runs on every PR inside ci.yml. sbt-compatibility.yml says so itself: "ci.yml runs the small blocking slice on every PR (agent 1.2.8 + 1.13.0 in coverage-docker, hosted + vendored 1.13.0 on ubuntu in e2e)".
Root cause
The full-matrix compatibility workflows run, without blocking anything, on every PR push that touches shared engine code. That is nearly every PR. They duplicate ci.yml's blocking slice and repeat the full matrix that push-to-main runs again after merge.
Proposed fix (M), the same shape as #1177 for Gradle
For each of sbt-, pdm-, vlt-, bun-, composer-, npm-, poetry-, go-, pnpm- and pipenv-compatibility.yml:
Narrow on.pull_request.paths to files that belong to that ecosystem only: the workflow file, its matrix/backtest script, its docs and Dockerfile, its tests/*<eco>* files, and its crawlers/<eco>, formats/<eco>, vendor/<eco> and patch/redirect/<eco> modules. Drop the shared-engine globs and Cargo.lock.
Keep push: branches: [main], which runs the full matrix after every merge. Add a nightly schedule where one is missing (only vlt has one today). Add workflow_dispatch, plus a PR label such as compat-full (pull_request: types: [labeled] gated on the label), so an engine PR can still ask for the full matrix before merge.
Optional: keep the cheapest smoke cell (the current-version hosted cell on Linux) on engine PRs, so the PR still gets an early warning.
Expected saving
≈ 60,000–70,000 Linux + 20,000 Windows job-min/day, assuming ~80% of today's PR runs stop triggering. No macOS change: these workflows' macOS legs already left the PR path in Skip draft PRs and run macOS legs off the PR path #1093.
PR feedback latency: none directly, because these workflows aren't required. Indirectly, ~250 fewer Linux jobs per PR push frees runner slots. Linux queue wait p90 was 3.7 min in the 24h window and 5.4 min in the busiest hour (20h UTC).
Coverage and risk
Every cell still runs on push to main (~40/day) and nightly. Engine PRs still run ci.yml's blocking per-ecosystem slice in both the PR run and the merge queue.
The cost is that an engine change breaking a non-blocking full-matrix cell is found after merge instead of before. In this window that already happened for every compatibility break (7 red push runs, all with green PRs). The opt-in label covers risky PRs.
The required checks ci-ok and clippy are untouched.
Effort
M. Ten workflow files with the same edit, plus a docs line in each docs/testing/*-compatibility.md.
ROI
Weighted saving, after the 80% assumption, in thousands of job-min/day: 65.6 (Linux) + 21 × 2 (Windows) ≈ 108. Then 108 × confidence 0.5 (small per-workflow sample, policy change) / effort 2 = 27. Same scale as the dashboard.
Measurement
Window: 2026-10-08 00:17 → 2026-10-09 00:17 UTC. Daily run counts are estimated from six 4-hour slices of the Actions runs list (4,850 runs in the window). Per-run job-minutes come from the job data of 2–3 sampled PR runs per workflow, so treat each workflow's figure as ±50%.
For comparison, the whole
CIworkflow on pull_request costs about 89,000 Linux + 16,000 Windows job-min/day.ci-okandclippy), so a red result would not block a merge anyway.CIPR runs/day. Theirpull_request.pathscover the shared engine (crates/socket-patch-core/src/{formats,crawlers,hosted,patch,manifest,api,utils,vendor,vex}/**,crates/socket-patch-cli/src/**,Cargo.toml/Cargo.lock; sbt-compatibility.yml alone lists 11 such globs), and almost every PR touches one of them.Where the time goes
imagebuild job (~12 min). Then 26 Docker cells, at 2–10 min each, across 6 sbt lines × {agent, hosted, vendored} plus Mill and scala-cli. Then 2 Windows native cells (11.4 + 7.3 min).ci.yml. sbt-compatibility.yml says so itself: "ci.yml runs the small blocking slice on every PR (agent 1.2.8 + 1.13.0 in coverage-docker, hosted + vendored 1.13.0 on ubuntu in e2e)".Root cause
The full-matrix compatibility workflows run, without blocking anything, on every PR push that touches shared engine code. That is nearly every PR. They duplicate ci.yml's blocking slice and repeat the full matrix that push-to-main runs again after merge.
Proposed fix (M), the same shape as #1177 for Gradle
For each of
sbt-,pdm-,vlt-,bun-,composer-,npm-,poetry-,go-,pnpm-andpipenv-compatibility.yml:on.pull_request.pathsto files that belong to that ecosystem only: the workflow file, its matrix/backtest script, its docs and Dockerfile, itstests/*<eco>*files, and itscrawlers/<eco>,formats/<eco>,vendor/<eco>andpatch/redirect/<eco>modules. Drop the shared-engine globs andCargo.lock.push: branches: [main], which runs the full matrix after every merge. Add a nightlyschedulewhere one is missing (only vlt has one today). Addworkflow_dispatch, plus a PR label such ascompat-full(pull_request: types: [labeled]gated on the label), so an engine PR can still ask for the full matrix before merge.Expected saving
Coverage and risk
ci-okandclippyare untouched.Effort
M. Ten workflow files with the same edit, plus a docs line in each
docs/testing/*-compatibility.md.ROI
Weighted saving, after the 80% assumption, in thousands of job-min/day: 65.6 (Linux) + 21 × 2 (Windows) ≈ 108. Then 108 × confidence 0.5 (small per-workflow sample, policy change) / effort 2 = 27. Same scale as the dashboard.
Generated by Claude Code