You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
CI perf: Gradle e2e shards unbalanced — b-p leg is the merge-queue critical path (~3 min/merge-group run) #1171
37828932401: the 8.14.3 b-p leg ends at 23.1 of 23.2 min.
37822515034: the 8.14.3 leg ends at 24.3 of 24.4 min.
37828927026: the 8.14.3 leg ends at 23.6 of 23.7 min.
PRs. Since 10-08 02:16Z, an e2e (ubuntu-latest, …) leg was the last job before ci-ok in 94 of 111 successful PR CI runs. PR push → ci-ok p50 is 41 min and p90 53 min.
Setup and checkout: 8 s. Downloading the e2e-bin artifact: 18 s. Gradle distribution: 2 s. Maven Central traffic: none (FakeCentral).
Test step: 784 s (96%), with libtest's default 4 threads saturating 4 vCPUs. Test time adds up to 3,064 test-seconds, about 3.9 tests running at once.
Slowest tests:
Test (Gradle version)
Seconds
hosted_fallback_snippet_compiles_kotlin (8.14.3)
363
hosted_config_cache_second_row (8.14.3)
348
hosted_vendored_takeover_and_eject (7.6.6)
287
hosted_multiproject_buildsrc_includebuild
275
hosted_stale_lock_fails
248
hosted_tamper_fails
242
Most other hosted tests take 140–190 s.
In the agent leg, three suites (12 s, 200 s, 373 s) run back-to-back in a for loop, so each suite's tail is serialized.
Root cause
The shards from #1133 split by test-name prefix (ci.yml around lines 1240–1255, enforced by scripts/test_ci_gradle_prefixes.py), not by measured duration. Per Gradle line the total is about 1,900 s across 4 legs, which would be about 480 s each if balanced, but the heaviest leg takes 700–800 s.
Each test is slow by design: it uses a fresh GRADLE_USER_HOME and --no-daemon (gradle_build_common/mod.rs:283). Caching ~/.gradle would not help.
Proposed fix
Rebalance by duration. Move about 700 test-seconds into the vendor leg: hosted_config_cache_second_row and hosted_fallback_snippet_compiles_kotlin from b-p, and hosted_vendored_takeover_and_eject from the catch-all. Update the skip/filter lists and scripts/test_ci_gradle_prefixes.py so that every test still runs in exactly one leg.
Run the agent leg's three suites concurrently (& + wait, or one cargo test invocation with multiple --test) instead of a sequential for loop. That saves about 1 min on that leg.
Optional, needs a cost decision: use an 8-vCPU larger runner with --test-threads=8 for the two heaviest Gradle legs. The tests are CPU-bound, so this is an estimated ~6 min off the leg.
Expected saving
Critical leg goes from about 13.5 to about 10.5 min. That cuts ~3 min of merge_group wall time (p50 ~23 → ~20 min) on every queue entry, and ~3 min of PR push→ci-ok latency on Gradle-heavy runs.
[agent] Triaged as priority:p3 (CI-only). No open PR references this yet; it is not a duplicate of the other CI-perf reports (#1170–#1178 each target a different workflow cost).
Measurement
Merge queue. After Cut merge-group CI from ~46 to ~20 min: shard Gradle e2e and test legs, skip test-release in queue, cancel orphaned runs #1133 merged (2026-10-08 18:55Z), merge_group CI reaches
ci-okin 22–25 min. In 8 of the 9 non-cancelled runs since then, the last job to finish beforeci-okwas a Gradle leg ofe2e (ubuntu-latest, e2e_redirect_gradle_build | e2e_gradle_discovery_build …). Examples:PRs. Since 10-08 02:16Z, an
e2e (ubuntu-latest, …)leg was the last job beforeci-okin 94 of 111 successful PR CI runs. PR push →ci-okp50 is 41 min and p90 53 min.Per-leg times. "Run e2e tests" step in three merge_group runs (37828932401, 37828929076, 37822515034), per Gradle line:
hosted_[b-p]Job links: 8.14.3 b-p, 784 s, 7.6.6 catch-all, 694 s, 8.14.3 b-p, 803 s, 7.6.6 b-p, 750 s.
Where the time goes
Breakdown of job 113492360128, 817 s in total:
Slowest tests:
hosted_fallback_snippet_compiles_kotlin(8.14.3)hosted_config_cache_second_row(8.14.3)hosted_vendored_takeover_and_eject(7.6.6)hosted_multiproject_buildsrc_includebuildhosted_stale_lock_failshosted_tamper_failsMost other hosted tests take 140–190 s.
In the agent leg, three suites (12 s, 200 s, 373 s) run back-to-back in a
forloop, so each suite's tail is serialized.Root cause
The shards from #1133 split by test-name prefix (ci.yml around lines 1240–1255, enforced by
scripts/test_ci_gradle_prefixes.py), not by measured duration. Per Gradle line the total is about 1,900 s across 4 legs, which would be about 480 s each if balanced, but the heaviest leg takes 700–800 s.Each test is slow by design: it uses a fresh
GRADLE_USER_HOMEand--no-daemon(gradle_build_common/mod.rs:283). Caching~/.gradlewould not help.Proposed fix
hosted_config_cache_second_rowandhosted_fallback_snippet_compiles_kotlinfrom b-p, andhosted_vendored_takeover_and_ejectfrom the catch-all. Update the skip/filter lists andscripts/test_ci_gradle_prefixes.pyso that every test still runs in exactly one leg.&+wait, or onecargo testinvocation with multiple--test) instead of a sequentialforloop. That saves about 1 min on that leg.--test-threads=8for the two heaviest Gradle legs. The tests are CPU-bound, so this is an estimated ~6 min off the leg.Expected saving
ci-oklatency on Gradle-heavy runs.e2e-buildis done) also shortens this path. With both changes, the Gradle legs stay the long pole, so this remains additive.Coverage and risk
test_ci_gradle_prefixes.pymust keep asserting full and disjoint coverage.ci-okandclippyare unchanged.Effort
S
ROI
ROI = critical-path saving × confidence / effort. Critical-path minutes are weighted at 1.0 per minute.