Repository navigation
ci: Frontend E2E hangs on Install Chromium and is killed by the job timeout, silently skipping every spec #1869
Description
Activity
- addedtriagedItem has been triagedItem has been triagedpriority/p1Next up; this sprintNext up; this sprintseverity/highSignificant harmSignificant harmurgency/this-sprintWithin the current sprintWithin the current sprintimpact/internalTeam-internal onlyTeam-internal onlyeffort/sHoursHourstype/choreMaintenance / non-user-visibleMaintenance / non-user-visible
on Aug 19, 2026 Correction to the counts in this issue
The hang is real and every step-level detail here holds, but the count is wrong and the durations are bounds rather than measurements. Recording it so the record is right.
gh run listandgh run viewreport only a run's latest attempt. Both affected runs were re-run manually to green, so the failed attempts vanish from the listing and what remains looks like separate successful runs. That is where "four times across two PRs" (and the later "six") came from.Walking
repos/LeanerCloud/CUDly/actions/runs/<id>/attempts/<n>/jobsinstead:run PR attempts Install Chromiumper attempt32273788976 #1864 4 cancelled 585s, cancelled 586s, cancelled 586s, success 101s 32293420755 #1868 3 cancelled 592s, cancelled 591s, success 27s So the correct figure is 5 cancelled attempts across 2 runs, not 4 or 6 runs.
Run Playwright tests: skippedon all five, exactly as described.The durations are lower bounds
All five were truncated by the job cap rather than self-terminating. Job wall clock is a constant 615-618s across all five against the then-600s
timeout-minutes: 10, so the install simply consumed whatever was left after the preamble:run attempt job wall clock Install Chromiumeverything else 32273788976 1 618s 585s 33s 32273788976 2 615s 586s 29s 32273788976 3 618s 586s 32s 32293420755 1 615s 592s 23s 32293420755 2 617s 591s 26s 585-592s is therefore a floor. Nothing here says whether the install would have cleared at 601s or never, which matters for diagnosis: a slow download and a wedged lock are indistinguishable in this data.
What the same walk established, which changed the fix
Over the last 40 runs and every attempt of each: 39 successful installs (36 at or under 39s, outliers at 101s, 140s, 381s) against 5 wedges, none under 585s.
That killed the retry this issue suggested. A retry capped below 381s turns healthy runs red, and two attempts capped above it do not fit the job budget alongside
npm ci, the build and the specs. The 22s figure quoted above is the best case, not the distribution.The fix instead leans on a step-level
timeout-minutes, because a step that busts its own timeout is markedfailed, while a job that busts the job-level one is markedcancelled(actions/runner,src/Runner.Worker/StepsRunner.cs~L322-337). That directly addresses this issue's central complaint thatcancelledis notfailure.The remaining half, that
Playwright Chromiumis not a required status check so a red X still gates nothing, is filed separately as LeanerCloud/cloud-commitments-platform#212 along with a lock-contention hypothesis for the wedge itself.- added a commit that references this issue
on Sep 27, 2026
Frontend E2E (PR)has hung onInstall Chromiumand been killed by the job timeout four times today across two PRs (#1864 three times, #1868 once). It cleared on the third re-run for #1864; #1868 is on its first re-run now.What happens
npx playwright install --with-deps chromiumhangs, consumes the job'stimeout-minutes: 10, and GitHub records the run ascancelled. Steps 1-5 (checkout, node, deps, build) succeed every time, so the hang is in the browser download and is upstream of anything in the diff under review.For contrast, a healthy run of this job completes end to end in 66 seconds, with
Install Chromiumtaking 22 (run 32213189459, 2026-08-19T03:44Z). This is a hang, not slowness, so raisingtimeout-minuteswould only change how long it takes to learn that.Why it matters more than a flaky job
A cancelled run skips every spec, and
cancelledis notfailure. Any check that decides "green" by denylisting known failure conclusions reads this run as passing. The Playwright suite is the only evidence that a whole class of frontend defect is fixed: jsdom does not resolve stylesheets into layout, so the jest suite is structurally incapable of catching the zero-height box in #1777 or the zero-separation collision in #1776. On both PRs the e2e signal was absent while the run list showed nothing marked failed.So the failure mode is a CI gate that silently does not run, which is the shape this repo has repeatedly paid for elsewhere (#1836, #1846, #1852).
Suggested fix
Either or both:
nick-fields/retrywrapper aroundnpx playwright install, so a hung download is retried rather than consuming the whole job budget.package-lock.json, so the common path does not download at all.Also worth considering: give
Install Chromiumits own shorter step-level timeout so the failure is attributable, and make any aggregate CI check treat green as an allowlist (success,skipped) rather than a denylist, so a future unfamiliar conclusion cannot read as passing.Not doing here
No change was made to
.github/workflows/frontend-e2e.ymlas part of #1864 or #1868; both were re-run manually instead. Filed so the recurrence is tracked rather than absorbed as folklore.