Skip to content

ci: Frontend E2E hangs on Install Chromium and is killed by the job timeout, silently skipping every spec #1869

Description

@cristim

Frontend E2E (PR) has hung on Install Chromium and been killed by the job timeout four times today across two PRs (#1864 three times, #1868 once). It cleared on the third re-run for #1864; #1868 is on its first re-run now.

What happens

6. Install Chromium         -> cancelled   9m52s
7. Run Playwright tests     -> skipped
8. Upload failure artifacts -> skipped

npx playwright install --with-deps chromium hangs, consumes the job's timeout-minutes: 10, and GitHub records the run as cancelled. Steps 1-5 (checkout, node, deps, build) succeed every time, so the hang is in the browser download and is upstream of anything in the diff under review.

For contrast, a healthy run of this job completes end to end in 66 seconds, with Install Chromium taking 22 (run 32213189459, 2026-08-19T03:44Z). This is a hang, not slowness, so raising timeout-minutes would only change how long it takes to learn that.

Why it matters more than a flaky job

A cancelled run skips every spec, and cancelled is not failure. Any check that decides "green" by denylisting known failure conclusions reads this run as passing. The Playwright suite is the only evidence that a whole class of frontend defect is fixed: jsdom does not resolve stylesheets into layout, so the jest suite is structurally incapable of catching the zero-height box in #1777 or the zero-separation collision in #1776. On both PRs the e2e signal was absent while the run list showed nothing marked failed.

So the failure mode is a CI gate that silently does not run, which is the shape this repo has repeatedly paid for elsewhere (#1836, #1846, #1852).

Suggested fix

Either or both:

  1. Retry the install step, e.g. a nick-fields/retry wrapper around npx playwright install, so a hung download is retried rather than consuming the whole job budget.
  2. Cache the browser keyed on the Playwright version from package-lock.json, so the common path does not download at all.

Also worth considering: give Install Chromium its own shorter step-level timeout so the failure is attributable, and make any aggregate CI check treat green as an allowlist (success, skipped) rather than a denylist, so a future unfamiliar conclusion cannot read as passing.

Not doing here

No change was made to .github/workflows/frontend-e2e.yml as part of #1864 or #1868; both were re-run manually instead. Filed so the recurrence is tracked rather than absorbed as folklore.

Activity

  1. cristim commented on Aug 20, 2026

    @cristim
    MemberAuthor

    Correction to the counts in this issue

    The hang is real and every step-level detail here holds, but the count is wrong and the durations are bounds rather than measurements. Recording it so the record is right.

    gh run list and gh run view report only a run's latest attempt. Both affected runs were re-run manually to green, so the failed attempts vanish from the listing and what remains looks like separate successful runs. That is where "four times across two PRs" (and the later "six") came from.

    Walking repos/LeanerCloud/CUDly/actions/runs/<id>/attempts/<n>/jobs instead:

    run PR attempts Install Chromium per attempt
    32273788976 #1864 4 cancelled 585s, cancelled 586s, cancelled 586s, success 101s
    32293420755 #1868 3 cancelled 592s, cancelled 591s, success 27s

    So the correct figure is 5 cancelled attempts across 2 runs, not 4 or 6 runs. Run Playwright tests: skipped on all five, exactly as described.

    The durations are lower bounds

    All five were truncated by the job cap rather than self-terminating. Job wall clock is a constant 615-618s across all five against the then-600s timeout-minutes: 10, so the install simply consumed whatever was left after the preamble:

    run attempt job wall clock Install Chromium everything else
    32273788976 1 618s 585s 33s
    32273788976 2 615s 586s 29s
    32273788976 3 618s 586s 32s
    32293420755 1 615s 592s 23s
    32293420755 2 617s 591s 26s

    585-592s is therefore a floor. Nothing here says whether the install would have cleared at 601s or never, which matters for diagnosis: a slow download and a wedged lock are indistinguishable in this data.

    What the same walk established, which changed the fix

    Over the last 40 runs and every attempt of each: 39 successful installs (36 at or under 39s, outliers at 101s, 140s, 381s) against 5 wedges, none under 585s.

    That killed the retry this issue suggested. A retry capped below 381s turns healthy runs red, and two attempts capped above it do not fit the job budget alongside npm ci, the build and the specs. The 22s figure quoted above is the best case, not the distribution.

    The fix instead leans on a step-level timeout-minutes, because a step that busts its own timeout is marked failed, while a job that busts the job-level one is marked cancelled (actions/runner, src/Runner.Worker/StepsRunner.cs ~L322-337). That directly addresses this issue's central complaint that cancelled is not failure.

    The remaining half, that Playwright Chromium is not a required status check so a red X still gates nothing, is filed separately as LeanerCloud/cloud-commitments-platform#212 along with a lock-contention hypothesis for the wedge itself.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions