Skip to content

ci(e2e): run only GPU scenarios on GPU lanes, smoke-only on PRs - #518

Draft
rominf wants to merge 2 commits into
test/e2e-simulated-machinesfrom
ci/e2e-gpu-smoke-on-prs
Draft

rominf wants to merge 2 commits into
test/e2e-simulated-machinesfrom
ci/e2e-gpu-smoke-on-prs

Conversation

@rominf

@rominf rominf commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Stacked on #517, which is stacked on #510 (so the diff is this PR alone). Draft until both merge.

Summary

On every pull request, and again in the merge queue, the self-hosted GPU lanes run almost the whole suite (~170 scenarios) to cover the dozen that need a GPU, and they're what sets merge latency. After #517, every GPU scenario that only needs the CLI to see a GPU runs on a simulated machine on the hosted lane, and the rest of the suite needs no GPU at all. So:

  • The five GPU lanes run only GPU scenarios. They set E2E_GPU_ONLY=1 on every event, so they run only scenarios that need the real hardware (@requires-real-gpu, @requires-gpu, @requires-gfx-target). Everything else resolves to a skip that names the reason, so the report grid shows it as not applicable there rather than missing. A skip that has already resolved keeps its own, more precise reason.
  • On PRs they run a smoke test. On pull_request and merge_group the lanes also set E2E_GPU_SMOKE_ONLY=1. That narrows the real-GPU set to the cheap per-engine canaries, now tagged @gpu-smoke: vLLM inference, lemonade inference, and the lemonade HF-checkpoint serve. In the queue the @merge-queue serves run too, as before. Pushes to main and release/**, dispatches and nightly run the full real-GPU set.
  • New GitHub-hosted E2E tests (Windows) job. It runs the suite on windows-latest. Until now the Strix Windows GPU lane was the only place the non-GPU scenarios met native Windows.
    • It reports under its own platform (mock-windows, rendered "Mock / Windows") and writes its own summary.
    • The required E2E consolidated report deliberately doesn't wait for it, so a cold Windows release build doesn't sit in front of every merge. That job now downloads only the Linux artifact by exact name, so whether Windows shows up no longer depends on which job finished first.
  • The WSL lane is unchanged. It has no usable GPU, so it keeps running the whole non-GPU suite inside a real WSL2 distribution, our only real-distro coverage.

Tradeoffs for maintainers

  • Real-GPU regressions are caught later. The merge queue no longer runs the full real-GPU suite. A regression outside the canaries and the @merge-queue serves shows up on the next push-to-main run or nightly, not before merge. The @gpu-smoke set is deliberately the unmasked per-engine checks: the lemonade default-engine serves are flaky-xfail on Linux lemonade (EAI-7423), so on their own they could never turn a lane red. Tag a scenario @gpu-smoke to keep it blocking.
  • Native-Windows non-GPU coverage is advisory for now. It moves from the required Strix Windows lane to the new hosted job, which isn't required yet. Once it has a green record, an admin should add E2E tests (Windows) to the required checks; a workflow edit can't do that.
  • Two chat scenarios lose a bonus real serve. chat-06 and chat-07 used to do a real serve on GPU lanes and a mock serve elsewhere. Now they only run with the mock. Their own comments note that real inference is covered by the serve-*-inference scenarios.
  • The Windows job starts immediately. It depends only on changes, not on the Linux build-and-test, and uses the sccache Azure backend under its own prefix (rocm-cli/e2e-windows), written only on push/merge_group, like every other job.

Making E2E tests (Windows) required

Checked against what can make a required check unsafe here:

  • It always reports. The job-level if: is the same always() form the required e2e job uses, with the work gated at step level. On a non-heavy PR it reports success after a few seconds, and on push/merge_group changes forces heavy, so the merge queue always runs it for real.
  • It runs in the merge queue. ci.yml triggers on merge_group.
  • No name collision. No other job is called E2E tests (Windows).
  • Fork-safe. No secret is required; without the cache credential it compiles without a shared cache.
  • Stable so far. 6 of 6 runs green on windows-latest (the PR run plus 5 parallel dispatches): identical 106 scenarios every time, 0 unexpected failures, and the same 2 deterministic xfails.
  • Latency. Cold (cache empty, as on every PR so far) it takes about 19–20 minutes, finishing about 6.5 minutes after the last required hosted check. About 14 of those minutes are compile, which the cache should cut once a push or merge_group run has written it. That's not measurable before merge, since PRs only read the cache. Either way it's well short of the self-hosted lanes, which are what bound merge latency.

Promote it only after this PR is on main. A required check that doesn't exist on the base blocks every PR that doesn't contain the job, and open PRs based before the merge need a rebase to produce it. I'd also let it build a short record on main first: the self-hosted lanes were promoted on dozens of runs, and six is a small sample for network-dependent scenarios. The admin step is POST /repos/ROCm/rocm-cli/branches/main/protection/required_status_checks/contexts with ["E2E tests (Windows)"], after backing up the current protection.

Scenario coverage (AGENTS.md §3)

CI plumbing only; no CLI behaviour changes. The selection rule is unit-tested (a_gpu_lane_runs_only_its_gpu_scenarios_narrowed_to_the_smoke_test_on_prs), and the lane wiring, WSLENV forwarding and report artifact names are pinned by xtask contract tests.

Test plan

  • cargo test -p e2e-cucumber --lib (148), cargo test -p xtask (252) and cargo test -p e2e-report (54) all pass. Clippy and fmt are clean.

  • E2E_GPU_ONLY=1 cargo xtask e2e locally on a GPU-less host: 0 scenarios run. All 224 resolve to skip; 166 carry the new reason, and the others keep their own (lifecycle, nightly, real GPU, docker).

  • The new Windows job: 6 of 6 green on windows-latest; see above. The GPU lanes ran under the new rule (Strix Windows: 2 smoke scenarios in 7 minutes).

  • If this PR fixes a bug, searched tests/e2e-cucumber/expectations.toml for the fixed ticket ID and removed/narrowed any now-stale xfail rows. (n/a)

  • If this PR adds a new subcommand or subsystem, its domain implementation lives in its own file per docs/architecture.md. (n/a)

  • Every new or changed user-facing message was read against the code path that runs after it. (n/a)

The self-hosted GPU lanes ran almost the whole suite (about 170
scenarios) on every pull request and again in the merge queue, to cover
the dozen that need a GPU. They set merge latency. Every GPU scenario
that only needs the CLI to see a GPU now runs on a simulated machine on
the hosted lane, and the rest of the suite needs no GPU at all.

- The five GPU lanes set E2E_GPU_ONLY=1 on every event. They run only
  scenarios that need the real hardware (@requires-real-gpu,
  @requires-gpu, @requires-gfx-target). Everything else resolves to a
  skip that names the reason, so the report shows it as not applicable
  there rather than missing. A skip that has already resolved keeps its
  own, more precise reason.
- On pull_request and merge_group they also set E2E_GPU_SMOKE_ONLY=1.
  This narrows the real-GPU set to the cheap per-engine canaries, now
  tagged @gpu-smoke: vLLM inference, lemonade inference and the lemonade
  HF-checkpoint serve. In the queue the @merge-queue serves run as well,
  as before. Pushes to main and release/**, dispatches and the nightly
  lanes run the full real-GPU set with E2E_HARDWARE=real.
- A new GitHub-hosted "E2E tests (Windows)" job runs the suite on
  windows-latest. Until now the Strix Windows GPU lane was the only place
  the non-GPU scenarios met native Windows. The job reports under its
  own platform slug (mock-windows) and summarizes itself. The required
  consolidated report does not wait for it, so a cold Windows build does
  not sit in front of every merge. The job is advisory until it has a
  green record.
- The WSL lane has no usable GPU and keeps running the whole non-GPU
  suite inside a real WSL2 distribution.

The workflow contract pins the lane variables on all six lanes, checks
that every E2E_* variable a WSL job sets is forwarded through WSLENV,
and registers the new report artifact.

Signed-off-by: Roman Inflianskas <Roman.Inflianskas@amd.com>
@rominf
rominf force-pushed the ci/e2e-gpu-smoke-on-prs branch from 3d98e24 to 95fc83b Compare October 5, 2026 07:52
@rominf rominf changed the title ci(e2e): narrow the GPU lanes to a smoke test on pull requests ci(e2e): run only GPU scenarios on GPU lanes, smoke-only on PRs Oct 5, 2026
On its first run the hosted Windows E2E job finished nine minutes after
the last required hosted check. Most of that time was not the suite:
- It waited for the Linux build-and-test, which says nothing about the
  Windows build.
- About 14 of its 17 minutes were a cold compile of xtask, the release
  binaries and the test harness. The 106 scenarios took about 3.

It now needs only `changes`. It also uses the same sccache Azure
backend as windows-build-and-test, under its own key prefix, because
this build carries e2e-test-hooks. Like every other job, it writes the
cache only on push and merge_group, so pull requests start reading a
warm cache once one of those has run.

Signed-off-by: Roman Inflianskas <Roman.Inflianskas@amd.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant