Conversation
3 tasks done
The self-hosted GPU lanes ran almost the whole suite (about 170 scenarios) on every pull request and again in the merge queue, to cover the dozen that need a GPU. They set merge latency. Every GPU scenario that only needs the CLI to see a GPU now runs on a simulated machine on the hosted lane, and the rest of the suite needs no GPU at all. - The five GPU lanes set E2E_GPU_ONLY=1 on every event. They run only scenarios that need the real hardware (@requires-real-gpu, @requires-gpu, @requires-gfx-target). Everything else resolves to a skip that names the reason, so the report shows it as not applicable there rather than missing. A skip that has already resolved keeps its own, more precise reason. - On pull_request and merge_group they also set E2E_GPU_SMOKE_ONLY=1. This narrows the real-GPU set to the cheap per-engine canaries, now tagged @gpu-smoke: vLLM inference, lemonade inference and the lemonade HF-checkpoint serve. In the queue the @merge-queue serves run as well, as before. Pushes to main and release/**, dispatches and the nightly lanes run the full real-GPU set with E2E_HARDWARE=real. - A new GitHub-hosted "E2E tests (Windows)" job runs the suite on windows-latest. Until now the Strix Windows GPU lane was the only place the non-GPU scenarios met native Windows. The job reports under its own platform slug (mock-windows) and summarizes itself. The required consolidated report does not wait for it, so a cold Windows build does not sit in front of every merge. The job is advisory until it has a green record. - The WSL lane has no usable GPU and keeps running the whole non-GPU suite inside a real WSL2 distribution. The workflow contract pins the lane variables on all six lanes, checks that every E2E_* variable a WSL job sets is forwarded through WSLENV, and registers the new report artifact. Signed-off-by: Roman Inflianskas <Roman.Inflianskas@amd.com>
rominf
force-pushed
the
ci/e2e-gpu-smoke-on-prs
branch
from
October 5, 2026 07:52
3d98e24 to
95fc83b
Compare
On its first run the hosted Windows E2E job finished nine minutes after the last required hosted check. Most of that time was not the suite: - It waited for the Linux build-and-test, which says nothing about the Windows build. - About 14 of its 17 minutes were a cold compile of xtask, the release binaries and the test harness. The 106 scenarios took about 3. It now needs only `changes`. It also uses the same sccache Azure backend as windows-build-and-test, under its own key prefix, because this build carries e2e-test-hooks. Like every other job, it writes the cache only on push and merge_group, so pull requests start reading a warm cache once one of those has run. Signed-off-by: Roman Inflianskas <Roman.Inflianskas@amd.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #517, which is stacked on #510 (so the diff is this PR alone). Draft until both merge.
Summary
On every pull request, and again in the merge queue, the self-hosted GPU lanes run almost the whole suite (~170 scenarios) to cover the dozen that need a GPU, and they're what sets merge latency. After #517, every GPU scenario that only needs the CLI to see a GPU runs on a simulated machine on the hosted lane, and the rest of the suite needs no GPU at all. So:
E2E_GPU_ONLY=1on every event, so they run only scenarios that need the real hardware (@requires-real-gpu,@requires-gpu,@requires-gfx-target). Everything else resolves to a skip that names the reason, so the report grid shows it as not applicable there rather than missing. A skip that has already resolved keeps its own, more precise reason.pull_requestandmerge_groupthe lanes also setE2E_GPU_SMOKE_ONLY=1. That narrows the real-GPU set to the cheap per-engine canaries, now tagged@gpu-smoke: vLLM inference, lemonade inference, and the lemonade HF-checkpoint serve. In the queue the@merge-queueserves run too, as before. Pushes tomainandrelease/**, dispatches and nightly run the full real-GPU set.E2E tests (Windows)job. It runs the suite onwindows-latest. Until now the Strix Windows GPU lane was the only place the non-GPU scenarios met native Windows.mock-windows, rendered "Mock / Windows") and writes its own summary.E2E consolidated reportdeliberately doesn't wait for it, so a cold Windows release build doesn't sit in front of every merge. That job now downloads only the Linux artifact by exact name, so whether Windows shows up no longer depends on which job finished first.Tradeoffs for maintainers
@merge-queueserves shows up on the next push-to-mainrun or nightly, not before merge. The@gpu-smokeset is deliberately the unmasked per-engine checks: the lemonade default-engine serves are flaky-xfail on Linux lemonade (EAI-7423), so on their own they could never turn a lane red. Tag a scenario@gpu-smoketo keep it blocking.E2E tests (Windows)to the required checks; a workflow edit can't do that.chat-06andchat-07used to do a real serve on GPU lanes and a mock serve elsewhere. Now they only run with the mock. Their own comments note that real inference is covered by theserve-*-inferencescenarios.changes, not on the Linuxbuild-and-test, and uses the sccache Azure backend under its own prefix (rocm-cli/e2e-windows), written only on push/merge_group, like every other job.Making
E2E tests (Windows)requiredChecked against what can make a required check unsafe here:
if:is the samealways()form the requirede2ejob uses, with the work gated at step level. On a non-heavy PR it reports success after a few seconds, and onpush/merge_groupchangesforcesheavy, so the merge queue always runs it for real.ci.ymltriggers onmerge_group.E2E tests (Windows).windows-latest(the PR run plus 5 parallel dispatches): identical 106 scenarios every time, 0 unexpected failures, and the same 2 deterministic xfails.merge_grouprun has written it. That's not measurable before merge, since PRs only read the cache. Either way it's well short of the self-hosted lanes, which are what bound merge latency.Promote it only after this PR is on
main. A required check that doesn't exist on the base blocks every PR that doesn't contain the job, and open PRs based before the merge need a rebase to produce it. I'd also let it build a short record onmainfirst: the self-hosted lanes were promoted on dozens of runs, and six is a small sample for network-dependent scenarios. The admin step isPOST /repos/ROCm/rocm-cli/branches/main/protection/required_status_checks/contextswith["E2E tests (Windows)"], after backing up the current protection.Scenario coverage (AGENTS.md §3)
CI plumbing only; no CLI behaviour changes. The selection rule is unit-tested (
a_gpu_lane_runs_only_its_gpu_scenarios_narrowed_to_the_smoke_test_on_prs), and the lane wiring, WSLENV forwarding and report artifact names are pinned byxtaskcontract tests.Test plan
cargo test -p e2e-cucumber --lib(148),cargo test -p xtask(252) andcargo test -p e2e-report(54) all pass. Clippy and fmt are clean.E2E_GPU_ONLY=1 cargo xtask e2elocally on a GPU-less host: 0 scenarios run. All 224 resolve to skip; 166 carry the new reason, and the others keep their own (lifecycle, nightly, real GPU, docker).The new Windows job: 6 of 6 green on
windows-latest; see above. The GPU lanes ran under the new rule (Strix Windows: 2 smoke scenarios in 7 minutes).If this PR fixes a bug, searched
tests/e2e-cucumber/expectations.tomlfor the fixed ticket ID and removed/narrowed any now-stale xfail rows. (n/a)If this PR adds a new subcommand or subsystem, its domain implementation lives in its own file per
docs/architecture.md. (n/a)Every new or changed user-facing message was read against the code path that runs after it. (n/a)