Repository navigation
Conversation
The E2E suite can only exercise GPU behaviour on the self-hosted lanes, because every hardware probe reads fixed host paths: /dev/kfd, the KFD topology, /sys/module/amdgpu, /dev/dxg, /proc/version, /etc/os-release and /usr/lib/wsl. A scenario cannot describe a host it is not running on. Route those reads through rocm_core::host_path. In a build with the new rocm-core e2e-test-hooks feature it re-roots absolute paths under ROCM_CLI_TEST_HOST_ROOT; without the feature it is the identity and never reads the environment, so a release build always probes the real machine. The CLI keeps printing the logical path, so the root changes what the probes find, not what users read. The feature is forwarded from rocm, rocmd and both engines so the existing `--features rocm/e2e-test-hooks` build turns it on everywhere. Signed-off-by: Roman Inflianskas <Roman.Inflianskas@amd.com>
GPU scenarios are about to stop depending on the runner's hardware: they will plant a simulated host instead, so they run on GitHub-hosted runners on every PR. A few premises cannot be simulated, such as a model actually generating tokens on the device, and those need a way to say so. @requires-real-gpu implies @requires-gpu and also skips unless the run sets E2E_HARDWARE=real. The default is `simulated`. Any other value aborts the run, because a misspelt `real` would otherwise skip every real-GPU scenario and still report green. Every self-hosted GPU lane, per-PR and nightly, sets E2E_HARDWARE=real now, and the WSL lanes forward it through WSLENV. Nothing carries the tag yet, so no scenario changes lanes in this commit. Signed-off-by: Roman Inflianskas <Roman.Inflianskas@amd.com>
GPU scenarios could only run where the runner had the right hardware. The CLI reads the real machine, so detection and refusal paths that never touch the device were confined to the GPU lanes, multi-GPU premises to the hosts with several cards, and WSL premises to the one WSL lane. Scenarios now name the machine they need in a Given step, for example "a machine with eight AMD Instinct GPUs", "a bare-metal Linux machine with an AMD GPU" or "a WSL machine with an AMD GPU passed through". The step plants a directory laid out the way the kernel lays out /dev, /sys and /proc, points ROCM_CLI_TEST_HOST_ROOT at it, and isolates the commands from the runner: - Commands get a PATH with the hardware tools removed (uname, lspci, ldconfig, rocminfo, amd-smi, the kernel log, the WSL interop shells and others). The tools the simulated machine has come back as a fake-host-tool stand-in that answers from canned output. - The runner's GPU visibility masks and ROCm variables are unset. - HOME is moved into the scenario. "A managed runtime is active" plants a runtime record whose vllm is a fake-vllm stand-in. A serve on a simulated machine therefore runs the real engine selection, runtime lookup, launch and readiness wait, and only the process at the end is simulated. Twenty scenarios move onto simulated machines and run on every Linux lane, including the GitHub-hosted one: examine-04/10/11/17, diagnose-08/13/15/16/17/18, driver-install-01..03 and serve-11/13/15/16/19/20/21. examine-18 repeats examine-04 against the real machine, so a real GPU lane still covers detection there, including the Windows lane. Every scenario left on @requires-gpu needs a physical device or a runtime really installed for one, so those become @requires-real-gpu. Signed-off-by: Roman Inflianskas <Roman.Inflianskas@amd.com>
Scenarios on a simulated machine fail only when the CLI misreads the fixture. If a kernel or driver update moves something the fixture assumes, those scenarios keep passing while the CLI fails on hardware. That is how the KFD target came to be read from a file no kernel has. Two @nightly @requires-real-gpu scenarios now hold the fixture against a real machine. On a bare-metal GPU host they check: - each KFD GPU node states gfx_target_version inside its properties file - the property keys the fixture writes are present - gpu_id files exist - render nodes match drm_render_minor - CPU nodes report a target of 0 - the AMD DRM cards have vendor and device files - `lspci -nn -D` lines have the expected shape On a WSL2 distribution they check /dev/dxg, libdxcore.so and the microsoft-standard-WSL2 kernel string. Each difference is reported along with what to update. Signed-off-by: Roman Inflianskas <Roman.Inflianskas@amd.com>
This was referenced Oct 2, 2026
rominf
force-pushed
the
test/e2e-host-root
branch
3 times, most recently
from
October 6, 2026 12:26
9ddd114 to
c766a25
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #510 (base branch
test/e2e-host-root, so the diff is this PR alone). Draft until #510 merges; it then moves tomain.Summary
Lets an E2E scenario describe the GPU machine it needs (eight Instinct GPUs, a WSL2 distribution with GPU passthrough, a box with no GPU) instead of depending on the runner's hardware. Using #510's
ROCM_CLI_TEST_HOST_ROOT, scenarios that only need the CLI to see a GPU now run on every Linux lane, including the GitHub-hostedE2E testslane on every PR. Three commits:E2E_HARDWARE=simulated|realand@requires-real-gpu.@requires-real-gpuimplies@requires-gpuand also skips unless the run setsE2E_HARDWARE=real.simulated. Any other value aborts the run, so a misspeltrealcan't quietly skip the real-GPU scenarios and still go green.real, and the WSL lanes forward it throughWSLENV.tests/e2e-cucumber/src/simulated_host.rs)./dev,/sysand/proc. That covers the KFD topology (gfx_target_versioninside each node'sproperties, as the kernel writes it), DRM cards,/dev/kfd//dev/dxg,/proc/versionand/etc/os-release.PATHwith the machine-describing tools removed (uname,lspci,ldconfig,rocminfo,amd-smi,mpirun, the kernel log, the WSL interop shells).fake-host-toolstand-in that answers from canned output; the rest are absent.HIP_VISIBLE_DEVICES,ROCR_VISIBLE_DEVICES,ROCM_PATH, ...) is unset, andHOMEis moved into the scenario.Given a managed runtime is activeplants a runtime record whosevllmis afake-vllmstand-in. A serve then runs the CLI's real engine selection, runtime lookup, managed launch and readiness wait; only the process at the end is simulated.@requires-gpu,@requires-wsl,@requires-bare-metal,@requires-no-gpuor@requires-multi-gputo@requires-os:linux:@requires-gpuneeds a physical device or a runtime really installed for one, so those become@requires-real-gpu./dev,/sys,/proc/versionandlspciof a real machine.Not simulated, because it is not hardware: package-manager state, container detection, and discovered ROCm installs (#510 deliberately doesn't re-root those). No simulated scenario asserts on them.
Risk: low for the product (no product code changes here); medium for the suite, since it changes which lane runs 20 scenarios. Each moved scenario still asserts the same behaviour.
Scenario coverage (AGENTS.md §3)
No CLI behaviour changes. This PR changes which machines existing scenarios run against, and adds examine-18/19/20. The moved scenarios gain coverage: they used to run only on the matching self-hosted lane (several only on the WSL lane or multi-GPU hosts) and now run on every Linux lane. The previous real-kernel cross-check is kept by examine-18 (per real GPU lane) and examine-19/20 (nightly).
Test plan
cargo xtask e2e, the full default suite on a WSL2 laptop with no GPU: 166 scenarios, 0 unexpected failures, 2 expected xfails.cargo xtask e2e -- --name ...: the 20 simulated scenarios plus serve-11 all pass on that WSL2 laptop. The bare-metal and Instinct ones therefore pass on a host that is itself WSL2 without a GPU, which shows the root and PATH isolation hold.cargo test -p e2e-cucumber --lib --bins: 148 + 1 tests, including fixture-layout tests and drift-check tests that reproduce the original "target read from a file no kernel has" bug and expect it to be reported.cargo test -p xtask workflow_contractpasses, and clippy is clean.examine-18/19/20 need real hardware, so they run only on the self-hosted lanes (examine-19/20 nightly). This PR's GPU lanes run examine-18.
If this PR fixes a bug, searched
tests/e2e-cucumber/expectations.tomlfor the fixed ticket ID and removed/narrowed any now-stale xfail rows. (n/a; none of the moved scenarios has an xfail row)If this PR adds a new subcommand or subsystem, its domain implementation lives in its own file per
docs/architecture.md. (n/a)Every new or changed user-facing message was read against the code path that runs after it. (n/a, no user-facing message changes)