Skip to content

test(e2e): describe GPU hosts with simulated machines - #517

Draft
rominf wants to merge 4 commits into
test/e2e-host-rootfrom
test/e2e-simulated-machines
Draft

rominf wants to merge 4 commits into
test/e2e-host-rootfrom
test/e2e-simulated-machines

Conversation

@rominf

@rominf rominf commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Stacked on #510 (base branch test/e2e-host-root, so the diff is this PR alone). Draft until #510 merges; it then moves to main.

Summary

Lets an E2E scenario describe the GPU machine it needs (eight Instinct GPUs, a WSL2 distribution with GPU passthrough, a box with no GPU) instead of depending on the runner's hardware. Using #510's ROCM_CLI_TEST_HOST_ROOT, scenarios that only need the CLI to see a GPU now run on every Linux lane, including the GitHub-hosted E2E tests lane on every PR. Three commits:

  1. E2E_HARDWARE=simulated|real and @requires-real-gpu.
    • @requires-real-gpu implies @requires-gpu and also skips unless the run sets E2E_HARDWARE=real.
    • The default is simulated. Any other value aborts the run, so a misspelt real can't quietly skip the real-GPU scenarios and still go green.
    • Every self-hosted GPU lane (per-PR and nightly) sets real, and the WSL lanes forward it through WSLENV.
  2. Simulated machines (tests/e2e-cucumber/src/simulated_host.rs).
    • What a Given step plants: a directory laid out the way the kernel lays out /dev, /sys and /proc. That covers the KFD topology (gfx_target_version inside each node's properties, as the kernel writes it), DRM cards, /dev/kfd//dev/dxg, /proc/version and /etc/os-release.
    • Isolation from the runner:
      • Commands get a PATH with the machine-describing tools removed (uname, lspci, ldconfig, rocminfo, amd-smi, mpirun, the kernel log, the WSL interop shells).
      • The tools the simulated machine has come back as a small fake-host-tool stand-in that answers from canned output; the rest are absent.
      • The runner's own GPU scoping (HIP_VISIBLE_DEVICES, ROCR_VISIBLE_DEVICES, ROCM_PATH, ...) is unset, and HOME is moved into the scenario.
      • So neither a GPU runner nor a WSL laptop can answer for the simulated machine.
    • Serving on a simulated machine: Given a managed runtime is active plants a runtime record whose vllm is a fake-vllm stand-in. A serve then runs the CLI's real engine selection, runtime lookup, managed launch and readiness wait; only the process at the end is simulated.
    • Twenty scenarios move onto simulated machines, from @requires-gpu, @requires-wsl, @requires-bare-metal, @requires-no-gpu or @requires-multi-gpu to @requires-os:linux:
      • examine-04/10/11/17
      • diagnose-08/13/15/16/17/18
      • driver-install-01..03
      • serve-11/13/15/16/19/20/21
    • Real hardware is still covered: examine-18 repeats examine-04's assertions against the real machine on every GPU lane, the Windows one included.
    • Retagging: every scenario left on @requires-gpu needs a physical device or a runtime really installed for one, so those become @requires-real-gpu.
  3. A nightly drift check.
    • examine-19 (bare-metal GPU host) and examine-20 (WSL2) hold every fixture assumption against the live /dev, /sys, /proc/version and lspci of a real machine.
    • If a kernel or driver update moves something, the simulated scenarios would otherwise keep passing while the CLI fails on hardware; these fail instead, naming the exact difference.

Not simulated, because it is not hardware: package-manager state, container detection, and discovered ROCm installs (#510 deliberately doesn't re-root those). No simulated scenario asserts on them.

Risk: low for the product (no product code changes here); medium for the suite, since it changes which lane runs 20 scenarios. Each moved scenario still asserts the same behaviour.

Scenario coverage (AGENTS.md §3)

No CLI behaviour changes. This PR changes which machines existing scenarios run against, and adds examine-18/19/20. The moved scenarios gain coverage: they used to run only on the matching self-hosted lane (several only on the WSL lane or multi-GPU hosts) and now run on every Linux lane. The previous real-kernel cross-check is kept by examine-18 (per real GPU lane) and examine-19/20 (nightly).

Test plan

  • cargo xtask e2e, the full default suite on a WSL2 laptop with no GPU: 166 scenarios, 0 unexpected failures, 2 expected xfails.

  • cargo xtask e2e -- --name ...: the 20 simulated scenarios plus serve-11 all pass on that WSL2 laptop. The bare-metal and Instinct ones therefore pass on a host that is itself WSL2 without a GPU, which shows the root and PATH isolation hold.

  • cargo test -p e2e-cucumber --lib --bins: 148 + 1 tests, including fixture-layout tests and drift-check tests that reproduce the original "target read from a file no kernel has" bug and expect it to be reported.

  • cargo test -p xtask workflow_contract passes, and clippy is clean.

  • examine-18/19/20 need real hardware, so they run only on the self-hosted lanes (examine-19/20 nightly). This PR's GPU lanes run examine-18.

  • If this PR fixes a bug, searched tests/e2e-cucumber/expectations.toml for the fixed ticket ID and removed/narrowed any now-stale xfail rows. (n/a; none of the moved scenarios has an xfail row)

  • If this PR adds a new subcommand or subsystem, its domain implementation lives in its own file per docs/architecture.md. (n/a)

  • Every new or changed user-facing message was read against the code path that runs after it. (n/a, no user-facing message changes)

rominf added 4 commits October 2, 2026 13:38
The E2E suite can only exercise GPU behaviour on the self-hosted lanes,
because every hardware probe reads fixed host paths: /dev/kfd, the KFD
topology, /sys/module/amdgpu, /dev/dxg, /proc/version, /etc/os-release
and /usr/lib/wsl. A scenario cannot describe a host it is not running on.

Route those reads through rocm_core::host_path. In a build with the new
rocm-core e2e-test-hooks feature it re-roots absolute paths under
ROCM_CLI_TEST_HOST_ROOT; without the feature it is the identity and never
reads the environment, so a release build always probes the real machine.
The CLI keeps printing the logical path, so the root changes what the
probes find, not what users read.

The feature is forwarded from rocm, rocmd and both engines so the
existing `--features rocm/e2e-test-hooks` build turns it on everywhere.

Signed-off-by: Roman Inflianskas <Roman.Inflianskas@amd.com>
GPU scenarios are about to stop depending on the runner's hardware: they
will plant a simulated host instead, so they run on GitHub-hosted runners
on every PR. A few premises cannot be simulated, such as a model actually
generating tokens on the device, and those need a way to say so.

@requires-real-gpu implies @requires-gpu and also skips unless the run
sets E2E_HARDWARE=real. The default is `simulated`. Any other value aborts
the run, because a misspelt `real` would otherwise skip every real-GPU
scenario and still report green.

Every self-hosted GPU lane, per-PR and nightly, sets E2E_HARDWARE=real
now, and the WSL lanes forward it through WSLENV. Nothing carries the tag
yet, so no scenario changes lanes in this commit.

Signed-off-by: Roman Inflianskas <Roman.Inflianskas@amd.com>
GPU scenarios could only run where the runner had the right hardware.
The CLI reads the real machine, so detection and refusal paths that never
touch the device were confined to the GPU lanes, multi-GPU premises to
the hosts with several cards, and WSL premises to the one WSL lane.

Scenarios now name the machine they need in a Given step, for example "a
machine with eight AMD Instinct GPUs", "a bare-metal Linux machine with
an AMD GPU" or "a WSL machine with an AMD GPU passed through". The step
plants a directory laid out the way the kernel lays out /dev, /sys and
/proc, points ROCM_CLI_TEST_HOST_ROOT at it, and isolates the commands
from the runner:
- Commands get a PATH with the hardware tools removed (uname, lspci,
  ldconfig, rocminfo, amd-smi, the kernel log, the WSL interop shells
  and others). The tools the simulated machine has come back as a
  fake-host-tool stand-in that answers from canned output.
- The runner's GPU visibility masks and ROCm variables are unset.
- HOME is moved into the scenario.

"A managed runtime is active" plants a runtime record whose vllm is a
fake-vllm stand-in. A serve on a simulated machine therefore runs the
real engine selection, runtime lookup, launch and readiness wait, and
only the process at the end is simulated.

Twenty scenarios move onto simulated machines and run on every Linux
lane, including the GitHub-hosted one: examine-04/10/11/17,
diagnose-08/13/15/16/17/18, driver-install-01..03 and
serve-11/13/15/16/19/20/21. examine-18 repeats examine-04 against the
real machine, so a real GPU lane still covers detection there,
including the Windows lane. Every scenario left on @requires-gpu needs a
physical device or a runtime really installed for one, so those become
@requires-real-gpu.

Signed-off-by: Roman Inflianskas <Roman.Inflianskas@amd.com>
Scenarios on a simulated machine fail only when the CLI misreads the
fixture. If a kernel or driver update moves something the fixture
assumes, those scenarios keep passing while the CLI fails on hardware.
That is how the KFD target came to be read from a file no kernel has.

Two @nightly @requires-real-gpu scenarios now hold the fixture against
a real machine.

On a bare-metal GPU host they check:
- each KFD GPU node states gfx_target_version inside its properties file
- the property keys the fixture writes are present
- gpu_id files exist
- render nodes match drm_render_minor
- CPU nodes report a target of 0
- the AMD DRM cards have vendor and device files
- `lspci -nn -D` lines have the expected shape

On a WSL2 distribution they check /dev/dxg, libdxcore.so and the
microsoft-standard-WSL2 kernel string.

Each difference is reported along with what to update.

Signed-off-by: Roman Inflianskas <Roman.Inflianskas@amd.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant