Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
152 changes: 146 additions & 6 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -1520,15 +1520,155 @@ jobs:
name: e2e-report
path: tests/e2e-cucumber/results/

# The same suite on a GitHub-hosted Windows runner. The self-hosted GPU lanes
# (e2e-selfhosted.yml) run only the scenarios that need their GPU
# (E2E_GPU_ONLY), so this is where everything else meets native Windows —
# path handling, the Windows probes, the PowerShell surfaces. Built with
# `rocm/e2e-test-hooks` by `cargo xtask e2e`, like the Linux lane above and
# unlike windows-build-and-test's lifecycle run, which must ship exactly what
# a release ships. Scenarios on a simulated machine are `@requires-os:linux`
# and skip here: the Windows probes ask the registry, not a filesystem.
e2e-windows:
name: E2E tests (Windows)
timeout-minutes: 60
runs-on: windows-latest
# Only `changes`: unlike the Linux `e2e` job this does not wait for the
# Linux `build-and-test`, which tells it nothing about the Windows build and
# would put that job's duration in front of this one's.
needs: [changes]
# Same compile cache as windows-build-and-test (see build-and-test for the
# design: Azure backend, written only by push and merge_group), under its
# own prefix — this build carries `rocm/e2e-test-hooks`, so its entries
# would never match that job's anyway.
env:
CARGO_INCREMENTAL: "0"
RUSTC_WRAPPER: sccache
SCCACHE_AZURE_CONNECTION_STRING: ${{ secrets.CACHE_AZURE_CONNECTION_STRING }}
SCCACHE_AZURE_BLOB_CONTAINER: cache
SCCACHE_AZURE_KEY_PREFIX: rocm-cli/e2e-windows
SCCACHE_AZURE_RW_MODE: ${{ (github.event_name == 'push' || github.event_name == 'merge_group') && 'READ_WRITE' || 'READ_ONLY' }}
if: >-
always()
&& needs.changes.result == 'success'
&& (github.event_name != 'workflow_dispatch'
|| inputs.platform == 'all' || inputs.platform == 'mock')
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1

# Whether the real work runs this time: on a dispatch, or when the change
# is heavy (`merge_group` forces heavy). The job itself always runs, so the
# check is always reported — required checks that never report stall the
# merge queue.
- name: Decide whether to run
id: gate
shell: bash
run: |
if [ "${{ github.event_name }}" = "workflow_dispatch" ] \
|| [ "${{ needs.changes.outputs.heavy }}" = "true" ]; then
echo "run=true" >> "$GITHUB_OUTPUT"
else
echo "run=false" >> "$GITHUB_OUTPUT"
fi

# Installs sccache. Runs BEFORE setup-rust-toolchain, whose rust-cache step
# runs `cargo metadata` and honours RUSTC_WRAPPER. `version` is pinned so
# the compiler wrapper does not float with the latest sccache release.
# continue-on-error for the same reason as build-and-test: a failed install
# is handled by the probe, not by failing the job.
- name: Set up sccache
if: steps.gate.outputs.run == 'true'
continue-on-error: true
uses: mozilla-actions/sccache-action@fc920bf0ec8de6ee65d409111f7ec508035751ba # v0.0.11
with:
version: v0.17.0

# Keep the remote backend only if it answers — see the equivalent step in
# build-and-test for why this is a probe, and why failure clears the Azure
# variables rather than the wrapper.
- name: Probe sccache cache backend
id: sccache_probe
if: steps.gate.outputs.run == 'true'
shell: pwsh
run: |
# This script decides whether to use the cache; a non-zero status from
# the commands it runs is an answer, not a job failure. Pin the native
# error preference off so a future runner default cannot turn a failed
# probe into a failed step.
$PSNativeCommandUseErrorActionPreference = $false
# Every Out-File below passes -Encoding utf8 explicitly. PowerShell 7
# (`shell: pwsh`) already defaults to UTF-8, but Windows PowerShell 5.1
# defaults to UTF-16LE, which the runner cannot parse — so an edit that
# switched this step to `shell: powershell` would silently corrupt the
# environment file rather than fail visibly.
function Disable-Backend {
"SCCACHE_AZURE_CONNECTION_STRING=" | Out-File -FilePath $env:GITHUB_ENV -Append -Encoding utf8
"SCCACHE_AZURE_BLOB_CONTAINER=" | Out-File -FilePath $env:GITHUB_ENV -Append -Encoding utf8
}
if (-not (Get-Command sccache -ErrorAction SilentlyContinue)) {
Write-Output "::warning::sccache is not installed; compiling without it."
"RUSTC_WRAPPER=" | Out-File -FilePath $env:GITHUB_ENV -Append -Encoding utf8
# Nothing will read the credential now; don't leave it in the job env.
Disable-Backend
exit 0
}
# sccache is the compiler wrapper from here on, remote backend or not,
# so there are statistics worth reporting either way.
"active=true" | Out-File -FilePath $env:GITHUB_OUTPUT -Append -Encoding utf8
if (-not $env:SCCACHE_AZURE_CONNECTION_STRING) {
Write-Output "::notice::No cache credential available (expected on forks); compiling without a shared cache."
Disable-Backend
exit 0
}
sccache --start-server
if ($LASTEXITCODE -ne 0) {
Write-Output "::warning::sccache could not reach the cache backend; compiling without a shared cache."
# Leave no half-started server behind for the cargo steps to find.
sccache --stop-server *> $null
Disable-Backend
}
exit 0

- uses: actions-rust-lang/setup-rust-toolchain@166cdcfd11aee3cb47222f9ddb555ce30ddb9659 # v1.17.0
if: steps.gate.outputs.run == 'true'
with:
cache-on-failure: true

- name: Run E2E tests
if: steps.gate.outputs.run == 'true'
shell: pwsh
run: cargo xtask e2e

- name: sccache stats
if: always() && steps.sccache_probe.outputs.active == 'true'
shell: pwsh
run: sccache --show-stats

# Its own grid in the run Summary; see e2e-report for why the consolidated
# (required) report does not wait for this lane.
- name: Report summary
if: always() && steps.gate.outputs.run == 'true'
shell: pwsh
run: cargo xtask e2e-report --artifacts-dir tests/e2e-cucumber/results --html-out tests/e2e-cucumber/results/consolidated.html >> $env:GITHUB_STEP_SUMMARY

- name: Upload E2E report
if: always() && steps.gate.outputs.run == 'true'
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: e2e-windows-report
path: tests/e2e-cucumber/results/

# Consolidate the mock (GitHub-hosted) platform's report into the cross-platform
# report: a platform × tier matrix in the run Summary plus a single merged HTML
# artifact. `if: always()` so a failing platform still appears; runs on
# GitHub-hosted ubuntu (no GPU needed — it only parses report.json).
#
# The self-hosted GPU platforms are consolidated by e2e-selfhosted.yml's own
# report job (they moved there so an offline runner can't stall this workflow's
# required checks). This job only depends on the mock `e2e` lane; the `*-report`
# glob still auto-includes any new GitHub-hosted platform added here.
# required checks). This job only depends on the mock `e2e` lane, and downloads
# only its artifact. `e2e-windows` summarizes itself instead: waiting on it here
# would put a cold Windows release build in front of this REQUIRED check, and
# matching it by glob would include it or not depending on which job finished
# first.
e2e-report:
name: E2E consolidated report
runs-on: ubuntu-latest
Expand All @@ -1549,9 +1689,9 @@ jobs:

- uses: actions-rust-lang/setup-rust-toolchain@166cdcfd11aee3cb47222f9ddb555ce30ddb9659 # v1.17.0

# Pull every e2e artifact matching the glob. `xtask e2e-report` turns each
# into one labeled platform and the `*-report` glob makes new platforms
# appear automatically. Layout note: download-artifact@v8 extracts a match
# Pull the mock lane's artifact (named exactly; see the job comment for why
# not a glob). `xtask e2e-report` turns it into one labeled platform.
# Layout note: download-artifact@v8 extracts a match
# into e2e-artifacts/<artifact-name>/ when several match, but flattens
# straight into e2e-artifacts/ when EXACTLY ONE matches (its source picks
# the root path when `artifacts.length === 1`). After the ci.yml ⇄
Expand All @@ -1561,7 +1701,7 @@ jobs:
- name: Download all E2E reports
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
with:
pattern: '*-report'
pattern: 'e2e-report'
path: e2e-artifacts

- name: Build consolidated report + step summary
Expand Down
35 changes: 30 additions & 5 deletions .github/workflows/e2e-selfhosted.yml
Original file line number Diff line number Diff line change
Expand Up @@ -181,14 +181,25 @@ jobs:
# a single nightly scenario (e.g. the 27B serve) without the full nightly run.
E2E_INCLUDE_NIGHTLY: "${{ inputs.include_nightly && '1' || '' }}"
# Opt-in @merge-queue serves: the heavy real serves (default-engine +
# readiness) run only in the merge queue, where a cheaper per-engine canary
# (scenarios 5 vLLM / 7 lemonade) has already guarded the PR. This keeps the
# per-PR GPU run short while still exercising the full serve matrix before a
# change lands. Only set on the `merge_group` event.
# readiness) run only in the merge queue, where the cheaper per-engine
# `@gpu-smoke` canaries have already guarded the PR. Only set on the
# `merge_group` event.
E2E_MERGE_QUEUE: "${{ github.event_name == 'merge_group' && '1' || '' }}"
# On a pull request and in the merge queue, of the scenarios that need a
# real GPU only the `@gpu-smoke` canaries (and, in the queue, the
# @merge-queue serves) run. Every GPU scenario that only needs the CLI to
# see a GPU runs on a simulated machine on the hosted `E2E tests` lane. The
# full real-GPU suite runs on pushes to main/release, on dispatch, and
# nightly. Keeps the GPU lanes, which set merge latency, short.
E2E_GPU_SMOKE_ONLY: "${{ (github.event_name == 'merge_group' || github.event_name == 'pull_request') && '1' || '' }}"
# A real GPU lane: run the `@requires-real-gpu` scenarios that a simulated
# host cannot stand in for. See HardwareMode in src/expectation.rs.
E2E_HARDWARE: real
# ...and ONLY those, on every event: everything that needs no GPU runs on
# the GitHub-hosted Linux and Windows lanes in ci.yml (`E2E tests`,
# `E2E tests (Windows)`), so a GPU runner's time is not spent on it. See
# restrict_to_lane in src/expectation.rs.
E2E_GPU_ONLY: "1"
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1

Expand Down Expand Up @@ -421,8 +432,11 @@ jobs:
E2E_INCLUDE_NIGHTLY: "${{ inputs.include_nightly && '1' || '' }}"
# Heavy @merge-queue serves run only in the merge queue; see e2e-gpu.
E2E_MERGE_QUEUE: "${{ github.event_name == 'merge_group' && '1' || '' }}"
# GPU smoke test only on pull requests and in the merge queue; see the e2e-gpu lane.
E2E_GPU_SMOKE_ONLY: "${{ (github.event_name == 'merge_group' || github.event_name == 'pull_request') && '1' || '' }}"
# See the e2e-gpu lane.
E2E_HARDWARE: real
E2E_GPU_ONLY: "1"
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1

Expand Down Expand Up @@ -676,8 +690,11 @@ jobs:
E2E_INCLUDE_NIGHTLY: "${{ inputs.include_nightly && '1' || '' }}"
# Heavy @merge-queue serves run only in the merge queue; see e2e-gpu.
E2E_MERGE_QUEUE: "${{ github.event_name == 'merge_group' && '1' || '' }}"
# GPU smoke test only on pull requests and in the merge queue; see the e2e-gpu lane.
E2E_GPU_SMOKE_ONLY: "${{ (github.event_name == 'merge_group' || github.event_name == 'pull_request') && '1' || '' }}"
# See the e2e-gpu lane.
E2E_HARDWARE: real
E2E_GPU_ONLY: "1"
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1

Expand Down Expand Up @@ -964,6 +981,8 @@ jobs:
E2E_INCLUDE_NIGHTLY: "${{ inputs.include_nightly && '1' || '' }}"
# Heavy @merge-queue serves run only in the merge queue; see e2e-gpu.
E2E_MERGE_QUEUE: "${{ github.event_name == 'merge_group' && '1' || '' }}"
# GPU smoke test only on pull requests and in the merge queue; see the e2e-gpu lane.
E2E_GPU_SMOKE_ONLY: "${{ (github.event_name == 'merge_group' || github.event_name == 'pull_request') && '1' || '' }}"
# See the e2e-gpu lane.
E2E_HARDWARE: real
# apps/rocm/build.rs falls back to GitHub's own GITHUB_HEAD_REF/
Expand Down Expand Up @@ -1364,7 +1383,7 @@ jobs:
# Forward the job-level E2E_* env vars into the guest: `wsl.exe`
# does not share the calling process's environment unless a
# variable is named in WSLENV.
$env:WSLENV = 'E2E_SERVE_TIMEOUT_SECS:E2E_TUI_TIMEOUT_SECS:E2E_INCLUDE_NIGHTLY:E2E_MERGE_QUEUE:E2E_HARDWARE:ROCM_CLI_RELEASE_REF'
$env:WSLENV = 'E2E_SERVE_TIMEOUT_SECS:E2E_TUI_TIMEOUT_SECS:E2E_INCLUDE_NIGHTLY:E2E_MERGE_QUEUE:E2E_GPU_SMOKE_ONLY:E2E_HARDWARE:ROCM_CLI_RELEASE_REF'
& "$env:GITHUB_WORKSPACE\.github\scripts\Invoke-WslBash.ps1" -PipeFail -Script @'
export PATH="$HOME/.cargo/bin:$PATH"
cd /root/work/rocm-cli
Expand Down Expand Up @@ -1489,8 +1508,11 @@ jobs:
E2E_INCLUDE_NIGHTLY: "${{ inputs.include_nightly && '1' || '' }}"
# Full serve matrix at the merge-queue gate, like every other lane (see e2e-gpu).
E2E_MERGE_QUEUE: "${{ github.event_name == 'merge_group' && '1' || '' }}"
# GPU smoke test only on pull requests and in the merge queue; see the e2e-gpu lane.
E2E_GPU_SMOKE_ONLY: "${{ (github.event_name == 'merge_group' || github.event_name == 'pull_request') && '1' || '' }}"
# See the e2e-gpu lane.
E2E_HARDWARE: real
E2E_GPU_ONLY: "1"
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1

Expand Down Expand Up @@ -1619,8 +1641,11 @@ jobs:
E2E_INCLUDE_NIGHTLY: "${{ inputs.include_nightly && '1' || '' }}"
# Full serve matrix at the merge-queue gate, like every other lane (see e2e-gpu).
E2E_MERGE_QUEUE: "${{ github.event_name == 'merge_group' && '1' || '' }}"
# GPU smoke test only on pull requests and in the merge queue; see the e2e-gpu lane.
E2E_GPU_SMOKE_ONLY: "${{ (github.event_name == 'merge_group' || github.event_name == 'pull_request') && '1' || '' }}"
# See the e2e-gpu lane.
E2E_HARDWARE: real
E2E_GPU_ONLY: "1"
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1

Expand Down
3 changes: 3 additions & 0 deletions crates/e2e-report/src/consolidated.rs
Original file line number Diff line number Diff line change
Expand Up @@ -67,6 +67,8 @@ fn parse_descriptor(name: &str) -> Descriptor {
let (platform, os) = match core {
// The bare mock expect-pass artifact is `e2e-report` → core "report" or "".
"" | "report" => ("Mock", "Linux"),
// The GitHub-hosted Windows lane: no GPU, native Windows.
"windows" => ("Mock", "Windows"),
"gpu" => ("MI300X", "Linux"),
"gpu-rad3" => ("R9700", "Linux"),
"gpu-mi350p" => ("MI350P", "Linux"),
Expand Down Expand Up @@ -1758,6 +1760,7 @@ mod tests {
fn parse_descriptor_maps_known_artifacts() {
for (name, platform, os) in [
("e2e-report", "Mock", "Linux"),
("e2e-windows-report", "Mock", "Windows"),
("e2e-gpu-report", "MI300X", "Linux"),
("e2e-gpu-rad3-report", "R9700", "Linux"),
("e2e-gpu-mi350p-report", "MI350P", "Linux"),
Expand Down
20 changes: 15 additions & 5 deletions docs/ci-hardware-testing.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,7 @@ separate tier flag or tag filter to maintain.
| Job | Workflow | Platform | Runner labels |
|---|---|---|---|
| `e2e` | `ci.yml` | Mock (no GPU) | GitHub-hosted `ubuntu-latest` |
| `e2e-windows` | `ci.yml` | Mock (no GPU) on Windows | GitHub-hosted `windows-latest` |
| `e2e-gpu` | `e2e-selfhosted.yml` | MI300X (AMD Instinct, bare-metal Linux) | self-hosted `[self-hosted, linux, mi300x]` |
| `e2e-gpu-strix-ubuntu` | `e2e-selfhosted.yml` | Strix Halo (gfx1151) on Ubuntu | self-hosted `[self-hosted, linux, devlab-dispatch, strix-halo]` |
| `e2e-gpu-strix-windows` | `e2e-selfhosted.yml` | Strix Halo (gfx1151) on Windows 11 | self-hosted `[self-hosted, windows, devlab-dispatch, strix-halo]` |
Expand Down Expand Up @@ -78,12 +79,21 @@ named for.

`e2e` is the blocking, GitHub-hosted mock job: `@requires-gpu` scenarios
resolve to skip here, and known bugs resolve to xfail from
`expectations.toml`. It is a required check and must stay green.
`expectations.toml`. It is a required check and must stay green. It also runs
every scenario that plants a simulated machine (see `docs/testing.md`).
`e2e-windows` runs the same suite on GitHub-hosted Windows, the platform where
the non-GPU scenarios meet the Windows probes and paths; it is advisory until
it has a record of being green.

The self-hosted jobs (`e2e-gpu`, `e2e-gpu-strix-ubuntu`, `e2e-gpu-strix-windows`,
`e2e-wsl`, `e2e-gpu-rad3`, and `e2e-gpu-mi350p`) run on AMD GPU systems, so they
exercise host/GPU detection, engine `detect`/`capabilities`, and live serving
scenarios that the mock job cannot. GPU availability is advisory in the WSL lane, as
`e2e-wsl`, `e2e-gpu-rad3`, and `e2e-gpu-mi350p`) run on AMD GPU systems. All
but `e2e-wsl` run ONLY the scenarios that need a real GPU (`E2E_GPU_ONLY=1`): live serving, SDK installs,
and detection on the real machine. Everything else runs on the two hosted jobs
above, so a GPU runner's time is not spent on it. On a pull request and in the
merge queue they narrow further to the `@gpu-smoke` canaries (plus, in the
queue, the `@merge-queue` serves); the full real-GPU set runs on pushes to
`main` and `release/**`, on dispatch, and nightly. `e2e-wsl` has no usable GPU
and runs the whole non-GPU suite inside a real WSL2 distribution. GPU availability is advisory in the WSL lane, as
described below.

`e2e-wsl` runs on the DevLab Dispatch pool. Its WSL2/Ubuntu-24.04 distro is
Expand Down Expand Up @@ -146,7 +156,7 @@ covers the mock platform;
reports — including partial or failed runs — by scenario id into one HTML report
and GitHub step summary.

The lane artifacts are named canonically (`e2e-report`, `e2e-gpu-report`,
The lane artifacts are named canonically (`e2e-report`, `e2e-windows-report`, `e2e-gpu-report`,
`e2e-gpu-rad3-report`, `e2e-gpu-mi350p-report`, `e2e-gpu-strix-ubuntu-report`,
`e2e-gpu-strix-windows-report`, `e2e-gpu-strix-wsl-report`) in `ci.yml` and
`e2e-selfhosted.yml`, because the report derives each platform's name and OS
Expand Down
12 changes: 11 additions & 1 deletion docs/testing.md
Original file line number Diff line number Diff line change
Expand Up @@ -182,12 +182,22 @@ The cucumber suite runs in one of two hardware modes, chosen by
usable AMD GPU.

Any other value aborts the run, so a misspelt `real` cannot silently skip the
real-GPU scenarios. To run them locally on a GPU host:
real-GPU scenarios.

To run the real-GPU scenarios locally on a GPU host:

```bash
E2E_HARDWARE=real cargo xtask e2e
```

The self-hosted GPU lanes run only the scenarios that need their GPU
(`E2E_GPU_ONLY=1`); everything else runs on the GitHub-hosted `E2E tests`
(Linux) and `E2E tests (Windows)` jobs, and inside a real WSL2 distribution on
the WSL lane. On a pull request and in the merge queue the GPU lanes narrow
further, to the cheap per-engine `@gpu-smoke` canaries (and, in the queue, the
`@merge-queue` serves) with `E2E_GPU_SMOKE_ONLY=1`; the full real-GPU set runs
nightly and on pushes to `main` and `release/**`.

A scenario describes its machine with a Given step — `a machine with an AMD
Instinct GPU`, `a machine with eight AMD Instinct GPUs`, `a bare-metal Linux
machine with an AMD GPU`, `a WSL machine with an AMD GPU passed through`, `a
Expand Down
Loading
Loading