Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
25 commits
Select commit Hold shift + click to select a range
eb694a6
feat(vllm): OOM detection, serve-summary memory guidance, and diagnos…
r0x0r Sep 3, 2026
7c09129
serve: address EAI-8059 re-review — validate --gpu before engine work…
r0x0r Sep 3, 2026
83ecfaa
fix(vllm): reconcile OOM detector with diagnose threshold and de-dup …
r0x0r Sep 8, 2026
4a108b9
serve: make the GPU fail-fast contract true and share the vLLM tail b…
r0x0r Sep 9, 2026
baeab31
Merge gpu-out-of-memory into eai-8059-oom-memory-knobs-note
r0x0r Sep 14, 2026
9eaa54d
Merge gpu-out-of-memory into eai-8059-oom-memory-knobs-note
r0x0r Sep 15, 2026
0be2dde
fix(serve): the serve summary's OOM note let log text break out of th…
r0x0r Sep 15, 2026
4d5c868
Merge gpu-out-of-memory into eai-8059-oom-memory-knobs-note
r0x0r Sep 15, 2026
40c3567
fix(serve): stop the pre-launch hint suppressing the whole OOM note (…
r0x0r Sep 15, 2026
e47d0d6
Merge gpu-out-of-memory into eai-8059-oom-memory-knobs-note
r0x0r Sep 15, 2026
89cd2f0
Merge gpu-out-of-memory into eai-8059-oom-memory-knobs-note
r0x0r Sep 17, 2026
cafc8b0
chore(rocm-core): give the new terminal module the license header CI …
r0x0r Sep 17, 2026
afef225
Merge gpu-out-of-memory into eai-8059-oom-memory-knobs-note
r0x0r Sep 17, 2026
a1619d8
Merge gpu-out-of-memory into eai-8059-oom-memory-knobs-note
r0x0r Sep 17, 2026
9106611
Merge gpu-out-of-memory into eai-8059-oom-memory-knobs-note
r0x0r Sep 17, 2026
6a7137b
Merge gpu-out-of-memory into eai-8059-oom-memory-knobs-note
r0x0r Sep 18, 2026
6c0d537
Merge remote-tracking branch 'origin/gpu-out-of-memory' into wip-284-m3
r0x0r Sep 25, 2026
5e87745
Merge remote-tracking branch 'origin/gpu-out-of-memory' into wip-284-m3
r0x0r Sep 29, 2026
653d5db
EAI-8059: adapt the OOM fault-injection expectation test to Included
r0x0r Sep 29, 2026
390f6cc
Merge remote-tracking branch 'origin/gpu-out-of-memory' into wip-284-m3
r0x0r Sep 29, 2026
90ce2a6
serve: keep runtime resolution behind the no-usable-GPU bail (eai-8059)
r0x0r Sep 30, 2026
d0fee30
test(e2e): cover the behaviour the model-keyed reuse pre-gate introduces
r0x0r Sep 30, 2026
37247be
test(vllm): narrow the OOM fixture-table failure message to admissibi…
r0x0r Sep 30, 2026
73135d5
Merge remote-tracking branch 'origin/gpu-out-of-memory' into wip-284-m3
r0x0r Oct 2, 2026
fe41981
Merge remote-tracking branch 'origin/gpu-out-of-memory' into wip-284-m3
r0x0r Oct 2, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions apps/rocm/Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,12 @@ workspace = true

[features]
e2e-test-hooks = ["rocm-engine-lemonade/e2e-test-hooks"]
# Test-only fault injection for the e2e suite. Enabled ONLY by the `xtask e2e`
# mock-lane build (never by the release build), it compiles in a `rocm serve`
# hook that simulates a managed vLLM launch which OOM'd and never became ready,
# so the positive OOM-guidance summary can be verified on a GPU-less CI host.
# See `e2e_simulate_oom_launch` / `simulate_oom_managed_launch` in `main.rs`.
e2e-oom-fault-injection = []

[dependencies]
anyhow.workspace = true
Expand Down
719 changes: 695 additions & 24 deletions apps/rocm/src/main.rs

Large diffs are not rendered by default.

311 changes: 311 additions & 0 deletions apps/rocm/src/serve_summary.rs
Original file line number Diff line number Diff line change
Expand Up @@ -182,6 +182,83 @@ pub(crate) fn render_summary(summary: &DeploymentSummary) -> String {
out
}

/// Whether a managed service's recorded status means it *failed to become
/// ready* — the only state the post-failure OOM note should fire on.
///
/// The status vocabulary from `status_for_readiness` is `ready` (serving),
/// `running` (endpoint up, model still loading — healthy, deliberately kept
/// distinct from `starting` so `rocmd` does not restart a slow-loading model),
/// and `starting` (endpoint never came up). Only `starting` is a failure; a
/// bare `!= "ready"` check would wrongly treat a healthy still-loading `running`
/// service as a failed launch.
pub(crate) fn serve_failed_to_become_ready(status: &str) -> bool {
status == "starting"
}

/// Builds the actionable memory-knob note for a serve that failed to become
/// ready with an out-of-memory signature in its engine log. Returns `None` when
/// the serve became ready or the log carries no OOM signature, so healthy
/// deployments and unrelated failures are never cluttered with memory advice.
///
/// The note names `--gpu-memory-utilization` and `--gpu` and is worded for the
/// shared-node case rather than as unconditional advice — vLLM reserves a
/// fraction of *total* VRAM, so a value good for a shared card would degrade a
/// dedicated one. When the model simply does not fit, it says so and points at a
/// smaller/quantized model instead of the knob (which would only trade an
/// earlier OOM for a later one), matching the `rocm diagnose` vLLM-OOM entry it
/// then routes the user to. It shares
/// [`rocm_core::VLLM_GPU_MEMORY_UTILIZATION_HINT`] verbatim with the pre-launch
/// low-VRAM note so both surfaces point at the same fix.
///
/// `hint_already_present` says that shared hint is already in the summary's
/// notes (the pre-launch low-VRAM warning added it), and de-duplicates *that
/// fragment only*: the rest of the note is new information the pre-launch guess
/// does not carry — that this attempt really did run out of GPU memory rather
/// than might, the "if the model doesn't fit, lowering the reservation won't
/// help" branch, and the `rocm diagnose --symptom` command with the user's own
/// failing line. Low VRAM leading to an OOM is the causal chain this note exists
/// for, so suppressing the whole note there would silence it on its most likely
/// trigger.
pub(crate) fn oom_memory_note(
status: &str,
log_tail: &str,
hint_already_present: bool,
) -> Option<String> {
if !serve_failed_to_become_ready(status) {
return None;
}
// The symptom is vLLM's own subprocess output and it lands inside a
// single-quoted shell word in a sentence that invites the user to paste the
// command, so it is only routed through verbatim when it can be rendered as
// one intact quoted argument; otherwise the canonical symptom stands in. See
// [`rocm_core::quotable_in_single_quotes`] for why this rejects rather than
// escapes. The engine's startup-failure hint guards the same text the same
// way.
//
// `None` is also the "this tail shows no OOM" answer, so it doubles as the
// OOM gate: asking `vllm_log_shows_oom` first would evaluate the same
// predicate over the same string twice.
let symptom = rocm_core::vllm_oom_diagnose_symptom(log_tail)?;
let symptom = if rocm_core::quotable_in_single_quotes(&symptom) {
symptom
} else {
rocm_core::VLLM_OOM_CANONICAL_SYMPTOM.to_owned()
};
// Printed as its own sentence, or omitted when the pre-launch warning
// already printed the identical text.
let hint = if hint_already_present {
String::new()
} else {
format!("{} ", rocm_core::VLLM_GPU_MEMORY_UTILIZATION_HINT)
};
Some(format!(
"the serve attempt ran out of GPU memory. {hint}If the model simply does not fit in this \
GPU's VRAM, lowering the reservation will not help — serve a smaller or quantized model \
instead (rocm-cli serves one model on a single GPU). To have the tool pick the \
case-appropriate fix, run `rocm diagnose --symptom '{symptom}'`."
))
}

/// Run a single small chat completion against the just-started local server and
/// measure time-to-first-token and generation throughput. Best-effort: any error
/// (server not OpenAI-compatible, refused, timed out) yields empty metrics rather
Expand Down Expand Up @@ -423,6 +500,240 @@ mod tests {
assert!(render_summary(&summary).contains("note: selected GPU 0 has low free VRAM"));
}

#[test]
fn oom_signatures_are_detected_case_insensitively() {
// The signatures the ticket calls out, plus casing variants the engine
// log can emit.
assert!(rocm_core::vllm_log_shows_oom(
"torch.OutOfMemoryError: HIP out of memory. Tried to allocate 7.21 GiB."
));
assert!(rocm_core::vllm_log_shows_oom("HIP OUT OF MEMORY"));
assert!(rocm_core::vllm_log_shows_oom(
"RuntimeError: CUDA out of memory"
));
}

#[test]
fn unrelated_failures_are_not_flagged_as_oom() {
assert!(!rocm_core::vllm_log_shows_oom(
"OSError: model weights not found; check the model id"
));
assert!(!rocm_core::vllm_log_shows_oom(""));
// vLLM's generic EngineCore wrapper is the terminal line for *any*
// startup crash, not just OOM; treating it as an OOM signature would
// misreport unrelated failures as memory exhaustion.
assert!(!rocm_core::vllm_log_shows_oom(
"ERROR Engine core initialization failed"
));
}

#[test]
fn oom_note_names_both_memory_knobs_on_a_failed_serve() {
let note = oom_memory_note(
"starting",
"torch.OutOfMemoryError: HIP out of memory. Tried to allocate 7.21 GiB.",
false,
)
.expect("an OOM failure must produce a note");
assert!(note.contains("--gpu-memory-utilization"), "{note}");
assert!(note.contains("--gpu <index>"), "{note}");
assert!(note.contains("ran out of GPU memory"), "{note}");
// The knob is not unconditional: when the model simply does not fit, the
// note must say lowering the reservation will not help and point at a
// smaller/quantized model, matching the `rocm diagnose` vLLM-OOM entry.
assert!(
note.contains("smaller or quantized model"),
"the note must carry the model-too-large caveat: {note}"
);
assert!(
note.contains("rocm diagnose --symptom"),
"the note must route the user to the conditional diagnose entry: {note}"
);
}

#[test]
fn oom_note_is_withheld_for_a_ready_serve_or_a_clean_log() {
// A serve that became ready is healthy even if the log mentions memory.
assert_eq!(
oom_memory_note("ready", "torch.OutOfMemoryError: HIP out of memory", false),
None
);
// A failure with no OOM signature must not be given memory advice.
assert_eq!(
oom_memory_note("starting", "OSError: model weights not found", false),
None
);
// ...and neither case becomes advisable just because the pre-launch
// low-VRAM warning already fired.
assert_eq!(
oom_memory_note("ready", "torch.OutOfMemoryError: HIP out of memory", true),
None
);
assert_eq!(
oom_memory_note("starting", "OSError: model weights not found", true),
None
);
}

#[test]
fn the_shared_hint_is_de_duplicated_without_losing_the_rest_of_the_oom_note() {
// The pre-launch low-VRAM warning already printed the shared hint
// verbatim, and low VRAM leading to an OOM is the causal chain this note
// exists for -- so what must be dropped is that one fragment, not the
// note. Everything else it carries is information the pre-launch guess
// does not have: that this attempt actually ran out of GPU memory, the
// "if the model doesn't fit, lowering the reservation won't help"
// branch, and the diagnose command with the user's real failing line.
let log_tail = "torch.OutOfMemoryError: HIP out of memory. Tried to allocate 7.21 GiB.";
let note = oom_memory_note("starting", log_tail, true)
.expect("an OOM failure must still carry a note when the hint was already printed");

// The hint the pre-launch warning printed appears zero further times...
assert!(
!note.contains(rocm_core::VLLM_GPU_MEMORY_UTILIZATION_HINT),
"the shared hint must not be printed a second time: {note}"
);
// ...and the content only this note has must all survive.
assert!(
note.contains("ran out of GPU memory"),
"the note must confirm this attempt really did OOM: {note}"
);
assert!(
note.contains("smaller or quantized model"),
"the model-too-large branch is absent from the pre-launch hint: {note}"
);
assert_eq!(
quoted_symptom_argument(&note),
"vllm: torch.OutOfMemoryError: HIP out of memory. Tried to allocate 7.21 GiB.",
"the diagnose command must carry the user's own failing line: {note}"
);

// Counted across the whole summary, the hint is printed exactly once:
// the pre-launch note keeps it, the OOM note does not repeat it.
let pre_launch = rocm_core::VLLM_GPU_MEMORY_UTILIZATION_HINT.to_owned();
let printed = [pre_launch, note]
.iter()
.filter(|line| line.contains(rocm_core::VLLM_GPU_MEMORY_UTILIZATION_HINT))
.count();
assert_eq!(printed, 1, "the shared hint must appear exactly once");
}

/// The `--symptom` value the note actually hands the user, read back out of
/// the rendered text exactly the way a shell would: the note's prose carries
/// its own apostrophes (`GPU's VRAM`), so the command is extracted from
/// between its backticks first, then the first `'...'` word inside it.
fn quoted_symptom_argument(note: &str) -> &str {
note.split("run `")
.nth(1)
.and_then(|rest| rest.split('`').next())
.expect("the note must print a runnable command")
.split("--symptom '")
.nth(1)
.and_then(|rest| rest.split('\'').next())
.expect("the command must carry a --symptom value")
}

#[test]
fn an_apostrophe_in_the_failing_line_cannot_break_out_of_the_printed_command() {
// A realistic vLLM traceback tail: apostrophes are routine in Python
// error text, and this line scores well above MIN_SCORE_FOR_MATCH
// (torch.OutOfMemoryError + HIP out of memory), so the "route the user's
// real line" branch selects it.
let log_tail = concat!(
" File \"/opt/vllm/worker.py\", line 212, in load_model\n",
"ERROR 09-14 12:00:01 engine.py:389] torch.OutOfMemoryError: HIP out of memory. ",
"Tried to allocate 7.21 GiB. GPU 0 can't allocate the model's weights.\n"
);
let note =
oom_memory_note("starting", log_tail, false).expect("an OOM failure must carry a note");

// The rendered command must be one intact single-quoted argument: no
// byte the vLLM subprocess printed may close the quote and land outside
// it in a command the note invites the user to paste.
let command = note
.split("run `")
.nth(1)
.and_then(|rest| rest.split('`').next())
.expect("the note must print a runnable command");
assert_eq!(
command.matches('\'').count(),
2,
"the --symptom argument must stay a single balanced quoted word: {command}"
);
// Pin which branch ran, not just that no apostrophe survived: "contains
// no `'`" also holds if the apostrophes were silently deleted from the
// user's line, which is the escaping-style behaviour the fallback exists
// to avoid. Rejection means the canonical symptom, exactly.
let symptom = quoted_symptom_argument(&note);
assert_eq!(
symptom,
rocm_core::VLLM_OOM_CANONICAL_SYMPTOM,
"a quote-bearing line must be rejected in favour of the canonical symptom, \
not silently rewritten: {symptom:?}"
);
// ...and the command it does print must still report a cause.
assert!(
rocm_core::vllm_oom_symptom_is_diagnosable(symptom),
"the fallback symptom must still be diagnosable: {symptom:?}"
);
}

#[test]
fn control_bytes_from_the_log_never_reach_the_printed_command() {
// vLLM's logger colourises; an ANSI-coloured OOM line must not repaint
// the user's terminal from inside rocm-cli's own serve summary.
let log_tail = "\u{1b}[31mRuntimeError: HIP out of memory\u{1b}[0m\u{7}";
let note =
oom_memory_note("starting", log_tail, false).expect("an OOM failure must carry a note");
assert!(
!note.chars().any(|c| c.is_control() && c != '\n'),
"no control byte may survive into the printed note: {note:?}"
);
// Pin the exact value rather than the absence of a few fragments: an
// absence check cannot fail for the defect it names, since a stripper
// that drops only the escape byte and the `[` leaves `31m`/`0m` behind,
// which contains neither `[31m` nor `[0m` and carries no control byte.
let symptom = quoted_symptom_argument(&note);
assert_eq!(
symptom,
rocm_core::VLLM_OOM_CANONICAL_SYMPTOM,
"a control-byte-bearing line must be rejected in favour of the canonical \
symptom, not stripped into a lookalike: {symptom:?}"
);
}

#[test]
fn oom_note_quotes_a_clean_failing_line_verbatim() {
// The guard must not cost the common case its own error text: a line
// with no quote and no control byte is still routed into the command.
let note = oom_memory_note(
"starting",
"torch.OutOfMemoryError: HIP out of memory. Tried to allocate 7.21 GiB.",
false,
)
.expect("an OOM failure must carry a note");
assert_eq!(
quoted_symptom_argument(&note),
"vllm: torch.OutOfMemoryError: HIP out of memory. Tried to allocate 7.21 GiB."
);
}

#[test]
fn oom_note_renders_in_the_summary_notes() {
let mut summary = base_summary();
summary.status = "starting".to_owned();
if let Some(note) = oom_memory_note(
&summary.status,
"torch.OutOfMemoryError: HIP out of memory. Tried to allocate 7.21 GiB.",
false,
) {
summary.notes.push(note);
}
let rendered = render_summary(&summary);
assert!(rendered.contains("note: the serve attempt ran out of GPU memory"));
assert!(rendered.contains("--gpu-memory-utilization"));
}

#[test]
fn api_key_client_config_shown_only_when_present() {
// Loopback / no key: the summary must not mention an api key at all.
Expand Down
14 changes: 14 additions & 0 deletions crates/e2e-report/src/lib.rs
Original file line number Diff line number Diff line change
Expand Up @@ -16,3 +16,17 @@ mod single_report;
pub use consolidated::{RunMeta, consolidated_summary_markdown, generate_consolidated};
pub use parse::{XfailReport, evaluate_xfail, scenario_results_by_id};
pub use single_report::generate;

/// Environment variable signalling the `e2e-oom-fault-injection` hook is present.
///
/// `cargo xtask e2e` sets it to `1` when it built the binary under test with the
/// `rocm/e2e-oom-fault-injection` feature, so the `e2e-cucumber` harness can tell
/// whether `@requires-oom-fault-injection` scenarios can run against it.
///
/// The single source of truth for this xtask ↔ harness contract. It lives in
/// this lean crate — the one both `xtask` and `e2e-cucumber` already depend on —
/// so the producer (`xtask::e2e`) and the consumer
/// (`e2e_cucumber::capability`) reference the same literal and cannot drift: a
/// typo previously would not fail to compile or fail a test, silently turning
/// `@requires-oom-fault-injection` into a skip (green suite, zero coverage).
pub const OOM_FAULT_INJECTION_ENV: &str = "ROCM_E2E_OOM_FAULT_INJECTION";
24 changes: 24 additions & 0 deletions crates/rocm-core/src/fix.rs
Original file line number Diff line number Diff line change
Expand Up @@ -1910,6 +1910,30 @@ mod tests {
);
}

#[test]
fn the_utilization_hint_states_the_bound_the_parser_accepts() {
// `rocm serve`'s `parse_gpu_memory_utilization` rejects `<= 0` and
// accepts `1`, so the domain is (0, 1]. `<0-1>` advertises `0`, and this
// hint prints in three places (pre-launch note, engine OOM hint, and the
// `fix-16-vllm-oom` summary), so a user who OOMs, reads the tool's own
// advice and passes `0` is rejected by the same tool.
//
// Pinned because the wording was silently reverted once: it was fixed on
// this branch, then a merge resolved the same line from a pre-fix tree.
// Nothing went red, because the only pin on this const checked the
// worked example (`0.5`) and not the range text.
const FLAG: &str = "--gpu-memory-utilization";
let hint = crate::VLLM_GPU_MEMORY_UTILIZATION_HINT;
assert!(
hint.contains(&format!("{FLAG} <fraction greater than 0 and at most 1>")),
"the hint must state the bound `{FLAG}` actually accepts:\n{hint}"
);
assert!(
!hint.contains("<0-1>"),
"`<0-1>` wrongly advertises 0, which `{FLAG}` rejects:\n{hint}"
);
}

#[test]
fn auto_applicable_recipes_have_a_runner() {
for r in RECIPES {
Expand Down
Loading
Loading