Symptom
On Windows, rocm serve exits 0 when the managed engine process dies during startup. Nothing is printed to stderr, and the command reports a launched service. On Linux and macOS the same failure is caught and reported with the child's log tailed into the error.
Root cause
spawn_managed_engine_child (apps/rocm/src/main.rs:6646) branches on platform after building the service record:
- non-Windows —
command.spawn(), then thread::sleep(200ms), then child.try_wait(). If the child has already exited it bail!s with managed_engine_startup_failure_detail(status, &record.log_path).
- Windows —
rocm_core::spawn_detached_no_inherit(..), which returns a bare u32 PID. There is no Child handle, and so no liveness or exit check of any kind. Control falls straight through to record.status = "running" and Ok(ManagedSpawn::Spawned { .. }).
The caller does not rescue it. start_managed_service waits up to 45s for HTTP readiness, then sets record.status = status_for_readiness(readiness) and returns Ok(ManagedLaunchReport { .. }) regardless of the readiness outcome — a timeout is recorded as a status, not raised as an error. So a dead child produces a successful command.
This has been the case since the initial import; it is not a recent regression.
There is no diagnostic either
On non-Windows the child's stdio is attached to the service log via attach_background_stdio(&mut command, Some(&record.log_path)). The Windows path reaches spawn_windows_no_inherit(..) with std_handles: None (crates/rocm-core/src/lib.rs:1167), so the detached child gets no stdout/stderr handles at all. A user therefore sees neither a non-zero exit nor a captured startup error.
A naive fix has a trap
The obvious repair — poll the PID for liveness — carries a PID-reuse race, and the guard that exists for exactly that is inert on Windows. record.supervisor_start_ticks comes from rocm_core::process_start_ticks, which reads /proc/{pid}/stat and returns None on every non-Linux target (crates/rocm-core/src/proc_lifecycle.rs:307 and :314). A correct fix needs a real process handle, or a Win32 process start-time identity, rather than a bare PID.
How it surfaced
PR #351 reworks the Lemonade recovery scenario to inject its failure inside the engine's Serve path — i.e. in the detached child. The seam it replaces injected the failure in the parent's Install RPC, which failed synchronously on every platform and so never exercised this gap. serve-18 consequently reports serve unexpectedly succeeded (rc 0, empty stderr) on the Windows lane while passing on the Linux GPU lanes. The scenario is correct; it is asserting behaviour the platform does not implement.
Observed vs inferred
Every file, line and branch cited above is read from the tree at b0d598d0. That the scenario's injected failure occurs inside the detached child and is invisible to the parent is inferred from those two branches plus the rc 0 / empty-parent-stderr evidence in the job log — I have no Windows host, and the job log does not include the child's service log.
Symptom
On Windows,
rocm serveexits 0 when the managed engine process dies during startup. Nothing is printed to stderr, and the command reports a launched service. On Linux and macOS the same failure is caught and reported with the child's log tailed into the error.Root cause
spawn_managed_engine_child(apps/rocm/src/main.rs:6646) branches on platform after building the service record:command.spawn(), thenthread::sleep(200ms), thenchild.try_wait(). If the child has already exited itbail!s withmanaged_engine_startup_failure_detail(status, &record.log_path).rocm_core::spawn_detached_no_inherit(..), which returns a bareu32PID. There is noChildhandle, and so no liveness or exit check of any kind. Control falls straight through torecord.status = "running"andOk(ManagedSpawn::Spawned { .. }).The caller does not rescue it.
start_managed_servicewaits up to 45s for HTTP readiness, then setsrecord.status = status_for_readiness(readiness)and returnsOk(ManagedLaunchReport { .. })regardless of the readiness outcome — a timeout is recorded as a status, not raised as an error. So a dead child produces a successful command.This has been the case since the initial import; it is not a recent regression.
There is no diagnostic either
On non-Windows the child's stdio is attached to the service log via
attach_background_stdio(&mut command, Some(&record.log_path)). The Windows path reachesspawn_windows_no_inherit(..)withstd_handles: None(crates/rocm-core/src/lib.rs:1167), so the detached child gets no stdout/stderr handles at all. A user therefore sees neither a non-zero exit nor a captured startup error.A naive fix has a trap
The obvious repair — poll the PID for liveness — carries a PID-reuse race, and the guard that exists for exactly that is inert on Windows.
record.supervisor_start_tickscomes fromrocm_core::process_start_ticks, which reads/proc/{pid}/statand returnsNoneon every non-Linux target (crates/rocm-core/src/proc_lifecycle.rs:307and:314). A correct fix needs a real process handle, or a Win32 process start-time identity, rather than a bare PID.How it surfaced
PR #351 reworks the Lemonade recovery scenario to inject its failure inside the engine's Serve path — i.e. in the detached child. The seam it replaces injected the failure in the parent's Install RPC, which failed synchronously on every platform and so never exercised this gap.
serve-18consequently reportsserve unexpectedly succeeded(rc 0, empty stderr) on the Windows lane while passing on the Linux GPU lanes. The scenario is correct; it is asserting behaviour the platform does not implement.Observed vs inferred
Every file, line and branch cited above is read from the tree at
b0d598d0. That the scenario's injected failure occurs inside the detached child and is invisible to the parent is inferred from those two branches plus the rc 0 / empty-parent-stderr evidence in the job log — I have no Windows host, and the job log does not include the child's service log.