Skip to content

serve: on Windows a managed engine that dies at startup is reported as success #467

Description

@rominf

Symptom

On Windows, rocm serve exits 0 when the managed engine process dies during startup. Nothing is printed to stderr, and the command reports a launched service. On Linux and macOS the same failure is caught and reported with the child's log tailed into the error.

Root cause

spawn_managed_engine_child (apps/rocm/src/main.rs:6646) branches on platform after building the service record:

  • non-Windows — command.spawn(), then thread::sleep(200ms), then child.try_wait(). If the child has already exited it bail!s with managed_engine_startup_failure_detail(status, &record.log_path).
  • Windows — rocm_core::spawn_detached_no_inherit(..), which returns a bare u32 PID. There is no Child handle, and so no liveness or exit check of any kind. Control falls straight through to record.status = "running" and Ok(ManagedSpawn::Spawned { .. }).

The caller does not rescue it. start_managed_service waits up to 45s for HTTP readiness, then sets record.status = status_for_readiness(readiness) and returns Ok(ManagedLaunchReport { .. }) regardless of the readiness outcome — a timeout is recorded as a status, not raised as an error. So a dead child produces a successful command.

This has been the case since the initial import; it is not a recent regression.

There is no diagnostic either

On non-Windows the child's stdio is attached to the service log via attach_background_stdio(&mut command, Some(&record.log_path)). The Windows path reaches spawn_windows_no_inherit(..) with std_handles: None (crates/rocm-core/src/lib.rs:1167), so the detached child gets no stdout/stderr handles at all. A user therefore sees neither a non-zero exit nor a captured startup error.

A naive fix has a trap

The obvious repair — poll the PID for liveness — carries a PID-reuse race, and the guard that exists for exactly that is inert on Windows. record.supervisor_start_ticks comes from rocm_core::process_start_ticks, which reads /proc/{pid}/stat and returns None on every non-Linux target (crates/rocm-core/src/proc_lifecycle.rs:307 and :314). A correct fix needs a real process handle, or a Win32 process start-time identity, rather than a bare PID.

How it surfaced

PR #351 reworks the Lemonade recovery scenario to inject its failure inside the engine's Serve path — i.e. in the detached child. The seam it replaces injected the failure in the parent's Install RPC, which failed synchronously on every platform and so never exercised this gap. serve-18 consequently reports serve unexpectedly succeeded (rc 0, empty stderr) on the Windows lane while passing on the Linux GPU lanes. The scenario is correct; it is asserting behaviour the platform does not implement.

Observed vs inferred

Every file, line and branch cited above is read from the tree at b0d598d0. That the scenario's injected failure occurs inside the detached child and is invisible to the parent is inferred from those two branches plus the rc 0 / empty-parent-stderr evidence in the job log — I have no Windows host, and the job log does not include the child's service log.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions