Repository navigation
fix(serve): wait on the engine's readiness budget, fail fast on a dead engine - #4
Closed
mikeroySoft wants to merge 1 commit into
Closed
mikeroySoft wants to merge 1 commit into
mikeroySoft wants to merge 1 commit into
Conversation
…d engine
`start_managed_service` and the restart path both waited a hardcoded 45 s
for a managed service to become ready, while vLLM's own startup budget is
5 minutes (`DEFAULT_VLLM_READY_TIMEOUT`, tunable via
`ROCM_CLI_VLLM_READY_TIMEOUT_SECS`). A cold vLLM start — torch import,
aiter/triton JIT, KV-cache warmup — routinely runs past a minute, so a
healthy launch was reported as "not ready" seconds before the server came
up, and the deployment summary read as a failure:
14:38:37 vllm boot
14:39:22 rocm serve gives up -> "Deployment summary (not ready yet)"
14:39:31 Application startup complete -> service actually ready
Take the budget from the engine instead, via a new public
`rocm_engine_vllm::ready_timeout()`, so the CLI cannot contradict the
engine it is waiting on.
A longer budget must not make a genuine crash slower to report, so the
progress callback can now end the wait early and `start_managed_service`
does so once the engine records a terminal status. Process liveness is not
usable here: the supervisor is our own child, so `process_is_running`
still reports it alive while it sits unreaped as a zombie. A launch the
engine gave up on is reported as `failed` with a note pointing at the
engine error, rather than "it may still be loading".
Verified on a Radeon AI PRO R9700 (gfx1201) with a TheRock nightly wheel
runtime:
- Qwen3-0.6B: `status ready`, TTFT 27 ms, 166.9 tok/s
- a model that crashes in engine-core init: reported `failed` after
7.9 s instead of spinning the full budget
- a model that crashes 81 s in: now waits past 45 s and reports the real
outcome instead of a misleading "not ready"
Signed-off-by: Michael Roy <1791194+mikeroySoft@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Base:
upstream-main— pinned toROCm/rocm-cli@fdfa620.Group:
rocm servelaunch reliability. Unrelated to Hyperloom-R, but it toucheswait_for_service_http_ready_with_progress, which the Hyperloom-R branch also adapts — see the conflict note below.Mirror of the PR opened upstream as ROCm#202.
Problem
rocm servereported a healthy vLLM launch as a failure. Both the launch and restart paths waited a hardcoded 45 s while vLLM's own budget is 5 minutes, and a cold start here takes ~54 s:Change
rocm_engine_vllm::ready_timeout().on_tickcan end the wait early, so a longer budget does not slow a real crash down. Process liveness is unusable as the signal — the supervisor is our own child andkill(pid,0)still reports it alive while it is an unreaped zombie, which cost a full 5-minute spin before I switched to the engine's state file.failedfromrunning/starting.Verification
On a Radeon AI PRO R9700 (gfx1201), TheRock nightly, vLLM 0.23.0:
ready, TTFT 27 ms, 166.9 tok/sfailed, note points at the engine errorFour tests added. Clippy, fmt,
scripts/smoke_local.pyclean.Conflict note
PR #1 (Hyperloom-R) adapts the same function's call in
bench_run.rs. Whichever of the two merges second needs a trivial rebase — the two changes are compatible, they just touch adjacent code.