fix(render,nodeswap): close six model-switch/deploy gaps found in the bigger-models study - #480
Merged
Merged
Conversation
… bigger-models study Family-scope every videogen_* weight key (videogen_families, opt-in, back- compat: no explicit entry falls back to the flat keys unchanged); add a native (no-MultiGPU) Wan 2.2 fp8-scaled loader alongside the GGUF/DisTorch2 path (videogen_wan_loader); make ensureComfy fail fast on a dead/never- started ComfyUI child instead of silently polling the full ~10min budget; give comfy-render.mjs the same self-managed GPU-slot lifecycle the other render runners use; add a per-request LTX-2.5 transformer override (--transformer / MCP transformer); and close four node-swap gaps found in the d520701 deploy (a Linux launcher script, a standalone node's GPU-lease wait, auto-resolved --health-url from a node's own config, and the doubled backup-suffix prefix). Evidence: bigger-models-2026-09-24.md ("Interim Phase 2 round 2") and deploy-d5207011.md. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
CI caught this on the previous commit: stopForSwap calls
Deps.FindProcessesByExe unconditionally, on every platform, before any
restart-mechanism check. The non-Windows platformDeps returned a hard error
("implemented for windows only"), so EVERY real node-swap run on Linux
failed at the very first step, even a standalone binary-only swap with
nothing to stop — undermining this PR's own new Linux launcher script.
A live Linux binary can be renamed/replaced out from under a running
process with no lock at all, so "no holders" is the correct, successful
answer on non-Windows, never an error. Fixed in deps_other.go; new
deps_other_test.go (!windows-tagged) pins it so CI's Linux run keeps
covering this directly.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Closes six model-switch/deploy gaps surfaced by the bigger-models study and the d520701
fleet deploy (evidence:
bigger-models-2026-09-24.md"Interim Phase 2 round 2" anddeploy-d5207011.md, both outside this repo).videogen_text_encoder(and siblings) were not family-scoped. Everyvideogen_*weight/behavior key was a single machine-wide value applied to whichever family actually
rendered, so a request's
modeloverride could silently inherit a different family'sfiles. Added
videogen_families(opt-in, mirrorsimagegen_familiesbut with nolicensing overlay — video family selection is not a licensing feature). Back-compat is
load-bearing: a family with no explicit
videogen_families[name]entry falls back tothe box's flat
videogen_*keys unchanged — this preserves a pre-existing, tested,deliberate pattern (Wan GGUF files intentionally left bound on
videogen_unet_high/lowon an
ltx25-seated box, as the fallback for a baremodel:"wan"override). An operatorcloses a specific leak by adding only the keys that need their own family-distinct value.
doctorandoffload_status(media.video_family_bindings) now report the resolvedper-family bindings by name.
alongside the GGUF/DisTorch2 path:
videogen_wan_loader(""/auto|native|gguf-distorch).autonow picks the plain, un-wrappedUNETLoaderfor a.safetensorsexpert (measured ~27.6% faster than DisTorch2/MultiGPU) and keeps the historical wrapper
for
.gguf;--fast/--upscalework on both.ensureComfyused to only poll the HTTP health check, so achild that failed to start or crashed early looked identical to a slow cold boot for up
to the full ~10min budget, GPU lease held the whole time. Now listens for the child's own
error/earlyexitand fails within one poll interval, naming the cause and the tail ofthis run's own ComfyUI console log (PR Fix three media-lane defects: ffmpeg PATH resolution, ComfyUI console capture, doctor size checks #471).
comfy-render.mjsself-manages its own GPU-slot lifecycle now (launch, reuse, GPUslot, teardown) — it used to require an already-running ComfyUI. A new
--no-lifecycleflag keeps
comfy-generate.mjs's existing wrapped-child invocation unchanged, so a--batchsession's warm-ComfyUI optimization is never torn down between jobs.--transformerCLI flag /transformerMCP param): wins over the config-bound file for one render (int8 default, bf16 opt-in).
(
setup/linux-node-swap-launch.sh, the bash sibling of the PowerShell one — drives thesame engine, never a hand-rolled health poll);
--health-urlis now also auto-resolvedfrom a node's own config
fleet_listenwhen left unset, never from a loopback/wildcardguess; (b) a standalone node (no
--health-url) now waits for its own GPU lease to clearbefore swapping; (c)
backupPathForstrips a redundant leadingbak-from anoperator-supplied
--backup-suffixso it no longer doubles intobak-bak-....New config keys
videogen_wan_loader(top-level, string):""/"auto"|"native"|"gguf-distorch".videogen_families(top-level,map[string]object), each entry:unet_high,unet_low,text_encoder,transformer,video_vae,audio_vae,latent_upscaler,wan_virtual_vram_gb,wan_loader,fps,width,height,frames,pool_vvram_gb,pool_compute,pool_donor,upscale_model,upscale_width,upscale_height.Example — the Qube's Wan fp8-native lane, scoped away from its ltx25 default:
{ "videogen_family": "ltx25", "videogen_families": { "wan22": { "unet_high": "wan2.2_i2v_high_noise_14B_fp8_scaled.safetensors", "unet_low": "wan2.2_i2v_low_noise_14B_fp8_scaled.safetensors", "text_encoder": "umt5_xxl_fp8_e4m3fn_scaled.safetensors", "wan_loader": "native" } } }Example — an edit-fp8 family binding, scoped away from the box's default edit family:
{ "gen_edit_family": "qwen-image-edit-2511", "videogen_families": { "hunyuan": { "text_encoder": "qwen_2.5_vl_7b_fp8_scaled.safetensors" } } }Example — the 2512 family's
videogen_wan_loaderexplicit override (forcing the historicalDisTorch2/MultiGPU path on a box too small for native streaming):
{ "videogen_wan_loader": "gguf-distorch" }OptiPlex stall root cause
The 7+ minute stall with 0% GPU activity was
ensureComfy(render/comfy-lifecycle.mjs)spawning
python main.pyand then ONLY polling the HTTP health check — with no check thatspawn()produced a live child, and no listener on the child's ownerror/exitevents. Achild that fails to start (a bad cwd/python path — the study's direct
node comfy-video.mjs --graph ...invocation bypassed the Go pipeline's usualCOMFY_DIR/COMFY_PYenv, unlikethe working GGUF-baseline run through the normal Go CLI) or crashes on the way up is
indistinguishable from a slow cold boot until the full
COMFY_START_WAIT_SECbudget (default10 min) elapses — the GPU lease stays held the whole time with zero diagnostic signal. Fixed:
ensureComfynow fails within one poll interval (~2s) on a dead/never-started child, namingthe spawn command, the failure detail, and the tail of this run's own ComfyUI console log.
How tested
go build ./...,go vet ./...clean.go test ./...— full suite, all ~90 packages, clean cache run, 0 failures.node --test render/*.test.mjs— 454/454 pass.gofmtclean on every touched Go file (tr -d '\r' < f | gofmt -l, since this repo commitsCRLF).
watchdog (confirmed pre-fix hangs to timeout with no signal), the standalone GPU-lease wait
(confirmed pre-fix skips the wait unconditionally),
comfy-render.mjs's self-managedlifecycle (confirmed pre-fix hangs indefinitely against an unreachable API with no lease
check).
setup/linux-node-swap-launch.shmanually smoke-tested end-to-end against a stub binary ina sandboxed temp dir (detach, liveness check on both a healthy and an immediately-crashing
child, log/result file writing) — not part of the automated suite (a shell script), but
exercised live.
pr-review-toolkit:code-reviewerpass (Sonnet) on the full diff; its findings (across-family footprint-quant leak in
videoFootprintQuant, and a job-control PID-captureedge case in the Linux launcher) are fixed in this PR, with new regression tests for the
first.
Risk
Medium (render + deploy paths, per the consequential-change gate) — mitigated by the full
test suite, the reviewer pass, and back-compat being load-bearing throughout (no default
behavior changes for a config that doesn't opt into the new keys). Not deployed and no node
config changed by this PR — a separate rollout pass does that next.
🤖 Generated with Claude Code