Skip to content

fix(render,nodeswap): close six model-switch/deploy gaps found in the bigger-models study - #480

Merged
dmmdea merged 2 commits into
mainfrom
fix/model-switch-gaps
Sep 24, 2026
Merged

dmmdea merged 2 commits into
mainfrom
fix/model-switch-gaps

Conversation

@dmmdea

@dmmdea dmmdea commented Sep 24, 2026

Copy link
Copy Markdown
Owner

Summary

Closes six model-switch/deploy gaps surfaced by the bigger-models study and the d520701
fleet deploy (evidence: bigger-models-2026-09-24.md "Interim Phase 2 round 2" and
deploy-d5207011.md, both outside this repo).

  1. videogen_text_encoder (and siblings) were not family-scoped. Every videogen_*
    weight/behavior key was a single machine-wide value applied to whichever family actually
    rendered, so a request's model override could silently inherit a different family's
    files. Added videogen_families (opt-in, mirrors imagegen_families but with no
    licensing overlay — video family selection is not a licensing feature). Back-compat is
    load-bearing
    : a family with no explicit videogen_families[name] entry falls back to
    the box's flat videogen_* keys unchanged — this preserves a pre-existing, tested,
    deliberate pattern (Wan GGUF files intentionally left bound on videogen_unet_high/low
    on an ltx25-seated box, as the fallback for a bare model:"wan" override). An operator
    closes a specific leak by adding only the keys that need their own family-distinct value.
    doctor and offload_status (media.video_family_bindings) now report the resolved
    per-family bindings by name.
  2. Wan 2.2 native loader (fp8-scaled safetensors, no MultiGPU) as a supported variant
    alongside the GGUF/DisTorch2 path: videogen_wan_loader (""/auto | native |
    gguf-distorch). auto now picks the plain, un-wrapped UNETLoader for a .safetensors
    expert (measured ~27.6% faster than DisTorch2/MultiGPU) and keeps the historical wrapper
    for .gguf; --fast/--upscale work on both.
  3. ComfyUI boot fail-fast: ensureComfy used to only poll the HTTP health check, so a
    child that failed to start or crashed early looked identical to a slow cold boot for up
    to the full ~10min budget, GPU lease held the whole time. Now listens for the child's own
    error/early exit and fails within one poll interval, naming the cause and the tail of
    this run's own ComfyUI console log (PR Fix three media-lane defects: ffmpeg PATH resolution, ComfyUI console capture, doctor size checks #471).
  4. comfy-render.mjs self-manages its own GPU-slot lifecycle now (launch, reuse, GPU
    slot, teardown) — it used to require an already-running ComfyUI. A new --no-lifecycle
    flag keeps comfy-generate.mjs's existing wrapped-child invocation unchanged, so a
    --batch session's warm-ComfyUI optimization is never torn down between jobs.
  5. LTX-2.5 transformer per-request override (--transformer CLI flag / transformer
    MCP param): wins over the config-bound file for one render (int8 default, bf16 opt-in).
  6. node-swap gaps (from the d520701 deploy): (a) a Linux detached launcher
    (setup/linux-node-swap-launch.sh, the bash sibling of the PowerShell one — drives the
    same engine, never a hand-rolled health poll); --health-url is now also auto-resolved
    from a node's own config fleet_listen when left unset, never from a loopback/wildcard
    guess; (b) a standalone node (no --health-url) now waits for its own GPU lease to clear
    before swapping; (c) backupPathFor strips a redundant leading bak- from an
    operator-supplied --backup-suffix so it no longer doubles into bak-bak-....

New config keys

  • videogen_wan_loader (top-level, string): "" / "auto" | "native" | "gguf-distorch".
  • videogen_families (top-level, map[string]object), each entry: unet_high, unet_low,
    text_encoder, transformer, video_vae, audio_vae, latent_upscaler,
    wan_virtual_vram_gb, wan_loader, fps, width, height, frames, pool_vvram_gb,
    pool_compute, pool_donor, upscale_model, upscale_width, upscale_height.

Example — the Qube's Wan fp8-native lane, scoped away from its ltx25 default:

{
  "videogen_family": "ltx25",
  "videogen_families": {
    "wan22": {
      "unet_high": "wan2.2_i2v_high_noise_14B_fp8_scaled.safetensors",
      "unet_low": "wan2.2_i2v_low_noise_14B_fp8_scaled.safetensors",
      "text_encoder": "umt5_xxl_fp8_e4m3fn_scaled.safetensors",
      "wan_loader": "native"
    }
  }
}

Example — an edit-fp8 family binding, scoped away from the box's default edit family:

{
  "gen_edit_family": "qwen-image-edit-2511",
  "videogen_families": {
    "hunyuan": {
      "text_encoder": "qwen_2.5_vl_7b_fp8_scaled.safetensors"
    }
  }
}

Example — the 2512 family's videogen_wan_loader explicit override (forcing the historical
DisTorch2/MultiGPU path on a box too small for native streaming):

{
  "videogen_wan_loader": "gguf-distorch"
}

OptiPlex stall root cause

The 7+ minute stall with 0% GPU activity was ensureComfy (render/comfy-lifecycle.mjs)
spawning python main.py and then ONLY polling the HTTP health check — with no check that
spawn() produced a live child, and no listener on the child's own error/exit events. A
child that fails to start (a bad cwd/python path — the study's direct node comfy-video.mjs --graph ... invocation bypassed the Go pipeline's usual COMFY_DIR/COMFY_PY env, unlike
the working GGUF-baseline run through the normal Go CLI) or crashes on the way up is
indistinguishable from a slow cold boot until the full COMFY_START_WAIT_SEC budget (default
10 min) elapses — the GPU lease stays held the whole time with zero diagnostic signal. Fixed:
ensureComfy now fails within one poll interval (~2s) on a dead/never-started child, naming
the spawn command, the failure detail, and the tail of this run's own ComfyUI console log.

How tested

  • go build ./..., go vet ./... clean.
  • go test ./... — full suite, all ~90 packages, clean cache run, 0 failures.
  • node --test render/*.test.mjs — 454/454 pass.
  • gofmt clean on every touched Go file (tr -d '\r' < f | gofmt -l, since this repo commits
    CRLF).
  • Each new guard broken once (red) and restored, per clean-ship: the ComfyUI dead-child
    watchdog (confirmed pre-fix hangs to timeout with no signal), the standalone GPU-lease wait
    (confirmed pre-fix skips the wait unconditionally), comfy-render.mjs's self-managed
    lifecycle (confirmed pre-fix hangs indefinitely against an unreachable API with no lease
    check).
  • setup/linux-node-swap-launch.sh manually smoke-tested end-to-end against a stub binary in
    a sandboxed temp dir (detach, liveness check on both a healthy and an immediately-crashing
    child, log/result file writing) — not part of the automated suite (a shell script), but
    exercised live.
  • ONE pr-review-toolkit:code-reviewer pass (Sonnet) on the full diff; its findings (a
    cross-family footprint-quant leak in videoFootprintQuant, and a job-control PID-capture
    edge case in the Linux launcher) are fixed in this PR, with new regression tests for the
    first.

Risk

Medium (render + deploy paths, per the consequential-change gate) — mitigated by the full
test suite, the reviewer pass, and back-compat being load-bearing throughout (no default
behavior changes for a config that doesn't opt into the new keys). Not deployed and no node
config changed by this PR — a separate rollout pass does that next.

🤖 Generated with Claude Code

dmmdea and others added 2 commits September 24, 2026 15:03
… bigger-models study

Family-scope every videogen_* weight key (videogen_families, opt-in, back-
compat: no explicit entry falls back to the flat keys unchanged); add a
native (no-MultiGPU) Wan 2.2 fp8-scaled loader alongside the GGUF/DisTorch2
path (videogen_wan_loader); make ensureComfy fail fast on a dead/never-
started ComfyUI child instead of silently polling the full ~10min budget;
give comfy-render.mjs the same self-managed GPU-slot lifecycle the other
render runners use; add a per-request LTX-2.5 transformer override
(--transformer / MCP transformer); and close four node-swap gaps found in
the d520701 deploy (a Linux launcher script, a standalone node's GPU-lease
wait, auto-resolved --health-url from a node's own config, and the doubled
backup-suffix prefix).

Evidence: bigger-models-2026-09-24.md ("Interim Phase 2 round 2") and
deploy-d5207011.md.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
CI caught this on the previous commit: stopForSwap calls
Deps.FindProcessesByExe unconditionally, on every platform, before any
restart-mechanism check. The non-Windows platformDeps returned a hard error
("implemented for windows only"), so EVERY real node-swap run on Linux
failed at the very first step, even a standalone binary-only swap with
nothing to stop — undermining this PR's own new Linux launcher script.

A live Linux binary can be renamed/replaced out from under a running
process with no lock at all, so "no holders" is the correct, successful
answer on non-Windows, never an error. Fixed in deps_other.go; new
deps_other_test.go (!windows-tagged) pins it so CI's Linux run keeps
covering this directly.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@dmmdea
dmmdea merged commit c23b5c6 into main Sep 24, 2026
5 checks passed
@dmmdea
dmmdea deleted the fix/model-switch-gaps branch September 24, 2026 20:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant