Skip to content

fix(pipeline): pin comfy_cuda_device for un-pooled video, matching the image route - #478

Merged
dmmdea merged 1 commit into
mainfrom
fix/video-device-pin-unpooled
Sep 24, 2026
Merged

dmmdea merged 1 commit into
mainfrom
fix/video-device-pin-unpooled

Conversation

@dmmdea

@dmmdea dmmdea commented Sep 24, 2026

Copy link
Copy Markdown
Owner

Summary

Found while running the ComfyUI-MultiGPU archival A/B (see PR #477 and infra/multigpu-archival-actions-2026-09-24.md on the operator's Drive): runGenerateVideo hard-coded comfyGenEnv(false) on the theory that "every seeded video seat is pooled and owns its placement through its pool keys." That is false once a box runs an un-pooled/native video seat (videogen_pool_vvram_gb<=0) -- exactly the shape the archival A/B is measuring as a faster, display-card-safe alternative to pooled DisTorch2.

  • Measured on the blackwell-3x16 reference box (2026-09-24): with videogen_pool_vvram_gb=0, comfy_dynamic_vram=on, comfy_cuda_device="2", a native LTX-2.5 1280x704 render landed on nvidia-smi index 1 -- the DISPLAY card -- at 15.7 GiB, while card 2 (the intended target) stayed idle. Confirmed via torch.cuda's own enumeration on the box: ComfyUI's fastest-first cuda:0 IS this box's RTX 5070 Ti display card. docs/tiers/blackwell-3x16.md already documents this ordering for the POOLED compute/donor keys ("PIN BY GPU UUID, NEVER INDEX... TWO DEVICE ORDERINGS ARE IN PLAY"); the un-pooled path had no equivalent protection.
  • This is exactly the class of incident already recorded in that same tier doc for 2026-09-04 (a starved display card forced a reboot, no crash / no TDR in the event log).

What changed

  • internal/config/families.go: added Config.VideoPooled(), mirroring the existing ImagePooled().
  • internal/pipeline/pipeline.go: runGenerateVideo now passes !p.cfg.VideoPooled() as the comfyGenEnv singleCard argument instead of the hardcoded false. A pooled video seat is unaffected (still no pin -- the blackwell-3x16 pool's compute must stay on ComfyUI's default device per MultiGPU changelog 0.112.0: two-card store number #220, a documented law exception); an un-pooled one now takes the same pin every other single-card route already gets. Doc comments on comfyLaunch updated to match.
  • Tests: TestLaunchProfileReachesEveryComfyRouteAndThePinOnlyTheSingleCardOnes (generate_video: false -> true; added a generate_video (pooled) case asserting the pin stays OFF there) and the two gpuLeaseCases()-driven completeness tests.

How tested

  • go build ./..., go vet ./... -- clean.
  • go test ./internal/pipeline/... ./internal/config/... -- all green, including the new pooled-video case.
  • Live re-measurement after deploying the fixed binary to the reference box: native LTX-2.5 now measures peak VRAM card0=604 MiB, card1 (display)=2,487 MiB (baseline, untouched), card2=15,462 MiB at 1280x704, and card0=0, card1=2,713 MiB, card2=15,430 MiB at the full 1920x1088 (fits, no OOM) -- both correctly isolated to card 2.

Risk

Low-medium: changes production GPU device placement for one route (un-pooled video generation). Scoped by a passing test asserting the pooled path is byte-for-byte unaffected. Live-verified on the actual 3-card box the risk applies to, not just unit tests.

🤖 Generated with Claude Code

…e image route

generate_video hard-coded comfyGenEnv(false), on the theory that "every seeded
video seat is pooled and owns its placement through its pool keys." That is
false in general: videogen_pool_vvram_gb<=0 (the native/streaming shape the
ComfyUI-MultiGPU archival A/B is measuring) is un-pooled, and an un-pooled
UNETLoader carries no device input at all -- it lands on ComfyUI's DEFAULT
device with nothing to place it, same as an un-pooled image render would
without the image route's existing !cfg.ImagePooled() pin logic.

Measured on the blackwell-3x16 reference box (2026-09-24): with
videogen_pool_vvram_gb=0, comfy_dynamic_vram=on, comfy_cuda_device="2", a
native LTX-2.5 1280x704 render landed on nvidia-smi index 1 -- the DISPLAY
card -- at 15.7 GiB, while card 2 stayed at idle baseline. Confirmed via
torch.cuda's own enumeration: ComfyUI's fastest-first cuda:0 IS this box's
RTX 5070 Ti display card (docs/tiers/blackwell-3x16.md already documents this
ordering for the POOLED compute/donor keys; the un-pooled path had no
equivalent protection at all). This is exactly the class of incident recorded
in that same tier doc for 2026-09-04 (a starved display card forced a reboot).

Fix: add Config.VideoPooled() (mirrors ImagePooled()), and pass
!cfg.VideoPooled() as runGenerateVideo's singleCard argument instead of the
hardcoded false. A pooled video seat is unaffected (still no pin -- the
blackwell-3x16 pool's compute must stay on ComfyUI's default device per
MultiGPU #220, a documented law exception); an un-pooled one now takes the
same comfy_cuda_device pin every other single-card route already gets.

Tests: TestLaunchProfileReachesEveryComfyRouteAndThePinOnlyTheSingleCardOnes
updated (generate_video: false -> true; added a "generate_video (pooled)"
case asserting the pin STAYS off there) and
TestEveryGPUTaskRunsUnderALease/TestEveryGpuSlotRunnerIsCoveredByTheLeaseTable
cover the new case. go build/vet/test (internal/pipeline, internal/config)
all green.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
dmmdea added a commit that referenced this pull request Sep 24, 2026
… pooled

A/B measured on the reference box (2026-09-24, same session as the
ComfyUI-MultiGPU archival pin): native single-card streaming
(videogen_pool_vvram_gb/imagegen_pool_vvram_gb unset, comfy_dynamic_vram=on,
comfy_cuda_device=2) beats the pooled DisTorch2 recipe this tier seeded since
0.59.0/0.113.x on every measured axis:

- krea2 image: 77.7s native vs 128.3s pooled (40% faster); output
  pixel-identical (same model/precision, deterministic sampler).
- LTX-2.5 video 1280x704: 107.0s native vs 178.7s pooled (40% faster);
  frames/audio match (mean -43.0..-43.7 dB, max -27.3..-30.8 dB, both
  runs -- no silence, no clipping).
- LTX-2.5 video 1920x1088 (the flagship resolution POOLING CANNOT REACH on
  this tier -- documented OOM at every virtual_vram_gb, per the prior note):
  native fits at 107.3s, peak 15.4/15.9 GiB on the compute card, zero OOM.
- Display-card safety (nvidia-smi index 1, this box's RTX 5070 Ti): pooled
  krea2 loaded 12.1 GiB there via the CLIPLoader "default" device (untargeted
  by any pool key -- a live, previously-undocumented risk on the same class
  of incident recorded in this tier's own notes for 2026-09-04); pooled
  LTX-2.5 computes there BY DESIGN (MultiGPU #220, unaffected/unchanged).
  Native isolates fully to the compute card in both cases (confirmed via
  torch.cuda's own device enumeration plus nvidia-smi 1 Hz logging): display
  card never exceeds its ~1.5-2.7 GiB desktop baseline.

Required a prerequisite fix (PR #478, merged first) -- generate_video had no
device-pin path for the un-pooled shape at all, so the very first native
measurement landed on the display card until that shipped.

Removed imagegen_pool_vvram_gb/compute/donor and videogen_pool_vvram_gb/
compute/donor from this tier's config_seed (absent = the builders' own
un-pooled UNETLoader branch); added comfy_dynamic_vram="on" (comfy_cuda_device
stays "2", already seeded). docs/tiers/blackwell-3x16.md regenerated
(go run ./cmd/gentiers). videogen_width/height stay 1280x704 as the seeded
default (conservative: 1920x1088's ~15.4/15.9 GiB peak is thin margin for a
one-size-fits-all seed) -- 1920x1088 is documented as available and measured
in docs/systems/media-generation.md.

The live reference box (Qube) already carries this exact config change,
applied and smoke-tested in the same session (see
infra/multigpu-archival-actions-2026-09-24.md on the operator's Drive).

go test ./internal/tierseed/... ./internal/config/... ./internal/mediacap/...
and setup/render.tests.ps1 (ALL PASS) both green.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@dmmdea
dmmdea merged commit 81113b6 into main Sep 24, 2026
5 checks passed
@dmmdea
dmmdea deleted the fix/video-device-pin-unpooled branch September 24, 2026 13:24
dmmdea added a commit that referenced this pull request Sep 24, 2026
dmmdea added a commit that referenced this pull request Sep 24, 2026
…6-09-30 archival (#477)

* chore(comfy): pin ComfyUI-MultiGPU to v2.6.4 + local fix ahead of 2026-09-30 archival

Upstream pollockjj/ComfyUI-MultiGPU archives 2026-09-30 (issue #223, no successor
endorsed), and neither native Dynamic VRAM nor SelectModelDevice/"MultiGPU Work
Units" reproduce its donor+compute virtual_vram_gb weight-sharding, which the Wan
GGUF lane and the pooled krea2/LTX-2.5 seats depend on. Pin rather than migrate
blind (research: comfyui-multigpu-archival-2026-09-24.md).

Pinned commit ed1ffaef7cec1a66f35106c6a4c7a40927c2dc83 = upstream v2.6.4's last
code commit (b51c99a5) plus one already-deployed local fix for upstream issue
#220's libcudart.so-on-Windows crash. Created a private fallback mirror
(dmmdea/ComfyUI-MultiGPU-mirror, full history + this commit) for when upstream
goes read-only, and named the pin + mirror in both node-pack tables that already
existed for this exact purpose (render/comfy-nodes.mjs NODE_PACKS,
internal/mediacap/routeneeds.go nodePacks/packHint) so a box missing the pack is
pointed at the frozen, patched source instead of a soon-archived URL.

Docs: media-generation.md "ComfyUI-MultiGPU archival" section records the pin,
the mirror, and the compute_device constraint the local fix satisfies. Fleet
alignment (Qube/OptiPlex/Aorus/Lenovo, all now on this commit) and the LTX-2.5/
krea2 native-vs-pooled A/B are tracked outside this repo, on the operator's Drive
(infra/multigpu-archival-actions-2026-09-24.md).

go vet ./..., go test ./... (root module) and node --test render/*.test.mjs all
green.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(tiers): blackwell-3x16 seeds native LTX-2.5/krea2 streaming, not pooled

A/B measured on the reference box (2026-09-24, same session as the
ComfyUI-MultiGPU archival pin): native single-card streaming
(videogen_pool_vvram_gb/imagegen_pool_vvram_gb unset, comfy_dynamic_vram=on,
comfy_cuda_device=2) beats the pooled DisTorch2 recipe this tier seeded since
0.59.0/0.113.x on every measured axis:

- krea2 image: 77.7s native vs 128.3s pooled (40% faster); output
  pixel-identical (same model/precision, deterministic sampler).
- LTX-2.5 video 1280x704: 107.0s native vs 178.7s pooled (40% faster);
  frames/audio match (mean -43.0..-43.7 dB, max -27.3..-30.8 dB, both
  runs -- no silence, no clipping).
- LTX-2.5 video 1920x1088 (the flagship resolution POOLING CANNOT REACH on
  this tier -- documented OOM at every virtual_vram_gb, per the prior note):
  native fits at 107.3s, peak 15.4/15.9 GiB on the compute card, zero OOM.
- Display-card safety (nvidia-smi index 1, this box's RTX 5070 Ti): pooled
  krea2 loaded 12.1 GiB there via the CLIPLoader "default" device (untargeted
  by any pool key -- a live, previously-undocumented risk on the same class
  of incident recorded in this tier's own notes for 2026-09-04); pooled
  LTX-2.5 computes there BY DESIGN (MultiGPU #220, unaffected/unchanged).
  Native isolates fully to the compute card in both cases (confirmed via
  torch.cuda's own device enumeration plus nvidia-smi 1 Hz logging): display
  card never exceeds its ~1.5-2.7 GiB desktop baseline.

Required a prerequisite fix (PR #478, merged first) -- generate_video had no
device-pin path for the un-pooled shape at all, so the very first native
measurement landed on the display card until that shipped.

Removed imagegen_pool_vvram_gb/compute/donor and videogen_pool_vvram_gb/
compute/donor from this tier's config_seed (absent = the builders' own
un-pooled UNETLoader branch); added comfy_dynamic_vram="on" (comfy_cuda_device
stays "2", already seeded). docs/tiers/blackwell-3x16.md regenerated
(go run ./cmd/gentiers). videogen_width/height stay 1280x704 as the seeded
default (conservative: 1920x1088's ~15.4/15.9 GiB peak is thin margin for a
one-size-fits-all seed) -- 1920x1088 is documented as available and measured
in docs/systems/media-generation.md.

The live reference box (Qube) already carries this exact config change,
applied and smoke-tested in the same session (see
infra/multigpu-archival-actions-2026-09-24.md on the operator's Drive).

go test ./internal/tierseed/... ./internal/config/... ./internal/mediacap/...
and setup/render.tests.ps1 (ALL PASS) both green.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant