fix(pipeline): pin comfy_cuda_device for un-pooled video, matching the image route - #478
Merged
Merged
Conversation
…e image route generate_video hard-coded comfyGenEnv(false), on the theory that "every seeded video seat is pooled and owns its placement through its pool keys." That is false in general: videogen_pool_vvram_gb<=0 (the native/streaming shape the ComfyUI-MultiGPU archival A/B is measuring) is un-pooled, and an un-pooled UNETLoader carries no device input at all -- it lands on ComfyUI's DEFAULT device with nothing to place it, same as an un-pooled image render would without the image route's existing !cfg.ImagePooled() pin logic. Measured on the blackwell-3x16 reference box (2026-09-24): with videogen_pool_vvram_gb=0, comfy_dynamic_vram=on, comfy_cuda_device="2", a native LTX-2.5 1280x704 render landed on nvidia-smi index 1 -- the DISPLAY card -- at 15.7 GiB, while card 2 stayed at idle baseline. Confirmed via torch.cuda's own enumeration: ComfyUI's fastest-first cuda:0 IS this box's RTX 5070 Ti display card (docs/tiers/blackwell-3x16.md already documents this ordering for the POOLED compute/donor keys; the un-pooled path had no equivalent protection at all). This is exactly the class of incident recorded in that same tier doc for 2026-09-04 (a starved display card forced a reboot). Fix: add Config.VideoPooled() (mirrors ImagePooled()), and pass !cfg.VideoPooled() as runGenerateVideo's singleCard argument instead of the hardcoded false. A pooled video seat is unaffected (still no pin -- the blackwell-3x16 pool's compute must stay on ComfyUI's default device per MultiGPU #220, a documented law exception); an un-pooled one now takes the same comfy_cuda_device pin every other single-card route already gets. Tests: TestLaunchProfileReachesEveryComfyRouteAndThePinOnlyTheSingleCardOnes updated (generate_video: false -> true; added a "generate_video (pooled)" case asserting the pin STAYS off there) and TestEveryGPUTaskRunsUnderALease/TestEveryGpuSlotRunnerIsCoveredByTheLeaseTable cover the new case. go build/vet/test (internal/pipeline, internal/config) all green. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
dmmdea
added a commit
that referenced
this pull request
Sep 24, 2026
… pooled A/B measured on the reference box (2026-09-24, same session as the ComfyUI-MultiGPU archival pin): native single-card streaming (videogen_pool_vvram_gb/imagegen_pool_vvram_gb unset, comfy_dynamic_vram=on, comfy_cuda_device=2) beats the pooled DisTorch2 recipe this tier seeded since 0.59.0/0.113.x on every measured axis: - krea2 image: 77.7s native vs 128.3s pooled (40% faster); output pixel-identical (same model/precision, deterministic sampler). - LTX-2.5 video 1280x704: 107.0s native vs 178.7s pooled (40% faster); frames/audio match (mean -43.0..-43.7 dB, max -27.3..-30.8 dB, both runs -- no silence, no clipping). - LTX-2.5 video 1920x1088 (the flagship resolution POOLING CANNOT REACH on this tier -- documented OOM at every virtual_vram_gb, per the prior note): native fits at 107.3s, peak 15.4/15.9 GiB on the compute card, zero OOM. - Display-card safety (nvidia-smi index 1, this box's RTX 5070 Ti): pooled krea2 loaded 12.1 GiB there via the CLIPLoader "default" device (untargeted by any pool key -- a live, previously-undocumented risk on the same class of incident recorded in this tier's own notes for 2026-09-04); pooled LTX-2.5 computes there BY DESIGN (MultiGPU #220, unaffected/unchanged). Native isolates fully to the compute card in both cases (confirmed via torch.cuda's own device enumeration plus nvidia-smi 1 Hz logging): display card never exceeds its ~1.5-2.7 GiB desktop baseline. Required a prerequisite fix (PR #478, merged first) -- generate_video had no device-pin path for the un-pooled shape at all, so the very first native measurement landed on the display card until that shipped. Removed imagegen_pool_vvram_gb/compute/donor and videogen_pool_vvram_gb/ compute/donor from this tier's config_seed (absent = the builders' own un-pooled UNETLoader branch); added comfy_dynamic_vram="on" (comfy_cuda_device stays "2", already seeded). docs/tiers/blackwell-3x16.md regenerated (go run ./cmd/gentiers). videogen_width/height stay 1280x704 as the seeded default (conservative: 1920x1088's ~15.4/15.9 GiB peak is thin margin for a one-size-fits-all seed) -- 1920x1088 is documented as available and measured in docs/systems/media-generation.md. The live reference box (Qube) already carries this exact config change, applied and smoke-tested in the same session (see infra/multigpu-archival-actions-2026-09-24.md on the operator's Drive). go test ./internal/tierseed/... ./internal/config/... ./internal/mediacap/... and setup/render.tests.ps1 (ALL PASS) both green. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
dmmdea
added a commit
that referenced
this pull request
Sep 24, 2026
dmmdea
added a commit
that referenced
this pull request
Sep 24, 2026
…6-09-30 archival (#477) * chore(comfy): pin ComfyUI-MultiGPU to v2.6.4 + local fix ahead of 2026-09-30 archival Upstream pollockjj/ComfyUI-MultiGPU archives 2026-09-30 (issue #223, no successor endorsed), and neither native Dynamic VRAM nor SelectModelDevice/"MultiGPU Work Units" reproduce its donor+compute virtual_vram_gb weight-sharding, which the Wan GGUF lane and the pooled krea2/LTX-2.5 seats depend on. Pin rather than migrate blind (research: comfyui-multigpu-archival-2026-09-24.md). Pinned commit ed1ffaef7cec1a66f35106c6a4c7a40927c2dc83 = upstream v2.6.4's last code commit (b51c99a5) plus one already-deployed local fix for upstream issue #220's libcudart.so-on-Windows crash. Created a private fallback mirror (dmmdea/ComfyUI-MultiGPU-mirror, full history + this commit) for when upstream goes read-only, and named the pin + mirror in both node-pack tables that already existed for this exact purpose (render/comfy-nodes.mjs NODE_PACKS, internal/mediacap/routeneeds.go nodePacks/packHint) so a box missing the pack is pointed at the frozen, patched source instead of a soon-archived URL. Docs: media-generation.md "ComfyUI-MultiGPU archival" section records the pin, the mirror, and the compute_device constraint the local fix satisfies. Fleet alignment (Qube/OptiPlex/Aorus/Lenovo, all now on this commit) and the LTX-2.5/ krea2 native-vs-pooled A/B are tracked outside this repo, on the operator's Drive (infra/multigpu-archival-actions-2026-09-24.md). go vet ./..., go test ./... (root module) and node --test render/*.test.mjs all green. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(tiers): blackwell-3x16 seeds native LTX-2.5/krea2 streaming, not pooled A/B measured on the reference box (2026-09-24, same session as the ComfyUI-MultiGPU archival pin): native single-card streaming (videogen_pool_vvram_gb/imagegen_pool_vvram_gb unset, comfy_dynamic_vram=on, comfy_cuda_device=2) beats the pooled DisTorch2 recipe this tier seeded since 0.59.0/0.113.x on every measured axis: - krea2 image: 77.7s native vs 128.3s pooled (40% faster); output pixel-identical (same model/precision, deterministic sampler). - LTX-2.5 video 1280x704: 107.0s native vs 178.7s pooled (40% faster); frames/audio match (mean -43.0..-43.7 dB, max -27.3..-30.8 dB, both runs -- no silence, no clipping). - LTX-2.5 video 1920x1088 (the flagship resolution POOLING CANNOT REACH on this tier -- documented OOM at every virtual_vram_gb, per the prior note): native fits at 107.3s, peak 15.4/15.9 GiB on the compute card, zero OOM. - Display-card safety (nvidia-smi index 1, this box's RTX 5070 Ti): pooled krea2 loaded 12.1 GiB there via the CLIPLoader "default" device (untargeted by any pool key -- a live, previously-undocumented risk on the same class of incident recorded in this tier's own notes for 2026-09-04); pooled LTX-2.5 computes there BY DESIGN (MultiGPU #220, unaffected/unchanged). Native isolates fully to the compute card in both cases (confirmed via torch.cuda's own device enumeration plus nvidia-smi 1 Hz logging): display card never exceeds its ~1.5-2.7 GiB desktop baseline. Required a prerequisite fix (PR #478, merged first) -- generate_video had no device-pin path for the un-pooled shape at all, so the very first native measurement landed on the display card until that shipped. Removed imagegen_pool_vvram_gb/compute/donor and videogen_pool_vvram_gb/ compute/donor from this tier's config_seed (absent = the builders' own un-pooled UNETLoader branch); added comfy_dynamic_vram="on" (comfy_cuda_device stays "2", already seeded). docs/tiers/blackwell-3x16.md regenerated (go run ./cmd/gentiers). videogen_width/height stay 1280x704 as the seeded default (conservative: 1920x1088's ~15.4/15.9 GiB peak is thin margin for a one-size-fits-all seed) -- 1920x1088 is documented as available and measured in docs/systems/media-generation.md. The live reference box (Qube) already carries this exact config change, applied and smoke-tested in the same session (see infra/multigpu-archival-actions-2026-09-24.md on the operator's Drive). go test ./internal/tierseed/... ./internal/config/... ./internal/mediacap/... and setup/render.tests.ps1 (ALL PASS) both green. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Found while running the ComfyUI-MultiGPU archival A/B (see PR #477 and infra/multigpu-archival-actions-2026-09-24.md on the operator's Drive):
runGenerateVideohard-codedcomfyGenEnv(false)on the theory that "every seeded video seat is pooled and owns its placement through its pool keys." That is false once a box runs an un-pooled/native video seat (videogen_pool_vvram_gb<=0) -- exactly the shape the archival A/B is measuring as a faster, display-card-safe alternative to pooled DisTorch2.docs/tiers/blackwell-3x16.mdalready documents this ordering for the POOLED compute/donor keys ("PIN BY GPU UUID, NEVER INDEX... TWO DEVICE ORDERINGS ARE IN PLAY"); the un-pooled path had no equivalent protection.What changed
internal/config/families.go: addedConfig.VideoPooled(), mirroring the existingImagePooled().internal/pipeline/pipeline.go:runGenerateVideonow passes!p.cfg.VideoPooled()as thecomfyGenEnvsingleCard argument instead of the hardcodedfalse. A pooled video seat is unaffected (still no pin -- the blackwell-3x16 pool's compute must stay on ComfyUI's default device per MultiGPU changelog 0.112.0: two-card store number #220, a documented law exception); an un-pooled one now takes the same pin every other single-card route already gets. Doc comments oncomfyLaunchupdated to match.TestLaunchProfileReachesEveryComfyRouteAndThePinOnlyTheSingleCardOnes(generate_video: false -> true; added agenerate_video (pooled)case asserting the pin stays OFF there) and the twogpuLeaseCases()-driven completeness tests.How tested
go build ./...,go vet ./...-- clean.go test ./internal/pipeline/... ./internal/config/...-- all green, including the new pooled-video case.Risk
Low-medium: changes production GPU device placement for one route (un-pooled video generation). Scoped by a passing test asserting the pooled path is byte-for-byte unaffected. Live-verified on the actual 3-card box the risk applies to, not just unit tests.
🤖 Generated with Claude Code