Skip to content

chore(comfy): pin ComfyUI-MultiGPU to v2.6.4 + local fix ahead of 2026-09-30 archival - #477

Merged
dmmdea merged 5 commits into
mainfrom
chore/mgpu-archival-pin
Sep 24, 2026
Merged

dmmdea merged 5 commits into
mainfrom
chore/mgpu-archival-pin

Conversation

@dmmdea

@dmmdea dmmdea commented Sep 24, 2026

Copy link
Copy Markdown
Owner

Summary

  • pollockjj/ComfyUI-MultiGPU archives 2026-09-30 (issue profiles: close, mark and synthesize every exemplar thread (phantom "Go version" digests) — 0.113.1 #223, no successor endorsed). Neither native Dynamic VRAM nor SelectModelDevice/"MultiGPU Work Units" reproduce its donor+compute virtual_vram_gb weight-sharding, which the Wan GGUF lane and the pooled krea2/LTX-2.5 seats depend on -- so this pins rather than migrates blind (full research: comfyui-multigpu-archival-2026-09-24.md, operator's Drive).
  • Pins the fleet to ed1ffaef7cec1a66f35106c6a4c7a40927c2dc83 = upstream v2.6.4's last code commit (b51c99a5) plus one already-deployed local fix (fix(p2p): platform-aware cudart load + fail-closed P2P on Windows/WDDM) for upstream issue changelog 0.112.0: two-card store number #220's libcudart.so-on-Windows crash.
  • Creates a private fallback mirror (dmmdea/ComfyUI-MultiGPU-mirror, full git clone --mirror history + this pinned commit) for when upstream goes read-only.
  • Names the pin + mirror in both existing node-pack tables (render/comfy-nodes.mjs NODE_PACKS, internal/mediacap/routeneeds.go nodePacks/packHint) so the MISSING_NODE hint points at the frozen, patched source instead of a soon-archived URL.
  • Docs: docs/systems/media-generation.md gets a "ComfyUI-MultiGPU archival" section recording the pin, the mirror, and the compute_device constraint the local fix satisfies (why the pooled LTX-2.5 seat stays on cuda:0).
  • All four fleet ComfyUI installs (Qube, OptiPlex, Aorus, Lenovo) are already aligned to this commit as of 2026-09-24 (git-checked, recorded outside this repo).

How tested

  • go build ./..., go vet ./..., go test ./... (root module) -- all green, before and after merging in origin/main's concurrent PR fix(tiers): seed blackwell-8 and ampere-16 with 2026-09-23/24 measured values #475 (tier seeds).
  • node --test render/*.test.mjs -- 439/439 green.
  • go test -run TestConflictMarker ./... -- clean (no leftover merge markers).
  • Diff scanned for tailnet/LAN/user-path leaks before pushing (this is a public repo) -- clean.

Risk

Low. Additive: a note field on an existing struct/table plus doc/changelog text; no route behavior changes, no config schema changes. The only executable-path change is the operator-facing hint string shown when ComfyUI-MultiGPU is missing.

Generated with Claude Code

dmmdea and others added 2 commits September 24, 2026 07:31
…6-09-30 archival

Upstream pollockjj/ComfyUI-MultiGPU archives 2026-09-30 (issue #223, no successor
endorsed), and neither native Dynamic VRAM nor SelectModelDevice/"MultiGPU Work
Units" reproduce its donor+compute virtual_vram_gb weight-sharding, which the Wan
GGUF lane and the pooled krea2/LTX-2.5 seats depend on. Pin rather than migrate
blind (research: comfyui-multigpu-archival-2026-09-24.md).

Pinned commit ed1ffaef7cec1a66f35106c6a4c7a40927c2dc83 = upstream v2.6.4's last
code commit (b51c99a5) plus one already-deployed local fix for upstream issue
#220's libcudart.so-on-Windows crash. Created a private fallback mirror
(dmmdea/ComfyUI-MultiGPU-mirror, full history + this commit) for when upstream
goes read-only, and named the pin + mirror in both node-pack tables that already
existed for this exact purpose (render/comfy-nodes.mjs NODE_PACKS,
internal/mediacap/routeneeds.go nodePacks/packHint) so a box missing the pack is
pointed at the frozen, patched source instead of a soon-archived URL.

Docs: media-generation.md "ComfyUI-MultiGPU archival" section records the pin,
the mirror, and the compute_device constraint the local fix satisfies. Fleet
alignment (Qube/OptiPlex/Aorus/Lenovo, all now on this commit) and the LTX-2.5/
krea2 native-vs-pooled A/B are tracked outside this repo, on the operator's Drive
(infra/multigpu-archival-actions-2026-09-24.md).

go vet ./..., go test ./... (root module) and node --test render/*.test.mjs all
green.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
dmmdea and others added 3 commits September 24, 2026 08:23
… pooled

A/B measured on the reference box (2026-09-24, same session as the
ComfyUI-MultiGPU archival pin): native single-card streaming
(videogen_pool_vvram_gb/imagegen_pool_vvram_gb unset, comfy_dynamic_vram=on,
comfy_cuda_device=2) beats the pooled DisTorch2 recipe this tier seeded since
0.59.0/0.113.x on every measured axis:

- krea2 image: 77.7s native vs 128.3s pooled (40% faster); output
  pixel-identical (same model/precision, deterministic sampler).
- LTX-2.5 video 1280x704: 107.0s native vs 178.7s pooled (40% faster);
  frames/audio match (mean -43.0..-43.7 dB, max -27.3..-30.8 dB, both
  runs -- no silence, no clipping).
- LTX-2.5 video 1920x1088 (the flagship resolution POOLING CANNOT REACH on
  this tier -- documented OOM at every virtual_vram_gb, per the prior note):
  native fits at 107.3s, peak 15.4/15.9 GiB on the compute card, zero OOM.
- Display-card safety (nvidia-smi index 1, this box's RTX 5070 Ti): pooled
  krea2 loaded 12.1 GiB there via the CLIPLoader "default" device (untargeted
  by any pool key -- a live, previously-undocumented risk on the same class
  of incident recorded in this tier's own notes for 2026-09-04); pooled
  LTX-2.5 computes there BY DESIGN (MultiGPU #220, unaffected/unchanged).
  Native isolates fully to the compute card in both cases (confirmed via
  torch.cuda's own device enumeration plus nvidia-smi 1 Hz logging): display
  card never exceeds its ~1.5-2.7 GiB desktop baseline.

Required a prerequisite fix (PR #478, merged first) -- generate_video had no
device-pin path for the un-pooled shape at all, so the very first native
measurement landed on the display card until that shipped.

Removed imagegen_pool_vvram_gb/compute/donor and videogen_pool_vvram_gb/
compute/donor from this tier's config_seed (absent = the builders' own
un-pooled UNETLoader branch); added comfy_dynamic_vram="on" (comfy_cuda_device
stays "2", already seeded). docs/tiers/blackwell-3x16.md regenerated
(go run ./cmd/gentiers). videogen_width/height stay 1280x704 as the seeded
default (conservative: 1920x1088's ~15.4/15.9 GiB peak is thin margin for a
one-size-fits-all seed) -- 1920x1088 is documented as available and measured
in docs/systems/media-generation.md.

The live reference box (Qube) already carries this exact config change,
applied and smoke-tested in the same session (see
infra/multigpu-archival-actions-2026-09-24.md on the operator's Drive).

go test ./internal/tierseed/... ./internal/config/... ./internal/mediacap/...
and setup/render.tests.ps1 (ALL PASS) both green.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@dmmdea
dmmdea merged commit 15b90ac into main Sep 24, 2026
5 checks passed
@dmmdea
dmmdea deleted the chore/mgpu-archival-pin branch September 24, 2026 13:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant