Fix Dragon inference with local loader processes - #91
Conversation
|
Important Review skippedAuto reviews are limited based on label configuration. 🏷️ Required labels (at least one) (1)
Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Repository: NVIDIA/cuPhoton/.coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Comment |
melo-gonzo
left a comment
There was a problem hiding this comment.
LGTM. All 35 lifecycle tests passed with Dragon 0.14.2, and the new spawn-queue regression fails on the base revision. The environment overlay preserves inherited variables and leaves the coordinator environment unchanged. Multi-node B200/GB200 validation was not independently repeated.
Signed-off-by: Trent Nelson <trentn@nvidia.com>
0611603 to
4047d64
Compare
Dragon xScan inference with
--num-workers 1stalls before READY because the local Python spawn child receives a Dragon-patched queue and cannot finish bootstrap. Keep Python multiprocessing unpatched when starting native Dragon workers, so their local DataLoader children use ordinary spawn queues while native Dragon coordination remains intact.Validated on x86-64 B200 and two-node aarch64 GB200 with two GPU workers and three repeated rounds. Positive-loader inference and the default loader path complete; scientific outputs match the original candidate's working loader-zero baseline exactly, loader children persist across rounds, and shutdown leaves no owned processes. Separate fix wheels preserve the original candidate evidence. All 35 focused lifecycle tests, including a real Dragon spawn-queue regression, and pre-commit checks pass.