Problem
The current launch topology mixes three user-facing concepts:
training.devices (host-visible CUDA indices)
- rank-local CUDA index remapping inside launcher children
- backend payload fields and inference/ring/learner device strings
Our required policy is simpler:
- one rank owns exactly one GPU;
- learner, collector public plane, inference ring, replay, and backend payload all share that GPU;
- CPU-physics backends bind each rank to a CPU block;
- no cross-device CUDA IPC inside one rank.
The strict same-device guard in resolve_inference_placement is correct and should remain; the topology source and configuration burden should change.
Proposed contract
Rank-local visibility
- Single-rank execution: an externally set single-entry
CUDA_VISIBLE_DEVICES is authoritative. The rank uses cuda:0 in its local namespace; no training.devices is required.
- Multi-rank launch: each rank child receives exactly one visible physical GPU (
CUDA_VISIBLE_DEVICES=<one entry>) plus existing rank metadata. All in-rank consumers use cuda:0; remapped global indices never reach EnvCfg/backend payloads.
- Preserve UUID-based parent
CUDA_VISIBLE_DEVICES entries when mapping physical devices.
Configuration precedence
- external rank-local single-entry
CUDA_VISIBLE_DEVICES: authoritative;
training.devices: compatibility multi-rank selection used to partition/remap visibility before children start;
- conflict between an in-rank explicit device and visibility fails closed.
training.devices remains supported during transition but is not required for ordinary single-rank training.
CPU-physics rank binding
Acceptance criteria
Problem
The current launch topology mixes three user-facing concepts:
training.devices(host-visible CUDA indices)Our required policy is simpler:
The strict same-device guard in
resolve_inference_placementis correct and should remain; the topology source and configuration burden should change.Proposed contract
Rank-local visibility
CUDA_VISIBLE_DEVICESis authoritative. The rank usescuda:0in its local namespace; notraining.devicesis required.CUDA_VISIBLE_DEVICES=<one entry>) plus existing rank metadata. All in-rank consumers usecuda:0; remapped global indices never reach EnvCfg/backend payloads.CUDA_VISIBLE_DEVICESentries when mapping physical devices.Configuration precedence
CUDA_VISIBLE_DEVICES: authoritative;training.devices: compatibility multi-rank selection used to partition/remap visibility before children start;training.devicesremains supported during transition but is not required for ordinary single-rank training.CPU-physics rank binding
dp_collector_cpu_ids/EnvCfg.cpu_idsremains the explicit CPU block owner.Acceptance criteria
CUDA_VISIBLE_DEVICES=<one GPU>withouttraining.devicesproduces env/ring/learner allcuda:0.cuda:0.training.devicesremains compatible.