Skip to content

feat(training): adopt CUDA_VISIBLE_DEVICES as the rank-local GPU contract #1794

Description

@TATP-233

Problem

The current launch topology mixes three user-facing concepts:

  1. training.devices (host-visible CUDA indices)
  2. rank-local CUDA index remapping inside launcher children
  3. backend payload fields and inference/ring/learner device strings

Our required policy is simpler:

  • one rank owns exactly one GPU;
  • learner, collector public plane, inference ring, replay, and backend payload all share that GPU;
  • CPU-physics backends bind each rank to a CPU block;
  • no cross-device CUDA IPC inside one rank.

The strict same-device guard in resolve_inference_placement is correct and should remain; the topology source and configuration burden should change.

Proposed contract

Rank-local visibility

  • Single-rank execution: an externally set single-entry CUDA_VISIBLE_DEVICES is authoritative. The rank uses cuda:0 in its local namespace; no training.devices is required.
  • Multi-rank launch: each rank child receives exactly one visible physical GPU (CUDA_VISIBLE_DEVICES=<one entry>) plus existing rank metadata. All in-rank consumers use cuda:0; remapped global indices never reach EnvCfg/backend payloads.
  • Preserve UUID-based parent CUDA_VISIBLE_DEVICES entries when mapping physical devices.

Configuration precedence

  1. external rank-local single-entry CUDA_VISIBLE_DEVICES: authoritative;
  2. training.devices: compatibility multi-rank selection used to partition/remap visibility before children start;
  3. conflict between an in-rank explicit device and visibility fails closed.

training.devices remains supported during transition but is not required for ordinary single-rank training.

CPU-physics rank binding

Acceptance criteria

  • CUDA_VISIBLE_DEVICES=<one GPU> without training.devices produces env/ring/learner all cuda:0.
  • Multi-rank children each see exactly one GPU.
  • Backend payload fields resolve to local cuda:0.
  • Cross-device requests fail closed.
  • training.devices remains compatible.
  • CPU backend partitioning retains packed HOST_BRIDGE behavior.
  • Launcher/placement/backend mapping contract tests pass.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions