Skip to content

Price HybridEP expert work and buffer growth at EP>1 - #954

Merged
bradhilton merged 4 commits into
mainfrom
dalinar/planner-ep-moe
Sep 25, 2026
Merged

bradhilton merged 4 commits into
mainfrom
dalinar/planner-ep-moe

Conversation

@bradhilton

@bradhilton bradhilton commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

Partially addresses #949. After #951, cold admission at EP>1 still priced the routed-expert working set at zero. _moe_output_bytes_per_token returned 0 unless TP, EP and ETP were all 1. The HybridEP flex dispatcher ART uses at EP>1 was not modeled, and neither were its persistent buffers. The CP2/EP2 cold OOM in #949 happened in that working set: expert FC2 LoRA add, BF16 [2021662, 2048].

Measurements

Two local H200s, Qwen3.6-35B-A3B with random weights, 8 layers, CP2, ten unique-token requests scaled to 0.25 of #949's first wave (194,753 tokens), forward plus backward. Values are per rank.

observed PyTorch peak device memory outside PyTorch predicted on main predicted with this PR
EP2, cold 33.5–34.7 GB +3.82 GB during the call 8.4 GB 42.5 GB
EP2, warm 24.8–25.2 GB unchanged 39.1 GB 40.3 GB
EP1, cold (reference) 46.0–51.9 GB — 29.4 GB 29.4 GB

The cold EP2 call used about 38.5 GB of device memory in total: the PyTorch peak plus the non-PyTorch growth. The new prediction covers that with about 10% margin.

The cold-versus-warm gap is not a cold effect. Allocator traces (torch.cuda.memory history) of both calls show the same per-call working set: allocations made during each call peak at 35.05 GB cold and 34.80 GB warm. The warm call's net peak above its baseline is only 25.5 GB because 9.28 GB left over from the previous call is freed during it. 8.4 GB of that is released inside Megatron's HybridEP dispatch, when each MoE layer replaces the previous call's dispatcher state (about 1 GB per layer here). The warm prediction of 40.3 GB therefore compares with a working set of about 34.8 GB.

Changes

  • Expert working set at EP>1. HybridEP delivers each rank the token-expert pairs routed to its local experts, already permuted. With balanced routing that is local tokens × top-k, the same as EP1. So when the dispatcher is Megatron's flex dispatcher with the HybridEP manager, at TP1/ETP1, the EP1 per-token coefficient applies, with routed rows scaled by 1.5. The allowance comes from TrainerRank CP2/EP1 normal admission reaches CUDA OOM in compiled MoE LoRA forward #949's pretrained CP2/EP2 run: one rank received 2.02M pairs against a balanced 3.12M, so its peer received about 1.35×. It also currently covers other work the static floors miss. With balanced routing (1.0×), the EP2 prediction above would be about 32 GB against about 38.5 GB used, the same kind of shortfall as at CP2/EP1. The worst case is EP× the balanced load. The EP1 accounting for dispatched inputs, which counts two H-wide copies where HybridEP keeps one, is reused as a conservative bound. Other flex backends (DeepEP), all-to-all at EP>1, and TP/ETP>1 still price zero.
  • HybridEP buffer growth. The dispatch and combine buffers are allocated outside the PyTorch allocator when a forward first needs a larger capacity, which is after admission. So neither the sampled free memory nor a learned peak sees them. _plan_hybridep_growth_bytes computes the capacity _configure_hybridep will request: the busiest CP rank's rows, and at least the near-balanced reservation. When the held buffer is smaller, it charges the replacement's full size, because the old buffer stays referenced while the new one is allocated. HybridEP pads rows per rank to TMA alignment, a 512-row minimum and 64-row combine chunks, over the ETP×EP group. The charge is that padded count times the group size times (2H + 5E + 4H/128) bytes, where E is the group's expert columns. The terms are BF16 tokens, FP32 probabilities plus a routing-map byte per expert column, and FP32 scaling factors, which are allocated even without FP8; dispatch outputs alias the combine inputs by default. The charge enters both the forward and checkpoint peaks of a plan, but not forward retention. A grown buffer persists into later split children, so a split charges the largest child's growth once, beside whichever child's forward or checkpoint peak is highest. It is recorded in planner-report arguments, so replay passes it through.

Not addressed

  • Unmodeled memory outside PyTorch. The cold call grew it by 3.82 GB against about 2.1 GB from the buffer formula. The remainder is consistent with NCCL communicators initializing lazily on the first call; that is a one-time cost and is not charged here.
  • Routing imbalance rests on one pretrained observation; the measurements above use random weights, whose routing is near-balanced. Multi-node HybridEP (internode buffers) is not modeled.
  • GDN pending floor is still CP1-only.
  • Dispatcher state held between calls. At EP>1, each MoE layer's HybridEP dispatcher keeps the previous call's state until the next call reaches that layer. The planner sees it as used memory at admission. Its size across layers and tokens in production is not yet established, and releasing it after combine is a separate change.
  • Static floors are still below the full working set at CP1 and CP2/EP1 (for example 29.4 GB predicted against 46–52 GB at CP2/EP1). Apply trainer-rank memory floors per context-parallel rank #951 attributed part of that to a cold first-call excess; these traces suggest the warm figures it compared against were understated in the same way.

Validation

  • CPU-only suite: tests/unit/test_trainer_rank_*.py plus test_prefix_tree_packing.py on the previous head: 1,255 passed. It had four failures. The replay-fixture failure is fixed on this head (planner-reports file 43/43). The other three also fail on current main in isolation: two variants of test_native_checkpoint_gather_after_recovery, which fails different variants each run, and the live-graph collective test's wall-clock deadline. The new tests cover:
    • the EP2/EP4 HybridEP coefficient (12 routed rows instead of 8), including the enclosing FC1 stage;
    • DeepEP, EP-mismatched and EP1 flex dispatchers staying unmodeled;
    • HybridEP row rounding and the buffer formula, including ETP×EP;
    • full-size growth against a smaller held buffer, and none against a sufficient one;
    • growth entering the forward and checkpoint peaks but not retention, and a three-way split paying it once;
    • a split whose largest growth and largest workspace come from different children, charged for both together.
  • GPU: the table above, run through the harness on two H200s. The earlier main runs used the same harness and interpreter. Allocator traces of both calls are the basis for the cold-versus-warm note.

🤖 Generated with Claude Code

At EP>1 the routed-expert working set was priced at zero, so cold
CP2/EP2 plans fell back to retention floors alone (8.4 GB predicted
against 34 GB observed on 8 layers). HybridEP hands each rank the pairs
routed to its local experts, so reuse the EP1 per-token coefficient with
a 1.5x routing-imbalance allowance (a pretrained CP2/EP2 run put 1.35x
the balanced load on one rank).

HybridEP's persistent buffers are allocated outside the PyTorch
allocator after admission, so charge any growth a plan triggers:
capacity x EP x (2H + 5E + 4H/128) bytes, less what is already held.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 25, 2026 00:45 — with GitHub Actions Error
Review follow-ups: the growth now enters the forward peak and, via the
max-combined checkpoint workspace, the checkpoint peak, but not
retention, so split plans pay it once. The replacement buffer is
charged in full because the old one stays referenced while it is
allocated. Rows per rank follow HybridEP's TMA alignment, 512-row
minimum and 64-row chunks, over the ETPxEP group and its expert
columns.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 25, 2026 01:04 — with GitHub Actions Error
A grown buffer persists into later split children, so a child with the
largest workspace can run beside another child's growth. Track growth as
its own cost field, exclude it from each child's ephemeral and workspace
terms, and add the largest growth once to both split aggregates.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 25, 2026 01:13 — with GitHub Actions Error
Replay compares recorded cost components with asdict(cost), which now
includes hybridep_growth; a tampered value must also fail to match.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

This branch was successfully deployed

1 active deployment
trainer-rank-gpu-validation — 587ae7e6 Deployed Sep 25, 2026 by bradhilton via Run on 2x H200 #659
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant