Price HybridEP expert work and buffer growth at EP>1 - #954
Merged
Merged
Conversation
At EP>1 the routed-expert working set was priced at zero, so cold CP2/EP2 plans fell back to retention floors alone (8.4 GB predicted against 34 GB observed on 8 layers). HybridEP hands each rank the pairs routed to its local experts, so reuse the EP1 per-token coefficient with a 1.5x routing-imbalance allowance (a pretrained CP2/EP2 run put 1.35x the balanced load on one rank). HybridEP's persistent buffers are allocated outside the PyTorch allocator after admission, so charge any growth a plan triggers: capacity x EP x (2H + 5E + 4H/128) bytes, less what is already held. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton
had a problem deploying
to
trainer-rank-gpu-validation
September 25, 2026 00:45 — with
GitHub Actions
Error
Review follow-ups: the growth now enters the forward peak and, via the max-combined checkpoint workspace, the checkpoint peak, but not retention, so split plans pay it once. The replacement buffer is charged in full because the old one stays referenced while it is allocated. Rows per rank follow HybridEP's TMA alignment, 512-row minimum and 64-row chunks, over the ETPxEP group and its expert columns. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton
had a problem deploying
to
trainer-rank-gpu-validation
September 25, 2026 01:04 — with
GitHub Actions
Error
A grown buffer persists into later split children, so a child with the largest workspace can run beside another child's growth. Track growth as its own cost field, exclude it from each child's ephemeral and workspace terms, and add the largest growth once to both split aggregates. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton
had a problem deploying
to
trainer-rank-gpu-validation
September 25, 2026 01:13 — with
GitHub Actions
Error
Replay compares recorded cost components with asdict(cost), which now includes hybridep_growth; a tampered value must also fail to match. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton
deployed
to
trainer-rank-gpu-validation
September 25, 2026 01:20 — with
GitHub Actions
Active
This was referenced Sep 25, 2026
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Partially addresses #949. After #951, cold admission at EP>1 still priced the routed-expert working set at zero.
_moe_output_bytes_per_tokenreturned 0 unless TP, EP and ETP were all 1. The HybridEP flex dispatcher ART uses at EP>1 was not modeled, and neither were its persistent buffers. The CP2/EP2 cold OOM in #949 happened in that working set: expert FC2 LoRA add, BF16[2021662, 2048].Measurements
Two local H200s, Qwen3.6-35B-A3B with random weights, 8 layers, CP2, ten unique-token requests scaled to 0.25 of #949's first wave (194,753 tokens), forward plus backward. Values are per rank.
mainThe cold EP2 call used about 38.5 GB of device memory in total: the PyTorch peak plus the non-PyTorch growth. The new prediction covers that with about 10% margin.
The cold-versus-warm gap is not a cold effect. Allocator traces (
torch.cuda.memoryhistory) of both calls show the same per-call working set: allocations made during each call peak at 35.05 GB cold and 34.80 GB warm. The warm call's net peak above its baseline is only 25.5 GB because 9.28 GB left over from the previous call is freed during it. 8.4 GB of that is released inside Megatron's HybridEPdispatch, when each MoE layer replaces the previous call's dispatcher state (about 1 GB per layer here). The warm prediction of 40.3 GB therefore compares with a working set of about 34.8 GB.Changes
_plan_hybridep_growth_bytescomputes the capacity_configure_hybridepwill request: the busiest CP rank's rows, and at least the near-balanced reservation. When the held buffer is smaller, it charges the replacement's full size, because the old buffer stays referenced while the new one is allocated. HybridEP pads rows per rank to TMA alignment, a 512-row minimum and 64-row combine chunks, over the ETP×EP group. The charge is that padded count times the group size times(2H + 5E + 4H/128)bytes, where E is the group's expert columns. The terms are BF16 tokens, FP32 probabilities plus a routing-map byte per expert column, and FP32 scaling factors, which are allocated even without FP8; dispatch outputs alias the combine inputs by default. The charge enters both the forward and checkpoint peaks of a plan, but not forward retention. A grown buffer persists into later split children, so a split charges the largest child's growth once, beside whichever child's forward or checkpoint peak is highest. It is recorded in planner-report arguments, so replay passes it through.Not addressed
Validation
tests/unit/test_trainer_rank_*.pyplustest_prefix_tree_packing.pyon the previous head: 1,255 passed. It had four failures. The replay-fixture failure is fixed on this head (planner-reports file 43/43). The other three also fail on currentmainin isolation: two variants oftest_native_checkpoint_gather_after_recovery, which fails different variants each run, and the live-graph collective test's wall-clock deadline. The new tests cover:mainruns used the same harness and interpreter. Allocator traces of both calls are the basis for the cold-versus-warm note.🤖 Generated with Claude Code