Skip to content

Try budgeted CUDA cache recovery before accepting internal splits - #972

Draft
bradhilton wants to merge 5 commits into
mainfrom
schulman/recovery-before-split-integration-20260925
Draft

bradhilton wants to merge 5 commits into
mainfrom
schulman/recovery-before-split-integration-20260925

Conversation

@bradhilton

@bradhilton bradhilton commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

When the minimum scheduling wave fails its unsplit memory check but an internal split fits, admission currently returns the split before trying unused CUDA cache recovery. Retain the exact unsplit check and allow one optional recovery through the existing first-call allowance and later 5% budget. An actual release triggers fresh search and admission; budget denial retains the fitting split fallback, subject to its normal refreshed check.

The existing WORLD outcome reduction coordinates split, healthy and empty peers. Successful returns discard the temporary recovery-plan reference, including oversized admission. Public APIs and art.megatron are unchanged. Fresh search can select a wider or differently packed plan, so this is a numerical/layout-significant change requiring Brad's decision before merge.

Validation:

  • Original candidate: 157 CPU cases plus 15 subtests; independent real-NCCL audit of four injected two-rank cases; six full-model flat forwards/backwards with two actual releases. The later 5% decision denominator was dominated by warmup/compilation, so total and steady-state overhead remain unqualified. No optimizer or model-quality claim.
  • Frozen f93b6d85 local two-H200 qualification: all five exact GPU-CI command groups passed (95 pytest cases, two skips, CP2/TP2 trainer checks). Independent audit and cleanup passed. This remains evidence for that source, with successor applicability assessed through source review; it is not a relabeled native run on the current head or general full-size OOM proof.
  • The subsequent typing fix uses one typing.cast identity call plus test narrowing. Production is otherwise AST-identical to f93b6d85. The following two test-only corrections update inherited mocks for the optional recovery path while preserving their original assertions.
  • Current c15b1f20 changes only the empty-DP regression. It models exact recovery reductions and the peer split vote, checks matching rank traces and rejects CPU cache release. Seven unchanged companion cases plus two subtests, the corrected affected case, two positive/nine negative mock probes, and affected-file type/lint/format checks pass. Three independent reviewers cleared this exact test delta; production source is unchanged from its parent.
  • A separate local run of the exact hosted unit selection, composed with merged ART Collect dynamo's cyclic garbage after compiling a checkpointed layer #955, stopped during collection because the existing environment's TRL imports missing vllm. No dependency installation, extra exclusions or retry occurred. This is not a broad-suite pass.

The preceding hosted quality run passed all hooks and 1,228 lightweight cases, then failed the empty-DP mock corrected here (1,949 unit cases passed). Its hosted GPU job failed capacity acquisition before tests. Current-head Prek36143581458 passed all hooks, 1,228 lightweight tests and 1,970 unit tests (34 skipped, two deselected), with container/network cleanup complete. The tested merge was 160c246b50c3bf17f75157c13bfe322f2797195b, combining this head with main 99e290cd (ART #955 and #967); it is not the bare PR-head tree. GPU validation36143581357 failed capacity acquisition before any tests. Its exact request was cancelled and the intended cluster was reported absent; no independent provider census was performed. No hosted gate is waived; this PR remains draft and held for the significant behavior decision. Independent CI receipt: 13dbacd41e0a98391123bd5ba8ca79572a72e7a113a2899dc4728f7b9b248a28.

@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 25, 2026 12:34 — with GitHub Actions Failure
@bradhilton

Copy link
Copy Markdown
Collaborator Author

Root integration record for exact head f93b6d85338b62c52c6e0e2e66b95ceddd8b7d24:

Three independent agent source reviews are CLEAR: recovery-before-split reviewer (0b19c6a4), collector/source reviewer (71b41b4b), and independent successor reviewer (ef36111f). All explicitly preserve the merge hold for changed packing, additional executed recovery/search collectives, and use of the shared first-release allowance/budget. Public APIs and art.megatron are unchanged.

Review reproduced an oversized-override lifetime defect in the original candidate. The final two-line correction clears the temporary recovery target on that successful return, preserving the selected plan, inputs, indices and refreshed check. The regression fails before the correction and passes afterward; the independent focused four-case run passed. The original failing review is preserved.

Native evidence is deliberately narrower than exact-final-head GPU qualification: the predecessor's ordinary path completed six flat forwards/backwards with two actual releases, and a separate four-case NCCL fixture passed. The final cleanup-only branch has CPU evidence; hosted CI for this head has just started. The later budget decision's denominator was 93.4% warmup/compilation-associated work, so no steady-state or total-overhead guarantee follows. All owned native processes and GPUs are physically closed.

Private exact review/result receipts are retained under /home/brad/.local/share/schulman/, indexed by krennic-packed-boundary-20260925/recovery-pr972-root-integration/REVIEWS.json. This remains a draft; no merge, deployment, frozen-run adoption or general OOM-resolution claim.

@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 25, 2026 12:59 — with GitHub Actions Failure
@bradhilton

Copy link
Copy Markdown
Collaborator Author

Published cc3591f68c62d2fac6bd1ac1a19de7a0b5861aab as a fast-forward type-check correction. Three independent exact-head source reviews clear the delta: 093db0e2, c08f4419, 01ee2e45. The complete production AST equals f93b6d85 after removing the single existing typing.cast identity call; there is no new branch, collective or changed error ordering. Test fixture casts and explicit narrowing preserve prior assertions.

The retained 13 type errors reproduce before and clear after; affected-file type/lint/format checks and 11 focused CPU cases pass. One reviewer additionally checked 13 small outcome/error cases without Torch. These do not replace fresh hosted CI. The currently running local GPU suite remains frozen at f93b6d85, with source equivalence documented separately. The feature remains draft and held for Brad because fresh recovery can change layout/packing. No API or art.megatron changes and no merge/deployment.

@bradhilton

Copy link
Copy Markdown
Collaborator Author

Independent local two-H200 qualification is complete on frozen f93b6d85338b62c52c6e0e2e66b95ceddd8b7d24: all five exact CI command groups exited 0; 95 pytest cases passed and 2 skipped, plus passing CP2/TP2 trainer checks. The two skip reasons are inferred from exact source/device count, since -rs was not enabled.

All 97 source pins matched. Root and independent review joined the 254 tracked identities, 10 groups and 9 sessions as retired, with both assigned H200s idle. Runtime 870.291 seconds; peak sampled process-family RSS 8.565 GiB. Audit receipt 512ec8db9cf2c7f0cd042bdafc7cd9bf7032303966ea7a142f8e7869fea33ac7, root closure 1348dba44bf494c885b7f20fed328b8a9e7d968baef2584d8c9deb246751253c.

The current cc3591f successor adds one reviewed identity cast and test narrowing; this supports relevance of the f93 result without relabeling the executed head. Both hosted GPU attempts failed acquisition before tests; the required check remains failed. No full-size OOM, throughput, numerical-equivalence or merge clearance is implied.

@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 25, 2026 13:21 — with GitHub Actions Failure
@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 25, 2026 13:51 — with GitHub Actions Failure

This branch had an error being deployed

1 failed deployment
trainer-rank-gpu-validation — c15b1f20 Deployed Sep 25, 2026 by bradhilton via Run on 2x H200 #700
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant