Try budgeted CUDA cache recovery before accepting internal splits - #972
bradhilton wants to merge 5 commits into
Conversation
|
Root integration record for exact head Three independent agent source reviews are CLEAR: recovery-before-split reviewer ( Review reproduced an oversized-override lifetime defect in the original candidate. The final two-line correction clears the temporary recovery target on that successful return, preserving the selected plan, inputs, indices and refreshed check. The regression fails before the correction and passes afterward; the independent focused four-case run passed. The original failing review is preserved. Native evidence is deliberately narrower than exact-final-head GPU qualification: the predecessor's ordinary path completed six flat forwards/backwards with two actual releases, and a separate four-case NCCL fixture passed. The final cleanup-only branch has CPU evidence; hosted CI for this head has just started. The later budget decision's denominator was 93.4% warmup/compilation-associated work, so no steady-state or total-overhead guarantee follows. All owned native processes and GPUs are physically closed. Private exact review/result receipts are retained under |
|
Published The retained 13 type errors reproduce before and clear after; affected-file type/lint/format checks and 11 focused CPU cases pass. One reviewer additionally checked 13 small outcome/error cases without Torch. These do not replace fresh hosted CI. The currently running local GPU suite remains frozen at |
|
Independent local two-H200 qualification is complete on frozen All 97 source pins matched. Root and independent review joined the 254 tracked identities, 10 groups and 9 sessions as retired, with both assigned H200s idle. Runtime 870.291 seconds; peak sampled process-family RSS 8.565 GiB. Audit receipt The current |
When the minimum scheduling wave fails its unsplit memory check but an internal split fits, admission currently returns the split before trying unused CUDA cache recovery. Retain the exact unsplit check and allow one optional recovery through the existing first-call allowance and later 5% budget. An actual release triggers fresh search and admission; budget denial retains the fitting split fallback, subject to its normal refreshed check.
The existing WORLD outcome reduction coordinates split, healthy and empty peers. Successful returns discard the temporary recovery-plan reference, including oversized admission. Public APIs and
art.megatronare unchanged. Fresh search can select a wider or differently packed plan, so this is a numerical/layout-significant change requiring Brad's decision before merge.Validation:
f93b6d85local two-H200 qualification: all five exact GPU-CI command groups passed (95 pytest cases, two skips, CP2/TP2 trainer checks). Independent audit and cleanup passed. This remains evidence for that source, with successor applicability assessed through source review; it is not a relabeled native run on the current head or general full-size OOM proof.typing.castidentity call plus test narrowing. Production is otherwise AST-identical tof93b6d85. The following two test-only corrections update inherited mocks for the optional recovery path while preserving their original assertions.c15b1f20changes only the empty-DP regression. It models exact recovery reductions and the peer split vote, checks matching rank traces and rejects CPU cache release. Seven unchanged companion cases plus two subtests, the corrected affected case, two positive/nine negative mock probes, and affected-file type/lint/format checks pass. Three independent reviewers cleared this exact test delta; production source is unchanged from its parent.vllm. No dependency installation, extra exclusions or retry occurred. This is not a broad-suite pass.The preceding hosted quality run passed all hooks and 1,228 lightweight cases, then failed the empty-DP mock corrected here (1,949 unit cases passed). Its hosted GPU job failed capacity acquisition before tests. Current-head Prek36143581458 passed all hooks, 1,228 lightweight tests and 1,970 unit tests (34 skipped, two deselected), with container/network cleanup complete. The tested merge was
160c246b50c3bf17f75157c13bfe322f2797195b, combining this head with main99e290cd(ART #955 and #967); it is not the bare PR-head tree. GPU validation36143581357 failed capacity acquisition before any tests. Its exact request was cancelled and the intended cluster was reported absent; no independent provider census was performed. No hosted gate is waived; this PR remains draft and held for the significant behavior decision. Independent CI receipt:13dbacd41e0a98391123bd5ba8ca79572a72e7a113a2899dc4728f7b9b248a28.