Skip to content

Add versioned trainer graphs and logical-rank APIs - #925

Draft
bradhilton wants to merge 101 commits into
mainfrom
stark/trainer-v1-20260917
Draft

bradhilton wants to merge 101 commits into
mainfrom
stark/trainer-v1-20260917

Conversation

@bradhilton

@bradhilton bradhilton commented Sep 18, 2026 •

Copy link
Copy Markdown
Collaborator

Adds trainer planning for retained/offloaded/recomputed backward state, bounded gradient staleness, managed output tensors, and checkpoint-owned live parameters/modules. Rank-zero and logical-DP facades support the paired Caladan client API.

Current head b745db5a includes reviewed size reductions, fixed-main reconciliation, a padded GPT-OSS export correction, and coordinated handling of distributed export-preparation failures. Independent tensor/file and two-rank CPU controls passed. Focused CPU checks and statics passed; root and fresh Astra/Fable reviews completed. Current-head CI is pending. The preceding published head's original CPU and GPU runs passed.

Earlier Qwen3-8B throughput deltas remain −12.1% at 4×1K and −4.5% at 4×4K; this update makes no new performance claim. An inherited exchange-allocation failure path remains under investigation. Draft for API review and further size reductions.

@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 18, 2026 00:36 — with GitHub Actions Failure
@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 18, 2026 01:04 — with GitHub Actions Error
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 18, 2026 01:24 — with GitHub Actions Active
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 19, 2026 21:25 — with GitHub Actions Active
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 20, 2026 02:55 — with GitHub Actions Active
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 27, 2026 00:31 — with GitHub Actions Active
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 27, 2026 01:10 — with GitHub Actions Active
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 27, 2026 01:53 — with GitHub Actions Active
Rely on the existing finalizer lock and preserve the guarded FIFO check, cleanup outcomes, and retry protocol. Remove the redundant transient map/wait protocol and assert lock release in the existing retry control.
Preserve all trainer-v1 and R44 blobs while inheriting the disjoint trajectory golden harness and staged tokenizer decomposition. Deterministic golden controls and incoming-path statics pass.
Recover retained native temporary files without rolling back the finished outcome or replaying serialization. Keep invalid terminal states and finish-after-abort rejected. Cover cleanup retries, terminal guards, committed bytes, and asymmetric finalization state.
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 27, 2026 03:31 — with GitHub Actions Active
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 28, 2026 09:32 — with GitHub Actions Active
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 28, 2026 11:08 — with GitHub Actions Active
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 28, 2026 12:05 — with GitHub Actions Active
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 28, 2026 13:53 — with GitHub Actions Active

This branch was successfully deployed

1 active deployment
trainer-rank-gpu-validation — b745db5a Deployed Sep 28, 2026 by bradhilton via Run on 2x H200 #909
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant