Add versioned trainer graphs and logical-rank APIs - #925
Draft
bradhilton wants to merge 101 commits into
Draft
bradhilton wants to merge 101 commits into
bradhilton wants to merge 101 commits into
Conversation
bradhilton
had a problem deploying
to
trainer-rank-gpu-validation
September 18, 2026 00:36 — with
GitHub Actions
Failure
bradhilton
had a problem deploying
to
trainer-rank-gpu-validation
September 18, 2026 01:04 — with
GitHub Actions
Error
bradhilton
deployed
to
trainer-rank-gpu-validation
September 18, 2026 01:24 — with
GitHub Actions
Active
…view-fixes-merge-20260919
bradhilton
deployed
to
trainer-rank-gpu-validation
September 19, 2026 21:25 — with
GitHub Actions
Active
bradhilton
deployed
to
trainer-rank-gpu-validation
September 20, 2026 02:55 — with
GitHub Actions
Active
bradhilton
deployed
to
trainer-rank-gpu-validation
September 27, 2026 00:31 — with
GitHub Actions
Active
bradhilton
deployed
to
trainer-rank-gpu-validation
September 27, 2026 01:10 — with
GitHub Actions
Active
bradhilton
deployed
to
trainer-rank-gpu-validation
September 27, 2026 01:53 — with
GitHub Actions
Active
Rely on the existing finalizer lock and preserve the guarded FIFO check, cleanup outcomes, and retry protocol. Remove the redundant transient map/wait protocol and assert lock release in the existing retry control.
Preserve all trainer-v1 and R44 blobs while inheriting the disjoint trajectory golden harness and staged tokenizer decomposition. Deterministic golden controls and incoming-path statics pass.
Recover retained native temporary files without rolling back the finished outcome or replaying serialization. Keep invalid terminal states and finish-after-abort rejected. Cover cleanup retries, terminal guards, committed bytes, and asymmetric finalization state.
bradhilton
deployed
to
trainer-rank-gpu-validation
September 27, 2026 03:31 — with
GitHub Actions
Active
bradhilton
deployed
to
trainer-rank-gpu-validation
September 28, 2026 09:32 — with
GitHub Actions
Active
bradhilton
deployed
to
trainer-rank-gpu-validation
September 28, 2026 11:08 — with
GitHub Actions
Active
bradhilton
deployed
to
trainer-rank-gpu-validation
September 28, 2026 12:05 — with
GitHub Actions
Active
bradhilton
deployed
to
trainer-rank-gpu-validation
September 28, 2026 13:53 — with
GitHub Actions
Active
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds trainer planning for retained/offloaded/recomputed backward state, bounded gradient staleness, managed output tensors, and checkpoint-owned live parameters/modules. Rank-zero and logical-DP facades support the paired Caladan client API.
Current head
b745db5aincludes reviewed size reductions, fixed-main reconciliation, a padded GPT-OSS export correction, and coordinated handling of distributed export-preparation failures. Independent tensor/file and two-rank CPU controls passed. Focused CPU checks and statics passed; root and fresh Astra/Fable reviews completed. Current-head CI is pending. The preceding published head's original CPU and GPU runs passed.Earlier Qwen3-8B throughput deltas remain −12.1% at 4×1K and −4.5% at 4×4K; this update makes no new performance claim. An inherited exchange-allocation failure path remains under investigation. Draft for API review and further size reductions.