feat(driver): a cold suspend produces the portable workspace checkpoint, and a resume can restore from it - #104
Merged
Merged
Conversation
…end and resume The workspace is an ext4 image the guest mounts; the checkpoint library needs a file tree. The host must never mount or parse a tenant's ext4 image, so the guest streams a tar of /workspace over the relay conn it already has, and the host feeds that stream to the checkpoint writer as an fs.FS. The note argues the six options (host loop-mount, host userspace ext4 reader, guest-streamed tar, opaque image blob, network filesystem, block diff), states the handshake and its new budgets, the limits the host enforces on the stream, what every failure leaves behind, the restore path through mkfs.ext4 -d, the self-hosted store and key, and what rainier-cloud owes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A cold suspend ends the VM, and the only copy of the user's work that survives it is the one the guest puts on the wire. sessiond now writes the workspace as a stream on the relay conn it already has, after the flush, the exec kill and the agent-home unmount, and before it answers suspend_ready — because the host may not terminate the VM until the tree is out of it. - internal/relay: FrameStream (type 5), carrying the suspend nonce as its stream id so a stream from an abandoned suspend is routed to nothing; ControlSender.SendStream, which shares the control writer so the end marker cannot overtake the chunk it ends; the workspace_end and checkpoint_committed kinds, and the two counts the end marker carries. - internal/wstream: the stream format and its write side. A magic line, an index of (kind, name) in fs.WalkDir order, then a tar of the same entries in the same order. The index exists because fs.WalkDir is not streamable and buffering entry CONTENT to paper over that is what this design refuses. The guest walks twice rather than buffering the index, and a digest over both passes catches a workspace that changed under it. - cmd/sessiond: the streamer, a chunk writer that holds exactly one chunk, and the end marker — counts on success, a stage_failed-shaped stage on failure, and never a path in either. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The host's end of the workspace stream: a cursor over a one-way stream, presented as the fs.FS checkpoint.Writer walks. The design decision, stated where it is made: fs.WalkDir is not streamable — it asks for every child of a directory before it visits the first of them — so the stream carries an index of (kind, name) ahead of the tar, and the host buffers that shape and never a file byte. Everything else about an entry comes from its tar header when the cursor arrives at it, so each fact has one source. The index is required to be in strict fs.WalkDir order (depth-first, lexical, parents before children, no duplicates), which is checked at the head of the stream rather than discovered half way through a checkpoint; the cursor only ever moves forward, and skipping an excluded entry's body is free. Every limit the design note names is enforced here, on the untrusted end's own numbers: entries, index bytes, one entry's bytes, the whole stream's bytes, names, symlink containment, entry kinds, mode bits, sparse entries, and the index and tar agreeing about what each entry is. Tests cover each limit and each malformed shape, the cursor's refusal to be read backwards, a truncated body, and the claim the package exists for: a workspace streamed out of a guest is checkpointed by the real library, verified, restored, and compared entry for entry against the tree it came from. FuzzFS runs the parser against arbitrary bytes and requires every refusal to be one of this package's own sentinels. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Suspend(warm=false) on the microVM driver now sends the cold notice, reads the workspace the guest streams back, runs the checkpoint writer against it into a blob store, waits for the manifest's put-if-absent, verifies the committed bytes by reading them back, tells the guest, and only then terminates the VM, discards the rootfs and releases the slot. That is ADR-0003 §4.4. Any failure before the commit returns with the VM running and the error naming the STAGE — config, stream, write, verify — and never a path, a name or a value. runnerd's settleFailedColdSuspend already rolls such an entry back to running. A failure after the commit still fails the suspend (a checkpoint that cannot be read back is not one) and leaves the committed objects for the sweeper rather than deleting a tenant's only copy. - internal/relay: the Hub routes FrameStream to a registered sink, keyed by the suspend nonce. One reader or none; a chunk nobody asked for is dropped. The sink runs on the read loop, which is what back-pressures the guest when the store is slow — so a workspace is never buffered on this host. - internal/runnerd: StreamWorkspace runs the handshake and the copy under three budgets (30 s to the first chunk, 60 s between chunks, 30 min total), checks the guest's byte count against what arrived, and refuses everything else. CheckpointCommitted is the best-effort "your work is durable" event. - internal/driver: MicrovmOpts.Checkpoint (store, keys, key ref, prefix, limits), the barrier itself, the generation and manifest key on the instance record, and a claim that refuses two concurrent cold suspends of one instance. A driver with no checkpoint configuration behaves exactly as before. - runnerd sends no cold notice of its own when the driver checkpoints: the notice and the stream are one handshake with one nonce. Tests: the happy path (two objects, a verified restore test, the record, the VM stopped only afterwards); the store refusing; the guest dying mid-stream; the handshake failing; verify failing; the VM still running and the rootfs still there in each; the generation advancing only on a commit; two concurrent suspends refused; and a walk of the whole state directory proving no plaintext workspace byte reached host disk. The KVM harness gained the same handshake, so a real host exercises the whole barrier. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ckpoint The deep-dormant tier of ADR-0003 §2.3: the workspace image has been deleted and the checkpoint is the only copy of the session's work left. A cold Resume that finds the image ABSENT — and only then — restores the newest committed checkpoint into a 0700 scratch directory under the state dir, gives the tree to the user the guest runs as, builds a fresh ext4 from it with `mkfs.ext4 -d` through the DiskFormatter seam, removes the scratch directory, attaches and boots. The host mounts nothing at either end of the session's life. It runs before the token mint, the slot and the rootfs clone, so a resume that cannot produce a workspace spends nothing finding that out. No image and no checkpoint is a refusal in as many words: an empty workspace that looks like a successful resume is the failure a person discovers by finding their work gone. The ownership step is the one thing a fake cannot show and a real host cannot do without. A checkpoint records modes and not owners — a uid is not portable and the format is — so a restored tree belongs to root, while the agent runs as the session image's user and the guest's /init chowns the mount point only when it is empty. CheckpointOpts.OwnerUID/OwnerGID carry it; 0,0 means "leave it". RemoveWorkspace still removes the image only: checkpoint deletion is a retention decision this driver does not have the policy for. Tests: the restore taken only when the image is absent (counted at the formatter, not inferred), the tree coming back entry for entry, the scratch directory gone on every path, a tampered checkpoint refusing and leaving no image behind, a failed mkfs leaving no half-built one, no checkpoint and no store each refusing clearly, and RemoveWorkspace keeping the checkpoint. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
--checkpoint-store-dir and --checkpoint-key-file, both required when --driver=microvm and both Fatal when they are wrong, for the reason the kernel and the rootfs are: a runner that started without them would accept placements and then fail every cold suspend, with a tenant's work still inside a VM the durability barrier will not let it terminate. DirBlobStore is a checkpoint.BlobStore in a directory. Its one load-bearing property is put-if-absent — the format's atomic commit IS that conditional write — so it writes a temporary file in the same directory, fsyncs it, and link(2)s it into place: link fails with EEXIST where rename would silently replace a committed manifest, which is the one operation this format has no answer for. Both fsyncs are the durability half of §4.4. A failed write leaves nothing at the key and nothing beside it. Keys are validated element by element before they become a path, because the last hop before os.Remove trusts nobody. LoadCheckpointKey takes 64 hex characters or 32 raw bytes, and refuses a file that is missing, not regular, readable by group or other, the wrong length, or all zeros — without quoting a byte of it in any refusal. --checkpoint-restore-uid/-gid (default 1000) carry the ownership a restored workspace is given before it becomes a filesystem. Tests: the round trip, the never-overwrites rule, eight concurrent writers with exactly one winner, a failed write leaving nothing, key-path traversal, the directory modes, a real checkpoint written and verified through the store, every key-file refusal, and the flag wiring with the sentence each refusal owes an operator. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The checkpoint library reduces a file system's error to a CATEGORY rather than wrapping it — tree.go's fsCategory exists so that no path can travel out in an error — so a stream that failed under its walk arrived at the barrier as ErrSource with the sentence flattened into text, and a sandbox that had gone away was reported as a failed write. The cause is now chosen between the two errors the barrier has: the stream's own error wins when the local one is about the bytes it was given, and a store that refused still names the write stage, because there the stream failed only BECAUSE this end stopped taking bytes. The mid-stream test now cuts the guest off in both places one can die: inside the index, where the tree is not described yet, and inside the bodies. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
An independent adversarial review of the branch found four things that would have lost or exposed a tenant's work, and several smaller ones. All are fixed here, each with the test that would have caught it. THE BARRIER - A committed-but-unverified checkpoint bricked every later cold suspend of that session. A manifest key has no attempt suffix and put-if-absent never overwrites, so a generation whose manifest committed is spent — but the record only advanced after Verify, so every retry recomputed the same spent generation, lost its manifest put to ErrExists, and left another orphan content object. The generation is now recorded AND persisted the moment Write returns, and the driver probes for a free generation before it streams, so a record behind the store heals itself rather than trapping the session. - The generation was only persisted by Suspend's own saveRecord, with two returns in between; a restart in that window restored an older checkpoint than the one that existed. PLAINTEXT - `mkfs.ext4 -d` reports a per-file failure by naming the file it was copying, and that output went into an error that reaches a session's error column. It is now counted, not quoted. - RemoveAll's and the chown walk's errors are *fs.PathErrors naming an entry inside the workspace; both are reduced to their cause now. - A crash mid-restore left a tenant's whole workspace in plaintext under the state directory with nothing to sweep it. The driver now clears <state>/restore at startup, beside the other reclaims. THE RESTORE - The image was created at its FINAL path and populated over minutes, and the restore's trigger is that path's absence — so a crash mid-mkfs left a file the next resume would see, skip the restore for, and hand the guest unformatted, with the checkpoint still in the store and nothing left to consult it. It is built at `.partial` and renamed. - The 1 MiB trailer bound turned a host-side-only exclusion into a permanently failing suspend; the drain now runs under the stream's own MaxTotalBytes. THE STREAM AND THE STORE - wstream.FS.ReadLink trusted seek's backwards guard and had none of its own: a caller that asked about a passed entry would have been handed the CURRENT entry's target, checkpointing a symlink with somebody else's destination. - A symlink with a body is refused explicitly rather than by relying on archive/tar's classing of it as header-only. - MaxIndexBytes is charged what the index COSTS (~96 bytes per entry beyond the name) rather than only what it reads, and raised to 64 MiB so the bound is honest and the entry ceiling is still reachable. Both ends charge it the same. - A session id that this driver accepts and checkpoint.Context refuses now fails the CREATE rather than every later cold suspend. - The key file must be owned by the runner, the store directory is tightened if it was pre-made group-readable, and a sandbox-supplied stage and tail are bounded before they reach a runner's error. TESTS FOR GUARDS THAT HAD NONE The review checked six guards by reverting them; three had no test behind them. Now: the ownership step is covered at its call site (a uid the test cannot give a file to, which must fail closed and leave no image), the host's own exclusions are covered by a checkpoint that must not carry them, the put-if-absent is covered deterministically (a winner committing DURING the loser's write, which os.Rename would not refuse), and the cursor is covered for readlink, stat and open. Each was confirmed by reverting the guard and watching the new test fail. The design note carries the corrected arithmetic, the generation rule, the startup sweep, and two open questions the review surfaced: nothing bounds how many concurrent checkpoints spend that memory, and a session whose sandbox is gone can never be cold-parked. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
jiashuoz
marked this pull request as ready for review
September 23, 2026 14:43
**The deep-dormant restore after a runnerd restart.** The review is right that `ensureWorkspaceForResume` sits behind the `!bootLive` refusal and that a session dormant long enough to have its disk deleted has almost certainly outlived the process — so in production the restore path is behind that refusal. Reordering the two does not fix it and is not faked here: a recovered record has no guest configuration (ADR-0003 §2.7 item 1 keeps it off host disk), so restoring first would spend minutes of copying and a plaintext scratch tree on a resume that is certain to fail a moment later, and would widen the exposure window §7 bounds. What the driver can do, and now does, is tell the two apart at the cost of one Stat and no tenant bytes: `coldResumeNeedsGuestConfig` names whether the workspace image is also gone and which committed generation holds the only copy of the work, so "re-dispatch through the control plane" and "this person's work is only in checkpoint generation 4" are not the same sentence. The precondition is written down in the note's §7 and the create-shaped resume that lifts it is now owed in §9. **Error sanitization.** `scratchDir`'s `RemoveAll` fails on an entry inside a previous restore of this workspace, so its `*fs.PathError` names a tenant's file; it goes through `pathFreeCause` like the deferred removal below it. **The untested guards.** `reclaimRestoreScratch` (a planted tree from a previous run is gone after the next start, and its bytes are nowhere else under the state directory), `checkKeyOwner` (a foreign uid through a faked `fs.FileInfo`, which is the only way to exercise a check a test has no privilege to set up), and `clampGuestText`. Each was confirmed by reverting the guard and watching the new test fail. **Docs.** The two required checkpoint flags are on the startup-refusal list in `docs/microvm-host-privileges.md`. In the note: §4.4's quote carries an ellipsis where "to regional GCS" was dropped; §6 says that the before-termination order is a deliberate inversion of §4.4's stated order, forced by §4.5's rule that the host never parses tenant ext4, with the consequence that a host dying mid-checkpoint cannot re-checkpoint that workspace afterwards (rainier-cloud is amending ADR-0003 to match); and §9 gains the host-wide checkpoint concurrency bound, the unowned disk-deletion freshness guard, the fact that `CheckpointAt` is a commit time with nothing distinguishing committed from verified, and `DirBlobStore`'s missing deletion path. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
A cold suspend of a microVM session now turns the session's workspace into a
committed, verified portable workspace checkpoint before the VM is
terminated, and a resume whose workspace image is gone rebuilds it from that
checkpoint. That is ADR-0003 §4.4's durability barrier and §2.3's deep-dormant
tier, on top of the format and library from #99.
The decision this implements, and the reason everything else follows from it:
the workspace is an ext4 image the guest mounts, the checkpoint library
takes an
fs.FS, and the host must never mount or parse a tenant's ext4image. So the guest streams a tar of
/workspaceover the vsock relayconnection it already has, and the host feeds that stream to the checkpoint
writer as a tar-backed
fs.FS. Nothing plaintext reaches host disk on thesuspend path: the bytes go from a socket, through a 32 KiB copy buffer, into an
AES-256-GCM frame, into a blob store.
Design note:
docs/design/2026-09-23-workspace-checkpoint-wiring.md.The commits
docs(design)— the note: the six options (host loop-mount, hostuserspace ext4 reader, guest-streamed tar, opaque image blob, network
filesystem, block diff) and why guest-streamed wins; the handshake, its
frames and its new budgets; the limits the host enforces; what each failure
leaves behind; the restore path; the self-hosted store and key; what
rainier-cloud owes.
feat(sessiond)—relay.FrameStream(type 5) carrying the suspendnonce as its stream id,
ControlSender.SendStreamsharing the controlwriter, the
workspace_endandcheckpoint_committedkinds; theinternal/wstreamformat and its write side; and the cold handler streaming/workspacebefore it answerssuspend_ready.feat(wstream)— the tar-backedfs.FSthe checkpoint writer walks: aone-way cursor over the stream, every limit, the walk-order checks, a round
trip through the real library (write → verify → restore → compare), and
FuzzFS.feat(driver)— the barrier:Suspend(warm=false)sends the notice,reads the stream, writes the checkpoint, waits for the manifest commit,
verifies, tells the guest, and only then terminates the VM.
feat(driver)— the deep-dormant resume: restore into a 0700 scratchdirectory, give the tree to the guest's user,
mkfs.ext4 -d, rename intoplace, boot.
feat(runnerd)—--checkpoint-store-dirand--checkpoint-key-file(plus
--checkpoint-restore-uid/-gid), a local-directory blob store whoseput-if-absent is
link(2), and a key loader that fails closed.Plus two fix commits from the review below.
The handshake
#98's cold suspend was 2 s + 30 s, which bought a flush, a ten-second exec kill
and an unmount; it cannot also buy a 10 GiB copy. The budgets are now a shape
rather than one stopwatch — 30 s to the first chunk, 60 s of progress between
chunks, 30 min total, 30 s for the ready, 10 min for the verify — worst case
≈ 41 minutes and only for a session that is genuinely stuck. The idle budget
covers this host's slowness too: the stream is back-pressured by the store
through an
io.Pipe, so a store that stopped taking bytes fails the suspend,which keeps the VM.
The stream is not only a tar.
fs.WalkDirasks for every child of adirectory before it visits the first one, and a child's subtree sits between it
and its next sibling — so a plain tar is not walkable without buffering content,
which is the one thing this design refuses. The stream is therefore a magic
line, an index of (kind, name) in walk order, then the tar. The index is the
tree's shape only; every other attribute comes from the tar header when the
cursor arrives, so each fact has one source.
The limits the host enforces
.././ NUL in a namefs.WalkDirorder, no duplicatesAll configurable (
wstream.Limits), all applied by the guest and by the host,because the host must not be the hop that trusts the guest — and re-applied
again by
checkpointitself on the way in and on the way out.What only a real host can verify
Everything here runs against the simulated engine, the pipe-backed fake
transport,
fstest.MapFS/temp directories and the in-memory blob store. Fourthings are consequently claimed and not evidenced:
the same handshake, so
RAINIER_MICROVM_KVM_TEST=1on a real KVM hostexercises a real sessiond streaming a real ext4 workspace into a real
checkpoint. It has never run here.
mkfs.ext4 -don a restore. The simulated formatter writes the tree as atar so a test can read back what the restore produced; only a real host runs
mke2fs. Two specifics to watch there: the default inode count for a 10 GiB
image (~655,000, above our entry ceiling — a tree past it fails the mkfs), and
the ownership step, which needs a runner running as root to give files to uid
1000.
~6 MiB/s, from arithmetic and not from a measurement of vsock plus a store.
create a file owned by somebody else.
What rainier-cloud owes
checkpoint.BlobStore(
x-goog-if-generation-match: 0, resumable upload for the content object).concrete key version, plus the regional key-readiness rule
Preflightis for.Specanduses the workspace volume name, so a cell carrying the control plane's id
should pass that instead.
sweep for objects whose manifest lost a race or failed its verify, and
deletion (tenancy §14.2).
The review
An independent adversarial review of the whole diff (a subagent on opus, given
the diff, the design note, ADR §4.4 and the tenancy rows) was asked for five
specific things: plaintext reaching host disk, a barrier that reports success
before the commit, tar-limit bypasses, a VM terminated on a failure path, and
vacuous tests (revert each guard, see what fails). It found eleven real problems
and confirmed the rest. Every finding is fixed, each with the test that would
have caught it (
deca963,d0e344c):High. (1) A committed-but-unverified checkpoint bricked every later cold
suspend of that session: a manifest key has no attempt suffix and
put-if-absent never overwrites, so a spent generation that the record never
advanced past made every retry lose its manifest put to
ErrExists. Thegeneration is now recorded and persisted the moment
Writereturns, and thedriver probes for a free generation before it streams, so a record behind the
store heals itself. (2) The generation was persisted only by
Suspend's ownsaveRecord, with two returns in between. (3)mkfs.ext4 -d's output names thefiles it was copying and was going into an error that reaches a session's error
column — counted now, not quoted. (4)
RemoveAll's error is an*fs.PathErrornaming an entry inside the workspace.
Medium. (5) The chown walk's error, same shape. (6) A crash mid-restore left
a tenant's whole workspace in plaintext under the state directory with nothing
to sweep it — the driver now clears
<state>/restoreat startup. (7) The imagewas created at its final path and populated over minutes, while the
restore's trigger is that path's absence: a crash there left a file the next
resume would see, skip the restore for, and hand the guest unformatted. Built at
.partialand renamed now. (8) Accepted, not fixed, and now an open questionin the note: a session whose sandbox is gone can never be cold-parked and pins
a slot. The narrower rule — park without a checkpoint, gate only the
deep-dormant disk deletion on one — is what §4.4 actually asks for and is a
change to what the runner records rather than to this barrier. (9) The 1 MiB
trailer bound turned a host-side-only exclusion into a permanently failing
suspend; the drain runs under the stream's own
MaxTotalBytesnow. (10) Theindex's memory was understated and unbounded across concurrent suspends: it is
now charged what it costs, and the concurrency bound is an open question with
the arithmetic written down. (11) A session id this driver accepts and
checkpoint.Contextrefuses failed every cold suspend; it fails the create now.Low. Key-file ownership, a pre-existing store directory's mode, a symlink
with a body, a stale comment, and an unbounded guest-supplied tail.
Vacuous tests. The review reverted six guards. Three had no test behind them:
the ownership step's call site (covered now by a uid the test cannot give a
file to, which must fail closed and leave no image), the driver's own
exclusions (covered by a checkpoint that must not carry them), and the
put-if-absent (covered deterministically by a winner committing during the
loser's write —
os.Renamepasses the sequential test and fails this one). Italso found that
wstream.FS.ReadLinktrusted the cursor's backwards guard andhad none of its own, so a caller asking about a passed entry would have been
handed the current entry's target: a symlink checkpointed with somebody
else's destination, verified and wrong. Fixed, with tests for readlink, stat and
open. Each fix was confirmed by reverting the guard and watching the new test
fail.
Review round 2
A second pass over this branch raised four things. What changed, and where it
did not:
ensureWorkspaceForResumesits behind the!bootLiverefusal inResume,and a session dormant long enough for its disk to be deleted has almost
certainly outlived the runnerd that created it — so in production the
restore path is behind that refusal. It is not reordered, and the reason
is not ordering. A recovered record has no guest configuration at all:
ADR-0003 §2.7 item 1 keeps the session id, the proxy, the boot chain and the
secret names off a shared host's disk, and this driver cannot invent them.
Restoring first would spend minutes of copying and leave a tenant's whole
workspace in plaintext under
<state>/restorefor a resume that is certainto fail a moment later — widening the very window §7 bounds — and then still
refuse. So the refusal stays first, and what it now does is tell the two
cases apart at the cost of one
Statand no tenant bytes(
internal/driver/microvm_checkpoint.go:477coldResumeNeedsGuestConfigand
deepDormantNote): whether the workspace image is also gone, andwhich committed generation holds the only copy of the work, or that no
checkpoint was ever committed and nothing can recover it. "Re-dispatch this
session through the control plane" and "this person's work exists only as
checkpoint generation 4" are different operator actions and are now
different sentences. The precondition — the restore is reachable only from
a record this process still holds the guest configuration for — is stated
in the note's §7, and the create-shaped resume that lifts it (the
control plane re-resolving the configuration) is now listed among what
rainier-cloud owes, in §9. Tests:
TestAColdResumeAfterARestartSaysWhichThingItIsWaitingForandTestAColdResumeAfterARestartWithNoCheckpointSaysThatToo(
internal/driver/microvm_restore_test.go:615and:666) build a second driver overthe state directory a first one left, resume the recovered record, and
assert the refusal names both halves — and that nothing was spent on it:
no
mkfs, no scratch tree, nothing at the workspace image path.scratchDir'sos.RemoveAllfails on an entryinside a previous restore of this workspace, so the
*fs.PathErroritreturns names a tenant's file, and this error reaches a session's error
column through
restoreErr. It goes throughpathFreeCausenow, like thedeferred removal ten lines below it
(
internal/driver/microvm_checkpoint.go:699). Test:TestTheScratchDirectorysRemovalErrorNamesNoPath(
internal/driver/microvm_restore_test.go:761) plants a directory thisprocess cannot clear and asserts the error carries neither the file name,
nor the state directory, nor a separator.
reclaimRestoreScratch:TestAPreviousRunsRestoreScratchIsRemovedAtStartup(
internal/driver/microvm_restore_test.go:706) plants<state>/restore/i-old/src/fbefore the driver is constructed and assertsit is gone after — and that its bytes are nowhere else under the state
directory.
checkKeyOwner:internal/driver/checkpointkey_unix_test.gofakes the
fs.FileInfo(a foreign uid in its*syscall.Stat_t), which isthe only way to exercise a check a test has no privilege to set up; a
matching uid and a filesystem that reports no ownership are both accepted,
so the refusal is about ownership and not about every file.
clampGuestText:internal/runnerd/workspacestream_test.go:355— thebound, the boundary, the truncation notice, and a 1 MiB guest string coming
back bounded. Each of the three was confirmed by reverting the guard and
watching the new test fail.
docs/microvm-host-privileges.mdcarries--checkpoint-store-dirand
--checkpoint-key-fileon the startup-refusal list, with a note thatthese two are checked in
cmd/runnerdrather than generated fromMicrovmHostRequirements()'s table. In the design note: §4.4's quote nowcarries an ellipsis where "to regional GCS" was dropped, and says what
was elided and why; §6 no longer calls the before-termination order
"stricter than §4.4 needs" — it is a deliberate inversion of §4.4's
stated order (checkpoint from the detached disk, after termination), forced
by §4.5's rule that the host never mounts or parses a tenant's ext4 image,
with the consequence stated plainly that a host dying mid-checkpoint
cannot re-checkpoint that workspace afterwards (rainier-cloud is amending
ADR-0003 §4.1, §4.3 and §4.4 to match); and §9 gains four things this change
leaves undone — a host-wide bound on concurrent checkpoints (the index
is
O(entries)per in-flight suspend, capped at 64 MiB, times the host'sslots: ~2 GB for 32), the "verified checkpoint newer than the disk's last
write" guard before deep-dormant disk deletion, which is owned by nobody
today and belongs in the
cell-workerpath that callsRemoveWorkspace,the fact that the record's
CheckpointAtis a commit time with nothingon the record distinguishing committed from verified, and that
DirBlobStorehas no deletion path at all.Gates
make verify; the gate set (./internal/driver/... ./cmd/sessiond/... ./internal/runnerd/... ./internal/relay/... ./checkpoint/...) under-race -count=1three times, each with the full top-level count and no panic;go test ./... -race -count=1;go vet ./...;GOOS=linux go build ./... && GOOS=linux go vet ./internal/driver/... ./cmd/...;git diff --check; a 30 sFuzzFSrun — all clean.
One thing worth recording rather than quietly re-running away: on the first
full
go test ./... -raceafter the review fixes,TestExecResizeAndSignalReachTheProcessininternal/sandboxexecfailed on a5 s timeout under the load of the whole suite on this 4-core box, then passed on
a targeted
-count=3re-run, on a full run of that package, and on a secondfull
./... -racerun. That package's only contact with this branch isrelay.Execer/relay.ExecAttachment, neither of which this change touches.🤖 Generated with Claude Code