Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
18 commits
Select commit Hold shift + click to select a range
21bfd95
The qboot experiment: a pinned build, and a probe that asserts before…
aledbf Sep 26, 2026
1d56d76
boot: where the kernel's own boot goes, initcall by initcall
aledbf Sep 26, 2026
301fe71
boot: two probes for early userspace, and three ways it loses data
aledbf Sep 26, 2026
ef96390
boot: what udev actually walks, and why there is no single win in it
aledbf Sep 26, 2026
3819a79
boot: udev is not on the critical path, and three rankings said it was
aledbf Sep 26, 2026
2e5b35a
boot: the initcall probe uses --profile instead of redefining it
aledbf Sep 26, 2026
4dc5053
boot: clarify that releases do not include software that owns a guest
aledbf Sep 26, 2026
e8c4c27
kernel: a floor, and what this machine's kernel configuration actuall…
aledbf Sep 26, 2026
e0254c5
boot: keep the conclusions near the code, not the sample tables
aledbf Sep 26, 2026
f85ae64
kernel: erofs goes; nothing has mounted one since qcow2 replaced that…
aledbf Sep 26, 2026
7af74bc
boot: measure the console swap from inside, and eliminate a third can…
aledbf Sep 26, 2026
662a170
image: TERM=dumb on the console unit, which is worth 331 ms
aledbf Sep 26, 2026
74dfd72
kernel: KSM and THP are not the remaining candidate either
aledbf Sep 26, 2026
c3582c1
boot: the number of unit files is not the boot, and the caches are al…
aledbf Sep 26, 2026
ea8e086
boot: ask systemd what unit loading costs, and it is 26 ms not 212
aledbf Sep 26, 2026
bfe9b17
boot: disable unnecessary systemd generators and virtual console supp…
aledbf Sep 26, 2026
19b6a5d
kernel: CONFIG_VT off as dead configuration, and a number withdrawn
aledbf Sep 26, 2026
d260061
boot: the two largest items in QEMU's startup are cold-boot only
aledbf Sep 26, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
83 changes: 55 additions & 28 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,29 +5,36 @@
A virtual machine: QEMU, a guest kernel, a base image, and the definition of the machine
they make. Four artefacts and one version.

**What it is not: anything that runs inside a guest.** No container runtime, no agent, no
supervisor, no RPC, and no init. A release is not bootable on its own and that is deliberate
— whoever runs guests brings the init. There is no exception; there was one, a debug
initramfs, and it was a second init doing what the consumer's already does. It went on
2026-09-10, when udev was turned back on and the machine reached a login prompt through
`root=/dev/vda init=/sbin/init` without it.

**It does not know about the projects that consume it.** No repository names, no file paths
into other trees, no ADR numbers. If a rationale can only be stated by naming a consumer,
it is either the wrong rationale or the wrong repository for the code. Say what is true of
the machine.
**A release does not carry software that owns a guest.** No container runtime, agent,
supervisor, RPC service, or init is part of the published machine. A release is not bootable
on its own: whoever runs a guest supplies its init.

Diagnostic input is different from release content. A test or experiment may use a caller
supplied initrd, a temporary guest helper, a disposable overlay, or a separately built
firmware image when it is isolated from `task build` and `task release`. State clearly what
is diagnostic, who supplies it, and what it is measuring. The qboot probe is the model:
it does not change a release tree and compares both variants with the same diagnostic initrd.

**Keep consumer-specific implementation out of this tree, but record real contracts.** Do
not copy consumer code, paths, or an ADR as a substitute for an explanation. It is correct to
name an external component when its protocol, provisioned file, lifecycle, or compatibility
contract affects this machine; describe the contract here and test the observable behaviour
where possible.

## The three rules that are load-bearing

Everything else is style. These are correctness, and each fails silently.
These are release invariants, not defaults. Everything below them is a strong engineering
default unless it says otherwise; an experiment may depart from a default when it is isolated
and makes its scope and result clear.

1. **A release is one machine.** `machine.Spec.Fingerprint` hashes the QEMU binary, the
kernel and the initrd by content, together with the four arguments that decide the
machine's shape. Two machines with the same fingerprint may exchange templates; two
without may not, and a restore across them is undefined rather than an error. Any change
to any of those three files invalidates every template in existence. That is the design,
not a bug to work around — but it means "I only changed a comment in `kernel/Dockerfile`"
is a fleet-wide event, and it has already happened once in this repository's history.
without may not, and a restore across them is undefined rather than an error. A content
change to any of those three files invalidates every template in existence. That is the
design, not a bug to work around. A source or recipe change is a possible fleet-wide event; whether
it is one is decided by the resulting artifact hashes and shape, not by the diff's apparent
size. Verify them before promotion.

2. **The base image is never written to.** Every VM maps it read-only through a qcow2
backing chain and many share one file. Anything that opens it for writing invalidates
Expand Down Expand Up @@ -59,6 +66,25 @@ what it makes clear.
systemd unit written by a Go program at run time was the wrong answer to a real problem;
the unit belongs in `image/`, and only the symlink that enables it belongs in the boot.

### Experiments

Experiments are welcome when they make a production decision cheaper to evaluate. They must
not silently become production behaviour.

- Keep experiment outputs, overlays, initrds, firmware, and caches separate from release
outputs. Do not modify a shared base image; use a disposable overlay or a new artifact.
- A probe may change one production default at a time, including a kernel option, unit,
QEMU sandbox setting, CPU affinity, or diagnostic initrd. Production defaults remain the
safe settings until the result is promoted deliberately.
- Record the baseline, environment, sample count, and the measured result. A fast exploratory
run is enough to decide what to investigate; a promotion needs a reproducible comparison
and the functional check that could reveal a false win.
- If a promoted change alters what a restored guest sees, update the machine identity and
prove that templates do not cross the boundary. If it is host-only (for example launcher
CPU affinity), document why it does not.
- Remove or quarantine an experiment once it no longer has an active question. Keep the
conclusion and durable evidence, not a permanent switch in the release path.

## Go

- **Clarity beats cleverness.** This code is read far more often than it is written, and
Expand All @@ -82,13 +108,13 @@ what it makes clear.
Comments here are longer than usual, on purpose. The rule is what makes them worth reading:

- **Say why, not what.** The code says what. If a comment restates it, delete the comment.
- **Carry the evidence.** When something is the way it is because of a measurement, put the
number in: `87.6 ms to 51.1 ms, kernel to init`, `+356 ms and rejected`, `−3.7 ms of
boot`. A claim without a number invites someone to undo it on a hunch.
- **Record what was tried and failed.** The comments that have paid for themselves most
here are the ones naming a dead end: the drop-in that did not lift the device dependency,
the flag that broke a link twenty minutes into a build, the check that passed for every
input because a tool exits 0 on failure. Someone will otherwise try it again.
- **Carry durable evidence.** When a production decision rests on a measurement, put a
concise result and its context near the decision, or link to a maintained measurement
record. Include a number when it is the reason for the decision; do not turn an exploratory
observation into a permanent claim.
- **Record failed approaches selectively.** Keep a failed attempt near code only when it
prevents a plausible, costly regression or repeats a non-obvious tool failure. Put raw
samples and short-lived exploration in a measurement record, issue, or commit instead.
- **Date a fact that could go stale.** "measured 2026-09-07" tells a reader whether to
re-check.
- **Do not write down that it works.** The README had a "Status" section listing what had
Expand All @@ -97,14 +123,15 @@ Comments here are longer than usual, on purpose. The rule is what makes them wor
justified — so it could only rot, and it did. A claim about the state of the tree
belongs in a workflow, a badge, or a commit message. Never in prose that nobody
re-reads.
- **Do not name the projects that consume this.** See above.
- **Name external contracts when relevant.** See above.

## Artefacts leave this machine

QEMU is linked statically, and anything added beside it should be too. An artefact that
needs the host to have the right libraries is not one artefact, and the failure it produces
arrives on somebody else's machine, at start-up, naming a library rather than a decision
made here.
Host-executed binaries published in a release should be statically linked unless there is a
documented deployment reason not to be. This rule does not apply to guest userspace, build
tools, or diagnostic artifacts. A released host binary that depends on host libraries needs a
compatibility and packaging story, because its failure otherwise arrives on someone else's
machine at start-up.

## Verifying

Expand Down
7 changes: 7 additions & 0 deletions NOTICE
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,13 @@ Copyright The spin-stack Authors
This product includes software developed by the spin-stack project.
Licensed under the Apache License, Version 2.0; see LICENSE.

The patch in qemu/qboot/write-pointer.patch includes context from qboot,
https://github.com/bonzini/qboot, commit
8ca302e86d685fa05b16e2b208888243da319941, under GPL-2.0 (see qemu/qboot/COPYING).
It modifies fw_cfg.c, include/fw_cfg.h and tables.c to implement WRITE_POINTER,
DMA writes and error/bounds checks. This experimental firmware is not shipped
in releases. The Apache-2.0 licence does not apply to that patch or COPYING.

--------------------------------------------------------------------------------
Third-party software in a release, which this licence does NOT cover
--------------------------------------------------------------------------------
Expand Down
12 changes: 7 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,9 +65,11 @@ Each part's targets live beside what they build — `qemu/Taskfile.yml`, `kernel
`image/Taskfile.yml` — so `task qemu:build` is next to `qemu/Dockerfile`. The root
`Taskfile.yml` holds the vars every part reads and the targets that cross all of them.

**What is deliberately not here: the software that runs inside a guest.** This repository
builds a machine. It knows nothing about what boots on it, and a release is not bootable on
its own by design — whoever runs guests brings the initrd. There is no exception to that.
**What is deliberately not in a release: software that owns a guest.** This repository
builds a machine, and a release is not bootable on its own by design — whoever runs guests
brings the initrd. Tests and performance probes may use caller-supplied diagnostic initrds
or disposable guest helpers, but those are not release artifacts and must stay isolated from
the published machine.

## What it publishes

Expand Down Expand Up @@ -324,8 +326,8 @@ the tools a workspace expects, and boot optimizations that were each measured
`image/mkosi.extra/usr/local/lib/spin-base/optimize-systemd.sh` names the milliseconds every
mask saved, and is an unmodified copy for that reason.

**ext4**, because the guest kernel has `EXT4_FS`, `EROFS_FS` and `OVERLAY_FS` and explicitly
not `XFS_FS`, `BTRFS_FS` or `SQUASHFS`. Anything else starts with a kernel config change one
**ext4**, because the guest kernel has `EXT4_FS` and `OVERLAY_FS` and explicitly not
`XFS_FS`, `BTRFS_FS`, `SQUASHFS` or `EROFS_FS`. Anything else starts with a kernel config change one
directory over — which `kernel/Dockerfile` now asserts.

**Partitionless**, not a bootable disk with an ESP and a GPT, which nothing here would read.
Expand Down
64 changes: 64 additions & 0 deletions Taskfile.yml
Original file line number Diff line number Diff line change
Expand Up @@ -121,6 +121,13 @@ vars:
QEMU_CACHE_TO: '{{if eq .CACHE_BACKEND "gha"}}type=gha,scope=qemu,mode=max{{else}}type=local,dest={{.BUILDKIT_CACHE_DIR}}/qemu,mode=max,compression=zstd{{end}}'
KERNEL_CACHE_FROM: '{{if eq .CACHE_BACKEND "gha"}}type=gha,scope=kernel{{else}}type=local,src={{.BUILDKIT_CACHE_DIR}}/kernel{{end}}'
KERNEL_CACHE_TO: '{{if eq .CACHE_BACKEND "gha"}}type=gha,scope=kernel,mode=max{{else}}type=local,dest={{.BUILDKIT_CACHE_DIR}}/kernel,mode=max,compression=zstd{{end}}'
# The firmware experiment. Its own scope like everything else, and it is never built by
# `task build`: nothing in a release comes from it.
# The floor-kernel experiment, like QBOOT_ above: never built by `task build`.
KMIN_CACHE_FROM: '{{if eq .CACHE_BACKEND "gha"}}type=gha,scope=kmin{{else}}type=local,src={{.BUILDKIT_CACHE_DIR}}/kmin{{end}}'
KMIN_CACHE_TO: '{{if eq .CACHE_BACKEND "gha"}}type=gha,scope=kmin,mode=max{{else}}type=local,dest={{.BUILDKIT_CACHE_DIR}}/kmin,mode=max,compression=zstd{{end}}'
QBOOT_CACHE_FROM: '{{if eq .CACHE_BACKEND "gha"}}type=gha,scope=qboot{{else}}type=local,src={{.BUILDKIT_CACHE_DIR}}/qboot{{end}}'
QBOOT_CACHE_TO: '{{if eq .CACHE_BACKEND "gha"}}type=gha,scope=qboot,mode=max{{else}}type=local,dest={{.BUILDKIT_CACHE_DIR}}/qboot,mode=max,compression=zstd{{end}}'
E2FSPROGS_CACHE_FROM: '{{if eq .CACHE_BACKEND "gha"}}type=gha,scope=e2fsprogs{{else}}type=local,src={{.BUILDKIT_CACHE_DIR}}/e2fsprogs{{end}}'
E2FSPROGS_CACHE_TO: '{{if eq .CACHE_BACKEND "gha"}}type=gha,scope=e2fsprogs,mode=max{{else}}type=local,dest={{.BUILDKIT_CACHE_DIR}}/e2fsprogs,mode=max,compression=zstd{{end}}'
# The container that builds the base image. Small — mkosi, qemu-utils — and cached like
Expand Down Expand Up @@ -198,6 +205,63 @@ tasks:
cmds:
- SPIN_LOGIND_TEST=1 go test ./boot/ -run '^TestLogindSessions$' -count=1 -v -timeout 5m

boot:initcalls:
desc: >-
Where the kernel's own boot goes, initcall by initcall, as a p50 over REPS boots.
Needs KVM, a built release and sudo. It boots with the console silent and reads the
ring buffer afterwards, because initcall_debug on a serial console puts a VM exit
inside every interval it reports. TOP= how many rows, REPS= how many boots.
cmds:
- SPIN_INITCALL_PROBE=1 go test ./boot/ -run TestKernelInitcalls -count=1 -v -timeout 30m

boot:unitload:
desc: >-
What systemd spends loading units and building the initial transaction, from its own
instrumentation: `systemd --test` reports it at LOG_INFO where a real boot logs it at
LOG_DEBUG. Many runs inside one boot, so there is no boot-to-boot noise, and it
measures the generators by masking two that cannot do anything on this machine. Needs
KVM, a built release and sudo.
cmds:
- SPIN_UNITLOAD_PROBE=1 go test ./boot/ -run TestUnitLoadCost -count=1 -v -timeout 15m

boot:userspace:
desc: >-
systemd's own account of the boot: its kernel/userspace split, `blame` and the
critical chain, read after the boot rather than during it. Needs KVM, a built release
and sudo. Read a critical chain as what finished last, never as what was blocking —
see the "console no dev" row in boot/bench_test.go for what that mistake cost.
cmds:
- SPIN_USERSPACE_PROBE=1 go test ./boot/ -run TestUserspaceCost -count=1 -v -timeout 15m

boot:systemd:
desc: >-
Inside early userspace, in systemd's own words: debug logging to the journal, read
afterwards, ranked by the gaps in it. Needs KVM, a built release and sudo. It inflates
the phase it measures, so the absolute numbers are not `task boot:bench`'s — the shape
is what it is for. GAP= the floor in ms, TOP= how many gaps.
cmds:
- SPIN_SYSTEMD_DEBUG=1 go test ./boot/ -run TestSystemdDebug -count=1 -v -timeout 15m

boot:floor:
desc: >-
What a kernel costs before anything is configured into it: the release kernel against
the one `task kernel:minimal` builds, interleaved, on a marker both can reach. Needs
KVM, a built release, `task kernel:minimal`, and a diagnostic initrd in
SPIN_PROBE_INITRD=. The answer bounds every kernel-config change that could follow.
cmds:
- SPIN_FLOOR_PROBE=1 go test ./boot/ -run TestKernelFloor -count=1 -v -timeout 20m

boot:firmware:
desc: >-
Compare firmware for this machine: SeaBIOS against the qboot build in qemu/qboot/.
Needs KVM, a built release, `task qemu:qboot`, and a diagnostic initrd in
SPIN_PROBE_INITRD=. Without the firmware or the initrd it says which is missing and
skips: neither is an artefact a release carries. It asserts a firmware can run this
machine at all before it times one, because a firmware with no WRITE_POINTER loses
vmgenid and a lost vmgenid has no symptom.
cmds:
- SPIN_FIRMWARE_PROBE=1 go test ./boot/ -run TestFirmwareCost -count=1 -v -timeout 20m

boot:trace:
desc: >-
One boot's console printed against the host's clock, for finding a gap that belongs to
Expand Down
Loading
Loading