Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 7 additions & 6 deletions .github/workflows/image.yml
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,11 @@ jobs:
- name: Set up Buildx
uses: docker/setup-buildx-action@v4

# Without this, `--cache-to type=gha` below is a silent no-op: the tokens BuildKit
# needs reach JavaScript actions and not `run:` steps. See the long note in qemu.yml.
- name: Expose the Actions cache to buildx
uses: crazy-max/ghaction-github-runtime@v4

# image/build.sh asserts what came out on the finished filesystem rather than
# trusting the build that made it: that there is an /sbin/init and a /bin/sh, that
# the directories the guest mounts over exist, that no identity is baked in, and
Expand All @@ -59,15 +64,11 @@ jobs:
# the build container's, and image/build.sh stops when it is not there.
- name: Build e2fsprogs
run: |
task e2fsprogs:build \
E2FSPROGS_CACHE_FROM=type=gha \
E2FSPROGS_CACHE_TO=type=gha,mode=max
CACHE_BACKEND=gha task e2fsprogs:build

- name: Build the base image
run: |
task image:build \
IMAGE_CACHE_FROM=type=gha \
IMAGE_CACHE_TO=type=gha,mode=max
CACHE_BACKEND=gha task image:build

- name: What this image is
run: |
Expand Down
23 changes: 12 additions & 11 deletions .github/workflows/kernel.yml
Original file line number Diff line number Diff line change
Expand Up @@ -44,27 +44,28 @@ jobs:
- name: Set up Buildx
uses: docker/setup-buildx-action@v4

# Without this, `--cache-to type=gha` below is a silent no-op: the tokens BuildKit
# needs reach JavaScript actions and not `run:` steps. See the long note in qemu.yml.
- name: Expose the Actions cache to buildx
uses: crazy-max/ghaction-github-runtime@v4

# Ends in `task kernel:verify`, which opens the ELF and checks the Xen PVH notes
# survived the strip. Without them QEMU has no entry point into this kernel and the
# failure is a VM that does not start, in whichever lane boots one next.
- name: Build the kernel
run: |
task kernel:build \
KERNEL_CACHE_FROM=type=gha \
KERNEL_CACHE_TO=type=gha,mode=max \
KERNEL_NPROC=4
CACHE_BACKEND=gha task kernel:build KERNEL_NPROC=4

# Compared, now, against the last released machine — which is the reference this step
# used to say it did not have. It does: every release publishes its machine.env as an
# asset, and the kernel's checksum is one of the two numbers in it that decide whether
# a template still matches.
# Compared against the last released machine, which every release publishes as a
# machine.env asset. The kernel's checksum is one of the two numbers in it that decide
# whether a template still matches.
#
# So a pull request that changes the config, or bumps the Debian toolchain the kernel
# is compiled with, says on its own summary that it invalidates every template in
# existence. That is not a reason to reject it — it is a release this repository makes
# deliberately, and the step exits 0 either way — but it was previously a sentence
# nobody wrote down, and the last time it happened it was a comment in this directory's
# Dockerfile that nobody expected to rebuild anything.
# deliberately, and the step exits 0 either way. It is a reason to know: the change
# that does this need not look like much, and a comment edited in this directory's
# Dockerfile is enough.
#
# It speaks only for the kernel: this workflow builds no QEMU, and the verdict says so
# rather than implying it checked the whole machine.
Expand Down
40 changes: 31 additions & 9 deletions .github/workflows/qemu.yml
Original file line number Diff line number Diff line change
Expand Up @@ -9,9 +9,9 @@ name: QEMU

on:
# qemu/** and nothing else. The version pin is the ARG in qemu/Dockerfile, so a bump is a
# change under qemu/ and fires this; the root Taskfile holds only how many jobs to compile
# with and which cache to use, and used to be listed here — which meant every edit to it
# rebuilt QEMU, the kernel and the base image for nothing.
# change under qemu/ and fires this. The root Taskfile is deliberately not listed: it holds
# only how many jobs to compile with and which cache to use, so listing it would rebuild
# QEMU, the kernel and the base image for an edit that cannot change any of them.
push:
branches: [main]
paths:
Expand Down Expand Up @@ -58,6 +58,24 @@ jobs:
- name: Set up Buildx
uses: docker/setup-buildx-action@v4

# What makes `--cache-to type=gha` do anything at all.
#
# BuildKit writes to GitHub's cache service with $ACTIONS_RUNTIME_TOKEN and
# $ACTIONS_RESULTS_URL, and those are handed to JavaScript actions, not to `run:`
# steps. Docker's documentation is explicit — "If you invoke the `docker buildx`
# command manually from an inline step, then the variables must be manually exposed"
# — and names this action for it. Every build here runs through `task`, which is a
# `run:` step.
#
# Without it the flags are accepted, nothing is stored, and nothing says so: the
# builds pass and take the cold time, every run. That failure is invisible in a log.
# Where it shows is `gh api repos/{owner}/{repo}/actions/cache/usage` — a repository
# that builds QEMU and a kernel and reports a handful of megabytes is caching neither.
#
# Remove this step and nothing breaks; everything just gets slow again.
- name: Expose the Actions cache to buildx
uses: crazy-max/ghaction-github-runtime@v4

# The pinned version comes from the Taskfile, never from a literal here: two places
# to bump is how CI and production end up on different QEMUs.
#
Expand All @@ -73,10 +91,7 @@ jobs:

- name: Build QEMU and verify what came out
run: |
task qemu:build \
QEMU_CACHE_FROM=type=gha \
QEMU_CACHE_TO=type=gha,mode=max \
QEMU_JOBS=4
CACHE_BACKEND=gha task qemu:build QEMU_JOBS=4

# vhost-vsock is a kernel module, and the machine's vsock device cannot be created
# without /dev/vhost-vsock. The runner has the node and gives the user no access to
Expand Down Expand Up @@ -145,15 +160,22 @@ jobs:
# Two tags: the QEMU version, which is what everything pins, and the commit, which is
# what makes a specific build reachable when the version tag has moved because a flag
# changed.
#
# Two --cache-from and one --cache-to, which is not symmetry gone wrong. This target
# sits on top of the builder stage the step above just compiled, so it reads the
# `qemu` scope to get it and writes only its own — sharing one scope would have the
# two builds overwriting each other's cache, which is the default behaviour that made
# every scope in this repository distinct in the first place.
- name: Publish the runtime image
if: github.event_name != 'pull_request'
run: |
docker buildx build \
--file qemu/Dockerfile \
--target runtime \
--platform linux/amd64 \
--cache-from type=gha \
--cache-to type=gha,mode=max \
--cache-from type=gha,scope=qemu \
--cache-from type=gha,scope=qemu-image \
--cache-to type=gha,scope=qemu-image,mode=max \
--build-arg QEMU_VERSION=${{ steps.qemu.outputs.version }} \
--build-arg JOBS=4 \
--label org.opencontainers.image.source=https://github.com/${{ github.repository }} \
Expand Down
47 changes: 27 additions & 20 deletions .github/workflows/release.yml
Original file line number Diff line number Diff line change
Expand Up @@ -26,12 +26,12 @@ permissions:
contents: write
packages: write

# One release at a time, across every ref. The group used to be per-ref, which serialised
# two runs on the same tag and nothing else: a workflow_dispatch on main and a tag push are
# different refs, so both counted the tags already taken today, both generated .01, and the
# loser found out at `git tag` — after six hours of building. There is no reason to build
# two releases at once, and the version is decided at the start of a run and pushed at the
# end, so the window is the whole build.
# One release at a time, across every ref, and not per-ref. Per-ref serialises two runs on
# the same tag and nothing else: a workflow_dispatch on main and a tag push are different
# refs, so both count the tags already taken today, both generate .01, and the loser finds
# out at `git tag` — after six hours of building. The version is decided at the start of a
# run and pushed at the end, so that window is the whole build, and there is no reason to
# build two releases at once anyway.
concurrency:
group: release
cancel-in-progress: false
Expand Down Expand Up @@ -63,6 +63,11 @@ jobs:
- name: Set up Buildx
uses: docker/setup-buildx-action@v4

# Without this, `--cache-to type=gha` below is a silent no-op: the tokens BuildKit
# needs reach JavaScript actions and not `run:` steps. See the long note in qemu.yml.
- name: Expose the Actions cache to buildx
uses: crazy-max/ghaction-github-runtime@v4

# Three sources, in order: what the caller typed, the tag that triggered this, or a
# new one for today. The generated form is deliberately not `date +%s` or a commit
# hash — it has to be readable, sortable, and the same shape as one typed by hand.
Expand Down Expand Up @@ -107,26 +112,28 @@ jobs:
- name: Build the release
env:
VERSION: ${{ steps.version.outputs.version }}
QEMU_CACHE_FROM: type=gha
QEMU_CACHE_TO: type=gha,mode=max
KERNEL_CACHE_FROM: type=gha
KERNEL_CACHE_TO: type=gha,mode=max
E2FSPROGS_CACHE_FROM: type=gha
E2FSPROGS_CACHE_TO: type=gha,mode=max
IMAGE_CACHE_FROM: type=gha
IMAGE_CACHE_TO: type=gha,mode=max
# The same scopes the per-artefact workflows write on main, decided in
# Taskfile.yml. That is the whole point of them being named there: a release
# reads what the lane on main already paid to compile, instead of starting cold
# because a workflow file spelled the scope differently.
#
# This run's own writes are mostly wasted — it is triggered by a tag, and a cache
# written under one tag ref cannot be read from another tag or from a branch. The
# reads are what matter, and those come from the default branch.
CACHE_BACKEND: gha
run: task release QEMU_JOBS=4 KERNEL_NPROC=4

# The machine's identity, and what it costs whoever holds templates.
#
# "why did every template in the fleet stop matching?" used to be answered by opening
# two release pages and comparing sixty-four hex digits by eye, which is a thing
# nobody does and therefore an answer nobody had. hack/fingerprint-diff fetches the
# previous release's machine.env and states the consequence in a sentence, at the top
# of the notes, where a consumer deciding whether to take this release reads it.
# The alternative to hack/fingerprint-diff is opening two release pages and comparing
# sixty-four hex digits by eye, which is a thing nobody does — so the answer to "why
# did every template in the fleet stop matching?" is one nobody has. The script
# fetches the previous release's machine.env and states the consequence in a sentence,
# at the top of the notes, where a consumer deciding whether to take this release
# reads it.
#
# It exits 0 on every verdict: a changed machine is a release this repository makes on
# purpose. The point is that it can no longer be made quietly.
# purpose. The point is not to prevent it but to keep it from being quiet.
- name: What this release is, and what it breaks
id: manifest
env:
Expand Down
31 changes: 15 additions & 16 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -69,9 +69,9 @@ One tarball:

`task build` writes that same tree into `_output/`, byte for byte the layout above, and
`machine.Open` reads either. There is one layout: nothing rearranges the files on the way
out of a build, into a tarball or into a consumer, because the three used to differ and
what fell out of the translation between them was a path that existed and held the
previous release's kernel.
out of a build, into a tarball or into a consumer. Let the three differ and what falls out
of the translation between them is a path that exists and holds the previous release's
kernel.

```go
rel, err := machine.Open("/usr/share/spin-stack") // says which file is missing, if one is
Expand Down Expand Up @@ -174,10 +174,9 @@ It boots this QEMU and this kernel over a throwaway qcow2 overlay on `rootfs.qco

It boots with no initrd at all: `root=/dev/vda rw init=/sbin/init`, which the kernel can
serve because virtio-blk and ext4 are built in and `image/build.sh` writes a partitionless
filesystem. There was a debug initramfs here until 2026-09-10 — static Go that mounted the
API filesystems, found the root disk and exec'd — and it was a second init, doing what the
init a consumer brings already does. What removed the need for it was turning `systemd-udevd`
back on.
filesystem. There is no debug initramfs: one that mounted the API filesystems, found the
root disk and exec'd would be a second init, doing what the init a consumer brings already
does. What removes the need for one is `systemd-udevd` being on.

Three things the shell has found, all of them true of the image and none of them visible
from outside:
Expand All @@ -190,12 +189,12 @@ from outside:
its own with the dependency removed. Masking bought nothing measurable and cost a
ten-second `dev-ttyS0.device` timeout on every boot that had no initrd; the numbers are in
`optimize-systemd.sh`, where the decision is.
- **`ssh.service` used to fail five times and give up.** The image ships no host keys — they
are identity — and the distribution's `sshd-keygen.service` carries
`ConditionFirstBoot=yes`, so it did not run before `sshd` was asked to validate a
configuration with no keys. Fixed by generating them at boot instead
(`spin-machine-sshd-keygen.service`): every boot here *is* a first boot, since the root
filesystem is a fresh overlay, so the keys live exactly as long as the VM does.
- **`sshd` needs host keys generated at boot.** The image ships none — they are identity —
and the distribution's `sshd-keygen.service` carries `ConditionFirstBoot=yes`, so it does
not run before `sshd` is asked to validate a configuration with no keys, and
`ssh.service` fails five times and gives up. `spin-machine-sshd-keygen.service` generates
them instead: every boot here *is* a first boot, since the root filesystem is a fresh
overlay, so the keys live exactly as long as the VM does.
- **Two thirds of the boot was systemd asking the console questions.** The kernel execs
`/sbin/init` at 59 ms and systemd's first log line arrived at 852, with nothing running in
between. It was a terminfo query and a terminal reset, 334 ms of timeout each, waiting for
Expand Down Expand Up @@ -283,9 +282,9 @@ which is a different fingerprint. It moves on purpose, not when Debian publishes
release.

And the build is reproducible: two builds from the same inputs produce the same `vmlinux`,
byte for byte. It was not until 2026-09-14: bookworm's pahole 1.24 encoded `.BTF` with make's
`-j` in whatever order its threads finished, and every release had a new fingerprint whether
or not the kernel had changed. trixie's 1.30 is deterministic with `-j`.
byte for byte. That needs trixie's pahole 1.30, which is deterministic with `-j`. Bookworm's
1.24 encodes `.BTF` in whatever order make's `-j` threads finish, which gives every release a
new fingerprint whether or not the kernel changed.

### Base image

Expand Down
Loading
Loading