Skip to content

Qboot experiment and initcall probe - #12

Merged
aledbf merged 18 commits into
mainfrom
qboot-experiment-and-initcall-probe
Sep 26, 2026
Merged

aledbf merged 18 commits into
mainfrom
qboot-experiment-and-initcall-probe

Conversation

@aledbf

@aledbf aledbf commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

aledbf and others added 18 commits September 26, 2026 11:53
… it times

qboot cannot run this machine as upstream ships it. Its ACPI linker-loader
implements ALLOCATE, ADD_POINTER and ADD_CHECKSUM and stops on anything else, and
vmgenid needs WRITE_POINTER to tell QEMU where the GUID lives. The symptom is a
vCPU halted inside the firmware with nothing on the console, not even earlyprintk.
write-pointer.patch adds that command and the fw_cfg DMA writes it needs; NOTICE
carries its GPL-2.0 provenance, since the repository's own licence does not cover
it and nothing here ships it.

The build is a container for the same reason kernel/Dockerfile is, and the reason
turned out to be load-bearing rather than tidy. The earlier recipe used whatever
GCC the host had — 15.2.0, producing 7a316e3c… for the patched binary. The pinned
Debian's GCC is 14.2.0 and produces bf7ddddc… from the same commit, the same patch
and the same flags, and the saving moves from 6.32 ms to 3.86 ms. The compiler
accounts for more than a third of what the firmware change itself buys, which is
the same result the SeaBIOS experiment recorded in boot/phases.go. A number in
qemu/qboot/README.md is only comparable to another built the same way, and it now
says which rows those are.

The probe moves from Python to Go, beside the rest of the boot harness and sharing
its release discovery, its percentiles and its interleaving. It asserts that the
stock firmware hangs with vmgenid on before it times anything: a stock qboot that
booted would mean the WRITE_POINTER story is wrong and every timing after it is
measuring something else. It then saves the guest, restores it with a fresh GUID
and requires the guest to log the reseed — a firmware that boots and quietly fails
to publish the table gives every VM restored from one template the same random
pool, and nothing on either of them reports a fault.

The firmware is about 10 ms of a boot that reaches a login in a few hundred, and
this decides whether a firmware can run the machine at all before it argues about
the milliseconds. It is also not adoptable yet for a reason that is not the number:
Spec.Fingerprint hashes the QEMU binary, the kernel and the initrd, and not the
firmware, so two releases differing only in firmware would exchange templates
across different ACPI and E820 layouts. Selectable firmware needs that fixed first.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The kernel is the largest block in a cold boot after early userspace, and there was
no breakdown of it, only a total. Getting one is not as simple as turning
initcall_debug on, and the two ways it goes wrong are both silent.

The first is the console. With the messages going to a serial port, every one is a
VM exit inside the interval being reported, so a gap between two adjacent stamps is
mostly the cost of printing the one before it and the biggest gap is wherever the
most was printed. This boots with the console quiet and reads the ring buffer
afterwards: `quiet` raises the console's threshold and does not touch what printk
stores, and initcall_debug's two lines are KERN_DEBUG printks either way.

The second is how the buffer gets out. The obvious unit — StandardOutput=journal+
console — routes the dump through the journal, which forwards it to the kernel log,
so every reprinted line arrives carrying a second, fresh stamp:

    [    0.199752] sh[106]: [    0.007375] IOAPIC[0]: apic_id 0, version 17, …

A parser reading the first stamp then reads the clock of the dump. The initcall
durations stay right, because those are text, and every timestamp-derived number
silently describes how long the console took to print a megabyte. It reported the
kernel reaching "Freeing unused kernel image" at 931 ms on a machine that boots in
about 200, which is the only reason it was caught. The dump now goes straight to
/dev/ttyS0, and a line carrying two stamps fails the test rather than being
averaged into it.

Everything is a p50 with a p95 beside it, because one boot is not a measurement:
three consecutive runs on one busy host read 59.5, 96.8 and 167.7 ms for the same
kernel. log_buf_len is 2M and not 8M for the same reason the console is quiet — 8M
allocates 37 MB and cost 4.9 ms of the thing being measured. Gaps are keyed by the
message after them with pids normalised away, because the largest gap in this
machine's boot came out as two entries with half the samples each when journald's
pid differed between runs.

What it says, p50 over 12 boots: 98.2 ms from kernel start to "Freeing unused
kernel image", 54.5 of it in 680 initcalls and 35.9 outside them, with the top ten
initcalls accounting for 84% of the initcall time. acpi_init is 14.0 ms and
inet_init 8.7, and neither can be a module. And the largest gap in the whole boot
is not in the kernel at all: 115-144 ms between the kernel handing over and
journald's first line.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`systemd-analyze` puts userspace at 139 ms against the kernel's 79, so the largest
block in a cold boot is on the far side of the kernel. Nothing could see into it:
the other probes read the kernel ring buffer, and the phase in question is before
journald exists.

boot:userspace asks systemd afterwards — its own kernel/userspace split, `blame`
and the critical chain. Type=simple rather than oneshot, and that is not a style
choice: startup is finished when every job of the initial transaction completes, a
unit in multi-user.target.wants is one of those jobs, and as oneshot it waits for a
timestamp that cannot be set until it exits. It retried for four seconds and gave
up while `blame`, which needs no finish timestamp, answered immediately — so it
looked like a flake rather than a deadlock.

boot:systemd turns systemd's debug log on and ranks the gaps in it. Getting a whole
log took three attempts, and each failure looked like an answer:

  - To the serial console, every line is a VM exit inside the phase being measured.
  - To kmsg, systemd rate-limits itself and says so in the log it is writing, so
    lines vanish in bursts — which is where a gap would be.
  - To the journal, journald's runtime cap is reached mid-boot and it vacuums the
    beginning, the part nothing else can see. It left 48 lines of some nine hundred,
    and a ranked list of gaps across 48 lines reads exactly like a finding.

So it logs at debug to the journal with the runtime cap raised, and asserts the log
is whole before ranking anything: it fails on the kmsg rate-limit message, on a
vacuum that freed a non-zero amount, and on a suspiciously small line count. It also
sorts by timestamp, because journalctl prints in arrival order and two writers'
streams interleave — treating a backwards stamp as the end of the window ranked
gaps over fifteen lines and called the largest gap in the boot 8.9 ms.

What it says, one boot: 2368 tagged lines over 96.8 ms, systemd-udevd the largest
writer at 34.5 ms with udevadm another 11.4, systemd itself 32.1, and udev
reporting "Maximum number (14) of children reached" before a 4.5 ms gap.

The bench gains two rows and both are negative results worth keeping. serial-getty@
ttyS0 carries BindsTo=dev-ttyS0.device and sits on the critical chain at 128 ms, so
masking it and using the image's own device-independent console unit should have
been free. It is 340 ms worse, and removing that unit's Type=idle does not bring it
back. dev-ttyS0.device on the chain is when udev got to the port, not something the
prompt waited for. Read a critical chain as what finished last.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`task boot:systemd` counted it: udev coldplugs 266 devices in a 51 ms window on
this machine, hitting "Maximum number (14) of children reached" eighteen times, and
every worker pays a failing dlopen of libnss_systemd.so.2 and a failing userdb group
lookup on the way. udevd and udevadm together are the largest writers in early
userspace, ~46 ms of it.

Most of those devices are for hardware this machine does not have. Counted:

    64 ttyN      16 memoryN     8 loopN      4 ttySN      4 cpuN      3 virtioN

The 64 virtual consoles are CONFIG_VT on a machine whose QEMU ships no VGA at all
and whose console is ttyS0. The 8 loop devices are never used. One of the four
16550s exists. udev also tries to open /dev/snd/timer and /dev/snd/seq, on a machine
with no sound card.

The new bench row switches off the two that are boot parameters, and it is a
negative: 235/584 against the baseline's 239/423 over 15 boots. That is the
arithmetic rather than a surprise — 11 devices of 266 is 4% of 46 ms, about 2 ms,
below what this harness resolves. Parameters were the right thing to try first
because a kernel config change invalidates every template in existence, and this way
nobody paid for one to find out.

What it settles is the shape, which is the useful part: the cost is per device and no
single device matters. CONFIG_VT=n removes 64 of the 266, so it should be worth 5 to
12 ms of the udev window with 14 workers in parallel — a real prediction to test, and
one worth knowing is ~10 ms of a 239 ms boot before spending a build on it. Early
userspace is 139 ms and it is not one expensive thing; it is 266 cheap ones.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The previous commit predicted that CONFIG_VT=n would be worth 5 to 12 ms: it removes
64 of the 266 devices udev coldplugs, a quarter of them, and udevd and udevadm are
the largest writers in early userspace at ~46 ms. The kernel was built and the
prediction is wrong. 15 boots each on a quiet host, p50/p95 to a login prompt:

    baseline           228/237        fewer udev rules   228/234
    fewer devices      225/238        kernel B (VT=n)    228/246

Zero. Not noise hiding it: at a p95 of 237 a 10 ms saving would be visible. Removing
a quarter of the devices, and nineteen rule files for hardware this machine cannot
have, changes the time to a login prompt by nothing at all.

So udev's 46 ms overlaps with something rather than delaying it. It runs up to 14
workers — it says so, eighteen times — and taking work away hands the wall clock to
whatever it was running alongside.

That is the same mistake three times in three disguises. `systemd-analyze blame`
ranks units by their own duration; a critical chain ranks by what finished last; and
`task boot:systemd`, added two commits ago, ranks writers by the time between their
log lines. Each looks like a list of savings and none of them is one. The comment on
the logind row has said so since 2026-09-10 about the first, and it took building a
kernel to learn it about the third.

The bench can compare kernels now, which is what made this answerable: SPIN_KERNEL_B
adds a row booting a kernel built elsewhere, interleaved with the rest rather than
run as a second pass, because three consecutive runs of one kernel on a busy host
read 59.5, 96.8 and 167.7 ms. editOverlay also refuses a mask that masks nothing — a
symlink to /dev/null whose name matches no rule file switches nothing off and reports
a row that looks like a measurement.

What is left after this: the image's current arrangement is at or near a local
optimum for removing things. logind out of the boot transaction is worth ~20 ms and
is already taken; the console unit that looks device-independent costs 340 ms; and
the remaining 228 ms is not attributable to anything that has been taken out and
measured yet.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`spin-machine boot --profile` already exists and machine.Cmdline.Profiling() is what
it means: initcall_debug, printk.time, a log_buf_len the default ring needs, and a
silent console. The probe added two commits ago built that command line itself, which
is a second definition of the gate — and the two had already drifted, since one used
log_buf_len=2M and the other 4M.

Profiling()'s own comment also names the mechanism the probe rediscovered the hard
way: a console registers during the device_initcall phase and registering it replays
the whole printk ring into it synchronously, inside that initcall, so a verbose boot
charges the measurement to whichever console driver registered. That was written down
before the probe existed.

The probe keeps what is its own: p50/p95 over interleaved boots, the gap ranking, and
the assertions that fail when the measurement is broken rather than averaging it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…y costs

Four removals in a row measured zero — CONFIG_VT, nineteen udev rule files, two
device-count boot parameters, smaller TCP hash tables — and each cost a build or a
bench run to learn. None of them was measured against anything: there was no answer
to how much of the kernel phase is even available.

kernel/minimal/ builds the smallest kernel that still boots: tinyconfig plus PVH, the
8250, an initramfs and the handful of symbols needed to exec an init. 573 symbols
against the release's 1379, 10.2 MiB against 37.5. From tinyconfig up rather than
from the release config down, because subtracting leaves whatever nobody thought to
name — the same reason qemu/devices.mak is an allowlist. It is isolated from
`task build` and `task release`, with its own cache scope, and it could not be a
release kernel: no BTF, BPF, ftrace, cgroups, ext4 or virtio, so it cannot mount the
base image or run systemd.

Which is why the comparison is against a diagnostic initrd's marker and not a login
prompt. That bounds what it measures — the kernel phase and the launch around it —
and it is also the point. 20 boots each, interleaved, 2026-09-26, i9-13900HK, KVM,
q35, 2 vCPUs, 2 GiB:

    release       97.95 / 104.91 ms
    floor         52.33 /  56.95 ms
    release+lz4   91.56 / 103.53 ms

The release kernel's configuration costs 45.6 ms. That is the budget every
kernel-config question has been spending against without knowing it, and it is not
noise: about 9 ms of it is BTF's one-time section parse, ~10 ms acpi_init, ~6 ms
ksm/kcompactd/hugepage and ~5 ms ftrace, by the measurements already recorded in
kernel/Dockerfile and by task boot:initcalls. BTF, ftrace and ACPI are not available
— BPF and hotplug need them — so the part in play is smaller than 45.6, and now it
has a size.

The third row answers a question the tree had only argued: kernel/Dockerfile requires
RD_LZ4 because unpacking the archive is one of the largest single items in kernel
boot. On this machine lz4 rather than gzip for the same 2.1 MB archive is 6.4 ms,
more than the qboot firmware experiment's 3.9. Nothing here ships an initrd, so that
is a number for whoever supplies one.

Two things the probe found on the way, both of which look like unrelated faults:
tinyconfig leaves CONFIG_BINFMT_SCRIPT off, so `rdinit=/init` on a shebang script
fails, the kernel falls through its default inits and starts an interactive BusyBox
shell. And the config assertions run before the compile, because a missing PVH note
is twenty minutes of build followed by a guest that halts with nothing on the console.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CLAUDE.md asks for durable evidence near a decision and for raw samples and
short-lived exploration to live in a commit or a measurement record instead. The rows
added over the last few commits went the other way: three blocks of p50/p95 tables
from one afternoon's exploration, pasted beside the variants they came from. The
numbers are in the commits that produced them; what belongs here is the claim and the
regression it prevents.

So the tables go and the warnings stay: that a critical chain is what finished last
and never a list of savings, which is what stops the console-unit swap being tried
again for 340 ms; and that udev's work overlaps the boot rather than delaying it.

The "console no idle" row goes entirely. It existed to tell two candidates apart, it
did, and CLAUDE.md asks that an experiment be removed once it has no active question.
Its answer — Type=idle is not what the swap costs — is in the commit that added it.

One correction, which is why this is not only a deletion. The comment said
CONFIG_VT=n changes the time to a login prompt by nothing. That is true and it reads
as "it is worth nothing", which is false: `task boot:floor` measures it at 2.09 ms of
the kernel phase. This harness cannot resolve that through ~139 ms of userspace, so a
kernel-config change belongs in boot:floor first and here only once a login prompt
could show it. Four zeroes in a row were the instrument, not the changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… path

CONFIG_EROFS_FS was there as the read-only half of a containerd path that qcow2
replaced. Nothing in this tree builds, mounts or names an erofs any more — the only
references left were the assertion that required it and two comments listing the
filesystem set. It brought CONFIG_EROFS_FS_ZIP with it, and with that LZMA at 16
streams, DEFLATE and ZSTD decompressors, for an image format this machine never reads.

It buys no boot time, and that is not why it goes. Measured with `task boot:floor`,
20 boots each interleaved, 2026-09-26: 100.46 ms against 102.22 with it, which is
noise in both directions. The stripped vmlinux is 48 KiB smaller, because the
decompressors it pulled in have other users and stay. Dead configuration is worth
removing for being dead; the measurement is here so nobody expects the change to pay
for itself in milliseconds.

The assertion is inverted rather than deleted, so the symbol coming back fails the
build and names the reason instead of arriving unnoticed in a config refresh.

This is a content change to the kernel, so it invalidates every template in
existence. It should not be released on its own: CONFIG_VT=n measures 2.09 ms by the
same probe and is not committed yet, and the fleet is better off paying for one
fingerprint change than for two.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…didate

The "console no dev" row costs 340 ms and nothing explained it. Two things here.

consoleSwap() is extracted from the row so `task boot:systemd` can measure the same
configuration instead of a second copy of it, which is what
SPIN_SYSTEMD_VARIANT=console-swap now does. What it shows is not a dependency being
waited on: early userspace runs about three times slower throughout — udevd 120 ms
against 34.5, udevadm 60 against 11, systemd-sysctl 39 against 3 — and systemd's
first logged line lands at 1009 ms instead of 409.

And a third candidate eliminated. Both units carry TTYReset and TTYVHangup, but
serial-getty@ waits for dev-%i.device where the replacement is ordered only after
systemd-user-sessions, so the replacement resets and hangs up /dev/ttyS0 early in the
boot rather than near the end of it, while systemd is still using that console. That
is the shape that would degrade everything after it. It is not the cause: with
TTYReset=no and TTYVHangup=no the swap measures 693 against 701.

So: not the device dependency it was meant to remove, not Type=idle, not the tty
handling. The cost is real and reproducible and the cause is unknown, which is worth
saying plainly rather than leaving three plausible explanations standing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
spin-machine-console.service cost 331 ms whenever it was used, and three explanations
were measured and wrong before the console named the fourth itself. On a swapped boot
the bytes before the login prompt are

    ESC P + q 6 E 6 1 6 D 6 5 ESC backslash

an XTGETTCAP asking for the capability whose name is the hex string 6E616D65, "name",
with nothing after it until the wait expires. That is the 334 ms terminfo timeout
machine/cmdline.go already puts TERM=dumb on the kernel command line to avoid: the far
end of this serial port is a file and never answers.

The parameter reaches systemd and not this unit. systemd sets TERM itself for a
service that declares its own TTYPath rather than passing its own environment down, so
the unit gets a terminal type that looks capable and the tty setup asks. One
Environment= line fixes it: 221/233 against 558/572 without, over 12 boots, with a
baseline of 227/236. Without it the unit is unusable; with it it costs nothing.

Which reopens something. spin-machine-console.service's own comment says the hole it
leaves is one this file is the wrong place to close, and the bench row that tried
closing it read 340 ms worse. That row now reads 224/230 against the baseline's
227/241. The device-independent console unit is viable; what stands between it and
production is the exposure question that comment raises, not a boot cost.

What the localisation took, since the wrong reading came first: with DEBUG=1 the bench
shows FIRMWARE, KERNEL, PID1 and STARTED identical between the two configurations and
only USABLE moving, so the cost was entirely after systemd finished. An earlier reading
of `task boot:systemd` had suggested the whole of early userspace running three times
slower; that was the debug probe inflating its own window, not the machine.

Three candidates are recorded as eliminated rather than kept as rows: the device
dependency the swap removes, the unit's Type=idle, and its TTYReset/TTYVHangup, which
both units carry and so could never have been a difference between them. The same
timeout was tried on the real getty — TERM=dumb on serial-getty@ttyS0 changes nothing,
1229 against 1232 — so the second that `as shipped` reads is agetty's own sleep.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`task boot:initcalls` puts ksm_init at 2.00 ms, hugepage_init at 2.00 and
kcompactd_init at 2.00, which made the memory-management options look like the largest
config change still available. A kernel with CONFIG_KSM and CONFIG_TRANSPARENT_HUGEPAGE
off, on top of the erofs removal, measures -0.50 ms against the release kernel at
`task boot:floor` — 93.09 against 93.60 over 20 interleaved boots, 2026-09-26. Nothing,
on the instrument that resolves the 2.09 ms CONFIG_VT=n is worth.

The round 2.00 values are the tell: that is what initcall_debug resolves at
CONFIG_HZ=1000, not what those initcalls cost. A per-initcall duration is no more a list
of savings than `systemd-analyze blame` is.

COMPACTION is not available at all, which the build found before compiling anything:
CONFIG_VIRTIO_MEM depends on CONTIG_ALLOC, and that is
def_bool (MEMORY_ISOLATION && COMPACTION) || CMA, so compaction is what this machine's
memory resize path rides on. The VIRTIO_MEM assertion is what caught it.

Recorded in kernel/Dockerfile beside the other candidates measured and rejected, because
finding this out costs a kernel build.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ready warm

A systemd issue about unit-file cache rescanning looked like it might be this
machine's: the message it reports — "Modification times have changed, need to update
cache" — is in this image's debug log, twice per boot, and systemd says it spends
212 ms loading units and determining the initial transaction.

It is not, and two checks say so.

The persistent caches are built at image time and the services that would rebuild
them at boot are masked in optimize-systemd.sh: ld.so.cache at 9975 bytes, the journal
catalog at 274807, locale-archive at 3064432, and systemd-update-done,
systemd-hwdb-update, systemd-journal-catalog-update, ldconfig and systemd-sysusers all
symlinked to /dev/null. There is no pre-warm missing. The unit cache is not one of
these: it is names and modification times held in memory for the length of one boot,
with no on-disk artefact to build, and it is rescanned because the generators write
into /run/systemd/generator.late, which is in the search path. That is how systemd
starts.

And the count is not the cost. A row adding 300 unit files that nothing wants —
against 282 in /usr/lib/systemd/system and 65 in /etc/systemd/system — measures 236
against the baseline's 246. The distinction worth keeping is that systemd's cache
holds names, not parsed units, so a unit nothing references is scanned and never
loaded: this measures the scan, and the scan is free. Parsing is only for units
something pulls in, and those cannot be added without starting them too.

The variant struct gains `setup`, one shell run with the overlay mounted, because
three hundred files through one `sudo tee` each would take longer than the boots they
prepare. It runs under `set -u` with $MNT passed through sudo's own assignment rather
than the child environment, which sudo resets — the first version failed on the unset
variable, and without -u it would have written 300 unit files into the host's
/etc/systemd/system.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A systemd issue about unit-cache rescanning pointed at "Loaded units and determined
initial transaction in 212ms", which this image's debug log reports. Reading where that
message comes from was cheaper than measuring around it. src/core/main.c brackets
manager_startup() and do_queue_default_job(), and logs at LOG_DEBUG — except under
`--test`, where it logs at LOG_INFO. So systemd will report the number without the
debug logging that produced it, and in test mode it does only that work: nothing is
started and manager_preset_all is skipped.

`systemd --test --system`, seven runs inside one boot, says 26 ms. The 212 was the
measurement's own logging, inflating by about eight times — the fourth time today that
observing a boot has been most of what a boot appeared to cost.

Masking two generators that cannot do anything on this machine — the tpm2 one, whose
libtss2-esys.so.0 the boot log says cannot be opened, and the factory-reset one, on a
machine the same log says is not booted in EFI mode — is 1.0 ms of that 26. Seven of
the thirteen the image carries are masked already.

And manager_preset_all does not run here, which the source explains and the logs
confirm: it needs first_boot, systemd decides that from /etc/machine-id, and this image
ships that file present and empty rather than absent or containing "uninitialized" —
so systemd reads an initialised system and skips the preset pass. An empty machine-id
is load-bearing for more than rule 3.

The probe runs as nobody because systemd refuses test mode for root, which also makes
its absolute number a floor for what PID 1 pays: no cgroup setup, no inherited
descriptors, nothing to deserialise. It is for the order of magnitude and for the
difference between two configurations, and it fails rather than measuring the wrong
user if privileges cannot be dropped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CONFIG_VT builds 64 tty devices on a machine whose QEMU ships no VGA at all and whose
console is one 16550 on ttyS0. Nothing opens them and udev walks them on every boot — a
quarter of the 266 devices it coldplugs. It goes for the same reason erofs went: it is
dead, and this release is already paying one fingerprint change for that removal, so it
costs nothing more to take both.

It is not off because it is slow, and the number that said otherwise is withdrawn. An
earlier 20-boot run of `task boot:floor` put CONFIG_VT=n at 2.09 ms, and two commits
quoted it — the erofs one as a reason to bundle the two, and the KSM one as the
instrument's resolution. It does not reproduce: 40 interleaved boots on an idle host
read 0.31 ms for erofs and CONFIG_VT together. The 2.09 was measured while this host was
building a kernel, and at this size a single run of that probe is a guess.

So the standing claim is narrower and duller: nothing found in this configuration is
worth more than about half a millisecond. The 45.6 ms between this kernel and the floor
one in kernel/minimal/ is BTF, ACPI and ftrace, and BPF and hotplug need all three.

Also masked, in image/: systemd-tpm2-generator and systemd-factory-reset-generator, the
two of the six still running that cannot do anything here — the boot log says
libtss2-esys.so.0 cannot be opened and that this machine is not booted in EFI mode.
Worth 1.0 ms of the 26 ms systemd spends loading units, by `task boot:unitload`. That one
is an image change and invalidates nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The ~27 ms between exec'ing QEMU and the guest's first instruction had a breakdown but
no attribution, and two attempts to find something reducible in it came back empty:
321 ioctls are issued before the first KVM_RUN, which cannot be the 12 ms of accel,
realize and ACPI build, and a perf profile bounded to the first 85 ms puts QEMU's own
code at 3.67%.

What it is, is the host. kernel_init_pages — zeroing pages as guest RAM is faulted in —
is 18.8% of a cold start over twelve launches. On a restore it is 3.0%, with
next_uptodate_folio and filemap_get_entry in its place: the template's memory arrives as
mapped file pages instead of fresh anonymous ones. rom_reset's 6 ms goes the same way and
QEMU says so in its own source rather than leaving it to be measured, skipping every ROM
under RUN_STATE_INMIGRATE "because we'll fill the data in during the next incoming
migration in all cases".

The caveat on the comparison: the cold window includes the guest executing and the
restore window does not, since a machine started with -incoming waits in inmigrate for a
cont. The item in dispute is in the setup phase of both.

So the largest cost in this phase is the host giving a machine memory it does not have
yet, which is not a configuration problem, and the thing that avoids it is a template —
at 2 GiB on disk each, not sparse, plus ~92 MiB of device state, fast to restore only
while that file is in the host's page cache, and coupled to exactly one fingerprint.
Recorded where the breakdown is so the next person reading those five lines knows which
two a restore does not pay.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@aledbf
aledbf merged commit a0efd65 into main Sep 26, 2026
5 checks passed
@aledbf
aledbf deleted the qboot-experiment-and-initcall-probe branch September 26, 2026 17:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant