From 159e10ce81ad59295b5199956b2faca55c31f693 Mon Sep 17 00:00:00 2001 From: jiashuoz Date: Mon, 28 Sep 2026 21:48:43 +0800 Subject: [PATCH 1/3] Require cross-UID signaling for non-root microVM lifecycle --- docs/microvm-host-privileges.md | 29 ++++++++++++++-------- internal/driver/microvm_caps.go | 6 ++++- internal/driver/microvm_caps_linux_test.go | 2 +- internal/driver/microvm_caps_test.go | 9 +++++++ internal/driver/microvm_kvm_test.go | 25 ++++++++++++------- scripts/microvm-kvm-test.sh | 5 +++- 6 files changed, 54 insertions(+), 22 deletions(-) diff --git a/docs/microvm-host-privileges.md b/docs/microvm-host-privileges.md index bcb7042..387eaf1 100644 --- a/docs/microvm-host-privileges.md +++ b/docs/microvm-host-privileges.md @@ -12,14 +12,14 @@ everything not on this list. ## Hardware qualification limitation -The capability list below is the intended non-root contract, not proof that the -upstream jailer can run with it. Firecracker jailer 1.17.0 unconditionally writes -`cgroup.subtree_control` in every ancestor up to the host cgroup root when a -`--cgroup` property is supplied. A systemd-delegated subtree does not authorize -those ancestor writes, so the non-root launch fails before the API socket opens. -Do not grant the runner write access to the host cgroup root to work around it. -A scoped privileged launch path or a jailer change needs separate qualification. -Root-run Phase A feasibility tests do not establish this non-root contract. +The capability list below requires a separately qualified jailer. Stock +Firecracker jailer 1.17.0 writes outside a delegated cgroup subtree, gives +away directory ownership before device setup is complete, and retains ambient +capabilities across a non-root-to-non-root UID change. The Rainier Cloud +`infra/microvm/jailer` patch addresses these boundaries and carries its own +source/compiler/binary pins and real launcher tests. Use the reviewed pin; +do not grant global cgroup write access or broad DAC powers as a workaround. +Root-run feasibility alone does not establish this non-root contract. The host must also prepare `/run/netns` for namespace mount-point creation: root-owned, runner-group-writable, mode `0770` on a host dedicated to that @@ -34,6 +34,7 @@ is transferred to the VM. A resume temporarily reclaims the inode using | | What it is for | |---|---| | `CAP_CHOWN` | Give each VM's jail directory, control socket and disk images to that VM's own uid. | +| `CAP_KILL` | Signal and reap VMM processes owned by different per-VM uids during stop, failed launch and delete. | | `CAP_SETGID` | Let the jailer drop the VMM to its per-VM gid. | | `CAP_SETUID` | Let the jailer drop the VMM to its per-VM uid. | | `CAP_NET_ADMIN` | Create each session's network namespace, veth pair and TAP device, and install its `nftables` ruleset. | @@ -77,8 +78,8 @@ thing to discover at startup and not at the first suspend. User=rainier Group=rainier SupplementaryGroups=kvm -AmbientCapabilities=CAP_CHOWN CAP_SETGID CAP_SETUID CAP_NET_ADMIN CAP_SYS_CHROOT CAP_SYS_ADMIN CAP_MKNOD -CapabilityBoundingSet=CAP_CHOWN CAP_SETGID CAP_SETUID CAP_NET_ADMIN CAP_SYS_CHROOT CAP_SYS_ADMIN CAP_MKNOD +AmbientCapabilities=CAP_CHOWN CAP_KILL CAP_SETGID CAP_SETUID CAP_NET_ADMIN CAP_SYS_CHROOT CAP_SYS_ADMIN CAP_MKNOD +CapabilityBoundingSet=CAP_CHOWN CAP_KILL CAP_SETGID CAP_SETUID CAP_NET_ADMIN CAP_SYS_CHROOT CAP_SYS_ADMIN CAP_MKNOD Delegate=yes ExecStart=/usr/local/bin/runnerd --driver=microvm ... ``` @@ -163,3 +164,11 @@ whichever filesystem `/tmp` happens to be. That a create on XFS with `reflink=1` takes microseconds and shares the image's extents — and that a guest writing to its root leaves the shared image untouched on a real block device rather than on a test's byte comparison — is host qualification too. + +## Non-root hardware tests + +The opt-in KVM harness accepts `RAINIER_MICROVM_TEST_CGROUP_PARENT` for the +service's delegated cgroup path (relative to `/sys/fs/cgroup`). Keep the test +process in a child leaf before enabling controllers, exactly as the runner's +service launcher does. The default remains `rainier-kvmtest` for existing +root-operated probes. No global cgroup write grant is needed. diff --git a/internal/driver/microvm_caps.go b/internal/driver/microvm_caps.go index 52c3463..9b5b734 100644 --- a/internal/driver/microvm_caps.go +++ b/internal/driver/microvm_caps.go @@ -50,6 +50,10 @@ type hostCapability struct { // microvmCapabilities is the whole list. The bit numbers are // 's. var microvmCapabilities = []hostCapability{ + { + Name: "CAP_KILL", Bit: 5, + Why: "signal and reap VMM processes running as distinct per-VM uids", + }, { Name: "CAP_CHOWN", Bit: 0, Why: "give each VM's jail directory, control socket and disk images to that VM's own uid (chownJailPath)", @@ -111,7 +115,7 @@ func capabilityNames() string { // // It fails closed and names the FIRST thing that is missing rather than // summarising: an operator fixing a unit file wants the next thing to add, -// and a list of seven is a list nobody reads to the end. +// and a list of eight is a list nobody reads to the end. func checkMicrovmPrivileges(opts MicrovmOpts) error { effective, err := effectiveCapabilities() if err != nil { diff --git a/internal/driver/microvm_caps_linux_test.go b/internal/driver/microvm_caps_linux_test.go index efa47fa..6ed6ad0 100644 --- a/internal/driver/microvm_caps_linux_test.go +++ b/internal/driver/microvm_caps_linux_test.go @@ -9,7 +9,7 @@ // namespace and no jailer on a Mac either. Running these there would assert // that a capability check fails for the reason the platform does not have // capabilities, which is not what any of them is about, and it is what made -// seven subtests fail on every developer machine. The portable half — the +// eight subtests fail on every developer machine. The portable half — the // requirements text, the host preparation — is in microvm_caps_test.go and // runs everywhere. package driver diff --git a/internal/driver/microvm_caps_test.go b/internal/driver/microvm_caps_test.go index 7089249..cd50276 100644 --- a/internal/driver/microvm_caps_test.go +++ b/internal/driver/microvm_caps_test.go @@ -97,3 +97,12 @@ func TestTheHostPreparationTextNamesBothHalves(t *testing.T) { } } } + +func TestMicrovmRequiresCrossUIDSignalCapability(t *testing.T) { + for _, capability := range microvmCapabilities { + if capability.Name == "CAP_KILL" && capability.Bit == 5 { + return + } + } + t.Fatal("runner must require CAP_KILL before accepting VMs it cannot stop") +} diff --git a/internal/driver/microvm_kvm_test.go b/internal/driver/microvm_kvm_test.go index aa8a688..5115769 100644 --- a/internal/driver/microvm_kvm_test.go +++ b/internal/driver/microvm_kvm_test.go @@ -121,6 +121,8 @@ const ( // kvmTimingsEnv, when set, is a file the harness writes its measurements // to as JSON, for the runbook's evidence table to cite. kvmTimingsEnv = "RAINIER_MICROVM_TEST_TIMINGS" + // Optional delegated parent for a non-root systemd test service. + kvmCgroupEnv = "RAINIER_MICROVM_TEST_CGROUP_PARENT" kvmDefaultKernel = "vmlinux" kvmDefaultRootfs = "rootfs.ext4" @@ -148,13 +150,14 @@ const ( // kvmFixture is the host, resolved, once every precondition has held. type kvmFixture struct { - images string - kernel string - rootfs string - stateParent string - connectWait time.Duration - proxyAddr string - proxyPort int + images string + kernel string + rootfs string + stateParent string + connectWait time.Duration + proxyAddr string + proxyPort int + cgroupParent string } // requireKVMHost skips unless this machine can actually boot a microVM, and @@ -242,6 +245,10 @@ func requireKVMHost(t *testing.T) kvmFixture { stateParent: os.Getenv(kvmStateParentEnv), connectWait: kvmConnectWait, } + fx.cgroupParent = os.Getenv(kvmCgroupEnv) + if fx.cgroupParent == "" { + fx.cgroupParent = kvmCgroupChild + } if raw := os.Getenv(kvmConnectEnv); raw != "" { d, err := time.ParseDuration(raw) if err != nil { @@ -351,7 +358,7 @@ func newKVMMicrovm(t *testing.T, fx kvmFixture) (*Microvm, *kvmRunner, CloneMeth EgressProxyAddr: fx.proxyAddr, EgressProxyPort: fx.proxyPort, - Jail: JailOpts{CgroupParent: kvmCgroupChild}, + Jail: JailOpts{CgroupParent: fx.cgroupParent}, }) if err != nil { // A Fatal and not a Skip: every precondition this constructor checks @@ -1222,7 +1229,7 @@ func TestMicrovmBootSmokeOnKVM(t *testing.T) { if err != nil { t.Fatalf("read /proc/%d/cgroup: %v", pid, err) } - if want := "/" + kvmCgroupChild + "/" + h.ID; !strings.Contains(string(procCgroup), want) { + if want := "/" + fx.cgroupParent + "/" + h.ID; !strings.Contains(string(procCgroup), want) { t.Fatalf("ADR-0003 §4.6: the VMM is not in its own cgroup (%q is not in %q)", want, strings.TrimSpace(string(procCgroup))) } diff --git a/scripts/microvm-kvm-test.sh b/scripts/microvm-kvm-test.sh index 4a3461f..83c19c9 100755 --- a/scripts/microvm-kvm-test.sh +++ b/scripts/microvm-kvm-test.sh @@ -29,6 +29,8 @@ # Put it on a filesystem that reflinks (XFS) or # the harness logs that every create is a full # copy of the environment image. +# RAINIER_MICROVM_TEST_CGROUP_PARENT optional delegated service cgroup; +# default rainier-kvmtest (root-operated harness). # RAINIER_MICROVM_TEST_PROXY ip:port of the egress proxy, the one # host-side destination a slot's firewall # allows. Empty means no exception, which is a @@ -140,7 +142,7 @@ BINS command -v "$go_bin" >/dev/null 2>&1 || note_missing "\`$go_bin\` is not on PATH. Set GO to the go binary, or install one." -# The capability set, from the same seven bits internal/driver checks. A host +# The capability set, from the same eight bits internal/driver checks. A host # with no /proc is not a microVM host and will have failed above; the guard is # here so this script's own test can run somewhere else. if [[ -r "$proc_status" ]]; then @@ -150,6 +152,7 @@ if [[ -r "$proc_status" ]]; then if (( ( 0x$capeff >> bit ) & 1 )); then continue; fi note_missing "this process does not hold $name, needed to ${why}. Re-run under \`sudo -E\`, or give runnerd's user the ambient set (see MicrovmHostRequirements)." done <<'CAPS' +5 CAP_KILL signal VMM processes owned by per-VM uids 0 CAP_CHOWN give each VM's jail, control socket and disk images to that VM's own uid 6 CAP_SETGID let the jailer drop the VMM to its per-VM gid 7 CAP_SETUID let the jailer drop the VMM to its per-VM uid From c0750e8e0a79afa6bcdcb953cd75786a491f1efa Mon Sep 17 00:00:00 2001 From: jiashuoz Date: Mon, 28 Sep 2026 22:21:34 +0800 Subject: [PATCH 2/3] Document the KVM lifecycle observer privilege boundary --- docs/microvm-host-privileges.md | 7 +++++++ 1 file changed, 7 insertions(+) diff --git a/docs/microvm-host-privileges.md b/docs/microvm-host-privileges.md index 387eaf1..f5f017b 100644 --- a/docs/microvm-host-privileges.md +++ b/docs/microvm-host-privileges.md @@ -172,3 +172,10 @@ service's delegated cgroup path (relative to `/sys/fs/cgroup`). Keep the test process in a child leaf before enabling controllers, exactly as the runner's service launcher does. The default remains `rainier-kvmtest` for existing root-operated probes. No global cgroup write grant is needed. + +Select `TestMicrovmSatisfiesContractOnKVM` for the non-root contract run. The +separate `TestMicrovmBootSmokeOnKVM` also reads another UID's +`/proc//ns/net`, which requires a root observer on the qualified host. +Keep that observer evidence separate; do not add tracing capabilities to the +production runner just to satisfy the test. Real non-root dev-cell lifecycle +checks and externally observed namespace/capability checks cover that boundary. From e72be58ad9266a8e7a5c7472757edaa9523c2175 Mon Sep 17 00:00:00 2001 From: jiashuoz Date: Mon, 28 Sep 2026 22:27:58 +0800 Subject: [PATCH 3/3] Clarify qualification privileges and final evidence --- docs/microvm-host-privileges.md | 7 ++++--- 1 file changed, 4 insertions(+), 3 deletions(-) diff --git a/docs/microvm-host-privileges.md b/docs/microvm-host-privileges.md index f5f017b..b875580 100644 --- a/docs/microvm-host-privileges.md +++ b/docs/microvm-host-privileges.md @@ -175,7 +175,8 @@ root-operated probes. No global cgroup write grant is needed. Select `TestMicrovmSatisfiesContractOnKVM` for the non-root contract run. The separate `TestMicrovmBootSmokeOnKVM` also reads another UID's -`/proc//ns/net`, which requires a root observer on the qualified host. -Keep that observer evidence separate; do not add tracing capabilities to the -production runner just to satisfy the test. Real non-root dev-cell lifecycle +`/proc//ns/net`, which requires root on the qualified host. Supplemental +probe runs execute the entire driver-level harness as root, including namespace +inspection. Keep that evidence separate from non-root qualification; do not add +tracing capabilities to the production runner just to satisfy the test. Real non-root dev-cell lifecycle checks and externally observed namespace/capability checks cover that boundary.