Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
37 changes: 27 additions & 10 deletions docs/microvm-host-privileges.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,14 +12,14 @@ everything not on this list.

## Hardware qualification limitation

The capability list below is the intended non-root contract, not proof that the
upstream jailer can run with it. Firecracker jailer 1.17.0 unconditionally writes
`cgroup.subtree_control` in every ancestor up to the host cgroup root when a
`--cgroup` property is supplied. A systemd-delegated subtree does not authorize
those ancestor writes, so the non-root launch fails before the API socket opens.
Do not grant the runner write access to the host cgroup root to work around it.
A scoped privileged launch path or a jailer change needs separate qualification.
Root-run Phase A feasibility tests do not establish this non-root contract.
The capability list below requires a separately qualified jailer. Stock
Firecracker jailer 1.17.0 writes outside a delegated cgroup subtree, gives
away directory ownership before device setup is complete, and retains ambient
capabilities across a non-root-to-non-root UID change. The Rainier Cloud
`infra/microvm/jailer` patch addresses these boundaries and carries its own
source/compiler/binary pins and real launcher tests. Use the reviewed pin;
do not grant global cgroup write access or broad DAC powers as a workaround.
Root-run feasibility alone does not establish this non-root contract.

The host must also prepare `/run/netns` for namespace mount-point creation:
root-owned, runner-group-writable, mode `0770` on a host dedicated to that
Expand All @@ -34,6 +34,7 @@ is transferred to the VM. A resume temporarily reclaims the inode using
| | What it is for |
|---|---|
| `CAP_CHOWN` | Give each VM's jail directory, control socket and disk images to that VM's own uid. |
| `CAP_KILL` | Signal and reap VMM processes owned by different per-VM uids during stop, failed launch and delete. |
| `CAP_SETGID` | Let the jailer drop the VMM to its per-VM gid. |
| `CAP_SETUID` | Let the jailer drop the VMM to its per-VM uid. |
| `CAP_NET_ADMIN` | Create each session's network namespace, veth pair and TAP device, and install its `nftables` ruleset. |
Expand Down Expand Up @@ -77,8 +78,8 @@ thing to discover at startup and not at the first suspend.
User=rainier
Group=rainier
SupplementaryGroups=kvm
AmbientCapabilities=CAP_CHOWN CAP_SETGID CAP_SETUID CAP_NET_ADMIN CAP_SYS_CHROOT CAP_SYS_ADMIN CAP_MKNOD
CapabilityBoundingSet=CAP_CHOWN CAP_SETGID CAP_SETUID CAP_NET_ADMIN CAP_SYS_CHROOT CAP_SYS_ADMIN CAP_MKNOD
AmbientCapabilities=CAP_CHOWN CAP_KILL CAP_SETGID CAP_SETUID CAP_NET_ADMIN CAP_SYS_CHROOT CAP_SYS_ADMIN CAP_MKNOD
CapabilityBoundingSet=CAP_CHOWN CAP_KILL CAP_SETGID CAP_SETUID CAP_NET_ADMIN CAP_SYS_CHROOT CAP_SYS_ADMIN CAP_MKNOD
Delegate=yes
ExecStart=/usr/local/bin/runnerd --driver=microvm ...
```
Expand Down Expand Up @@ -163,3 +164,19 @@ whichever filesystem `/tmp` happens to be. That a create on XFS with
`reflink=1` takes microseconds and shares the image's extents — and that a
guest writing to its root leaves the shared image untouched on a real block
device rather than on a test's byte comparison — is host qualification too.

## Non-root hardware tests

The opt-in KVM harness accepts `RAINIER_MICROVM_TEST_CGROUP_PARENT` for the
service's delegated cgroup path (relative to `/sys/fs/cgroup`). Keep the test
process in a child leaf before enabling controllers, exactly as the runner's
service launcher does. The default remains `rainier-kvmtest` for existing
root-operated probes. No global cgroup write grant is needed.

Select `TestMicrovmSatisfiesContractOnKVM` for the non-root contract run. The
separate `TestMicrovmBootSmokeOnKVM` also reads another UID's
`/proc/<pid>/ns/net`, which requires root on the qualified host. Supplemental
probe runs execute the entire driver-level harness as root, including namespace
inspection. Keep that evidence separate from non-root qualification; do not add
tracing capabilities to the production runner just to satisfy the test. Real non-root dev-cell lifecycle
checks and externally observed namespace/capability checks cover that boundary.
6 changes: 5 additions & 1 deletion internal/driver/microvm_caps.go
Original file line number Diff line number Diff line change
Expand Up @@ -50,6 +50,10 @@ type hostCapability struct {
// microvmCapabilities is the whole list. The bit numbers are
// <linux/capability.h>'s.
var microvmCapabilities = []hostCapability{
{
Name: "CAP_KILL", Bit: 5,
Why: "signal and reap VMM processes running as distinct per-VM uids",
},
{
Name: "CAP_CHOWN", Bit: 0,
Why: "give each VM's jail directory, control socket and disk images to that VM's own uid (chownJailPath)",
Expand Down Expand Up @@ -111,7 +115,7 @@ func capabilityNames() string {
//
// It fails closed and names the FIRST thing that is missing rather than
// summarising: an operator fixing a unit file wants the next thing to add,
// and a list of seven is a list nobody reads to the end.
// and a list of eight is a list nobody reads to the end.
func checkMicrovmPrivileges(opts MicrovmOpts) error {
effective, err := effectiveCapabilities()
if err != nil {
Expand Down
2 changes: 1 addition & 1 deletion internal/driver/microvm_caps_linux_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@
// namespace and no jailer on a Mac either. Running these there would assert
// that a capability check fails for the reason the platform does not have
// capabilities, which is not what any of them is about, and it is what made
// seven subtests fail on every developer machine. The portable half — the
// eight subtests fail on every developer machine. The portable half — the
// requirements text, the host preparation — is in microvm_caps_test.go and
// runs everywhere.
package driver
Expand Down
9 changes: 9 additions & 0 deletions internal/driver/microvm_caps_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -97,3 +97,12 @@ func TestTheHostPreparationTextNamesBothHalves(t *testing.T) {
}
}
}

func TestMicrovmRequiresCrossUIDSignalCapability(t *testing.T) {
for _, capability := range microvmCapabilities {
if capability.Name == "CAP_KILL" && capability.Bit == 5 {
return
}
}
t.Fatal("runner must require CAP_KILL before accepting VMs it cannot stop")
}
25 changes: 16 additions & 9 deletions internal/driver/microvm_kvm_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -121,6 +121,8 @@ const (
// kvmTimingsEnv, when set, is a file the harness writes its measurements
// to as JSON, for the runbook's evidence table to cite.
kvmTimingsEnv = "RAINIER_MICROVM_TEST_TIMINGS"
// Optional delegated parent for a non-root systemd test service.
kvmCgroupEnv = "RAINIER_MICROVM_TEST_CGROUP_PARENT"

kvmDefaultKernel = "vmlinux"
kvmDefaultRootfs = "rootfs.ext4"
Expand Down Expand Up @@ -148,13 +150,14 @@ const (

// kvmFixture is the host, resolved, once every precondition has held.
type kvmFixture struct {
images string
kernel string
rootfs string
stateParent string
connectWait time.Duration
proxyAddr string
proxyPort int
images string
kernel string
rootfs string
stateParent string
connectWait time.Duration
proxyAddr string
proxyPort int
cgroupParent string
}

// requireKVMHost skips unless this machine can actually boot a microVM, and
Expand Down Expand Up @@ -242,6 +245,10 @@ func requireKVMHost(t *testing.T) kvmFixture {
stateParent: os.Getenv(kvmStateParentEnv),
connectWait: kvmConnectWait,
}
fx.cgroupParent = os.Getenv(kvmCgroupEnv)
if fx.cgroupParent == "" {
fx.cgroupParent = kvmCgroupChild
}
if raw := os.Getenv(kvmConnectEnv); raw != "" {
d, err := time.ParseDuration(raw)
if err != nil {
Expand Down Expand Up @@ -351,7 +358,7 @@ func newKVMMicrovm(t *testing.T, fx kvmFixture) (*Microvm, *kvmRunner, CloneMeth
EgressProxyAddr: fx.proxyAddr,
EgressProxyPort: fx.proxyPort,

Jail: JailOpts{CgroupParent: kvmCgroupChild},
Jail: JailOpts{CgroupParent: fx.cgroupParent},
})
if err != nil {
// A Fatal and not a Skip: every precondition this constructor checks
Expand Down Expand Up @@ -1222,7 +1229,7 @@ func TestMicrovmBootSmokeOnKVM(t *testing.T) {
if err != nil {
t.Fatalf("read /proc/%d/cgroup: %v", pid, err)
}
if want := "/" + kvmCgroupChild + "/" + h.ID; !strings.Contains(string(procCgroup), want) {
if want := "/" + fx.cgroupParent + "/" + h.ID; !strings.Contains(string(procCgroup), want) {
t.Fatalf("ADR-0003 §4.6: the VMM is not in its own cgroup (%q is not in %q)", want, strings.TrimSpace(string(procCgroup)))
}

Expand Down
5 changes: 4 additions & 1 deletion scripts/microvm-kvm-test.sh
Original file line number Diff line number Diff line change
Expand Up @@ -29,6 +29,8 @@
# Put it on a filesystem that reflinks (XFS) or
# the harness logs that every create is a full
# copy of the environment image.
# RAINIER_MICROVM_TEST_CGROUP_PARENT optional delegated service cgroup;
# default rainier-kvmtest (root-operated harness).
# RAINIER_MICROVM_TEST_PROXY ip:port of the egress proxy, the one
# host-side destination a slot's firewall
# allows. Empty means no exception, which is a
Expand Down Expand Up @@ -140,7 +142,7 @@ BINS
command -v "$go_bin" >/dev/null 2>&1 ||
note_missing "\`$go_bin\` is not on PATH. Set GO to the go binary, or install one."

# The capability set, from the same seven bits internal/driver checks. A host
# The capability set, from the same eight bits internal/driver checks. A host
# with no /proc is not a microVM host and will have failed above; the guard is
# here so this script's own test can run somewhere else.
if [[ -r "$proc_status" ]]; then
Expand All @@ -150,6 +152,7 @@ if [[ -r "$proc_status" ]]; then
if (( ( 0x$capeff >> bit ) & 1 )); then continue; fi
note_missing "this process does not hold $name, needed to ${why}. Re-run under \`sudo -E\`, or give runnerd's user the ambient set (see MicrovmHostRequirements)."
done <<'CAPS'
5 CAP_KILL signal VMM processes owned by per-VM uids
0 CAP_CHOWN give each VM's jail, control socket and disk images to that VM's own uid
6 CAP_SETGID let the jailer drop the VMM to its per-VM gid
7 CAP_SETUID let the jailer drop the VMM to its per-VM uid
Expand Down
Loading