ateom-microvm: measure the guest agent for GetWorkloadStats - #832
Draft
Tim Bai (baizhenyu) wants to merge 1 commit into
Draft
ateom-microvm: measure the guest agent for GetWorkloadStats#832Tim Bai (baizhenyu) wants to merge 1 commit into
Tim Bai (baizhenyu) wants to merge 1 commit into
Conversation
Tim Bai (baizhenyu)
force-pushed
the
stats-microvm-guest-agent
branch
from
August 10, 2026 20:39
208656c to
c8a2f2c
Compare
Fills in the measurement half of GetWorkloadStats for the micro-VM runtime, so it returns real numbers instead of Unimplemented. The gVisor runtime's half is an independent change; this one does not touch it. The sample comes from inside the guest, not from a host cgroup. The host cgroup here holds cloud-hypervisor, whose memory is the guest RAM allocation it took at boot: near-constant, and near-identical for an idle actor and a saturated one. The numbers that move with the workload are the ones the guest kernel keeps, and the kata-agent's StatsContainer is what can read them out. AgentClient grows that call; it is safe alongside the stdout/stderr forwarding, since ttrpc multiplexes the one connection, which is already what those goroutines rely on. What gets summed is the actor's overlay WORKLOADS, one per container. Their carriers are deliberately absent: a carrier is created and never started (see CreateCarrier), so it runs no process and its cgroup has nothing in it to add. Summing the containers is what turns per-container guest accounting into the one per-actor figure the proto reports. Summing the peaks is an upper bound on the peak of the sum rather than the figure itself -- two containers need not have peaked at the same moment -- and for the single-container actors this runtime mostly serves it is exact. The conversion lives in a new cmd/ateom-microvm/internal/agentstats. It is pure, taking an already-fetched CgroupStats and never talking to a guest, which keeps it testable without a live micro-VM and, unlike the rest of the micro-VM ateom, without the linux build tag. It never fails: every field the agent left out reads as zero, and nil stats -- what the agent answers for a container it has no accounting for -- is a zero sample rather than a panic on a path polled for the life of every workload. Working set subtracts the guest's reclaimable page cache, saturating at zero, and accepts both the v2 name (inactive_file) and the v1 hierarchical one (total_inactive_file). CPU time is divided by 1000: the agent reports nanoseconds, matching the runc stats struct its own is modeled on, and the proto wants microseconds. A container the agent cannot report contributes nothing instead of failing the sample, and for the common way that happens zero is the correct contribution rather than a fallback: a container that has exited took its guest cgroup with it and consumes nothing from here on. Failing an actor's telemetry because one sidecar is gone would be the wrong answer. The sample fails only when no container could be read at all, which is the guest as a whole not answering rather than one container being gone. Status codes follow the NOT_FOUND / FAILED_PRECONDITION split the proto documents, and the handler is a deliberate mirror of the gVisor one so the two runtimes answer a poller the same way. Available and a UID mismatch are both NOT_FOUND: each says the actor is not here, and the caller's worker-to-actor mapping wants re-resolving. No guest to ask yet is the transient FAILED_PRECONDITION -- a poll landing in the boot or the restore, since attribution is retained from the moment the ateom accepts the actor, or one landing mid-teardown. A guest that answers nothing at all is the same code and not Internal: the sandbox going away is a routine state here, and the next CheckpointWorkload turns it into the NOT_FOUND above. The handler takes no lock. That is a stronger constraint than on the gVisor side, where a cgroup path is a constant: the agent client lives in AteomService.running, which lock guards, so reading it from the handler would be a data race whatever the read is for. Hence AteomService.guestStats, an atomic holding the agent client and the container ids, published once the containers are up and cleared by teardownActor before it closes anything. Clearing there rather than alongside the attribution is what keeps a poll landing mid-teardown on the "no numbers right now" path instead of surfacing a closed connection as a failed read. TestGetWorkloadStatsDoesNotTakeLock pins the whole property: it holds s.lock across the call, so a handler that reached for it -- or that looked the agent up in running -- deadlocks. After the read the handler reloads activeActor and compares pointer identity, so a checkpoint plus a fresh run completing underneath it is NOT_FOUND rather than misattributed. One gap worth knowing about: a restore whose post-restore agent dial fails answers FAILED_PRECONDITION for the rest of that activation. Telemetry rides on the connection log forwarding already keeps open, and that dial is best-effort by design (a failed dial must not fail a restore whose actor is already running). A second dial of its own would not help -- whatever kept the agent from answering a 15s retry loop would keep it from answering that one too. The two accumulating fields, memory_peak_bytes and cpu_usage_usec, are scoped to an epoch that does not begin where the gVisor source's does. That is why the proto's epoch note, which lands with the gVisor change, is written per source rather than as one rule, and this is the source it is being careful about. There a restore ends the epoch: runsc delete destroys the sandbox cgroup and its counters with it. Here the counters live in the guest kernel's own memory, so restoreFullScope -- relaunch cloud-hypervisor with --restore, then resume -- brings them back with the guest RAM rather than restarting them, and a FULL restore continues the epoch across what the caller sees as a gap. DATA has no guest to resume and cold-boots, so that scope does restart at zero. DATA_ON_GOLDEN is the case that defeats the obvious detection: it resumes the template's golden guest, so an actor's first sample can begin at whatever the golden had accumulated when it was snapshotted -- an epoch beginning above the value last reported, with no decrease anywhere for a caller to notice. Nothing here can hide that. A caller wanting a lifetime figure has to accumulate one itself and treat a scope transition as a boundary rather than trusting the numbers to announce one. The field mapping is not covered by a test here, since it is a property of the guest kernel and image rather than of this code. It was instead measured against the pinned assets themselves -- the kata-static 4.0.0 kernel and rootfs that hack/microvm-assets/assemble.sh stages and ateom fetches. Which cgroup version the guest gives a container is not a question the guest can answer two ways. Its kernel is 6.18.35, built CONFIG_MEMCG=y with CONFIG_MEMCG_V1 unset, so a v1 memory controller cannot be mounted there at all; its init is systemd 255.4, which reports default-hierarchy=unified when run from that rootfs, and buildVMConfig passes no hierarchy override on the cmdline. Containers therefore get v2, and inactive_file is the spelling that appears. Whether a high-water mark is reported resolves the same way. The agent links cgroups-rs 0.5.1, whose v2 path reads memory.current into usage and memory.peak into max_usage, and passes memory.stat through verbatim. memory.peak has existed since 5.19, so on a 6.18 guest the peak is a real figure rather than the zero this defends against. The same source settles the unit conversion: v2 has no cpuacct controller, so the agent falls through to cpu.stat and multiplies usage_usec by 1000, which is the nanosecond reading the division here turns back into microseconds. The v1 spellings stay anyway. The guest image is a SandboxConfig asset, so a cluster can be pointed at a different one, and a wrong guess there should cost a low working-set figure rather than a failed sample. Still not covered: a live guest answering StatsContainer. That needs a micro-VM worker, which needs a nested-virtualization node. Part of agent-substrate#594
Tim Bai (baizhenyu)
force-pushed
the
stats-microvm-guest-agent
branch
from
August 11, 2026 13:49
c8a2f2c to
bf3a015
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Third and last of the PRs for Phase 0 of #550, and what completes #594. #739 does the gVisor half; this one is independent of it and does not touch it, so the two can land in either order.
Fills in the measurement half of
GetWorkloadStatsfor the micro-VM runtime, so it returns real numbers instead ofUnimplemented.Where the numbers come from
Inside the guest, not from a host cgroup.
The host cgroup here holds cloud-hypervisor, and its memory is the guest RAM allocation it took at boot — near-constant, and near-identical for an idle actor and a saturated one. The numbers that move with the workload are the ones the guest kernel keeps, and the kata-agent's
StatsContaineris what reads them out.AgentClientgrows that call; it is safe alongside the stdout/stderr forwarding, since ttrpc multiplexes the one connection, which is already what those goroutines rely on.What gets summed is the actor's overlay workloads, one per container. Their carriers are deliberately absent: a carrier is created and never started (see
CreateCarrier), so it runs no process and its cgroup has nothing in it to add. Summing the containers is what turns per-container guest accounting into the one per-actor figure the proto reports.Summing the peaks is an upper bound on the peak of the sum rather than the figure itself — two containers need not have peaked at the same moment — and for the single-container actors this runtime mostly serves it is exact. Flagging it in case you'd rather report the largest single peak instead; I think the bound is the more useful of two imperfect answers.
The conversion
New
cmd/ateom-microvm/internal/agentstats. Pure: it takes an already-fetchedCgroupStatsand never talks to a guest, which keeps it testable without a live micro-VM and — unlike the rest of the micro-VM ateom — without thelinuxbuild tag.It never fails. Every field the agent left out reads as zero, and nil stats (what the agent answers for a container it has no accounting for) is a zero sample rather than a panic on a path polled for the life of every workload.
Working set subtracts the guest's reclaimable page cache, saturating at zero, and accepts both the v2 name (
inactive_file) and the v1 hierarchical one (total_inactive_file). CPU time is divided by 1000: the agent reports nanoseconds, matching the runc stats struct its own is modeled on, and the proto wants microseconds.A container the agent cannot report contributes nothing instead of failing the sample, and for the common way that happens zero is the correct contribution rather than a fallback — a container that has exited took its guest cgroup with it and consumes nothing from here on. Failing an actor's telemetry because one sidecar is gone would be the wrong answer. The sample fails only when no container could be read at all, which is the guest as a whole not answering rather than one container being gone.
Status codes
A deliberate mirror of the gVisor handler, so the two runtimes answer a poller the same way.
NOT_FOUND— available, or a UID mismatch. Each says the actor is not here, and the caller's worker-to-actor mapping wants re-resolving.FAILED_PRECONDITION— no guest to ask yet. A poll landing in the boot or the restore (attribution is retained from the moment the ateom accepts the actor), or one landing mid-teardown. A guest that answers nothing at all is this code and notInternal: the sandbox going away is a routine state here, and the nextCheckpointWorkloadturns it into theNOT_FOUNDabove.Locking
The handler takes no lock, and that is a stronger constraint here than on the gVisor side, where the cgroup path is a constant. The agent client lives in
AteomService.running, whichlockguards, so reading it from the handler would be a data race whatever the read is for.Hence
AteomService.guestStats: an atomic holding the agent client and the container ids, published once the containers are up and cleared byteardownActorbefore it closes anything. Clearing it there rather than alongside the attribution is what keeps a poll landing mid-teardown on the "no numbers right now" path instead of surfacing a closed connection as a failed read.TestGetWorkloadStatsDoesNotTakeLockpins the whole property: it holdss.lockacross the call, so a handler that reached for it — or that looked the agent up inrunning— deadlocks. After the read the handler reloadsactiveActorand compares pointer identity, so a checkpoint plus a fresh run completing underneath it isNOT_FOUNDrather than misattributed.Epoch semantics differ from the cgroup source
memory_peak_bytesandcpu_usage_usecaccumulate, and the epoch they accumulate over does not begin where the gVisor source's does. This is the source the proto's epoch note (landing in #739) is being careful about.There, a restore ends the epoch:
runsc deletedestroys the sandbox cgroup and its counters with it. Here the counters live in the guest kernel's own memory, sorestoreFullScope— relaunch cloud-hypervisor with--restore, then resume — brings them back with the guest RAM rather than restarting them, and a FULL restore continues the epoch across what the caller sees as a gap. DATA has no guest to resume and cold-boots, so that scope does restart at zero.DATA_ON_GOLDEN is the case that defeats the obvious detection: it resumes the template's golden guest, so an actor's first sample can begin at whatever the golden had accumulated when it was snapshotted — an epoch beginning above the value last reported, with no decrease anywhere for a caller to notice. Nothing here can hide that. A caller wanting a lifetime figure has to accumulate one itself and treat a scope transition as a boundary.
Proto
Two comment-only edits, on different lines than #739 touches:
STATS_SOURCE_GUEST_AGENTnow says what it counts and what it cannot see — the guest kernel, the agent, and the host VMM process are overhead outside the workload's own containers. The cgroup source is the other way round, since the sandbox's host process is one process and its runtime's overhead is charged along with the workload's. The two sources are not comparable figures and the enum should say so.A gap worth knowing about
A restore whose post-restore agent dial fails answers
FAILED_PRECONDITIONfor the rest of that activation. Telemetry rides on the connection log forwarding already keeps open, and that dial is best-effort by design — a failed dial must not fail a restore whose actor is already running. A second dial of its own would not help: whatever kept the agent from answering a 15s retry loop would keep it from answering that one too.Testing
agentstats: a fifteen-case table overFromCgroupStats(v2 guest, v1total_inactive_file, both keys present, reclaimable cache at and above usage, missingmemory.stat, no peak, sub-microsecond truncation, memory-without-cpu and the converse, nil and empty stats), plusTestSamplePlusfor the summation including saturation.GetWorkloadStats: happy path,TestGetWorkloadStatsSumsContainers,TestGetWorkloadStatsSkipsUnreadableContainer,TestGetWorkloadStatsCountsAnsweredContainer, a seven-case error table, and the no-lock regression test.cmd/ateom-microvmtests are//go:build linuxand were run rather than only compile-checked:go test ./...is green for the whole repo inside agolang:1.26.3container, and theGetWorkloadStatstests pass under-race.Draft, because one thing is measured rather than exercised. The field mapping is a property of the guest kernel and image rather than of this code, and a live guest answering
StatsContainerneeds a micro-VM worker, which needs a nested-virtualization node — neither the GKE dev cluster (c3-standard-8withadvancedMachineFeaturesabsent) nor a local Mac provides one today. So it was measured against the pinned assets themselves — the kata-static 4.0.0 kernel and rootfs thathack/microvm-assets/assemble.shstages and ateom fetches:CONFIG_MEMCG=ywithCONFIG_MEMCG_V1unset, so a v1 memory controller cannot be mounted there at all. Its init is systemd 255.4, which reportsdefault-hierarchy=unifiedwhen run from that rootfs, andbuildVMConfigpasses no hierarchy override on the cmdline. Containers get v2, andinactive_fileis the spelling that appears.memory.currentintousageandmemory.peakintomax_usage, and passesmemory.statthrough verbatim.memory.peakhas existed since 5.19, so on a 6.18 guest the peak is a real figure rather than the zero this defends against.cpuacctcontroller, so the agent falls through tocpu.statand multipliesusage_usecby 1000 — the nanosecond reading the division here turns back into microseconds.The v1 spellings stay anyway. The guest image is a
SandboxConfigasset, so a cluster can be pointed at a different one, and a wrong guess there should cost a low working-set figure rather than a failed sample.Happy to take it out of draft as-is if reviewers are content with the assets being measured instead of a guest being run; otherwise it waits on a nested-virt node.
Part of #594