Report the actual cause when a sandbox bundle fails mid-session - #8
Merged
Merged
Conversation
A bundle that became Ready and later turned unhealthy mid-session (an OOM-killed container, an evicted pod) lands in the bundle-ready deadline backstop, because the deadline is measured from the bundle's CreationTimestamp and mid-session it is trivially exceeded on the first not-ready pass. The terminal message then blamed a slow/cold image pull; in a live incident the images were cached, the bundle had been Ready in 13 seconds, and the real cause was a cgroup OOM kill 35 minutes in. The deadline branch now consults the session's durable BundlesReady condition: when boot already succeeded, the message says the bundle became unhealthy after running instead of using provisioning framing. When kubelet has recorded the container's termination (OOMKilled, exit code), both the mid-session and boot-time messages name that observed cause instead of guessing. The BundleFailed reason token is unchanged, so lifecycle's transient-failure classification is preserved.
9 of 11 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
A production session failed with:
Every claim in that message pointed the wrong way. The cluster events showed the bundle's images were already cached on the node, all containers started within seconds, and the session's own status recorded
BundlesReady=True13 seconds after creation. The session then ran normally for ~35 minutes — until the kernel OOM-killed the bundle pod's cgroup (a multi-GB compile inside a 4 Gi bundle), taking down the sandbox container's PID 1. The pod went NotReady, and two seconds later the session was declared BundleFailed with the provisioning message above.The wording comes from the bundle-ready deadline backstop in the AgentSession reconciler: the deadline is measured from the bundle SpiceboxSession's
CreationTimestamp, so once a bundle is older than 8 minutes, any not-Ready observation — including a mid-session container death — trips it on the first reconcile pass and gets described as a boot/provisioning failure. Operators chasing an image pull for an OOM kill is exactly the failure mode the earlier unschedulable-message fix (bundle_timeout_message_test.go) was written to prevent, one branch over.Summary of changes
All in
pkg/controllers/agentsession/:BundlesReadycondition. When boot had already succeeded, the terminal message readssandbox bundle(s) [X] became unhealthy after running (Ready since <time>): …instead of the provisioning/image-pull framing. The condition Reason stays the exactBundleFailedtoken, so lifecycle's transient-failure classification (recoverable on a follow-up message) is unchanged.firstContainerTerminationhelper (besidefirstSchedulingStall, same fail-safe shape): reads the not-ready pods' container statuses — a currently-terminated container, or a restarting container'sLastTerminationState— and skips exit-code-0Completedcontainers. When kubelet has recorded the death, the message names it instead of guessing, e.g.container "sandbox" in pod "<session>-<bundle>-pod" terminated: OOMKilled (exit code 137).exit code 1, CrashLoopBackOff) replaces the "likely a slow/cold image pull" guess.bundle_midsession_unready_message_test.go, following the existing fake-client fixtures frombundle_epoch_reconcile_test.go/bundle_unschedulable_surface_test.go. Each was written first and observed failing with the old message.Out of scope, deliberately: whether a mid-session bundle death should be terminal at all (the OOM-killed container was restarting and would likely have recovered; today the first not-Ready pass fails the session with zero grace). That is a behavior change, not a message fix, and is left for a follow-up discussion.
Alternatives considered
status.bundleSessions(aFirstReadyAtonResolvedBundle): more precise in the compound case where one bundle is a retry replacement mid-provisioning while another dies, but requires a CRD field + regeneration for a message fix. The session-levelBundlesReadycondition is already durable, already written, and correct for the observed failure class; the compound case falls back to the old (then partially-accurate) message.OOMKillingevents: kubelet's containerLastTerminationStatecarries the same fact (OOMKilled/137) on an object the reconciler already reads, without event-list plumbing.Provenance
Ship gate
All three suites are green in this PR's CI run
(actions/runs/36625295328):
mage test:unit— CIunitjob: pass (35m28s)mage test:integration— CIintegrationjob: pass (42m58s)mage test:e2e— CIe2ejob: pass (29m36s)Results:
A local
mage test:unitalso ran (553 packages ok) with one environmentalflake —
TestPickStablePort_SkipsOccupiedPort, which assumes base+1 is freeafter occupying base; re-run in isolation per the AGENTS.md flake protocol
(
go test -race -count=1 -run 'TestPickStablePort' ./cmd/oap/internal/desktop/)and green. Unrelated to this diff, and CI's unit job passed it outright.
Regeneration
+kubebuilder:rbacmarker or CRD-shaping field changedconfig/**changedmage fmt:check: "All matched files use the correct format." (347 files)Coverage
Before requesting review