Skip to content

Report the actual cause when a sandbox bundle fails mid-session - #8

Merged
josephschorr merged 1 commit into
mainfrom
fix/bundle-midsession-failure-message
Sep 29, 2026
Merged

josephschorr merged 1 commit into
mainfrom
fix/bundle-midsession-failure-message

Conversation

@josephschorr

@josephschorr josephschorr commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Motivation

A production session failed with:

BundleFailed: sandbox bundle(s) [<session>-<bundle>] did not become Ready within 8m0s of provisioning start; the pod(s) scheduled but never became Ready, so the likely cause is a slow/cold image pull or a crashing container

Every claim in that message pointed the wrong way. The cluster events showed the bundle's images were already cached on the node, all containers started within seconds, and the session's own status recorded BundlesReady=True 13 seconds after creation. The session then ran normally for ~35 minutes — until the kernel OOM-killed the bundle pod's cgroup (a multi-GB compile inside a 4 Gi bundle), taking down the sandbox container's PID 1. The pod went NotReady, and two seconds later the session was declared BundleFailed with the provisioning message above.

The wording comes from the bundle-ready deadline backstop in the AgentSession reconciler: the deadline is measured from the bundle SpiceboxSession's CreationTimestamp, so once a bundle is older than 8 minutes, any not-Ready observation — including a mid-session container death — trips it on the first reconcile pass and gets described as a boot/provisioning failure. Operators chasing an image pull for an OOM kill is exactly the failure mode the earlier unschedulable-message fix (bundle_timeout_message_test.go) was written to prevent, one branch over.

Summary of changes

All in pkg/controllers/agentsession/:

  • The deadline branch now consults the session's durable BundlesReady condition. When boot had already succeeded, the terminal message reads sandbox bundle(s) [X] became unhealthy after running (Ready since <time>): … instead of the provisioning/image-pull framing. The condition Reason stays the exact BundleFailed token, so lifecycle's transient-failure classification (recoverable on a follow-up message) is unchanged.
  • New firstContainerTermination helper (beside firstSchedulingStall, same fail-safe shape): reads the not-ready pods' container statuses — a currently-terminated container, or a restarting container's LastTerminationState — and skips exit-code-0 Completed containers. When kubelet has recorded the death, the message names it instead of guessing, e.g. container "sandbox" in pod "<session>-<bundle>-pod" terminated: OOMKilled (exit code 137).
  • The never-became-Ready path gets the same honesty: it keeps the provisioning-deadline framing, but a recorded crash (exit code 1, CrashLoopBackOff) replaces the "likely a slow/cold image pull" guess.
  • Three new reconcile-level tests in bundle_midsession_unready_message_test.go, following the existing fake-client fixtures from bundle_epoch_reconcile_test.go / bundle_unschedulable_surface_test.go. Each was written first and observed failing with the old message.

Out of scope, deliberately: whether a mid-session bundle death should be terminal at all (the OOM-killed container was restarting and would likely have recovered; today the first not-Ready pass fails the session with zero grace). That is a behavior change, not a message fix, and is left for a follow-up discussion.

Alternatives considered

  • Per-bundle "ever ready" marker in status.bundleSessions (a FirstReadyAt on ResolvedBundle): more precise in the compound case where one bundle is a retry replacement mid-provisioning while another dies, but requires a CRD field + regeneration for a message fix. The session-level BundlesReady condition is already durable, already written, and correct for the observed failure class; the compound case falls back to the old (then partially-accurate) message.
  • Detecting OOM from node OOMKilling events: kubelet's container LastTerminationState carries the same fact (OOMKilled/137) on an object the reconciler already reads, without event-list plumbing.

Provenance

  • Author (person, or model and version):
  • Harness or tooling, with version:
  • Person who read the diff:

Ship gate

All three suites are green in this PR's CI run
(actions/runs/36625295328):

  • mage test:unit — CI unit job: pass (35m28s)
  • mage test:integration — CI integration job: pass (42m58s)
  • mage test:e2e — CI e2e job: pass (29m36s)

Results:

$ gh pr checks 8
e2e          pass  29m36s
format       pass  49s
frontend     pass  7m46s
integration  pass  42m58s
unit         pass  35m28s

A local mage test:unit also ran (553 packages ok) with one environmental
flake — TestPickStablePort_SkipsOccupiedPort, which assumes base+1 is free
after occupying base; re-run in isolation per the AGENTS.md flake protocol
(go test -race -count=1 -run 'TestPickStablePort' ./cmd/oap/internal/desktop/)
and green. Unrelated to this diff, and CI's unit job passed it outright.

Regeneration

  • n/a — no +kubebuilder:rbac marker or CRD-shaping field changed
  • n/a — nothing under config/** changed
  • n/a — no cobra command or CRD schema changed
  • Ran mage fmt:check: "All matched files use the correct format." (347 files)

Coverage

  • No new user-visible behavior: this corrects the wording of an existing failure message. The three message variants are pinned by reconcile-level tests; a bronzethread bundle cannot express a mid-session pod death (the scripted harness has no real kubelet to kill a container under a running session).
  • Touches no authorization, SpiceDB schema, or approver-model code.

Before requesting review

  • The PR holds a single change.
  • A person has read every line of the diff.

A bundle that became Ready and later turned unhealthy mid-session (an
OOM-killed container, an evicted pod) lands in the bundle-ready deadline
backstop, because the deadline is measured from the bundle's
CreationTimestamp and mid-session it is trivially exceeded on the first
not-ready pass. The terminal message then blamed a slow/cold image pull;
in a live incident the images were cached, the bundle had been Ready in
13 seconds, and the real cause was a cgroup OOM kill 35 minutes in.

The deadline branch now consults the session's durable BundlesReady
condition: when boot already succeeded, the message says the bundle
became unhealthy after running instead of using provisioning framing.
When kubelet has recorded the container's termination (OOMKilled, exit
code), both the mid-session and boot-time messages name that observed
cause instead of guessing. The BundleFailed reason token is unchanged,
so lifecycle's transient-failure classification is preserved.
@josephschorr
josephschorr merged commit eb6b559 into main Sep 29, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant