Skip to content

Runner capacity: replace the fixed slot count with resource-aware admission and idle auto-suspend #85

Description

@jiashuoz

What happened

A hosted runner (one 4 vCPU / 16 GB machine) refused new sessions with "no free capacity" while only three agents were actually working. Fourteen other sessions had finished hours earlier; their sandboxes were still up, each holding a slot. runnerd starts with -slots (default 16), and every session whose sandbox is up occupies one until it is stopped or deleted, whether or not its child process has exited or anyone is attached. Stopping the idle sessions by hand freed the runner immediately.

Why a fixed count is the wrong admission control

It answers "can this runner take one more session?" badly in both directions: it refuses when the machine is idle, and it would admit a sixteenth active agent onto a machine that saturates at about three. What actually runs out, in order: CPU (~3 active coding agents per 4 vCPU), memory (1–2 GB per active agent), disk (workspaces, the only thing idle sessions consume in quantity), then kernel/daemon limits (netns, fds, container count) far behind.

Proposal to discuss and finalise

The lifecycle vocabulary already has the states: running, suspended_warm, suspended_cold, queued.

  1. Admission on resources, not a count. Admit while there is memory headroom for one more active sandbox and disk for its workspace; otherwise the session is queued, not refused.
  2. Idle sessions suspend themselves. Child exited, or no attachment and no activity for a configurable period → suspended_warm (container stopped, files kept, resumes on attach in seconds). The definition of "idle" should be the same one the activity signal uses.
  3. Disk bounds the idle set. When disk is short, the oldest warm sessions go suspended_cold (snapshotted off the runner). Cold resume is slower; that is the honest trade.
  4. One hard cap stays as a safety valve, set high, for kernel and daemon limits.
  5. rainier status says what is held: active, warm, cold, and headroom, rather than "no free capacity".

Near-term relief (agreed, small)

  • Raise the default the hosted runner startup passes to something large, so the count is a safety valve rather than the effective limit.
  • Idle auto-stop in runnerd: a sandbox whose child has exited and that has had no attachment for N minutes (configurable) is stopped, files kept, exactly what rainier stop does today.

Open questions

  • Who owns the policy: the runner (local resource view) or the cell's placement loop (fleet view)? Probably admission locally, cold-migration centrally.
  • Idle thresholds for warm and cold, and whether they are per environment.
  • How queued sessions are surfaced to the CLI and browser, and whether they time out.
  • Interaction with the controller lease: a suspended session has no controller; resume should not auto-claim.

Acceptance for the design

A written contract for the four states and their transitions, the admission rule, the status report, and tests that drive a runner to each limit and observe the right transition rather than a refusal.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions