Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -71,6 +71,14 @@ so authorization, grants and scope checks are one code path.
`agentdrive.scripts.apply_schema`: an empty database takes `schema.sql` alone;
anything else replays pending migrations. Never edit a shipped migration;
fold a new one's end state into `schema.sql` in the same change.
- **Maintenance jobs** — `jobs/`: garbage collection (`jobs/gc.py`, semantics
in `core/gc.py`) and usage maintenance (`jobs/usage_snapshot.py`), on the
cadence in `jobs/schedule.py`. A self-hosted install runs them with
`python -m agentdrive.jobs.scheduler`, which the process supervisor starts
inside the API container when `SCHEDULER_ENABLED=true` (as
`compose.selfhost.yml` sets it) — never as a separate container, so the jobs
that delete content cannot be configured apart from the API. Without it,
deleted content is never reclaimed.
- **Formal model** — `specs/tla/ArtifactGC.tla`, the garbage collector's
safety argument; its TLC configs run in CI.

Expand Down
10 changes: 6 additions & 4 deletions Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -2,10 +2,12 @@
#
# Everything here runs against your own checkout or install: the local stack,
# the generated stylesheet, the model checker, the OpenAPI policy, the image,
# and the maintenance jobs a deployment schedules. How you schedule those jobs
# (cron, a Kubernetes CronJob, your cloud's scheduler) is your deployment's
# business; each one is `python -m agentdrive.jobs.<name>` against the same
# DATABASE_URL and store the API uses.
# and the maintenance jobs a deployment must schedule. compose.selfhost.yml
# runs them on their cadence inside the API container
# (SCHEDULER_ENABLED=true starts `python -m agentdrive.jobs.scheduler`
# there); anything else runs the commands
# `python -m agentdrive.jobs.scheduler --list` prints, on that cadence,
# against the same settings as the API. The targets below run one job now.

SHELL := /usr/bin/env bash

Expand Down
21 changes: 21 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -53,6 +53,27 @@ reads:
- `COMPOSE_PROJECT_NAME` for a second install on the same machine; each
project gets its own containers and volumes.

**Maintenance runs itself.** The API container also runs the maintenance
jobs on the hosted product's schedule (UTC): garbage collection of abandoned
uploads hourly, the full sweep — purging expired deletes and unreferenced
content — daily at 03:00, the same plus an orphan sweep on Sundays at 04:00,
and usage maintenance every 15 minutes. Without them, deleted content is
never reclaimed and storage only grows. The scheduler remembers the last
minute it handled, so a job that came due while the stack was down runs once
when it is back. See the schedule and its state, or run one job now:

```bash
docker compose -f compose.selfhost.yml exec api python -m agentdrive.jobs.scheduler --list
docker compose -f compose.selfhost.yml exec api python -m agentdrive.jobs.scheduler --status
docker compose -f compose.selfhost.yml exec api python -m agentdrive.jobs.scheduler --run gc-daily
```

A job that is already running holds its lock, so a second `--run` of it
reports `"skipped": true` and does nothing. Not using this compose file? Set
`SCHEDULER_ENABLED=true` (and `SCHEDULER_STATE_FILE` to a writable path) on
the API container, or run the commands `--list` prints, on the same
schedule, from your own scheduler with the API's settings.

Data lives in the `agentdrive_pgdata` and `agentdrive_data` volumes.
**Upgrade** with `git pull && docker compose -f compose.selfhost.yml up -d --build`
(without `--build`, `up` keeps running the old image). `docker compose -f
Expand Down
18 changes: 15 additions & 3 deletions compose.selfhost.yml
Original file line number Diff line number Diff line change
Expand Up @@ -5,8 +5,10 @@
# from this file). Three services: `postgres` (plain postgres:16 — schema.sql
# uses only built-in tsvector), `migrate` (the image running apply_schema
# once, after Postgres is healthy), and `api` (the image under
# AUTH_MODE=local with the filesystem object store on a named volume and the
# MCP transport served at /mcp by the sidecar inside the same container).
# AUTH_MODE=local with the filesystem object store on a named volume, the
# MCP transport served at /mcp by the sidecar inside the same container, and
# the job scheduler running garbage collection and usage maintenance there
# too — without it deleted content is never reclaimed).
#
# Settings come from a `.env` beside this file, which every `docker compose`
# command reads:
Expand Down Expand Up @@ -80,6 +82,12 @@ services:
# loopback; the supervisor puts the sidecar into api-key mode.
MCP_PROXY_URL: http://127.0.0.1:8081
PORT: "8080"
# Run the maintenance jobs on their schedule (src/agentdrive/jobs/
# schedule.py) inside this container, so they can never be pointed at
# a different database or store than the API. The last handled minute
# persists on the data volume; a restart catches up what came due.
SCHEDULER_ENABLED: "true"
SCHEDULER_STATE_FILE: /data/.agentdrive-scheduler.json
ports:
- "${AGENTDRIVE_BIND:-127.0.0.1}:${AGENTDRIVE_PORT:-8080}:8080"
volumes:
Expand All @@ -94,8 +102,12 @@ services:
timeout: 5s
retries: 30
# The supervisor gives its children 10 s to exit; Docker's default grace
# period is also 10 s, so give the supervisor room to finish first.
# period is also 10 s, so give the supervisor room to finish first. A
# running job is terminated with its process group; its advisory lock is
# released with its database connection and the next run resumes.
stop_grace_period: 20s
# A real PID 1 that reaps orphaned processes.
init: true
restart: unless-stopped

volumes:
Expand Down
37 changes: 37 additions & 0 deletions scripts/selfhost-smoke.sh
Original file line number Diff line number Diff line change
Expand Up @@ -168,6 +168,43 @@ else
fi
rm -f "$AFTER_BODY"

echo "== the maintenance jobs run on this install"
# Without the scheduler, deleted content is never reclaimed. It runs inside
# the api container; its loop records the minute it last handled, so a state
# file that appears proves the loop itself is alive (it ticks every minute).
LOOP_ALIVE=""
for _ in $(seq 1 45); do
if $COMPOSE exec -T api python -m agentdrive.jobs.scheduler --status > /dev/null; then LOOP_ALIVE=1; break; fi
sleep 2
done
if [ -n "$LOOP_ALIVE" ]; then ok "the scheduler loop is ticking"; else bad "the scheduler loop never recorded a tick"; fi
# Then each scheduled job runs once, now, against this install's database and
# store — the rows the steps above created included; the sweeps' age rules
# mean nothing that young is deleted, which the test suite proves on the same
# store. A run that met the loop's own run of the same job reports
# `"skipped": true` (lock held); that is retried, never counted as a pass.
JOBS=$($COMPOSE exec -T api python -m agentdrive.jobs.scheduler --list | grep -v '^ ' | awk '{print $1}')
RUN_OUT=$(mktemp)
for name in gc-hourly gc-daily gc-weekly usage-snapshot; do
if ! echo "$JOBS" | grep -qx "$name"; then bad "--list does not show $name"; continue; fi
result=""
for _ in 1 2 3; do
if $COMPOSE exec -T api python -m agentdrive.jobs.scheduler --run "$name" > "$RUN_OUT" 2>&1; then
if grep -q '"skipped": true' "$RUN_OUT"; then result="skipped"; sleep 5; continue; fi
if grep -q "Traceback" "$RUN_OUT"; then result="traceback"; break; fi
result="ok"; break
fi
result="failed"; break
done
if [ "$result" = "ok" ]; then
ok "scheduled job $name exits 0"
else
bad "scheduled job $name: $result"
sed 's/^/ /' "$RUN_OUT" | tail -20
fi
done
rm -f "$RUN_OUT"

echo "== the container logs for THIS run are clean"
LOGS=$($COMPOSE logs --no-color --since "$START" api 2>/dev/null)
if echo "$LOGS" | grep -q "Traceback"; then bad "api log has a traceback"; else ok "api log has no traceback"; fi
Expand Down
10 changes: 6 additions & 4 deletions src/agentdrive/jobs/gc.py
Original file line number Diff line number Diff line change
@@ -1,11 +1,13 @@
"""Cloud Run Job entrypoint for the GC sweeper.
"""Job entrypoint for the GC sweeper.

Thin wrapper around ``GCSweeper(...).run()``. Cloud Scheduler invokes
Thin wrapper around ``GCSweeper(...).run()``, run as
``python -m agentdrive.jobs.gc`` hourly with ``--sessions-only``, daily at
03:00 UTC for the full sweep (session reconciliation + transfer cleanup +
purge + CAS mark-sweep + scratch sweep), and weekly on Sunday 04:00 UTC with
``--orphan-sweep`` appended. Operator runbook lives in ``deploy/README.md``;
the sweep semantics live in ``agentdrive.core.gc``.
``--orphan-sweep`` appended. That cadence is ``agentdrive.jobs.schedule``:
a self-hosted install runs it with ``python -m agentdrive.jobs.scheduler``,
a hosted deployment with its platform's scheduler. The sweep semantics live
in ``agentdrive.core.gc``.

CLI (the scheduled contract — do not change argument meanings):
--dry-run Preview without persisting: all database work runs inside
Expand Down
72 changes: 72 additions & 0 deletions src/agentdrive/jobs/schedule.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,72 @@
"""When the maintenance jobs run: the one record of the cadence.

A deployment must run these, or it never reclaims storage: deleted artifacts
stay in the database and their bytes stay in the store, abandoned uploads keep
their reserved bytes, and usage is never finalized. A self-hosted install runs
them with `python -m agentdrive.jobs.scheduler`, which the process
supervisor starts inside the API container when `SCHEDULER_ENABLED=true`
(`compose.selfhost.yml` sets it); a deployment with its own scheduler (cron, a
Kubernetes CronJob, a cloud scheduler) runs the same commands on the same
cadence, listed by `python -m agentdrive.jobs.scheduler --list`.

Times are UTC, in five-field cron syntax. Each job is
`python -m <module> <args>` against the same settings as the API.
"""

from __future__ import annotations

from dataclasses import dataclass


@dataclass(frozen=True)
class ScheduledJob:
name: str
cron: str
module: str
args: tuple[str, ...]
# The scheduler kills a run that outlives this. The GC stops itself at
# `GCSweeper.HARD_TIMEOUT_S` (50 min), so its kill comes after that.
timeout_s: int
description: str = ""


SCHEDULE: tuple[ScheduledJob, ...] = (
ScheduledJob(
name="gc-hourly",
cron="0 * * * *",
module="agentdrive.jobs.gc",
args=("--sessions-only",),
timeout_s=3600,
description=(
"Session phases only: terminalize abandoned uploads and release "
"their reserved bytes. No object-store listing."
),
),
ScheduledJob(
name="gc-daily",
cron="0 3 * * *",
module="agentdrive.jobs.gc",
args=(),
timeout_s=3600,
description=(
"Full sweep: purge expired soft-deletes, mark-sweep unreferenced "
"content blobs, clean up scratch."
),
),
ScheduledJob(
name="gc-weekly",
cron="0 4 * * 0",
module="agentdrive.jobs.gc",
args=("--orphan-sweep",),
timeout_s=3600,
description="The full sweep plus the orphan sweep over purged drives' prefixes.",
),
ScheduledJob(
name="usage-snapshot",
cron="*/15 * * * *",
module="agentdrive.jobs.usage_snapshot",
args=("--no-notify",),
timeout_s=900,
description="Finalize usage reservations and prune bounded usage history.",
),
)
Loading
Loading