Skip to content

feat(sending): retention janitors for the sending ledger - #1062

Merged
jiashuoz merged 3 commits into
mainfrom
feat/sending-ledger-janitors
Sep 30, 2026
Merged

jiashuoz merged 3 commits into
mainfrom
feat/sending-ledger-janitors

Conversation

@jiashuoz

@jiashuoz jiashuoz commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Provider-bound sends leave behind operation, reservation, counter, notice, and audit records. This adds bounded hourly retention cleanup, also run at startup, while preserving records needed by in-flight sends and later abuse decisions.

Behavior

  • Delete expired operations with their expired attempts. Keep reserved attempts, current/future provider authorizations, live River jobs, pre-terminal messages, and pending notices. Superseded authorizations cannot redeem and no longer keep an otherwise terminal operation forever.
  • Keep today's and yesterday's budget counters; delete older closed-day counters without refunding current exposure.
  • Delete terminal notice events/deliveries after 30 days. Delete control/access audit at its stamped expiry, except abuse history needed while an account still exists for purge-time identity holds.
  • Keep customer feedback provenance independent of operation cleanup. System/internal synthetic traffic gets a 30-day feedback horizon; customer account-lifetime retention is unchanged.
  • Never delete append-only policy, attestation, registry, or grandfathering authority.

Each sweep uses batches of 1,000, at most 50 batches per table/run, FOR UPDATE SKIP LOCKED, a 2-second lock timeout, and a 30-second statement timeout. Table failures are isolated; remaining backlog continues on later runs. Migrations 126–128 add concurrent indexes. Migration 129 backfills synthetic feedback expiry with 5-second lock waits and 30-second per-statement limits.

Deletion counts extend the existing janitor counter. The new e2a_sending_ledger_retention_runs_total{outcome} exports zero baselines for complete/partial/failed before the first run. Startup observer and periodic-job wiring have regression coverage. Retention and metrics documentation is updated.

Review and verification

Independent correctness and adversarial reviews completed, then re-reviewed the fixes. Regressions failed before and passed after correcting superseded-authorization retention and absent initial metric series. Removing the scheduled-job keep guard also caused the expected regression failure; mutation reverted.

Fresh verification on the followup commit:

  • go build ./...
  • go vet ./internal/sendingpolicy ./internal/telemetry ./cmd/e2a
  • go test -tags integration -p 1 ./internal/sendingpolicy ./internal/telemetry ./internal/sqlguard ./cmd/e2a -count=1
  • Real built-server smoke with a disposable local Postgres database: migrations and HTTP health pass; River startup cleanup deletes an expired synthetic operation and retains a fresh one; /metrics exposes one complete run, zero failed runs, and one operation deletion. Server stopped and database removed afterward.
  • Initial PR CI passed; new checks run on this revision.

The original implementation also tested retention boundaries, all keep guards, batching, failure isolation, abuse-history preservation, synthetic-account classification, and feedback independence. Its scratch-data query plans showed index-backed batches; those timings are not a production guarantee.

Rollout and remaining risks

The new metric needs companion ops alert coverage in the promotion carrying this release. Monitoring has been prepared separately; it has not been applied or fire-drilled. Failures before the first scrape remain invisible to counter increases.

Migration 129 is a single-transaction backfill. If its backlog cannot fit the statement budget, stop rollout and arrange a reviewed batched backfill; do not remove timeouts. Genuinely stranded reserved/current-authorized attempts remain conservatively retained pending a separate recovery design. No customer 180-day retention cap is introduced.

No public API change; no client regeneration required. No merge or deployment performed.

jiashuoz and others added 3 commits September 30, 2026 00:21
Nothing deleted the sending-ledger rows every provider-bound send has
written since 1.9.0. Add a bounded, batched River periodic
(sending_ledger_retention, hourly and on start) that removes:

- provider operations with all their attempts 30 days after their stamped
  expiry, never while an attempt is reserved/authorized or unexpired, a
  runnable River job names the operation, its customer message is
  pre-terminal, or a pending notice is bound to it;
- budget day counters once the day and the next have closed (today and
  yesterday always kept, on the gate's own DB clock);
- terminal notice event/delivery pairs 30 days after event creation;
- control and external-access audit at their stamped expiry, keeping an
  abuse-class pause event while its account exists (the purge reads it).

Batches are LIMIT 1000 FOR UPDATE SKIP LOCKED under a 2s lock_timeout and
30s statement_timeout, capped at 50 per table per run; a failing table is
reported and retried next run. Deletions count on
e2a_janitor_rows_deleted_total{table}; each run on the new
e2a_sending_ledger_retention_runs_total{outcome}. Migrations 126-128 add
the missing predicate indexes concurrently. B8 feedback provenance and
the policy authority tables are untouched.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018tVLxUHk3fqQuq8C3wqyHW
System- and internal-class accounts (prober, monitors, conformance) are
never deleted, so the account-lifetime rule kept their feedback
correlations forever; the standing prober makes them nearly the whole
table. The gate now stamps their customer-purpose correlations with the
post-account horizon (30 days) at creation, as it already does for
non-customer purposes, and migration 129 backfills existing rows of those
accounts with created_at + 30 days. Standard and demo accounts are
unchanged. B8's existing janitor removes the expired rows.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018tVLxUHk3fqQuq8C3wqyHW
@jiashuoz
jiashuoz merged commit ed7f060 into main Sep 30, 2026
29 checks passed
@jiashuoz
jiashuoz deleted the feat/sending-ledger-janitors branch September 30, 2026 04:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant