feat(sending): retention janitors for the sending ledger - #1062
Merged
Merged
Conversation
Nothing deleted the sending-ledger rows every provider-bound send has
written since 1.9.0. Add a bounded, batched River periodic
(sending_ledger_retention, hourly and on start) that removes:
- provider operations with all their attempts 30 days after their stamped
expiry, never while an attempt is reserved/authorized or unexpired, a
runnable River job names the operation, its customer message is
pre-terminal, or a pending notice is bound to it;
- budget day counters once the day and the next have closed (today and
yesterday always kept, on the gate's own DB clock);
- terminal notice event/delivery pairs 30 days after event creation;
- control and external-access audit at their stamped expiry, keeping an
abuse-class pause event while its account exists (the purge reads it).
Batches are LIMIT 1000 FOR UPDATE SKIP LOCKED under a 2s lock_timeout and
30s statement_timeout, capped at 50 per table per run; a failing table is
reported and retried next run. Deletions count on
e2a_janitor_rows_deleted_total{table}; each run on the new
e2a_sending_ledger_retention_runs_total{outcome}. Migrations 126-128 add
the missing predicate indexes concurrently. B8 feedback provenance and
the policy authority tables are untouched.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018tVLxUHk3fqQuq8C3wqyHW
System- and internal-class accounts (prober, monitors, conformance) are never deleted, so the account-lifetime rule kept their feedback correlations forever; the standing prober makes them nearly the whole table. The gate now stamps their customer-purpose correlations with the post-account horizon (30 days) at creation, as it already does for non-customer purposes, and migration 129 backfills existing rows of those accounts with created_at + 30 days. Standard and demo accounts are unchanged. B8's existing janitor removes the expired rows. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018tVLxUHk3fqQuq8C3wqyHW
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Provider-bound sends leave behind operation, reservation, counter, notice, and audit records. This adds bounded hourly retention cleanup, also run at startup, while preserving records needed by in-flight sends and later abuse decisions.
Behavior
Each sweep uses batches of 1,000, at most 50 batches per table/run,
FOR UPDATE SKIP LOCKED, a 2-second lock timeout, and a 30-second statement timeout. Table failures are isolated; remaining backlog continues on later runs. Migrations 126–128 add concurrent indexes. Migration 129 backfills synthetic feedback expiry with 5-second lock waits and 30-second per-statement limits.Deletion counts extend the existing janitor counter. The new
e2a_sending_ledger_retention_runs_total{outcome}exports zero baselines for complete/partial/failed before the first run. Startup observer and periodic-job wiring have regression coverage. Retention and metrics documentation is updated.Review and verification
Independent correctness and adversarial reviews completed, then re-reviewed the fixes. Regressions failed before and passed after correcting superseded-authorization retention and absent initial metric series. Removing the scheduled-job keep guard also caused the expected regression failure; mutation reverted.
Fresh verification on the followup commit:
go build ./...go vet ./internal/sendingpolicy ./internal/telemetry ./cmd/e2ago test -tags integration -p 1 ./internal/sendingpolicy ./internal/telemetry ./internal/sqlguard ./cmd/e2a -count=1/metricsexposes one complete run, zero failed runs, and one operation deletion. Server stopped and database removed afterward.The original implementation also tested retention boundaries, all keep guards, batching, failure isolation, abuse-history preservation, synthetic-account classification, and feedback independence. Its scratch-data query plans showed index-backed batches; those timings are not a production guarantee.
Rollout and remaining risks
The new metric needs companion ops alert coverage in the promotion carrying this release. Monitoring has been prepared separately; it has not been applied or fire-drilled. Failures before the first scrape remain invisible to counter increases.
Migration 129 is a single-transaction backfill. If its backlog cannot fit the statement budget, stop rollout and arrange a reviewed batched backfill; do not remove timeouts. Genuinely stranded reserved/current-authorized attempts remain conservatively retained pending a separate recovery design. No customer 180-day retention cap is introduced.
No public API change; no client regeneration required. No merge or deployment performed.