From 76423a84b8522cd791a3a87a1d7a9f0efb69d23d Mon Sep 17 00:00:00 2001 From: jiashuoz Date: Tue, 29 Sep 2026 11:16:53 +0800 Subject: [PATCH 1/8] docs(design): generic feature packs, neutral event vocabulary, declarative custom features Splits the built-in features into namespaced core/email/brand packs enabled per tenant, adds a neutral delivery.sent event with a read-side view over content.sent, product-declared resource kinds, channels and custom types with kind-driven redaction, and a closed declarative feature DSL (count, distinct, share, peak, time_between, history-relative modifier) with mandatory caps and a static cost model. A frozen canonical-key alias table keeps input hashes, local-scorer summation order and version hashes identical, proven by a golden replay of every fixture. Adds a pointer in the main design and a plan row. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_014cdM7WyRc3mD3vQNXMTDB8 --- docs/design/2026-09-27-abusekit-design.md | 5 + .../2026-09-29-generic-feature-packs.md | 1211 +++++++++++++++++ docs/plans/2026-09-27-v0-plan.md | 1 + 3 files changed, 1217 insertions(+) create mode 100644 docs/design/2026-09-29-generic-feature-packs.md diff --git a/docs/design/2026-09-27-abusekit-design.md b/docs/design/2026-09-27-abusekit-design.md index f74f16e..2c90006 100644 --- a/docs/design/2026-09-27-abusekit-design.md +++ b/docs/design/2026-09-27-abusekit-design.md @@ -6,6 +6,11 @@ see account churn, tiers failed open when a scorer was unavailable, and request cover reads or replay. This revision fixes those and tightens every place an implementer would have had to guess. Changes from r1 are marked **[r2]**. +**Amendment (proposed 2026-09-29):** [`2026-09-29-generic-feature-packs.md`](2026-09-29-generic-feature-packs.md) +splits the built-in features into namespaced `core`/`email`/`brand` packs enabled per tenant, adds a +neutral `delivery.sent` event and product-declared vocabularies, and adds declarative custom +features in YAML. It amends §4.2, §4.3, §4.5, §4.6 and §4.10 below, with bit-identical scores for e2a. + ## 1. Problem statement In September 2026 a single operator ran two phishing campaigns through e2a. The first churned diff --git a/docs/design/2026-09-29-generic-feature-packs.md b/docs/design/2026-09-29-generic-feature-packs.md new file mode 100644 index 0000000..3c72af9 --- /dev/null +++ b/docs/design/2026-09-29-generic-feature-packs.md @@ -0,0 +1,1211 @@ +# Generic feature packs, a neutral event vocabulary, and declarative custom features + +Status: proposed, 2026-09-29 · owner: Josh Zhang · amends +[`2026-09-27-abusekit-design.md`](2026-09-27-abusekit-design.md) (§4.2, §4.3, §4.5, §4.6, §4.10). +Written against `main` at S3 plus the two open PRs treated as merged: #5 (S4 evaluation harness) +and #7 (S2b send-volume, webmail, recipient and subject-brand features). Section numbers below +refer to this document; `main §x` refers to the main design. + +## 1. Problem statement + +abusekit's current feature set assumes an email platform. The main design promises a scoring +service for any product that mints accounts, takes payments and lets users create resources. The +built code does not keep that promise: + +- **The vocabulary is email-shaped.** `content.sent` carries `subject_line`, `recipient_domain` + and `recipient_is_own_identity`, fields that only mean something for mail. Resource kinds are an + undeclared convention: `internal/feature` counts `kind == "key"` (plus S2b's spelling aliases) + and treats everything else as a generic resource. +- **Half the features are email features.** Of the 25 features on `main` + #7, + `first_day_distinct_domains`, `self_send_before_external`, `sends_10m_max`, `sends_1h`, + `sends_first_day`, `webmail_recipient_share`, `webmail_sends_1h`, `distinct_recipients_1h` and + `subject_brand_match` read `content.sent` and only mean something for mail. `key_velocity_1h` + and `key_total` assume API keys. A file-sharing, payments or chat product would get these at + zero. +- **Custom event types are dead weight.** Main §4.3 says unknown types are "stored, available to + Go-registered features only". No Go feature reads them, so the only way for a product to add a + signal is to write Go in this repo. +- **Everything is global.** One `config/rules.yaml`, one `config/local_weights.yaml` and one flat + feature namespace serve every tenant. A second tenant can't turn features on or off, and a new + feature name can collide with an existing one. + +**Desired outcome.** A product that is not an email platform can onboard with YAML only. It +enables the packs that fit its domain, declares its event types and field kinds, defines its +product-specific signals as declarative features, and runs in shadow. The first consumer (e2a) +keeps identical scores, bit for bit. + +### Success criteria (measurable) + +1. **Bit-identical migration.** A golden replay covers every committed fixture + (`eval/fixtures/*.jsonl`, `eval/fixtures/synthetic/`, and #7's fixtures). It scores after every + event and at every scheduled rescore instant. Each recorded feature value, `NextRescoreAt`, + per-rule input hash, risk (compared as `math.Float64bits`), tier and local-scorer `Version()` + must be identical before and after every slice in §9. The same holds for the harness's + `run.json` metrics once timestamps and git sha are removed. +2. **Zero-Go onboarding.** The three fictional products in §7 load from a tenant YAML file with + no Go changes. With uniform priors (§5.9), each product's abusive fixture scores above every + one of its benign fixtures. +3. **Enablement is enforced.** For a tenant without the `email` pack, no `email.*` feature is + computed, stored or rendered. A rule that references one fails the config load with + `feature_not_enabled`. +4. **Declarative features are correct and bounded.** For every operation in §5.5, the compiled + evaluator equals a naive O(n²) reference implementation on 10,000 randomized histories, + including shuffled arrival order and future-dated events. Per-subject extraction p99 is ≤ 50 ms + on 2 vCPU for a tenant at the limits in §5.7: 64 custom features over a history of 50,000 + events. +5. **Privacy by construction.** A property test shows that no stored value of a declared `text` + field matches the email-shape matcher. Values in undeclared fields are stored only as keyed + hashes, and config can never introduce a regex or executable code. The loader fuzz test finds + no profile that passes validation and breaks a limit in §5.7. + +## 2. Goals and non-goals + +**Goals** +- Split the built-in features into three packs: `core` (product-neutral), `email`, and `brand` + (display-name impersonation). Each pack is registered in Go and enabled per tenant. +- Namespace every feature (`core.subject_age_h`, `email.sends_10m_max`). Keep a frozen alias + table for today's flat names, so stored verdicts, corpora, cassettes, floors and weights stay + valid. +- Add a neutral delivery event (`delivery.sent`) and product-declared resource kinds, channels + and custom event types. `content.sent` stays accepted forever. +- Add declarative custom features in YAML. The set is closed: count, distinct, share, peak, + time-between, and a history-relative modifier. Every feature carries a mandatory cap and + transform, has deterministic semantics, and is validated and versioned at load. +- Per-tenant redaction for declared fields, driven by field kinds. +- Per-pack starter weights, fixtures, floors and golden-sign/mutation tests. A shadow-only + uniform-prior mode for a tenant that has no labels yet. + +**Non-goals** +- Arbitrary code or expressions in config: no CEL, no regex, no WASM, no Go plugins (§5.12). +- Learning weights from labels (`abusekit fit`). Bootstrapping stays in shadow until an operator + hand-tunes weights or a later design adds fitting (§12 Q8). +- Cross-tenant feature sharing or linking. Main §2 already defers this. +- Changes to scoring math. `core.Plan`, `core.Combine` and the local logistic model are + unchanged. Only their feature keys and quantization metadata move to a table (§5.2). +- A runtime API for declaring vocabularies. Declarations are reviewed config, like rules (§5.6). + +## 3. Relevant context and constraints + +**Code this design must fit (on `main` + #5 + #7):** +- `internal/event`: structural `Validate`, plus a static redaction `schema` map keyed by type. + Its `RedactionSchemaVersion` is 2 after #7. Unknown types keep every key after a recursive + leak scan; unlisted keys of known types are dropped. +- `internal/feature`: a monolithic `Extract(ctx, tenant, subject, events, neighbors, windows, + brands, webmail)` returns one fixed struct, `Features`, with 25 fields and a hand-written + `Map()`. `Names` feeds `config.FeatureSet`. `nextRescoreAt` hard-codes the windowed types + (`resource.created`, `content.sent`). +- `internal/core`: `inputHash` JSON-encodes the rule's feature map keyed by name. + `quantizeAgeFeaturesForHash` special-cases `subject_age_h` and `upgrade_delay_min` **by name**. +- `internal/model/local`: sums `weight × feature` in **sorted feature-name order** (S10). Its + `Version()` is a SHA-256 of the JSON-encoded `Weights`, whose weight map is keyed by name. + Renaming a feature therefore changes both the floating-point summation order and the version + hash. The migration must neutralise both (§5.2). +- `eval` (#5): corpus-v1 rows carry `input.features` keyed by flat name. Cassettes are keyed on + `(scorer, scorer_version, model, prompt_version, input_hash)`. `floors.yaml` entries are keyed + on `(rule, scorer, slice)`. `eval/gen` generates the synthetic corpus. +- Worker and config: rules and weights are global, and keys already carry a `tenant`. + +**Patterns to reuse:** load-time validation that rejects the whole document (main §4.5). Data +files loaded at boot (`brands.yaml`, `webmail.yaml`). Pure core with injected dependencies. The +compare-and-clear worker queue. Expand-only migrations. + +**Assumptions** (unconfirmed ones are repeated in §12): +- A1. Neither vendor adapter (S5) nor the hosted deploy (S8) has shipped. No production verdicts + and no vendor cassettes exist yet. The design still keeps them stable (§5.2) in case the order + changes. +- A2. At most about 100 tenants and about 64 custom features per tenant. Retention of 90 days + for event text (main §4.11) bounds any lookback. +- A3. Products can compute keyed hashes for identifiers they want to count distinctly, as e2a + already does for `recipient_hash`. As a fallback, abusekit hashes undeclared fields with its + own per-tenant key (§5.6). + +## 4. Proposed design: overview + +``` + tenant profile (config/tenants/.yaml) + ├─ packs: [core@1, email@1, brand@1] + ├─ vocabulary: resource kinds, channels, custom types + field kinds + ├─ features: declarative custom.* definitions + ├─ rules: inputs reference namespaced features + └─ weights: per rule (file, or `uniform`) + │ compile + validate (per tenant, atomic) + ▼ +ingest ──▶ vocab.Redact(tenant) ──▶ store (wire form + vocab_version) + │ +worker ──▶ vocab.View (content.sent → delivery view) ──▶ feature.Extract(profile) + │ for each enabled pack, in fixed order: + │ core (Go) · email (Go) · brand (Go) · custom (compiled YAML) + ▼ + Vector{values, canonical keys, quanta} + │ + core.Plan / Combine (unchanged math) +``` + +| Module | Interface | Deletion test | +| --- | --- | --- | +| `internal/vocab` (new) | `Compile(builtin, decl) (*Vocabulary, error)`; `(*Vocabulary).Redact(*event.Event) error`; `(*Vocabulary).View(event.Event) event.View` | Without it, redaction, kind declarations and the `content.sent` → delivery mapping spread across ingest and every pack. Keep. | +| `internal/pack` (new) | `Pack` interface (§5.1) + registry; adapters `core`, `email`, `brand` (Go) and `custom` (compiled YAML) | Four adapters, so the seam is real. Without it, `feature.Extract` stays a monolith that only grows. | +| `internal/pack/custom` (new) | `Compile(tenant, []Def, *Vocabulary) (Pack, error)` | Holds all DSL semantics. Deleting it leaves products writing Go again. Keep. | +| `internal/feature` | `Extract(ctx, profile, subject, events, now) (Result, error)`: runs enabled packs, merges, min-reduces rescore instants | Becomes a thin orchestrator. It stays because the worker, evaluate and eval all call exactly this one function. | +| `internal/config` | `LoadProfiles(dir, deps) (map[tenant]*Profile, []error)` | Absorbs `rules.yaml`; per-tenant atomic reload. | +| `internal/core` | unchanged signatures; `Plan` takes `Vector` instead of `map[string]float64` | The only change is that canonical keys and quanta come from data (§5.2). | + +The existing S2b brand, webmail and window helpers move into the packs that own them, with their +code unchanged. That is what keeps the migration bit-for-bit. + +## 5. Proposed design: detail + +### 5.1 Packs: registration, enablement, dependencies + +A **pack** is a named, versioned bundle of feature definitions and computation. It may also +carry data files (brand lists, webmail lists), starter weights, fixtures and floors. Packs are +compiled into the binary and registered in `internal/pack/registry.go`. A tenant **enables** +packs; it never supplies pack code. + +```go +// Pack computes a namespaced group of features for one subject. +type Pack interface { + ID() ID // {Name: "email", Version: 1} + Requires() []string // packs whose features/views it reads, e.g. email@1 → [core] + Features() []FeatureDef // static metadata, see below + // Extract is pure: same Input → same Output, independent of event order. + // It must exclude events with At > Input.Now from every window (§5.5 notes one frozen + // legacy exception). + Extract(ctx context.Context, in Input) (Output, error) +} + +type FeatureDef struct { + Name string // "email.sends_10m_max"; grammar §5.3 + Legacy string // frozen flat alias ("sends_10m_max"), "" for features born namespaced + HashQuantum float64 // input-hash bucket; 0 = exact (§5.2) + Bound float64 // max value after transform; used by uniform priors (§5.9) + Reads []string // view types read ("delivery", "resource.created", ...), for dispatch + // and rescore scheduling + RequiresPacks []string // feature-level dependency, e.g. email.subject_brand_match → [brand] +} + +type Input struct { + Tenant, Subject string + Events []event.View // stored events projected through the tenant vocabulary + Now time.Time + Start time.Time // account start: subject.created.account_created_at if present, + // else earliest event At (S2b R8) + Neighbors NeighborEvidence // resolved once by the orchestrator, only when core is enabled + Params PackParams // this tenant's validated per-pack settings (lists, channels, kinds) +} + +type Output struct { + Values map[string]float64 // exactly the names in Features(); missing = bug, rejected + Rescore []time.Time // candidate instants at which some value changes with no new event +} +``` + +**Enablement** lives in the tenant profile: `packs: [core@1, email@1, brand@1]`. The rules are: +- `core` is always enabled and is implied if omitted. Every other pack is opt-in. +- A pin names a major version. A pack changes feature semantics only by shipping a new major + version (`email@2`) next to the old one for at least one release. Semantic changes never + happen in place. Every verdict records the pinned versions in `profile_sha` (§5.10). +- `Requires` must be satisfied by the enabled set. Otherwise the tenant fails to load with + `pack_requires`. +- A feature whose `RequiresPacks` are not all enabled is **not registered** for the tenant. For + example, `email.subject_brand_match` needs `brand`; enable `email` without `brand` and that one + feature doesn't exist for the tenant. +- The orchestrator runs packs in fixed order: `core`, `email`, `brand`, then `custom`. It merges + their `Values` and takes the earliest `Rescore` instant, coalesced to the existing 5-minute + buckets. No two packs share a namespace, so collisions are impossible by construction. + +**Pack contents after the split** (legacy name → namespaced name): + +| Pack | Features | +| --- | --- | +| `core@1`, account and onboarding | `subject_age_h`→`core.subject_age_h` (quantum 1) | +| `core@1`, payment | `upgrade_delay_min`→`core.upgrade_delay_min` (quantum 60), `upgraded`→`core.upgraded`, `declines_before_first_success`→`core.declines_before_first_success`, `first_funding_prepaid`→`core.first_funding_prepaid`, `fingerprint_seen_on_other_subjects`→`core.fingerprint_seen_on_other_subjects` | +| `core@1`, resource velocity | `resource_velocity_1h`→`core.resource_velocity_1h`, `resource_total`→`core.resource_total`, `key_velocity_1h`→`core.credential_velocity_1h`, `key_total`→`core.credential_total` (a declared kind with `role: credential`, §5.4) | +| `core@1`, linked subjects and neighbour evidence | `linked_deleted_n`→`core.linked_deleted_n`, `neighbors_truncated`→`core.neighbors_truncated` | +| `core@1`, labels | `linked_labelled_abusive_n`→`core.linked_labelled_abusive_n` | +| `core@1`, behaviour change | `burst_ratio_24h_vs_lifetime`→`core.burst_ratio_24h_vs_lifetime`. Its activity set is `resource.created` ∪ delivery views, which is exactly today's `resource.created` ∪ `content.sent`. The shared `burstFactor`/`ageDecayFactor` functions are also exposed to the DSL as `relative_to_history` (§5.5). | +| `email@1`, requires `core` | `sends_10m_max`, `sends_1h`, `sends_first_day`, `webmail_recipient_share`, `webmail_sends_1h`, `distinct_recipients_1h`, `first_day_distinct_domains`, `self_send_before_external`, `subject_brand_match` (the last also requires `brand`). All become `email.`. They read delivery views whose `channel` is in the pack's `channels` param (default `[email]`). | +| `brand@1` | `name_brand_match`→`brand.name_match`, `name_has_at`→`brand.name_has_at`; new `brand.title_match` (§5.1.1). Data: `brands.yaml` plus the tenant's optional private `extra` list (S2b `brands_extra`). | + +The `core.credential_*` rename is the only name that moves beyond adding a prefix. "Key" is an +e2a word; "credential" is the product-neutral role (§5.4). The legacy alias keeps e2a's hash and +weight keys unchanged. + +New core features are **additive**. They are registered but not in e2a's rule, so e2a's golden +replay can't move: +- `core.email_domain_class_disposable`: main §8 Q7 deferred this. It reads the existing + `subject.created.email_domain_class` and is neutral (sign-up identity quality, not mail sending). +- `core.verdict_max_24h`: the maximum `content.verdict.score` in the trailing 24 h. +- `core.history_truncated`: see §5.7. + +#### 5.1.1 The brand pack + +The matcher is S2b's `BrandSet`, moved unchanged: NFKC and confusables skeleton, word and token +boundaries, Unicode punctuation tokenizing, case-sensitive entries, the integration and +community gates, and Cf stripping. It applies to any product with user-chosen display names: +workspace names, storefront names, profile names, community names. + +- `brand.name_match` and `brand.name_has_at` read `resource.created` and `resource.deleted` + `name` (and `name_skeleton`) across **all** declared kinds, as they do today. +- `brand.title_match` is the neutral sibling of `email.subject_brand_match`. It counts distinct + brands in non-self delivery titles over a trailing 1 h, capped at 3, and has none of the + email-specific exemptions. Non-email tenants use it. e2a keeps `email.subject_brand_match`, + whose S2b integration-name exemption and double-count rule are specific to mail. +- The brand list is pack data: `config/packs/brand/brands.yaml` (public), plus a per-tenant + private `brand.extra` path merged at boot with `MergeBrandSets`. + +### 5.2 Namespacing and the canonical key (how migration stays bit-for-bit) + +Every feature has two identities: +- **Name**: the namespaced name, used everywhere a human or config refers to the feature: rules, + weights files, corpus v2, reasons, docs. +- **Canonical key**: `Legacy` if the feature has one, else `Name`. It is used in exactly three + places: + 1. `core.inputHash`: the rule's feature map is re-keyed by canonical key before JSON encoding. + Quantization comes from each feature's `HashQuantum`, which replaces the name switch in + `quantizeAgeFeaturesForHash`. `core.subject_age_h` has quantum 1 and + `core.upgrade_delay_min` has quantum 60, the same values applied under the same keys. + 2. The local scorer's summation order: `sortedFeatures` is sorted by canonical key. + 3. The local scorer's `Version()`: it hashes the weights map re-keyed by canonical key. + +For a migrated feature, the canonical key is the flat name today's code uses. Summation order, +version hash, input hashes and therefore cassette keys are byte-identical, so no subject is +rescored at cutover and no vendor call is repeated. A feature born namespaced (every new pack +feature and every `custom.*`) has canonical key = name, so nothing legacy leaks into new work. + +**The alias table is frozen.** It is a Go literal of exactly the 25 flat names on `main` + #7, +and CI enforces it: the table can never gain an entry, and no new `Name` may equal any alias. + +**Where flat names are still accepted (read-side only):** + +| Artifact | Behaviour | +| --- | --- | +| Rules (`inputs`) | Resolved through the alias table, but only if the owning pack is enabled for the tenant. Otherwise the load fails with `feature_not_enabled`. A load warning says the name is deprecated. | +| Weights files | Same resolution. A file must not mix a flat name and its namespaced twin (`duplicate_feature`). | +| Corpus v1 rows | `eval.LoadSnapshotCorpus` maps `input.features` keys through the table. An unknown flat key fails the load. Export always writes corpus-v2 (§5.9). | +| Text inputs | Aliases `subject_line_skeleton`→`delivery.title_skeleton` and `first_link_host`→`delivery.link_host`. Their canonical keys stay the legacy names, for the same hashing reason. | +| Floors | Entries with no `profile:` default to `tenant:e2a` (§5.9). | +| Stored verdicts | Untouched. They store `input_hash`, risk and reason, and never feature names. | + +### 5.3 Name grammar and reserved namespaces + +- Feature names: `^[a-z][a-z0-9]*\.[a-z][a-z0-9_]{0,55}$`, at most 64 bytes. +- Reserved feature namespaces: `core`, `email`, `brand`, `custom`, plus any future pack name. + Tenants may define only `custom.*`. The `custom.` namespace is per tenant, so two tenants may + each have a `custom.invites_1h` that means different things (§12 Q11). +- Event type names keep the wire grammar `^[a-z_.]+$`, at most 64 bytes, and must contain a dot. + Reserved type prefixes (built-ins, tenants may not declare types under them): `subject.`, + `payment.`, `subscription.`, `resource.`, `content.`, `delivery.`, `abusekit.`. + +### 5.4 Neutral event vocabulary + +#### Built-in types after this change + +The wire contract (`POST /v1/events`, main §4.3) is unchanged in shape. Every change below is +additive: a new built-in type, new optional fields, and tenant-declared types. + +| Type | Change | +| --- | --- | +| `subject.created`, `subject.deleted`, `subject.class`, `payment.attempt`, `subscription.changed`, `content.verdict` | none (keeps #7's optional `account_created_at`) | +| `resource.created` / `resource.deleted` | `kind` is interpreted through the tenant's declared **resource kinds** (below). On the wire it stays free text. | +| `content.sent` | **Legacy email delivery, accepted forever, with #7's redaction unchanged.** Projected to a delivery view (below). | +| `delivery.sent` (new) | Neutral delivery: N recipients at destination D, with optional title text. | + +`delivery.sent.data`: + +| Field | Kind | Rule | +| --- | --- | --- | +| `channel` | enum, declared per tenant | Optional. A tenant with exactly one declared channel may omit it; otherwise it is required. An undeclared value → `redaction_failed`. | +| `destination` | the kind the channel declares: `domain` \| `hash` \| `enum` | Optional. Where the delivery lands: a recipient domain, a keyed community id, a region code. | +| `destination_class` | enum, declared per channel | Optional, for example `own_community` \| `other_community`. | +| `recipient_count` | positive integer | Optional, default 1 for every feature that sums it. | +| `recipient_hash` | hash (`^[A-Za-z0-9_:+/=-]{8,128}$`) | Optional. Exactly one recipient, so paired with `recipient_count > 1` → `redaction_failed` (#7's rule). | +| `to_self` | bool | Optional. The recipient is the subject's own identity. | +| `title` | text ≤ 200 | Optional. NFKC, email-shaped substrings masked, skeleton stored as `title_skeleton`. | +| `link_host` | domain | Optional. First link host in the delivered content. | + +#### The delivery view (read-side projection) + +Features never read `content.sent` or `delivery.sent` directly. They read `event.View`, which +`vocab.View` produces from the stored row. For `content.sent`: + +| Delivery view field | Taken from `content.sent` | +| --- | --- | +| `channel` | constant `email` | +| `destination` (kind `domain`) | `recipient_domain` | +| `recipient_count` | `recipient_count` | +| `recipient_hash` | `recipient_hash` | +| `to_self` | `recipient_is_own_identity` | +| `title` / `title_skeleton` | `subject_line` / `subject_line_skeleton` | +| `link_host` | `first_link_host` | + +A `delivery.sent` row maps to the same view field for field. The email pack's features are the +S2b functions with one mechanical edit: `e.Type != "content.sent"` becomes +`v.Kind != event.ViewDelivery || !channels[v.Channel]`, and the field reads are renamed. +Everything else is untouched, including missing-field defaults, `recipientCountOf`'s default of +1, `normalizeToken` domain folding and `isSelfSend`. So an e2a stream of `content.sent` and the +same stream rewritten as `delivery.sent{channel: email}` yield identical features. §8's +translation test proves this on every fixture. + +**Why read-side and not a rewrite at ingest.** Rewriting would change the stored type and the +`BodyHash` that separates `duplicate` from `conflict`. It would also need a data migration for +stored rows and would still leave old rows to interpret. The read-side view needs no migration +and makes both forms equivalent by construction. + +#### Product-declared resource kinds and channels + +```yaml +vocabulary: + version: 3 # monotonically increasing; see §5.6 for compatibility rules + resource_kinds: + agent: {role: identity} + key: {role: credential, aliases: [keys, api_key, api_keys, api-key, apikey, "api key"]} + channels: + email: {destination: domain} +``` + +- `role` is a closed enum: `credential`, `identity`, `workspace`, `content`, `other`. + `core.credential_*` counts kinds with `role: credential`; `core.resource_*` counts every kind. + Matching folds case and trims whitespace, then applies `aliases`. This is the S2b behaviour + moved into data: e2a's declaration above reproduces `resourceKindAliases` exactly, and the + golden replay proves it. +- An **undeclared kind** is stored, counts in `core.resource_*` (role `other`), and increments + `abusekit_undeclared_kind_total{tenant}`. That matches today's behaviour for kinds that aren't + keys. With `vocabulary.strict: true`, an undeclared kind is instead rejected with the existing + `redaction_failed` code, so no new per-item code is needed. +- A tenant with no `resource_kinds` block gets the **implicit legacy declaration** above. That is + how today's global config keeps working (§8). + +#### Versioning + +The wire format stays `/v1`. Every addition is optional, and producers already handle unknown +per-item codes because the code list is closed and unchanged. The one semantic change is to +how **undeclared** types are stored (§5.6). It is flagged, and it is licensed because the +service is pre-GA: main §4.3 documents unknown types as "stored", and nothing reads them yet. +`RedactionSchemaVersion` for built-ins stays at 2 (#7). Each stored row gains +`vocab_version text` (`"@"`, NULL for built-in-only schemas) next to +`redaction_version`, in an expand-only migration. + +### 5.5 Declarative custom features + +#### Schema + +```yaml +features: + - name: custom.public_links_1h # custom.* only + version: 1 # bump on ANY change to the definition (enforced, §5.7) + description: public share links created in the last hour # required, shown in reasons + count: # exactly one op key: count | distinct | share | peak | time_between + type: share.link_created # a declared type, or a built-in type/view + where: {field: visibility, eq: public} + sum: {field: size_class_weight, default: 1, cap_each: 10} # optional; count = sum of this + window: 1h + transform: {log1p: true, cap: 50} # cap mandatory; log1p optional; applied log1p → cap +``` + +**Windows.** Either `window: ` (trailing, `(now − dur, now]`) or `first: ` (anchored, +`[start, start + dur)`). `` is a whole number of minutes, hours or days: `1m`–`30d`. +`lifetime` means `(−∞, now]`. `start` is the account start defined in §5.1 `Input.Start`. + +**Predicates (`where`).** A closed set of operators. There is no regex, no arithmetic and no +user functions. + +| Leaf | Meaning | Allowed field kinds | +| --- | --- | --- | +| `{field: f, eq: v}` / `{field: f, ne: v}` | equality after the kind's normalisation | enum, bool, number, domain, hash | +| `{field: f, in: [..]}` / `not_in` | membership, ≤ 256 values; for enums every value is checked against the declared enum at load, so a typo is a load error | enum, number, domain, hash | +| `{field: f, in_set: }` | membership in a tenant- or pack-provided set file (for example, `email`'s webmail list), hashed at load; ≤ 100,000 entries | domain, hash, enum | +| `{field: f, suffix_in_set: }` | domain-label-boundary suffix match (`a.b.example.test` ⊂ `example.test`), linear time | domain | +| `{field: f, gte: n}` / `lte` / `gt` / `lt` | numeric compare | number | +| `{field: f, exists: bool}` | presence | any | +| `{all: [..]}` / `{any: [..]}` / `{not: leaf}` | combinators, depth ≤ 2, ≤ 8 leaves in total | | + +An absent field or a type mismatch makes a leaf false; `{exists: false}` is the only leaf that +is true on absence. `text` fields can't appear in `where` or as a `distinct` field. Text is only +available to text-accepting scorers through a rule's `text:` list. That keeps free text out of +every aggregation. + +**Operations.** + +| Op | Value at `now` | +| --- | --- | +| `count` | The number of matching events in the window, or, with `sum`, the sum of the field over them. `default` covers absence, and `cap_each` clamps each event's addend before summing, as S2b does with `recipient_count`. | +| `distinct` | `{type, where, field}`: the number of distinct normalised values of `field` among matching events in the window. `field` must be an `enum`, `number`, `domain` or `hash` field. Tracking stops at `track_max`, the smallest count whose transformed value reaches `transform.cap` (`cap` itself without `log1p`, `ceil(expm1(cap / scale))` with it). The loader rejects a definition whose `track_max` exceeds 10,000, so memory is O(track_max). | +| `share` | `{type, where (denominator), match (numerator predicate), sum?}`: numerator ÷ denominator over the window; `if_empty` (default 0) when the denominator is 0. | +| `peak` | `{type, where, sum?, size: }`: the maximum count or sum in any window `(t − size, t]` with `t` an event instant inside the outer window. `size` ≤ window and window ÷ size ≤ 1440. A two-pointer scan over the time-sorted matches. | +| `time_between` | `{from: {type, where}, to: {type, where}, until_now: bool, if_absent: n}`: minutes from `t_A` (the earliest matching `from`) to `t_B` (the earliest matching `to` with `t_B ≥ t_A`). No `from` event → `if_absent`. A `from` but no `to` → minutes since `t_A` when `until_now`, else `if_absent`. `if_absent` is mandatory. Negative results can't occur, and the value is floored at 0 as a guard. | + +**The history-relative modifier.** It may be added to `count`, `distinct` and `peak`, and it +generalises S2b's `burstFactor × ageDecayFactor`: + +```yaml + relative_to_history: + lookback: 30d # ≤ 30d + exclude_recent: 24h # the current burst never serves as its own baseline + age_decay: {full_until: 3d, zero_at: 30d, floor: 0.2} # optional; these are the defaults +``` + +`baseline` is the maximum of the same op, with the same width and predicate, over sliding +windows whose end lies in `(now − lookback, now − exclude_recent]`. The value is then +`min(current / max(baseline, 1), transform.cap)`. If `age_decay` is set, the result is +multiplied by `clamp(1 − (age_days − full_until)/(zero_at − full_until), floor, 1)`. With the +defaults, `full_until = 3d` and `zero_at = 30d` make the denominator 27, which is exactly S2b's +`ageDecayFactor`. `floor > 0` is mandatory, so a decayed signal is never a hard zero. + +**Transform.** `cap` is mandatory for every feature: a finite value > 0, and at most the op's +natural bound (1 for `share`). `log1p` is optional and takes one of three forms, applied before +the cap: `log1p: true` gives `ln(1 + v)`; `log1p: {scale: s}` gives `s · ln(1 + v)`; and +`log1p: {anchored_at: n}` sets `s = n / ln(1 + n)`, so `v = n` maps to `n`. The last form +reproduces `first_day_distinct_domains`'s S2b shape. The transformed value always lies in `[0, cap]`, and +that is the feature's `Bound`. + +#### Evaluation semantics + +- **Pure and order-independent.** A feature is a function of the set of stored events and + `now`, never of arrival order. Where "first" is ambiguous, ties at the same instant break by + `(at, event id)`. Late events bump `dirty_seq` as they do today. +- **Half-open windows.** Trailing windows are `(now − W, now]`; anchored windows are + `[start, start + W)`; `peak` sub-windows are `(t − S, t]`. +- **Future events are excluded everywhere,** `lifetime` included: an event with `at > now` + contributes to no custom feature until `now` reaches it. It does contribute a rescore + candidate at its own `at` (the S2b B4 behaviour, generalised). + *Frozen legacy exception:* the migrated Go features `core.resource_total` and + `core.credential_total` count future-dated events in their lifetime totals (the S2 behaviour), + and `email.first_day_distinct_domains` uses an inclusive `[start, start + 24h]`. Both are kept + for bit-for-bit parity and documented in each feature's docstring. Harmonising them is a + `core@2`/`email@2` change (§12 Q5). +- **Rescore candidates.** For each windowed feature, the compiler emits: the exit instant of the + oldest in-window match (`at + W`); anchored window ends still in the future; for `peak`, the + exit of the current maximum's sub-window; for `relative_to_history`, the instants where an + event crosses `now − exclude_recent` or `now − lookback`; and every future-dated match. The + orchestrator takes the minimum and coalesces it to the 5-minute bucket. This generalises + `isWindowedEventType`: the set of windowed types is the union of every feature's `Reads`. +- **Determinism of floats.** Sums accumulate in time order `(at, id)` with a single `float64` + accumulator, so the same event set always gives the same bits. + +#### Compilation and cost model + +`pack/custom.Compile` validates each definition and produces a **per-tenant plan**: an index +from view type to the list of (feature, predicate program) pairs that read it. Extraction makes +one pass over the subject's events, dispatches each event to the features for its type, and +evaluates the flat predicate programs. `peak` and `relative_to_history` keep per-feature +matched-instant slices, and a final pass per feature runs the two-pointer scans. + +Static cost units, checked at load: + +| Op | Units | +| --- | --- | +| `count`, `share`, `time_between` | 1 | +| `distinct`, `peak` | 2 | +| `relative_to_history` | ×2 on top of the op's own cost | +| Each predicate leaf beyond the first | +0.25 | + +Per-tenant budget: **256 units**. With the 50,000-event scan bound (§5.7), the worst case is +about 50,000 events × 8 dispatched features per type × 8 leaves ≈ 3.2M predicate steps, plus +O(n) scans. That is tens of milliseconds, which is what criterion 4 measures. At runtime a +per-subject extraction deadline (default 250 ms) applies to the custom pack. If it expires, the +pack fails as described in §6. + +### 5.6 Redaction for declared types and fields + +Ingest never consults rules (main §4.3). It consults the tenant **vocabulary**, which is +config, reviewed like code, and compiled into `vocab.Vocabulary`. The built-in `schema` map +becomes the built-in half of every vocabulary, unchanged. + +```yaml +vocabulary: + version: 1 + types: + share.link_created: + fields: + visibility: {kind: enum, values: [public, org, private]} + file_kind: {kind: enum, values: [document, archive, executable, image, other]} + size_bytes: {kind: number, min: 0, integer: true} + folder_title: {kind: text, max_len: 120, skeleton: true} + share.downloaded: + fields: + link_hash: {kind: hash} + downloader_ip24: {kind: hash} + downloader_is_owner: {kind: bool} +``` + +**Field kinds and the rule for each:** + +| Kind | Accepts | On violation | +| --- | --- | --- | +| `text` | string; NFKC; control characters and invalid UTF-8 rejected; **email-shaped substrings masked to `@`** (#7's `subject_line` rule, generalised); truncated at `max_len` (≤ 500, default 200); optional `skeleton` companion | reject for control characters or bad UTF-8; mask for an email shape | +| `number` | finite float64; optional `min`, `max`, `integer` | reject | +| `bool` | bool | reject | +| `enum` | string in `values` (≤ 64 values, each ≤ 64 bytes, `[a-z0-9_.-]+`) | reject | +| `hash` | `^[A-Za-z0-9_:+/=-]{8,128}$` (no `@`, `%` or whitespace) | reject (never truncated) | +| `domain` | lower-cased, IDNA to ASCII, hostname grammar (labels 1–63, total ≤ 253); a user part is impossible by grammar | reject | +| `timestamp` | RFC 3339 (#7's `account_created_at` pattern) | reject | + +**Rules that make it privacy by construction:** +1. The existing recursive leak scan runs first, over every key and value of every event. + Email-shaped content outside a `text` field is still rejected. Declared `text` fields are the + only exemption, and they are masked instead of rejected. +2. **Undeclared fields of a declared type are hashed, not stored.** A string becomes + `hk1:` + hex(HMAC-SHA256(tenant redaction key, type ‖ field ‖ value))[:32]. Numbers and bools + pass through. Objects and arrays are dropped. A hashed field can be used only in + `distinct`, `eq`/`in` on a hash (the producer would have to compute the same HMAC, which it + can't, so in practice only `distinct` and `exists`). It can never be read back. +3. **Undeclared types** get the same treatment, with every top-level key treated as undeclared. + This replaces "kept as-is" (§5.4 versioning note). Undeclared types remain unusable by + features until they are declared. +4. **The redaction key** is a per-tenant secret held by abusekit (Secret Manager, + `abusekit--redaction-key`). It is separate from the producer-held link-hash key. + Rotating it breaks equality across the rotation boundary for hashed fields, which is + acceptable because they only feed windowed `distinct`. The key id is recorded in + `vocab_version`. +5. Caps: ≤ 32 declared types per tenant, ≤ 32 fields per type, `data` ≤ 8 KiB after redaction + (unchanged). + +**Vocabulary compatibility.** Stored rows are immutable, and features must be able to read old +rows. So a vocabulary may only **widen**: add a type, add a field, add enum values, raise +`max_len` or `max`, or switch `strict` off. Changing a field's kind, removing or narrowing enum +values, or lowering a cap requires a new field name. The loader enforces this against the latest +accepted vocabulary for the tenant, recorded in a new `tenant_vocabularies(tenant, version, +sha, body, accepted_at)` table. The version must increase with any change, and an incompatible +change fails the load with `vocab_incompatible`. The harness reports how many rows were stored +under each `vocab_version`. + +### 5.7 Limits (validated at load; also the fuzz oracle) + +| Limit | Value | +| --- | --- | +| Custom features per tenant | 64 | +| Cost units per tenant | 256 | +| Predicate depth / leaves per feature | 2 / 8 | +| `in` list size / set file entries | 256 / 100,000 | +| Window, lookback | ≤ 30 d; `lifetime` allowed only for `count`, `distinct`, `share`, `time_between` | +| `peak` window ÷ size | ≤ 1440 | +| `distinct` cap | ≤ 10,000 | +| Events scanned per subject | 50,000 newest by `(at, id)`, plus the earliest event and `subject.created` for `start` | +| Extraction deadline (custom pack) | 250 ms default, per-tenant override ≤ 1 s | + +Above the scan bound, the orchestrator sets `core.history_truncated = 1`. No committed fixture +comes near the bound, so the golden replay is unaffected. Custom-feature versioning: the loader +keeps a SHA-256 of each normalised definition per `(tenant, name, version)` in +`tenant_feature_defs`. Redefining an existing `(name, version)` with a different body fails with +`feature_version_reused`. + +### 5.8 Rules and validation at load + +A tenant profile is `config/tenants/.yaml`: + +```yaml +tenant: e2a +packs: [core@1, email@1, brand@1] +pack_params: + brand: {extra: /run/secrets/brands_extra.yaml} # optional, private + email: {channels: [email], webmail_set: default} +vocabulary: {...} # §5.4, §5.6 +features: [...] # §5.5 +tiers: {medium: 0.4, high: 0.8} +min_scored_advise: 1 +rules: + - name: new_account_velocity + mode: advise + scorer: local + weights: tenants/e2a/weights.yaml # or `uniform` (shadow only, §5.9) + inputs: [core.subject_age_h, ...] +``` + +Validation adds these checks to main §4.5's list. The whole tenant profile is rejected with a +collected error list, and the codes are machine-readable in `/healthz`: +- `pack_unknown`, `pack_requires`, `pack_version_unknown` +- `feature_unknown`: a name exists in no pack and not in the tenant's `custom.*` set +- `feature_not_enabled`: the name, or its alias, exists but its pack or `RequiresPacks` isn't + enabled +- `feature_namespace`: a tenant tries to define a feature outside `custom.*` +- `duplicate_feature`, `feature_version_reused` +- `dsl_invalid` (with a JSON-pointer path): an unknown op, a missing cap, a text field in a + predicate, an enum value not declared, a window out of range, too many leaves, a cost overrun +- `vocab_invalid`, `vocab_incompatible` +- `weights_unknown_feature`: every weight must name an input of the rule it serves +- `uniform_not_shadow`: a rule that uses `weights: uniform` must be in `mode: shadow` + +**Reload is atomic per tenant.** A rejected profile keeps that tenant's previous profile live +and has no effect on other tenants; `/healthz` reports `config_error{tenant}`. This replaces +today's whole-config rejection, which would let one tenant's typo freeze every tenant's rule +changes. Before any tenant file exists, `config/rules.yaml` + `config/local_weights.yaml` load +as the **default profile**. That profile has implicit `packs: [core@1, email@1, brand@1]`, the +implicit legacy vocabulary, and flat-name resolution, and it applies to every tenant that has +keys but no file. This is the zero-change path for e2a until slice G7 (§9). + +### 5.9 Weights, scorers, eval and bootstrap per pack + +- **Weights are per rule**, keyed by namespaced name (flat names are accepted through the alias + table). The scorer name stays `local`. At profile compile time the loader binds a local scorer + instance per weights file. `Version()` is content-derived over canonical keys (§5.2), so + different weights produce different versions automatically. Registry lookups become + `(tenant, scorer)`, and vendor scorers stay global. +- **The local scorer's math is unchanged:** `sigmoid(bias + Σ w·x)`, summed in canonical-key + order. +- **Starter weights per pack.** Each pack ships `config/packs//starter.yaml`: one rule, + `_starter`, with its own bias and weights over that pack's features only. Every weight + carries `sign: +|-` (the golden-sign contract). Starter rules compose: a tenant can enable + `core_starter` and `brand_starter` as separate shadow rules, and `Combine`'s + `max(risk)` handles them without inventing a joint model. The `core` starter is derived from + today's e2a weights restricted to core features. The `email` and `brand` starters are derived + the same way. All are placeholders until labelled data exists for a second product. +- **Uniform priors, shadow only.** `weights: uniform` binds a local scorer with, for `k` inputs, + `x̂ᵢ = xᵢ / Boundᵢ ∈ [0, 1]` and `wᵢ = ±4/k` on `x̂ᵢ`. The sign defaults to `+`, and a feature + can declare `prior_sign: -` (for example, an account-age or time-to-first-action feature). The + bias is `−2 + (4/k) × (number of negative-sign inputs)`, so risk always ranges from + sigmoid(−2) ≈ 0.12 to sigmoid(2) ≈ 0.88. The result is a ranking device + for operator review, never a tier driver. The loader rejects it in `advise` + (`uniform_not_shadow`), and promotion (main §4.5) requires a real weights file plus a gate run. +- **Eval is scoped to a profile.** `abusekit eval --profile pack:email` or + `--profile tenant:e2a` selects the rule, weights, fixtures and floors. Layout: + `eval/packs//{fixtures/, floors.yaml}` and `eval/tenants//{fixtures/, + floors.yaml, golden/}`. `floors.yaml` entries gain `profile:`; a missing value means + `tenant:e2a`, so #5's file keeps working unchanged. The manifest gains `profile`, + `profile_sha` and `pack_versions`. +- **Cassettes** are unchanged. Their key already includes `input_hash`, and canonical keys keep + it stable (§5.2). +- **Corpus v2** (`eval/schema/corpus-v2.schema.json`) is v1 plus a required `profile` and + `vocab_version`, with `input.features` keyed by namespaced name. `LoadSnapshotCorpus` reads + both versions. +- **Golden-sign and mutation tests per pack** become a reusable harness, + `internal/pack/packtest.Run(t, pack)`, and every pack must pass it in CI, like the adapter + contract test. It checks: + 1. Every starter weight's sign matches `sign:`. + 2. Zeroing each weight moves at least one of the pack's fixture bands or an isolated scenario + (today's `mutation_test.go` logic, parameterised). + 3. Determinism: same bits under shuffled event order and repeated runs. + 4. No leakage from the future: adding an event at `now + ε` changes no value except rescore + candidates. + 5. Every emitted name is in `Features()` and in the pack's namespace, with `Bound` respected. + + e2a's tenant profile keeps its own golden-sign and mutation suite over its composed rule, with + namespaced keys. +- **How a second product bootstraps:** + 1. Enable packs, declare the vocabulary, and write custom features. Run shadow rules + `core_starter` (plus `brand_starter` if relevant) and a `custom_uniform` rule over the + custom features. + 2. Collect labels through `POST /v1/labels` (main §4.9). Corpus rows accrue per profile. + 3. Once a labelled set passes the harness, hand-tune a real weights file, set floors, and + promote through the normal shadow → advise path. Fitting is §12 Q8. + +### 5.10 Provenance and storage changes (expand-only) + +- `events.vocab_version text NULL`. +- `verdicts.profile_sha text NULL`: SHA-256 of the pack versions, the normalised custom-feature + definitions and the vocabulary version. It is recorded, not part of the input hash, so + editing an unrelated custom feature doesn't force rescoring. A changed value still changes the + hash through the value. +- New tables `tenant_vocabularies` and `tenant_feature_defs` (§5.6, §5.7). +- No change to `links`, `subjects`, `labels` or `corpus_examples`. + +### 5.11 API surface summary + +| Surface | Change | Compatibility | +| --- | --- | --- | +| `POST /v1/events` | `delivery.sent` built-in; tenant-declared types redacted by kind; undeclared fields hashed | Additive on the wire. Storage of undeclared types changes (pre-GA, §5.4). | +| `GET /v1/subjects/{id}`, `evaluate` | Signal `reason` templates may name namespaced features in prose | Additive; the shape is unchanged. | +| Per-item codes | none new (`redaction_failed`, `bad_type` reused) | unchanged | +| Config YAML | tenant profiles; `rules.yaml` still loads as the default profile | Backward compatible. Flat names are deprecated with a warning. | +| Weights, floors, corpus files | namespaced keys; `profile:`; corpus-v2 | v1 read forever via the alias table | +| `pkg/abusekit` client | `DeliverySent` event helper; no removals | additive | + +**Rejected API alternative:** a `PUT /v1/vocabulary` endpoint that would let producers declare +schemas at runtime. It would let a producer key widen its own redaction boundary, which is a +privilege escalation. It would also move a privacy decision out of code review. + +### 5.12 Alternatives considered + +- **CEL (cel-go) for predicates and features.** It is sandboxed, has cost estimation, and is a + known quantity. It lost for four reasons: + 1. Features aggregate over time-windowed sequences of events. CEL has no windowed aggregates, + so we would still have to write count, distinct, peak, time-between and history-relative as + custom functions. CEL would only wrap the predicate, which is the easy part. + 2. It pulls in a large dependency (cel-go plus protobuf), against the repo's minimal-dependency + convention. + 3. CEL's cost estimate is per expression. Ours has to be per subject history, which needs our + own model anyway. + 4. Its error messages and semantics (`has()`, dynamic types) are harder for a product engineer + to get right than a closed YAML schema whose load errors carry JSON pointers. + + CEL remains the fallback **for `where` only** if the closed predicate set proves too small + (§12 Q9). +- **A home-grown expression language.** It would bring a parser, a grammar, precedence rules and + an injection surface, all needing a security review, for no coverage beyond the six closed + operations the target signals need. +- **A plugin ABI.** Go `plugin` needs an identical toolchain and build flags and has no sandbox. + WASM (wazero) is sandboxed with fuel metering, but it deploys arbitrary code disguised as + config, reviewers can't read it, and float determinism depends on the guest. Products that + need code contribute a Go pack upstream. The `Pack` seam is where that code goes, and + `packtest` gates it. +- **SQL-defined features against the store.** They couple to the schema, their cost is + unbounded, and they are a tenant-isolation hazard. Rejected. +- **Keep flat names and prefix only new features.** No migration, but also no enablement + boundary, and two naming styles forever. Rejected in favour of the canonical-key bridge, which + costs one frozen table. +- **Translate `content.sent` to `delivery.sent` at ingest.** See §5.4. Rejected for the + body-hash and migration costs. +- **Rename everything and rescore once.** This gives a simpler hash with no canonical keys. It + lost because summation order changes the last bits of risk (breaking the bit-for-bit + criterion), and any vendor cassette recorded before cutover would go stale. + +## 6. Edge cases and failure handling + +- **A pack's `Extract` errors or the custom pack exceeds its deadline.** The orchestrator + records the failure per pack, not per subject. Rules whose inputs include any feature of that + pack become `unscored` with `error_code: feature_error` or `feature_timeout`, and + `degraded: true` is set. Rules that don't read the pack still score. This fails closed: + `unknown` or `degraded`, never `low` by absence (main §5). A pack that fails for every subject + of a tenant pages through the existing metric. +- **A profile is rejected on reload.** The previous profile stays live for that tenant only. + On a cold start with no valid profile, the tenant's subjects stay unscored (`unknown`) and + `/healthz` is red. The service never falls back to a different tenant's rules. +- **Events of a type arrive before its declaration (deploy ordering).** They are stored under + the undeclared-type rule (strings hashed). Features declared later can't read those strings, + but they can still count the events. The runbook says to declare first, and the harness + reports the row counts per `vocab_version`. +- **A declared kind or channel is missing on an event.** Undeclared kinds count as `other`, as + in §5.4. For channels, `delivery.sent` with an undeclared channel is rejected + (`redaction_failed`), because silently counting it under a guessed channel would corrupt the + email features. +- **Absent optional fields in custom features.** Predicates are false. `sum` uses `default`. + `share` with a zero denominator uses `if_empty`. `time_between` with no `from` event uses + `if_absent`. There is never a NaN: the compiler proves every op total, and the orchestrator + rejects a non-finite value as a pack error. +- **Duplicates and out-of-order events.** Ingest idempotency is unchanged. Features are + set-functions with `(at, id)` tie-breaks. +- **Clock skew and future events.** Excluded from custom windows, but they schedule a rescore + at their `at`. The frozen legacy exceptions are listed in §5.5. +- **Disabling a pack that rules still reference.** The load fails with `feature_not_enabled`. + Past verdicts stay; they are provenance. +- **A tenant enables `email` for a non-email channel.** `email.channels` must name declared + channels whose `destination` kind is `domain`. Otherwise the load fails (`pack_params_invalid`). + This stops webmail and domain logic running over hashes. +- **Alias misuse.** A rule lists both `sends_1h` and `email.sends_1h` → `duplicate_feature`. A + custom feature named after an alias is impossible because of the `custom.` prefix. +- **Hostile config.** There is no code and no regex. Every string set is hashed at load. Every + size is capped. The loader is fuzzed with the §5.7 limits as the oracle. +- **Hostile events against a custom feature.** An attacker can't exceed the per-subject scan or + deadline bounds. Flooding one subject raises only that subject's cost, which the existing + per-subject budgets cap. `distinct` memory is O(cap). +- **Brand pack on a product whose display names are routinely brand-adjacent**, such as a + marketplace reselling branded goods. `brand.*` stays shadow until that tenant's own labels + justify a weight. That tenant can also point `brand.extra` at an empty list and a narrowed + public list through `pack_params.brand.list` (§12 Q10). + +## 7. Worked examples (fictional) + +All three products, their names, ids and domains are invented. Timestamps use the fictional +2031 convention. + +### 7a. File sharing: malware-distribution burst ("Driftbox") + +The pattern: a fresh account uploads an executable or archive, creates many public share links +quickly, and those links are downloaded from many distinct networks within an hour. + +```yaml +tenant: driftbox +packs: [core@1, brand@1] # no email pack: Driftbox doesn't deliver mail +vocabulary: + version: 1 + resource_kinds: + workspace: {role: workspace} + api_token: {role: credential} + folder: {role: content} + types: + share.link_created: + fields: + visibility: {kind: enum, values: [public, org, private]} + file_kind: {kind: enum, values: [document, archive, executable, image, other]} + size_bytes: {kind: number, min: 0, integer: true} + share.downloaded: + fields: + link_hash: {kind: hash} + downloader_ip24: {kind: hash} # producer-keyed hash; never a raw IP + downloader_is_owner: {kind: bool} +features: + - name: custom.public_links_1h + version: 1 + description: public share links created in the last hour + count: {type: share.link_created, where: {field: visibility, eq: public}} + window: 1h + transform: {log1p: true, cap: 6} + - name: custom.risky_file_link_share_24h + version: 1 + description: share of new links pointing at executables or archives + share: + type: share.link_created + match: {field: file_kind, in: [executable, archive]} + window: 24h + transform: {cap: 1} + - name: custom.distinct_downloader_nets_1h + version: 1 + description: distinct downloader /24 networks, excluding the owner + distinct: + type: share.downloaded + where: {field: downloader_is_owner, eq: false} + field: downloader_ip24 + window: 1h + transform: {log1p: true, cap: 9} # track_max = ceil(expm1(9)) = 8103 ≤ 10,000 + - name: custom.download_peak_10m_vs_history + version: 1 + description: 10-minute download peak relative to the account's own past + peak: {type: share.downloaded, where: {field: downloader_is_owner, eq: false}, size: 10m} + window: 24h + relative_to_history: {lookback: 30d, exclude_recent: 24h, age_decay: {}} + transform: {log1p: true, cap: 6} + - name: custom.signup_to_first_public_link_min + version: 1 + description: minutes from sign-up to the first public link + time_between: + from: {type: subject.created} + to: {type: share.link_created, where: {field: visibility, eq: public}} + until_now: true + if_absent: 1440 + transform: {cap: 1440} + prior_sign: "-" # faster = riskier +rules: + - name: core_starter + mode: shadow + scorer: local + weights: packs/core/starter.yaml + inputs: [core.subject_age_h, core.credential_velocity_1h, core.resource_velocity_1h, + core.declines_before_first_success, core.first_funding_prepaid, + core.linked_deleted_n, core.linked_labelled_abusive_n, core.burst_ratio_24h_vs_lifetime] + labels: [benign, abusive] + benign_label: benign + threshold: 0.6 + - name: malware_burst + mode: shadow + scorer: local + weights: uniform + inputs: [custom.public_links_1h, custom.risky_file_link_share_24h, + custom.distinct_downloader_nets_1h, custom.download_peak_10m_vs_history, + custom.signup_to_first_public_link_min, brand.name_match] + labels: [benign, abusive] + benign_label: benign + threshold: 0.7 +``` + +Sample events: + +```json +{"id":"db-001","subject":"acct_example_db_1","type":"subject.created","at":"2031-03-02T09:00:00Z","links":{"email_hash":"<64-hex>"},"data":{"channel":"signup","email_domain_class":"disposable"}} +{"id":"db-002","subject":"acct_example_db_1","type":"resource.created","at":"2031-03-02T09:01:10Z","data":{"kind":"workspace","name":"Official Document Center"}} +{"id":"db-003","subject":"acct_example_db_1","type":"share.link_created","at":"2031-03-02T09:03:00Z","data":{"visibility":"public","file_kind":"archive","size_bytes":812345}} +{"id":"db-004","subject":"acct_example_db_1","type":"share.link_created","at":"2031-03-02T09:03:20Z","data":{"visibility":"public","file_kind":"executable","size_bytes":402112}} +{"id":"db-005","subject":"acct_example_db_1","type":"share.downloaded","at":"2031-03-02T09:05:02Z","data":{"link_hash":"lk_4f1c9a0e7b2d","downloader_ip24":"ip_9a1b2c3d4e5f","downloader_is_owner":false}} +``` + +The benign counterpart fixture: an older workspace sharing documents with an organisation. Its +links are `org`-visibility, with a few downloads from two networks. + +### 7b. Payments or marketplace: card testing ("Tallyport") + +The pattern: a merchant account (the subject) pushes many small charge attempts across many +distinct cards, most of them declined, in short bursts. + +```yaml +tenant: tallyport +packs: [core@1, brand@1] # core also scores the merchant's own onboarding payments +vocabulary: + version: 1 + resource_kinds: + api_key: {role: credential} + storefront: {role: workspace} + types: + charge.attempted: + fields: + outcome: {kind: enum, values: [succeeded, declined, blocked]} + decline_code: {kind: enum, values: [insufficient_funds, do_not_honor, incorrect_cvc, expired_card, fraudulent, other]} + amount_minor: {kind: number, min: 0, integer: true} + card_hash: {kind: hash} +features: + - name: custom.declines_10m_peak + version: 1 + description: largest number of declined charges in any 10 minutes today + peak: {type: charge.attempted, where: {field: outcome, eq: declined}, size: 10m} + window: 24h + transform: {log1p: true, cap: 7} + - name: custom.distinct_cards_1h + version: 1 + description: distinct cards charged in the last hour + distinct: {type: charge.attempted, field: card_hash} + window: 1h + transform: {log1p: true, cap: 7} + - name: custom.small_charge_share_1h + version: 1 + description: share of charges at or under 2.00 in minor units + share: {type: charge.attempted, match: {field: amount_minor, lte: 200}} + window: 1h + transform: {cap: 1} + - name: custom.decline_share_1h + version: 1 + description: share of charges declined + share: {type: charge.attempted, match: {field: outcome, in: [declined, blocked]}} + window: 1h + transform: {cap: 1} + - name: custom.cvc_declines_vs_history + version: 1 + description: CVC/expiry declines this hour vs the merchant's own past + count: + type: charge.attempted + where: {all: [{field: outcome, eq: declined}, {field: decline_code, in: [incorrect_cvc, expired_card]}]} + window: 1h + relative_to_history: {lookback: 30d, exclude_recent: 24h, age_decay: {}} + transform: {log1p: true, cap: 6} +rules: + - name: card_testing + mode: shadow + scorer: local + weights: uniform + inputs: [custom.declines_10m_peak, custom.distinct_cards_1h, custom.small_charge_share_1h, + custom.decline_share_1h, custom.cvc_declines_vs_history, core.credential_velocity_1h] + labels: [benign, abusive] + benign_label: benign + threshold: 0.7 +``` + +```json +{"id":"tp-101","subject":"acct_example_tp_7","type":"charge.attempted","at":"2031-06-10T02:14:01Z","data":{"outcome":"declined","decline_code":"incorrect_cvc","amount_minor":100,"card_hash":"cd_1a2b3c4d5e6f"}} +{"id":"tp-102","subject":"acct_example_tp_7","type":"charge.attempted","at":"2031-06-10T02:14:04Z","data":{"outcome":"declined","decline_code":"expired_card","amount_minor":100,"card_hash":"cd_7f8e9d0c1b2a"}} +{"id":"tp-103","subject":"acct_example_tp_7","type":"charge.attempted","at":"2031-06-10T02:14:09Z","data":{"outcome":"succeeded","amount_minor":100,"card_hash":"cd_0f1e2d3c4b5a"}} +``` + +The benign counterparts: a storefront with steady larger charges and an ordinary decline rate, +and a storefront whose flash sale has high volume but few distinct-card declines. The second +exercises `relative_to_history`, because the merchant's own past peaks raise the baseline. + +`payment.attempt` is **not** used for the charges. In the core vocabulary, `payment.attempt` +means the subject paying the product, and it feeds onboarding facts. Card testing is the +merchant's product activity, so it is a custom type. The example makes that distinction +explicit. + +### 7c. Chat or community: spam invites ("Hearthchat") + +The pattern: new accounts with brand-like display names send large volumes of invites to +people outside their own communities, with external links in the invite text. + +This product uses the **neutral built-in** `delivery.sent`, which is what it is for. + +```yaml +tenant: hearthchat +packs: [core@1, brand@1] +vocabulary: + version: 1 + resource_kinds: + profile: {role: identity} + community: {role: workspace} + bot_token: {role: credential} + channels: + invite: + destination: hash # keyed community id + destination_class: [own_community, other_community] + direct_message: + destination: hash +features: + - name: custom.invites_10m_peak + version: 1 + description: largest invite fan-out in any 10 minutes today + peak: {type: delivery.sent, where: {field: channel, eq: invite}, sum: {field: recipient_count, default: 1, cap_each: 50}, size: 10m} + window: 24h + transform: {log1p: true, cap: 7} + - name: custom.distinct_invitees_1h + version: 1 + description: distinct invitees in the last hour + distinct: {type: delivery.sent, where: {field: channel, eq: invite}, field: recipient_hash} + window: 1h + transform: {log1p: true, cap: 7} + - name: custom.external_invite_share_24h + version: 1 + description: share of invites to communities the sender doesn't own + share: + type: delivery.sent + where: {field: channel, eq: invite} + match: {field: destination_class, eq: other_community} + window: 24h + transform: {cap: 1} + - name: custom.linked_invite_share_24h + version: 1 + description: share of invites carrying an external link + share: + type: delivery.sent + where: {field: channel, eq: invite} + match: {field: link_host, exists: true} + window: 24h + transform: {cap: 1} + - name: custom.signup_to_first_invite_min + version: 1 + description: minutes from sign-up to the first invite + time_between: + from: {type: subject.created} + to: {type: delivery.sent, where: {field: channel, eq: invite}} + until_now: true + if_absent: 1440 + transform: {cap: 1440} + prior_sign: "-" +rules: + - name: invite_spam + mode: shadow + scorer: local + weights: uniform + inputs: [custom.invites_10m_peak, custom.distinct_invitees_1h, custom.external_invite_share_24h, + custom.linked_invite_share_24h, custom.signup_to_first_invite_min, + brand.name_match, brand.title_match, core.linked_deleted_n] + labels: [benign, abusive] + benign_label: benign + threshold: 0.7 +``` + +```json +{"id":"hc-201","subject":"acct_example_hc_3","type":"resource.created","at":"2031-08-01T18:00:05Z","data":{"kind":"profile","name":"Support Team - Official"}} +{"id":"hc-202","subject":"acct_example_hc_3","type":"delivery.sent","at":"2031-08-01T18:02:11Z","data":{"channel":"invite","destination":"cm_5e6f7a8b9c0d","destination_class":"other_community","recipient_hash":"iv_0a1b2c3d4e5f","title":"You have been selected - claim now","link_host":"claim-prize.example.test"}} +``` + +The benign counterpart: a community organiser inviting 20 people over an evening to their own +community, with no links. It exercises `external_invite_share_24h` = 0 and a low peak. + +**What the three examples demonstrate:** none needs the email pack. All three reuse `core` and +`brand`. Every product-specific signal is declarative. 7c uses the neutral delivery type +directly, and 7a and 7b show that custom types cover what `delivery.sent` doesn't. Slice G6 +commits each example as a loadable profile with fixtures, which is success criterion 2. + +## 8. Migration plan for e2a + +**Emitter: no change required.** e2a keeps emitting `content.sent` and the rest of main §4.12, +and S6 is built as currently specified. Switching to `delivery.sent{channel: email}` later is +optional and equivalent by construction; the translation test (below) proves it. + +**Config mapping.** The default profile (§5.8) serves e2a until G7. G7 then commits +`config/tenants/e2a.yaml`: +- `packs: [core@1, email@1, brand@1]` +- `vocabulary`: the implicit legacy declaration from §5.4, made explicit: `agent` → + `identity`, `key` → `credential` with S2b's aliases, and `email` → `{destination: domain}`. +- `pack_params`: the brand `extra` path (the private list, as `--brands-extra` today) and + `email.webmail_set: default` (`config/packs/email/webmail.yaml`, moved from + `config/webmail.yaml`). +- The `new_account_velocity` inputs rewritten through the alias table. The weights file moves + to `config/tenants/e2a/weights.yaml` with namespaced keys and the same values. +- Floors move to `eval/tenants/e2a/floors.yaml` with `profile: tenant:e2a`, with numbers + unchanged. + +**Proof of identical scores: the golden replay.** +1. **G0 runs first, on the pre-migration code.** It adds `cmd/abusekit golden` (test-only + build tag), which replays every fixture. Each subject is scored by the real `feature.Extract` + → `core.Plan` → local scorer → `core.Combine` path after every event instant and at every + `NextRescoreAt` the replay produces. The command writes `eval/tenants/e2a/golden/v0.jsonl`, + one row per (fixture, subject, instant): + `{features: {canonical_key: float64-bits-hex}, next_rescore_at, input_hash{rule}, + risk_bits{rule}, score_bits, tier, local_version}`. + It also records the harness `run.json` for the synthetic corpus with volatile fields removed. + Fixtures come from `eval/fixtures/*.jsonl` (including all of #7's), and the generated corpus + comes from `eval/fixtures/synthetic/`. +2. **Every later slice** runs `TestGoldenReplay_BitIdentical`, which compares exactly, row for + row, including the row count. Any diff fails CI and prints the first differing feature. +3. **The translation test** rewrites every `content.sent` in every fixture as + `delivery.sent{channel: email, ...}` (field mapping in §5.4), replays it, and asserts the + same golden file. +4. **A corpus round-trip** loads the corpus-v1 synthetic corpus, exports it as v2, reloads it, + and requires identical `run.json` metrics. + +**Stored state at cutover, if S8 has already shipped** (assumption A1 says it hasn't): +- Verdict `input_hash`: unchanged (canonical keys), so no subject is rescored and no vendor call + repeats. +- Local `Version()`: unchanged (canonical keys). +- Cassettes: unchanged keys. +- Rows stored before G5: `vocab_version` is NULL, which reads as the built-in-only schema. + +**Rollback.** Every slice up to G7 is behaviour-neutral for e2a, and the golden replay guards +it. G7 is a config move. Reverting it restores the default profile, which yields the same +golden. + +## 9. Slices + +These fit after #5 and #7 merge. Each is its own PR with the usual review. + +| # | Slice | Contents | Done when | +| --- | --- | --- | --- | +| G0 | Golden capture | `cmd/abusekit golden` (test build tag); `eval/tenants/e2a/golden/v0.jsonl` generated from `main` after #5 and #7; `TestGoldenReplay_BitIdentical` | Golden committed, generated by pre-migration code; the test passes on `main`; perturbing one weight's last bit fails it | +| G1 | Canonical keys | `FeatureDef` metadata table (name, legacy, quantum, bound, reads) for today's 25 features; `core.Vector`; `inputHash` over canonical keys with quanta from data (the name switch removed); local scorer order and version by canonical key; alias resolution in rules, weights and corpus loaders; the frozen-table CI check | Golden bit-identical; a rules file in flat names and one in namespaced names load to equal `Config`s; `quantizeAgeFeaturesForHash` deleted | +| G2 | Vocabulary and delivery view | `internal/vocab` (built-in schema moved unchanged); `delivery.sent` built-in; `event.View`; `content.sent` projection; declared resource kinds with roles and aliases (implicit legacy declaration); `pkg/abusekit` `DeliverySent` | Golden bit-identical; translation test green; redaction tests moved and green; `delivery.sent` contract tests (happy, `recipient_hash`+count>1 rejection, undeclared channel rejection) | +| G3 | Packs and tenant profiles | `internal/pack` registry, `core`/`email`/`brand` adapters (code moved, not rewritten); `feature.Extract(profile, …)` orchestrator with per-pack failure isolation; `config/tenants/*.yaml` loader; default profile from `rules.yaml`; per-tenant atomic reload; `/healthz` per tenant; `brand.title_match`; new additive core features | Golden bit-identical; `feature_not_enabled` and `pack_requires` load tests; a tenant with only `core` computes no `email.*`; one tenant's bad profile leaves another's reload applied | +| G4 | Declarative features | `internal/pack/custom` compiler and evaluator, all six ops plus the modifier, predicates, limits and cost model, rescore candidates, `tenant_feature_defs`; naive reference evaluator in tests; property tests; loader fuzz; benchmark | Criterion 4 met (equality on 10k randomized histories; p99 ≤ 50 ms at the limits); a declarative `custom.key_velocity_1h` and `custom.resource_total_lifetime` match their Go twins bit for bit on every fixture (lifetime excluding future events, documented); fuzzing finds no over-limit profile that loads | +| G5 | Declared types and redaction | field kinds; masking of `text`; hashing of undeclared fields and types with the abusekit-held per-tenant key; `vocab_version` column; `tenant_vocabularies` ledger and the widening-only check; `strict` mode | Criterion 5 property tests; `vocab_incompatible` tests; e2a golden bit-identical (e2a declares no custom types) | +| G6 | Per-pack eval and bootstrap | `packtest` harness; per-pack starter weights with `sign:`; pack fixtures and floors; `--profile`; corpus-v2 schema and export; uniform priors plus `uniform_not_shadow`; the three §7 profiles committed as fixtures | Every pack passes `packtest`; `make gate` runs e2a and every pack profile; each §7 abusive fixture outranks its benign fixtures under uniform priors (criterion 2); #5's floors file still loads unchanged | +| G7 | e2a explicit profile | `config/tenants/e2a.yaml`, weights and floors moved; flat-name deprecation warning on; docs updated (main §4.5 and §4.12 pointers) | Golden bit-identical against the explicit profile; the default profile is used by no tenant in the hosted config; the reverting diff also passes golden | + +G0 must land before any other G slice. G1 → G2 → G3 are sequential. G4 and G5 can run in +parallel after G3. G6 needs G4. G7 needs G3 (G5 and G6 are optional for it). **Recommended +ordering against the v0 plan:** G0–G3 before S5, so that vendor render templates are born with +namespaced names, and before S8, so that nothing stored needs the bridge in anger. S3b and S6 +are independent of all G slices. + +## 10. Scalability and extensibility + +- **Custom features per tenant.** Bounded by 64 features and 256 cost units. Compiled plans are + cached per `profile_sha`, and a reload recompiles one tenant only. +- **Extraction cost.** One pass per subject: O(E) dispatch plus O(E) per windowed scan, with E + ≤ 50,000. `distinct` memory is O(cap). The existing per-subject budget and the priority queue + are unchanged. What grows is `EventsForSubject`'s load. Later the store can bound the query to + `max(lookback, window) + exclude_recent`, plus `subjects.first_seen` for `start`. That is a + store-only change the pack interface already permits, because `Input.Start` is explicit. +- **Tenants.** About 100 tenants × 64 features is a config and plan-cache concern, not a + database one. Metrics are labelled `{tenant, pack}`, not `{feature}`, to bound cardinality. +- **Made easier later.** + - A new domain pack (`sms`, `marketplace`) is one Go package plus `packtest`. + - A popular custom feature can be promoted into a pack upstream, keeping its name through a + per-pack alias. + - Cross-tenant linking is untouched by this design. + - Weight fitting consumes corpus-v2 per profile. + - CEL for `where` alone, if ever needed, slots in as a new leaf kind. + +## 11. Verification strategy + +The seams tested are the ones callers cross: the config loader (profiles in, errors out), +`feature.Extract` (events in, vector out), `POST /v1/events` (redaction), and the harness +(`--profile`). + +1. The golden replay and the translation test (§8), in every slice's CI. +2. `packtest` for every pack (§5.9). +3. The DSL conformance suite: the naive reference versus the compiled evaluator on randomized + histories (shuffled order, future events, ties, empty windows, anchored windows before + `start`); table tests for every predicate leaf and each op's edge (window boundary + inclusivity, `if_absent`, `if_empty`, cap and log1p order, `distinct` early stop). +4. Loader tests for every §5.8 code, plus the fuzzer, with limits as the oracle. +5. Redaction property tests: mask versus reject per kind; undeclared fields hashed; no email + shape stored outside masked text. +6. HTTP contract tests: `delivery.sent`, declared types, `strict` mode, per-tenant `/healthz`. +7. A benchmark at the limits (criterion 4). +8. **Most likely regressions:** summation order (caught by golden), a quantum lost for the age + features (golden input hashes), the S2b alias list drifting in the vocabulary move (golden + `core.credential_*`), and window-boundary off-by-one in a DSL op (conformance suite). +9. **Manual checks:** load each §7 profile in a local instance, post its sample events, and read + the shadow signals and reasons. + +## 12. Open questions (owner decisions) + +1. **Ordering.** Land G0–G3 before S5 (vendor adapters) and S8 (hosted deploy)? Recommended: + yes. +2. **What S6 emits.** `content.sent` as designed (recommended, no churn), or the neutral + `delivery.sent{channel: email}`? +3. **Undeclared-type storage change.** Approve moving unknown types from "kept as-is" to + "strings keyed-hashed" (a pre-GA semantic change, §5.4 and §5.6)? +4. **abusekit-held per-tenant redaction key.** This is a new secret per tenant, separate from + the producer's link key. Approve? +5. **Legacy window quirks.** `core.resource_total` and `core.credential_total` count + future-dated events, and `email.first_day_distinct_domains` has an inclusive end. Keep them + frozen in `@1` and harmonise in `@2` later (recommended), or harmonise now and accept a golden + diff? +6. **The `credential` rename.** `key_*` → `core.credential_*`: accept the neutral role vocabulary + (`credential`, `identity`, `workspace`, `content`, `other`)? +7. **Limits.** 64 features, 256 units, 30-day max window, 50,000 events scanned, 250 ms deadline. + Confirm or adjust. +8. **Bootstrap policy.** Uniform priors are shadow-only, with no fitting in scope. Confirm that + advise always needs a hand-set or fitted weights file plus a passing gate. +9. **CEL fallback.** Pre-approve CEL for `where` only if the closed predicate set proves + insufficient, or require a new design pass? +10. **Brand list scope.** Is `config/packs/brand/brands.yaml` one global public list, or may a + tenant narrow it (`pack_params.brand.list`) as well as extend it (`extra`)? +11. **The custom namespace.** `custom.*` scoped per tenant (proposed), or `.*` so names + are globally unique in logs and corpora? +12. **Per-tenant reload isolation.** Replaces the main design's whole-config rejection. Confirm. diff --git a/docs/plans/2026-09-27-v0-plan.md b/docs/plans/2026-09-27-v0-plan.md index a50c22e..85471e9 100644 --- a/docs/plans/2026-09-27-v0-plan.md +++ b/docs/plans/2026-09-27-v0-plan.md @@ -35,6 +35,7 @@ design pass) rather than deferred to v1 outright. | S7 | Billing events | ops sidecar: `payment.attempt` (with `card_fingerprint_hash` under the tenant key) and `subscription.changed` | staging checkout produces events | | S8 | Hosted deploy | ops: compose service, Secret Manager keys, tenant config, Terraform alerts for queue depth / budget / drops | abusekit running on prod in shadow | | S9 | Incident evaluation | private backfill of the incident accounts and a benign sample into a private corpus; harness run; report precision/recall/lead-time before first send per account | report reviewed; floors set; decision on `advise` for the local rule | +| G0–G7 | Generic feature packs (own design: `docs/design/2026-09-29-generic-feature-packs.md`) | After S4 (#5) and S2b (#7) merge: golden capture (G0), canonical feature keys (G1), neutral vocabulary and delivery view (G2), packs and tenant profiles (G3), declarative custom features (G4), declared-type redaction (G5), per-pack eval and bootstrap (G6), explicit e2a profile (G7). Recommended before S5 and S8. | per-slice "Done when" in that design's §9; every slice keeps e2a's golden replay bit-identical | S1–S4 (including S3b) are pure abusekit and can run back to back; S5 needs vendor keys; S6–S8 are e2a/ops work that can start after S3 (S3b is not a blocker for them — nothing in S6–S8 depends on From 3c1c32cf9076c85a1a6b09c3334a5236c2881206 Mon Sep 17 00:00:00 2001 From: jiashuoz Date: Tue, 29 Sep 2026 11:38:19 +0800 Subject: [PATCH 2/8] docs(design): rework generic feature packs after adversarial review (rev 2) Drops the canonical-key bridge for a one-time rename with a per-consumer test list and a semantic-identity golden (derived ulp bound). Replaces the event-count bound and wall-clock deadline with time-bounded loading, full onboarding loads, byte caps, a deterministic step budget and truncation as a positive signal. Closes redaction channels: re-HMAC of every hash (length-prefixed, join domains), undeclared values dropped, PSL-checked domains, card/IP/phone scanning, skeleton-only custom text. Adds group_by, sequence, ratio, neighbours with declared link kinds, before_first, and subject kinds, and walks five fictional scenarios. Drops delivery.sent for declared types with field roles. Re-slices into P0-P7. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_014cdM7WyRc3mD3vQNXMTDB8 --- docs/design/2026-09-27-abusekit-design.md | 9 +- .../2026-09-29-generic-feature-packs.md | 2198 +++++++++-------- docs/plans/2026-09-27-v0-plan.md | 2 +- 3 files changed, 1168 insertions(+), 1041 deletions(-) diff --git a/docs/design/2026-09-27-abusekit-design.md b/docs/design/2026-09-27-abusekit-design.md index 2c90006..d11946e 100644 --- a/docs/design/2026-09-27-abusekit-design.md +++ b/docs/design/2026-09-27-abusekit-design.md @@ -6,10 +6,11 @@ see account churn, tiers failed open when a scorer was unavailable, and request cover reads or replay. This revision fixes those and tightens every place an implementer would have had to guess. Changes from r1 are marked **[r2]**. -**Amendment (proposed 2026-09-29):** [`2026-09-29-generic-feature-packs.md`](2026-09-29-generic-feature-packs.md) -splits the built-in features into namespaced `core`/`email`/`brand` packs enabled per tenant, adds a -neutral `delivery.sent` event and product-declared vocabularies, and adds declarative custom -features in YAML. It amends §4.2, §4.3, §4.5, §4.6 and §4.10 below, with bit-identical scores for e2a. +**Amendment (proposed 2026-09-29, revision 2):** [`2026-09-29-generic-feature-packs.md`](2026-09-29-generic-feature-packs.md) +renames the built-in features once into namespaced `core`/`email`/`brand` packs enabled per tenant, adds +product-declared vocabularies (types, field kinds, link kinds, subject kinds) with pseudonymising +redaction, and adds bounded declarative custom features in YAML. It amends §4.2, §4.3, §4.5, §4.6, §4.8 +and §4.10 below, and keeps e2a's feature values and tiers identical. ## 1. Problem statement diff --git a/docs/design/2026-09-29-generic-feature-packs.md b/docs/design/2026-09-29-generic-feature-packs.md index 3c72af9..31ee3d9 100644 --- a/docs/design/2026-09-29-generic-feature-packs.md +++ b/docs/design/2026-09-29-generic-feature-packs.md @@ -1,390 +1,357 @@ -# Generic feature packs, a neutral event vocabulary, and declarative custom features +# Generic feature packs, product vocabularies, and declarative custom features -Status: proposed, 2026-09-29 · owner: Josh Zhang · amends -[`2026-09-27-abusekit-design.md`](2026-09-27-abusekit-design.md) (§4.2, §4.3, §4.5, §4.6, §4.10). -Written against `main` at S3 plus the two open PRs treated as merged: #5 (S4 evaluation harness) -and #7 (S2b send-volume, webmail, recipient and subject-brand features). Section numbers below -refer to this document; `main §x` refers to the main design. +Status: proposed, revision 2 (after adversarial review), 2026-09-29 · owner: Josh Zhang · amends +[`2026-09-27-abusekit-design.md`](2026-09-27-abusekit-design.md) §4.2, §4.3, §4.5, §4.6, §4.8 and +§4.10. Written against `main` at S3, with the two open PRs treated as merged: #5 (S4 evaluation +harness) and #7 (S2b send-volume, webmail, recipient and subject-brand features). `main §x` refers +to the main design. §13 lists what changed from revision 1. ## 1. Problem statement -abusekit's current feature set assumes an email platform. The main design promises a scoring -service for any product that mints accounts, takes payments and lets users create resources. The -built code does not keep that promise: - -- **The vocabulary is email-shaped.** `content.sent` carries `subject_line`, `recipient_domain` - and `recipient_is_own_identity`, fields that only mean something for mail. Resource kinds are an - undeclared convention: `internal/feature` counts `kind == "key"` (plus S2b's spelling aliases) - and treats everything else as a generic resource. -- **Half the features are email features.** Of the 25 features on `main` + #7, - `first_day_distinct_domains`, `self_send_before_external`, `sends_10m_max`, `sends_1h`, - `sends_first_day`, `webmail_recipient_share`, `webmail_sends_1h`, `distinct_recipients_1h` and - `subject_brand_match` read `content.sent` and only mean something for mail. `key_velocity_1h` - and `key_total` assume API keys. A file-sharing, payments or chat product would get these at - zero. -- **Custom event types are dead weight.** Main §4.3 says unknown types are "stored, available to - Go-registered features only". No Go feature reads them, so the only way for a product to add a - signal is to write Go in this repo. -- **Everything is global.** One `config/rules.yaml`, one `config/local_weights.yaml` and one flat - feature namespace serve every tenant. A second tenant can't turn features on or off, and a new - feature name can collide with an existing one. - -**Desired outcome.** A product that is not an email platform can onboard with YAML only. It -enables the packs that fit its domain, declares its event types and field kinds, defines its -product-specific signals as declarative features, and runs in shadow. The first consumer (e2a) -keeps identical scores, bit for bit. +abusekit's current feature set assumes an email platform. The main design promises scoring for +any product that mints accounts, takes payments and lets users create resources. The built code +falls short of that: + +- **The vocabulary is email-shaped.** `content.sent` carries `subject_line`, `recipient_domain` and + `recipient_is_own_identity`. Resource kinds are an undeclared convention: `internal/feature` + counts `kind == "key"` (plus #7's spelling aliases) and treats every other kind as generic. +- **Much of the feature set is email-specific.** Nine of the 25 features on `main` + #7 are email + features, and two more assume API keys. A file-sharing, payments, chat, developer-API or + AI-inference product would get these at zero. +- **Custom event types are dead weight.** Main §4.3 says unknown types are stored "for + Go-registered features only". No such feature exists, so the only way to add a signal is to + write Go in this repo. +- **Everything is global and flat.** There is one `rules.yaml`, one `local_weights.yaml`, one + feature namespace, and one subject kind: the account. + +**Desired outcome.** A product that is not an email platform onboards with a private YAML profile. +In that profile it: +- declares its event types, field kinds, link kinds and subject kinds; +- enables the packs that fit its domain; +- defines its product-specific signals as declarative features; +- runs them in shadow. + +Every new value is pseudonymised or dropped at ingest. For e2a, feature values, tiers and rescore +times stay identical, and risk moves by no more than a stated floating-point bound. ### Success criteria (measurable) -1. **Bit-identical migration.** A golden replay covers every committed fixture - (`eval/fixtures/*.jsonl`, `eval/fixtures/synthetic/`, and #7's fixtures). It scores after every - event and at every scheduled rescore instant. Each recorded feature value, `NextRescoreAt`, - per-rule input hash, risk (compared as `math.Float64bits`), tier and local-scorer `Version()` - must be identical before and after every slice in §9. The same holds for the harness's - `run.json` metrics once timestamps and git sha are removed. -2. **Zero-Go onboarding.** The three fictional products in §7 load from a tenant YAML file with - no Go changes. With uniform priors (§5.9), each product's abusive fixture scores above every - one of its benign fixtures. -3. **Enablement is enforced.** For a tenant without the `email` pack, no `email.*` feature is - computed, stored or rendered. A rule that references one fails the config load with +1. **Semantically identical migration.** A golden replay covers every committed fixture: all of + `eval/fixtures/*.jsonl` (including #7's) and the synthetic corpus. It scores after every event + and at every scheduled rescore instant. Across the rename and every later slice: + - every feature value is identical as `math.Float64bits`, compared under the rename map; + - every `NextRescoreAt` and every tier is identical; + - every risk and score satisfies `|Δ| ≤ 1e-12` (§5.2 derives the bound). + + Input hashes, cassette keys, the `run.json` rule/weights SHAs and the local `Version()` may + change **only in the rename slice**. The golden records the before and after values. +2. **Five generic scenarios.** Each scenario in §7 loads as a fictional profile, with zero Go + changes for everything §7 marks as declarative. Held-out fixtures are authored after the + features and never used while writing them. On those fixtures, with uniform priors (§5.9), + each scenario's abusive fixtures outrank every one of its benign fixtures. +3. **Enablement is enforced.** For a tenant without a pack, none of that pack's features is + computed, stored or rendered. A rule that references one fails to load with `feature_not_enabled`. -4. **Declarative features are correct and bounded.** For every operation in §5.5, the compiled - evaluator equals a naive O(n²) reference implementation on 10,000 randomized histories, - including shuffled arrival order and future-dated events. Per-subject extraction p99 is ≤ 50 ms - on 2 vCPU for a tenant at the limits in §5.7: 64 custom features over a history of 50,000 - events. -5. **Privacy by construction.** A property test shows that no stored value of a declared `text` - field matches the email-shape matcher. Values in undeclared fields are stored only as keyed - hashes, and config can never introduce a regex or executable code. The loader fuzz test finds - no profile that passes validation and breaks a limit in §5.7. +4. **The DSL is correct and bounded.** + - Every operation equals a naive reference implementation on 10,000 randomized histories, + including shuffled arrival, ties and future-dated events. + - Per-subject time is measured **end to end**: store load, JSON decode, view projection, pack + extraction and orchestration. For a profile at the §5.7 limits with history at its byte cap, + it is p99 ≤ 50 ms on the reference 2-vCPU host. + - The step budget is calibrated from that benchmark (§5.7), so the budget cuts off work before + the latency target is breached. +5. **Privacy by construction.** Property tests show that, at rest: + - no stored value matches an email, card (Luhn), IP or phone shape outside a masked `text` + field; + - every stored `hash` value is an abusekit-keyed HMAC; + - undeclared data keeps only type, time and field names. + + The loader fuzzer finds no profile that passes validation while breaking a §5.7 limit. +6. **Flooding doesn't evade.** For every built-in and DSL feature, a property test adds cheap + events totalling up to 10 times the byte cap to an abusive fixture. Risk must not fall (§5.7). ## 2. Goals and non-goals **Goals** -- Split the built-in features into three packs: `core` (product-neutral), `email`, and `brand` - (display-name impersonation). Each pack is registered in Go and enabled per tenant. -- Namespace every feature (`core.subject_age_h`, `email.sends_10m_max`). Keep a frozen alias - table for today's flat names, so stored verdicts, corpora, cassettes, floors and weights stay - valid. -- Add a neutral delivery event (`delivery.sent`) and product-declared resource kinds, channels - and custom event types. `content.sent` stays accepted forever. -- Add declarative custom features in YAML. The set is closed: count, distinct, share, peak, - time-between, and a history-relative modifier. Every feature carries a mandatory cap and - transform, has deterministic semantics, and is validated and versioned at load. -- Per-tenant redaction for declared fields, driven by field kinds. -- Per-pack starter weights, fixtures, floors and golden-sign/mutation tests. A shadow-only - uniform-prior mode for a tenant that has no labels yet. +- Rename the built-in features into namespaced names **once**, while no production verdicts or + vendor cassettes exist. There is no compatibility bridge (§5.2). +- Per-tenant profiles, kept in a private config mount. Only fictional example profiles live in + this repo. +- Product-declared vocabularies: + - custom event types, with a field kind and optional role for each field; + - extension fields on built-in types; + - an `activity` role for types; + - resource-kind roles; + - declared link kinds; + - subject kinds beyond `account`. +- Redaction: + - pseudonymise every declared hash with an abusekit-held key; + - drop every undeclared value; + - validate domains against the public suffix list (PSL); + - mask text or store it as a skeleton; + - scan for card, IP and phone shapes. +- A closed declarative feature DSL covering the five scenarios in §7. Every feature has a + mandatory cap, deterministic semantics and a deterministic step budget. +- Packs (`core`, `email`, `brand`) enabled per tenant, each with starter weights, fixtures, floors + and a `packtest` harness. Uniform priors, shadow-only, until a tenant has labels. **Non-goals** -- Arbitrary code or expressions in config: no CEL, no regex, no WASM, no Go plugins (§5.12). -- Learning weights from labels (`abusekit fit`). Bootstrapping stays in shadow until an operator - hand-tunes weights or a later design adds fitting (§12 Q8). -- Cross-tenant feature sharing or linking. Main §2 already defers this. -- Changes to scoring math. `core.Plan`, `core.Combine` and the local logistic model are - unchanged. Only their feature keys and quantization metadata move to a table (§5.2). -- A runtime API for declaring vocabularies. Declarations are reviewed config, like rules (§5.6). +- Code or expressions in config: no CEL, regex, WASM or plugins (§5.12). +- Weight fitting (§12 Q8) and cross-tenant linking (main §2). +- Changes to `Plan`/`Combine` math, apart from the stage-gate fix (§5.8) and per-tenant rule sets. +- Runtime vocabulary declaration over HTTP (§5.11). +- A neutral built-in "delivery" event. Revision 1 proposed `delivery.sent`; revision 2 drops it + (§5.4, §13). ## 3. Relevant context and constraints -**Code this design must fit (on `main` + #5 + #7):** -- `internal/event`: structural `Validate`, plus a static redaction `schema` map keyed by type. - Its `RedactionSchemaVersion` is 2 after #7. Unknown types keep every key after a recursive - leak scan; unlisted keys of known types are dropped. -- `internal/feature`: a monolithic `Extract(ctx, tenant, subject, events, neighbors, windows, - brands, webmail)` returns one fixed struct, `Features`, with 25 fields and a hand-written - `Map()`. `Names` feeds `config.FeatureSet`. `nextRescoreAt` hard-codes the windowed types - (`resource.created`, `content.sent`). -- `internal/core`: `inputHash` JSON-encodes the rule's feature map keyed by name. - `quantizeAgeFeaturesForHash` special-cases `subject_age_h` and `upgrade_delay_min` **by name**. -- `internal/model/local`: sums `weight × feature` in **sorted feature-name order** (S10). Its - `Version()` is a SHA-256 of the JSON-encoded `Weights`, whose weight map is keyed by name. - Renaming a feature therefore changes both the floating-point summation order and the version - hash. The migration must neutralise both (§5.2). -- `eval` (#5): corpus-v1 rows carry `input.features` keyed by flat name. Cassettes are keyed on - `(scorer, scorer_version, model, prompt_version, input_hash)`. `floors.yaml` entries are keyed - on `(rule, scorer, slice)`. `eval/gen` generates the synthetic corpus. -- Worker and config: rules and weights are global, and keys already carry a `tenant`. - -**Patterns to reuse:** load-time validation that rejects the whole document (main §4.5). Data -files loaded at boot (`brands.yaml`, `webmail.yaml`). Pure core with injected dependencies. The -compare-and-clear worker queue. Expand-only migrations. - -**Assumptions** (unconfirmed ones are repeated in §12): -- A1. Neither vendor adapter (S5) nor the hosted deploy (S8) has shipped. No production verdicts - and no vendor cassettes exist yet. The design still keeps them stable (§5.2) in case the order - changes. -- A2. At most about 100 tenants and about 64 custom features per tenant. Retention of 90 days - for event text (main §4.11) bounds any lookback. -- A3. Products can compute keyed hashes for identifiers they want to count distinctly, as e2a - already does for `recipient_hash`. As a fallback, abusekit hashes undeclared fields with its - own per-tenant key (§5.6). +**Code this design touches (`main` + #5 + #7):** +- `internal/event`: + - `Validate`, and a static `schema` for redaction (`RedactionSchemaVersion` is 2 after #7). + - Unknown types keep every key after a leak scan; unlisted keys of known types are dropped. +- `internal/feature`: + - a monolithic `Extract` returns a fixed 25-field `Features`, with `Map()` and `Names`; + - `nextRescoreAt` hard-codes `resource.created` and `content.sent`; + - `firstSeenAt` is the minimum event `at`. +- `internal/core`: + - `inputHash` JSON-encodes the rule's features keyed by name; + - `quantizeAgeFeaturesForHash` switches on the literals `subject_age_h` and `upgrade_delay_min`; + - `stageSkip` reads the literal `features["subject_age_h"]`; + - `maxRiskByScorer` takes the maximum over **every** local rule, shadow rules included. +- `internal/model/local`: sums in sorted feature-name order, and `Version()` hashes the whole + JSON-encoded `Weights` struct. +- `internal/worker`: + - `renderReason` prints flat feature names into every stored verdict reason; + - `computeVerdict` builds `currentRuleNames` from a global config. +- `internal/serve`: `currentRuleNames()` reads the global config. +- `internal/store`: + - `EventsForSubject` loads every event, ordered by `(at, seq)`; + - `corpus_examples.features` is JSON keyed by feature name; + - `subjects.first_seen_at` keeps `LEAST(existing, new)`. +- `eval` (#5): + - `hashScoreRequest` keys cassettes over `req.Features` by name; + - `run.json` records `dataset_sha`, `rule_sha` and `weights_sha`; + - corpus-v1 rows key features by name. + +**Assumptions** (unconfirmed ones repeat in §12): +- A1. S5 (vendor adapters) and S8 (hosted deploy) have not shipped. No production verdicts, corpus + rows or vendor cassettes exist yet. That is what makes a one-time rename safe, and why the rename + slice must land before S5 and S8. +- A2. At most about 100 tenants and at most 64 custom features per tenant. Event text retention is + 90 days (main §4.11). +- A3. Producers can compute a keyed hash of any identifier they want counted. abusekit re-hashes + it anyway (§5.6), so a producer's mistake can't turn into stored personal data. ## 4. Proposed design: overview ``` - tenant profile (config/tenants/.yaml) - ├─ packs: [core@1, email@1, brand@1] - ├─ vocabulary: resource kinds, channels, custom types + field kinds - ├─ features: declarative custom.* definitions - ├─ rules: inputs reference namespaced features - └─ weights: per rule (file, or `uniform`) - │ compile + validate (per tenant, atomic) - ▼ -ingest ──▶ vocab.Redact(tenant) ──▶ store (wire form + vocab_version) - │ -worker ──▶ vocab.View (content.sent → delivery view) ──▶ feature.Extract(profile) - │ for each enabled pack, in fixed order: - │ core (Go) · email (Go) · brand (Go) · custom (compiled YAML) - ▼ - Vector{values, canonical keys, quanta} - │ - core.Plan / Combine (unchanged math) +private config mount: tenants//profile.yaml + history/NNNN.yaml (append-only, CI-checked) + packs · vocabulary (types, fields+kinds+roles, extensions, link kinds, subject kinds) + features (custom.*) · rules · weights + │ compile + validate per tenant (atomic, isolated) + ▼ +ingest ─▶ vocab.Redact(tenant): leak scan · kind rules · re-HMAC hashes · drop undeclared + ─▶ store: event row + vocab_version + (subject_kind, subject) index rows (≤ 4 per event) +worker (per-tenant fair queue) ─▶ loader: time-bounded + onboarding-full + byte cap + ─▶ feature.Extract(profile): core · email · brand (Go) · custom (compiled DSL), under a + deterministic step budget + ─▶ core.Plan / Combine (math unchanged) ``` | Module | Interface | Deletion test | | --- | --- | --- | -| `internal/vocab` (new) | `Compile(builtin, decl) (*Vocabulary, error)`; `(*Vocabulary).Redact(*event.Event) error`; `(*Vocabulary).View(event.Event) event.View` | Without it, redaction, kind declarations and the `content.sent` → delivery mapping spread across ingest and every pack. Keep. | -| `internal/pack` (new) | `Pack` interface (§5.1) + registry; adapters `core`, `email`, `brand` (Go) and `custom` (compiled YAML) | Four adapters, so the seam is real. Without it, `feature.Extract` stays a monolith that only grows. | -| `internal/pack/custom` (new) | `Compile(tenant, []Def, *Vocabulary) (Pack, error)` | Holds all DSL semantics. Deleting it leaves products writing Go again. Keep. | -| `internal/feature` | `Extract(ctx, profile, subject, events, now) (Result, error)`: runs enabled packs, merges, min-reduces rescore instants | Becomes a thin orchestrator. It stays because the worker, evaluate and eval all call exactly this one function. | -| `internal/config` | `LoadProfiles(dir, deps) (map[tenant]*Profile, []error)` | Absorbs `rules.yaml`; per-tenant atomic reload. | -| `internal/core` | unchanged signatures; `Plan` takes `Vector` instead of `map[string]float64` | The only change is that canonical keys and quanta come from data (§5.2). | - -The existing S2b brand, webmail and window helpers move into the packs that own them, with their -code unchanged. That is what keeps the migration bit-for-bit. +| `internal/vocab` (new) | `Compile(builtin, decl) (*Vocabulary, error)`; `(*Vocabulary).Redact(*event.Event, Keys) error` | Without it, redaction, kinds, roles and extensions spread across ingest and the packs. Keep. | +| `internal/pack` (new) | `Pack` interface + registry; adapters `core`, `email`, `brand` (Go) and `custom` (compiled DSL) | Four adapters, so the seam is real. Without it, `feature.Extract` stays a monolith. | +| `internal/pack/custom` (new) | `Compile(tenant, []Def, *Vocabulary, CostTable) (Pack, error)` | Holds every DSL semantic. Keep. | +| `internal/feature` | `Extract(ctx, profile, subject, History, now) (Result, error)` | A thin orchestrator. The worker, evaluate and eval all call this one function. | +| `internal/store` | `LoadHistory(ctx, tenant, subjectRef, LoadPlan) (History, error)` replaces `EventsForSubject` for scoring | Bounded loading lives in one place. | +| `internal/secret` (new) | `Keys` interface (§5.6), provider-agnostic | Two adapters (a file, a cloud secret manager). Keep. | +| `internal/config` | `LoadProfiles(mount, deps) (map[tenant]*Profile, map[tenant][]error)` | Absorbs `rules.yaml`. Per-tenant atomic reload. | ## 5. Proposed design: detail ### 5.1 Packs: registration, enablement, dependencies -A **pack** is a named, versioned bundle of feature definitions and computation. It may also -carry data files (brand lists, webmail lists), starter weights, fixtures and floors. Packs are -compiled into the binary and registered in `internal/pack/registry.go`. A tenant **enables** -packs; it never supplies pack code. - ```go -// Pack computes a namespaced group of features for one subject. type Pack interface { - ID() ID // {Name: "email", Version: 1} - Requires() []string // packs whose features/views it reads, e.g. email@1 → [core] - Features() []FeatureDef // static metadata, see below - // Extract is pure: same Input → same Output, independent of event order. - // It must exclude events with At > Input.Now from every window (§5.5 notes one frozen - // legacy exception). + ID() ID // {Name: "email", Version: 1} + Requires() []string // e.g. email@1 → [core] + Features() []FeatureDef + // Extract is pure: same input → same output, whatever the arrival order. Every step it + // takes is charged to in.Budget (§5.7). Extract(ctx context.Context, in Input) (Output, error) } type FeatureDef struct { - Name string // "email.sends_10m_max"; grammar §5.3 - Legacy string // frozen flat alias ("sends_10m_max"), "" for features born namespaced - HashQuantum float64 // input-hash bucket; 0 = exact (§5.2) - Bound float64 // max value after transform; used by uniform priors (§5.9) - Reads []string // view types read ("delivery", "resource.created", ...), for dispatch - // and rescore scheduling - RequiresPacks []string // feature-level dependency, e.g. email.subject_brand_match → [brand] + Name string // "email.sends_10m_max"; grammar §5.3 + HashQuantum float64 // input-hash bucket (replaces the name switch in core) + Bound float64 // max value; used by uniform priors (§5.9) + PriorSign int8 // +1 / -1; uniform-prior and golden-sign default + TruncationDir int8 // -1 if the value can only fall when history is truncated, else 0/+1 (§5.7) + Reads []string // types/roles read: dispatch, load plan, rescore scheduling + Lookback Lookback // history the feature needs (§5.7 load plan) + RequiresPacks []string // e.g. email.subject_brand_match → [brand] + SubjectKinds []string // kinds this feature is defined for; default [account] } type Input struct { - Tenant, Subject string - Events []event.View // stored events projected through the tenant vocabulary - Now time.Time - Start time.Time // account start: subject.created.account_created_at if present, - // else earliest event At (S2b R8) - Neighbors NeighborEvidence // resolved once by the orchestrator, only when core is enabled - Params PackParams // this tenant's validated per-pack settings (lists, channels, kinds) + Tenant string + Subject SubjectRef // {Kind, ID} + History History // bounded, ordered (at, producer, id); truncation flags + Now time.Time + Start time.Time // subject.created.account_created_at, else subjects.first_seen_at + Neighbors NeighborResolver // store-backed, budgeted, cached per extraction + Params PackParams + Budget *StepBudget } type Output struct { - Values map[string]float64 // exactly the names in Features(); missing = bug, rejected - Rescore []time.Time // candidate instants at which some value changes with no new event + Values map[string]float64 + Rescore []time.Time + Truncated bool // this pack's budget or the history's truncation was hit } ``` -**Enablement** lives in the tenant profile: `packs: [core@1, email@1, brand@1]`. The rules are: -- `core` is always enabled and is implied if omitted. Every other pack is opt-in. -- A pin names a major version. A pack changes feature semantics only by shipping a new major - version (`email@2`) next to the old one for at least one release. Semantic changes never - happen in place. Every verdict records the pinned versions in `profile_sha` (§5.10). -- `Requires` must be satisfied by the enabled set. Otherwise the tenant fails to load with - `pack_requires`. -- A feature whose `RequiresPacks` are not all enabled is **not registered** for the tenant. For - example, `email.subject_brand_match` needs `brand`; enable `email` without `brand` and that one - feature doesn't exist for the tenant. -- The orchestrator runs packs in fixed order: `core`, `email`, `brand`, then `custom`. It merges - their `Values` and takes the earliest `Rescore` instant, coalesced to the existing 5-minute - buckets. No two packs share a namespace, so collisions are impossible by construction. - -**Pack contents after the split** (legacy name → namespaced name): - -| Pack | Features | -| --- | --- | -| `core@1`, account and onboarding | `subject_age_h`→`core.subject_age_h` (quantum 1) | -| `core@1`, payment | `upgrade_delay_min`→`core.upgrade_delay_min` (quantum 60), `upgraded`→`core.upgraded`, `declines_before_first_success`→`core.declines_before_first_success`, `first_funding_prepaid`→`core.first_funding_prepaid`, `fingerprint_seen_on_other_subjects`→`core.fingerprint_seen_on_other_subjects` | -| `core@1`, resource velocity | `resource_velocity_1h`→`core.resource_velocity_1h`, `resource_total`→`core.resource_total`, `key_velocity_1h`→`core.credential_velocity_1h`, `key_total`→`core.credential_total` (a declared kind with `role: credential`, §5.4) | -| `core@1`, linked subjects and neighbour evidence | `linked_deleted_n`→`core.linked_deleted_n`, `neighbors_truncated`→`core.neighbors_truncated` | -| `core@1`, labels | `linked_labelled_abusive_n`→`core.linked_labelled_abusive_n` | -| `core@1`, behaviour change | `burst_ratio_24h_vs_lifetime`→`core.burst_ratio_24h_vs_lifetime`. Its activity set is `resource.created` ∪ delivery views, which is exactly today's `resource.created` ∪ `content.sent`. The shared `burstFactor`/`ageDecayFactor` functions are also exposed to the DSL as `relative_to_history` (§5.5). | -| `email@1`, requires `core` | `sends_10m_max`, `sends_1h`, `sends_first_day`, `webmail_recipient_share`, `webmail_sends_1h`, `distinct_recipients_1h`, `first_day_distinct_domains`, `self_send_before_external`, `subject_brand_match` (the last also requires `brand`). All become `email.`. They read delivery views whose `channel` is in the pack's `channels` param (default `[email]`). | -| `brand@1` | `name_brand_match`→`brand.name_match`, `name_has_at`→`brand.name_has_at`; new `brand.title_match` (§5.1.1). Data: `brands.yaml` plus the tenant's optional private `extra` list (S2b `brands_extra`). | - -The `core.credential_*` rename is the only name that moves beyond adding a prefix. "Key" is an -e2a word; "credential" is the product-neutral role (§5.4). The legacy alias keeps e2a's hash and -weight keys unchanged. - -New core features are **additive**. They are registered but not in e2a's rule, so e2a's golden -replay can't move: -- `core.email_domain_class_disposable`: main §8 Q7 deferred this. It reads the existing - `subject.created.email_domain_class` and is neutral (sign-up identity quality, not mail sending). -- `core.verdict_max_24h`: the maximum `content.verdict.score` in the trailing 24 h. -- `core.history_truncated`: see §5.7. - -#### 5.1.1 The brand pack - -The matcher is S2b's `BrandSet`, moved unchanged: NFKC and confusables skeleton, word and token -boundaries, Unicode punctuation tokenizing, case-sensitive entries, the integration and -community gates, and Cf stripping. It applies to any product with user-chosen display names: -workspace names, storefront names, profile names, community names. - -- `brand.name_match` and `brand.name_has_at` read `resource.created` and `resource.deleted` - `name` (and `name_skeleton`) across **all** declared kinds, as they do today. -- `brand.title_match` is the neutral sibling of `email.subject_brand_match`. It counts distinct - brands in non-self delivery titles over a trailing 1 h, capped at 3, and has none of the - email-specific exemptions. Non-email tenants use it. e2a keeps `email.subject_brand_match`, - whose S2b integration-name exemption and double-count rule are specific to mail. -- The brand list is pack data: `config/packs/brand/brands.yaml` (public), plus a per-tenant - private `brand.extra` path merged at boot with `MergeBrandSets`. - -### 5.2 Namespacing and the canonical key (how migration stays bit-for-bit) - -Every feature has two identities: -- **Name**: the namespaced name, used everywhere a human or config refers to the feature: rules, - weights files, corpus v2, reasons, docs. -- **Canonical key**: `Legacy` if the feature has one, else `Name`. It is used in exactly three - places: - 1. `core.inputHash`: the rule's feature map is re-keyed by canonical key before JSON encoding. - Quantization comes from each feature's `HashQuantum`, which replaces the name switch in - `quantizeAgeFeaturesForHash`. `core.subject_age_h` has quantum 1 and - `core.upgrade_delay_min` has quantum 60, the same values applied under the same keys. - 2. The local scorer's summation order: `sortedFeatures` is sorted by canonical key. - 3. The local scorer's `Version()`: it hashes the weights map re-keyed by canonical key. - -For a migrated feature, the canonical key is the flat name today's code uses. Summation order, -version hash, input hashes and therefore cassette keys are byte-identical, so no subject is -rescored at cutover and no vendor call is repeated. A feature born namespaced (every new pack -feature and every `custom.*`) has canonical key = name, so nothing legacy leaks into new work. - -**The alias table is frozen.** It is a Go literal of exactly the 25 flat names on `main` + #7, -and CI enforces it: the table can never gain an entry, and no new `Name` may equal any alias. - -**Where flat names are still accepted (read-side only):** - -| Artifact | Behaviour | -| --- | --- | -| Rules (`inputs`) | Resolved through the alias table, but only if the owning pack is enabled for the tenant. Otherwise the load fails with `feature_not_enabled`. A load warning says the name is deprecated. | -| Weights files | Same resolution. A file must not mix a flat name and its namespaced twin (`duplicate_feature`). | -| Corpus v1 rows | `eval.LoadSnapshotCorpus` maps `input.features` keys through the table. An unknown flat key fails the load. Export always writes corpus-v2 (§5.9). | -| Text inputs | Aliases `subject_line_skeleton`→`delivery.title_skeleton` and `first_link_host`→`delivery.link_host`. Their canonical keys stay the legacy names, for the same hashing reason. | -| Floors | Entries with no `profile:` default to `tenant:e2a` (§5.9). | -| Stored verdicts | Untouched. They store `input_hash`, risk and reason, and never feature names. | +**Enablement.** A profile lists `packs: [core@1, email@1, brand@1]`. The rules: +- `core` is always enabled; every other pack is opt-in. +- A pin names a major version, and a pack's semantics change only by shipping a new major version. +- `Requires` must hold, or the profile fails with `pack_requires`. +- A feature whose `RequiresPacks` or `SubjectKinds` don't match is not registered for that tenant + or kind. +- Packs run in the fixed order `core`, `email`, `brand`, `custom`. Each has its own namespace, so + two packs can't produce the same feature name. -### 5.3 Name grammar and reserved namespaces +**Pack contents after the one-time rename:** -- Feature names: `^[a-z][a-z0-9]*\.[a-z][a-z0-9_]{0,55}$`, at most 64 bytes. -- Reserved feature namespaces: `core`, `email`, `brand`, `custom`, plus any future pack name. - Tenants may define only `custom.*`. The `custom.` namespace is per tenant, so two tenants may - each have a `custom.invites_1h` that means different things (§12 Q11). -- Event type names keep the wire grammar `^[a-z_.]+$`, at most 64 bytes, and must contain a dot. - Reserved type prefixes (built-ins, tenants may not declare types under them): `subject.`, - `payment.`, `subscription.`, `resource.`, `content.`, `delivery.`, `abusekit.`. +| Pack | Features (old flat name → new name) | +| --- | --- | +| `core@1`, account | `subject_age_h`→`core.subject_age_h` | +| `core@1`, payment | `upgrade_delay_min`→`core.upgrade_delay_min`, `upgraded`→`core.upgraded`, `declines_before_first_success`→`core.declines_before_first_success`, `first_funding_prepaid`→`core.first_funding_prepaid`, `fingerprint_seen_on_other_subjects`→`core.fingerprint_seen_on_other_subjects` | +| `core@1`, resource velocity | `resource_velocity_1h`→`core.resource_velocity_1h`, `resource_total`→`core.resource_total`, `key_velocity_1h`→`core.credential_velocity_1h`, `key_total`→`core.credential_total` | +| `core@1`, linked subjects | `linked_deleted_n`→`core.linked_deleted_n`, `neighbors_truncated`→`core.neighbors_truncated` | +| `core@1`, labels | `linked_labelled_abusive_n`→`core.linked_labelled_abusive_n` | +| `core@1`, behaviour change | `burst_ratio_24h_vs_lifetime`→`core.burst_ratio_24h_vs_lifetime` (counts `resource.created` + `content.sent` + declared `activity` types, §5.4); new `core.history_truncated` (§5.7) | +| `email@1`, requires `core` | `sends_10m_max`, `sends_1h`, `sends_first_day`, `webmail_recipient_share`, `webmail_sends_1h`, `distinct_recipients_1h`, `first_day_distinct_domains`, `self_send_before_external`, `subject_brand_match` (also requires `brand`), each becoming `email.`. The pack owns the built-in `content.sent` type. | +| `brand@1` | `name_brand_match`→`brand.name_match`, `name_has_at`→`brand.name_has_at`; new `brand.title_match` over any declared field with `role: title` (§5.4) | -### 5.4 Neutral event vocabulary +In the early slices every built-in feature is available to every tenant (P2, §9). Pack **gating** +comes later, in P5. -#### Built-in types after this change +New `core` features are additive and are not in e2a's rule: +- `core.email_domain_class_disposable` (deferred in main §8 Q7); +- `core.verdict_max_24h`; +- `core.history_truncated`. -The wire contract (`POST /v1/events`, main §4.3) is unchanged in shape. Every change below is -additive: a new built-in type, new optional fields, and tenant-declared types. +### 5.2 The one-time rename (replaces revision 1's canonical-key bridge) -| Type | Change | -| --- | --- | -| `subject.created`, `subject.deleted`, `subject.class`, `payment.attempt`, `subscription.changed`, `content.verdict` | none (keeps #7's optional `account_created_at`) | -| `resource.created` / `resource.deleted` | `kind` is interpreted through the tenant's declared **resource kinds** (below). On the wire it stays free text. | -| `content.sent` | **Legacy email delivery, accepted forever, with #7's redaction unchanged.** Projected to a delivery view (below). | -| `delivery.sent` (new) | Neutral delivery: N recipients at destination D, with optional title text. | +**Decision.** Rename in one slice (P1) and re-baseline the golden replay once. No alias table +survives past that slice. The bridge existed only to keep stored verdict hashes and vendor +cassettes stable. Under A1 there are none, and a bridge would keep two names for everything alive +forever. -`delivery.sent.data`: +**Every consumer the rename touches, with its test:** -| Field | Kind | Rule | +| Consumer | Change | Test | | --- | --- | --- | -| `channel` | enum, declared per tenant | Optional. A tenant with exactly one declared channel may omit it; otherwise it is required. An undeclared value → `redaction_failed`. | -| `destination` | the kind the channel declares: `domain` \| `hash` \| `enum` | Optional. Where the delivery lands: a recipient domain, a keyed community id, a region code. | -| `destination_class` | enum, declared per channel | Optional, for example `own_community` \| `other_community`. | -| `recipient_count` | positive integer | Optional, default 1 for every feature that sums it. | -| `recipient_hash` | hash (`^[A-Za-z0-9_:+/=-]{8,128}$`) | Optional. Exactly one recipient, so paired with `recipient_count > 1` → `redaction_failed` (#7's rule). | -| `to_self` | bool | Optional. The recipient is the subject's own identity. | -| `title` | text ≤ 200 | Optional. NFKC, email-shaped substrings masked, skeleton stored as `title_skeleton`. | -| `link_host` | domain | Optional. First link host in the delivered content. | +| `core.stageSkip`, which reads the literal `features["subject_age_h"]` | Reads `core.subject_age_h` instead. The `Rule.Stage` **keys** (`max_subject_age_h`, `min_subject_age_h`) are stage names, not feature names, and keep their names. | Table test: a staged rule skips at the age bounds with a namespaced vector. A grep test fails if any `core` source still contains a flat feature literal. | +| `core.quantizeAgeFeaturesForHash`, which switches on names | Deleted. Quantization comes from `FeatureDef.HashQuantum` (1 h for `core.subject_age_h`, 60 min for `core.upgrade_delay_min`), passed to `Plan` in `core.Vector`. | The existing "hash stable under sub-hour drift" tests, re-run with namespaced names. | +| `core.inputHash` | Changes once (the names inside it change). | Golden records both hashes. A test asserts every rule's hash changes in P1 and is stable in every later slice. | +| Eval cassette key (`hashScoreRequest` over `req.Features`) | Changes once. P1 re-records every committed cassette (fake and local only, per A1) and stamps `feature_key_space: ns-v1` in the cassette header. | Loading a cassette with a mismatched `feature_key_space` fails with a clear error rather than silently missing. | +| `run.json`: `dataset_sha`, `rule_sha`, `weights_sha` | `rule_sha` and `weights_sha` change once. `dataset_sha` changes only for corpus-v1 snapshot files, which P1 rewrites to v2. The manifest gains `feature_key_space`. | Golden compares metrics, not SHAs. A test checks `feature_key_space: ns-v1` is in the manifest. | +| `worker.renderReason` and stored verdict reasons | Prints namespaced names. Verdict rows gain `reason_version: 2`. Old stored reasons are history and are not rewritten. | A reason snapshot test, plus a test that the verdict row carries `reason_version`. | +| `corpus_examples.features` | New column `feature_key_space text NOT NULL DEFAULT 'flat-v0'`. P1's migration rewrites the JSON keys of existing rows (none in production, per A1; this touches local dev databases only) and sets `ns-v1`. Once migrated, the corpus loader refuses `flat-v0`. | Migration test on a seeded DB: keys rewritten and marker set. A second test checks a `flat-v0` row is refused. | +| `local.Version()`, which hashes the whole `Weights` struct | Changes once. The weights file's `version:` is bumped to `v2`. | `Version()` differs from the pre-rename golden value and is stable afterwards. | +| `feature.Names`, `Features.Map`, `config.FeatureSet`, `config/rules.yaml`, `config/local_weights.yaml`, `mutation_test.go`, `ablation_test.go`, golden-sign tables, `eval/fixtures/README.md` | Mechanical rename. | The full existing suite passes with renamed literals, and the grep test. | +| `abusekit score --jsonl` input rows (an external contract) | A row keyed by an old flat name is rejected with `feature_renamed`, and the error names the new name. It fails loudly and never translates silently. | CLI contract test. | + +**The floating-point bound.** The local scorer sums weight × feature in sorted feature-name order. +Renaming the features changes the sort order, and with it the order of the additions. For `n` +terms, recursive summation error is at most `(n−1)·u·Σ|wᵢxᵢ|` with `u = 2⁻⁵³`. For e2a's rule: +- `n = 25` and `Σ|wᵢxᵢ| < 100` on any bounded vector, so `|Δlinear| < 2.7e-13`; +- the sigmoid's slope is at most ¼, so `|Δrisk| < 6.7e-14`. + +The golden asserts `|Δrisk| ≤ 1e-12`, which leaves margin. So that tiers are provably unchanged, it +also asserts that no recorded score lies within `1e-12` of a tier cut point or a rule threshold. If +one ever does, the fixture is flagged rather than passing silently. -#### The delivery view (read-side projection) +### 5.3 Name grammar and reserved namespaces -Features never read `content.sent` or `delivery.sent` directly. They read `event.View`, which -`vocab.View` produces from the stored row. For `content.sent`: +- **Features** match `^[a-z][a-z0-9]*\.[a-z][a-z0-9_]{0,55}$`. The namespaces `core`, `email`, + `brand` and `custom` are reserved, as is any future pack name. Tenants define only `custom.*`, + which is scoped to the tenant. +- **Event types** keep `^[a-z_.]+$`, are at most 64 bytes, and must contain a dot. The prefixes + `subject.`, `payment.`, `subscription.`, `resource.`, `content.` and `abusekit.` are reserved. +- **Extension fields** on built-in types must be named `x_`, so they can never collide with + a future built-in field. -| Delivery view field | Taken from `content.sent` | -| --- | --- | -| `channel` | constant `email` | -| `destination` (kind `domain`) | `recipient_domain` | -| `recipient_count` | `recipient_count` | -| `recipient_hash` | `recipient_hash` | -| `to_self` | `recipient_is_own_identity` | -| `title` / `title_skeleton` | `subject_line` / `subject_line_skeleton` | -| `link_host` | `first_link_host` | - -A `delivery.sent` row maps to the same view field for field. The email pack's features are the -S2b functions with one mechanical edit: `e.Type != "content.sent"` becomes -`v.Kind != event.ViewDelivery || !channels[v.Channel]`, and the field reads are renamed. -Everything else is untouched, including missing-field defaults, `recipientCountOf`'s default of -1, `normalizeToken` domain folding and `isSelfSend`. So an e2a stream of `content.sent` and the -same stream rewritten as `delivery.sent{channel: email}` yield identical features. §8's -translation test proves this on every fixture. - -**Why read-side and not a rewrite at ingest.** Rewriting would change the stored type and the -`BodyHash` that separates `duplicate` from `conflict`. It would also need a data migration for -stored rows and would still leave old rows to interpret. The read-side view needs no migration -and makes both forms equivalent by construction. - -#### Product-declared resource kinds and channels +### 5.4 Product-declared vocabulary + +There is no new built-in delivery type. The wire contract (`POST /v1/events`) keeps its shape; +every change below is additive and optional. ```yaml vocabulary: - version: 3 # monotonically increasing; see §5.6 for compatibility rules + version: 4 # must increase with any change; history in §5.6 + subject_kinds: # default [account] + account: {} + api_key: {parent: account} # a key's events also mark its parent account dirty + card: {} resource_kinds: - agent: {role: identity} - key: {role: credential, aliases: [keys, api_key, api_keys, api-key, apikey, "api key"]} - channels: - email: {destination: domain} + key: {role: credential, aliases: [keys, api_key, api_keys, api-key, apikey, "api key"]} + agent: {role: other} + link_kinds: # in addition to the six built-in kinds + phone_hash: {evidence: true} # counts as neighbour evidence + oauth_sub_hash: {evidence: true} + types: + invite.sent: + role: activity # included by core velocity/burst features + fields: + invitee_hash: {kind: hash, join_domain: member} + target_class: {kind: enum, values: [own_community, other_community]} + preview: {kind: text, role: title, max_len: 200} # stored as skeleton by default + link_host: {kind: domain} + extend: # extension fields on built-in types + resource.created: + x_visibility: {kind: enum, values: [public, private]} ``` -- `role` is a closed enum: `credential`, `identity`, `workspace`, `content`, `other`. - `core.credential_*` counts kinds with `role: credential`; `core.resource_*` counts every kind. - Matching folds case and trims whitespace, then applies `aliases`. This is the S2b behaviour - moved into data: e2a's declaration above reproduces `resourceKindAliases` exactly, and the - golden replay proves it. -- An **undeclared kind** is stored, counts in `core.resource_*` (role `other`), and increments - `abusekit_undeclared_kind_total{tenant}`. That matches today's behaviour for kinds that aren't - keys. With `vocabulary.strict: true`, an undeclared kind is instead rejected with the existing - `redaction_failed` code, so no new per-item code is needed. -- A tenant with no `resource_kinds` block gets the **implicit legacy declaration** above. That is - how today's global config keeps working (§8). - -#### Versioning - -The wire format stays `/v1`. Every addition is optional, and producers already handle unknown -per-item codes because the code list is closed and unchanged. The one semantic change is to -how **undeclared** types are stored (§5.6). It is flagged, and it is licensed because the -service is pre-GA: main §4.3 documents unknown types as "stored", and nothing reads them yet. -`RedactionSchemaVersion` for built-ins stays at 2 (#7). Each stored row gains -`vocab_version text` (`"@"`, NULL for built-in-only schemas) next to -`redaction_version`, in an expand-only migration. +**Roles.** Only these ship: +- resource kinds: `credential` and `other` (the review's minimal set; an undeclared kind is + `other`); +- types: `activity`; +- fields: `title`, which `brand.title_match` reads, and `self`, a bool marking a self-directed + event that such features exclude. + +`brand.title_match` counts distinct brands across the `title` fields of non-self activity events in +a trailing 1 h, capped at 3. A `title` field stored skeleton-only is matched through +`BrandSet.MatchesSkeleton`, which compares already-folded tokens to the brands' skeletons and +skips the second fold. (#7 found that double folding is not idempotent for leetspeak.) + +**Subject kinds.** Events gain two optional fields: +- `subject_kind` (default `account`); +- `also: [{kind, id}]`, with at most 3 entries. + +`also` lets one event (a charge attempt, say) be indexed under the merchant account, the card and +the customer at once: +- the event is stored once, and idempotency is still keyed on `(tenant, producer, id)`; +- one index row is written per subject; +- each named subject, and the declared `parent` of the primary subject, is marked dirty. + +Subjects are keyed `(tenant, kind, id)`, and the API follows: +- `GET /v1/subjects/{subject}` and `POST .../evaluate` take an optional `?kind=` (default + `account`); +- the list endpoint (S3b) gains a `kind` filter; +- rules declare `applies_to: [kinds]` (default `[account]`); +- a feature is computed only for the kinds in its `SubjectKinds`. For example, the `core` + onboarding features exist only for `account`. + +**Link kinds.** The `links` object gains `custom: {"": ""}`, with at most 8 +entries. Each value is re-HMACed like every hash (§5.6). Built-in and declared evidence kinds feed +both `Neighbors` and the new `neighbours` op (§5.5). + +**Legacy behaviour by declaration.** A profile with no `resource_kinds` gets the implicit +declaration `key: credential`, with #7's aliases. The golden replay proves this reproduces +`resourceKindAliases`. ### 5.5 Declarative custom features @@ -392,820 +359,979 @@ service is pre-GA: main §4.3 documents unknown types as "stored", and nothing r ```yaml features: - - name: custom.public_links_1h # custom.* only - version: 1 # bump on ANY change to the definition (enforced, §5.7) - description: public share links created in the last hour # required, shown in reasons - count: # exactly one op key: count | distinct | share | peak | time_between - type: share.link_created # a declared type, or a built-in type/view + - name: custom.public_links_1h # custom.* only + version: 1 # bump on any change (checked against history, §5.6) + description: public share links created in the last hour + subject_kinds: [account] # default [account] + count: # exactly one op key + type: share.link_created where: {field: visibility, eq: public} - sum: {field: size_class_weight, default: 1, cap_each: 10} # optional; count = sum of this - window: 1h - transform: {log1p: true, cap: 50} # cap mandatory; log1p optional; applied log1p → cap + sum: {field: size_class_weight, default: 1, cap_each: 10} # optional + window: 1h # or first: | lifetime | before_first: {type, where} + transform: {log1p: true, cap: 50} # cap mandatory + prior_sign: "+" # optional, for uniform priors ``` -**Windows.** Either `window: ` (trailing, `(now − dur, now]`) or `first: ` (anchored, -`[start, start + dur)`). `` is a whole number of minutes, hours or days: `1m`–`30d`. -`lifetime` means `(−∞, now]`. `start` is the account start defined in §5.1 `Input.Start`. - -**Predicates (`where`).** A closed set of operators. There is no regex, no arithmetic and no -user functions. +**Windows.** -| Leaf | Meaning | Allowed field kinds | -| --- | --- | --- | -| `{field: f, eq: v}` / `{field: f, ne: v}` | equality after the kind's normalisation | enum, bool, number, domain, hash | -| `{field: f, in: [..]}` / `not_in` | membership, ≤ 256 values; for enums every value is checked against the declared enum at load, so a typo is a load error | enum, number, domain, hash | -| `{field: f, in_set: }` | membership in a tenant- or pack-provided set file (for example, `email`'s webmail list), hashed at load; ≤ 100,000 entries | domain, hash, enum | -| `{field: f, suffix_in_set: }` | domain-label-boundary suffix match (`a.b.example.test` ⊂ `example.test`), linear time | domain | -| `{field: f, gte: n}` / `lte` / `gt` / `lt` | numeric compare | number | -| `{field: f, exists: bool}` | presence | any | -| `{all: [..]}` / `{any: [..]}` / `{not: leaf}` | combinators, depth ≤ 2, ≤ 8 leaves in total | | - -An absent field or a type mismatch makes a leaf false; `{exists: false}` is the only leaf that -is true on absence. `text` fields can't appear in `where` or as a `distinct` field. Text is only -available to text-accepting scorers through a rule's `text:` list. That keeps free text out of -every aggregation. - -**Operations.** - -| Op | Value at `now` | +| Window | Range | | --- | --- | -| `count` | The number of matching events in the window, or, with `sum`, the sum of the field over them. `default` covers absence, and `cap_each` clamps each event's addend before summing, as S2b does with `recipient_count`. | -| `distinct` | `{type, where, field}`: the number of distinct normalised values of `field` among matching events in the window. `field` must be an `enum`, `number`, `domain` or `hash` field. Tracking stops at `track_max`, the smallest count whose transformed value reaches `transform.cap` (`cap` itself without `log1p`, `ceil(expm1(cap / scale))` with it). The loader rejects a definition whose `track_max` exceeds 10,000, so memory is O(track_max). | -| `share` | `{type, where (denominator), match (numerator predicate), sum?}`: numerator ÷ denominator over the window; `if_empty` (default 0) when the denominator is 0. | -| `peak` | `{type, where, sum?, size: }`: the maximum count or sum in any window `(t − size, t]` with `t` an event instant inside the outer window. `size` ≤ window and window ÷ size ≤ 1440. A two-pointer scan over the time-sorted matches. | -| `time_between` | `{from: {type, where}, to: {type, where}, until_now: bool, if_absent: n}`: minutes from `t_A` (the earliest matching `from`) to `t_B` (the earliest matching `to` with `t_B ≥ t_A`). No `from` event → `if_absent`. A `from` but no `to` → minutes since `t_A` when `until_now`, else `if_absent`. `if_absent` is mandatory. Negative results can't occur, and the value is floored at 0 as a guard. | +| `window: ` | `(now − dur, now]`, where `dur` is 1 minute to 30 days, in whole minutes, hours or days | +| `first: ` | `[start, start + dur)` | +| `lifetime` | `(now − 90d, now]`: the retention horizon. The docs call this "retained lifetime" so nobody reads it as all-time. | +| `before_first: {type, where}` | `(now − 90d, t_B]`, where `t_B` is the first matching event at or before `now`. If there is no such event, it falls back to `(now − 90d, now]`. The end is inclusive, matching #7's "at or before". | -**The history-relative modifier.** It may be added to `count`, `distinct` and `peak`, and it -generalises S2b's `burstFactor × ageDecayFactor`: +Events with `at > now` are excluded from every window. See "Future events" under Semantics. -```yaml - relative_to_history: - lookback: 30d # ≤ 30d - exclude_recent: 24h # the current burst never serves as its own baseline - age_decay: {full_until: 3d, zero_at: 30d, floor: 0.2} # optional; these are the defaults -``` +**Predicates.** A closed set, with no regex: +- equality and membership: `eq`, `ne`, `in` / `not_in` (at most 256 values; enum values are checked + against the declaration); +- set files: `in_set` / `suffix_in_set` (named files of at most 100,000 entries, hashed at load); +- numeric comparison: `gt`, `gte`, `lt`, `lte`; +- presence: `exists`; +- combinators: `all`, `any`, `not`, with nesting depth at most 2 and at most 8 leaves. -`baseline` is the maximum of the same op, with the same width and predicate, over sliding -windows whose end lies in `(now − lookback, now − exclude_recent]`. The value is then -`min(current / max(baseline, 1), transform.cap)`. If `age_decay` is set, the result is -multiplied by `clamp(1 − (age_days − full_until)/(zero_at − full_until), floor, 1)`. With the -defaults, `full_until = 3d` and `zero_at = 30d` make the denominator 27, which is exactly S2b's -`ageDecayFactor`. `floor > 0` is mandatory, so a decayed signal is never a hard zero. - -**Transform.** `cap` is mandatory for every feature: a finite value > 0, and at most the op's -natural bound (1 for `share`). `log1p` is optional and takes one of three forms, applied before -the cap: `log1p: true` gives `ln(1 + v)`; `log1p: {scale: s}` gives `s · ln(1 + v)`; and -`log1p: {anchored_at: n}` sets `s = n / ln(1 + n)`, so `v = n` maps to `n`. The last form -reproduces `first_day_distinct_domains`'s S2b shape. The transformed value always lies in `[0, cap]`, and -that is the feature's `Bound`. - -#### Evaluation semantics - -- **Pure and order-independent.** A feature is a function of the set of stored events and - `now`, never of arrival order. Where "first" is ambiguous, ties at the same instant break by - `(at, event id)`. Late events bump `dirty_seq` as they do today. -- **Half-open windows.** Trailing windows are `(now − W, now]`; anchored windows are - `[start, start + W)`; `peak` sub-windows are `(t − S, t]`. -- **Future events are excluded everywhere,** `lifetime` included: an event with `at > now` - contributes to no custom feature until `now` reaches it. It does contribute a rescore - candidate at its own `at` (the S2b B4 behaviour, generalised). - *Frozen legacy exception:* the migrated Go features `core.resource_total` and - `core.credential_total` count future-dated events in their lifetime totals (the S2 behaviour), - and `email.first_day_distinct_domains` uses an inclusive `[start, start + 24h]`. Both are kept - for bit-for-bit parity and documented in each feature's docstring. Harmonising them is a - `core@2`/`email@2` change (§12 Q5). -- **Rescore candidates.** For each windowed feature, the compiler emits: the exit instant of the - oldest in-window match (`at + W`); anchored window ends still in the future; for `peak`, the - exit of the current maximum's sub-window; for `relative_to_history`, the instants where an - event crosses `now − exclude_recent` or `now − lookback`; and every future-dated match. The - orchestrator takes the minimum and coalesces it to the 5-minute bucket. This generalises - `isWindowedEventType`: the set of windowed types is the union of every feature's `Reads`. -- **Determinism of floats.** Sums accumulate in time order `(at, id)` with a single `float64` - accumulator, so the same event set always gives the same bits. - -#### Compilation and cost model - -`pack/custom.Compile` validates each definition and produces a **per-tenant plan**: an index -from view type to the list of (feature, predicate program) pairs that read it. Extraction makes -one pass over the subject's events, dispatches each event to the features for its type, and -evaluates the flat predicate programs. `peak` and `relative_to_history` keep per-feature -matched-instant slices, and a final pass per feature runs the two-pointer scans. - -Static cost units, checked at load: - -| Op | Units | -| --- | --- | -| `count`, `share`, `time_between` | 1 | -| `distinct`, `peak` | 2 | -| `relative_to_history` | ×2 on top of the op's own cost | -| Each predicate leaf beyond the first | +0.25 | +A leaf is false when its field is absent or has the wrong type. `text` fields are banned from +predicates, `distinct`, `group_by` and `on`. -Per-tenant budget: **256 units**. With the 50,000-event scan bound (§5.7), the worst case is -about 50,000 events × 8 dispatched features per type × 8 leaves ≈ 3.2M predicate steps, plus -O(n) scans. That is tens of milliseconds, which is what criterion 4 measures. At runtime a -per-subject extraction deadline (default 250 ms) applies to the custom pack. If it expires, the -pack fails as described in §6. +**Operations.** `cap(v)` below is the transform. -### 5.6 Redaction for declared types and fields - -Ingest never consults rules (main §4.3). It consults the tenant **vocabulary**, which is -config, reviewed like code, and compiled into `vocab.Vocabulary`. The built-in `schema` map -becomes the built-in half of every vocabulary, unchanged. +| Op | Definition at `now` | Bound / cost class | +| --- | --- | --- | +| `count` | The number of matching events in the window, or the sum of `sum.field` over them. An absent field counts as `default`, and each value is clamped to `cap_each`. | O(n) | +| `distinct` | `{field}`: the number of distinct values of `field` among matches. Tracking stops at `track_max`, the smallest count whose transformed value reaches `cap`; `track_max` must be ≤ 10,000. | O(n); memory O(track_max) | +| `share` | `{where, match, sum?}`: the numerator over the `match` events divided by the denominator over the `where` events. When the denominator is 0, the value is `if_empty` (default 0). | O(n) | +| `peak` | `{size}`: the maximum of `agg(E ∩ (t − size, t])`, taken over `t` = the instants of matching events inside the outer window. Sub-windows are **clipped** to the outer window, so events outside it never count even if they fall inside `(t − size, t]`. Requires `size` ≤ window and window ÷ size ≤ 1440. | O(n), two-pointer | +| `time_between` | `{from: {type, where, anchor: first\|last}, to: {type, where}, until_now, if_absent}`. `t_A` is the first (or last) matching `from`; `t_B` is the first matching `to` with `t_B ≥ t_A`. The value is minutes from `t_A` to `t_B`. With `until_now` and no `t_B`, it is `now − t_A`. `if_absent` is mandatory. Floored at 0. | O(n) | +| `sequence` | `{a: {type, where}, b: {type, where}, within: , on: {a: field, b: field}?}`: the number of `b` events in the window that have at least one `a` event with `t_a ∈ (t_b − within, t_b]` and, if `on` is set, `a.on == b.on`. Both `on` fields must be `hash` fields with the **same `join_domain`** (§5.6), or equality across them is meaningless; the loader checks this. The evaluator keeps, per `on` value, the latest `a` instant ≤ `t_b` in an LRU capped at 10,000 keys. Eviction sets the feature's truncated flag. Requires `within` ≤ 24 h. | O(n); memory O(keys) | +| `group_by` | A modifier on `count`, `distinct` or `sum`: `{field, reduce: max\|{count_gte: k}, max_groups ≤ 1000}`. Matches are grouped by `field` and the op is applied per group. `max` returns the largest group value; `count_gte: k` returns how many groups have a value ≥ k. Groups are created in `(at, producer, id)` order. Once `max_groups` is reached, new groups are ignored and the truncated flag is set. | O(n); memory O(groups) | +| `ratio` | `{num: custom.a, den: custom.b, if_empty}`: `num ÷ den` over two **non-ratio** custom features, forming a depth-1 DAG. The loader rejects cycles and ratio-of-ratio. | O(1): the inputs are computed anyway | +| `neighbours` | `{via: [link kinds], where: {deleted: permanent} \| {labelled: abusive} \| {created_within: } \| {}}`: the number of distinct other same-tenant, same-kind subjects that share any `via` key and satisfy the condition. Fan-in is capped at 50 per key and 200 in total, as in main §4.2, and hitting a cap sets `core.neighbors_truncated`. | One store query per distinct `via` set, cached per extraction. At most 4 `neighbours` features per tenant. | + +**`relative_to_history`.** A modifier on `count`, `distinct` and `peak`. It is the exact pipeline +of #7's `burstFactor × ageDecayFactor`: -```yaml -vocabulary: - version: 1 - types: - share.link_created: - fields: - visibility: {kind: enum, values: [public, org, private]} - file_kind: {kind: enum, values: [document, archive, executable, image, other]} - size_bytes: {kind: number, min: 0, integer: true} - folder_title: {kind: text, max_len: 120, skeleton: true} - share.downloaded: - fields: - link_hash: {kind: hash} - downloader_ip24: {kind: hash} - downloader_is_owner: {kind: bool} +``` +cur = op over the feature's window at now +E_b = { matching events e : now − lookback < e.at ≤ now − exclude_recent } +base = baseline_op over E_b (default: the same op; for count/distinct over window W it is + the max of that op over sliding windows (t − W, t], t ∈ instants of E_b, clipped to E_b; + for peak it is peak over E_b with the same size) +v1 = min( cur / max(base, 1), ratio_cap ) ratio_cap defaults to transform.cap +age_days = (now − start) / 24h +d = clamp( 1 − (age_days − full_until)/(zero_at − full_until), floor, 1 ) if age_decay +v2 = v1 · d (v1 if no age_decay) +value = transform(v2) log1p (optional), then cap ``` -**Field kinds and the rule for each:** - -| Kind | Accepts | On violation | +Parameters: +- `lookback` ≤ 30 d, `exclude_recent` < `lookback`, and `floor` > 0. +- `age_decay` defaults to `full_until: 3d, zero_at: 30d, floor: 0.2`, which are #7's constants. +- `baseline:` may override the baseline op, e.g. `baseline: {peak: {size: 10m}}`. That lets a 1 h + sum be compared against a prior 10-minute peak, which is how #7's `sends_1h` works. + +**Which #7 and S2 features the DSL can express:** +- **Expressible:** + - `sends_10m_max`, `sends_1h`, `webmail_sends_1h`: `relative_to_history` with a `baseline` + override, `sum: recipient_count`, `cap_each: 300`, `ratio_cap: 300`; + - `sends_first_day`; + - `webmail_recipient_share`: `share` + `in_set: webmail`; + - `declines_before_first_success`, and `self_send_before_external` with `cap: 2`: via + `before_first`; + - `resource_*`, `credential_*`; + - `upgrade_delay_min`: `time_between` + `cap`. +- **Not expressible:** + - `distinct_recipients_1h`, a mixed aggregate: distinct hashes, falling back to summing + `recipient_count` for events that have no hash; + - `subject_brand_match` and `brand.*`, which need the brand matcher and its integration and + community gates; + - `first_day_distinct_domains`, which has an inclusive end where DSL `first:` windows are + half-open; + - `burst_ratio_24h_vs_lifetime`, which counts future-dated events in its lifetime denominator; + - `linked_*` and `fingerprint_*`, which carry specific deleted-and-labelled evidence semantics; + the `neighbours` op covers the generic cases; + - `subject_age_h`, which reads `start` rather than events. + +The non-expressible features stay Go features in their packs. P4b adds a test that re-expresses +every "expressible" feature in the DSL and checks it matches the Go value bit for bit on every +fixture. + +**Transform.** `cap` is mandatory: a finite value > 0, and at most 1 for `share`. `log1p` is +optional and is applied before the cap: +- `log1p: true` gives `ln(1+v)`; +- `log1p: {scale: s}` gives `s·ln(1+v)`; +- `log1p: {anchored_at: n}` sets `s = n/ln(1+n)`. + +The output always lies in `[0, cap]`, and that range is the feature's `Bound`. + +#### Semantics + +- **Pure and order-independent.** A feature is a function of the loaded event set and `now`. + Ordering, ties and "first" all use `(at, producer, id)`. P0 switches the store's scoring loader + from `(at, seq)` to this order, and the golden captures it before the rename. Any fixture whose + ties now resolve differently is listed as a baseline change: a tie resolved by arrival order was + a latent non-determinism. +- **Half-open windows.** `(now − W, now]`, `[start, start + W)`, and clipped `peak` sub-windows. + `before_first` is the one intentionally inclusive end. +- **Future events.** Events with `at > now` are excluded from every custom window, and schedule a + rescore at their `at`. Some Go features keep frozen legacy behaviour for semantic identity, + documented on each `FeatureDef`: + - `core.resource_total`, `core.credential_total` and `core.burst_ratio_24h_vs_lifetime` count + future-dated events; + - `email.first_day_distinct_domains` has an inclusive end. + + Harmonising them is `core@2`/`email@2` work (§12 Q5). +- **Deterministic floats.** Sums accumulate in `(at, producer, id)` order in one `float64`. +- **Rescore candidates.** Each DSL feature emits: + - the exit time of its oldest in-window match; + - the end of any anchored window; + - for `peak`, the exit of the current maximum's sub-window; + - for `relative_to_history`, the crossings of `now − exclude_recent` and `now − lookback`; + - for `before_first`, the first `to` event; + - future-dated matches. + + Rescore-storm control (§5.8) then filters and coalesces them. + +### 5.6 Redaction, pseudonymisation and vocabulary history + +Ingest never consults rules. It consults the tenant's **vocabulary**. What abusekit stores is +**pseudonymised**, not anonymous: anyone holding both the key and a candidate value can recompute +a keyed hash. Retention and erasure (main §4.4, §4.11) therefore apply to it. + +**Field kinds:** + +| Kind | Stored as | Rejected when | | --- | --- | --- | -| `text` | string; NFKC; control characters and invalid UTF-8 rejected; **email-shaped substrings masked to `@`** (#7's `subject_line` rule, generalised); truncated at `max_len` (≤ 500, default 200); optional `skeleton` companion | reject for control characters or bad UTF-8; mask for an email shape | -| `number` | finite float64; optional `min`, `max`, `integer` | reject | -| `bool` | bool | reject | -| `enum` | string in `values` (≤ 64 values, each ≤ 64 bytes, `[a-z0-9_.-]+`) | reject | -| `hash` | `^[A-Za-z0-9_:+/=-]{8,128}$` (no `@`, `%` or whitespace) | reject (never truncated) | -| `domain` | lower-cased, IDNA to ASCII, hostname grammar (labels 1–63, total ≤ 253); a user part is impossible by grammar | reject | -| `timestamp` | RFC 3339 (#7's `account_created_at` pattern) | reject | - -**Rules that make it privacy by construction:** -1. The existing recursive leak scan runs first, over every key and value of every event. - Email-shaped content outside a `text` field is still rejected. Declared `text` fields are the - only exemption, and they are masked instead of rejected. -2. **Undeclared fields of a declared type are hashed, not stored.** A string becomes - `hk1:` + hex(HMAC-SHA256(tenant redaction key, type ‖ field ‖ value))[:32]. Numbers and bools - pass through. Objects and arrays are dropped. A hashed field can be used only in - `distinct`, `eq`/`in` on a hash (the producer would have to compute the same HMAC, which it - can't, so in practice only `distinct` and `exists`). It can never be read back. -3. **Undeclared types** get the same treatment, with every top-level key treated as undeclared. - This replaces "kept as-is" (§5.4 versioning note). Undeclared types remain unusable by - features until they are declared. -4. **The redaction key** is a per-tenant secret held by abusekit (Secret Manager, - `abusekit--redaction-key`). It is separate from the producer-held link-hash key. - Rotating it breaks equality across the rotation boundary for hashed fields, which is - acceptable because they only feed windowed `distinct`. The key id is recorded in - `vocab_version`. -5. Caps: ≤ 32 declared types per tenant, ≤ 32 fields per type, `data` ≤ 8 KiB after redaction - (unchanged). - -**Vocabulary compatibility.** Stored rows are immutable, and features must be able to read old -rows. So a vocabulary may only **widen**: add a type, add a field, add enum values, raise -`max_len` or `max`, or switch `strict` off. Changing a field's kind, removing or narrowing enum -values, or lowering a cap requires a new field name. The loader enforces this against the latest -accepted vocabulary for the tenant, recorded in a new `tenant_vocabularies(tenant, version, -sha, body, accepted_at)` table. The version must increase with any change, and an incompatible -change fails the load with `vocab_incompatible`. The harness reports how many rows were stored -under each `vocab_version`. - -### 5.7 Limits (validated at load; also the fuzz oracle) +| `text` | NFKC. Email, card, IP and phone shapes are **masked** (`@`, `#card`, `#ip`, `#phone`), and the value is truncated at `max_len` (≤ 500). **`store: skeleton` is the default for custom text**: only the confusables skeleton is kept. `store: raw` is opt-in, for text-accepting scorers only. The built-in `content.sent.subject_line` stays raw (main §4.3). | Control characters, invalid UTF-8 | +| `number` | A finite float64, checked against the optional `min`, `max` and `integer`. Integers of 13–19 digits also get the Luhn check. | Out of range, non-finite, or Luhn-valid | +| `bool` | As-is | Not a bool | +| `enum` | One of the declared values (at most 64, each matching `[a-z0-9_.-]{1,64}`) | Undeclared value | +| `hash` | **Re-HMACed at ingest:** `hk:` + hex(HMAC-SHA256(k_tenant, input))[:32]. `input` is `lp(type) ‖ lp(field) ‖ lp(value)`, or `lp("join:" ‖ join_domain) ‖ lp(value)` when the field declares a `join_domain`; `lp` is a u32 length prefix. A `join_domain` makes the same identifier equal across fields and types (for `sequence.on` and cross-type `distinct`); without one, hashes are separated per field. The producer's value is never stored. | Doesn't match `^[A-Za-z0-9_:+/=-]{8,128}$` | +| `domain` | Lower-cased and IDNA-encoded to ASCII. It must end in a public suffix from the embedded, versioned PSL snapshot, or in an RFC 6761 special-use name (`.test`, `.example`, `.invalid`, `.localhost`) so synthetic fixtures stay valid. With the optional `reduce: etld1`, only the registrable domain is stored. | IP literal, all-numeric label, unknown suffix, `@` | +| `timestamp` | RFC 3339 (#7's pattern) | Anything else | + +**Rules, in order:** +1. **Leak scan.** It runs over every key and value of every event, and now detects: + - email addresses; + - Luhn-valid runs of 13–19 digits (separators allowed); + - IPv4 and IPv6 literals; + - phone shapes: `+` followed by 8–15 digits, or grouped national formats of 10 or more digits. + + What happens on a match depends on the field: + - a declared `text` field is masked; + - a declared `hash` field is exempt, because its value is replaced by the HMAC; + - anywhere else, the event is rejected with `redaction_failed`. +2. **Declared fields** are handled by kind, as in the table above. +3. **Undeclared fields** of any type, built-in or declared, are **dropped whatever their kind**, + numbers included. The row keeps `x_undeclared: [sorted field names]`; a name that fails the + leak scan is rejected. **Undeclared types** keep only `type`, `at` and that list of names. + Revision 1 hashed undeclared strings instead; that is removed. +4. **Built-in hash values are re-HMACed too:** `content.sent.recipient_hash` and every `links` + value, both the six built-in link kinds and declared ones. Every `links` kind has its own + `join_domain`, which is the kind's name. Equality is preserved, so `distinct` counts, + neighbour joins and every golden feature value are unchanged; only the stored bytes change. A + producer that mistakenly sends a raw phone number as a "hash" never has it stored. +5. **Built-in domain fields** (`recipient_domain`, `address_domain`, `first_link_host`) get the + `domain` rules, and `RedactionSchemaVersion` becomes 3. Every committed fixture uses `.test` + and passes. + +**Keys and rotation.** The key interface is provider-agnostic: + +```go +type Keys interface { + // Current returns the key used to write new values, and its id. + Current(ctx context.Context, tenant string, purpose Purpose) (id string, key []byte, err error) + // ReadSet returns every key readers must accept right now (current, plus the previous key + // during a rotation). + ReadSet(ctx context.Context, tenant string, purpose Purpose) ([]KeyRef, error) +} +``` + +There are two adapters: a file adapter for dev and tests, and a cloud secret-manager adapter. +Neither the names nor the config mention a provider. Rotation is **dual-key**: +- For `max_lookback` (at most 30 days), ingest writes each hash field twice: `` under the + new key and `__prev` under the old one. Links get one row per key id. +- Until the rotation's `read_flip_at`, features read the `__prev` values and neighbour joins match + either key id. +- After `read_flip_at`, readers use only the new values and `__prev` writes stop. +- The key id is part of `vocab_version` (`"@/k"`), so every row can be traced to + the key that hashed it. + +**Vocabulary history lives in the config tree.** The mount holds `tenants//history/NNNN.yaml`, +an append-only list of accepted vocabulary and custom-feature versions, each with an +`effective_at`. CI in the private config repo runs `abusekit config check` over the full history +and enforces three rules: +- **Widening only.** Adding a type, field, enum value, `max_len` or `max` is fine. Changing a + field's kind, removing an enum value or narrowing a limit needs a new field name. +- **No redefinition.** A custom feature's `(name, version)` is never redefined. +- **Monotonic time.** `effective_at` never goes backwards. + +At runtime, the `tenant_config_versions` table records what was actually loaded and refuses a +profile that contradicts it. It is a guard, never the source of truth. + +**Warm-up.** A feature is **cold** from its `effective_at` until +`effective_at + max(window, lookback + exclude_recent)`. That applies to a new custom feature, a +new version of one, and any feature over a newly declared field or type. An `advise` rule with a +cold input is scored and stored as shadow for that period: it reports `mode: shadow` with +`warming_until`. The value is still computed; it just can't drive a tier on partial history. + +### 5.7 Bounded loading, step budget, and truncation as a signal + +This replaces revision 1's bound of the 50,000 newest events and its 250 ms wall-clock deadline. + +**Load plan (per profile and subject kind, computed at compile time):** +1. **Onboarding types, in full.** `subject.*`, `payment.*` and `subscription.*` are loaded + oldest-first, up to `onboarding_bytes` (1 MiB decoded by default). Onboarding facts such as + "first success" come from the *earliest* events, so recent activity can never push them out. +2. **Anchored range.** `[start, start + A)` is loaded oldest-first, where `A` is the maximum + `first:` duration (24 h for e2a). +3. **Trailing range.** `(now − L, now + 24h]` is loaded **newest-first** until the total decoded + size reaches `history_bytes` (8 MiB by default). + - `L = max over features of max(window, lookback + exclude_recent + baseline width)`. + - `L` is 90 d if any feature uses `lifetime` or `before_first`. + - The `+24h` covers future-dated events inside the skew allowance, which schedule rescores. +4. **`start`** comes from `subject.created.account_created_at` if present, else from + `subjects.first_seen_at`, never from loaded events. Truncation therefore can't move it. + +Events are deduplicated across the three ranges by `(producer, id)`, and the combined history is +ordered `(at, producer, id)`. + +**Step budget, instead of wall-clock time.** A step is one event dispatched to one feature, plus +one step per predicate leaf. The per-subject budget is +`steps_max = history_bytes / avg_event_bytes × per_type_fanout_max × leaves_max`. With the defaults +that is 8 MiB / 256 B × 16 × 8 ≈ 4.2M. Every pack charges its steps through `Input.Budget`, the Go +packs included. + +The P4a benchmark calibrates the ns-per-step figure end to end (store load, decode, projection, +extraction, orchestration). The committed `cost_table.yaml` records it, and CI fails if a profile +at the limits exceeds p99 50 ms. Because the budget counts work, when it runs out doesn't depend +on how fast the host is. + +**Exhaustion and truncation.** Exhausting the byte cap, the step budget, `max_groups` or the +`sequence` keys never makes a rule unscored: +- features are computed over what was processed, in a deterministic order (newest-first for + trailing features, earliest-first for onboarding and anchored ones); +- `core.history_truncated = 1` is set, and `core` requires it to carry a **positive** weight. + +A pack **error** (a bug) still marks the rules that read that pack as unscored and degraded. +Truncation never does. + +**Why flooding with cheap events can't evade:** +- (a) Onboarding facts and `start` are loaded separately and are never displaced. +- (b) Newest-first loading keeps every trailing window, up to the byte cap. The flood is itself the + most recent activity, so it is counted, and it raises every volume or velocity feature it + matches. +- (c) Truncation drops only the *oldest* trailing events, which feed three things: + - baselines: a smaller baseline gives a larger `v1`, because `cur / max(base, 1)` is monotone; + - lifetime denominators: a smaller denominator gives a larger ratio (as with `burst_ratio`); + - lifetime totals: a smaller total gives a smaller value. This is the only direction that can + lower risk. +- (d) An invariant closes the third case. `packtest` checks it for every weights file that + includes `core.history_truncated`: + `w(core.history_truncated) ≥ Σ over features f with TruncationDir(f) = −1 of |w_f| · Bound_f`. + Truncation therefore never lowers the linear sum. Two Go features have no natural bound + (`core.resource_total`, `core.credential_total`). They get `Bound` from a cap of 1,000, which no + fixture reaches, so semantic identity holds. +- (e) Forcing truncation sets the abusive subject's own `history_truncated` signal. + +Criterion 6's property test covers every built-in and DSL feature. + +**Limits** (validated at load; also the fuzz oracle): | Limit | Value | | --- | --- | | Custom features per tenant | 64 | -| Cost units per tenant | 256 | -| Predicate depth / leaves per feature | 2 / 8 | -| `in` list size / set file entries | 256 / 100,000 | -| Window, lookback | ≤ 30 d; `lifetime` allowed only for `count`, `distinct`, `share`, `time_between` | +| Features per event type (fan-out) | 16 | +| Predicate depth / leaves | 2 / 8 | +| `in` list size / set-file entries | 256 / 100,000 | +| Window / lookback | ≤ 30 d (`lifetime` and `before_first` are 90 d: retention) | | `peak` window ÷ size | ≤ 1440 | -| `distinct` cap | ≤ 10,000 | -| Events scanned per subject | 50,000 newest by `(at, id)`, plus the earliest event and `subject.created` for `start` | -| Extraction deadline (custom pack) | 250 ms default, per-tenant override ≤ 1 s | - -Above the scan bound, the orchestrator sets `core.history_truncated = 1`. No committed fixture -comes near the bound, so the golden replay is unaffected. Custom-feature versioning: the loader -keeps a SHA-256 of each normalised definition per `(tenant, name, version)` in -`tenant_feature_defs`. Redefining an existing `(name, version)` with a different body fails with -`feature_version_reused`. - -### 5.8 Rules and validation at load - -A tenant profile is `config/tenants/.yaml`: - -```yaml -tenant: e2a -packs: [core@1, email@1, brand@1] -pack_params: - brand: {extra: /run/secrets/brands_extra.yaml} # optional, private - email: {channels: [email], webmail_set: default} -vocabulary: {...} # §5.4, §5.6 -features: [...] # §5.5 -tiers: {medium: 0.4, high: 0.8} -min_scored_advise: 1 -rules: - - name: new_account_velocity - mode: advise - scorer: local - weights: tenants/e2a/weights.yaml # or `uniform` (shadow only, §5.9) - inputs: [core.subject_age_h, ...] -``` - -Validation adds these checks to main §4.5's list. The whole tenant profile is rejected with a -collected error list, and the codes are machine-readable in `/healthz`: -- `pack_unknown`, `pack_requires`, `pack_version_unknown` -- `feature_unknown`: a name exists in no pack and not in the tenant's `custom.*` set -- `feature_not_enabled`: the name, or its alias, exists but its pack or `RequiresPacks` isn't - enabled -- `feature_namespace`: a tenant tries to define a feature outside `custom.*` -- `duplicate_feature`, `feature_version_reused` -- `dsl_invalid` (with a JSON-pointer path): an unknown op, a missing cap, a text field in a - predicate, an enum value not declared, a window out of range, too many leaves, a cost overrun +| `distinct` track_max / `group_by` groups / `sequence` keys | 10,000 / 1,000 / 10,000 | +| `neighbours` features | 4 | +| `also` subjects per event / declared link kinds | 3 / 8 | +| `history_bytes` / `onboarding_bytes` | 8 MiB / 1 MiB decoded | +| Steps per subject | Calibrated; default ≈ 4.2M | +| Static cost units per tenant | 256. A unit is the benchmarked cost of one O(n) op at the byte cap. | + +### 5.8 Profiles, rules, scheduling + +**Profiles live in a private mount** (`--profiles /etc/abusekit/tenants/`, mounted by the hosted +deploy from the operator's private config repo). This repo ships only: +- `examples/tenants/*.yaml`: the five fictional §7 profiles; +- `examples/tenants/reference/`: today's `config/rules.yaml` + `local_weights.yaml`, renamed. The + golden replay runs this profile, and e2a's private profile starts as a copy of it. + +**Validation.** It adds these codes to main §4.5's list. Any failure rejects the whole profile, +with every error collected: +- `pack_unknown`, `pack_requires` +- `feature_unknown`, `feature_not_enabled`, `feature_namespace`, `duplicate_feature`, + `feature_version_reused` +- `dsl_invalid` (with a JSON-pointer path) - `vocab_invalid`, `vocab_incompatible` -- `weights_unknown_feature`: every weight must name an input of the rule it serves -- `uniform_not_shadow`: a rule that uses `weights: uniform` must be in `mode: shadow` - -**Reload is atomic per tenant.** A rejected profile keeps that tenant's previous profile live -and has no effect on other tenants; `/healthz` reports `config_error{tenant}`. This replaces -today's whole-config rejection, which would let one tenant's typo freeze every tenant's rule -changes. Before any tenant file exists, `config/rules.yaml` + `config/local_weights.yaml` load -as the **default profile**. That profile has implicit `packs: [core@1, email@1, brand@1]`, the -implicit legacy vocabulary, and flat-name resolution, and it applies to every tenant that has -keys but no file. This is the zero-change path for e2a until slice G7 (§9). - -### 5.9 Weights, scorers, eval and bootstrap per pack - -- **Weights are per rule**, keyed by namespaced name (flat names are accepted through the alias - table). The scorer name stays `local`. At profile compile time the loader binds a local scorer - instance per weights file. `Version()` is content-derived over canonical keys (§5.2), so - different weights produce different versions automatically. Registry lookups become - `(tenant, scorer)`, and vendor scorers stay global. -- **The local scorer's math is unchanged:** `sigmoid(bias + Σ w·x)`, summed in canonical-key - order. -- **Starter weights per pack.** Each pack ships `config/packs//starter.yaml`: one rule, - `_starter`, with its own bias and weights over that pack's features only. Every weight - carries `sign: +|-` (the golden-sign contract). Starter rules compose: a tenant can enable - `core_starter` and `brand_starter` as separate shadow rules, and `Combine`'s - `max(risk)` handles them without inventing a joint model. The `core` starter is derived from - today's e2a weights restricted to core features. The `email` and `brand` starters are derived - the same way. All are placeholders until labelled data exists for a second product. -- **Uniform priors, shadow only.** `weights: uniform` binds a local scorer with, for `k` inputs, - `x̂ᵢ = xᵢ / Boundᵢ ∈ [0, 1]` and `wᵢ = ±4/k` on `x̂ᵢ`. The sign defaults to `+`, and a feature - can declare `prior_sign: -` (for example, an account-age or time-to-first-action feature). The - bias is `−2 + (4/k) × (number of negative-sign inputs)`, so risk always ranges from - sigmoid(−2) ≈ 0.12 to sigmoid(2) ≈ 0.88. The result is a ranking device - for operator review, never a tier driver. The loader rejects it in `advise` - (`uniform_not_shadow`), and promotion (main §4.5) requires a real weights file plus a gate run. -- **Eval is scoped to a profile.** `abusekit eval --profile pack:email` or - `--profile tenant:e2a` selects the rule, weights, fixtures and floors. Layout: - `eval/packs//{fixtures/, floors.yaml}` and `eval/tenants//{fixtures/, - floors.yaml, golden/}`. `floors.yaml` entries gain `profile:`; a missing value means - `tenant:e2a`, so #5's file keeps working unchanged. The manifest gains `profile`, - `profile_sha` and `pack_versions`. -- **Cassettes** are unchanged. Their key already includes `input_hash`, and canonical keys keep - it stable (§5.2). -- **Corpus v2** (`eval/schema/corpus-v2.schema.json`) is v1 plus a required `profile` and - `vocab_version`, with `input.features` keyed by namespaced name. `LoadSnapshotCorpus` reads - both versions. -- **Golden-sign and mutation tests per pack** become a reusable harness, - `internal/pack/packtest.Run(t, pack)`, and every pack must pass it in CI, like the adapter - contract test. It checks: - 1. Every starter weight's sign matches `sign:`. - 2. Zeroing each weight moves at least one of the pack's fixture bands or an isolated scenario - (today's `mutation_test.go` logic, parameterised). - 3. Determinism: same bits under shuffled event order and repeated runs. - 4. No leakage from the future: adding an event at `now + ε` changes no value except rescore - candidates. - 5. Every emitted name is in `Features()` and in the pack's namespace, with `Bound` respected. - - e2a's tenant profile keeps its own golden-sign and mutation suite over its composed rule, with - namespaced keys. -- **How a second product bootstraps:** - 1. Enable packs, declare the vocabulary, and write custom features. Run shadow rules - `core_starter` (plus `brand_starter` if relevant) and a `custom_uniform` rule over the - custom features. - 2. Collect labels through `POST /v1/labels` (main §4.9). Corpus rows accrue per profile. - 3. Once a labelled set passes the harness, hand-tune a real weights file, set floors, and - promote through the normal shadow → advise path. Fitting is §12 Q8. - -### 5.10 Provenance and storage changes (expand-only) - -- `events.vocab_version text NULL`. -- `verdicts.profile_sha text NULL`: SHA-256 of the pack versions, the normalised custom-feature - definitions and the vocabulary version. It is recorded, not part of the input hash, so - editing an unrelated custom feature doesn't force rescoring. A changed value still changes the - hash through the value. -- New tables `tenant_vocabularies` and `tenant_feature_defs` (§5.6, §5.7). -- No change to `links`, `subjects`, `labels` or `corpus_examples`. +- `weights_unknown_feature`, `uniform_not_shadow`, `truncation_weight_insufficient` +- `subject_kind_unknown` + +**Per-tenant everything:** +- **Reload.** Atomic per tenant. A rejected profile keeps that tenant's previous profile live and + never affects another tenant. `/healthz` reports `config_error{tenant}`. +- **Rule sets.** `worker.computeVerdict` and `serve.currentRuleNames()` both read the subject's + tenant profile, not a global config. P2's test: tenant A's retired rule never appears in tenant + B's view. +- **Scorer version.** Each tenant's local scorer is bound to its own weights file. Its + content-derived version (`local@`) appears in the verdict `model` field and in the + signals of `GET /v1/subjects/{id}`. + +**Stage gates consider only advise-mode local rules.** `maxRiskByScorer` excludes shadow rules, so +a shadow experiment can never open or close a vendor call's gate. This changes behaviour on +`main`, but e2a has no staged rules, so the golden replay is unaffected. + +**Scheduling:** +- **Per-tenant fair queue.** The worker claims dirty subjects round-robin across tenants, with + per-tenant weights (equal by default) and a per-tenant concurrency cap (4 of the batch by + default). One tenant's backlog can't starve another. Main §4.8's priority order applies within a + tenant. +- **Rescore-storm control.** Timer rescores (as opposed to event-driven dirty marks) have three + limits: + - they are scheduled only from features that feed at least one non-shadow rule; + - DSL features coalesce into buckets of `max(5 min, W / 12)`, while built-in Go features keep + the fixed 5-minute bucket for semantic identity; + - a per-tenant timer-rescore budget (default 20 × active subjects per hour) is enforced by the + queue. When it is exhausted, timer rescores defer to the next hour and a metric counts them. + + Event-driven scoring is never budgeted. + +### 5.9 Weights, eval and bootstrap per pack + +- **Weights.** Per rule, keyed by namespaced name. The local math is unchanged. +- **Starter weights.** Each pack has `config/packs//starter.yaml`, with one rule named + `_starter`. Every weight has a `sign:`. `core`'s starter includes + `core.history_truncated`, set so the §5.7 invariant holds. +- **Uniform priors, shadow only.** `weights: uniform` binds a local scorer with, for `k` inputs: + - normalisation `x̂ᵢ = min(max(xᵢ, 0), Boundᵢ) / Boundᵢ`, which clamps to `[0, 1]`; + - `wᵢ = sᵢ · 4/k`, where `sᵢ` is the input's prior sign: `FeatureDef.PriorSign` for pack + features, `prior_sign` for custom ones, default `+`; + - `bias = −2 + (4/k) · |{i : sᵢ = −1}|`. + + So `risk ∈ [sigmoid(−2), sigmoid(2)] ≈ [0.12, 0.88]`. The loader rejects uniform weights in + advise (`uniform_not_shadow`). Promotion needs a real weights file plus a gate run. +- **Eval by profile.** `abusekit eval --profile examples/tenants/` or `--pack

`. + - Floor entries gain `profile:`; an entry without one belongs to the reference profile. + - The manifest gains `profile_sha`, `pack_versions` and `feature_key_space`. + - Corpus v2 is corpus-v1 plus `profile`, `vocab_version` and `feature_key_space: ns-v1`. +- **`packtest`** runs for every pack in CI and checks: + - starter-weight signs match their `sign:`; + - zeroing any weight moves a pack fixture band or an isolated scenario; + - results are deterministic under shuffled arrival; + - no future leakage; + - the pack stays in its namespace and within `Bound`; + - the truncation invariant holds; + - the flood property (criterion 6). +- **Held-out fixtures** for the bootstrap criterion live in + `examples/tenants//fixtures/{dev,heldout}/`. + - Held-out fixtures are authored in a separate commit after the features are frozen, with + non-overlapping generator seeds. + - CI evaluates criterion 2 only on the held-out set. + - A PR that changes a scenario's features and its held-out fixtures together fails a CI check; + held-out fixtures change only in their own PR. +- **Bootstrap.** A new product: + 1. writes its profile; + 2. runs `core_starter` (plus `brand_starter` if it has display names) and a `custom_uniform` + rule, all in shadow; + 3. collects labels through `POST /v1/labels`; + 4. once the harness passes on its labelled set, hand-tunes a weights file and promotes it + through the normal path. + +### 5.10 Storage (expand-only) + +- `events`: add `vocab_version text NULL` and `subject_kind text NOT NULL DEFAULT 'account'`. +- New table `event_subjects(tenant, subject_kind, subject, event_seq)`. It indexes `also` and + parent rows, and the scoring loader reads through it. +- `subjects`: the key becomes `(tenant, kind, subject)`, via a new `kind` column (default + `account`) and a unique index. +- `links`: add `key_id` for rotation. +- `corpus_examples`: add `feature_key_space`. +- `verdicts`: add `profile_sha` and `reason_version`. +- New table `tenant_config_versions`, the runtime guard (§5.6). + +**S3b erasure must be vocabulary-aware.** Which stored fields are raw text, skeleton or +pseudonymised hash depends on each row's `vocab_version`, and the erasure ledger records the +version it applied. S3b's design pass must include this. ### 5.11 API surface summary | Surface | Change | Compatibility | | --- | --- | --- | -| `POST /v1/events` | `delivery.sent` built-in; tenant-declared types redacted by kind; undeclared fields hashed | Additive on the wire. Storage of undeclared types changes (pre-GA, §5.4). | -| `GET /v1/subjects/{id}`, `evaluate` | Signal `reason` templates may name namespaced features in prose | Additive; the shape is unchanged. | -| Per-item codes | none new (`redaction_failed`, `bad_type` reused) | unchanged | -| Config YAML | tenant profiles; `rules.yaml` still loads as the default profile | Backward compatible. Flat names are deprecated with a warning. | -| Weights, floors, corpus files | namespaced keys; `profile:`; corpus-v2 | v1 read forever via the alias table | -| `pkg/abusekit` client | `DeliverySent` event helper; no removals | additive | +| `POST /v1/events` | Optional `subject_kind`, `also`, `links.custom` and `x_` extensions; declared types; stricter domain kind; card/IP/phone scan; re-HMAC; undeclared values dropped | Wire additive. Storage semantics change (pre-GA; decisions Q3, Q14). | +| `GET /v1/subjects/{id}`, `evaluate` | Optional `?kind=`; per-tenant `model` version; reasons use namespaced names; `warming_until` on warming signals | Additive | +| Per-item codes | None new | Unchanged | +| `abusekit score --jsonl` | Namespaced names; flat names return `feature_renamed` | Breaks once, pre-GA (P1) | +| Config | Per-tenant profiles in a private mount; `rules.yaml` replaced by `examples/tenants/reference` | Breaks once, pre-GA | -**Rejected API alternative:** a `PUT /v1/vocabulary` endpoint that would let producers declare -schemas at runtime. It would let a producer key widen its own redaction boundary, which is a -privilege escalation. It would also move a privacy decision out of code review. +**Rejected: runtime vocabulary declaration over HTTP** (`PUT /v1/vocabulary`). It would let a +producer key widen its own redaction boundary. ### 5.12 Alternatives considered -- **CEL (cel-go) for predicates and features.** It is sandboxed, has cost estimation, and is a - known quantity. It lost for four reasons: - 1. Features aggregate over time-windowed sequences of events. CEL has no windowed aggregates, - so we would still have to write count, distinct, peak, time-between and history-relative as - custom functions. CEL would only wrap the predicate, which is the easy part. - 2. It pulls in a large dependency (cel-go plus protobuf), against the repo's minimal-dependency - convention. - 3. CEL's cost estimate is per expression. Ours has to be per subject history, which needs our - own model anyway. - 4. Its error messages and semantics (`has()`, dynamic types) are harder for a product engineer - to get right than a closed YAML schema whose load errors carry JSON pointers. - - CEL remains the fallback **for `where` only** if the closed predicate set proves too small - (§12 Q9). -- **A home-grown expression language.** It would bring a parser, a grammar, precedence rules and - an injection surface, all needing a security review, for no coverage beyond the six closed - operations the target signals need. -- **A plugin ABI.** Go `plugin` needs an identical toolchain and build flags and has no sandbox. - WASM (wazero) is sandboxed with fuel metering, but it deploys arbitrary code disguised as - config, reviewers can't read it, and float determinism depends on the guest. Products that - need code contribute a Go pack upstream. The `Pack` seam is where that code goes, and - `packtest` gates it. -- **SQL-defined features against the store.** They couple to the schema, their cost is - unbounded, and they are a tenant-isolation hazard. Rejected. -- **Keep flat names and prefix only new features.** No migration, but also no enablement - boundary, and two naming styles forever. Rejected in favour of the canonical-key bridge, which - costs one frozen table. -- **Translate `content.sent` to `delivery.sent` at ingest.** See §5.4. Rejected for the - body-hash and migration costs. -- **Rename everything and rescore once.** This gives a simpler hash with no canonical keys. It - lost because summation order changes the last bits of risk (breaking the bit-for-bit - criterion), and any vendor cassette recorded before cutover would go stale. +- **CEL.** It has no windowed aggregates, so we would still write every op. It is also a large + dependency, costs expressions rather than histories, and gives product engineers worse errors. + It remains a possible later leaf kind for `where` only (§12 Q9). +- **A home-grown expression language.** A parser, a grammar and an injection surface, with no + coverage beyond the closed ops. +- **A plugin ABI.** Go `plugin` is fragile and unsandboxed; WASM is code disguised as config, and + reviewers can't read it. Code belongs in a Go pack upstream, gated by `packtest`. +- **SQL features.** They would couple features to the store's schema, give unbounded cost, and + put tenant isolation at risk. +- **A canonical-key bridge (revision 1).** It lost because nothing stored needs it (A1), and it + would keep two names alive forever. +- **A neutral built-in `delivery.sent` (revision 1).** It lost because "delivery" is still an + email-shaped abstraction with channels bolted on. Declared types with field roles are genuinely + neutral, and `content.sent` stays the email pack's own type. +- **Hashing undeclared strings (revision 1).** It lost because the hash of an unreviewed field is + still pseudonymous personal data nobody asked for. Dropping it is strictly safer, and declaring + a field is cheap. +- **Wall-clock deadlines (revision 1).** They lost because they are non-deterministic and + host-dependent, and because a timeout that marks a rule unscored rewards flooding. A step + budget plus truncation-as-signal has neither problem. ## 6. Edge cases and failure handling -- **A pack's `Extract` errors or the custom pack exceeds its deadline.** The orchestrator - records the failure per pack, not per subject. Rules whose inputs include any feature of that - pack become `unscored` with `error_code: feature_error` or `feature_timeout`, and - `degraded: true` is set. Rules that don't read the pack still score. This fails closed: - `unknown` or `degraded`, never `low` by absence (main §5). A pack that fails for every subject - of a tenant pages through the existing metric. -- **A profile is rejected on reload.** The previous profile stays live for that tenant only. - On a cold start with no valid profile, the tenant's subjects stay unscored (`unknown`) and - `/healthz` is red. The service never falls back to a different tenant's rules. -- **Events of a type arrive before its declaration (deploy ordering).** They are stored under - the undeclared-type rule (strings hashed). Features declared later can't read those strings, - but they can still count the events. The runbook says to declare first, and the harness - reports the row counts per `vocab_version`. -- **A declared kind or channel is missing on an event.** Undeclared kinds count as `other`, as - in §5.4. For channels, `delivery.sent` with an undeclared channel is rejected - (`redaction_failed`), because silently counting it under a guessed channel would corrupt the - email features. -- **Absent optional fields in custom features.** Predicates are false. `sum` uses `default`. - `share` with a zero denominator uses `if_empty`. `time_between` with no `from` event uses - `if_absent`. There is never a NaN: the compiler proves every op total, and the orchestrator - rejects a non-finite value as a pack error. -- **Duplicates and out-of-order events.** Ingest idempotency is unchanged. Features are - set-functions with `(at, id)` tie-breaks. -- **Clock skew and future events.** Excluded from custom windows, but they schedule a rescore - at their `at`. The frozen legacy exceptions are listed in §5.5. -- **Disabling a pack that rules still reference.** The load fails with `feature_not_enabled`. - Past verdicts stay; they are provenance. -- **A tenant enables `email` for a non-email channel.** `email.channels` must name declared - channels whose `destination` kind is `domain`. Otherwise the load fails (`pack_params_invalid`). - This stops webmail and domain logic running over hashes. -- **Alias misuse.** A rule lists both `sends_1h` and `email.sends_1h` → `duplicate_feature`. A - custom feature named after an alias is impossible because of the `custom.` prefix. -- **Hostile config.** There is no code and no regex. Every string set is hashed at load. Every - size is capped. The loader is fuzzed with the §5.7 limits as the oracle. -- **Hostile events against a custom feature.** An attacker can't exceed the per-subject scan or - deadline bounds. Flooding one subject raises only that subject's cost, which the existing - per-subject budgets cap. `distinct` memory is O(cap). -- **Brand pack on a product whose display names are routinely brand-adjacent**, such as a - marketplace reselling branded goods. `brand.*` stays shadow until that tenant's own labels - justify a weight. That tenant can also point `brand.extra` at an empty list and a narrowed - public list through `pack_params.brand.list` (§12 Q10). - -## 7. Worked examples (fictional) - -All three products, their names, ids and domains are invented. Timestamps use the fictional -2031 convention. +- **A pack errors (a bug).** Rules that read that pack are unscored with `feature_error` and + marked degraded; other rules still score. Hitting a budget or truncating is not an error (§5.7). +- **A profile is rejected.** The tenant's previous profile stays live. On a cold start with no + valid profile, the tenant's scores are `unknown` and `/healthz` is red. Another tenant's rules + are never borrowed. +- **Events arrive before their declaration.** Undeclared data is dropped (§5.6 rule 3). A feature + declared later sees only the type and time of those rows. Warm-up keeps rules that use it in + shadow until its window is fully covered. +- **Unknown values.** + - An undeclared resource kind is treated as `other` and counted in a metric; in `strict` mode it + is rejected. + - An undeclared `subject_kind` or `also` kind is rejected with `redaction_failed`. + - An undeclared link kind is rejected with `bad_links`. +- **Absent fields.** A predicate on an absent field is false, and `sum` uses `default`. + `share`/`ratio` use `if_empty`, and `time_between` uses `if_absent`. Every op is total; a + non-finite value is a pack error. +- **Duplicates, out-of-order arrival, ties.** Idempotency is unchanged. Features are set functions + over the history ordered `(at, producer, id)`. +- **Clock skew and future events.** They are excluded from custom windows, but loaded (within + 24 h) so they can schedule a rescore. The legacy exceptions are listed in §5.5. +- **Key rotation during a burst.** Dual keys (§5.6) keep `distinct` and neighbour equality exact. +- **`also` abuse.** A producer naming arbitrary subjects is limited to 3 per event. Each dirty + mark counts against the per-tenant rescore and scoring budgets. +- **Parent fan-in.** Many `api_key` subjects mark the same parent account dirty. Dirty marks + coalesce per subject through `dirty_seq`, so the parent costs O(1) per scoring round. +- **Hostile config.** No code or regex, every set hashed, every size capped, and the loader is + fuzzed. +- **Hostile events.** Byte caps, the step budget, and capped groups and keys bound the cost. + Truncation raises risk rather than lowering it (§5.7). +- **The brand pack on a marketplace that resells branded goods.** `brand.*` stays in shadow until + the tenant's own labels justify it. The tenant may also narrow the brand list (§12 Q10). + +## 7. Genericity walk: five scenarios (fictional) + +All products, ids and domains are invented; timestamps use the 2031 convention. Each profile is +committed in P6b as `examples/tenants//`, with dev and held-out fixtures. ### 7a. File sharing: malware-distribution burst ("Driftbox") -The pattern: a fresh account uploads an executable or archive, creates many public share links -quickly, and those links are downloaded from many distinct networks within an hour. +The pattern: a fresh account uploads executables or archives and creates many public links, which +are then downloaded from many distinct networks within the hour. ```yaml -tenant: driftbox -packs: [core@1, brand@1] # no email pack: Driftbox doesn't deliver mail +packs: [core@1, brand@1] vocabulary: version: 1 - resource_kinds: - workspace: {role: workspace} - api_token: {role: credential} - folder: {role: content} + resource_kinds: {api_token: {role: credential}} types: share.link_created: + role: activity fields: visibility: {kind: enum, values: [public, org, private]} file_kind: {kind: enum, values: [document, archive, executable, image, other]} - size_bytes: {kind: number, min: 0, integer: true} + folder_title: {kind: text, role: title, max_len: 120} # skeleton-only share.downloaded: fields: link_hash: {kind: hash} - downloader_ip24: {kind: hash} # producer-keyed hash; never a raw IP - downloader_is_owner: {kind: bool} + downloader_ip24: {kind: hash} + downloader_is_owner: {kind: bool, role: self} features: - - name: custom.public_links_1h - version: 1 - description: public share links created in the last hour - count: {type: share.link_created, where: {field: visibility, eq: public}} - window: 1h - transform: {log1p: true, cap: 6} - - name: custom.risky_file_link_share_24h - version: 1 - description: share of new links pointing at executables or archives - share: - type: share.link_created - match: {field: file_kind, in: [executable, archive]} - window: 24h - transform: {cap: 1} - - name: custom.distinct_downloader_nets_1h - version: 1 - description: distinct downloader /24 networks, excluding the owner - distinct: - type: share.downloaded - where: {field: downloader_is_owner, eq: false} - field: downloader_ip24 - window: 1h - transform: {log1p: true, cap: 9} # track_max = ceil(expm1(9)) = 8103 ≤ 10,000 - - name: custom.download_peak_10m_vs_history - version: 1 - description: 10-minute download peak relative to the account's own past - peak: {type: share.downloaded, where: {field: downloader_is_owner, eq: false}, size: 10m} - window: 24h - relative_to_history: {lookback: 30d, exclude_recent: 24h, age_decay: {}} - transform: {log1p: true, cap: 6} - - name: custom.signup_to_first_public_link_min - version: 1 - description: minutes from sign-up to the first public link - time_between: - from: {type: subject.created} - to: {type: share.link_created, where: {field: visibility, eq: public}} - until_now: true - if_absent: 1440 - transform: {cap: 1440} - prior_sign: "-" # faster = riskier + - {name: custom.public_links_1h, version: 1, description: public links in the last hour, + count: {type: share.link_created, where: {field: visibility, eq: public}}, + window: 1h, transform: {log1p: true, cap: 6}} + - {name: custom.risky_link_share_24h, version: 1, description: links to executables or archives, + share: {type: share.link_created, match: {field: file_kind, in: [executable, archive]}}, + window: 24h, transform: {cap: 1}} + - {name: custom.distinct_downloader_nets_1h, version: 1, description: distinct downloader networks, + distinct: {type: share.downloaded, where: {field: downloader_is_owner, eq: false}, field: downloader_ip24}, + window: 1h, transform: {log1p: true, cap: 9}} # track_max 8103 + - {name: custom.max_downloads_per_link_1h, version: 1, description: busiest link's downloads, + count: {type: share.downloaded, where: {field: downloader_is_owner, eq: false}, + group_by: {field: link_hash, reduce: max, max_groups: 1000}}, + window: 1h, transform: {log1p: true, cap: 8}} + - {name: custom.download_peak_vs_history, version: 1, description: 10-min download peak vs own past, + peak: {type: share.downloaded, size: 10m}, window: 24h, + relative_to_history: {lookback: 30d, exclude_recent: 24h, age_decay: {}}, + transform: {log1p: true, cap: 6}} + - {name: custom.signup_to_first_public_link_min, version: 1, description: minutes to first public link, + time_between: {from: {type: subject.created}, to: {type: share.link_created, where: {field: visibility, eq: public}}, + until_now: true, if_absent: 1440}, + transform: {cap: 1440}, prior_sign: "-"} rules: - - name: core_starter - mode: shadow - scorer: local - weights: packs/core/starter.yaml - inputs: [core.subject_age_h, core.credential_velocity_1h, core.resource_velocity_1h, - core.declines_before_first_success, core.first_funding_prepaid, - core.linked_deleted_n, core.linked_labelled_abusive_n, core.burst_ratio_24h_vs_lifetime] - labels: [benign, abusive] - benign_label: benign - threshold: 0.6 - - name: malware_burst - mode: shadow - scorer: local - weights: uniform - inputs: [custom.public_links_1h, custom.risky_file_link_share_24h, - custom.distinct_downloader_nets_1h, custom.download_peak_10m_vs_history, - custom.signup_to_first_public_link_min, brand.name_match] - labels: [benign, abusive] - benign_label: benign - threshold: 0.7 + - {name: malware_burst, mode: shadow, scorer: local, weights: uniform, + inputs: [custom.public_links_1h, custom.risky_link_share_24h, custom.distinct_downloader_nets_1h, + custom.max_downloads_per_link_1h, custom.download_peak_vs_history, + custom.signup_to_first_public_link_min, brand.title_match, core.history_truncated], + labels: [benign, abusive], benign_label: benign, threshold: 0.7} ``` -Sample events: - ```json -{"id":"db-001","subject":"acct_example_db_1","type":"subject.created","at":"2031-03-02T09:00:00Z","links":{"email_hash":"<64-hex>"},"data":{"channel":"signup","email_domain_class":"disposable"}} -{"id":"db-002","subject":"acct_example_db_1","type":"resource.created","at":"2031-03-02T09:01:10Z","data":{"kind":"workspace","name":"Official Document Center"}} -{"id":"db-003","subject":"acct_example_db_1","type":"share.link_created","at":"2031-03-02T09:03:00Z","data":{"visibility":"public","file_kind":"archive","size_bytes":812345}} -{"id":"db-004","subject":"acct_example_db_1","type":"share.link_created","at":"2031-03-02T09:03:20Z","data":{"visibility":"public","file_kind":"executable","size_bytes":402112}} -{"id":"db-005","subject":"acct_example_db_1","type":"share.downloaded","at":"2031-03-02T09:05:02Z","data":{"link_hash":"lk_4f1c9a0e7b2d","downloader_ip24":"ip_9a1b2c3d4e5f","downloader_is_owner":false}} +{"id":"db-003","subject":"acct_example_db_1","type":"share.link_created","at":"2031-03-02T09:03:00Z","data":{"visibility":"public","file_kind":"archive","folder_title":"Invoice Center"}} +{"id":"db-005","subject":"acct_example_db_1","type":"share.downloaded","at":"2031-03-02T09:05:02Z","data":{"link_hash":"lk_4f1c9a0e7b2d11","downloader_ip24":"ip_9a1b2c3d4e5f66","downloader_is_owner":false}} ``` -The benign counterpart fixture: an older workspace sharing documents with an organisation. Its -links are `org`-visibility, with a few downloads from two networks. +**Declarative:** everything above. **Go-only:** file-content verdicts, such as a malware hash or a +sandbox result. The product emits these as `content.verdict`, and `core.verdict_max_24h` reads +them. -### 7b. Payments or marketplace: card testing ("Tallyport") +### 7b. Marketplace: card testing ("Tallyport") -The pattern: a merchant account (the subject) pushes many small charge attempts across many -distinct cards, most of them declined, in short bursts. +The pattern: a merchant account pushes many small charge attempts across many cards, and most are +declined. A card that turns up across many merchants is suspicious in itself. ```yaml -tenant: tallyport -packs: [core@1, brand@1] # core also scores the merchant's own onboarding payments +packs: [core@1, brand@1] vocabulary: version: 1 - resource_kinds: - api_key: {role: credential} - storefront: {role: workspace} + subject_kinds: {account: {}, card: {}} + resource_kinds: {api_key: {role: credential}} types: charge.attempted: + role: activity fields: outcome: {kind: enum, values: [succeeded, declined, blocked]} decline_code: {kind: enum, values: [insufficient_funds, do_not_honor, incorrect_cvc, expired_card, fraudulent, other]} amount_minor: {kind: number, min: 0, integer: true} - card_hash: {kind: hash} + card_hash: {kind: hash, join_domain: card} features: - - name: custom.declines_10m_peak - version: 1 - description: largest number of declined charges in any 10 minutes today - peak: {type: charge.attempted, where: {field: outcome, eq: declined}, size: 10m} - window: 24h - transform: {log1p: true, cap: 7} - - name: custom.distinct_cards_1h - version: 1 - description: distinct cards charged in the last hour - distinct: {type: charge.attempted, field: card_hash} - window: 1h - transform: {log1p: true, cap: 7} - - name: custom.small_charge_share_1h - version: 1 - description: share of charges at or under 2.00 in minor units - share: {type: charge.attempted, match: {field: amount_minor, lte: 200}} - window: 1h - transform: {cap: 1} - - name: custom.decline_share_1h - version: 1 - description: share of charges declined - share: {type: charge.attempted, match: {field: outcome, in: [declined, blocked]}} - window: 1h - transform: {cap: 1} - - name: custom.cvc_declines_vs_history - version: 1 - description: CVC/expiry declines this hour vs the merchant's own past - count: - type: charge.attempted - where: {all: [{field: outcome, eq: declined}, {field: decline_code, in: [incorrect_cvc, expired_card]}]} - window: 1h - relative_to_history: {lookback: 30d, exclude_recent: 24h, age_decay: {}} - transform: {log1p: true, cap: 6} + - {name: custom.max_declines_per_card_1h, version: 1, description: most declines on one card, + count: {type: charge.attempted, where: {field: outcome, eq: declined}, + group_by: {field: card_hash, reduce: max, max_groups: 1000}}, + window: 1h, transform: {cap: 20}} + - {name: custom.cards_with_3plus_declines_1h, version: 1, description: cards declined 3+ times, + count: {type: charge.attempted, where: {field: outcome, eq: declined}, + group_by: {field: card_hash, reduce: {count_gte: 3}, max_groups: 1000}}, + window: 1h, transform: {log1p: true, cap: 6}} + - {name: custom.distinct_cards_1h, version: 1, description: distinct cards, + distinct: {type: charge.attempted, field: card_hash}, window: 1h, transform: {log1p: true, cap: 7}} + - {name: custom.small_charges_1h, version: 1, description: charges at or under 200 minor units, + count: {type: charge.attempted, where: {field: amount_minor, lte: 200}}, window: 1h, transform: {cap: 500}} + - {name: custom.charges_1h, version: 1, description: all charges, + count: {type: charge.attempted}, window: 1h, transform: {cap: 500}} + - {name: custom.small_charge_ratio_1h, version: 1, description: small ÷ all charges, + ratio: {num: custom.small_charges_1h, den: custom.charges_1h, if_empty: 0}, transform: {cap: 1}} + - {name: custom.declines_10m_peak, version: 1, description: declines in busiest 10 min today, + peak: {type: charge.attempted, where: {field: outcome, eq: declined}, size: 10m}, + window: 24h, transform: {log1p: true, cap: 7}} + - {name: custom.card_merchants_24h, version: 1, subject_kinds: [card], + description: distinct merchants that charged this card, + distinct: {type: charge.attempted, field: x_primary_subject_hash}, window: 24h, transform: {cap: 50}} rules: - - name: card_testing - mode: shadow - scorer: local - weights: uniform - inputs: [custom.declines_10m_peak, custom.distinct_cards_1h, custom.small_charge_share_1h, - custom.decline_share_1h, custom.cvc_declines_vs_history, core.credential_velocity_1h] - labels: [benign, abusive] - benign_label: benign - threshold: 0.7 + - {name: card_testing, mode: shadow, scorer: local, weights: uniform, applies_to: [account], + inputs: [custom.max_declines_per_card_1h, custom.cards_with_3plus_declines_1h, custom.distinct_cards_1h, + custom.small_charge_ratio_1h, custom.declines_10m_peak, core.credential_velocity_1h, + core.history_truncated], labels: [benign, abusive], benign_label: benign, threshold: 0.7} + - {name: tested_card, mode: shadow, scorer: local, weights: uniform, applies_to: [card], + inputs: [custom.card_merchants_24h], labels: [benign, abusive], benign_label: benign, threshold: 0.7} ``` +Each charge is also indexed under the card subject, through `also`: + ```json -{"id":"tp-101","subject":"acct_example_tp_7","type":"charge.attempted","at":"2031-06-10T02:14:01Z","data":{"outcome":"declined","decline_code":"incorrect_cvc","amount_minor":100,"card_hash":"cd_1a2b3c4d5e6f"}} -{"id":"tp-102","subject":"acct_example_tp_7","type":"charge.attempted","at":"2031-06-10T02:14:04Z","data":{"outcome":"declined","decline_code":"expired_card","amount_minor":100,"card_hash":"cd_7f8e9d0c1b2a"}} -{"id":"tp-103","subject":"acct_example_tp_7","type":"charge.attempted","at":"2031-06-10T02:14:09Z","data":{"outcome":"succeeded","amount_minor":100,"card_hash":"cd_0f1e2d3c4b5a"}} +{"id":"tp-101","subject":"acct_example_tp_7","also":[{"kind":"card","id":"card_example_c1"}],"type":"charge.attempted","at":"2031-06-10T02:14:01Z","data":{"outcome":"declined","decline_code":"incorrect_cvc","amount_minor":100,"card_hash":"cd_1a2b3c4d5e6f77"}} ``` -The benign counterparts: a storefront with steady larger charges and an ordinary decline rate, -and a storefront whose flash sale has high volume but few distinct-card declines. The second -exercises `relative_to_history`, because the merchant's own past peaks raise the baseline. +`x_primary_subject_hash` is a **derived** field. Ingest adds it to every row indexed through +`also`, as the re-HMACed id of the primary subject (join domain `subject:`). It is declared +implicitly for any type used with `also`, so a secondary subject can count distinct primaries. -`payment.attempt` is **not** used for the charges. In the core vocabulary, `payment.attempt` -means the subject paying the product, and it feeds onboarding facts. Card testing is the -merchant's product activity, so it is a custom type. The example makes that distinction -explicit. +**Declarative:** everything above. **Go-only:** issuer and BIN intelligence, and velocity seen by +external card networks. Products can supply these as `content.verdict` or enum fields. ### 7c. Chat or community: spam invites ("Hearthchat") -The pattern: new accounts with brand-like display names send large volumes of invites to -people outside their own communities, with external links in the invite text. - -This product uses the **neutral built-in** `delivery.sent`, which is what it is for. +The pattern: new accounts with brand-like names invite people outside their own communities, +often with links, and the recipients block them soon after. ```yaml -tenant: hearthchat packs: [core@1, brand@1] vocabulary: version: 1 - resource_kinds: - profile: {role: identity} - community: {role: workspace} - bot_token: {role: credential} - channels: - invite: - destination: hash # keyed community id - destination_class: [own_community, other_community] - direct_message: - destination: hash + link_kinds: {phone_hash: {evidence: true}} + types: + invite.sent: + role: activity + fields: + invitee_hash: {kind: hash, join_domain: member} + target_class: {kind: enum, values: [own_community, other_community]} + preview: {kind: text, role: title, max_len: 200} + link_host: {kind: domain, reduce: etld1} + block.received: + fields: {blocker_hash: {kind: hash, join_domain: member}} +features: + - {name: custom.invites_10m_peak, version: 1, description: invites in busiest 10 min today, + peak: {type: invite.sent, size: 10m}, window: 24h, transform: {log1p: true, cap: 7}} + - {name: custom.external_invite_share_24h, version: 1, description: invites outside own communities, + share: {type: invite.sent, match: {field: target_class, eq: other_community}}, window: 24h, transform: {cap: 1}} + - {name: custom.invites_blocked_within_10m, version: 1, description: invitees who blocked within 10 min, + sequence: {a: {type: invite.sent}, b: {type: block.received}, within: 10m, on: {a: invitee_hash, b: blocker_hash}}, + window: 24h, transform: {log1p: true, cap: 6}} + - {name: custom.linked_invite_share_24h, version: 1, description: invites with links, + share: {type: invite.sent, match: {field: link_host, exists: true}}, window: 24h, transform: {cap: 1}} + - {name: custom.phone_siblings_7d, version: 1, description: accounts sharing a phone created this week, + neighbours: {via: [phone_hash], where: {created_within: 7d}}, transform: {cap: 20}} +rules: + - {name: invite_spam, mode: shadow, scorer: local, weights: uniform, + inputs: [custom.invites_10m_peak, custom.external_invite_share_24h, custom.invites_blocked_within_10m, + custom.linked_invite_share_24h, custom.phone_siblings_7d, brand.name_match, + brand.title_match, core.linked_deleted_n, core.history_truncated], + labels: [benign, abusive], benign_label: benign, threshold: 0.7} +``` + +`invitee_hash` and `blocker_hash` share `join_domain: member`. The same member id therefore hashes +identically in both types, and `sequence.on` can join them. + +**Declarative:** everything above. **Go-only:** classifying message text. The product emits that +as `content.verdict`. + +### 7d. Developer API: credential stuffing through customer API keys ("Keyforge") + +The pattern: a customer's API key drives many end-user login attempts across many distinct +usernames. Most fail, with the occasional success shortly after a failure on the same username. + +```yaml +packs: [core@1] +vocabulary: + version: 1 + subject_kinds: {account: {}, api_key: {parent: account}} + resource_kinds: {api_key: {role: credential}} + types: + auth.attempted: + role: activity + fields: + outcome: {kind: enum, values: [succeeded, failed, locked]} + login_hash: {kind: hash, join_domain: login} + client_ip24: {kind: hash} + client_asn: {kind: enum, values: [residential, mobile, hosting, unknown]} features: - - name: custom.invites_10m_peak - version: 1 - description: largest invite fan-out in any 10 minutes today - peak: {type: delivery.sent, where: {field: channel, eq: invite}, sum: {field: recipient_count, default: 1, cap_each: 50}, size: 10m} - window: 24h - transform: {log1p: true, cap: 7} - - name: custom.distinct_invitees_1h - version: 1 - description: distinct invitees in the last hour - distinct: {type: delivery.sent, where: {field: channel, eq: invite}, field: recipient_hash} - window: 1h - transform: {log1p: true, cap: 7} - - name: custom.external_invite_share_24h - version: 1 - description: share of invites to communities the sender doesn't own - share: - type: delivery.sent - where: {field: channel, eq: invite} - match: {field: destination_class, eq: other_community} - window: 24h - transform: {cap: 1} - - name: custom.linked_invite_share_24h - version: 1 - description: share of invites carrying an external link - share: - type: delivery.sent - where: {field: channel, eq: invite} - match: {field: link_host, exists: true} - window: 24h - transform: {cap: 1} - - name: custom.signup_to_first_invite_min - version: 1 - description: minutes from sign-up to the first invite - time_between: - from: {type: subject.created} - to: {type: delivery.sent, where: {field: channel, eq: invite}} - until_now: true - if_absent: 1440 - transform: {cap: 1440} - prior_sign: "-" + - {name: custom.distinct_logins_10m, version: 1, subject_kinds: [api_key], + description: distinct usernames tried, distinct: {type: auth.attempted, field: login_hash}, + window: 10m, transform: {log1p: true, cap: 8}} + - {name: custom.failures_1h, version: 1, subject_kinds: [api_key, account], description: failed logins, + count: {type: auth.attempted, where: {field: outcome, eq: failed}}, window: 1h, transform: {cap: 10000}} + - {name: custom.attempts_1h, version: 1, subject_kinds: [api_key, account], description: all logins, + count: {type: auth.attempted}, window: 1h, transform: {cap: 10000}} + - {name: custom.failure_ratio_1h, version: 1, subject_kinds: [api_key, account], description: failed ÷ all, + ratio: {num: custom.failures_1h, den: custom.attempts_1h, if_empty: 0}, transform: {cap: 1}} + - {name: custom.success_after_failure_10m, version: 1, subject_kinds: [api_key], + description: successes shortly after a failure on the same username, + sequence: {a: {type: auth.attempted, where: {field: outcome, eq: failed}}, + b: {type: auth.attempted, where: {field: outcome, eq: succeeded}}, + within: 10m, on: {a: login_hash, b: login_hash}}, + window: 24h, transform: {log1p: true, cap: 6}} + - {name: custom.hosting_share_1h, version: 1, subject_kinds: [api_key], description: attempts from hosting networks, + share: {type: auth.attempted, match: {field: client_asn, eq: hosting}}, window: 1h, transform: {cap: 1}} + - {name: custom.max_attempts_per_ip_1h, version: 1, subject_kinds: [api_key], description: busiest client network, + count: {type: auth.attempted, group_by: {field: client_ip24, reduce: max, max_groups: 1000}}, + window: 1h, transform: {log1p: true, cap: 9}} + - {name: custom.attempts_vs_history, version: 1, subject_kinds: [api_key], description: attempts vs own past, + count: {type: auth.attempted}, window: 1h, + relative_to_history: {lookback: 30d, exclude_recent: 24h}, transform: {log1p: true, cap: 6}} rules: - - name: invite_spam - mode: shadow - scorer: local - weights: uniform - inputs: [custom.invites_10m_peak, custom.distinct_invitees_1h, custom.external_invite_share_24h, - custom.linked_invite_share_24h, custom.signup_to_first_invite_min, - brand.name_match, brand.title_match, core.linked_deleted_n] - labels: [benign, abusive] - benign_label: benign - threshold: 0.7 + - {name: stuffing_key, mode: shadow, scorer: local, weights: uniform, applies_to: [api_key], + inputs: [custom.distinct_logins_10m, custom.failure_ratio_1h, custom.success_after_failure_10m, + custom.hosting_share_1h, custom.max_attempts_per_ip_1h, custom.attempts_vs_history, + core.history_truncated], labels: [benign, abusive], benign_label: benign, threshold: 0.7} ``` ```json -{"id":"hc-201","subject":"acct_example_hc_3","type":"resource.created","at":"2031-08-01T18:00:05Z","data":{"kind":"profile","name":"Support Team - Official"}} -{"id":"hc-202","subject":"acct_example_hc_3","type":"delivery.sent","at":"2031-08-01T18:02:11Z","data":{"channel":"invite","destination":"cm_5e6f7a8b9c0d","destination_class":"other_community","recipient_hash":"iv_0a1b2c3d4e5f","title":"You have been selected - claim now","link_host":"claim-prize.example.test"}} +{"id":"kf-9001","subject":"key_example_k3","subject_kind":"api_key","type":"auth.attempted","at":"2031-09-04T11:00:01Z","data":{"outcome":"failed","login_hash":"lg_8c7b6a5f4e3d21","client_ip24":"ip_1f2e3d4c5b6a77","client_asn":"hosting"}} ``` -The benign counterpart: a community organiser inviting 20 people over an evening to their own -community, with no links. It exercises `external_invite_share_24h` = 0 and a low peak. +**Declarative:** everything above. `client_asn` is an enum the product classifies. **Go-only:** +whether a username appears in a breach corpus (this needs an external lookup), and ASN reputation +finer than the product's own enum. + +### 7e. AI inference: free-tier farming ("Lumenloop") -**What the three examples demonstrate:** none needs the email pack. All three reuse `core` and -`brand`. Every product-specific signal is declarative. 7c uses the neutral delivery type -directly, and 7a and 7b show that custom types cover what `delivery.sent` doesn't. Slice G6 -commits each example as a loadable profile with fixtures, which is success criterion 2. +The pattern: many free accounts share a device or an OAuth identity. Each one exhausts its free +token quota soon after sign-up, favours the most expensive models, and is then abandoned. + +```yaml +packs: [core@1] +vocabulary: + version: 1 + link_kinds: {oauth_sub_hash: {evidence: true}} + types: + usage.recorded: + role: activity + fields: + model_tier: {kind: enum, values: [small, medium, large]} + tokens: {kind: number, min: 0, integer: true} + quota_state: {kind: enum, values: [ok, near_limit, exhausted]} +features: + - {name: custom.signup_to_quota_exhausted_min, version: 1, description: minutes to exhaust free quota, + time_between: {from: {type: subject.created}, to: {type: usage.recorded, where: {field: quota_state, eq: exhausted}}, + until_now: false, if_absent: 10080}, transform: {cap: 10080}, prior_sign: "-"} + - {name: custom.large_model_token_share_24h, version: 1, description: tokens spent on large models, + share: {type: usage.recorded, match: {field: model_tier, eq: large}, sum: {field: tokens, default: 0, cap_each: 200000}}, + window: 24h, transform: {cap: 1}} + - {name: custom.tokens_10m_peak, version: 1, description: tokens in busiest 10 min, + peak: {type: usage.recorded, sum: {field: tokens, default: 0, cap_each: 200000}, size: 10m}, + window: 24h, transform: {log1p: true, cap: 15}} + - {name: custom.minutes_since_last_use, version: 1, description: idle time after last use, + time_between: {from: {type: usage.recorded, anchor: last}, to: {type: abusekit.never}, until_now: true, if_absent: 0}, + transform: {cap: 10080}} + - {name: custom.oauth_siblings_deleted, version: 1, description: deleted accounts sharing the OAuth identity, + neighbours: {via: [oauth_sub_hash, device_hash], where: {deleted: permanent}}, transform: {cap: 20}} + - {name: custom.oauth_siblings_new_7d, version: 1, description: accounts sharing identity created this week, + neighbours: {via: [oauth_sub_hash, device_hash], where: {created_within: 7d}}, transform: {cap: 20}} +rules: + - {name: free_tier_farm, mode: shadow, scorer: local, weights: uniform, + inputs: [custom.signup_to_quota_exhausted_min, custom.large_model_token_share_24h, custom.tokens_10m_peak, + custom.oauth_siblings_deleted, custom.oauth_siblings_new_7d, core.first_funding_prepaid, + core.linked_deleted_n, core.history_truncated], + labels: [benign, abusive], benign_label: benign, threshold: 0.7} +``` + +`abusekit.never` is a reserved type that never occurs. With `anchor: last` and `until_now`, +`time_between` becomes "minutes since the last event of type A". That gives the idle-time +primitive with no new op. + +`custom.minutes_since_last_use` is deliberately left out of the rule. Abandonment only means +something alongside the neighbour counts, and uniform priors can't express that combination. The +feature is kept for hand-tuned weights later. + +**Declarative:** everything above. **Go-only:** +- prompt-content similarity across accounts, which needs cross-subject text clustering and a text + scorer; +- feature interactions ("abandoned **and** has farmed siblings"). A uniform-prior logistic model + can't capture these; they need tuned or fitted weights, or a multiplicative custom feature. §12 + Q15 asks whether `ratio` should gain a `product` form. + +### 7f. What the walk shows + +| Need | Primitive | +| --- | --- | +| Per-entity maxima and counts (per card, per link, per IP) | `group_by` | +| Cause, then effect within a time limit (fail → success, invite → block, signup → exhaustion) | `sequence`, `time_between` | +| Rates | `ratio` | +| Cross-account farms | Declared link kinds + `neighbours` | +| Non-account actors (cards, API keys) | Subject kinds + `also` + `parent` | + +**Still Go-only across all five:** +- content understanding (files, messages, prompts); +- external reputation lookups; +- text similarity across subjects; +- non-linear feature interactions under uniform priors. + +The first two already have a channel: products emit `content.verdict` or enum fields. ## 8. Migration plan for e2a -**Emitter: no change required.** e2a keeps emitting `content.sent` and the rest of main §4.12, -and S6 is built as currently specified. Switching to `delivery.sent{channel: email}` later is -optional and equivalent by construction; the translation test (below) proves it. - -**Config mapping.** The default profile (§5.8) serves e2a until G7. G7 then commits -`config/tenants/e2a.yaml`: -- `packs: [core@1, email@1, brand@1]` -- `vocabulary`: the implicit legacy declaration from §5.4, made explicit: `agent` → - `identity`, `key` → `credential` with S2b's aliases, and `email` → `{destination: domain}`. -- `pack_params`: the brand `extra` path (the private list, as `--brands-extra` today) and - `email.webmail_set: default` (`config/packs/email/webmail.yaml`, moved from - `config/webmail.yaml`). -- The `new_account_velocity` inputs rewritten through the alias table. The weights file moves - to `config/tenants/e2a/weights.yaml` with namespaced keys and the same values. -- Floors move to `eval/tenants/e2a/floors.yaml` with `profile: tenant:e2a`, with numbers - unchanged. - -**Proof of identical scores: the golden replay.** -1. **G0 runs first, on the pre-migration code.** It adds `cmd/abusekit golden` (test-only - build tag), which replays every fixture. Each subject is scored by the real `feature.Extract` - → `core.Plan` → local scorer → `core.Combine` path after every event instant and at every - `NextRescoreAt` the replay produces. The command writes `eval/tenants/e2a/golden/v0.jsonl`, - one row per (fixture, subject, instant): - `{features: {canonical_key: float64-bits-hex}, next_rescore_at, input_hash{rule}, - risk_bits{rule}, score_bits, tier, local_version}`. - It also records the harness `run.json` for the synthetic corpus with volatile fields removed. - Fixtures come from `eval/fixtures/*.jsonl` (including all of #7's), and the generated corpus - comes from `eval/fixtures/synthetic/`. -2. **Every later slice** runs `TestGoldenReplay_BitIdentical`, which compares exactly, row for - row, including the row count. Any diff fails CI and prints the first differing feature. -3. **The translation test** rewrites every `content.sent` in every fixture as - `delivery.sent{channel: email, ...}` (field mapping in §5.4), replays it, and asserts the - same golden file. -4. **A corpus round-trip** loads the corpus-v1 synthetic corpus, exports it as v2, reloads it, - and requires identical `run.json` metrics. - -**Stored state at cutover, if S8 has already shipped** (assumption A1 says it hasn't): -- Verdict `input_hash`: unchanged (canonical keys), so no subject is rescored and no vendor call - repeats. -- Local `Version()`: unchanged (canonical keys). -- Cassettes: unchanged keys. -- Rows stored before G5: `vocab_version` is NULL, which reads as the built-in-only schema. - -**Rollback.** Every slice up to G7 is behaviour-neutral for e2a, and the golden replay guards -it. G7 is a config move. Reverting it restores the default profile, which yields the same -golden. +- **Emitter.** No change. S6 emits `content.sent` and the rest of main §4.12 as designed. +- **Profile.** e2a's private profile starts as a byte copy of `examples/tenants/reference/`: + - `packs: [core@1, email@1, brand@1]`; + - the implicit legacy vocabulary, made explicit: `key: credential` with #7's aliases, and + `agent: other`; + - `new_account_velocity` with namespaced inputs. + + The weights move to the private mount. The private brand list stays private (`brand.extra`). + Floors for e2a's real corpus live privately; the public reference floors stay here. +- **Golden (P0).** `abusekit eval --golden out.jsonl` extends the existing eval replay rather than + adding a new command. It runs on `main` after #5 and #7 merge, and its output is committed as + `eval/golden/reference-flat.jsonl`. For every fixture, subject, event instant and scheduled + rescore instant, it records: + - feature values, as `Float64bits`; + - `NextRescoreAt`; + - per-rule input hashes, risks and tiers; + - the score; + - the local `Version()`. +- **Rename (P1).** P1 re-baselines under the semantic-identity rules of criterion 1: + - identical feature bits under the rename map; + - identical `NextRescoreAt` and tiers; + - `|Δrisk| ≤ 1e-12`, with no score near a cut point; + - every hash recorded as changed, exactly once. + + The result is committed as `eval/golden/reference-ns.jsonl`. Every later slice asserts + **exact** equality with that file. +- **Other one-time baseline changes, each isolated in its own slice:** + - the `(at, producer, id)` tie order, in P0, before capture; + - re-HMAC of hash values and links, in P3: feature values unchanged, stored bytes changed; + - the domain kind at `RedactionSchemaVersion` 3, also in P3. Fixtures use `.test`, so nothing + changes. +- **Rollback.** Before S8, each slice can be reverted on its own. After S8, the rename can't be + undone without re-scoring, which is why it lands first. ## 9. Slices -These fit after #5 and #7 merge. Each is its own PR with the usual review. - -| # | Slice | Contents | Done when | -| --- | --- | --- | --- | -| G0 | Golden capture | `cmd/abusekit golden` (test build tag); `eval/tenants/e2a/golden/v0.jsonl` generated from `main` after #5 and #7; `TestGoldenReplay_BitIdentical` | Golden committed, generated by pre-migration code; the test passes on `main`; perturbing one weight's last bit fails it | -| G1 | Canonical keys | `FeatureDef` metadata table (name, legacy, quantum, bound, reads) for today's 25 features; `core.Vector`; `inputHash` over canonical keys with quanta from data (the name switch removed); local scorer order and version by canonical key; alias resolution in rules, weights and corpus loaders; the frozen-table CI check | Golden bit-identical; a rules file in flat names and one in namespaced names load to equal `Config`s; `quantizeAgeFeaturesForHash` deleted | -| G2 | Vocabulary and delivery view | `internal/vocab` (built-in schema moved unchanged); `delivery.sent` built-in; `event.View`; `content.sent` projection; declared resource kinds with roles and aliases (implicit legacy declaration); `pkg/abusekit` `DeliverySent` | Golden bit-identical; translation test green; redaction tests moved and green; `delivery.sent` contract tests (happy, `recipient_hash`+count>1 rejection, undeclared channel rejection) | -| G3 | Packs and tenant profiles | `internal/pack` registry, `core`/`email`/`brand` adapters (code moved, not rewritten); `feature.Extract(profile, …)` orchestrator with per-pack failure isolation; `config/tenants/*.yaml` loader; default profile from `rules.yaml`; per-tenant atomic reload; `/healthz` per tenant; `brand.title_match`; new additive core features | Golden bit-identical; `feature_not_enabled` and `pack_requires` load tests; a tenant with only `core` computes no `email.*`; one tenant's bad profile leaves another's reload applied | -| G4 | Declarative features | `internal/pack/custom` compiler and evaluator, all six ops plus the modifier, predicates, limits and cost model, rescore candidates, `tenant_feature_defs`; naive reference evaluator in tests; property tests; loader fuzz; benchmark | Criterion 4 met (equality on 10k randomized histories; p99 ≤ 50 ms at the limits); a declarative `custom.key_velocity_1h` and `custom.resource_total_lifetime` match their Go twins bit for bit on every fixture (lifetime excluding future events, documented); fuzzing finds no over-limit profile that loads | -| G5 | Declared types and redaction | field kinds; masking of `text`; hashing of undeclared fields and types with the abusekit-held per-tenant key; `vocab_version` column; `tenant_vocabularies` ledger and the widening-only check; `strict` mode | Criterion 5 property tests; `vocab_incompatible` tests; e2a golden bit-identical (e2a declares no custom types) | -| G6 | Per-pack eval and bootstrap | `packtest` harness; per-pack starter weights with `sign:`; pack fixtures and floors; `--profile`; corpus-v2 schema and export; uniform priors plus `uniform_not_shadow`; the three §7 profiles committed as fixtures | Every pack passes `packtest`; `make gate` runs e2a and every pack profile; each §7 abusive fixture outranks its benign fixtures under uniform priors (criterion 2); #5's floors file still loads unchanged | -| G7 | e2a explicit profile | `config/tenants/e2a.yaml`, weights and floors moved; flat-name deprecation warning on; docs updated (main §4.5 and §4.12 pointers) | Golden bit-identical against the explicit profile; the default profile is used by no tenant in the hosted config; the reverting diff also passes golden | - -G0 must land before any other G slice. G1 → G2 → G3 are sequential. G4 and G5 can run in -parallel after G3. G6 needs G4. G7 needs G3 (G5 and G6 are optional for it). **Recommended -ordering against the v0 plan:** G0–G3 before S5, so that vendor render templates are born with -namespaced names, and before S8, so that nothing stored needs the bridge in anger. S3b and S6 -are independent of all G slices. +These come after #5 and #7 merge. The early slices are small, and pack gating arrives only after +the machinery exists. + +| # | Slice | Contents | Depends on | Done when | +| --- | --- | --- | --- | --- | +| P0 | Golden replay | `abusekit eval --golden`; scoring loader ordered `(at, producer, id)`; `eval/golden/reference-flat.jsonl` | #5, #7 | Golden committed; test green; flipping one weight's last bit fails it | +| P1 | One-time rename | Every §5.2 consumer, each with its test; the `FeatureDef` metadata table (quantum, bound, prior sign, truncation direction, reads); `core.Vector`; reason v2; the corpus key-space column and migration; the cassette header; `feature_renamed`; stage gate limited to advise-mode local rules | P0 | `reference-ns.jsonl` meets criterion 1; the grep test finds no flat literals; all existing suites green | +| P2 | Tenant profiles | Private-mount loader; `examples/tenants/reference`; per-tenant atomic reload and `/healthz`; per-tenant rule sets (`computeVerdict`, `currentRuleNames`); per-tenant local scorer and version; per-tenant fair queue and concurrency cap. Every built-in feature is available to every tenant. | P1 | Golden exact; tenant-isolation tests (rules, reload, view); fairness test | +| P3 | Declared types and kind redaction | `internal/vocab`; `internal/secret` (`Keys`, file adapter); field kinds; roles (`credential`/`other`, `activity`, `title`, `self`); `x_` extensions; the PSL-based domain kind; card/IP/phone scan; re-HMAC of all hash fields and links, with `join_domain`; undeclared values dropped; skeleton-only custom text; `vocab_version`; config history plus `abusekit config check` | P2 | Criterion 5 property tests; golden exact (features unchanged); `vocab_incompatible` and history CI tests | +| P4a | DSL core: `count`, `distinct`, `share`, `peak` | Compiler; windows (`window`, `first`, `lifetime`); predicates; transforms; caps and limits; the bounded loader (§5.7) and step budget; `core.history_truncated` with its truncation invariant; end-to-end benchmark and `cost_table.yaml` | P3 | Reference equality on 10k histories for these four ops; criterion 4 at P4a limits; criterion 6 flood property; loader fuzz | +| P4b | DSL extended: `time_between`, `before_first`, `relative_to_history`, `group_by`, `sequence`, `ratio` | Plus proportional rescore coalescing, the per-tenant rescore budget, and warm-up | P4a | Reference equality for every op; every "expressible" #7/S2 feature equals its Go twin bit for bit; warm-up test | +| P5 | Pack gating | `internal/pack` registry; `core`/`email`/`brand` adapters (code moved); enablement and dependency validation; `brand.title_match`; `packtest`; starter weights | P2 (P4a for `packtest`'s flood check) | Golden exact; `feature_not_enabled`; every pack passes `packtest` | +| P6a | Link kinds, `neighbours`, subject kinds | `links.custom`; `neighbours`; `subject_kind`, `also`, `parent`; `event_subjects`; `?kind=`; `applies_to`; derived `x_primary_subject_hash` | P4a, P5 | Contract tests for kinds and `also`; neighbour caps; golden exact | +| P6b | Scenarios and bootstrap | The five `examples/tenants/*` profiles, with dev and held-out fixtures; uniform priors; `--profile`; corpus v2; floors with `profile:` | P4b, P6a | Criterion 2 on held-out fixtures; the held-out isolation CI check | +| P7 | e2a cutover | e2a's profile in the ops repo's private mount (outside this repo); hosted-config CI runs `abusekit config check`; the `rules.yaml` path removed | P5 (and S8's mount) | Golden exact against the private copy; the hosted deploy loads it | + +- **Ordering against the v0 plan.** P0 and P1 must land before S5 and S8, because assumption A1 + is what makes a rename without a bridge safe. +- **S3b (erasure)** must be vocabulary-aware (§5.10), and is easiest to build after P3. +- **S6** is independent of every P slice. ## 10. Scalability and extensibility -- **Custom features per tenant.** Bounded by 64 features and 256 cost units. Compiled plans are - cached per `profile_sha`, and a reload recompiles one tenant only. -- **Extraction cost.** One pass per subject: O(E) dispatch plus O(E) per windowed scan, with E - ≤ 50,000. `distinct` memory is O(cap). The existing per-subject budget and the priority queue - are unchanged. What grows is `EventsForSubject`'s load. Later the store can bound the query to - `max(lookback, window) + exclude_recent`, plus `subjects.first_seen` for `start`. That is a - store-only change the pack interface already permits, because `Input.Start` is explicit. -- **Tenants.** About 100 tenants × 64 features is a config and plan-cache concern, not a - database one. Metrics are labelled `{tenant, pack}`, not `{feature}`, to bound cardinality. -- **Made easier later.** - - A new domain pack (`sms`, `marketplace`) is one Go package plus `packtest`. - - A popular custom feature can be promoted into a pack upstream, keeping its name through a - per-pack alias. - - Cross-tenant linking is untouched by this design. - - Weight fitting consumes corpus-v2 per profile. - - CEL for `where` alone, if ever needed, slots in as a new leaf kind. +- **Per-subject cost** is bounded by bytes and steps, not by event count. The loader fetches only + as far back as the profile's largest lookback, capped at 8 MiB of history plus 1 MiB of + onboarding events. Compiled plans are cached per `profile_sha`. +- **Tenants.** The fair queue and the per-tenant concurrency cap let about 100 tenants share one + worker pool without starving each other. Metrics are labelled `{tenant, pack}`. +- **Rescores.** + - Timer rescores come only from features that feed non-shadow rules. + - Their coalescing buckets scale with the feature's window. + - Each tenant has an hourly budget. +- **Neighbours.** At most 4 queries per extraction; each is indexed, fan-in capped, and cached for + the extraction. +- **Made easier later:** + - new Go packs, gated by `packtest`; + - promoting a popular custom feature into a pack upstream; + - weight fitting on corpus v2; + - CEL as a `where` leaf; + - cross-tenant linking, which this design leaves untouched. ## 11. Verification strategy -The seams tested are the ones callers cross: the config loader (profiles in, errors out), -`feature.Extract` (events in, vector out), `POST /v1/events` (redaction), and the harness -(`--profile`). - -1. The golden replay and the translation test (§8), in every slice's CI. -2. `packtest` for every pack (§5.9). -3. The DSL conformance suite: the naive reference versus the compiled evaluator on randomized - histories (shuffled order, future events, ties, empty windows, anchored windows before - `start`); table tests for every predicate leaf and each op's edge (window boundary - inclusivity, `if_absent`, `if_empty`, cap and log1p order, `distinct` early stop). -4. Loader tests for every §5.8 code, plus the fuzzer, with limits as the oracle. -5. Redaction property tests: mask versus reject per kind; undeclared fields hashed; no email - shape stored outside masked text. -6. HTTP contract tests: `delivery.sent`, declared types, `strict` mode, per-tenant `/healthz`. -7. A benchmark at the limits (criterion 4). -8. **Most likely regressions:** summation order (caught by golden), a quantum lost for the age - features (golden input hashes), the S2b alias list drifting in the vocabulary move (golden - `core.credential_*`), and window-boundary off-by-one in a DSL op (conformance suite). -9. **Manual checks:** load each §7 profile in a local instance, post its sample events, and read - the shadow signals and reasons. +Tests sit at the seams callers actually cross: +- the profile loader: profiles in, errors out; +- `LoadHistory` + `feature.Extract`: history in, vector out; +- `POST /v1/events`: redaction; +- `abusekit eval --profile`. + +The checks: +1. **Golden replay.** Criterion 1 in P1, then exact equality in every later slice. +2. **`packtest`** for every pack (§5.9). +3. **DSL conformance.** The naive reference against the compiled evaluator. Table tests for each + op's edge cases, clipping, `before_first` inclusivity, `track_max`, and the group and key caps. +4. **Redaction property tests.** + - Mask versus reject, by kind. + - Re-HMAC, including `join_domain` equality and separation. + - Undeclared values dropped. + - Domain PSL and IP checks. + - Luhn, IP and phone detection, with explicit false-positive fixtures. +5. **Loader fuzzer**, with the §5.7 limits as the oracle, plus config-history CI tests. +6. **Performance and flooding.** The end-to-end benchmark (criterion 4) and the flood property + (criterion 6). +7. **Tenant isolation.** Rules, reload, view, the fair queue and the rescore budget. +8. **HTTP contract tests.** `subject_kind`, `also`, `links.custom`, `?kind=` and `feature_renamed`. +9. **Most likely regressions, and what catches each:** + +| Regression | Caught by | +| --- | --- | +| A stage lookup missed by the rename | grep test + stage test | +| A lost hash quantum | hash-drift test | +| Tie order | golden | +| A weights edit that breaks the truncation invariant | `packtest` | +| PSL snapshot drift | pinned version + test | ## 12. Open questions (owner decisions) -1. **Ordering.** Land G0–G3 before S5 (vendor adapters) and S8 (hosted deploy)? Recommended: - yes. -2. **What S6 emits.** `content.sent` as designed (recommended, no churn), or the neutral - `delivery.sent{channel: email}`? -3. **Undeclared-type storage change.** Approve moving unknown types from "kept as-is" to - "strings keyed-hashed" (a pre-GA semantic change, §5.4 and §5.6)? -4. **abusekit-held per-tenant redaction key.** This is a new secret per tenant, separate from - the producer's link key. Approve? -5. **Legacy window quirks.** `core.resource_total` and `core.credential_total` count - future-dated events, and `email.first_day_distinct_domains` has an inclusive end. Keep them - frozen in `@1` and harmonise in `@2` later (recommended), or harmonise now and accept a golden - diff? -6. **The `credential` rename.** `key_*` → `core.credential_*`: accept the neutral role vocabulary - (`credential`, `identity`, `workspace`, `content`, `other`)? -7. **Limits.** 64 features, 256 units, 30-day max window, 50,000 events scanned, 250 ms deadline. - Confirm or adjust. -8. **Bootstrap policy.** Uniform priors are shadow-only, with no fitting in scope. Confirm that - advise always needs a hand-set or fitted weights file plus a passing gate. -9. **CEL fallback.** Pre-approve CEL for `where` only if the closed predicate set proves - insufficient, or require a new design pass? -10. **Brand list scope.** Is `config/packs/brand/brands.yaml` one global public list, or may a - tenant narrow it (`pack_params.brand.list`) as well as extend it (`extra`)? -11. **The custom namespace.** `custom.*` scoped per tenant (proposed), or `.*` so names - are globally unique in logs and corpora? -12. **Per-tenant reload isolation.** Replaces the main design's whole-config rejection. Confirm. +Where the review's answer differs from revision 1's recommendation, both are shown. + +1. **Bridge vs rename.** Revision 1: a frozen canonical-key bridge. Review, and now recommended: + rename once in P1 and re-baseline under semantic identity. Reason: nothing stored needs a + bridge yet (A1), and a bridge would keep two names alive forever. Approve? +2. **Ordering.** Revision 1: G0–G3 before S5 and S8. Now: P0 and P1 **must** land before S5 and + S8, because the no-bridge rename depends on it. Approve? +3. **Undeclared data.** Revision 1: keyed-hash undeclared strings. Review, and now: drop every + undeclared value and keep only type, time and field names. Reason: the hash of an unreviewed + field is still personal data. Approve? +4. **Built-in re-HMAC.** Re-HMAC `recipient_hash` and every `links` value at ingest. Stored bytes + change; feature values don't. Approve? +5. **Legacy window quirks.** Freeze them in `@1` and harmonise in `@2` (recommendation unchanged). + Confirm? +6. **Roles.** Revision 1: five resource roles. Review, and now: only `credential` and `other`, + plus the `activity` type role and the `title`/`self` field roles. Approve? +7. **Load bounds.** Revision 1: the 50,000 newest events and a 250 ms deadline. Review, and now: a + time-bounded load, onboarding types in full, byte caps of 8 MiB and 1 MiB, a calibrated step + budget, and truncation as a positive signal. Confirm the defaults? +8. **Bootstrap.** Uniform priors, shadow only, no fitting, and held-out fixtures. Confirm? +9. **CEL.** Add it later as a `where` leaf only, or require a new design pass? (Unchanged.) +10. **Brand list.** May a tenant narrow the list as well as extend it? (Unchanged.) +11. **Custom namespace.** `custom.*` per tenant, or `.*`? (Unchanged.) +12. **Config history.** Revision 1: history held in the DB. Review, and now: history in the config + tree, checked in CI, with the DB as a runtime guard only; atomic reload per tenant. Approve? +13. **What S6 emits.** Revision 1 offered a choice between `content.sent` and `delivery.sent`. + Review, and now: `delivery.sent` is dropped, so S6 emits `content.sent`. Settled unless you + object. +14. **Domain kind.** Require a PSL suffix or an RFC 6761 special-use name, and reject IP literals + and all-numeric labels, **including on built-in domain fields** (`RedactionSchemaVersion` 3). + Declared fields also get an optional `reduce: etld1`. Approve? +15. **Feature interactions.** Should `ratio` gain a `product` form (depth-1 DAG, capped) for the + interactions that §7e shows uniform priors can't capture, or should that wait for fitted + weights? +16. **Subject kinds and `also`.** Add the wire fields `subject_kind` and `also` (at most 3), with + subjects keyed `(tenant, kind, id)`. Both are additive. Approve? +17. **Stage-gate fix.** Stage gates consider only advise-mode local rules. This changes behaviour + on `main`. Approve? +18. **Profiles and secrets.** Tenant profiles live in a private mount, with only fictional + examples in the repo, and the key interface is provider-agnostic. Approve? +19. **Rescore control.** Proportional coalescing applies to DSL features only (built-ins keep + 5 minutes for semantic identity). Only non-shadow rules schedule timer rescores, under a + per-tenant hourly budget. Confirm the default of 20 × active subjects per hour? + +## 13. Changes from revision 1 + +**Blockers:** +- **B1:** the canonical-key bridge is dropped. The rename happens once, with a test for every + consumer it touches (§5.2), and the golden checks semantic identity against a derived + floating-point bound. +- **B2:** the event-count bound and the wall-clock deadline are replaced (§5.7) by: + - a time-bounded load, with onboarding types loaded in full; + - byte caps and a deterministic step budget; + - truncation as a positively weighted signal, with a checked invariant and an argument that + flooding can't evade. +- **B3 (§5.6):** + - every hash is re-HMACed at ingest, with length prefixes and `join_domain`; + - undeclared values are dropped; + - domains are checked against the PSL and rejected if they are IP literals; + - the leak scan covers card, IP and phone shapes; + - custom text is stored skeleton-only; + - stored values are described as pseudonymised throughout. +- **B4:** new primitives `group_by`, `sequence`, `ratio`, `neighbours` with declared link kinds, + `before_first`, `anchor: last`, and subject kinds with `also` and `parent`. All five scenarios + are walked, and what remains Go-only is stated (§7). + +**Should-fix:** +- `delivery.sent` is dropped in favour of declared types with `title`/`self` roles and an + `activity` type role; `x_` extension fields are added. +- Warm-up, and dual-key rotation. +- Exact `relative_to_history` equations, with a list of what the DSL can't express. +- A benchmark-calibrated cost table and a per-type fan-out cap. +- Scheduling: rescore-storm control, a fair queue, and stage gates limited to advise-mode rules. +- Config history in the config tree; profiles in a private mount. +- `(at, producer, id)` tie-breaks; held-out fixtures; `PriorSign` on `FeatureDef`; per-tenant + rule names. +- Re-slicing into P0 → P7; S3b erasure made vocabulary-aware. + +**Nits:** +- Only the `credential`/`other` resource roles. +- A provider-agnostic `Keys` interface. +- A per-tenant scorer version. +- `peak` clipping specified. +- Uniform-prior normalisation written out. diff --git a/docs/plans/2026-09-27-v0-plan.md b/docs/plans/2026-09-27-v0-plan.md index 85471e9..26e5936 100644 --- a/docs/plans/2026-09-27-v0-plan.md +++ b/docs/plans/2026-09-27-v0-plan.md @@ -35,7 +35,7 @@ design pass) rather than deferred to v1 outright. | S7 | Billing events | ops sidecar: `payment.attempt` (with `card_fingerprint_hash` under the tenant key) and `subscription.changed` | staging checkout produces events | | S8 | Hosted deploy | ops: compose service, Secret Manager keys, tenant config, Terraform alerts for queue depth / budget / drops | abusekit running on prod in shadow | | S9 | Incident evaluation | private backfill of the incident accounts and a benign sample into a private corpus; harness run; report precision/recall/lead-time before first send per account | report reviewed; floors set; decision on `advise` for the local rule | -| G0–G7 | Generic feature packs (own design: `docs/design/2026-09-29-generic-feature-packs.md`) | After S4 (#5) and S2b (#7) merge: golden capture (G0), canonical feature keys (G1), neutral vocabulary and delivery view (G2), packs and tenant profiles (G3), declarative custom features (G4), declared-type redaction (G5), per-pack eval and bootstrap (G6), explicit e2a profile (G7). Recommended before S5 and S8. | per-slice "Done when" in that design's §9; every slice keeps e2a's golden replay bit-identical | +| P0–P7 | Generic feature packs (own design: `docs/design/2026-09-29-generic-feature-packs.md`) | After S4 (#5) and S2b (#7) merge: golden replay via `abusekit eval --golden` (P0), one-time namespaced rename (P1), tenant profiles (P2), declared types and kind redaction (P3), DSL core (P4a) and extended ops (P4b), pack gating (P5), link/subject kinds (P6a), example scenarios and bootstrap (P6b), e2a private-profile cutover (P7). P0 and P1 must land before S5 and S8; S3b erasure must be vocabulary-aware. | per-slice "Done when" in that design's §9; P1 re-baselines the golden under semantic identity, every later slice keeps it exact | S1–S4 (including S3b) are pure abusekit and can run back to back; S5 needs vendor keys; S6–S8 are e2a/ops work that can start after S3 (S3b is not a blocker for them — nothing in S6–S8 depends on From 427d462db1e5b3f217ab2786aac10db44e0523ba Mon Sep 17 00:00:00 2001 From: jiashuoz Date: Tue, 29 Sep 2026 12:02:10 +0800 Subject: [PATCH 3/8] docs(design): generic feature packs revision 3 after re-review Onboarding from ingest-maintained subject_facts and lifetime totals from subject_counters; evaluation classes F/N/A/R/G/D with a fixed pass order and per-feature/per-pack budgets; exact per-feature aggregates; truncation becomes a one-sided `partial` flag, not a weighted feature; flood property restated against an unbounded reference with a specified generator. Rename is bit-exact via registry-order summation; fake scorer and corpus loader covered. Redaction: author-trusted bounded numbers, domain eTLD+1 with allowlist-or-HMAC, name grammar for undeclared fields, pseudonymised non-account ids, HKDF per-tenant keys, egress scan. DSL: absence indicators, pre-transform ratio, exact group_by, hash_quantum, as-of neighbours. Slices re-split (P1s, P3a-d, P4a-d, P5b). Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_014cdM7WyRc3mD3vQNXMTDB8 --- docs/design/2026-09-27-abusekit-design.md | 4 +- .../2026-09-29-generic-feature-packs.md | 2057 +++++++++-------- docs/plans/2026-09-27-v0-plan.md | 2 +- 3 files changed, 1109 insertions(+), 954 deletions(-) diff --git a/docs/design/2026-09-27-abusekit-design.md b/docs/design/2026-09-27-abusekit-design.md index d11946e..47ef66c 100644 --- a/docs/design/2026-09-27-abusekit-design.md +++ b/docs/design/2026-09-27-abusekit-design.md @@ -6,11 +6,11 @@ see account churn, tiers failed open when a scorer was unavailable, and request cover reads or replay. This revision fixes those and tightens every place an implementer would have had to guess. Changes from r1 are marked **[r2]**. -**Amendment (proposed 2026-09-29, revision 2):** [`2026-09-29-generic-feature-packs.md`](2026-09-29-generic-feature-packs.md) +**Amendment (proposed 2026-09-29, revision 3):** [`2026-09-29-generic-feature-packs.md`](2026-09-29-generic-feature-packs.md) renames the built-in features once into namespaced `core`/`email`/`brand` packs enabled per tenant, adds product-declared vocabularies (types, field kinds, link kinds, subject kinds) with pseudonymising redaction, and adds bounded declarative custom features in YAML. It amends §4.2, §4.3, §4.5, §4.6, §4.8 -and §4.10 below, and keeps e2a's feature values and tiers identical. +and §4.10 below, and keeps e2a's feature values, risks and tiers bit-identical. ## 1. Problem statement diff --git a/docs/design/2026-09-29-generic-feature-packs.md b/docs/design/2026-09-29-generic-feature-packs.md index 31ee3d9..b903576 100644 --- a/docs/design/2026-09-29-generic-feature-packs.md +++ b/docs/design/2026-09-29-generic-feature-packs.md @@ -1,842 +1,960 @@ # Generic feature packs, product vocabularies, and declarative custom features -Status: proposed, revision 2 (after adversarial review), 2026-09-29 · owner: Josh Zhang · amends -[`2026-09-27-abusekit-design.md`](2026-09-27-abusekit-design.md) §4.2, §4.3, §4.5, §4.6, §4.8 and -§4.10. Written against `main` at S3, with the two open PRs treated as merged: #5 (S4 evaluation -harness) and #7 (S2b send-volume, webmail, recipient and subject-brand features). `main §x` refers -to the main design. §13 lists what changed from revision 1. +Status: proposed, revision 3, 2026-09-29. This revision answers the re-review of revision 2, which +returned "approve after listed changes". Owner: Josh Zhang. + +This document amends [`2026-09-27-abusekit-design.md`](2026-09-27-abusekit-design.md), §4.2, +§4.3, §4.5, §4.6, §4.8 and §4.10. It is written against `main` at S3, treating two open PRs as +merged: #5 (the S4 evaluation harness) and #7 (the S2b send-volume, webmail, recipient and +subject-brand features). `main §x` refers to the main design. §13 lists what changed in each +revision. ## 1. Problem statement -abusekit's current feature set assumes an email platform. The main design promises scoring for -any product that mints accounts, takes payments and lets users create resources. The built code -falls short of that: - -- **The vocabulary is email-shaped.** `content.sent` carries `subject_line`, `recipient_domain` and - `recipient_is_own_identity`. Resource kinds are an undeclared convention: `internal/feature` - counts `kind == "key"` (plus #7's spelling aliases) and treats every other kind as generic. -- **Much of the feature set is email-specific.** Nine of the 25 features on `main` + #7 are email - features, and two more assume API keys. A file-sharing, payments, chat, developer-API or - AI-inference product would get these at zero. -- **Custom event types are dead weight.** Main §4.3 says unknown types are stored "for - Go-registered features only". No such feature exists, so the only way to add a signal is to - write Go in this repo. -- **Everything is global and flat.** There is one `rules.yaml`, one `local_weights.yaml`, one - feature namespace, and one subject kind: the account. - -**Desired outcome.** A product that is not an email platform onboards with a private YAML profile. -In that profile it: -- declares its event types, field kinds, link kinds and subject kinds; -- enables the packs that fit its domain; -- defines its product-specific signals as declarative features; -- runs them in shadow. - -Every new value is pseudonymised or dropped at ingest. For e2a, feature values, tiers and rescore -times stay identical, and risk moves by no more than a stated floating-point bound. +abusekit's current feature set assumes an email platform. The main design promises scoring for any +product that mints accounts, takes payments and lets users create resources. The code built so far +doesn't deliver that: +- **Email-shaped vocabulary.** `content.sent` carries `subject_line`, `recipient_domain` and + `recipient_is_own_identity`. Resource kinds are an undeclared convention. +- **Email-shaped features.** Of the 25 features on `main` + #7, nine are email features and two + more assume API keys. +- **Custom types go unused.** No feature reads custom types, so a new signal means writing Go in + this repo. +- **Global, flat configuration.** There is one `rules.yaml`, one `local_weights.yaml`, one feature + namespace and one subject kind. + +**Desired outcome.** A product that isn't an email platform onboards with a private YAML profile. +In it, the product: +1. declares its event types, field kinds, link kinds and subject kinds; +2. enables the packs that fit its domain; +3. defines product-specific signals as declarative features; +4. runs them in shadow. + +Three properties hold throughout: +- Every new value is pseudonymised or dropped at ingest. +- Every feature is either exact or explicitly flagged `partial`. +- Flooding a subject with events can't lower its risk. + +For e2a, feature values, risks, tiers and rescore times stay bit-identical. ### Success criteria (measurable) -1. **Semantically identical migration.** A golden replay covers every committed fixture: all of - `eval/fixtures/*.jsonl` (including #7's) and the synthetic corpus. It scores after every event - and at every scheduled rescore instant. Across the rename and every later slice: - - every feature value is identical as `math.Float64bits`, compared under the rename map; - - every `NextRescoreAt` and every tier is identical; - - every risk and score satisfies `|Δ| ≤ 1e-12` (§5.2 derives the bound). - - Input hashes, cassette keys, the `run.json` rule/weights SHAs and the local `Version()` may - change **only in the rename slice**. The golden records the before and after values. -2. **Five generic scenarios.** Each scenario in §7 loads as a fictional profile, with zero Go - changes for everything §7 marks as declarative. Held-out fixtures are authored after the - features and never used while writing them. On those fixtures, with uniform priors (§5.9), - each scenario's abusive fixtures outrank every one of its benign fixtures. -3. **Enablement is enforced.** For a tenant without a pack, none of that pack's features is - computed, stored or rendered. A rule that references one fails to load with - `feature_not_enabled`. -4. **The DSL is correct and bounded.** - - Every operation equals a naive reference implementation on 10,000 randomized histories, - including shuffled arrival, ties and future-dated events. - - Per-subject time is measured **end to end**: store load, JSON decode, view projection, pack - extraction and orchestration. For a profile at the §5.7 limits with history at its byte cap, - it is p99 ≤ 50 ms on the reference 2-vCPU host. - - The step budget is calibrated from that benchmark (§5.7), so the budget cuts off work before - the latency target is breached. -5. **Privacy by construction.** Property tests show that, at rest: - - no stored value matches an email, card (Luhn), IP or phone shape outside a masked `text` - field; - - every stored `hash` value is an abusekit-keyed HMAC; - - undeclared data keeps only type, time and field names. - - The loader fuzzer finds no profile that passes validation while breaking a §5.7 limit. -6. **Flooding doesn't evade.** For every built-in and DSL feature, a property test adds cheap - events totalling up to 10 times the byte cap to an abusive fixture. Risk must not fall (§5.7). +1. **Bit-exact migration.** A golden replay covers every committed fixture, #7's included, plus the + synthetic corpus. It scores after every event and at every scheduled rescore instant. + - Across the rename and every later slice, these stay identical under the rename map, compared + with `math.Float64bits`: every feature value, `NextRescoreAt`, risk, score and tier. + - These may change, and only in the rename slice: input hashes, cassette keys, fake-scorer + probabilities, the `run.json` rule and weights SHAs, and the local `Version()`. The golden + records their before and after values. +2. **Five generic scenarios.** + - Each scenario in §7 loads as a fictional profile, with zero Go changes for everything §7 + marks as declarative. + - Held-out fixtures are authored only after the features are frozen. + - On those fixtures, with uniform priors, every abusive fixture outranks every benign fixture of + the same scenario. +3. **Enablement enforced.** Features from a pack that isn't enabled are never computed, stored or + rendered. A rule that references one fails with `feature_not_enabled`. +4. **Correct, independent, bounded.** + - Every DSL op equals a naive reference implementation on 10,000 randomized histories. The + histories include shuffled arrival, ties, future-dated events and floods. + - Changing or removing one feature never changes another feature's value for the same subject. + - End-to-end p99 is at most 50 ms per subject at the §5.7 limits. +5. **Privacy by construction.** + - At rest, nothing matches an email, card (Luhn), IP or phone shape, even after NFKC and + Unicode-digit folding. This is checked on the serialised row before insert. The only + exception is a masked marker inside a text field. + - Every hash, every non-allowlisted domain and every non-account subject id is keyed per tenant. + - No two tenants share a key. +6. **Flooding can't lower risk.** + - `risk(flooded, bounded evaluator) ≥ risk(flooded, unbounded reference)` holds for every + fixture under the §5.7 flood generator. + - A flood of 10,000 tiny `blocked` payments moves no onboarding feature. ## 2. Goals and non-goals **Goals** -- Rename the built-in features into namespaced names **once**, while no production verdicts or - vendor cassettes exist. There is no compatibility bridge (§5.2). -- Per-tenant profiles, kept in a private config mount. Only fictional example profiles live in - this repo. +- A one-time rename to namespaced features that keeps the golden replay bit-exact (§5.2). +- Per-tenant profiles in a private config mount. Only fictional examples live in this repo. - Product-declared vocabularies: - - custom event types, with a field kind and optional role for each field; - - extension fields on built-in types; - - an `activity` role for types; - - resource-kind roles; - - declared link kinds; - - subject kinds beyond `account`. -- Redaction: - - pseudonymise every declared hash with an abusekit-held key; - - drop every undeclared value; - - validate domains against the public suffix list (PSL); - - mask text or store it as a skeleton; - - scan for card, IP and phone shapes. -- A closed declarative feature DSL covering the five scenarios in §7. Every feature has a - mandatory cap, deterministic semantics and a deterministic step budget. -- Packs (`core`, `email`, `brand`) enabled per tenant, each with starter weights, fixtures, floors - and a `packtest` harness. Uniform priors, shadow-only, until a tenant has labels. + - types, with field kinds and roles; + - `x_` extension fields; + - the `activity` role; + - resource roles; + - link kinds; + - subject kinds. +- Pseudonymising redaction: + - re-HMAC under keys derived per tenant; + - domain allowlist or HMAC; + - drop undeclared data; + - bounded numbers; + - masked or skeleton-only text; + - an egress scan. +- A closed DSL covering five scenarios, with mandatory caps and per-feature budgets. +- Evaluation over facts and counters maintained at ingest, plus exact per-feature aggregates. + Hitting a bound produces a flag, never a weight. +- Packs, each with starter weights, fixtures, floors and a `packtest`. Uniform priors in shadow for + tenants that have no labels yet. **Non-goals** -- Code or expressions in config: no CEL, regex, WASM or plugins (§5.12). -- Weight fitting (§12 Q8) and cross-tenant linking (main §2). -- Changes to `Plan`/`Combine` math, apart from the stage-gate fix (§5.8) and per-tenant rule sets. -- Runtime vocabulary declaration over HTTP (§5.11). -- A neutral built-in "delivery" event. Revision 1 proposed `delivery.sent`; revision 2 drops it - (§5.4, §13). +- Code or expressions in config (§5.12). +- Weight fitting (§12 Q8). +- Cross-tenant linking. +- Changes to `Plan`/`Combine` math, other than the stage-gate fix (slice P1s) and per-tenant rule + sets. +- Declaring vocabularies at runtime. +- A built-in "delivery" event (dropped in revision 2). ## 3. Relevant context and constraints -**Code this design touches (`main` + #5 + #7):** -- `internal/event`: - - `Validate`, and a static `schema` for redaction (`RedactionSchemaVersion` is 2 after #7). - - Unknown types keep every key after a leak scan; unlisted keys of known types are dropped. -- `internal/feature`: - - a monolithic `Extract` returns a fixed 25-field `Features`, with `Map()` and `Names`; - - `nextRescoreAt` hard-codes `resource.created` and `content.sent`; - - `firstSeenAt` is the minimum event `at`. -- `internal/core`: - - `inputHash` JSON-encodes the rule's features keyed by name; - - `quantizeAgeFeaturesForHash` switches on the literals `subject_age_h` and `upgrade_delay_min`; - - `stageSkip` reads the literal `features["subject_age_h"]`; - - `maxRiskByScorer` takes the maximum over **every** local rule, shadow rules included. -- `internal/model/local`: sums in sorted feature-name order, and `Version()` hashes the whole - JSON-encoded `Weights` struct. -- `internal/worker`: - - `renderReason` prints flat feature names into every stored verdict reason; - - `computeVerdict` builds `currentRuleNames` from a global config. -- `internal/serve`: `currentRuleNames()` reads the global config. -- `internal/store`: - - `EventsForSubject` loads every event, ordered by `(at, seq)`; - - `corpus_examples.features` is JSON keyed by feature name; - - `subjects.first_seen_at` keeps `LEAST(existing, new)`. -- `eval` (#5): - - `hashScoreRequest` keys cassettes over `req.Features` by name; - - `run.json` records `dataset_sha`, `rule_sha` and `weights_sha`; - - corpus-v1 rows key features by name. - -**Assumptions** (unconfirmed ones repeat in §12): -- A1. S5 (vendor adapters) and S8 (hosted deploy) have not shipped. No production verdicts, corpus - rows or vendor cassettes exist yet. That is what makes a one-time rename safe, and why the rename - slice must land before S5 and S8. -- A2. At most about 100 tenants and at most 64 custom features per tenant. Event text retention is - 90 days (main §4.11). -- A3. Producers can compute a keyed hash of any identifier they want counted. abusekit re-hashes - it anyway (§5.6), so a producer's mistake can't turn into stored personal data. +This design touches the following code on `main` plus #5 and #7. + +- `internal/event` + - Static redaction schema, v2. + - Unknown types keep every key after the leak scan. + - ASN travels in the clear. +- `internal/feature` + - A monolithic `Extract` over *every* stored event. + - `nextRescoreAt` hard-codes the windowed types. + - `firstSeenAt` is the minimum `at`. +- `internal/core` + - `inputHash` keys features by name. + - `quantizeAgeFeaturesForHash` and `stageSkip` switch on literal feature names. + - `maxRiskByScorer` includes shadow rules. +- `internal/model/local` sums weights in sorted-name order. Its `Version()` hashes the `Weights` + struct. +- `internal/model/fake`: `deterministicProbs` hashes the sorted feature names and values. +- `internal/worker` / `internal/serve` + - `renderReason` prints flat names. + - `computeVerdict` and `currentRuleNames()` read one global config. +- `internal/store` + - `EventsForSubject` loads everything, ordered by `(at, seq)`. + - `corpus_examples.features` is keyed by name. + - `subjects.first_seen_at` uses `LEAST`. +- `eval` (#5) + - `LoadSnapshotCorpus` reads corpus-v1. + - `hashScoreRequest` keys cassettes by feature names. + - `run.json` records SHAs. + - `eval/neighbors.go` evaluates evidence as of the scoring instant. + +**Assumptions** +- A1. S5 and S8 haven't shipped, so no production verdicts, corpus rows or vendor cassettes exist + yet. +- A2. At most about 100 tenants and 64 custom features per tenant. Text is retained for 90 days. +- A3. Producers hash any identifier they want counted, and abusekit re-hashes it anyway. ## 4. Proposed design: overview ``` -private config mount: tenants//profile.yaml + history/NNNN.yaml (append-only, CI-checked) - packs · vocabulary (types, fields+kinds+roles, extensions, link kinds, subject kinds) - features (custom.*) · rules · weights +private config mount: tenants//profile.yaml + history/NNNN.yaml (append-only, CI-checked) │ compile + validate per tenant (atomic, isolated) ▼ -ingest ─▶ vocab.Redact(tenant): leak scan · kind rules · re-HMAC hashes · drop undeclared - ─▶ store: event row + vocab_version + (subject_kind, subject) index rows (≤ 4 per event) -worker (per-tenant fair queue) ─▶ loader: time-bounded + onboarding-full + byte cap - ─▶ feature.Extract(profile): core · email · brand (Go) · custom (compiled DSL), under a - deterministic step budget - ─▶ core.Plan / Combine (math unchanged) +ingest ─▶ vocab.Redact(tenant): leak scan · kinds · re-HMAC · domain allowlist/HMAC · drop undeclared + ─▶ store (one transaction): event row · event_subjects index rows (≤ 8) + · subject_facts upsert · subject_counters increments · dirty marks +worker (per-tenant fair queue) + ─▶ feature.Extract(profile), in pass order (§5.7): + F facts + counters (O(1)) + N anchored windows (frozen facts after close) + A exact per-feature aggregate queries + R row features with per-feature budgets (peak, relative baselines) + G Go-pack residue, per-pack budget + D derived (ratio, absent indicators) + ─▶ core.Plan / Combine (math unchanged) ─▶ verdict {…, partial: [features]} ``` | Module | Interface | Deletion test | | --- | --- | --- | -| `internal/vocab` (new) | `Compile(builtin, decl) (*Vocabulary, error)`; `(*Vocabulary).Redact(*event.Event, Keys) error` | Without it, redaction, kinds, roles and extensions spread across ingest and the packs. Keep. | -| `internal/pack` (new) | `Pack` interface + registry; adapters `core`, `email`, `brand` (Go) and `custom` (compiled DSL) | Four adapters, so the seam is real. Without it, `feature.Extract` stays a monolith. | -| `internal/pack/custom` (new) | `Compile(tenant, []Def, *Vocabulary, CostTable) (Pack, error)` | Holds every DSL semantic. Keep. | -| `internal/feature` | `Extract(ctx, profile, subject, History, now) (Result, error)` | A thin orchestrator. The worker, evaluate and eval all call this one function. | -| `internal/store` | `LoadHistory(ctx, tenant, subjectRef, LoadPlan) (History, error)` replaces `EventsForSubject` for scoring | Bounded loading lives in one place. | -| `internal/secret` (new) | `Keys` interface (§5.6), provider-agnostic | Two adapters (a file, a cloud secret manager). Keep. | -| `internal/config` | `LoadProfiles(mount, deps) (map[tenant]*Profile, map[tenant][]error)` | Absorbs `rules.yaml`. Per-tenant atomic reload. | +| `internal/vocab` (new) | `Compile(builtin, decl) (*Vocabulary, error)`; `Redact(*event.Event, Keys) (Stored, error)` | Without it, redaction, kinds, roles and allowlists spread across ingest and packs. | +| `internal/secret` (new) | `Keys` plus HKDF `Derive` (§5.6) | Two adapters: file and cloud secret manager. | +| `internal/facts` (new) | `Apply(tx, Stored)` at ingest; `Read(tenant, subject)` | Onboarding and lifetime values live in one place. | +| `internal/pack` (new) | `Pack` plus a registry; adapters `core`, `email`, `brand` (Go) and `custom` (DSL) | Four adapters, so the seam is real. | +| `internal/pack/custom` (new) | `Compile(tenant, []Def, *Vocabulary) (Pack, error)` | Holds every DSL semantic. | +| `internal/evalengine` (new) | `Aggregate(ctx, Query) (Result, error)`; Postgres and in-memory adapters | Two adapters (store and eval replay), proven equal by conformance tests. | +| `internal/feature` | `Extract(ctx, profile, subject, now) (Result, error)` | Thin orchestrator called by the worker, evaluate and eval. | +| `internal/config` | `LoadProfiles(mount, deps)` | Absorbs `rules.yaml`; reloads atomically per tenant. | ## 5. Proposed design: detail -### 5.1 Packs: registration, enablement, dependencies +### 5.1 Packs ```go type Pack interface { - ID() ID // {Name: "email", Version: 1} - Requires() []string // e.g. email@1 → [core] - Features() []FeatureDef - // Extract is pure: same input → same output, whatever the arrival order. Every step it - // takes is charged to in.Budget (§5.7). - Extract(ctx context.Context, in Input) (Output, error) + ID() ID // {Name: "email", Version: 1} + Requires() []string // e.g. email@1 → [core] + Features() []FeatureDef // in registry order + Extract(ctx context.Context, in Input) (Output, error) // pure; passes in §5.7 } type FeatureDef struct { Name string // "email.sends_10m_max"; grammar §5.3 + Order int // local-scorer summation order (§5.2), fixed forever + Class Class // F | N | A | R | G | D (§5.7) HashQuantum float64 // input-hash bucket (replaces the name switch in core) - Bound float64 // max value; used by uniform priors (§5.9) - PriorSign int8 // +1 / -1; uniform-prior and golden-sign default - TruncationDir int8 // -1 if the value can only fall when history is truncated, else 0/+1 (§5.7) - Reads []string // types/roles read: dispatch, load plan, rescore scheduling - Lookback Lookback // history the feature needs (§5.7 load plan) + Bound float64 // normalisation ceiling for uniform priors (§5.9); never a value cap + PriorSign int8 // +1 / -1 + Reads []string // types/roles read: dispatch, rescore scheduling, dirty marks RequiresPacks []string // e.g. email.subject_brand_match → [brand] - SubjectKinds []string // kinds this feature is defined for; default [account] -} - -type Input struct { - Tenant string - Subject SubjectRef // {Kind, ID} - History History // bounded, ordered (at, producer, id); truncation flags - Now time.Time - Start time.Time // subject.created.account_created_at, else subjects.first_seen_at - Neighbors NeighborResolver // store-backed, budgeted, cached per extraction - Params PackParams - Budget *StepBudget + SubjectKinds []string // default [account] } type Output struct { - Values map[string]float64 - Rescore []time.Time - Truncated bool // this pack's budget or the history's truncation was hit + Values map[string]float64 + Rescore []time.Time + Partial []string // features whose own budget was hit (§5.7) } ``` -**Enablement.** A profile lists `packs: [core@1, email@1, brand@1]`. The rules: -- `core` is always enabled; every other pack is opt-in. -- A pin names a major version, and a pack's semantics change only by shipping a new major version. -- `Requires` must hold, or the profile fails with `pack_requires`. -- A feature whose `RequiresPacks` or `SubjectKinds` don't match is not registered for that tenant +**Enablement** +- A profile lists its packs: `packs: [core@1, email@1, brand@1]`. +- `core` is always on. Other packs are opt-in. +- A pack's major version pins its semantics. +- Unmet `Requires` fails with `pack_requires`. +- A feature whose `RequiresPacks` or `SubjectKinds` don't match isn't registered for that tenant or kind. -- Packs run in the fixed order `core`, `email`, `brand`, `custom`. Each has its own namespace, so - two packs can't produce the same feature name. +- Each pack owns its namespace, so feature names can't collide. -**Pack contents after the one-time rename:** +**Pack contents.** The *Class* column says where each feature's inputs come from (§5.7). -| Pack | Features (old flat name → new name) | +| New name (old name) | Class | Source | +| --- | --- | --- | +| `core.subject_age_h` | F | `start` (§5.7) | +| `core.upgrade_delay_min`, `core.upgraded` | F | `subject_facts.first_paid_upgrade_at` | +| `core.declines_before_first_success` | F | `subject_facts.declines_before_first_success` | +| `core.first_funding_prepaid` | F | `subject_facts.first_success_funding` | +| `core.fingerprint_seen_on_other_subjects`, `core.linked_deleted_n`, `core.linked_labelled_abusive_n`, `core.neighbors_truncated` | F | neighbour evidence, built-in link kinds only (§5.5) | +| `core.resource_velocity_1h`, `core.credential_velocity_1h` (`key_velocity_1h`) | A | exact 1 h counts | +| `core.resource_total`, `core.credential_total` (`key_total`) | F | `subject_counters` (lifetime) | +| `core.burst_ratio_24h_vs_lifetime` | A + F | 24 h count (A) ÷ lifetime counters | +| `email.sends_1h`, `email.webmail_sends_1h`, `email.distinct_recipients_1h` | A + R | exact 1 h current value (A); history baseline (R) | +| `email.sends_10m_max` | R | `peak` with a saturation limit (exact); history baseline (R) | +| `email.sends_first_day`, `email.first_day_distinct_domains` | N | anchored; frozen into facts after day one | +| `email.webmail_recipient_share` | F | lifetime counters: webmail recipients vs all non-self recipients | +| `email.self_send_before_external` | F | `subject_facts.self_sends_before_first_external` | +| `email.subject_brand_match` (requires `brand`) | A + F | distinct ingest-matched brand ids in 1 h, minus the exempt set held in facts | +| `brand.name_match`, `brand.name_has_at` | F | `subject_facts.name_brands`, `name_has_at` (matched at ingest) | +| `brand.title_match` (new) | A | distinct brand ids across non-self `title` fields in 1 h | +| `core.email_domain_class_disposable`, `core.verdict_max_24h` (new) | F, A | not part of e2a's rule | + +For the rows above, the code changes from scanning every event to reading facts, counters and +exact aggregates. The arithmetic doesn't change: it operates on the same inputs and sums integers, +which float64 represents exactly below 2⁵³. The golden replay proves this. + +**Bound** is only a normalisation ceiling for uniform priors, which clamp to it. It never caps a +value. + +| Feature | Bound | | --- | --- | -| `core@1`, account | `subject_age_h`→`core.subject_age_h` | -| `core@1`, payment | `upgrade_delay_min`→`core.upgrade_delay_min`, `upgraded`→`core.upgraded`, `declines_before_first_success`→`core.declines_before_first_success`, `first_funding_prepaid`→`core.first_funding_prepaid`, `fingerprint_seen_on_other_subjects`→`core.fingerprint_seen_on_other_subjects` | -| `core@1`, resource velocity | `resource_velocity_1h`→`core.resource_velocity_1h`, `resource_total`→`core.resource_total`, `key_velocity_1h`→`core.credential_velocity_1h`, `key_total`→`core.credential_total` | -| `core@1`, linked subjects | `linked_deleted_n`→`core.linked_deleted_n`, `neighbors_truncated`→`core.neighbors_truncated` | -| `core@1`, labels | `linked_labelled_abusive_n`→`core.linked_labelled_abusive_n` | -| `core@1`, behaviour change | `burst_ratio_24h_vs_lifetime`→`core.burst_ratio_24h_vs_lifetime` (counts `resource.created` + `content.sent` + declared `activity` types, §5.4); new `core.history_truncated` (§5.7) | -| `email@1`, requires `core` | `sends_10m_max`, `sends_1h`, `sends_first_day`, `webmail_recipient_share`, `webmail_sends_1h`, `distinct_recipients_1h`, `first_day_distinct_domains`, `self_send_before_external`, `subject_brand_match` (also requires `brand`), each becoming `email.`. The pack owns the built-in `content.sent` type. | -| `brand@1` | `name_brand_match`→`brand.name_match`, `name_has_at`→`brand.name_has_at`; new `brand.title_match` over any declared field with `role: title` (§5.4) | +| `core.resource_velocity_1h`, `core.credential_velocity_1h` | 50 | +| `core.resource_total`, `core.credential_total` | 1,000 | +| `core.linked_labelled_abusive_n` | 10 | +| `core.subject_age_h` | 24 (already clamped) | +| `core.upgrade_delay_min` | 1,440 (already clamped) | -In the early slices every built-in feature is available to every tenant (P2, §9). Pack **gating** -comes later, in P5. +Revision 2's 1,000 value cap on totals is dropped, because totals now come from counters. -New `core` features are additive and are not in e2a's rule: -- `core.email_domain_class_disposable` (deferred in main §8 Q7); -- `core.verdict_max_24h`; -- `core.history_truncated`. +### 5.2 The one-time rename, bit-exact -### 5.2 The one-time rename (replaces revision 1's canonical-key bridge) +The rename happens once, in P1. No alias table outlives that slice. -**Decision.** Rename in one slice (P1) and re-baseline the golden replay once. No alias table -survives past that slice. The bridge existed only to keep stored verdict hashes and vendor -cassettes stable. Under A1 there are none, and a bridge would keep two names for everything alive -forever. +**Summation order doesn't change.** The local scorer used to sum in sorted feature-name order, +which a rename would reshuffle. It now sums in `FeatureDef.Order`: +- legacy features keep the index of their old flat name in the old sorted order (0–24); +- new pack features take indices from 100 up; +- custom features follow, in sorted-name order. -**Every consumer the rename touches, with its test:** +The additions happen in exactly the same sequence as before, so risks and tiers are bit-exact. +Revision 2's 1e-12 tolerance and cut-point proximity check are deleted. + +**Every consumer of feature names, and its test** | Consumer | Change | Test | | --- | --- | --- | -| `core.stageSkip`, which reads the literal `features["subject_age_h"]` | Reads `core.subject_age_h` instead. The `Rule.Stage` **keys** (`max_subject_age_h`, `min_subject_age_h`) are stage names, not feature names, and keep their names. | Table test: a staged rule skips at the age bounds with a namespaced vector. A grep test fails if any `core` source still contains a flat feature literal. | -| `core.quantizeAgeFeaturesForHash`, which switches on names | Deleted. Quantization comes from `FeatureDef.HashQuantum` (1 h for `core.subject_age_h`, 60 min for `core.upgrade_delay_min`), passed to `Plan` in `core.Vector`. | The existing "hash stable under sub-hour drift" tests, re-run with namespaced names. | -| `core.inputHash` | Changes once (the names inside it change). | Golden records both hashes. A test asserts every rule's hash changes in P1 and is stable in every later slice. | -| Eval cassette key (`hashScoreRequest` over `req.Features`) | Changes once. P1 re-records every committed cassette (fake and local only, per A1) and stamps `feature_key_space: ns-v1` in the cassette header. | Loading a cassette with a mismatched `feature_key_space` fails with a clear error rather than silently missing. | -| `run.json`: `dataset_sha`, `rule_sha`, `weights_sha` | `rule_sha` and `weights_sha` change once. `dataset_sha` changes only for corpus-v1 snapshot files, which P1 rewrites to v2. The manifest gains `feature_key_space`. | Golden compares metrics, not SHAs. A test checks `feature_key_space: ns-v1` is in the manifest. | -| `worker.renderReason` and stored verdict reasons | Prints namespaced names. Verdict rows gain `reason_version: 2`. Old stored reasons are history and are not rewritten. | A reason snapshot test, plus a test that the verdict row carries `reason_version`. | -| `corpus_examples.features` | New column `feature_key_space text NOT NULL DEFAULT 'flat-v0'`. P1's migration rewrites the JSON keys of existing rows (none in production, per A1; this touches local dev databases only) and sets `ns-v1`. Once migrated, the corpus loader refuses `flat-v0`. | Migration test on a seeded DB: keys rewritten and marker set. A second test checks a `flat-v0` row is refused. | -| `local.Version()`, which hashes the whole `Weights` struct | Changes once. The weights file's `version:` is bumped to `v2`. | `Version()` differs from the pre-rename golden value and is stable afterwards. | -| `feature.Names`, `Features.Map`, `config.FeatureSet`, `config/rules.yaml`, `config/local_weights.yaml`, `mutation_test.go`, `ablation_test.go`, golden-sign tables, `eval/fixtures/README.md` | Mechanical rename. | The full existing suite passes with renamed literals, and the grep test. | -| `abusekit score --jsonl` input rows (an external contract) | A row keyed by an old flat name is rejected with `feature_renamed`, and the error names the new name. It fails loudly and never translates silently. | CLI contract test. | - -**The floating-point bound.** The local scorer sums weight × feature in sorted feature-name order. -Renaming the features changes the sort order, and with it the order of the additions. For `n` -terms, recursive summation error is at most `(n−1)·u·Σ|wᵢxᵢ|` with `u = 2⁻⁵³`. For e2a's rule: -- `n = 25` and `Σ|wᵢxᵢ| < 100` on any bounded vector, so `|Δlinear| < 2.7e-13`; -- the sigmoid's slope is at most ¼, so `|Δrisk| < 6.7e-14`. - -The golden asserts `|Δrisk| ≤ 1e-12`, which leaves margin. So that tiers are provably unchanged, it -also asserts that no recorded score lies within `1e-12` of a tier cut point or a rule threshold. If -one ever does, the fixture is flagged rather than passing silently. - -### 5.3 Name grammar and reserved namespaces - -- **Features** match `^[a-z][a-z0-9]*\.[a-z][a-z0-9_]{0,55}$`. The namespaces `core`, `email`, - `brand` and `custom` are reserved, as is any future pack name. Tenants define only `custom.*`, - which is scoped to the tenant. -- **Event types** keep `^[a-z_.]+$`, are at most 64 bytes, and must contain a dot. The prefixes - `subject.`, `payment.`, `subscription.`, `resource.`, `content.` and `abusekit.` are reserved. -- **Extension fields** on built-in types must be named `x_`, so they can never collide with - a future built-in field. +| `core.stageSkip` literal `features["subject_age_h"]` | Reads `core.subject_age_h`. `Rule.Stage` **keys** (`max_subject_age_h`, `min_subject_age_h`) are stage names and stay unchanged. | Table test with a namespaced vector. A grep test fails on flat literals in `core`. | +| `core.quantizeAgeFeaturesForHash` | Deleted. Buckets come from `FeatureDef.HashQuantum`, plus `hash_quantum` for custom features (§5.5), passed in `core.Vector`. | The hash-drift tests are re-run. | +| `core.inputHash` | Changes once. | The golden records both values; stable afterwards. | +| `local.Scorer` summation and `Version()` | Order follows `FeatureDef.Order`. `Version()` changes once, and the weights file's `version:` becomes `v2`. | Golden risk bits are identical; `Version()` is recorded before and after. | +| `fake.deterministicProbs` (hashes sorted feature names) | Output changes once. P1 re-baselines every fake-based expectation: contract tests, eval tests and fake cassettes. | The committed re-baseline is stable afterwards. | +| Eval cassette key (`hashScoreRequest`) | Changes once. Cassettes (fake and local only, per A1) are re-recorded, and the header gains `feature_key_space: ns-v1`. | A cassette with a mismatched key space fails loudly. | +| `eval.LoadSnapshotCorpus` | The schema bumps to `corpus-v2` (namespaced keys plus `feature_key_space`). A v1 row with a flat name fails with `feature_renamed`, which names the new feature. Nothing is translated silently. | Loader tests. | +| `run.json` SHAs | `rule_sha` and `weights_sha` change once. The manifest gains `feature_key_space`. | The golden compares metrics, not SHAs. | +| `worker.renderReason` / stored reasons | Namespaced names, plus `reason_version: 2`. Old rows are left as they are. | Snapshot test. | +| `corpus_examples.features` | New column `feature_key_space`, defaulting to `flat-v0`. A migration rewrites keys to `ns-v1` (production has no rows, per A1). The loader refuses `flat-v0`. | Migration test on a seeded DB. | +| `feature.Names`, `Map`, `FeatureSet`, `rules.yaml`, weights, mutation/ablation/golden-sign tests, fixtures README | Mechanical rename. | Full suite plus the grep test. | +| `abusekit score --jsonl` | Flat names fail with `feature_renamed`. | CLI contract test. | + +**Enforcement.** P1 lands before any S5 or S8 PR. The v0 plan marks S5 and S8 "blocked on P1". +`TestNoVendorAdapterBeforeRename` fails if `internal/model/{gemini,jev,laya}` exists while +`feature.KeySpace != "ns-v1"`. + +### 5.3 Name grammar + +- **Features:** `^[a-z][a-z0-9]*\.[a-z][a-z0-9_]{0,55}$`. + - Reserved namespaces: `core`, `email`, `brand`, `custom`, and any future pack. + - Tenants define only `custom.*`, scoped to the tenant. + - The suffix `__absent` is reserved for derived indicators (§5.5). +- **Types:** `^[a-z_.]+$`, at most 64 bytes, with a dot. + - Reserved prefixes: `subject.`, `payment.`, `subscription.`, `resource.`, `content.`, + `abusekit.`. +- **Extension fields:** `x_`. +- **Stored undeclared names:** see §5.6. ### 5.4 Product-declared vocabulary -There is no new built-in delivery type. The wire contract (`POST /v1/events`) keeps its shape; -every change below is additive and optional. - ```yaml vocabulary: - version: 4 # must increase with any change; history in §5.6 - subject_kinds: # default [account] - account: {} - api_key: {parent: account} # a key's events also mark its parent account dirty - card: {} + version: 4 + subject_kinds: + account: {} + api_key: {parent: account} + card: {} resource_kinds: - key: {role: credential, aliases: [keys, api_key, api_keys, api-key, apikey, "api key"]} + key: {role: credential, aliases: [keys, api_key, api_keys, api-key, apikey, "api key"]} agent: {role: other} - link_kinds: # in addition to the six built-in kinds - phone_hash: {evidence: true} # counts as neighbour evidence + link_kinds: + phone_hash: {evidence: true} oauth_sub_hash: {evidence: true} types: invite.sent: - role: activity # included by core velocity/burst features + role: activity fields: - invitee_hash: {kind: hash, join_domain: member} - target_class: {kind: enum, values: [own_community, other_community]} - preview: {kind: text, role: title, max_len: 200} # stored as skeleton by default - link_host: {kind: domain} - extend: # extension fields on built-in types + invitee_hash: {kind: hash, join_domain: member} + target_class: {kind: enum, values: [own_community, other_community]} + preview: {kind: text, role: title, max_len: 200} # skeleton-only by default + link_host: {kind: domain} # reduce: etld1 by default + extend: resource.created: x_visibility: {kind: enum, values: [public, private]} ``` -**Roles.** Only these ship: -- resource kinds: `credential` and `other` (the review's minimal set; an undeclared kind is - `other`); -- types: `activity`; -- fields: `title`, which `brand.title_match` reads, and `self`, a bool marking a self-directed - event that such features exclude. - -`brand.title_match` counts distinct brands across the `title` fields of non-self activity events in -a trailing 1 h, capped at 3. A `title` field stored skeleton-only is matched through -`BrandSet.MatchesSkeleton`, which compares already-folded tokens to the brands' skeletons and -skips the second fold. (#7 found that double folding is not idempotent for leetspeak.) - -**Subject kinds.** Events gain two optional fields: -- `subject_kind` (default `account`); -- `also: [{kind, id}]`, with at most 3 entries. - -`also` lets one event (a charge attempt, say) be indexed under the merchant account, the card and -the customer at once: -- the event is stored once, and idempotency is still keyed on `(tenant, producer, id)`; -- one index row is written per subject; -- each named subject, and the declared `parent` of the primary subject, is marked dirty. - -Subjects are keyed `(tenant, kind, id)`, and the API follows: -- `GET /v1/subjects/{subject}` and `POST .../evaluate` take an optional `?kind=` (default - `account`); -- the list endpoint (S3b) gains a `kind` filter; -- rules declare `applies_to: [kinds]` (default `[account]`); -- a feature is computed only for the kinds in its `SubjectKinds`. For example, the `core` - onboarding features exist only for `account`. - -**Link kinds.** The `links` object gains `custom: {"": ""}`, with at most 8 -entries. Each value is re-HMACed like every hash (§5.6). Built-in and declared evidence kinds feed -both `Neighbors` and the new `neighbours` op (§5.5). - -**Legacy behaviour by declaration.** A profile with no `resource_kinds` gets the implicit -declaration `key: credential`, with #7's aliases. The golden replay proves this reproduces -`resourceKindAliases`. +**Roles.** Only these roles ship: +- resource kinds: `credential` and `other` (an undeclared kind counts as `other`); +- event type: `activity`; +- fields: `title` and `self`. + +`brand.title_match` counts distinct brand ids in the `title` fields of non-self activity events +over the trailing hour, capped at 3. +- The brand ids are matched at ingest, on the value before skeleton reduction, using #7's matcher. +- They are stored as the derived field `x_brands`. + +**Subject kinds.** Two optional wire fields are added: `subject_kind` (default `account`) and +`also` (a list of up to 3 `{kind, id}` entries). +- The event row is stored once. +- `event_subjects` gets one row for the primary subject and one per `also` entry, at most 4. +- It also gets one `via_parent` row for each of those subjects that has a declared parent. + Parents can't have parents, so an event has at most 8 index rows. +- Each indexed subject is marked dirty. The parent rows let account-level features see events from + child keys. +- Non-account ids and `also` ids are pseudonymised (§5.6). +- Subjects are keyed `(tenant, kind, id)`. +- `GET /v1/subjects/{subject}?kind=` and `evaluate` derive the key from the path id in the same + way. +- Rules declare `applies_to`. + +**Link kinds.** `links.custom` holds up to 8 declared kinds, and every value is re-HMACed. +Declared kinds feed only `neighbours` (§5.5). + +**Legacy default.** A profile without `resource_kinds` gets `key: credential`, with #7's aliases. ### 5.5 Declarative custom features -#### Schema - ```yaml features: - - name: custom.public_links_1h # custom.* only - version: 1 # bump on any change (checked against history, §5.6) + - name: custom.public_links_1h + version: 1 description: public share links created in the last hour - subject_kinds: [account] # default [account] - count: # exactly one op key + subject_kinds: [account] + count: type: share.link_created where: {field: visibility, eq: public} - sum: {field: size_class_weight, default: 1, cap_each: 10} # optional - window: 1h # or first: | lifetime | before_first: {type, where} - transform: {log1p: true, cap: 50} # cap mandatory - prior_sign: "+" # optional, for uniform priors + sum: {field: size_class, default: 1, cap_each: 10} # optional; integer field only + window: 1h # or first: | lifetime | before_first: {type, where} + transform: {log1p: true, cap: 50} + hash_quantum: 0 # optional; defaults below + prior_sign: "+" ``` -**Windows.** +**Windows** | Window | Range | | --- | --- | -| `window: ` | `(now − dur, now]`, where `dur` is 1 minute to 30 days, in whole minutes, hours or days | -| `first: ` | `[start, start + dur)` | -| `lifetime` | `(now − 90d, now]`: the retention horizon. The docs call this "retained lifetime" so nobody reads it as all-time. | -| `before_first: {type, where}` | `(now − 90d, t_B]`, where `t_B` is the first matching event at or before `now`. If there is no such event, it falls back to `(now − 90d, now]`. The end is inclusive, matching #7's "at or before". | - -Events with `at > now` are excluded from every window. See "Future events" under Semantics. +| `window: ` | `(now − dur, now]`, for any duration from 1 minute to 30 days | +| `first: ` | `[start, start + dur)`, class N | +| `lifetime` | every ingested event with `at ≤ now`, served from counters (class F). Only valid on `count`/`sum`. | +| `before_first: {type, where}` | `(−∞, t_B]`, where `t_B` is the first matching event at or before now, or `(−∞, now]` if there is none. Served from facts, only valid on `count`. The end is inclusive, as in #7. | -**Predicates.** A closed set, with no regex: -- equality and membership: `eq`, `ne`, `in` / `not_in` (at most 256 values; enum values are checked - against the declaration); -- set files: `in_set` / `suffix_in_set` (named files of at most 100,000 entries, hashed at load); -- numeric comparison: `gt`, `gte`, `lt`, `lte`; -- presence: `exists`; -- combinators: `all`, `any`, `not`, with nesting depth at most 2 and at most 8 leaves. +**Predicates.** The set is closed and has no regex: +- `eq`, `ne`, `in`, `not_in` (at most 256 values; enum values are checked at load); +- `in_set`, `suffix_in_set` (set files of at most 100k entries, hashed and leak-scanned at profile + load); +- `gt`, `gte`, `lt`, `lte`, `exists`; +- `all`, `any`, `not`, with nesting depth at most 2 and at most 8 leaves. -A leaf is false when its field is absent or has the wrong type. `text` fields are banned from -predicates, `distinct`, `group_by` and `on`. +A leaf over an absent or mistyped field is false. Every predicate is a parameterised query shape, +never user SQL, so all of them push down to the aggregate engine. `text` fields can't be used in +predicates, `distinct`, `group_by` or `on`. -**Operations.** `cap(v)` below is the transform. +**Operations** -| Op | Definition at `now` | Bound / cost class | +| Op | Definition at `now` | Class and cost | | --- | --- | --- | -| `count` | The number of matching events in the window, or the sum of `sum.field` over them. An absent field counts as `default`, and each value is clamped to `cap_each`. | O(n) | -| `distinct` | `{field}`: the number of distinct values of `field` among matches. Tracking stops at `track_max`, the smallest count whose transformed value reaches `cap`; `track_max` must be ≤ 10,000. | O(n); memory O(track_max) | -| `share` | `{where, match, sum?}`: the numerator over the `match` events divided by the denominator over the `where` events. When the denominator is 0, the value is `if_empty` (default 0). | O(n) | -| `peak` | `{size}`: the maximum of `agg(E ∩ (t − size, t])`, taken over `t` = the instants of matching events inside the outer window. Sub-windows are **clipped** to the outer window, so events outside it never count even if they fall inside `(t − size, t]`. Requires `size` ≤ window and window ÷ size ≤ 1440. | O(n), two-pointer | -| `time_between` | `{from: {type, where, anchor: first\|last}, to: {type, where}, until_now, if_absent}`. `t_A` is the first (or last) matching `from`; `t_B` is the first matching `to` with `t_B ≥ t_A`. The value is minutes from `t_A` to `t_B`. With `until_now` and no `t_B`, it is `now − t_A`. `if_absent` is mandatory. Floored at 0. | O(n) | -| `sequence` | `{a: {type, where}, b: {type, where}, within: , on: {a: field, b: field}?}`: the number of `b` events in the window that have at least one `a` event with `t_a ∈ (t_b − within, t_b]` and, if `on` is set, `a.on == b.on`. Both `on` fields must be `hash` fields with the **same `join_domain`** (§5.6), or equality across them is meaningless; the loader checks this. The evaluator keeps, per `on` value, the latest `a` instant ≤ `t_b` in an LRU capped at 10,000 keys. Eviction sets the feature's truncated flag. Requires `within` ≤ 24 h. | O(n); memory O(keys) | -| `group_by` | A modifier on `count`, `distinct` or `sum`: `{field, reduce: max\|{count_gte: k}, max_groups ≤ 1000}`. Matches are grouped by `field` and the op is applied per group. `max` returns the largest group value; `count_gte: k` returns how many groups have a value ≥ k. Groups are created in `(at, producer, id)` order. Once `max_groups` is reached, new groups are ignored and the truncated flag is set. | O(n); memory O(groups) | -| `ratio` | `{num: custom.a, den: custom.b, if_empty}`: `num ÷ den` over two **non-ratio** custom features, forming a depth-1 DAG. The loader rejects cycles and ratio-of-ratio. | O(1): the inputs are computed anyway | -| `neighbours` | `{via: [link kinds], where: {deleted: permanent} \| {labelled: abusive} \| {created_within: } \| {}}`: the number of distinct other same-tenant, same-kind subjects that share any `via` key and satisfy the condition. Fan-in is capped at 50 per key and 200 in total, as in main §4.2, and hitting a cap sets `core.neighbors_truncated`. | One store query per distinct `via` set, cached per extraction. At most 4 `neighbours` features per tenant. | - -**`relative_to_history`.** A modifier on `count`, `distinct` and `peak`. It is the exact pipeline -of #7's `burstFactor × ageDecayFactor`: +| `count` | Matching events in the window. With `sum`, the sum of an integer field: `default` if absent, clamped to `cap_each`. | A; one aggregate | +| `distinct` | `{field}`: the number of distinct values. | A; `count(DISTINCT)` | +| `share` | `{where, match, sum?}`: numerator over `where ∧ match` ÷ denominator over `where`. With `sum`, **both** numerator and denominator sum the field. An empty denominator gives `if_empty`. | A; one aggregate | +| `peak` | `{size, sum?}`: the maximum of `agg(E ∩ (t − size, t])`, over `t` at the instants of matching events inside the outer window. `agg` is a count, or with `sum`, a sum. Sub-windows are clipped to the outer window. | R; exact under a saturation limit (§5.7) | +| `time_between` | `{from: {type, where, anchor: first\|last}, to: {type, where}, until_now, if_absent}`. `t_A` is the first (or last) matching `from` with `at ≤ now`. `t_B` is the first matching `to` with `t_A ≤ t_B ≤ now`. The value is minutes from `t_A` to `t_B`, or `now − t_A` when `until_now` is set and there is no `t_B`. Absence semantics are below. | A; indexed first/last-match queries | +| `sequence` | `{a, b, within ≤ 24h, on?: {a: f, b: g}}`: the number of `b` events in the window with an `a` event where `t_a ∈ (t_b − within, t_b]` and, if `on` is set, `a.f == b.g`. Both `on` fields must be `hash` fields with the same `join_domain`. | A; aggregate with a correlated existence test | +| `group_by` | A modifier on `count`, `distinct` or `sum`: `{field, reduce: max \| {count_gte: k}, max_groups}`. It groups by `field`, applies the op per group, then reduces: `max` takes the largest group value, `count_gte` counts groups at or above `k`. **Exact**: the aggregate engine groups every matching event, so there is no first-come admission for decoys to exploit. `max_groups` (at most 1,000) bounds only the in-memory adapter. Past it, the in-memory adapter uses space-saving (Metwally) with `k = max_groups` and flags the result `partial`. Postgres is always exact. | A; `GROUP BY` | +| `ratio` | `{num, den, if_empty}` over the **pre-transform** values of two non-ratio custom features. Depth 1: no cycles, no ratio of ratios. | D; O(1) | +| `neighbours` | `{via: [declared link kinds], where: {deleted: permanent} \| {labelled: abusive} \| {created_within: } \| {}}`: the number of distinct other same-tenant, same-kind subjects that share a `via` key and meet the condition. **As of `now`**: links with `first_seen ≤ now`, subjects created ≤ now, deletions and labels ≤ now. `created_within` is relative to `now`. This matches `evidenceAsOf` in `eval/neighbors.go`. Fan-in is capped at 50 per key and 200 in total; hitting the cap flags the result `partial`. | F/A; one indexed query per `via` set; at most 4 per tenant | + +**Absence semantics** (`time_between`, `sequence`) +- `if_absent` is a mandatory number. +- The loader also derives a 0/1 indicator, `__absent`, and `absent_sign` is mandatory. +- This matters for abandonment. Without the indicator, an account that signs up and never uses the + product would read as the most benign value on the duration scale. With it, a rule can weight + abandonment separately. +- Uniform priors use `absent_sign` as the indicator's sign. +- Durations should use `log1p`, for example `transform: {log1p: true, cap: 9.3}` for up to a week. + All examples do. + +**Neighbour evidence stays separate** +- `core.linked_*` and `core.fingerprint_*` read only built-in link kinds admitted by the tenant's + `feature.Config`. +- Declared kinds feed only `neighbours`, so they never double-feed `core.linked_deleted_n`. +- **Dirty-mark propagation** covers built-in *and* declared evidence kinds. When a subject gains a + link key, is permanently deleted, or is labelled, every subject sharing any evidence key with it + is marked dirty, up to the fan-in cap. + +**`relative_to_history`** applies to `count`, `distinct` and `peak`: ``` -cur = op over the feature's window at now -E_b = { matching events e : now − lookback < e.at ≤ now − exclude_recent } -base = baseline_op over E_b (default: the same op; for count/distinct over window W it is - the max of that op over sliding windows (t − W, t], t ∈ instants of E_b, clipped to E_b; - for peak it is peak over E_b with the same size) -v1 = min( cur / max(base, 1), ratio_cap ) ratio_cap defaults to transform.cap -age_days = (now − start) / 24h -d = clamp( 1 − (age_days − full_until)/(zero_at − full_until), floor, 1 ) if age_decay -v2 = v1 · d (v1 if no age_decay) -value = transform(v2) log1p (optional), then cap +cur = op over its window at now (class A, or R for peak: exact) +E_b = { matching e : now − lookback < e.at ≤ now − exclude_recent } +base = baseline_op over E_b (default: same op; count/distinct over W → max over sliding + windows (t − W, t], t ∈ instants of E_b, clipped to E_b; peak → peak over E_b) +v1 = min( cur / max(base, 1), ratio_cap ) ratio_cap defaults to transform.cap +d = clamp( 1 − (age_days − full_until)/(zero_at − full_until), floor, 1 ) if age_decay +value = transform( v1 · d ) ``` -Parameters: -- `lookback` ≤ 30 d, `exclude_recent` < `lookback`, and `floor` > 0. -- `age_decay` defaults to `full_until: 3d, zero_at: 30d, floor: 0.2`, which are #7's constants. -- `baseline:` may override the baseline op, e.g. `baseline: {peak: {size: 10m}}`. That lets a 1 h - sum be compared against a prior 10-minute peak, which is how #7's `sends_1h` works. - -**Which #7 and S2 features the DSL can express:** -- **Expressible:** - - `sends_10m_max`, `sends_1h`, `webmail_sends_1h`: `relative_to_history` with a `baseline` - override, `sum: recipient_count`, `cap_each: 300`, `ratio_cap: 300`; - - `sends_first_day`; - - `webmail_recipient_share`: `share` + `in_set: webmail`; - - `declines_before_first_success`, and `self_send_before_external` with `cap: 2`: via - `before_first`; - - `resource_*`, `credential_*`; - - `upgrade_delay_min`: `time_between` + `cap`. -- **Not expressible:** - - `distinct_recipients_1h`, a mixed aggregate: distinct hashes, falling back to summing - `recipient_count` for events that have no hash; - - `subject_brand_match` and `brand.*`, which need the brand matcher and its integration and - community gates; - - `first_day_distinct_domains`, which has an inclusive end where DSL `first:` windows are - half-open; - - `burst_ratio_24h_vs_lifetime`, which counts future-dated events in its lifetime denominator; - - `linked_*` and `fingerprint_*`, which carry specific deleted-and-labelled evidence semantics; - the `neighbours` op covers the generic cases; - - `subject_age_h`, which reads `start` rather than events. - -The non-expressible features stay Go features in their packs. P4b adds a test that re-expresses -every "expressible" feature in the DSL and checks it matches the Go value bit for bit on every -fixture. - -**Transform.** `cap` is mandatory: a finite value > 0, and at most 1 for `share`. `log1p` is -optional and is applied before the cap: -- `log1p: true` gives `ln(1+v)`; -- `log1p: {scale: s}` gives `s·ln(1+v)`; -- `log1p: {anchored_at: n}` sets `s = n/ln(1+n)`. - -The output always lies in `[0, cap]`, and that range is the feature's `Bound`. - -#### Semantics - -- **Pure and order-independent.** A feature is a function of the loaded event set and `now`. - Ordering, ties and "first" all use `(at, producer, id)`. P0 switches the store's scoring loader - from `(at, seq)` to this order, and the golden captures it before the rename. Any fixture whose - ties now resolve differently is listed as a baseline change: a tie resolved by arrival order was - a latent non-determinism. -- **Half-open windows.** `(now − W, now]`, `[start, start + W)`, and clipped `peak` sub-windows. - `before_first` is the one intentionally inclusive end. -- **Future events.** Events with `at > now` are excluded from every custom window, and schedule a - rescore at their `at`. Some Go features keep frozen legacy behaviour for semantic identity, - documented on each `FeatureDef`: - - `core.resource_total`, `core.credential_total` and `core.burst_ratio_24h_vs_lifetime` count - future-dated events; - - `email.first_day_distinct_domains` has an inclusive end. - - Harmonising them is `core@2`/`email@2` work (§12 Q5). -- **Deterministic floats.** Sums accumulate in `(at, producer, id)` order in one `float64`. -- **Rescore candidates.** Each DSL feature emits: - - the exit time of its oldest in-window match; - - the end of any anchored window; - - for `peak`, the exit of the current maximum's sub-window; - - for `relative_to_history`, the crossings of `now − exclude_recent` and `now − lookback`; - - for `before_first`, the first `to` event; +- `base` is class R, with its own row budget. +- If that budget is hit, `base` is computed over the newest rows of `E_b`. That can only + under-estimate the maximum, so `v1` can only rise. The value is flagged `partial`. +- These features must have `prior_sign: +`. A negative weight on one is rejected + (`baseline_weight_negative`), so truncation can't lower risk. All four of e2a's history-relative + features have positive weights. +- `age_decay` defaults to `{full_until: 3d, zero_at: 30d, floor: 0.2}`. +- `baseline: {peak: {size: 10m}}` overrides the baseline op. + +**Which #7 and S2 features the DSL can express** + +Expressible: +- `sends_10m_max`, `sends_1h` and `webmail_sends_1h`, via `baseline`, `sum: recipient_count`, + `cap_each: 300` and `ratio_cap: 300`; +- `sends_first_day`; +- `webmail_recipient_share`; +- `declines_before_first_success`; +- `self_send_before_external`, with `cap: 2`; +- `resource_*` and `credential_*`; +- `upgrade_delay_min`. + +Not expressible: +- `distinct_recipients_1h`: distinct hashes plus a fallback sum; +- `subject_brand_match` and `brand.*`: they need the matcher's gates; +- `first_day_distinct_domains`: its window end is inclusive; +- `burst_ratio_24h_vs_lifetime`: its lifetime denominator counts future-dated events; +- `linked_*`: built-in evidence semantics; +- `subject_age_h`: it reads `start`. + +P5b re-expresses every expressible feature in the DSL and checks it bit-for-bit against the Go +implementation. + +**Transform.** `cap` is mandatory: finite, greater than 0, and at most 1 for `share`. `log1p` is +optional and applied before the cap: +- `true` gives `ln(1+v)`; +- `{scale: s}` gives `s·ln(1+v)`; +- `{anchored_at: n}` sets `s = n/ln(1+n)`. + +The result lies in `[0, cap]`, and the feature's `Bound` is `cap`. + +**Semantics** +- **Pure.** A feature is a function of the stored events, the facts and counters, and `now`. + - "First" and ties resolve by `(at, producer, id)`. + - P0 switches the scoring order from `(at, seq)` to that key before capturing the golden. +- **Half-open windows.** `(now − W, now]` and `[start, start + W)`, and `peak` clips its + sub-windows. `before_first` is the one intentionally inclusive end. +- **Future events.** Events with `at > now` are excluded from custom features and schedule a + rescore at their `at`. + - Frozen legacy exception: `core.resource_total`, `core.credential_total` and the + `burst_ratio` denominator count future-dated events. Counters include them naturally. + - Frozen legacy exception: `email.first_day_distinct_domains` has an inclusive end. + - Harmonising these is `@2` work (§12 Q5). +- **Counter exactness.** Custom lifetime counters exclude events with `at > now`. Those can only + sit inside the ±24 h skew window, so the class A engine subtracts them exactly with a query over + `(now, now + 24h]`. Counter `sum` fields must be integers, so the order of addition can't matter. +- **`hash_quantum`.** This lets `SkipInputUnchanged` skip unchanged inputs. The default is 60 + minutes (pre-transform) for `time_between` with `until_now`, and 0 otherwise. +- **Rescore candidates.** A feature proposes: + - the exit of its oldest in-window match; + - the end of its anchored window; + - the exit of the peak sub-window; + - baseline boundary crossings; - future-dated matches. - Rescore-storm control (§5.8) then filters and coalesces them. + §5.8 filters and coalesces them. -### 5.6 Redaction, pseudonymisation and vocabulary history +### 5.6 Redaction, pseudonymisation, keys, vocabulary history -Ingest never consults rules. It consults the tenant's **vocabulary**. What abusekit stores is -**pseudonymised**, not anonymous: anyone holding both the key and a candidate value can recompute -a keyed hash. Retention and erasure (main §4.4, §4.11) therefore apply to it. +Ingest consults the tenant **vocabulary**, never rules. What it stores is **pseudonymised**: anyone +holding a key can recompute a hash from a candidate value. Retention and erasure therefore apply to +hashed values too. -**Field kinds:** +**Field kinds** | Kind | Stored as | Rejected when | | --- | --- | --- | -| `text` | NFKC. Email, card, IP and phone shapes are **masked** (`@`, `#card`, `#ip`, `#phone`), and the value is truncated at `max_len` (≤ 500). **`store: skeleton` is the default for custom text**: only the confusables skeleton is kept. `store: raw` is opt-in, for text-accepting scorers only. The built-in `content.sent.subject_line` stays raw (main §4.3). | Control characters, invalid UTF-8 | -| `number` | A finite float64, checked against the optional `min`, `max` and `integer`. Integers of 13–19 digits also get the Luhn check. | Out of range, non-finite, or Luhn-valid | -| `bool` | As-is | Not a bool | -| `enum` | One of the declared values (at most 64, each matching `[a-z0-9_.-]{1,64}`) | Undeclared value | -| `hash` | **Re-HMACed at ingest:** `hk:` + hex(HMAC-SHA256(k_tenant, input))[:32]. `input` is `lp(type) ‖ lp(field) ‖ lp(value)`, or `lp("join:" ‖ join_domain) ‖ lp(value)` when the field declares a `join_domain`; `lp` is a u32 length prefix. A `join_domain` makes the same identifier equal across fields and types (for `sequence.on` and cross-type `distinct`); without one, hashes are separated per field. The producer's value is never stored. | Doesn't match `^[A-Za-z0-9_:+/=-]{8,128}$` | -| `domain` | Lower-cased and IDNA-encoded to ASCII. It must end in a public suffix from the embedded, versioned PSL snapshot, or in an RFC 6761 special-use name (`.test`, `.example`, `.invalid`, `.localhost`) so synthetic fixtures stay valid. With the optional `reduce: etld1`, only the registrable domain is stored. | IP literal, all-numeric label, unknown suffix, `@` | -| `timestamp` | RFC 3339 (#7's pattern) | Anything else | - -**Rules, in order:** -1. **Leak scan.** It runs over every key and value of every event, and now detects: +| `text` | NFKC-folded. Email, card, IP and phone shapes are **masked** as `@`, `#card`, `#ip`, `#phone`. Truncated at `max_len` (at most 500). Custom text defaults to `store: skeleton`; `store: raw` is opt-in, for text-accepting scorers only. **Built-in `content.sent.subject_line`** stays raw (main §4.3). From P3a it gets the same masking: email (already in #7), card, IP and phone. | Control characters or invalid UTF-8 | +| `number` | **Trusted as declared.** `max` is **required**; `min` and `integer` are optional. There's no Luhn check on numbers: the field is declared and reviewed, and Luhn only falsely rejects amounts and counts. | Out of range, non-finite, or declared without `max` (a load error) | +| `bool` | As is. | Not a bool | +| `enum` | A declared value: at most 64 values, each matching `[a-z0-9_.-]{1,64}`. Leak-scanned at profile load. | Undeclared value | +| `hash` | **Re-HMACed:** `hk:` + hex(HMAC-SHA256(k, input))[:32], with `input = lp(type) ‖ lp(field) ‖ lp(value)`. If the field declares `join_domain`, `input = lp("join:" ‖ join_domain) ‖ lp(value)`. `lp` is a u32 length prefix. | Doesn't match `^[A-Za-z0-9_:+/=-]{8,128}$` | +| `domain` | Lower-cased and IDNA-encoded to ASCII. Must end in a public suffix from the pinned PSL snapshot, or in an RFC 6761 special-use name. **Defaults to `reduce: etld1`**; `none` opts out. Stored **in cleartext only if allowlisted**: the email pack's webmail list, the disposable list, or the vocabulary's `popular_domains` set file (top-N). Anything else is stored as `hd:` + HMAC. | IP literals, all-numeric labels, unknown suffixes, or an `@` | +| `timestamp` | RFC 3339. | Anything else | + +**Built-in fields in P3b** (`RedactionSchemaVersion` becomes 3) + +| Field | Reduction | Stored as | Feature impact | +| --- | --- | --- | --- | +| `content.sent.recipient_domain` | **`none`**: `etld1` would merge `target1.example.test` and `target2.example.test`, which would change `email.first_day_distinct_domains` | Cleartext if allowlisted, else HMAC of the normalised value | HMAC is injective, so distinct counts don't change. Webmail membership is checked on cleartext. | +| `content.sent.first_link_host` | `etld1` | Cleartext if allowlisted, else HMAC | Text scorers see a token for unknown hosts (§12 Q14) | +| `resource.*.address_domain` | `none` | Cleartext if allowlisted, else HMAC | No feature reads it | +| `content.sent.recipient_hash` and every `links` value (built-in and declared) | Re-HMAC. Each link kind's `join_domain` is its kind name. | — | Equality is preserved, so no value changes | +| `links.asn` | **Exempt** from the hash format check and from re-HMAC. It's coarse routing metadata in the clear (main §4.2), and its grammar tightens to `^(AS)?[0-9]{1,10}$`. | Cleartext | None | + +**Subject ids** +- `account` ids stay as the tenant chose them, but they are now leak-scanned. An email, card or + phone shape is rejected with `bad_subject`. +- Every other kind, and every `also` id, is stored as `hs:` + HMAC(lp(kind) ‖ lp(id)). +- Lookups apply the same derivation. + +**Ingest rules, in order** +1. **Input leak scan.** Every key and value is checked for: - email addresses; - - Luhn-valid runs of 13–19 digits (separators allowed); + - Luhn-valid runs of 13–19 digits, with separators; - IPv4 and IPv6 literals; - - phone shapes: `+` followed by 8–15 digits, or grouped national formats of 10 or more digits. - - What happens on a match depends on the field: - - a declared `text` field is masked; - - a declared `hash` field is exempt, because its value is replaced by the HMAC; - - anywhere else, the event is rejected with `redaction_failed`. -2. **Declared fields** are handled by kind, as in the table above. -3. **Undeclared fields** of any type, built-in or declared, are **dropped whatever their kind**, - numbers included. The row keeps `x_undeclared: [sorted field names]`; a name that fails the - leak scan is rejected. **Undeclared types** keep only `type`, `at` and that list of names. - Revision 1 hashed undeclared strings instead; that is removed. -4. **Built-in hash values are re-HMACed too:** `content.sent.recipient_hash` and every `links` - value, both the six built-in link kinds and declared ones. Every `links` kind has its own - `join_domain`, which is the kind's name. Equality is preserved, so `distinct` counts, - neighbour joins and every golden feature value are unchanged; only the stored bytes change. A - producer that mistakenly sends a raw phone number as a "hash" never has it stored. -5. **Built-in domain fields** (`recipient_domain`, `address_domain`, `first_link_host`) get the - `domain` rules, and `RedactionSchemaVersion` becomes 3. Every committed fixture uses `.test` - and passes. - -**Keys and rotation.** The key interface is provider-agnostic: + - phone shapes. + + The scan runs after NFKC and after folding every Unicode `Nd` digit to ASCII. A hit in a `text` + field is masked. A hit in a `hash` field is exempt, because the value is replaced. Anywhere else, + the event is rejected with `redaction_failed`. +2. **Declared fields** are handled by kind (table above). +3. **Undeclared fields** are dropped, numbers included. + - The stored field names must match `^[a-z0-9_]{1,64}$`, at most 32 per event. + - Names that fail the grammar or the leak scan are dropped and counted in + `x_undeclared_dropped_n`. + - An undeclared type keeps only `type`, `at` and `x_undeclared`. +4. **Egress scan.** The serialised row is scanned again before insert, with the same detectors, + skipping the known `hk`, `hd` and `hs` prefixes and the mask markers. A hit rejects the event + with `redaction_failed` and increments `abusekit_egress_scan_hits`; any hit means a redaction + bug. +5. **Profile load** leak-scans every enum value, `in` list and set file. A hit rejects the profile + with `vocab_invalid`. + +**Keys: distinct per tenant by construction** ```go type Keys interface { - // Current returns the key used to write new values, and its id. - Current(ctx context.Context, tenant string, purpose Purpose) (id string, key []byte, err error) - // ReadSet returns every key readers must accept right now (current, plus the previous key - // during a rotation). - ReadSet(ctx context.Context, tenant string, purpose Purpose) ([]KeyRef, error) + // Master returns the root secret for key id `id`, or the current one if id == "". + Master(ctx context.Context, id string) (keyID string, secret []byte, err error) + ReadIDs(ctx context.Context) ([]string, error) // current + previous during rotation } + +// Derive is the only way to obtain a tenant key: +// k = HKDF-SHA256(secret, salt = "abusekit/v1", info = lp(tenant) ‖ lp(purpose)) +func Derive(master []byte, tenant string, purpose Purpose) []byte ``` -There are two adapters: a file adapter for dev and tests, and a cloud secret-manager adapter. -Neither the names nor the config mention a provider. Rotation is **dual-key**: -- For `max_lookback` (at most 30 days), ingest writes each hash field twice: `` under the - new key and `__prev` under the old one. Links get one row per key id. -- Until the rotation's `read_flip_at`, features read the `__prev` values and neighbour joins match - either key id. -- After `read_flip_at`, readers use only the new values and `__prev` writes stop. -- The key id is part of `vocab_version` (`"@/k"`), so every row can be traced to - the key that hashed it. - -**Vocabulary history lives in the config tree.** The mount holds `tenants//history/NNNN.yaml`, -an append-only list of accepted vocabulary and custom-feature versions, each with an -`effective_at`. CI in the private config repo runs `abusekit config check` over the full history -and enforces three rules: -- **Widening only.** Adding a type, field, enum value, `max_len` or `max` is fine. Changing a - field's kind, removing an enum value or narrowing a limit needs a new field name. -- **No redefinition.** A custom feature's `(name, version)` is never redefined. -- **Monotonic time.** `effective_at` never goes backwards. - -At runtime, the `tenant_config_versions` table records what was actually loaded and refuses a -profile that contradicts it. It is a guard, never the source of truth. - -**Warm-up.** A feature is **cold** from its `effective_at` until -`effective_at + max(window, lookback + exclude_recent)`. That applies to a new custom feature, a -new version of one, and any feature over a newly declared field or type. An `advise` rule with a -cold input is scored and stored as shadow for that period: it reports `mode: shadow` with -`warming_until`. The value is still computed; it just can't drive a tier on partial history. - -### 5.7 Bounded loading, step budget, and truncation as a signal - -This replaces revision 1's bound of the 50,000 newest events and its 250 ms wall-clock deadline. - -**Load plan (per profile and subject kind, computed at compile time):** -1. **Onboarding types, in full.** `subject.*`, `payment.*` and `subscription.*` are loaded - oldest-first, up to `onboarding_bytes` (1 MiB decoded by default). Onboarding facts such as - "first success" come from the *earliest* events, so recent activity can never push them out. -2. **Anchored range.** `[start, start + A)` is loaded oldest-first, where `A` is the maximum - `first:` duration (24 h for e2a). -3. **Trailing range.** `(now − L, now + 24h]` is loaded **newest-first** until the total decoded - size reaches `history_bytes` (8 MiB by default). - - `L = max over features of max(window, lookback + exclude_recent + baseline width)`. - - `L` is 90 d if any feature uses `lifetime` or `before_first`. - - The `+24h` covers future-dated events inside the skew allowance, which schedule rescores. -4. **`start`** comes from `subject.created.account_created_at` if present, else from - `subjects.first_seen_at`, never from loaded events. Truncation therefore can't move it. - -Events are deduplicated across the three ranges by `(producer, id)`, and the combined history is -ordered `(at, producer, id)`. - -**Step budget, instead of wall-clock time.** A step is one event dispatched to one feature, plus -one step per predicate leaf. The per-subject budget is -`steps_max = history_bytes / avg_event_bytes × per_type_fanout_max × leaves_max`. With the defaults -that is 8 MiB / 256 B × 16 × 8 ≈ 4.2M. Every pack charges its steps through `Input.Budget`, the Go -packs included. - -The P4a benchmark calibrates the ns-per-step figure end to end (store load, decode, projection, -extraction, orchestration). The committed `cost_table.yaml` records it, and CI fails if a profile -at the limits exceeds p99 50 ms. Because the budget counts work, when it runs out doesn't depend -on how fast the host is. - -**Exhaustion and truncation.** Exhausting the byte cap, the step budget, `max_groups` or the -`sequence` keys never makes a rule unscored: -- features are computed over what was processed, in a deterministic order (newest-first for - trailing features, earliest-first for onboarding and anchored ones); -- `core.history_truncated = 1` is set, and `core` requires it to carry a **positive** weight. - -A pack **error** (a bug) still marks the rules that read that pack as unscored and degraded. -Truncation never does. - -**Why flooding with cheap events can't evade:** -- (a) Onboarding facts and `start` are loaded separately and are never displaced. -- (b) Newest-first loading keeps every trailing window, up to the byte cap. The flood is itself the - most recent activity, so it is counted, and it raises every volume or velocity feature it - matches. -- (c) Truncation drops only the *oldest* trailing events, which feed three things: - - baselines: a smaller baseline gives a larger `v1`, because `cur / max(base, 1)` is monotone; - - lifetime denominators: a smaller denominator gives a larger ratio (as with `burst_ratio`); - - lifetime totals: a smaller total gives a smaller value. This is the only direction that can - lower risk. -- (d) An invariant closes the third case. `packtest` checks it for every weights file that - includes `core.history_truncated`: - `w(core.history_truncated) ≥ Σ over features f with TruncationDir(f) = −1 of |w_f| · Bound_f`. - Truncation therefore never lowers the linear sum. Two Go features have no natural bound - (`core.resource_total`, `core.credential_total`). They get `Bound` from a cap of 1,000, which no - fixture reaches, so semantic identity holds. -- (e) Forcing truncation sets the abusive subject's own `history_truncated` signal. - -Criterion 6's property test covers every built-in and DSL feature. - -**Limits** (validated at load; also the fuzz oracle): +- **Purposes:** `hash`, `domain` and `subject`. +- **Adapters:** file (P3b) and cloud secret manager (P3d). +- **Test:** for every pair of example tenants and every purpose, the derived keys differ, and the + same input yields different stored values. + +**Rotation (P3d) is dual-key** +- For `max_lookback` (at most 30 days), ingest writes both `` (new key) and `__prev` + (old key). +- `links` gets one row per key id. +- Readers use `__prev` until `read_flip_at`. +- The key id is recorded in `vocab_version` (`"@/k"`). + +**Dev and staging re-HMAC migration (P3b).** Production has no data yet (A1). +- abusekit never saw the raw values, so re-HMAC means HMAC-ing the stored producer value. That is + exactly what ingest does from P3b on, so migrated rows equal newly ingested ones. +- A Go migration job runs under the tenant key and can resume by `seq`. It: + - rewrites `links.hash` and the hash, domain and subject values in `events.links` and + `events.data`; + - recomputes `events.body_hash`, so duplicate/conflict detection still matches; + - sets `vocab_version`. + +**Vocabulary history lives in the config tree.** Each tenant has an append-only +`tenants//history/NNNN.yaml` in the private config repo, and CI runs `abusekit config check`. +The check enforces three rules: +- the vocabulary can only widen; +- a feature's `(name, version)` is never redefined; +- `effective_at` is monotonic. + +At runtime, `tenant_config_versions` is only a guard. + +**Warm-up.** A newly introduced feature is cold from `effective_at` until `effective_at` plus its +horizon (§5.7). This covers: +- a new custom feature, or a new version of one; +- a feature over a newly declared field or type; +- a new counter spec that is still backfilling. + +While any input is cold, an advise rule is stored as shadow, with `warming_until`. + +### 5.7 Evaluation classes, pass order, budgets, flags and flooding + +Revision 2's shared byte and step budget, and its weighted truncation feature, are replaced by the +following. + +**(a) Onboarding facts are maintained at ingest.** `subject_facts(tenant, kind, subject)` is +updated in the same transaction as the event insert. Every update is monotone and +order-independent: + +| Fact | Update | +| --- | --- | +| `first_seen_at`, `account_created_at` | `LEAST(existing, new)`; `account_created_at` comes from `subject.created` | +| `first_success_at`, `first_success_key`, `first_success_funding` | On `payment.attempt{succeeded}`: replace when the event's `(at, producer, id)` is smaller | +| `payment_counts` | `{succeeded, declined, blocked}`: increment | +| `declines_before_first_success` | On `declined` with `at ≤ first_success_at` (or no success yet): increment. When `first_success_at` moves earlier: recompute with an **indexed count** on `(tenant, kind, subject, type, at)` where `outcome = declined` and `at ≤ first_success_at`. | +| `first_paid_upgrade_at` | `LEAST` over `subscription.changed{status: active, amount_minor > 0}` | +| `subscription_change_count` | Increment | +| `first_external_at`, `self_sends_before_first_external` | The same pattern for `content.sent`; capped at 2 when read | +| `name_brands`, `name_has_at`, `exempt_subject_brands` | Brand ids matched at ingest on `resource.*` names, with the integration-token gate applied | + +- **No onboarding scan is ever byte-capped.** Scoring reads one facts row. +- A brand-list change bumps the facts spec version. A backfill job then recomputes from retained + events, and the brand features stay cold until it finishes. + +**The blocked-payment flood can't move onboarding.** Take 10,000 tiny `payment.attempt` events with +`outcome: blocked`. +- Each one only increments `payment_counts.blocked`, which no built-in feature reads. +- It leaves `first_success_*` unchanged, because that only moves on `succeeded`. +- It leaves `declines_before_first_success` unchanged, because that only moves on `declined`. +- It leaves `first_paid_upgrade_at` unchanged, because that only moves on `subscription.changed`. +- It leaves `start` unchanged. + +So `core.upgrade_delay_min`, `core.upgraded`, `core.declines_before_first_success` and +`core.first_funding_prepaid` are bit-identical with and without the flood. Criterion 6 asserts +this. The same argument covers any flood of any type the update rules above don't name. + +**(b) Lifetime totals come from counters.** Ingest increments +`subject_counters(tenant, kind, subject, spec_id, day, n, sum)`. +- Specs are compiled from the profile and pushed into the vocabulary, so ingest still consults only + the vocabulary. They cover: + - `resource.created`, all kinds; + - `resource.created` with role `credential`; + - `resource.created` + `content.sent` + activity types (for `burst_ratio`); + - non-self `content.sent` recipients, all and webmail-only; + - custom `lifetime` counts. +- A new spec backfills from retained events and stays cold until the backfill completes. +- Values are integers, so every sum is exact. +- No feature comes from a truncated scan of lifetime data, so **no feature can fall when history + is bounded.** `TruncationDir`, its invariant and its packtest check are deleted. + +**Evaluation classes and pass order** (per subject) + +| Pass | Class | What runs | Budget | +| --- | --- | --- | --- | +| 1 | **F** | Read `subject_facts` and `subject_counters`. Evaluate neighbour evidence (indexed, fan-in capped). | None; O(1) rows plus capped neighbour queries. | +| 2 | **N** | Anchored `first:` features. While `now < start + A + 24h`, each runs its own class A query over `[start, start + A)`. After that, its value is **frozen** into `subject_facts.anchored`. Later events fall outside the skew allowance, so the frozen value can't go stale. | **Reserved**: runs before A, R and G, and shares with nothing. | +| 3 | **A** | One exact aggregate query per feature: `count`, `sum`, `share`, `distinct`, `group_by`, `time_between`, `sequence`, the `cur` part of `relative_to_history`, and the class A parts of Go features. | **Per feature.** Returns O(1) or O(groups) rows. DB cost is O(rows in that feature's window), via the `(tenant, kind, subject, type, at)` index. | +| 4 | **R** | `peak` and `relative_to_history` baselines. Rows stream newest-first under the feature's own budget. | **Per feature.** `peak`: LIMIT `cap·⌈W/S⌉`. Baselines: 50,000 rows by default. | +| 5 | **G** | The remaining Go-pack computation that isn't expressible as A or R. For e2a after §5.1, this is only `email.distinct_recipients_1h`'s fallback sum, which is itself a class A aggregate. | **Per pack.** The pack's own row budget, not shared with any other pack or feature. | +| 6 | **D** | `ratio` and the `__absent` indicators. | O(1) | + +**Why `peak` is exact under its limit.** Stream the matching rows newest-first with LIMIT +`N = cap·⌈W/S⌉`. +- Split `W` into `⌈W/S⌉` slots of width `S`. +- If the window holds at least `N` matching units, some slot holds at least `cap` units. +- The sub-window `(t − S, t]` that ends at the last event in that slot covers the whole slot. So + the true peak is at least `cap`, and the transformed value saturates at `cap`, which is what the + limited computation reports. +- If the window holds fewer than `N` units, the stream is complete and exact. +- With `sum`, the units are integer addends. Addends of 0 or less are excluded by the pushed-down + predicate, so `N` rows carry at least `N` units. + +**(c) Hitting a bound sets a flag; it's never a weight.** +- `Output.Partial` lists features whose own budget was hit: + - a `relative_to_history` baseline over budget; + - the in-memory `group_by` adapter past `max_groups`; + - `neighbours` at its fan-in cap. +- `core.history_truncated` no longer exists. +- In the API, each signal gets `partial: ["custom.x", …]` (omitted when empty). The subject gets + `partial: true` when any advise rule's signal has partial features. +- `partial` does **not** set `degraded`, because the rule was scored. Every partial computation is + one-sided toward *higher* risk: + - baselines only shrink; + - space-saving over-estimates; + - fan-in caps apply only to features required to be positively weighted. + + So callers can read `partial` as "risk may be overstated by these features, never understated". +- **Uniform rules** record `partial` like any rule. They are shadow-only, so it never affects a tier. + +**(d) Budgets are per feature and per pack. Changing one feature never moves another.** +- A feature's value depends only on the facts, the counters, its own queries and its own budget. +- Anchored work is reserved and runs first; facts need no budget. +- No pass reads another feature's intermediate state. The one exception is class D, which reads its + declared inputs' pre-transform values. +- **`TestFeatureIndependence`:** for every example profile and every feature `f`, delete, add or + redefine every other custom feature `g`, one at a time, and assert that `f`'s bits are unchanged + on every fixture. A Go-pack variant does the same by swapping pack versions. + +**(e) The flood property: `risk(flooded, bounded evaluator) ≥ risk(flooded, unbounded reference)`.** + +The *unbounded reference* is the naive evaluator from criterion 4: it scans every stored event, with +no limits and no facts. The generator injects floods of minimum-size events (the smallest valid +encoding of each type): + +| Dimension | Values | +| --- | --- | +| Read types | Every type any feature reads, with field values that match and don't match its predicates | +| Unread types | Declared types no feature reads, and undeclared types | +| Onboarding types | `payment.attempt` with every outcome (including 10,000 tiny `blocked`), `subscription.changed`, duplicate `subject.created` | +| Placement | Entirely before, interleaved with, and after the real events. Timestamps land inside, at the edges of, and outside every feature window, including future-dated events within skew. | +| Volume | 1×, 10× and 100× each feature's saturation bound | + +The property holds by construction: +- classes F, N and A are exact, so bounded equals reference; +- `peak` is exact by saturation; +- a partial `relative_to_history` value is never below the reference, and its weight is never + negative; +- space-saving never under-counts. + +The test also checks it empirically for every built-in and DSL feature on every fixture. + +**(f) Rollout.** e2a keeps today's unbounded `EventsForSubject` evaluation until all of P4a–P4c have +landed: facts, counters, classes A and R, flags, independence and the flood property. Only then is +bounded evaluation switched on for e2a (§9). + +**Horizons feed the plan.** For each feature, the compiler derives the store range it queries: + +| Feature | Store range | +| --- | --- | +| `window: W` | `(now − W, now]` | +| `first: A` | `[start, start + A)` | +| `time_between` | `(−∞, now]`, via indexed first/last-match queries | +| `sequence` | `a` rows from `(now − W − within, now]`; `b` rows from `(now − W, now]` | +| `relative_to_history` baseline | `(now − lookback, now − exclude_recent]` | + +**Limits** (validated at load; also the fuzz oracle) | Limit | Value | | --- | --- | | Custom features per tenant | 64 | -| Features per event type (fan-out) | 16 | +| Features per event type | 16 | | Predicate depth / leaves | 2 / 8 | -| `in` list size / set-file entries | 256 / 100,000 | -| Window / lookback | ≤ 30 d (`lifetime` and `before_first` are 90 d: retention) | -| `peak` window ÷ size | ≤ 1440 | -| `distinct` track_max / `group_by` groups / `sequence` keys | 10,000 / 1,000 / 10,000 | +| `in` values / set-file entries | 256 / 100,000 | +| Window / lookback | ≤ 30 d | +| `peak` window ÷ size | ≤ 1,440 | +| `group_by` `max_groups` (in-memory adapter) | ≤ 1,000 | +| Baseline row budget | 50,000 | | `neighbours` features | 4 | -| `also` subjects per event / declared link kinds | 3 / 8 | -| `history_bytes` / `onboarding_bytes` | 8 MiB / 1 MiB decoded | -| Steps per subject | Calibrated; default ≈ 4.2M | -| Static cost units per tenant | 256. A unit is the benchmarked cost of one O(n) op at the byte cap. | +| `also` entries / declared link kinds | 3 / 8 | +| Declared `number` without `max` | rejected | + +**Cost calibration.** +- The P4c benchmark measures cost end to end: store queries, decode, extraction and orchestration. +- It runs at the limits, with fixtures flooded to 100× saturation, and writes `cost_table.yaml`. CI + fails if p99 exceeds 50 ms. +- Database work grows with flood volume through index range scans. §5.8's fair queue bounds its + share of worker time. +- A subject whose last pass exceeded its budget is rescored at most every 10 minutes, with a + metric. This changes latency, never values. ### 5.8 Profiles, rules, scheduling -**Profiles live in a private mount** (`--profiles /etc/abusekit/tenants/`, mounted by the hosted -deploy from the operator's private config repo). This repo ships only: -- `examples/tenants/*.yaml`: the five fictional §7 profiles; -- `examples/tenants/reference/`: today's `config/rules.yaml` + `local_weights.yaml`, renamed. The - golden replay runs this profile, and e2a's private profile starts as a copy of it. - -**Validation.** It adds these codes to main §4.5's list. Any failure rejects the whole profile, -with every error collected: -- `pack_unknown`, `pack_requires` -- `feature_unknown`, `feature_not_enabled`, `feature_namespace`, `duplicate_feature`, - `feature_version_reused` -- `dsl_invalid` (with a JSON-pointer path) -- `vocab_invalid`, `vocab_incompatible` -- `weights_unknown_feature`, `uniform_not_shadow`, `truncation_weight_insufficient` -- `subject_kind_unknown` - -**Per-tenant everything:** -- **Reload.** Atomic per tenant. A rejected profile keeps that tenant's previous profile live and - never affects another tenant. `/healthz` reports `config_error{tenant}`. -- **Rule sets.** `worker.computeVerdict` and `serve.currentRuleNames()` both read the subject's - tenant profile, not a global config. P2's test: tenant A's retired rule never appears in tenant - B's view. -- **Scorer version.** Each tenant's local scorer is bound to its own weights file. Its - content-derived version (`local@`) appears in the verdict `model` field and in the - signals of `GET /v1/subjects/{id}`. - -**Stage gates consider only advise-mode local rules.** `maxRiskByScorer` excludes shadow rules, so -a shadow experiment can never open or close a vendor call's gate. This changes behaviour on -`main`, but e2a has no staged rules, so the golden replay is unaffected. - -**Scheduling:** -- **Per-tenant fair queue.** The worker claims dirty subjects round-robin across tenants, with - per-tenant weights (equal by default) and a per-tenant concurrency cap (4 of the batch by - default). One tenant's backlog can't starve another. Main §4.8's priority order applies within a - tenant. -- **Rescore-storm control.** Timer rescores (as opposed to event-driven dirty marks) have three - limits: - - they are scheduled only from features that feed at least one non-shadow rule; - - DSL features coalesce into buckets of `max(5 min, W / 12)`, while built-in Go features keep - the fixed 5-minute bucket for semantic identity; - - a per-tenant timer-rescore budget (default 20 × active subjects per hour) is enforced by the - queue. When it is exhausted, timer rescores defer to the next hour and a metric counts them. - - Event-driven scoring is never budgeted. +**Profiles live in a private mount** (`--profiles`). This repo ships only: +- `examples/tenants/*`: the five fictional §7 profiles; +- `examples/tenants/reference/`: today's config, renamed. The golden replay runs it, and e2a's + private profile starts as a copy of it. + +**New validation codes.** These add to main §4.5's list. The whole profile is rejected, with every +error collected. + +| Area | Codes | +| --- | --- | +| Packs | `pack_unknown`, `pack_requires` | +| Features | `feature_unknown`, `feature_not_enabled`, `feature_namespace`, `duplicate_feature`, `feature_version_reused` | +| DSL | `dsl_invalid` (with a JSON pointer) | +| Vocabulary | `vocab_invalid`, `vocab_incompatible`, `subject_kind_unknown` | +| Weights | `weights_unknown_feature`, `uniform_not_shadow`, `baseline_weight_negative` | + +**Per tenant** +- Reload is atomic and isolated per tenant; `/healthz` reports `config_error{tenant}`. +- Rule sets are per tenant: `computeVerdict` and `currentRuleNames()` read the subject's tenant + profile (tested in P2). +- The local scorer's version (`local@`) is per tenant, in the verdict `model` field and + in API signals. + +**Stage gate (P1s).** `maxRiskByScorer` considers only advise-mode local rules. This changes +behaviour on `main`. e2a has no staged rules, so its golden doesn't change. + +**Scheduling** +- **Fair queue.** Claims go round-robin across tenants, with weights (equal by default) and a + per-tenant concurrency cap (4 by default). Main §4.8's priorities apply within a tenant. +- **Rescore-storm control:** + - timer rescores come only from features that feed a non-shadow rule; + - DSL features coalesce into buckets of `max(5 min, W/12)`, while built-ins keep 5 minutes; + - each tenant has a timer budget, by default 20 × active subjects per hour. Once it runs out, + timer rescores are deferred and counted; + - event-driven scoring is never budgeted. ### 5.9 Weights, eval and bootstrap per pack -- **Weights.** Per rule, keyed by namespaced name. The local math is unchanged. -- **Starter weights.** Each pack has `config/packs//starter.yaml`, with one rule named - `_starter`. Every weight has a `sign:`. `core`'s starter includes - `core.history_truncated`, set so the §5.7 invariant holds. -- **Uniform priors, shadow only.** `weights: uniform` binds a local scorer with, for `k` inputs: - - normalisation `x̂ᵢ = min(max(xᵢ, 0), Boundᵢ) / Boundᵢ`, which clamps to `[0, 1]`; - - `wᵢ = sᵢ · 4/k`, where `sᵢ` is the input's prior sign: `FeatureDef.PriorSign` for pack - features, `prior_sign` for custom ones, default `+`; - - `bias = −2 + (4/k) · |{i : sᵢ = −1}|`. - - So `risk ∈ [sigmoid(−2), sigmoid(2)] ≈ [0.12, 0.88]`. The loader rejects uniform weights in - advise (`uniform_not_shadow`). Promotion needs a real weights file plus a gate run. -- **Eval by profile.** `abusekit eval --profile examples/tenants/` or `--pack

`. - - Floor entries gain `profile:`; an entry without one belongs to the reference profile. - - The manifest gains `profile_sha`, `pack_versions` and `feature_key_space`. - - Corpus v2 is corpus-v1 plus `profile`, `vocab_version` and `feature_key_space: ns-v1`. -- **`packtest`** runs for every pack in CI and checks: - - starter-weight signs match their `sign:`; - - zeroing any weight moves a pack fixture band or an isolated scenario; - - results are deterministic under shuffled arrival; - - no future leakage; - - the pack stays in its namespace and within `Bound`; - - the truncation invariant holds; - - the flood property (criterion 6). -- **Held-out fixtures** for the bootstrap criterion live in - `examples/tenants//fixtures/{dev,heldout}/`. - - Held-out fixtures are authored in a separate commit after the features are frozen, with - non-overlapping generator seeds. - - CI evaluates criterion 2 only on the held-out set. - - A PR that changes a scenario's features and its held-out fixtures together fails a CI check; - held-out fixtures change only in their own PR. -- **Bootstrap.** A new product: - 1. writes its profile; - 2. runs `core_starter` (plus `brand_starter` if it has display names) and a `custom_uniform` - rule, all in shadow; - 3. collects labels through `POST /v1/labels`; - 4. once the harness passes on its labelled set, hand-tunes a weights file and promotes it - through the normal path. +**Weights** +- Weights are per rule and keyed by namespaced name. +- The local scorer's math is unchanged; it sums in `FeatureDef.Order`. +- Starter weights ship at `config/packs//starter.yaml`, with `sign:` on every weight. + +**Uniform priors (shadow only)** +- Normalise each input: `x̂ᵢ = min(max(xᵢ, 0), Boundᵢ)/Boundᵢ`. +- Weight: `wᵢ = sᵢ·4/k`, where `sᵢ` is `PriorSign`, `prior_sign` or `absent_sign`. +- Bias: `bias = −2 + (4/k)·|{sᵢ = −1}|`. +- Risk therefore stays in `[0.12, 0.88]`. +- Using uniform priors in advise is rejected with `uniform_not_shadow`. + +**Eval** +- `abusekit eval --profile

` or `--pack

`. +- Floors gain a `profile:` field. +- The manifest gains `profile_sha`, `pack_versions` and `feature_key_space`. +- Replay uses the in-memory adapters for facts, counters and aggregates, which are + conformance-tested against the Postgres adapters. + +**`packtest`** runs for every pack in CI. It checks: +- signs; +- that zeroing a weight moves a fixture band or an isolated scenario; +- determinism under shuffled arrival; +- no leakage from the future; +- namespace and `Bound`; +- independence (§5.7 d); +- the flood property (§5.7 e). + +**Held-out fixtures** live in `examples/tenants//fixtures/heldout/`. +- They're authored in a separate commit after the features are frozen, with non-overlapping seeds. +- CI runs criterion 2 on the held-out set only. +- A PR that touches both a scenario's features and its held-out fixtures fails CI. + +**Bootstrap** for a new tenant: +1. Start with `core_starter`, `brand_starter` if relevant, and `custom_uniform`, all in shadow. +2. Collect labels. +3. Hand-tune the weights. +4. Pass the gate. +5. Promote. ### 5.10 Storage (expand-only) -- `events`: add `vocab_version text NULL` and `subject_kind text NOT NULL DEFAULT 'account'`. -- New table `event_subjects(tenant, subject_kind, subject, event_seq)`. It indexes `also` and - parent rows, and the scoring loader reads through it. -- `subjects`: the key becomes `(tenant, kind, subject)`, via a new `kind` column (default - `account`) and a unique index. -- `links`: add `key_id` for rotation. -- `corpus_examples`: add `feature_key_space`. -- `verdicts`: add `profile_sha` and `reason_version`. -- New table `tenant_config_versions`, the runtime guard (§5.6). +**`events`** +- New columns `vocab_version` and `subject_kind`. +- New index on `(tenant, subject_kind, subject, type, at)`. +- Partial expression indexes per declared enum field that a feature reads, built with + `CREATE INDEX CONCURRENTLY`. + +**New tables** +- `event_subjects(tenant, subject_kind, subject, event_seq, via_parent)`: at most 8 rows per event. +- `subject_facts`. +- `subject_counters(…, day, n, sum)`: expires with main §4.11's numeric retention. +- `tenant_config_versions`. + +**Changed tables** +- `subjects` is keyed by `(tenant, kind, subject)`. +- `links` gains `key_id`. +- `corpus_examples` gains `feature_key_space`. +- `verdicts` gains `profile_sha`, `reason_version` and `partial`. -**S3b erasure must be vocabulary-aware.** Which stored fields are raw text, skeleton or -pseudonymised hash depends on each row's `vocab_version`, and the erasure ledger records the -version it applied. S3b's design pass must include this. +**S3b erasure must be vocabulary-aware.** It uses each row's `vocab_version` to know which fields +are raw text, skeletons or pseudonyms. It also clears `subject_facts` and `subject_counters`, apart +from the numeric retention for abusive subjects (main §4.4). ### 5.11 API surface summary | Surface | Change | Compatibility | | --- | --- | --- | -| `POST /v1/events` | Optional `subject_kind`, `also`, `links.custom` and `x_` extensions; declared types; stricter domain kind; card/IP/phone scan; re-HMAC; undeclared values dropped | Wire additive. Storage semantics change (pre-GA; decisions Q3, Q14). | -| `GET /v1/subjects/{id}`, `evaluate` | Optional `?kind=`; per-tenant `model` version; reasons use namespaced names; `warming_until` on warming signals | Additive | +| `POST /v1/events` | Optional `subject_kind`, `also`, `links.custom` and `x_` fields; declared types; the redaction in §5.6; account-subject leak scan (`bad_subject`) | The wire format is additive. Storage semantics change (pre-GA). | +| `GET /v1/subjects/{id}`, `evaluate` | `?kind=`; per-tenant `model`; namespaced reasons; `warming_until`; `partial` | Additive | | Per-item codes | None new | Unchanged | -| `abusekit score --jsonl` | Namespaced names; flat names return `feature_renamed` | Breaks once, pre-GA (P1) | -| Config | Per-tenant profiles in a private mount; `rules.yaml` replaced by `examples/tenants/reference` | Breaks once, pre-GA | +| `score --jsonl`, corpus loader | `feature_renamed`; `corpus-v2` | Breaks once, pre-GA (P1) | +| Config | Per-tenant profiles in a private mount | Breaks once, pre-GA | -**Rejected: runtime vocabulary declaration over HTTP** (`PUT /v1/vocabulary`). It would let a -producer key widen its own redaction boundary. +**Rejected: declaring vocabularies at runtime over HTTP.** It would let a producer key widen its own +redaction boundary. ### 5.12 Alternatives considered -- **CEL.** It has no windowed aggregates, so we would still write every op. It is also a large - dependency, costs expressions rather than histories, and gives product engineers worse errors. - It remains a possible later leaf kind for `where` only (§12 Q9). -- **A home-grown expression language.** A parser, a grammar and an injection surface, with no - coverage beyond the closed ops. -- **A plugin ABI.** Go `plugin` is fragile and unsandboxed; WASM is code disguised as config, and - reviewers can't read it. Code belongs in a Go pack upstream, gated by `packtest`. -- **SQL features.** They would couple features to the store's schema, give unbounded cost, and - put tenant isolation at risk. -- **A canonical-key bridge (revision 1).** It lost because nothing stored needs it (A1), and it - would keep two names alive forever. -- **A neutral built-in `delivery.sent` (revision 1).** It lost because "delivery" is still an - email-shaped abstraction with channels bolted on. Declared types with field roles are genuinely - neutral, and `content.sent` stays the email pack's own type. -- **Hashing undeclared strings (revision 1).** It lost because the hash of an unreviewed field is - still pseudonymous personal data nobody asked for. Dropping it is strictly safer, and declaring - a field is cheap. -- **Wall-clock deadlines (revision 1).** They lost because they are non-deterministic and - host-dependent, and because a timeout that marks a rule unscored rewards flooding. A step - budget plus truncation-as-signal has neither problem. +- **CEL.** It has no windowed aggregates, is a heavy dependency, costs per expression rather than + per history, and gives poorer error messages. It may come back later as a `where` leaf (§12 Q9). +- **A custom expression language, a plugin ABI, or SQL features.** These bring a parser, + unreviewable code, unbounded cost, or weak isolation. +- **Canonical-key bridge (revision 1).** Nothing stored needs it. Registry order gives bit-exactness + without keeping two names alive. +- **Neutral `delivery.sent` (revision 1).** It is still email-shaped. Declared roles are neutral. +- **Hashing undeclared strings (revision 1).** A hash of unreviewed data is still personal data. +- **A shared byte and step budget with a weighted truncation feature (revision 2).** It had four + problems: + - onboarding facts could be displaced; + - totals could fall under truncation; + - one feature's change could move others; + - truncation mixed a data-quality signal into risk. + + Facts, counters, exact aggregates and flags remove all four. +- **First-come `group_by` admission (revision 2).** Early decoys could take every slot. The review + proposed space-saving; revision 3 goes further with exact aggregation, and keeps space-saving + only for the in-memory adapter's memory bound. +- **Luhn on declared numbers (revision 2).** It rejected valid amounts. Declared numbers are + reviewed and must carry `max` instead. ## 6. Edge cases and failure handling -- **A pack errors (a bug).** Rules that read that pack are unscored with `feature_error` and - marked degraded; other rules still score. Hitting a budget or truncating is not an error (§5.7). -- **A profile is rejected.** The tenant's previous profile stays live. On a cold start with no - valid profile, the tenant's scores are `unknown` and `/healthz` is red. Another tenant's rules - are never borrowed. -- **Events arrive before their declaration.** Undeclared data is dropped (§5.6 rule 3). A feature - declared later sees only the type and time of those rows. Warm-up keeps rules that use it in - shadow until its window is fully covered. -- **Unknown values.** - - An undeclared resource kind is treated as `other` and counted in a metric; in `strict` mode it - is rejected. - - An undeclared `subject_kind` or `also` kind is rejected with `redaction_failed`. +- **Pack error (a bug).** Rules that read the pack are unscored with `feature_error` and marked + `degraded`. Hitting a bound isn't an error; it sets `partial` instead. +- **Rejected profile.** The previous profile stays live. On a cold start the tenant's subjects are + `unknown`, and `/healthz` is red. +- **Events that arrive before their declaration.** Their undeclared data is dropped. Counters and + facts start from the declaration, and features stay cold until their horizon passes or a + backfill finishes. +- **Undeclared kinds.** + - An undeclared resource kind becomes `other`; `strict` rejects it instead. + - An undeclared subject or `also` kind is rejected with `redaction_failed`. - An undeclared link kind is rejected with `bad_links`. -- **Absent fields.** A predicate on an absent field is false, and `sum` uses `default`. - `share`/`ratio` use `if_empty`, and `time_between` uses `if_absent`. Every op is total; a - non-finite value is a pack error. -- **Duplicates, out-of-order arrival, ties.** Idempotency is unchanged. Features are set functions - over the history ordered `(at, producer, id)`. -- **Clock skew and future events.** They are excluded from custom windows, but loaded (within - 24 h) so they can schedule a rescore. The legacy exceptions are listed in §5.5. -- **Key rotation during a burst.** Dual keys (§5.6) keep `distinct` and neighbour equality exact. -- **`also` abuse.** A producer naming arbitrary subjects is limited to 3 per event. Each dirty - mark counts against the per-tenant rescore and scoring budgets. -- **Parent fan-in.** Many `api_key` subjects mark the same parent account dirty. Dirty marks - coalesce per subject through `dirty_seq`, so the parent costs O(1) per scoring round. -- **Hostile config.** No code or regex, every set hashed, every size capped, and the loader is - fuzzed. -- **Hostile events.** Byte caps, the step budget, and capped groups and keys bound the cost. - Truncation raises risk rather than lowering it (§5.7). -- **The brand pack on a marketplace that resells branded goods.** `brand.*` stays in shadow until - the tenant's own labels justify it. The tenant may also narrow the brand list (§12 Q10). +- **Absent fields.** + - Predicates on absent fields are false. + - `sum` uses `default`. + - `share` and `ratio` use `if_empty`. + - `time_between` and `sequence` use `if_absent` plus the `__absent` indicator. + - A non-finite result is a pack error. +- **Out-of-order arrival.** Fact updates are monotone, with a recount triggered when the first + success moves earlier. Counters are commutative. +- **Duplicates.** Facts and counters are updated only for a newly accepted event. +- **Future-dated events.** They're counted (legacy) or subtracted exactly (custom), and they + schedule a rescore. +- **Key rotation.** Dual keys keep equality exact. +- **Parent and `also` fan-out.** At most 8 index rows and 8 dirty marks per event, coalesced by + `dirty_seq`. +- **Hostile config or events.** There's no code or regex. Limits, per-feature budgets and exactness + apply, and partial computations are one-sided. +- **Brand-list change.** The facts spec version is bumped and a backfill runs; the affected + features stay cold until it finishes. ## 7. Genericity walk: five scenarios (fictional) -All products, ids and domains are invented; timestamps use the 2031 convention. Each profile is -committed in P6b as `examples/tenants//`, with dev and held-out fixtures. +The products, ids and domains below are invented, and the timestamps are in 2031. They are +committed in P6b with dev and held-out fixtures. Throughout: +- durations use `log1p`; +- every `time_between` declares `if_absent` and `absent_sign`; +- every number declares `max`. ### 7a. File sharing: malware-distribution burst ("Driftbox") -The pattern: a fresh account uploads executables or archives and creates many public links, which -are then downloaded from many distinct networks within the hour. - ```yaml packs: [core@1, brand@1] vocabulary: @@ -846,9 +964,9 @@ vocabulary: share.link_created: role: activity fields: - visibility: {kind: enum, values: [public, org, private]} - file_kind: {kind: enum, values: [document, archive, executable, image, other]} - folder_title: {kind: text, role: title, max_len: 120} # skeleton-only + visibility: {kind: enum, values: [public, org, private]} + file_kind: {kind: enum, values: [document, archive, executable, image, other]} + folder_title: {kind: text, role: title, max_len: 120} share.downloaded: fields: link_hash: {kind: hash} @@ -863,7 +981,7 @@ features: window: 24h, transform: {cap: 1}} - {name: custom.distinct_downloader_nets_1h, version: 1, description: distinct downloader networks, distinct: {type: share.downloaded, where: {field: downloader_is_owner, eq: false}, field: downloader_ip24}, - window: 1h, transform: {log1p: true, cap: 9}} # track_max 8103 + window: 1h, transform: {log1p: true, cap: 9}} - {name: custom.max_downloads_per_link_1h, version: 1, description: busiest link's downloads, count: {type: share.downloaded, where: {field: downloader_is_owner, eq: false}, group_by: {field: link_hash, reduce: max, max_groups: 1000}}, @@ -874,13 +992,14 @@ features: transform: {log1p: true, cap: 6}} - {name: custom.signup_to_first_public_link_min, version: 1, description: minutes to first public link, time_between: {from: {type: subject.created}, to: {type: share.link_created, where: {field: visibility, eq: public}}, - until_now: true, if_absent: 1440}, - transform: {cap: 1440}, prior_sign: "-"} + until_now: false, if_absent: 1440}, + absent_sign: "-", transform: {log1p: true, cap: 7.3}, prior_sign: "-"} rules: - {name: malware_burst, mode: shadow, scorer: local, weights: uniform, inputs: [custom.public_links_1h, custom.risky_link_share_24h, custom.distinct_downloader_nets_1h, custom.max_downloads_per_link_1h, custom.download_peak_vs_history, - custom.signup_to_first_public_link_min, brand.title_match, core.history_truncated], + custom.signup_to_first_public_link_min, custom.signup_to_first_public_link_min__absent, + brand.title_match], labels: [benign, abusive], benign_label: benign, threshold: 0.7} ``` @@ -889,15 +1008,12 @@ rules: {"id":"db-005","subject":"acct_example_db_1","type":"share.downloaded","at":"2031-03-02T09:05:02Z","data":{"link_hash":"lk_4f1c9a0e7b2d11","downloader_ip24":"ip_9a1b2c3d4e5f66","downloader_is_owner":false}} ``` -**Declarative:** everything above. **Go-only:** file-content verdicts, such as a malware hash or a -sandbox result. The product emits these as `content.verdict`, and `core.verdict_max_24h` reads -them. +- **Declarative:** everything above. +- **Go-only:** file-content verdicts. The product emits these as `content.verdict`, and + `core.verdict_max_24h` reads them. ### 7b. Marketplace: card testing ("Tallyport") -The pattern: a merchant account pushes many small charge attempts across many cards, and most are -declined. A card that turns up across many merchants is suspicious in itself. - ```yaml packs: [core@1, brand@1] vocabulary: @@ -910,7 +1026,7 @@ vocabulary: fields: outcome: {kind: enum, values: [succeeded, declined, blocked]} decline_code: {kind: enum, values: [insufficient_funds, do_not_honor, incorrect_cvc, expired_card, fraudulent, other]} - amount_minor: {kind: number, min: 0, integer: true} + amount_minor: {kind: number, min: 0, max: 100000000, integer: true} card_hash: {kind: hash, join_domain: card} features: - {name: custom.max_declines_per_card_1h, version: 1, description: most declines on one card, @@ -927,7 +1043,7 @@ features: count: {type: charge.attempted, where: {field: amount_minor, lte: 200}}, window: 1h, transform: {cap: 500}} - {name: custom.charges_1h, version: 1, description: all charges, count: {type: charge.attempted}, window: 1h, transform: {cap: 500}} - - {name: custom.small_charge_ratio_1h, version: 1, description: small ÷ all charges, + - {name: custom.small_charge_ratio_1h, version: 1, description: small ÷ all charges (pre-transform), ratio: {num: custom.small_charges_1h, den: custom.charges_1h, if_empty: 0}, transform: {cap: 1}} - {name: custom.declines_10m_peak, version: 1, description: declines in busiest 10 min today, peak: {type: charge.attempted, where: {field: outcome, eq: declined}, size: 10m}, @@ -938,30 +1054,25 @@ features: rules: - {name: card_testing, mode: shadow, scorer: local, weights: uniform, applies_to: [account], inputs: [custom.max_declines_per_card_1h, custom.cards_with_3plus_declines_1h, custom.distinct_cards_1h, - custom.small_charge_ratio_1h, custom.declines_10m_peak, core.credential_velocity_1h, - core.history_truncated], labels: [benign, abusive], benign_label: benign, threshold: 0.7} + custom.small_charge_ratio_1h, custom.declines_10m_peak, core.credential_velocity_1h], + labels: [benign, abusive], benign_label: benign, threshold: 0.7} - {name: tested_card, mode: shadow, scorer: local, weights: uniform, applies_to: [card], inputs: [custom.card_merchants_24h], labels: [benign, abusive], benign_label: benign, threshold: 0.7} ``` -Each charge is also indexed under the card subject, through `also`: - ```json {"id":"tp-101","subject":"acct_example_tp_7","also":[{"kind":"card","id":"card_example_c1"}],"type":"charge.attempted","at":"2031-06-10T02:14:01Z","data":{"outcome":"declined","decline_code":"incorrect_cvc","amount_minor":100,"card_hash":"cd_1a2b3c4d5e6f77"}} ``` -`x_primary_subject_hash` is a **derived** field. Ingest adds it to every row indexed through -`also`, as the re-HMACed id of the primary subject (join domain `subject:`). It is declared -implicitly for any type used with `also`, so a secondary subject can count distinct primaries. - -**Declarative:** everything above. **Go-only:** issuer and BIN intelligence, and velocity seen by -external card networks. Products can supply these as `content.verdict` or enum fields. +- `card_example_c1` is stored as an `hs…` pseudonym. +- `x_primary_subject_hash` is derived at ingest for rows reached through `also`. It is the primary + subject's pseudonym, under the join domain `subject:`. +- **Declarative:** everything above. +- **Go-only:** issuer/BIN intelligence and external network velocity. The product can emit these as + `content.verdict` or enum fields. ### 7c. Chat or community: spam invites ("Hearthchat") -The pattern: new accounts with brand-like names invite people outside their own communities, -often with links, and the recipients block them soon after. - ```yaml packs: [core@1, brand@1] vocabulary: @@ -974,7 +1085,7 @@ vocabulary: invitee_hash: {kind: hash, join_domain: member} target_class: {kind: enum, values: [own_community, other_community]} preview: {kind: text, role: title, max_len: 200} - link_host: {kind: domain, reduce: etld1} + link_host: {kind: domain} # etld1; cleartext only if allowlisted block.received: fields: {blocker_hash: {kind: hash, join_domain: member}} features: @@ -984,7 +1095,7 @@ features: share: {type: invite.sent, match: {field: target_class, eq: other_community}}, window: 24h, transform: {cap: 1}} - {name: custom.invites_blocked_within_10m, version: 1, description: invitees who blocked within 10 min, sequence: {a: {type: invite.sent}, b: {type: block.received}, within: 10m, on: {a: invitee_hash, b: blocker_hash}}, - window: 24h, transform: {log1p: true, cap: 6}} + if_absent: 0, absent_sign: "+", window: 24h, transform: {log1p: true, cap: 6}} - {name: custom.linked_invite_share_24h, version: 1, description: invites with links, share: {type: invite.sent, match: {field: link_host, exists: true}}, window: 24h, transform: {cap: 1}} - {name: custom.phone_siblings_7d, version: 1, description: accounts sharing a phone created this week, @@ -993,21 +1104,15 @@ rules: - {name: invite_spam, mode: shadow, scorer: local, weights: uniform, inputs: [custom.invites_10m_peak, custom.external_invite_share_24h, custom.invites_blocked_within_10m, custom.linked_invite_share_24h, custom.phone_siblings_7d, brand.name_match, - brand.title_match, core.linked_deleted_n, core.history_truncated], + brand.title_match, core.linked_deleted_n], labels: [benign, abusive], benign_label: benign, threshold: 0.7} ``` -`invitee_hash` and `blocker_hash` share `join_domain: member`. The same member id therefore hashes -identically in both types, and `sequence.on` can join them. - -**Declarative:** everything above. **Go-only:** classifying message text. The product emits that -as `content.verdict`. +- **Declarative:** everything above. +- **Go-only:** message-text classification, which the product emits as `content.verdict`. ### 7d. Developer API: credential stuffing through customer API keys ("Keyforge") -The pattern: a customer's API key drives many end-user login attempts across many distinct -usernames. Most fail, with the occasional success shortly after a failure on the same username. - ```yaml packs: [core@1] vocabulary: @@ -1037,7 +1142,7 @@ features: sequence: {a: {type: auth.attempted, where: {field: outcome, eq: failed}}, b: {type: auth.attempted, where: {field: outcome, eq: succeeded}}, within: 10m, on: {a: login_hash, b: login_hash}}, - window: 24h, transform: {log1p: true, cap: 6}} + if_absent: 0, absent_sign: "+", window: 24h, transform: {log1p: true, cap: 6}} - {name: custom.hosting_share_1h, version: 1, subject_kinds: [api_key], description: attempts from hosting networks, share: {type: auth.attempted, match: {field: client_asn, eq: hosting}}, window: 1h, transform: {cap: 1}} - {name: custom.max_attempts_per_ip_1h, version: 1, subject_kinds: [api_key], description: busiest client network, @@ -1049,23 +1154,20 @@ features: rules: - {name: stuffing_key, mode: shadow, scorer: local, weights: uniform, applies_to: [api_key], inputs: [custom.distinct_logins_10m, custom.failure_ratio_1h, custom.success_after_failure_10m, - custom.hosting_share_1h, custom.max_attempts_per_ip_1h, custom.attempts_vs_history, - core.history_truncated], labels: [benign, abusive], benign_label: benign, threshold: 0.7} + custom.hosting_share_1h, custom.max_attempts_per_ip_1h, custom.attempts_vs_history], + labels: [benign, abusive], benign_label: benign, threshold: 0.7} ``` ```json {"id":"kf-9001","subject":"key_example_k3","subject_kind":"api_key","type":"auth.attempted","at":"2031-09-04T11:00:01Z","data":{"outcome":"failed","login_hash":"lg_8c7b6a5f4e3d21","client_ip24":"ip_1f2e3d4c5b6a77","client_asn":"hosting"}} ``` -**Declarative:** everything above. `client_asn` is an enum the product classifies. **Go-only:** -whether a username appears in a breach corpus (this needs an external lookup), and ASN reputation -finer than the product's own enum. +- The account-level `failures_1h` sees events from the account's keys through `via_parent` rows. +- **Declarative:** everything above. +- **Go-only:** breach-corpus checks, and ASN reputation beyond the enum. ### 7e. AI inference: free-tier farming ("Lumenloop") -The pattern: many free accounts share a device or an OAuth identity. Each one exhausts its free -token quota soon after sign-up, favours the most expensive models, and is then abandoned. - ```yaml packs: [core@1] vocabulary: @@ -1075,14 +1177,15 @@ vocabulary: usage.recorded: role: activity fields: - model_tier: {kind: enum, values: [small, medium, large]} - tokens: {kind: number, min: 0, integer: true} - quota_state: {kind: enum, values: [ok, near_limit, exhausted]} + model_tier: {kind: enum, values: [small, medium, large]} + tokens: {kind: number, min: 0, max: 2000000, integer: true} + quota_state: {kind: enum, values: [ok, near_limit, exhausted]} features: - {name: custom.signup_to_quota_exhausted_min, version: 1, description: minutes to exhaust free quota, time_between: {from: {type: subject.created}, to: {type: usage.recorded, where: {field: quota_state, eq: exhausted}}, - until_now: false, if_absent: 10080}, transform: {cap: 10080}, prior_sign: "-"} - - {name: custom.large_model_token_share_24h, version: 1, description: tokens spent on large models, + until_now: false, if_absent: 10080}, + absent_sign: "-", transform: {log1p: true, cap: 9.3}, prior_sign: "-"} + - {name: custom.large_model_token_share_24h, version: 1, description: tokens on large models (both sides summed), share: {type: usage.recorded, match: {field: model_tier, eq: large}, sum: {field: tokens, default: 0, cap_each: 200000}}, window: 24h, transform: {cap: 1}} - {name: custom.tokens_10m_peak, version: 1, description: tokens in busiest 10 min, @@ -1090,248 +1193,300 @@ features: window: 24h, transform: {log1p: true, cap: 15}} - {name: custom.minutes_since_last_use, version: 1, description: idle time after last use, time_between: {from: {type: usage.recorded, anchor: last}, to: {type: abusekit.never}, until_now: true, if_absent: 0}, - transform: {cap: 10080}} + absent_sign: "+", transform: {log1p: true, cap: 9.3}} # hash_quantum defaults to 60 (until_now) - {name: custom.oauth_siblings_deleted, version: 1, description: deleted accounts sharing the OAuth identity, - neighbours: {via: [oauth_sub_hash, device_hash], where: {deleted: permanent}}, transform: {cap: 20}} + neighbours: {via: [oauth_sub_hash], where: {deleted: permanent}}, transform: {cap: 20}} - {name: custom.oauth_siblings_new_7d, version: 1, description: accounts sharing identity created this week, - neighbours: {via: [oauth_sub_hash, device_hash], where: {created_within: 7d}}, transform: {cap: 20}} + neighbours: {via: [oauth_sub_hash], where: {created_within: 7d}}, transform: {cap: 20}} rules: - {name: free_tier_farm, mode: shadow, scorer: local, weights: uniform, - inputs: [custom.signup_to_quota_exhausted_min, custom.large_model_token_share_24h, custom.tokens_10m_peak, - custom.oauth_siblings_deleted, custom.oauth_siblings_new_7d, core.first_funding_prepaid, - core.linked_deleted_n, core.history_truncated], + inputs: [custom.signup_to_quota_exhausted_min, custom.signup_to_quota_exhausted_min__absent, + custom.large_model_token_share_24h, custom.tokens_10m_peak, custom.minutes_since_last_use, + custom.oauth_siblings_deleted, custom.oauth_siblings_new_7d, core.linked_deleted_n], labels: [benign, abusive], benign_label: benign, threshold: 0.7} ``` -`abusekit.never` is a reserved type that never occurs. With `anchor: last` and `until_now`, -`time_between` becomes "minutes since the last event of type A". That gives the idle-time -primitive with no new op. - -`custom.minutes_since_last_use` is deliberately left out of the rule. Abandonment only means -something alongside the neighbour counts, and uniform priors can't express that combination. The -feature is kept for hand-tuned weights later. - -**Declarative:** everything above. **Go-only:** -- prompt-content similarity across accounts, which needs cross-subject text clustering and a text - scorer; -- feature interactions ("abandoned **and** has farmed siblings"). A uniform-prior logistic model - can't capture these; they need tuned or fitted weights, or a multiplicative custom feature. §12 - Q15 asks whether `ratio` should gain a `product` form. +- `abusekit.never` is a reserved type that never occurs. Combined with `anchor: last` and + `until_now`, it measures time since the last use. +- Absence has its own indicator, so an account that never exhausts its quota isn't read as the most + benign case. +- `neighbours` uses only the declared `oauth_sub_hash`, because `device_hash` already feeds + `core.linked_deleted_n`. +- **Declarative:** everything above. +- **Go-only:** prompt similarity across accounts, and non-linear interactions under uniform priors + (§12 Q15). ### 7f. What the walk shows | Need | Primitive | | --- | --- | -| Per-entity maxima and counts (per card, per link, per IP) | `group_by` | -| Cause, then effect within a time limit (fail → success, invite → block, signup → exhaustion) | `sequence`, `time_between` | -| Rates | `ratio` | -| Cross-account farms | Declared link kinds + `neighbours` | -| Non-account actors (cards, API keys) | Subject kinds + `also` + `parent` | +| Per-entity maxima and counts | `group_by` (exact) | +| Cause, then effect within a time bound | `sequence`, and `time_between` with `__absent` | +| Rates | `ratio` over pre-transform values | +| Cross-account farms | Declared link kinds and `neighbours` (as of now) | +| Non-account actors | Subject kinds, `also` and `parent` | -**Still Go-only across all five:** -- content understanding (files, messages, prompts); -- external reputation lookups; -- text similarity across subjects; -- non-linear feature interactions under uniform priors. +Still Go-only across all five scenarios: +- content understanding; +- external reputation; +- cross-subject text similarity; +- non-linear interactions under uniform priors. -The first two already have a channel: products emit `content.verdict` or enum fields. +The first two already have a channel: `content.verdict` and enum fields. ## 8. Migration plan for e2a -- **Emitter.** No change. S6 emits `content.sent` and the rest of main §4.12 as designed. -- **Profile.** e2a's private profile starts as a byte copy of `examples/tenants/reference/`: - - `packs: [core@1, email@1, brand@1]`; - - the implicit legacy vocabulary, made explicit: `key: credential` with #7's aliases, and - `agent: other`; - - `new_account_velocity` with namespaced inputs. - - The weights move to the private mount. The private brand list stays private (`brand.extra`). - Floors for e2a's real corpus live privately; the public reference floors stay here. -- **Golden (P0).** `abusekit eval --golden out.jsonl` extends the existing eval replay rather than - adding a new command. It runs on `main` after #5 and #7 merge, and its output is committed as - `eval/golden/reference-flat.jsonl`. For every fixture, subject, event instant and scheduled - rescore instant, it records: - - feature values, as `Float64bits`; - - `NextRescoreAt`; - - per-rule input hashes, risks and tiers; - - the score; - - the local `Version()`. -- **Rename (P1).** P1 re-baselines under the semantic-identity rules of criterion 1: - - identical feature bits under the rename map; - - identical `NextRescoreAt` and tiers; - - `|Δrisk| ≤ 1e-12`, with no score near a cut point; - - every hash recorded as changed, exactly once. - - The result is committed as `eval/golden/reference-ns.jsonl`. Every later slice asserts - **exact** equality with that file. -- **Other one-time baseline changes, each isolated in its own slice:** - - the `(at, producer, id)` tie order, in P0, before capture; - - re-HMAC of hash values and links, in P3: feature values unchanged, stored bytes changed; - - the domain kind at `RedactionSchemaVersion` 3, also in P3. Fixtures use `.test`, so nothing - changes. -- **Rollback.** Before S8, each slice can be reverted on its own. After S8, the rename can't be - undone without re-scoring, which is why it lands first. +- **Emitter:** no change. S6 emits `content.sent`. +- **Profile:** a byte copy of `examples/tenants/reference/`. + - Packs: `packs: [core@1, email@1, brand@1]`. + - The legacy vocabulary made explicit: `key: credential` with #7's aliases, and `agent: other`. + - Weights and the private brand list stay in the private mount. +- **P0:** `abusekit eval --golden` extends the existing replay. + - It runs the current unbounded evaluator, ordered by `(at, producer, id)`. + - It records bits for feature values, `NextRescoreAt`, per-rule hashes, risks, tiers, the score + and `Version()`. + - The output goes to `eval/golden/reference-flat.jsonl`. +- **P1:** values, risks and tiers stay bit-exact under the rename map. Hashes, cassette keys, + fake-scorer outputs and SHAs are recorded as changed once. The output goes to + `reference-ns.jsonl`, and every later slice must match it exactly. +- **Later slices, each checked against that golden:** + - P3a: masking `subject_line` for card, IP and phone. + - P3b: re-HMAC and domain HMAC. Both are injective, and webmail domains stay in cleartext. + - P4a: facts and counters replace scans. + - P4c: bounded evaluation. The fixtures sit far below every saturation limit. + + Any fixture that changes is listed and justified in its slice. +- **Bounded evaluation reaches e2a only after P4a–P4c** (§5.7 f). +- **Rollback:** before S8, every slice can be reverted on its own. After S8, only the rename is + irreversible without a rescore, which is why P1 comes first. ## 9. Slices -These come after #5 and #7 merge. The early slices are small, and pack gating arrives only after -the machinery exists. +These come after #5 and #7 merge. P1 lands before any S5 or S8 PR (enforced as described in §5.2). | # | Slice | Contents | Depends on | Done when | | --- | --- | --- | --- | --- | -| P0 | Golden replay | `abusekit eval --golden`; scoring loader ordered `(at, producer, id)`; `eval/golden/reference-flat.jsonl` | #5, #7 | Golden committed; test green; flipping one weight's last bit fails it | -| P1 | One-time rename | Every §5.2 consumer, each with its test; the `FeatureDef` metadata table (quantum, bound, prior sign, truncation direction, reads); `core.Vector`; reason v2; the corpus key-space column and migration; the cassette header; `feature_renamed`; stage gate limited to advise-mode local rules | P0 | `reference-ns.jsonl` meets criterion 1; the grep test finds no flat literals; all existing suites green | -| P2 | Tenant profiles | Private-mount loader; `examples/tenants/reference`; per-tenant atomic reload and `/healthz`; per-tenant rule sets (`computeVerdict`, `currentRuleNames`); per-tenant local scorer and version; per-tenant fair queue and concurrency cap. Every built-in feature is available to every tenant. | P1 | Golden exact; tenant-isolation tests (rules, reload, view); fairness test | -| P3 | Declared types and kind redaction | `internal/vocab`; `internal/secret` (`Keys`, file adapter); field kinds; roles (`credential`/`other`, `activity`, `title`, `self`); `x_` extensions; the PSL-based domain kind; card/IP/phone scan; re-HMAC of all hash fields and links, with `join_domain`; undeclared values dropped; skeleton-only custom text; `vocab_version`; config history plus `abusekit config check` | P2 | Criterion 5 property tests; golden exact (features unchanged); `vocab_incompatible` and history CI tests | -| P4a | DSL core: `count`, `distinct`, `share`, `peak` | Compiler; windows (`window`, `first`, `lifetime`); predicates; transforms; caps and limits; the bounded loader (§5.7) and step budget; `core.history_truncated` with its truncation invariant; end-to-end benchmark and `cost_table.yaml` | P3 | Reference equality on 10k histories for these four ops; criterion 4 at P4a limits; criterion 6 flood property; loader fuzz | -| P4b | DSL extended: `time_between`, `before_first`, `relative_to_history`, `group_by`, `sequence`, `ratio` | Plus proportional rescore coalescing, the per-tenant rescore budget, and warm-up | P4a | Reference equality for every op; every "expressible" #7/S2 feature equals its Go twin bit for bit; warm-up test | -| P5 | Pack gating | `internal/pack` registry; `core`/`email`/`brand` adapters (code moved); enablement and dependency validation; `brand.title_match`; `packtest`; starter weights | P2 (P4a for `packtest`'s flood check) | Golden exact; `feature_not_enabled`; every pack passes `packtest` | -| P6a | Link kinds, `neighbours`, subject kinds | `links.custom`; `neighbours`; `subject_kind`, `also`, `parent`; `event_subjects`; `?kind=`; `applies_to`; derived `x_primary_subject_hash` | P4a, P5 | Contract tests for kinds and `also`; neighbour caps; golden exact | -| P6b | Scenarios and bootstrap | The five `examples/tenants/*` profiles, with dev and held-out fixtures; uniform priors; `--profile`; corpus v2; floors with `profile:` | P4b, P6a | Criterion 2 on held-out fixtures; the held-out isolation CI check | -| P7 | e2a cutover | e2a's profile in the ops repo's private mount (outside this repo); hosted-config CI runs `abusekit config check`; the `rules.yaml` path removed | P5 (and S8's mount) | Golden exact against the private copy; the hosted deploy loads it | - -- **Ordering against the v0 plan.** P0 and P1 must land before S5 and S8, because assumption A1 - is what makes a rename without a bridge safe. -- **S3b (erasure)** must be vocabulary-aware (§5.10), and is easiest to build after P3. -- **S6** is independent of every P slice. +| P0 | Golden replay | `abusekit eval --golden`; `(at, producer, id)` order; `reference-flat.jsonl` | #5, #7 | Golden committed; flipping one weight's last bit fails it | +| P1 | One-time rename | Every §5.2 consumer and its test; `FeatureDef`; `core.Vector`; registry-order summation; fake-scorer re-baseline; `corpus-v2` + `feature_renamed` in `LoadSnapshotCorpus` and `score --jsonl`; reason v2; corpus key-space migration; cassette header; `TestNoVendorAdapterBeforeRename` | P0 | `reference-ns.jsonl` bit-exact for values, risks and tiers; grep test clean | +| P1s | Stage gate | `maxRiskByScorer` limited to advise-mode local rules | P1 | Stage tests pass; golden exact | +| P2 | Tenant profiles | Private-mount loader; reference profile; per-tenant reload, `/healthz`, rule sets and scorer version; fair queue and concurrency cap. All built-in features available to every tenant. | P1 | Golden exact; isolation and fairness tests | +| P3a | Vocabulary and scans | `internal/vocab`; kinds (numbers require `max`); roles; `x_` fields; declared-domain PSL, `etld1` and IP checks; card/IP/phone scans with digit folding; `subject_line` masking; egress scan; drop-undeclared with name grammar; account-subject leak scan; ASN grammar; profile-load scans | P2 | Criterion 5 (non-key parts); golden exact or deviations justified | +| P3b | Keys and re-HMAC | `internal/secret` (HKDF, file adapter); re-HMAC of hashes and links with `join_domain`; domain allowlist/HMAC (built-in and declared); pseudonymised non-account and `also` ids; `RedactionSchemaVersion` 3; dev/staging migration job; cross-tenant key test | P3a | Criterion 5 complete; golden exact; migration test | +| P3c | Config history | `history/`; `abusekit config check`; `tenant_config_versions` | P2 | History CI tests | +| P3d | Rotation and cloud keys | Dual-key write, read and flip; `__prev` fields; `key_id`; cloud secret-manager adapter | P3b | Equality exact across a simulated rotation; adapter contract test | +| P4a | Facts and counters | `subject_facts`, `subject_counters`; ingest transaction; onboarding, brand and self-send facts; anchored freeze; backfill and cold state; Go features moved onto classes F and N | P3a | Golden exact; blocked-payment flood test; out-of-order fact tests | +| P4b | Aggregate engine (class A) | `internal/evalengine` (Postgres and in-memory); `count`, `sum`, `share`, `distinct`, `group_by`, `time_between` + `__absent`, `sequence`, `ratio`; pushdown and indexes; conformance | P4a | Reference equality on 10k histories; adapter equivalence | +| P4c | Class R, budgets, flags, flood | `peak` (with `sum`) under the saturation limit; `relative_to_history` with its baseline budget; `partial` in verdicts and the API; pass order; `TestFeatureIndependence`; flood generator and property; cost benchmark. **Then enable bounded evaluation for e2a.** | P4b | Criteria 4 and 6; golden exact under the bounded evaluator | +| P4d | Rescore control and warm-up | Proportional coalescing; timers only from non-shadow rules; per-tenant budget; warm-up | P4c, P3c | Storm and warm-up tests | +| P5 | Pack gating | Registry; `core`, `email` and `brand` adapters; enablement; `brand.title_match` (needs the `title` role); `packtest`; starter weights | P3a, P4c | Golden exact; `feature_not_enabled`; every pack passes `packtest` | +| P5b | DSL parity | Every expressible #7/S2 feature re-expressed in the DSL | P4c | Bit-exact against the Go feature on every fixture | +| P6a | Link kinds, `neighbours`, subject kinds | `links.custom`; `neighbours` (as of now); dirty propagation for declared kinds; `subject_kind`, `also`, `parent`, `event_subjects` (≤ 8); `?kind=`; `applies_to`; `x_primary_subject_hash` | P3b, P4b, P5 | Contract tests; propagation and no-double-feed tests; golden exact | +| P6b | Scenarios and bootstrap | Five example profiles with held-out fixtures; uniform priors; `--profile`; floors with `profile:` | P4d, P5, P6a | Criterion 2 on held-out fixtures; isolation check | +| P7 | e2a cutover | Private profile in the ops mount; hosted `config check` | P5, P4c, S8's mount | Golden exact against the private copy | + +Two v0 slices interact with this plan: +- S3b must be vocabulary-aware and must clear facts and counters. It is easiest after P4a. +- S6 is independent of all of the above. ## 10. Scalability and extensibility -- **Per-subject cost** is bounded by bytes and steps, not by event count. The loader fetches only - as far back as the profile's largest lookback, capped at 8 MiB of history plus 1 MiB of - onboarding events. Compiled plans are cached per `profile_sha`. -- **Tenants.** The fair queue and the per-tenant concurrency cap let about 100 tenants share one - worker pool without starving each other. Metrics are labelled `{tenant, pack}`. -- **Rescores.** - - Timer rescores come only from features that feed non-shadow rules. - - Their coalescing buckets scale with the feature's window. - - Each tenant has an hourly budget. -- **Neighbours.** At most 4 queries per extraction; each is indexed, fan-in capped, and cached for - the extraction. -- **Made easier later:** +- **Per subject:** facts and counters are O(1). Each feature runs one aggregate, costing O(window + rows) in the database. Row features run under per-feature budgets. Plans are cached per + `profile_sha`. +- **Floods:** a flood raises database scan cost, never values. The fair queue and the slow-subject + rescore limit keep one subject from monopolising workers. +- **Tenants:** about 100 tenants share the fair queue. Metrics are labelled `{tenant, pack}`. +- **Storage:** one facts row per subject, one counter row per spec per active day, and at most 8 + index rows per event. +- **Extensibility:** - new Go packs, gated by `packtest`; - - promoting a popular custom feature into a pack upstream; + - promoting a custom feature into a pack; - weight fitting on corpus v2; - - CEL as a `where` leaf; - - cross-tenant linking, which this design leaves untouched. + - a CEL `where` leaf. ## 11. Verification strategy -Tests sit at the seams callers actually cross: -- the profile loader: profiles in, errors out; -- `LoadHistory` + `feature.Extract`: history in, vector out; -- `POST /v1/events`: redaction; +**Seams under test:** +- the profile loader; +- ingest (redaction, facts, counters); +- `feature.Extract`, with both aggregate-engine adapters; - `abusekit eval --profile`. -The checks: -1. **Golden replay.** Criterion 1 in P1, then exact equality in every later slice. -2. **`packtest`** for every pack (§5.9). -3. **DSL conformance.** The naive reference against the compiled evaluator. Table tests for each - op's edge cases, clipping, `before_first` inclusivity, `track_max`, and the group and key caps. -4. **Redaction property tests.** - - Mask versus reject, by kind. - - Re-HMAC, including `join_domain` equality and separation. - - Undeclared values dropped. - - Domain PSL and IP checks. - - Luhn, IP and phone detection, with explicit false-positive fixtures. -5. **Loader fuzzer**, with the §5.7 limits as the oracle, plus config-history CI tests. -6. **Performance and flooding.** The end-to-end benchmark (criterion 4) and the flood property - (criterion 6). -7. **Tenant isolation.** Rules, reload, view, the fair queue and the rescore budget. -8. **HTTP contract tests.** `subject_kind`, `also`, `links.custom`, `?kind=` and `feature_renamed`. -9. **Most likely regressions, and what catches each:** +**Checks:** +1. The golden replay: bit-exact in P1 (hashes change once), then exact in every later slice. +2. `packtest` for every pack. +3. DSL conformance: the naive reference against both adapters. Edge tables cover: + - clipping; + - `before_first`; + - the `peak` saturation limit; + - the space-saving fallback; + - `__absent`; + - counter subtraction of future events. +4. Ingest: + - fact monotonicity under shuffled arrival; + - the recount when the first success moves; + - duplicates never touching facts; + - the blocked-payment flood. +5. Redaction: + - mask vs reject; + - digit folding; + - the egress scan; + - undeclared-name handling; + - re-HMAC equality and `join_domain` separation; + - cross-tenant key inequality; + - domain allowlist vs HMAC; + - subject-id pseudonyms. +6. `TestFeatureIndependence` and the flood property. +7. The loader fuzzer (limits as oracle) and config-history CI. +8. Tenant isolation and scheduling. +9. HTTP contracts: `subject_kind`, `also`, `?kind=`, `partial`, `bad_subject`, `feature_renamed`. + +**Likely regressions and what catches them** | Regression | Caught by | | --- | --- | -| A stage lookup missed by the rename | grep test + stage test | -| A lost hash quantum | hash-drift test | -| Tie order | golden | -| A weights edit that breaks the truncation invariant | `packtest` | -| PSL snapshot drift | pinned version + test | +| Summation-order drift | The golden | +| A stage literal missed in the rename | The grep test and the stage test | +| Facts diverging from a scan | The P4a golden and the naive reference | +| Pushdown SQL diverging from the in-memory adapter | Adapter equivalence | +| PSL drift | A pinned snapshot and its test | ## 12. Open questions (owner decisions) -Where the review's answer differs from revision 1's recommendation, both are shown. - -1. **Bridge vs rename.** Revision 1: a frozen canonical-key bridge. Review, and now recommended: - rename once in P1 and re-baseline under semantic identity. Reason: nothing stored needs a - bridge yet (A1), and a bridge would keep two names alive forever. Approve? -2. **Ordering.** Revision 1: G0–G3 before S5 and S8. Now: P0 and P1 **must** land before S5 and - S8, because the no-bridge rename depends on it. Approve? -3. **Undeclared data.** Revision 1: keyed-hash undeclared strings. Review, and now: drop every - undeclared value and keep only type, time and field names. Reason: the hash of an unreviewed - field is still personal data. Approve? -4. **Built-in re-HMAC.** Re-HMAC `recipient_hash` and every `links` value at ingest. Stored bytes - change; feature values don't. Approve? -5. **Legacy window quirks.** Freeze them in `@1` and harmonise in `@2` (recommendation unchanged). - Confirm? -6. **Roles.** Revision 1: five resource roles. Review, and now: only `credential` and `other`, - plus the `activity` type role and the `title`/`self` field roles. Approve? -7. **Load bounds.** Revision 1: the 50,000 newest events and a 250 ms deadline. Review, and now: a - time-bounded load, onboarding types in full, byte caps of 8 MiB and 1 MiB, a calibrated step - budget, and truncation as a positive signal. Confirm the defaults? -8. **Bootstrap.** Uniform priors, shadow only, no fitting, and held-out fixtures. Confirm? -9. **CEL.** Add it later as a `where` leaf only, or require a new design pass? (Unchanged.) -10. **Brand list.** May a tenant narrow the list as well as extend it? (Unchanged.) -11. **Custom namespace.** `custom.*` per tenant, or `.*`? (Unchanged.) -12. **Config history.** Revision 1: history held in the DB. Review, and now: history in the config - tree, checked in CI, with the DB as a runtime guard only; atomic reload per tenant. Approve? -13. **What S6 emits.** Revision 1 offered a choice between `content.sent` and `delivery.sent`. - Review, and now: `delivery.sent` is dropped, so S6 emits `content.sent`. Settled unless you - object. -14. **Domain kind.** Require a PSL suffix or an RFC 6761 special-use name, and reject IP literals - and all-numeric labels, **including on built-in domain fields** (`RedactionSchemaVersion` 3). - Declared fields also get an optional `reduce: etld1`. Approve? -15. **Feature interactions.** Should `ratio` gain a `product` form (depth-1 DAG, capped) for the - interactions that §7e shows uniform priors can't capture, or should that wait for fitted - weights? -16. **Subject kinds and `also`.** Add the wire fields `subject_kind` and `also` (at most 3), with - subjects keyed `(tenant, kind, id)`. Both are additive. Approve? -17. **Stage-gate fix.** Stage gates consider only advise-mode local rules. This changes behaviour - on `main`. Approve? -18. **Profiles and secrets.** Tenant profiles live in a private mount, with only fictional - examples in the repo, and the key interface is provider-agnostic. Approve? -19. **Rescore control.** Proportional coalescing applies to DSL features only (built-ins keep - 5 minutes for semantic identity). Only non-shadow rules schedule timer rescores, under a - per-tenant hourly budget. Confirm the default of 20 × active subjects per hour? - -## 13. Changes from revision 1 - -**Blockers:** -- **B1:** the canonical-key bridge is dropped. The rename happens once, with a test for every - consumer it touches (§5.2), and the golden checks semantic identity against a derived - floating-point bound. -- **B2:** the event-count bound and the wall-clock deadline are replaced (§5.7) by: - - a time-bounded load, with onboarding types loaded in full; - - byte caps and a deterministic step budget; - - truncation as a positively weighted signal, with a checked invariant and an argument that - flooding can't evade. -- **B3 (§5.6):** - - every hash is re-HMACed at ingest, with length prefixes and `join_domain`; - - undeclared values are dropped; - - domains are checked against the PSL and rejected if they are IP literals; - - the leak scan covers card, IP and phone shapes; - - custom text is stored skeleton-only; - - stored values are described as pseudonymised throughout. -- **B4:** new primitives `group_by`, `sequence`, `ratio`, `neighbours` with declared link kinds, - `before_first`, `anchor: last`, and subject kinds with `also` and `parent`. All five scenarios - are walked, and what remains Go-only is stated (§7). - -**Should-fix:** -- `delivery.sent` is dropped in favour of declared types with `title`/`self` roles and an - `activity` type role; `x_` extension fields are added. -- Warm-up, and dual-key rotation. -- Exact `relative_to_history` equations, with a list of what the DSL can't express. -- A benchmark-calibrated cost table and a per-type fan-out cap. -- Scheduling: rescore-storm control, a fair queue, and stage gates limited to advise-mode rules. -- Config history in the config tree; profiles in a private mount. -- `(at, producer, id)` tie-breaks; held-out fixtures; `PriorSign` on `FeatureDef`; per-tenant - rule names. -- Re-slicing into P0 → P7; S3b erasure made vocabulary-aware. - -**Nits:** -- Only the `credential`/`other` resource roles. -- A provider-agnostic `Keys` interface. -- A per-tenant scorer version. -- `peak` clipping specified. -- Uniform-prior normalisation written out. +Where the re-review's (R3) answer differs from the earlier recommendation, both are shown. + +1. **Rename vs bridge:** rename once in P1. The R3 review agrees and asks for a bit-exact golden. + **Now:** registry-order summation makes it bit-exact, and the 1e-12 tolerance is deleted. + Approve? +2. **Ordering:** P0 and P1 before any S5 or S8 PR, enforced in the plan and by a test. Approve? +3. **Undeclared data:** drop the values and keep only the names. + - R3: also constrain the names. + - **Now:** names must match `^[a-z0-9_]{1,64}$`, at most 32 per event. + + Approve? +4. **Re-HMAC of built-in fields:** `recipient_hash` and every link, with a dev/staging migration + job. Approve? +5. **Legacy window quirks:** freeze them in `@1` and harmonise in `@2`. Unchanged. Confirm? +6. **Roles:** `credential`, `other`, `activity`, `title`, `self`. Unchanged. Confirm? +7. **Bounding:** + - Revision 2: byte caps, a shared step budget and a weighted truncation feature. + - R3: ingest facts, counters, per-feature budgets, and truncation as a flag. + - **Now:** R3, plus exact per-feature aggregates. The only bounded pieces left are the + one-sided `relative_to_history` baselines and the in-memory `group_by` fallback. + + Confirm the 50,000-row baseline budget and `max_groups` of 1,000? +8. **Bootstrap:** uniform priors, shadow-only, no fitting, held-out fixtures, and `__absent` + indicators with `absent_sign`. Confirm? +9. **CEL:** later as a `where` leaf only, or a new design pass? Unchanged. +10. **Brand list:** can a tenant narrow it as well as extend it? Unchanged. +11. **Custom namespace:** `custom.*` per tenant, or `.*`? Unchanged. +12. **Config history:** in the config tree, checked in CI; the database is only a guard. Approve? +13. **S6:** emits `content.sent`. Settled. +14. **Domains:** + - Revision 2: public-suffix check, stored in cleartext everywhere. + - R3: default to eTLD+1, keep cleartext only for allowlisted or popular domains, and HMAC the + rest, including built-in `recipient_domain` and `first_link_host`. + - **Now:** R3, **except** built-in `recipient_domain` uses `reduce: none` plus HMAC. Reducing + it to eTLD+1 would merge distinct fixture domains and change + `email.first_day_distinct_domains`, whereas HMAC is injective. + + Also: accept that text scorers see a token instead of an unknown `first_link_host`? +15. **Feature interactions:** should `ratio` gain a capped `product` form, or should that wait for + fitted weights? +16. **Subject kinds:** `subject_kind` and `also` (at most 3), `via_parent` rows (at most 8 per + event), and pseudonymised non-account ids. Approve? +17. **Stage gate:** only advise-mode local rules count, as its own slice (P1s). Approve? +18. **Profiles and keys:** a private mount, and a provider-agnostic `Keys` interface with HKDF per + tenant and purpose. Approve? +19. **Rescore control:** proportional coalescing for DSL features only, timers only from non-shadow + rules, and 20 × active subjects per hour. Confirm? +20. **`group_by` admission:** + - Revision 2: first-come. + - R3: space-saving. + - **Now:** an exact `GROUP BY` in the aggregate engine, with space-saving only as the in-memory + adapter's memory bound, flagged `partial`. + + Approve? +21. **Declared numbers:** + - Revision 2: a Luhn check on integers. + - R3: author-trusted, with `max` required and no Luhn. + - **Now:** R3. + + Approve? +22. **`partial` flag:** it doesn't set `degraded`, because every partial computation is one-sided + toward higher risk. Should callers treat it like `degraded` anyway? + +## 13. Changes from earlier revisions + +### Revision 2 (after the adversarial review) + +- Dropped the canonical-key bridge in favour of a one-time rename. +- Replaced the event-count bound and the wall-clock deadline. +- Added re-HMAC, dropping of undeclared data, public-suffix domain checks, card/IP/phone scans and + skeleton-only custom text. +- Added `group_by`, `sequence`, `ratio`, `neighbours` and subject kinds, with a five-scenario walk. +- Dropped `delivery.sent`. +- Should-fixes: + - warm-up; + - dual keys; + - rescore control; + - a fair queue; + - advise-only stage gates; + - config history in the tree; + - private profiles. +- Re-sliced the plan into P0–P7. + +### Revision 3 (addendum, after the re-review of revision 2) + +**B2: bounded evaluation.** +- (a) Onboarding values come from ingest-maintained `subject_facts`, with monotone updates and an + indexed recount. No onboarding scan is byte-capped, and the blocked-payment flood provably moves + nothing. +- (b) Lifetime totals come from `subject_counters`, so no feature can fall under truncation. + `TruncationDir` is deleted. +- (c) Hitting a bound sets a `partial` flag on the signal and subject; it's never a weight. + `core.history_truncated` and its invariant and `packtest` check are deleted. +- (d) Evaluation classes F, N, A, R, G and D run in a fixed pass order, with reserved anchored + work and per-feature or per-pack budgets. `TestFeatureIndependence` guards this. +- (e) The flood property is restated against an unbounded reference, with a specified generator + and a per-class exactness argument, including `peak` saturation. +- (f) Bounded evaluation reaches e2a only after P4a–P4c. + +**B1: rename.** +- `feature_renamed` in `LoadSnapshotCorpus`, and `corpus-v2`. +- The fake scorer is re-baselined in P1. +- Registry-order summation makes the golden bit-exact, so the 1e-12 tolerance is gone. +- `Bound` is defined for velocity features and is never a cap. +- The 1,000 cap on totals is dropped. + +**B3: redaction.** +- Declared numbers are author-trusted, with `max` required and no Luhn. +- Domains default to eTLD+1, stored in cleartext if allowlisted and HMACed otherwise, built-in + fields included. The `recipient_domain` exception is explained in §12 Q14. +- Undeclared names have a grammar and a cap. +- Non-account and `also` ids are pseudonymised, and account ids are leak-scanned. +- Keys are HKDF-derived per tenant and purpose, with a cross-tenant test. +- An egress scan runs after digit folding. +- `subject_line` masking is clarified. +- ASN is exempt from hashing and gets a tighter grammar. +- A dev/staging migration job is added. +- Enum values and set files are scanned when a profile loads. + +**B4: the DSL.** +- Horizons feed the store range. +- `__absent` indicators with `absent_sign`, and `log1p` durations. +- `ratio` uses pre-transform values. +- `group_by` is exact, with a space-saving fallback. +- `hash_quantum`, with a 60-minute default for `until_now`. +- `peak` accepts `sum`, and `share` sums both sides; both are stated explicitly. +- `neighbours` is evaluated as of now, matching `eval/neighbors.go`. +- Dirty-mark propagation covers declared kinds. +- Declared evidence no longer feeds `core.linked_*`. +- `event_subjects` is capped at 8 rows per event, with `via_parent`. + +**Slices.** +- P1s is split out. +- P3 is split into P3a, P3b, P3c and P3d. +- P4 is split into P4a–P4d, and P5b is added. +- P5 now needs P3a and P4c, and P6a needs P3b. +- A test enforces that P1 lands first. diff --git a/docs/plans/2026-09-27-v0-plan.md b/docs/plans/2026-09-27-v0-plan.md index 26e5936..52ec100 100644 --- a/docs/plans/2026-09-27-v0-plan.md +++ b/docs/plans/2026-09-27-v0-plan.md @@ -35,7 +35,7 @@ design pass) rather than deferred to v1 outright. | S7 | Billing events | ops sidecar: `payment.attempt` (with `card_fingerprint_hash` under the tenant key) and `subscription.changed` | staging checkout produces events | | S8 | Hosted deploy | ops: compose service, Secret Manager keys, tenant config, Terraform alerts for queue depth / budget / drops | abusekit running on prod in shadow | | S9 | Incident evaluation | private backfill of the incident accounts and a benign sample into a private corpus; harness run; report precision/recall/lead-time before first send per account | report reviewed; floors set; decision on `advise` for the local rule | -| P0–P7 | Generic feature packs (own design: `docs/design/2026-09-29-generic-feature-packs.md`) | After S4 (#5) and S2b (#7) merge: golden replay via `abusekit eval --golden` (P0), one-time namespaced rename (P1), tenant profiles (P2), declared types and kind redaction (P3), DSL core (P4a) and extended ops (P4b), pack gating (P5), link/subject kinds (P6a), example scenarios and bootstrap (P6b), e2a private-profile cutover (P7). P0 and P1 must land before S5 and S8; S3b erasure must be vocabulary-aware. | per-slice "Done when" in that design's §9; P1 re-baselines the golden under semantic identity, every later slice keeps it exact | +| P0–P7 | Generic feature packs (own design: `docs/design/2026-09-29-generic-feature-packs.md`) | After S4 (#5) and S2b (#7) merge: golden replay (P0), one-time rename (P1), stage gate (P1s), tenant profiles (P2), vocabulary and scans (P3a), keys and re-HMAC (P3b), config history (P3c), rotation and cloud keys (P3d), facts and counters (P4a), aggregate engine (P4b), row features, flags and flood property (P4c), rescore control and warm-up (P4d), pack gating (P5), DSL parity (P5b), link and subject kinds (P6a), scenarios (P6b), e2a cutover (P7). **P1 lands before any S5/S8 PR** (S5 and S8 are blocked on P1; enforced by `TestNoVendorAdapterBeforeRename`). Bounded evaluation reaches e2a only after P4a–P4c. S3b erasure must be vocabulary-aware and clear facts/counters. | per-slice "Done when" in that design's §9; P1 golden bit-exact for values, risks and tiers, exact after every later slice | S1–S4 (including S3b) are pure abusekit and can run back to back; S5 needs vendor keys; S6–S8 are e2a/ops work that can start after S3 (S3b is not a blocker for them — nothing in S6–S8 depends on From a2346234199a18bcf754c1b025ad1513b0f080d2 Mon Sep 17 00:00:00 2001 From: jiashuoz Date: Tue, 29 Sep 2026 12:16:33 +0800 Subject: [PATCH 4/8] docs(design): generic feature packs revision 4 after verification Peak saturation sized in raw units per transform (and after the baseline), with partial+degraded when the row budget binds first; neighbours exact by saturation with where-before-limit; partial/degraded direction table (fixes the backwards fan-in claim); start defined by precedence (account_created_at, first accepted subject.created, server first_received_at) with anchored-fact invalidation; lock-first fact updates and bounded decline recount on inf->t and earlier moves; webmail counters subtract (now, +inf); facts/counters subject assignment for also/via_parent and type/at on event_subjects; ratio partial propagation. Plus the listed text fixes, slice fixes and a section 13 revision-4 addendum. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_014cdM7WyRc3mD3vQNXMTDB8 --- docs/design/2026-09-27-abusekit-design.md | 2 +- .../2026-09-29-generic-feature-packs.md | 404 ++++++++++++++---- docs/plans/2026-09-27-v0-plan.md | 2 +- 3 files changed, 329 insertions(+), 79 deletions(-) diff --git a/docs/design/2026-09-27-abusekit-design.md b/docs/design/2026-09-27-abusekit-design.md index 47ef66c..36aff45 100644 --- a/docs/design/2026-09-27-abusekit-design.md +++ b/docs/design/2026-09-27-abusekit-design.md @@ -6,7 +6,7 @@ see account churn, tiers failed open when a scorer was unavailable, and request cover reads or replay. This revision fixes those and tightens every place an implementer would have had to guess. Changes from r1 are marked **[r2]**. -**Amendment (proposed 2026-09-29, revision 3):** [`2026-09-29-generic-feature-packs.md`](2026-09-29-generic-feature-packs.md) +**Amendment (proposed 2026-09-29, revision 4):** [`2026-09-29-generic-feature-packs.md`](2026-09-29-generic-feature-packs.md) renames the built-in features once into namespaced `core`/`email`/`brand` packs enabled per tenant, adds product-declared vocabularies (types, field kinds, link kinds, subject kinds) with pseudonymising redaction, and adds bounded declarative custom features in YAML. It amends §4.2, §4.3, §4.5, §4.6, §4.8 diff --git a/docs/design/2026-09-29-generic-feature-packs.md b/docs/design/2026-09-29-generic-feature-packs.md index b903576..59efa4b 100644 --- a/docs/design/2026-09-29-generic-feature-packs.md +++ b/docs/design/2026-09-29-generic-feature-packs.md @@ -1,7 +1,7 @@ # Generic feature packs, product vocabularies, and declarative custom features -Status: proposed, revision 3, 2026-09-29. This revision answers the re-review of revision 2, which -returned "approve after listed changes". Owner: Josh Zhang. +Status: proposed, revision 4, 2026-09-29. This revision answers the verification of revision 3, +which returned "approve after listed changes". Owner: Josh Zhang. This document amends [`2026-09-27-abusekit-design.md`](2026-09-27-abusekit-design.md), §4.2, §4.3, §4.5, §4.6, §4.8 and §4.10. It is written against `main` at S3, treating two open PRs as @@ -65,10 +65,18 @@ For e2a, feature values, risks, tiers and rescore times stay bit-identical. exception is a masked marker inside a text field. - Every hash, every non-allowlisted domain and every non-account subject id is keyed per tenant. - No two tenants share a key. -6. **Flooding can't lower risk.** - - `risk(flooded, bounded evaluator) ≥ risk(flooded, unbounded reference)` holds for every - fixture under the §5.7 flood generator. - - A flood of 10,000 tiny `blocked` payments moves no onboarding feature. +6. **Bounding can't lower risk relative to the unbounded reference.** + - For every fixture under the §5.7 flood generator, + `risk(flooded, bounded evaluator) ≥ risk(flooded, unbounded reference)`. If that can't be + guaranteed for some feature, the verdict carries `degraded`. + - This is a relative guarantee. A flood can still legitimately move a feature in the reference + itself. Frozen legacy dilution is the known case: extra activity lowers + `core.burst_ratio_24h_vs_lifetime`'s share, and extra non-webmail sends lower + `email.webmail_recipient_share`, in both evaluators alike. + - **6b.** Neither `start` nor any class F or N onboarding fact moves under either of two floods: + - a flood of any type other than `subject.created`, from non-backfill producer keys, however + its `at` is chosen within the skew allowance; + - 10,000 tiny `blocked` payments. ## 2. Goals and non-goals @@ -225,7 +233,7 @@ type Output struct { | `core.resource_velocity_1h`, `core.credential_velocity_1h` (`key_velocity_1h`) | A | exact 1 h counts | | `core.resource_total`, `core.credential_total` (`key_total`) | F | `subject_counters` (lifetime) | | `core.burst_ratio_24h_vs_lifetime` | A + F | 24 h count (A) ÷ lifetime counters | -| `email.sends_1h`, `email.webmail_sends_1h`, `email.distinct_recipients_1h` | A + R | exact 1 h current value (A); history baseline (R) | +| `email.sends_1h`, `email.webmail_sends_1h`, `email.distinct_recipients_1h` | A + R | exact 1 h current value (A); history baseline (R). For `distinct_recipients_1h`, the current value is two class A aggregates: `count(DISTINCT recipient_hash)` plus the sum of `recipient_count` over rows without a hash. | | `email.sends_10m_max` | R | `peak` with a saturation limit (exact); history baseline (R) | | `email.sends_first_day`, `email.first_day_distinct_domains` | N | anchored; frozen into facts after day one | | `email.webmail_recipient_share` | F | lifetime counters: webmail recipients vs all non-self recipients | @@ -282,9 +290,12 @@ Revision 2's 1e-12 tolerance and cut-point proximity check are deleted. | `feature.Names`, `Map`, `FeatureSet`, `rules.yaml`, weights, mutation/ablation/golden-sign tests, fixtures README | Mechanical rename. | Full suite plus the grep test. | | `abusekit score --jsonl` | Flat names fail with `feature_renamed`. | CLI contract test. | -**Enforcement.** P1 lands before any S5 or S8 PR. The v0 plan marks S5 and S8 "blocked on P1". -`TestNoVendorAdapterBeforeRename` fails if `internal/model/{gemini,jev,laya}` exists while -`feature.KeySpace != "ns-v1"`. +**Enforcement.** P1 lands before any S5 or S8 PR, and the v0 plan marks S5 and S8 "blocked on P1". +`TestNoProductionBeforeRename` guards both sides: +- **S5:** it fails if `internal/model/{gemini,jev,laya}` exists while `feature.KeySpace != "ns-v1"`. +- **S8:** it asserts that `abusekit serve` refuses to start with `ABUSEKIT_ENV=production` unless + `feature.KeySpace == "ns-v1"`. The hosted deploy (S8) always sets that variable, so a pre-rename + binary can't be deployed. ### 5.3 Name grammar @@ -384,8 +395,16 @@ features: **Predicates.** The set is closed and has no regex: - `eq`, `ne`, `in`, `not_in` (at most 256 values; enum values are checked at load); -- `in_set`, `suffix_in_set` (set files of at most 100k entries, hashed and leak-scanned at profile - load); +- `in_set`, `suffix_in_set` (set files of at most 100k entries, leak-scanned at profile load). + - **On a `domain` field**, both are evaluated **at ingest** against the cleartext value, before + any HMAC. The result goes into a derived bool `x___in_`, which the predicate then + reads. + - Changing a set file bumps the vocabulary version and triggers a backfill. Features that use + the set stay cold until the backfill completes. + - `suffix_in_set` on any other kind is rejected at load. +- **Literals on pseudonymised fields.** `eq`/`in` literals and set files on `hash` fields are + HMACed at load, the same way ingest hashes values. During a rotation they are hashed under + **both** keys, and a predicate matches either one. - `gt`, `gte`, `lt`, `lte`, `exists`; - `all`, `any`, `not`, with nesting depth at most 2 and at most 8 leaves. @@ -403,9 +422,9 @@ predicates, `distinct`, `group_by` or `on`. | `peak` | `{size, sum?}`: the maximum of `agg(E ∩ (t − size, t])`, over `t` at the instants of matching events inside the outer window. `agg` is a count, or with `sum`, a sum. Sub-windows are clipped to the outer window. | R; exact under a saturation limit (§5.7) | | `time_between` | `{from: {type, where, anchor: first\|last}, to: {type, where}, until_now, if_absent}`. `t_A` is the first (or last) matching `from` with `at ≤ now`. `t_B` is the first matching `to` with `t_A ≤ t_B ≤ now`. The value is minutes from `t_A` to `t_B`, or `now − t_A` when `until_now` is set and there is no `t_B`. Absence semantics are below. | A; indexed first/last-match queries | | `sequence` | `{a, b, within ≤ 24h, on?: {a: f, b: g}}`: the number of `b` events in the window with an `a` event where `t_a ∈ (t_b − within, t_b]` and, if `on` is set, `a.f == b.g`. Both `on` fields must be `hash` fields with the same `join_domain`. | A; aggregate with a correlated existence test | -| `group_by` | A modifier on `count`, `distinct` or `sum`: `{field, reduce: max \| {count_gte: k}, max_groups}`. It groups by `field`, applies the op per group, then reduces: `max` takes the largest group value, `count_gte` counts groups at or above `k`. **Exact**: the aggregate engine groups every matching event, so there is no first-come admission for decoys to exploit. `max_groups` (at most 1,000) bounds only the in-memory adapter. Past it, the in-memory adapter uses space-saving (Metwally) with `k = max_groups` and flags the result `partial`. Postgres is always exact. | A; `GROUP BY` | -| `ratio` | `{num, den, if_empty}` over the **pre-transform** values of two non-ratio custom features. Depth 1: no cycles, no ratio of ratios. | D; O(1) | -| `neighbours` | `{via: [declared link kinds], where: {deleted: permanent} \| {labelled: abusive} \| {created_within: } \| {}}`: the number of distinct other same-tenant, same-kind subjects that share a `via` key and meet the condition. **As of `now`**: links with `first_seen ≤ now`, subjects created ≤ now, deletions and labels ≤ now. `created_within` is relative to `now`. This matches `evidenceAsOf` in `eval/neighbors.go`. Fan-in is capped at 50 per key and 200 in total; hitting the cap flags the result `partial`. | F/A; one indexed query per `via` set; at most 4 per tenant | +| `group_by` | A modifier on `count`, `distinct` or `sum`: `{field, reduce: max \| {count_gte: k}, max_groups}`. It groups by `field`, applies the op per group, then reduces: `max` takes the largest group value, `count_gte` counts groups at or above `k`. **Exact**: the aggregate engine groups every matching event, so there is no first-come admission for decoys to exploit. `max_groups` (at most 1,000) bounds only the in-memory adapter. Past it, the in-memory adapter uses space-saving (Metwally) with `k = max_groups`. Space-saving over-estimates tracked counts and loses evicted groups. `reduce: max` with a positive sign stays one-sided upward and is flagged `partial` only. `count_gte` can undercount groups whose true count is at or above `k` but that were evicted. A negative `prior_sign` or weight inverts the direction. Those combinations are flagged `partial` **and** `degraded`. Postgres is always exact. | A; `GROUP BY` | +| `ratio` | `{num, den, if_empty}` over the **pre-transform** values of two non-ratio custom features. Depth 1: no cycles, no ratio of ratios. **Partial propagation:** a `partial` input makes the ratio `partial`. `den` may not be a feature that can go partial (`relative_to_history`, `peak`, `group_by` or `neighbours`); the load fails with `ratio_den_partial_capable`. A partial `num` inherits its own `degraded` status. | D; O(1) | +| `neighbours` | `{via: [declared link kinds], where: {deleted: permanent} \| {labelled: abusive} \| {created_within: } \| {}}`: the number of distinct other same-tenant, same-kind subjects that share a `via` key and meet the condition. **As of `now`**: links with `first_seen ≤ now`, subjects created ≤ now, deletions and labels ≤ now. `created_within` is relative to `now`. This matches `evidenceAsOf` in `eval/neighbors.go`. **`where` is applied before any limit.** The query returns distinct matching subjects up to `K = ⌈T⁻¹(cap)⌉ + 1` (§5.7, saturation), so the count is exact up to saturation. The per-key fan-in cap no longer truncates the counted set. If the query's examined-row budget (default 50,000) runs out before it reaches `K` or finishes, the value is `partial` **and** `degraded`, because an undercount could lower risk. | F/A; one indexed query per `via` set; at most 4 per tenant | **Absence semantics** (`time_between`, `sequence`) - `if_absent` is a mandatory number. @@ -423,7 +442,16 @@ predicates, `distinct`, `group_by` or `on`. - Declared kinds feed only `neighbours`, so they never double-feed `core.linked_deleted_n`. - **Dirty-mark propagation** covers built-in *and* declared evidence kinds. When a subject gains a link key, is permanently deleted, or is labelled, every subject sharing any evidence key with it - is marked dirty, up to the fan-in cap. + is marked dirty. + - The first 50 per key are marked in the ingest transaction. + - The rest are marked by a background job that pages through them. It runs under a per-tenant + rate budget, and a metric counts the marks still pending. + - A subject whose mark is still pending picks up the change on its next event or timer rescore. + Its evidence is always read as of `now` and is never cached, so a late rescore sees the change + in full. +- **Legacy `core.linked_*`** keep their frozen caps (50 per key, 200 in total) and + `core.neighbors_truncated`, for bit-exactness. When a cap is hit, the verdict now also carries + `partial` and `degraded`, because the undercount lowers a positively weighted value. **`relative_to_history`** applies to `count`, `distinct` and `peak`: @@ -489,11 +517,19 @@ The result lies in `[0, cap]`, and the feature's `Bound` is `cap`. `burst_ratio` denominator count future-dated events. Counters include them naturally. - Frozen legacy exception: `email.first_day_distinct_domains` has an inclusive end. - Harmonising these is `@2` work (§12 Q5). -- **Counter exactness.** Custom lifetime counters exclude events with `at > now`. Those can only - sit inside the ±24 h skew window, so the class A engine subtracts them exactly with a query over - `(now, now + 24h]`. Counter `sum` fields must be integers, so the order of addition can't matter. -- **`hash_quantum`.** This lets `SkipInputUnchanged` skip unchanged inputs. The default is 60 - minutes (pre-transform) for `time_between` with `until_now`, and 0 otherwise. +- **Counter exactness.** Custom lifetime counters and the built-in webmail-share counters exclude + events with `at > now`. The built-in counters match the legacy behaviour, where + `webmailRecipientShare` excludes future-dated events. Backfill-scope keys are exempt from the + skew check, so future-dated events aren't limited to +24 h. The class A engine therefore + subtracts them exactly with an indexed query over `(now, +∞)`. + - The legacy total counters (`resource_total`, `credential_total`, the `burst_ratio` + denominator) do not subtract anything, matching their frozen behaviour. + - Counter `sum` fields must be integers, so the order of addition can't matter. +- **`hash_quantum`.** This lets `SkipInputUnchanged` skip unchanged inputs. Defaults: + - 60 minutes (pre-transform) for `time_between` with `until_now`; + - `0.01 × cap` (post-transform) for any feature with `age_decay`, whose value drifts with age on + every tick; + - 0 otherwise. - **Rescore candidates.** A feature proposes: - the exit of its oldest in-window match; - the end of its anchored window; @@ -527,7 +563,7 @@ hashed values too. | --- | --- | --- | --- | | `content.sent.recipient_domain` | **`none`**: `etld1` would merge `target1.example.test` and `target2.example.test`, which would change `email.first_day_distinct_domains` | Cleartext if allowlisted, else HMAC of the normalised value | HMAC is injective, so distinct counts don't change. Webmail membership is checked on cleartext. | | `content.sent.first_link_host` | `etld1` | Cleartext if allowlisted, else HMAC | Text scorers see a token for unknown hosts (§12 Q14) | -| `resource.*.address_domain` | `none` | Cleartext if allowlisted, else HMAC | No feature reads it | +| `resource.*.address_domain` | `etld1` (the default; revision 3's `none` exception is removed because no feature needs the full host) | Cleartext if allowlisted, else HMAC | No feature reads it | | `content.sent.recipient_hash` and every `links` value (built-in and declared) | Re-HMAC. Each link kind's `join_domain` is its kind name. | — | Equality is preserved, so no value changes | | `links.asn` | **Exempt** from the hash format check and from re-HMAC. It's coarse routing metadata in the clear (main §4.2), and its grammar tightens to `^(AS)?[0-9]{1,10}$`. | Cleartext | None | @@ -538,11 +574,16 @@ hashed values too. - Lookups apply the same derivation. **Ingest rules, in order** -1. **Input leak scan.** Every key and value is checked for: +1. **Input leak scan.** Every key and every **string** value is checked for: - email addresses; - Luhn-valid runs of 13–19 digits, with separators; - IPv4 and IPv6 literals; - - phone shapes. + - phone shapes: a leading `+` followed by 8–15 digits, or a digit run broken by space, `-`, `.` + or parentheses into the grouped national formats of 10 or more digits. + + A bare digit run with no `+` and no grouping is **not** a phone shape, so a 10-digit ASN + passes. Declared `number` values are JSON numbers, not strings, and are exempt, as §5.6 says + for `number`. Undeclared numbers are dropped before they're stored. The scan runs after NFKC and after folding every Unicode `Nd` digit to ASCII. A hit in a `text` field is masked. A hit in a `hash` field is exempt, because the value is replaced. Anywhere else, @@ -592,7 +633,11 @@ func Derive(master []byte, tenant string, purpose Purpose) []byte - A Go migration job runs under the tenant key and can resume by `seq`. It: - rewrites `links.hash` and the hash, domain and subject values in `events.links` and `events.data`; - - recomputes `events.body_hash`, so duplicate/conflict detection still matches; + - leaves `events.body_hash` untouched. From P3b on, `body_hash` is a SHA-256 over a + **key-independent canonical form**: the redacted event *before* pseudonymisation (declared + fields after masking and reduction, hash and domain values as the producer sent them). Neither + rotation nor the migration can break duplicate/conflict detection. The digest covers the whole + body, so it can't be used to test a single field; - sets `vocab_version`. **Vocabulary history lives in the config tree.** Each tenant has an append-only @@ -617,22 +662,81 @@ While any input is cold, an advise rule is stored as shadow, with `warming_until Revision 2's shared byte and step budget, and its weighted truncation feature, are replaced by the following. +**`start`, defined.** The subject's anchor instant is the first of these that exists: +1. `subject.created.account_created_at`, when the producer supplied it; +2. the `at` of the **first accepted** `subject.created` event, ordered by `received_at`. Later + `subject.created` events never move it; +3. `first_received_at`: the minimum **server-assigned** `received_at` over accepted events. + +The third fallback replaces revision 3's `LEAST(at)`. With `LEAST(at)`, a flood dated in the past +could move `start` earlier by up to the skew allowance, or without bound under backfill scope. In +replay, `received_at = at`, so fixtures without `subject.created` keep their legacy `start`. P4a +lists and justifies any fixture whose `start` changes because events precede its +`subject.created`. + **(a) Onboarding facts are maintained at ingest.** `subject_facts(tenant, kind, subject)` is -updated in the same transaction as the event insert. Every update is monotone and -order-independent: +updated in the same transaction as the event insert. **The fact row is a deterministic function of +the accepted event set:** any order of arrival produces the same row. The permutation test in +§5.7 (a2) checks this. The update rules: | Fact | Update | | --- | --- | -| `first_seen_at`, `account_created_at` | `LEAST(existing, new)`; `account_created_at` comes from `subject.created` | +| `first_seen_at` (legacy, informational), `first_received_at` | `LEAST(existing, new)` over `at` and over server `received_at`, respectively | +| `account_created_at`, `first_subject_created_at` | Taken from the first accepted `subject.created`, by `received_at`. Later `subject.created` events never replace them. These feed `start` (above). | | `first_success_at`, `first_success_key`, `first_success_funding` | On `payment.attempt{succeeded}`: replace when the event's `(at, producer, id)` is smaller | | `payment_counts` | `{succeeded, declined, blocked}`: increment | -| `declines_before_first_success` | On `declined` with `at ≤ first_success_at` (or no success yet): increment. When `first_success_at` moves earlier: recompute with an **indexed count** on `(tenant, kind, subject, type, at)` where `outcome = declined` and `at ≤ first_success_at`. | +| `declines_before_first_success` | On a `declined` event with `at ≤ first_success_at` (or no success yet): increment. **Recount** whenever `first_success_at` changes, both from absent to `t` and from an earlier move. The recount is Σ of the daily decline counters (a built-in `subject_counters` spec) for days before `day(t)`, plus one boundary-day query using a partial index on declined payment attempts: `(tenant, kind, subject, at) WHERE type = 'payment.attempt' AND data->>'outcome' = 'declined'`, bounded to `day(t)` and `at ≤ t`. Cost: O(days + declines on the boundary day). | | `first_paid_upgrade_at` | `LEAST` over `subscription.changed{status: active, amount_minor > 0}` | | `subscription_change_count` | Increment | | `first_external_at`, `self_sends_before_first_external` | The same pattern for `content.sent`; capped at 2 when read | | `name_brands`, `name_has_at`, `exempt_subject_brands` | Brand ids matched at ingest on `resource.*` names, with the integration-token gate applied | +- **Locking (a1).** Every ingest transaction runs in a fixed order: + 1. **Lock the facts row first.** `INSERT … ON CONFLICT (tenant, kind, subject) DO UPDATE SET + seq = subject_facts.seq + 1 RETURNING *` takes the row lock (or `SELECT … FOR UPDATE` when the + row already exists). + 2. Insert the event. + 3. Apply the updates, and recount when required. + + Counter-example to the unlocked version: transaction T1 moves `first_success_at` earlier and + recounts, while T2 concurrently inserts a decline dated before the new `t`. Without the lock, T1's + recount can't see T2's uncommitted decline, and T2 compares against the old `first_success_at`. + Either way, one decline is lost or counted twice. With the lock, T2 blocks until T1 commits, then + reads the new `t` and increments correctly. T2's event row can't commit before T2 holds the lock, + so it's never counted twice. +- **Permutation test (a2).** For every fixture, and for 1,000 random permutations and concurrent + interleavings of the fixture's events (run on the Postgres adapter with parallel transactions), + the final fact row must be bit-identical. - **No onboarding scan is ever byte-capped.** Scoring reads one facts row. +- **Frozen class N facts are invalidated and recomputed** whenever: + - `start` changes; the row records the `anchored_start` its values used; + - a backfill-scope event with `at` inside `[start, start + A)` is accepted, which sets + `anchored_dirty`. +- **Which subjects get facts and counters.** + - Every indexed subject (primary, `also`, `via_parent`) gets `subjects` row updates and + `first_received_at`. + - Onboarding facts (payment, subscription, brand, self-send) are updated **only for the primary + subject**. They describe the acting account. + - Counters are incremented for **every index row** whose `subject_kind` is in the counter spec's + `subject_kinds`. That includes `via_parent` rows, so parent-level lifetime totals see child + events, matching what class A queries see through `event_subjects`. +- **Custom `before_first` features** compile to a generic fact spec + `{count: {type, where}, before: {type, where}}`, kept in the facts row with the same rules: + lock-first, recount on `∞ → t` and on earlier moves, and daily counters plus a boundary-day query. + The recount uses the partial expression index for the spec's `count` predicate. +- **Erasure and re-signup.** + - A legal erasure (S3b) deletes the subject's facts row and counter rows, except that numeric + facts and counters of `abusive`-labelled subjects are kept under main §4.4's 24-month basis. + - A re-signup is a new subject id, with a new facts row and fresh counters. Churn evidence + carries across only through retained link hashes (`core.linked_*`, `neighbours`), never through + facts. +- **Counter expiry vs lifetime totals.** Counter day rows expire under the **same retention rule + as the event rows they count** (main §4.11). "Lifetime" therefore always means "over retained + events", which is exactly what the legacy scans computed. + - A backfill computes from the events retained when it runs and records a `retained_from` + watermark. + - The counter spec is not warm until the backfill completes, and its lifetime values are defined + over `[retained_from, now]`. - A brand-list change bumps the facts spec version. A backfill job then recomputes from retained events, and the brand features stay cold until it finishes. @@ -666,40 +770,83 @@ this. The same argument covers any flood of any type the update rules above don' | Pass | Class | What runs | Budget | | --- | --- | --- | --- | -| 1 | **F** | Read `subject_facts` and `subject_counters`. Evaluate neighbour evidence (indexed, fan-in capped). | None; O(1) rows plus capped neighbour queries. | -| 2 | **N** | Anchored `first:` features. While `now < start + A + 24h`, each runs its own class A query over `[start, start + A)`. After that, its value is **frozen** into `subject_facts.anchored`. Later events fall outside the skew allowance, so the frozen value can't go stale. | **Reserved**: runs before A, R and G, and shares with nothing. | +| 1 | **F** | Read `subject_facts` and `subject_counters`. Evaluate neighbour evidence: `neighbours` exact by saturation (§5.5), and legacy `core.linked_*` with their frozen caps and flags. | O(1) rows, plus neighbour queries under their own examined-row budget. | +| 2 | **N** | Anchored `first:` features. While `now < start + A + 24h`, each runs its own dedicated anchored-range query over `[start, start + A)`. In P4a that query runs the legacy Go code over the anchored rows; from P4b it can use the class A engine. After that window, the value is **frozen** into `subject_facts.anchored`, stamped with `anchored_start`. It is recomputed when `start` changes or a backfill-scope event lands in the anchor (see (a)). | **Reserved.** Runs before A, R and G, and shares with nothing. Its row budget is 50,000; if hit, the feature is `partial` + `degraded`. | | 3 | **A** | One exact aggregate query per feature: `count`, `sum`, `share`, `distinct`, `group_by`, `time_between`, `sequence`, the `cur` part of `relative_to_history`, and the class A parts of Go features. | **Per feature.** Returns O(1) or O(groups) rows. DB cost is O(rows in that feature's window), via the `(tenant, kind, subject, type, at)` index. | -| 4 | **R** | `peak` and `relative_to_history` baselines. Rows stream newest-first under the feature's own budget. | **Per feature.** `peak`: LIMIT `cap·⌈W/S⌉`. Baselines: 50,000 rows by default. | -| 5 | **G** | The remaining Go-pack computation that isn't expressible as A or R. For e2a after §5.1, this is only `email.distinct_recipients_1h`'s fallback sum, which is itself a class A aggregate. | **Per pack.** The pack's own row budget, not shared with any other pack or feature. | +| 4 | **R** | `relative_to_history` baselines first, then `peak`. Rows stream newest-first under each feature's own budget. | **Per feature.** Baselines: 50,000 rows by default. `peak`: LIMIT `N` from the saturation sizing below, capped by a 50,000-row budget. If that budget binds before `N`, the feature is `partial` + `degraded`. | +| 5 | **G** | Go-pack computation that isn't expressible as A or R. For e2a after §5.1 this is **empty**: every built-in feature is F, N, A or R. The class exists for future Go packs. | **Per feature.** Each G feature declares its own row budget in its `FeatureDef`. Hitting it sets the feature `partial` + `degraded`, because a Go feature's direction under truncation isn't proven. | | 6 | **D** | `ratio` and the `__absent` indicators. | O(1) | -**Why `peak` is exact under its limit.** Stream the matching rows newest-first with LIMIT -`N = cap·⌈W/S⌉`. -- Split `W` into `⌈W/S⌉` slots of width `S`. -- If the window holds at least `N` matching units, some slot holds at least `cap` units. -- The sub-window `(t − S, t]` that ends at the last event in that slot covers the whole slot. So - the true peak is at least `cap`, and the transformed value saturates at `cap`, which is what the - limited computation reports. -- If the window holds fewer than `N` units, the stream is complete and exact. -- With `sum`, the units are integer addends. Addends of 0 or less are excluded by the pushed-down - predicate, so `N` rows carry at least `N` units. +**Saturation sizing: why `peak` is exact under its limit, per transform.** Revision 3 sized the +limit in *post-transform* units, which is wrong. For example, `custom.declines_10m_peak` +(`log1p`, cap 7) needed `N = 7·144 = 1,008` rows under that sizing. Those rows could give a loaded +peak of 9, whose `ln 10 ≈ 2.3`, while the true value was `ln 1001 ≈ 6.9`. The limit must be sized +in **raw units**. + +Let `x_sat` be the smallest raw peak at which the scored value reaches its maximum. + +| Transform | `x_sat` | +| --- | --- | +| cap `C` only | `C` | +| `log1p: true`, cap `C` | `⌈e^C − 1⌉` | +| `log1p: {scale: s}` or `{anchored_at: n}`, cap `C` | `⌈e^(C/s) − 1⌉` | +| with `relative_to_history` (after its baseline `B` and age factor `d` are computed) | Scored value `T(min(P/max(B,1), rc)·d)`, where `T` is the transform above, so the maximum is `T(rc·d)`, reached when `P ≥ max(B,1)·min(rc, T⁻¹(C)/d)`. Hence `x_sat = ⌈max(B,1)·min(rc, T⁻¹(C)/d)⌉`. | + +Every `T` is monotone non-decreasing, and so is `P ↦ min(P/B, rc)·d`. A raw peak at or above +`x_sat` therefore scores exactly the maximum, and below it the value is exact whenever the stream +is complete. + +Stream matching rows newest-first with LIMIT `N = x_sat·⌈W/S⌉`: +- Every loaded row carries at least one raw unit. Rows with addends ≤ 0 are excluded in the + query. Rows are loaded but units summed, so if the limit binds, the loaded rows carry at least + `N` units. +- Split `W` into `⌈W/S⌉` slots of width `S`. By pigeonhole, some slot holds at least `x_sat` + loaded units. +- The sub-window `(t − S, t]` ending at that slot's last loaded event covers the whole slot. So + the **loaded** peak is at least `x_sat`, and the scored value equals the maximum, which is also + the true value, because the true peak is at least the loaded peak. +- If the limit doesn't bind, the stream is complete and the value is exact. + +For a history-relative `peak`, `B` is computed first (pass order). If `B` is partial, it is a +lower bound on the true baseline, so the `x_sat` sized from it is smaller than the true one: +- if the limit binds, `v1` saturates at `rc`, which is at least the true `v1`; +- if not, `P` is exact, and dividing by a smaller `B` only raises `v1`. + +Either way the result is one-sided upward, flagged `partial`. + +The second counter-example is `email.sends_10m_max` with `B = 50`: 5,000 units in one old slot and +about 302 units in each newer slot. Here `x_sat = 50·300/d`. With `d = 1`, `N = 15,000·144 ≈ 2.2M` +rows, far above the 50,000-row budget. So the budget binds first, and the feature is `partial` + +`degraded` rather than silently reporting `v1 ≈ 6`. For fixtures, the budget never binds. In +production this is an honest degradation, not a wrong value. **(c) Hitting a bound sets a flag; it's never a weight.** -- `Output.Partial` lists features whose own budget was hit: - - a `relative_to_history` baseline over budget; - - the in-memory `group_by` adapter past `max_groups`; - - `neighbours` at its fan-in cap. -- `core.history_truncated` no longer exists. -- In the API, each signal gets `partial: ["custom.x", …]` (omitted when empty). The subject gets - `partial: true` when any advise rule's signal has partial features. -- `partial` does **not** set `degraded`, because the rule was scored. Every partial computation is - one-sided toward *higher* risk: - - baselines only shrink; - - space-saving over-estimates; - - fan-in caps apply only to features required to be positively weighted. - - So callers can read `partial` as "risk may be overstated by these features, never understated". -- **Uniform rules** record `partial` like any rule. They are shadow-only, so it never affects a tier. +- `Output.Partial` lists features whose own budget was hit. `core.history_truncated` no longer + exists. +- Every partial source is classified by direction: + +| Partial source | Direction | Flags | +| --- | --- | --- | +| `relative_to_history` baseline budget hit | Value can only rise | `partial` | +| History-relative `peak` whose limit binds after a partial baseline | Value can only rise | `partial` | +| In-memory `group_by` space-saving with `reduce: max` and positive sign | Value can only rise | `partial` | +| `peak` or anchored (class N) row budget hit before saturation | Value may undercount | `partial` + `degraded` | +| `neighbours` examined-row budget hit before `K` | Value may undercount | `partial` + `degraded` | +| Legacy `core.linked_*` fan-in cap hit | Value undercounts | `partial` + `degraded` | +| Space-saving with `count_gte`, or with a negative sign or weight | Value may undercount | `partial` + `degraded` | +| Class G budget hit | Direction unproven | `partial` + `degraded` | +| `ratio` with a partial `num` | Inherits `num`'s flags | — | + + Revision 3 claimed that fan-in caps apply only to positively weighted features and therefore err + upward. **That had the direction backwards:** an undercount lowers a positively weighted value. + It is corrected above. +- **API.** Each signal gets `partial: ["custom.x", …]`, omitted when empty. `degraded` follows + main §4.4 and is also set when any advise rule has a feature that may undercount. The subject + gets `partial: true` when any advise rule's signal has partial features. + - Callers can read a bare `partial` as "risk may be overstated, never understated". + - `degraded` keeps its existing meaning: "don't trust a low score". +- **Uniform rules** record the flags like any rule. They are shadow-only, so the flags never + affect a tier. **(d) Budgets are per feature and per pack. Changing one feature never moves another.** - A feature's value depends only on the facts, the counters, its own queries and its own budget. @@ -724,12 +871,17 @@ encoding of each type): | Placement | Entirely before, interleaved with, and after the real events. Timestamps land inside, at the edges of, and outside every feature window, including future-dated events within skew. | | Volume | 1×, 10× and 100× each feature's saturation bound | -The property holds by construction: -- classes F, N and A are exact, so bounded equals reference; -- `peak` is exact by saturation; -- a partial `relative_to_history` value is never below the reference, and its weight is never - negative; -- space-saving never under-counts. +The property holds by construction, or the verdict is `degraded`: +- Classes F and A are exact, so the bounded value equals the reference. Class N is exact until + its budget binds, which sets `degraded`. +- `peak` is exact by raw-unit saturation, or `degraded` when its budget binds first. +- `neighbours` is exact by saturation, or `degraded`. +- A partial `relative_to_history` value is never below the reference, and its weight is never + negative. +- Space-saving errs upward only for `max` with a positive sign. Every other combination is + `degraded`. +- `ratio` can't take a partial-capable `den`. +- Criterion 6b covers the anchor: `start` ignores `at` on everything except `subject.created`. The test also checks it empirically for every built-in and DSL feature on every fixture. @@ -746,6 +898,9 @@ bounded evaluation switched on for e2a (§9). | `time_between` | `(−∞, now]`, via indexed first/last-match queries | | `sequence` | `a` rows from `(now − W − within, now]`; `b` rows from `(now − W, now]` | | `relative_to_history` baseline | `(now − lookback, now − exclude_recent]` | +| `before_first` | Facts row (§5.7a), plus one boundary-day query on recount | +| `lifetime` | `subject_counters` rows for the spec, minus an indexed `(now, +∞)` query for specs that exclude future-dated events | +| `neighbours` | `links` index `(tenant, kind, hash)` for each `via` key, with `where` applied, up to `K` distinct subjects, under the examined-row budget | **Limits** (validated at load; also the fuzz oracle) @@ -861,7 +1016,10 @@ behaviour on `main`. e2a has no staged rules, so its golden doesn't change. `CREATE INDEX CONCURRENTLY`. **New tables** -- `event_subjects(tenant, subject_kind, subject, event_seq, via_parent)`: at most 8 rows per event. +- `event_subjects(tenant, subject_kind, subject, type, at, event_seq, via_parent)`: at most 8 + rows per event. Its index `(tenant, subject_kind, subject, type, at)` lets class A queries + through `also` and `via_parent` rows run as index-bounded range scans before the join to + `events` for pushed-down data predicates. - `subject_facts`. - `subject_counters(…, day, n, sum)`: expires with main §4.11's numeric retention. - `tenant_config_versions`. @@ -941,7 +1099,7 @@ redaction boundary. - **Parent and `also` fan-out.** At most 8 index rows and 8 dirty marks per event, coalesced by `dirty_seq`. - **Hostile config or events.** There's no code or regex. Limits, per-feature budgets and exactness - apply, and partial computations are one-sided. + apply. Each partial source is either one-sided upward (`partial`) or flagged `degraded` (§5.7c). - **Brand-list change.** The facts spec version is bumped and a backfill runs; the affected features stay cold until it finishes. @@ -1095,7 +1253,7 @@ features: share: {type: invite.sent, match: {field: target_class, eq: other_community}}, window: 24h, transform: {cap: 1}} - {name: custom.invites_blocked_within_10m, version: 1, description: invitees who blocked within 10 min, sequence: {a: {type: invite.sent}, b: {type: block.received}, within: 10m, on: {a: invitee_hash, b: blocker_hash}}, - if_absent: 0, absent_sign: "+", window: 24h, transform: {log1p: true, cap: 6}} + if_absent: 0, absent_sign: "-", window: 24h, transform: {log1p: true, cap: 6}} # absent = no `a` events at all: benign - {name: custom.linked_invite_share_24h, version: 1, description: invites with links, share: {type: invite.sent, match: {field: link_host, exists: true}}, window: 24h, transform: {cap: 1}} - {name: custom.phone_siblings_7d, version: 1, description: accounts sharing a phone created this week, @@ -1142,7 +1300,7 @@ features: sequence: {a: {type: auth.attempted, where: {field: outcome, eq: failed}}, b: {type: auth.attempted, where: {field: outcome, eq: succeeded}}, within: 10m, on: {a: login_hash, b: login_hash}}, - if_absent: 0, absent_sign: "+", window: 24h, transform: {log1p: true, cap: 6}} + if_absent: 0, absent_sign: "-", window: 24h, transform: {log1p: true, cap: 6}} # absent = no `a` events at all: benign - {name: custom.hosting_share_1h, version: 1, subject_kinds: [api_key], description: attempts from hosting networks, share: {type: auth.attempted, match: {field: client_asn, eq: hosting}}, window: 1h, transform: {cap: 1}} - {name: custom.max_attempts_per_ip_1h, version: 1, subject_kinds: [api_key], description: busiest client network, @@ -1267,16 +1425,16 @@ These come after #5 and #7 merge. P1 lands before any S5 or S8 PR (enforced as d | # | Slice | Contents | Depends on | Done when | | --- | --- | --- | --- | --- | | P0 | Golden replay | `abusekit eval --golden`; `(at, producer, id)` order; `reference-flat.jsonl` | #5, #7 | Golden committed; flipping one weight's last bit fails it | -| P1 | One-time rename | Every §5.2 consumer and its test; `FeatureDef`; `core.Vector`; registry-order summation; fake-scorer re-baseline; `corpus-v2` + `feature_renamed` in `LoadSnapshotCorpus` and `score --jsonl`; reason v2; corpus key-space migration; cassette header; `TestNoVendorAdapterBeforeRename` | P0 | `reference-ns.jsonl` bit-exact for values, risks and tiers; grep test clean | +| P1 | One-time rename | Every §5.2 consumer and its test; `FeatureDef`; `core.Vector`; registry-order summation; fake-scorer re-baseline; `corpus-v2` + `feature_renamed` in `LoadSnapshotCorpus` and `score --jsonl`; reason v2; corpus key-space migration; cassette header; `TestNoProductionBeforeRename` (guards S5 and S8) | P0 | `reference-ns.jsonl` bit-exact for values, risks and tiers; grep test clean | | P1s | Stage gate | `maxRiskByScorer` limited to advise-mode local rules | P1 | Stage tests pass; golden exact | | P2 | Tenant profiles | Private-mount loader; reference profile; per-tenant reload, `/healthz`, rule sets and scorer version; fair queue and concurrency cap. All built-in features available to every tenant. | P1 | Golden exact; isolation and fairness tests | | P3a | Vocabulary and scans | `internal/vocab`; kinds (numbers require `max`); roles; `x_` fields; declared-domain PSL, `etld1` and IP checks; card/IP/phone scans with digit folding; `subject_line` masking; egress scan; drop-undeclared with name grammar; account-subject leak scan; ASN grammar; profile-load scans | P2 | Criterion 5 (non-key parts); golden exact or deviations justified | | P3b | Keys and re-HMAC | `internal/secret` (HKDF, file adapter); re-HMAC of hashes and links with `join_domain`; domain allowlist/HMAC (built-in and declared); pseudonymised non-account and `also` ids; `RedactionSchemaVersion` 3; dev/staging migration job; cross-tenant key test | P3a | Criterion 5 complete; golden exact; migration test | | P3c | Config history | `history/`; `abusekit config check`; `tenant_config_versions` | P2 | History CI tests | | P3d | Rotation and cloud keys | Dual-key write, read and flip; `__prev` fields; `key_id`; cloud secret-manager adapter | P3b | Equality exact across a simulated rotation; adapter contract test | -| P4a | Facts and counters | `subject_facts`, `subject_counters`; ingest transaction; onboarding, brand and self-send facts; anchored freeze; backfill and cold state; Go features moved onto classes F and N | P3a | Golden exact; blocked-payment flood test; out-of-order fact tests | +| P4a | Facts and counters | `start` precedence; `subject_facts` and `subject_counters`; the lock-first ingest transaction; recount on `∞ → t` and on earlier moves (daily decline counters plus a boundary-day query on a partial index); onboarding, brand and self-send facts; generic `before_first` fact specs; webmail counters subtracting `(now, +∞)`; anchored freeze using its **own** anchored-range query (legacy Go code over `[start, start + A)`), with invalidation on `start` change or backfill; fact and counter subject assignment for `also`/`via_parent`; retention-aligned counter expiry; backfill with a `retained_from` watermark | P3a | Golden exact (or `start` deviations justified); blocked-payment and pre-dated flood tests (6b); the permutation and concurrent-interleaving test (a2); the lost-decline race test | | P4b | Aggregate engine (class A) | `internal/evalengine` (Postgres and in-memory); `count`, `sum`, `share`, `distinct`, `group_by`, `time_between` + `__absent`, `sequence`, `ratio`; pushdown and indexes; conformance | P4a | Reference equality on 10k histories; adapter equivalence | -| P4c | Class R, budgets, flags, flood | `peak` (with `sum`) under the saturation limit; `relative_to_history` with its baseline budget; `partial` in verdicts and the API; pass order; `TestFeatureIndependence`; flood generator and property; cost benchmark. **Then enable bounded evaluation for e2a.** | P4b | Criteria 4 and 6; golden exact under the bounded evaluator | +| P4c | Class R, budgets, flags, flood | `peak` (with `sum`) under the raw-unit saturation limit per transform (§5.7); `neighbours` exact by saturation; the partial/degraded direction table; `ratio` partial propagation and `ratio_den_partial_capable`; `relative_to_history` with its baseline budget; `partial` in verdicts and the API; pass order; `TestFeatureIndependence`; flood generator and property; cost benchmark. **Then enable bounded evaluation for e2a.** | P4b | Criteria 4 and 6; golden exact under the bounded evaluator | | P4d | Rescore control and warm-up | Proportional coalescing; timers only from non-shadow rules; per-tenant budget; warm-up | P4c, P3c | Storm and warm-up tests | | P5 | Pack gating | Registry; `core`, `email` and `brand` adapters; enablement; `brand.title_match` (needs the `title` role); `packtest`; starter weights | P3a, P4c | Golden exact; `feature_not_enabled`; every pack passes `packtest` | | P5b | DSL parity | Every expressible #7/S2 feature re-expressed in the DSL | P4c | Bit-exact against the Go feature on every fixture | @@ -1285,7 +1443,10 @@ These come after #5 and #7 merge. P1 lands before any S5 or S8 PR (enforced as d | P7 | e2a cutover | Private profile in the ops mount; hosted `config check` | P5, P4c, S8's mount | Golden exact against the private copy | Two v0 slices interact with this plan: -- S3b must be vocabulary-aware and must clear facts and counters. It is easiest after P4a. +- **S3b** must be vocabulary-aware. Its **done-when** includes a test that erasing a non-abusive + subject deletes its `subject_facts` and `subject_counters` rows. A second test checks that an + `abusive`-labelled subject keeps only numeric facts and counters under the 24-month basis. S3b + is easiest after P4a. - S6 is independent of all of the above. ## 10. Scalability and extensibility @@ -1371,10 +1532,13 @@ Where the re-review's (R3) answer differs from the earlier recommendation, both 7. **Bounding:** - Revision 2: byte caps, a shared step budget and a weighted truncation feature. - R3: ingest facts, counters, per-feature budgets, and truncation as a flag. - - **Now:** R3, plus exact per-feature aggregates. The only bounded pieces left are the - one-sided `relative_to_history` baselines and the in-memory `group_by` fallback. + - **Now:** R3, plus exact per-feature aggregates. + - **Revision 4:** `peak` is sized in raw units per transform, and `neighbours` is exact by + saturation. Anything that can undercount (row, anchored, neighbour and G budgets; + space-saving outside `max`+positive) sets `partial` + `degraded`. - Confirm the 50,000-row baseline budget and `max_groups` of 1,000? + Confirm the 50,000-row budgets (baseline, `peak`, anchored, neighbour examined rows) and + `max_groups` of 1,000? 8. **Bootstrap:** uniform priors, shadow-only, no fitting, held-out fixtures, and `__absent` indicators with `absent_sign`. Confirm? 9. **CEL:** later as a `where` leaf only, or a new design pass? Unchanged. @@ -1413,8 +1577,22 @@ Where the re-review's (R3) answer differs from the earlier recommendation, both - **Now:** R3. Approve? -22. **`partial` flag:** it doesn't set `degraded`, because every partial computation is one-sided - toward higher risk. Should callers treat it like `degraded` anyway? +22. **`partial` flag:** + - Revision 3: `partial` never set `degraded`. + - Verification: that was wrong for undercounting sources. + - **Now (revision 4):** a bare `partial` means the value can only have risen. Any source that + may undercount also sets `degraded` (§5.7c table). + + Should callers still treat a bare `partial` like `degraded`? +23. **`start` precedence (revision 4, new):** `account_created_at`, then the first accepted + `subject.created` `at`, then server `first_received_at`. This replaces `LEAST(at)`, so a + pre-dated flood can't age an account. It may change `start` for fixtures whose events precede + `subject.created`; P4a lists them. Approve? +24. **Key-independent `body_hash` (revision 4, new):** SHA-256 over the redacted body before + pseudonymisation, so rotation and the re-HMAC migration never break duplicate/conflict + detection. Approve? +25. **Dirty marks beyond the fan-in cap (revision 4, new):** the first 50 per key are marked in + the ingest transaction, and the rest by a rate-budgeted background job. Confirm? ## 13. Changes from earlier revisions @@ -1490,3 +1668,75 @@ Where the re-review's (R3) answer differs from the earlier recommendation, both - P4 is split into P4a–P4d, and P5b is added. - P5 now needs P3a and P4c, and P6a needs P3b. - A test enforces that P1 lands first. + +### Revision 4 (addendum, after verification of revision 3) + +**Blocking fixes (P4a, P4c)** +1. **`peak` exactness.** The limit is now sized in **raw units** per transform (§5.7 saturation + sizing): + - `x_sat` is `C`, `⌈e^C − 1⌉` or `⌈e^(C/s) − 1⌉`. + - With `relative_to_history`, `x_sat = ⌈max(B,1)·min(rc, T⁻¹(C)/d)⌉`, computed after the + baseline. + - `N = x_sat·⌈W/S⌉` rows. The pigeonhole argument is restated over loaded units. + - If the row budget binds before `N`, the feature is `partial` + `degraded`. + - Both verification counter-examples are worked through. +2. **`neighbours`.** + - `where` is applied before any limit, and the query counts up to `K = ⌈T⁻¹(cap)⌉ + 1`, so the + count is exact by saturation. An examined-row budget hit sets `degraded`. + - Legacy `core.linked_*` caps now also set `partial` + `degraded`. + - The backwards direction claim in §5.7c is corrected. +3. **`start` is defined** by precedence: `account_created_at`, then the first accepted + `subject.created` by `received_at`, then server `first_received_at`. + - Criterion 6b is restated. + - Frozen class N facts carry `anchored_start` and are recomputed when `start` changes or a + backfill-scope event lands in the anchor. +4. **Recount locking.** + - The facts row lock is taken first (`INSERT … ON CONFLICT … RETURNING` / `FOR UPDATE`). The + lost-decline race is given as the counter-example. + - The recount costs O(days + boundary-day declines), using daily decline counters plus a + partial index. + - It fires on `∞ → t` and on earlier moves. +5. **Webmail counters** subtract future-dated events over `(now, +∞)`, matching legacy and + covering backfill keys that are exempt from the skew check. +6. **`also`/`via_parent`.** + - Onboarding facts are updated for the primary subject only. + - Counters are updated for every index row whose kind is in the spec. + - `event_subjects` gains `type` and `at`, with a matching index, so class A queries through + parents are index-bounded. +7. **`ratio`.** Partial status propagates into class D. A partial-capable `den` is rejected at + load (`ratio_den_partial_capable`). + +**Text fixes** +- Space-saving direction: only `max` with a positive sign is upward. `count_gte` and negative signs + are `degraded`. +- The facts row is restated as a deterministic function of the accepted event set, backed by a + permutation and concurrent-interleaving test. +- The leak scan covers strings only, so declared numbers are exempt. Phone shapes need `+` or + grouping, so 10-digit ASNs pass. +- `in_set`/`suffix_in_set` on domain fields are evaluated at ingest into a derived bool. +- Literals and set files on hash fields are HMACed at load, under both keys during rotation. +- `body_hash` is computed over a key-independent canonical form, so the migration no longer + recomputes it. +- Class G budgets are per feature, with `partial` + `degraded` on a hit. G is empty for e2a. +- `distinct_recipients_1h` is consistently A (current) + R (baseline). +- §1 criterion 6 is restated relative to the unbounded reference, with the frozen dilution + semantics (`burst_ratio`, webmail share) noted. +- Erasure and re-signup intent for facts and counters is specified. Counter expiry is aligned with + event retention, and backfills carry a `retained_from` watermark. +- Load-plan rows are added for `before_first`, `lifetime` and `neighbours`, with a generic fact + spec for custom `before_first`. +- `age_decay` features default to `hash_quantum` `0.01 × cap`. +- Dirty marks beyond the fan-in cap are handled by a rate-budgeted background job. +- The `address_domain` `none` exception is removed. +- The 7c/7d sequence `absent_sign` is `-`. + +**Slices** +- P4a's anchored freeze uses its own anchored-range query, not the P4b engine. +- S3b's done-when includes the facts/counters erasure tests. +- `TestNoProductionBeforeRename` guards S8 as well as S5, by refusing production startup before + the rename. + +**Decisions** +- Q7 and Q22 are revised. +- Q23 (`start` precedence), Q24 (key-independent `body_hash`) and Q25 (dirty marks beyond the cap) + are new. diff --git a/docs/plans/2026-09-27-v0-plan.md b/docs/plans/2026-09-27-v0-plan.md index 52ec100..fe85e79 100644 --- a/docs/plans/2026-09-27-v0-plan.md +++ b/docs/plans/2026-09-27-v0-plan.md @@ -35,7 +35,7 @@ design pass) rather than deferred to v1 outright. | S7 | Billing events | ops sidecar: `payment.attempt` (with `card_fingerprint_hash` under the tenant key) and `subscription.changed` | staging checkout produces events | | S8 | Hosted deploy | ops: compose service, Secret Manager keys, tenant config, Terraform alerts for queue depth / budget / drops | abusekit running on prod in shadow | | S9 | Incident evaluation | private backfill of the incident accounts and a benign sample into a private corpus; harness run; report precision/recall/lead-time before first send per account | report reviewed; floors set; decision on `advise` for the local rule | -| P0–P7 | Generic feature packs (own design: `docs/design/2026-09-29-generic-feature-packs.md`) | After S4 (#5) and S2b (#7) merge: golden replay (P0), one-time rename (P1), stage gate (P1s), tenant profiles (P2), vocabulary and scans (P3a), keys and re-HMAC (P3b), config history (P3c), rotation and cloud keys (P3d), facts and counters (P4a), aggregate engine (P4b), row features, flags and flood property (P4c), rescore control and warm-up (P4d), pack gating (P5), DSL parity (P5b), link and subject kinds (P6a), scenarios (P6b), e2a cutover (P7). **P1 lands before any S5/S8 PR** (S5 and S8 are blocked on P1; enforced by `TestNoVendorAdapterBeforeRename`). Bounded evaluation reaches e2a only after P4a–P4c. S3b erasure must be vocabulary-aware and clear facts/counters. | per-slice "Done when" in that design's §9; P1 golden bit-exact for values, risks and tiers, exact after every later slice | +| P0–P7 | Generic feature packs (own design: `docs/design/2026-09-29-generic-feature-packs.md`) | After S4 (#5) and S2b (#7) merge: golden replay (P0), one-time rename (P1), stage gate (P1s), tenant profiles (P2), vocabulary and scans (P3a), keys and re-HMAC (P3b), config history (P3c), rotation and cloud keys (P3d), facts and counters (P4a), aggregate engine (P4b), row features, flags and flood property (P4c), rescore control and warm-up (P4d), pack gating (P5), DSL parity (P5b), link and subject kinds (P6a), scenarios (P6b), e2a cutover (P7). **P1 lands before any S5/S8 PR** (S5 and S8 are blocked on P1; enforced by `TestNoProductionBeforeRename`, which guards vendor adapters and production startup). Bounded evaluation reaches e2a only after P4a–P4c. S3b erasure must be vocabulary-aware and clear facts/counters. | per-slice "Done when" in that design's §9; P1 golden bit-exact for values, risks and tiers, exact after every later slice | S1–S4 (including S3b) are pure abusekit and can run back to back; S5 needs vendor keys; S6–S8 are e2a/ops work that can start after S3 (S3b is not a blocker for them — nothing in S6–S8 depends on From b74c595bafbe69dd83d48865219717aa94a1d23e Mon Sep 17 00:00:00 2001 From: jiashuoz Date: Tue, 29 Sep 2026 12:21:27 +0800 Subject: [PATCH 5/8] docs(design): generic feature packs revision 5 after final check Recount-backing counters (decline, before_first) are primary-subject-only, with a parent/child recount test; "first accepted" is the smallest (received_at, producer, id) and determinism holds with received_at fixed; peak worked examples corrected (streams complete; binding is decided at run time; x_sat fixed); READ COMMITTED ingest with bounded retry and hourly decline counters plus a boundary-hour recount; class N features rejected as ratio den and negative-sign ratios over partial-capable nums flagged degraded; start clamped to first_received_at. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_014cdM7WyRc3mD3vQNXMTDB8 --- docs/design/2026-09-27-abusekit-design.md | 2 +- .../2026-09-29-generic-feature-packs.md | 113 +++++++++++++----- 2 files changed, 87 insertions(+), 28 deletions(-) diff --git a/docs/design/2026-09-27-abusekit-design.md b/docs/design/2026-09-27-abusekit-design.md index 36aff45..be9b7d6 100644 --- a/docs/design/2026-09-27-abusekit-design.md +++ b/docs/design/2026-09-27-abusekit-design.md @@ -6,7 +6,7 @@ see account churn, tiers failed open when a scorer was unavailable, and request cover reads or replay. This revision fixes those and tightens every place an implementer would have had to guess. Changes from r1 are marked **[r2]**. -**Amendment (proposed 2026-09-29, revision 4):** [`2026-09-29-generic-feature-packs.md`](2026-09-29-generic-feature-packs.md) +**Amendment (proposed 2026-09-29, revision 5):** [`2026-09-29-generic-feature-packs.md`](2026-09-29-generic-feature-packs.md) renames the built-in features once into namespaced `core`/`email`/`brand` packs enabled per tenant, adds product-declared vocabularies (types, field kinds, link kinds, subject kinds) with pseudonymising redaction, and adds bounded declarative custom features in YAML. It amends §4.2, §4.3, §4.5, §4.6, §4.8 diff --git a/docs/design/2026-09-29-generic-feature-packs.md b/docs/design/2026-09-29-generic-feature-packs.md index 59efa4b..b1026e4 100644 --- a/docs/design/2026-09-29-generic-feature-packs.md +++ b/docs/design/2026-09-29-generic-feature-packs.md @@ -1,7 +1,6 @@ # Generic feature packs, product vocabularies, and declarative custom features -Status: proposed, revision 4, 2026-09-29. This revision answers the verification of revision 3, -which returned "approve after listed changes". Owner: Josh Zhang. +Status: proposed, revision 5, 2026-09-29. This revision answers the final check of revision 4. Owner: Josh Zhang. This document amends [`2026-09-27-abusekit-design.md`](2026-09-27-abusekit-design.md), §4.2, §4.3, §4.5, §4.6, §4.8 and §4.10. It is written against `main` at S3, treating two open PRs as @@ -423,7 +422,7 @@ predicates, `distinct`, `group_by` or `on`. | `time_between` | `{from: {type, where, anchor: first\|last}, to: {type, where}, until_now, if_absent}`. `t_A` is the first (or last) matching `from` with `at ≤ now`. `t_B` is the first matching `to` with `t_A ≤ t_B ≤ now`. The value is minutes from `t_A` to `t_B`, or `now − t_A` when `until_now` is set and there is no `t_B`. Absence semantics are below. | A; indexed first/last-match queries | | `sequence` | `{a, b, within ≤ 24h, on?: {a: f, b: g}}`: the number of `b` events in the window with an `a` event where `t_a ∈ (t_b − within, t_b]` and, if `on` is set, `a.f == b.g`. Both `on` fields must be `hash` fields with the same `join_domain`. | A; aggregate with a correlated existence test | | `group_by` | A modifier on `count`, `distinct` or `sum`: `{field, reduce: max \| {count_gte: k}, max_groups}`. It groups by `field`, applies the op per group, then reduces: `max` takes the largest group value, `count_gte` counts groups at or above `k`. **Exact**: the aggregate engine groups every matching event, so there is no first-come admission for decoys to exploit. `max_groups` (at most 1,000) bounds only the in-memory adapter. Past it, the in-memory adapter uses space-saving (Metwally) with `k = max_groups`. Space-saving over-estimates tracked counts and loses evicted groups. `reduce: max` with a positive sign stays one-sided upward and is flagged `partial` only. `count_gte` can undercount groups whose true count is at or above `k` but that were evicted. A negative `prior_sign` or weight inverts the direction. Those combinations are flagged `partial` **and** `degraded`. Postgres is always exact. | A; `GROUP BY` | -| `ratio` | `{num, den, if_empty}` over the **pre-transform** values of two non-ratio custom features. Depth 1: no cycles, no ratio of ratios. **Partial propagation:** a `partial` input makes the ratio `partial`. `den` may not be a feature that can go partial (`relative_to_history`, `peak`, `group_by` or `neighbours`); the load fails with `ratio_den_partial_capable`. A partial `num` inherits its own `degraded` status. | D; O(1) | +| `ratio` | `{num, den, if_empty}` over the **pre-transform** values of two non-ratio custom features. Depth 1: no cycles, no ratio of ratios. **Partial propagation:** a `partial` input makes the ratio `partial`. `den` may not be a feature that can go partial (`relative_to_history`, `peak`, `group_by`, `neighbours` or any class N `first:` feature); the load fails with `ratio_den_partial_capable`. A partial `num` inherits its own `degraded` status. If a partial-capable `num` meets a **negative** ratio sign (the ratio's `prior_sign`, or its weight in any weights file), an upward error in `num` would lower risk. The ratio is then flagged `partial` + `degraded` whenever `num` is partial, the same treatment space-saving gets. | D; O(1) | | `neighbours` | `{via: [declared link kinds], where: {deleted: permanent} \| {labelled: abusive} \| {created_within: } \| {}}`: the number of distinct other same-tenant, same-kind subjects that share a `via` key and meet the condition. **As of `now`**: links with `first_seen ≤ now`, subjects created ≤ now, deletions and labels ≤ now. `created_within` is relative to `now`. This matches `evidenceAsOf` in `eval/neighbors.go`. **`where` is applied before any limit.** The query returns distinct matching subjects up to `K = ⌈T⁻¹(cap)⌉ + 1` (§5.7, saturation), so the count is exact up to saturation. The per-key fan-in cap no longer truncates the counted set. If the query's examined-row budget (default 50,000) runs out before it reaches `K` or finishes, the value is `partial` **and** `degraded`, because an undercount could lower risk. | F/A; one indexed query per `via` set; at most 4 per tenant | **Absence semantics** (`time_between`, `sequence`) @@ -664,10 +663,17 @@ following. **`start`, defined.** The subject's anchor instant is the first of these that exists: 1. `subject.created.account_created_at`, when the producer supplied it; -2. the `at` of the **first accepted** `subject.created` event, ordered by `received_at`. Later - `subject.created` events never move it; +2. the `at` of the **first accepted** `subject.created` event. Later `subject.created` events never + move it; 3. `first_received_at`: the minimum **server-assigned** `received_at` over accepted events. +The chosen value is then clamped: `start = min(chosen, first_received_at)`. A producer-supplied +anchor can therefore never place `start` after the moment abusekit first saw the subject. + +**"First accepted"** means the event with the smallest `(received_at, producer, id)`. `received_at` +is assigned by the server once, at acceptance, and never changes. Event-time orderings (`at`-based +facts such as `first_success_at`, and every DSL "first") keep `(at, producer, id)`. + The third fallback replaces revision 3's `LEAST(at)`. With `LEAST(at)`, a flood dated in the past could move `start` earlier by up to the skew allowance, or without bound under backfill scope. In replay, `received_at = at`, so fixtures without `subject.created` keep their legacy `start`. P4a @@ -676,21 +682,26 @@ lists and justifies any fixture whose `start` changes because events precede its **(a) Onboarding facts are maintained at ingest.** `subject_facts(tenant, kind, subject)` is updated in the same transaction as the event insert. **The fact row is a deterministic function of -the accepted event set:** any order of arrival produces the same row. The permutation test in -§5.7 (a2) checks this. The update rules: +the accepted event set, with each event's `received_at` held fixed.** Given the same events with +the same `received_at` values, any processing order or interleaving produces the same row. The +permutation test in §5.7 (a2) checks this. The update rules: | Fact | Update | | --- | --- | | `first_seen_at` (legacy, informational), `first_received_at` | `LEAST(existing, new)` over `at` and over server `received_at`, respectively | -| `account_created_at`, `first_subject_created_at` | Taken from the first accepted `subject.created`, by `received_at`. Later `subject.created` events never replace them. These feed `start` (above). | +| `account_created_at`, `first_subject_created_at` | Taken from the first accepted `subject.created`: the smallest `(received_at, producer, id)`. A later-processed event with a smaller key replaces them, so the result doesn't depend on processing order. These feed `start` (above). | | `first_success_at`, `first_success_key`, `first_success_funding` | On `payment.attempt{succeeded}`: replace when the event's `(at, producer, id)` is smaller | | `payment_counts` | `{succeeded, declined, blocked}`: increment | -| `declines_before_first_success` | On a `declined` event with `at ≤ first_success_at` (or no success yet): increment. **Recount** whenever `first_success_at` changes, both from absent to `t` and from an earlier move. The recount is Σ of the daily decline counters (a built-in `subject_counters` spec) for days before `day(t)`, plus one boundary-day query using a partial index on declined payment attempts: `(tenant, kind, subject, at) WHERE type = 'payment.attempt' AND data->>'outcome' = 'declined'`, bounded to `day(t)` and `at ≤ t`. Cost: O(days + declines on the boundary day). | +| `declines_before_first_success` | On a `declined` event with `at ≤ first_success_at` (or no success yet): increment. **Recount** whenever `first_success_at` changes, both from absent to `t` and from an earlier move. The recount is the Σ of the **hourly** decline counters for hours before `hour(t)`, plus one boundary-hour query over `[hour(t), t]`. The counters are a built-in `subject_counters` spec, **primary-subject-only** (see H-A below). The query uses a partial index on declined payment attempts: `(tenant, kind, subject, at) WHERE type = 'payment.attempt' AND data->>'outcome' = 'declined'`. Cost: one indexed SUM over at most 2,160 hourly rows (90-day retention), plus the declines inside a single hour. That bounds the part an attacker controls to one hour's events. | | `first_paid_upgrade_at` | `LEAST` over `subscription.changed{status: active, amount_minor > 0}` | | `subscription_change_count` | Increment | | `first_external_at`, `self_sends_before_first_external` | The same pattern for `content.sent`; capped at 2 when read | | `name_brands`, `name_has_at`, `exempt_subject_brands` | Brand ids matched at ingest on `resource.*` names, with the integration-token gate applied | +- **Isolation.** Ingest transactions run at **READ COMMITTED**. On a serialization failure or a + deadlock (`40001`, `40P01`), the transaction is retried up to 3 times with jittered backoff. If it + still fails, the item is rejected as a whole-request `5xx`, which the producer's outbox retries. + Idempotency on `(tenant, producer, id)` makes the retries safe. - **Locking (a1).** Every ingest transaction runs in a fixed order: 1. **Lock the facts row first.** `INSERT … ON CONFLICT (tenant, kind, subject) DO UPDATE SET seq = subject_facts.seq + 1 RETURNING *` takes the row lock (or `SELECT … FOR UPDATE` when the @@ -704,9 +715,11 @@ the accepted event set:** any order of arrival produces the same row. The permut Either way, one decline is lost or counted twice. With the lock, T2 blocks until T1 commits, then reads the new `t` and increments correctly. T2's event row can't commit before T2 holds the lock, so it's never counted twice. -- **Permutation test (a2).** For every fixture, and for 1,000 random permutations and concurrent - interleavings of the fixture's events (run on the Postgres adapter with parallel transactions), - the final fact row must be bit-identical. +- **Permutation test (a2).** Each fixture event's `received_at` is taken from the fixture, so it + is fixed rather than wall-clock. For every fixture, and for 1,000 random permutations and + concurrent interleavings of its events, the final fact row must be bit-identical. The runs use + the Postgres adapter with parallel transactions, each event inserted with its assigned + `received_at`. - **No onboarding scan is ever byte-capped.** Scoring reads one facts row. - **Frozen class N facts are invalidated and recomputed** whenever: - `start` changes; the row records the `anchored_start` its values used; @@ -717,12 +730,22 @@ the accepted event set:** any order of arrival produces the same row. The permut `first_received_at`. - Onboarding facts (payment, subscription, brand, self-send) are updated **only for the primary subject**. They describe the acting account. - - Counters are incremented for **every index row** whose `subject_kind` is in the counter spec's - `subject_kinds`. That includes `via_parent` rows, so parent-level lifetime totals see child - events, matching what class A queries see through `event_subjects`. + - **Lifetime** counters are incremented for **every index row** whose `subject_kind` is in the + counter spec's `subject_kinds`. That includes `via_parent` rows, so parent-level lifetime + totals see child events, matching what class A queries see through `event_subjects`. + - **H-A: recount-backing counters are primary-subject-only.** Two kinds of counter spec back a + recount: the built-in hourly decline spec, and the counters behind every `before_first` fact + spec. These are incremented **only** for the primary subject's index row. `also` and + `via_parent` rows never touch them, because the facts they feed are primary-only too. So the + incremental path and the recount path always count the same set of events. + - **Test (`TestRecountParentChild`):** child `C` (an `api_key` whose parent is `P`) emits 5 + declines, then `P` gets its first success. Run once with and once without forcing a recount. + `P.declines_before_first_success` must be identical both times (0: the declines are `C`'s, not + `P`'s), and `C`'s own facts must be unaffected. - **Custom `before_first` features** compile to a generic fact spec `{count: {type, where}, before: {type, where}}`, kept in the facts row with the same rules: - lock-first, recount on `∞ → t` and on earlier moves, and daily counters plus a boundary-day query. + lock-first, recount on `∞ → t` and on earlier moves, and primary-only hourly counters plus a + boundary-hour query. The recount uses the partial expression index for the spec's `count` predicate. - **Erasure and re-signup.** - A legal erasure (S3b) deletes the subject's facts row and counter rows, except that numeric @@ -797,9 +820,11 @@ Every `T` is monotone non-decreasing, and so is `P ↦ min(P/B, rc)·d`. A raw p is complete. Stream matching rows newest-first with LIMIT `N = x_sat·⌈W/S⌉`: -- Every loaded row carries at least one raw unit. Rows with addends ≤ 0 are excluded in the - query. Rows are loaded but units summed, so if the limit binds, the loaded rows carry at least - `N` units. +- **"The limit binds"** is determined at run time, never from the static size of `N`: the query + asks for `min(N, budget) + 1` rows, and it binds only if it returns more than `min(N, budget)`, + meaning matching rows remained beyond the limit. +- Every loaded row carries at least one raw unit, because rows with addends ≤ 0 are excluded in + the query. So if `N` binds (rather than the budget), the loaded rows carry at least `N` units. - Split `W` into `⌈W/S⌉` slots of width `S`. By pigeonhole, some slot holds at least `x_sat` loaded units. - The sub-window `(t − S, t]` ending at that slot's last loaded event covers the whole slot. So @@ -814,11 +839,18 @@ lower bound on the true baseline, so the `x_sat` sized from it is smaller than t Either way the result is one-sided upward, flagged `partial`. -The second counter-example is `email.sends_10m_max` with `B = 50`: 5,000 units in one old slot and -about 302 units in each newer slot. Here `x_sat = 50·300/d`. With `d = 1`, `N = 15,000·144 ≈ 2.2M` -rows, far above the 50,000-row budget. So the budget binds first, and the feature is `partial` + -`degraded` rather than silently reporting `v1 ≈ 6`. For fixtures, the budget never binds. In -production this is an honest degradation, not a wrong value. +**Both verification counter-examples are exact under this sizing, because their streams run to +completion:** +- **`custom.declines_10m_peak`** (log1p, cap 7): `x_sat = ⌈e^7 − 1⌉ = 1,096`, so + `N = 1,096 · 144 = 157,824`. The counter-example's stream is about 2,200 rows, which is below + both `N` and the 50,000-row budget, so it completes. The peak is exact: `ln 1001 ≈ 6.9`. +- **`email.sends_10m_max`** (`B = 50`, `rc = 300`, no log1p, cap 300): + `x_sat = 50·min(300, T⁻¹(C)/d) = 50·min(300, 300/d)`, which is 15,000 at `d = 1`, so + `N = 2.16M`. The stream is 5,000 + 302·143 = 48,186 rows, again below `N` and the budget, so it + completes. The value is exact: `v1 = 5,000/50 = 100`. + +A budget can bind only when a real stream is longer than the budget at run time. Only then is the +feature `partial` + `degraded`. **(c) Hitting a bound sets a flag; it's never a weight.** - `Output.Partial` lists features whose own budget was hit. `core.history_truncated` no longer @@ -835,7 +867,7 @@ production this is an honest degradation, not a wrong value. | Legacy `core.linked_*` fan-in cap hit | Value undercounts | `partial` + `degraded` | | Space-saving with `count_gte`, or with a negative sign or weight | Value may undercount | `partial` + `degraded` | | Class G budget hit | Direction unproven | `partial` + `degraded` | -| `ratio` with a partial `num` | Inherits `num`'s flags | — | +| `ratio` with a partial `num` | Inherits `num`'s flags; with a negative ratio sign or weight, always `partial` + `degraded` | — | Revision 3 claimed that fan-in caps apply only to positively weighted features and therefore err upward. **That had the direction backwards:** an undercount lowers a positively weighted value. @@ -898,7 +930,7 @@ bounded evaluation switched on for e2a (§9). | `time_between` | `(−∞, now]`, via indexed first/last-match queries | | `sequence` | `a` rows from `(now − W − within, now]`; `b` rows from `(now − W, now]` | | `relative_to_history` baseline | `(now − lookback, now − exclude_recent]` | -| `before_first` | Facts row (§5.7a), plus one boundary-day query on recount | +| `before_first` | Facts row (§5.7a), plus the hourly-counter sum and one boundary-hour query on recount | | `lifetime` | `subject_counters` rows for the spec, minus an indexed `(now, +∞)` query for specs that exclude future-dated events | | `neighbours` | `links` index `(tenant, kind, hash)` for each `via` key, with `where` applied, up to `K` distinct subjects, under the examined-row budget | @@ -1432,7 +1464,7 @@ These come after #5 and #7 merge. P1 lands before any S5 or S8 PR (enforced as d | P3b | Keys and re-HMAC | `internal/secret` (HKDF, file adapter); re-HMAC of hashes and links with `join_domain`; domain allowlist/HMAC (built-in and declared); pseudonymised non-account and `also` ids; `RedactionSchemaVersion` 3; dev/staging migration job; cross-tenant key test | P3a | Criterion 5 complete; golden exact; migration test | | P3c | Config history | `history/`; `abusekit config check`; `tenant_config_versions` | P2 | History CI tests | | P3d | Rotation and cloud keys | Dual-key write, read and flip; `__prev` fields; `key_id`; cloud secret-manager adapter | P3b | Equality exact across a simulated rotation; adapter contract test | -| P4a | Facts and counters | `start` precedence; `subject_facts` and `subject_counters`; the lock-first ingest transaction; recount on `∞ → t` and on earlier moves (daily decline counters plus a boundary-day query on a partial index); onboarding, brand and self-send facts; generic `before_first` fact specs; webmail counters subtracting `(now, +∞)`; anchored freeze using its **own** anchored-range query (legacy Go code over `[start, start + A)`), with invalidation on `start` change or backfill; fact and counter subject assignment for `also`/`via_parent`; retention-aligned counter expiry; backfill with a `retained_from` watermark | P3a | Golden exact (or `start` deviations justified); blocked-payment and pre-dated flood tests (6b); the permutation and concurrent-interleaving test (a2); the lost-decline race test | +| P4a | Facts and counters | `start` precedence; `subject_facts` and `subject_counters`; the lock-first ingest transaction; READ COMMITTED with bounded retry; recount on `∞ → t` and on earlier moves (primary-only hourly decline counters plus a boundary-hour query on a partial index); `start` clamp and `(received_at, producer, id)` "first accepted"; onboarding, brand and self-send facts; generic `before_first` fact specs; webmail counters subtracting `(now, +∞)`; anchored freeze using its **own** anchored-range query (legacy Go code over `[start, start + A)`), with invalidation on `start` change or backfill; fact and counter subject assignment for `also`/`via_parent`; retention-aligned counter expiry; backfill with a `retained_from` watermark | P3a | Golden exact (or `start` deviations justified); blocked-payment and pre-dated flood tests (6b); the permutation and concurrent-interleaving test with fixture-assigned `received_at` (a2); the lost-decline race test; `TestRecountParentChild` | | P4b | Aggregate engine (class A) | `internal/evalengine` (Postgres and in-memory); `count`, `sum`, `share`, `distinct`, `group_by`, `time_between` + `__absent`, `sequence`, `ratio`; pushdown and indexes; conformance | P4a | Reference equality on 10k histories; adapter equivalence | | P4c | Class R, budgets, flags, flood | `peak` (with `sum`) under the raw-unit saturation limit per transform (§5.7); `neighbours` exact by saturation; the partial/degraded direction table; `ratio` partial propagation and `ratio_den_partial_capable`; `relative_to_history` with its baseline budget; `partial` in verdicts and the API; pass order; `TestFeatureIndependence`; flood generator and property; cost benchmark. **Then enable bounded evaluation for e2a.** | P4b | Criteria 4 and 6; golden exact under the bounded evaluator | | P4d | Rescore control and warm-up | Proportional coalescing; timers only from non-shadow rules; per-tenant budget; warm-up | P4c, P3c | Storm and warm-up tests | @@ -1740,3 +1772,30 @@ Where the re-review's (R3) answer differs from the earlier recommendation, both - Q7 and Q22 are revised. - Q23 (`start` precedence), Q24 (key-independent `body_hash`) and Q25 (dirty marks beyond the cap) are new. + +### Revision 5 (addendum, after the final check of revision 4) + +- **H-A: parent/child recount.** + - The built-in decline counters, and every counter behind a `before_first` recount, are now + primary-subject-only. `also` and `via_parent` rows never feed a recount. + - `TestRecountParentChild` covers it: child `C` emits 5 declines, then parent `P` gets its first + success, and `P`'s value is identical with and without a recount. +- **H-B: "first accepted" and determinism.** + - "First accepted" means the smallest `(received_at, producer, id)`. + - The determinism claim is now stated with each event's `received_at` held fixed. + - The a2 permutation and concurrency tests take `received_at` from the fixture. +- **Fix 1: `peak` worked examples.** + - Both counter-examples are exact, because their streams complete (about 2,200 and 48,186 rows). + - "The limit binds" is decided at run time (more rows remained beyond the limit), never from the + static size of `N`. + - `x_sat` is corrected to `50·min(300, T⁻¹(C)/d)`. +- **Fix 4: ingest transactions and the decline recount.** + - Ingest transactions run at READ COMMITTED, with up to 3 jittered retries on serialization + failure or deadlock. + - The boundary-day recount is replaced by hourly decline counters plus one boundary-hour query, + so the cost an attacker controls is bounded to one hour's declines. +- **Fix 7: `ratio`.** + - (a) Class N `first:` features join the `ratio_den_partial_capable` rejection list. + - (b) A partial-capable `num` combined with a negative ratio sign or weight is flagged + `partial` + `degraded`. +- **Recommended change adopted:** `start = min(chosen, first_received_at)`. From a80527d0eb7487f932203b613fb9fe719a64092a Mon Sep 17 00:00:00 2001 From: jiashuoz Date: Tue, 29 Sep 2026 14:01:38 +0800 Subject: [PATCH 6/8] docs(design): generic feature packs revision 6, domain-neutral binary Owner decision: the compiled binary carries no domain knowledge. Email becomes a YAML reference pack (packs/email) loaded like custom features; content.sent, email_hash, email_domain_class and address_domain move into its declared vocabulary with byte-compatible wire handling (pack extensions of built-in types, flat declared links map). New generic DSL primitives (baseline override, distinct.on_missing, versioned compat options, lifetime share, brand_match, cross-field constraints) give a bit-for-bit parity table for every PR #7 email feature; P-E1 loads the pack in shadow, P-E2 proves parity and deletes the Go email code. Core audit, brand pack lists as data, neutrality CI, non-email reference packs, new decisions on pack location, visibility and pinning. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_014cdM7WyRc3mD3vQNXMTDB8 --- docs/design/2026-09-27-abusekit-design.md | 4 +- .../2026-09-29-generic-feature-packs.md | 345 +++++++++++++++--- docs/plans/2026-09-27-v0-plan.md | 2 +- 3 files changed, 295 insertions(+), 56 deletions(-) diff --git a/docs/design/2026-09-27-abusekit-design.md b/docs/design/2026-09-27-abusekit-design.md index be9b7d6..2a3208b 100644 --- a/docs/design/2026-09-27-abusekit-design.md +++ b/docs/design/2026-09-27-abusekit-design.md @@ -6,8 +6,8 @@ see account churn, tiers failed open when a scorer was unavailable, and request cover reads or replay. This revision fixes those and tightens every place an implementer would have had to guess. Changes from r1 are marked **[r2]**. -**Amendment (proposed 2026-09-29, revision 5):** [`2026-09-29-generic-feature-packs.md`](2026-09-29-generic-feature-packs.md) -renames the built-in features once into namespaced `core`/`email`/`brand` packs enabled per tenant, adds +**Amendment (proposed 2026-09-29, revision 6):** [`2026-09-29-generic-feature-packs.md`](2026-09-29-generic-feature-packs.md) +makes the binary domain-neutral (email becomes a YAML reference pack), renames the built-in features once into namespaced packs enabled per tenant, adds product-declared vocabularies (types, field kinds, link kinds, subject kinds) with pseudonymising redaction, and adds bounded declarative custom features in YAML. It amends §4.2, §4.3, §4.5, §4.6, §4.8 and §4.10 below, and keeps e2a's feature values, risks and tiers bit-identical. diff --git a/docs/design/2026-09-29-generic-feature-packs.md b/docs/design/2026-09-29-generic-feature-packs.md index b1026e4..bb1850e 100644 --- a/docs/design/2026-09-29-generic-feature-packs.md +++ b/docs/design/2026-09-29-generic-feature-packs.md @@ -1,6 +1,8 @@ # Generic feature packs, product vocabularies, and declarative custom features -Status: proposed, revision 5, 2026-09-29. This revision answers the final check of revision 4. Owner: Josh Zhang. +Status: proposed, revision 6, 2026-09-29. This revision applies an owner decision: abusekit is a +generic framework, and **the compiled binary contains no domain-specific knowledge** (no email and +no e2a concepts). Domain knowledge lives only in declared packs (YAML plus data files). Owner: Josh Zhang. This document amends [`2026-09-27-abusekit-design.md`](2026-09-27-abusekit-design.md), §4.2, §4.3, §4.5, §4.6, §4.8 and §4.10. It is written against `main` at S3, treating two open PRs as @@ -36,6 +38,15 @@ Three properties hold throughout: For e2a, feature values, risks, tiers and rescore times stay bit-identical. +**Governing principle (revision 6).** The binary is domain-neutral. Everything specific to a +product domain is configuration: +- event types such as `content.sent`, and email-shaped fields and link kinds; +- webmail lists; +- email features. + +It ships as **reference packs** in YAML and is loaded through the same mechanism as a product's own +custom features. §5.0 defines exactly what stays compiled. + ### Success criteria (measurable) 1. **Bit-exact migration.** A golden replay covers every committed fixture, #7's included, plus the @@ -81,7 +92,13 @@ For e2a, feature values, risks, tiers and rescore times stay bit-identical. **Goals** - A one-time rename to namespaced features that keeps the golden replay bit-exact (§5.2). -- Per-tenant profiles in a private config mount. Only fictional examples live in this repo. +- A domain-neutral binary (§5.0): + - The email features become a YAML **reference pack**, proven bit-identical to #7's Go features + and then deleted from Go. + - A CI check keeps the core neutral. + - A second, non-email reference pack proves the framework end to end. +- Per-tenant profiles in a private config mount. Only fictional examples and reference packs live + in this repo. - Product-declared vocabularies: - types, with field kinds and roles; - `x_` extension fields; @@ -174,7 +191,7 @@ worker (per-tenant fair queue) | `internal/vocab` (new) | `Compile(builtin, decl) (*Vocabulary, error)`; `Redact(*event.Event, Keys) (Stored, error)` | Without it, redaction, kinds, roles and allowlists spread across ingest and packs. | | `internal/secret` (new) | `Keys` plus HKDF `Derive` (§5.6) | Two adapters: file and cloud secret manager. | | `internal/facts` (new) | `Apply(tx, Stored)` at ingest; `Read(tenant, subject)` | Onboarding and lifetime values live in one place. | -| `internal/pack` (new) | `Pack` plus a registry; adapters `core`, `email`, `brand` (Go) and `custom` (DSL) | Four adapters, so the seam is real. | +| `internal/pack` (new) | `Pack` plus a registry. Two Go adapters, `core` and `brand`, both domain-neutral (§5.0). A third adapter, `declared`, compiles any YAML pack: reference packs such as `email` and `card-testing`, and a tenant's `custom` features. | Three adapters, so the seam is real. | | `internal/pack/custom` (new) | `Compile(tenant, []Def, *Vocabulary) (Pack, error)` | Holds every DSL semantic. | | `internal/evalengine` (new) | `Aggregate(ctx, Query) (Result, error)`; Postgres and in-memory adapters | Two adapters (store and eval replay), proven equal by conformance tests. | | `internal/feature` | `Extract(ctx, profile, subject, now) (Result, error)` | Thin orchestrator called by the worker, evaluate and eval. | @@ -182,6 +199,108 @@ worker (per-tenant fair queue) ## 5. Proposed design: detail +### 5.0 The domain-neutral binary (revision 6) + +**What stays compiled into the binary**, and only this: + +1. **The engine.** Ingest and redaction by field kind, keys and pseudonymisation, facts, counters, + the aggregate engine, the DSL compiler and evaluator, scoring, and the HTTP surface. + + The leak scanners (email, card, IP and phone shapes) are privacy detectors, not domain features. + They exist to *reject or mask* personal data in any product's events, and they appear on the + neutrality allowlist with that justification. +2. **Built-in core vocabulary.** After revision 6, "built-in" means types, fields and link kinds + whose schema is compiled into the binary and needs no pack. There are exactly: + - **Subject lifecycle:** + - `subject.created`: `channel`, `identity_kind`, `account_created_at`; + - `subject.deleted`: `mode`; + - `subject.class`: `class`. + - **Payment:** `payment.attempt`: `outcome`, `reason`, `funding`, `amount_minor`, `currency`. + - **Subscription:** `subscription.changed`: `plan`, `status`, `amount_minor`. + - **Resource:** `resource.created` / `resource.deleted`: `kind`, and `name`, which carries the + role `display_name`. + - **Verdict:** `content.verdict`: `source`, `category`, `score`. + - **Label:** not an event type. Labels arrive through `POST /v1/labels`, whose vocabulary + (`benign`/`abusive` plus each rule's labels) stays built-in. + - **Link kinds:** `card_fingerprint_hash`, `device_hash`, `ip24_hash`, `ua_hash` and `asn`. +3. **Two Go packs, both domain-neutral:** + - `core@1`. It reads only built-in core types, declared roles (`activity`, `credential`, + `self`) and link kinds marked `evidence`. + - `brand@1`, the impersonation matcher described below. + +**Everything else is declared.** Declarations live in a YAML pack or in a tenant profile, and they +include: +- `content.sent`: its fields `recipient_domain`, `recipient_is_own_identity`, `recipient_hash`, + `recipient_count`, `subject_line` and `first_link_host`, its cross-field rule, and its + masking and raw-storage settings; +- the `email_hash` link kind; +- the `subject.created.email_domain_class` and `resource.*.address_domain` fields; +- the webmail list; +- every `email.*` feature. + +The **email reference pack** (`packs/email/`) declares all of these. The wire contract stays +byte-compatible: e2a keeps sending `content.sent`, `links.email_hash` and `email_domain_class` +exactly as today. + +**Two generic vocabulary mechanisms make that possible:** +- **Pack extensions of built-in types.** A *pack* (not a tenant) may add fields to a built-in type + under unprefixed names, for example `email_domain_class` on `subject.created` and + `address_domain` on `resource.*`. Two enabled packs extending the same name fail with + `extension_conflict`. Tenant extensions keep the `x_` prefix. +- **A flat `links` map.** `links` is a map of **declared** link kinds, built-in or pack-declared, + so `links.email_hash` stays a top-level key. Revision 2's `links.custom` is dropped. An + undeclared key is rejected with `bad_links`. Which kinds count as neighbour evidence is set by + `evidence: true` on the declaration, never hard-coded. The email pack marks `email_hash` as + evidence, which reproduces today's default of `{email_hash, card_fingerprint_hash, + device_hash}` when the email pack is enabled. + +**Core neutrality audit.** Every revision-5 core feature was checked: + +| Feature | Hidden email assumption | Revision 6 | +| --- | --- | --- | +| `core.burst_ratio_24h_vs_lifetime` | It counted `content.sent` by type name. | Generalised: it counts `resource.created` plus every type with role `activity`. The email pack declares `content.sent` as `activity`, so values are identical. | +| `core.email_domain_class_disposable` (new in revision 2) | It reads `email_domain_class`. | Moved to `email.domain_class_disposable`. | +| `core.linked_*`, `core.fingerprint_seen_on_other_subjects` | The default evidence kinds included `email_hash`. | The evidence set is declared: `evidence: true` on link kinds. The fingerprint feature keeps reading the built-in `card_fingerprint_hash`. | +| `core.subject_age_h`, `upgrade_*`, `declines_*`, `first_funding_prepaid`, `resource_*`, `credential_*`, `verdict_max_24h`, `neighbors_truncated` | None | Unchanged | +| Self-send, webmail and destination logic | Previously Go facts and counters | Now entirely in the email pack, as declared `before_first` fact specs and lifetime share counters over the roles `self`, `recipient` and `destination` | + +**Roles (field level).** +- `title`: text matched by `brand.title_match`. +- `display_name`: an actor-chosen name matched by `brand.name_match`. The core + `resource.*.name` field carries it. +- `self`: bool; a self-directed event. +- `destination`: domain or hash; where an activity lands. +- `recipient`: hash; a single counterparty. Pairs with `recipient_count`. + +Roles only let **neutral** code (core and brand) find fields generically. The DSL always names +fields explicitly. + +**The brand pack: what remains Go, and why that is neutral.** +- **In Go:** the matcher. It applies NFKC and a confusables skeleton, strips `Cf` characters, + tokenises on Unicode punctuation and symbols, splits camelCase, and matches on token boundaries + (including multi-word sequences and case-sensitive entries). It also runs two gate mechanisms, + the integration-token gate and the community-phrase gate. All of this is string processing over + *any* field with role `title` or `display_name`, and it knows nothing about mail or any product. +- **Moved from Go literals to data:** the brand list (`packs/brand/brands.yaml`, plus a tenant's + private `extra`); the integration tokens (`packs/brand/integration_tokens.yaml`, formerly + `buildIntegrationTokens`); and the community phrases (`packs/brand/community_phrases.yaml`, + formerly `buildCommunityPhrases`). +- **Exposed** as features `brand.name_match`, `brand.name_has_at` and `brand.title_match`, and as + one generic DSL op, **`brand_match`** (§5.5), which declared packs may use. + +**Neutrality CI.** P-N (§9) adds three checks, which then gate every slice: +1. `TestCoreIsDomainNeutral`. It walks the Go AST of every package outside `packs/`, `examples/` + and `testdata/`, and fails on any identifier or string literal on + `internal/neutrality/denylist.txt` (for example `email`, `webmail`, `recipient`, `subject_line`, + `smtp`, `mailbox`, `inbox`, `e2a`). Exceptions come only from `allowlist.txt`, one justified + entry per line; the privacy leak detector is the canonical one. +2. `TestNonEmailProfileEndToEnd`. The binary loads the card-testing reference pack + (`examples/tenants/tallyport`) with **no email pack**. It ingests that pack's fixtures through + HTTP and scores them. It asserts that no `email.*` feature is registered, that `content.sent` is + an undeclared type for that tenant, and that verdicts match the pack's golden replay. +3. `TestGoRegistryLint`. Every Go `FeatureDef` may read only built-in core types or roles + (`Reads`). Any type name outside the built-in list fails. + ### 5.1 Packs ```go @@ -231,16 +350,11 @@ type Output struct { | `core.fingerprint_seen_on_other_subjects`, `core.linked_deleted_n`, `core.linked_labelled_abusive_n`, `core.neighbors_truncated` | F | neighbour evidence, built-in link kinds only (§5.5) | | `core.resource_velocity_1h`, `core.credential_velocity_1h` (`key_velocity_1h`) | A | exact 1 h counts | | `core.resource_total`, `core.credential_total` (`key_total`) | F | `subject_counters` (lifetime) | -| `core.burst_ratio_24h_vs_lifetime` | A + F | 24 h count (A) ÷ lifetime counters | -| `email.sends_1h`, `email.webmail_sends_1h`, `email.distinct_recipients_1h` | A + R | exact 1 h current value (A); history baseline (R). For `distinct_recipients_1h`, the current value is two class A aggregates: `count(DISTINCT recipient_hash)` plus the sum of `recipient_count` over rows without a hash. | -| `email.sends_10m_max` | R | `peak` with a saturation limit (exact); history baseline (R) | -| `email.sends_first_day`, `email.first_day_distinct_domains` | N | anchored; frozen into facts after day one | -| `email.webmail_recipient_share` | F | lifetime counters: webmail recipients vs all non-self recipients | -| `email.self_send_before_external` | F | `subject_facts.self_sends_before_first_external` | -| `email.subject_brand_match` (requires `brand`) | A + F | distinct ingest-matched brand ids in 1 h, minus the exempt set held in facts | -| `brand.name_match`, `brand.name_has_at` | F | `subject_facts.name_brands`, `name_has_at` (matched at ingest) | +| `core.burst_ratio_24h_vs_lifetime` | A + F | 24 h count (A) ÷ lifetime counters, over `resource.created` + `activity` types | +| `email.*`: all nine #7 email features, plus `email.domain_class_disposable` | per DSL op | **Declared** in `packs/email/pack.yaml`, not Go (§5.5 parity table). Until the parity slice deletes it, the #7 Go code is transitional. | +| `brand.name_match`, `brand.name_has_at` | F | `subject_facts.name_brands`, `name_has_at`, matched at ingest over `display_name` fields | | `brand.title_match` (new) | A | distinct brand ids across non-self `title` fields in 1 h | -| `core.email_domain_class_disposable`, `core.verdict_max_24h` (new) | F, A | not part of e2a's rule | +| `core.verdict_max_24h` (new) | A | not part of e2a's rule | For the rows above, the code changes from scanning every event to reading facts, counters and exact aggregates. The arithmetic doesn't change: it operates on the same inputs and sums integers, @@ -339,7 +453,7 @@ vocabulary: **Roles.** Only these roles ship: - resource kinds: `credential` and `other` (an undeclared kind counts as `other`); - event type: `activity`; -- fields: `title` and `self`. +- fields: `title`, `display_name`, `self`, `destination` and `recipient` (§5.0). `brand.title_match` counts distinct brand ids in the `title` fields of non-self activity events over the trailing hour, capped at 3. @@ -360,8 +474,11 @@ over the trailing hour, capped at 3. way. - Rules declare `applies_to`. -**Link kinds.** `links.custom` holds up to 8 declared kinds, and every value is re-HMACed. -Declared kinds feed only `neighbours` (§5.5). +**Link kinds.** `links` is a flat map of declared kinds (§5.0), holding at most 8 non-built-in +kinds, and every value is re-HMACed. Kinds declared with `evidence: true` feed the `core.linked_*` +features when their pack declares them as core evidence (the email pack does this for +`email_hash`). They also feed `neighbours`. A tenant's own declared kinds feed only `neighbours`, +so core evidence is never double-fed. **Legacy default.** A profile without `resource_kinds` gets `key: credential`, with #7's aliases. @@ -436,9 +553,8 @@ predicates, `distinct`, `group_by` or `on`. All examples do. **Neighbour evidence stays separate** -- `core.linked_*` and `core.fingerprint_*` read only built-in link kinds admitted by the tenant's - `feature.Config`. -- Declared kinds feed only `neighbours`, so they never double-feed `core.linked_deleted_n`. +- `core.linked_*` read only link kinds that a *pack* declares as core evidence (`evidence: true`), plus the built-in kinds the tenant's `feature.Config` admits. `core.fingerprint_*` reads the built-in `card_fingerprint_hash`. +- Kinds a *tenant* declares feed only `neighbours`, so they never double-feed `core.linked_deleted_n`. - **Dirty-mark propagation** covers built-in *and* declared evidence kinds. When a subject gains a link key, is permanently deleted, or is labelled, every subject sharing any evidence key with it is marked dirty. @@ -473,28 +589,46 @@ value = transform( v1 · d ) - `age_decay` defaults to `{full_until: 3d, zero_at: 30d, floor: 0.2}`. - `baseline: {peak: {size: 10m}}` overrides the baseline op. -**Which #7 and S2 features the DSL can express** - -Expressible: -- `sends_10m_max`, `sends_1h` and `webmail_sends_1h`, via `baseline`, `sum: recipient_count`, - `cap_each: 300` and `ratio_cap: 300`; -- `sends_first_day`; -- `webmail_recipient_share`; -- `declines_before_first_success`; -- `self_send_before_external`, with `cap: 2`; -- `resource_*` and `credential_*`; -- `upgrade_delay_min`. - -Not expressible: -- `distinct_recipients_1h`: distinct hashes plus a fallback sum; -- `subject_brand_match` and `brand.*`: they need the matcher's gates; -- `first_day_distinct_domains`: its window end is inclusive; -- `burst_ratio_24h_vs_lifetime`: its lifetime denominator counts future-dated events; -- `linked_*`: built-in evidence semantics; -- `subject_age_h`: it reads `start`. - -P5b re-expresses every expressible feature in the DSL and checks it bit-for-bit against the Go -implementation. +**Generic primitives added in revision 6, so the email pack needs no Go.** Each one is +domain-neutral and usable by any pack: + +| Primitive | Semantics | +| --- | --- | +| **Baseline override** `baseline: {op, size?, where?}` | The baseline uses its own op, sub-window width and predicate, independent of `cur`. For example, a 1 h sum compared against a prior 10 min peak over *all* non-self rows. | +| **`distinct` with `on_missing`** `{field, on_missing: ignore \| count \| {sum: {field, default}}}` | Rows without `field` are ignored (the default), counted as one unit each, or contribute a summed field. The value is `count(DISTINCT field) + fallback`. | +| **Compat options** `compat: {version: 0, window_end: closed, include_future: true}` | Explicit, versioned reproductions of legacy quirks. `window_end: closed` makes an anchored window `[start, start + A]`. `include_future: true` stops excluding `at > now` for that feature. The loader rejects `compat` in tenant `custom.*` features, so only reference packs may freeze a legacy quirk, and each quirk is listed in the pack's changelog. `compat.version` is bumped if a quirk's semantics ever change. | +| **`share` over `lifetime`** | Numerator and denominator come from two counter specs, with the `(now, +∞)` subtraction (unless `include_future`). | +| **`brand_match`** `{role: title \| display_name, window, exclude_self, exempt_on_display_name_token: integration, exclude_brands_of: display_name, cap}` | Counts distinct brand ids matched in `role` fields, provided by the brand pack. It can exclude brands whose display-name mention sits next to an integration token, and brands already credited through `display_name`. | +| **Cross-field constraint** (vocabulary) `constraints: [{if_present: f, then: {field: g, lte: 1}}]` | Validated at redaction. A violation is rejected with `redaction_failed`. | + +**Parity: how every #7 email feature is declared** (`packs/email/pack.yaml`, all over +`content.sent`, with `self` = `recipient_is_own_identity`): + +| Feature | Declaration | +| --- | --- | +| `email.sends_10m_max` | `peak: {size: 10m, sum: {field: recipient_count, default: 1, cap_each: 300}, where: not self}`, `window: 24h`; `relative_to_history: {lookback: 30d, exclude_recent: 24h, age_decay: {}, ratio_cap: 300}`; `cap: 300` | +| `email.sends_1h` | `count` with `sum` (as above) over 1 h, not self; baseline override `{peak, size: 10m, where: not self}`; `ratio_cap: 300`; `age_decay` | +| `email.webmail_sends_1h` | As `sends_1h`, with `where: {all: [not self, {field: recipient_domain, in_set: webmail}]}`. The baseline override keeps `where: not self`. | +| `email.distinct_recipients_1h` | `distinct: {field: recipient_hash, on_missing: {sum: {field: recipient_count, default: 1, cap_each: 300}}}` over 1 h, not self; baseline override `{peak, size: 10m, sum: recipient_count}`; `age_decay` | +| `email.sends_first_day` | `count` with `sum` over `first: 24h`, not self; `compat: {version: 0, window_end: closed}` | +| `email.first_day_distinct_domains` | `distinct: {field: recipient_domain}` over `first: 24h`; `log1p: {anchored_at: 10}`; `compat: {version: 0, window_end: closed, include_future: true}` | +| `email.webmail_recipient_share` | `share` over `lifetime`, `sum: recipient_count`, `where: not self`, `match: in_set webmail` | +| `email.self_send_before_external` | `count: {where: self}`, `before_first: {type: content.sent, where: not self}`, `cap: 2` | +| `email.subject_brand_match` | `brand_match: {role: title, window: 1h, exclude_self: true, exempt_on_display_name_token: integration, exclude_brands_of: display_name, cap: 3}`; `age_decay` | + +**Wire constants and edge conventions.** +- `cap_each: 300` and `default: 1` reproduce `recipientCountOf`: an invalid or non-positive count + reads as 1, and each value is capped at 300. +- `normalizeToken` is the `domain` kind's normalisation. +- `log1p: {anchored_at: 10}` computes `s = 10/math.Log1p(10)` exactly as the Go constant does. +- Parity for each feature is **bit-for-bit on the golden replay** (slice P-E2). Any mismatch + blocks deletion of the Go code. + +**Features that stay Go, all neutral:** +- `core.subject_age_h`, which reads `start`; +- `core.linked_*` and `core.fingerprint_*`, which use the evidence semantics; +- `core.burst_ratio_24h_vs_lifetime`, a frozen core quirk that counts future-dated events; +- the `brand.*` features. **Transform.** `cap` is mandatory: finite, greater than 0, and at most 1 for `share`. `log1p` is optional and applied before the cap: @@ -556,7 +690,9 @@ hashed values too. | `domain` | Lower-cased and IDNA-encoded to ASCII. Must end in a public suffix from the pinned PSL snapshot, or in an RFC 6761 special-use name. **Defaults to `reduce: etld1`**; `none` opts out. Stored **in cleartext only if allowlisted**: the email pack's webmail list, the disposable list, or the vocabulary's `popular_domains` set file (top-N). Anything else is stored as `hd:` + HMAC. | IP literals, all-numeric labels, unknown suffixes, or an `@` | | `timestamp` | RFC 3339. | Anything else | -**Built-in fields in P3b** (`RedactionSchemaVersion` becomes 3) +**Email-pack fields in P3b** (`RedactionSchemaVersion` becomes 3). These fields' schema lives in the +email pack from P-E1. Until then it sits in the transitional built-in schema; the handling is +identical either way. | Field | Reduction | Stored as | Feature impact | | --- | --- | --- | --- | @@ -695,7 +831,7 @@ permutation test in §5.7 (a2) checks this. The update rules: | `declines_before_first_success` | On a `declined` event with `at ≤ first_success_at` (or no success yet): increment. **Recount** whenever `first_success_at` changes, both from absent to `t` and from an earlier move. The recount is the Σ of the **hourly** decline counters for hours before `hour(t)`, plus one boundary-hour query over `[hour(t), t]`. The counters are a built-in `subject_counters` spec, **primary-subject-only** (see H-A below). The query uses a partial index on declined payment attempts: `(tenant, kind, subject, at) WHERE type = 'payment.attempt' AND data->>'outcome' = 'declined'`. Cost: one indexed SUM over at most 2,160 hourly rows (90-day retention), plus the declines inside a single hour. That bounds the part an attacker controls to one hour's events. | | `first_paid_upgrade_at` | `LEAST` over `subscription.changed{status: active, amount_minor > 0}` | | `subscription_change_count` | Increment | -| `first_external_at`, `self_sends_before_first_external` | The same pattern for `content.sent`; capped at 2 when read | +| Declared `before_first` fact specs | The same pattern, generically. The email pack's `self_send_before_external` is one of these specs (from P-E1; before that, a transitional Go fact) | | `name_brands`, `name_has_at`, `exempt_subject_brands` | Brand ids matched at ingest on `resource.*` names, with the integration-token gate applied | - **Isolation.** Ingest transactions run at **READ COMMITTED**. On a serialization failure or a @@ -781,8 +917,8 @@ this. The same argument covers any flood of any type the update rules above don' the vocabulary. They cover: - `resource.created`, all kinds; - `resource.created` with role `credential`; - - `resource.created` + `content.sent` + activity types (for `burst_ratio`); - - non-self `content.sent` recipients, all and webmail-only; + - `resource.created` + `activity` types (for `burst_ratio`); + - declared lifetime `share` specs (the email pack's non-self recipients, all and webmail-only); - custom `lifetime` counts. - A new spec backfills from retained events and stays cold until the backfill completes. - Values are integers, so every sum is exact. @@ -1070,7 +1206,7 @@ from the numeric retention for abusive subjects (main §4.4). | Surface | Change | Compatibility | | --- | --- | --- | -| `POST /v1/events` | Optional `subject_kind`, `also`, `links.custom` and `x_` fields; declared types; the redaction in §5.6; account-subject leak scan (`bad_subject`) | The wire format is additive. Storage semantics change (pre-GA). | +| `POST /v1/events` | Optional `subject_kind`, `also`, declared keys in the flat `links` map, and `x_` fields; pack-declared types such as `content.sent` (unchanged on the wire); declared types; the redaction in §5.6; account-subject leak scan (`bad_subject`) | The wire format is additive. Storage semantics change (pre-GA). | | `GET /v1/subjects/{id}`, `evaluate` | `?kind=`; per-tenant `model`; namespaced reasons; `warming_until`; `partial` | Additive | | Per-item codes | None new | Unchanged | | `score --jsonl`, corpus loader | `feature_renamed`; `corpus-v2` | Breaks once, pre-GA (P1) | @@ -1428,7 +1564,8 @@ The first two already have a channel: `content.verdict` and enum fields. - **Emitter:** no change. S6 emits `content.sent`. - **Profile:** a byte copy of `examples/tenants/reference/`. - - Packs: `packs: [core@1, email@1, brand@1]`. + - Packs: `packs: [core@1, brand@1, email@1]`. `email@1` is the YAML reference pack from P-E1 on. + Before P-E1, email features come from the transitional Go code in #7. - The legacy vocabulary made explicit: `key: credential` with #7's aliases, and `agent: other`. - Weights and the private brand list stay in the private mount. - **P0:** `abusekit eval --golden` extends the existing replay. @@ -1447,12 +1584,22 @@ The first two already have a channel: `content.verdict` and enum fields. Any fixture that changes is listed and justified in its slice. - **Bounded evaluation reaches e2a only after P4a–P4c** (§5.7 f). +- **Email becomes configuration (P-E1, P-E2).** + - P-E1: the YAML email pack is loaded for e2a *alongside* the transitional Go email features, + under shadow names (`email_yaml.*`). + - P-E2: the golden replay asserts that each YAML feature is bit-identical to its Go twin at every + recorded instant. Only then does it switch e2a to the YAML names and delete the Go email code. + e2a's scores are bit-identical throughout. +- **#7 merges first**, as the transitional Go implementation. It protects e2a before the + migration, and it is the parity oracle for P-E2. - **Rollback:** before S8, every slice can be reverted on its own. After S8, only the rename is irreversible without a rescore, which is why P1 comes first. ## 9. Slices -These come after #5 and #7 merge. P1 lands before any S5 or S8 PR (enforced as described in §5.2). +These come after #5 and #7 merge. **#7 merges first**, as the transitional Go implementation of the +email features, so e2a is protected before the migration. P1 lands before any S5 or S8 PR, enforced +as described in §5.2. | # | Slice | Contents | Depends on | Done when | | --- | --- | --- | --- | --- | @@ -1468,10 +1615,12 @@ These come after #5 and #7 merge. P1 lands before any S5 or S8 PR (enforced as d | P4b | Aggregate engine (class A) | `internal/evalengine` (Postgres and in-memory); `count`, `sum`, `share`, `distinct`, `group_by`, `time_between` + `__absent`, `sequence`, `ratio`; pushdown and indexes; conformance | P4a | Reference equality on 10k histories; adapter equivalence | | P4c | Class R, budgets, flags, flood | `peak` (with `sum`) under the raw-unit saturation limit per transform (§5.7); `neighbours` exact by saturation; the partial/degraded direction table; `ratio` partial propagation and `ratio_den_partial_capable`; `relative_to_history` with its baseline budget; `partial` in verdicts and the API; pass order; `TestFeatureIndependence`; flood generator and property; cost benchmark. **Then enable bounded evaluation for e2a.** | P4b | Criteria 4 and 6; golden exact under the bounded evaluator | | P4d | Rescore control and warm-up | Proportional coalescing; timers only from non-shadow rules; per-tenant budget; warm-up | P4c, P3c | Storm and warm-up tests | -| P5 | Pack gating | Registry; `core`, `email` and `brand` adapters; enablement; `brand.title_match` (needs the `title` role); `packtest`; starter weights | P3a, P4c | Golden exact; `feature_not_enabled`; every pack passes `packtest` | -| P5b | DSL parity | Every expressible #7/S2 feature re-expressed in the DSL | P4c | Bit-exact against the Go feature on every fixture | -| P6a | Link kinds, `neighbours`, subject kinds | `links.custom`; `neighbours` (as of now); dirty propagation for declared kinds; `subject_kind`, `also`, `parent`, `event_subjects` (≤ 8); `?kind=`; `applies_to`; `x_primary_subject_hash` | P3b, P4b, P5 | Contract tests; propagation and no-double-feed tests; golden exact | -| P6b | Scenarios and bootstrap | Five example profiles with held-out fixtures; uniform priors; `--profile`; floors with `profile:` | P4d, P5, P6a | Criterion 2 on held-out fixtures; isolation check | +| P5 | Pack gating | Registry; the Go `core` and `brand` adapters and the `declared` (YAML) adapter; enablement; `brand.title_match` (needs the `title` role); brand lists moved to data; `packtest`; starter weights | P3a, P4c | Golden exact; `feature_not_enabled`; every pack passes `packtest` | +| P-E1 | YAML email pack | `packs/email/` (`pack.yaml`, `webmail.txt`, starter weights, fixtures, floors). The new generic primitives: baseline override, `distinct.on_missing`, compat options, lifetime `share`, `brand_match`, cross-field constraints, pack extensions of built-in types, and the flat declared `links` map. The `content.sent` schema, `email_hash`, `email_domain_class` and `address_domain` move into the pack, wire-compatible. Loaded for e2a as shadow `email_yaml.*`. | P4b, P4c, P5 | The pack loads; wire replay of every fixture is byte-compatible (same accepted, duplicate and conflict results); golden exact (Go features still drive scores) | +| P-E2 | Parity, then deletion | The golden replay compares each `email_yaml.*` value to its Go twin, bit for bit, at every instant. Then: switch the names to `email.*`, **delete** the Go email features and the transitional built-in `content.sent` schema, and turn on the neutrality CI (§5.0) | P-E1 | Parity passes for all nine features; golden exact after deletion; `TestCoreIsDomainNeutral`, `TestNonEmailProfileEndToEnd` and `TestGoRegistryLint` green | +| P5b | Core DSL parity | Every expressible *core* S2 feature re-expressed in the DSL (a conformance check; the core Go features stay, being neutral) | P4c | Bit-exact against the Go feature on every fixture | +| P6a | Link kinds, `neighbours`, subject kinds | declared link kinds in the flat `links` map; `neighbours` (as of now); dirty propagation for declared kinds; `subject_kind`, `also`, `parent`, `event_subjects` (≤ 8); `?kind=`; `applies_to`; `x_primary_subject_hash` | P3b, P4b, P5 | Contract tests; propagation and no-double-feed tests; golden exact | +| P6b | Scenarios, reference packs and bootstrap | Five example profiles with held-out fixtures. **Non-email reference packs** `packs/card-testing/` (from 7b) and `packs/api-credential-abuse/` (from 7d), each in YAML with fixtures, floors and starter weights. Uniform priors; `--profile`; floors with `profile:` | P4d, P5, P6a, P-E2 | Criterion 2 on held-out fixtures; isolation check; both non-email reference packs pass `packtest` with no email pack loaded | | P7 | e2a cutover | Private profile in the ops mount; hosted `config check` | P5, P4c, S8's mount | Golden exact against the private copy | Two v0 slices interact with this plan: @@ -1560,7 +1709,8 @@ Where the re-review's (R3) answer differs from the earlier recommendation, both 4. **Re-HMAC of built-in fields:** `recipient_hash` and every link, with a dev/staging migration job. Approve? 5. **Legacy window quirks:** freeze them in `@1` and harmonise in `@2`. Unchanged. Confirm? -6. **Roles:** `credential`, `other`, `activity`, `title`, `self`. Unchanged. Confirm? +6. **Roles:** revision 5 had `credential`, `other`, `activity`, `title` and `self`. Revision 6 adds + the field roles `display_name`, `destination` and `recipient` (§5.0). Confirm? 7. **Bounding:** - Revision 2: byte caps, a shared step budget and a weighted truncation feature. - R3: ingest facts, counters, per-feature budgets, and truncation as a flag. @@ -1577,7 +1727,8 @@ Where the re-review's (R3) answer differs from the earlier recommendation, both 10. **Brand list:** can a tenant narrow it as well as extend it? Unchanged. 11. **Custom namespace:** `custom.*` per tenant, or `.*`? Unchanged. 12. **Config history:** in the config tree, checked in CI; the database is only a guard. Approve? -13. **S6:** emits `content.sent`. Settled. +13. **S6:** emits `content.sent`, unchanged on the wire. From P-E1, the email reference pack + declares it, not the binary. Settled. 14. **Domains:** - Revision 2: public-suffix check, stored in cleartext everywhere. - R3: default to eTLD+1, keep cleartext only for allowlisted or popular domains, and HMAC the @@ -1625,6 +1776,36 @@ Where the re-review's (R3) answer differs from the earlier recommendation, both detection. Approve? 25. **Dirty marks beyond the fan-in cap (revision 4, new):** the first 50 per key are marked in the ingest transaction, and the rest by a rate-budgeted background job. Confirm? +26. **Where reference packs live and how they are versioned (revision 6, new).** + - Proposed: `packs//` in this repo. Each pack has a `pack.yaml` (name, semver version, + `requires`), data files, starter weights, fixtures and floors, plus a `CHANGELOG.md` that + lists every `compat` quirk. + - Major versions change semantics, minor versions only add, and patches change data only. + - A pack's content SHA (over its whole directory) is recorded in every verdict's + `profile_sha`. + - Packs ship with abusekit releases, but the loader also accepts them from a tenant's mount. + Approve? +27. **Are reference packs part of the public repo? (revision 6, new)** Proposed: yes. `email`, + `brand` lists, `card-testing` and `api-credential-abuse` hold only public facts (provider and + brand names) and synthetic fixtures. Private additions (brand extras, tenant packs) live only + in the tenant mount. Approve? +28. **How a product pins a pack version (revision 6, new).** + - `packs: [email@1]` pins a major and takes the newest compatible minor or patch on the search + path. + - `email@1.4.2` pins exactly. + - `email@1.4.2#sha256:` also pins the content. + - `abusekit config check` records the resolved version and SHA in the tenant's history, and a + change of either is a history entry. + - A patch or minor must pass the pack's golden replay before release. + Approve? +29. **Tenant-private YAML packs (revision 6, new).** May a product ship its own YAML pack under + its own namespace from the private mount, not only `custom.*` features? Proposed: yes, with + reserved-name and collision checks. Such packs may not use `compat` options. Confirm? +30. **Frozen quirks as `compat` options (revision 6, new).** Legacy quirks move out of Go and into + explicit, versioned `compat` options usable only by reference packs: + `email.first_day_distinct_domains` gets a closed window end and includes future-dated events, + and `email.sends_first_day` gets a closed end. Harmonising them later means a new pack major + with the option removed. Approve? ## 13. Changes from earlier revisions @@ -1799,3 +1980,61 @@ Where the re-review's (R3) answer differs from the earlier recommendation, both - (b) A partial-capable `num` combined with a negative ratio sign or weight is flagged `partial` + `degraded`. - **Recommended change adopted:** `start = min(chosen, first_received_at)`. + +### Revision 6 (addendum: owner decision, a domain-neutral binary) + +**Principle.** The compiled binary contains no domain-specific knowledge: no email and no e2a +concepts. §5.0 defines exactly what stays compiled: +- the engine; +- the built-in core vocabulary (subject lifecycle, payment, subscription, resource, verdict, + label); +- the built-in neutral link kinds; +- the neutral Go packs `core` and `brand`. + +**The email pack becomes configuration.** +- It becomes `packs/email/`: `pack.yaml` plus `webmail.txt`, starter weights, fixtures and + floors. It is loaded by the same `declared` adapter as tenant `custom.*` features. +- The Go email features from #7 are transitional: + - **P-E1** loads the YAML pack in shadow; + - **P-E2** proves bit-for-bit parity on the golden replay, then deletes the Go email code in the + same slice. +- New generic DSL primitives, so no Go special cases remain: + - baseline override (own op, width and predicate); + - `distinct.on_missing`; + - versioned `compat` options (closed window end, include future-dated events); + - lifetime `share`; + - a `brand_match` op; + - vocabulary cross-field constraints. + +**Core vocabulary.** +- `content.sent` and its fields, the `email_hash` link kind, `email_domain_class` and + `address_domain` all move into the email pack's declared vocabulary. The wire stays + byte-compatible, via pack extensions of built-in types and a flat declared `links` map, which + replaces `links.custom`. +- New field roles: `display_name`, `destination` and `recipient`. + +**Core audit.** +- `core.burst_ratio` is generalised to `activity` types. +- `core.email_domain_class_disposable` moves to `email.domain_class_disposable`. +- The neighbour evidence set is declared rather than hard-coded. +- Self-send and webmail logic is now declared only. + +**Brand pack.** +- The matcher, including its gates, stays Go. It is neutral string processing over `title` and + `display_name` fields. +- The brand list, integration tokens and community phrases become data files. + +**Neutrality CI.** +- `TestCoreIsDomainNeutral`: an AST denylist, with a justified allowlist. +- `TestNonEmailProfileEndToEnd`: the card-testing profile, with no email pack loaded. +- `TestGoRegistryLint`. + +**Reference packs.** Besides `email`, P6b ships `card-testing` and `api-credential-abuse` in YAML. + +**Slices.** +- #7 merges first, as the transitional implementation. +- P-E1 and P-E2 are added, and P5, P5b, P6a and P6b are revised. +- e2a's scores stay bit-identical throughout. + +**Decisions.** Q6 and Q13 are revised. Q26–Q30 are new: where reference packs live and how they +are versioned, whether they are public, pack pinning, tenant-private packs, and `compat` options. diff --git a/docs/plans/2026-09-27-v0-plan.md b/docs/plans/2026-09-27-v0-plan.md index fe85e79..428f5aa 100644 --- a/docs/plans/2026-09-27-v0-plan.md +++ b/docs/plans/2026-09-27-v0-plan.md @@ -35,7 +35,7 @@ design pass) rather than deferred to v1 outright. | S7 | Billing events | ops sidecar: `payment.attempt` (with `card_fingerprint_hash` under the tenant key) and `subscription.changed` | staging checkout produces events | | S8 | Hosted deploy | ops: compose service, Secret Manager keys, tenant config, Terraform alerts for queue depth / budget / drops | abusekit running on prod in shadow | | S9 | Incident evaluation | private backfill of the incident accounts and a benign sample into a private corpus; harness run; report precision/recall/lead-time before first send per account | report reviewed; floors set; decision on `advise` for the local rule | -| P0–P7 | Generic feature packs (own design: `docs/design/2026-09-29-generic-feature-packs.md`) | After S4 (#5) and S2b (#7) merge: golden replay (P0), one-time rename (P1), stage gate (P1s), tenant profiles (P2), vocabulary and scans (P3a), keys and re-HMAC (P3b), config history (P3c), rotation and cloud keys (P3d), facts and counters (P4a), aggregate engine (P4b), row features, flags and flood property (P4c), rescore control and warm-up (P4d), pack gating (P5), DSL parity (P5b), link and subject kinds (P6a), scenarios (P6b), e2a cutover (P7). **P1 lands before any S5/S8 PR** (S5 and S8 are blocked on P1; enforced by `TestNoProductionBeforeRename`, which guards vendor adapters and production startup). Bounded evaluation reaches e2a only after P4a–P4c. S3b erasure must be vocabulary-aware and clear facts/counters. | per-slice "Done when" in that design's §9; P1 golden bit-exact for values, risks and tiers, exact after every later slice | +| P0–P7 | Generic feature packs (own design: `docs/design/2026-09-29-generic-feature-packs.md`) | After S4 (#5) and S2b (#7) merge: golden replay (P0), one-time rename (P1), stage gate (P1s), tenant profiles (P2), vocabulary and scans (P3a), keys and re-HMAC (P3b), config history (P3c), rotation and cloud keys (P3d), facts and counters (P4a), aggregate engine (P4b), row features, flags and flood property (P4c), rescore control and warm-up (P4d), pack gating (P5), YAML email reference pack in shadow (P-E1), email parity then deletion of the Go email code (P-E2), core DSL parity (P5b), link and subject kinds (P6a), scenarios and non-email reference packs (P6b), e2a cutover (P7). **P1 lands before any S5/S8 PR** (S5 and S8 are blocked on P1; enforced by `TestNoProductionBeforeRename`, which guards vendor adapters and production startup). Bounded evaluation reaches e2a only after P4a–P4c. S3b erasure must be vocabulary-aware and clear facts/counters. #7 merges first as the transitional Go email implementation; the binary becomes domain-neutral at P-E2. | per-slice "Done when" in that design's §9; P1 golden bit-exact for values, risks and tiers, exact after every later slice | S1–S4 (including S3b) are pure abusekit and can run back to back; S5 needs vendor keys; S6–S8 are e2a/ops work that can start after S3 (S3b is not a blocker for them — nothing in S6–S8 depends on From 46d5a8195b4d8176bd0a93e784c715276015a1b5 Mon Sep 17 00:00:00 2001 From: jiashuoz Date: Tue, 29 Sep 2026 14:17:21 +0800 Subject: [PATCH 7/8] docs(design): generic feature packs revision 7 after focused review brand_match gains declarable exemptions computed at score time from a display-name fact table, per-role matcher variants, match after masking, and standalone age_decay; self-send uses eq:false with before_first include_future; one explicit monotone baseline shared by the four history-relative email features; embedded SHA-addressed reference packs and fail-closed 503 ingest on config_error; hash_quantum 0 and a rescore legacy_v0 mode for bit-exact NextRescoreAt; canonical links serialisation with a byte-identity test; derived column; completed neutrality audit with a P-N0 cleanup slice; stricter neutrality tests; closed compat enum; shadow-mismatch gate before P-E2; emailshadow namespace bound at load; decisions updated. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_014cdM7WyRc3mD3vQNXMTDB8 --- docs/design/2026-09-27-abusekit-design.md | 2 +- .../2026-09-29-generic-feature-packs.md | 401 ++++++++++++++---- docs/plans/2026-09-27-v0-plan.md | 2 +- 3 files changed, 329 insertions(+), 76 deletions(-) diff --git a/docs/design/2026-09-27-abusekit-design.md b/docs/design/2026-09-27-abusekit-design.md index 2a3208b..1a21e2b 100644 --- a/docs/design/2026-09-27-abusekit-design.md +++ b/docs/design/2026-09-27-abusekit-design.md @@ -6,7 +6,7 @@ see account churn, tiers failed open when a scorer was unavailable, and request cover reads or replay. This revision fixes those and tightens every place an implementer would have had to guess. Changes from r1 are marked **[r2]**. -**Amendment (proposed 2026-09-29, revision 6):** [`2026-09-29-generic-feature-packs.md`](2026-09-29-generic-feature-packs.md) +**Amendment (proposed 2026-09-29, revision 7):** [`2026-09-29-generic-feature-packs.md`](2026-09-29-generic-feature-packs.md) makes the binary domain-neutral (email becomes a YAML reference pack), renames the built-in features once into namespaced packs enabled per tenant, adds product-declared vocabularies (types, field kinds, link kinds, subject kinds) with pseudonymising redaction, and adds bounded declarative custom features in YAML. It amends §4.2, §4.3, §4.5, §4.6, §4.8 diff --git a/docs/design/2026-09-29-generic-feature-packs.md b/docs/design/2026-09-29-generic-feature-packs.md index bb1850e..3ef883a 100644 --- a/docs/design/2026-09-29-generic-feature-packs.md +++ b/docs/design/2026-09-29-generic-feature-packs.md @@ -1,7 +1,8 @@ # Generic feature packs, product vocabularies, and declarative custom features -Status: proposed, revision 6, 2026-09-29. This revision applies an owner decision: abusekit is a -generic framework, and **the compiled binary contains no domain-specific knowledge** (no email and +Status: proposed, revision 7, 2026-09-29. Revision 7 applies the focused review of revision 6. +Revision 6 applied an owner decision: abusekit is a generic +framework, and **the compiled binary contains no domain-specific knowledge** (no email and no e2a concepts). Domain knowledge lives only in declared packs (YAML plus data files). Owner: Josh Zhang. This document amends [`2026-09-27-abusekit-design.md`](2026-09-27-abusekit-design.md), §4.2, @@ -248,11 +249,60 @@ exactly as today. `address_domain` on `resource.*`. Two enabled packs extending the same name fail with `extension_conflict`. Tenant extensions keep the `x_` prefix. - **A flat `links` map.** `links` is a map of **declared** link kinds, built-in or pack-declared, - so `links.email_hash` stays a top-level key. Revision 2's `links.custom` is dropped. An - undeclared key is rejected with `bad_links`. Which kinds count as neighbour evidence is set by - `evidence: true` on the declaration, never hard-coded. The email pack marks `email_hash` as - evidence, which reproduces today's default of `{email_hash, card_fingerprint_hash, - device_hash}` when the email pack is enabled. + so `links.email_hash` stays a top-level key. Revision 2's `links.custom` is dropped. + - **Neighbour evidence.** Which kinds count as evidence is set by `evidence: true` on the + declaration, never hard-coded. The email pack marks `email_hash` as evidence, which + reproduces today's default of `{email_hash, card_fingerprint_hash, device_hash}` when the + email pack is enabled. + - **Canonical serialisation (M1), pinned.** + - The legacy kinds are written first, in today's `event.Links` field order: `email_hash`, + `card_fingerprint_hash`, `ip24_hash`, `asn`, `ua_hash`, `device_hash`, each omitted when + empty. + - Any other declared kinds follow, sorted by name. + - An empty map serialises as `{}`. + + This keeps the stored `links` bytes and `body_hash` unchanged for every pre-P-E1 event. + - **Validation is unchanged for legacy kinds.** + - `email_hash` keeps its 64-lowercase-hex check. It is declared as a link kind with + `format: hex64`. + - A malformed legacy value is still a per-item `bad_links`. + - An **unknown link key** is still a whole-request `400 bad_request`. Today strict JSON decoding + produces that, and from P-E1 the tenant's declared-kind set does. + - **Byte-identity test (`TestWireByteIdentity`).** It replays the whole pre-P-E1 corpus (every + committed fixture plus the synthetic corpus, and every error-path contract test) through the + old and new ingest paths. For each event it asserts identical stored `links` and `data` bytes, + identical `body_hash`, and an identical outcome: accepted, `duplicate`, `conflict`, or the + exact error code. +- **Derived fields (M2).** Fields that ingest computes, such as `x___in_` and + `x_brands`, are stored in a separate `events.derived jsonb` column and never in `data`. They are + excluded from the `body_hash` canonical form and from the 8 KiB `data` limit. A set or matcher + change recomputes them by backfill (§5.6 warm-up). +- **`display_name` is opt-in per declared resource kind (L3).** `resource_kinds: + {agent: {role: other, display_name: true}}`. The implicit legacy declaration sets + `display_name: true` for every kind, including undeclared ones, because the Go `nameBrandMatch` + reads every resource name. + +**Reference packs are embedded and addressed by content SHA (H4, M5).** +- The binary embeds `packs/**` (`go:embed`) together with a **release manifest** + (`packs/MANIFEST`) that lists each reference pack's `name@x.y.z` and the SHA-256 of its + directory. +- The embedded packs are data, never Go. Loading a reference pack never depends on a mount, so a + missing file can't disable one. +- **A pack counts as a reference pack only if its SHA is in the embedded manifest.** Only those + packs may use `compat` options. +- Reserved pack names (`core`, `brand`, `email`, `card-testing`, `api-credential-abuse`, `custom`, + and every manifest name) **can't be shadowed from a tenant mount**. A same-name directory + in the mount fails with `pack_shadowing`. + +**Fail closed on configuration errors (H4).** +- If a tenant is in `config_error` (no valid profile: a cold start with a bad profile, or an + embedded pack SHA mismatch), `POST /v1/events` for that tenant returns **`503` with + `Retry-After` for the whole request**, and code `tenant_config_unavailable`. +- Ingest never drops fields, never rejects single items, and never stores events under a partial + vocabulary. The producer's durable outbox retries. +- **Test (`TestIngestFailsClosedOnConfigError`):** with the tenant's profile invalid, a valid + batch returns 503 and nothing is stored. Once the profile is fixed, the retried batch is stored + with bytes identical to a batch that never failed. **Core neutrality audit.** Every revision-5 core feature was checked: @@ -264,10 +314,27 @@ exactly as today. | `core.subject_age_h`, `upgrade_*`, `declines_*`, `first_funding_prepaid`, `resource_*`, `credential_*`, `verdict_max_24h`, `neighbors_truncated` | None | Unchanged | | Self-send, webmail and destination logic | Previously Go facts and counters | Now entirely in the email pack, as declared `before_first` fact specs and lifetime share counters over the roles `self`, `recipient` and `destination` | +**Completed audit of non-feature code (M3).** Every item below is fixed in slice P-N0 (the +cleanup that runs before the neutrality CI) or deleted in P-E2: + +| Location | Domain knowledge | Fix | Slice | +| --- | --- | --- | --- | +| `internal/worker/reason.go` `renderReason` | Prints email-named features by hand. | Reason is generated generically from the rule's inputs: the top contributing features by `|w·x|`, with their values. `reason_version: 3`. | P-N0 | +| `pkg/abusekit` `Links.EmailHash` (public Go SDK) | Email link kind in a typed struct. The flat map is a **breaking SDK change**. | Add the module path `pkg/abusekit/v2`, where `Links` is `map[string]string`. v1 is frozen and deprecated but still emits identical wire bytes. e2a's emitter (S6) moves to v2 at its own pace, and v1 is removed one release after that. The v1 identifier sits on the allowlist, keyed by `(file, identifier)`, until then. | P-N0 (v2), P-E2 (allowlist expiry date) | +| `eval/neighbors.go` `defaultNeighborKinds` | Hard-codes `email_hash`. | Uses the profile's declared evidence kinds. | P-N0 | +| `eval/gen` | Generates `content.sent` families in Go. | Becomes a data-driven generator. Families are YAML in each pack (`packs/

/gen/*.yaml`) or example profile; the Go generator stays neutral. | P-N0 | +| `internal/feature/store_neighbors.go` `allLinkKinds` | Lists `email_hash`. | Derived from the tenant's declared link kinds. | P-N0 | +| Brand matcher literals `buildIntegrationTokens` and `buildCommunityPhrases` | Word lists in Go. | Moved to `packs/brand/*.yaml`. | P-N0 | +| YAML config: `config/rules.yaml`, `local_weights.yaml`, `webmail.yaml`, `brands.yaml` | Email feature names and lists. | Moved to `examples/tenants/reference/`, `packs/email/` and `packs/brand/`. | P-N0 (moves; the golden stays exact) | +| #7 `--webmail` flag and `ABUSEKIT_WEBMAIL_CONFIG` | An email concept in the CLI and environment. | Deleted in P-E2. For one release the flag and variable are accepted, ignored and logged as deprecated. **Ops migration note:** the hosted compose (S8) removes them in the same release that bumps the email pack pin. | P-E2 | +| #7 `WebmailSet` dependencies in `serve`, `labels` and `worker` | Webmail threaded through the core. | Deleted; the webmail list is pack data. | P-E2 | +| #7 `"content.sent"` matches in `windows.go` (`isWindowedEventType`, sends helpers) | Type name in core code. | Core windowed types derive from `Reads`/`activity`. The sends helpers are deleted with the Go email code. | P-E2 | +| #7 `"agent"` match in `exemptSubjectBrands` (and the `feature.go` comment) | e2a resource-kind name. | Becomes the declared exemption `where: {field: kind, in: [agent]}` in the email pack (§5.5, H1). | P-E2 | + **Roles (field level).** - `title`: text matched by `brand.title_match`. - `display_name`: an actor-chosen name matched by `brand.name_match`. The core - `resource.*.name` field carries it. + `resource.*.name` field carries it for every resource kind that opts in (L3). - `self`: bool; a self-directed event. - `destination`: domain or hash; where an activity lands. - `recipient`: hash; a single counterparty. Pairs with `recipient_count`. @@ -288,18 +355,32 @@ fields explicitly. - **Exposed** as features `brand.name_match`, `brand.name_has_at` and `brand.title_match`, and as one generic DSL op, **`brand_match`** (§5.5), which declared packs may use. -**Neutrality CI.** P-N (§9) adds three checks, which then gate every slice: -1. `TestCoreIsDomainNeutral`. It walks the Go AST of every package outside `packs/`, `examples/` - and `testdata/`, and fails on any identifier or string literal on - `internal/neutrality/denylist.txt` (for example `email`, `webmail`, `recipient`, `subject_line`, - `smtp`, `mailbox`, `inbox`, `e2a`). Exceptions come only from `allowlist.txt`, one justified - entry per line; the privacy leak detector is the canonical one. +**Neutrality CI.** Switched on at the end of P-E2 (§9), after the P-N0 cleanup. It then gates +every slice. +1. `TestCoreIsDomainNeutral` (M4). + - **Scope:** it walks the Go AST of **every** Go file in the module, including Go under + `packs/` and `_test.go` files. Test files are scanned too, so domain knowledge can't hide in + test helpers. Data files (YAML, JSON, JSONL, text) are not scanned. + - **Tokenising:** each identifier and string literal is split on camelCase and snake_case + boundaries and on non-alphanumerics, then lowercased. Every token is compared against + `internal/neutrality/denylist.txt`: `email`, `webmail`, `recipient`, `subject`+`line`, + `smtp`, `mailbox`, `inbox`, `e2a`, `agent`. + - **`agent`** is denied except as the second token of `user agent` (`UserAgent`, `user_agent`, + `ua_hash` is unaffected). The rule is explicit in the denylist syntax: `agent !after user`. + - **Allowlist:** `allowlist.txt` is keyed by `(file, identifier)`, one justified entry per line, + with an optional expiry date. The leak scanner gets exactly three entries, all in + `internal/vocab/leakscan.go`: `emailRe`, `looksLikeEmail` and `maskEmails`. + `emailMaskExemptKey` is deleted. Which field is masked instead of rejected is now the pack's + `text` field declaration. 2. `TestNonEmailProfileEndToEnd`. The binary loads the card-testing reference pack (`examples/tenants/tallyport`) with **no email pack**. It ingests that pack's fixtures through HTTP and scores them. It asserts that no `email.*` feature is registered, that `content.sent` is an undeclared type for that tenant, and that verdicts match the pack's golden replay. 3. `TestGoRegistryLint`. Every Go `FeatureDef` may read only built-in core types or roles (`Reads`). Any type name outside the built-in list fails. +4. **Runtime enforcement of `Reads`.** A Go pack's `Extract` receives a **filtered event view** + that contains only events whose type is in its `Reads`, directly or by role. A Go pack + therefore can't observe a pack-declared type even through a generic loop over events. ### 5.1 Packs @@ -589,37 +670,116 @@ value = transform( v1 · d ) - `age_decay` defaults to `{full_until: 3d, zero_at: 30d, floor: 0.2}`. - `baseline: {peak: {size: 10m}}` overrides the baseline op. -**Generic primitives added in revision 6, so the email pack needs no Go.** Each one is -domain-neutral and usable by any pack: +**Generic primitives added in revisions 6 and 7, so the email pack needs no Go.** Each one is +domain-neutral and usable by any pack. | Primitive | Semantics | | --- | --- | -| **Baseline override** `baseline: {op, size?, where?}` | The baseline uses its own op, sub-window width and predicate, independent of `cur`. For example, a 1 h sum compared against a prior 10 min peak over *all* non-self rows. | -| **`distinct` with `on_missing`** `{field, on_missing: ignore \| count \| {sum: {field, default}}}` | Rows without `field` are ignored (the default), counted as one unit each, or contribute a summed field. The value is `count(DISTINCT field) + fallback`. | -| **Compat options** `compat: {version: 0, window_end: closed, include_future: true}` | Explicit, versioned reproductions of legacy quirks. `window_end: closed` makes an anchored window `[start, start + A]`. `include_future: true` stops excluding `at > now` for that feature. The loader rejects `compat` in tenant `custom.*` features, so only reference packs may freeze a legacy quirk, and each quirk is listed in the pack's changelog. `compat.version` is bumped if a quirk's semantics ever change. | -| **`share` over `lifetime`** | Numerator and denominator come from two counter specs, with the `(now, +∞)` subtraction (unless `include_future`). | -| **`brand_match`** `{role: title \| display_name, window, exclude_self, exempt_on_display_name_token: integration, exclude_brands_of: display_name, cap}` | Counts distinct brand ids matched in `role` fields, provided by the brand pack. It can exclude brands whose display-name mention sits next to an integration token, and brands already credited through `display_name`. | +| **Baseline override** `baseline: {op: count \| sum \| distinct \| peak, size?, where?, sum?: {field, default, cap_each}}` | The baseline uses its own op, sub-window width, predicate and summed field, independent of `cur`. **Only monotone ops are allowed** (count, sum, distinct, peak), so a truncated baseline can only under-estimate (§5.7). `lookback` and `exclude_recent` stay on `relative_to_history`. | +| **`distinct` with `on_missing`** `{field, on_missing: ignore \| count \| {sum: {field, default, cap_each}}}` | Rows without `field` are ignored (the default), counted as one unit each, or contribute a summed field. The value is `count(DISTINCT field) + fallback`. | +| **Compat options** (M5): `compat: [

/gen/*.yaml`) or example profile; the Go generator stays neutral. | P-N0 | | `internal/feature/store_neighbors.go` `allLinkKinds` | Lists `email_hash`. | Derived from the tenant's declared link kinds. | P-N0 | @@ -1800,7 +1800,7 @@ as described in §5.2. | P4c | Class R, budgets, flags, flood | `peak` (with `sum`) under the raw-unit saturation limit per transform (§5.7); `neighbours` exact by saturation; the partial/degraded direction table; `ratio` partial propagation and `ratio_den_partial_capable`; `relative_to_history` with its baseline budget; `partial` in verdicts and the API; pass order; `TestFeatureIndependence`; flood generator and property; cost benchmark. **Then enable bounded evaluation for e2a.** | P4b | Criteria 4 and 6; golden exact under the bounded evaluator | | P4d | Rescore control and warm-up | Proportional coalescing; timers only from non-shadow rules; per-tenant budget; warm-up | P4c, P3c | Storm and warm-up tests | | P5 | Pack gating | Registry; the Go `core` and `brand` adapters and the `declared` (YAML) adapter; enablement; `brand.title_match` (needs the `title` role); brand lists moved to data; `packtest`; starter weights | P3a, P4c | Golden exact; `feature_not_enabled`; every pack passes `packtest` | -| P-N0 | Neutral cleanup (M3) | `reason.go` made generic; `pkg/abusekit/v2` with a map `Links` (v1 frozen); `eval/neighbors.go` and `allLinkKinds` driven by declared evidence; data-driven `eval/gen`; brand word lists moved to data; config YAML moved to `examples/tenants/reference`, `packs/email` and `packs/brand` | P5 | Golden exact; SDK v1 and v2 produce identical wire bytes | +| P-N0 | Neutral cleanup (M3) | `reason.go` made generic; `pkg/abusekit` `Links` replaced in place by a map (a breaking change, noted in the release notes); `eval/neighbors.go` and `allLinkKinds` driven by declared evidence; data-driven `eval/gen`; brand word lists moved to data; config YAML moved to `examples/tenants/reference`, `packs/email` and `packs/brand` | P5 | Golden exact. The SDK's contract tests are updated to the map `Links`, and a wire test shows the new SDK emits bytes identical to the pre-change SDK for every legacy link kind. The changelog carries a breaking-change entry. | | P-E1 | YAML email pack, in shadow | Embedded `packs/email/` plus the manifest. Primitives: baseline override (monotone ops, `sum`/`cap_each`), `distinct.on_missing`, the closed `compat` enum with conformance tests, lifetime `share`, `brand_match` with declarable exemptions and the per-role matcher variants, standalone `age_decay`, cross-field constraints, `rescore: legacy_v0`. The `subject_display_names` fact table; pack extensions of built-in types; the flat `links` map with its canonical serialisation; the `derived` column; `store: raw+skeleton`. The `content.sent` schema, `email_hash`, `email_domain_class` and `address_domain` move into the pack. Loaded as `emailshadow.*`, with the production mismatch metric. | P-N0, P4b, P4c | The pack loads from the embedded manifest. **`TestWireByteIdentity`**: stored `links` and `data` bytes, `body_hash` and error codes identical over the whole pre-P-E1 corpus. Golden exact (Go still drives scores). Every shadow feature bit-identical to its Go twin on the golden, including `NextRescoreAt` and input hashes. The H1 counter-example fixture passes. | | P-E2 | Parity gate, switch, deletion | Preconditions:
• embedded packs;
• fail-closed ingest (`TestIngestFailsClosedOnConfigError`);
• the shadow-mismatch gate: N clean days, divergences labelled (M6).
Then:
• switch the binding `as: emailshadow` to `as: email`; the pack SHA is unchanged;
• delete the Go email features, `WebmailSet`, the `--webmail` flag, the `content.sent` matches, and the transitional built-in `content.sent` schema;
• turn on the neutrality CI and the runtime `Reads` filter.
**Ordered ops step:** the hosted profile pin (`email@1.0.0#sha256:…`) and the compose cleanup (the `--webmail` flag and its environment variable) ship **in the same release** as the deletion. | P-E1 and the N-day gate | Parity oracle is the post-P4a Go code (M7). Golden exact after deletion. `TestCoreIsDomainNeutral`, `TestNonEmailProfileEndToEnd`, `TestGoRegistryLint` and `Reads` enforcement green. | | P5b | Core DSL parity | Every expressible *core* S2 feature re-expressed in the DSL (a conformance check; the core Go features stay, being neutral) | P4c | Bit-exact against the Go feature on every fixture | @@ -1987,8 +1987,9 @@ Where the re-review's (R3) answer differs from the earlier recommendation, both `rescore_legacy_v0` (H5). 31. **The P-E2 clean-days gate (new).** How many consecutive days of zero unlabelled shadow mismatches in production before the Go email code is deleted? Proposed: 7. -32. **Go SDK v2 (new).** A breaking `pkg/abusekit/v2`, with `Links` as a map, and v1 frozen and - removed one release after S6 moves. Approve the plan and v1's end-of-life? +32. **Go SDK links change. Decided (owner):** replace `pkg/abusekit` in place, with `Links` as a + map. It's a breaking change on `main`, called out in the release notes and changelog. No + parallel versions. 33. **`recipient_count` bounds (new, M2).** The email pack declares `recipient_count` with `min: 1, integer: true, max: 1000000`. The max must be at least e2a's largest per-message recipient count. Confirm the value against e2a's send limits before P-E1. @@ -2255,7 +2256,7 @@ are versioned, whether they are public, pack pinning, tenant-private packs, and - **M2:** derived fields live in a separate `derived` column, outside `body_hash` and the 8 KiB limit. Adds `store: raw+skeleton`. `recipient_count` declares `min: 1, integer: true, max: 1000000` (Q33). -- **M3:** the audit of non-feature code is complete: `reason.go`, SDK v2 for `EmailHash`, +- **M3:** the audit of non-feature code is complete: `reason.go`, the SDK's `EmailHash`, `eval/neighbors.go`, `eval/gen`, #7's `--webmail` flag and env var (with an ops note), `WebmailSet`, the `content.sent` matches, `allLinkKinds`, `"agent"`, and the config YAML. The cleanup slice P-N0 is added. @@ -2290,4 +2291,8 @@ are versioned, whether they are public, pack pinning, tenant-private packs, and **Decisions** - Q26–Q30 now carry the review's answers. -- New: Q31 (clean days), Q32 (SDK v2) and Q33 (`recipient_count` max). +- New: Q31 (clean days), Q32 (SDK links change) and Q33 (`recipient_count` max). + +### Revision 7a + +- Owner decision on Q32: the Go SDK is replaced in place. `pkg/abusekit` `Links` becomes a map as a breaking change, noted in the release notes and changelog. There's no v2 module path and no retirement window.