Skip to content

design: generic feature packs, neutral event vocabulary, declarative custom features - #8

Merged
jiashuoz merged 8 commits into
mainfrom
design/generic-feature-packs
Sep 29, 2026
Merged

jiashuoz merged 8 commits into
mainfrom
design/generic-feature-packs

Conversation

@jiashuoz

@jiashuoz jiashuoz commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Summary

abusekit's current feature set assumes an email platform. This design, docs/design/2026-09-29-generic-feature-packs.md (now at revision 7), makes abusekit a generic abuse-detection system for any SaaS product. It assumes #5 (S4 eval harness) and #7 (S2b features) have merged.

  • One-time rename. Built-in features are renamed once into namespaced core, email and brand packs, and each tenant enables the packs it needs. The local scorer now sums features in a fixed registry order, so every score, tier and feature value stays bit-for-bit identical after the rename. Every consumer of feature names has a test.

  • Per-tenant profiles live in a private config mount. A product declares its own event types, field kinds and roles, extension fields on built-in types, link kinds, and subject kinds beyond accounts (cards, API keys, customers).

  • Pseudonymising redaction:

    • every hash is re-HMACed at ingest, with per-tenant keys derived by HKDF;
    • domains are reduced to eTLD+1 by default and stored in cleartext only if allowlisted, otherwise HMACed;
    • undeclared values are dropped;
    • declared numbers must declare a max;
    • text is masked, or stored only as a skeleton;
    • stored bytes are scanned on the way out.
  • Evaluation combines:

    • facts maintained at ingest (onboarding values);
    • counters (lifetime totals);
    • exact aggregate queries, one per feature;
    • row features with their own per-feature budgets.

    Hitting a bound sets a one-sided partial flag on the verdict and is never used as a weight. Floods cannot lower risk: this holds against an unbounded reference and is tested with a specified flood generator.

  • A closed DSL: count, distinct, share, peak, time_between (with absence indicators), sequence, exact group_by, ratio, neighbours, and relative_to_history (with exact equations).

  • Walkthroughs of five fictional scenarios, each stating plainly what remains Go-only.

  • Slices P0–P7. P1 (the rename) must land before any S5 or S8 PR. Bounded evaluation does not reach e2a until P4a–P4c have landed.

This PR is docs only: the design doc, a pointer line in the main design, and a row in the plan.

Decisions needed from the owner

Where the re-review's (R3) answer differs from the earlier recommendation, both are shown.

  1. Rename vs bridge. Rename once in P1. R3 agrees and asks for a bit-exact golden replay. Now: summing in registry order makes it bit-exact, and the 1e-12 tolerance is deleted. Approve?

  2. Ordering. P0 and P1 land before any S5 or S8 PR, enforced by the plan and a test. Approve?

  3. Undeclared data. Drop the values and keep only the field names. R3: also constrain the names. Now: names must match ^[a-z0-9_]{1,64}$, at most 32 per event. Approve?

  4. Re-HMAC of built-in fields. Re-HMAC recipient_hash and every link, with a dev/staging migration job. Approve?

  5. Legacy window quirks. Freeze them in @1, harmonise them in @2. Unchanged. Confirm?

  6. Roles. Revision 6 adds the field roles display_name, destination and recipient to credential, other, activity, title and self. Confirm?

  7. Bounding.

    • Revision 2: byte caps, a shared step budget, and a weighted truncation feature.
    • R3: ingest facts, counters, per-feature budgets, and truncation as a flag.
    • Now: R3 plus exact per-feature aggregates.

    Revision 4 adds raw-unit peak sizing and saturation-exact neighbours. Anything that can undercount is degraded. Confirm the 50,000-row budgets (baseline, peak, anchored, neighbour) and max_groups of 1,000?

  8. Bootstrap. Uniform priors, shadow-only, no fitting, held-out fixtures, and __absent indicators. Confirm?

  9. CEL. Allow it later as a predicate leaf only, or require a new design pass? Unchanged.

  10. Brand list. Can a tenant narrow the list as well as extend it? Unchanged.

  11. Custom namespace. custom.* per tenant, or <tenant>.*? Unchanged.

  12. Config history. Keep it in the config tree and check it in CI; the DB is only a guard. Approve?

  13. S6. Emits content.sent, unchanged on the wire. From P-E1, the email reference pack declares it, not the binary. Settled.

  14. Domains.

    • Revision 2: public-suffix check, stored in cleartext.
    • R3: default to eTLD+1, cleartext only if allowlisted, otherwise HMAC, including on built-in fields.
    • Now: R3, except built-in recipient_domain keeps the full domain (reduce: none) and is HMACed. Reducing to eTLD+1 would merge distinct fixture domains and change email.first_day_distinct_domains, whereas HMAC is injective.

    Also: accept that text scorers see a token for unknown link hosts?

  15. ratio product form. Add a capped product form, or wait for fitted weights?

  16. Subject kinds. subject_kind and also (at most 3), via_parent index rows (at most 8 per event), and pseudonymised non-account ids. Approve?

  17. Stage gate. Advise-mode local rules only, as its own slice (P1s). Approve?

  18. Profiles and keys. Profiles in a private mount; a provider-agnostic Keys interface with HKDF per tenant and purpose. Approve?

  19. Rescore control. Proportional coalescing for DSL features only, and timers only from non-shadow rules. Confirm a budget of 20 × active subjects per hour?

  20. group_by admission.

    • Revision 2: first-come.
    • R3: space-saving.
    • Now: an exact GROUP BY in the aggregate engine, with space-saving only as the in-memory adapter's memory bound, flagged partial.

    Approve?

  21. Declared numbers.

    • Revision 2: Luhn check on integers.
    • R3: author-trusted, max required, no Luhn.
    • Now: R3.

    Approve?

  22. partial flag. Revised in revision 4: a bare partial means the value can only have risen, and any source that may undercount also sets degraded. Should callers still treat a bare partial like degraded?

  23. start precedence (new). account_created_at, then the first accepted subject.created, then server first_received_at, replacing LEAST(at). Approve?

  24. Key-independent body_hash (new). Computed over the redacted body before pseudonymisation, so rotation and migration never break dedupe. Approve?

  25. Dirty marks beyond the fan-in cap (new). The first 50 are marked in the ingest transaction; the rest by a rate-budgeted background job. Confirm?

  26. Reference pack location. Adopting the review's answer: packs are embedded in the binary and addressed by content SHA through a release manifest. Tenants may add packs only in non-reserved namespaces and can never shadow a reference pack.

  27. Public reference packs. Yes. Starter weights and floors come from synthetic fixtures only.

  28. Pinning. Revision 6 allowed floating pins everywhere. Now advise-mode profiles must pin name@x.y.z#sha256:…; floating pins are allowed only in shadow-only profiles and in dev.

  29. Tenant packs. No compat, no shadowing, must follow the naming grammar, and may extend built-in types only with x_-prefixed names.

  30. compat. Revision 6 had free-form options. Now it's a closed enum in the engine, with a conformance test per option: window_end_closed, include_future, before_first_include_future (new) and rescore_legacy_v0.

  31. P-E2 clean-days gate (new). How many days of zero unlabelled shadow mismatches are required before the Go email code is deleted? Proposed: 7.

  32. Go SDK links change. DECIDED (owner). Replace pkg/abusekit in place: Links becomes a map, as a breaking change on main noted in the release notes and changelog. No v2 module path and no parallel versions.

  33. recipient_count bounds (new). min 1, integer, max 1000000. Please confirm this is at least e2a's largest per-message recipient count.

Rework (revision 2, after adversarial review)

Blockers

  • B1:
    • Dropped the canonical-key bridge. The rename now happens once, in P1.
    • §5.2 lists every consumer of feature names, each with a test: core.stageSkip (and how Rule.Stage keys are handled), the quantization switch, inputHash, the eval cassette key and header marker, the run.json SHAs and key-space marker, renderReason with reason_version, corpus_examples.features with a feature_key_space column and migration, local.Version(), and score --jsonl, which now returns feature_renamed.
    • The golden replay was re-baselined under semantic identity, with a derived bound of |Δrisk| ≤ 1e-12.
  • B2:
    • Loading is by time, over max(window, lookback + exclude_recent). Onboarding types (subject.*, payment.*, subscription.*) always load in full, and byte caps apply.
    • A deterministic step budget, calibrated by an end-to-end benchmark, replaces the wall-clock deadline.
    • Hitting the byte cap or the step budget sets core.history_truncated, which carries positive weight. Rules are never marked unscored for this. packtest enforces the truncation invariant.
    • §5.7 gives the argument that flooding cheap events can't evade, and criterion 6 is the matching property test.
  • B3:
    • Every hash, built-in or declared, is re-HMACed at ingest (length-prefixed, with join domains).
    • Undeclared values of every kind are dropped.
    • Domains must end in a public suffix; IP literals and all-numeric labels are rejected; eTLD+1 reduction is optional.
    • The leak scan covers Luhn-checked card numbers, IPv4/IPv6 and phone shapes.
    • Custom text is stored as a skeleton only.
    • Terminology is now "pseudonymisation" throughout.
  • B4:
    • New primitives: group_by (max or count_gte reduce, with capped groups); sequence (with an on join); ratio (a depth-1 DAG); declared link kinds plus a neighbours op; before_first; anchor: last; and subject kinds with also and parent.
    • All five scenarios are re-walked in §7 with exact YAML and sample events, and each scenario's Go-only remainder is stated.

Should-fix

  • delivery.sent dropped. Replaced by declared types with title and self field roles, an activity type role that core velocity and burst features include, and x_ extension fields on built-in types.
  • Warm-up and key rotation. New features warm up in shadow. Key rotation reads both keys during the rotation window, and the key id is recorded in vocab_version.
  • relative_to_history. The doc gives the exact equations and lists which S2b features the DSL can't express.
  • Cost model. Cost units are calibrated from a benchmark, and features per event type are capped.
  • Scheduling.
    • Rescore-storm control: proportional coalescing, rescores only for non-shadow rules, and a per-tenant budget.
    • Per-tenant fair queue with a concurrency cap.
    • Stage gates consider only advise-mode rules.
  • Config and determinism.
    • Config history lives in the config tree and is checked in CI.
    • Profiles live in a private mount.
    • Tie-breaks use (at, producer, id).
  • Fixtures and plumbing.
    • Held-out fixtures for the bootstrap criterion.
    • PriorSign on FeatureDef.
    • Rule names are resolved per tenant.
  • Re-sliced into P0–P7. P0 is the golden replay, built by extending eval. The dependency errors in revision 1's G4/G5/G6 are fixed. S3b erasure is now required to be vocabulary-aware.
  • Nits. Only the credential and other roles ship. The key interface is provider-agnostic. Scorer versions are reported per tenant. peak clips sub-windows to its outer window. The uniform-prior normalisation is written out.

Revision 3 (after re-review of 3c1c32c)

B2 (open in revision 2):

  • (a) Onboarding values come from subject_facts, maintained at ingest with monotone updates and an indexed recount when the first success moves. No onboarding scan is byte-capped. A flood of 10,000 tiny blocked payments provably moves no onboarding feature.
  • (b) Lifetime totals come from subject_counters, so no feature can fall when history is truncated. TruncationDir is deleted.
  • (c) Hitting a bound sets a partial flag on the signal and the subject. It is not a weighted feature and does not set degraded. The flag only ever errs toward higher risk; uniform rules just record it. The truncation invariant and its packtest are deleted.
  • (d) Evaluation classes F, N, A, R, G and D run in a fixed pass order. Anchored work is reserved, and budgets are per feature or per pack. TestFeatureIndependence checks that changing one feature never moves another.
  • (e) Flood property: risk(flooded, bounded) ≥ risk(flooded, unbounded reference). The generator covers read and unread types, onboarding types, placements (before, interleaved, after, edges, future-dated), minimum-size events, and 1×/10×/100× saturation. Each evaluation class carries an exactness argument, including why peak is exact under its saturation limit.
  • (f) Bounded evaluation reaches e2a only after P4a–P4c. The slice plan says so.

B1:

  • feature_renamed in LoadSnapshotCorpus, plus the corpus-v2 schema.
  • The fake scorer's deterministicProbs is re-baselined in P1.
  • The local scorer sums in FeatureDef.Order (the pre-rename flat order), so the golden replay is bit-exact and the 1e-12 machinery is deleted.
  • Bound is defined for the velocity features as a normalisation ceiling, never a cap.
  • The 1,000 cap on totals is dropped, since totals now come from counters.

B3:

  • Declared numbers are author-trusted: max is required and there is no Luhn check.
  • Domains default to eTLD+1, stored in cleartext only if allowlisted and HMACed otherwise, including first_link_host and recipient_domain. The recipient_domain exception to eTLD+1 is explained in decision 14.
  • Undeclared names must match ^[a-z0-9_]{1,64}$, at most 32 per event.
  • Non-account and also ids are pseudonymised, and account ids are leak-scanned.
  • Keys derives a distinct key per tenant and purpose with HKDF; a cross-tenant inequality test covers it.
  • An egress scan checks the serialised row, with Unicode digits folded.
  • subject_line masking is clarified.
  • ASN is exempt from hashing and gets a tighter grammar.
  • A dev/staging re-HMAC migration job is added; it recomputes body_hash.
  • Enum values and set files are leak-scanned at profile load.

B4:

  • time_between and sequence horizons feed each feature's store range.
  • if_absent plus a derived __absent indicator with a mandatory absent_sign, so abandonment is never the benign extreme.
  • Durations use log1p.
  • ratio uses pre-transform values.
  • group_by is exact (space-saving is only a fallback).
  • Custom features get hash_quantum, defaulting to 60 minutes for until_now.
  • peak accepts sum, and share sums both numerator and denominator.
  • neighbours is evaluated as of now, matching eval/neighbors.go.
  • Dirty-mark propagation covers declared link kinds.
  • Declared evidence no longer feeds core.linked_*.
  • event_subjects is reconciled: at most 8 rows per event, with via_parent.

Slices:

  • The stage gate is split out as P1s.
  • P3 is split into P3a (vocabulary and scans), P3b (Keys, re-HMAC, link migration), P3c (config history) and P3d (dual-key rotation and the cloud Keys adapter).
  • P4 is split into P4a (facts and counters), P4b (aggregates), P4c (row features, flags and flood; this is where bounded evaluation is enabled for e2a) and P4d (rescore control and warm-up). P5b is added.
  • Dependencies are fixed: P5 needs P3a (for the title role) and P4c; P6a needs P3b.
  • Enforcement: S5 and S8 are marked blocked on P1, and TestNoVendorAdapterBeforeRename checks it.

Revision 4 (after verification of 427d462)

Blocking fixes (P4a, P4c):

  1. peak saturation is now sized in raw units per transform: C, ⌈e^C − 1⌉ or ⌈e^(C/s) − 1⌉. For history-relative features it becomes ⌈max(B,1)·min(rc, T⁻¹(C)/d)⌉, computed after the baseline. The row limit is x_sat·⌈W/S⌉, and if the budget binds first the feature is flagged partial + degraded. Both counter-examples are worked through.
  2. neighbours applies where before any limit and counts up to ⌈T⁻¹(cap)⌉ + 1 subjects, so the value is exact by saturation. A budget hit sets degraded. Legacy core.linked_* caps now also set degraded. The backwards direction claim in §5.7c is corrected with a per-source direction table.
  3. start is defined by precedence: account_created_at, then the first accepted subject.created, then server first_received_at. Criterion 6b is restated. Frozen class-N facts carry anchored_start and are recomputed when start changes or a backfill event lands in the anchor.
  4. Recount locking. The fact-row lock is taken first. The lost-decline race is the counter-example. The recount costs O(days + declines on the boundary day), using daily decline counters plus a partial index, and fires on ∞ → t and on earlier moves.
  5. Webmail counters subtract future-dated events over (now, +∞), which covers backfill keys that are exempt from the skew check.
  6. also / via_parent. Onboarding facts are kept for the primary subject only. Counters are kept for every index row of a matching kind. event_subjects gains type and at, with a matching index.
  7. ratio. Partial status propagates. A partial-capable den is rejected at load (ratio_den_partial_capable).

Text fixes:

  • Space-saving direction: only max with a positive sign stays upward. Everything else is degraded.
  • Facts are a deterministic function of the event set, with a permutation and interleaving test.
  • The leak scan covers strings only (declared numbers are exempt). Phone shapes need + or grouping, so 10-digit ASNs pass.
  • in_set on domain fields is evaluated at ingest.
  • Hash literals and set files are HMACed at load, under both keys during rotation.
  • body_hash is key-independent.
  • Class G has per-feature budgets. distinct_recipients_1h is consistently A+R.
  • The §1 flood criterion is restated as relative to the reference, with dilution noted.
  • Erasure, re-signup, counter expiry and backfill semantics are specified.
  • The load plan gains rows for before_first, lifetime and neighbours, plus a generic before_first fact spec.
  • age_decay gets a hash_quantum default.
  • Dirty marks beyond the fan-in cap are covered.
  • The address_domain exception is removed.
  • The sequence absent_sign is -.

Slices:

  • P4a gets its own anchored-range query.
  • S3b's done-when now covers clearing facts and counters.
  • TestNoProductionBeforeRename guards S8 as well as S5.

Decisions: Q7 and Q22 are revised. Q23–Q25 are new.

Revision 5 (after the final check of a234623)

  • H-A: Recount-backing counters (declines, before_first) now count only the primary subject, so also and via_parent rows never feed a recount. TestRecountParentChild covers this.
  • H-B: "First accepted" is the smallest (received_at, producer, id). Determinism now holds with each event's received_at fixed, and the a2 tests take received_at from the fixtures.
  • Fix 1: Both peak worked examples are exact: their streams complete at about 2,200 and 48,186 rows. Whether the limit binds is decided at run time, and x_sat is corrected to 50·min(300, T⁻¹(C)/d).
  • Fix 4: Ingest runs at READ COMMITTED with bounded retry. The recount uses hourly decline counters plus one boundary-hour query.
  • Fix 7: Class N features can't be a ratio denominator. A partial-capable numerator combined with a negative sign is flagged partial + degraded.
  • Also: start is clamped so it is never later than first_received_at.
  • Decisions: no new items. Q23 now includes the start clamp.

Revision 6 (owner decision: domain-neutral binary)

  • No domain knowledge in the binary. It carries only the engine, the built-in neutral vocabulary (subject lifecycle, payment, subscription, resource, verdict, label, and neutral link kinds), and the neutral Go packs core and brand. §5.0 defines what "built-in" now means.
  • Email is a YAML reference pack. packs/email/ (pack.yaml + webmail.txt) is loaded like custom features. content.sent, email_hash, email_domain_class and address_domain move into its declared vocabulary. The wire stays byte-compatible, via pack extensions of built-in types and a flat declared links map.
  • New generic primitives, so no Go special cases remain: baseline override, distinct.on_missing, versioned compat options, lifetime share, a brand_match op, and cross-field constraints. A parity table covers all nine S2b: send-volume, webmail, recipient and subject-brand features #7 email features.
  • Core audit. burst_ratio is generalised to activity types. email_domain_class_disposable moves to the email pack. The evidence link set is declared.
  • Brand. The matcher stays Go, as neutral string processing over title and display_name. The brand list, integration tokens and community phrases are now data.
  • Neutrality CI: TestCoreIsDomainNeutral (an AST denylist), TestNonEmailProfileEndToEnd (card testing with no email pack loaded), and TestGoRegistryLint.
  • Slices. S2b: send-volume, webmail, recipient and subject-brand features #7 merges first as the transitional implementation. P-E1 loads the YAML email pack in shadow. P-E2 proves bit-for-bit parity on the golden replay, then deletes the Go email code and switches the CI on. P6b ships the non-email reference packs card-testing and api-credential-abuse. e2a stays bit-identical throughout.
  • Decisions. Q6 and Q13 are revised. Q26–Q30 are new: pack location and versioning, public reference packs, pinning, tenant-private packs, and compat options.

Revision 7 (after focused review of a80527d)

  • H1: brand_match gets declarable exemptions (exempt: {type, where, require_token, match_variant, live_unless}), computed at score time from a display-name fact table. The engine pins the matcher variant per role: for titles the community gate is on and the integration gate off; for display names it's the reverse. Matching runs after masking. Standalone age_decay is defined as T(min(raw, T⁻¹(cap))·d). The "Stripe API Key" counter-example is a committed parity fixture and gives 1·d.

  • H2: "external" in self-send now means eq: false, so a missing field doesn't count as external. before_first_include_future is added to compat.

  • H3: the four history-relative email features share one explicit baseline, B10. The baseline override gains sum/cap_each and accepts only monotone ops.

  • H4: reference packs are embedded in the binary and addressed by SHA. A tenant in config_error gets a retryable whole-request 503; nothing is dropped. There's a new test for this.

  • H5: all nine features set hash_quantum: 0, and a rescore: legacy_v0 mode reproduces Go's rescore candidate set. Together these make NextRescoreAt and the input hashes bit-exact.

  • M1: the links map has a canonical serialisation. TestWireByteIdentity checks stored bytes, body_hash and error codes. An unknown link key still returns a whole-request 400.

  • M2: derived fields move to a separate column, outside body_hash and the 8 KiB limit. Adds store: raw+skeleton, and recipient_count gets bounds.

  • M3: the audit now covers everything outside the features themselves:

    A new cleanup slice, P-N0, handles it.

  • M4: the neutrality test now:

    • covers all Go files, including packs/ and _test.go;
    • splits identifiers into tokens before matching;
    • keys the allowlist by (file, identifier), with an explicit rule for "user agent";
    • enforces Reads at runtime;
    • scopes the leak-scanner allowlist to three identifiers.
  • M5: only packs listed in the manifest count as reference packs. Reserved names can't be shadowed, and compat is a closed enum.

  • M6: a production shadow-mismatch metric, a clean-days gate before P-E2, and a list of production-only divergences.

  • M7: the P-E2 parity oracle is the Go code after P4a.

  • M8: the shadow namespace is emailshadow, bound at load time, so the pack SHA stays byte-identical all the way to ship.

  • L1–L4: the exact age_decay float expression; an unreachable cap; if_empty: 0; an IDNA note; display_name is opt-in per resource kind.

  • Slices: P-N0 comes first. P-E1's done-when adds byte identity. P-E2 requires embedded packs, fail-closed ingest, the ops pin and compose change in the same release, and the clean-days gate.

  • Decisions: Q26–Q30 now use the review's answers. Q31–Q33 are new. Revision 7a: Q32 is decided, so the SDK is replaced in place.

🤖 Generated with Claude Code

https://claude.ai/code/session_014cdM7WyRc3mD3vQNXMTDB8

jiashuoz and others added 8 commits September 29, 2026 11:16
…ative custom features

Splits the built-in features into namespaced core/email/brand packs enabled
per tenant, adds a neutral delivery.sent event with a read-side view over
content.sent, product-declared resource kinds, channels and custom types with
kind-driven redaction, and a closed declarative feature DSL (count, distinct,
share, peak, time_between, history-relative modifier) with mandatory caps and
a static cost model. A frozen canonical-key alias table keeps input hashes,
local-scorer summation order and version hashes identical, proven by a golden
replay of every fixture. Adds a pointer in the main design and a plan row.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014cdM7WyRc3mD3vQNXMTDB8
…rev 2)

Drops the canonical-key bridge for a one-time rename with a per-consumer test
list and a semantic-identity golden (derived ulp bound). Replaces the event-count
bound and wall-clock deadline with time-bounded loading, full onboarding loads,
byte caps, a deterministic step budget and truncation as a positive signal.
Closes redaction channels: re-HMAC of every hash (length-prefixed, join domains),
undeclared values dropped, PSL-checked domains, card/IP/phone scanning,
skeleton-only custom text. Adds group_by, sequence, ratio, neighbours with
declared link kinds, before_first, and subject kinds, and walks five fictional
scenarios. Drops delivery.sent for declared types with field roles. Re-slices
into P0-P7.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014cdM7WyRc3mD3vQNXMTDB8
Onboarding from ingest-maintained subject_facts and lifetime totals from
subject_counters; evaluation classes F/N/A/R/G/D with a fixed pass order and
per-feature/per-pack budgets; exact per-feature aggregates; truncation becomes a
one-sided `partial` flag, not a weighted feature; flood property restated against
an unbounded reference with a specified generator. Rename is bit-exact via
registry-order summation; fake scorer and corpus loader covered. Redaction:
author-trusted bounded numbers, domain eTLD+1 with allowlist-or-HMAC, name
grammar for undeclared fields, pseudonymised non-account ids, HKDF per-tenant
keys, egress scan. DSL: absence indicators, pre-transform ratio, exact group_by,
hash_quantum, as-of neighbours. Slices re-split (P1s, P3a-d, P4a-d, P5b).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014cdM7WyRc3mD3vQNXMTDB8
Peak saturation sized in raw units per transform (and after the baseline), with
partial+degraded when the row budget binds first; neighbours exact by saturation
with where-before-limit; partial/degraded direction table (fixes the backwards
fan-in claim); start defined by precedence (account_created_at, first accepted
subject.created, server first_received_at) with anchored-fact invalidation;
lock-first fact updates and bounded decline recount on inf->t and earlier moves;
webmail counters subtract (now, +inf); facts/counters subject assignment for
also/via_parent and type/at on event_subjects; ratio partial propagation. Plus
the listed text fixes, slice fixes and a section 13 revision-4 addendum.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014cdM7WyRc3mD3vQNXMTDB8
Recount-backing counters (decline, before_first) are primary-subject-only, with
a parent/child recount test; "first accepted" is the smallest (received_at,
producer, id) and determinism holds with received_at fixed; peak worked examples
corrected (streams complete; binding is decided at run time; x_sat fixed);
READ COMMITTED ingest with bounded retry and hourly decline counters plus a
boundary-hour recount; class N features rejected as ratio den and negative-sign
ratios over partial-capable nums flagged degraded; start clamped to
first_received_at.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014cdM7WyRc3mD3vQNXMTDB8
Owner decision: the compiled binary carries no domain knowledge. Email becomes a
YAML reference pack (packs/email) loaded like custom features; content.sent,
email_hash, email_domain_class and address_domain move into its declared
vocabulary with byte-compatible wire handling (pack extensions of built-in
types, flat declared links map). New generic DSL primitives (baseline override,
distinct.on_missing, versioned compat options, lifetime share, brand_match,
cross-field constraints) give a bit-for-bit parity table for every PR #7 email
feature; P-E1 loads the pack in shadow, P-E2 proves parity and deletes the Go
email code. Core audit, brand pack lists as data, neutrality CI, non-email
reference packs, new decisions on pack location, visibility and pinning.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014cdM7WyRc3mD3vQNXMTDB8
brand_match gains declarable exemptions computed at score time from a
display-name fact table, per-role matcher variants, match after masking, and
standalone age_decay; self-send uses eq:false with before_first include_future;
one explicit monotone baseline shared by the four history-relative email
features; embedded SHA-addressed reference packs and fail-closed 503 ingest on
config_error; hash_quantum 0 and a rescore legacy_v0 mode for bit-exact
NextRescoreAt; canonical links serialisation with a byte-identity test; derived
column; completed neutrality audit with a P-N0 cleanup slice; stricter
neutrality tests; closed compat enum; shadow-mismatch gate before P-E2;
emailshadow namespace bound at load; decisions updated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014cdM7WyRc3mD3vQNXMTDB8
The owner decided to replace pkg/abusekit Links with a map in place. It is
pre-GA with no external consumers, so this ships as a breaking change noted in
the release notes, with no v2 module path and no retirement window. Updates the
audit row, P-N0 done-when, decision list and a revision-7a note.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014cdM7WyRc3mD3vQNXMTDB8
@jiashuoz
jiashuoz merged commit 9c2b055 into main Sep 29, 2026
2 checks passed
@jiashuoz
jiashuoz deleted the design/generic-feature-packs branch September 29, 2026 09:35
jiashuoz added a commit that referenced this pull request Sep 29, 2026
Merges origin/main (PR #5's S4 eval harness + PR #8's design doc) into
this branch. The merge itself was clean except for a doc-only conflict
in eval/fixtures/README.md (both sides appended a new section); resolved
by keeping both sides' content, generic wording only.

The merge combined cleanly at the text level but not at the type level:
eval/replay.go's feature.Extract call predates S2b's WebmailSet
parameter, so the merged tree didn't build. Fixed by threading a
feature.WebmailSet parameter through LoadReplayDataset (added right
after brands, mirroring internal/feature.Extract's own parameter order)
and every call site, test helper, and CLI flag that constructs one:

- eval.LoadReplayDataset(in, brands, webmail, benignLabel) — webmail
  flows straight into every feature.Extract call, no package-level
  global, no panic on a zero value (feature.WebmailSet{} just matches no
  domain, same as passing no webmail config at all).
- `abusekit eval` gains --brands-extra and --webmail, both named,
  env-var'd, and defaulted exactly like `serve`'s own flags (main.go's
  parseServeFlags) — the harness now loads brands/brands-extra/webmail
  the identical way a real deployment does, not a silently different
  subset.
- New test: TestLoadReplayDataset_WebmailRecipientShareIsNonZero proves
  a webmail-heavy replay subject scores a non-zero
  webmail_recipient_share through the harness, and that the SAME subject
  scored against an empty WebmailSet reads back to 0 — proving the
  parameter is load-bearing, not merely accepted (verified by temporarily
  reverting the wiring and confirming the test catches it).

F9 TODO: added two new eval/gen families (abusive_webmail_blast,
abusive_subject_lure) — every other family's recipient_domain is a
synthetic .example.test name and no family ever sets subject_line at
all, so webmail_recipient_share/webmail_sends_1h/subject_brand_match
otherwise read 0 across the ENTIRE synthetic corpus regardless of the
harness wiring above. webmail_blast sends to real consumer webmail
domains (config/webmail.yaml's own list, a public fact); subject_lure
uses a fictional brand in its subject line (eval/fixtures/
test_brands.yaml, never a real one — this repo's hygiene rule for
fabricated lure prose, a stricter bar than a bare resource name).
Regenerated eval/fixtures/synthetic/{events,labels}.jsonl deterministically
from the documented seed (20260927) — 20 families, 297 subjects (was 18
families, 286). Makefile's gate target now passes --brands-extra
eval/fixtures/test_brands.yaml so the fictional lure brand is recognized
when scoring this corpus.

The new features and families change scores, so eval/floors.yaml was
re-derived against a fresh run, same margin policy the file documents,
weights untouched:

  precision 0.8036->0.8169, recall 0.8333->0.8788, AUROC 0.9735->0.9819,
  high-tier recall 0.7222->0.7879 (all IMPROVED or held family-steady:
  burst 0.8333, churn_incarnation_ge3 1.0, dormant_then_blast 1.0
  unchanged) -- min_precision/min_recall/min_auroc/min_high_tier_recall/
  family_min_high_tier_recall floors are UNCHANGED, now with MORE margin,
  not less.

  ONLY max_ece broke: baseline ECE moved 0.1108->0.1265 (an expected
  calibration cost of seven new hand-set, unfitted weight dimensions,
  not a regression), already past the old 0.115 floor before any weight
  was touched. Re-derived 0.115->0.132 (baseline+0.0055, same tight-margin
  policy T5 documented), confirmed to still catch subject_age_h
  (zeroed ECE 0.1509) and upgrade_delay_min (zeroed ECE 0.1378) --
  TestGate_NegativeWeightRegressionCaughtByTightECEFloor passes
  unmodified.

Hygiene check clean.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014cdM7WyRc3mD3vQNXMTDB8
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant