Skip to content

feat(eval): exact scoring baseline for generic-feature migration - #9

Merged
jiashuoz merged 6 commits into
mainfrom
feat/generic-baseline-rename
Sep 30, 2026
Merged

jiashuoz merged 6 commits into
mainfrom
feat/generic-baseline-rename

Conversation

@jiashuoz

@jiashuoz jiashuoz commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

The generic-feature migration needs an exact baseline before feature names and storage contracts change. Adds abusekit eval --golden and a committed oracle covering 8,671 event/timer points across the 28 synthetic event fixtures and synthetic corpus. Features, scores, risks, hashes, versions, tiers, and rescore times are pinned without vendor or database calls; --golden-check fails on drift.

Scoring histories now resolve timestamp ties by (at, producer, id) in both Postgres and offline replay. Neighbor evidence includes only already-accepted events, including within ties. This is P0 of the approved generic-feature design; P1 (namespaced rename) remains next, and vendor adapters/hosted deployment remain blocked on P1.

Validation:

  • Full suite with ABUSEKIT_REQUIRE_DB=1 go test ./....
  • Focused race checks for replay and database ordering.
  • make lint, make gate, and built-binary golden comparison.
  • Built service on a disposable local database: signed ingest, duplicate handling, invalid-signature rejection, synchronous evaluation; clean server logs and database cleanup.
  • One-bit weight mutation fails the golden comparison; permutation, timer-drain, malformed-input, and label-identity regression checks.

Review findings addressed:

  • Independent review: source/reference overwrite through directory members or file aliases; fixed with resolved-path and inode checks, regression tests, and atomic output writes. Re-review passed.
  • Adversarial review: the same overwrite issue, nondeterministic neighbor selection above the total cap, and trailing JSON values silently omitted. Fixed with deterministic kind/hash traversal and strict single-value JSONL parsing; regression tests reproduced all three before the fixes. Re-review passed.

Exact comparison exposed existing CPU arithmetic differences. The PR keeps three explicit numeric profiles: ARM64, AMD64 with FMA, and AMD64 without FMA. Selection uses CPU capabilities independently of replay outputs. All 8,671 point identities, features, hashes, versions, timers, tiers, and flags agree; 499 risks/scores differ between ARM64 and non-FMA AMD64, and 94 between the two AMD64 paths. Disabling FMA on the same Linux CI runner reproduced the non-FMA oracle exactly. No scorer changes or tolerance were introduced.

Both reviewers independently verified all three reference files and passed the final implementation. CI checks native FMA, forced non-FMA, and ARM64 with explicit profile assertions. Local verification also exercised the CI Go 1.23.0 AMD64 toolchain. Future migration slices must preserve all three exact oracles. No deployment changes.

Final head 5b8f5d9: all four CI jobs pass (build/lint/unit including both AMD64 paths, DB-backed suite, ARM64 exact replay, synthetic evaluation gate).

@jiashuoz
jiashuoz merged commit 050ebad into main Sep 30, 2026
4 checks passed
@jiashuoz
jiashuoz deleted the feat/generic-baseline-rename branch September 30, 2026 04:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant