feat(eval): exact scoring baseline for generic-feature migration - #9
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The generic-feature migration needs an exact baseline before feature names and storage contracts change. Adds
abusekit eval --goldenand a committed oracle covering 8,671 event/timer points across the 28 synthetic event fixtures and synthetic corpus. Features, scores, risks, hashes, versions, tiers, and rescore times are pinned without vendor or database calls;--golden-checkfails on drift.Scoring histories now resolve timestamp ties by
(at, producer, id)in both Postgres and offline replay. Neighbor evidence includes only already-accepted events, including within ties. This is P0 of the approved generic-feature design; P1 (namespaced rename) remains next, and vendor adapters/hosted deployment remain blocked on P1.Validation:
ABUSEKIT_REQUIRE_DB=1 go test ./....make lint,make gate, and built-binary golden comparison.Review findings addressed:
Exact comparison exposed existing CPU arithmetic differences. The PR keeps three explicit numeric profiles: ARM64, AMD64 with FMA, and AMD64 without FMA. Selection uses CPU capabilities independently of replay outputs. All 8,671 point identities, features, hashes, versions, timers, tiers, and flags agree; 499 risks/scores differ between ARM64 and non-FMA AMD64, and 94 between the two AMD64 paths. Disabling FMA on the same Linux CI runner reproduced the non-FMA oracle exactly. No scorer changes or tolerance were introduced.
Both reviewers independently verified all three reference files and passed the final implementation. CI checks native FMA, forced non-FMA, and ARM64 with explicit profile assertions. Local verification also exercised the CI Go 1.23.0 AMD64 toolchain. Future migration slices must preserve all three exact oracles. No deployment changes.
Final head
5b8f5d9: all four CI jobs pass (build/lint/unit including both AMD64 paths, DB-backed suite, ARM64 exact replay, synthetic evaluation gate).