Skip to content

fix(cache): exact-match index replaces the global similarity walk - #327

Merged
TAIPANBOX merged 3 commits into
mainfrom
fix/cache-lookup-without-a-global-walk
Sep 24, 2026
Merged

TAIPANBOX merged 3 commits into
mainfrom
fix/cache-lookup-without-a-global-walk

Conversation

@TAIPANBOX

Copy link
Copy Markdown
Owner

What

SemanticCache::get/put (crates/core/src/cache.rs) are rewritten
around a per-partition exact-match index instead of one global Mutex
plus a linear walk:

  1. Exact-match fast path. A cheap 64-bit hash of the normalized core
    text maps to the (small) set of candidate entries; each is verified
    against a SHA-256 digest of its own normalized core text before ever
    being served, so a 64-bit bucket collision can never serve the wrong
    response. An identical core is found in O(1) average with no embedding
    call and no similarity walk at all.
  2. put replaces, not appends. A core already present in its
    partition is refreshed in place (new response, cost, created_millis,
    same id) instead of growing the partition with a duplicate.
  3. Embeddings stored L2-normalized, query normalized once per lookup;
    similarity is a plain dot product. A zero vector normalizes to itself
    and dots to 0.0 against anything, never NaN.
  4. No retain() inside get. Expired entries are skipped while
    walking, not swept; the sweep is partition-local and runs inside put.
  5. RwLock instead of Mutex, so concurrent get callers are never
    serialized against each other. put takes the write lock.
  6. O(1) eviction via a VecDeque of insertion order, with the
    exact-match index kept consistent on both eviction and replacement.

This is T3 (a wrong hit silently serves another request's response in
On mode); correctness is prioritized over speed throughout.

Why

Issue #319, same root cause as PR #326: SemanticCache::get ran under
one global lock with a full linear walk on every eligible call.

Tests

14 new, all in core::cache::tests, plus every pre-existing test (the
original 7) unchanged and green:

  • identical_core_twice_replaces_not_appends
  • an_exact_hit_never_walks_similarity
  • a_miss_still_falls_through_to_the_similarity_walk
  • hash_collision_never_serves_the_wrong_response (a test-only hook
    forces two different cores into one bucket, since a real 64-bit SipHash
    collision cannot be hand-crafted, then queries through the same
    find_exact the real get calls)
  • an_expired_exact_entry_is_not_served
  • expired_entries_disappear_after_a_put
  • zero_vector_embedding_never_matches
  • eviction_keeps_the_exact_index_in_step
  • entity_and_length_guards_still_apply_on_the_similarity_path
  • threshold_boundary_is_inclusive_on_the_similarity_path
  • a_similarity_just_under_the_threshold_never_hits
  • the_query_is_normalized_before_the_similarity_walk
  • concurrent_get_and_put_never_deadlock_or_cross_wires (8 threads x 200
    put/get pairs; asserts every response served is the one its own thread
    just wrote)
  • shadow_mode_still_records_and_reports_would_hits (replaces
    shadow_mode_still_never_serves: my first version of that test
    encoded a wrong assumption about get's mode contract — Shadow
    still finds and reports a hit, only Off skips lookup entirely, and
    the caller in proxy.rs is what decides not to serve it. Caught by
    running it, not by review.)

347 tests in tokenfuse-core, 1500 across the workspace, all green
(cargo test --all).

Mutation table (7 mutants, hand-planted and reverted)

# Mutant Caught by
m1 find_exact's digest check skipped (any candidate in the bucket accepted) hash_collision_never_serves_the_wrong_response
m2 Exact path's TTL check dropped an_expired_exact_entry_is_not_served, ttl_expires_entries
m3 put always appends, never checks for an existing entry identical_core_twice_replaces_not_appends
m4a Threshold >= -> > threshold_boundary_is_inclusive_on_the_similarity_path
m4b Threshold >= -> <= a_similarity_just_under_the_threshold_never_hits
m5 Entity guard removed eviction_keeps_the_exact_index_in_step (NOT entity_guard_blocks_number_mismatch as I expected — that pair's HashEmbedder similarity happens to stay under its own 0.9 threshold even with the guard gone, so it's a weaker witness than it looks. Recorded rather than quietly relabelled.)
m6 Partition::remove drops the exact-index cleanup on eviction eviction_keeps_the_exact_index_in_step
m7 Query's l2_normalize call removed from get the_query_is_normalized_before_the_similarity_walk

Every mutant was applied, the cache test suite run to confirm red, then
reverted and re-confirmed green.

Benchmark

Ignored, timing-only test (get_cost_at_10k_entries, run with --ignored --nocapture), 10,000 entries/partition (the default cap), release build:

before (old code) after (this PR)
exact-repeat lookup 79.751us 717ns (~111x faster)
genuine miss (similarity walk) 83.864us 207.854us (~2.5x slower)

The miss-path regression is real and reported rather than hidden: the
walk now iterates a HashMap<u64, Entry> instead of a contiguous
Vec<Entry>, which costs more per entry than it saves. Recorded as
CLAUDE.md invariant 64's own "not closed" note — a slab/Vec-backed
partition indexed by id would recover it, not done here. The hot path
this PR exists for is repeated/near-repeated traffic (most calls become
exact repeats once a cache has run for a while), and even at 207.854us
the miss path is three orders of magnitude below a live provider call.

Gates

cargo fmt --all -- --check, cargo clippy --all-targets and
cargo clippy --all-targets --all-features -- -D warnings, cargo test --all, cargo test -p tokenfuse-gateway --features cluster --test cluster_backend, every scripts/*.sh in CLAUDE.md's gate list
(core-deps.sh confirms no new dependency: sha2 is already on
tokenfuse-core's allowed list, invariant 1), and gates-have-teeth.sh
after committing: 65 cases, all as expected.

Also added: features/the-cache-stops-walking-everything.feature (8
scenarios, each bound) and CLAUDE.md invariant 64.

NOT proven

The miss-path regression (see benchmark above) is not closed, only
measured and recorded. This PR does not touch TOKENFUSE_CACHE's
default (that's PR #326) or anything about Shadow/On mode's
behavior beyond the lookup mechanics. cosine (the old, general-purpose
similarity primitive) is kept public but is no longer called from inside
this file; nothing outside it calls it either, at least as far as this
repository's own source goes.

Refs #319

🤖 Generated with Claude Code

TAIPANBOX and others added 3 commits September 24, 2026 11:48
SemanticCache::get held one global Mutex over the whole store, ran a
retain() TTL sweep over the whole partition, then a linear cosine walk
against every stored entry, on every eligible call. Issue #319 measured
79-92% of gateway CPU in SemanticCache::get and throughput dropping from
1102 req/s to 177 req/s (with 403s) once the cache carried load, in
Shadow as much as On.

Each partition now keeps an exact-match index: a cheap 64-bit hash of
the normalized core text maps to entries that might share it, each
candidate verified against a SHA-256 digest of its own normalized core
text before ever being served, so a 64-bit bucket collision can never
serve the wrong response. An identical core is found in O(1) average
with no embedding call and no similarity walk; only a core this
partition has never seen falls through to the unchanged linear walk.
put() of a core that already exists replaces that entry in place
instead of appending a duplicate. Eviction is O(1) via a VecDeque of
insertion order, with the exact-match index kept consistent on both
eviction and replacement.

Embeddings are stored L2-normalized (the query is normalized once per
lookup), so similarity is a plain dot product; a zero vector never
matches and never produces NaN. get() takes a read lock and removes
nothing, so concurrent readers are never serialized against each
other; put() takes the write lock and does the partition-local TTL
sweep.

Measured at 10,000 entries/partition, release build: an exact-repeat
lookup went from 79.751us to 717ns (about 111x). A genuine miss (falls
through to the similarity walk) went from 83.864us to 207.854us, about
2.5x slower, because the walk now iterates a HashMap<u64, Entry>
instead of a contiguous Vec<Entry>. Recorded as a known, not-closed
tradeoff (CLAUDE.md invariant 64) rather than hidden.

14 new tests in core::cache, all pre-existing tests unchanged and
green. Seven mutants planted by hand in the product code and reverted,
each caught by a named test (see CLAUDE.md invariant 64 for the table).

Adds CLAUDE.md invariant 64 and
features/the-cache-stops-walking-everything.feature (8 scenarios, each
bound). README.md and PROGRESS.md test counts updated.

Refs #319

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Architect review on #327 found two real problems in round 1:

1. put() was still O(n) under the write lock: sweep_expired ran a
   full retain() over every entry in the partition on every call, and
   order was rebuilt with a HashSet each time. In shadow mode and on
   every miss that shape reintroduces #319 itself, just moved from get
   to put.

2. The miss path (falls through to the similarity walk) regressed
   2.5x (83.9us -> 207.9us) because entries lived in a HashMap, whose
   scattered buckets cost more per entry to walk than a contiguous
   Vec.

Both fixed together, since the fix for one shapes the fix for the
other:

- entries now live in a dense Vec<Entry>, with id_to_pos: HashMap<u64,
  usize> mapping a stable id to its current slot; removal is O(1) via
  swap_remove with the displaced entry's slot updated. The similarity
  walk iterates Vec::iter(), contiguous memory.

- Eviction and TTL sweeping are amortised O(1): order is now a
  VecDeque<OrderRecord> (id + the created_millis it was written or
  refreshed with). A record is stale the moment its entry is gone or
  has since been refreshed to a different created_millis; stale
  records are skipped, never chased down early. sweep_expired and
  evict_over_cap each pop from the front, skip stale records, and stop
  at the first genuinely live one (TTL) or once back under the cap.
  maybe_compact_order rebuilds order from the live entries, sorted by
  created_millis, once it has grown past 2*entries.len()+16 -
  amortised O(1), since that only fires after O(n) stale records have
  accumulated.

- A refresh now pushes a fresh age-order record to the back instead of
  leaving the entry's only record in its original position: without
  that push, the refreshed entry becomes permanently un-evictable
  (nothing in order ever names its current state again) while a
  genuinely newer entry gets evicted in its place, which is exactly
  what a straight refresh-in-place looked like before this round.

Measured at 10,000 entries/partition, release build, before vs this
state: exact-repeat get 79.751us -> 633ns (~126x); miss get 83.864us
-> 72.101us (now FASTER than the old code, closing round 1's
regression); put (mixed refresh/new) 29.923us -> 3.77us (~7.9x);
steady-state shadow loop (get+put) 7421 -> 12622 calls/s (~1.7x).

17 tests in core::cache now (3 new this round): a dedicated
entity-guard witness whose two cores' similarity-above-threshold is
asserted as a precondition in the test itself
(entity_guard_blocks_a_near_identical_pair_that_would_otherwise_hit,
since the pre-existing entity_guard_blocks_number_mismatch's pair
happens to fall under threshold anyway and so was only an incidental
witness for the entity-guard mutant), and two FIFO-refresh tests
(a_refreshed_entry_is_not_evicted_as_if_it_were_still_the_oldest,
a_refreshed_entry_moves_to_the_back_of_fifo_order_not_out_of_it).

Mutation table grows to 9: m1-m7 re-verified against the rewritten
data structures, plus m8 (a refresh's fresh order record never
pushed) and m9 (is_live_record checking only presence, not
created_millis, so a stale record is treated as live), each caught by
name (see CLAUDE.md invariant 64).

Also added: put_cost_at_10k_entries and
steady_state_shadow_loop_throughput (both ignored, timing-only).

Refs #319

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Conflicts were documentation only: invariants 63 and 64 kept in order, the
test count is the sum of both branches (1509, core 359, gateway 840), checked
by stated-numbers.sh after cargo test --all.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant