fix(cache): exact-match index replaces the global similarity walk - #327
Merged
Merged
Conversation
SemanticCache::get held one global Mutex over the whole store, ran a retain() TTL sweep over the whole partition, then a linear cosine walk against every stored entry, on every eligible call. Issue #319 measured 79-92% of gateway CPU in SemanticCache::get and throughput dropping from 1102 req/s to 177 req/s (with 403s) once the cache carried load, in Shadow as much as On. Each partition now keeps an exact-match index: a cheap 64-bit hash of the normalized core text maps to entries that might share it, each candidate verified against a SHA-256 digest of its own normalized core text before ever being served, so a 64-bit bucket collision can never serve the wrong response. An identical core is found in O(1) average with no embedding call and no similarity walk; only a core this partition has never seen falls through to the unchanged linear walk. put() of a core that already exists replaces that entry in place instead of appending a duplicate. Eviction is O(1) via a VecDeque of insertion order, with the exact-match index kept consistent on both eviction and replacement. Embeddings are stored L2-normalized (the query is normalized once per lookup), so similarity is a plain dot product; a zero vector never matches and never produces NaN. get() takes a read lock and removes nothing, so concurrent readers are never serialized against each other; put() takes the write lock and does the partition-local TTL sweep. Measured at 10,000 entries/partition, release build: an exact-repeat lookup went from 79.751us to 717ns (about 111x). A genuine miss (falls through to the similarity walk) went from 83.864us to 207.854us, about 2.5x slower, because the walk now iterates a HashMap<u64, Entry> instead of a contiguous Vec<Entry>. Recorded as a known, not-closed tradeoff (CLAUDE.md invariant 64) rather than hidden. 14 new tests in core::cache, all pre-existing tests unchanged and green. Seven mutants planted by hand in the product code and reverted, each caught by a named test (see CLAUDE.md invariant 64 for the table). Adds CLAUDE.md invariant 64 and features/the-cache-stops-walking-everything.feature (8 scenarios, each bound). README.md and PROGRESS.md test counts updated. Refs #319 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Architect review on #327 found two real problems in round 1: 1. put() was still O(n) under the write lock: sweep_expired ran a full retain() over every entry in the partition on every call, and order was rebuilt with a HashSet each time. In shadow mode and on every miss that shape reintroduces #319 itself, just moved from get to put. 2. The miss path (falls through to the similarity walk) regressed 2.5x (83.9us -> 207.9us) because entries lived in a HashMap, whose scattered buckets cost more per entry to walk than a contiguous Vec. Both fixed together, since the fix for one shapes the fix for the other: - entries now live in a dense Vec<Entry>, with id_to_pos: HashMap<u64, usize> mapping a stable id to its current slot; removal is O(1) via swap_remove with the displaced entry's slot updated. The similarity walk iterates Vec::iter(), contiguous memory. - Eviction and TTL sweeping are amortised O(1): order is now a VecDeque<OrderRecord> (id + the created_millis it was written or refreshed with). A record is stale the moment its entry is gone or has since been refreshed to a different created_millis; stale records are skipped, never chased down early. sweep_expired and evict_over_cap each pop from the front, skip stale records, and stop at the first genuinely live one (TTL) or once back under the cap. maybe_compact_order rebuilds order from the live entries, sorted by created_millis, once it has grown past 2*entries.len()+16 - amortised O(1), since that only fires after O(n) stale records have accumulated. - A refresh now pushes a fresh age-order record to the back instead of leaving the entry's only record in its original position: without that push, the refreshed entry becomes permanently un-evictable (nothing in order ever names its current state again) while a genuinely newer entry gets evicted in its place, which is exactly what a straight refresh-in-place looked like before this round. Measured at 10,000 entries/partition, release build, before vs this state: exact-repeat get 79.751us -> 633ns (~126x); miss get 83.864us -> 72.101us (now FASTER than the old code, closing round 1's regression); put (mixed refresh/new) 29.923us -> 3.77us (~7.9x); steady-state shadow loop (get+put) 7421 -> 12622 calls/s (~1.7x). 17 tests in core::cache now (3 new this round): a dedicated entity-guard witness whose two cores' similarity-above-threshold is asserted as a precondition in the test itself (entity_guard_blocks_a_near_identical_pair_that_would_otherwise_hit, since the pre-existing entity_guard_blocks_number_mismatch's pair happens to fall under threshold anyway and so was only an incidental witness for the entity-guard mutant), and two FIFO-refresh tests (a_refreshed_entry_is_not_evicted_as_if_it_were_still_the_oldest, a_refreshed_entry_moves_to_the_back_of_fifo_order_not_out_of_it). Mutation table grows to 9: m1-m7 re-verified against the rewritten data structures, plus m8 (a refresh's fresh order record never pushed) and m9 (is_live_record checking only presence, not created_millis, so a stale record is treated as live), each caught by name (see CLAUDE.md invariant 64). Also added: put_cost_at_10k_entries and steady_state_shadow_loop_throughput (both ignored, timing-only). Refs #319 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Conflicts were documentation only: invariants 63 and 64 kept in order, the test count is the sum of both branches (1509, core 359, gateway 840), checked by stated-numbers.sh after cargo test --all. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Sep 24, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
SemanticCache::get/put(crates/core/src/cache.rs) are rewrittenaround a per-partition exact-match index instead of one global
Mutexplus a linear walk:
text maps to the (small) set of candidate entries; each is verified
against a SHA-256 digest of its own normalized core text before ever
being served, so a 64-bit bucket collision can never serve the wrong
response. An identical core is found in O(1) average with no embedding
call and no similarity walk at all.
putreplaces, not appends. A core already present in itspartition is refreshed in place (new response, cost,
created_millis,same id) instead of growing the partition with a duplicate.
similarity is a plain dot product. A zero vector normalizes to itself
and dots to
0.0against anything, neverNaN.retain()insideget. Expired entries are skipped whilewalking, not swept; the sweep is partition-local and runs inside
put.RwLockinstead ofMutex, so concurrentgetcallers are neverserialized against each other.
puttakes the write lock.VecDequeof insertion order, with theexact-match index kept consistent on both eviction and replacement.
This is T3 (a wrong hit silently serves another request's response in
Onmode); correctness is prioritized over speed throughout.Why
Issue #319, same root cause as PR #326:
SemanticCache::getran underone global lock with a full linear walk on every eligible call.
Tests
14 new, all in
core::cache::tests, plus every pre-existing test (theoriginal 7) unchanged and green:
identical_core_twice_replaces_not_appendsan_exact_hit_never_walks_similaritya_miss_still_falls_through_to_the_similarity_walkhash_collision_never_serves_the_wrong_response(a test-only hookforces two different cores into one bucket, since a real 64-bit SipHash
collision cannot be hand-crafted, then queries through the same
find_exactthe realgetcalls)an_expired_exact_entry_is_not_servedexpired_entries_disappear_after_a_putzero_vector_embedding_never_matcheseviction_keeps_the_exact_index_in_stepentity_and_length_guards_still_apply_on_the_similarity_paththreshold_boundary_is_inclusive_on_the_similarity_patha_similarity_just_under_the_threshold_never_hitsthe_query_is_normalized_before_the_similarity_walkconcurrent_get_and_put_never_deadlock_or_cross_wires(8 threads x 200put/get pairs; asserts every response served is the one its own thread
just wrote)
shadow_mode_still_records_and_reports_would_hits(replacesshadow_mode_still_never_serves: my first version of that testencoded a wrong assumption about
get's mode contract —Shadowstill finds and reports a hit, only
Offskips lookup entirely, andthe caller in
proxy.rsis what decides not to serve it. Caught byrunning it, not by review.)
347 tests in
tokenfuse-core, 1500 across the workspace, all green(
cargo test --all).Mutation table (7 mutants, hand-planted and reverted)
find_exact's digest check skipped (any candidate in the bucket accepted)hash_collision_never_serves_the_wrong_responsean_expired_exact_entry_is_not_served,ttl_expires_entriesputalways appends, never checks for an existing entryidentical_core_twice_replaces_not_appends>=->>threshold_boundary_is_inclusive_on_the_similarity_path>=-><=a_similarity_just_under_the_threshold_never_hitseviction_keeps_the_exact_index_in_step(NOTentity_guard_blocks_number_mismatchas I expected — that pair's HashEmbedder similarity happens to stay under its own 0.9 threshold even with the guard gone, so it's a weaker witness than it looks. Recorded rather than quietly relabelled.)Partition::removedrops the exact-index cleanup on evictioneviction_keeps_the_exact_index_in_stepl2_normalizecall removed fromgetthe_query_is_normalized_before_the_similarity_walkEvery mutant was applied, the cache test suite run to confirm red, then
reverted and re-confirmed green.
Benchmark
Ignored, timing-only test (
get_cost_at_10k_entries, run with--ignored --nocapture), 10,000 entries/partition (the default cap), release build:The miss-path regression is real and reported rather than hidden: the
walk now iterates a
HashMap<u64, Entry>instead of a contiguousVec<Entry>, which costs more per entry than it saves. Recorded asCLAUDE.md invariant 64's own "not closed" note — a slab/
Vec-backedpartition indexed by id would recover it, not done here. The hot path
this PR exists for is repeated/near-repeated traffic (most calls become
exact repeats once a cache has run for a while), and even at 207.854us
the miss path is three orders of magnitude below a live provider call.
Gates
cargo fmt --all -- --check,cargo clippy --all-targetsandcargo clippy --all-targets --all-features -- -D warnings,cargo test --all,cargo test -p tokenfuse-gateway --features cluster --test cluster_backend, everyscripts/*.shin CLAUDE.md's gate list(
core-deps.shconfirms no new dependency:sha2is already ontokenfuse-core's allowed list, invariant 1), andgates-have-teeth.shafter committing: 65 cases, all as expected.
Also added:
features/the-cache-stops-walking-everything.feature(8scenarios, each bound) and CLAUDE.md invariant 64.
NOT proven
The miss-path regression (see benchmark above) is not closed, only
measured and recorded. This PR does not touch
TOKENFUSE_CACHE'sdefault (that's PR #326) or anything about
Shadow/Onmode'sbehavior beyond the lookup mechanics.
cosine(the old, general-purposesimilarity primitive) is kept public but is no longer called from inside
this file; nothing outside it calls it either, at least as far as this
repository's own source goes.
Refs #319
🤖 Generated with Claude Code