schema: index search instead of scanning every symbol - #32
Merged
Merged
Conversation
Bundle and Strand search now build an inverted index on first use (cached per weighted/l1Only/identifier option combination) instead of scoring every symbol, argument and return type on every query. - Entries are grouped by term, so shared strings like int, void and Handle are scored once per query. - Substring candidates come from the posting list of the needle's rarest lowercase bigram. - Dice candidates come from summing shared bigrams over the needle's posting lists, which gives the exact coefficient, and are pruned against the 0.5 cutoff minus the largest weight bonus. - Survivors are rescored with calculateScore, so results, order and recent additions match the previous full scan. Per-symbol search is refactored into searchEntries (unscored entries) plus a shared scorer, so the index and search() share one definition of what is searchable. This also drops the JSON deep clone of options per symbol. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLCPZBzDzF1cP5qKEEhEsg
This was referenced Sep 25, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Bundle.searchandStrand.searchused to score every symbol, argument and return type on every query. Each symbol also deep-cloned the options withJSON.parse(JSON.stringify(...))and went through its ownawait. They now use an inverted index. The index is built on the first search and cached perweighted/l1Only/identifiercombination.parentsis applied to each query, so it doesn't need its own index.How the index narrows a query (
schema/src/classes/search_index.ts):int,void,Handle, ...). Each distinct term is scored at most once per query.includes.0.5 - max weight bonusare dropped.calculateScore, so the results are the same as a full scan. Needles too short for bigrams (0 or 1 characters) fall back to scoring every distinct term.Refactor: each symbol class now implements
searchEntries(options), which returns unscored entries.Declaration.searchscores those entries. The index and the per-symbolsearch()now share one definition of what can be searched, including existing quirks such as typeset arguments reportingarg.namebut being scored onarg.type.Recent additions: results still feed
processAdditionuntilsetRecentFinalization(true)is called. The order is the one the old async traversal produced: every declaration first, then nested methods round-robin across their parents. That order decides which item is kept whencreated.countvalues tie at the 20-item cutoff.Behaviour
These are identical to before:
search()outputgetRecentAddtions()One new caveat: changes to
bundle.strandsafter the first search are not searched until a newBundleis built.Verification
Real bundles. I ran every bundle on the
bundlesbranch ofsourcemod-dev/manifest(22 bundles, includingcore) through the previous implementation and this one side by side. Needles were drawn from each bundle's own names: exact names, lowercased names, substrings, single-character deletions and mixed prefixes, plus fixed edge cases (empty string, single characters, whitespace,const char[]). Each needle ran with all 5 option combinations: default, with parents,weighted: false,l1Onlyand anidentifieroverride. Forcore, 60 needles ran with every option set and 240 more with the default options only.getRecentAddtions()was identical for every bundle.Speed on the real
core.bundle(9,271 entries, 3,459 distinct terms):ArrayListGetClientNameinta(no bigram, falls back to scoring every term)Building the index plus the first search took 51 ms.
Synthetic bundle. On a generated bundle with 50k entries, the average search went from 127 ms to 0.97 ms, again with zero differences over about 1.1M results.
Committed test.
schema/src/tests/search_index.test.tscomparesBundle.searchandStrand.searchagainst a linear scan over per-symbolsearch(). It covers every symbol kind, all option combinations, whitespace and case variants, and 0 or 1 character needles.Package dependencies could not be installed in the environment I used, so I typechecked with
tscand ran everything withbun, not Jest. The package version is not bumped.🤖 Generated with Claude Code
https://claude.ai/code/session_01NLCPZBzDzF1cP5qKEEhEsg