Skip to content

schema: index search instead of scanning every symbol - #32

Merged
rumblefrog merged 1 commit into
masterfrom
claude/gifted-pascal-nf5o38
Sep 25, 2026
Merged

rumblefrog merged 1 commit into
masterfrom
claude/gifted-pascal-nf5o38

Conversation

@rumblefrog

@rumblefrog rumblefrog commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

Summary

Bundle.search and Strand.search used to score every symbol, argument and return type on every query. Each symbol also deep-cloned the options with JSON.parse(JSON.stringify(...)) and went through its own await. They now use an inverted index. The index is built on the first search and cached per weighted / l1Only / identifier combination. parents is applied to each query, so it doesn't need its own index.

How the index narrows a query (schema/src/classes/search_index.ts):

  • Entries grouped by term. Many entries share a string (int, void, Handle, ...). Each distinct term is scored at most once per query.
  • Substring candidates. A term containing the needle must contain every lowercase bigram of the needle. So the index only checks the posting list of the needle's rarest bigram with includes.
  • Dice candidates. Summing shared bigrams over the needle's posting lists gives the exact Sørensen-Dice intersection. Terms that can't clear 0.5 - max weight bonus are dropped.
  • Exact rescoring. Candidates are rescored with the existing calculateScore, so the results are the same as a full scan. Needles too short for bigrams (0 or 1 characters) fall back to scoring every distinct term.

Refactor: each symbol class now implements searchEntries(options), which returns unscored entries. Declaration.search scores those entries. The index and the per-symbol search() now share one definition of what can be searched, including existing quirks such as typeset arguments reporting arg.name but being scored on arg.type.

Recent additions: results still feed processAddition until setRecentFinalization(true) is called. The order is the one the old async traversal produced: every declaration first, then nested methods round-robin across their parents. That order decides which item is kept when created.count values tie at the 20-item cutoff.

Behaviour

These are identical to before:

  • search results, including scores and order
  • the per-symbol search() output
  • getRecentAddtions()

One new caveat: changes to bundle.strands after the first search are not searched until a new Bundle is built.

Verification

Real bundles. I ran every bundle on the bundles branch of sourcemod-dev/manifest (22 bundles, including core) through the previous implementation and this one side by side. Needles were drawn from each bundle's own names: exact names, lowercased names, substrings, single-character deletions and mixed prefixes, plus fixed edge cases (empty string, single characters, whitespace, const char[]). Each needle ran with all 5 option combinations: default, with parents, weighted: false, l1Only and an identifier override. For core, 60 needles ran with every option set and 240 more with the default options only.

  • 38,605 searches and 587,946 results, with zero differences.
  • getRecentAddtions() was identical for every bundle.

Speed on the real core.bundle (9,271 entries, 3,459 distinct terms):

query before after
average over 200 needles 27.7 ms 0.16 ms
ArrayList 27.9 ms 0.05 ms
GetClientName 28.8 ms 0.37 ms
int 23.5 ms 0.86 ms
a (no bigram, falls back to scoring every term) 9.5 ms 2.9 ms

Building the index plus the first search took 51 ms.

Synthetic bundle. On a generated bundle with 50k entries, the average search went from 127 ms to 0.97 ms, again with zero differences over about 1.1M results.

Committed test. schema/src/tests/search_index.test.ts compares Bundle.search and Strand.search against a linear scan over per-symbol search(). It covers every symbol kind, all option combinations, whitespace and case variants, and 0 or 1 character needles.

Package dependencies could not be installed in the environment I used, so I typechecked with tsc and ran everything with bun, not Jest. The package version is not bumped.

🤖 Generated with Claude Code

https://claude.ai/code/session_01NLCPZBzDzF1cP5qKEEhEsg

Bundle and Strand search now build an inverted index on first use
(cached per weighted/l1Only/identifier option combination) instead of
scoring every symbol, argument and return type on every query.

- Entries are grouped by term, so shared strings like int, void and
  Handle are scored once per query.
- Substring candidates come from the posting list of the needle's rarest
  lowercase bigram.
- Dice candidates come from summing shared bigrams over the needle's
  posting lists, which gives the exact coefficient, and are pruned
  against the 0.5 cutoff minus the largest weight bonus.
- Survivors are rescored with calculateScore, so results, order and
  recent additions match the previous full scan.

Per-symbol search is refactored into searchEntries (unscored entries)
plus a shared scorer, so the index and search() share one definition of
what is searchable. This also drops the JSON deep clone of options per
symbol.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NLCPZBzDzF1cP5qKEEhEsg
@rumblefrog
rumblefrog merged commit c758457 into master Sep 25, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants