feat(indexes): add configurable Bigram/Trigram tokenizer for CJK search - #37
Open
harehare wants to merge 1 commit into
Open
feat(indexes): add configurable Bigram/Trigram tokenizer for CJK search#37harehare wants to merge 1 commit into
harehare wants to merge 1 commit into
Conversation
The Word tokenizer lumps a punctuation-free CJK run (e.g. Japanese text)
into a single token, so match()/score()/bm25() were effectively unusable
on Japanese content. Add TokenizerKind::{Word,Bigram,Trigram}, fixed per
store at creation time and persisted in the catalog (additive trailing
section, no file-format bump), with Bigram/Trigram splitting CJK runs
into overlapping n-grams while leaving ASCII word tokenization unchanged.
Exposed via `mq-db index --tokenizer <word|bigram|trigram>`.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The Word tokenizer lumps a punctuation-free CJK run (e.g. Japanese text) into a single token, so match()/score()/bm25() were effectively unusable on Japanese content. Add TokenizerKind::{Word,Bigram,Trigram}, fixed per store at creation time and persisted in the catalog (additive trailing section, no file-format bump), with Bigram/Trigram splitting CJK runs into overlapping n-grams while leaving ASCII word tokenization unchanged.
Exposed via
mq-db index --tokenizer <word|bigram|trigram>.