Repository navigation
Conversation
EmbeddingGemma 2 ships as one ONNX graph that also declares image, video and audio feature inputs, so BuiltinSentenceEmbedding now feeds them as zero-row tensors when the model is used for text only. The onnx-community export's tokenizer_config.json carries transformers' "unknown" model_max_length sentinel, which load_tokenizer rejects. Add an optional default_max_length that a model class can declare for such exports; usable config keys still take precedence. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (4)
Included review availability: This review used your included allowance. Your plan provides up to 8 included reviews per hour; 7 remain after this review. 📝 WalkthroughWalkthroughThe tokenizer utilities now accept a default maximum length when tokenizer configuration has no usable limit. The supported model list adds full-precision and quantized Priority: ⬇️ Low Estimated code review effort: 3 (Moderate) | ~20 minutes Change: Feature Merge Risk: ⚪ Minimal · up to No actionable issue remains; the change is mergeable after normal checks. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Adds EmbeddingGemma 2 (Apache 2.0, 768-dim, 8192-token context, multilingual) as two text models backed by
onnx-community/embeddinggemma-2-ONNX:google/embeddinggemma-2onnx/model.onnx(+.onnx_data)google/embeddinggemma-2-Qonnx/model_quantized.onnx(+.onnx_data), int8Both are
BuiltinSentenceEmbeddingmodels, likegoogle/embeddinggemma-300m: the graph's second output is the pooled, projected and normalizedsentence_embedding. Only the text graph is downloaded; the vision and audio encoders are separate files and are not fetched.Two changes beyond the model entries
image_features,video_featuresandaudio_featuresinputs ([num_*_tokens, 512]).BuiltinSentenceEmbedding._preprocess_onnx_inputnow feeds any missing*_featuresinput as a zero-row tensor. It's a no-op for graphs without such inputs, so the existing models are unaffected.tokenizer_config.jsonhas transformers'model_max_lengthsentinel (1e30), whichload_tokenizerrightly rejects. I added an optionaldefault_max_lengthtoload_tokenizer/_resolve_max_context. A model class can declare it for exports like this one, and a usablemodel_max_length/max_lengthin the config still wins.BuiltinSentenceEmbeddingpasses 8192 for the two new models only. If you'd rather not grow theload_tokenizersignature, the alternative is fixingmodel_max_lengthin the HF repo (or hosting a mirror, as withQdrant/Qwen3-Embedding-0.6B-onnx);DEFAULT_MAX_LENGTHScould then go.Canonical values
Computed with the reference implementation:
SentenceTransformer("google/embeddinggemma-2", model_kwargs={"dtype": torch.float32})on sentence-transformers 6.1.0 and transformers 5.19.0, using the same prefixed inputs as the existing EmbeddingGemma tests. Against those values:hello worldmax abs diff (first 5)google/embeddinggemma-2google/embeddinggemma-2-QBoth models therefore share the reference values in
CANONICAL_VECTOR_VALUES/CANONICAL_QUERY_VECTOR_VALUES(within the tests'atol=1e-3).Tests run locally
tests/test_preprocessor_utils.py: 43 passed, including the newtest_default_max_length_fills_unusable_config(sentinel → default, config key still wins).test_embedding/test_query_embeddingassertions forgoogle/embeddinggemma-2-Qvia a fresh Hugging Face download: pass,truncation.max_length == 8192.google/embeddinggemma-2(fp32), loaded from the same repo files with the unmodifiedtokenizer_config.json: matches the reference as above.ruff check/ruff format --check(v0.3.4, per pre-commit) andmypy fastembed --disallow-incomplete-defs --disallow-untyped-defs --disable-error-code=import-untyped: clean.The fp32 model is 1.08 GB, so it's skipped by the local
size_in_GB > 1rule and covered by the manual/weekly CI run.🤖 Generated with Claude Code