Skip to content

new: add google/embeddinggemma-2 and its int8 variant - #790

Open
nova28 wants to merge 1 commit into
qdrant:mainfrom
nova28:add-embeddinggemma-2
Open

nova28 wants to merge 1 commit into
qdrant:mainfrom
nova28:add-embeddinggemma-2

Conversation

@nova28

@nova28 nova28 commented Oct 8, 2026

Copy link
Copy Markdown

Adds EmbeddingGemma 2 (Apache 2.0, 768-dim, 8192-token context, multilingual) as two text models backed by onnx-community/embeddinggemma-2-ONNX:

model file size
google/embeddinggemma-2 onnx/model.onnx (+ .onnx_data) 1.08 GB
google/embeddinggemma-2-Q onnx/model_quantized.onnx (+ .onnx_data), int8 0.31 GB

Both are BuiltinSentenceEmbedding models, like google/embeddinggemma-300m: the graph's second output is the pooled, projected and normalized sentence_embedding. Only the text graph is downloaded; the vision and audio encoders are separate files and are not fetched.

Two changes beyond the model entries

  1. Empty modality inputs. The text graph also declares image_features, video_features and audio_features inputs ([num_*_tokens, 512]). BuiltinSentenceEmbedding._preprocess_onnx_input now feeds any missing *_features input as a zero-row tensor. It's a no-op for graphs without such inputs, so the existing models are unaffected.
  2. Max context length. The export's tokenizer_config.json has transformers' model_max_length sentinel (1e30), which load_tokenizer rightly rejects. I added an optional default_max_length to load_tokenizer / _resolve_max_context. A model class can declare it for exports like this one, and a usable model_max_length / max_length in the config still wins. BuiltinSentenceEmbedding passes 8192 for the two new models only. If you'd rather not grow the load_tokenizer signature, the alternative is fixing model_max_length in the HF repo (or hosting a mirror, as with Qdrant/Qwen3-Embedding-0.6B-onnx); DEFAULT_MAX_LENGTHS could then go.

Canonical values

Computed with the reference implementation: SentenceTransformer("google/embeddinggemma-2", model_kwargs={"dtype": torch.float32}) on sentence-transformers 6.1.0 and transformers 5.19.0, using the same prefixed inputs as the existing EmbeddingGemma tests. Against those values:

doc hello world max abs diff (first 5) query max abs diff (first 5) cosine to reference
google/embeddinggemma-2 < 1e-6 < 1e-6 1.00000
google/embeddinggemma-2-Q 2.8e-4 3.6e-4 0.99996–0.99997

Both models therefore share the reference values in CANONICAL_VECTOR_VALUES / CANONICAL_QUERY_VECTOR_VALUES (within the tests' atol=1e-3).

Tests run locally

  • tests/test_preprocessor_utils.py: 43 passed, including the new test_default_max_length_fills_unusable_config (sentinel → default, config key still wins).
  • test_embedding / test_query_embedding assertions for google/embeddinggemma-2-Q via a fresh Hugging Face download: pass, truncation.max_length == 8192.
  • google/embeddinggemma-2 (fp32), loaded from the same repo files with the unmodified tokenizer_config.json: matches the reference as above.
  • ruff check / ruff format --check (v0.3.4, per pre-commit) and mypy fastembed --disallow-incomplete-defs --disallow-untyped-defs --disable-error-code=import-untyped: clean.

The fp32 model is 1.08 GB, so it's skipped by the local size_in_GB > 1 rule and covered by the manual/weekly CI run.

🤖 Generated with Claude Code

EmbeddingGemma 2 ships as one ONNX graph that also declares image, video
and audio feature inputs, so BuiltinSentenceEmbedding now feeds them as
zero-row tensors when the model is used for text only.

The onnx-community export's tokenizer_config.json carries transformers'
"unknown" model_max_length sentinel, which load_tokenizer rejects. Add an
optional default_max_length that a model class can declare for such
exports; usable config keys still take precedence.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@nova28
nova28 requested a review from joein as a code owner October 8, 2026 04:10
@coderabbitai

coderabbitai Bot commented Oct 8, 2026

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 91288028-4729-4a7a-bda8-42f095b6e9a9
📥 Commits

Reviewing files that changed from the base of the PR and between 539499b and d47ea52.

📒 Files selected for processing (4)
  • fastembed/common/preprocessor_utils.py
  • fastembed/text/builtin_sentence_embedding.py
  • tests/test_preprocessor_utils.py
  • tests/test_text_onnx_embeddings.py

Included review availability: This review used your included allowance. Your plan provides up to 8 included reviews per hour; 7 remain after this review.


📝 Walkthrough

Walkthrough

The tokenizer utilities now accept a default maximum length when tokenizer configuration has no usable limit. The supported model list adds full-precision and quantized google/embeddinggemma-2 entries. Tokenizer loading applies an 8192-token fallback for these models, and ONNX preprocessing supplies zero-row tensors for missing _features inputs. Tests cover tokenizer limit selection and document and query embeddings for both model variants.

Priority: ⬇️ Low

Estimated code review effort: 3 (Moderate) | ~20 minutes

Change: Feature

Merge Risk: ⚪ Minimal · up to d47ea

No actionable issue remains; the change is mergeable after normal checks.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 40.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 10 functions across 4 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly identifies the main change: adding google/embeddinggemma-2 and its int8 variant.
Description check ✅ Passed The description directly explains the two new models, their implementation, supporting changes, and validation results.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
  • Fix all pre-merge checks with AI
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant