Skip to content

fix: reject case-insensitive duplicate cross-encoder registrations - #795

Open
wToTw123 wants to merge 1 commit into
qdrant:mainfrom
wToTw123:codex/fix-reranker-case-registration
Open

wToTw123 wants to merge 1 commit into
qdrant:mainfrom
wToTw123:codex/fix-reranker-case-registration

Conversation

@wToTw123

@wToTw123 wToTw123 commented Oct 9, 2026

Copy link
Copy Markdown

TextCrossEncoder.add_custom_model checks duplicate names case-sensitively, while model selection and description lookup compare lowercase names. Registering a case-only variant is accepted but resolves to the earlier entry, silently selecting its original source. Case-only aliases of built-in names are affected as well.

Use the same lowercase comparison during registration and raise the existing duplicate-name ValueError. Document that model names must be unique regardless of case. Tests cover built-in and custom names, four case variants, registry preservation on rejection, and actual lazy construction of distinct custom models with case-variant names and the expected sources. This aligns with the existing TextEmbedding.add_custom_model behavior rather than changing model lookup semantics.

Validation

  • Before the production fix: 6 failed, 3 passed in the new regression module.
  • After the fix: 20 passed, 5 deselected, including existing common/custom-model unit tests, with sockets disabled. The five deselected tests download and run models.
  • mypy: no issues in 66 source files; pyright: zero errors/warnings with the isolated interpreter.
  • Ruff 0.3.4 lint and format checks, plus git diff --check: passed.
  • Full downloaded-model suite, GPU runs and the full CI matrix were not run locally. Focused offline regressions are the alternative validation.

All Submissions

  • Followed the Contributing guidelines.
  • Checked all open PRs and related closed PRs for the same change as of 2026-10-09.

New-feature and new-model checklist items are not applicable to this bug fix. No pre-commit hook was installed for this uncommitted preview; its configured Ruff checks were run directly.

AI assistance: the patch and tests were prepared with Codex and reviewed and approved for submission by the contributor.

No separate issue has been opened.

Additional offline public-API verification: register differently cased names with different sources, then construct TextCrossEncoder with lazy_load and a local model path. The baseline accepts the alias but selects the first source; the patch rejects the duplicate and still constructs the original model correctly. No ONNX inference was run.

@wToTw123
wToTw123 requested a review from joein as a code owner October 9, 2026 09:43
@coderabbitai

coderabbitai Bot commented Oct 9, 2026

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 205006e8-448a-4aec-871b-8fb907a4a08c

📥 Commits

Reviewing files that changed from the base of the PR and between d076f08 and 2a859a5.


📒 Files selected for processing (2)
  • fastembed/rerank/cross_encoder/text_cross_encoder.py
  • tests/test_cross_encoder_registration.py

Included review availability: This review used your included allowance. Your plan provides up to 8 included reviews per hour; 7 remain after this review.



📝 Walkthrough

Walkthrough

TextCrossEncoder.add_custom_model now checks registered model names case-insensitively and raises the existing ValueError for duplicates. Tests verify that rejected registrations do not change the supported-model list, and that distinct custom models can be constructed with uppercase names while retaining their configured Hugging Face sources.

Priority: ⬇️ Low

Estimated code review effort: 2 (Simple) | ~10 minutes

Change: Bug fix


Merge Risk: ⚪ Minimal · up to 2a859

Case-only duplicate names are rejected, and no issue requiring a fix before merge is identified.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check Passed The title clearly and concisely describes the main change: rejecting case-insensitive duplicate cross-encoder registrations.
Description check Passed The description directly explains the registration bug, the fix, affected behavior, tests, and validation results.
Docstring Coverage Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 3 functions across 2 files.
Linked Issues check Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check Passed Check skipped because no linked issues were found for this pull request.


✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create a new PR


  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant