Fix #2446: split_continuous_references corrupts plain bracketed lists (non-reference text) - #2450
Open
Memtensor-AI wants to merge 2 commits into
Open
Memtensor-AI wants to merge 2 commits into
Memtensor-AI wants to merge 2 commits into
Conversation
… bracketed prose split_continuous_references previously rewrote any single [...] block that contained a comma into pseudo reference tags, so ordinary prose like "The set is [apple, banana] here" was corrupted into "The set is [apple][banana] here". Add a shape guard: only split when every comma-separated item in the block matches <int>:<id> (the actual reference tag shape, e.g. [1:92ff35fb, 4:bfe6f044]); otherwise return the text unchanged. Adds tests/mem_os/test_reference_utils.py covering the regression case, mixed / non-reference items, and boundary conditions (empty, multiple brackets, single item). Fixes MemTensor#2446
Collaborator
Author
🤖 Open Code ReviewTarget: PR #2450 ✅ OpenCodeReview: Review complete: 0 finding(s) across 2 selected item(s). Generated by cloud-assistant via Open Code Review. |
Collaborator
Author
🔧 Open Code Review requested Agent fixOpen Code Review found 1 issue(s). I have resumed the development Agent to fix them.
The Agent will push a new commit to this PR branch. OCR will recheck after the commit is pushed. |
The previous two-step str.replace approach in split_continuous_references was broken for blocks that mixed ", " and bare "," separators: Input: "[1:aa, 2:bb,3:cc]" content_between_brackets = "1:aa, 2:bb,3:cc" After the first replace, text became "[1:aa][2:bb,3:cc]"; the second replace searched for the original substring, could no longer find it, and the bare comma between 2:bb and 3:cc was never split. Rebuild the bracketed block in a single pass by joining the already- computed items with "][" and stripping surrounding whitespace, so all separator styles collapse identically. Adds two regression tests for mixed separator styles (PR MemTensor#2450 review).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Fix #2446: split_continuous_references now guards against plain bracketed prose. Previously the function rewrote any single [...] block containing a comma into split tags, so "The set is [apple, banana] here" was corrupted into "The set is [apple][banana] here". Added a shape guard that only rewrites blocks whose every comma-separated item matches the reference tag shape
<int>:<id>(e.g.1:92ff35fb); otherwise the block is returned unchanged.Changes:
src/memos/mem_os/utils/reference_utils.pygains a module-level_REFERENCE_ITEM_REregex and a validation step before the split; docstring updated to describe the new condition. New test filetests/mem_os/test_reference_utils.pycovers the regression case from the issue, mixed-content brackets (e.g.[1:92ff35fb, banana]), numeric-only lists ([1, 2, 3]), non-integer prefixes ([a:1, b:2]), and the original happy path ([1:92ff35fb, 4:bfe6f044]→ split, with and without space after comma).Verification:
pytest tests/mem_os/test_reference_utils.py -q→ 11 passed.pytest tests/mem_os/ -q→ 47 passed.ruff check+ruff format --check→ clean. Task file archived to memos-autodev-specs main.Related Issue (Required): Fixes #2446
Type of change
Please delete options that are not relevant.
How Has This Been Tested?
Automated tests are pending.
Checklist
@WeiminLee please review this PR.
Reviewer Checklist