Skip to content

Fix #2446: split_continuous_references corrupts plain bracketed lists (non-reference text) - #2450

Open
Memtensor-AI wants to merge 2 commits into
MemTensor:dev-v2.0.36from
Memtensor-AI:bugfix/autodev-2446-20260930155806113
Open

Memtensor-AI wants to merge 2 commits into
MemTensor:dev-v2.0.36from
Memtensor-AI:bugfix/autodev-2446-20260930155806113

Conversation

@Memtensor-AI

Copy link
Copy Markdown
Collaborator

Description

Fix #2446: split_continuous_references now guards against plain bracketed prose. Previously the function rewrote any single [...] block containing a comma into split tags, so "The set is [apple, banana] here" was corrupted into "The set is [apple][banana] here". Added a shape guard that only rewrites blocks whose every comma-separated item matches the reference tag shape <int>:<id> (e.g. 1:92ff35fb); otherwise the block is returned unchanged.

Changes: src/memos/mem_os/utils/reference_utils.py gains a module-level _REFERENCE_ITEM_RE regex and a validation step before the split; docstring updated to describe the new condition. New test file tests/mem_os/test_reference_utils.py covers the regression case from the issue, mixed-content brackets (e.g. [1:92ff35fb, banana]), numeric-only lists ([1, 2, 3]), non-integer prefixes ([a:1, b:2]), and the original happy path ([1:92ff35fb, 4:bfe6f044] → split, with and without space after comma).

Verification: pytest tests/mem_os/test_reference_utils.py -q → 11 passed. pytest tests/mem_os/ -q → 47 passed. ruff check + ruff format --check → clean. Task file archived to memos-autodev-specs main.

Related Issue (Required): Fixes #2446

Type of change

Please delete options that are not relevant.

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Refactor (does not change functionality, e.g. code style improvements, linting)
  • Documentation update

How Has This Been Tested?

Automated tests are pending.

  • Unit Test
  • Test Script Or Test Steps (please provide)
  • Pipeline Automated API Test (please provide)

Checklist

  • I have performed a self-review of my own code
  • I have commented my code in hard-to-understand areas
  • I have added tests that prove my fix is effective or that my feature works
  • I have created related documentation issue/PR in MemOS-Docs (if applicable)
  • I have linked the issue to this PR (if applicable)
  • I have mentioned the person who will review this PR

@WeiminLee please review this PR.

Reviewer Checklist

… bracketed prose

split_continuous_references previously rewrote any single [...] block that
contained a comma into pseudo reference tags, so ordinary prose like
"The set is [apple, banana] here" was corrupted into
"The set is [apple][banana] here".

Add a shape guard: only split when every comma-separated item in the block
matches <int>:<id> (the actual reference tag shape, e.g. [1:92ff35fb,
4:bfe6f044]); otherwise return the text unchanged.

Adds tests/mem_os/test_reference_utils.py covering the regression case,
mixed / non-reference items, and boundary conditions (empty, multiple
brackets, single item).

Fixes MemTensor#2446
@Memtensor-AI Memtensor-AI added ai:generated Generated or modified by AI | 由 AI 生成或修改 area:core MOS 编排层 / 框架底座 / 跨模块问题 status:in-progress Someone or AI is working on it | 人工或 AI 正在处理 labels Sep 30, 2026
@Memtensor-AI

Memtensor-AI commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator Author

🤖 Open Code Review

Target: PR #2450
Task: 9201e950886deeb7
Base: dev-v2.0.36
Head: bugfix/autodev-2446-20260930155806113
Head SHA: 907c95e7a14a56feb5b724e16e00b299e699a314

✅ OpenCodeReview: Review complete: 0 finding(s) across 2 selected item(s).

Generated by cloud-assistant via Open Code Review.

@Memtensor-AI

Copy link
Copy Markdown
Collaborator Author

🔧 Open Code Review requested Agent fix

Open Code Review found 1 issue(s). I have resumed the development Agent to fix them.

  • Task: 9201e950886deeb7
  • Fix attempt: 1/2
  • Finding delta: 0 repeated / 1 new / 0 likely resolved

The Agent will push a new commit to this PR branch. OCR will recheck after the commit is pushed.

The previous two-step str.replace approach in split_continuous_references
was broken for blocks that mixed ", " and bare "," separators:

  Input: "[1:aa, 2:bb,3:cc]"
  content_between_brackets = "1:aa, 2:bb,3:cc"
  After the first replace, text became "[1:aa][2:bb,3:cc]"; the second
  replace searched for the original substring, could no longer find it,
  and the bare comma between 2:bb and 3:cc was never split.

Rebuild the bracketed block in a single pass by joining the already-
computed items with "][" and stripping surrounding whitespace, so all
separator styles collapse identically.

Adds two regression tests for mixed separator styles (PR MemTensor#2450 review).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ai:generated Generated or modified by AI | 由 AI 生成或修改 area:core MOS 编排层 / 框架底座 / 跨模块问题 status:in-progress Someone or AI is working on it | 人工或 AI 正在处理

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants