Skip to content

Fix #2417: split_continuous_references leaves references merged when comma styles are mixed - #2419

Open
Memtensor-AI wants to merge 2 commits into
MemTensor:dev-v2.0.36from
Memtensor-AI:bugfix/autodev-2417-20260928032002157
Open

Memtensor-AI wants to merge 2 commits into
MemTensor:dev-v2.0.36from
Memtensor-AI:bugfix/autodev-2417-20260928032002157

Conversation

@Memtensor-AI

Copy link
Copy Markdown
Collaborator

Description

Fixes #2417: split_continuous_references in src/memos/mem_os/utils/reference_utils.py no longer strands bare commas when a reference tag mixes , and , separator styles. The two sequential str.replace passes (one for ", ", one for ",") mutated the text on the first pass, so the second pass could not find the original content substring and any surviving bare commas were left un-split — producing merged references like [1:92ff35fb,4:bfe6f044][7:abcd1234] that render as literal text on the streaming chat path.

The fix replaces the two passes with a single re.sub(r",\s*", "][", content_between_brackets), so every comma between the brackets becomes a split boundary regardless of trailing whitespace. All guard clauses (single [, single ], comma present between brackets) are preserved bit-for-bit; multi-tag and non-reference inputs are unaffected.

Tests: added tests/mem_os/test_reference_utils.py with 17 cases covering the exact reproduction from the issue plus other separator variants (bare, comma-space, multi-space, tab) and the existing guard clauses. All 17 pass; the full tests/mem_os/ regression suite (53 tests) also passes; ruff check and ruff format are clean on the touched files. Before the fix, 5 of the new cases (including the issue's exact reproduction) failed red as expected.

Related Issue (Required): Fixes #2417

Type of change

Please delete options that are not relevant.

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Refactor (does not change functionality, e.g. code style improvements, linting)
  • Documentation update

How Has This Been Tested?

Automated tests are pending.

  • Unit Test
  • Test Script Or Test Steps (please provide)
  • Pipeline Automated API Test (please provide)

Checklist

  • I have performed a self-review of my own code
  • I have commented my code in hard-to-understand areas
  • I have added tests that prove my fix is effective or that my feature works
  • I have created related documentation issue/PR in MemOS-Docs (if applicable)
  • I have linked the issue to this PR (if applicable)
  • I have mentioned the person who will review this PR

@WeiminLee please review this PR.

Reviewer Checklist

…ensor#2417)

split_continuous_references used two sequential str.replace passes to
convert `[1:a, 2:b]` into `[1:a][2:b]`: one for `", "` and one for
`","`. The first pass mutated the text, so the second pass could no
longer find the original content substring and any surviving bare
commas were stranded. For inputs mixing both styles in a single tag
(e.g. `[1:aaa,4:bbb, 7:ccc]`) only the last boundary was split,
leaving merged references that rendered as literal text on the
streaming chat path.

Replace the two passes with a single `re.sub(r",\s*", "][", ...)` on
the content substring. Every comma between the brackets becomes a
split boundary regardless of whether it is followed by whitespace.

Add tests/mem_os/test_reference_utils.py (17 cases) covering the
exact reproduction from the issue plus other separator variants and
the guard clauses.
@Memtensor-AI Memtensor-AI added ai:generated Generated or modified by AI | 由 AI 生成或修改 area:core MOS 编排层 / 框架底座 / 跨模块问题 status:in-progress Someone or AI is working on it | 人工或 AI 正在处理 labels Sep 28, 2026
@Memtensor-AI

Memtensor-AI commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator Author

🤖 Open Code Review

Target: PR #2419
Task: 75c34c73ae1d5ded
Base: dev-v2.0.36
Head: bugfix/autodev-2417-20260928032002157
Head SHA: 935e59b8e13e92c89091c8ff63b605dc6e6d0efd

✅ OpenCodeReview: Review complete: 0 finding(s) across 2 selected item(s).

Generated by cloud-assistant via Open Code Review.

@Memtensor-AI

Copy link
Copy Markdown
Collaborator Author

🔧 Open Code Review requested Agent fix

Open Code Review found 1 issue(s). I have resumed the development Agent to fix them.

  • Task: 75c34c73ae1d5ded
  • Fix attempt: 1/2
  • Finding delta: 0 repeated / 1 new / 0 likely resolved

The Agent will push a new commit to this PR branch. OCR will recheck after the commit is pushed.

test_content_containing_no_comma_returned_unchanged exercised the same
early-return branch as test_single_reference_unchanged. The regex has no
digit:token constraint, so [abc] and [1:aaa] are indistinguishable at
the logic level.
@Memtensor-AI

Copy link
Copy Markdown
Collaborator Author

✅ Automated Test Results: PASSED

All tests passed (17/17 executed). memos_github_open_source/smoke: 1/1, memos_python_core/changed-repo-python: 16/16. Duration: 10s [advisory, non-gating] AI-generated tests on branch test/auto-gen-75c34c73ae1d5ded-20260928113727: 33/34 passed, 1 failed — these do NOT affect the PR verdict; review the branch manually.

Branch: bugfix/autodev-2417-20260928032002157

@Memtensor-AI Memtensor-AI added status:ready Ready for implementation; waiting for assignee or AI dispatch | 可进入实现,等待认领或派发 and removed status:in-progress Someone or AI is working on it | 人工或 AI 正在处理 labels Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ai:generated Generated or modified by AI | 由 AI 生成或修改 area:core MOS 编排层 / 框架底座 / 跨模块问题 status:ready Ready for implementation; waiting for assignee or AI dispatch | 可进入实现,等待认领或派发

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants