Fix #2417: split_continuous_references leaves references merged when comma styles are mixed - #2419
Memtensor-AI wants to merge 2 commits into
Conversation
…ensor#2417) split_continuous_references used two sequential str.replace passes to convert `[1:a, 2:b]` into `[1:a][2:b]`: one for `", "` and one for `","`. The first pass mutated the text, so the second pass could no longer find the original content substring and any surviving bare commas were stranded. For inputs mixing both styles in a single tag (e.g. `[1:aaa,4:bbb, 7:ccc]`) only the last boundary was split, leaving merged references that rendered as literal text on the streaming chat path. Replace the two passes with a single `re.sub(r",\s*", "][", ...)` on the content substring. Every comma between the brackets becomes a split boundary regardless of whether it is followed by whitespace. Add tests/mem_os/test_reference_utils.py (17 cases) covering the exact reproduction from the issue plus other separator variants and the guard clauses.
🤖 Open Code ReviewTarget: PR #2419 ✅ OpenCodeReview: Review complete: 0 finding(s) across 2 selected item(s). Generated by cloud-assistant via Open Code Review. |
🔧 Open Code Review requested Agent fixOpen Code Review found 1 issue(s). I have resumed the development Agent to fix them.
The Agent will push a new commit to this PR branch. OCR will recheck after the commit is pushed. |
test_content_containing_no_comma_returned_unchanged exercised the same early-return branch as test_single_reference_unchanged. The regex has no digit:token constraint, so [abc] and [1:aaa] are indistinguishable at the logic level.
✅ Automated Test Results: PASSEDAll tests passed (17/17 executed). memos_github_open_source/smoke: 1/1, memos_python_core/changed-repo-python: 16/16. Duration: 10s [advisory, non-gating] AI-generated tests on branch test/auto-gen-75c34c73ae1d5ded-20260928113727: 33/34 passed, 1 failed — these do NOT affect the PR verdict; review the branch manually. Branch: |
Description
Fixes #2417:
split_continuous_referencesinsrc/memos/mem_os/utils/reference_utils.pyno longer strands bare commas when a reference tag mixes,and,separator styles. The two sequentialstr.replacepasses (one for", ", one for",") mutated the text on the first pass, so the second pass could not find the original content substring and any surviving bare commas were left un-split — producing merged references like[1:92ff35fb,4:bfe6f044][7:abcd1234]that render as literal text on the streaming chat path.The fix replaces the two passes with a single
re.sub(r",\s*", "][", content_between_brackets), so every comma between the brackets becomes a split boundary regardless of trailing whitespace. All guard clauses (single[, single], comma present between brackets) are preserved bit-for-bit; multi-tag and non-reference inputs are unaffected.Tests: added
tests/mem_os/test_reference_utils.pywith 17 cases covering the exact reproduction from the issue plus other separator variants (bare, comma-space, multi-space, tab) and the existing guard clauses. All 17 pass; the fulltests/mem_os/regression suite (53 tests) also passes;ruff checkandruff formatare clean on the touched files. Before the fix, 5 of the new cases (including the issue's exact reproduction) failed red as expected.Related Issue (Required): Fixes #2417
Type of change
Please delete options that are not relevant.
How Has This Been Tested?
Automated tests are pending.
Checklist
@WeiminLee please review this PR.
Reviewer Checklist