Skip to content

fix(cartesia): keep plain-string search results and skip malformed messages - #1683

Open
Agnik47 wants to merge 1 commit into
supermemoryai:mainfrom
Agnik47:fix/cartesia-utils-parity
Open

Agnik47 wants to merge 1 commit into
supermemoryai:mainfrom
Agnik47:fix/cartesia-utils-parity

Conversation

@Agnik47

@Agnik47 Agnik47 commented Sep 18, 2026

Copy link
Copy Markdown
Contributor

Summary

src/supermemory_cartesia/utils.py and src/supermemory_pipecat/utils.py are copies of the same file. #1434 hardened two spots in the pipecat copy and missed the cartesia one. This brings the two hunks back to parity and adds regression tests.

Bug 1 — plain-string search results are silently dropped

deduplicate_memories → unique_search calls _field(r, "memory", "chunk", "content", default="") on every result. For a plain str that finds no attribute, falls back to "", and the entry never reaches the prompt.

format_memories_to_text in the same file already has an isinstance(item, str) branch for search results, so that branch is currently unreachable — nothing can get past dedup as a string.

Reproduced against the pipecat copy as the reference:

cartesia  search_results -> []
cartesia  rendered       -> ''
pipecat   search_results -> ['User prefers tea']
pipecat   rendered       -> 'Based on previous conversations, I recall:\n\n\n## Relevant Memories\n\n- User prefers tea'

Bug 2 — get_last_user_message raises KeyError

It indexed msg["role"] and msg["content"] positionally, so a tool-call turn or provider event missing either key raised instead of being skipped. Non-string content was also returned verbatim despite the -> str | None annotation.

user with no content   cartesia  -> KeyError: 'content'
user with no content   pipecat   -> 'hello'
entry with no role     cartesia  -> KeyError: 'role'
entry with no role     pipecat   -> 'hello'
non-string content     cartesia  -> [{'type': 'text', 'text': 'hello'}]
non-string content     pipecat   -> None

Changes

  • src/supermemory_cartesia/utils.py — both hunks now match the pipecat implementation exactly.
  • tests/test_utils.py — new, 6 tests.

Test plan

  • 3 of the 6 new tests fail on unmodified main and pass with the fix:
    FAILED tests/test_utils.py::TestGetLastUserMessage::test_skips_entries_without_role_or_content
    FAILED tests/test_utils.py::TestGetLastUserMessage::test_skips_non_string_content
    FAILED tests/test_utils.py::TestDeduplicateMemories::test_keeps_plain_string_search_results
    3 failed, 5 passed
    
  • Full suite with the fix: 8 passed.
  • ci-python.yml already path-filters packages/cartesia-sdk-python/**, so these run in CI with no workflow change.

No behaviour change for dict or Stainless-model search results — covered by test_still_reads_the_memory_field_of_models_and_dicts.

…ssages

`utils.py` in cartesia-sdk-python and pipecat-sdk-python are copies of the
same file. supermemoryai#1434 hardened two spots in the pipecat copy and missed cartesia.

`unique_search` calls `_field(r, ...)` on every search result. For a plain
string that finds no attribute, falls back to `""`, and the entry is dropped
before it reaches the prompt. `format_memories_to_text` in the same file has
an `isinstance(item, str)` branch for search results, so that branch is
currently unreachable.

`get_last_user_message` indexed `msg["role"]` and `msg["content"]`
positionally, so a tool-call turn or provider event missing either key raised
KeyError instead of being skipped, and non-string content was returned
verbatim despite the `str | None` annotation.

Both hunks now match the pipecat implementation.
@Agnik47

Agnik47 commented Sep 20, 2026

Copy link
Copy Markdown
Contributor Author

Two things a reviewer would reasonably want to know, both checked:

Do the other SDKs have the same two bugs? No. agent-framework-python and openai-sdk-python already handle both cases — plain-string search results survive dedup, and their get_last_user_message uses .get() with an isinstance guard. cartesia was the only outlier, so this PR closes the class rather than one instance of it.

Is the rest of the file in sync with pipecat now? Yes. The two utils.py copies are behaviourally identical after this change — I differential-tested them across 288 dedup/format input combinations (string / dict / model / empty / unicode / [recent] and [YYYY-MM-DD] prefixed / tag-injection payloads), 8 message shapes, and 4 timestamps: 0 differences.

One thing that audit did turn up, which is not fixed here: both voice SDKs diverge from the TypeScript getMemoryText on a hit shaped {"memory": "", "chunk": "..."} — _field stops at the first non-None value, so it resolves to "" and never falls through to chunk. That one affects cartesia and pipecat equally, so it is out of scope for a cartesia-parity PR. Filed separately as #1693, stacked on this branch.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant