Skip to content

fix(cartesia,pipecat): fall through to chunk when memory is empty - #1693

Open
Agnik47 wants to merge 2 commits into
supermemoryai:mainfrom
Agnik47:fix/voice-sdk-chunk-fallback
Open

Agnik47 wants to merge 2 commits into
supermemoryai:mainfrom
Agnik47:fix/voice-sdk-chunk-fallback

Conversation

@Agnik47

@Agnik47 Agnik47 commented Sep 20, 2026

Copy link
Copy Markdown
Contributor

Stacked on #1683. This branch has that PR's commit as its parent, so the diff below shows both. Only the second commit belongs to this PR: 3d19805. Happy to rebase onto main once #1683 lands.

Summary

_field() returns the first value that is not None. A v4 hybrid-search hit shaped {"memory": "", "chunk": "..."} therefore resolves to the empty string and never falls through to chunk — the exact fallback the call site's own comment says it is there for:

# v4 search.memories/hybrid uses `memory` or `chunk`.
memory = _field(r, "memory", "chunk", "content", default="")

It bites on both paths:

  • deduplicate_memories — the empty text produces an empty comparison key, so the hit is dropped and the memory never reaches the prompt.
  • format_memories_to_text — when such an item does survive, the same _field call renders it as a bare - [16 hrs ago] .

Why this is the odd one out

The TypeScript getMemoryText in packages/tools/src/tools-shared.ts takes the first non-empty field. I differential-tested all four Python SDKs against it:

input TS (reference) agent-framework openai cartesia pipecat
{"memory": "", "chunk": "text"} ['text'] ['text'] ['text'] [] []
{"memory": " ", "chunk": "text"} ['text'] ['text'] ['text'] [] []
SimpleNamespace(memory="", chunk="text") ['text'] ['text'] ['text'] [] []
{"memory": None, "chunk": "text"} ['text'] ['text'] ['text'] ['text'] ['text']
{"chunk": "text"} ['text'] ['text'] ['text'] ['text'] ['text']
{"memory": "mem", "chunk": "chk"} ['mem'] ['mem'] ['mem'] ['mem'] ['mem']

agent-framework-python and openai-sdk-python already match the reference — they filter on isinstance(value, str) and value.strip(). The two voice SDKs were the only implementations that disagreed, which is why #1683 (a cartesia↔pipecat parity fix) did not surface it.

Changes

  • _memory_text() in both voice SDKs, mirroring the TypeScript getMemoryText semantics, used on the dedup and the render path. Keeping both on one helper is what stops a rescued hit from rendering as an empty bullet.
  • packages/pipecat-sdk-python/tests/conftest.py — pipecat's import stubs lived inside test_empty_profile.py, so a second test module could only import the package when collection happened to run that file first. Moving them to conftest.py makes tests/test_utils.py importable on its own. test_empty_profile.py is untouched; its own guarded copy becomes a no-op.

Test plan

  • 4 new tests per package fail without the source change and pass with it:
    FAILED tests/test_utils.py::TestSearchResultExtraction::test_falls_through_to_chunk_when_memory_is_empty
    FAILED tests/test_utils.py::TestSearchResultExtraction::test_falls_through_to_chunk_when_memory_is_whitespace
    FAILED tests/test_utils.py::TestSearchResultExtraction::test_falls_through_on_models_too
    FAILED tests/test_utils.py::TestSearchResultExtraction::test_never_renders_an_empty_bullet
    
  • With the fix: cartesia 14 passed, pipecat 8 passed. Pipecat verified isolated and in reverse collection order.
  • Differential harness re-run after the change: divergences from the TypeScript reference 0, cartesia↔pipecat behavioural differences 0 across 288 dedup/format input combinations.
  • test_memory_still_wins_over_chunk and test_entry_with_no_usable_text_is_still_dropped pin the behaviour that should not change.

…ssages

`utils.py` in cartesia-sdk-python and pipecat-sdk-python are copies of the
same file. supermemoryai#1434 hardened two spots in the pipecat copy and missed cartesia.

`unique_search` calls `_field(r, ...)` on every search result. For a plain
string that finds no attribute, falls back to `""`, and the entry is dropped
before it reaches the prompt. `format_memories_to_text` in the same file has
an `isinstance(item, str)` branch for search results, so that branch is
currently unreachable.

`get_last_user_message` indexed `msg["role"]` and `msg["content"]`
positionally, so a tool-call turn or provider event missing either key raised
KeyError instead of being skipped, and non-string content was returned
verbatim despite the `str | None` annotation.

Both hunks now match the pipecat implementation.
`_field` returns the first value that is not None. A v4 hybrid-search hit
shaped `{"memory": "", "chunk": "..."}` therefore resolves to the empty
string and never falls through to `chunk` — the exact fallback the call
site's comment says it is there for.

The consequences are in both directions. In `deduplicate_memories` the hit
produces an empty comparison key and is dropped, so the memory never reaches
the prompt. When such an item does survive, `format_memories_to_text` renders
it through the same `_field` call and emits a bare `- [16 hrs ago] `.

The TypeScript `getMemoryText` takes the first non-empty field instead, and
`agent-framework-python` and `openai-sdk-python` both already match it; the
two voice SDKs were the only implementations that disagreed. Add a
`_memory_text` helper mirroring the TypeScript semantics and use it on both
the dedup and the render path.

Also add tests/conftest.py to pipecat so the import stubs are installed
regardless of collection order — test_utils.py cannot import the package
on its own otherwise.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant