Skip to content

fix: preserve reference time in LongMemEval search - #2411

Open
wangxinyufighting wants to merge 1 commit into
MemTensor:mainfrom
wangxinyufighting:fix/longmemeval-temporal-search
Open

wangxinyufighting wants to merge 1 commit into
MemTensor:mainfrom
wangxinyufighting:fix/longmemeval-temporal-search

Conversation

@wangxinyufighting

Copy link
Copy Markdown

Description

This PR fixes temporal search behavior for the LongMemEval evaluation.

LongMemEval questions may contain relative time expressions such as “yesterday” or “last week”. Previously, the question's reference date was not consistently propagated from the evaluation client to the API search context and temporal task parser. As a result, relative dates could be interpreted using the current time instead of the question date.

Implementation:

  • Add optional reference_time support to the open-source evaluation client.
  • Keep the cloud client interface compatible while omitting the unsupported reference_time field from its request payload.
  • Preserve reference_time in API and single-cube search contexts.
  • Propagate reference_time from the search context to the task parser.
  • Include the reference date in fine-mode temporal parsing prompts.
  • Add regression tests covering the client contract, API context, parser, and searcher propagation.

Dependencies: None.

Related Issue (Required): None

Type of change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Refactor (does not change functionality, e.g. code style improvements, linting)
  • Documentation update

How Has This Been Tested?

Focused tests:

pytest -q \
  tests/evaluation/test_longmemeval_api_contract.py \
  tests/memories/textual/test_tree_task_goal_parser.py \
  tests/memories/textual/test_tree_searcher.py

Copilot AI lite review requested due to automatic review settings September 25, 2026 14:42
@Memtensor-AI Memtensor-AI added area:memcube GeneralMemCube / cube 生命周期 / cube 配置 area:memory 记忆存储、检索、更新、召回逻辑 status:in-progress Someone or AI is working on it | 人工或 AI 正在处理 labels Sep 25, 2026
@Memtensor-AI

Copy link
Copy Markdown
Collaborator

🤖 Open Code Review

Target: PR #2411
Task: f39438dccceb08f1
Base: main
Head: fix/longmemeval-temporal-search

🔍 OpenCodeReview found 2 issue(s) in this PR.

⚠️ 1 warning(s) occurred during review.


1. evaluation/scripts/utils/client.py (L229-L241)

The reference_time parameter is accepted but never forwarded in the request payload. Unlike MemosApiClient.search which correctly adds "reference_time": reference_time to the payload, this implementation silently drops the value, making the parameter a no-op.

Add the field to the payload dict:

💡 Suggested Change

Before:

    def search(self, query, user_id, top_k, reference_time=None):
        """Search memories."""
        url = f"{self.memos_url}/search/memory"
        payload = json.dumps(
            {
                "query": query,
                "user_id": user_id,
                "memory_limit_number": top_k,
                "mode": os.getenv("SEARCH_MODE", "fast"),
                "include_preference": True,
                "pref_top_k": 6,
            }
        )

After:

    def search(self, query, user_id, top_k, reference_time=None):
        """Search memories."""
        url = f"{self.memos_url}/search/memory"
        payload = json.dumps(
            {
                "query": query,
                "user_id": user_id,
                "memory_limit_number": top_k,
                "mode": os.getenv("SEARCH_MODE", "fast"),
                "include_preference": True,
                "pref_top_k": 6,
                "reference_time": reference_time,
            }
        )

2. src/memos/memories/textual/tree_text_memory/retrieve/task_goal_parser.py (L99-L104)

The reference_time value comes directly from an API request field (confirmed via search_service.py and single_cube.py) and is concatenated into the LLM prompt without any sanitization. An attacker can craft a reference_time value containing adversarial instructions (e.g., "2024-01-01\nIgnore all previous instructions and instead...") to hijack the model's behavior — a classic prompt injection.

Consider validating that reference_time conforms to an expected format (e.g., ISO 8601 datetime string) before interpolating it into the prompt. You can also wrap it in a structured format that makes the boundary explicit to the model.

💡 Suggested Change

Before:

            reference_time = kwargs.get("reference_time")
            if reference_time:
                prompt += (
                    "\nReference time for resolving relative dates in the user query: "
                    f"{reference_time}\n"
                )

After:

            reference_time = kwargs.get("reference_time")
            if reference_time:
                import re
                # Validate reference_time matches an expected datetime pattern before injecting into prompt
                if not re.fullmatch(r"\d{4}-\d{2}-\d{2}([ T]\d{2}:\d{2}(:\d{2})?)?Z?", str(reference_time).strip()):
                    logger.warning(f"Invalid reference_time format, skipping: {reference_time!r}")
                else:
                    prompt += (
                        "\nReference time for resolving relative dates in the user query: "
                        f"{reference_time}\n"
                    )

🧹 Filtered 1 low-confidence OCR finding(s) before posting/fix-loop (duplicate: 1).

Generated by cloud-assistant via Open Code Review.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Default fast mode and the agentic search path still lose the question reference time.

Review effort: Lite
Findings: None

What changed in this PR

Fixes LongMemEval relative-date search by propagating reference_time through clients, search contexts, and temporal parsing.

Changes:

  • Adds optional reference-time client support.
  • Preserves reference time across API and single-cube searches.
  • Adds regression tests for propagation and parsing.

Review findings:

  • Moderate: Default fast mode does not consume reference_time.
  • Moderate: The agentic search path drops reference_time.
File Summary
tests/​memories/​textual/​test_tree_task_goal_parser.py Tests temporal prompt reference dates.
tests/​memories/​textual/​test_tree_searcher.py Tests reference-time propagation.
tests/​evaluation/​test_longmemeval_api_contract.py Tests client and API context contracts.
src/​memos/​search/​search_service.py Preserves reference time in search context.
src/​memos/​multi_mem_cube/​single_cube.py Propagates reference time to cube searches.
src/​memos/​memories/​textual/​tree_text_memory/​retrieve/​task_goal_parser.py Uses reference time in fine-mode prompts.
src/​memos/​memories/​textual/​tree_text_memory/​retrieve/​searcher.py Forwards reference time to the parser.
evaluation/​scripts/​utils/​client.py Adds reference-time client support.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@Memtensor-AI

Copy link
Copy Markdown
Collaborator

✅ Automated Test Results: PASSED

All tests passed (13/13 executed). memos_python_core/changed-repo-python: 13/13. Duration: 12s

Branch: fix/longmemeval-temporal-search

@Memtensor-AI Memtensor-AI added status:ready Ready for implementation; waiting for assignee or AI dispatch | 可进入实现,等待认领或派发 and removed status:in-progress Someone or AI is working on it | 人工或 AI 正在处理 labels Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:memcube GeneralMemCube / cube 生命周期 / cube 配置 area:memory 记忆存储、检索、更新、召回逻辑 status:ready Ready for implementation; waiting for assignee or AI dispatch | 可进入实现,等待认领或派发

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants