Skip to content

bug(memos-local-plugin): L3 world-model creation always fails with "This operation was aborted" — background LLM call aborted by an inherited signal #2412

Description

@DCAMAR

Pre-submission checklist | 提交前检查

  • I have searched existing issues and this hasn't been mentioned before | 我已搜索现有问题,确认此问题尚未被提及
  • I have read the project documentation and confirmed this issue doesn't already exist | 我已阅读项目文档并确认此问题尚未存在
  • This issue is specific to MemOS and not a general software issue | 该问题是针对 MemOS 的,而不是一般软件问题

Bug Description | 问题描述

On my deployment (@memtensor/memos-local-plugin@2.0.19, DeepSeek Harness adapter, Windows), every L3 "new world model" run fails — i.e. exactly the runs that actually need an LLM call:

l3.failed / stage=abstract / error.code=llm_failed
error.message = "openai_compatible request failed: This operation was aborted"
clusterKey    = "git|git-cli"

Meanwhile the merge path (which needs no LLM call) works perfectly, so L3 itself is alive:

run.start = 4   run.done = 4   l3.merge = 31   (successes)
run.done ... clusters=8 created=0 merged=8 skipped=0

中文简述:L3 里凡是真正要调 LLM 的新建簇,100% 失败,错误是 This operation was aborted(HTTP 请求被中止);而不需要 LLM 的合并路径完全正常(本次启动 31 次合并全成功)。

这不是 #2335 / #2374:那两个报的是 'title' must be a non-empty string(LLM 有返回但字段不合法)。我这条更早——请求本身被 abort,LLM 根本没返回。(而且我把 L3 的 maxTokens 提到 16384 之后,title / must be an array 那一族已经消失,剩下的就是这个 abort。)

How to Reproduce | 如何重现

Environment-agnostic: run the plugin inside an agent host, let it work in the background after a turn, and wait for an l2.policy.induced event whose cluster is not mergeable (i.e. requires a new world model).

Observed sequence (host log, 2026-09-21 03:45 UTC):

[core.memory.l3] run.start trigger="l2.policy.induced" episodeId="..."
[core.memory.l3.cluster] clusters.built eligiblePolicies=113 clusters=8
[core.pipeline.bootstrap] l3Llm.llm_error provider="openai_compatible" model="glm-4.5-air"
   message="llm_unavailable: openai_compatible request failed: This operation was aborted"
[core.memory.l3.abstract] abstract.llm_failed clusterKey="git|git-cli"
   err="openai_compatible request failed: This operation was aborted"
[core.memory.l3] run.done clusters=8 created=0 merged=7 skipped=1

What I ruled out (with evidence)

  1. Not the endpoint or the key — I called the same endpoint (open.bigmodel.cn, same key, same glm-4.5-air) directly with curl/fetch: HTTP 200 and valid JSON in ~12–30 s.
  2. Not a timeout — dist/core/llm/fetcher.js:42 does mergeSignals(opts.signal, AbortSignal.timeout(opts.timeoutMs)). An expiry of AbortSignal.timeout surfaces as TimeoutError; what we see is AbortError: This operation was aborted. So the aborting signal is the caller-provided one (dist/core/llm/client.js:254 threads opts?.signal down to the provider). My guess is that post-turn work inherits a signal that is already aborted by the time the background run executes — and every retry inherits it again, hence a permanent 0% success rate.
  3. Not the DeepSeek Harness host-LLM bridge — I explicitly disabled it (hostLlmEnabled: false) and re-tested: the abort is unchanged. (Also, that bridge can never serve background work anyway: dist/adapters/deepseek-harness/host-llm.js:58-62 requires an AsyncLocalStorage route that only exists inside a turn, and its own budget is 2000 ms.)
  4. Not maxTokens — already raised to 16384 (see below).

Environment | 环境信息

  • Operating System: Windows 10.0.26200 (x64)
  • Host: DeepSeek Harness 0.1.6-alpha.1 (dsh web), Windows
  • Plugin: @memtensor/memos-local-plugin@2.0.19 — DeepSeek Harness adapter (dist/adapters/deepseek-harness/index.js)
  • LLM: provider: openai_compatible → https://open.bigmodel.cn/api/paas/v4/chat/completions, model glm-4.5-air
  • L3 slot (l3Llm): same model, maxTokens: 16384, timeoutMs: 180000, temperature: 0
  • Storage: local sqlite (3k+ traces, 200+ policies, 100 skills, 6 world models)

Additional Context | 其他信息

1. What I fixed locally, in case it is worth upstreaming: the L3 maxTokens budget

l3Llm is empty by default, so core/pipeline/memory-core.js:307 (l3Llm ?? llm) makes L3 silently inherit the shared llm.maxTokens — which is 1024 in the default config. With a real L3 stress prompt I measured:

max_tokens result (same prompt, stream, temperature 0.15)
1024 content = 0 bytes (the model burns 1024 completion tokens and returns nothing)
4096 truncated JSON → 'constraints' must be an array / 'inference' must be an array
16384 valid JSON, all three facets (ℰ/ℐ/𝒞) present

This matches #1896 (add maxTokens to LlmSchema). After I set l3Llm.maxTokens: 16384, the whole title / must be an array failure family disappeared in my deployment — so I suspect #2335 / #2374 are at least partly a budget problem rather than only an algorithm-layer one.

2. Questions

  • For post-turn / background L3 work, how is the LLM call supposed to obtain a signal and a lifecycle budget? Right now it seems to inherit a signal that is already aborted, and each retry inherits it again → permanent failure with no way to recover.
  • Should l3Llm.maxTokens default to something non-empty (e.g. the 4096 that dist/core/config/defaults.js:60-76 already documents) instead of falling back to the shared 1024?

2026-09-25 Re-confirmation on the same deployment | 同机复现确认(4 天后)

Same host, same plugin version, same L3 slot config — the abort is still there, and it is wider than L3:

  • L3 create path still produces nothing: 2026-09-25 16:03 [core.memory.l3] run.done trigger="l2.policy.induced" clusters=8 created=0 merged=8 skipped=0(新建 0、合并 8 成功 —— 与报告主症状一致)。
  • world_model 表最新一条停在 2026-09-18 —— 此后 7 天零新建(合并路径照常工作)。
  • 同一 abort 现在也打在检索链上(本报告原本只归因到 L3,这里更正/加强):同一份 host 日志里 This operation was aborted 共 38 次,最近两次为
    WARN [core.retrieval] llm_filter.failed err="openai_compatible request failed: This operation was aborted" candidateCount=10(16:41 / 16:53)。
    → 也就是说凡是"回合外/后台"发起的 openai_compatible 调用都可能拿到一个已 abort 的 signal,不止 L3 抽象一处。
  • 排除项复核:宿主 LLM 桥保持关闭(hostLlmEnabled: false),abort 依旧 —— 与报告"排除 3"一致;l3Llm.maxTokens = 16384 也保持,title/must be an array 那一族确实没再出现(报告"排除 4"成立)。

(新增段落由同一部署在 2026-09-25 复核,仅追加证据,未改动原结论。)

Willingness to Implement | 实现意愿

  • I'm willing to implement this myself | 我愿意自己解决
  • I would like someone else to implement this | 我希望其他人来解决

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

ai:taskDispatched to AI coding agent | 已派发给 AI 编码任务ai:testingAI agent is running tests | AI 正在运行测试area:pluginOpenClaw & Hermesstatus:in-progressSomeone or AI is working on it | 人工或 AI 正在处理types:bugSomething isn't working | 功能异常

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions