From be7aac382039776e283336b11d447464e35c128a Mon Sep 17 00:00:00 2001 From: Thiago Rotta Date: Tue, 4 Aug 2026 16:16:43 -0500 Subject: [PATCH] docs(memory): expand memory chapter with LTM, patterns and pipelines Evolve the memory section from an intro stub plus a storage-focused STM page into a complete chapter, reusing the existing STM content. - Memory.md: memory types (STM, LTM, working memory), cognitive sub-types (semantic, episodic, procedural), memory vs. knowledge sources, key principles and chapter map. - Memory-Architecture-Patterns.md (new): storage strategies, retrieval patterns, vendor-neutral pattern archetypes, memory scoping, trade-offs including the fading memory problem, token economics and metrics. - Short-Term-Memory.md: keep existing design approaches and retention, add context window management, working memory assembly, session end handling and multi-channel considerations. - Long-Term-Memory.md (new): what to persist, memory entry data model, lifecycle, relevance scoring and decay with tier demotion, scope decision framework, governance, security risks, operations at scale and a phased adoption roadmap. - Memory-Pipelines.md (new): mermaid diagrams for the retrieval, session end and STM to LTM promotion pipelines. - SUMMARY.md: add the new pages and fill the empty long-term memory entry. - References.md: add the Mem0 long-term memory benchmark paper. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --- SUMMARY.md | 4 +- docs/References.md | 3 +- docs/memory/Long-Term-Memory.md | 506 ++++++++++++++++++++ docs/memory/Memory-Architecture-Patterns.md | 306 ++++++++++++ docs/memory/Memory-Pipelines.md | 227 +++++++++ docs/memory/Memory.md | 157 +++++- docs/memory/Short-Term-Memory.md | 168 ++++++- 7 files changed, 1361 insertions(+), 10 deletions(-) create mode 100644 docs/memory/Long-Term-Memory.md create mode 100644 docs/memory/Memory-Architecture-Patterns.md create mode 100644 docs/memory/Memory-Pipelines.md diff --git a/SUMMARY.md b/SUMMARY.md index 72e735a..e5f113b 100644 --- a/SUMMARY.md +++ b/SUMMARY.md @@ -12,8 +12,10 @@ - [Message-driven](./docs/agents-communication/Message-Driven.md) - [Data exchange protocols]() - [Memory](./docs/memory/Memory.md) + - [Memory architecture patterns](./docs/memory/Memory-Architecture-Patterns.md) - [Short-term memory](./docs/memory/Short-Term-Memory.md) - - [Long-term memory]() + - [Long-term memory](./docs/memory/Long-Term-Memory.md) + - [Memory pipelines](./docs/memory/Memory-Pipelines.md) - [Observability](./docs/observability/Observability.md) - [Evaluation](./docs/evaluation/Evaluation.md) - [Tool Call Evaluation](./docs/evaluation/ToolCall.md) diff --git a/docs/References.md b/docs/References.md index 5ae5e54..73f25c5 100644 --- a/docs/References.md +++ b/docs/References.md @@ -25,6 +25,7 @@ _Last updated: 2025-06-30_ ## Memory - [How Microsoft Copilot scales to millions of users with Azure Cosmos DB](https://devblogs.microsoft.com/cosmosdb/how-microsoft-copilot-scales-to-millions-of-users-with-azure-cosmos-db/) +- [Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory](https://arxiv.org/abs/2504.19413) ## Observability @@ -88,4 +89,4 @@ _Last updated: 2025-06-30_ --- -{{ #include ../components/discuss-button.hbs }} \ No newline at end of file +{{ #include ../components/discuss-button.hbs }} diff --git a/docs/memory/Long-Term-Memory.md b/docs/memory/Long-Term-Memory.md new file mode 100644 index 0000000..3ff804f --- /dev/null +++ b/docs/memory/Long-Term-Memory.md @@ -0,0 +1,506 @@ + + +# Long-term Memory + +_Last updated: 2026-08-04_ + +Long-Term Memory (LTM) is what allows an agent to recognize a returning user, +recall a decision made three weeks ago, and avoid asking the same question +twice. Unlike [short-term memory](./Short-Term-Memory.md), which holds the raw +session and disappears with it, LTM holds a **compressed, distilled +representation** of what mattered, persisted across sessions, channels, and +agents. + +LTM is not a transcript archive and it is not a knowledge base. It is a curated +set of durable statements about a subject — a customer, an employee, a project — +that the system actively recalls to make future interactions better. + +This topic covers: + +- [What qualifies as long-term memory](#what-qualifies-as-long-term-memory) +- [Long-term memory stores](#long-term-memory-stores) +- [Memory entry data model](#memory-entry-data-model) +- [Memory lifecycle](#memory-lifecycle) +- [Relevance scoring and decay](#relevance-scoring-and-decay) +- [Memory scope decision framework](#memory-scope-decision-framework) +- [Safety, privacy and governance](#safety-privacy-and-governance) +- [Security risks](#security-risks) +- [Operating memory at scale](#operating-memory-at-scale) +- [Adoption roadmap](#adoption-roadmap) + +--- + +## What Qualifies as Long-Term Memory + +### What to remember + +- **Preferences and working style:** tone, format, language, preferred tools and + channels. +- **Durable attributes:** role, team, account tier, recurring projects or + topics. +- **Decisions and commitments:** what was agreed, what was promised, what was + ruled out and why. +- **Entities and relationships:** systems, products, tickets, and people that + keep reappearing. +- **Outcomes and resolution patterns:** what actually solved a recurring + problem. + +### What not to remember + +- **Every conversation detail.** Raw history belongs in STM and in the archive, + not in LTM. +- **Sensitive facts that were not explicitly offered for retention.** Health, + financial, and personal details require explicit intent. +- **Business or transactional records** that already live in a system of record. + Retrieve them; do not duplicate them into memory, where they will go stale. +- **Credentials, tokens, and secrets.** Never, under any policy. + +### When to write + +Two signals justify creating a memory: + +- **Explicit intent** — the user asked the system to remember something + ("remember that I prefer bullet points"). These memories carry the highest + importance. +- **Repeated implicit signal** — the same preference or fact appears across + multiple sessions with consistent phrasing and no contradiction. + +A single incidental mention is usually noise. Requiring either explicit intent +or repetition is the cheapest defense against memory pollution. + +--- + +## Long-Term Memory Stores + +Each memory sub-type has a different shape, and therefore a different natural +storage strategy. See +[storage strategies](./Memory-Architecture-Patterns.md#storage-strategies) for +the full comparison. + +| Sub-type | What is stored | Fitting storage | +| -------------- | --------------------------------------------------- | ----------------------------------------------------------------- | +| **Semantic** | Structured profile of durable facts and preferences | Document or relational store, small and directly injectable | +| **Episodic** | Session summaries and events, timestamped | Vector store with metadata filters for semantic recall | +| **Procedural** | Learned workflows and resolution patterns | Structured records, optionally a graph when steps relate entities | + +Practical guidance: + +- **Keep semantic memory small enough to inject wholesale.** If the profile no + longer fits comfortably in the prompt, it needs consolidation, not a bigger + budget. +- **Keep episodic memory searchable, not injected.** It grows without bound and + should be reached through on-demand retrieval. +- **Introduce graph storage only when questions require traversal.** Entity + relationships are a Phase 3 concern for most systems. + +--- + +## Memory Entry Data Model + +Every memory entry should carry enough metadata to be retrieved, ranked, +governed, decayed, audited, and deleted. + +| Field | Purpose | +| ------------------- | ------------------------------------------------------------------- | +| `memory_id` | Unique identifier of the entry | +| `subject_id` | Who or what the memory is about (user, customer, project, team) | +| `scope` | Visibility boundary: session, global, project, channel, role, org | +| `type` | `semantic`, `episodic`, or `procedural` | +| `content` | The statement itself, in a compact and self-contained form | +| `embedding` | Vector representation used for semantic retrieval | +| `source_session_id` | Provenance: the session the memory was extracted from | +| `source_type` | How it was created: explicit user request, extraction, or import | +| `confidence` | How certain the extraction is that the statement is true | +| `importance` | How much this memory matters, independent of how often it is used | +| `retrieval_count` | How many times the entry has been retrieved into working memory | +| `last_retrieved_at` | Timestamp of the most recent retrieval — the input to decay | +| `created_at` | When the memory was first written | +| `updated_at` | When it was last modified or reinforced | +| `version` | Monotonic version, enabling change tracking and rollback | +| `tier` | Lifecycle tier: `hot`, `warm`, `cold`, or `archived` | +| `sensitivity` | Classification: public, internal, confidential, restricted | +| `expires_at` | Hard expiration for policy-bound or short-lived memories | +| `pinned` | Whether the memory is exempt from decay (user-pinned or compliance) | + +> `last_retrieved_at` and `retrieval_count` are not optional bookkeeping. They +> are the fields that make decay and relevance scoring possible, and they must +> be written back by the retrieval pipeline. + +--- + +## Memory Lifecycle + +A memory is not written once and kept forever. It moves through a lifecycle, and +each stage needs an owner in the architecture. + +```mermaid +flowchart LR + Extraction[Extraction
what to capture] --> Consolidation[Consolidation
merge and resolve conflicts] + Consolidation --> Reinforcement[Reinforcement
strengthen what is used] + Reinforcement --> Decay[Decay
fade what is not] + Decay --> Deletion[Deletion
policy or user driven] + Consolidation -.-> Versioning[Versioning
track changes, allow rollback] + Reinforcement -.-> Versioning +``` + +- **Extraction:** decide what in a session is durable. Runs asynchronously after + a session ends, never on the inference critical path. See the + [session end pipeline](./Memory-Pipelines.md#stm-session-end-pipeline). +- **Consolidation:** merge duplicates and resolve contradictions. A newer, + higher-confidence statement should supersede an older one rather than coexist + with it — two contradictory memories are worse than none. +- **Reinforcement:** every retrieval that proves useful increases confidence and + resets the decay clock. +- **Decay:** memories that are never retrieved lose relevance over time. See + [relevance scoring and decay](#relevance-scoring-and-decay). +- **Versioning:** keep the previous value when a memory changes. This supports + rollback, explains behavior changes, and is often a compliance requirement. +- **Deletion:** user-initiated or policy-driven, and it must be genuinely + effective — including in the vector index, the archive, and any derived + summaries. + +--- + +## Relevance Scoring and Decay + +The most valuable memory is the one that keeps surfacing when the agent looks +for context. That behavior can be measured, and it should be, with an explicit +**score** attached to every entry. + +### Scoring + +In practice, teams start by combining **how often a memory has been retrieved** +with **how recently it was last accessed**, and then compose additional +variables into the same score. The most important addition is **explicit +importance**: a memory created because the user said _"I want you to remember +that I have three children"_ deserves far more weight than one incidentally +extracted from a transcript. + +A workable starting formula: + +```text +score = w_f · normalize(retrieval_count) + + w_r · recency(last_retrieved_at) + + w_i · importance + + w_c · confidence +``` + +Where: + +- `normalize(retrieval_count)` dampens raw counts (for example `log(1 + count)`) + so a handful of very old, frequently used memories cannot dominate forever. +- `recency(last_retrieved_at)` is an exponential decay function of the time + since the last access, with a half-life tuned per memory type — days for + volatile operational context, months for stable profile facts. +- `importance` is highest for explicitly requested memories, moderate for + repeated implicit signals, and lowest for single-mention extractions. +- `confidence` reflects the extraction and validation quality. + +Tuning guidance: + +- **Start with frequency and recency only**, and add importance and confidence + once the pipeline is producing enough data to calibrate them. +- **Weight importance high enough that explicit user memories never decay out** + of the store through simple disuse. +- **Score at write time and refresh on access**, storing the score so ranking + does not require recomputation across the whole store. + +The score is used in two places: to rank candidates during retrieval (see the +[reranker stage](./Memory-Pipelines.md#memory-retrieval-pipeline)) and to decide +which memories survive. + +### Decay: use it or lose it + +Decay follows a simple rule: **memories that nobody accesses fade over time, and +every access resets the counter.** This is precisely why `last_retrieved_at` and +`retrieval_count` must be updated by the retrieval pipeline — without that +write-back, there is no signal to decay against, and the store only ever grows. + +Implementation notes: + +- **Write back asynchronously.** Updating access metadata must not add latency + to the read path; batch the updates or emit them as events. +- **Count real usage, not candidate generation.** Ideally, increment when a + memory actually enters the assembled context, not merely when it appears in a + candidate list. +- **Access frequency is the primary survival criterion.** When in doubt about + whether a memory still matters, how often it is retrieved is the strongest + available signal. + +### Never delete directly — demote across tiers + +What works best in practice is to **never hard-delete a memory as a decay +outcome**. Instead, demote it through tiers as its score drops, and let a batch +job walk the store, identify facts that are no longer relevant, and move them +down — only removing entries at the end of that journey. + +```mermaid +stateDiagram-v2 + [*] --> Hot: created / promoted + Hot --> Warm: score below hot threshold + Warm --> Cold: prolonged non-retrieval + Cold --> Archived: batch curation marks as irrelevant + Archived --> [*]: policy-driven purge + + Warm --> Hot: retrieved and reinforced + Cold --> Warm: retrieved and reinforced + Archived --> Cold: explicitly restored +``` + +| Tier | Retrieval behavior | Typical treatment | +| ------------ | ------------------------------------------------------ | -------------------------------------- | +| **Hot** | Always eligible; semantic profile may be auto-injected | Indexed, low-latency store | +| **Warm** | Retrieved on demand only | Indexed, lower ranking priority | +| **Cold** | Retrieved only with explicit, targeted queries | Cheaper storage, optionally de-indexed | +| **Archived** | Not retrieved; retained for audit and restoration | Cold storage, purged by policy | + +Key properties of this model: + +- **Demotion is reversible.** A retrieval promotes a memory back up a tier, + which is exactly the reinforcement behavior you want: rarely used but still + relevant facts recover instead of disappearing. +- **The batch job is a curation job.** It scans memories, evaluates scores, + detects facts contradicted by newer entries, and demotes or expires them. It + is also the natural place to run consolidation. +- **Deletion becomes deliberate.** Hard deletion is reserved for user-initiated + removal, policy-driven purges, and clear violations (secrets, restricted + data), not for ordinary aging. + +### Exceptions to decay + +- **User-pinned memories** — explicitly requested facts stay until the user + removes them. +- **Compliance-mandated retention** — governed by policy, not by usage. +- **Legal hold** — decay and purge are suspended entirely. + +--- + +## Memory Scope Decision Framework + +Scope is the single most consequential decision in a memory design, and it is a +function of three dimensions: + +```text +scope = f(use_case, duration, sensitivity) +``` + +### Dimension 1: use case type + +| | Customer-facing agent | Internal employee agent | +| --------------- | ----------------------------------------- | ------------------------------------ | +| **Scope** | Per-customer and per-channel | Per-employee and per-project or team | +| **Retention** | Governed by data policy (GDPR, LGPD) | Aligned with HR and IT policies | +| **Content** | Service history, preferences, open issues | Workflows, tools, past resolutions | +| **Sensitivity** | High — PII and financial data | Moderate — internal procedures | + +### Dimension 2: memory duration + +- **Ephemeral:** single session only; appropriate for sensitive queries. +- **Short-lived:** hours to days; an active ticket or case. +- **Persistent:** weeks to months; the user profile. +- **Permanent:** compliance-mandated retention. + +### Dimension 3: sensitivity classification + +- **Public:** non-sensitive general preferences. +- **Internal:** business processes and workflows. +- **Confidential:** PII, financial, or health data — encrypted, tightly scoped, + minimal retention. +- **Restricted:** credentials and access tokens — **never stored**, at any tier. + +### Applying the framework + +```mermaid +flowchart TD + Start[Candidate memory] --> Sensitive{Restricted
data?} + Sensitive -- Yes --> Drop[Do not store] + Sensitive -- No --> Confidential{Confidential?} + Confidential -- Yes --> Narrow[Narrowest scope
+ shortest retention
+ encryption] + Confidential -- No --> UseCase{Use case} + UseCase -- Customer-facing --> CustScope[Scope: subject + channel
Retention: data policy] + UseCase -- Internal --> EmpScope[Scope: subject + project
Retention: HR/IT policy] + Narrow --> Duration{Expected
duration} + CustScope --> Duration + EmpScope --> Duration + Duration -- Ephemeral --> STMOnly[Keep in STM only] + Duration -- Short-lived --> Expiring[LTM with expires_at] + Duration -- Persistent --> Profile[LTM, decay enabled] + Duration -- Permanent --> Pinned[LTM, pinned, audited] +``` + +Decide the scope **at write time** and store it on the entry. Retrofitting scope +onto an existing memory store is expensive and rarely complete. + +--- + +## Safety, Privacy and Governance + +Memory changes the risk profile of an agentic system: data that used to vanish +with the session now persists, accumulates, and can resurface in unexpected +contexts. + +### User control + +- **Explicit opt-in or opt-out** for memory features. +- **Transparency:** users can see exactly what the agent remembers about them. +- **Editability:** users can add, correct, or delete individual memories. +- **Incognito or temporary mode** for sensitive interactions, where nothing is + written. +- **Clear ownership:** in enterprise deployments memory typically belongs to the + organization, not the individual — which also means a user who leaves the + organization loses those memories. In consumer contexts, the memory belongs to + the user. Whichever model applies, it must be explicit and communicated. + +### Data protection + +- **Encryption at rest and in transit** for all memory stores and indexes. +- **Regulatory compliance** (GDPR, LGPD, and sector-specific regimes) for any + memory holding customer data, including the right to erasure. +- **Retention and deletion policies** defined per scope and sensitivity, and + enforced by an automated job rather than by convention. +- **Never store credentials, tokens, or passwords**, and scan extraction output + for them before writing. +- **Audit trail** for every memory access and modification: who, when, which + entry, and why. +- **Memory must not feed model training.** + +### Governance model + +- **Admin controls** for enabling memory per team, per agent, and per channel. +- **Sensitive topic exclusion rules** that prevent extraction on defined + categories. +- **Regular memory hygiene reviews** — sampling stored memories to check + precision, scope correctness, and policy compliance. +- **Clear data ownership:** a named owner for the memory data of each scope. + +See [Security](../security/Security.md) and +[Governance](../governance/Governance.md) for the broader controls this plugs +into. + +--- + +## Security Risks + +| Risk | What happens | Mitigation | +| ------------------------------- | -------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------- | +| **Prompt injection via memory** | Malicious content in a conversation is stored and later re-injected as trusted context | Treat memory content as untrusted input; strip instruction-like content at extraction; validate on write | +| **Memory poisoning** | False facts are deliberately planted to alter future behavior | Require explicit intent or repetition; confidence thresholds; provenance on every entry; anomaly monitoring | +| **Context collapse** | Memories from one domain, customer, or channel surface in another | Hard scope boundaries enforced as query filters, not as ranking hints | +| **Hallucinated memories** | Over-compressed summaries introduce details that never occurred | Preserve key details verbatim; validate extractions against the source; keep provenance so entries can be re-verified | +| **Silent data retention** | Sensitive data persists past its policy window | `expires_at` enforcement, automated purge jobs, hygiene reviews | + +The unifying mitigation is **provenance**: every memory entry should be +traceable to the session, turn, and mechanism that produced it. Without +provenance, none of the above can be investigated after the fact. + +--- + +## Operating Memory at Scale + +### Infrastructure + +- **Separate conversation storage from memory processing.** The transactional + path that serves conversations and the analytical path that builds memory have + different scaling and availability profiles. +- **Extract asynchronously.** Memory extraction must never block inference; + drive it from session events (for example Azure Event Grid or Service Bus) + into background workers. +- **Rate limit memory operations** per subject and per agent to contain runaway + extraction cost. +- **Enforce memory size budgets** per user and per project, with consolidation + triggered when a budget is exceeded. +- **Use a vector store with metadata filtering** so scope and sensitivity + filters are applied inside the query, not after it. Azure examples: Azure AI + Search or Azure Cosmos DB vector search. +- **Log memory updates in a relational store.** A simple, queryable activity log + of memory writes, merges, demotions, and deletions is invaluable for auditing + and debugging. + +### Observability and quality + +- Track **retrieval precision and recall** for memory queries. +- Monitor **token cost per query** with and without memory. +- Measure the **latency impact** of memory retrieval on end-to-end response + time. +- Compare **user satisfaction with memory on versus off**. +- Watch for **fading memory**: a downward trend in retrieval precision as the + store grows is a signal to tighten scoring, decay, and consolidation. + +See [Observability](../observability/Observability.md) for instrumentation +practices. + +### Cross-channel operation + +- Maintain a **unified profile layer** shared across channels, so a subject is + recognized consistently. +- Keep **episodic memory channel-scoped** by default to prevent leakage. +- Allow **channel-specific policies** for retention and extraction, since an + asynchronous messaging conversation and an internal project session have very + different lifecycles. + +--- + +## Adoption Roadmap + +Memory capability is best delivered in phases, each independently valuable. + +### Phase 1 — Short-term memory (quick wins) + +**Objective:** session continuity within conversations. + +**Implementation:** + +- Maintain a conversation window with a sliding buffer. +- Inject the existing customer or employee profile at session start. +- Summarize older turns to extend the effective context. +- Leverage CRM and profile data that is already mapped but underused — this + requires no new infrastructure and delivers immediate personalization. + +**Use cases:** maintaining context within a support ticket, multi-step IT +troubleshooting, preserving context in asynchronous channels. + +**Expected outcome:** reduced repetition and faster resolution. + +### Phase 2 — Cross-session long-term memory (core value) + +**Objective:** remember the user across sessions. + +**Implementation:** + +- Build the fact extraction pipeline (preferences, decisions, history). +- Deploy a vector store for semantic retrieval. +- Create the structured user memory profile. +- Implement memory injection into the agent context with an explicit budget. +- Define memory scope per use case and per channel. + +**Use cases:** returning customer recognition with history recall, persistent +project context for employees, proactive suggestions grounded in past +interactions. + +**Expected outcome:** personalized, proactive experiences. + +### Phase 3 — Intelligent memory at scale + +**Objective:** adaptive, governed, enterprise-grade memory. + +**Implementation:** + +- Add a knowledge graph for entity relationships. +- Introduce importance-weighted storage and retrieval. +- Operate decay and lifecycle policies, including tier demotion. +- Enable cross-agent memory sharing where appropriate. +- Build analytics on memory patterns, coverage, and quality. + +**Use cases:** complex multi-agent workflows, insight mining across the user +base, predictive assistance based on behavioral patterns. + +**Expected outcome:** emergent knowledge graphs and insight mining on top of a +governed memory layer. + +> Sequence matters more than speed. Phase 1 is achievable with existing +> infrastructure, while a full knowledge graph is a substantial investment. +> Start simple and let measured impact justify each subsequent phase. + +--- + +{{ #include ../../components/discuss-button.hbs }} diff --git a/docs/memory/Memory-Architecture-Patterns.md b/docs/memory/Memory-Architecture-Patterns.md new file mode 100644 index 0000000..6d7a50e --- /dev/null +++ b/docs/memory/Memory-Architecture-Patterns.md @@ -0,0 +1,306 @@ +# Memory Architecture Patterns + +_Last updated: 2026-08-04_ + +Before choosing a database or a framework, a memory architecture requires a +small number of explicit decisions. This section covers those decisions, the +storage and retrieval options available for each, and the trade-offs that +production systems consistently run into. + +This topic covers: + +- [The four decisions](#the-four-decisions) +- [Storage strategies](#storage-strategies) +- [Retrieval patterns](#retrieval-patterns) +- [Pattern archetypes](#pattern-archetypes) +- [Memory scoping](#memory-scoping) +- [Trade-offs](#trade-offs) +- [Token economics](#token-economics) +- [Metrics to track](#metrics-to-track) + +--- + +## The Four Decisions + +Every memory design answers these four questions, explicitly or by accident: + +1. **What to remember.** Which signals from an interaction are durable enough to + outlive the session, and which are noise. +2. **Where to store it.** The storage strategy for each memory sub-type. +3. **How to retrieve it.** Whether memory is pushed into every prompt or pulled + on demand, and how candidates are ranked. +4. **Which boundaries share memory.** Which channels, agents, projects, and + tenants share a global long-term memory, and which must stay isolated. + +The last one is the most consequential and the hardest to reverse. See +[memory scoping](#memory-scoping) and the scope decision framework in +[Long-term memory](./Long-Term-Memory.md). + +--- + +## Storage Strategies + +| Strategy | Mechanics | Best for | Risk | +| ----------------------- | -------------------------------------------------------------------- | ------------------------------------------- | --------------------------------------------- | +| **Vector store** | Embed memory entries, retrieve by semantic similarity | Open-ended recall, fuzzy matching | No explicit structure or relationships | +| **Knowledge graph** | Model entities and explicit relationships between them | Multi-hop reasoning, entity tracking | Schema rigidity, higher setup and upkeep cost | +| **Hybrid vector+graph** | Vector for fuzzy recall, graph for relations, metadata for filtering | Production systems with entity-rich domains | Two systems to operate and keep consistent | +| **Simple persistence** | Files, key-value, or relational rows; often plain Markdown | Small scopes, transparency, fast start | Scalability limits, no semantic search | + +Notes on each: + +- **Vector store.** The default for episodic memory and session summaries. + Metadata filtering is as important as the vector itself: scope, owner, + sensitivity, and tier must be filterable at query time, not post-filtered. + Azure examples: Azure AI Search vector indexes, or the vector search + capabilities of Azure Cosmos DB co-located with the conversation documents. +- **Knowledge graph.** Worth adding when questions require traversing + relationships ("which tickets from this account touched the same subsystem"). + Usually introduced in a later phase, not at the start. +- **Hybrid.** The most successful production pattern: vectors handle recall, the + graph handles relations, and metadata tags drive filtering and ranking. +- **Simple persistence.** Underrated. A structured profile in a relational table + or a small Markdown file is transparent, cheap, auditable, and sufficient for + semantic memory in many use cases. + +> Storage strategy should be chosen per memory sub-type, not for memory as a +> whole. A very common and effective combination is a relational or document +> profile for semantic memory plus a vector index for episodic memory. + +--- + +## Retrieval Patterns + +| Pattern | How it works | Strengths | Weaknesses | +| ------------------------------- | -------------------------------------------------------------------------- | ---------------------------------------------------- | ------------------------------------------------- | +| **RAG over history** | Chunk conversation history, embed, retrieve top-k | Standard approach for episodic memory | Retrieval noise, chunking artifacts | +| **Summarization buffer** | Compress older turns into rolling summaries | Large token reduction, keeps continuity | Lossy; compression can produce false details | +| **Fact extraction + injection** | Extract durable facts, inject as a structured profile in the system prompt | Cheap, deterministic, always present | Lossy; unbounded growth if not curated | +| **On-demand search** | Agent calls a memory search tool only when it needs to | Lowest token overhead, higher precision, transparent | Misses context when the agent fails to trigger it | + +These patterns compose. A typical production configuration injects an extracted +profile on every turn, keeps a summarization buffer for the active session, and +exposes an on-demand search tool for episodic recall. + +--- + +## Pattern Archetypes + +Publicly documented assistant and agent memory implementations converge on a +small set of archetypes. They are described here as neutral patterns, since the +underlying mechanics — not the products — are what transfers to your +architecture. + +### Auto-Injected Layered Memory + +Memory is always present. A set of layers is assembled into every request +without the agent asking for it: session metadata, explicitly saved facts, +lightweight summaries of recent conversations, and the current conversation +window. There is no retrieval step over raw history; summaries are pre-computed +asynchronously. + +- **Pros:** seamless continuity, zero user effort, no dependence on the model + deciding to look something up. +- **Cons:** higher token cost on every request, less user control, and a real + risk of **context collapse**, where unrelated domains (work and personal, or + two different customers' contexts) mix. Compressed summaries can also produce + hallucinated memories. +- **Choose when:** continuity is the product, sessions are short, and the memory + footprint per user is small and well curated. + +### On-Demand Tool-Based Memory + +Every conversation starts as a blank slate. Memory is activated only when the +agent explicitly calls a tool, typically one that searches past conversations +and one that lists recent sessions. The search runs over raw history rather than +pre-computed summaries, and the tool calls are visible to the user. + +- **Pros:** low baseline token cost, high precision, strong transparency (the + user sees exactly when memory was consulted), and natural scoping by project + or workspace so memory banks do not collapse into each other. +- **Cons:** sacrifices automatic continuity — if the agent does not decide to + search, relevant context is simply absent. +- **Choose when:** precision, auditability, and cost control matter more than + effortless continuity; or when memory volume per user is large. + +### Retrieval on Demand Over a Permission-Trimmed Index + +Nothing organizational is memorized at all. The agent holds only an explicit, +small personal memory layer (preferences, working style, recurring topics), and +everything else is retrieved live from an index that respects role-based access +control and tenant boundaries. + +- **Pros:** prevents data leakage, guarantees freshness, keeps the agent + stateless with respect to enterprise data, and makes deletion requirements + tractable. +- **Cons:** requires a mature, permission-aware index; recall quality depends + entirely on that index. +- **Choose when:** the content already lives in governed enterprise systems. + This is the recommended default for enterprise knowledge. + +### Extract-and-Update Memory Layer + +A dedicated memory service sits beside the agent. An incremental pipeline +extracts candidate facts from each exchange, decides whether to add, update, +merge, or delete existing entries, generates conversation summaries +asynchronously, and serves retrieval through vector search — optionally +augmented with a graph variant for relationships. + +- **Pros:** strong accuracy-to-cost ratio, memory logic decoupled from the + agent, reusable across agents and channels. +- **Cons:** an extra service to operate; extraction quality becomes a + first-class concern requiring its own evaluation. +- **Choose when:** building production-grade, cost-sensitive deployments with + multiple agents sharing a memory layer. + +### Self-Managed Tiered Memory + +Inspired by operating system virtual memory: the model has a limited "main +context" and an unlimited external store, and it manages the movement between +them itself through function calls. The agent decides what to page in and page +out. + +- **Pros:** highly autonomous, adapts to unusual tasks without hand-tuned + retrieval rules. +- **Cons:** unpredictable cost and latency, harder to audit, and memory quality + depends on model behavior rather than on a deterministic pipeline. +- **Choose when:** building autonomous, long-running agents where no fixed + retrieval policy can be defined up front. + +### Recommendation + +For enterprise multi-agent systems, a **hybrid** is the most reliable starting +point: + +- **Auto-inject** a small, curated profile (semantic memory) — cheap, always + useful. +- **Retrieve on demand** episodic memory through a memory search tool — precise + and cost controlled. +- **Never memorize** governed enterprise content; retrieve it live through a + permission-trimmed index. + +```mermaid +flowchart LR + Turn[User turn] --> Assemble{Context assembly} + Profile[(Semantic profile
auto-injected)] --> Assemble + Assemble --> Agent((Agent)) + Agent -- when needed --> EpisodicTool[Memory search tool] + EpisodicTool --> Episodic[(Episodic memory
vector store)] + Agent -- when needed --> KnowledgeTool[Knowledge retrieval] + KnowledgeTool --> Index[(Permission-trimmed
enterprise index)] + Episodic --> Agent + Index --> Agent +``` + +--- + +## Memory Scoping + +Scope defines who and what a memory is visible to. It must be an explicit +attribute of every memory entry, not an emergent property of where it was +stored. + +| Scope | Boundary | Typical content | +| --------------------- | ---------------------------------------------- | ----------------------------------------- | +| **Session** | Current conversation only | Working state, transient decisions | +| **Global** | All conversations for a subject | User profile, durable preferences | +| **Project/Workspace** | A single domain or work item | Project decisions, domain-specific facts | +| **Channel-specific** | One interaction channel | Channel preferences, channel-only history | +| **Role-based** | A class of subject (customer, employee, admin) | Policy defaults, role instructions | + +Two rules follow from this: + +- **Isolating by scope prevents context collapse.** Separate memory banks per + project or domain are the simplest mitigation for cross-domain leakage. +- **A unified profile can span channels while episodic memory stays + channel-scoped.** Sharing durable preferences across channels is usually + desirable; sharing raw conversation history across channels usually is not. + +--- + +## Trade-Offs + +### Auto-Inject vs On-Demand + +| | Auto-inject | On-demand | +| ------- | ------------------------------------------------------ | ------------------------------------------------- | +| **Pro** | Seamless, no user effort | Lower token cost, transparent, precise | +| **Con** | Higher token cost, less control, context collapse risk | May miss context if the search is never triggered | + +### Depth vs Breadth + +- **Deep memory** keeps full episodic history: expensive and noisy. +- **Shallow memory** keeps extracted facts only: cheap but lossy. +- **Balance:** facts for the durable profile, episodic detail for active cases. + +### Complexity vs Time-to-Value + +Short-term memory delivers value almost immediately; a full knowledge graph does +not. Start simple, measure impact, and add sophistication where the measurements +justify it. See the phased roadmap in +[Long-term memory](./Long-Term-Memory.md#adoption-roadmap). + +### The Fading Memory Problem + +As the memory store grows, retrieval precision drops: the signal gets lost in +the noise, and the system appears to "forget" despite storing more than ever. +This is the single most common failure mode of long-lived memory. + +Mitigations: importance weighting, access-frequency scoring, decay and tier +demotion, consolidation of duplicates, and active curation. These are detailed +in +[Relevance scoring and decay](./Long-Term-Memory.md#relevance-scoring-and-decay). + +--- + +## Token Economics + +Memory is a token budget decision as much as an accuracy decision. + +| Approach | Relative cost per query | Notes | +| -------------------- | ----------------------- | --------------------------------------------------------------------------------------------------- | +| **Full context** | Highest | Sending the entire history; establishes the accuracy baseline but is unaffordable and slow at scale | +| **Summarization** | Moderate | Roughly a 43% token reduction while retaining most of the context | +| **Fact extraction** | Low | Around 2K tokens per query in published benchmarks | +| **On-demand search** | Lowest baseline | Cost is only incurred when memory is actually consulted | + +Published benchmark comparisons of memory approaches — such as +[Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory](https://arxiv.org/abs/2504.19413), +which evaluates memory strategies on long-conversation benchmarks — consistently +report the same shape of result: extraction-based memory layers reach higher +accuracy than naive stored-summary memory at a fraction of the latency and token +cost of full-context, while graph-augmented variants trade some latency and +tokens for additional accuracy on relational questions. + +> Fact extraction is not free: its cost is the sum of the input chunk being +> analyzed, the extraction instruction prompt, and the structured output tokens. +> Because it runs asynchronously, it does not sit on the inference critical +> path, but it does appear on the bill. + +Define an explicit **memory token budget per agent type** and enforce it during +context assembly. See the budget allocation stage in +[Memory pipelines](./Memory-Pipelines.md#memory-retrieval-pipeline). + +--- + +## Metrics to Track + +- **Memory retrieval accuracy:** are the relevant facts actually surfaced? +- **Retrieval precision and recall:** how much of what is injected is used, and + how much of what was needed was missed. +- **Context retention rate across sessions:** target above 90%. +- **Token cost per query**, measured with and without memory. +- **p95 latency** for memory retrieval plus inference. +- **Repetition reduction rate:** how often users have to re-explain themselves. +- **Resolution time improvement** for customer-facing scenarios. +- **Satisfaction delta:** A/B comparison of memory-on versus memory-off. +- **Fading memory detection:** is retrieval precision degrading as the store + grows? + +See [Observability](../observability/Observability.md) and +[Evaluation](../evaluation/Evaluation.md) for how to instrument and evaluate +these signals. + +--- + +{{ #include ../../components/discuss-button.hbs }} diff --git a/docs/memory/Memory-Pipelines.md b/docs/memory/Memory-Pipelines.md new file mode 100644 index 0000000..47de985 --- /dev/null +++ b/docs/memory/Memory-Pipelines.md @@ -0,0 +1,227 @@ +# Memory Pipelines + +_Last updated: 2026-08-04_ + +Memory is not a database that agents read and write directly. It is a set of +pipelines: one on the read path that assembles context for an inference, and two +on the write path that turn raw conversations into durable, governed memory. + +This topic covers: + +- [How the pipelines connect](#how-the-pipelines-connect) +- [Memory retrieval pipeline](#memory-retrieval-pipeline) +- [STM session end pipeline](#stm-session-end-pipeline) +- [STM to LTM promotion pipeline](#stm-to-ltm-promotion-pipeline) + +--- + +## How the Pipelines Connect + +Three pipelines operate on different triggers and different timescales: + +| Pipeline | Trigger | Timescale | Path | +| --------------- | ------------------------ | ------------------------------ | ----- | +| **Retrieval** | Every user turn | Synchronous, milliseconds | Read | +| **Session end** | Session close or timeout | Asynchronous, seconds | Write | +| **Promotion** | Scheduled batch | Asynchronous, minutes to hours | Write | + +```mermaid +flowchart LR + User([User turn]) --> Retrieval[Memory retrieval
synchronous] + Retrieval --> STM[(Short-term
memory)] + Retrieval --> LTM[(Long-term
memory)] + Retrieval --> Context[Assembled context] --> Agent((Agent)) + Agent --> STM + + SessionEnd([Session end event]) --> EndPipeline[Session end pipeline
summarize + extract] + EndPipeline --> STM + + Schedule([Scheduled trigger]) --> Promotion[Promotion pipeline
validate + dedupe + scope] + STM --> Promotion --> LTM + + Retrieval -. access metadata
write-back .-> LTM +``` + +Two properties keep this design safe: + +- **Only the retrieval pipeline is on the critical path.** Summarization, + extraction, and promotion run asynchronously and never add latency to a user + turn. +- **The read path feeds the lifecycle.** Retrieval updates `retrieval_count` and + `last_retrieved_at`, which is what makes + [decay and reinforcement](./Long-Term-Memory.md#relevance-scoring-and-decay) + possible. + +--- + +## Memory Retrieval Pipeline + +Runs on every turn that requires context. Its job is to produce the best +possible context that fits the available token budget. + +```mermaid +flowchart TD + Query([Agent query
user turn requires context]) --> Window[Context window
calculate available tokens
system prompt, history] + Window --> MemQuery[Memory query
parallel dispatch to all tiers] + + MemQuery --> Tier1[Tier 1: Short-term memory
vector search] + MemQuery --> Tier2[Tier 2: Long-term memory
vector search] + + Tier1 --> Aggregator[Aggregator
merge results from all tiers] + Tier2 --> Aggregator + Aggregator --> Dedup[Deduplicate
remove redundant entries] + Dedup --> Rerank[Reranker
score by relevance to query
+ recency + confidence] + Rerank --> Budget[Budget allocation
fit top-ranked results
into available token budget] + Budget --> Assembled([Assembled context
injected into LLM prompt]) + Budget -.->|access metadata| WriteBack[(Update retrieval_count
and last_retrieved_at)] +``` + +### Stages + +| Stage | Purpose | Key logic | +| --------------------- | -------------------- | ----------------------------------------------------------------------------------------------------------- | +| **Context window** | Establish the budget | Subtract system prompt, injected profile, recent turns, and response headroom from the model context size | +| **Memory query** | Fetch candidates | Dispatch the query to STM and LTM **in parallel**; apply scope and sensitivity as query filters | +| **Aggregator** | Merge tiers | Combine candidate sets while preserving tier and provenance metadata | +| **Deduplicate** | Remove redundancy | Drop near-identical entries — a fact restated in both tiers wastes budget | +| **Reranker** | Order by usefulness | Combine semantic relevance to the query with recency, confidence, and the stored memory score | +| **Budget allocation** | Fit the budget | Fill the reserved retrieval budget with top-ranked entries; truncate rather than overflow | +| **Write-back** | Feed the lifecycle | Asynchronously increment `retrieval_count` and set `last_retrieved_at` for entries that entered the context | + +### Design notes + +- **Parallel dispatch is what makes multi-tier retrieval affordable.** Querying + tiers sequentially multiplies latency for no accuracy gain. +- **Filter, do not post-filter.** Scope, tenant, and sensitivity must be applied + inside the vector query. Post-filtering silently reduces recall and risks + leaking entries into ranking traces. +- **Reranking is where the memory score earns its keep.** Pure vector similarity + favors superficially similar text; adding recency, confidence, and access + frequency favors what has actually proven useful. +- **Count usage, not candidacy.** Increment access metadata for entries that + reach the assembled context, not for everything the search returned. +- **Fail open, degrade gracefully.** If a memory tier is unavailable, the turn + should proceed with reduced context rather than fail. + +--- + +## STM Session End Pipeline + +Runs when a session ends, either explicitly or through an inactivity timeout. It +converts a raw conversation into structured, searchable short-term memory +entries that are ready for later promotion. + +```mermaid +flowchart TD + Event([Session end event]) --> Fetch[Fetch full
conversation history] + Fetch --> Summarize[Summarization
LLM generates session summary] + Summarize --> Extract[Fact extraction
LLM extracts durable facts
preferences · decisions · entities] + Extract --> Entities[Entity recognition
link to known entities] + Entities --> Embed[Embedding generation
for summaries and facts] + Embed --> Conflict{Conflict resolution
compare against existing STM} + + Conflict -->|Conflict| Update[Update in-place
replace contradicted entry
preserve audit trail] + Conflict -->|No conflict| Write[Write to STM
vector index + KV store] + Update --> Write + Write --> Ack([Acknowledge
mark session as processed]) +``` + +### How it works + +1. **When a session ends, the full conversation history is fetched** from the + STM store — the raw turns, not the trimmed context that was sent to the + model. +2. **An LLM summarizes the session and extracts durable facts** — preferences, + decisions, and entities. Extraction should return structured output with a + confidence value per fact, so downstream stages can filter on it. +3. **Named entities are recognized and linked** to known entities in the system + (customers, products, tickets, systems). This is what later enables + relationship queries and, eventually, a knowledge graph. +4. **Vector embeddings are generated** for the summary and for each fact, so + they become semantically searchable. +5. **Conflicts with existing STM entries are resolved** — either updating an + entry in place, preserving the previous value as an audit trail, or writing a + new entry when there is no contradiction. +6. **Data is written to STM** as a vector index entry plus a key-value record. +7. **The session is acknowledged** and marked as processed via + `promotion_state`, which makes the pipeline idempotent and safe to retry. + +> Consider logging every memory update in a relational database. A queryable +> activity trail of writes, in-place updates, and conflicts is what makes memory +> behavior explainable after the fact. + +### Design notes + +- **Never run this synchronously with the last user turn.** It is triggered by + an event and processed by background workers. +- **Extraction is a quality-critical component and deserves evaluation.** Treat + its prompts as versioned artifacts and measure precision of extracted facts. + See [Evaluation](../evaluation/Evaluation.md). +- **Preserve the previous value on in-place updates.** Overwriting without + history makes contradictions impossible to investigate. +- **Guard against injected instructions.** Content extracted from a conversation + is untrusted input; strip instruction-like text before it becomes memory. + +--- + +## STM to LTM Promotion Pipeline + +Runs on a schedule. Its job is to decide which short-term facts have earned +durable, cross-session status — and to make sure that promoting them does not +create duplicates, contradictions, or scope violations. + +```mermaid +flowchart TD + Trigger([Pipeline trigger
scheduled]) --> Scan[Scan STM
find promotion candidates] + Scan --> Validation[Validation
verify fact consistency
cross-reference sources] + + Validation -->|Invalid| Discard[Discard or flag
remove or queue for review] + Validation -->|Valid| Dedup[Deduplication
check for existing LTM entries
semantic similarity search] + + Dedup -->|Duplicate| Merge[Merge
update existing LTM entry
increase confidence score] + Dedup -->|New| Scope[Scope assignment
determine visibility
instance / domain / org] + + Merge --> VectorWrite[Vector write
index in LTM vector store] + Scope --> VectorWrite + VectorWrite --> Audit([Audit entry
log promotion decision
source · reason · scope]) +``` + +### Stages + +| Stage | Purpose | Key logic | +| -------------------- | ----------------------------- | ------------------------------------------------------------------------------------------------------ | +| **Scan** | Identify promotion candidates | Query STM for entries with `confidence >= threshold` **and** `age >= minimum_age` | +| **Validation** | Verify factual consistency | Cross-reference the fact against other STM and LTM entries; LLM-based consistency check | +| **Deduplication** | Prevent duplicate LTM entries | Semantic similarity search in LTM; if similarity exceeds the threshold, merge instead of creating | +| **Scope assignment** | Determine sharing visibility | Apply the memory policies of the originating instance configuration (instance / domain / organization) | +| **Vector write** | Enable semantic search | Index the entry in the LTM vector store together with its scope metadata | +| **Audit** | Maintain a compliance trail | Log the promotion decision with source session, pipeline version, and scope rationale | + +### Design notes + +- **The minimum age requirement is deliberate.** A fact stated once at the end + of a session has not yet proven durable; waiting filters out transient + statements and lets contradicting evidence arrive first. +- **Validation is what prevents memory poisoning.** Promotion is the point where + a statement becomes trusted across sessions, so it is the right place to spend + compute on verification. +- **Merging beats accumulating.** Repeated evidence for the same fact should + raise its confidence and importance, not create a second entry — duplicate + entries are a direct cause of the + [fading memory problem](./Memory-Architecture-Patterns.md#the-fading-memory-problem). +- **Scope is assigned once, explicitly, at promotion time**, from the policy of + the originating configuration. This is the enforcement point that prevents + cross-domain leakage later. +- **Invalid candidates are flagged, not silently dropped**, when they concern + sensitive scopes: a queue for human review is often a governance requirement. +- **Everything is audited.** Source, reason, and scope for every promotion make + the memory store explainable and are typically mandatory for regulated + workloads. + +Promoted entries then enter the +[memory lifecycle](./Long-Term-Memory.md#memory-lifecycle), where retrieval +reinforces them and disuse gradually demotes them across tiers. + +--- + +{{ #include ../../components/discuss-button.hbs }} diff --git a/docs/memory/Memory.md b/docs/memory/Memory.md index b0fee9f..1198951 100644 --- a/docs/memory/Memory.md +++ b/docs/memory/Memory.md @@ -1,23 +1,166 @@ # Memory -_Last updated: 2025-05-18_ +_Last updated: 2026-08-04_ Memory is a foundational aspect of multi-agent systems, shaping how agents -understand context, make decisions, and collaborate effectively. This chapter -introduces two core memory types in multi-agent systems: +understand context, make decisions, and collaborate effectively. Without memory, +every interaction starts from zero: users re-explain themselves, agents repeat +work already done, and the system is unable to build on what it learned before. + +Memory is what turns a stateless request/response assistant into a system that +accumulates context over time. It is also one of the riskiest parts of the +architecture, because everything an agent remembers is data that must be scoped, +governed, secured, and eventually forgotten. + +This chapter covers: + +- [Memory types](#memory-types) +- [Memory sub-types](#memory-sub-types-cognitive-model) +- [Memory is not a knowledge base](#memory-is-not-a-knowledge-base) +- [Key principles](#key-principles) +- [Chapter map](#chapter-map) + +--- + +## Memory Types + +Three types of memory show up in almost every agentic architecture. They are +complementary, not alternatives. - **Short-term Memory (STM):** Enables agents to maintain recent context within an active session (also known as conversation history), supporting coherent - interaction and task coordination across agents. + interaction and task coordination across agents. It is session-scoped, + token-limited by the model context size, and is trimmed, summarized, or + discarded as the session evolves. - **Long-term Memory (LTM):** Provides persistence of information across sessions, allowing agents to recall knowledge, preferences, and outcomes over - time to provide personalized experiences. + time to provide personalized experiences. It stores a compressed, distilled + representation of what happened, not the raw transcript, and requires an + extraction and retrieval pipeline to be useful. +- **Working Memory:** The context actually assembled for a single inference + call. It is not a store, it is a _composition_: the system prompt, the + relevant portion of STM, and the facts retrieved from LTM, all fitted into the + available token budget. + +A useful analogy: STM is what is happening in this conversation, LTM is what is +stored in your database, and working memory is what is open on your desk right +now. + +```mermaid +flowchart LR + subgraph Stores + STM[(Short-term memory
session scoped)] + LTM[(Long-term memory
cross-session)] + end + SystemPrompt[/System prompt
+ instructions/] + Working[Working memory
assembled per request] + Model((LLM inference)) + + SystemPrompt --> Working + STM -- recent turns
+ summaries --> Working + LTM -- retrieved facts
+ past episodes --> Working + Working a1@==> Model + Model -- new turns --> STM + STM -. promotion .-> LTM + a1@{ animate: true } +``` + +> Working memory is the only thing the model ever sees. STM and LTM are design +> decisions about _what gets to be there_ and _at what cost_. + +--- + +## Memory Sub-Types (Cognitive Model) + +Within long-term memory, it helps to separate what is being remembered. Each +sub-type has different extraction logic, different storage needs, and different +retention rules. + +| Sub-type | What it holds | Typical examples | +| -------------- | -------------------------------------------------- | ------------------------------------------------------------------------------ | +| **Semantic** | Extracted facts and attributes | User prefers email communication; account tier; domain business rules | +| **Episodic** | Complete past interactions and events, timestamped | "On Jan 15 the user reported a billing issue"; a prior troubleshooting session | +| **Procedural** | Learned workflows and methods | Password reset follows a 3-step flow; escalation rules | + +- **Semantic memory** is the cheapest and highest-signal form of memory. It is + what usually gets injected directly into the system prompt as a compact + profile. +- **Episodic memory** matters for multi-touch journeys, where a customer or + employee interacts across days and channels. It is what allows an agent to + respond correctly when someone says _"I was told it would be resolved by + Friday"_, or when an employee needs context from a prior session handled by a + different agent. +- **Procedural memory** only becomes a distinct concern when the procedure was + never written down by anyone. If the workflow already exists as documentation, + a runbook, or code, it belongs in a knowledge source or in a tool, not in + memory. + +--- + +## Memory Is Not a Knowledge Base + +A frequent design mistake is to treat enterprise content as memory. Knowledge +sources such as document repositories, search indexes, and RAG corpora act as +long-term **recall**, not memory: they are authoritative, shared, permission +controlled, and change independently of any conversation. + +A robust pattern is to keep enterprise content out of memory entirely and +retrieve it on demand through a permission-trimmed index. This approach: + +- **Prevents data leakage**, because access control is evaluated at query time + rather than frozen at the moment a memory was written. +- **Enables real-time freshness**, because the agent always sees the current + version of the content. +- **Keeps the agent stateless** with respect to enterprise data, which + dramatically simplifies compliance and deletion requirements. + +Memory, in contrast, holds what is true about _this user, this session, and this +collaboration_ and would otherwise be lost: preferences, decisions, open issues, +and interaction history. + +--- + +## Key Principles + +The following principles apply regardless of the storage technology or the +agentic framework in use: + +- **Not all memories are equal.** Memory needs importance weighting; a fact the + user explicitly asked the agent to remember is worth far more than an + incidental detail extracted from a transcript. +- **Reinforcement through repetition strengthens recall.** Facts that keep being + retrieved and confirmed should become more prominent, not less. +- **Retrieval must be contextual.** Relevance to the current query, not raw + recency, decides what enters working memory. +- **Memory decays.** Outdated or unused information should fade or expire rather + than accumulate indefinitely. +- **User control and transparency are non-negotiable.** Users must be able to + see, edit, and delete what the system remembers about them. +- **Memory scope must match use case boundaries.** A memory captured in one + channel, project, or tenant should not silently surface in another. + +--- + +## Chapter Map Designing STM and LTM in multi-agent systems brings unique challenges around synchronization, ownership, privacy, and data consistency. The following sections outline key patterns and trade-offs for integrating effective memory into your architecture. ---- +- **[Memory architecture patterns](./Memory-Architecture-Patterns.md):** storage + strategies, retrieval patterns, memory scoping, and the trade-offs between + automatically injecting memory and retrieving it on demand. +- **[Short-term memory](./Short-Term-Memory.md):** session storage design + approaches, context window management, and what happens when a session ends. +- **[Long-term memory](./Long-Term-Memory.md):** what to persist, how to score + and decay it, scope decision framework, governance, and operating memory at + scale. +- **[Memory pipelines](./Memory-Pipelines.md):** the retrieval, session-end, and + promotion pipelines that connect the two tiers. -{{ #include ../../components/discuss-button.hbs }} +Memory is also the largest single input to context assembly. See +[Context Engineering](../context-engineering/Context-Engineering.md) for how the +assembled context is optimized once memory has produced it. + +--- diff --git a/docs/memory/Short-Term-Memory.md b/docs/memory/Short-Term-Memory.md index 4e8db01..f7c625e 100644 --- a/docs/memory/Short-Term-Memory.md +++ b/docs/memory/Short-Term-Memory.md @@ -2,13 +2,18 @@ # Short-term Memory -_Last updated: 2025-05-26_ +_Last updated: 2026-08-04_ When building the STM (Short-Term Memory) layer for a multi-agent system, the storage engine choice is critical. STM typically serves to persist conversation history, contextual state, and intermediary data that agents use to manage context and continuity during workflows. +STM is **session-scoped** and **token-limited**: it lives for as long as the +session is active, and only a subset of it ever reaches the model. Everything +that must survive the session belongs in +[Long-term memory](./Long-Term-Memory.md). + This topic covers: - [Key requirements](#key-requirements) @@ -16,6 +21,11 @@ This topic covers: - [Shared memory](#1-shared-memory) - [Distributed memory](#2-distributed-memory) - [Hybrid memory](#3-hybrid-memory) +- [Recommended practices](#recommended-practices) +- [Context window management](#context-window-management) +- [Working memory assembly](#working-memory-assembly) +- [Session end](#session-end) +- [Multi-channel considerations](#multi-channel-considerations) - [Data retention](#data-retention) --- @@ -98,6 +108,20 @@ Consider designing the documents to capture: communication channel, tags). - **Timestamp:** When the message was written. +Additionally, consider the fields consumed by the downstream memory pipelines: + +- **Session Status:** Whether the session is active, idle, closed, or expired. + The session end event that triggers summarization and extraction is derived + from this transition. +- **Summary:** The rolling summary of older turns, so it does not have to be + recomputed on every request. +- **Token Count:** Tokens per message and accumulated per session, used to + enforce the context budget and decide when to summarize. +- **Promotion State:** Whether the session has already been processed by the + [session end](#session-end) and + [promotion](./Memory-Pipelines.md#stm-to-ltm-promotion-pipeline) pipelines. + This makes both pipelines idempotent and safely retryable. + > **Leverage chat history message objects from agentic frameworks rather than > normalizing data when possible**. Frameworks such as Semantic Kernel and > LangChain already provide rich, extensible and tested chat history objects. @@ -253,6 +277,145 @@ the isolated collections. --- +## Context Window Management + +Storage keeps the whole session; the model does not. The context window is a +fixed token budget shared by the system prompt, the injected user profile, the +retrieved long-term facts, and the recent turns. A **token-limit mechanism is +mandatory** — without it, sessions fail abruptly once they outgrow the model +context size. + +### Sliding Window + +The simplest mechanism keeps the last _N_ turns verbatim and drops everything +older. It is cheap and predictable, but it discards decisions and constraints +agreed earlier in the session, which is exactly the context users assume the +agent still has. + +### Progressive Summarization + +As the session approaches the context limit, older turns are compressed into a +rolling summary that is carried forward while recent turns stay verbatim. + +```mermaid +flowchart LR + Turns[Full session history] --> Check{Approaching
token limit?} + Check -- No --> Window[Recent turns verbatim] + Check -- Yes --> Summarize[Summarize oldest turns] + Summarize --> Summary[(Rolling summary)] + Summary --> Window + Window --> Context[Context for next inference] +``` + +Recommendations: + +- **Trigger on a threshold, not on the hard limit** (for example, at 70-80% of + the budget), so summarization never competes with the current request. +- **Persist the summary in the session document** and update it incrementally, + rather than re-summarizing the entire history each time. +- **Preserve decisions, constraints, identifiers, and open items verbatim.** + Summaries lose exactly the details that later turns depend on, and + over-compression is a common source of fabricated context. +- **Keep the raw history in STM storage** even when it is no longer sent to the + model — it is required for auditing and for the session end pipeline. + +### Thread Resumption + +When a user returns to an existing conversation thread after a gap, the session +context no longer exists in the process memory. Before inferencing, the thread +must be reconstructed: load the session document, produce (or reuse) a thread +summary, and send that summary plus the most recent turns to the model. + +For long-lived threads, treat resumption as a first-class path: it is where +context loss is most visible to the user, and where a stale or missing summary +turns into an obvious regression. + +--- + +## Working Memory Assembly + +Working memory is what the model actually sees for a single request. It is +assembled per turn from four sources, each with its own share of the token +budget: + +| Section | Source | Typical budget behavior | +| ---------------------- | -------------------------------- | ---------------------------------------- | +| System prompt | Static configuration | Fixed, reserved first | +| Semantic profile | Long-term memory (auto-injected) | Small and capped; curated, not unbounded | +| Retrieved memories | Long-term memory (on demand) | Variable, filled by relevance ranking | +| Recent turns + summary | Short-term memory | Remainder of the budget | + +Guidelines: + +- **Reserve the budget in priority order** — instructions first, then durable + profile, then retrieved memories, then recent history — and let the lowest + priority section absorb the truncation. +- **Cap each section explicitly.** An uncapped profile or an uncapped retrieval + result set will silently starve the conversation history. +- **Label the provenance of each block** in the prompt (profile, retrieved + memory, conversation) so the model can weigh them differently and so the + assembled context stays debuggable. +- **Leave headroom for the response.** The output tokens come out of the same + window. + +The full read path, including ranking and budget allocation across both memory +tiers, is described in +[Memory pipelines](./Memory-Pipelines.md#memory-retrieval-pipeline). + +--- + +## Session End + +When a session closes, expires, or goes idle beyond a threshold, three distinct +decisions must be made about its content: + +1. **Discard** transient working state that has no value beyond the session + (intermediate tool payloads, retry scaffolding, partial plans). +2. **Archive** the raw conversation to cold storage for cost, analytics, and + compliance — see [Data retention](#data-retention). +3. **Promote** the durable signal (preferences, decisions, entities, outcomes) + toward long-term memory. + +> Archiving and promotion are different paths and must not be conflated. +> Archiving moves the _transcript_ to cheaper storage for later human or +> analytical use. Promotion extracts _facts_ that the agent will actively recall +> in future sessions. A conversation can be archived and never promoted. + +The session end event triggers summarization, fact extraction, embedding +generation, and conflict resolution against existing entries. That flow is +described in detail in +[Memory pipelines](./Memory-Pipelines.md#stm-session-end-pipeline), and the +subsequent promotion into long-term memory in +[STM to LTM promotion](./Memory-Pipelines.md#stm-to-ltm-promotion-pipeline). + +Because sessions may never be explicitly closed, define an **inactivity +timeout** that emits the same event, and make the downstream processing +idempotent using the `promotion_state` field. + +--- + +## Multi-Channel Considerations + +Session boundaries are channel-specific, and getting them wrong is a frequent +source of both context loss and context leakage. + +| Channel | Session characteristics | Implications | +| -------------------------- | ------------------------------------------- | ----------------------------------------------------------------------- | +| **Asynchronous messaging** | Long gaps between messages, no explicit end | Long inactivity timeouts, thread summaries, resumption path is the norm | +| **Web chat** | Synchronous, short, explicit start and end | Short timeouts, simple sliding window is often enough | +| **Internal tools** | Persistent project or ticket context | Session bound to the work item rather than to time | + +Recommendations: + +- **Share the durable profile across channels**, so a user is recognized + everywhere. +- **Do not leak raw conversation history across channels.** Keep episodic STM + channel-scoped unless there is an explicit reason to share it. +- **Record the channel on every session** so retention, summarization, and + promotion policies can differ per channel. + +--- + ## Data Retention While STM is suitable for hot, fast-access storage, it is recommended to archive @@ -274,6 +437,9 @@ a set number of days. This practice serves several key purposes: documents from STM to a cold storage. 3. **Indexing for Analytics:** Optionally, batch-load archived data into an analytics warehouse for reporting and custom queries. +4. **Promote Before Archiving:** Ensure the session end and promotion pipelines + have processed a session (check `promotion_state`) before its documents leave + hot storage, otherwise durable facts are lost with the transcript. ---