MongoDB-backed persistent memory for the Vercel AI SDK and a first-class memory provider for Vercel Eve. Gives your agent five structured memory tiers — Session, Semantic, Procedural, Episodic, and Scratchpad — all stored in MongoDB Atlas with automatic vector search indexes and per-type retention policies. Bring your own embedding model, or let Atlas Automated Embedding (Voyage AI) generate vectors server-side with zero client embedding pipeline.
- 🧠 5 memory types — session history, entity knowledge, how-to procedures, event episodes, and a temporary scratchpad
- 🤖 Vercel Eve provider — drop-in
mongodbMemory()for Eve's memory slots via the/evesubpath: automatic recall before every turn, automatic or agent-driven capture, andremember/forgettools — all partitioned by Eve's trusted scope key - ⚡ Atlas Automated Embedding — opt in with
autoEmbed: { model: 'voyage-4' }and Atlas embeds your memories server-side (Voyage AI); no embedding API key or client pipeline required - 🔍 Atlas Vector Search — semantic retrieval powered by any Vercel AI SDK embedding model
- 📏 Auto-detected dimensions — no need to configure vector dimensions; the package probes them at startup
- ⏱️ Flexible retention — per-type policies:
none,ttl,ttl+importance, or fullydynamic(forgetting curve) - 🗑️ Agent-driven
memory_forget— let the model mark individual memories for immediate deletion - 🧩 Multi-tenant ready — custom collection/index names, extra filter fields (e.g.
tenant_id), or disable types entirely - 🔄 Temporal versioning — semantic and procedural memories are versioned; history is preserved
- 🏗️ Zero lock-in — works with any AI SDK model and any embedding provider (VoyageAI, OpenAI, Cohere, Google, etc.)
- 🔌 Lazy connection — MongoDB connects on first tool use; call
.connect()early if you prefer
npm install @mongodb-developer/vercel-ai-memory
# or
pnpm add @mongodb-developer/vercel-ai-memoryPeer dependencies (install separately):
npm install ai mongodb zodThis package supports AI SDK v6 and v7 (ai@^6.0.0 || ^7.0.0). The examples
use APIs such as ToolLoopAgent and isLoopFinished() that are available in
both major versions.
import { createMongoDBMemory } from '@mongodb-developer/vercel-ai-memory'
import { openai } from '@ai-sdk/openai'
import { ToolLoopAgent } from 'ai'
// ── 1. Create the memory instance (once, at module/server level) ──────────────
const mongodbMemory = createMongoDBMemory({
uri: process.env.MONGODB_URI!,
embedder: openai.embedding('text-embedding-3-small'),
})
// ── 2. Use per-request, scoped to a user and session ─────────────────────────
const agent = new ToolLoopAgent({
model: openai('gpt-4.1'),
tools: mongodbMemory({ userId: 'alice', sessionId: 'sess-001' }),
})
const result = await agent.generate({
prompt: 'My name is Alice and I love hiking. Remember that.',
})mongodbMemory({ userId, sessionId }) returns the tools record directly — no spreading needed.
The package exposes session (short-term conversation) memory in two modes, both backed by the same Mongo collection. Pick based on how much determinism you need — they can't be combined for the same session.
Shown in the Quick Start above. The LLM sees memory with session_append / session_recent commands and decides when to call them. This is the simplest setup but not deterministic: the model sometimes skips writes (only user or only assistant turns get persisted) or skips reads (agent forgets earlier turns).
Use this for demos, prototypes, or agents where the LLM genuinely curates what belongs in the transcript.
Hide the session tool commands from the LLM and let the runtime read/write the transcript on every turn via Vercel AI SDK hooks. Every user, assistant, and tool turn is captured exactly once per generation, regardless of what the LLM decides.
import { createMongoDBMemory } from '@mongodb-developer/vercel-ai-memory'
import { openai } from '@ai-sdk/openai'
import { ToolLoopAgent, isLoopFinished } from 'ai'
import { z } from 'zod'
const mongodbMemory = createMongoDBMemory({
uri: process.env.MONGODB_URI!,
embedder: openai.embedding('text-embedding-3-small'),
// Hide session_append / session_recent from the LLM tool surface, but keep
// the session collection + store methods live for the runtime hooks below.
// (Use `disable: ['session']` instead if you want the collection entirely off.)
topology: { hideToolCommands: ['session'] },
})
const agent = new ToolLoopAgent({
model: openai('gpt-4o-mini'),
callOptionsSchema: z.object({
userId: z.string(),
sessionId: z.string(),
prompt: z.string(),
}),
// ── PRE hook: read history + scope onFinish ─────────────────────────────────
// Strip `prompt` / `messages` from the incoming settings — the AI SDK enforces
// `prompt` XOR `messages`, and we're replacing them with our restored history.
prepareCall: async ({ options, prompt: _p, messages: _m, ...settings }) => {
const { userId, sessionId, prompt } = options!
const history = await mongodbMemory.loadSession({ userId, sessionId })
return {
...settings,
tools: mongodbMemory({ userId, sessionId }),
messages: [...history, { role: 'user', content: prompt }],
// per-call state flows to onFinish via experimental_context
experimental_context: { userId, sessionId, prompt },
}
},
// ── POST hook: write every turn exactly once ────────────────────────────────
onFinish: mongodbMemory.onFinish(),
stopWhen: isLoopFinished(),
})
// Usage
await agent.generate({
prompt: 'Hi! My name is Alex.',
options: { userId: 'alice', sessionId: 'sess-001', prompt: 'Hi! My name is Alex.' },
})What the hooks do:
| Hook | Method | When it runs | What it does |
|---|---|---|---|
| Pre | mongodbMemory.loadSession({ userId, sessionId }) |
Inside prepareCall, before every LLM call |
Reads prior turns from Mongo, returns ModelMessage[] you prepend to messages. |
| Post | mongodbMemory.onFinish() |
After the full tool loop finishes | Persists the user prompt + every assistant & tool message across all steps. |
loadSession() restores user/assistant turns only. Tool turns remain stored in Mongo as
transcript/audit records, but they are not replayed as provider tool-result blocks because
OpenAI Responses and Anthropic Messages reject orphan tool results without the matching prior
assistant tool call.
Why experimental_context? ToolLoopAgent only accepts onFinish at construction time, but you still need per-call scope (userId, sessionId, prompt). The hook reads those from event.experimental_context, which you set inside prepareCall. You can also call mongodbMemory.onFinish({ userId, sessionId, prompt }) with a baked-in closure when using one-shot generateText / streamText.
See examples/deterministic-agent.ts for the full runnable example.
Note: The other memory tiers (semantic, procedural, episodic, scratchpad) are intentionally left under LLM control — they should be selective and content-dependent, not every-turn.
MongoDB is also available as a first-class Vercel Eve memory provider via the additive /eve subpath — layered on top of the exact same store, embeddings, vector search, retention, and forgetting logic used by the AI SDK integration.
npm install @mongodb-developer/vercel-ai-memory eve
eveis an optional peer dependency. If you only use the AI SDK integration (the root import), you don't need to installeve, and it is never bundled into the root entrypoint.
// agent/memory/profile.ts
import { defineMemory } from 'eve/memory'
import { byPrincipal } from 'eve/memory/scope'
import { mongodbMemory } from '@mongodb-developer/vercel-ai-memory/eve'
export default defineMemory({
scope: byPrincipal,
provider: mongodbMemory({
uri: process.env.MONGODB_URI!,
embedder,
capture: 'automatic',
}),
})That's the whole integration. Eve owns the agent runtime and lifecycle; MongoDB provides persistent, personalized, cross-session memory. You don't write remember.ts, recall.ts, semantic_search.ts, episodic_search.ts, or forget.ts — the provider handles all of it.
| Eve surface | What MongoDB does |
|---|---|
recall["turn.started"] |
Parallel semantic / episodic / procedural vector search, returned as recalled messages before the model responds — automatically. |
capture["turn.completed"] |
(When capture: "automatic") extracts & classifies durable memories and stores them. |
tools |
Exposes remember and forget (Eve qualifies them as profile__remember, profile__forget). |
The MongoDB owner/user id is never supplied by the model. The adapter uses Eve's trusted ctx.memory.scope.key as the partition key for every read and write:
const scope = ctx.memory.scope.key
store.semanticSearch(scope, query) // recall
store.semanticSave(scope, name, content) // capture / rememberTwo principals therefore can never retrieve each other's memories.
Eve's partition key is a digest of namespace + scope. Both must be stable across runs or the agent will write to a fresh, empty partition every time and never recall anything:
scope—byPrincipaluses the authenticated caller. In the local dev TUI andeve invokethere is no real auth, so the anonymous principal changes between runs. Pin a stable dev id.namespace— when omitted, Eve's default namespace includes a hash of the build root (.eve/builds/<id>/…), which changes on everyeve dev/eve invokestart. So even with a pinned scope, the key still changes each run. Set an explicit namespace.
const scope = (ctx) =>
process.env.EVE_MEMORY_PROD ? byPrincipal(ctx) : [process.env.DEV_PRINCIPAL_ID ?? 'dev-user']
export default defineMemory({
namespace: 'my-agent/profile', // stable across builds & restarts
scope, // stable in dev, real per-user in prod
provider: mongodbMemory({ /* … */ }),
})You can confirm the partition is stable with eve traces: gen_ai.memory.store.id should be identical across invocations, and gen_ai.memory.record.count should be > 0 in a brand-new session once you've saved something. See examples/eve-grove for the full runnable setup.
| Mode | Behavior |
|---|---|
"automatic" (default) |
After each turn, extract & classify durable memories and store them. Provide captureModel for full LLM classification; without one, a lightweight heuristic fallback is used and a one-time notice is logged. |
"agent" |
The model decides via the profile__remember / profile__forget tools. |
false |
Recall only; no automatic capture. |
mongodbMemory({
uri: process.env.MONGODB_URI!,
embedder, // or: autoEmbed: { model: 'voyage-3-large' }
capture: 'automatic', // 'automatic' | 'agent' | false
captureModel, // optional LanguageModel for automatic capture
memory: { semantic: true, episodic: true, procedural: true },
recall: { limit: 8, minSimilarity: 0.72 },
retention: { episodic: { mode: 'ttl', ttlSeconds: 60 * 60 * 24 * 90 } },
})Session and scratchpad tiers are intentionally disabled inside the Eve adapter because Eve owns the session/runtime lifecycle. This is Eve-adapter-only and does not change the AI SDK integration.
Both the AI SDK integration and the Eve adapter can let MongoDB Atlas generate and manage embeddings server-side (Voyage AI) — no client embedding pipeline. Pass autoEmbed instead of embedder:
createMongoDBMemory({
uri: process.env.MONGODB_URI!,
autoEmbed: { model: 'voyage-3-large' },
})
// or in Eve:
mongodbMemory({ uri: process.env.MONGODB_URI!, autoEmbed: { model: 'voyage-3-large' } })In this mode the store persists the source text (embed_text), bootstraps autoEmbed vector indexes, and issues $vectorSearch with a query string. Provide exactly one of embedder / autoEmbed.
A full runnable agent example lives in examples/auto-embed-agent.ts — it's the auto-embedding twin of examples/basic-agent.ts (same multi-tenant ToolLoopAgent, just autoEmbed: { model: 'voyage-4' } instead of a client embedder). Run it with npx tsx examples/auto-embed-agent.ts.
Requires MongoDB Atlas with Voyage AI enabled for your organization. On dedicated (
M10+) clusters, enable storage auto-scaling. See the Automated Embedding docs.
Verifying it works. An Atlas-gated integration test exercises this path end-to-end (
tests/integration/auto-embed.test.ts): it bootstrapstype:"autoEmbed"indexes (modelvoyage-4, pathembed_text), asserts each*Savepersistsembed_text(and no clientembeddingarray), and confirms each*Searchranks the relevant doc via server-side embeddings — including through the AI SDKmemorytool. Run it withMONGODB_URI="mongodb+srv://…" npm run test:integration. It runs against an isolated database (agent_memory_autoembed_test) so itsautoEmbedindexes never collide with the default client-modevectorindexes. If Voyage AI / Automated Embedding (Preview) is not enabled for your org,createSearchIndexesfails — the store logs[mongodb-memory] Failed to create vector index …, and the index/search assertions are skipped with a clear reason (the persistence assertions still run).
All options are passed to createMongoDBMemory(options).
| Option | Type | Required | Default | Description |
|---|---|---|---|---|
uri |
string |
✅ | — | MongoDB Atlas connection string |
embedder |
EmbeddingModel |
✅ | — | Any Vercel AI SDK embedding model |
userId |
string |
— | 'default' |
Default userId (override per-call) |
sessionId |
string |
— | 'default' |
Default sessionId (override per-call) |
dbName |
string |
— | 'agent_memory' |
Database name. (Deprecated — prefer topology.dbName.) |
topology |
TopologyOptions |
— | {} |
Where data lives — db, collections, indexes, disabled types |
retention |
RetentionOptions |
— | see below | Per-type decay / TTL policies |
filtering |
FilteringOptions |
— | see below | Retrieval-time filters on vector search results |
defaults |
DefaultsOptions |
— | see below | Small defaults (importance, limits, similarity) |
| Option | Type | Default | Description |
|---|---|---|---|
topology.dbName |
string |
'agent_memory' |
Database name (takes precedence over the legacy top-level dbName) |
topology.collections |
Partial<Record<MemoryType, string>> |
see Collection defaults below | Override collection names per memory type |
topology.vectorIndexNames |
Partial<Record<VectorMemoryType, string>> |
see Index defaults below | Override Atlas Vector Search index names |
topology.disable |
MemoryType[] |
[] |
Disable memory types entirely — no bootstrap, tool commands removed, store methods throw |
topology.hideToolCommands |
MemoryType[] |
[] |
Hide a memory type's commands from the tool schema, but keep its collection + store methods live (use with runtime hooks) |
topology.extraFilterFields |
Partial<Record<VectorMemoryType, string[]>> |
{} |
Extra scalar fields to index as Atlas Search filters (e.g. ['tenant_id']) |
Collection defaults:
| Memory type | Default collection name |
|---|---|
session |
session_memory |
semantic |
semantic_memory |
procedural |
procedural_memory |
episodic |
episodic_memory |
scratchpad |
scratchpad_memory |
Index defaults:
| Memory type | Default vector index name |
|---|---|
semantic |
semantic_vector_index |
procedural |
procedural_vector_index |
episodic |
episodic_vector_index |
Every memory type accepts a DecayPolicy, one of four modes:
| Mode | Shape | What it does |
|---|---|---|
none |
{ mode: 'none' } |
Memories never auto-expire. |
ttl |
{ mode: 'ttl', ttlSeconds, field? } |
Classic Mongo TTL index on a Date field. |
ttl+importance |
{ mode: 'ttl+importance', ttlSeconds, minImportance, field? } |
TTL only applies to docs with importance < minImportance — important memories are immune. |
dynamic |
{ mode: 'dynamic', computeExpireAt, refreshOnRead? } |
Per-doc expire_at computed on write (and recomputed on read when refreshOnRead: true, default). Backed by an expireAfterSeconds: 0 TTL index — perfect for forgetting-curve semantics. |
computeExpireAt(input) receives { importance, stats, createdAt } and returns a Date (or null to never expire).
Default retention policies:
| Memory type | Default policy |
|---|---|
session |
{ mode: 'ttl', ttlSeconds: 86_400 } — 24 h on created_at |
scratchpad |
{ mode: 'ttl', ttlSeconds: 3_600 } — 1 h on created_at |
episodic |
{ mode: 'ttl', ttlSeconds: 31_536_000, field: 'stats.last_retrieved' } — 1 yr of inactivity |
semantic |
{ mode: 'none' } |
procedural |
{ mode: 'none' } |
Examples:
retention: {
// Keep only important semantic facts after a week
semantic: { mode: 'ttl+importance', ttlSeconds: 604_800, minImportance: 7 },
// Forgetting curve — important episodes live longer, rarely-read ones decay fast
episodic: {
mode: 'dynamic',
refreshOnRead: true,
computeExpireAt: ({ importance, stats, createdAt }) => {
const hoursFromNow = Math.pow(2, importance) // 2^importance hours
return new Date(Date.now() + hoursFromNow * 3600 * 1000)
},
},
// Disable session auto-expiry entirely
session: { mode: 'none' },
}Applied to every *_search vector query.
| Option | Type | Default | Description |
|---|---|---|---|
filtering.minImportance |
number (1–10) |
0 |
Drop results with importance < minImportance |
filtering.recencyWindowHours |
number |
0 (disabled) |
Only return memories retrieved within the last N hours |
filtering.numCandidatesMultiplier |
number |
10 |
$vectorSearch.numCandidates = limit * multiplier. Higher → better recall, slower query. |
| Option | Type | Default | Description |
|---|---|---|---|
defaults.importance |
number (1–10) |
5 |
Default importance when the agent doesn't supply one |
defaults.sessionRecentLimit |
number |
40 |
Default limit for session_recent |
defaults.searchLimit |
number |
5 |
Default limit for all *_search commands |
defaults.similarity |
'cosine' | 'dotProduct' | 'euclidean' |
'cosine' |
Vector similarity used when creating Atlas Search indexes |
createMongoDBMemory({
uri: process.env.MONGODB_URI!,
embedder: openai.embedding('text-embedding-3-small'),
})createMongoDBMemory({
uri: process.env.MONGODB_URI!,
embedder: openai.embedding('text-embedding-3-small'),
topology: {
extraFilterFields: {
semantic: ['tenant_id'],
procedural: ['tenant_id'],
episodic: ['tenant_id'],
},
},
})createMongoDBMemory({
uri: process.env.MONGODB_URI!,
embedder: openai.embedding('text-embedding-3-small'),
topology: {
disable: ['semantic', 'procedural', 'episodic', 'scratchpad'],
},
})createMongoDBMemory({
uri: process.env.MONGODB_URI!,
embedder: openai.embedding('text-embedding-3-small'),
topology: {
dbName: 'my_app',
collections: {
session: 'agent_sessions',
semantic: 'agent_facts',
},
vectorIndexNames: {
semantic: 'agent_facts_vs',
},
},
})createMongoDBMemory({
uri: process.env.MONGODB_URI!,
embedder: openai.embedding('text-embedding-3-small'),
retention: {
semantic: { mode: 'ttl+importance', ttlSeconds: 30 * 86_400, minImportance: 6 },
},
filtering: {
minImportance: 3, // never surface low-importance memories
recencyWindowHours: 24 * 30, // only last 30 days of reads
},
})createMongoDBMemory({
uri: process.env.MONGODB_URI!,
embedder: openai.embedding('text-embedding-3-small'),
retention: {
episodic: {
mode: 'dynamic',
refreshOnRead: true,
computeExpireAt: ({ importance, stats }) => {
// Half-life grows with importance and retrieval count
const baseHours = 24 * Math.pow(1.5, importance)
const boost = (stats?.retrieval_ct ?? 0) * 12
return new Date(Date.now() + (baseHours + boost) * 3600 * 1000)
},
},
},
})filtering: {
numCandidatesMultiplier: 25, // higher recall, slower
}Any model that implements the Vercel AI SDK EmbeddingModel interface. Dimensions are auto-detected.
import { openai } from '@ai-sdk/openai'
import { cohere } from '@ai-sdk/cohere'
import { google } from '@ai-sdk/google'
embedder: openai.embedding('text-embedding-3-small') // 1536
embedder: cohere.embedding('embed-english-v3.0') // 1024
embedder: google.textEmbeddingModel('text-embedding-004') // 768Creates a MongoDB memory provider. Returns a callable MongoDBMemoryInstance.
Callable — returns a { memory: Tool } record scoped to the given userId and sessionId.
tools: mongodbMemory({ userId: req.userId, sessionId: req.sessionId })
tools: mongodbMemory() // use defaults set at creation timeExplicitly connect and bootstrap indexes. Called automatically on first tool use.
await mongodbMemory.connect() // pre-warm on server startupGracefully close the MongoDB connection.
process.on('SIGTERM', () => mongodbMemory.close())Raw MongoMemoryStore for advanced direct access:
await mongodbMemory.store.semanticSave('alice', 'Preference', 'Loves hiking', { importance: 8 })
const results = await mongodbMemory.store.semanticSearch('alice', 'outdoor activities')
await mongodbMemory.store.forget('semantic', someMemoryId) // immediate agent-driven deleteThe single memory tool accepts a command field and routes to the right memory type.
Per-session conversation turns.
session_append {role, content} — Save a turn
session_recent {limit?} — Get last N turns (default: defaults.sessionRecentLimit)
Long-term knowledge about people, entities, and user preferences. Temporally versioned.
semantic_save {name, content, importance?, tags?} — Save/update entity knowledge
semantic_search {query, limit?} — Vector search
How-to knowledge: tasks, workflows, agent instructions. Temporally versioned.
procedural_save {task, content, importance?, source?} — Save/update a procedure
procedural_search {query, limit?} — Vector search
Records of key events and outcomes.
episodic_save {event_type, content, importance?, context?} — Record an event
episodic_search {query, limit?} — Vector search
Temporary working notes. Can be promoted to Episodic memory.
scratchpad_write {content} — Write a temporary note
scratchpad_read — Read current session notes
scratchpad_promote {scratchpad_id, event_type} — Promote note → Episodic
Always available; lets the LLM explicitly forget a specific memory (e.g. when the user says "forget that").
memory_forget {memory_type, id, reason?}
Internally sets expire_at to "now", leveraging the expire_at TTL index. The doc is removed on the next TTL sweep (usually within ~60 s).
Disabled types are removed from the tool's command enum and skipped during bootstrap, so the agent can't attempt to use them.
With defaults, the package creates:
| Collection | Retention | Vector Index |
|---|---|---|
session_memory |
24 h on created_at |
— |
semantic_memory |
none | ✅ cosine (auto-dims) |
procedural_memory |
none | ✅ cosine (auto-dims) |
episodic_memory |
1 yr on stats.last_retrieved |
✅ cosine (auto-dims) |
scratchpad_memory |
1 h on created_at |
— |
All collections also get an expire_at TTL index (expireAfterSeconds: 0) to power memory_forget and dynamic retention.
Note: Atlas Vector Search indexes are created asynchronously. Allow a few seconds for them to build on a fresh database.
// app/api/chat/route.ts
import { createMongoDBMemory } from '@mongodb-developer/vercel-ai-memory'
import { openai } from '@ai-sdk/openai'
import { ToolLoopAgent, createAgentUIStreamResponse } from 'ai'
const mongodbMemory = createMongoDBMemory({
uri: process.env.MONGODB_URI!,
embedder: openai.embedding('text-embedding-3-small'),
})
export async function POST(req: Request) {
const { messages, userId, sessionId } = await req.json()
const agent = new ToolLoopAgent({
model: openai('gpt-4.1'),
tools: mongodbMemory({ userId, sessionId }),
instructions: `You are a helpful assistant with persistent memory.
At the start of each session, call session_recent to restore context.
Save important facts with semantic_save. If the user asks you to forget
something, call memory_forget with the matching memory_type + id.`,
})
return createAgentUIStreamResponse({ agent, uiMessages: messages })
}const { store } = mongodbMemory
// Seed procedural knowledge
await store.proceduralSave(
'system',
'Onboarding Flow',
'1. Greet user by name. 2. Ask about their goals. 3. Set up preferences.',
{ source: 'human_expert', importance: 9 }
)
// Promote a scratchpad note to episodic memory
const scratchId = await store.scratchpadWrite('alice', 'sess-001', 'User mentioned they dislike emails')
await store.scratchpadPromote(scratchId.toString(), 'alice', 'preference', { importance: 7 })
// Immediately forget a memory
await store.forget('semantic', '507f1f77bcf86cd799439011')MONGODB_URI=mongodb+srv://user:pass@cluster.mongodb.net/?retryWrites=true&w=majority
OPENAI_API_KEY=sk-... # or whichever embedding provider you useApache 2.0 — see LICENSE