AI Agent Memory is an open-source, multi-tenant agent memory server. Your agents connect over MCP, file verbatim memories, and recall them with hybrid semantic search — so every session builds on the last instead of starting from zero.
Free plan · 1,000 requests / month · no card required · social sign-in available
am_search · hybrid recall
$am_search"auth middleware decision"→ wing backend · room decisions drawer 0xC4F1…score 0.91"Token expiry check uses <= not <; off-by-one at the boundary."→ fused: vec 0.62 · bm25 0.27 · closet +0.12
36 / 37
MCP tools shipped
3-way
hybrid recall: vector · BM25 · closet
per-team
isolated vector store
€0
to start — 1,000 requests / month
The idea
What is AI agent memory?
An LLM forgets everything the moment its context window closes. AI agent memory fixes that: it is durable, long-term storage that an AI agent reads from and writes to across sessions — the decisions it made, the facts it learned, the threads it left open.
AI Agent Memory serves that store as a remote MCP server. Agents file verbatim memories — never lossy summaries — and recall the right ones on demand with hybrid semantic search. The next run picks up exactly where the last one stopped.
Why it matters
Why your AI agents need memory — even before you notice
Every large language model begins each session with total amnesia. The cost is easy to miss because it hides inside work you are already doing — re-explaining, re-deciding, re-discovering. Here is what that amnesia actually costs, and how to tell it is happening to you.
The default state of every AI agent is forgetting. A language model has no memory of its own: it reasons only over the tokens in its context window, and the instant that window fills or the chat ends, everything it "knew" is gone. It did not learn — it was briefly told, and then reset.
You rarely log this as a bug, because the model still sounds fluent and helpful in the moment. So the missing piece stays invisible: instead of noticing the amnesia, you quietly work around it — pasting the same context back in, re-stating the same rules, re-answering the same questions, session after session.
But the cost compounds. Every minute spent re-establishing what the agent already figured out yesterday is a minute not spent on the actual work — and the agent never gets better at your project, only repeatedly re-introduced to it. Long-term memory is the piece that turns those repeated introductions into accumulated knowledge.
You explain your project again every session
A stateless model remembers nothing once its window closes, so each new chat starts blank and you re-describe the same stack, conventions and goals you explained yesterday. With a shared memory the agent recalls that context itself — you brief it once, not once per session.
It re-opens a decision you already settled
Last week the agent helped you choose an approach and reasoned through the trade-offs; today, with no record of it, it proposes the option you already rejected. Filing the decision verbatim means the next run recalls not just what you chose but why, and builds on it instead of relitigating it.
The best answer disappeared when the window filled
A long session produces a sharp insight, then scrolls out of the context window and is gone — the model never actually learned it. Durable memory captures that insight the moment it happens and keeps it recallable forever, so a full window stops erasing your best work.
Two agents reach two contradictory conclusions
Run several agents — or several sessions of one — on the same problem and each works from its own private scratchpad, so they duplicate effort and disagree. One shared, multi-agent memory lets each build on what the others already found, so the fleet converges instead of colliding.
Switching models means onboarding from zero
Move from one model to another — or just open a fresh chat — and all the hard-won context stays trapped in the old conversation. Because memory lives outside the model in a portable store, any agent that connects inherits the whole history; the knowledge is yours, not the chat's.
It repeats a mistake it already learned from
An agent hits the same wrong turn it hit before — the flaky command, the misread config — because nothing carried the lesson forward. Write the correction down once and every later run recalls it, so a mistake gets made once instead of on a loop.
Give an agent durable memory and the pattern inverts: it recalls your stack, honours the decisions you already made, keeps its sharpest insights and learns from each mistake once. That is the difference between an assistant that resets every morning and one that compounds — and it is exactly what AI Agent Memory adds over one MCP endpoint. Start free and connect your first agent; it takes minutes, no card required.
In depth
Long-term, multi-agent, open source
The three ideas behind AI Agent Memory, explained in full.
What is long-term agent memory?
Long-term agent memory is durable storage that outlives a single model context window. A large language model is stateless between calls: when its context fills or the session ends, everything it "knew" is discarded. Long-term memory fixes that by externalising what the agent learns into a persistent store it reads from and writes to across sessions — the decisions it made, the facts it verified, the threads it left open.
AI Agent Memory keeps every memory verbatim — the exact text of what happened, never a lossy summary — and recalls the most relevant ones on demand with hybrid semantic search. So an agent on its hundredth run still has everything the first run learned, and nothing is silently dropped when the window closes.
What is multi-agent memory?
Multi-agent memory is a single shared memory that a whole fleet of agents reads and writes, instead of a private notebook per agent. When several agents — or many runs of one agent — collaborate on the same work, each one should build on what the others already discovered.
AI Agent Memory is multi-tenant: every team owns one isolated workspace, and every agent that connects with that team's token shares the same wings, drawers and knowledge graph. One agent files a decision; the next agent — a different model, a different session — recalls it by meaning. Memory stays isolated between teams, since each workspace has its own physically separate vector store, but it is fully shared within one.
Open-source agent memory you can self-host
AI Agent Memory is open source. The entire server is on GitHub, so you can read exactly how memories are stored, embedded and ranked — with no proprietary core to trust. A single Go binary migrates its schema, seeds a demo workspace and prints a one-time MCP bearer token; point any MCP client at POST /mcp and you have long-term memory running in two commands.
Use the hosted service or self-host the identical code — either way your agents' memory is portable and never locked in. An existing local memory palace imports over /import, and everything you file can be read back out.
The data model
The memory palace is the schema
Seven primitives compose every memory — borrowed from how humans file what they want to remember.
Wing
A project or context namespace — one isolated workspace per team.
Room
An aspect within a wing, like backend or decisions.
Drawer
One verbatim memory chunk plus rich metadata. Never summarised.
Closet
A topic and quote pointer index that boosts ranking — never a gate.
Hallway
A within-wing link between entities that co-occur in drawers.
Tunnel
A cross-wing link — authored, or auto-derived from a shared topic.
Knowledge graph
Temporal subject→predicate→object facts with validity windows.
How it works
One endpoint, fully isolated
Stateless Streamable-HTTP MCP: every request re-resolves its tenant from the bearer token, so the service scales out behind a load balancer.
1
Connect over MCP
Point any MCP client — Claude, your own agent — at POST /mcp with an Authorization: Bearer token.
2
Resolve the tenant
The token becomes a workspace in exactly one place. Every tool reads that tenant off the context and fails closed without it.
3
File and recall
Write verbatim drawers that get embedded and indexed, then recall them with hybrid search across the whole team's memory.
4
Stay isolated
SQLite is the relational source of truth; Qdrant holds per-tenant vectors, rebuildable from it. The transport is stateless, so it scales out.
Capabilities
Everything an agent needs to remember
36 of 37 MCP tools shipped — the write, recall, graph, knowledge-graph and skill families.
Recall
Hybrid semantic search
Vector similarity, BM25 lexical match and a closet boost, fused into one ranking — so agents recall by meaning and by exact term.
Isolation
Memory that can't leak
Every workspace gets its own Qdrant collection, named by a hash of the team id. A missing filter can't cross tenants — the data isn't even colocated.
Skills
Centralised, versioned skills
One shared source of truth for prompts and skills. Agents pull the latest with am_load_skill instead of copy-pasting local files.
Diary
An append-only agent diary
A timestamped journal per agent. Sessions thread across time, so the next run reads what the last one learned.
Knowledge
Temporal knowledge graph
Subject→predicate→object facts with validity windows, queryable as-of any point in time. Know what was true then, not just now.
Mining
Idempotent mining pipeline
am_mine turns raw text into chunked, embedded drawers plus a closet index — keyed by source, so re-running finishes rather than duplicates.
Graph
A navigable memory graph
Hallways link co-occurring entities; tunnels bridge wings. Traverse the graph to surface context a flat search would miss.
Migrate
Bring your mempalace
A read-only exporter streams an existing local mempalace into your workspace over /import — re-embedded server-side, graph rebuilt, fully idempotent.
Export
Own and export your data
Download everything a workspace holds as one self-contained SQLite file — scoped to your tenant, secrets redacted. Your BDAR/GDPR right of access and data portability, in one click.
Install
Build your install line
One command downloads the aiagentmemory binary and wires our MCP, the /M and /am commands, and the Stop hook into your agent — then prompts for your workspace token. Choose where it lands and what it brings; the command rewrites itself.
The installer downloads the aiagentmemory binary and wires the kit into your agent. Works on macOS, Linux and WSL.
There is no Windows binary — but the memory is a remote MCP server, so the assistant in your editor sets it up for you. Also the route for VS Code, Cursor and Claude Desktop on any OS.
Wires the kit into the agent you already run. One config, every project — start here if you are not sure.
A complete config of its own under ~/.sandboxes/ — commands, settings, MCP servers and token. Your global agent stays exactly as it was.
Claude Code is the default. The sandbox guide covers what changes for Codex and pi.
Names the folder under ~/.sandboxes/. Usually your project's name — it stays on your machine.
It will ask for your workspace token.Create a free workspace and copy the token first — leaving the prompt blank still installs the kit, but the memory itself stays disconnected until you add the token. Already have it in your shell? Set AGENTSMEMORY_TOKEN and the installer never asks.
prompt for your assistant
Read https://aiagentmemory.dev/install-memory-mcp and set up agentsmemory for me globally in this editor. Ask me for my workspace API token when you need it.
It reads the no-CLI guide, asks you for your workspace token, and writes the global MCP config for VS Code, Cursor or Claude Desktop. Everything lands in your own user config — there are no sandboxes without the CLI. Want the full kit on Windows? Install it inside WSL, where the bash installer runs unchanged.
Then work in it
Three commands, in the order you meet them: open the sandbox, pin the project to it, and the one line everyone on the team runs from then on.
open it
aiagentmemory run myproject
Starts the agent with this sandbox's config, commands, MCP and token. Your global install is untouched.
pin this repo
aiagentmemory init --sandbox myproject
Run it inside the project. It writes .aiagentmemory — commit that file. Append -- --model opus to record agent flags too.
every day after
aiagentmemory load
Opens the pinned project from any subdirectory. Teammates run this same line against their own sandbox.
The split is the point: .aiagentmemory holds the agent and its flags and gets committed, while your sandbox name is recorded in ~/.sandboxes/agents on this machine alone. A teammate whose sandbox is called something else runs the same load and gets the same setup.
With no sandbox recorded on any layer, load stops and tells you to run init. It never falls back to your global config: launching unpinned would defeat the point while looking like success.
A sandbox installed with --agent all is one directory any of the three CLIs can open — add --agent codex or --agent pi to run to open it as one of the others.
Or let Claude install it
Paste this into Claude Code (or any agent). It reads the install guide, asks you for your workspace token, and runs the install itself — nothing to copy by hand.
Read https://aiagentmemory.dev/claude-guide and install the agentsmemory Claude Code kit for me. When you need my workspace API token, ask me — I'll create one in the dashboard.
No CLI? Windows, VS Code, Cursor, Claude Desktop
The installer is a bash script, so it needs macOS, Linux or WSL — but the memory itself is a remote MCP server. Paste this into the assistant you already use and it reads the no-CLI guide, asks for your token, and writes the global config for whichever editor it is running in.
Read https://aiagentmemory.dev/install-memory-mcp and set up agentsmemory for me globally in this editor. Ask me for my workspace API token when you need it.
aiagentmemory install
Global — wire the kit into your existing ~/.claude.
aiagentmemory install --sandbox <name>
Isolated — a self-contained config under ~/.sandboxes/<name>.
aiagentmemory install --agent codex|pi|all
Same kit, other agent CLIs — see the sandbox guide.
aiagentmemory install --token <key>
Skip the prompt — same as exporting AGENTSMEMORY_TOKEN. Without either, the kit installs but the memory MCP is not registered.
aiagentmemory install --sandbox <name> --copy
Seed it from your global config — logins, MCP servers, plugins, skills, settings.
--copy brings your logins, MCP servers, plugins, skills and settings
History, logs and caches stay behind; nothing already there is overwritten
--shared-auth links credentials instead — log in once, every sandbox sees it
Per project — commit it
aiagentmemory init --sandbox acme -- --model opus
Records the agent and its flags in .aiagentmemory — commit that file
Your sandbox name stays on your machine, never in the repository
Everyone then runs load: same agent and flags, their own sandbox
Quick start
Running in two commands
Go 1.25+. The binary migrates an embedded schema, seeds a demo workspace, and prints a one-time MCP bearer token. Point any MCP client at POST /mcp with that token.
Search that understands meaning isn't free to run. Behind every stored drawer and every recall are three recurring costs — the compute, the electricity and the isolation that keep your agents' memory fast and private.
Compute
Every memory runs a model on a GPU
There is no keyword shortcut to meaning. Each drawer you file and each search you run is embedded by the bge-m3 model — a neural network whose matrix maths is only fast on a GPU, and a busy GPU draws hundreds of watts. That electricity is spent on every write and every recall, not once at signup.
Always-on
Vectors live in memory, day and night
So recall stays fast, Qdrant keeps each team's vectors hot in RAM on a server that runs 24/7 — powered, cooled and standing by whether you query once an hour or a thousand times a minute. You are renting a slice of always-on hardware, not just the seconds you spend searching.
Isolation
Your memory can't share a bill
Every workspace gets its own physically separate vector store, named by a hash of the team id — the guarantee that one team's memory can never leak into another's. That isolation is the point, and it means we can't amortise one giant shared index across everyone: your compute is genuinely yours.
The Free plan absorbs this for small, everyday use; Pro covers the always-on GPU and power for teams that lean on it. And because the server is open source, you can always self-host and pay your own hardware bill instead — your memory is never locked in.
Migrate
Bring your existing memory palace
Already running the local Python mempalace? A read-only exporter streams every drawer, diary entry, closet, knowledge-graph fact and tunnel into your workspace over /import. The server re-embeds each memory and rebuilds the graph — and the import is idempotent, so a re-run finishes rather than duplicates.
AI agent memory is persistent, long-term storage that lets an AI agent remember context across sessions — past decisions, facts and learnings — instead of starting cold every run. AI Agent Memory provides it as a remote MCP server: agents file verbatim drawers of memory and recall them later with semantic search.
Long-term agent memory is persistent storage that outlives a model's context window, letting an AI agent keep what it learned — decisions, facts, open threads — across sessions instead of forgetting when the window closes. AI Agent Memory stores each memory verbatim and recalls it later with hybrid semantic search, so later runs build on earlier ones.
Multi-agent memory is one shared memory store that a whole team of agents reads and writes, rather than a private notebook per agent. AI Agent Memory is multi-tenant: every agent connecting with a team's token shares the same wings, drawers and knowledge graph, so one agent recalls what another filed — while memory stays isolated between teams in physically separate vector stores.
Yes. AI Agent Memory is open-source software — the full Go server is on GitHub, so you can read exactly how memories are stored, embedded and ranked, and self-host it with no proprietary core. Run the hosted service or your own copy; your agents' memory is portable and never locked in.
An MCP (Model Context Protocol) memory server exposes memory operations — write, search, recall — as tools any MCP-compatible agent can call over HTTP. agentsmemory speaks stateless Streamable HTTP MCP, so Claude and other agents read and write memory with a bearer token.
Because a large language model is stateless: it only 'knows' what currently fits in its context window, and the moment that window fills or the session ends, everything is discarded — the model itself never learns from the conversation. That is why an agent re-asks about your project, re-opens settled decisions and repeats old mistakes. AI Agent Memory fixes it by storing what matters outside the model, in a persistent store the agent writes to and recalls from over MCP, so each session starts with everything the last one learned instead of a blank slate.
They externalise memory to a store outside the model's context window. agentsmemory embeds each memory with the bge-m3 model and indexes it in Qdrant, then ranks recall with a hybrid of vector similarity, BM25 and a closet boost — so agents retrieve the most relevant past context on demand.
Yes. Each workspace gets its own physically separate Qdrant collection, named by a hash of the team id. There is no shared collection to mis-filter, so memory cannot leak across tenants.
No, and that is deliberate. `aiagentmemory init` splits the record in two: the agent and its flags go into a .aiagentmemory file you commit, while your sandbox name is written to ~/.sandboxes/agents on your machine alone. A teammate clones the repository, runs init once with whatever they call their own sandbox, and `aiagentmemory load` then opens the same agent with the same flags inside their own isolated config. Nothing about your machine travels in the repository, so a committed launch config can never point someone at a sandbox that does not exist for them.
Yes. A read-only exporter streams an existing local Python mempalace — drawers, diary, closets, knowledge-graph facts and tunnels — into your workspace over /import. The server re-embeds each memory and rebuilds the graph, and the import is idempotent.
Yes. Any workspace member can download everything the workspace holds — drawers, diary, closets, knowledge-graph facts, tunnels, skills and account details — as a single self-contained SQLite file, scoped to your own tenant with credentials redacted. It is the BDAR (the EU GDPR) right of access and data portability: one click from the project page, and the file opens in any SQLite tool.
The Free plan is free forever with 1,000 requests per month. Teams running agents in production upgrade to Pro at €50 per month, or €500 per year (two months free).
Because hybrid semantic recall runs on real hardware. Every memory you file and every search you run is embedded by the bge-m3 model on a GPU that draws hundreds of watts, and each team's vectors are kept hot in a Qdrant store on a server that runs 24/7 — physically isolated per team, so the compute can't be shared. The Free plan absorbs that cost for small use; Pro covers the always-on GPU and electricity for teams that lean on it. And because the server is open source, you can always self-host and pay your own hardware bill instead.
Give your agents a memory that lasts.
Spin up a free workspace and connect your first agent in minutes.