Agent Memory Architecture
"Give the agent memory" is where most agent projects go off the rails. The word memory gets used for at least four different things, each of which has different storage, retrieval, and expiration semantics. Conflating them produces agents that hallucinate their own history.
Most of what looks like "the agent needs to remember X" turns out to mean "the agent needs to look up X." Use memory for this user's history with this agent. Use retrieval for everything else.
The four layers
Layer 1: Context window. Whatever's in the prompt on this turn. Strictly speaking, it's input, not memory. Budget this space carefully, because every token displaces another; Context Engineering is about that budget.
Layer 2: Working memory. The current session's state. Conversation turns so far, tool call results, intermediate scratch. Lives in your application's memory or a short-lived cache. Redis with a TTL works fine. When the session ends, this can go.
Layer 3: Episodic memory. Summarized records of past sessions. What did the user ask last week? What did the agent do? How did it go? These are compressed at end-of-session and indexed for retrieval. You retrieve them into Layer 1 when they're relevant to the current query.
Layer 4: Semantic memory. Durable, curated facts and patterns. This user prefers concise answers. This customer's account is enterprise-tier. Last time we tried tool X for this workflow, it failed for reason Y. Long-lived, deliberately edited, and accumulated. "Semantic" here is the cognitive-science term for general knowledge; it has nothing to do with semantic search.
The mistake to stop making
Dumping every conversation into a vector database and calling it memory. The database grows unboundedly. Retrieval quality degrades as noise accumulates. The agent starts recalling things that were never meant to be persisted (a user's frustrated aside, an incorrect early guess).
Production memory architecture is opinionated about what gets remembered. Layer 3 gets end-of-session summaries. Layer 4 gets facts explicitly promoted from Layer 3, either by rules or by a curator agent that runs periodically.
The retrieval flow
Retrieve from L3 and L4 concurrently. Rank by relevance. Fit into the context budget. Pass to the model with clear labels (Relevant past interaction:, Known fact about user:) so the model knows what kind of memory it's looking at.
Implementation choices per layer
Layer 2 (working): in-memory during the session, Redis with 24h TTL if you need multi-request continuity. Do not vector-embed working memory. It's not there long enough to earn the cost.
Layer 3 (episodic): vector store (pgvector, Qdrant, Pinecone) with summaries as documents and metadata for filtering (user_id, timestamp, session_outcome). Summarize at session end using a cheap model. Store the summary; the transcript can go.
Layer 4 (semantic): structured storage. Postgres tables for facts, with a schema. Do not vector-store this layer. You want deterministic retrieval (what do we know about user X?) rather than similarity search. Semantic memory that's just a vector store is memory you can't audit.
The curator pattern
Layer 4 grows through explicit promotion. A curator process (agent or scheduled job) reads recent Layer 3 entries and decides: is this a durable fact, a preference, a correction? If so, promote it to Layer 4 with a schema-typed record. If not, let it stay in Layer 3 and eventually expire.
Without curation, memory is an ever-growing bucket of stale strings.
Scoping and deletion
Layers 3 and 4 are per tenant and per user, always. A memory store shared across tenants is a leak waiting for the right prompt; Multi-Tenant AI covers the isolation side. They also need a deletion story. When a user asks to be forgotten, delete their Layer 3 entries and Layer 4 facts by key, which is one more reason Layer 4 lives in a table with a schema and a user id on every row.
What to log
Every retrieval: which layers, which documents, what relevance scores. Every context window assembly: what was included, what was dropped for budget. Every Layer 4 write: who or what promoted this fact, what's the source, when does it expire.
When the agent says something surprising, this trail tells you which memory it drew from.
Working on something like this?
Start a Conversation