Part of Language AI Handbook
Explains how AI agents manage short-term context and long-term memory. Examines vector databases, retrieval algorithms.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Agent Memory Systems: Architecture and Implementation
An agent without memory is like a goldfish swimming through a castle, rediscovering the same towers and turrets with every lap around the bowl. It can execute tools and respond to immediate queries, but it cannot learn from past interactions, maintain continuity across conversations, or accumulate knowledge over time. Memory turns a stateless function caller into a persistent, learning system able to complex, multi-step tasks that span hours, days, or even years.
Consider the frustration of explaining your project requirements to a coding assistant, only to have it forget the tech stack you specified moments later. Or imagine a personal assistant that cannot remember your dietary preferences from one day to the next. Without memory, every interaction resets to zero, forcing you to repeat context and rebuild state constantly. This friction makes advanced collaboration impossible, limiting agents to simple, atomic tasks that fit entirely within a single context window. The agent is perpetually a stranger to you, no matter how many hours you have spent working together.
In this chapter, we explore how modern AI agents manage information across different timescales. We begin with short-term memory, the immediate context window that holds the current conversation and recent tool outputs. We then examine long-term memory systems that compress and index information for retrieval from thousands of past interactions. Along the way, we investigate the algorithms that decide what to remember, what to forget, and how to synthesize scattered observations into coherent knowledge.
Building on our understanding of function calling from the previous chapter, we will see how memory lets agents to maintain state across tool invocations and learn from their own execution history. We will trace the journey from raw sensory input to consolidated knowledge, examining the engineering trade-offs at each layer of the memory stack. By the end, you will have both the conceptual framework and a working implementation to build memory-enabled agents from scratch.
The Memory Hierarchy
Human cognition organizes memory into distinct temporal layers: sensory memory lasting milliseconds, working memory holding a handful of items for seconds, and long-term memory storing a lifetime of experiences. AI agents mirror this hierarchy, though their implementation differs materially from biological neural systems. Where human memory relies on synaptic plasticity, protein synthesis, and hippocampal consolidation, artificial memory depends on vector embeddings, database indexes, and compression algorithms.
This parallel structure emerges not from biological mimicry alone, but from computational necessity. Just as humans cannot consciously attend to every sensory stimulus simultaneously, language models cannot process infinite context in parallel. The hierarchy solves a basic resource allocation problem: how to make the most relevant information available at the lowest latency while preserving large archives of potentially useful history at higher latency and lower cost.
Understanding this hierarchy is the key to designing effective agents. Each layer of the stack differs in capacity and latency as well as fidelity and cost. Choosing which information belongs at which layer, and how information flows between layers, determines whether an agent is snappy and accurate or slow and confused. This is the central engineering problem of agent memory, and it mirrors problems that operating system designers and database architects have grappled with for decades.
Short-Term Memory: The Context Window
Short-term memory in language models is deceptively simple yet fundamentally constrained. It consists of the tokens currently fed into the model's context window, typically ranging from 4,000 to 200,000 tokens depending on the architecture. As we discussed in Part XVIII on long context, extending this window remains an active research area fraught with computational and theoretical challenges.
For agents, the context window is working memory, holding:
- The current user query and conversation history
- Recent tool outputs and observations
- Active plans and subgoals
- Temporary variables and scratchpad calculations
However, this memory is ephemeral. Once a token slides beyond the context window, it is effectively forgotten unless explicitly stored elsewhere. This creates a necessary challenge for long-running agents: the "amnesia problem," where early conversation details or initial task specifications vanish as new information arrives. Imagine an agent tasked with refactoring a codebase over several hours. If the initial instructions about coding standards fall out of the context window, the agent might inadvertently violate those standards in later edits, undermining the entire task.
The constraints extend beyond mere token limits. Even within the context window, information suffers from positional bias. Studies show that language models often struggle to retrieve facts located in the middle of long contexts, exhibiting a U-shaped attention curve where information at the beginning and end of the window is most accessible. This "lost in the middle" phenomenon means that necessary instructions buried in the middle of a long conversation may be effectively invisible to the model, even if they technically fit within the token budget. A model might confidently ignore a constraint stated in the middle of a ten-page context while correctly citing instructions from the opening paragraph and the most recent message.

The practical implication is significant for agent designers. Placing the most necessary instructions both at the start of the system prompt and immediately before the current task increases the probability the model will respect them. Some production systems implement "instruction pinning," where necessary constraints are automatically duplicated at the end of the context window before each generation call, exploiting the recency effect to counteract positional degradation. Others rotate important instructions through the context, surfacing them periodically rather than letting them sink to the forgotten middle.
Long-Term Memory: Beyond the Context Window
When context windows prove insufficient, agents turn to external memory systems. These store information outside the model's parameters, typically in vector databases, knowledge graphs, or traditional structured storage. Long-term memory persists across sessions, letting agents to recognize you, recall facts from previous conversations, and build cumulative expertise. Unlike the fixed weights of the neural network, which encode general world knowledge during training, long-term memory stores personal details and temporally grounded, specific information acquired during deployment.
Long-term memory systems generally fall into three categories.
Semantic memory stores factual knowledge about the world as dense vector embeddings in vector databases. This mirrors human declarative memory, encoding concepts like "Paris is the capital of France" or "Python uses indentation for code blocks." Semantic memory captures the "what" of knowledge: definitions and relationships together with attributes and categories. When an agent retrieves the API documentation for a specific library or recalls that a user prefers dark mode interfaces, it draws upon semantic memory. The key property of semantic memory is generalizability: a fact stored in semantic memory can be retrieved in response to many different queries that relate to the same concept, even if those queries use different words.
Episodic memory preserves specific experiences and interaction histories as timestamped events with contextual metadata. This includes past conversations, previous tool executions, and their outcomes. Episodic memory captures the "when" and "where" of experience: the specific debugging session where a particular error was resolved, or the conversation last Tuesday where the user changed their project requirements. This temporal grounding allows agents to maintain narrative continuity and learn from specific historical trajectories rather than just general facts. Unlike semantic memory, which is atemporal, episodic memory is always anchored to a moment in time, which makes it particularly useful for reasoning about causality and sequence.
Procedural memory encodes learned skills and behavioral patterns, often stored as few-shot examples or refined system prompts based on successful past trajectories. This corresponds to human "muscle memory" or expertise: the knowledge of how to do things rather than what things are. Procedural memory might manifest as a library of successful tool-calling patterns, refined prompting strategies for specific types of queries, or learned heuristics for planning complex tasks. Unlike semantic memory, which stores explicit facts, procedural memory shapes behavior through pattern recognition and conditioned responses. An agent with strong procedural memory for debugging Python code will approach a traceback differently from one that lacks this history, even if both agents have identical semantic knowledge of Python syntax.
The distinction between these three types matters for system design because each type demands different storage formats, retrieval mechanisms, and update strategies. A vector database excels for semantic and episodic retrieval, but procedural memory is better represented as parameterized examples or updated few-shot demonstrations. Effective agent architectures recognize these distinctions and route different types of information to appropriate storage back-ends.
Memory Retrieval: Finding the Right Information
Storing memories is straightforward; retrieving the relevant ones at the right moment remains the central challenge. Unlike database queries with precise keys, memory retrieval in agents is approximate, context-dependent, and often probabilistic. The system must infer relevance from the current situation, sifting through thousands of potential memories to retrieve the handful that matter for the present query.
This challenge mirrors the human experience of tip-of-the-tongue states or involuntary memory. Sometimes we retrieve exactly what we need, sometimes we are flooded with irrelevant associations, and sometimes we cannot access information we know we possess. For agents, these failures manifest as hallucinations based on wrong memories, missed opportunities to use past experience, or costly retrieval of useless information that crowds out the relevant context.
The basic difficulty is that relevance is not a fixed property of a memory. The same memory might be highly relevant for one query and completely irrelevant for another. A record of a user's preference for dark-mode interfaces is irrelevant when designing an algorithm but needed when generating a dashboard. A retrieval system must therefore evaluate each memory relative to the current query, not in isolation. This relational nature of relevance is why naive keyword search fails for agent memory: a query about "rendering" might need to retrieve memories about the user's display preferences, their project's rendering engine, their past frustration with slow load times, or all three simultaneously.
Vector Similarity Search
The dominant retrieval mechanism for semantic memory uses dense vector embeddings. As we explored in the retrieval-augmented generation chapters, we encode memories and queries into the same latent space, then retrieve the nearest neighbors. The intuition is simple: if two pieces of text have similar meanings, their vector representations should point in similar directions in the high-dimensional embedding space.
For an agent, this works as follows: when you ask, "What did we decide about the database schema last week?", the agent embeds this query, searches the episodic memory store for semantically similar stored interactions, and retrieves the relevant conversation fragments. The embedding model captures keyword overlap alongside conceptual similarity: a query about "database schema" might retrieve memories mentioning "table structure" or "SQL design" even if the exact words do not match.
However, pure similarity search has important limitations. The most similar memory is not always the most useful one. Consider an agent managing a complex coding project: retrieving the most semantically similar past bug fix might miss a more recent architectural decision that superseded the old approach. Similarly, similarity search struggles with negation and temporal reasoning. A query about "solutions that don't use recursion" might retrieve memories about recursion simply because the topic is semantically related. The embedding space encodes what content is about, not whether it should be used or avoided.
There is also the problem of semantic drift over time. If a user's preferences or requirements change, the vector store will contain both old and new preferences, and similarity search alone cannot distinguish which is current. This is compounded by the fact that embedding models tend to cluster semantically related content tightly regardless of temporal distance, meaning a preference stated six months ago and one stated yesterday may receive nearly identical similarity scores for a relevant query. Temporal metadata becomes indispensable for breaking these ties.
Multi-Factor Scoring
Sophisticated memory systems combine multiple signals to rank candidates, recognizing that relevance is multidimensional. A memory might be relevant because it is semantically similar to the current query, because it occurred recently in the conversation, or because it was marked as important when first stored. No single signal works for all scenarios.
The multi-factor scoring formula combines these signals with tunable weights:
where:
- : the memory candidate being scored
- : the current query embedding
- : the cosine similarity between memory and query embeddings, measuring conceptual overlap between the candidate and the current need
- : a time-decay factor, typically computed as where controls how fast older memories lose relevance
- : a learned or heuristic score showing how significant the memory is, irrespective of the current query
- : weight for semantic similarity, controlling how much conceptual match matters
- : weight for recency, controlling how strongly the system favors recent memories
- : weight for importance, controlling how much stored significance raises a candidate's score
The weights , , and tune the retrieval behavior for different contexts. High creates a recency bias suitable for ongoing task tracking. This keeps recent context is not drowned out by older but semantically similar memories. This is important for maintaining conversational coherence: if you just mentioned switching from Python to JavaScript, recent memories about JavaScript should outrank older Python-related memories even if the query is about "coding best practices."
Conversely, high favors semantic relevance for knowledge retrieval, appropriate when searching for factual information where temporal proximity matters less than conceptual alignment. When you ask, "What is our company's vacation policy?", the answer from two years ago is likely still valid regardless of more recent conversations about project deadlines. The weight configuration in effect encodes assumptions about how quickly knowledge in the domain expires.


The recency term uses exponential decay because both human and artificial memory exhibit forgetting curves where information loses accessibility over time. The decay rate can be tuned to match the expected volatility of the domain. For rapidly changing environments like software development, a faster decay (higher ) ensures that stale architecture decisions do not compete with current ones. For stable domains like legal documentation, slower decay preserves access to older but still valid precedents. Setting is therefore a domain-modeling decision as well as a hyperparameter choice.

Importance Estimation
Determining which memories deserve long-term retention requires judging importance at storage time. Not all experiences are equally worthy of permanent storage. Just as humans naturally forget mundane details but retain emotionally significant events, agents must triage memories to avoid filling storage with noise.
Several heuristics are effective for estimating importance:
-
Surprise: Memories where the outcome diverged materially from expectations carry high information content. If an agent attempts a tool call and receives an unexpected error, or if a user reacts with surprise to a suggestion, these moments signal learning opportunities. Surprise indicates a gap in the agent's world model, making the memory useful for future error correction. Technically, surprise can be estimated by comparing the model's predicted outcome with the observed outcome, rewarding large discrepancies with higher importance scores.
-
Emotional valence: In human-facing agents, emotionally charged interactions often warrant preservation. Explicit expressions of frustration, delight, or urgency indicate that something significant occurred. While AI systems do not experience emotions, they can detect emotional language in user responses and increase the importance of those interactions. A user saying "this is exactly what I needed, save this!" is a strong signal that the current context should be preserved.
-
Action consequences: Memories preceding significant state changes or tool executions deserve retention. If a memory describes the command that deployed code to production or the query that modified a database, the stakes of forgetting are high. These causal links between actions and outcomes form the basis of procedural learning. An agent that remembers which action caused which outcome can avoid repeating costly mistakes.
-
User annotations: Explicit user feedback showing "remember this" or bookmarking gives ground truth for importance. Some systems allow users to star or flag specific information, creating a supervised signal that can train importance classifiers. This human-in-the-loop approach ensures that the importance model shows actual user needs rather than proxy heuristics.
-
Frequency of reference: Information that appears repeatedly across multiple turns, or that the user circles back to after discussing other things, is probably important. A topic that keeps resurfacing is one the user cares deeply about, even if they never explicitly mark it as important.
Some systems train a small classifier to predict importance scores based on these features, while others use the LLM itself to evaluate importance during the storage phase. The LLM-based approach might prompt the model with the memory content and ask it to rate importance on a scale of 1 to 5, or to generate a brief justification for why the memory should be kept. This uses the model's pre-trained understanding of what information is likely to be useful in the future, though it adds latency to the storage process. The classification approach is faster but requires labeled training data and may miss nuances that a language model captures naturally.
Importance is a property of the memory itself, estimated at storage time without knowledge of future queries. Relevance is a property of the memory-query pair, computed at retrieval time. A memory can be highly important but temporarily irrelevant (e.g., the user's emergency contact is necessary information but rarely retrieved), or highly relevant but unimportant (e.g., a casual mention of coffee preferences that happens to match the current query). Effective memory systems track both separately, using importance to guide retention decisions and relevance to guide retrieval ranking.
Memory Summarization: Compression and Consolidation
Raw conversation logs and tool outputs consume tokens rapidly. A detailed hour-long interaction might exhaust the context window of even the most capable models. Memory summarization compresses this history into dense representations that preserve needed information while discarding noise. This process is analogous to sleep-dependent memory consolidation in humans, where the brain transfers information from the hippocampus to the cortex, extracting patterns and discarding ephemeral details.
Without summarization, agents face a painful trade-off: either truncate the conversation history, losing early context entirely, or include everything and hit token limits that prevent new reasoning. Summarization offers a middle path, distilling the essence of long interactions into compact forms that fit within constraints while retaining actionable information. A well-summarized interaction might compress ten thousand tokens of dialogue into two hundred tokens that capture key decisions and preferences alongside relevant facts, freeing the remaining budget for current reasoning.
Progressive Summarization
The simplest approach maintains a rolling summary. After every turns or tokens, the agent processes the recent history through a summarization prompt:
Summarize the following conversation segment, preserving:
- Key facts and decisions
- Active tasks and their statuses
- User preferences expressed
- Any errors or unexpected outcomes
Segment: [conversation text]
This summary replaces the raw text in the context window, freeing space for new interactions. The agent might keep the most recent five turns in full detail while summarizing everything before that point into a condensed paragraph. This creates a sliding window of high-fidelity recent memory backed by increasingly abstracted historical context.
However, aggressive summarization risks losing necessary details, creating a trade-off between compression ratio and information fidelity. A summary stating "we discussed authentication" loses the nuance that "we specifically decided to use JWT with a custom refresh strategy rather than session cookies." Some systems address this by maintaining multiple summary levels: a detailed summary of the last ten turns, a medium-level summary of the last fifty turns, and a high-level abstract of everything before that. When retrieving context, the agent can zoom in or out depending on the specificity required. This multi-resolution approach mirrors how human memory stores both the gist of an experience (we went on vacation to Italy) and specific vivid details (the taste of gelato in a particular piazza).
The quality of summarization depends heavily on the instructions given to the summarizer. Generic prompts produce generic summaries that might omit domain-specific details the agent needs. A coding agent benefits from prompts that explicitly ask to preserve function signatures, error messages, and architectural decisions. A customer service agent needs prompts that capture complaint categories, resolution status, and customer sentiment. Tailoring the summarization prompt to the domain can sharply improve how much actionable information survives compression.
Hierarchical Memory Structures
More advanced systems organize memory hierarchically, similar to how computer memory uses caches and RAM. MemGPT, for instance, implements a tiered architecture that explicitly manages information flow between storage layers with different capacity and latency characteristics.
The tiers form a spectrum from fastest-and-smallest to slowest-and-largest:
Core memory holds always-present system instructions and user preferences, analogous to CPU registers. This small, high-speed store contains the most necessary information that the agent needs for every turn: the user's name, the current project context, and any hard constraints that must never be violated. Core memory is never evicted and is always included in the context window without requiring any retrieval operation.
Working context contains the recent conversation within the context window, analogous to RAM. This holds the active conversation thread and recent tool outputs in their original form. It gives high-fidelity access to recent history but is limited by the context window size. When working context fills up, older entries must be either discarded or moved to longer-term storage.
External storage archives conversations and documents on disk-equivalent storage. This large, slow repository holds the full history of past interactions, reference documents, and background knowledge. Retrieval from this tier requires explicit search operations using the vector similarity and multi-factor scoring techniques described earlier.
The agent explicitly manages movement between tiers through function calls: search_archives(query), add_to_core(key, value), and summarize_and_store(segment). This explicit memory management, while adding complexity, allows the agent to reason about its own memory constraints and optimize information placement strategically. For example, when learning a necessary API key or user constraint, the agent can explicitly move it into core memory for immediate access, while archiving routine conversation to external storage. This meta-cognitive capacity, the ability to think about and manage one's own memory, is one of the more intriguing capabilities of advanced agent architectures.
Knowledge Graph Extraction
Rather than compressing text into shorter text, some systems extract structured knowledge graphs from conversation history. The agent parses conversations into subject-relation-object triples. "Alice prefers Python over Java" becomes the triple (Alice, prefers, Python) and the triple (Alice, avoids, Java). This transformation converts unstructured narrative into structured data that supports precise querying and logical inference.
These triples populate a graph database where nodes stand for entities and edges stand for relationships. Retrieval then involves graph traversal: finding Alice's node and following preference edges yields her language choices without scanning conversation logs. If you ask, "What programming languages does Alice like?", the system queries the graph for all objects connected to "Alice" via "prefers" edges, returning a clean list without wading through paragraphs of dialogue.
This approach excels for factual, relational knowledge but struggles with procedural information and fine-grained context that resists triple extraction. Statements like "Alice prefers Python for data science but uses JavaScript for web development" decompose cleanly into conditional triples, but procedural knowledge like "When debugging async code, Alice starts by checking for race conditions" loses its conditional structure when flattened into simple relations. The temporal dimension is also lost: a knowledge graph stores that Alice prefers Python, but cannot naturally stand for that she used to prefer Ruby and switched six months ago after a frustrating project.
Hybrid systems often combine graph extraction for entity relationships with narrative summarization for procedural and contextual information. The graph handles the "who knows what" questions efficiently, while summarized narrative handles the "how and why things happened" questions that require more context to answer correctly. This combination uses the precision of structured data where that precision is achievable, and falls back to the flexibility of natural language where it is not.
Advanced Memory Architectures
Modern agent systems combine these primitives into advanced memory management schemes that rival the complexity of operating systems managing virtual memory. These architectures address specific limitations of simple vector stores, letting agents to learn from mistakes, integrate retrieval into the generative process itself, and combine structured and unstructured knowledge smoothly.
Reflexion: Self-Reflective Memory
The Reflexion architecture introduces a feedback loop where the agent evaluates its own performance and stores verbal critiques in episodic memory. When a task fails, the agent generates a textual reflection: "I should have checked the API documentation before assuming the parameter format. Next time, verify the schema first."
These self-critiques accumulate in a memory stream. Before attempting similar tasks, the agent retrieves relevant past reflections, effectively learning from mistakes without gradient updates. This creates a form of in-context learning that persists across sessions. Unlike fine-tuning, which permanently modifies model weights, Reflexion keeps the learning explicit and inspectable: a human can read the memory stream and understand exactly what the agent has learned from its failures. This transparency is a real advantage in production settings where debugging agent behavior requires understanding the agent's actions and its reasons for choosing them.
The process typically works in three stages. First, the agent attempts a task and receives feedback on whether it succeeded. Second, if it failed, the agent generates a reflection explaining what went wrong and how to avoid the same mistake. Third, this reflection is stored in memory with high importance weighting. On subsequent attempts at similar tasks, the retrieval system surfaces these reflections as part of the context, conditioning the agent to avoid previous mistakes.
This design mimics human metacognition, where we explicitly think about our thinking and encode lessons learned from experience. A student who shows on why they got an exam question wrong, writes a note to themselves about the conceptual gap, and reviews that note before the next exam is performing the same cognitive operation that Reflexion operationalizes for AI agents. The key insight is that explicit, verbalized self-critique is a more sample-efficient learning signal than implicit reward signals, because the critique contains semantic information about what specifically went wrong and how to fix it.
Memory Attention Mechanisms
Rather than retrieving memories before generation as a preprocessing step, some architectures integrate retrieval into the attention mechanism itself. The model attends to the current context and a large external memory bank, treating memory retrieval as a differentiable operation within the forward pass.
This approach, seen in models like Longformer and BigBird, treats memory as an extension of the context window. The retrieval process concatenates keys and values from both the current context and external memory before applying attention:
where:
- : the concatenated key vectors from both the context window and external memory
- : the concatenated value vectors from both the context window and external memory
- : key vectors from the immediate context window tokens
- : key vectors retrieved from external memory and appended to the sequence
- : value vectors from the immediate context window
- : value vectors retrieved from external memory
The attention weights are then computed over this combined set using scaled dot-product attention:
where:
- : the query matrix derived from the current input token
- : the dimension of the key vectors, used as a scaling factor to prevent the dot products from growing too large before the softmax
By concatenating memory vectors with context vectors, the model treats retrieved memories as extensions of the context window, letting it to dynamically focus on relevant past information during generation through the standard attention mechanism. This blurs the line between short-term and long-term memory, creating a unified attention surface over both recent and historical information.
The key advantage is differentiability: the retrieval process becomes part of the forward pass, letting end-to-end training of both the memory representation and the attention weights. However, this comes at significant computational cost. Attending over large memory banks requires specialized sparse attention patterns or approximate methods to remain tractable. The full attention cost over millions of memory vectors is prohibitive, motivating approximations like locality-sensitive hashing or hierarchical clustering to restrict attention to a promising subset of candidates.
Vector Database vs. Symbolic Memory
The choice between vector-based semantic memory and symbolic structured storage depends on the retrieval pattern and the nature of the information being stored. Each approach offers distinct advantages that make them suitable for different types of information and query patterns.
Vector databases excel when:
- Queries are natural language questions lacking precise structure
- Memories are unstructured text such as conversation logs or documents
- Approximate matching is acceptable, and the goal is conceptual similarity rather than exact identity
- The retrieval goal is "find similar experiences" or "find relevant context"
Symbolic stores (SQL, knowledge graphs) excel when:
- Queries involve precise constraints (e.g., "find tasks completed yesterday")
- Memories have clear relational structure with defined schemas
- Exact matching is required, particularly for identifiers, dates, or categorical values
- The retrieval goal is "find specific facts" or "filter by exact criteria"
Hybrid architectures use vectors for initial retrieval and symbolic filters for precision, combining the semantic flexibility of embeddings with the exactness of structured queries. For example, an agent might use vector similarity to find memories related to "database performance issues," then apply a SQL filter to select only those from the last month involving PostgreSQL specifically. This two-stage retrieval uses the strengths of both approaches: the vector stage captures semantic nuance that SQL queries would miss, while the symbolic stage filters out false positives that semantic similarity might erroneously include.
The architectural decision often depends on update frequency as well. Symbolic stores typically enforce schemas that make them rigid but consistent, while vector stores accommodate arbitrary new content types without structural changes. For rapidly evolving domains where the structure of knowledge changes frequently, vector stores offer flexibility. For stable domains with well-defined ontologies, symbolic stores give reliability and interpretability. Many production systems start with vector stores for their flexibility and selectively introduce symbolic components as patterns in the data stabilize and precise querying becomes important.
Worked Example: A Coding Assistant with Memory
Consider a coding agent helping you refactor a large codebase over several sessions spanning multiple days. Each session builds on the previous one, requiring advanced memory management to maintain continuity and avoid repeating earlier mistakes.
Session 1 starts fresh. The agent analyzes the authentication module, noting the use of JWT tokens with a custom refresh strategy. It stores this in episodic memory with high importance due to security implications. The agent also extracts semantic knowledge as triples: (authentication, uses, JWT) and (JWT, has_property, custom_refresh). This structured extraction allows precise retrieval later when questions arise about authentication specifically, even if the query uses different terminology like "login system" or "session management."
Session 2, three days later: you ask, "How should we handle the API rate limiting?" The agent retrieves memories about the authentication system, recognizing that rate limiting rules often interact with authentication tiers. By recalling the JWT refresh patterns from the previous session, it suggests rate limits that respect the authentication state transitions, avoiding scenarios where aggressive rate limiting interferes with legitimate token refreshes. Without this memory, the agent might suggest naive rate limiting that breaks the authentication flow whenever a client needs to refresh an expired token. The multi-factor scoring here works in favor of the correct retrieval: the authentication memories have high importance scores (security-necessary decisions), and they score well on similarity despite not containing the words "rate limiting," because the embedding model captures the conceptual connection between authentication and API access control.
Session 3: the context window fills with current file contents. The agent summarizes the previous refactoring steps, maintaining only the high-level plan in working memory while storing detailed function signatures in the vector database. This compression ensures that the agent can still see the overall architecture (decoupling the user service from the payment service) without the token overhead of every specific function change. When you ask for details about a specific function modified earlier, the agent retrieves the full signature from the vector store on demand. This gives high fidelity where it matters without burning tokens on information not currently needed.
Session 4: an error occurs during testing. The agent retrieves similar past errors from episodic memory, finds a previous fix involving asynchronous token refresh, and applies the same pattern, avoiding a bug it previously encountered. This shows procedural learning: the agent stores the fact that an error occurred, learns the pattern of the solution, and applies it to a new but structurally related context.
This continuity shows how memory turns isolated tool calls into coherent, learning assistance. Without memory, each session would begin with the agent re-exploring the codebase, re-learning the architecture, and potentially repeating previous mistakes. With memory, the agent accumulates expertise specific to this codebase and this developer's preferences, becoming more effective over time. The improvement is not from retraining the underlying model but from intelligently managing the information available at inference time.
Code Implementation: Building a Memory-Enabled Agent
Let's implement a simple but functional memory system for an agent. We will use a vector store for semantic retrieval and implement importance scoring and summarization. The goal is to show the core mechanisms described above in working code that you can extend.
We start with the foundational data structures: an embedding function and a memory entry class.
from datetime import datetime
from typing import Dict, Optional
import numpy as np
# We'll use a simple embedding function for demonstration.
# In production, use sentence-transformers or OpenAI embeddings.
def simple_embed(text: str, dim: int = 64) -> np.ndarray:
"""Simple bag-of-characters embedding for demonstration."""
vec = np.zeros(dim)
for i, char in enumerate(text.lower()):
vec[ord(char) % dim] += 1
# Normalize to unit length for cosine similarity
norm = np.linalg.norm(vec)
return vec / norm if norm > 0 else vec
class Memory:
"""Single memory entry with all metadata needed for retrieval scoring."""
def __init__(
self,
content: str,
memory_type: str = "episodic",
importance: float = 1.0,
metadata: Optional[Dict] = None,
):
self.content = content
self.memory_type = memory_type # 'episodic', 'semantic', 'procedural'
self.importance = importance
self.created_at = datetime.now()
self.embedding = simple_embed(content)
self.metadata = metadata or {}
self.access_count = 0
self.last_accessed = datetime.now()
def to_dict(self) -> Dict:
return {
"content": self.content[:100] + "..."
if len(self.content) > 100
else self.content,
"type": self.memory_type,
"importance": self.importance,
"age_hours": (datetime.now() - self.created_at).total_seconds()
/ 3600,
"access_count": self.access_count,
}The Memory class stores both the raw content and a pre-computed embedding, avoiding the cost of re-embedding at retrieval time. The access_count field tracks how often a memory has been retrieved, letting promotion heuristics where frequently accessed memories receive importance boosts before consolidation.
Next, we implement the AgentMemory class that manages the two-tier hierarchy of working memory and long-term storage:
from typing import Dict, List, Optional, Tuple
class AgentMemory:
"""Hierarchical memory system implementing working memory and long-term storage."""
def __init__(self, capacity: int = 1000, context_size: int = 5):
self.long_term_memory: List[Memory] = []
self.working_memory: List[Memory] = [] # Recent, unconsolidated
self.capacity = capacity
self.context_size = context_size # Items to keep in working memory
# Retrieval scoring weights (must sum to 1.0)
self.alpha = 0.6 # Similarity weight
self.beta = 0.3 # Recency weight
self.gamma = 0.1 # Importance weight
def add_memory(
self,
content: str,
memory_type: str = "episodic",
importance: float = 1.0,
metadata: Optional[Dict] = None,
):
"""Add a new memory to working memory, triggering consolidation if needed."""
mem = Memory(content, memory_type, importance, metadata)
self.working_memory.append(mem)
# Consolidate working memory if it grows too large
if len(self.working_memory) > self.context_size * 2:
self._consolidate()
def _consolidate(self):
"""Move older working memories to long-term storage."""
# Preserve the most recent items in working memory
to_archive = self.working_memory[: -self.context_size]
self.working_memory = self.working_memory[-self.context_size :]
for mem in to_archive:
# Boost importance for frequently accessed memories before archiving
if mem.access_count > 2:
mem.importance *= 1.5
self.long_term_memory.append(mem)
# Evict from long-term memory if over capacity
if len(self.long_term_memory) > self.capacity:
self._evict_memories()
def _evict_memories(self):
"""Remove lowest-scoring memories to respect capacity constraints."""
now = datetime.now()
scores = []
for i, mem in enumerate(self.long_term_memory):
age_days = (now - mem.created_at).days
recency_score = np.exp(-0.1 * age_days)
# Eviction score weights retention factors: importance most, then access, then recency
retention_score = (
0.3 * recency_score
+ 0.5 * min(mem.importance / 5.0, 1.0)
+ 0.2 * min(mem.access_count / 10.0, 1.0)
)
scores.append((retention_score, i))
# Keep the highest-scoring memories (capacity - 100 to create buffer)
scores.sort()
keep_indices = set(idx for _, idx in scores[-(self.capacity - 100) :])
self.long_term_memory = [
m for i, m in enumerate(self.long_term_memory) if i in keep_indices
]
def retrieve(self, query: str, k: int = 3) -> List[Tuple[Memory, float]]:
"""Retrieve top-k memories using multi-factor scoring."""
query_vec = simple_embed(query)
now = datetime.now()
scored_memories = []
all_memories = self.working_memory + self.long_term_memory
for mem in all_memories:
# Semantic similarity via cosine dot product on unit vectors
similarity = np.dot(query_vec, mem.embedding)
# Recency: exponential decay based on age in days
age_days = (now - mem.created_at).total_seconds() / (24 * 3600)
recency = np.exp(-0.1 * age_days)
# Importance: normalized to [0, 1]
importance = min(mem.importance / 5.0, 1.0)
# Combined score from the multi-factor formula
score = (
self.alpha * similarity
+ self.beta * recency
+ self.gamma * importance
)
scored_memories.append((mem, score))
scored_memories.sort(key=lambda x: x[1], reverse=True)
# Update access statistics for retrieved memories
for mem, _ in scored_memories[:k]:
mem.access_count += 1
mem.last_accessed = now
return scored_memories[:k]The _evict_memories method implements a retention policy that favors important and frequently accessed memories over purely recent ones. This mimics how human long-term memory reinforces frequently recalled information while letting less-accessed material to fade. The buffer of 100 slots below capacity prevents constant eviction when the store hovers near its limit.
Now we implement the agent layer that adds importance estimation and periodic summarization:
import re
class SummarizingAgent:
"""Agent with memory capabilities including importance scoring and summarization."""
def __init__(self):
self.memory = AgentMemory(capacity=100, context_size=3)
self.conversation_turns = 0
def interact(self, user_input: str, system_response: str = None):
"""Process a conversation turn, storing both sides with estimated importance."""
self.memory.add_memory(
content=f"User: {user_input}",
memory_type="episodic",
importance=self._estimate_importance(user_input),
)
if system_response:
self.memory.add_memory(
content=f"Assistant: {system_response}",
memory_type="episodic",
importance=1.0,
)
self.conversation_turns += 1
# Every 5 turns, create a compressed summary of recent context
if self.conversation_turns % 5 == 0:
self._create_summary()
def _estimate_importance(self, text: str) -> float:
"""Assign importance scores using heuristic signals."""
importance = 1.0
# Questions indicate information gaps that should be preserved
if "?" in text:
importance += 1.0
# Task-framing language suggests high-stakes information
if any(
word in text.lower()
for word in ["need", "must", "should", "task", "goal"]
):
importance += 1.5
# Personal information is useful for personalization
if any(
word in text.lower()
for word in ["i like", "i prefer", "my name", "i am"]
):
importance += 2.0
return min(importance, 5.0)
def _create_summary(self):
"""Compress recent context into a single high-importance summary memory."""
recent = self.memory.working_memory[-5:]
if len(recent) < 3:
return
# In production, this would call an LLM with a summarization prompt.
# Here we extract salient terms as a proxy for summarization.
topics = set()
for mem in recent:
words = re.findall(r'"([^"]*)"|\b([A-Z][a-z]+)\b', mem.content)
topics.update(w for pair in words for w in pair if w)
summary = f"Summary: Discussed {', '.join(sorted(topics)[:3])}. Key points from last 5 turns."
self.memory.add_memory(
content=summary,
memory_type="semantic",
importance=2.5,
metadata={"is_summary": True, "turns_covered": 5},
)
def recall(self, query: str) -> str:
"""Retrieve and format relevant memories for a given query."""
memories = self.memory.retrieve(query, k=3)
if not memories:
return "No relevant memories found."
results = []
for mem, score in memories:
results.append(
f"[{mem.memory_type}, score={score:.2f}] {mem.content[:80]}..."
)
return "\n".join(results)The _estimate_importance method implements the heuristic signals discussed earlier. In a production system, you would augment these heuristics with a learned classifier trained on human-labeled examples of important versus unimportant memories. The classification signals used here, questions, task-framing language, and personal information, are proxies that work surprisingly well for many conversational agents without requiring labeled data.
Now let's show this system with a simulated multi-session interaction:
# Initialize agent
agent = SummarizingAgent()
# Session 1: Initial project setup conversation
session1_inputs = [
(
"We need to build a Python API for our e-commerce platform",
"I'll help you design a Python API for e-commerce. What specific features do you need?",
),
(
"It needs to handle user authentication with JWT tokens",
"JWT authentication is a good choice. I'll note we need secure token handling.",
),
(
"Also implement a shopping cart with Redis caching",
"Redis caching for the shopping cart will ensure fast performance.",
),
(
"My name is Sarah and I prefer FastAPI over Flask",
"Noted, Sarah. FastAPI is excellent for this use case with its async support.",
),
]
for user_msg, agent_msg in session1_inputs:
agent.interact(user_msg, agent_msg)
# Simulate time passing with less important conversation
for i in range(20):
agent.interact(f"Filler conversation point {i}", "Acknowledged.")
# Session 2: User returns with a specific question
query = "How should we handle the JWT refresh strategy?"
recall_result = agent.recall(query)=== Session 2: Authentication Question === Query: How should we handle the JWT refresh strategy? Retrieved memories: [episodic, score=0.88] User: It needs to handle user authentication with JWT tokens... [episodic, score=0.84] User: We need to build a Python API for our e-commerce platform... [episodic, score=0.83] Assistant: I'll help you design a Python API for e-commerce. What specific featu...
The retrieval successfully returns the earlier conversation about JWT tokens despite the intervening filler conversations. Notice how the system combines the semantic similarity to the authentication query with the high importance score assigned to technical requirements. The multi-factor scoring ensures that necessary technical details from the initial session compete effectively against recent but less relevant filler content, even though the filler memories are newer and therefore score higher on recency.
# Session 3: Checking personal preferences
query2 = "What framework should I use for the web API?"
recall_result2 = agent.recall(query2)=== Session 3: Personalization Check === Query: What framework should I use for the web API? Retrieved memories: [episodic, score=0.89] User: We need to build a Python API for our e-commerce platform... [episodic, score=0.88] User: My name is Sarah and I prefer FastAPI over Flask... [episodic, score=0.86] Assistant: I'll help you design a Python API for e-commerce. What specific featu...
The system retrieves Sarah's preference for FastAPI. This shows how importance weighting preserves personal details even when they occur early in the conversation history. The personal preference received an importance boost of 2.0 additional points in our heuristic, bringing its total importance to 4.0 out of 5. This higher importance compensates for its reduced recency score and keeps the preference accessible long after it was stated.
Let's examine the memory statistics to understand how the consolidation worked:
def analyze_memory(agent):
"""Analyze the distribution of memories across tiers and types."""
stats = {}
stats["working_size"] = len(agent.memory.working_memory)
stats["long_term_size"] = len(agent.memory.long_term_memory)
types = {}
for mem in agent.memory.long_term_memory + agent.memory.working_memory:
types[mem.memory_type] = types.get(mem.memory_type, 0) + 1
stats["type_distribution"] = types
all_mems = agent.memory.long_term_memory + agent.memory.working_memory
top_accessed = sorted(all_mems, key=lambda x: x.access_count, reverse=True)[
:3
]
stats["top_accessed"] = [
(mem.content[:60], mem.access_count)
for mem in top_accessed
if mem.access_count > 0
]
return stats
memory_stats = analyze_memory(agent){'working_size': 4,
'long_term_size': 48,
'type_distribution': {'episodic': 48, 'semantic': 4},
'top_accessed': [('User: We need to build a Python API for our e-commerce platf',
2),
("Assistant: I'll help you design a Python API for e-commerce.", 2),
('User: It needs to handle user authentication with JWT tokens', 1)]}


The statistics reveal how working memory maintains the most recent interactions while long-term memory stores the consolidated history. The access counts show which memories the retrieval system found most relevant to recent queries. This shows the feedback loop where frequently accessed memories gain importance over time. This dynamic reweighting mimics the human phenomenon of rehearsal, where frequently recalled memories become more entrenched and easier to access.
Scaling Considerations: Memory in Production
The implementation above shows the core mechanisms, but production systems face scaling challenges that require additional engineering. Understanding these challenges is needed before deploying memory-enabled agents at scale.
Approximate Nearest Neighbor Search
Exhaustive similarity search scans every memory for each query, taking time where is the number of stored memories. For a database of millions of memories, this latency becomes unacceptable for real-time applications. Approximate nearest neighbor (ANN) algorithms reduce this to or amortized time by trading exact correctness for speed.
The most widely deployed ANN algorithm for agent memory is Hierarchical Navigable Small World (HNSW) graphs. HNSW builds a multi-layer graph where higher layers contain long-range connections between distant points and lower layers contain fine-grained local connections. Queries traverse the hierarchy from coarse to fine, following greedy paths toward the nearest neighbor with high probability. The error rate, the fraction of true nearest neighbors missed, can be controlled via construction parameters, letting a speed-accuracy trade-off calibrated to the application's requirements.
For agent memory specifically, the approximate nature of ANN search is usually acceptable. If the true nearest neighbor is a memory about "Python authentication libraries" and the ANN returns a memory about "Python JWT implementation" instead, the agent still retrieves something useful. The necessary information is not lost; it is just slightly less optimal than the exact nearest neighbor would have been. In contrast, for exact key-value lookups (like finding a specific user ID), approximate search is inappropriate and you should use symbolic stores with exact indexing.
Memory Partitioning and Sharding
As memory stores grow, a single index becomes a bottleneck. Production systems partition memories by user, session, or domain, distributing storage and retrieval across multiple shards. Each shard handles a subset of memories, and the query router determines which shards to search based on metadata filters applied before vector search.
For example, a customer service agent might partition memories by customer ID. When serving a query for Customer A, only Customer A's memory shard is searched, reducing both latency and the risk of retrieving memories from the wrong customer. This partitioning improves performance while giving the strong isolation guarantees needed for privacy compliance.
Consistency and Freshness
Vector databases use indexing structures that must be rebuilt or updated when new memories are added. Most systems use "eventually consistent" indexing: new memories are available for exact search immediately but may not appear in the ANN index until the next index rebuild. For long-running agents that continuously accumulate memories, this lag means very recent memories might not be retrievable via approximate search for a few seconds to minutes after creation.
One pragmatic solution is a two-stage retrieval: first search the recent buffer (last N minutes) using exact search, then search the ANN index for older memories, and merge the results. This ensures freshness for very recent memories while maintaining the speed of approximate search for the bulk of the archive.
Key Parameters
The key parameters for the AgentMemory system are:
- capacity: Maximum number of memories stored in long-term memory. When exceeded, the system evicts the lowest-scoring memories based on recency and importance together with access frequency.
- context_size: Number of recent items maintained in working memory before consolidation into long-term storage.
- alpha: Weight for semantic similarity in retrieval scoring. Higher values favor conceptually relevant memories regardless of age.
- beta: Weight for recency in retrieval scoring. Higher values favor recent memories, appropriate for fast-moving conversational contexts.
- gamma: Weight for importance in retrieval scoring. Higher values raise the scores of memories marked as significant, which helps preserve necessary decisions.
- lambda (decay rate): Controls how quickly memories lose their recency score. Higher values create faster decay, suitable for rapidly changing domains.
These weights sum to 1.0 and together determine the retrieval "personality" of the agent. A conservative, accuracy-focused agent might use high alpha with slow decay. A context-sensitive assistant might use high beta to stay focused on the current conversational thread. An agent managing high-stakes decisions might use high gamma to ensure that previously flagged necessary constraints always surface when relevant.
Limitations and Practical Challenges
Despite elegant theoretical frameworks, agent memory systems face significant practical constraints that limit their effectiveness in production environments. These challenges span technical and cognitive domains alongside ethical concerns, requiring careful architectural decisions to mitigate.
The retrieval noise problem plagues vector-based systems. Similar embeddings do not guarantee relevance. A query about "Python" might retrieve memories about snakes, the programming language, or a comedy group, depending on the embedding model's training distribution. This ambiguity forces agents to retrieve generously and filter aggressively, increasing both latency and the risk of distracting the model with irrelevant context. Hybrid approaches combining vector similarity with keyword matching or metadata filtering help, but add architectural complexity and require maintaining multiple indices that must stay synchronized.
Memory staleness presents another persistent challenge. Unlike human memory, which updates and reconciles contradictions semi-automatically through a lifetime of experience, agent memory stores often accumulate inconsistent facts. If a user first states they live in New York, then later mentions moving to San Francisco, both facts persist in the vector store as separate memories. The agent might retrieve the outdated location, leading to confusion or embarrassing errors. Solutions include timestamp weighting (preferring newer memories), explicit contradiction detection using natural language inference models, or periodic memory garbage collection that removes superseded facts. However, detecting contradictions automatically is difficult: distinguishing between "I used to live in New York" (superseded) and "I visit New York frequently" (not superseded) requires deep contextual understanding that current systems cannot reliably give.
The summarization fidelity trade-off creates tension between compression and accuracy. Aggressive summarization, necessary for long conversations, inevitably loses nuance. Critical caveats like "this approach works only for version 2.0" get compressed into "use this approach," leading to errors when the agent applies advice to version 3.0. The challenge is that summarization models optimize for coherence and fluency, not for preserving the specific technical qualifications that make a piece of advice correct in context. Some systems address this by maintaining multiple summary granularities, but this increases storage requirements and complicates retrieval logic considerably.
Computational cost scales poorly with memory size. Exhaustive similarity search across millions of memories requires unacceptable latency for real-time agents. ANN indices help but introduce approximation errors and operational complexity. Memory tiering becomes needed, with hot memories in fast in-memory stores, warm memories in vector databases, and cold memories in cheap object storage, connected by promotion and demotion policies that must be tuned carefully to avoid cache thrashing.
Finally, privacy and security concerns complicate long-term memory materially. Storing conversation histories creates surveillance risks and data protection obligations under regulations like GDPR and CCPA. Episodic memory containing personal details requires encryption, access controls, and clearly defined retention policies. Some applications implement "memory sanitization" that strips personally identifiable information before storage, but this risks removing details that make the memory useful in the first place. The "right to be forgotten" requires mechanisms to expunge specific memories without corrupting the remaining knowledge base, which is technically challenging when memories are embedded in dense vector spaces that cannot be surgically edited.
Despite these challenges, memory-enabled agents stand for a significant advance over stateless function callers. They transform tools from utilities into collaborators that learn and adapt while accumulating expertise over time. The limitations above are engineering problems, not basic barriers, and each is an active area of research and product development.
Summary
Memory distinguishes ephemeral chatbots from persistent intelligent agents. We have explored the architectural layers that let this persistence, from immediate context to archival storage.
Short-term memory uses the model's context window as working memory, holding the immediate conversation and active tool outputs. Its finite nature necessitates careful management through sliding windows and compression. We examined how positional bias creates the "lost in the middle" phenomenon, why token limits force prioritization decisions, and how instruction pinning can partially counteract positional degradation.
Long-term memory extends agents beyond context limits through external vector stores and structured databases. Semantic memory captures factual knowledge via dense embeddings, while episodic memory preserves interaction histories with temporal metadata. Procedural memory encodes behavioral patterns and learned skills, shaping how agents approach tasks without explicit instructions. Each type demands different storage formats and retrieval strategies.
Retrieval mechanisms combine vector similarity with recency and importance weighting to surface relevant memories. Multi-factor scoring balances semantic relevance against temporal proximity and estimated significance, controlled by tunable weights that encode domain assumptions about how quickly knowledge expires. The decay rate in the recency term is a domain-modeling decision as much as a hyperparameter.
Summarization and consolidation compress raw interactions into dense representations, managing the trade-off between detail preservation and token efficiency. Hierarchical architectures organize memory into tiers with explicit management policies, letting agents to optimize information placement based on access patterns. Knowledge graph extraction converts unstructured narrative into structured triples that support precise relational queries.
Advanced architectures like Reflexion store explicit self-critiques to let in-context learning from failure, while memory attention mechanisms integrate retrieval directly into the transformer forward pass for differentiable end-to-end training. The choice between vector and symbolic stores depends on query structure, update frequency, and whether approximate or exact matching is appropriate.
As we continue through the tool use and agents section, these memory foundations let increasingly advanced agent behaviors. Memory allows agents to maintain state across complex tool workflows, learning from previous invocations to improve future ones. In upcoming chapters, we will explore how agents use memory to evaluate their own performance, correct their mistakes, and eventually coordinate in multi-agent systems where shared and individual memories create emergent collective intelligence. The architectures we have built here form the substrate upon which persistent, learning, adaptive agency becomes possible.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about agent memory systems and architectures.
Agent Memory Systems
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!