Part of AI Agent Handbook
Explains how AI agents store and retrieve information across sessions using vector databases, embeddings, and semantic search.
Toggle tooltip visibility. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Long-Term Knowledge Storage and Retrieval
In the previous subchapter, we gave our assistant the ability to remember recent conversations. This short-term memory works well for ongoing chats, but what happens when the conversation ends? What if you want your assistant to remember your birthday, your favorite restaurants, or important project details days or weeks later?
This is where long-term knowledge storage comes in. Think of it like the difference between remembering what someone just said versus writing important information in a notebook you can reference later. Our assistant needs both capabilities to be truly useful.
Why Long-Term Memory Matters
Imagine asking your assistant, "Remember that my birthday is July 20." A few days later, you ask, "When is my birthday?" Without long-term memory, the assistant has no way to recall this information. The conversation from days ago is gone.
Long-term memory solves this problem by storing important information persistently. Your assistant can save facts, preferences, and knowledge, then retrieve them when needed. With that memory, it can personalize its responses instead of treating every conversation as brand new.
Let's see what this looks like in practice.
Storing Information for Later
The simplest form of long-term memory is a key-value store. You save information with a label (the key) and retrieve it later using that same label.
Here's a basic example using a Python dictionary that persists to a file:
import json
import os
class SimpleMemory:
def __init__(self, memory_file="assistant_memory.json"):
self.memory_file = memory_file
self.memory = self._load_memory()
def _load_memory(self):
"""Load memory from file if it exists"""
if os.path.exists(self.memory_file):
with open(self.memory_file, "r") as f:
return json.load(f)
return {}
def _save_memory(self):
"""Save memory to file"""
with open(self.memory_file, "w") as f:
json.dump(self.memory, f, indent=2)
def store(self, key, value):
"""Store a fact in long-term memory"""
self.memory[key] = value
self._save_memory()
return f"Remembered: {key}"
def retrieve(self, key):
"""Retrieve a fact from long-term memory"""
return self.memory.get(key, "I don't have that information stored.")
## Example usage
memory = SimpleMemory()
memory.store("user_birthday", "July 20")
memory.store("favorite_restaurant", "Luigi's Italian Kitchen")
## Later, even after restarting the program
print(memory.retrieve("user_birthday")) # Output: July 20July 20
This works, but it has a limitation: you need to know the exact key to retrieve information. What if you want to ask, "What do you know about my food preferences?" The assistant would need to search through all stored information to find relevant facts.
Searching Your Knowledge Base
A more powerful approach is to store information in a way that allows searching by content, not just by exact keys. This is where the concept of a knowledge base comes in.
Let's expand our memory system to support searching:
class SearchableMemory:
def __init__(self, memory_file="assistant_knowledge.json"):
self.memory_file = memory_file
self.facts = self._load_facts()
def _load_facts(self):
"""Load facts from file"""
if os.path.exists(self.memory_file):
with open(self.memory_file, "r") as f:
return json.load(f)
return []
def _save_facts(self):
"""Save facts to file"""
with open(self.memory_file, "w") as f:
json.dump(self.facts, f, indent=2)
def add_fact(self, fact, category=None):
"""Add a fact to the knowledge base"""
entry = {
"fact": fact,
"category": category,
"timestamp": str(datetime.now()),
}
self.facts.append(entry)
self._save_facts()
return f"Stored: {fact}"
def search(self, query):
"""Search for facts containing the query text"""
results = []
query_lower = query.lower()
for entry in self.facts:
# Simple keyword search
if query_lower in entry["fact"].lower():
results.append(entry["fact"])
elif entry["category"] and query_lower in entry["category"].lower():
results.append(entry["fact"])
return results if results else ["No matching information found."]
## Example usage
from datetime import datetime
kb = SearchableMemory()
kb.add_fact("User's birthday is July 20", category="personal")
kb.add_fact("User prefers Italian food", category="preferences")
kb.add_fact("User is allergic to peanuts", category="preferences")
## Search by keyword
print(kb.search("birthday"))
## Output: ["User's birthday is July 20"]
print(kb.search("food"))
## Output: ["User prefers Italian food"]
print(kb.search("preferences"))
## Output: ["User prefers Italian food", "User is allergic to peanuts"]["User's birthday is July 20", "User's birthday is July 20", "User's birthday is July 20", "User's birthday is July 20", "User's birthday is July 20"] ['User prefers Italian food', 'User prefers Italian food', 'User prefers Italian food', 'User prefers Italian food', 'User prefers Italian food'] ['User prefers Italian food', 'User is allergic to peanuts', 'User prefers Italian food', 'User is allergic to peanuts', 'User prefers Italian food', 'User is allergic to peanuts', 'User prefers Italian food', 'User is allergic to peanuts', 'User prefers Italian food', 'User is allergic to peanuts']
This is better. Now you can search for information without knowing the exact key. But there's still a problem: this only finds exact keyword matches. What if you ask, "What should I avoid eating?" The assistant needs to understand that this relates to allergies, even though the word "avoid" doesn't appear in the stored facts.
Understanding Meaning, Not Just Keywords
This is where vector stores and embeddings become useful. Instead of matching exact words, we can represent the meaning of text as numbers (vectors) and find information that's semantically similar.
Here's the concept: when you store a fact like "User is allergic to peanuts," the system converts this into a vector that represents its meaning. Later, when you ask "What should I avoid eating?", that question also gets converted to a vector. The system then finds stored facts whose vectors are close to the question's vector, meaning they're semantically related.
Let's see this in action using a simple example with sentence embeddings:
from sentence_transformers import SentenceTransformer
import numpy as np
class VectorMemory:
def __init__(self):
# Using Claude Sonnet 4.5 for the agent, but a local model for embeddings
# to keep costs down for frequent similarity searches
self.encoder = SentenceTransformer("all-MiniLM-L6-v2")
self.facts = []
self.vectors = []
def add_fact(self, fact):
"""Add a fact and its vector representation"""
vector = self.encoder.encode(fact)
self.facts.append(fact)
self.vectors.append(vector)
return f"Stored: {fact}"
def search(self, query, top_k=3):
"""Find the most relevant facts for a query"""
if not self.facts:
return ["No information stored yet."]
# Convert query to vector
query_vector = self.encoder.encode(query)
# Calculate similarity with all stored facts
similarities = []
for i, fact_vector in enumerate(self.vectors):
# Cosine similarity
similarity = np.dot(query_vector, fact_vector) / (
np.linalg.norm(query_vector) * np.linalg.norm(fact_vector)
)
similarities.append((similarity, self.facts[i]))
# Sort by similarity and return top results
similarities.sort(reverse=True, key=lambda x: x[0])
return [fact for score, fact in similarities[:top_k] if score > 0.3]
## Example usage
memory = VectorMemory()
memory.add_fact("User's birthday is July 20")
memory.add_fact("User prefers Italian food")
memory.add_fact("User is allergic to peanuts")
memory.add_fact("User lives in San Francisco")
memory.add_fact("User works as a software engineer")
## Semantic search
print(memory.search("What should I avoid eating?"))
## Output: ["User is allergic to peanuts", "User prefers Italian food"]
print(memory.search("Where does the user live?"))
## Output: ["User lives in San Francisco"]
print(memory.search("What is the user's profession?"))
## Output: ["User works as a software engineer"]Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
Loading weights: 0%| | 0/103 [00:00<?, ?it/s]
['User prefers Italian food'] ['User lives in San Francisco', 'User works as a software engineer', "User's birthday is July 20"] ['User works as a software engineer', 'User lives in San Francisco']
Notice how the search understands meaning. When you ask "What should I avoid eating?", it finds the allergy information even though the words don't match exactly. The vector representation captures that avoiding food relates to allergies and food preferences.
Integrating Long-Term Memory with Our Assistant
Now let's connect this to our personal assistant. We want the assistant to:
- Recognize when you're telling it something to remember
- Store that information in long-term memory
- Retrieve relevant information when answering questions
- Combine retrieved knowledge with its language model capabilities
Here's how this works with Claude Sonnet 4.5 (Example: Claude Sonnet 4.5):
import os
from anthropic import Anthropic
class AssistantWithMemory:
def __init__(self):
self.client = Anthropic(api_key=os.getenv("ANTHROPIC_API_KEY"))
# Using Claude Sonnet 4.5 for its superior agent reasoning capabilities
self.model = "claude-sonnet-4-5"
self.memory = VectorMemory()
self.conversation_history = []
def process_message(self, user_message):
"""Process a user message, using memory when relevant"""
# First, check if this is a request to remember something
if self._is_memory_request(user_message):
return self._store_memory(user_message)
# Search for relevant facts
relevant_facts = self.memory.search(user_message, top_k=3)
# Build context with conversation history and retrieved facts
context = self._build_context(relevant_facts)
# Get response from Claude
response = self.client.messages.create(
model=self.model,
max_tokens=1024,
system=context,
messages=self.conversation_history
+ [{"role": "user", "content": user_message}],
)
assistant_message = response.content[0].text
# Update conversation history
self.conversation_history.append(
{"role": "user", "content": user_message}
)
self.conversation_history.append(
{"role": "assistant", "content": assistant_message}
)
return assistant_message
def _is_memory_request(self, message):
"""Detect if user wants to store information"""
memory_keywords = [
"remember",
"store",
"save",
"keep in mind",
"note that",
]
return any(keyword in message.lower() for keyword in memory_keywords)
def _store_memory(self, message):
"""Extract and store information from user message"""
# Use Claude to extract the fact to remember
response = self.client.messages.create(
model=self.model,
max_tokens=256,
system="Extract the key fact the user wants you to remember. Return only the fact as a clear statement.",
messages=[{"role": "user", "content": message}],
)
fact = response.content[0].text
self.memory.add_fact(fact)
return f"Got it! I'll remember that {fact.lower()}"
def _build_context(self, relevant_facts):
"""Build system context with retrieved facts"""
if not relevant_facts or relevant_facts == [
"No information stored yet."
]:
return "You are a helpful personal assistant."
facts_text = "\n".join(f"- {fact}" for fact in relevant_facts)
return f"""You are a helpful personal assistant. You have access to the following information about the user:
{facts_text}
Use this information when relevant to provide personalized responses."""
## Example conversation
assistant = AssistantWithMemory()
print(assistant.process_message("Remember that my birthday is July 20"))
## Output: Got it! I'll remember that your birthday is july 20
print(assistant.process_message("Remember that I'm allergic to peanuts"))
## Output: Got it! I'll remember that you're allergic to peanuts
print(
assistant.process_message("What should I be careful about when eating out?")
)
## Output: Based on what I know, you should be careful about peanuts since you're
## allergic to them. When eating out, make sure to inform the restaurant staff about
## your peanut allergy and ask about ingredients in dishes...Loading weights: 0%| | 0/103 [00:00<?, ?it/s]
[31m---------------------------------------------------------------------------[39m
[31mTypeError[39m Traceback (most recent call last)
[36mCell[39m[36m [39m[32mIn[6][39m[32m, line 89[39m
[32m 86[39m [38;5;66;03m## Example conversation[39;00m
[32m 87[39m assistant = AssistantWithMemory()
[32m---> [39m[32m89[39m [38;5;28mprint[39m([43massistant[49m[43m.[49m[43mprocess_message[49m[43m([49m[33;43m"[39;49m[33;43mRemember that my birthday is July 20[39;49m[33;43m"[39;49m[43m)[49m)
[32m 90[39m [38;5;66;03m## Output: Got it! I'll remember that your birthday is july 20[39;00m
[32m 92[39m [38;5;28mprint[39m(assistant.process_message([33m"[39m[33mRemember that I[39m[33m'[39m[33mm allergic to peanuts[39m[33m"[39m))
[36mCell[39m[36m [39m[32mIn[6][39m[32m, line 17[39m, in [36mAssistantWithMemory.process_message[39m[34m(self, user_message)[39m
[32m 15[39m [38;5;66;03m# First, check if this is a request to remember something[39;00m
[32m 16[39m [38;5;28;01mif[39;00m [38;5;28mself[39m._is_memory_request(user_message):
[32m---> [39m[32m17[39m [38;5;28;01mreturn[39;00m [38;5;28;43mself[39;49m[43m.[49m[43m_store_memory[49m[43m([49m[43muser_message[49m[43m)[49m
[32m 19[39m [38;5;66;03m# Search for relevant facts[39;00m
[32m 20[39m relevant_facts = [38;5;28mself[39m.memory.search(user_message, top_k=[32m3[39m)
[36mCell[39m[36m [39m[32mIn[6][39m[32m, line 60[39m, in [36mAssistantWithMemory._store_memory[39m[34m(self, message)[39m
[32m 58[39m [38;5;250m[39m[33;03m"""Extract and store information from user message"""[39;00m
[32m 59[39m [38;5;66;03m# Use Claude to extract the fact to remember[39;00m
[32m---> [39m[32m60[39m response = [38;5;28;43mself[39;49m[43m.[49m[43mclient[49m[43m.[49m[43mmessages[49m[43m.[49m[43mcreate[49m[43m([49m
[32m 61[39m [43m [49m[43mmodel[49m[43m=[49m[38;5;28;43mself[39;49m[43m.[49m[43mmodel[49m[43m,[49m
[32m 62[39m [43m [49m[43mmax_tokens[49m[43m=[49m[32;43m256[39;49m[43m,[49m
[32m 63[39m [43m [49m[43msystem[49m[43m=[49m[33;43m"[39;49m[33;43mExtract the key fact the user wants you to remember. Return only the fact as a clear statement.[39;49m[33;43m"[39;49m[43m,[49m
[32m 64[39m [43m [49m[43mmessages[49m[43m=[49m[43m[[49m[43m{[49m[33;43m"[39;49m[33;43mrole[39;49m[33;43m"[39;49m[43m:[49m[43m [49m[33;43m"[39;49m[33;43muser[39;49m[33;43m"[39;49m[43m,[49m[43m [49m[33;43m"[39;49m[33;43mcontent[39;49m[33;43m"[39;49m[43m:[49m[43m [49m[43mmessage[49m[43m}[49m[43m][49m[43m,[49m
[32m 65[39m [43m[49m[43m)[49m
[32m 67[39m fact = response.content[[32m0[39m].text
[32m 68[39m [38;5;28mself[39m.memory.add_fact(fact)
[36mFile [39m[32m/private/tmp/mb-ai-style-scanner/.venv/lib/python3.11/site-packages/anthropic/_utils/_utils.py:282[39m, in [36mrequired_args.<locals>.inner.<locals>.wrapper[39m[34m(*args, **kwargs)[39m
[32m 280[39m msg = [33mf[39m[33m"[39m[33mMissing required argument: [39m[38;5;132;01m{[39;00mquote(missing[[32m0[39m])[38;5;132;01m}[39;00m[33m"[39m
[32m 281[39m [38;5;28;01mraise[39;00m [38;5;167;01mTypeError[39;00m(msg)
[32m--> [39m[32m282[39m [38;5;28;01mreturn[39;00m [43mfunc[49m[43m([49m[43m*[49m[43margs[49m[43m,[49m[43m [49m[43m*[49m[43m*[49m[43mkwargs[49m[43m)[49m
[36mFile [39m[32m/private/tmp/mb-ai-style-scanner/.venv/lib/python3.11/site-packages/anthropic/resources/messages/messages.py:996[39m, in [36mMessages.create[39m[34m(self, max_tokens, messages, model, cache_control, container, inference_geo, metadata, output_config, service_tier, stop_sequences, stream, system, temperature, thinking, tool_choice, tools, top_k, top_p, extra_headers, extra_query, extra_body, timeout)[39m
[32m 989[39m [38;5;28;01mif[39;00m model [38;5;129;01min[39;00m MODELS_TO_WARN_WITH_THINKING_ENABLED [38;5;129;01mand[39;00m thinking [38;5;129;01mand[39;00m thinking[[33m"[39m[33mtype[39m[33m"[39m] == [33m"[39m[33menabled[39m[33m"[39m:
[32m 990[39m warnings.warn(
[32m 991[39m [33mf[39m[33m"[39m[33mUsing Claude with [39m[38;5;132;01m{[39;00mmodel[38;5;132;01m}[39;00m[33m and [39m[33m'[39m[33mthinking.type=enabled[39m[33m'[39m[33m is deprecated. Use [39m[33m'[39m[33mthinking.type=adaptive[39m[33m'[39m[33m instead which results in better model performance in our testing: https://platform.claude.com/docs/en/build-with-claude/adaptive-thinking[39m[33m"[39m,
[32m 992[39m [38;5;167;01mUserWarning[39;00m,
[32m 993[39m stacklevel=[32m3[39m,
[32m 994[39m )
[32m--> [39m[32m996[39m [38;5;28;01mreturn[39;00m [38;5;28;43mself[39;49m[43m.[49m[43m_post[49m[43m([49m
[32m 997[39m [43m [49m[33;43m"[39;49m[33;43m/v1/messages[39;49m[33;43m"[39;49m[43m,[49m
[32m 998[39m [43m [49m[43mbody[49m[43m=[49m[43mmaybe_transform[49m[43m([49m
[32m 999[39m [43m [49m[43m{[49m
[32m 1000[39m [43m [49m[33;43m"[39;49m[33;43mmax_tokens[39;49m[33;43m"[39;49m[43m:[49m[43m [49m[43mmax_tokens[49m[43m,[49m
[32m 1001[39m [43m [49m[33;43m"[39;49m[33;43mmessages[39;49m[33;43m"[39;49m[43m:[49m[43m [49m[43mmessages[49m[43m,[49m
[32m 1002[39m [43m [49m[33;43m"[39;49m[33;43mmodel[39;49m[33;43m"[39;49m[43m:[49m[43m [49m[43mmodel[49m[43m,[49m
[32m 1003[39m [43m [49m[33;43m"[39;49m[33;43mcache_control[39;49m[33;43m"[39;49m[43m:[49m[43m [49m[43mcache_control[49m[43m,[49m
[32m 1004[39m [43m [49m[33;43m"[39;49m[33;43mcontainer[39;49m[33;43m"[39;49m[43m:[49m[43m [49m[43mcontainer[49m[43m,[49m
[32m 1005[39m [43m [49m[33;43m"[39;49m[33;43minference_geo[39;49m[33;43m"[39;49m[43m:[49m[43m [49m[43minference_geo[49m[43m,[49m
[32m 1006[39m [43m [49m[33;43m"[39;49m[33;43mmetadata[39;49m[33;43m"[39;49m[43m:[49m[43m [49m[43mmetadata[49m[43m,[49m
[32m 1007[39m [43m [49m[33;43m"[39;49m[33;43moutput_config[39;49m[33;43m"[39;49m[43m:[49m[43m [49m[43moutput_config[49m[43m,[49m
[32m 1008[39m [43m [49m[33;43m"[39;49m[33;43mservice_tier[39;49m[33;43m"[39;49m[43m:[49m[43m [49m[43mservice_tier[49m[43m,[49m
[32m 1009[39m [43m [49m[33;43m"[39;49m[33;43mstop_sequences[39;49m[33;43m"[39;49m[43m:[49m[43m [49m[43mstop_sequences[49m[43m,[49m
[32m 1010[39m [43m [49m[33;43m"[39;49m[33;43mstream[39;49m[33;43m"[39;49m[43m:[49m[43m [49m[43mstream[49m[43m,[49m
[32m 1011[39m [43m [49m[33;43m"[39;49m[33;43msystem[39;49m[33;43m"[39;49m[43m:[49m[43m [49m[43msystem[49m[43m,[49m
[32m 1012[39m [43m [49m[33;43m"[39;49m[33;43mtemperature[39;49m[33;43m"[39;49m[43m:[49m[43m [49m[43mtemperature[49m[43m,[49m
[32m 1013[39m [43m [49m[33;43m"[39;49m[33;43mthinking[39;49m[33;43m"[39;49m[43m:[49m[43m [49m[43mthinking[49m[43m,[49m
[32m 1014[39m [43m [49m[33;43m"[39;49m[33;43mtool_choice[39;49m[33;43m"[39;49m[43m:[49m[43m [49m[43mtool_choice[49m[43m,[49m
[32m 1015[39m [43m [49m[33;43m"[39;49m[33;43mtools[39;49m[33;43m"[39;49m[43m:[49m[43m [49m[43mtools[49m[43m,[49m
[32m 1016[39m [43m [49m[33;43m"[39;49m[33;43mtop_k[39;49m[33;43m"[39;49m[43m:[49m[43m [49m[43mtop_k[49m[43m,[49m
[32m 1017[39m [43m [49m[33;43m"[39;49m[33;43mtop_p[39;49m[33;43m"[39;49m[43m:[49m[43m [49m[43mtop_p[49m[43m,[49m
[32m 1018[39m [43m [49m[43m}[49m[43m,[49m
[32m 1019[39m [43m [49m[43mmessage_create_params[49m[43m.[49m[43mMessageCreateParamsStreaming[49m
[32m 1020[39m [43m [49m[38;5;28;43;01mif[39;49;00m[43m [49m[43mstream[49m
[32m 1021[39m [43m [49m[38;5;28;43;01melse[39;49;00m[43m [49m[43mmessage_create_params[49m[43m.[49m[43mMessageCreateParamsNonStreaming[49m[43m,[49m
[32m 1022[39m [43m [49m[43m)[49m[43m,[49m
[32m 1023[39m [43m [49m[43moptions[49m[43m=[49m[43mmake_request_options[49m[43m([49m
[32m 1024[39m [43m [49m[43mextra_headers[49m[43m=[49m[43mextra_headers[49m[43m,[49m[43m [49m[43mextra_query[49m[43m=[49m[43mextra_query[49m[43m,[49m[43m [49m[43mextra_body[49m[43m=[49m[43mextra_body[49m[43m,[49m[43m [49m[43mtimeout[49m[43m=[49m[43mtimeout[49m
[32m 1025[39m [43m [49m[43m)[49m[43m,[49m
[32m 1026[39m [43m [49m[43mcast_to[49m[43m=[49m[43mMessage[49m[43m,[49m
[32m 1027[39m [43m [49m[43mstream[49m[43m=[49m[43mstream[49m[43m [49m[38;5;129;43;01mor[39;49;00m[43m [49m[38;5;28;43;01mFalse[39;49;00m[43m,[49m
[32m 1028[39m [43m [49m[43mstream_cls[49m[43m=[49m[43mStream[49m[43m[[49m[43mRawMessageStreamEvent[49m[43m][49m[43m,[49m
[32m 1029[39m [43m[49m[43m)[49m
[36mFile [39m[32m/private/tmp/mb-ai-style-scanner/.venv/lib/python3.11/site-packages/anthropic/_base_client.py:1364[39m, in [36mSyncAPIClient.post[39m[34m(self, path, cast_to, body, content, options, files, stream, stream_cls)[39m
[32m 1355[39m warnings.warn(
[32m 1356[39m [33m"[39m[33mPassing raw bytes as `body` is deprecated and will be removed in a future version. [39m[33m"[39m
[32m 1357[39m [33m"[39m[33mPlease pass raw bytes via the `content` parameter instead.[39m[33m"[39m,
[32m 1358[39m [38;5;167;01mDeprecationWarning[39;00m,
[32m 1359[39m stacklevel=[32m2[39m,
[32m 1360[39m )
[32m 1361[39m opts = FinalRequestOptions.construct(
[32m 1362[39m method=[33m"[39m[33mpost[39m[33m"[39m, url=path, json_data=body, content=content, files=to_httpx_files(files), **options
[32m 1363[39m )
[32m-> [39m[32m1364[39m [38;5;28;01mreturn[39;00m cast(ResponseT, [38;5;28;43mself[39;49m[43m.[49m[43mrequest[49m[43m([49m[43mcast_to[49m[43m,[49m[43m [49m[43mopts[49m[43m,[49m[43m [49m[43mstream[49m[43m=[49m[43mstream[49m[43m,[49m[43m [49m[43mstream_cls[49m[43m=[49m[43mstream_cls[49m[43m)[49m)
[36mFile [39m[32m/private/tmp/mb-ai-style-scanner/.venv/lib/python3.11/site-packages/anthropic/_base_client.py:1058[39m, in [36mSyncAPIClient.request[39m[34m(self, cast_to, options, stream, stream_cls)[39m
[32m 1055[39m options = [38;5;28mself[39m._prepare_options(options)
[32m 1057[39m remaining_retries = max_retries - retries_taken
[32m-> [39m[32m1058[39m request = [38;5;28;43mself[39;49m[43m.[49m[43m_build_request[49m[43m([49m[43moptions[49m[43m,[49m[43m [49m[43mretries_taken[49m[43m=[49m[43mretries_taken[49m[43m)[49m
[32m 1059[39m [38;5;28mself[39m._prepare_request(request)
[32m 1061[39m kwargs: HttpxSendArgs = {}
[36mFile [39m[32m/private/tmp/mb-ai-style-scanner/.venv/lib/python3.11/site-packages/anthropic/_base_client.py:521[39m, in [36mBaseClient._build_request[39m[34m(self, options, retries_taken)[39m
[32m 518[39m [38;5;28;01melse[39;00m:
[32m 519[39m [38;5;28;01mraise[39;00m [38;5;167;01mRuntimeError[39;00m([33mf[39m[33m"[39m[33mUnexpected JSON data type, [39m[38;5;132;01m{[39;00m[38;5;28mtype[39m(json_data)[38;5;132;01m}[39;00m[33m, cannot merge with `extra_body`[39m[33m"[39m)
[32m--> [39m[32m521[39m headers = [38;5;28;43mself[39;49m[43m.[49m[43m_build_headers[49m[43m([49m[43moptions[49m[43m,[49m[43m [49m[43mretries_taken[49m[43m=[49m[43mretries_taken[49m[43m)[49m
[32m 522[39m params = _merge_mappings([38;5;28mself[39m.default_query, options.params)
[32m 523[39m content_type = headers.get([33m"[39m[33mContent-Type[39m[33m"[39m)
[36mFile [39m[32m/private/tmp/mb-ai-style-scanner/.venv/lib/python3.11/site-packages/anthropic/_base_client.py:451[39m, in [36mBaseClient._build_headers[39m[34m(self, options, retries_taken)[39m
[32m 441[39m custom_headers = options.headers [38;5;129;01mor[39;00m {}
[32m 442[39m headers_dict = _merge_mappings(
[32m 443[39m {
[32m 444[39m [33m"[39m[33mx-stainless-timeout[39m[33m"[39m: [38;5;28mstr[39m(options.timeout.read)
[32m (...)[39m[32m 449[39m custom_headers,
[32m 450[39m )
[32m--> [39m[32m451[39m [38;5;28;43mself[39;49m[43m.[49m[43m_validate_headers[49m[43m([49m[43mheaders_dict[49m[43m,[49m[43m [49m[43mcustom_headers[49m[43m)[49m
[32m 453[39m [38;5;66;03m# headers are case-insensitive while dictionaries are not.[39;00m
[32m 454[39m headers = httpx.Headers(headers_dict)
[36mFile [39m[32m/private/tmp/mb-ai-style-scanner/.venv/lib/python3.11/site-packages/anthropic/_client.py:196[39m, in [36mAnthropic._validate_headers[39m[34m(self, headers, custom_headers)[39m
[32m 193[39m [38;5;28;01mif[39;00m headers.get([33m"[39m[33mAuthorization[39m[33m"[39m) [38;5;129;01mor[39;00m [38;5;28misinstance[39m(custom_headers.get([33m"[39m[33mAuthorization[39m[33m"[39m), Omit):
[32m 194[39m [38;5;28;01mreturn[39;00m
[32m--> [39m[32m196[39m [38;5;28;01mraise[39;00m [38;5;167;01mTypeError[39;00m(
[32m 197[39m [33m'[39m[33m"[39m[33mCould not resolve authentication method. Expected either api_key or auth_token to be set. Or for one of the `X-Api-Key` or `Authorization` headers to be explicitly omitted[39m[33m"[39m[33m'[39m
[32m 198[39m )
[31mTypeError[39m: "Could not resolve authentication method. Expected either api_key or auth_token to be set. Or for one of the `X-Api-Key` or `Authorization` headers to be explicitly omitted"Let's trace through what happens when you ask "What should I be careful about when eating out?":
- The assistant searches its vector memory for relevant facts
- It finds "You're allergic to peanuts" as highly relevant
- It includes this fact in the system context when calling Claude
- Claude uses this information to provide a personalized, helpful response
The assistant now has a persistent memory that survives across sessions. You can close the program, restart it days later, and it will still remember your allergy.
When to Use Long-Term Memory
Not everything needs to be stored in long-term memory. Here's a practical guide:
Store in long-term memory:
- Personal facts (birthday, location, occupation)
- Preferences (favorite foods, music, work style)
- Important information (allergies, constraints, requirements)
- Project details that span multiple sessions
- Learned facts about recurring topics
Keep in short-term memory:
- Current conversation context
- Temporary working information
- Details specific to this session only
- Information that will become outdated quickly
You can even ask the assistant to decide what's worth remembering:
def should_remember(self, information):
"""Use Claude to decide if information is worth storing long-term"""
response = self.client.messages.create(
model=self.model,
max_tokens=128,
system="Decide if this information should be stored in long-term memory. Answer only 'yes' or 'no' with a brief reason.",
messages=[
{
"role": "user",
"content": f"Should I remember this: {information}",
}
],
)
decision = response.content[0].text.lower()
return "yes" in decisionPractical Considerations
As you build long-term memory into your assistant, keep these points in mind:
Storage limits: Vector databases can grow large. Consider setting limits on how many facts to store, or implementing a way to archive or remove old information.
Privacy: Long-term memory means persistent data. Be thoughtful about what you store and how you protect it. Never store sensitive information like passwords or financial details without proper encryption.
Retrieval quality: The quality of your retrieval depends on your embedding model. Better embeddings lead to better semantic search. The example above uses a small, fast model, but you might want a more powerful one for production use.
Context window limits: Even with long-term memory, you can only include so many retrieved facts in each request to the language model. Prioritize the most relevant information.
Updating facts: What happens when information changes? You might need a way to update or delete facts. For example, if the user moves to a new city, you want to update that fact rather than having two conflicting locations stored.
Here's a simple update mechanism:
def update_fact(self, old_fact, new_fact):
"""Replace an old fact with updated information"""
# Remove old fact
if old_fact in self.memory.facts:
idx = self.memory.facts.index(old_fact)
self.memory.facts.pop(idx)
self.memory.vectors.pop(idx)
# Add new fact
self.memory.add_fact(new_fact)
return f"Updated: {new_fact}"Combining Short-Term and Long-Term Memory
Your assistant now has both types of memory:
- Short-term memory: Recent conversation history (from the previous subchapter)
- Long-term memory: Persistent facts and knowledge (this subchapter)
The most effective assistants use both together. Short-term memory provides immediate context for the current conversation. Long-term memory provides background knowledge and personalization.
Here's how they work together:
User: "I'm planning a dinner party next week"
Assistant: [Uses short-term memory to track this conversation]
"That sounds great! What kind of cuisine are you thinking?"
User: "Maybe Italian?"
Assistant: [Retrieves from long-term memory: "User prefers Italian food"]
[Uses short-term memory: knows we're discussing a dinner party]
"Perfect choice! I know you love Italian food. Are you thinking
of going to Luigi's Italian Kitchen, your favorite restaurant,
or cooking at home?"
User: "Cooking at home. Can you suggest a menu?"
Assistant: [Retrieves from long-term memory: "User is allergic to peanuts"]
[Uses short-term memory: dinner party, Italian, cooking at home]
"I'll suggest a menu that avoids peanuts since you're allergic.
How about starting with a Caprese salad, followed by homemade
pasta with marinara sauce, and tiramisu for dessert?"Notice how the assistant seamlessly blends information from both memory systems. It remembers the ongoing conversation (short-term) and applies personal knowledge (long-term) to provide helpful, customized suggestions.
Building Your Own Knowledge Base
You now understand the core concepts of long-term memory for AI agents. The examples above use simple implementations to illustrate the ideas. In practice, you might use specialized tools:
Vector databases like Pinecone, Weaviate, or Chroma provide optimized storage and retrieval for embeddings. They handle large-scale data better than our simple list-based approach.
Document stores like Elasticsearch or MongoDB work well when you need to store structured information with multiple fields and complex queries.
Graph databases like Neo4j excel when relationships between facts matter (for example, "Alice is Bob's manager" and "Bob works on Project X" implies "Alice oversees Project X").
The choice depends on your needs. For a personal assistant with a few hundred facts, the simple approach we've shown works fine. For a system serving thousands of users with millions of facts, you'll want more robust infrastructure.
What We've Built
Your assistant can now:
- Store facts persistently across sessions
- Search for information by meaning, not just keywords
- Retrieve relevant knowledge when answering questions
- Combine retrieved facts with language model capabilities
- Decide what information is worth remembering long-term
With long-term memory, your assistant no longer behaves like a stateless question-answering system. It can remember what matters to you and use that knowledge to provide more relevant help.
In the next chapter, we'll explore how to organize all these pieces into a coherent agent architecture, showing how memory, reasoning, and tools work together in a unified system.
Glossary
Embedding: A numerical representation (vector) of text that captures its semantic meaning. Similar meanings produce similar vectors, enabling semantic search.
Key-Value Store: A simple storage system where data is saved with a label (key) and retrieved using that same label. Like a dictionary or hash map.
Knowledge Base: A structured collection of information that an agent can search and retrieve from. More sophisticated than simple key-value storage.
Semantic Search: Finding information based on meaning rather than exact keyword matches. Uses embeddings to understand that "What should I avoid eating?" relates to allergy information.
Vector Database: A specialized database optimized for storing and searching embeddings. It supports fast similarity search across large collections of data.
Vector Store: Another term for vector database. A system that stores embeddings and supports similarity-based retrieval.
Cosine Similarity: A mathematical measure of how similar two vectors are, ranging from 0 (completely different) to 1 (identical). Used to find relevant information in vector search.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about long-term knowledge storage and retrieval for AI agents.
Long-Term Knowledge Storage and Retrieval
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of AI Agent Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore AI Agent HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!