"Memory" in an agent means different things. Three distinct types; each solves different problems.
Type 1 — Short-term (working memory / context)
Messages in the current conversation. Lives in the LLM's context window.
- Stored: in the conversation's message list (state).
- Lifetime: current session / thread.
- Access: every LLM call sees all messages.
This is what LangGraph's checkpointer persists. Multi-turn conversations naturally have short-term memory if you use a consistent thread_id.
Challenges:
- Grows with conversation length → cost grows linearly.
- Context window limit (e.g., 200K tokens for Claude 4 standard).
- "Lost in the middle" — older messages get less attention.
Mitigations:
- Trim: keep only last N messages.
- Summarize: replace older messages with a summary.
- Compress: re-prompt to be more concise.
def trim_messages(state):
messages = state["messages"]
if len(messages) > 20:
# Keep system message + last 19
return {"messages": [messages[0]] + messages[-19:]}
return state
# Or summarize
def summarize_old_messages(state):
if len(state["messages"]) > 30:
old_messages = state["messages"][:-10]
summary = llm.invoke([
SystemMessage("Summarize this conversation in 200 words."),
*old_messages
])
return {"messages": [SystemMessage(f"Earlier conversation summary: {summary.content}")] + state["messages"][-10:]}
return state
For production multi-turn agents: trim or summarize beyond 20-30 messages.
Type 2 — Long-term (persistent across sessions)
Facts the agent should remember about a user across conversations:
- User preferences ("call me Anuj").
- Past purchases.
- Decisions made in prior sessions.
This requires explicit storage outside the conversation.
Implementation options:
Database-backed
Store user facts in your own database. Retrieve at conversation start:
def load_user_context(user_id):
profile = db.query("SELECT * FROM user_profiles WHERE user_id = ?", user_id)
history = db.query("SELECT * FROM user_facts WHERE user_id = ? ORDER BY relevance DESC LIMIT 10", user_id)
return {
"profile": profile,
"facts": history,
}
def start_conversation(state, user_id):
context = load_user_context(user_id)
system_msg = f'''You're helping {context['profile']['name']}.
Known facts about this user:
{format_facts(context['facts'])}
'''
return {"messages": [SystemMessage(system_msg)] + state["messages"]}
Update during the conversation:
def extract_and_save_facts(state, user_id):
response = llm.invoke([
SystemMessage("Extract any new facts about the user from this conversation. JSON output."),
*state["messages"]
])
new_facts = json.loads(response.content)
for fact in new_facts:
db.insert("user_facts", user_id=user_id, **fact)
Memory tools
Expose memory as tools the agent can call:
@tool
def remember_about_user(user_id: str, fact: str) -> str:
'''Save a fact about the user for future conversations.'''
db.insert("user_facts", user_id=user_id, fact=fact)
return f"Remembered: {fact}"
@tool
def recall_about_user(user_id: str, query: str) -> list:
'''Look up what we know about the user related to a topic.'''
return db.query("SELECT fact FROM user_facts WHERE user_id = ? AND fact MATCH ?", user_id, query)
The agent decides when to remember and when to recall.
Dedicated memory libraries
Mem0 is a popular memory layer for agents — extraction + storage + retrieval in one package:
from mem0 import Memory
memory = Memory()
# At each turn:
memory.add(user_message, user_id=user_id)
# Retrieve relevant memories:
related = memory.search(query="payment preferences", user_id=user_id)
For complex memory needs, libraries like Mem0 (or Letta, formerly MemGPT) save building from scratch.
Type 3 — Semantic (knowledge memory)
Facts about the world, not the user. Effectively RAG over a knowledge base:
- Product specifications.
- Company policies.
- Historical records.
Implementation: vector database + retrieval. Module 3 of the LLMs course covered this in depth.
For agents:
@tool
def search_company_knowledge(query: str) -> str:
'''Search internal company docs (policies, products, FAQs).'''
chunks = vector_store.similarity_search(query, k=5)
return "\n\n".join(c.page_content for c in chunks)
The agent retrieves on demand. Different from user memory — same for all users, doesn't change per conversation.
Combining all three
A production agent often uses all three memory types:
USER ASKS: "Did the issue I mentioned last time get resolved?"
1. SHORT-TERM: see current conversation context.
2. LONG-TERM: look up this user's past tickets ("issue from last time").
3. SEMANTIC: check current product status / known bugs.
4. RESPOND: combine all three into the answer.
The agent decides (or the graph encodes) which memory layer to query.
Memory consolidation
Over time, long-term memory grows. Strategies:
Decay
Forget old facts unless they're re-encountered. Like human memory.
Summarization
Periodically condense many small facts into broader summaries.
Importance scoring
Tag facts by importance; prune low-importance over time.
User-driven editing
Let users review and edit their stored facts.
For most agents, simple "last N facts" or "facts updated in last 90 days" suffices.
Privacy and memory
Storing facts about users has implications:
- GDPR: users have the right to view and delete their data.
- Sensitive data: PII, financial, medical info needs special handling.
- Cross-user contamination: never share facts across users.
Always:
- Scope memory by user_id strictly.
- Allow user to view/edit/delete their stored facts.
- Don't store secrets (passwords, payment details).
- Comply with data protection regulations.
Common memory mistakes
- No memory beyond conversation. Agents that can't remember user across sessions feel dumb.
- Storing everything. Memory bloat = retrieval noise + privacy risk.
- No way to forget. Facts persist incorrectly forever.
- Mixing memory types. Storing world knowledge in user memory or vice versa.
- Cross-user leakage. Memory retrieval that doesn't strictly scope by user_id.
Takeaway
Short-term: conversation context (trim or summarize at length). Long-term: facts about the user (DB or memory library). Semantic: world knowledge (RAG). Production agents use all three. Scope strictly by user. Allow forgetting. Mem0/Letta save you from building memory infrastructure if your needs are complex.