LangGraph saves state at every node transition through checkpointers. With checkpointing, agents become:
- Resumable (continue from where you stopped).
- Multi-turn (state persists across user inputs).
- Human-in-the-loop ready (pause, get approval, resume).
- Debuggable (inspect state at any point).
- Replayable (re-run from any checkpoint).
Without checkpointing
graph = builder.compile()
graph.invoke({"messages": [user_msg]})
# State exists only during invoke(). Returns final state. State lost.
Stateless. Each invocation is fresh. Fine for single-turn workflows.
With checkpointing
from langgraph.checkpoint.memory import MemorySaver
graph = builder.compile(checkpointer=MemorySaver())
config = {"configurable": {"thread_id": "user-123"}}
graph.invoke({"messages": [user_msg]}, config=config)
# State saved under thread_id "user-123"
# Later (different request, same user):
graph.invoke({"messages": [next_user_msg]}, config=config)
# Resumes from saved state — accumulated messages, all prior decisions
The thread_id is the key. Each thread has its own state.
Checkpointer choices
MemorySaver
In-process memory. Lost on restart.
from langgraph.checkpoint.memory import MemorySaver
checkpointer = MemorySaver()
Good for: development, testing, single-process apps.
SQLite checkpointer
Persistent on disk; single-process.
from langgraph.checkpoint.sqlite import SqliteSaver
checkpointer = SqliteSaver.from_conn_string("checkpoints.db")
Good for: simple production deployments, single-instance services.
Postgres checkpointer
Persistent, multi-process, production-grade.
from langgraph.checkpoint.postgres import PostgresSaver
checkpointer = PostgresSaver.from_conn_string("postgresql://...")
Good for: multi-instance production services, sharing state across replicas, leveraging existing Postgres.
Redis checkpointer
In-memory speed, multi-process.
from langgraph.checkpoint.redis import RedisSaver
checkpointer = RedisSaver.from_conn_string("redis://localhost:6379")
Good for: high-throughput agents, short-lived state (with TTL).
Custom checkpointer
Implement the BaseCheckpointSaver interface for custom storage (S3, DynamoDB, etc.).
What gets checkpointed
Every state transition (node execution) creates a checkpoint with:
- Full state at that point.
- Which node ran.
- Timestamp.
- Parent checkpoint reference (for the history tree).
Storage cost is real. A complex agent with 20 node transitions creates 20 checkpoints. Most checkpointers offer cleanup policies.
Reading state from a thread
# Get current state of thread
state = graph.get_state(config)
print(state.values) # the actual state
print(state.next) # next node(s) to execute (if interrupted)
print(state.config) # the config including checkpoint_id
# Get state history
for state in graph.get_state_history(config):
print(state.values["iteration"], state.next)
Useful for:
- Building admin UIs that show conversation state.
- Debugging: "what was state at step 5?"
- Auditing.
Updating state directly
You can mutate state from outside the graph:
graph.update_state(
config,
{"user_preference": "concise"},
)
Useful for:
- Injecting external context.
- Recovery from human-in-the-loop decisions.
- Admin overrides.
Replaying from a checkpoint
# Get a specific checkpoint
all_states = list(graph.get_state_history(config))
target_checkpoint = all_states[3] # 4th state in history
# Replay from there
new_config = target_checkpoint.config
graph.invoke(None, config=new_config)
# Re-runs from checkpoint 4 onward
Useful for:
- Debugging: re-run from where things went wrong.
- A/B testing: re-run with different prompts from a specific state.
- Recovery: resume after a crash.
Threading concerns
The thread_id is YOUR responsibility to manage. Common patterns:
- Per-user:
thread_id=user_id. One conversation per user. - Per-conversation:
thread_id=conversation_id. User can have multiple parallel conversations. - Per-session:
thread_id=session_id. Resets per session. - Per-task:
thread_id=task_id. For long-running tasks.
Pick based on your data model. Many apps use thread_id=f"{user_id}:{conversation_id}".
Cleaning up old checkpoints
Production agents accumulate state. Add cleanup:
# Manually delete a thread
graph.delete_thread(config)
# Or use TTL-based checkpointer (Redis with expiry)
checkpointer = RedisSaver.from_conn_string("redis://...", ttl=86400) # 24h TTL
For Postgres-backed: schedule a periodic DELETE on old checkpoints based on your retention policy.
What checkpointing enables
- Multi-turn conversations. Same thread_id across user turns.
- Long-running agents. Pause overnight, resume next day.
- Human-in-the-loop. Interrupt, persist, wait for human, resume.
- Crash recovery. Process dies mid-execution? Resume from last checkpoint.
- Multi-instance scaling. Multiple workers process the same thread (with Postgres/Redis).
- Conversation history without manual management. No separate "store messages" code.
Production deployment checklist
- Use Postgres or Redis checkpointer (not MemorySaver).
- Manage thread_ids deliberately (don't auto-generate per request).
- Set up checkpoint cleanup (TTL or scheduled).
- Monitor checkpoint storage growth.
- Test crash recovery (kill the process mid-execution, restart, resume).
- Have an admin path to inspect / update / delete threads.
Common checkpointing mistakes
- Using MemorySaver in production. State lost on restart.
- Random
thread_idper call. Every call starts fresh; benefits lost. - No cleanup policy. Storage grows indefinitely.
- Not testing resumption. "It works" until a deploy restarts the process.
- Sharing thread_id across users. Privacy violation: User B sees User A's state.
Takeaway
Checkpointing is the substrate of production agents. MemorySaver for dev; Postgres or Redis for prod. Pick thread_ids deliberately. Set up cleanup. Test resumption. Once you have it, multi-turn conversations, HITL, and crash recovery come almost for free.
Production Deep Dive: PostgreSQL Checkpointing & Connection Pooling
In production web applications running under multiple Uvicorn worker processes:
- Never use
MemorySaverin production: In-memory checkpointers are isolated to a single Python process. Load balancers routing requests across workers will cause subsequent turns to lose context. - Use
psycopg_pool.ConnectionPool: Initializing a new database connection on every checkpointer write adds 30–50ms of network overhead. Use connection pooling to keep connections warm. - Namespace isolation: Group checkpoints using
thread_id(representing the conversation session) andcheckpoint_ns(representing subgraphs) to prevent cross-tenant state collisions.