A 1-LLM-call app is debuggable from logs. A 20-step agent is not. Production agents need three layers of observability.
Layer 1 — Streaming state changes
Stream each node's state change as it happens:
for chunk in graph.stream({"messages": [user_msg]}, config):
print(chunk)
# {"agent": {"messages": [AIMessage(...)]}}
# {"tools": {"messages": [ToolMessage(...)]}}
# {"agent": {"messages": [AIMessage(...)]}}
For UX: stream to the user too. "Researching... Found 3 documents... Generating answer..." → better than a 30-second silent wait.
For debugging: see exactly where execution is at any moment.
In production:
- Stream events to a queue (Kafka, Redis Streams).
- Log each event.
- Power a real-time debug UI.
Layer 2 — LangSmith tracing
LangSmith automatically traces every LLM call, tool invocation, and chain/graph execution:
import os
os.environ["LANGCHAIN_TRACING_V2"] = "true"
os.environ["LANGCHAIN_API_KEY"] = "ls-..."
os.environ["LANGCHAIN_PROJECT"] = "support-agent-prod"
# Now every call is traced automatically
LangSmith UI shows:
- Per-trace view: each step of an agent execution.
- Inputs, outputs, timing, token counts per step.
- Error stack traces.
- Searchable history across all runs.
- Cost breakdown.
For multi-step agents, this is THE debugging tool. Without it, you're squinting at logs trying to reconstruct what happened.
Tracing structure
Each "run" in LangSmith corresponds to a top-level invoke. Nested runs (LLM calls, tool calls, sub-graph executions) appear as children.
Example trace:
Run: graph.invoke (8.3s, $0.034)
├── agent (1.2s, $0.008)
│ └── llm call (1.1s, $0.008)
├── tools (0.5s, $0)
│ └── get_order_status (0.4s, $0)
├── agent (1.0s, $0.006)
│ └── llm call (0.9s, $0.006)
└── (3 more iterations...)
Click any step → see its input, output, prompt, tools available, raw API response.
Searching traces
Find specific runs:
- By user_id (custom metadata).
- By thread_id.
- By error.
- By model version.
- By latency or cost outliers.
graph.invoke(
{"messages": [user_msg]},
config={
"configurable": {"thread_id": "user-123"},
"metadata": {
"user_id": "user-123",
"version": "v1.2.0",
"feature_flag": "new_routing",
}
}
)
Tags and metadata make traces searchable in LangSmith UI.
Datasets and eval
LangSmith doubles as an eval platform:
- Build a dataset of (input, expected_output) examples.
- Run your agent on the dataset.
- Score outputs (deterministic, LLM-as-judge, human).
- Compare runs (e.g., before and after a prompt change).
from langsmith import Client
client = Client()
dataset = client.create_dataset("support-agent-eval")
client.create_examples(
inputs=[{"messages": [...]} for ex in eval_data],
outputs=[{"answer": ex.expected} for ex in eval_data],
dataset_id=dataset.id,
)
# Run evals
from langsmith.evaluation import evaluate
results = evaluate(
lambda inputs: graph.invoke(inputs)["messages"][-1].content,
data=dataset,
evaluators=[my_judge_evaluator],
)
CI/CD integration: run evals on every PR; fail if scores drop.
Layer 3 — Structured logging
Independent of LangSmith, log key events to your standard logging infra:
import structlog
logger = structlog.get_logger()
def call_model(state):
logger.info("agent_decide", thread_id=state.get("thread_id"), iteration=state["iteration"], message_count=len(state["messages"]))
response = llm_with_tools.invoke(state["messages"])
logger.info("agent_decided", tool_calls=len(getattr(response, "tool_calls", [])))
return {"messages": [response], "iteration": state["iteration"] + 1}
For dashboards (Grafana, Datadog):
- Requests per minute.
- p50/p95/p99 latency.
- Error rate.
- Cost per request (aggregate from token usage).
- Hit rate of specific tools.
LangSmith is your "look at one specific run". Logs + dashboards are your "is the system healthy overall?".
Monitoring production agents
What to watch:
Latency
p99 latency. Multi-step agents have variable latency; alert when p99 spikes.
Error rates
- LLM API errors.
- Tool execution errors.
- Validation failures.
- Max-iteration-reached signals.
Cost
Per request and aggregate. Spikes mean something changed (longer prompts, more iterations, model swap).
Tool usage
Which tools fire most? Which never fire (dead code)? Which fire too often (loop bug)?
Conversation length
Average and outliers. Long conversations might be a UX issue or a memory bug.
User feedback
Thumbs up/down, complaints, escalations. Direct quality signal.
Debugging a specific failure
User reports: "The agent gave me a wrong answer."
Process:
- Get the thread_id from your app's logs.
- Open LangSmith, search by thread_id.
- Look at the trace: what tools fired? what did the LLM see at each step? what went wrong?
- Reproduce locally: replay from the offending checkpoint.
- Fix: prompt change, new tool, validation, etc.
Without LangSmith, step 3 is "stare at print() output and guess". With it, the failure mode is usually obvious in 5 minutes.
Setting up LangSmith
pip install langsmith
import os
os.environ["LANGCHAIN_TRACING_V2"] = "true"
os.environ["LANGCHAIN_API_KEY"] = "ls-..."
os.environ["LANGCHAIN_PROJECT"] = "your-project-name"
Get an API key from smith.langchain.com. Free tier covers ~5K traces/month. Paid tier scales up.
For self-hosted environments: LangSmith on-prem exists (Enterprise). Same UX, your infra.
Common observability mistakes
- No tracing at all. Production debugging is then impossible at non-trivial scale.
- Tracing in dev only. Production failures are the ones you need to debug.
- No metadata on traces. Hard to search and aggregate.
- Ignoring cost dashboards. Cost surprises happen.
- No eval in CI. Prompt changes ship without regression testing.
- Treating LangSmith as the only monitor. It's great for individual traces; combine with app-level metrics.
Takeaway
Streaming for per-step visibility. LangSmith for trace-level debugging. Structured logs + dashboards for system-level monitoring. All three together = debuggable, observable agents. Set up LangSmith from day one; the cost is minimal and the productivity gain is enormous.