Evaluating agents is harder than evaluating single LLM calls. A correct final answer might have been reached via wrong reasoning. A wrong answer might have been due to one bad step in an otherwise sound trajectory. Production-grade eval needs multiple layers.
Level 1 — Final-answer eval
Simplest: did the agent produce the right answer?
eval_set = [
{
"input": {"messages": [HumanMessage("What's the weather in Mumbai?")]},
"expected": "Should include the weather condition and temperature for Mumbai",
},
{
"input": {"messages": [HumanMessage("Refund order #99999 from 2 years ago")]},
"expected": "Should refuse (outside policy) and offer alternative",
},
# ...
]
def eval_final_answer(graph, eval_set):
results = []
for case in eval_set:
output = graph.invoke(case["input"])
final_msg = output["messages"][-1].content
# Use LLM-as-judge to evaluate
judgment = judge_llm.invoke([
SystemMessage(f"Evaluate if this answer meets the expectation:\nExpected: {case['expected']}\nActual: {final_msg}\nOutput JSON: {{'meets': bool, 'reason': str}}")
])
results.append(json.loads(judgment.content))
return results
Works well when there's a clear "right answer". Less useful for open-ended tasks.
Level 2 — Trajectory eval
Did the agent take a reasonable path? Even with the right final answer, a wasteful trajectory is a bug.
def eval_trajectory(graph, eval_set):
results = []
for case in eval_set:
# Stream to get all steps
trajectory = []
for chunk in graph.stream(case["input"]):
for node_name, state_update in chunk.items():
trajectory.append({"node": node_name, "update": state_update})
# Check expected trajectory properties
tool_calls = [step for step in trajectory if "tools" in step["node"]]
results.append({
"total_steps": len(trajectory),
"tool_calls_made": len(tool_calls),
"tools_used": [tc["update"]["messages"][0].name for tc in tool_calls if "messages" in tc["update"]],
"loops_detected": detect_loops(trajectory),
})
return results
Things to check:
- Did it call the expected tools? (For known queries, expected tool usage.)
- Total steps within budget? (Agents that take 30 steps for a 3-step problem.)
- Loops? (Calling the same tool repeatedly with same args = bug.)
- Dead ends? (Calling a tool, ignoring result, calling another.)
- Cost? (Total tokens used.)
Level 3 — Component eval
Test individual nodes / sub-graphs in isolation.
def eval_tool_selection_node(test_cases):
# Test that the LLM correctly selects tools given a query
results = []
for case in test_cases:
# Just test the agent node, not the whole graph
state = {"messages": [HumanMessage(case["query"])]}
result = call_model(state)
last_msg = result["messages"][-1]
# Check which tool was called
if hasattr(last_msg, "tool_calls") and last_msg.tool_calls:
chosen_tool = last_msg.tool_calls[0]["name"]
else:
chosen_tool = None
results.append({
"expected_tool": case["expected_tool"],
"actual_tool": chosen_tool,
"correct": chosen_tool == case["expected_tool"],
})
return results
Component evals isolate failure modes:
- Is the LLM picking wrong tools?
- Are tool descriptions misleading?
- Are sub-graphs returning wrong state?
Easier to debug than whole-agent failures.
Building an eval dataset
Start with:
- Real production examples you've seen succeed (regression: don't break what works).
- Real production examples you've seen fail (regression: don't repeat known bugs).
- Edge cases (ambiguous queries, off-topic, adversarial).
- Happy paths for each major use case.
Size: 30-100 examples for initial eval; 200-500 for high-confidence releases.
Maintain it like code: version control, code review, document each example's rationale.
LLM-as-judge for agent outputs
Open-ended outputs need LLM judges:
def llm_judge_response(query, response, criteria):
judge_prompt = f'''Evaluate this agent response on:
- Accuracy: factually correct given the query
- Helpfulness: addresses the user's actual need
- Completeness: covers all parts of the request
- Tone: appropriate for customer support
Query: {query}
Response: {response}
Score each 1-5 with brief reasoning. Output JSON.'''
return judge_llm.invoke([SystemMessage(judge_prompt)])
Use a more capable model as judge than the one being evaluated (frontier judging mid-tier).
Calibrate: rate 30-50 examples by hand, compare to LLM-judge scores. If correlation is low, refine the judge prompt.
Running evals in CI
# .github/workflows/eval.yml
on:
pull_request:
paths: ['agents/**', 'prompts/**']
jobs:
eval:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v3
- run: pip install -r requirements.txt
- run: python scripts/run_evals.py
- run: python scripts/check_regressions.py
check_regressions.py fails the PR if scores drop below threshold. Like unit tests for agent behavior.
Production sample-eval
Sample 1-5% of production runs daily; eval them. Catches drift, prompt regressions, model deprecation effects.
def daily_production_eval():
sample = sample_production_runs(n=100, since="24h ago")
for run in sample:
# Run through judge
score = judge_run(run)
log_to_warehouse({
"run_id": run.id,
"score": score,
"timestamp": now(),
})
Track score over time. Sudden drops = something changed (data, model, prompt, downstream).
Common agent eval mistakes
- Only eyeball testing. Doesn't scale; doesn't catch regressions.
- Final-answer eval only. Misses trajectory issues.
- Static eval set. Production failures don't get incorporated.
- Same judge model for all evals. Same bias affects all scores.
- No regression eval in CI. Ship prompt changes blind.
- No production sample eval. Drift goes unnoticed for months.
Eval-driven development
The agent development cycle:
1. Build initial agent.
2. Build eval set covering known cases + edge cases.
3. Iterate on prompts / tools / graph until eval passes.
4. Deploy.
5. Monitor production; add new failures to eval set.
6. Iterate.
The eval set IS the spec. Production agents that survive long-term have careful eval discipline.
Takeaway
Three levels: final-answer (did it succeed?), trajectory (did it succeed efficiently?), component (which piece failed?). Build an eval set early; treat it as code; run in CI. Sample production traces daily. Use LLM-as-judge with calibration. The teams that ship reliable agents do this; the teams that don't, ship surprises.