Evaluation is the most-neglected part of production LLM work. Without it, every prompt change is faith-based; regressions are silent.
Three layers of eval, ordered from easy to sophisticated:
Layer 1 — Deterministic checks
For tasks with known-correct outputs:
def eval_extraction_task(model, eval_set):
correct = 0
for example in eval_set:
output = model.generate(example.input)
expected = example.expected_output
if output == expected:
correct += 1
return correct / len(eval_set)
Works for:
- Classification (predicted label vs ground truth).
- Structured extraction (output JSON vs expected JSON).
- Code that compiles and passes tests.
- Math with exact answers.
When you can write if output == expected, you have deterministic eval. Fast, cheap, reliable.
Layer 2 — LLM-as-judge
For open-ended outputs where there's no single right answer:
def llm_judge(model_output, expected_qualities, reference_answer=None):
judge_prompt = f'''
Evaluate this model output on these criteria:
{expected_qualities}
{f"Reference answer: {reference_answer}" if reference_answer else ""}
Model output: {model_output}
Score 1-5 for each criterion. Output JSON.
'''
return judge_model.generate(judge_prompt)
Works for:
- Chat responses (helpful? accurate? appropriate tone?).
- Summarization (covers key points? concise?).
- Translation (preserves meaning? natural?).
- Generated content quality.
Best practices:
- Use a more capable model as judge than the model being evaluated (frontier model judging mid-tier output).
- Calibrate by checking that the judge agrees with human ratings on a sample.
- Use multiple criteria; produce a structured score.
Limitations:
- Judges have biases (length, verbosity, format).
- Cost: each eval call costs API tokens.
- Not perfectly aligned with human preference.
But for open-ended tasks, LLM-as-judge is the best automatable option.
Layer 3 — Human evaluation
Have humans rate outputs.
When to use:
- High-stakes deployments (medical, legal, financial).
- Calibrating LLM-as-judge.
- Edge cases LLM-judge handles poorly.
How to do it:
- 50-200 examples covering the variety of real inputs.
- Multiple raters when feasible.
- Clear rubric (don't say "rate it good or bad" — specify criteria).
- Track inter-rater agreement.
Cost: human time. Not scalable for every release but essential for grounding automatable evals.
RAG-specific eval
RAG has its own dimensions:
Retrieval eval
- Recall@K: of the relevant chunks for this query, what fraction did we retrieve in the top K?
- Precision@K: of the top K retrieved chunks, what fraction are actually relevant?
- MRR (Mean Reciprocal Rank): how high in the ranking is the first relevant chunk?
These require labeled data: for each query, which chunks ARE relevant.
Generation eval
- Faithfulness: does the answer stay grounded in the retrieved context (no hallucination)?
- Answer relevance: does the answer actually address the question?
- Context relevance: was the retrieved context relevant to the question?
Tools like Ragas (open-source) automate these with LLM-as-judge.
Building an eval set
The eval set is your eval. Good eval sets cover:
- Common cases — the bulk of real requests.
- Edge cases — the weird ones that often fail.
- Known regressions — cases that broke in production; lock them in tests.
- Adversarial cases — prompt injection attempts, jailbreaks, ambiguity.
Size: 30-100 for fast iteration; 200-500 for higher confidence releases.
Maintain the eval set like code: version it, review changes, document the rationale for each example.
When to run evals
In an LLM development cycle:
- Locally during prompt iteration: run quick eval (50 examples) after each prompt change.
- In CI: run full eval (200+ examples) before merging.
- In production monitoring: sample real production outputs and run eval continuously.
# .github/workflows/llm-eval.yml
on: pull_request
jobs:
eval:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v3
- run: python scripts/run_eval.py
- run: python scripts/check_eval_regression.py
The PR fails if eval scores drop below threshold. Like unit tests, but for LLM behavior.
Hallucination detection
A specific eval concern: did the model make up facts?
Approaches:
- Fact-checking against source: for RAG, check if claims appear in retrieved context.
- Cross-verification: ask the same question multiple times; flag inconsistent answers.
- Citation-required prompting: prompt forces citations; flag uncited claims.
- Specialized hallucination detection: tools like Vectara HHEM, RAGAS faithfulness.
For trust-sensitive apps, hallucination eval is non-negotiable.
Latency and cost as eval dimensions
Don't just eval quality. Also track:
- Latency: p50, p95, p99 per request.
- Cost: tokens consumed per request.
A more accurate model at 10x cost and 5x latency might lose to a slightly less accurate model on UX and unit economics. Eval the full picture.
Common eval mistakes
- No eval at all. Most common; biggest mistake.
- Only eyeball testing. Inconsistent; doesn't catch regressions.
- Single-metric eval. Quality has many dimensions.
- Eval set leaked into prompt iteration. Same content in prompt examples AND eval — inflated scores.
- Not updating eval set. Production failures don't get codified into tests.
- Judge model biases ignored. Verbosity bias, length bias, format bias — calibrate.
What good LLM eval looks like
- Eval set of 50-200 examples, versioned in git.
- Deterministic eval where possible, LLM-as-judge where not, human sample for calibration.
- Run in CI on every prompt change.
- Track results over time; alert on regressions.
- Include latency and cost alongside quality.
- Hallucination check for RAG apps.
- Production sample monitoring (eval a slice of real outputs daily).
The teams that build this iterate quickly and confidently. The teams that don't ship invisible regressions.
Takeaway
Three layers: deterministic > LLM-as-judge > human. Build an eval set early; treat it like a test suite. Run on every change. Track regressions. Without eval, every prompt change is a guess.