Fine-tuning = training a pretrained LLM further on your specific data. Two questions: when is it worth it, and how to do it efficiently.
When fine-tuning makes sense
Three conditions to justify fine-tuning over prompt engineering + RAG:
- The model can mostly do the task but not reliably.
- You need consistent style/format that prompting alone doesn't produce.
- Volume is high enough that fine-tuning's cost amortizes.
If any condition is missing, prompt + RAG is probably better.
When fine-tuning does NOT make sense
- Adding facts to the model. Facts go in RAG. Fine-tuning poorly absorbs factual updates.
- You have less than ~50-100 high-quality training examples. Below this, you get worse performance than prompting.
- The base model can't do the task at all. Fine-tuning doesn't add capabilities the base model lacks; it specializes existing capability.
- Frequent updates required. Each retrain is expensive; RAG updates instantly.
- Low-volume use case. API costs don't justify the fine-tuning investment.
Common fine-tuning use cases (where it shines)
- Format consistency — always produce JSON with a specific schema, no deviations.
- Tone/style — corporate voice, brand-specific writing.
- Specialized classification — classify into domain-specific categories.
- Structured extraction — pull specific fields from messy text reliably.
- Code generation in your codebase's style.
- Distillation — fine-tune a small model to behave like a frontier model on your task.
Data requirements
Quality matters more than quantity.
| Data quality | Minimum examples |
|---|---|
| High-quality, diverse, representative | 100-500 |
| Mixed quality | 500-2000 |
| Auto-generated / synthetic | 2000-10000 |
Synthetic data (generated by a frontier model) is increasingly common — distill GPT-5/Claude 4 Opus outputs into a smaller model.
What "high quality" means:
- Input/output pairs that match real production usage.
- Diverse: covers the distribution of real cases.
- Clean: no errors in the expected outputs.
- Format-consistent.
Building a good dataset is most of the fine-tuning work.
Fine-tuning methods
Full fine-tuning
Update all model parameters. Cost: prohibitive for large models. Rare today.
LoRA (Low-Rank Adaptation)
Train small "adapter" matrices on top of frozen base model parameters. ~100-1000x cheaper than full fine-tuning.
QLoRA (Quantized LoRA)
LoRA + 4-bit quantization. Even cheaper. Now the default for cost-conscious fine-tuning.
Provider-managed fine-tuning
OpenAI, Anthropic (limited), Google all offer managed fine-tuning APIs. You provide data, they handle the rest. Convenient but more expensive.
Costs (2026 rough estimates)
Closed model fine-tuning
- OpenAI GPT-4o-mini fine-tune: ~$3 per 1M training tokens + ~10% premium per inference token.
- Anthropic (limited beta): similar pricing.
For 10K training examples × 500 tokens each = 5M training tokens × $3 = $15. Cheap.
Open model fine-tuning (LoRA)
- Self-hosted on a single GPU: ~$5-50 in compute for a few-hour LoRA fine-tune on a 7-13B model.
- Managed services (Together AI, Modal, Replicate): $20-200 for similar.
Open model fine-tunes cost less, but you need to host them after. Hosting a 7B model: ~$0.20 per 1M tokens on Together AI; or self-hosted on a $200/month GPU.
The fine-tuning workflow
1. Collect training data (input/output pairs).
2. Split into train (90%) and eval (10%).
3. Validate data quality (manual sample, schema check).
4. Run fine-tuning (LoRA on a base model).
5. Evaluate fine-tuned model vs base model on eval set.
6. If better → deploy. If worse → iterate on data.
Most failures happen at step 1 (bad data) or step 5 (no proper eval).
Iteration matters
First fine-tune is rarely great. Production teams iterate:
- Add more examples for edge cases.
- Remove bad examples (manual review).
- Adjust hyperparameters (learning rate, epochs).
- Try different base models.
Plan for 3-5 iterations to reach production quality.
Distillation as a fine-tuning strategy
The pattern:
- Use a frontier model (Claude 4 Opus, GPT-5) to produce high-quality outputs on your task.
- Collect those outputs as training data.
- Fine-tune a smaller, cheaper model (Llama 3 8B, GPT-4o-mini) on this data.
- Deploy the small model in production.
Result: production performance close to frontier at 10-100x lower cost.
This is increasingly the standard path for high-volume narrow tasks. Don't run frontier in production if you can distill.
Evaluating a fine-tuned model
Before promoting to production:
- Held-out eval set — never seen during training.
- Task-specific metrics — accuracy, F1, exact match, BLEU, or LLM-as-judge.
- Compare to base model on the same eval set.
- Compare to base model + good prompt on the same eval set.
If fine-tuned ≠ better than base+prompt, don't deploy. The fine-tuning didn't help.
Common fine-tuning mistakes
- Too little data. Below 100 examples, prompting beats fine-tuning.
- Poor data quality. Garbage in, garbage out.
- No eval set. Can't tell if fine-tuning worked.
- Not comparing to base+prompt baseline. Maybe a good prompt was enough.
- Fine-tuning for facts. Use RAG instead.
- Skipping iteration. Production fine-tuning is iterative.
- Hosting cost surprise. Fine-tuned closed models are 10% more per token; fine-tuned open models need self-hosting infra.
Takeaway
Fine-tuning is for style/format/specialization, not for adding facts. LoRA / QLoRA make it 100-1000x cheaper than full fine-tuning. Data quality matters more than quantity. Always have an eval set and compare to base + prompt baseline. Most production fine-tuning is distillation: frontier model generates training data, small model serves in production.