Every major LLM provider now offers prompt caching: cache the static prefix of your prompt so subsequent calls reuse it at a fraction of the cost. For most production apps, this is a 5-10x cost reduction. It's the highest-ROI optimization you can make.
How it works
Your prompt has two parts:
- Static prefix: system prompt, knowledge base, examples, tool definitions. Same across many requests.
- Dynamic suffix: user message, conversation turn. Different each request.
Cache the prefix. The provider:
- First call: process the prefix as normal, then cache it. Small overhead.
- Subsequent calls (within TTL): reuse the cached prefix at ~10% the input cost.
# Anthropic example
response = client.messages.create(
model="claude-sonnet-4-6",
system=[
{
"type": "text",
"text": "You are a customer support agent for Acme. " + LARGE_KNOWLEDGE_BASE,
"cache_control": {"type": "ephemeral"}
}
],
messages=[
{"role": "user", "content": user_message}
],
)
The system prompt + knowledge base are cached. Subsequent calls with the same prefix hit the cache.
What to cache
Best caching candidates:
- System prompts that don't change per user.
- Large knowledge bases that stay constant.
- Few-shot examples repeated across calls.
- Tool definitions (often 2-5K tokens).
- Long documents in document Q&A apps.
In short: anything static that you'd be sending in every request.
What NOT to cache
- Truly dynamic content (user messages, retrieved RAG snippets per query).
- Very short prefixes (caching overhead > savings).
- Content that changes often.
Cost math
Example RAG app:
- Total input per call: 6000 tokens.
- Of which static (cacheable): 4500 tokens (system + examples + tools).
- Dynamic: 1500 tokens (user message + retrieved chunks).
Without caching at Claude Sonnet ($3 input / $15 output per 1M):
- Input: 6000 × $3/1M = $0.018.
- 100K requests/month: $1800.
With caching:
- First call: $0.018 + small caching overhead.
- Cached calls: 4500 × $0.30/1M (cached price) + 1500 × $3/1M = $0.00135 + $0.0045 = $0.00585.
- 100K requests/month: $585 (after cache warmup).
3x cost reduction. For longer cached prefixes (10K+ tokens), the multiplier is even bigger.
TTL and cache management
Caches typically expire in 5 minutes (varies by provider). For high-traffic apps, this is fine — calls land within the TTL and hit cache.
For lower-traffic apps:
- Build a keepalive: ping the cache every 4 minutes with a no-op request.
- Or use longer TTL options (some providers offer extended TTL at higher cost).
Cache management strategies
Strategy 1: Pure static prefix
Same prefix for every user. Cache once, all users benefit.
SYSTEM_PROMPT = load_static_system_prompt() # cached
USER_MESSAGE = request.message # dynamic
Simplest case. Hugely effective.
Strategy 2: Per-user caches
If user-specific data is large and stable (per-user knowledge base, preferences), cache that per user.
def get_response(user_id, message):
user_context = load_user_context(user_id) # cacheable, ~5K tokens
return llm.messages.create(
system=[
{"type": "text", "text": user_context, "cache_control": {"type": "ephemeral"}}
],
...
)
Each user gets their own cache slot. Effective for sticky-user apps.
Strategy 3: Tool/template prefix
Cache the tool definitions and instructions; vary the user-specific data after.
Other cost optimization techniques
Beyond caching:
1. Right-sized model
Don't use Claude Opus when Sonnet would do. Don't use Sonnet when Haiku would do. Test cheaper models on your actual task.
2. Tight RAG retrieval
5 chunks instead of 20. Better reranking. Smaller chunk size if topic-focused.
3. Output token capping
Set max_tokens. Prompt for brevity.
4. Batch processing
For non-time-sensitive work, use batch APIs (50% discount, async).
5. Distillation
Generate examples with frontier model; fine-tune a small open model on them. Production at 100x lower cost.
6. Prompt compression
Tools like LLMLingua compress prompts while preserving meaning. Underused; works.
Monitoring caching
Track in production:
- Cache hit rate per endpoint.
- Tokens saved per day.
- Cost saved per day.
SELECT
DATE(request_at) AS day,
SUM(input_tokens) AS total_input,
SUM(cached_input_tokens) AS cached_input,
ROUND(100.0 * SUM(cached_input_tokens) / SUM(input_tokens), 1) AS cache_hit_pct,
SUM(estimated_cost) AS daily_cost
FROM llm_requests
GROUP BY day;
If cache hit rate drops, investigate — usually a static prefix changed inadvertently.
Common caching mistakes
- Not using caching at all — leaving 5-10x savings on the table.
- Putting dynamic content in the cached portion — invalidates the cache constantly.
- Cache key changes — a one-character difference in system prompt breaks caching. Be deliberate about the static prefix.
- No monitoring — degrading cache hit rate goes unnoticed.
- Caching tiny prefixes — caching overhead > savings below ~1K tokens.
The cost-optimization sequence
In order of impact for most production apps:
- Prompt caching for static prefixes. (5-10x reduction)
- Right-sized model for the task. (2-10x reduction)
- Tight RAG retrieval + reranking. (2-3x reduction)
- Output capping + brevity prompts. (1.5-2x reduction)
- Batch processing for async workloads. (2x reduction)
- Fine-tune small model for high-volume narrow tasks. (10-100x reduction)
Each compounds. Apps that do all of these have unit economics that work; apps that don't, don't.
Takeaway
Prompt caching is the single highest-ROI optimization in 2026 LLM apps. Cache static prefixes; monitor hit rates; combine with right-sized models and tight RAG. Production-grade LLM apps live and die by cost discipline — and caching is where most of it happens.