In a synchronous ML service, latency is what determines whether your model can be used at all. If the page load times out waiting for your prediction, the prediction is worthless.
Setting a latency budget
Define the target explicitly. For an online ML service:
- p50 (median): typical user experience.
- p95: most users' experience including some delays.
- p99: rare slow requests.
A typical fraud-detection budget:
- p50 < 30ms
- p95 < 80ms
- p99 < 150ms
A web page loading a recommendation:
- p50 < 100ms
- p95 < 300ms
Slower budgets are OK for less-critical use cases. Faster budgets need engineering work.
What contributes to latency
Break down a prediction request:
Total = network_in + feature_lookup + feature_compute + model_predict + network_out + logging
For a typical service:
- Network in: ~5-10ms (client to server).
- Feature lookup (warehouse / cache): 1-50ms.
- Feature compute: 1-10ms.
- Model predict: 1-50ms depending on model.
- Network out: ~5-10ms.
- Logging: 1-5ms (if async, doesn't count).
Total: 13-135ms for a single prediction. p99 is dominated by feature lookups and model prediction time.
Optimizing each stage
Network
Use HTTP/2, keep-alive connections, gzip compression. Most modern frameworks handle this. Don't fight network.
Feature lookups
Most expensive part for many systems. Options:
- In-memory cache (Redis, Memcached): sub-millisecond lookups.
- Async pre-fetch: while user is doing other things, fetch features in background.
- Pre-compute and store: batch-compute features hourly, store in fast-lookup table.
For low-cardinality feature lookups (e.g., user-level features for 1M users), Redis costs almost nothing and saves dozens of milliseconds.
Feature computation
- Avoid Python loops in feature code: vectorize with NumPy/Pandas where possible.
- Pre-compute slow features: anything based on time-windows over a user's history should be in a feature store, not computed per request.
Model inference
- Lightweight models (logistic regression, small GBMs): sub-millisecond.
- Big GBM models (XGBoost 1000 trees): few milliseconds.
- Deep networks: 5-50ms on CPU, sub-millisecond on GPU.
Optimization:
- Quantization: int8 instead of float32 → 4× smaller, often 2-4× faster.
- Distillation: train a smaller model to mimic the big one.
- Tree pruning: fewer trees with little accuracy loss.
- ONNX Runtime / TensorRT: optimized inference engines.
For most tabular ML, a tuned XGBoost runs in 1-5ms. Don't optimize prematurely.
Logging
- Async logging (background tasks): doesn't block prediction.
- Batch writes to logging system.
Caching predictions
If the same input often produces the same prediction, cache:
from functools import lru_cache
@lru_cache(maxsize=10000)
def cached_predict(features_tuple):
return model.predict_proba([features_tuple])[0][1]
Or Redis-backed cache with TTL.
When this helps: features that don't change much (user_id-based predictions where the user's features are stable for the day). When this doesn't help: features that change per request (transaction context).
Autoscaling
Traffic varies. Autoscaling adjusts the number of service instances.
Horizontal scaling (more instances)
The standard pattern. Cloud Run / Kubernetes scales up replicas based on CPU or QPS:
# Cloud Run example
min_instances: 1
max_instances: 100
concurrency: 80 # requests per instance simultaneously
Add instances as traffic grows. Remove during quiet periods.
Vertical scaling (bigger instances)
More CPU/memory per instance. Useful for memory-hungry models (large embeddings). Costlier per request than horizontal.
Scale-to-zero
Some platforms (Cloud Run, Lambda) can scale to zero when no traffic. Cold-start penalty (loading the model on first request after a cold start). Trade-off: cost savings vs first-request latency.
For sporadic traffic: scale-to-zero is fine. For constant traffic: keep min_instances ≥ 1 to avoid cold starts.
Concurrency vs replication
Most Python ML services use Gunicorn / Uvicorn with multiple workers:
uvicorn app:app --workers 4
4 workers per instance, 4 instances = 16 concurrent requests. Each worker has its own copy of the model in memory.
Trade-off: workers consume memory (4× model size per instance). For big models, prefer instances over workers.
Load testing
Before deploying, simulate traffic:
import locust
from locust import HttpUser, task
class FraudUser(HttpUser):
@task
def predict(self):
self.client.post('/predict', json={
'transaction_id': 'test',
'amount': 100.0,
# ...
})
Run with locust -f loadtest.py --users 100 --spawn-rate 10. Watch:
- p99 latency at target QPS.
- Error rate.
- CPU/memory usage.
If you hit the latency budget at expected QPS, you can ship. If not, optimize the bottleneck.
When latency budget is unmeetable
Sometimes the budget genuinely can't be met:
- Model takes too long.
- Feature lookups are slow.
- Single-request budget too tight for any ML.
Options:
- Smaller model (lower accuracy, faster).
- Pre-compute predictions (batch, served from cache).
- Skip ML on the synchronous path; do it offline.
- Set a more realistic budget.
It's better to deliver a 50ms prediction with 0.84 AUC than a 200ms prediction with 0.85 AUC that breaks the page experience.
Common latency mistakes
- No budget set — "fast" is subjective.
- Load testing only locally — production has network and concurrency.
- Synchronous logging — adds latency for no benefit.
- Loading model per request — many seconds added.
- No cache for stable features — querying warehouse on every request.
- Over-using GPUs — for tabular ML, CPU is often as fast and much cheaper.
Takeaway
Set explicit latency budget (p50/p95/p99). Identify the bottleneck (usually feature lookup or model inference). Cache aggressively. Async-log. Autoscale horizontally. Load-test before deploying. If you can't meet the budget, choose a smaller model or a different serving pattern.