Two ways to serve a model:
Batch inference
Score many examples at once, on a schedule. Save predictions to a database. Downstream systems read those predictions.
# Daily batch scoring job
all_users = load_users_to_score()
features = build_features_batch(all_users)
predictions = model.predict_proba(features)
save_to_db({
'user_id': all_users.user_id,
'prediction': predictions[:, 1],
'scored_at': datetime.utcnow(),
'model_version': 'v1.2.0',
})
Run via cron, Airflow, GitHub Actions, dbt — whatever orchestrates your existing data pipelines.
Pros:
- Simple: one Python script.
- No latency concerns.
- Same infrastructure as ETL.
- Easy to log all predictions.
- Cheap.
Cons:
- Predictions are stale (as old as the last batch).
- Can't react to events as they happen.
Online inference
Serve predictions in real-time, one at a time, in response to requests.
# FastAPI endpoint
@app.post('/predict')
def predict(req: TransactionRequest):
features = build_features_online(req)
prediction = model.predict_proba([features])[0][1]
return {'fraud_probability': prediction}
Pros:
- Up-to-date predictions on every request.
- Required for use cases where prediction must precede action (fraud blocking, ad bidding, recommendations on a live page).
Cons:
- Need an always-up service.
- Latency budget (often <100ms).
- Scaling, monitoring, deployment overhead.
- More expensive to operate.
The honest question: do you need online?
Most ML systems don't. The cases that genuinely require online inference:
- Sub-second decisions: fraud blocking, ad bidding, real-time recommendations.
- Personalization on live pages: showing a user-specific product list as they browse.
- Action-trigger systems: chatbot, voice assistant, autopilot.
The cases that don't:
- Churn prediction (act on it next email cycle — batch is fine).
- LTV forecasting (used in monthly planning — batch).
- Customer segmentation (slow-moving — daily batch).
- Lead scoring (sales team works in business hours — batch).
- Most "analytics-adjacent" ML.
For batch-OK use cases, building an online service is a self-imposed tax.
The middle ground: near-real-time
Some systems need fresh predictions but not synchronously:
- Predictions refreshed when underlying data changes (event-triggered batch).
- Cached predictions with TTL — refresh in background, serve cached.
- Streaming jobs that score events as they arrive (Kafka + Flink / Spark Streaming).
A daily batch is fine if the prediction is stable for a day. If it changes by the hour, hourly batch. If by the minute, streaming or online.
Cost comparison
For 10 million predictions per day:
- Batch: one job at 2am, scoring 10M rows, takes ~10 minutes on a small cluster. $5/day in compute. Storage of predictions: minimal.
- Online: 100 QPS sustained. Need 3-5 instances of a web service for redundancy + autoscaling. $50-200/day depending on cloud provider. Plus monitoring infrastructure.
Online is 10-40× more expensive operationally for the same volume. Only pay this when you need it.
Hybrid: batch + online
Common pattern: batch-pre-compute the slow parts; online-compute the fast parts.
Example: recommendation system:
- Batch (nightly): train embeddings, pre-score candidate items per user.
- Online (per request): filter candidates by current context (page, query), re-rank with a lightweight model.
Best of both: heavy ML done offline, fast personalization done online.
Building for online
If you must go online:
- Latency budget — set it explicitly. p99 < 100ms is a common target.
- Cache aggressively — predictions, feature lookups, both.
- Limit feature complexity — every feature is computed on the request path.
- Failover plan — if the model service is down, what does the API return? A cached default; a simple fallback model; a graceful error.
Common pattern mistakes
- Building online for batch-OK use cases — self-imposed complexity.
- Latency budgets only discovered at load test — design with latency in mind.
- Forgetting that batch can serve "real-time" via cache — cache last batch's predictions; serve them online instantly. Refresh in background.
- No fallback plan — model service goes down, the whole product goes down.
Decision tree
Does the consumer need the prediction within 1 second of the event?
├─ No → batch
└─ Yes
├─ Can a 1-hour-old prediction (cached) work? → cached batch
└─ Needs prediction reflecting current state → online inference
Takeaway
Batch covers 90% of ML use cases at a fraction of the operational cost. Reserve online inference for genuinely synchronous decisions (fraud, ads, live personalization). Hybrid pre-compute + online filter is the sophisticated middle ground. Always ask: does this need to be online?