Software incidents are usually loud (500 errors, latency spikes, services down). ML incidents are often silent — predictions are wrong but the service runs fine. The playbook needs to handle both.
Step 1 — Detect
For each ML system, set up alerts on:
Loud signals (service health)
- Service down / 503s.
- High latency (p99 > budget).
- Error rate > threshold.
These fire fast and clearly.
Quiet signals (model behavior)
- Prediction distribution shift > threshold.
- Action rate (e.g., % flagged) outside normal range.
- Business KPI dropping.
- Drift in input features.
These are slower, often only noticed when business stakeholders raise concerns.
A good ML alerting strategy combines both. Alerts route to:
- Engineer on-call for loud signals.
- Slack channel for quiet signals (investigation, not immediate action).
Step 2 — Investigate
When an alert fires:
5-minute checks
- Is the service up?
- Are predictions being made?
- Has anything been deployed recently?
15-minute checks
- Has the input data changed (upstream pipeline issue)?
- Has the feature engineering changed?
- Has the model version changed?
- Has the model started predicting weird values (distribution shift)?
Going deeper
- Pull recent predictions, look at examples.
- Compare predicted distributions across last 7 days, 30 days, 90 days.
- Spot-check a few requests end-to-end.
Most incidents are resolved at the 5-15 minute level.
Step 3 — Stop the bleeding
Determine the right response:
Severity 1: Service down or producing garbage
- Roll back to previous model version (config change).
- If pipeline issue, switch to fallback prediction (baseline, last cached, simple rule).
Severity 2: Service up but degraded
- Continue investigating.
- Communicate to stakeholders.
- Consider rollback if degradation persists.
Severity 3: Suspicious signals but no clear harm
- Monitor closely.
- Don't take drastic action yet.
Step 4 — Root cause
Once stable, find the actual cause. Common categories:
Pipeline bugs
- Upstream data source changed schema.
- Feature engineering function has a bug.
- A new dependency broke compatibility.
Drift
- Real data drift or concept drift.
- Out-of-distribution requests at scale.
Deployment issues
- New model version has a subtle issue not caught in pre-deployment testing.
- Environment changed (Docker image, dependencies).
Calibration / threshold issues
- Model probabilities shifted; default threshold no longer right.
Step 5 — Fix and verify
Fix the root cause. Critically:
- Don't just fix the symptom (e.g., increase threshold to compensate for drift).
- Fix the underlying issue.
- Re-test in staging.
- Roll out carefully (shadow → canary → full).
Step 6 — Document
Write a postmortem:
- What happened (timeline, observed symptoms).
- Root cause (technical).
- How it was detected.
- How it was resolved.
- What we'd do differently.
- Action items to prevent recurrence.
Even short postmortems compound team knowledge. The fifth time something similar happens, you'll have a playbook.
Common ML incident scenarios
Scenario A: Model predictions stuck at one value
Model predicting 0.05 for everything. Probably:
- Feature pipeline broken (returning NaN, zero, default).
- Model file corrupted at startup.
- Wrong model version loaded.
Fix: roll back, investigate.
Scenario B: Suddenly more false positives
Model flagging way more transactions as fraud. Probably:
- Threshold drift (model probabilities recalibrated).
- Real fraud spike (validate against ground truth).
- Bug in scoring logic.
Fix: raise threshold temporarily, investigate, retrain or recalibrate.
Scenario C: Service is fine but business KPI dropped
Conversion rate down 5%. Could be the model, could be many other things.
- A/B test against holdout group: if holdout has same drop, not the model.
- If only the new model's cohort has the drop, roll back.
Fix: rule out the model first via holdout comparison.
Scenario D: Slow degradation over weeks
AUC silently dropping. Classic concept drift.
- Retrain with recent data.
- Investigate whether features have changed.
Fix: retrain, validate, deploy.
The rollback path
You should be able to roll back any deployed model in under 5 minutes. The recipe:
- Update config to point at the previous version.
- Reload / restart service.
- Verify it's serving the old model (check
/healthreturns the right version). - Communicate the rollback.
If rollback requires a code deploy (build, test, push, deploy), that's hours. Configure rollback to be a config change.
Pre-incident preparation
The best incident response is one you've prepared for:
- Runbooks: written instructions for common scenarios.
- Dashboards: monitoring already set up before incident.
- Alerts: tuned (not too noisy, not too quiet).
- Stakeholder contacts: who to call when business impact is suspected.
- Rollback procedures: tested in non-emergency conditions.
Run a "game day" exercise: simulate an incident in staging, follow the runbook, find the gaps. Cheaper than learning during a real incident.
On-call for ML systems
In smaller teams, ML engineers do their own on-call. Issues:
- Most rotations have ML engineers who don't speak ops.
- Most ops people don't know how to diagnose model issues.
Mitigations:
- Cross-train: ops people learn enough to know "the model looks weird, escalate to ML on-call".
- Document clearly: "if the prediction mean drifts more than X, escalate."
- Don't make every ML engineer carry a pager; identify a small ML-on-call rotation.
Common incident response mistakes
- No runbook — every incident is improvised, takes longer than necessary.
- No rollback procedure — emergency code deploys during incidents are dangerous.
- No postmortem — same incidents recur.
- Treating ML incidents like software incidents — ignoring quiet signals.
- Over-rolling-back — rolling back when investigation would have been better.
Takeaway
Detect via loud + quiet signals. Investigate quickly (most issues resolve in 15 min of triage). Stop bleeding (rollback or fallback). Root-cause and fix the underlying issue. Document via postmortem. Prepare runbooks before incidents. Rollback must be config-only, sub-5-minute. The team that practices incident response handles them; the team that doesn't, panics.