This lesson on Monitoring / Drift (Latency, Errors, PSI-Style) is hands-on and example-driven. You will be able to design and implement operational dashboards tracking model latency and errors. Crucially, you will configure PSI-style drift detection and set up alert rules using burn-in windows to ensure model reliability in production.
What You'll Be Able To Do
- Design operational dashboards for model latency and error rates.
- Calculate Population Stability Index (PSI) for production feature distributions.
- Configure alert thresholds based on statistical drift metrics (e.g., PSI).
- Implement burn-in windows to prevent alert storms during model deployment.
- Sketch a production drift dashboard integrating latency, errors, and PSI.
- Define an alert rule that triggers a postmortem process.
Detailed Concept Walkthrough
1. Operational Model Health Monitoring
Tracking inference request metrics (latency, error codes) provides real-time insight into the model serving infrastructure's performance, separate from model quality. High latency or error rates often indicate infrastructure bottlenecks or deployment failures.
- Mechanism: Model servers emit metrics (e.g., Prometheus/CloudWatch format) for every request, logging total duration and HTTP status codes (2xx success, 4xx client error, 5xx server error). These metrics are essential for Service Level Objective (SLO) adherence.
- Under the Hood: Latency is tracked using histograms or percentiles (P95, P99) rather than simple averages. Averages hide critical tail performance issues where a small percentage of users experience very slow responses, violating user experience standards.
- Best Practice: Always monitor 5xx errors (server-side) as critical infrastructure failures, and track 4xx errors (client-side) to identify potential data format or input validation issues from upstream services that need correction.
# Example using a hypothetical metrics library in the model server
REQUEST_LATENCY = Histogram(
'model_inference_latency_seconds',
'Inference request latency in seconds',
buckets=(.01, .05, .1, .5, 1.0, 2.0) # Track common percentiles
)
REQUEST_ERRORS = Counter(
'model_inference_errors_total',
'Total count of inference errors by status code',
['status_code']
)
Key Takeaway: Monitor P95/P99 latency and 5xx error rates to ensure the model infrastructure is stable and responsive.
2. Population Stability Index (PSI)
PSI is a statistical measure used to quantify how much a feature's distribution has changed between a baseline (training/reference) and a current production window. It provides a single score indicating the magnitude of data drift.
- Mechanism: The feature range is divided into 10-20 bins (buckets). For each bin, the percentage of observations in the baseline ($P_{baseline}$) and the current period ($P_{current}$) are calculated to determine the relative shift in density.
- Under the Hood: PSI is calculated as $\sum_{i=1}^{N} ((P_{current, i} - P_{baseline, i}) \times \ln(\frac{P_{current, i}}{P_{baseline, i}}))$. This formula weights the difference by the log of the ratio, penalizing larger relative shifts more heavily.
- Best Practice: PSI is most effective for numerical and ordinal categorical features where binning is straightforward. For high-cardinality categorical features, consider grouping rare values into an 'Other' bin before calculating PSI to maintain statistical relevance.
# PSI Thresholds (common industry standards)
PSI_WARNING_THRESHOLD = 0.10 # Investigate
PSI_CRITICAL_THRESHOLD = 0.25 # Alert/Retrain
# Conceptual PSI calculation logic (requires numpy for log)
# psi_score = sum((P_current - P_baseline) * np.log(P_current / P_baseline))
Key Takeaway: PSI quantifies data distribution shift; scores above 0.25 usually mandate immediate investigation or model retraining.
3. Alerting Thresholds and Burn-in
Alert thresholds define the critical limits for monitored metrics (latency, errors, PSI). Burn-in windows are temporary suppression periods applied immediately after a deployment to prevent false alarms caused by initial system stabilization or traffic ramp-up.
- Mechanism: Alerts are triggered only when a metric exceeds a defined threshold (e.g., P99 latency > 500ms) for a specified, sustained duration (e.g., 5 minutes). This duration prevents transient spikes from causing alert noise.
- Under the Hood: A burn-in window (e.g., 30 minutes) is configured in the alerting system (e.g., Alertmanager). During this window, any alert related to the newly deployed model version is suppressed or routed to a lower-priority channel.
- Best Practice: Set separate warning and critical thresholds for drift metrics. A warning (PSI 0.1) might trigger automated data logging, while a critical alert (PSI 0.25) triggers human intervention and feeds into the postmortem process (M5-L18).
# Prometheus Alertmanager Rule Example (YAML)
- alert: CriticalModelDrift
expr: model_psi_score{feature="age"} > 0.25
for: 10m # Must be sustained for 10 minutes
labels:
severity: critical
runbook: M5-L18_Drift_Postmortem
annotations:
summary: "Critical PSI drift detected on feature 'age'."
Key Takeaway: Use sustained thresholds (
for: Xm) to filter noise, and implement burn-in windows post-deployment to avoid unnecessary alerts.
Topics Covered in Monitoring / Drift (Latency, Errors, PSI-Style)
- Operational Monitoring (0:00 - 0:45) — Operational monitoring tracks infrastructure performance metrics like latency and errors.
- Latency Percentiles (0:45 - 1:45) — P95 and P99 latency are crucial indicators of poor user experience and tail performance issues.
- Introduction to Drift (1:45 - 3:00) — Drift monitoring detects changes in data characteristics or model behavior over time.
- Calculating PSI (3:00 - 4:30) — Population Stability Index quantifies distribution change by comparing feature bins between reference and current data.
- PSI Thresholds (4:30 - 5:45) — Industry standard thresholds (0.1 warning, 0.25 critical) guide automated response actions.
- Alerting Mechanics (5:45 - 6:45) — Alerts require sustained metric breaches over a time window to filter out transient noise.
- Burn-in Windows (6:45 - 8:00) — Burn-in windows suppress alerts immediately following deployment to allow system stabilization.
MLOps + Cloud Deploy (AWS-First) Cheat Sheet
-
P95 Latency— 95% of requests complete faster than this valuemodel_latency_seconds{quantile="0.95"} -
Population Stability Index— Quantifies distribution shift between two datasetsPSI = sum((P_c - P_b) * ln(P_c / P_b)) -
Burn-in Window— Time period suppressing alerts post-deploymentalert_suppress_duration: 30m -
5xx Error Rate— Measures server-side infrastructure failuressum(rate(http_requests_total{status="5xx"}[5m])) -
PSI Critical Threshold— Score indicating severe data drift requiring action0.25 -
Alertfor:`` — Specifies how long a condition must be sustainedfor: 5m
Comparison Table
| Metric Type | Focus | Alert Trigger |
|---|---|---|
| Operational Health | Infrastructure Stability | High 5xx Rate |
| Data Drift | Input Distribution Change | PSI > 0.25 |
| Model Quality (Proxy) | Prediction Shift | Sudden Confidence Drop |
Common Pitfalls
- Mistake: Alerting on average latency. Avoid: Use P95 or P99 latency to capture tail performance.
- Mistake: Setting PSI thresholds too low (e.g., 0.05). Avoid: Use 0.1 (warning) and 0.25 (critical) standards.
- Mistake: Alerting immediately after deployment. Avoid: Implement a 30-minute burn-in window for stabilization.
- Mistake: Using PSI on high-cardinality features. Avoid: Group rare categories or use specialized divergence metrics.
FAQs
- Why is PSI preferred over simpler metrics like mean shift? PSI accounts for changes across the entire distribution shape, not just the central tendency, making it robust to subtle shifts.
- What is the difference between data drift and concept drift? Data drift is when the input features (X) change; concept drift is when the relationship between X and the target (Y) changes.
- How do burn-in windows relate to CI/CD gates? CI/CD gates ensure pre-deployment quality; burn-in windows manage post-deployment instability before full alerting is active.
- If PSI is high, does the model need retraining? Not always immediately; high PSI means the input data changed, which might degrade performance, requiring investigation first.