Production models decay. Sometimes fast (a feature pipeline breaks), sometimes slow (market conditions shift). Two distinct kinds of decay:
Data drift (covariate shift)
The distribution of input features changes. Example: a fraud model trained on India transactions now serves US transactions; feature distributions shift.
The relationship between features and target (P(y|X)) hasn't necessarily changed — but the inputs the model sees are different from what it trained on.
Symptom: feature distribution monitoring shows shifts. Model accuracy may or may not degrade (depends on whether the model generalizes to the new feature region).
Causes:
- Pipeline upstream changed.
- Customer base shifted.
- Seasonality (holiday traffic, ad campaigns).
- New product launches.
Concept drift
The relationship between features and target changes. Example: fraud patterns evolve — new attack methods, criminals adapt to detection.
The features look the same, but the target's response to them is different. P(y|X) has changed.
Symptom: accuracy degrades even though feature distributions look stable.
Causes:
- Adversarial adaptation (fraud, security).
- Behavior changes (post-COVID consumer patterns).
- Policy changes (new pricing, new product tiers).
- Regulatory or market shifts.
Why distinguishing matters
Different problems, different fixes:
- Data drift: retrain on the new distribution (the model is fine, the inputs are new).
- Concept drift: retrain with newer labels (the model's mapping is stale).
- Both: more comprehensive retraining.
If you don't distinguish, you might retrain when nothing's actually broken (wasteful) or fail to retrain when you should (model decays further).
Detecting data drift
Compare current feature distribution to training distribution.
Per-feature comparison
For each feature:
- Kolmogorov-Smirnov test: compares distributions, returns a p-value for "are they different?"
- Population Stability Index (PSI): a divergence measure.
def psi(expected, actual, buckets=10):
expected_pct = pd.cut(expected, buckets).value_counts(normalize=True)
actual_pct = pd.cut(actual, buckets).value_counts(normalize=True)
# Avoid log(0)
expected_pct = expected_pct.where(expected_pct > 0, 0.0001)
actual_pct = actual_pct.where(actual_pct > 0, 0.0001)
return ((actual_pct - expected_pct) * np.log(actual_pct / expected_pct)).sum()
psi_score = psi(training_data['amount'], today_data['amount'])
PSI scale:
- < 0.1: no drift.
- 0.1 - 0.25: moderate drift, investigate.
-
0.25: significant drift, retrain.
Multivariate detection
Joint distributions can shift even if marginals look fine. Use:
- Maximum Mean Discrepancy (MMD) or kernel methods: detect joint distribution differences.
- Adversarial classifier: train a classifier to distinguish "training set" vs "today's data". If AUC > 0.6, there's drift.
adversarial_data = pd.concat([
training_data.assign(label=0),
today_data.assign(label=1)
])
adversarial_model = XGBClassifier()
score = cross_val_score(adversarial_model, X, y).mean()
if score > 0.6:
print('Significant distribution shift detected.')
Detecting concept drift (harder)
Requires labels. Often labels arrive late.
When labels are available
Compute model performance on recent labeled data. Compare to baseline.
recent_predictions = predictions_table.query('date > now() - INTERVAL 7 days')
recent_predictions = recent_predictions.merge(labels_table, on='id')
recent_auc = roc_auc_score(recent_predictions['label'], recent_predictions['prediction'])
baseline_auc = 0.85 # what we expect
if recent_auc < baseline_auc - 0.05:
print('Possible concept drift, AUC dropped from 0.85 to', recent_auc)
When labels are delayed
You can't immediately verify accuracy. Proxy signals:
- Prediction distribution stability: if model predictions look weird, performance likely is too.
- Confidence distribution: model less confident than usual? Concept drift signal.
- Business metric drift: if downstream KPIs change, the model may be a cause.
For systems where labels arrive in weeks/months (churn, LTV), proxies are essential.
What to monitor (the checklist)
Daily / hourly checks:
| Category | Metric | What it signals |
|---|---|---|
| Input | Feature distributions (PSI, mean, stddev) | Data drift |
| Input | Missing value rate per feature | Pipeline failures |
| Input | New categorical values | New patterns or data quality |
| Output | Prediction distribution (mean, percentiles) | Major shift |
| Output | Prediction count, latency | Service health |
| Outcome | Action rate (e.g., % flagged) | Threshold drift |
| Outcome | Realized accuracy on labeled data | Concept drift |
| Business | KPI affected by model | Real-world impact |
Tools: Evidently (open-source, Python-native), WhyLabs, Arize, custom dashboards on top of logged predictions.
Alerting
Two main patterns:
Threshold-based
"PSI > 0.25 on any feature → alert."
Simple. False positives when natural variance crosses thresholds.
Trend-based
"AUC drops more than 0.05 over the rolling 7-day average compared to the prior 30-day average → alert."
Captures sustained drift, ignores one-day noise.
Use both. Threshold for acute issues; trend for slow degradation.
What to do when drift is detected
-
Investigate the cause — is it real drift or a data quality issue?
- Check upstream pipelines (Fivetran sync issues, source table changes).
- Spot-check a few examples.
- Compare drift across features: if everything moved, probably pipeline; if one feature, probably real change.
-
Decide on response:
- Pipeline bug → fix the pipeline.
- Real drift → retrain.
- Marginal drift → continue monitoring.
-
Document the incident — for trend analysis later.
Common drift detection mistakes
- No drift detection at all — model dies invisibly.
- Per-feature alerts only — miss joint distribution shifts.
- Threshold too tight — false alarms.
- Threshold too loose — miss drift until it's bad.
- Treating drift detection as one-time — should be ongoing.
- Reacting to a single day's drift — daily fluctuations are normal; need sustained shift before retraining.
Takeaway
Data drift = inputs changed; concept drift = input-target relationship changed. Detect data drift via PSI/KS/adversarial validation; detect concept drift via labeled performance monitoring or proxies when labels are delayed. Investigate before reacting — pipeline bugs and real drift look similar. Use threshold + trend alerts together.