Models decay; retraining restores them. The question is when. Three strategies:
Strategy 1: Scheduled (cadence-based)
Retrain on a fixed schedule, regardless of conditions.
# Airflow DAG / GitHub Actions cron
schedule: '0 3 1 * *' # 1st of every month at 3am
Pros:
- Predictable.
- Easy to plan around.
- Simple ops: run, deploy, monitor.
Cons:
- May retrain when nothing's broken (waste).
- May fail to retrain when drift is severe.
Use when:
- Drift is gradual and predictable.
- Most production ML systems start here.
Common cadences:
- Daily: high-velocity domains (ads, fraud at peak).
- Weekly: most B2C ML.
- Monthly: stable B2B or analytical models.
- Quarterly: very stable, like LTV forecasting on a SaaS business.
Strategy 2: Triggered (event-based)
Retrain when monitoring detects degradation.
def check_and_retrain():
if proxies_show_drift() or recent_auc_dropped():
trigger_retraining_pipeline()
# Check daily
Pros:
- Retrains only when needed.
- Responsive to actual drift.
Cons:
- Triggers can fire from false alarms (data bugs, transient blips).
- Less predictable for operations.
- Requires reliable detection.
Use when:
- Drift is variable; sometimes weeks of stability, sometimes sudden change.
- Monitoring infrastructure is mature.
Hybrid is common: scheduled monthly + triggered on serious alerts.
Strategy 3: Continual / online learning
The model updates with each new labeled example (or small mini-batch).
# Pseudo-code for online learning
for new_example in stream:
prediction = model.predict(new_example.features)
if new_example.label_available:
model.partial_fit(new_example.features, new_example.label)
Pros:
- Always up-to-date.
- No "retraining cliff".
Cons:
- Hard to evaluate (no fixed training set).
- Subject to runaway drift if label noise is high.
- Complex monitoring (the model is different every minute).
- Not all libraries support online updates.
Use when:
- Labels arrive quickly.
- Domain changes faster than batch retraining can keep up.
- Examples: ad bidding, fraud at adversarial scale, recommendation systems.
For most teams, this is overkill. Stick with scheduled / triggered.
What "retraining" actually involves
A retraining pipeline is a workflow, not a script:
1. Fetch fresh training data (with new labels).
2. Preprocess (using the same feature engineering as production).
3. Split into train/val/test (time-based).
4. Train the candidate model.
5. Evaluate against the current production model.
6. If candidate is better: register as new version, promote through stages.
7. If candidate is worse: log, alert, don't replace.
8. Update documentation.
Each step needs to be automated. Manual retraining drift between code versions; automated pipelines don't.
Comparing candidate vs current
A new training doesn't automatically replace the old. Compare:
- Offline metrics: AUC, PR-AUC, calibration, etc.
- Cohort metrics: does the new model perform similarly on each customer segment?
- Edge cases: does it handle the same edge cases the old one did?
- Calibration: are predicted probabilities consistent with actuals?
candidate_auc = roc_auc_score(y_test, candidate.predict_proba(X_test)[:, 1])
incumbent_auc = roc_auc_score(y_test, incumbent.predict_proba(X_test)[:, 1])
if candidate_auc > incumbent_auc + 0.01: # meaningful improvement
register_for_canary(candidate)
elif candidate_auc > incumbent_auc - 0.005: # not meaningfully worse
register_as_backup(candidate)
else:
log_failure_and_alert(candidate_auc, incumbent_auc)
Many retrainings should NOT replace the current model. The current was already vetted; the new one isn't until tested.
Data freshness
How recent should training data be?
For high-drift domains (fraud, ads):
- Use the last 90-180 days.
- Drop data older than the regime that no longer applies.
For stable domains (LTV, churn in mature businesses):
- 2-3 years of data fine.
Watch for "stale" data:
- If the product launched a new feature in 2024 that changes user behavior, training data from 2022 may not apply.
- If pricing changed, old LTV behavior is irrelevant.
Maintain a "training data window" config that's reviewed quarterly.
Cold-start: retraining when patterns change
Sometimes the data shift is so big that incremental retraining isn't enough. Cold-start scenarios:
- New product launch (no historical data for the new product).
- Major UX change (old behavioral patterns invalid).
- Geographic expansion (no India data for the India launch).
Approaches:
- Start with a hand-crafted rule-based system; collect data; train as enough accumulates.
- Use transfer learning from a related model.
- Use a simpler model initially (logistic with few features); upgrade as data grows.
Documenting retraining
For each retraining cycle:
- Data range used.
- Feature pipeline version.
- Validation metrics (vs baseline).
- Decision (deploy / skip / investigate).
- Approval / sign-off (in regulated industries).
This documentation answers "why is this version live?" months later.
Common retraining mistakes
- Manual retraining only — slow, error-prone, drifts from automated standards.
- Auto-deploy every retrain — no validation step; bad models reach production.
- No comparison to incumbent — retrains "win" by default even when worse.
- Stale training data window — model trained on patterns no longer relevant.
- Retraining too often — wastes compute, introduces churn.
- Retraining too rarely — drift accumulates.
Takeaway
Pick a retraining strategy by drift rate: scheduled for predictable, triggered for variable, continual for fast-moving. Automate the pipeline. Always compare candidate to incumbent; don't auto-replace. Maintain a sensible training-data window. Document every retraining cycle.