Replacing a production model is risky. The new model might be:
- Worse in production despite winning offline.
- Wrong in some specific case the test set didn't cover.
- Slower than the old one.
- Less calibrated.
Don't just swap. Roll out in stages.
Stage 1: Shadow mode
The new model (challenger) runs alongside the existing one (champion). The challenger's predictions are computed but NOT used for decisions. They're logged for comparison.
@app.post('/predict')
def predict(req: TransactionRequest):
features = build_features(req.dict())
# Champion (in use)
champion_pred = champion_model.predict_proba([features])[0][1]
# Challenger (shadow)
challenger_pred = challenger_model.predict_proba([features])[0][1]
# Log both for offline comparison
log_prediction({
'transaction_id': req.transaction_id,
'champion_pred': champion_pred,
'challenger_pred': challenger_pred,
'champion_version': CHAMPION_VERSION,
'challenger_version': CHALLENGER_VERSION,
})
# Decision uses champion only
return {'fraud_probability': float(champion_pred)}
What you check in shadow:
- Are challenger predictions in a sensible range?
- Are they correlated with champion predictions where they should be?
- Where do they diverge? Are the divergences justified by features?
- Does the challenger handle edge cases (high amounts, rare merchants) gracefully?
Run shadow for ~1 week. Catches obvious bugs before any user is affected.
Stage 2: Canary (small-percentage rollout)
Once shadow looks good, route a small fraction of traffic to the challenger:
@app.post('/predict')
def predict(req: TransactionRequest):
features = build_features(req.dict())
# Decide which model based on hash of request ID
use_challenger = (hash(req.transaction_id) % 100) < CANARY_PCT # e.g., 5%
if use_challenger:
prediction = challenger_model.predict_proba([features])[0][1]
version = CHALLENGER_VERSION
else:
prediction = champion_model.predict_proba([features])[0][1]
version = CHAMPION_VERSION
log_prediction({
'transaction_id': req.transaction_id,
'prediction': prediction,
'model_version': version,
'is_canary': use_challenger,
})
return {'fraud_probability': float(prediction)}
CANARY_PCT starts at 1-5%. Monitor:
- Business metrics for canary cohort vs champion cohort.
- Latency.
- Error rate.
If anything's worse, halt the rollout. Investigate before increasing.
Stage 3: Gradual ramp
Increase canary_pct over days:
- Day 1: 5%
- Day 3: 10%
- Day 5: 25%
- Day 7: 50%
- Day 10: 100%
At each step, hold for at least a day to see business impact (which often lags behind prediction by hours-days). If anything regresses, roll back.
What to monitor during rollout
Three categories:
Prediction-level
- Distribution of predictions (should match historical baseline).
- Confidence calibration if applicable.
Action-level
- Decision rate (how often is the model flagging fraud?).
- Action mix (auto-block vs review vs ignore).
Business-level
- Real outcomes (when labeled): precision, recall.
- KPIs (revenue lost to fraud, false-positive customer complaints).
The business-level metrics often lag — fraud that happened today shows up as a chargeback in 30 days. You can't validate the new model's business impact in real-time. Monitor proxies (decision rate, distribution stability) and expect a slower confirmation cycle.
What can go wrong
1. Champion-challenger correlation
If both models share the same flawed feature, they both predict wrong together. Comparing them looks fine, but both are bad. Detect this with absolute monitoring (not just relative).
2. Trafficking the wrong way
Hash-based assignment by transaction_id keeps the same transaction on the same model — needed for retries to be consistent. Random per-request would create flickering predictions.
3. Confounded comparisons
The challenger gets seen primarily by Tuesday's traffic; champion by Monday's. Day-of-week effects bias the comparison. Mitigate by running long enough.
4. Calibration drift between models
Champion outputs are calibrated at threshold 0.5; challenger outputs at 0.3. You can't directly compare. Either align thresholds or compare AUC / ranking quality.
Holdout (control) groups for unbiased comparison
A common refinement: even at 100% rollout, keep a 5% holdout on the old model permanently. Lets you measure the new model's incremental value over time.
@app.post('/predict')
def predict(req: TransactionRequest):
# Permanent holdout: 5% always on champion (old)
is_holdout = (hash(req.user_id) % 100) < 5
use_new = not is_holdout
if use_new:
prediction = new_model.predict_proba(...)
else:
prediction = old_model.predict_proba(...)
After a month, compare metrics for new vs holdout cohort. If new model is genuinely better, the holdout will show worse business outcomes. If new isn't better, kill it.
This is the analog of a "control group" in clinical trials. Cost: 5% suboptimal performance on the holdout (worth it for the inference value).
Rollback procedure
# Update CHAMPION_VERSION and CHALLENGER_VERSION in config
CHAMPION_VERSION = 'v1.2.0' # was v1.3.0, rolled back
CHALLENGER_VERSION = None # nothing in challenger slot
# Restart service or hot-reload config
Critical that this is a config change, not a code deploy. Code deploys have their own risk (compilation, dependency changes); rollback should be near-instantaneous.
Document the rollout
For each model deployment:
- Champion version, challenger version.
- Shadow start date, canary start date, canary percentages over time.
- Metrics observed at each stage.
- Decision to promote / roll back.
This documentation is what you need when the new model degrades a month later and you're asking "what changed?"
Common rollout mistakes
- Direct swap (no shadow / canary) — bug in new model affects all traffic at once.
- No business metric monitoring — predictions look fine, business outcomes degrade silently.
- Rapid ramp without holds — bug shows up at 50%, by then thousands of users affected.
- No permanent holdout — can't compare new vs old over time.
- No documented rollback procedure — emergency code changes during outages.
Takeaway
Three stages: shadow (no user impact) → canary (small %, monitored) → full. Hash-based assignment for consistency. Monitor prediction-level, action-level, and business-level. Keep a permanent holdout for ongoing comparison. Rollback is a config change, not a deploy. Skip stages at your peril.
Deep Dive: Shadow Traffic Architecture (Envoy / NGINX Duplication vs. Async Shadowing)
Running a candidate model in shadow mode can be implemented at two architectural layers:
1. Ingress Proxy Layer (Envoy / NGINX Shadowing)
- The API gateway or ingress controller asynchronously duplicates incoming HTTP request payloads and forwards a mirror copy to the shadow service endpoint (
mirror_percent: 100). - The response from the shadow service is completely ignored, ensuring zero latency impact on the primary client.
- Advantage: Completely decouples primary service code from candidate service code.
2. Application Layer (Async BackgroundTasks)
- The primary service processes the request using the champion model, returns the response to the user, and enqueues an asynchronous coroutine via
BackgroundTasksto send the payload to the challenger model. - Advantage: Simpler to implement on basic cloud platforms (e.g. Cloud Run, AWS ECS) without configuring complex service meshes.