Back to M4 — Monitoring/Drift + K8s Reading

Monitoring / Drift (Latency, Errors, PSI-Style)

Outcome: Build drift dashboard sketch + alert rule Curated video (1littlecoder): Machine Learning Model Drift - Concept Drift & Data Drift in ML - Explanation — https://www.youtube.com/watch?v=QJTRNxUxmuc (verified live via yt-dlp 2026-09-24). Pointer: dashboard skeleton (starter); shell: courses/video-scripts/mlops-cloud-deploy/04.md.

10 minutesVideo LessonPDF notes
🎯 Free Guest Mode: You are learning for free. Sign in to save your completion progress and quiz answers.

Ready to continue?

Mark this lesson as complete when you're ready to proceed.

Key moments

  1. Operational Monitoring — Operational monitoring tracks infrastructure performance metrics like latency and errors.
  2. Latency Percentiles — P95 and P99 latency are crucial indicators of poor user experience and tail performance issues.
  3. Introduction to Drift — Drift monitoring detects changes in data characteristics or model behavior over time.
  4. Calculating PSI — Population Stability Index quantifies distribution change by comparing feature bins between reference and current data.
  5. PSI Thresholds — Industry standard thresholds (0.1 warning, 0.25 critical) guide automated response actions.
  6. Alerting Mechanics — Alerts require sustained metric breaches over a time window to filter out transient noise.
  7. Burn-in Windows — Burn-in windows suppress alerts immediately following deployment to allow system stabilization.
PDF notes

Frequently asked questions

Why is PSI preferred over simpler metrics like mean shift?

PSI accounts for changes across the entire distribution shape, not just the central tendency, making it robust to subtle shifts.

What is the difference between data drift and concept drift?

Data drift is when the input features (X) change; concept drift is when the relationship between X and the target (Y) changes.

How do burn-in windows relate to CI/CD gates?

CI/CD gates ensure pre-deployment quality; burn-in windows manage post-deployment instability before full alerting is active.

If PSI is high, does the model need retraining?

Not always immediately; high PSI means the input data changed, which might degrade performance, requiring investigation first.

How was this lesson?

Your feedback helps us refine explanations and catch bugs.