A model that wins on Kaggle dies in production. Not because the model is wrong — because the gaps between the notebook and the production system are wider than the model's accuracy is large.
There are three gaps. Each one quietly destroys ML projects.
Gap 1: Training-serving skew
The model trains on data shape A. At serving time, it gets data shape B. Same code, different inputs, silently wrong outputs.
Common causes:
- Different feature pipelines (
build_features()in notebook vsbuild_features_v2()in the API). - Time-zone mismatches.
- Missing-value handling that differs.
- Categorical encoders fit on training that don't see new categories.
- Different data sources (training pulled from data warehouse; serving pulls from API).
Symptom: model accuracy collapses in production despite passing all offline tests.
Fix: one source of truth for feature computation. A feature store, or at minimum a shared library imported by both training and serving.
Gap 2: Environment drift
The model was trained with scikit-learn==1.2.0 and xgboost==1.7.0. The serving environment runs scikit-learn==1.4.0 and xgboost==2.0.0. Some predictions differ slightly. Sometimes it crashes (deprecated API).
Causes:
- No pinned dependency versions.
- "It works on my machine."
- Operating system or hardware differences.
Fix: pin every dependency. Use Docker containers. Test in an environment matching production.
Gap 3: No monitoring
The model is deployed. It's "working" — the endpoint returns 200s. Six months later, an analyst notices the dashboard hasn't moved. The model has been silently predicting the mean for weeks because a feature pipeline broke and started returning all zeros.
Fix: monitor predictions, inputs, and business outcomes. Alert on drift. Schedule retraining.
The trap: ML engineers prioritize accuracy
Notebook-to-production is hard. So researchers focus on the easier problem: accuracy gains. They tune from 0.85 to 0.87 AUC, then ship, then the model dies because of one of the three gaps.
The right hierarchy:
- Working at all (no leakage, reproducible).
- Working in production (no training-serving skew, no environment drift).
- Monitored (drift detection, performance tracking).
- Then optimize (squeeze more accuracy).
Most teams invert this. Lots of accuracy work, no production discipline. Predictable result.
What "MLOps" actually is
MLOps = DevOps + ML. The discipline of:
- Reproducible training pipelines.
- Versioning models, data, features.
- Reliable serving.
- Continuous monitoring.
- Safe deployment.
- Incident response.
Lightweight MLOps (this course): the minimum to make ML production-grade. No platform engineering team required.
Heavy MLOps (separate course material): feature stores, model registries, online experimentation platforms, real-time labeling, fairness/bias audits, A/B testing infrastructure. For mature ML organizations.
Tools to learn
In rough order:
- MLflow — experiment tracking + model registry. Lightweight, easy to start.
- DVC — data versioning. Git for datasets.
- FastAPI — building Python web services.
- Docker — containerization.
- GitHub Actions — CI/CD for ML.
- Optional: Feast (feature store), Evidently (drift detection), Streamlit (model dashboards).
This course covers MLflow + DVC + FastAPI + Docker + GitHub Actions + Evidently. Enough to deploy and maintain a model professionally.
What this course is NOT about
- Kubernetes operators and complex platform infrastructure.
- Custom feature store implementation.
- Distributed training (multi-GPU / multi-node).
- Real-time stream processing.
These are advanced topics. For lightweight production ML at most companies, you don't need them.
The 80/20 of MLOps
If you do 5 things well:
- Pin dependencies (Docker / requirements.txt with versions).
- Use the same code for feature computation in training and serving.
- Version your models (MLflow registry or just dated S3 paths).
- Log every prediction.
- Set up basic monitoring (prediction distribution + business metric).
…you've solved 80% of the production-ML problem. The other 20% (auto-retraining, online experimentation, fairness audits) is nice-to-have but not blocking.
Takeaway
The three gaps between notebook and production: training-serving skew, environment drift, no monitoring. Most ML failures live here, not in model accuracy. Lightweight MLOps closes the gaps with simple discipline; this course teaches that discipline.