Back to Evaluation, Leakage, and Unsupervised

Data Leakage — The #1 ML Bug

The reason your model is too good to be true is almost always leakage. Spot it before deployment. FIND_VIDEO: search 'data leakage machine learning examples' — recommended channel: Kaggle / StatQuest. Aim for 11 min or under.

9 minutesVideo LessonPDF notes
🎯 Free Guest Mode: You are learning for free. Sign in to save your completion progress and quiz answers.

Ready to continue?

Mark this lesson as complete when you're ready to proceed.

Key moments

  1. Defining Data Leakage — Data leakage occurs when the model uses information during training that is unavailable in production.
  2. Feature Leakage (TFIDF) — Calculating TFIDF across the entire corpus before splitting allows test set vocabulary characteristics to bleed into training features.
  3. Scaling Leakage — Computing normalization parameters globally allows the test set's distribution to influence the scaling applied to the training set.
  4. The Golden Rule — The primary mitigation strategy is to ensure all transformations are calculated only on the training set after the initial split.
  5. CV and Duplicates — Leakage awareness must be applied to cross-validation and data hygiene to prevent redundant rows from crossing data boundaries.
  6. Non-IID Structural Leakage — Time leakage and Group leakage require structural splitting methods instead of simple random splitting.
  7. Impact Summary — Data leakage drastically overestimates a model's real-world utility, leading to failure upon deployment.
PDF notes

Frequently asked questions

Why is random splitting insufficient for medical imaging data?

Medical data often involves Group Leakage; all images from a single patient must be kept together in either the train or test set to prevent the model from memorizing subject-specific noise.

If I use cross-validation, do I still need to worry about leakage?

Yes. The transformer must be refit or cloned and fitted separately within each CV fold to maintain the isolation of the validation set from the training data for that specific fold.

What is 'Premature Featurization'?

It is calculating summary statistics or features (like mean, variance, or TFIDF) on the entire dataset before the train/test split, thereby leaking test set characteristics into the training features.

How was this lesson?

Your feedback helps us refine explanations and catch bugs.