Back to ML Foundations: The Mental Model

Train/Validation/Test — The Three-Split Discipline

The single piece of ML discipline that separates honest models from theatre. FIND_VIDEO: search 'train validation test split machine learning' — recommended channel: StatQuest / Andrew Ng. Aim for 10 min or under.

4 minutesVideo LessonPDF notes
🎯 Free Guest Mode: You are learning for free. Sign in to save your completion progress and quiz answers.

Ready to continue?

Mark this lesson as complete when you're ready to proceed.

Key moments

  1. Problem Context — Machine learning requires labeled historical data (features and target) to perform predictive tasks.
  2. Data Split Rule — The 80/20 train/test split discipline is established to ensure unbiased model evaluation.
  3. Model Training — The model learns patterns exclusively on the larger Training Set, risking training set specificity.
  4. Evaluating Generalization — The Test Set objectively measures the model's ability to generalize to unseen data before approval.
  5. Overfitting Check — Comparing training performance versus test performance reveals if the model learned general rules or merely memorized the data.
PDF notes

Frequently asked questions

Why is the 80/20 split standard?

It provides enough data (80%) for the model to learn complex patterns while reserving a sufficiently large sample (20%) for reliable evaluation.

What if my model performs perfectly on the Training Set?

Perfect training performance is often a sign of overfitting, meaning the model has memorized noise rather than generalized rules.

Does the Test Set ever influence the model training?

No, the Test Set must remain completely separate and unseen; using it during training invalidates the generalization assessment.

How was this lesson?

Your feedback helps us refine explanations and catch bugs.