Back to Supervised Learning

Decision Trees and Random Forests

The model that dominates tabular ML. Trees, then forests, then why bagging works. FIND_VIDEO: search 'decision tree random forest tutorial' — recommended channel: StatQuest. Aim for 11 min or under.

10 minutesVideo LessonPDF notes
🎯 Free Guest Mode: You are learning for free. Sign in to save your completion progress and quiz answers.

Ready to continue?

Mark this lesson as complete when you're ready to proceed.

Key moments

  1. DT High Variance — Decision Trees suffer from high variance and overfitting, necessitating the use of ensemble methods.
  2. Bootstrap Aggregating Data — Bootstrap samples are created by resampling the original training data with replacement to train individual trees.
  3. Feature Subsetting — Only a random subset of variables is considered at each split point to ensure the resulting trees are decorrelated.
  4. RF Assembly and Voting — The Random Forest aggregates predictions from all trees using majority voting to reach a final ensemble classification.
  5. Defining Bagging — Bagging is the formal term describing the combined technique of bootstrapping the data and aggregating the decisions.
  6. OOB Error Validation — Data samples not included in a tree's bootstrap sample are used as an internal validation set to estimate generalization error.
  7. Hyperparameter Tuning — The accuracy of the model is optimized by tuning the number of variables used at each split to minimize the OOB error.
PDF notes

Frequently asked questions

Why do we need two types of randomness?

Bootstrapping diversifies the data, while feature subsetting ensures the resulting trees are decorrelated, which is necessary for variance reduction.

What happens if I don't use feature subsetting?

The trees will be highly correlated because they will all split on the strongest predictor, negating the variance reduction benefits of the ensemble.

How do Random Forests make a final prediction?

For classification, they use majority voting across all individual tree predictions; for regression, they average the numerical outputs.

How should I choose the optimal feature subset size?

Start with the standard $\sqrt{P}$ (classification) or $P/3$ (regression) and tune around that value to minimize the OOB error rate.

How was this lesson?

Your feedback helps us refine explanations and catch bugs.