This lesson on Linear Regression — The Math, Not Just the API is hands-on and example-driven. You will be able to mathematically derive and calculate the core components of a linear regression model, moving beyond API calls to understand the underlying statistical mechanics. You will master the Least Squares Method for fitting the line and quantify model performance using the R-squared metric based on sums of squares.
What You'll Be Able To Do
- Calculate the residual error for any given data point relative to a fitted line.
- Apply the Least Squares Method to determine the optimal slope and intercept parameters.
- Define and calculate the Sum of Squared Residuals (SS Fit) and the total variation (SS Mean).
- Derive the R-squared value using the ratio of explained variation to total variation.
- Generalize the concept of Sums of Squares to evaluate complex non-linear models.
Detailed Concept Walkthrough
1. Least Squares Fitting and Residuals
The Least Squares Method (LSM) finds the line that minimizes the total vertical distance between the line and all data points. This distance, called the residual, represents the error of the model's prediction.
- Mechanism: LSM iteratively adjusts the slope and intercept until the Sum of Squared Residuals (SSR or SS Fit) reaches its absolute minimum. Squaring the residuals ensures positive errors do not cancel out negative errors.
- Under the Hood: The optimal parameters (slope $m$ and intercept $b$) are found by taking the partial derivatives of the SSR function with respect to $m$ and $b$, setting them to zero, and solving the resulting system of equations.
- Best Practice: Always visualize the residuals (e.g., in a scatter plot of residuals vs. predicted values) to check for non-linear patterns, which indicate the linear model is inappropriate.
import numpy as np
# Assume X and Y are data arrays, m and b are parameters
def calculate_ssr(X, Y, m, b):
# Calculate predicted Y values
Y_pred = m * X + b
# Calculate residuals (actual - predicted)
residuals = Y - Y_pred
# Return the Sum of Squared Residuals
return (residuals ** 2).sum()
Key Takeaway: The "best fit" line is mathematically defined as the one that yields the lowest possible Sum of Squared Residuals.
2. Sums of Squares: Baseline and Residual
Regression performance is measured by comparing the total variation in the dependent variable (Y) against the variation remaining after the model is applied. This comparison uses Sums of Squares (SS).
- Mechanism: SS Mean (Total Sum of Squares) establishes the baseline error by measuring the squared distance of every Y value from the overall mean of Y ($\bar{Y}$). This represents the variation if the predictor X were ignored.
- Execution Flow: SS Fit (Residual Sum of Squares) measures the squared distance of every Y value from the fitted line ($Y_{pred}$). This quantifies the variation that the model failed to explain.
- Best Practice: The difference ($ ext{SS Mean} - ext{SS Fit}$) represents the variation explained by the model, which forms the numerator in the R-squared calculation.
import numpy as np
# Assume Y is the dependent variable array and Y_pred is the fitted line
Y_mean = np.mean(Y)
# SS Mean: Variation around the mean (baseline)
SS_Mean = np.sum((Y - Y_mean)**2)
# SS Fit: Variation around the fitted line (residual error)
SS_Fit = np.sum((Y - Y_pred)**2)
Key Takeaway: SS Mean is the total variation to be explained, and SS Fit is the unexplained variation remaining after fitting the line.
3. R-Squared Calculation and Interpretation
R-squared ($R^2$) quantifies the proportion of the total variance in the dependent variable (Y) that is predictable from the independent variable(s) (X). It is the primary measure of model goodness-of-fit.
- Mechanism: $R^2$ is calculated as $1 - (\text{SS Fit} / \text{SS Mean})$, which expresses the explained variation as a percentage of the total variation.
- Boundary Cases: If the fit is perfect, $\text{SS Fit} = 0$, resulting in $R^2 = 1$ (100% explained); if the fit is no better than the mean, $R^2 \approx 0$.
- Nuance: While R-squared is often defined using variance, the raw Sums of Squares (SS) can be used directly because the sample size ($N$) cancels out in the numerator and denominator of the ratio.
# Assuming SS_Mean and SS_Fit have been calculated
SS_Explained = SS_Mean - SS_Fit
R_squared = SS_Explained / SS_Mean
print(f"R-squared: {R_squared:.2f}")
# Example output: R-squared: 0.60 (60% of variance explained)
Key Takeaway: R-squared measures the reduction in error achieved by using the regression line instead of simply using the mean of Y.
4. Multivariate Regression Generalization
The core principles of minimizing squared residuals and calculating R-squared extend directly to models involving multiple predictors (Multiple Linear Regression).
- Mechanism: Instead of fitting a 2D line ($Y = mX + b$), multivariate regression fits a hyperplane defined by $Y = m_1X_1 + m_2X_2 + \dots + b$.
- Visualization: With two predictors ($X_1, X_2$), the model is visualized as a plane in a 3D space, where the residuals are the vertical distances from the data points to this fitted plane.
- Best Practice: The R-squared calculation remains mathematically identical regardless of the number of predictors or the complexity of the model equation, provided the model is fitted by minimizing the Sum of Squared Residuals.
Key Takeaway: The mathematical foundation (Least Squares and SS calculation) is universal for fitting and evaluating any model that minimizes squared error.
Topics Covered in Linear Regression — The Math, Not Just the API
- Core Regression Steps (0:20 - 0:58) — Regression analysis involves fitting a line, quantifying performance, and testing statistical significance.
- Least Squares Method (1:20 - 2:36) — The best-fit line minimizes the Sum of Squared Residuals (SSR), which are the errors between data and the line.
- Model Parameters (2:37 - 3:00) — The fitted line is defined by the estimated slope and Y-intercept, which determine its predictive power.
- Baseline Variation (3:15 - 4:14) — SS Mean establishes the total variation in Y that must be explained, calculated around the average Y value.
- Residual Variation (4:15 - 4:53) — SS Fit quantifies the unexplained variation remaining after the least squares line has been fitted to the data.
- R-Squared Formula (4:54 - 6:32) — R-squared is calculated as the ratio of explained variation to total variation, showing goodness-of-fit.
- Boundary Cases (6:33 - 8:18) — R-squared ranges from 0% (no fit) to 100% (perfect fit) and applies universally across model types.
- Multiple Predictors (8:19 - 9:13) — Linear regression extends to multiple variables by fitting a plane or hyperplane in higher dimensions.
Statistics for Data Science Cheat Sheet
-
Residual— Distance from data point to the fitted lineresidual = Y_actual - Y_predicted -
Least Squares Method— Minimizes the total Sum of Squared Residuals (SSR)min(sum((Y - (m*X + b))**2)) -
SS Mean— Total variation in Y around the mean of Ynp.sum((Y - Y.mean())**2) -
SS Fit— Unexplained variation remaining around the fitted linenp.sum((Y - Y_pred)**2) -
R-squared— Proportion of Y variance explained by the modelR2 = 1 - (SS_Fit / SS_Mean) -
Multivariate Regression— Linear model using two or more independent variablesY = m1*X1 + m2*X2 + b
Comparison Table
| SS Metric | Definition | Purpose in R-squared |
|---|---|---|
| SS Mean | Total squared distance from Y to $\bar{Y}$. | Denominator (Total Variation) |
| SS Fit | Total squared distance from Y to $Y_{pred}$. | Numerator (Unexplained Variation) |
| Variance | Average sum of squares ($ ext{SS}/N$). | Used conceptually, cancels out in ratio. |
Common Pitfalls
- Mistake: Confusing residual (error) with the model parameters (slope/intercept). Avoid: Residuals are vertical distances; parameters define the line's position.
- Mistake: Assuming a high R-squared guarantees a good predictive model. Avoid: R-squared only measures fit; check residual plots for linearity assumptions.
- Mistake: Calculating R-squared using raw variance values instead of SS values. Avoid: Use SS values directly since the sample size $N$ cancels out in the ratio.
- Mistake: Forgetting to square the residuals before summing them. Avoid: Squaring prevents positive and negative errors from erroneously canceling out.
FAQs
- Why do we square the residuals instead of just summing the absolute values? Squaring penalizes larger errors more heavily, ensuring the line is closer to all points. It also makes the error function differentiable, which is necessary for optimization algorithms.
- Does R-squared work for non-linear models? Yes, the R-squared calculation based on SS Mean and SS Fit is universal for quantifying the goodness-of-fit of any model equation, provided it minimizes squared error.
- What does a non-zero slope imply? A non-zero slope suggests that the independent variable (X) has predictive power over the dependent variable (Y), meaning the fitted line is better than just using the mean.