This lesson on Logistic Regression and the GLM Family is hands-on and example-driven. You will learn the fundamental differences between linear and logistic regression, enabling you to select the correct model for continuous prediction versus binary classification tasks. You will gain conceptual comfort with Maximum Likelihood Estimation (MLE) as the core technique for fitting classification models and assessing feature significance.
What You'll Be Able To Do
- Distinguish suitable applications for linear regression (continuous) versus logistic regression (binary).
- Apply the Sigmoid function concept to transform linear output into a probability estimate (0-1).
- Select and justify the use of Maximum Likelihood Estimation (MLE) for parameter fitting.
- Assess feature importance in a logistic model using statistical tests like Wald's test.
- Construct a multi-variable logistic model using mixed data types (continuous and discrete).
Detailed Concept Walkthrough
1. Linear vs. Logistic Regression
Linear Regression predicts continuous values (e.g., price) using a straight line, while Logistic Regression predicts the probability of a binary outcome (0 or 1) using an S-shaped curve. The choice depends entirely on the nature of the target variable.
- Mechanism: Linear models use the identity link function, mapping predictors directly to the outcome. Logistic models use the logit link function, mapping predictors to the log-odds of the outcome, which is then transformed into a probability.
- Best Practice: Always confirm the target variable type before modeling: continuous targets require Linear Regression (or variants), while categorical/binary targets require Logistic Regression (or classifiers).
- Under the Hood: Linear Regression minimizes the Sum of Squared Residuals (Least Squares). Logistic Regression maximizes the likelihood of observing the data points given the model parameters (Maximum Likelihood Estimation).
import statsmodels.api as sm
# Linear Regression (Continuous Target)
# model_lin = sm.OLS(y_continuous, X).fit()
# Logistic Regression (Binary Target: 0 or 1)
model_log = sm.Logit(y_binary, X).fit()
# Note the distinct fitting methods required by the target type.
Key Takeaway: Logistic Regression is a classification algorithm that estimates probability, not a regression algorithm that predicts a continuous value.
2. Sigmoid Function and Classification
The Sigmoid (or Logistic) function takes any real number input (the linear combination of predictors) and squashes it into an output between 0 and 1, representing the probability P(Y=1|X). This S-curve is essential for modeling probabilities.
- Mechanism: The function $f(z) = 1 / (1 + e^{-z})$ transforms the linear predictor $z$ into a probability. As $z$ approaches positive infinity, the output approaches 1; as $z$ approaches negative infinity, the output approaches 0.
- Execution Flow: After calculating the probability $P$, a threshold (typically 0.5) is applied to convert the continuous probability estimate into a hard binary classification (e.g., if $P > 0.5$, predict 'True').
- Best Practice: Logistic models are highly flexible regarding input data types; they can simultaneously handle continuous features (age, weight) and discrete/categorical features (genotype, gender) as predictors.
import numpy as np
def sigmoid(z):
"""Calculates the Sigmoid function output."""
return 1 / (1 + np.exp(-z))
z_val = 2.0 # Linear output (log-odds)
probability = sigmoid(z_val)
# probability will be ~0.88. If threshold is 0.5, classify as 1.
Key Takeaway: The Sigmoid function converts the model's linear output (log-odds) into a meaningful probability between 0 and 1.
3. Maximum Likelihood Estimation (MLE)
Since Logistic Regression does not rely on minimizing squared errors (residuals), it uses MLE to find the set of parameters (coefficients) that maximize the joint probability (likelihood) of observing the actual outcomes in the training data.
- Under the Hood: MLE iteratively adjusts the model parameters. For each data point, the model calculates the probability of the observed outcome, and the likelihood function is the product of these probabilities across all data points.
- Mechanism: The goal is to find the curve (defined by the parameters) that makes the observed data most probable. This is fundamentally different from Least Squares, which seeks to minimize the distance between the predicted line and the data points.
- Best Practice: Conceptual comfort with MLE is crucial because it is the standard fitting technique for Generalized Linear Models (GLMs) and many other classification algorithms where minimizing residuals is inappropriate.
Key Takeaway: MLE fits the S-curve by maximizing the likelihood of the observed data, replacing the Least Squares minimization used in linear models.
4. Feature Significance and Assessment
Standard metrics like $R^2$ and residual analysis are invalid for logistic regression because the target is binary, not continuous. Feature importance must be assessed using statistical tests that evaluate if a coefficient is significantly different from zero.
- Mechanism: Wald’s test is commonly used to test the null hypothesis that a specific predictor's coefficient ($eta_i$) is zero. If the test rejects the null hypothesis, the feature is considered statistically significant and useful for prediction.
- Limitation: Comparing the overall fit of complicated logistic models to simpler ones is challenging. You cannot rely on a single generalized measure like $R^2$ to determine which model is globally better.
- Best Practice: Feature selection in logistic regression relies on these significance tests (like Wald's test) to prune non-contributing variables, ensuring the final model is parsimonious and predictive.
import statsmodels.api as sm
# Fit the model
model = sm.Logit(y_binary, X).fit()
# Review the summary table to find P>|z| column
# This P-value is derived from Wald's test for each coefficient.
print(model.summary())
Key Takeaway: Use significance tests (e.g., Wald's test) for feature selection, as standard residual analysis and $R^2$ are not applicable to logistic models.
Topics Covered in Logistic Regression and the GLM Family
- Linear Regression Review (0:18 - 2:27) — Linear regression models continuous outcomes by fitting a straight line and uses R^2 for performance assessment.
- Defining Logistic Regression (2:25 - 3:20) — Logistic regression predicts binary outcomes using the S-shaped Sigmoid function to output probabilities between 0 and 1.
- Classification and Flexibility (3:22 - 4:06) — A probability threshold converts the continuous probability estimate into a hard binary classification, and the model accepts mixed data types.
- Feature Significance Tests (4:08 - 5:15) — Feature importance is determined using statistical tests like Wald’s test because standard residual analysis and R^2 are not applicable.
- Maximum Likelihood Fitting (5:17 - 6:06) — Model parameters are estimated by iteratively maximizing the joint probability (likelihood) of observing the given data points.
Statistics for Data Science Cheat Sheet
-
Logistic Regression— Predicts probability for binary outcomes (0 or 1)model = sm.Logit(y_binary, X).fit() -
Sigmoid Function— Maps linear output to probability (0 to 1)1 / (1 + np.exp(-z)) -
Maximum Likelihood Estimation— Finds parameters maximizing data observation probability -
Wald’s Test— Assesses if a predictor's coefficient is significantprint(model.summary()) # Look for P>|z| -
Logit Link Function— Links linear predictor to the log-odds of the outcomelog(P / (1-P)) = B0 + B1*X -
Classification Threshold— Converts probability into a hard binary predictionif P > 0.5: return 1
Comparison Table
| Linear Regression | Logistic Regression | Model Assessment |
|---|---|---|
| Continuous (e.g., size, price) | Binary (0 or 1, True/False) | Target Variable |
| Least Squares (Minimize SSE) | Maximum Likelihood Estimation (MLE) | Fitting Method |
| $R^2$, Residual Plots | Wald’s Test, Feature P-values | Performance Metric |
Common Pitfalls
- Mistake: Trying to use R^2 or standard residuals to evaluate logistic model fit. Avoid: Rely on feature significance tests (Wald's) and overall likelihood metrics.
- Mistake: Interpreting the raw output of the logistic model as a continuous value. Avoid: Remember the output is a probability (0-1) that requires thresholding for classification.
- Mistake: Assuming Least Squares can be used to fit the S-curve parameters. Avoid: Always use Maximum Likelihood Estimation (MLE) for fitting logistic and GLM models.
- Mistake: Only using continuous predictors in a logistic model. Avoid: Logistic models are flexible and handle continuous and discrete predictors simultaneously.
FAQs
- Why can't I use R^2 to compare logistic models? R^2 relies on minimizing squared residuals, which is invalid when the outcome is binary and the model estimates probability.
- What does the Sigmoid function output represent? It represents the estimated probability that the observation belongs to the positive class (Y=1), ranging from 0 to 1.
- How do I know if a predictor is important in a logistic model? You use statistical tests like Wald's test to check if the predictor's coefficient is significantly different from zero.
- Is Logistic Regression a type of Linear Regression? No. While it uses a linear combination of predictors, it is fundamentally a classification algorithm using a non-linear link function (Sigmoid).