This lesson on Chi-Square and Categorical Data is hands-on and example-driven. You will learn how to apply the Chi-Square statistical test to evaluate the relationship between categorical features and a classification target. You will be able to preprocess string-based categorical data into the required non-negative numerical format. Finally, you will use Scikit-learn to rank features by their predictive power based on P-values.
What You'll Be Able To Do
- Apply the Chi-Square test to measure feature dependence on a target variable.
- Transform string-based categorical features into non-negative numerical inputs using encoding.
- Execute the
sklearn.feature_selection.chi2function on training data partitions. - Interpret the resulting F-scores and P-values to determine feature significance.
- Rank features using Pandas based on the lowest P-values for selection.
Detailed Concept Walkthrough
1. Chi-Square Test Fundamentals
The Chi-Square test assesses if two categorical variables are statistically independent. In feature selection, it measures the strength of the relationship between a feature and the classification target.
- Mechanism: The test compares the observed frequencies of categories in the data against the frequencies expected if the feature and target were completely independent (the null hypothesis).
- Data Requirement: Chi-Square specifically requires non-negative numerical inputs; therefore, categorical features must be encoded into positive integers representing counts or indices.
- Interpretation Duality: Feature importance is determined by considering both a high Chi-Square F-score (statistic magnitude) AND a low P-value (statistical significance, typically < 0.05).
Key Takeaway: High F-score and low P-value indicate the feature is strongly dependent on the target and should be selected.
2. Encoding for Non-Negative Inputs
String-based categorical features must be converted to positive integers because the Chi-Square calculation relies on non-negative frequency counts. Label Encoding or Ordinal Encoding are suitable methods.
- Execution Flow: For binary features (e.g., 'sex'),
numpy.whereprovides a fast, vectorized way to map strings to 0 and 1, ensuring non-negative values. - Best Practice: For multi-class features (e.g., 'embarked'), use dictionary mapping or Label Encoding to assign unique positive integer indices to each category level.
- Syntax Rule: Ensure the resulting encoded column contains only integers greater than or equal to zero; negative values will invalidate the statistical test.
import numpy as np
import pandas as pd
# Example: Binary encoding using numpy.where
df['sex_encoded'] = np.where(df['sex'] == 'male', 1, 0)
# Example: Multi-class ordinal encoding using dictionary
embarked_map = {'S': 1, 'C': 2, 'Q': 3}
df['embarked_encoded'] = df['embarked'].map(embarked_map)
Key Takeaway: All input features (X) must be encoded as non-negative integers before applying the Chi-Square test.
3. Chi2 Execution and Ranking
Scikit-learn's chi2 function calculates the statistic and P-value for each feature relative to the target, allowing for direct comparison and ranking of feature importance.
- Under the Hood: The
chi2function returns two arrays: the Chi-Square statistic (F-score) and the corresponding P-value for every input feature column. - Execution Flow: Apply
train_test_splitfirst, then runchi2only on the training data (X_train,Y_train) to prevent data leakage during feature selection. - Best Practice: Convert the P-value array into a Pandas Series, indexed by the original feature names, and sort ascendingly to identify the most important features.
from sklearn.feature_selection import chi2
import pandas as pd
# Assume X_train and Y_train are prepared and encoded
chi_scores, p_values = chi2(X_train, Y_train)
# Create a Series for ranking and sort by P-value
feature_ranking = pd.Series(p_values, index=X_train.columns)
print(feature_ranking.sort_values(ascending=True))
Key Takeaway: Features are ranked by P-value; the lowest P-values indicate the strongest evidence against independence (highest importance).
Topics Covered in Chi-Square and Categorical Data
- Chi-Square Definition (00:00 - 01:22) — The test evaluates dependence between a non-negative feature and a categorical target class.
- Data Loading (01:22 - 02:22) — Load the dataset and identify the specific categorical features requiring transformation.
- Label Encoding (02:22 - 03:55) — String features must be converted to non-negative numerical inputs using techniques like numpy.where or dictionary mapping.
- Setup and Import (03:55 - 04:33) — Split the data into training sets and import the Chi2 function from Scikit-learn.
- Execute Chi2 (04:33 - 05:13) — Apply the chi2 function to the training data to generate the F-scores and P-values.
- Feature Ranking (05:13 - 06:30) — Convert the P-values into a Pandas Series and sort them to rank features by predictive significance.
Statistics for Data Science Cheat Sheet
-
from sklearn.feature_selection import chi2— Imports the statistical test functionfrom sklearn.feature_selection import chi2 -
chi2(X, y)— Calculates F-score and P-valuesscores, pvals = chi2(X_train, Y_train) -
np.where(condition, 1, 0)— Efficiently performs binary encodingdf['enc'] = np.where(df['cat'] == 'A', 1, 0) -
pd.Series(p_values, index=cols)— Maps P-values back to feature namesranking = pd.Series(p_values, index=X.columns) -
Series.sort_values(ascending=True)— Ranks features by lowest P-valueranking.sort_values(ascending=True) -
P-value < 0.05— Threshold for statistical significance
Comparison Table
| Metric | Goal | Interpretation |
|---|---|---|
| Chi-Square F-score | Measures magnitude of dependence. | Higher score means stronger relationship. |
| P-value | Measures statistical significance. | Lower value means relationship is real. |
| Label/Ordinal Encoding | Multi-class features. | Positive integer indices. |
| Binary Encoding | Two-class features. | 0 or 1. |
Common Pitfalls
- Mistake: Using negative numbers after encoding. Avoid: Ensure all encoded values are >= 0.
- Mistake: Running Chi2 on string features. Avoid: Always encode categorical features first.
- Mistake: Selecting features based only on F-score. Avoid: Use P-value as the primary determinant.
- Mistake: Running Chi2 on the full dataset. Avoid: Split data first, run selection only on X_train.
FAQs
- What is the null hypothesis in this test? The null hypothesis is that the feature and the target variable are statistically independent. We seek to reject this hypothesis.
- What P-value is considered 'low' or significant? Typically, a P-value below 0.05 is used to reject the null hypothesis, indicating the feature is significant.
- Can I use One-Hot Encoding instead of Label Encoding? Yes, but Label or Ordinal Encoding is often preferred here as Chi2 works well with integer representations of categories.