This lesson on Python Regular Expressions (Regex) is hands-on and example-driven. You will be able to construct regular expression patterns in Python to search, validate, and manipulate text strings using the built-in re module. You will master character sets, quantifiers, capture groups, and functions like re.search and re.sub for data cleaning pipelines.
What You'll Be Able To Do
- Match numeric and non-numeric characters using shorthand metacharacters like \d and \D
- Construct character classes and quantifiers to validate structured text formats
- Validate input strings against regex patterns using re.search()
- Capture substring components using parenthesized group syntax
- Reformat and clean structured text by referencing numbered capture groups with re.sub()
Detailed Concept Walkthrough
1. Character Classes and Quantifiers
Regular expressions define textual search patterns using shorthand tokens and character sets to match single or repeated sequences. They allow precise matching beyond static substring searches.
- Shorthand Metacharacters: Tokens like
\dmatch any decimal digit from 0 to 9, while\Dinverts the rule to match any non-digit character. These metacharacters act as universal wildcards for specific character categories. - Character Sets and Quantifiers: Square brackets
[...]define a custom set of permitted characters, such as[a-zA-Z0-9]for alphanumeric ranges. Appending the+quantifier matches one or more consecutive occurrences of any character in the preceding set. - Literal Escaping: Characters that serve syntactic purposes in regex (such as
.,+, and?) must be escaped with a backslash\.to match their literal character representations in strings.
import re
# Match one or more alphanumeric characters followed by literal dot and domain
pattern = r"[a-zA-Z0-9]+@(gmail|edu|net)\.com"
sample_text = "user123@gmail.com"
# Confirm whether the sample matches the defined character sets
is_match = bool(re.search(pattern, sample_text))
print(f"Pattern match: {is_match}") # Outputs: True
Key Takeaway: Character classes define allowable characters while quantifiers specify repetition counts, requiring backslash escapes for literal symbols.
2. Input Validation with re.search
The re.search() function scans an entire string to locate the first location where a regular expression pattern produces a match. It provides a boolean-like check for validating structured text.
- Match Evaluation:
re.search(pattern, string)scans through the target string and returns a Match object if the pattern is found, orNoneif no match occurs. In conditional evaluations, a Match object resolves toTrueandNoneresolves toFalse. - Raw String Prefixing: Prefixing regex string literals with
r(such asr"\d+") instructs Python to treat backslashes as raw literals rather than language-level escape sequences. This prevents unintended string parsing issues before the regex engine compiles the pattern. - Validation Logic: Wrapping
re.search()in conditional branches allows applications to accept or reject inputs such as email addresses, IDs, or account numbers based on structural validity.
import re
email_pattern = r"[a-zA-Z0-9]+@[a-zA-Z]+\.(com|edu|net)"
user_email = "student@university.edu"
# Validate input format using re.search
if re.search(email_pattern, user_email):
print("Valid email")
else:
print("Invalid email")
Key Takeaway: Use raw string literals with re.search() to validate string structure by checking for non-None Match objects.
3. Capture Groups and re.sub
Parentheses group sub-patterns into positional capture groups, enabling extraction and dynamic replacement through re.sub(). This allows targeted string transformation without altering surrounding text.
- Group Definition: Placing parentheses
(...)around portions of a regex pattern registers each segment as a distinct numbered capture group, indexed left-to-right starting at 1. These groups can be referenced later during substitution. - Substitution Engine: The
re.sub(pattern, repl, string)function searches for all occurrences ofpatterninstringand replaces them withrepl. The replacement string can contain plain text or backreferences to captured groups. - Backreference Assembly: Positional references like
\1,\2, and\3insert the exact text captured by corresponding parenthetical groups into the replacement string, allowing selective delimiter removal or reordering.
import re
# Isolate three groups of digits separated by hyphens
phone_pattern = r"(\d{3})-(\d{3})-(\d{4})"
replacement_pattern = r"\1\2\3"
raw_phone = "Call support at 800-555-0199 or 888-555-0100."
# Strip hyphens only from phone numbers while retaining digits
cleaned_text = re.sub(phone_pattern, replacement_pattern, raw_phone)
print(cleaned_text) # Outputs: Call support at 8005550199 or 8885550100.
Key Takeaway: Group sub-patterns with parentheses to capture segments and reassemble them cleanly using numbered backreferences in re.sub().
Topics Covered in Python Regular Expressions (Regex)
- Regex Overview (0:00 - 0:38) — Explains regular expressions as a universal pattern-matching mechanism across programming languages.
- Editor Testing and Metacharacters (0:39 - 1:05) — Demonstrates interactive regex search and contrasts digit matching using \d versus non-digit matching using \D.
- Email Validation Pattern (1:06 - 2:35) — Builds a regular expression for email validation using character classes, quantifiers, and escaped domain dots.
- Validation with re.search (2:36 - 3:24) — Implements re.search in Python conditional logic to verify valid and invalid user email inputs.
- String Transformation with re.sub (3:25 - 4:45) — Uses parenthesized capture groups and backreferences in re.sub to reformat hyphenated phone numbers.
Python Cheat Sheet
-
\d— Matches any single numeric digit between 0 and 9re.search(r"\d+", "Order 504") -
\D— Matches any single non-digit characterre.search(r"\D+", "504 Order") -
[...]— Matches any single character enclosed in the setre.search(r"[a-zA-Z]+", "DataEng") -
+— Matches one or more occurrences of preceding tokenre.search(r"\d+", "12345") -
\.— Matches a literal period characterre.search(r"\.com", "site.com") -
re.search(pattern, string)— Scans string for first pattern match returning Match/Nonematch = re.search(r"\d+", "ID: 99") -
re.sub(pattern, repl, string)— Replaces pattern occurrences with replacement stringclean = re.sub(r"(\d+)-(\d+)", r"\1\2", "12-34")
Comparison Table
| Feature | Primary Operation | Return Value |
|---|---|---|
| \d | Matches single decimal digits | Matched digit character |
| \D | Matches single non-digit characters | Matched non-digit character |
| re.search() | Locates first pattern occurrence | Match object or None |
| re.sub() | Replaces matched string patterns | Modified output string |
Common Pitfalls
- Mistake: Omitting the raw string r prefix before regex strings containing backslashes. Avoid: Prefix all regex strings with r to prevent Python escape sequence interpretation.
- Mistake: Using an unescaped dot when trying to match a literal period. Avoid: Always escape literal dots as backslash dot to prevent wildcard matching.
- Mistake: Using hardcoded string replacements when transforming dynamic text patterns. Avoid: Use parenthesized capture groups with numbered backreferences in re.sub replacement strings.
FAQs
- Why is it important to use raw strings (r"...") when defining regex in Python? Raw strings tell Python not to interpret backslashes as standard string escape sequences, ensuring metacharacters like \d reach the regex engine intact.
- What happens if re.search() does not find a match in the target string? It returns None, which evaluates to False in conditional statements without raising an exception.
- How do numbered backreferences like \1 work inside re.sub()? They refer to parenthesized capture groups in the search pattern, ordered sequentially from left to right starting at 1.