Understanding Regular Expressions
Regular expressions (regex) are powerful pattern-matching tools used across virtually every programming language. They enable you to search, validate, extract, and manipulate text with precision. Whether you're validating email addresses, parsing log files, or cleaning data, regex provides a concise syntax for describing complex text patterns.
How Regular Expressions Work
A regex engine processes your pattern character by character, attempting to match it against the input text. The engine uses backtracking to try different possibilities when a match fails. Understanding this behavior helps you write efficient patterns that avoid catastrophic backtracking on large inputs.
Regex Syntax Quick Reference
Character Classes
| Syntax | Matches | Example |
|---|---|---|
| . | Any character except newline | a.c matches "abc", "a1c" |
| \\d | Any digit (0-9) | \\d{3} matches "123" |
| \\D | Any non-digit | \\D+ matches "abc" |
| \\w | Word character (a-z, A-Z, 0-9, _) | \\w+ matches "hello_123" |
| \\W | Non-word character | \\W matches "@", "#" |
| \\s | Whitespace (space, tab, newline) | \\s+ matches " " |
| \\S | Non-whitespace | \\S+ matches "word" |
| [abc] | Any of a, b, or c | [aeiou] matches vowels |
| [^abc] | Not a, b, or c | [^0-9] matches non-digits |
| [a-z] | Range: a through z | [A-Za-z] matches letters |
Quantifiers
| Syntax | Meaning | Example |
|---|---|---|
| \* | 0 or more (greedy) | a* matches "", "a", "aaa" |
| + | 1 or more (greedy) | a+ matches "a", "aaa" |
| ? | 0 or 1 (optional) | colou?r matches "color", "colour" |
| {n} | Exactly n times | \\d{4} matches "2024" |
| {n,} | n or more times | \\w{3,} matches 3+ chars |
| {n,m} | Between n and m times | \\d{2,4} matches "12", "123", "1234" |
| \*? | 0 or more (lazy) | a*? matches minimum |
| +? | 1 or more (lazy) | a+? matches single "a" |
Anchors and Boundaries
| Syntax | Matches | Use Case |
|---|---|---|
| ^ | Start of string/line | ^Hello matches "Hello World" at start |
| $ | End of string/line | end$ matches "the end" |
| \\b | Word boundary | \\bcat\\b matches "cat" not "catalog" |
| \\B | Non-word boundary | \\Bcat matches "catalog" |
Groups and Lookarounds
| Syntax | Purpose | Example |
|---|---|---|
| (...) | Capturing group | (\\d{3})-(\\d{4}) captures area code |
| (?:...) | Non-capturing group | (?:https?://) groups without capturing |
| (?=...) | Positive lookahead | \\d(?=px) matches digit before "px" |
| (?!...) | Negative lookahead | \\d(?!px) matches digit not before "px" |
| (?<=...) | Positive lookbehind | (?<=\\$)\\d+ matches digits after "$" |
Common Regex Patterns
| Use Case | Pattern | Matches |
|---|---|---|
[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\\.[a-zA-Z]{2,} | user@example.com | |
| US Phone | \\(?\\d{3}\\)?[-. ]?\\d{3}[-. ]?\\d{4} | (555) 123-4567 |
| URL | https?://[\\w.-]+(?:/[\\w.-]*)* | https://example.com/path |
| IP Address | \\b(?:\\d{1,3}\\.){3}\\d{1,3}\\b | 192.168.1.1 |
| Date (ISO) | \\d{4}-\\d{2}-\\d{2} | 2024-12-25 |
| Hex Color | #(?:[0-9a-fA-F]{3}){1,2}\\b | #fff, #a1b2c3 |
| ZIP Code | \\d{5}(?:-\\d{4})? | 12345, 12345-6789 |
Tips for Writing Effective Regex
1. Start simple and build incrementally - Test each part of your pattern before combining them. Complex patterns are easier to debug when built step by step.
2. Be as specific as possible - Use character classes like \\d instead of . when you know what type of character to expect. This prevents unexpected matches.
3. Use non-capturing groups when you don't need the capture - (?:...) is more efficient than (...) when you only need grouping, not extraction.
4. Escape special characters - Characters like . * + ? ^ $ { } [ ] \\ | ( ) have special meaning. Escape them with \\ when matching literally.
5. Avoid catastrophic backtracking - Patterns like (a+)+ can cause exponential time complexity. Use atomic groups or possessive quantifiers when available.
6. Use anchors for validation - When validating entire strings, always use ^ and $ anchors to ensure the pattern matches the complete input.
7. Test with edge cases - Empty strings, very long inputs, special characters, and unicode should all be tested against your pattern.