Regular Expressions (RegEx) in Python
Regular Expressions (RegEx)
A regular expression (often shortened to RegEx or regex) is a sequence of characters that defines a search pattern. Instead of searching for one exact piece of text, a regex describes the shape of the text you want — for example, "any 10-digit number", "a word that starts with a capital E", or "anything that looks like an email address".
With regular expressions, you can:
- Validate input — check that a phone number, PIN, or email has the right format
- Search for patterns — find the first date or order number in a document
- Extract data — pull every hashtag or price out of a block of text
- Replace text — mask digits, clean extra spaces, or reformat values
String methods such as find() and replace() work with fixed text. Regular expressions work with patterns, which makes them far more flexible.
The re Module
Python provides a built-in module called re for working with regular expressions.
import re
Once imported, you can use its functions to search and manipulate strings.
Use Raw Strings for Patterns
Regex patterns often contain backslashes, such as \d (a digit) or \b (a word boundary). In a normal Python string, a backslash also starts an escape sequence — "\b" means "backspace". To avoid this conflict, write patterns as raw strings by adding an r prefix:
pattern = r"\d+"
In a raw string, backslashes are passed to the regex engine unchanged. Using raw strings for every pattern is a good habit.
A First Example
The following example checks whether a sentence starts with "Hello" and ends with "World".
import re
text = "Hello Python World"
pattern = r"^Hello.*World$"
result = re.search(pattern, text)
print(result)
Expected output:
<re.Match object; span=(0, 18), match='Hello Python World'>
How the pattern works:
| Part | Meaning |
|---|---|
^ | The match must start at the beginning of the string |
Hello | The literal text "Hello" |
.* | Any characters (.), repeated zero or more times (*) |
World | The literal text "World" |
$ | The match must end at the end of the string |
If a match is found, re.search() returns a Match object; otherwise, it returns None. Because None is falsy, the result can be used directly in an if statement:
if re.search(pattern, text):
print("Pattern found")
RegEx Functions
The re module provides several functions:
| Function | Description | Returns |
|---|---|---|
findall() | Finds all matches | A list of strings |
search() | Finds the first match anywhere in the string | A Match object or None |
split() | Splits the string wherever the pattern matches | A list of strings |
sub() | Replaces matches with new text | A new string |
Two other useful functions are re.match(), which only checks for a match at the beginning of the string, and re.fullmatch(), which requires the entire string to match — ideal for validation.
Metacharacters
Metacharacters are characters with special meanings in a pattern.
| Symbol | Meaning | Example | Matches | ||
|---|---|---|---|---|---|
[] | A set of characters | [a-z] | Any one lowercase letter | ||
| `` | Starts a special sequence, or escapes a metacharacter | \d, 0 | A digit; a literal dot | ||
. | Any character except a newline | c.t | cat, cut, c9t | ||
^ | Starts with | ^Hi | "Hi" at the start of the string | ||
$ | Ends with | end$ | "end" at the end of the string | ||
* | Zero or more occurrences | go* | g, go, goo | ||
+ | One or more occurrences | go+ | go, goo (not g) | ||
? | Zero or one occurrence | colou?r | color, colour | ||
{} | An exact number (or range) of occurrences | \d{3} | Exactly three digits | ||
| ` | ` | Either or | `cat | dog` | cat or dog |
() | Groups part of a pattern and captures it | (abc)+ | abc, abcabc |
To match a metacharacter literally, escape it with a backslash. For example, r"0" matches an actual dot, and r"\$" matches a dollar sign.
Special Sequences
Special sequences start with a backslash and represent common character groups or positions.
| Sequence | Description | Example |
|---|---|---|
\A | Start of the string | r"\AHello" |
\b | Word boundary (the start or end of a word) | r"\bcat" |
\B | Not a word boundary | r"\Bcat" |
\d | A digit (0–9) | r"\d+" |
\D | A non-digit | r"\D" |
\s | A whitespace character (space, tab, newline) | r"\s" |
\S | A non-whitespace character | r"\S" |
\w | A word character (letter, digit, or underscore) | r"\w+" |
\W | A non-word character | r"\W" |
\Z | End of the string | r"end\Z" |
Example: Word Boundaries
import re
text = "cat concat category"
print(re.findall(r"\bcat", text))
print(re.findall(r"\Bcat", text))
Expected output:
['cat', 'cat']
['cat']
\bcatmatches "cat" only at the start of a word: incatandcategory.\Bcatmatches "cat" only when it is not at a word start: insideconcat.
Sets
A set is written in square brackets and matches one character from the group.
| Set | Matches |
|---|---|
[xyz] | x, y, or z |
[a-z] | Any lowercase letter |
[^abc] | Any character except a, b, or c |
[0-9] | Any digit |
[A-Za-z] | Any uppercase or lowercase letter |
[+] | A literal + |
Inside a set, most metacharacters lose their special meaning. [+] matches a plus sign, and [.] matches a dot. The ^ has special meaning only as the first character in the set, where it means "not".
import re
print(re.findall("[+]", "1+2+3"))
Expected output:
['+', '+']
Flags
Flags modify how a pattern behaves. Pass them with the flags argument.
| Flag | Short Form | Description |
|---|---|---|
re.IGNORECASE | re.I | Case-insensitive matching |
re.MULTILINE | re.M | ^ and $ match at the start and end of each line |
re.DOTALL | re.S | . also matches newline characters |
re.ASCII | re.A | \w, \d, \s match ASCII characters only |
re.VERBOSE | re.X | Allows spaces and comments in patterns for readability |
Example: Case-Insensitive Search
import re
text = "Python PYTHON python"
print(re.findall("python", text, flags=re.IGNORECASE))
Expected output:
['Python', 'PYTHON', 'python']
Multiple flags can be combined with |, for example flags=re.I | re.M.
The findall() Function
findall() returns all matches as a list.
import re
sentence = "Python is powerful and Python is easy"
matches = re.findall("Python", sentence)
print(matches)
Expected output:
['Python', 'Python']
If there are no matches, findall() returns an empty list [].
Practical Example: Extracting Numbers
import re
invoice = "Items: 3, Subtotal: 1250, Tax: 225"
numbers = re.findall(r"\d+", invoice)
print(numbers)
Expected output:
['3', '1250', '225']
\d+ means "one or more digits". The results are strings; convert them with int() if you need to calculate with them.
The search() Function
search() scans the string and returns the first match as a Match object.
import re
data = "Order number is 4598"
match = re.search(r"\d+", data)
print(match.start())
print(match.group())
Expected output:
16
4598
match.start()gives the index where the match begins.match.group()gives the matched text.
If no match exists, search() returns None. Calling .group() on None raises an AttributeError, so check the result first:
match = re.search(r"\d+", "No numbers here")
if match:
print(match.group())
else:
print("No match found")
Expected output:
No match found
The split() Function
split() splits a string wherever the pattern matches.
import re
text = "apple,banana,orange"
result = re.split(",", text)
print(result)
Expected output:
['apple', 'banana', 'orange']
The real power of re.split() is splitting on several different separators at once, which str.split() cannot do:
import re
text = "a, b;c d"
print(re.split(r"[,;\s]+", text))
Expected output:
['a', 'b', 'c', 'd']
The pattern [,;\s]+ matches one or more commas, semicolons, or whitespace characters.
Limiting the Number of Splits
The maxsplit argument limits how many splits are performed.
import re
text = "2024-05-20"
result = re.split("-", text, maxsplit=1)
print(result)
Expected output:
['2024', '05-20']
Only the first - is used as a split point.
Note: Pass
maxsplit(andflags) as keyword arguments. Passing them positionally, as inre.split("-", text, 1), is deprecated since Python 3.13.
The sub() Function
sub() ("substitute") replaces every match with new text.
import re
message = "My pin is 1234"
result = re.sub(r"\d", "*", message)
print(result)
Expected output:
My pin is ****
Each digit (\d) is replaced with *.
Limiting the Number of Replacements
The count argument limits how many replacements are made.
import re
message = "Call 9876543210 now"
result = re.sub(r"\d", "X", message, count=4)
print(result)
Expected output:
Call XXXX543210 now
Only the first four digits are replaced. As with split(), pass count as a keyword argument.
Practical Example: Cleaning Extra Spaces
import re
messy = "Python is fun"
print(re.sub(r"\s+", " ", messy))
Expected output:
Python is fun
The Match Object
A Match object contains information about a successful match.
import re
text = "Sun rises in the East"
match = re.search(r"\bE\w+", text)
print(match)
Expected output:
<re.Match object; span=(17, 21), match='East'>
The pattern \bE\w+ means: a word boundary, then a capital E, then one or more word characters — in other words, a word starting with "E".
Match Object Methods and Attributes
span() — Start and End Positions
print(match.span())
Expected output:
(17, 21)
The match starts at index 17 and ends before index 21.
string — The Original String
print(match.string)
Expected output:
Sun rises in the East
group() — The Matched Text
print(match.group())
Expected output:
East
Capturing Groups
Parentheses in a pattern create groups, which let you extract specific parts of a match.
import re
text = "Contact: asha@gocourse.com"
match = re.search(r"(\w+)@(\w+)\.com", text)
print(match.group(1))
print(match.group(2))
print(match.groups())
Expected output:
asha
gocourse
('asha', 'gocourse')
group(1)is the text matched by the first pair of parentheses (the user name).group(2)is the text matched by the second pair (the domain name).groups()returns all groups as a tuple.
Practical Example: Validating a Mobile Number
Indian mobile numbers have 10 digits and start with 6, 7, 8, or 9.
import re
def is_valid_mobile(number):
return re.fullmatch(r"[6-9]\d{9}", number) is not None
print(is_valid_mobile("9876543210"))
print(is_valid_mobile("12345"))
Expected output:
True
False
Pattern breakdown:
[6-9]— the first digit must be 6, 7, 8, or 9.\d{9}— followed by exactly nine more digits.re.fullmatch()— the whole string must match, so extra characters are rejected.
Using re.search() here would be a mistake: it would accept "abc9876543210xyz" because a valid number appears somewhere inside it.
Common Mistakes
| Mistake | Problem |
|---|---|
| Not using raw strings | "\b" becomes a backspace character; use r"\b" |
Calling .group() without checking for None | AttributeError: 'NoneType' object has no attribute 'group' |
Using search() for validation | Matches anywhere in the string; use fullmatch() |
Forgetting to escape . | . matches any character; use 0 for a literal dot |
Greedy .* matching too much | .* grabs as much as possible; use .*? for the shortest match |
Passing count/maxsplit positionally | Deprecated in Python 3.13+; use keyword arguments |
Related Concepts
- Python Strings and Python String Methods — simpler text operations
- Python Modules — importing the
remodule - Python User Input — validating input with regular expressions