Think in building blocks
Regular expressions compose from a small vocabulary: literals (abc), character classes ([a-z], d, w, s and their negations), quantifiers (* 0+, + 1+, ? 0–1, {2,5}), alternation (a|b), groups ((…)), anchors (^ $ ). Every regex you'll ever need is these blocks arranged to describe the shape of what you're matching.
Recipes worth memorizing
# ISO date, with groups
(d{4})-(d{2})-(d{2})
# hex color (case-insensitive, optional #)
#?([0-9a-f]{3}|[0-9a-f]{6})
# key=value pairs in a config line
^(w+)s*=s*(.*)$
# a URL-safe slug
^[a-z0-9]+(?:-[a-z0-9]+)*$
# whole-word match (avoid cat matching category)
cat
# negative lookahead: lines not starting with #
^(?!#).+$
The four classic mistakes
- Greedy by default:
<.*>on<a>x</a>matches the whole string. Use lazy.*?or a negated class<[^>]*>. - Catastrophic backtracking: nested quantifiers like
(a+)+$against a near-miss string take exponential time — the classic ReDoS vector. Un-nest, or use atomic groups/possessive quantifiers where supported. - Unanchored surprises: without
^…$, the regex matches anywhere —d{4}finds "2048" inside "12048". Anchor, or use. - Dot matches almost nothing:
.excludes newlines unless you enable dot-all (sflag). Parse HTML with a parser, not regex; parse JSON with a JSON parser. Regex is for text shapes, not structure.
A testing discipline
Write your test set first: strings that must match, must-not-match, and near-misses (extra character, wrong case, empty). Iterate against the set, then keep the tests. When a regex survives ten near-misses, it deserves to ship. Named groups ((?<year>d{4})) make the final expression self-documenting.
Build and test against live matches with highlight and capture groups in the Regex Tester.