Regex Best Practices: Writing Patterns That Don't Explode
A regex is a program in a very terse language, and like any program it has complexity, failure modes, and reviewers. These practices keep patterns small, safe, and fast in JavaScript specifically — where the engine is backtracking, the g flag has hidden state, and one bad quantifier can hang the event loop.
Anchor on purpose, not habit
/^email:/ and /END$/ cost almost nothing and stop the engine from trying the pattern at every position. But anchoring is also a semantic choice: test() on /foo/ matches foo inside barfoobaz — is that intended? Decide explicitly whether you're validating a whole field (^...$ or a full-match API) or searching within one (unanchored, and you'd better mean it). The email recipe is anchored because validation means "the whole string"; the phone recipe is alternation because formats differ — see the regex routes for both shapes.
Prefer explicit classes to .
. matches anything except newline, which is almost never what a delimiter-scan means: /<(.*)>/ on <b>hi</b> <i>there</i> captures b>hi</b> <i>there — the greedy . swallowed the middle tags. Three fixes, in order of preference:
- Character class of what's legal:
/<([^>]+)>/— can't cross a>by construction, and it's faster because failure is earlier. - Lazy with a real boundary:
/<(.*?)>/— still rescans on failure, but stops at the first>. .only when "anything, including the mess" is genuinely intended.
Know the flags you're combining
gontest()/exec()is a state machine:lastIndexpersists between calls on the same regex object, soif (/x/g.test(s))can return true then false for identical input. Use it only in a loop that consumes matches, or drop the flag.mchanges^$to line boundaries — pair it with intent, because/^SELECT/gmmatches line-lead keywords in a multi-line comment you forgot existed.u(unicode) changes semantics, not just display: it makes., quantifiers, and ranges code-point-aware, and it's what lets\p{L}property escapes work at all.s(dotAll) makes.match newlines — the honest way to say "across lines," better than the[.\n]folklore.
The ReDoS trap: nested quantifiers
(a+)+ style constructs — a quantifier containing a quantifier, or alternation inside a quantified group with overlapping branches — can take exponential time because each failure retries every partition. JS engines are backtracking engines (V8, SpiderMonkey) with no polynomial guarantee; catastrophic backtracking is the subject of its own guide and the one regex failure mode that crashes production rather than returning a wrong answer. The defense is structural, not a timeout:
bad: /^(\w+\s?)+$/ good: /^\w+(?:\s\w+)*$/
Same language, one pass, no nested loops over the same text. Before deploying any user-input-facing pattern, paste it into the tester with adversarial strings — 30 characters of aaaaaaaaaaaaaaaaaaaaaaaaaaaa! against the two patterns above will show you the difference in wall-clock, not just theory.
Build once, name everything
Literal regexes in hot code compile per parse but allocate per match; new RegExp in a loop is worse. Hoist module constants, and use named groups where readability pays:
const DATE = /^(?<y>\d{4})-(?<m>\d{2})-(?<d>\d{2})$/;
const { y, m } = DATE.exec(input)?.groups ?? {};
Numbered groups are fine for two groups; past four, (?: the non-capturing ones — every capture is memory and a renumbering hazard when someone inserts a group.
FAQ
Is .* with lazy .*? safe? Linear in the failure case (it scans to the boundary and extends), so usually fine; nested around another quantifier is where it stops being fine. The gotchas guide has the exact shapes.
Should I compile once and share across requests? Yes for literals (V8 caches them anyway); be careful with g-flagged shared instances because of lastIndex — reset .lastIndex = 0 or use String.matchAll which gives you a fresh iterator per call.
Are lookbehinds safe in Node? Since V8 shipped them (Node 8.3+/all modern browsers), yes, with one caveat: JS lookbehind must be bounded width — (?<=a*) is a syntax error. RE2 doesn't support lookarounds at all, which matters if your regex ever needs to run in Go: see the engines compared guide.