RegexCalc Regex Engine

Regex Best Practices: Writing Patterns That Don't Explode

RegexCalc — Clean, Modern Regular Expression Calculator & Tester Guides · Updated 2026-10-01 · All guides

A regex is a program in a very terse language, and like any program it has complexity, failure modes, and reviewers. These practices keep patterns small, safe, and fast in JavaScript specifically — where the engine is backtracking, the g flag has hidden state, and one bad quantifier can hang the event loop.

Anchor on purpose, not habit

/^email:/ and /END$/ cost almost nothing and stop the engine from trying the pattern at every position. But anchoring is also a semantic choice: test() on /foo/ matches foo inside barfoobaz — is that intended? Decide explicitly whether you're validating a whole field (^...$ or a full-match API) or searching within one (unanchored, and you'd better mean it). The email recipe is anchored because validation means "the whole string"; the phone recipe is alternation because formats differ — see the regex routes for both shapes.

Prefer explicit classes to .

. matches anything except newline, which is almost never what a delimiter-scan means: /<(.*)>/ on <b>hi</b> <i>there</i> captures b>hi</b> <i>there — the greedy . swallowed the middle tags. Three fixes, in order of preference:

  1. Character class of what's legal: /<([^>]+)>/ — can't cross a > by construction, and it's faster because failure is earlier.
  2. Lazy with a real boundary: /<(.*?)>/ — still rescans on failure, but stops at the first >.
  3. . only when "anything, including the mess" is genuinely intended.

Know the flags you're combining

  • g on test()/exec() is a state machine: lastIndex persists between calls on the same regex object, so if (/x/g.test(s)) can return true then false for identical input. Use it only in a loop that consumes matches, or drop the flag.
  • m changes ^$ to line boundaries — pair it with intent, because /^SELECT/gm matches line-lead keywords in a multi-line comment you forgot existed.
  • u (unicode) changes semantics, not just display: it makes ., quantifiers, and ranges code-point-aware, and it's what lets \p{L} property escapes work at all.
  • s (dotAll) makes . match newlines — the honest way to say "across lines," better than the [.\n] folklore.

The ReDoS trap: nested quantifiers

(a+)+ style constructs — a quantifier containing a quantifier, or alternation inside a quantified group with overlapping branches — can take exponential time because each failure retries every partition. JS engines are backtracking engines (V8, SpiderMonkey) with no polynomial guarantee; catastrophic backtracking is the subject of its own guide and the one regex failure mode that crashes production rather than returning a wrong answer. The defense is structural, not a timeout:

bad:   /^(\w+\s?)+$/      good: /^\w+(?:\s\w+)*$/

Same language, one pass, no nested loops over the same text. Before deploying any user-input-facing pattern, paste it into the tester with adversarial strings — 30 characters of aaaaaaaaaaaaaaaaaaaaaaaaaaaa! against the two patterns above will show you the difference in wall-clock, not just theory.

Build once, name everything

Literal regexes in hot code compile per parse but allocate per match; new RegExp in a loop is worse. Hoist module constants, and use named groups where readability pays:

const DATE = /^(?<y>\d{4})-(?<m>\d{2})-(?<d>\d{2})$/;
const { y, m } = DATE.exec(input)?.groups ?? {};

Numbered groups are fine for two groups; past four, (?: the non-capturing ones — every capture is memory and a renumbering hazard when someone inserts a group.

FAQ

Is .* with lazy .*? safe? Linear in the failure case (it scans to the boundary and extends), so usually fine; nested around another quantifier is where it stops being fine. The gotchas guide has the exact shapes.

Should I compile once and share across requests? Yes for literals (V8 caches them anyway); be careful with g-flagged shared instances because of lastIndex — reset .lastIndex = 0 or use String.matchAll which gives you a fresh iterator per call.

Are lookbehinds safe in Node? Since V8 shipped them (Node 8.3+/all modern browsers), yes, with one caveat: JS lookbehind must be bounded width — (?<=a*) is a syntax error. RE2 doesn't support lookarounds at all, which matters if your regex ever needs to run in Go: see the engines compared guide.

Developer Sponsor / Partner
Copied to clipboard!