The Regex Cheatsheet: JS vs PCRE2 vs RE2, Plus 20 Recipes That Actually Work
One screen: the syntax you re-Googles weekly, the constructs your engine may not have, and twenty patterns we ran on Node v24.21.0 before publishing them. Every output below is real, from this box. When you need the tool, paste and tweak in the tester.
Syntax primer
| Construct | Meaning | Notes |
|---|---|---|
^ / $ | Start / end anchors | $ matches before a trailing \n without m; anchors cut the scan, not the backtrack |
\b | Word boundary | Zero-width; \B is its complement |
\d \w \s | Digit / word / whitespace classes | \w is [0-9A-Za-z_] in JS — it lies about accents (see below) |
[abc] [^abc] | Class / negation | .-in-class and $-in-class are literals; see gotchas |
. | Any char except newline | Use [\s\S] for "everything", or s flag |
* + ? | 0+, 1+, 0/1 | Greedy by default; append ? to make lazy |
{n,m} | Counted repeat | {n,} = n or more |
(?<name>...) | Named capture | Read .groups.name; immune to renumbering |
(?:...) | Non-capturing group | Default to this; capture on purpose |
(?=) (?!) | Lookahead / neg | Zero-width assertions |
(?<=) (?<!) | Lookbehind | JS: yes, any length. PCRE2: yes. RE2: no |
Flags g i m s u y d | Global, ignore-case, multiline, dotAll, unicode, sticky, indices | y and d are the under-used ones — see hacks |
Possessive and atomic quantifiers (a++, (?>a+)) do not exist in JavaScript. Verified on this box: /a++/ throws Nothing to repeat and /(?>a)/ throws Invalid group — even under the v flag on Node 24. They are PCRE2-only tools for killing backtracking, and the JS rewrite is a disjoint quantifier, not a syntax flag.
Engine support matrix (the tricky bits)
| Feature | JS (V8) | PCRE2 | RE2 (Go/Rust crate) |
|---|---|---|---|
Lookbehind (?<=…) | Yes, variable length too | Yes (bounded repeats) | No — compile error |
Atomic group (?>…) | No | Yes (grep -P '(?>a)a' → aa) | No (no backtracking, so moot) |
Possessive a++ | No (SyntaxError) | Yes | No (moot) |
\p{L} property classes | Yes, with u or v flag | Yes | Yes |
\K (reset start) | No — \Ka tests false, silently | Yes (grep -P '\Kab' → ab) | No |
Backreferences \1 | Yes | Yes | No — compile error |
Named backref \k<name> | Yes | Yes | No |
Verified here: JS lookbehind — '€5'.replace(/(?<=€)\d/, '9') → €9; variable-length — 'xxabc'.replace(/(?<=x{2})abc/, 'ABC') → xxABC. PCRE2 column spot-checked with GNU grep 3.11 (grep -P, PCRE2-backed). The RE2 column is structural: its engine simulates all states in lockstep, so constructs that need history (backrefs, lookaround) fail at compile time, not match time. Consequences for ReDoS in the dedicated deep dive.
Unicode reality check, verified: 'café'.match(/\w+/) returns ["caf"] — with or without u, \w stops at é. 'naïve café'.replace(/\p{L}+/gu, '[$&]') returns [naïve] [café]. If your user base has accents, \w is a bug, not a style choice.
Top 20 recipes (tested on Node v24)
Anchored validation patterns first — each line is pattern + what it did:
- Email-ish:
/^[^\s@]+@[^\s@]+\.[^\s@]+$/— accepts[email protected], rejectsada@example. Read the rant below before using it. - URL (loose extract):
/\bhttps?:\/\/\S+/gi— onsee https://example.com/a?b=1 and http://x.io.returns exactly["https://example.com/a?b=1","http://x.io."]. Trailing-punctuation trimming is your job; no regex parses URLs correctly, usenew URL(). - IPv4 (octet-exact):
/^(25[0-5]|2[0-4]\d|1\d{2}|[1-9]?\d)(\.(25[0-5]|2[0-4]\d|1\d{2}|[1-9]?\d)){3}$/—255.192.0.1true,256.1.1.1false,1.2.3false,01.1.1.1false. - Semver (strict, spec grammar):
/^(0|[1-9]\d*)\.(0|[1-9]\d*)\.(0|[1-9]\d*)(?:-((?:0|[1-9]\d*|\d*[a-zA-Z-][\-0-9a-zA-Z]*)(?:\.(?:0|[1-9]\d*|\d*[a-zA-Z-][\-0-9a-zA-Z]*))*))?(?:\+([\-0-9a-zA-Z]+(?:\.[\-0-9a-zA-Z]+)*))?$/—1.2.3true,1.2.3-beta.1+build.5true,01.2.3false (leading zeros are invalid, and lazy patterns accept them). - Time HH:MM(:SS):
/^(?:[01]\d|2[0-3]):[0-5]\d(?::[0-5]\d)?$/—23:59true,09:30:45true,24:00false. - ISO date:
/^\d{4}-(?:0[1-9]|1[0-2])-(?:0[1-9]|[12]\d|3[01])$/—2026-10-01true,2026-13-01false. (No leap-year logic;Date.parsefor that.) - Quoted CSV field:
/^(?:"(?:[^"]|"")*"|[^",\r\n]*)$/—"he said ""hi"""true, bareplain valuetrue,"openfalse. - CSV row split (quote-aware):
line.split(/,(?=(?:[^"]*"[^"]*")*[^"]*$)/)—a,"b,c",d→["a","\"b,c\"","d"]. - Hex color:
/^#(?:[0-9a-fA-F]{3}|[0-9a-fA-F]{6})$/—#ffftrue,#FFA042true,#FFA042cfalse. Extend with{4}|{8}for alpha. - Slug validate:
/^[a-z0-9]+(?:-[a-z0-9]+)*$/—hello-world-2true,-dashfalse (no leading/trailing/double dashes). - Slugify (transform):
s.toLowerCase().replace(/[^a-z0-9]+/g, '-').replace(/^-+|-+$/g, '')—Hello, World: Second Edition!→hello-world-second-edition. - Password rule combo (all-of, via lookahead-as-AND):
/^(?=.*[a-z])(?=.*[A-Z])(?=.*\d)(?=.*[!@#$%^&*]).{12,}$/—Str0ng!Passwtrue,Str0ngPassw0rdfalse (no symbol). Mechanics in hacks. - Extract digits:
'a1 b22'.matchAll(/\d+/g)— iterator, yields["1",1] ["22",4]with indices. - Thousands separators:
'1234567'.replace(/\d(?=(?:\D*\d){3}$)/g, d => d + ',')→1234,567. (Or justIntl.NumberFormat, honestly.) - Trailing-space killer:
text.replace(/[ \t]+$/gm, '')—mmakes$per-line. - Collapse whitespace runs:
s.replace(/\s+/g, ' ').trim(). - Match repeated word (case-insensitive):
/\b(\w+)\s+\1\b/gi— the backref family RE2 can never have. - Strip HTML tags (the good-enough one):
html.replace(/<\/?[^>]+>/g, '')— fine for stripping, cursed for parsing. - Key=value pairs:
/\b[\w.-]+=(?:\w+|"[^"]*")/gona=1 b="x y"yields both pairs. - Hex color extract from design file:
/#(?:[0-9a-fA-F]{3,4}|[0-9a-fA-F]{6}|[0-9a-fA-F]{8})\b/gi.
The email pattern, honestly
Recipe 1 is not "email validation". RFC 5322 permits "a b"@example.com, comments, IP literals — and recipe 1 rejects all of it (verified: "a b"@example.com → false). Meanwhile it happily accepts [email protected] (verified true), a domain that can receive nothing. Here is the industry's dirty secret: no serious team validates email with regex. The RFC-conformant pattern is ~64KB of pain, still doesn't prove deliverability, and the only universal check is the confirmation link you already send. Use the 20-char version as a typo filter (catches @@, missing dot, spaces) and move on. The only regex email that ever shipped correctly is the one next to a send-and-verify step.
Test it here
Every pattern above auto-fills into the RegexCalc tester with its test string — flags, groups, positions, and replace preview included. When a recipe's engine column matters (lookbehind on a RE2 backend, \K anywhere), the matrix above is the one-page answer; regex alternatives covers when to switch engines entirely.
FAQ
Why doesn't JS support possessive quantifiers at all? V8's backtracking engine could add them, but a disjoint-quantifier rewrite ((a+)b → aa*b style fixes, or atomic-emulating lookahead tricks) covers the fixable cases, and the unfixable ones are ReDoS anyway. See ReDoS protection.
Do I need the u flag? For \p{…}, correct case-folding, and astral-safe .: yes, always pair them. Omitting u is worse than an error: /\p{L}+/g without the flag is parsed as the literal p, so 'café'.replace(/\p{L}+/g, 'X') returns café unchanged — silent failure. And . without u counts surrogate pairs twice: '😀a'.replace(/./g, '_') gives ___, with u gives __.
Are these patterns performance-safe? All 20 are linear-friendly: no nested quantifiers, no overlapping alternations. That's a deliberate constraint — see best practices for the rules they follow.