Lookaround Guide: Lookaheads, Lookbehinds, and What Each Engine Actually Allows
Lookaround answers "match X only if Y is (not) next to it" without putting Y in the result. That single property — the assertion consumes nothing — is what makes it both the most powerful and most misused construct in regex. This is the engine-by-engine truth, verified on Node v24.21.0, GNU grep with PCRE2, and RE2 semantics (no lookaround at all), plus a myth-busting section you will want when someone tells you JS lookbehind "must be fixed-length".
Support matrix, verified
| Construct | JavaScript (ES2018+) | PCRE2 (grep -P, php -r) | RE2 (Go, ripgrep's default engine) |
|---|---|---|---|
(?=Y) positive lookahead | yes | yes | no |
(?!Y) negative lookahead | yes | yes (verified grep -P '(?!a)b') | no |
(?<=Y) positive lookbehind | yes, variable length allowed | yes, but fixed length required | no |
(?<!Y) negative lookbehind | yes | yes (fixed length) | no |
| Lookaround containing captures usable later | yes | yes | n/a |
The middle two columns disagree, and the disagreement is the portability bug that bites teams moving patterns between JS and PHP/grep. Node 24 accepts and matches all of: /(?<=a+)b/ on "aab", /(?<=a|abc)d/ on "abcd", /(?<=ab*)c/ on "abbc" — variable-length lookbehinds, no complaints. The same shape under grep -P dies at compile time:
$ echo abbbc | grep -P '(?<=ab{1,3})c'
grep: lookbehind assertion is not fixed length
$ echo abbbc | grep -P '(?<=abb)c' # same match, fixed-length spelling
abbbc
So the correct statement is: PCRE lookbehind must be fixed-length (PCRE2, pcre2test 10.x, behaves the same); JavaScript since ES2018 (Node 8.3+, all evergreen browsers) allows arbitrary lookbehind bodies including +, *, alternation, and nested groups, because the engine backtracks the lookbehind like everything else. If your pattern ported from JS throws assertion is not fixed length in grep/PHP, rewrite the lookbehind to enumerate lengths: (?<=ab{1,3}) → (?<=ab|abb|abbb).
RE2 (Go, ripgrep's default, Google Code Search, RE2-backed API validators) has no lookaround whatsoever — lookahead or lookbehind is a parse error. The escape hatches, in order of preference: rewrite as alternation ((?=foo)bar with (?:foo)?bar only when the optional-ness is genuinely local), use \b and character classes which RE2 does support, or drop to a backtracking engine for that one pattern. The ReDoS guide covers why the RE2 constraint is a feature; here, accept the cost and see the cheatsheet for the same matrix in one screen.
Six patterns that earn their complexity
All run on Node 24; outputs shown.
1. Thousands separator, zero capture groups. The workhorse: insert a comma at any digit position followed by a multiple of three digits to end-of-string.
'1234567'.replace(/(?<=\d)(?=(\d{3})+$)/g, ',') // '1,234,567'
Two adjacent zero-width assertions: after a digit, before the 3n-boundary. A single lookahead can't express "after a digit" — that's why the pair exists.
2. Money without eating the dollar sign.
'price $42.50'.match(/(?<=\$)[\d.]+/)[0] // '42.50'
With (?<=\$) this captures only the number. The pre-ES2018 equivalent was capture-and-index: /\$([\d.]+)/ then [1] — fine here, but the lookbehind version composes: /(?<=\$)[\d.]+(?=\b)/ reads left to right with no mental offset math. grep -oP '(?<=\$)[\d.]+' on the same string prints 42.50 — PCRE handles it identically, fixed-length and all.
3. Stacked lookaheads as AND. Password rules without five if statements:
/^(?=.*\d)(?=.*[a-z])(?=.*[A-Z])\S{8,}$/.test('Str0ngPass') // true
/^(?=.*\d)(?=.*[a-z])(?=.*[A-Z])\S{8,}$/.test('short1A') // false
Each lookahead asserts a separate condition at the same position; the engine consumes nothing until all pass, then \S{8,}$ consumes the match. Stacked negative lookaheads are the same trick inverted — (?!TODO) rejects the whole match if the line contains a marker. (For production password UX, do the validation in five readable checks and show them; the one-liner is the worst error reporter ever written. It's shown because it's the canonical lookahead-as-AND teaching example, and the one every stackoverflow asks about.)
4. Escape-aware quote matching. Find the " characters that actually delimit fields, not the \" inside one:
'"a\\"b","c"'.match(/(?<!\\)"/g) // 4 matches — the two delimiters, twice each pair
Negative lookbehind on the backslash. In PCRE2, the same pattern works ((?<!\\)" — one character, trivially fixed-length); in JS, it also composes with the variable-length freedom if you need (?<!\\\\*)"-style parity logic. This is the pattern people otherwise solve with a state machine.
5. Split that keeps delimiters. Zero-width split points defined by an ahead-look, for tokenizer-ish jobs:
'a1b22'.split(/(?=\d)/) // [ 'a', '1b', '2', '2' ]
The 2 and 2 split apart because the second digit also has a digit ahead — (?=\d) saw the run boundary at every digit, not just run starts. (?<=\D)(?=\d) gives run starts only. Lookahead-as-split-point is how you tokenize without writing a lexer.
6. Word-boundary-adjacent word swap. Replace a word only when it's not followed by a particular word:
'foo bar foo'.replace(/\bfoo\b(?! bar)/g, 'x') // 'foo bar x'
The negative lookahead prunes candidates after position, so you never pay for a second pass filtering results. Same shape as pattern 3's rejection trick, scoped to one token.
Replace with lookarounds: the actual productivity case
Lookaround shines in replace because it lets the assertion do the context work that otherwise ends up in $1/$2 gymnastics or a callback:
'1,2,3'.replace(/,(?=\d)/g, ';') // '1;2;3' — commas between digits only
'v1.2.10'.replace(/(?<=\.\d+)\.(?=\d+$)/, '.') // touch the LAST dotted separator only
And the interplay with matchAll, because zero-width matches behave specially there — verified:
[...'abc'.matchAll(/(?=[bc])/g)] // two matches: ['', index 1], ['', index 2]
Empty-string matches land at their assertion positions with correct .index — so matchAll(/(?=[bc])/g) is an index finder, and that's a feature for highlighter/sourcemap work. Two rules: matchAll requires the g flag (no g is a TypeError, not a silent single match), and a zero-width pattern without g in .test/.exec cannot advance — an infinite while(m=rx.exec(...)) loop on a zero-width /g regex is the classic hang; bump lastIndex by one when m[0] === ''.
FAQ
Do I need to memorize the matrix? Only the disagreement: JS lookbehinds flex, PCRE ones must be fixed-width, RE2 has none. Everything else is either universal or a compile error you'll see immediately.
"Variable-length lookbehind" — didn't JS reject that until recently? It still does not reject it. ES2018 adopted the variable-length-capable form (Annex B / TC39 history: the spec requires engines to handle arbitrary bodies). Safari shipped it late (16.4, 2023) — that's the last holdout; if you must support iOS < 16.4, enumerate lengths for portability.
Why does my RE2 engine reject (?<=x)y but accept (?<name>x)? (?<name>...) (named capture) and (?<=...) (lookbehind) share a prefix and mean different things. RE2 supports named captures, not lookbehind — read the full construct, not the first two characters.
Are lookarounds slow? Zero-width assertions backtrack like any other subexpression; the pathological cases in ReDoS anatomy are mostly nested-quantifier shapes, and a single lookahead at a stable position costs roughly one probe. Profile before contorting, but keep patterns 3-style stacks on cold paths where an engine like RE2 would refuse them anyway.