RegexCalc Regex Engine

Regex Hacks: Sticky Tokens, Lookahead-AND, and the lastIndex Clipboard Walk

RegexCalc — Clean, Modern Regular Expression Calculator & Tester Guides · Updated 2026-10-01 · All guides

Not syntax basics — those are in the cheatsheet. These are the eight tricks that change what you can build with a JS regex. All outputs below were produced on Node v24.21.0 before publishing; where a trick has a footgun, the footgun is printed too.

1. Named groups inside replace callbacks

$<name> works in the replacement string, and the callback gets a groups object second-to-last:

const re = /(?<y>\d{4})-(?<m>\d{2})-(?<d>\d{2})/;
'2026-10-01'.replace(re, '$<m>/$<d>/$<y>');          // '10/01/2026'
'2026-10-01'.replace(re, (...a) => {
  const g = a.at(-2);                                 // the groups object
  return `${g.m}/${g.d}/${g.y}`;
});                                                    // '10/01/2026'

Why it matters: the callback's positional args shift the moment anyone adds a group (the renumbering trap). at(-2) is positional-proof. Pair with the d flag when you also need spans: /(?<x>a)/d.exec('za').indices → [[1,2],[1,2]] — whole-match and named-group offsets, ready for sourcemaps and syntax highlighters.

2. String.matchAll is an iterator — treat it like one

[...'a1 b22'.matchAll(/\d+/g)].map(m => [m[0], m.index]);
// [['1',1], ['22',4]]

Three properties that make it strictly better than a hand-rolled exec loop: it's a real iterator (Symbol.iterator, verified), it clones the regex so your lastIndex survives untouched ('12'.matchAll(r); r.lastIndex → 0), and it throws immediately if you forget g — THROWS: String.prototype.matchAll called with a non-global RegExp argument — instead of silently matching once like match does. Named groups ride along: [...'2026-10-01'.matchAll(re)].map(m => m.groups) → [{y:'2026',m:'10',d:'01'}].

3. Lookahead-as-AND: all-of-N rules on one line

Lookaheads are zero-width, so stacking them at ^ means "at position 0, all of these must each be true":

const pw = /^(?=.*[a-z])(?=.*[A-Z])(?=.*\d)(?=.*[!@#$%^&*]).{12,}$/;
pw.test('Str0ng!Passw');  // true
pw.test('Str0ngPassw0rd'); // false — no symbol

Each lookahead scans independently from position 0 — a logical AND without capture groups or state. Constrain the tail to a class and you get "exact length, only these chars, must contain both": /^(?=.*[a-z])(?=.*\d)[a-z\d]{8}$/.test('ab12cd34') → true, 'abcdefgh' → false. The same shape works for feature-flag combinations, password rules, and "must contain both ids" log filters.

4. Replace-with-function: the offset param is a scalpel

The callback signature is (match, ...groups, offset, wholeString):

'x **a** y **bb** z'.replace(/\*\*(.+?)\*\*/g, (m, g, off) => `[${off}->${g}]`);
// 'x [2->a] y [10->bb] z'

With offset you can edit by position instead of content — annotate every markdown emphasis with its byte range, or number matches in-place without a counter. One caution the printed args make obvious: with zero capture groups the callback is (match, offset, string), so label your params (m, off, str) and never index a[2] hoping it's the string — a.at(-3) is the flag-proof read.

5. The y (sticky) flag: continue exactly where you left off

Sticky means must match at lastIndex, nowhere later — the missing primitive for hand-parsers:

const sty = /\w+=\w+/y;
sty.lastIndex = 0;   sty.exec('key=value key2=value2'); // 'key=value', lastIndex → 9
sty.lastIndex = 10;  sty.exec('key=value key2=value2'); // 'key2=value2'
sty.lastIndex = 4;   sty.exec('key=value key2=value2'); // null — position 4 isn't a token start

A tokenizer in six lines:

const src = '42 + 17';
const tok = /\d+|[+]/y;
const toks = [];
for (let pos = 0; pos < src.length;) {
  tok.lastIndex = pos;
  const m = tok.exec(src);
  if (m) { toks.push(m[0]); pos = tok.lastIndex; } else pos++;
}
// toks = ['42', '+', '17']

Footgun, verified: a failed sticky exec resets lastIndex to 0, not to "unchanged" — so while (tok.lastIndex < src.length) with lastIndex++ on miss loops forever and fills your array with garbage (RangeError: Invalid array length). Drive the loop from your own pos variable and assign lastIndex every iteration, as above. That's also the honest JS answer to PCRE's \G: sticky is "anchor at lastIndex", exactly continue-where-left-off.

6. split with a capture group keeps the delimiters

'a1b2c'.split(/(\d)/);   // ['a','1','b','2','c']
'a1b2c'.split(/\d/);     // ['a','b','c']

The captures become array elements, so nothing is lost — the trick for tokenizers that must preserve operators, or markdown splitters that keep the headers. Want a delimiter split that skips quoted commas? Combine with the AND trick: 'a,"b,c",d'.split(/,(?=(?:[^"]*"[^"]*")*[^"]*$)/) → ['a','"b,c"','d'] — split only at commas with an even count of quotes ahead, i.e. outside quotes.

7. The clipboard walk: match.index + lastIndex

Every match carries the exact [m.index, re.lastIndex) span of the source string — the coordinates any clipboard editor, syntax highlighter, or codemod needs:

const re = /\w+@\w+\.\w+/g; const walk = []; let m;
while ((m = re.exec('mail [email protected] or [email protected] now')))
  walk.push([m[0], m.index, re.lastIndex]);
// [['[email protected]',5,12], ['[email protected]',16,23]]
text.slice(5, 12);   // '[email protected]' — cuttable, redactable, replaceable in place

Pii-redactor in one loop: collect spans, splice from the end so earlier offsets stay valid. (matchAll version: same spans via m.index/m[0].length, without the lastIndex state bug.)

8. Modulo by position, via lookahead math

"Every 3rd digit" has no regex operator, but a lookahead can count what's ahead — {n}$ means exactly n remain:

'123456789'.replace(/\d(?=(?:\D*\d){2}$)/g, 'X');       // '123456X89'  (3rd from end)
'123456789'.replace(/(?<=(?:\d{3})+\d{2})\d/g, 'X');    // lookbehind twin, counts from start
'123456789'.replace(/\d/g, (d, off) => (off+1) % 3 ? d : 'X'); // '12X45X78X' (every 3rd, honest)
'1234567'.replace(/\d(?=(?:\D*\d){3}$)/g, d => d + ','); // '1234,567'

Two verified lessons: (?:\D*\d){2}*$ (quantifier-then-star) is a Nothing to repeat SyntaxError — the star belongs inside; and for pure stride work the function form beats the lookahead form the moment the stride isn't counted from an end. Use the lookahead form where anchoring to $ is the actual rule (thousands separators, right-to-left checksums like Luhn).

Try them live

Every snippet drops into the tester — toggle the y/d flags, watch lastIndex and group indices update as you type. For which of these survive porting to a RE2 backend (spoiler: lookbehind doesn't), see the engine matrix; for making them fast, best practices.

FAQ

Is matchAll faster than an exec loop? Comparable; you buy correctness (clone semantics, spec'd zero-width advance), not speed. For huge inputs the iterator also lets you break/take lazily instead of materializing all matches like match with g.

Does the y flag exist in other engines? PCRE2 has \G, same anchor-at-position semantics, and it combines with lookaround (which JS sticky can't distinguish from a plain match-at-position — same result, different syntax). RE2 supports neither, because it has no positions to resume from in its lockstep simulation.

Why not just use a real CSV/URL parser? Always do. These hacks are for the 90% of cases where the input is your own logs, your own config, your generated markdown — and a dependency would be overkill. Quoted-CSV split #6 is exactly the boundary where you stop and reach for papaparse.

Developer Sponsor / Partner
Copied to clipboard!