Files
regexx/examples/redos_atomic.c
T
retoorandClaude Sonnet 5 fa07622671 Add a Pike VM (regular engine, concept.md 7.2), dispatched automatically
Researched first (Cox's regexp2 article, the submatch/tagged-NFA
follow-up crediting Laurikari, and rust-lang/regex's PikeVM source for
the exact leftmost-first mid-search Match handling), then planned in
concept.md 7.7 before writing any code, per the explicit instruction to
plan before implementing.

The existing bytecode compiler turned out to already produce valid
Thompson-construction NFA bytecode (OP_SPLIT/OP_JMP/OP_SAVE match
Cox's instruction set almost exactly), so no new compiler was needed:
Prog gained one field (no_repeat1) and compile_repeat one condition,
letting the same compile_node produce a second, NFA-only Prog from the
same AST for any pattern containing none of OP_BACKREF/OP_LOOKAHEAD/
OP_LOOKBEHIND/OP_ATOMIC (a possessive quantifier already desugars to
the last of these at parse time). That second Prog is matched by a new
Pike VM (pike_addthread/pike_step/pike_find): a breadth-first thread
list simulation with per-thread capture arrays, epsilon closure
implemented with an explicit heap stack rather than C recursion so a
pattern with many alternations cannot recurse the C stack, every
allocation checked and failing through the existing -1 error
convention rather than a NULL dereference. do_one, Pattern_finditer,
and Pattern_split each gained an impl->has_nfa branch to this engine,
sharing one small helper (find_next) for the "unanchored scan from a
position" versus "single anchored attempt at a position" distinction
finditer/split's empty-match retry needs.

Two real bugs, both found and precisely localized by the existing
3,252-case CPython-derived suite without writing a single test
specifically for this engine: OP_MATCH not writing group 0's end
position (it has no OP_SAVE; run()'s own OP_MATCH handler sets it
directly, and this engine's first version missed replicating that),
and an unconditional "thread list empty -> stop" early exit that is
wrong for unanchored search specifically, since a freshly injected
start thread can die immediately in its own epsilon closure (a
leading \b failing outright, repeatedly, inside a longer word like
"catalog" for \bcat\b) without that meaning every later position
would too. Both fixed; full suite passes, three clean AddressSanitizer/
UndefinedBehaviorSanitizer runs, plus hand-written whitebox checks
(a 20,000-branch alternation, UTF8 named groups, BINARY matching
across an embedded NUL, greedy/lazy and alternation priority).

Measured result: (a+)+b, this project's own running example of the
backtracking engine's remaining weak spot, is Pike VM eligible and now
measures as genuinely linear (0.0018s to 0.0308s, n=10,000 to
160,000), not merely improved. Measured cost: re-running
bench_vs_posix.c's three ordinary scenarios (all now Pike VM eligible
too) found the gap to POSIX <regex.h> widened, from roughly 2x-25x
before this engine existed to roughly 8x-70x now, the direct,
expected cost of this engine's performance axis being explicitly
deferred (no literal prefilter, no lazy DFA state caching, no
allocation pooling) in favor of correctness first, per the instruction
this was built under. Both results, and the reasoning behind
deferring the second, are recorded in concept.md 7.7/7.8.

README.md, docs/API.md, USAGE.md, and examples/redos_atomic.c and
bench_vs_posix.c are updated throughout to describe the new two-engine
dispatch accurately, including this real trade-off, rather than
leaving the previous single-engine description in place; redos_atomic.c
specifically now shows both that (a+)+b no longer needs an atomic group
at all and a backreference-forced variant where one still does.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
2026-09-14 11:38:44 +00:00

115 lines
5.8 KiB
C

/* redos_atomic - demonstrates atomic groups (?>...) as a defense against
* catastrophic backtracking (ReDoS), on the textbook (a+)+b shape, and
* shows precisely when that defense is still needed versus when it no
* longer is.
*
* (a+)+b by itself is no longer a useful demonstration of the problem:
* it has no backreference, lookahead, lookbehind, or atomic group, so it
* is eligible for and automatically matched by the Pike VM (concept.md
* 7.2/7.7), which runs it in genuinely linear time, not merely improved
* time, with no atomic group needed at all (README.md "The Pike VM").
* The backtracking engine's own remaining ceiling for this exact shape
* (empirically quadratic, not exponential, because the inner a+'s
* OP_REPEAT1 is followed by the group's closing save rather than a
* simple atom, so the skip-ahead table that makes a*b linear does not
* apply to it) is real and still documented in README.md "Implementation
* status", but only actually applies to a pattern that keeps a
* construct forcing it onto that engine. This program shows both halves
* directly: first, that the plain (a+)+b form is now fast on its own,
* with no atomic group; second, a variant with a trailing backreference
* added specifically to keep it on the backtracking engine, where the
* same nested-quantifier ambiguity is still exponential (memoization is
* unsound in the presence of a backreference and is disabled whenever
* one is present anywhere in the pattern, README.md "Implementation
* status"), and where the atomic group fix from this file's original
* version still matters exactly as much as it always did.
*/
#define _POSIX_C_SOURCE 199309L
#include "regexx.h"
#include <stdio.h>
#include <string.h>
#include <stdlib.h>
#include <time.h>
static double now_seconds(void) {
struct timespec ts;
clock_gettime(CLOCK_MONOTONIC, &ts);
return (double)ts.tv_sec + (double)ts.tv_nsec / 1e9;
}
static void run_case(const char *label, const char *pattern, const char *subject, size_t subjlen) {
PatternError err; memset(&err, 0, sizeof err);
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
if (!pat) {
fprintf(stderr, "compile error for %s: %s\n", pattern, err.msg);
PatternError_free(&err);
return;
}
Input *in = Input_from_buffer((const uint8_t *)subject, subjlen);
Match m; memset(&m, 0, sizeof m);
double t0 = now_seconds();
int rc = Pattern_search(pat, in, 0, -1, &m);
double t1 = now_seconds();
printf("%-28s pattern=%-18s n=%-8zu rc=%-2d time=%.4fs\n",
label, pattern, subjlen, rc, t1 - t0);
if (rc == 1) Match_free(&m);
Input_free(in);
Pattern_free(pat);
}
int main(void) {
printf("=== Part 1: (a+)+b, backreference-free, now Pike VM eligible ===\n\n");
{
size_t n = 20000;
char *subject = malloc(n + 1);
memset(subject, 'a', n);
subject[n] = '\0';
printf("Searching %zu characters of 'a' with no trailing 'b' (never matches):\n\n", n);
run_case("plain nested quantifier", "(a+)+b", subject, n);
run_case("atomic inner group", "(?>a+)+b", subject, n);
run_case("possessive inner quant.", "(a++)+b", subject, n);
printf("\nAll three finish in comparable, close-to-instant time: the plain form\n"
"is no longer the slow one. It has no backreference, lookahead,\n"
"lookbehind, or atomic group, so it runs on the Pike VM (README.md\n"
"\"The Pike VM\"), which holds a bounded set of live parse states\n"
"instead of backtracking, and is linear in n regardless of this\n"
"pattern's nested-quantifier shape. The atomic and possessive forms\n"
"were not what fixed this; the engine dispatch did.\n\n");
free(subject);
}
printf("=== Part 2: the same shape, forced onto the backtracking engine ===\n\n");
{
/* \2 (a backreference to the literal "b") is enough to make this
* pattern ineligible for the Pike VM (no backreference can ever
* be, README.md "Implementation status"), which is the only
* change from Part 1 that matters here: it puts the same
* (a+)+-shaped ambiguity back onto the engine that still pays
* for it. n is deliberately tiny (not the 20000 above): a
* backreference disables memoization entirely (it is unsound
* whenever one is present anywhere in the pattern, not only
* where it appears), so this shape is exponential here, not
* merely quadratic, and n = 24 already takes most of a second. */
size_t n = 24;
char *subject = malloc(n + 1);
memset(subject, 'a', n);
subject[n] = '\0';
printf("Searching %zu characters of 'a' with no trailing \"bb\" (never matches):\n\n", n);
run_case("plain, backreference", "(a+)+(b)\\2", subject, n);
run_case("atomic, backreference", "(?>(a+))+(b)\\2", subject, n);
printf("\nThe plain form is slow again (exponential, not merely quadratic,\n"
"since the backreference disables memoization outright), and the\n"
"atomic group fixes it again, exactly as it always did: wrapping the\n"
"capturing group itself, (?>(a+))+, keeps group 1 capturing (unlike\n"
"(?>a+)+, which would make the inner run non-capturing) while still\n"
"forbidding the atomic sub-program from ever giving back characters\n"
"it already consumed to try a shorter run instead. This is the\n"
"pattern author's own tool for the minority of patterns that still\n"
"need it, not something the engine can fix automatically the way it\n"
"now does for Part 1's plain, backreference-free shape.\n");
free(subject);
}
return 0;
}