Add a Pike VM (regular engine, concept.md 7.2), dispatched automatically
Researched first (Cox's regexp2 article, the submatch/tagged-NFA follow-up crediting Laurikari, and rust-lang/regex's PikeVM source for the exact leftmost-first mid-search Match handling), then planned in concept.md 7.7 before writing any code, per the explicit instruction to plan before implementing. The existing bytecode compiler turned out to already produce valid Thompson-construction NFA bytecode (OP_SPLIT/OP_JMP/OP_SAVE match Cox's instruction set almost exactly), so no new compiler was needed: Prog gained one field (no_repeat1) and compile_repeat one condition, letting the same compile_node produce a second, NFA-only Prog from the same AST for any pattern containing none of OP_BACKREF/OP_LOOKAHEAD/ OP_LOOKBEHIND/OP_ATOMIC (a possessive quantifier already desugars to the last of these at parse time). That second Prog is matched by a new Pike VM (pike_addthread/pike_step/pike_find): a breadth-first thread list simulation with per-thread capture arrays, epsilon closure implemented with an explicit heap stack rather than C recursion so a pattern with many alternations cannot recurse the C stack, every allocation checked and failing through the existing -1 error convention rather than a NULL dereference. do_one, Pattern_finditer, and Pattern_split each gained an impl->has_nfa branch to this engine, sharing one small helper (find_next) for the "unanchored scan from a position" versus "single anchored attempt at a position" distinction finditer/split's empty-match retry needs. Two real bugs, both found and precisely localized by the existing 3,252-case CPython-derived suite without writing a single test specifically for this engine: OP_MATCH not writing group 0's end position (it has no OP_SAVE; run()'s own OP_MATCH handler sets it directly, and this engine's first version missed replicating that), and an unconditional "thread list empty -> stop" early exit that is wrong for unanchored search specifically, since a freshly injected start thread can die immediately in its own epsilon closure (a leading \b failing outright, repeatedly, inside a longer word like "catalog" for \bcat\b) without that meaning every later position would too. Both fixed; full suite passes, three clean AddressSanitizer/ UndefinedBehaviorSanitizer runs, plus hand-written whitebox checks (a 20,000-branch alternation, UTF8 named groups, BINARY matching across an embedded NUL, greedy/lazy and alternation priority). Measured result: (a+)+b, this project's own running example of the backtracking engine's remaining weak spot, is Pike VM eligible and now measures as genuinely linear (0.0018s to 0.0308s, n=10,000 to 160,000), not merely improved. Measured cost: re-running bench_vs_posix.c's three ordinary scenarios (all now Pike VM eligible too) found the gap to POSIX <regex.h> widened, from roughly 2x-25x before this engine existed to roughly 8x-70x now, the direct, expected cost of this engine's performance axis being explicitly deferred (no literal prefilter, no lazy DFA state caching, no allocation pooling) in favor of correctness first, per the instruction this was built under. Both results, and the reasoning behind deferring the second, are recorded in concept.md 7.7/7.8. README.md, docs/API.md, USAGE.md, and examples/redos_atomic.c and bench_vs_posix.c are updated throughout to describe the new two-engine dispatch accurately, including this real trade-off, rather than leaving the previous single-engine description in place; redos_atomic.c specifically now shows both that (a+)+b no longer needs an atomic group at all and a backreference-forced variant where one still does. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
This commit is contained in:
+1
-1
@@ -145,7 +145,7 @@ Mirror `re.Pattern.match`/`.fullmatch`/`.search` exactly, including the `pos`/`e
|
||||
- `fullmatch`: anchored at `pos`, must also reach exactly `endpos`.
|
||||
- `search`: tries every start position from `pos` to `endpos` inclusive, left to right, and reports the first that admits any match (ordinary backtracking priority decides which match that is at that position, `concept.md` 2.5).
|
||||
|
||||
**Performance:** `match`/`fullmatch` do one anchored attempt and cost time proportional to that attempt alone. `search` tries each candidate start position as a separate attempt, but, unlike a naive backtracking search, does not redo the same work at every one: for a backreference-free pattern, `run_memo` caches every proven failure at the (instruction, position) level, and `OP_REPEAT1` (a quantifier over a single character, class, or `.`) additionally uses precomputed run-length and skip-ahead tables so its own internal work is `O(1)` amortized per position rather than `O(remaining length)`. Together these make `search` linear, not quadratic, for the common case (a pattern with no backreference, built from simple repeated atoms). What is not fixed: a repeat over a *compound* body (`(ab)*`) gets the failure-memoization but not the skip-ahead table, so a pathological compound-repeat pattern can still be worse than linear; and any pattern with a backreference disables memoization entirely (unsound there, see Section 5) and can still be worst-case exponential, exactly as in CPython. See README.md "Implementation status" for the measured numbers on both the fixed case and the remaining one, and for citations to the published techniques this uses.
|
||||
**Performance:** which of two engines runs a given `Pattern` is decided once, at `re_compile` time, and is never a caller's choice (README.md "Implementation status", `concept.md` 7.2/7.7). A pattern with no backreference, lookahead, lookbehind, or atomic group (a possessive quantifier already desugars to the last of these) runs on the Pike VM, a Thompson-NFA simulation with no backtracking at all: every operation in this section is genuinely linear in input length on that engine, including `search` against an adversarial pattern shape like `(a+)+b` that would otherwise invite catastrophic backtracking, with no atomic group needed (README.md "The Pike VM"). Everything else (a backreference anywhere, or a lookaround/atomic construct) runs on the recursive backtracking engine instead: `match`/`fullmatch` there do one anchored attempt at time proportional to that attempt; `search` tries each candidate start position as a separate attempt but, unlike a naive backtracking search, does not redo the same work at every one, since `run_memo` caches every proven failure at the (instruction, position) level for a backreference-free pattern on this engine, and `OP_REPEAT1` additionally uses precomputed skip-ahead tables, together making `search` linear rather than quadratic for a pattern built from simple repeated atoms; a repeat over a *compound* body only gets the failure-memoization, and a pattern with a backreference disables memoization entirely (unsound there, Section 5) and can still be worst-case exponential, exactly as in CPython. The Pike VM currently has none of the backtracking engine's own constant-factor optimizations (no literal prefilter, no lazy DFA state caching), so an ordinary, non-adversarial pattern that happens to be Pike VM eligible can measure slower in absolute terms than the same pattern would have on the backtracking engine, a real, documented trade-off (`concept.md` 7.8), not an oversight. See README.md "Implementation status" and "The Pike VM" for the measured numbers on both engines and citations to the published techniques each uses.
|
||||
|
||||
**Memory:** the tables above cost real, measured memory, not just time complexity: `search`/`finditer`/`split`/`sub` on a pattern using `OP_REPEAT1` peak at roughly 8.4x the input length (measured with Valgrind/Massif; `match`/`fullmatch` do not allocate these tables at all and use proportionally less). README.md "Memory footprint" has the full measured breakdown, including a fix that was tried, measured, and deliberately reverted because it traded a fast allocation failure for an effectively-unbounded hang, which is a worse failure mode, not a better one.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user