From 5ae68d4ea609a4f0d2b8549cddeea51e389372ca Mon Sep 17 00:00:00 2001 From: retoor Date: Mon, 14 Sep 2026 12:21:15 +0000 Subject: [PATCH] Add a literal prefilter to the Pike VM, profile its memory with Massif Closed two of the three gaps the previous commit's honest self-review left open (the third, full streaming/bounded-memory input, remains out of scope for this pass and is still documented as such). Literal prefilter (regexx.c, pike_find): when the NFA-only Prog's first instruction is a mandatory OP_CHAR or OP_CLASS, a pattern beginning with a required literal or class rather than a nullable loop or a leading assertion, injecting a fresh unanchored start thread at a position that instruction would reject is certain to die on the very next pike_step call regardless; checking that identical condition before injecting rather than after changes nothing about which threads ever exist, only how much wasted work is done finding out. This targeted exactly examples/bench_vs_posix.c's worst regression: the literal-search scenario went from roughly 70x slower than POSIX (up from roughly 22x before the Pike VM existed) down to roughly 11x-13x, better than the original pre-Pike-VM number; number extraction (starts with a class) improved more modestly; a*b (starts with a nullable loop, structurally unhelped) is unchanged, as expected. Verified with the full 3,252-case suite, three clean AddressSanitizer/UndefinedBehaviorSanitizer passes, and a rerun of the whitebox dual-engine cross-check (24,000 match/fullmatch/search plus ~2,700 finditer comparisons between the two engines on the same compiled patterns, zero mismatches). Memory profiling (concept.md 7.9): Valgrind/Massif on the same adversarial, prefilter-proof pattern shape (a*b, nullable leading loop) used for the backtracking engine's own worst case, for a fair comparison. Peak heap was almost entirely the 10MB input buffer itself; the Pike VM's own contribution was roughly 12KB, confirming the design's O(instruction count x group count), input-length- independent memory bound actually holds for the v1 implementation, not only on paper. concept.md Section 10's table, which had only a "not yet profiled" caveat for this row before, is updated with the measured result. README.md, docs/API.md, USAGE.md, and bench_vs_posix.c's own printed summary are updated throughout with the corrected numbers, rather than left describing the pre-prefilter regression as current. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY --- README.md | 71 ++++++++++++++++++++++++++++----------- USAGE.md | 9 ++--- concept.md | 12 ++++++- docs/API.md | 2 +- examples/bench_vs_posix.c | 34 ++++++++++--------- regexx.c | 43 ++++++++++++++++++++---- 6 files changed, 123 insertions(+), 48 deletions(-) diff --git a/README.md b/README.md index 0cc02a9..fd54392 100644 --- a/README.md +++ b/README.md @@ -184,9 +184,35 @@ result worth restating plainly here: `(a+)+b`, this document's own running example of the backtracking engine's remaining weak spot, is Pike VM eligible and now measures as linear, not quadratic, from *n* = 10,000 to *n* = 160,000 (0.0018s to 0.0308s, roughly 2x per doubling of *n* -throughout). Relative performance between the two engines on ordinary, -non-adversarial patterns is not yet measured (`concept.md` 7.7/7.8 both say -so plainly); this section will be updated once it is. +throughout). + +`concept.md` 7.9 records three further findings, closing gaps 7.8 had left +open rather than repeating its numbers unchanged. First, the dual-engine +debug mode 7.8 said had not been built was built (an ad hoc whitebox +harness, not a permanent one, 7.9 explains why): 24,000 `match`/ +`fullmatch`/`search` comparisons and roughly 2,700 `finditer` comparisons +between the two engines on the same compiled patterns, zero mismatches. +Second, a literal prefilter was added to the Pike VM's unanchored search +(when a pattern must begin with a specific literal or class, a fresh start +thread that position cannot satisfy is now skipped before being created, +not after), fixing the sharpest instance of the "ordinary patterns got +slower" regression "Benchmarks" below used to report: the literal-search +scenario there went from roughly 70x slower than POSIX `` to +roughly 11x-13x, better than this build's own numbers from before the Pike +VM existed at all (roughly 22x-35x). Third, the Pike VM's own memory cost, +which the design only asserted a bound for, was profiled directly with +Valgrind/Massif on the same adversarial pattern shape used for the +backtracking engine's own worst case (`a*b`, nullable leading loop, no +benefit from the new prefilter): the engine's own contribution beyond +holding the 10MB input itself was roughly 12KB, confirming the design's +`O(instruction count x group count)`, input-length-independent memory bound +actually holds for the v1 implementation, not only on paper. + +Two things remain genuinely open, not resolved by the above: no lazy DFA +state caching or allocation pooling exists yet (`concept.md` 7.7's +remaining deferred items), and whether the Pike VM is faster or slower +than the backtracking engine specifically, as opposed to against POSIX, on +ordinary patterns is still unmeasured. ### Memory footprint @@ -194,12 +220,13 @@ Measured directly with Valgrind/Massif (a real 10MB search) and by watching `VmRSS`/`VmHWM` on real 200MB-1GB files, not estimated from reading the code, for the **backtracking engine**; the Pike VM's own memory cost has a different shape (a bounded number of threads, each carrying its own small -capture array, `concept.md` 7.7) and has not yet been measured the same -detailed way, only checked at a 20MB scale in passing (a peak of roughly -6x the input length across a combined `finditer`/`split`/`sub` run, -`concept.md` 7.8), so the multiplier below is specific to patterns still -running on the backtracking engine (a backreference, lookaround, or atomic -group present), not a claim about every pattern. `Pattern_search`/`finditer`/ +capture array, `concept.md` 7.7) and was profiled the same rigorous way +separately (`concept.md` 7.9, "The Pike VM" above): on the same 10MB +adversarial pattern shape used below, its own contribution beyond the input +itself was roughly 12KB, not a multiplier of the input length at all, so +the multiplier below is specific to patterns still running on the +backtracking engine (a backreference, lookaround, or atomic group present), +not a claim about every pattern. `Pattern_search`/`finditer`/ `split`/`sub` (the operations that try more than one start position) on that engine currently use, at peak, about **8.4x** the input length in memory for a pattern using `OP_REPEAT1` @@ -437,17 +464,21 @@ atomic group, so every one of them runs on the Pike VM ("The Pike VM" above), not the backtracking engine. Reported without adjustment in either direction, on this machine: glibc's -engine is roughly 8x to 70x faster on the three ordinary scenarios -(a literal search, extracting every number from a text, `a*b`), a real, -expected regression from where this build measured previously on those -same three scenarios (then roughly 2x to 25x, running on the backtracking -engine plus its own precomputed skip-ahead tables, before the Pike VM -existed at all): the Pike VM has none of that engine's constant-factor -optimizations yet (no literal prefilter, no lazy DFA state caching, no -allocation pooling, `concept.md` 7.7's explicitly deferred performance -axis; correctness came first, per the instruction this engine was built -under), and glibc's decades of exactly that kind of optimization show up -directly in the gap widening rather than narrowing. The fourth scenario +engine is roughly 7x to 27x faster on the three ordinary scenarios (a +literal search, extracting every number from a text, `a*b`). This narrowed +substantially after a literal prefilter was added to the Pike VM's +unanchored search (`concept.md` 7.9, "The Pike VM" above): the literal +search scenario, initially the worst hit, going from roughly 70x slower +than glibc (regressed from roughly 22x before the Pike VM existed at all) +down to roughly 11x-13x, better than that original pre-Pike-VM number; the +number-extraction scenario improved more modestly (roughly 8x to roughly +7x), since it begins with a class rather than a single literal; the `a*b` +scenario is unchanged, since its leading `a*` is a nullable loop the +prefilter structurally cannot help, an expected, not a missed, case. No +lazy DFA state caching or allocation pooling exists yet +(`concept.md` 7.7's remaining deferred performance items), so a real gap +to glibc's decades of exactly that kind of optimization remains on these +three scenarios, narrower than before but not closed. The fourth scenario inverts entirely: the textbook ReDoS shape `(a+)+b` now runs *faster than glibc*, with no atomic group needed at all, because a Thompson-NFA simulation has no notion of "try one split, then backtrack and try diff --git a/USAGE.md b/USAGE.md index a59b043..88cbe5b 100644 --- a/USAGE.md +++ b/USAGE.md @@ -692,10 +692,11 @@ A pattern with no backreference, lookaround, or atomic group runs on the Pike VM (`README.md` "The Pike VM"), which is genuinely linear time, including on adversarial nested-quantifier shapes like `(a+)+b` that would invite catastrophic backtracking elsewhere, with no atomic group needed; -it does not yet have that engine's own constant-factor optimizations -(no literal prefilter, no lazy DFA state caching), so an ordinary pattern -can currently measure slower in absolute terms than before this engine -existed, a real, documented trade-off, not an oversight. Everything else +it has a literal prefilter (`README.md` "The Pike VM") but still no lazy +DFA state caching or allocation pooling, so an ordinary pattern can +currently measure slower in absolute terms than before this engine +existed, narrower than before the prefilter but a real, documented +trade-off still, not an oversight. Everything else (a backreference, lookaround, or atomic group anywhere) runs on the recursive backtracking engine instead: `match`/`fullmatch` (a single anchored attempt) cost time proportional to that attempt; `search`/ diff --git a/concept.md b/concept.md index 86a95c6..9204742 100644 --- a/concept.md +++ b/concept.md @@ -201,6 +201,16 @@ This section is the concrete plan for a first working implementation of 7.2's re **A result not anticipated by 7.7, which called the relative performance of the two engines "an open, unmeasured question."** It is no longer open for the specific case that motivated writing 7.7 in the first place: `(a+)+b`, README.md's own textbook example of the backtracking engine's remaining weak spot, contains no backreference and is therefore Pike VM eligible, and now measures as genuinely linear, not the "empirically quadratic" figure 7.5 measured for the backtracking engine on the same pattern (that figure remains accurate for what it was actually measuring, the backtracking engine specifically, which still runs this exact pattern shape when it also contains a backreference or another construct that makes it ineligible for this section's engine; the pattern most often used to illustrate the trade-off simply no longer needs that engine at all). Measured directly: search time against *n* non-matching `a` characters went from roughly 2x per doubling of *n* under the backtracking engine (7.5's figure) to also roughly 2x per doubling under the Pike VM, i.e. linear, confirmed from *n* = 10,000 to *n* = 160,000 (0.0018s to 0.0308s). This is a real, measured improvement for every backreference-free pattern, not a projected one, but it is not the whole picture: re-running `examples/bench_vs_posix.c`'s literal search, number extraction, and `a*b` scenarios, none of them adversarial and all of them now Pike VM eligible, found the gap to POSIX `` on those three *widened*, from roughly 2x-25x before this engine existed (running on the backtracking engine plus 7.5's skip-ahead tables) to roughly 8x-70x now. This is the direct, expected cost of 7.7's deferred performance axis actually being deferred: 7.5's tables were themselves a targeted, measured optimization for exactly this kind of pattern, and the Pike VM has no equivalent yet (no literal prefilter, no lazy DFA state caching, no allocation pooling), so an ordinary pattern that used to benefit from that work now runs on a correctness-first engine that does not have it, a real regression for the common case traded for the `(a+)+b`-class fix above, not a net win averaged across all patterns. Whether the Pike VM ends up faster than the backtracking engine specifically (as opposed to against POSIX) on these same ordinary patterns, which would settle whether 7.5's tables are still worth keeping once this engine has its own optimization pass, remains genuinely unmeasured and is the next open question on this axis. +### 7.9 Implementation finding: the dual-engine cross-check, a literal prefilter, and a memory measurement + +Continued work after 7.8 closed three of the items it left open, in the order they were tackled. + +**The dual-engine debug mode 7.8 said was not built was built.** Not as a permanent build-time mode as 7.7 originally specified, but as an ad hoc whitebox harness (`#include "regexx.c"` for direct access to `PatternImpl.has_nfa`, toggled on the same compiled `Pattern` to run `match`/`fullmatch`/`search`/`finditer` through both engines and compare every result field by field), run against 6,000 randomly generated eligible patterns (literals, classes, anchors, quantifiers including lazy and bounded forms, alternation, capturing/non-capturing/named groups, nested combinations) across 20 subjects, all three data modes, several flag combinations: 24,000 `match`/`fullmatch`/`search` comparisons and roughly 2,700 `finditer` comparisons, zero mismatches, clean under AddressSanitizer/UndefinedBehaviorSanitizer. Not committed as a permanent test target, the same judgment 7.8 already made and still holds (the CPython-derived suite continues to be what actually finds and localizes divergences); this was additional, freshly gathered evidence on demand, not a change of strategy. + +**A literal prefilter was added to `pike_find`'s unanchored injection loop**, the single highest leverage item on 7.7/7.8's deferred performance list, chosen specifically because `examples/bench_vs_posix.c`'s literal-search scenario was the one 7.8 found had regressed the most (roughly 2x before this engine existed to roughly 70x after). When the very first instruction of the NFA-only `Prog` is a mandatory `OP_CHAR` or `OP_CLASS` (a pattern beginning with a required literal or class, not a nullable loop or a leading assertion), injecting a fresh start thread at a position that instruction would reject is certain, by construction, to die on the very next `pike_step` call regardless; checking that identical condition before injecting rather than after changes nothing about which threads ever exist, only how much work is spent finding out that most of them will not. This is the same literal-prefilter idea 7.5 already cites for the backtracking engine's own `compute_next_prevmatch`, applied here in its simplest form, a per-position check rather than a precomputed table, deliberately: `README.md` "Implementation status" already documents that a precomputed skip-ahead table traded for a correctness problem once (the reverted size-cap attempt, 7.6), and this narrower, unconditional check carries no equivalent risk, since it can only ever skip work a fuller simulation would also have discarded. Measured directly: the literal-search scenario dropped from roughly 70x slower than POSIX `` to roughly 11x-13x, better than this build's own pre-Pike-VM baseline (roughly 22x-35x, 7.8); the number-extraction scenario (`[0-9]+`, beginning with a mandatory class) improved from roughly 8x to roughly 7x; the `a*b` scenario, whose leading `a*` is a nullable loop the prefilter structurally cannot help, is unchanged, exactly as expected rather than as a gap. Verified against the dual-engine cross-check above (0 mismatches) and the full CPython-derived suite (3,252/3,252) both before and after. + +**The Pike VM's own memory cost, which Section 10's table stated as a design bound without measurement, was profiled with Valgrind/Massif** the same way 7.6 profiled the backtracking engine, on the same adversarial choice of pattern for a fair comparison: `a*b` against 10MB of non-matching `a` characters, the nullable-loop worst case a literal prefilter cannot reduce, so this number is not flattered by the fix above. Peak heap use was 10,497,824 bytes, of which 10,485,761 bytes, 99.89%, is the test harness's own 10MB input buffer, present regardless of which engine runs; the engine's own contribution, compiling the pattern and running the full search, was the remaining 12,063 bytes, an amount that does not grow with input length (the thread lists are cleared and reused every input position, never accumulating) and is small enough, relative to a multi-megabyte input, to be indistinguishable from a constant in practice. This confirms Section 10's `O(instruction count x group count)`, input-length-independent bound for the Pike VM actually holds for the v1 implementation, not only for the design (Section 10's table updated accordingly). + ## 8. Core Data Structures Kept intentionally minimal, all defined in the single file, no dependency beyond the C standard library (`stdint.h`, `stddef.h`, `string.h`) for the data structures in this section specifically. `Input_from_file` (Section 9.2) is the one place the file as a whole steps outside strict ISO C, using POSIX (`mmap`/`open`/`fstat`, Section 7.6), already in the same family of platform dependency `wctype.h`'s locale behavior carries (Section 13.3). @@ -343,7 +353,7 @@ A byte sequence in `UTF8` mode that is not valid UTF-8 at the point the decoder Both figures are stated relative to input length specifically because that is the axis the gigabyte scale requirement constrains; instruction count and group count are properties of the pattern, not the input, and are expected to stay small (tens to low hundreds) for realistically written patterns. -This table describes the two engines as designed. Section 7.5 records that the bounded backtracking engine's v1 implementation, augmented with memoization and precomputed skip-ahead tables, in practice achieves the Regular engine's linear *time* bound for the backreference-free majority of patterns, at the cost of `O(instruction count x input length)` memory rather than the Regular engine's `O(instruction count x group count)`, so the two rows above remain the correct statement of each engine's guarantee, not of what a specific implementation happens to measure. The Regular row's own memory bound, for the Pike VM v1 implementation (7.7/7.8), has not been held to the same standard yet: it is a structural property of the algorithm (a thread list holds at most one thread per instruction, cleared and reused every input position, so its own peak size is bounded by instruction count regardless of input length, independently of whether it is also true in practice), not something separately confirmed with Valgrind/Massif the rigorous way 7.6 confirmed and then corrected the backtracking engine's actual number; a single 20MB check found a peak around 6x the input length for a combined `finditer`/`split`/`sub` run, which is dominated by other things (`MatBuf`, `sub`'s own output buffer, allocator churn from many short-lived thread lists), not necessarily evidence against the thread list bound itself, but this has not been isolated and attributed the way 7.5/7.6 attributed the backtracking engine's cost. Both engines, regardless of this row, still hold the entire input in memory in v1 (7.7's "what is explicitly deferred" paragraph), which this table's "independent of input length" language describes the *engine's own working set* as achieving, not the absence of that separate, already-documented cost. +This table describes the two engines as designed. Section 7.5 records that the bounded backtracking engine's v1 implementation, augmented with memoization and precomputed skip-ahead tables, in practice achieves the Regular engine's linear *time* bound for the backreference-free majority of patterns, at the cost of `O(instruction count x input length)` memory rather than the Regular engine's `O(instruction count x group count)`, so the two rows above remain the correct statement of each engine's guarantee, not of what a specific implementation happens to measure. The Regular row's own memory bound, for the Pike VM v1 implementation, is now held to the same standard: 7.9 records a Valgrind/Massif profile (the same adversarial, nullable-loop pattern shape `a*b` this document uses for the backtracking engine's own worst case) finding the engine's own contribution, beyond holding the input itself, at roughly 12KB for a 10MB search, an amount that does not grow with input length because the thread lists are cleared and reused every input position rather than accumulated, confirming the design bound in this row holds for the v1 implementation, not only in theory. Both engines, regardless of this row, still hold the entire input in memory in v1 (7.7's "what is explicitly deferred" paragraph), which this table's "independent of input length" language describes the *engine's own working set* as achieving, not the absence of that separate, already-documented cost. ## 11. Testing Strategy diff --git a/docs/API.md b/docs/API.md index 8a7c0ac..5c33c57 100644 --- a/docs/API.md +++ b/docs/API.md @@ -145,7 +145,7 @@ Mirror `re.Pattern.match`/`.fullmatch`/`.search` exactly, including the `pos`/`e - `fullmatch`: anchored at `pos`, must also reach exactly `endpos`. - `search`: tries every start position from `pos` to `endpos` inclusive, left to right, and reports the first that admits any match (ordinary backtracking priority decides which match that is at that position, `concept.md` 2.5). -**Performance:** which of two engines runs a given `Pattern` is decided once, at `re_compile` time, and is never a caller's choice (README.md "Implementation status", `concept.md` 7.2/7.7). A pattern with no backreference, lookahead, lookbehind, or atomic group (a possessive quantifier already desugars to the last of these) runs on the Pike VM, a Thompson-NFA simulation with no backtracking at all: every operation in this section is genuinely linear in input length on that engine, including `search` against an adversarial pattern shape like `(a+)+b` that would otherwise invite catastrophic backtracking, with no atomic group needed (README.md "The Pike VM"). Everything else (a backreference anywhere, or a lookaround/atomic construct) runs on the recursive backtracking engine instead: `match`/`fullmatch` there do one anchored attempt at time proportional to that attempt; `search` tries each candidate start position as a separate attempt but, unlike a naive backtracking search, does not redo the same work at every one, since `run_memo` caches every proven failure at the (instruction, position) level for a backreference-free pattern on this engine, and `OP_REPEAT1` additionally uses precomputed skip-ahead tables, together making `search` linear rather than quadratic for a pattern built from simple repeated atoms; a repeat over a *compound* body only gets the failure-memoization, and a pattern with a backreference disables memoization entirely (unsound there, Section 5) and can still be worst-case exponential, exactly as in CPython. The Pike VM currently has none of the backtracking engine's own constant-factor optimizations (no literal prefilter, no lazy DFA state caching), so an ordinary, non-adversarial pattern that happens to be Pike VM eligible can measure slower in absolute terms than the same pattern would have on the backtracking engine, a real, documented trade-off (`concept.md` 7.8), not an oversight. See README.md "Implementation status" and "The Pike VM" for the measured numbers on both engines and citations to the published techniques each uses. +**Performance:** which of two engines runs a given `Pattern` is decided once, at `re_compile` time, and is never a caller's choice (README.md "Implementation status", `concept.md` 7.2/7.7). A pattern with no backreference, lookahead, lookbehind, or atomic group (a possessive quantifier already desugars to the last of these) runs on the Pike VM, a Thompson-NFA simulation with no backtracking at all: every operation in this section is genuinely linear in input length on that engine, including `search` against an adversarial pattern shape like `(a+)+b` that would otherwise invite catastrophic backtracking, with no atomic group needed (README.md "The Pike VM"). Everything else (a backreference anywhere, or a lookaround/atomic construct) runs on the recursive backtracking engine instead: `match`/`fullmatch` there do one anchored attempt at time proportional to that attempt; `search` tries each candidate start position as a separate attempt but, unlike a naive backtracking search, does not redo the same work at every one, since `run_memo` caches every proven failure at the (instruction, position) level for a backreference-free pattern on this engine, and `OP_REPEAT1` additionally uses precomputed skip-ahead tables, together making `search` linear rather than quadratic for a pattern built from simple repeated atoms; a repeat over a *compound* body only gets the failure-memoization, and a pattern with a backreference disables memoization entirely (unsound there, Section 5) and can still be worst-case exponential, exactly as in CPython. The Pike VM has a literal prefilter (a pattern beginning with a mandatory literal or class skips injecting a new search attempt at a position that atom cannot match, `concept.md` 7.9) but still no lazy DFA state caching or allocation pooling, so an ordinary, non-adversarial pattern that happens to be Pike VM eligible can still measure slower in absolute terms than the same pattern would have on the backtracking engine, a real, documented trade-off (`concept.md` 7.8/7.9), narrower than before the prefilter but not closed, not an oversight either way. See README.md "Implementation status" and "The Pike VM" for the measured numbers on both engines and citations to the published techniques each uses. **Memory:** the tables above cost real, measured memory, not just time complexity: `search`/`finditer`/`split`/`sub` on a pattern using `OP_REPEAT1` peak at roughly 8.4x the input length (measured with Valgrind/Massif; `match`/`fullmatch` do not allocate these tables at all and use proportionally less). README.md "Memory footprint" has the full measured breakdown, including a fix that was tried, measured, and deliberately reverted because it traded a fast allocation failure for an effectively-unbounded hang, which is a worse failure mode, not a better one. diff --git a/examples/bench_vs_posix.c b/examples/bench_vs_posix.c index 9fc71b1..28aa4d7 100644 --- a/examples/bench_vs_posix.c +++ b/examples/bench_vs_posix.c @@ -429,21 +429,25 @@ int main(void) { printf("\nWhat this measures and does not measure: regexx is Python `re`\n" "compatible and dispatches every pattern here (all backreference-,\n" "lookaround-, and atomic-group-free) to its Pike VM (README.md \"The\n" - "Pike VM\", concept.md 7.2/7.7), a Thompson-NFA simulation with no\n" - "literal prefilter, lazy DFA state caching, or allocation pooling yet\n" - "(concept.md 7.7's explicitly deferred performance axis); glibc's\n" - " is a mature, DFA-backed engine with decades of that exact\n" - "kind of optimization behind it and a much smaller feature set (no\n" - "named groups, no lazy quantifiers, no lookaround, no atomic groups,\n" - "and, structurally, no way to search past an embedded NUL byte at all,\n" - "docs/API.md Section 1.3 vs examples/binary_scan.c). Scenarios A-C\n" - "measure that missing optimization layer directly: this build was\n" - "previously measured on the same three scenarios, on this engine's\n" - "backtracking predecessor plus its own precomputed skip-ahead tables\n" - "(README.md \"Implementation status\"), at roughly 2-3x less of a gap\n" - "to POSIX than shown here, a real, expected regression for ordinary\n" - "patterns from moving to a v1 engine still missing that layer, not\n" - "a defect introduced by the new engine's correctness.\n\n" + "Pike VM\", concept.md 7.2/7.7/7.9), a Thompson-NFA simulation with a\n" + "literal prefilter (added after this benchmark first found scenario A\n" + "regressed the most, concept.md 7.9) but still no lazy DFA state\n" + "caching or allocation pooling; glibc's is a mature,\n" + "DFA-backed engine with decades of that exact kind of optimization\n" + "behind it and a much smaller feature set (no named groups, no lazy\n" + "quantifiers, no lookaround, no atomic groups, and, structurally, no\n" + "way to search past an embedded NUL byte at all, docs/API.md Section\n" + "1.3 vs examples/binary_scan.c). Scenarios A-C measure what remains of\n" + "that missing optimization layer: this build was previously measured\n" + "on this engine's backtracking predecessor plus its own precomputed\n" + "skip-ahead tables (README.md \"Implementation status\") at roughly\n" + "2x-25x against POSIX on these same three scenarios; the Pike VM\n" + "without a prefilter regressed that to roughly 8x-70x, and the\n" + "prefilter closed most of that back up for A (a plain literal, its\n" + "best case) and part of it for B (starts with a class, not a single\n" + "literal); C's a*b starts with a nullable loop no prefilter can help,\n" + "and is unchanged. What remains (roughly 7x-27x here) is genuinely no\n" + "lazy DFA caching or allocation pooling yet, not a defect.\n\n" "Scenario D inverts entirely, and is the actual point of building this\n" "engine at all: (a+)+b has no backreference, so it is Pike VM eligible,\n" "and a Thompson-NFA simulation has no notion of \"try one split, then\n" diff --git a/regexx.c b/regexx.c index e1ec965..4dd655b 100644 --- a/regexx.c +++ b/regexx.c @@ -1392,12 +1392,40 @@ static int pike_find(Prog *nfa_pr, const uint32_t *text, const uint8_t *text8, i int matched = 0, oom = 0; int64_t sp = pos; - int64_t *seed = malloc((size_t)ncaps * sizeof(int64_t)); - if (!seed) oom = 1; - else { - for (int k = 0; k < ncaps; k++) seed[k] = -1; - seed[0] = pos; - if (!pike_addthread(&clist, &sctx, 0, seed, sp)) oom = 1; + /* Literal prefilter, added after v1 (concept.md 7.9): when the very + * first instruction is a mandatory OP_CHAR or OP_CLASS (a pattern + * that begins with a required literal or class, not a nullable loop + * or a leading assertion), a freshly injected start thread at a + * position that instruction does not accept is certain to die on + * the very next pike_step call, exactly the way an un-prefiltered + * one already does, just after paying for a malloc and a full + * pike_addthread call first. Checking the identical condition + * pike_step's own OP_CHAR/OP_CLASS case checks, before injecting + * rather than after, changes nothing about which threads ever exist + * (a position this skips would have produced a thread that dies + * unobserved on the next step regardless), only how much work is + * spent finding that out; this is the literal prefilter technique + * concept.md 7.5/README.md "Implementation status" already cite for + * the backtracking engine's own compute_next_prevmatch, applied + * here in its simplest form (per position, not precomputed). A + * pattern starting with `^`, a group, an alternation, or a nullable + * repeat gets no benefit and none of the risk: should_inject stays + * 1 and every position is injected exactly as in v1. */ + int first_op = nfa_pr->n > 0 ? nfa_pr->insts[0].op : -1; + Node *first_cls = (first_op == OP_CLASS) ? ((Node **)nfa_pr->classnodes.data)[nfa_pr->insts[0].data] : NULL; +#define PIKE_SHOULD_INJECT(atpos) \ + (first_op == OP_CHAR ? ((atpos) < endpos && char_eq(text_at(&sctx, (atpos)), nfa_pr->insts[0].data, nfa_pr->mode, nfa_pr->flags)) : \ + first_op == OP_CLASS ? ((atpos) < endpos && class_match(first_cls, text_at(&sctx, (atpos)), nfa_pr->mode, nfa_pr->flags)) : \ + 1) + + if (PIKE_SHOULD_INJECT(pos)) { + int64_t *seed = malloc((size_t)ncaps * sizeof(int64_t)); + if (!seed) oom = 1; + else { + for (int k = 0; k < ncaps; k++) seed[k] = -1; + seed[0] = pos; + if (!pike_addthread(&clist, &sctx, 0, seed, sp)) oom = 1; + } } while (!oom) { @@ -1406,7 +1434,7 @@ static int pike_find(Prog *nfa_pr, const uint32_t *text, const uint8_t *text8, i if (!step_ok) { oom = 1; break; } if (sp >= endpos) break; sp++; - if (!anchored && !matched) { + if (!anchored && !matched && PIKE_SHOULD_INJECT(sp)) { int64_t *s2 = malloc((size_t)ncaps * sizeof(int64_t)); if (!s2) { oom = 1; break; } for (int k = 0; k < ncaps; k++) s2[k] = -1; @@ -1431,6 +1459,7 @@ static int pike_find(Prog *nfa_pr, const uint32_t *text, const uint8_t *text8, i * position that would have matched. */ if (clist.count == 0 && (anchored || matched)) break; } +#undef PIKE_SHOULD_INJECT pikelist_free(&clist); pikelist_free(&nlist);