Add a literal prefilter to the Pike VM, profile its memory with Massif

Closed two of the three gaps the previous commit's honest self-review
left open (the third, full streaming/bounded-memory input, remains
out of scope for this pass and is still documented as such).

Literal prefilter (regexx.c, pike_find): when the NFA-only Prog's
first instruction is a mandatory OP_CHAR or OP_CLASS, a pattern
beginning with a required literal or class rather than a nullable
loop or a leading assertion, injecting a fresh unanchored start
thread at a position that instruction would reject is certain to die
on the very next pike_step call regardless; checking that identical
condition before injecting rather than after changes nothing about
which threads ever exist, only how much wasted work is done finding
out. This targeted exactly examples/bench_vs_posix.c's worst
regression: the literal-search scenario went from roughly 70x slower
than POSIX <regex.h> (up from roughly 22x before the Pike VM existed)
down to roughly 11x-13x, better than the original pre-Pike-VM number;
number extraction (starts with a class) improved more modestly; a*b
(starts with a nullable loop, structurally unhelped) is unchanged, as
expected. Verified with the full 3,252-case suite, three clean
AddressSanitizer/UndefinedBehaviorSanitizer passes, and a rerun of
the whitebox dual-engine cross-check (24,000 match/fullmatch/search
plus ~2,700 finditer comparisons between the two engines on the same
compiled patterns, zero mismatches).

Memory profiling (concept.md 7.9): Valgrind/Massif on the same
adversarial, prefilter-proof pattern shape (a*b, nullable leading
loop) used for the backtracking engine's own worst case, for a fair
comparison. Peak heap was almost entirely the 10MB input buffer
itself; the Pike VM's own contribution was roughly 12KB, confirming
the design's O(instruction count x group count), input-length-
independent memory bound actually holds for the v1 implementation,
not only on paper. concept.md Section 10's table, which had only a
"not yet profiled" caveat for this row before, is updated with the
measured result.

README.md, docs/API.md, USAGE.md, and bench_vs_posix.c's own printed
summary are updated throughout with the corrected numbers, rather
than left describing the pre-prefilter regression as current.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
This commit is contained in:
2026-09-14 12:21:15 +00:00
co-authored by Claude Sonnet 5
parent f54cda3311
commit 5ae68d4ea6
6 changed files with 123 additions and 48 deletions
+51 -20
View File
@@ -184,9 +184,35 @@ result worth restating plainly here: `(a+)+b`, this document's own running
example of the backtracking engine's remaining weak spot, is Pike VM
eligible and now measures as linear, not quadratic, from *n* = 10,000 to
*n* = 160,000 (0.0018s to 0.0308s, roughly 2x per doubling of *n*
throughout). Relative performance between the two engines on ordinary,
non-adversarial patterns is not yet measured (`concept.md` 7.7/7.8 both say
so plainly); this section will be updated once it is.
throughout).
`concept.md` 7.9 records three further findings, closing gaps 7.8 had left
open rather than repeating its numbers unchanged. First, the dual-engine
debug mode 7.8 said had not been built was built (an ad hoc whitebox
harness, not a permanent one, 7.9 explains why): 24,000 `match`/
`fullmatch`/`search` comparisons and roughly 2,700 `finditer` comparisons
between the two engines on the same compiled patterns, zero mismatches.
Second, a literal prefilter was added to the Pike VM's unanchored search
(when a pattern must begin with a specific literal or class, a fresh start
thread that position cannot satisfy is now skipped before being created,
not after), fixing the sharpest instance of the "ordinary patterns got
slower" regression "Benchmarks" below used to report: the literal-search
scenario there went from roughly 70x slower than POSIX `<regex.h>` to
roughly 11x-13x, better than this build's own numbers from before the Pike
VM existed at all (roughly 22x-35x). Third, the Pike VM's own memory cost,
which the design only asserted a bound for, was profiled directly with
Valgrind/Massif on the same adversarial pattern shape used for the
backtracking engine's own worst case (`a*b`, nullable leading loop, no
benefit from the new prefilter): the engine's own contribution beyond
holding the 10MB input itself was roughly 12KB, confirming the design's
`O(instruction count x group count)`, input-length-independent memory bound
actually holds for the v1 implementation, not only on paper.
Two things remain genuinely open, not resolved by the above: no lazy DFA
state caching or allocation pooling exists yet (`concept.md` 7.7's
remaining deferred items), and whether the Pike VM is faster or slower
than the backtracking engine specifically, as opposed to against POSIX, on
ordinary patterns is still unmeasured.
### Memory footprint
@@ -194,12 +220,13 @@ Measured directly with Valgrind/Massif (a real 10MB search) and by watching
`VmRSS`/`VmHWM` on real 200MB-1GB files, not estimated from reading the
code, for the **backtracking engine**; the Pike VM's own memory cost has a
different shape (a bounded number of threads, each carrying its own small
capture array, `concept.md` 7.7) and has not yet been measured the same
detailed way, only checked at a 20MB scale in passing (a peak of roughly
6x the input length across a combined `finditer`/`split`/`sub` run,
`concept.md` 7.8), so the multiplier below is specific to patterns still
running on the backtracking engine (a backreference, lookaround, or atomic
group present), not a claim about every pattern. `Pattern_search`/`finditer`/
capture array, `concept.md` 7.7) and was profiled the same rigorous way
separately (`concept.md` 7.9, "The Pike VM" above): on the same 10MB
adversarial pattern shape used below, its own contribution beyond the input
itself was roughly 12KB, not a multiplier of the input length at all, so
the multiplier below is specific to patterns still running on the
backtracking engine (a backreference, lookaround, or atomic group present),
not a claim about every pattern. `Pattern_search`/`finditer`/
`split`/`sub` (the operations that try more than one start position) on
that engine currently use, at peak, about **8.4x** the input length in
memory for a pattern using `OP_REPEAT1`
@@ -437,17 +464,21 @@ atomic group, so every one of them runs on the Pike VM ("The Pike VM"
above), not the backtracking engine.
Reported without adjustment in either direction, on this machine: glibc's
engine is roughly 8x to 70x faster on the three ordinary scenarios
(a literal search, extracting every number from a text, `a*b`), a real,
expected regression from where this build measured previously on those
same three scenarios (then roughly 2x to 25x, running on the backtracking
engine plus its own precomputed skip-ahead tables, before the Pike VM
existed at all): the Pike VM has none of that engine's constant-factor
optimizations yet (no literal prefilter, no lazy DFA state caching, no
allocation pooling, `concept.md` 7.7's explicitly deferred performance
axis; correctness came first, per the instruction this engine was built
under), and glibc's decades of exactly that kind of optimization show up
directly in the gap widening rather than narrowing. The fourth scenario
engine is roughly 7x to 27x faster on the three ordinary scenarios (a
literal search, extracting every number from a text, `a*b`). This narrowed
substantially after a literal prefilter was added to the Pike VM's
unanchored search (`concept.md` 7.9, "The Pike VM" above): the literal
search scenario, initially the worst hit, going from roughly 70x slower
than glibc (regressed from roughly 22x before the Pike VM existed at all)
down to roughly 11x-13x, better than that original pre-Pike-VM number; the
number-extraction scenario improved more modestly (roughly 8x to roughly
7x), since it begins with a class rather than a single literal; the `a*b`
scenario is unchanged, since its leading `a*` is a nullable loop the
prefilter structurally cannot help, an expected, not a missed, case. No
lazy DFA state caching or allocation pooling exists yet
(`concept.md` 7.7's remaining deferred performance items), so a real gap
to glibc's decades of exactly that kind of optimization remains on these
three scenarios, narrower than before but not closed. The fourth scenario
inverts entirely: the textbook ReDoS shape `(a+)+b` now runs *faster than
glibc*, with no atomic group needed at all, because a Thompson-NFA
simulation has no notion of "try one split, then backtrack and try
+5 -4
View File
@@ -692,10 +692,11 @@ A pattern with no backreference, lookaround, or atomic group runs on the
Pike VM (`README.md` "The Pike VM"), which is genuinely linear time,
including on adversarial nested-quantifier shapes like `(a+)+b` that would
invite catastrophic backtracking elsewhere, with no atomic group needed;
it does not yet have that engine's own constant-factor optimizations
(no literal prefilter, no lazy DFA state caching), so an ordinary pattern
can currently measure slower in absolute terms than before this engine
existed, a real, documented trade-off, not an oversight. Everything else
it has a literal prefilter (`README.md` "The Pike VM") but still no lazy
DFA state caching or allocation pooling, so an ordinary pattern can
currently measure slower in absolute terms than before this engine
existed, narrower than before the prefilter but a real, documented
trade-off still, not an oversight. Everything else
(a backreference, lookaround, or atomic group anywhere) runs on the
recursive backtracking engine instead: `match`/`fullmatch` (a single
anchored attempt) cost time proportional to that attempt; `search`/
+11 -1
View File
@@ -201,6 +201,16 @@ This section is the concrete plan for a first working implementation of 7.2's re
**A result not anticipated by 7.7, which called the relative performance of the two engines "an open, unmeasured question."** It is no longer open for the specific case that motivated writing 7.7 in the first place: `(a+)+b`, README.md's own textbook example of the backtracking engine's remaining weak spot, contains no backreference and is therefore Pike VM eligible, and now measures as genuinely linear, not the "empirically quadratic" figure 7.5 measured for the backtracking engine on the same pattern (that figure remains accurate for what it was actually measuring, the backtracking engine specifically, which still runs this exact pattern shape when it also contains a backreference or another construct that makes it ineligible for this section's engine; the pattern most often used to illustrate the trade-off simply no longer needs that engine at all). Measured directly: search time against *n* non-matching `a` characters went from roughly 2x per doubling of *n* under the backtracking engine (7.5's figure) to also roughly 2x per doubling under the Pike VM, i.e. linear, confirmed from *n* = 10,000 to *n* = 160,000 (0.0018s to 0.0308s). This is a real, measured improvement for every backreference-free pattern, not a projected one, but it is not the whole picture: re-running `examples/bench_vs_posix.c`'s literal search, number extraction, and `a*b` scenarios, none of them adversarial and all of them now Pike VM eligible, found the gap to POSIX `<regex.h>` on those three *widened*, from roughly 2x-25x before this engine existed (running on the backtracking engine plus 7.5's skip-ahead tables) to roughly 8x-70x now. This is the direct, expected cost of 7.7's deferred performance axis actually being deferred: 7.5's tables were themselves a targeted, measured optimization for exactly this kind of pattern, and the Pike VM has no equivalent yet (no literal prefilter, no lazy DFA state caching, no allocation pooling), so an ordinary pattern that used to benefit from that work now runs on a correctness-first engine that does not have it, a real regression for the common case traded for the `(a+)+b`-class fix above, not a net win averaged across all patterns. Whether the Pike VM ends up faster than the backtracking engine specifically (as opposed to against POSIX) on these same ordinary patterns, which would settle whether 7.5's tables are still worth keeping once this engine has its own optimization pass, remains genuinely unmeasured and is the next open question on this axis.
### 7.9 Implementation finding: the dual-engine cross-check, a literal prefilter, and a memory measurement
Continued work after 7.8 closed three of the items it left open, in the order they were tackled.
**The dual-engine debug mode 7.8 said was not built was built.** Not as a permanent build-time mode as 7.7 originally specified, but as an ad hoc whitebox harness (`#include "regexx.c"` for direct access to `PatternImpl.has_nfa`, toggled on the same compiled `Pattern` to run `match`/`fullmatch`/`search`/`finditer` through both engines and compare every result field by field), run against 6,000 randomly generated eligible patterns (literals, classes, anchors, quantifiers including lazy and bounded forms, alternation, capturing/non-capturing/named groups, nested combinations) across 20 subjects, all three data modes, several flag combinations: 24,000 `match`/`fullmatch`/`search` comparisons and roughly 2,700 `finditer` comparisons, zero mismatches, clean under AddressSanitizer/UndefinedBehaviorSanitizer. Not committed as a permanent test target, the same judgment 7.8 already made and still holds (the CPython-derived suite continues to be what actually finds and localizes divergences); this was additional, freshly gathered evidence on demand, not a change of strategy.
**A literal prefilter was added to `pike_find`'s unanchored injection loop**, the single highest leverage item on 7.7/7.8's deferred performance list, chosen specifically because `examples/bench_vs_posix.c`'s literal-search scenario was the one 7.8 found had regressed the most (roughly 2x before this engine existed to roughly 70x after). When the very first instruction of the NFA-only `Prog` is a mandatory `OP_CHAR` or `OP_CLASS` (a pattern beginning with a required literal or class, not a nullable loop or a leading assertion), injecting a fresh start thread at a position that instruction would reject is certain, by construction, to die on the very next `pike_step` call regardless; checking that identical condition before injecting rather than after changes nothing about which threads ever exist, only how much work is spent finding out that most of them will not. This is the same literal-prefilter idea 7.5 already cites for the backtracking engine's own `compute_next_prevmatch`, applied here in its simplest form, a per-position check rather than a precomputed table, deliberately: `README.md` "Implementation status" already documents that a precomputed skip-ahead table traded for a correctness problem once (the reverted size-cap attempt, 7.6), and this narrower, unconditional check carries no equivalent risk, since it can only ever skip work a fuller simulation would also have discarded. Measured directly: the literal-search scenario dropped from roughly 70x slower than POSIX `<regex.h>` to roughly 11x-13x, better than this build's own pre-Pike-VM baseline (roughly 22x-35x, 7.8); the number-extraction scenario (`[0-9]+`, beginning with a mandatory class) improved from roughly 8x to roughly 7x; the `a*b` scenario, whose leading `a*` is a nullable loop the prefilter structurally cannot help, is unchanged, exactly as expected rather than as a gap. Verified against the dual-engine cross-check above (0 mismatches) and the full CPython-derived suite (3,252/3,252) both before and after.
**The Pike VM's own memory cost, which Section 10's table stated as a design bound without measurement, was profiled with Valgrind/Massif** the same way 7.6 profiled the backtracking engine, on the same adversarial choice of pattern for a fair comparison: `a*b` against 10MB of non-matching `a` characters, the nullable-loop worst case a literal prefilter cannot reduce, so this number is not flattered by the fix above. Peak heap use was 10,497,824 bytes, of which 10,485,761 bytes, 99.89%, is the test harness's own 10MB input buffer, present regardless of which engine runs; the engine's own contribution, compiling the pattern and running the full search, was the remaining 12,063 bytes, an amount that does not grow with input length (the thread lists are cleared and reused every input position, never accumulating) and is small enough, relative to a multi-megabyte input, to be indistinguishable from a constant in practice. This confirms Section 10's `O(instruction count x group count)`, input-length-independent bound for the Pike VM actually holds for the v1 implementation, not only for the design (Section 10's table updated accordingly).
## 8. Core Data Structures
Kept intentionally minimal, all defined in the single file, no dependency beyond the C standard library (`stdint.h`, `stddef.h`, `string.h`) for the data structures in this section specifically. `Input_from_file` (Section 9.2) is the one place the file as a whole steps outside strict ISO C, using POSIX (`mmap`/`open`/`fstat`, Section 7.6), already in the same family of platform dependency `wctype.h`'s locale behavior carries (Section 13.3).
@@ -343,7 +353,7 @@ A byte sequence in `UTF8` mode that is not valid UTF-8 at the point the decoder
Both figures are stated relative to input length specifically because that is the axis the gigabyte scale requirement constrains; instruction count and group count are properties of the pattern, not the input, and are expected to stay small (tens to low hundreds) for realistically written patterns.
This table describes the two engines as designed. Section 7.5 records that the bounded backtracking engine's v1 implementation, augmented with memoization and precomputed skip-ahead tables, in practice achieves the Regular engine's linear *time* bound for the backreference-free majority of patterns, at the cost of `O(instruction count x input length)` memory rather than the Regular engine's `O(instruction count x group count)`, so the two rows above remain the correct statement of each engine's guarantee, not of what a specific implementation happens to measure. The Regular row's own memory bound, for the Pike VM v1 implementation (7.7/7.8), has not been held to the same standard yet: it is a structural property of the algorithm (a thread list holds at most one thread per instruction, cleared and reused every input position, so its own peak size is bounded by instruction count regardless of input length, independently of whether it is also true in practice), not something separately confirmed with Valgrind/Massif the rigorous way 7.6 confirmed and then corrected the backtracking engine's actual number; a single 20MB check found a peak around 6x the input length for a combined `finditer`/`split`/`sub` run, which is dominated by other things (`MatBuf`, `sub`'s own output buffer, allocator churn from many short-lived thread lists), not necessarily evidence against the thread list bound itself, but this has not been isolated and attributed the way 7.5/7.6 attributed the backtracking engine's cost. Both engines, regardless of this row, still hold the entire input in memory in v1 (7.7's "what is explicitly deferred" paragraph), which this table's "independent of input length" language describes the *engine's own working set* as achieving, not the absence of that separate, already-documented cost.
This table describes the two engines as designed. Section 7.5 records that the bounded backtracking engine's v1 implementation, augmented with memoization and precomputed skip-ahead tables, in practice achieves the Regular engine's linear *time* bound for the backreference-free majority of patterns, at the cost of `O(instruction count x input length)` memory rather than the Regular engine's `O(instruction count x group count)`, so the two rows above remain the correct statement of each engine's guarantee, not of what a specific implementation happens to measure. The Regular row's own memory bound, for the Pike VM v1 implementation, is now held to the same standard: 7.9 records a Valgrind/Massif profile (the same adversarial, nullable-loop pattern shape `a*b` this document uses for the backtracking engine's own worst case) finding the engine's own contribution, beyond holding the input itself, at roughly 12KB for a 10MB search, an amount that does not grow with input length because the thread lists are cleared and reused every input position rather than accumulated, confirming the design bound in this row holds for the v1 implementation, not only in theory. Both engines, regardless of this row, still hold the entire input in memory in v1 (7.7's "what is explicitly deferred" paragraph), which this table's "independent of input length" language describes the *engine's own working set* as achieving, not the absence of that separate, already-documented cost.
## 11. Testing Strategy
+1 -1
View File
@@ -145,7 +145,7 @@ Mirror `re.Pattern.match`/`.fullmatch`/`.search` exactly, including the `pos`/`e
- `fullmatch`: anchored at `pos`, must also reach exactly `endpos`.
- `search`: tries every start position from `pos` to `endpos` inclusive, left to right, and reports the first that admits any match (ordinary backtracking priority decides which match that is at that position, `concept.md` 2.5).
**Performance:** which of two engines runs a given `Pattern` is decided once, at `re_compile` time, and is never a caller's choice (README.md "Implementation status", `concept.md` 7.2/7.7). A pattern with no backreference, lookahead, lookbehind, or atomic group (a possessive quantifier already desugars to the last of these) runs on the Pike VM, a Thompson-NFA simulation with no backtracking at all: every operation in this section is genuinely linear in input length on that engine, including `search` against an adversarial pattern shape like `(a+)+b` that would otherwise invite catastrophic backtracking, with no atomic group needed (README.md "The Pike VM"). Everything else (a backreference anywhere, or a lookaround/atomic construct) runs on the recursive backtracking engine instead: `match`/`fullmatch` there do one anchored attempt at time proportional to that attempt; `search` tries each candidate start position as a separate attempt but, unlike a naive backtracking search, does not redo the same work at every one, since `run_memo` caches every proven failure at the (instruction, position) level for a backreference-free pattern on this engine, and `OP_REPEAT1` additionally uses precomputed skip-ahead tables, together making `search` linear rather than quadratic for a pattern built from simple repeated atoms; a repeat over a *compound* body only gets the failure-memoization, and a pattern with a backreference disables memoization entirely (unsound there, Section 5) and can still be worst-case exponential, exactly as in CPython. The Pike VM currently has none of the backtracking engine's own constant-factor optimizations (no literal prefilter, no lazy DFA state caching), so an ordinary, non-adversarial pattern that happens to be Pike VM eligible can measure slower in absolute terms than the same pattern would have on the backtracking engine, a real, documented trade-off (`concept.md` 7.8), not an oversight. See README.md "Implementation status" and "The Pike VM" for the measured numbers on both engines and citations to the published techniques each uses.
**Performance:** which of two engines runs a given `Pattern` is decided once, at `re_compile` time, and is never a caller's choice (README.md "Implementation status", `concept.md` 7.2/7.7). A pattern with no backreference, lookahead, lookbehind, or atomic group (a possessive quantifier already desugars to the last of these) runs on the Pike VM, a Thompson-NFA simulation with no backtracking at all: every operation in this section is genuinely linear in input length on that engine, including `search` against an adversarial pattern shape like `(a+)+b` that would otherwise invite catastrophic backtracking, with no atomic group needed (README.md "The Pike VM"). Everything else (a backreference anywhere, or a lookaround/atomic construct) runs on the recursive backtracking engine instead: `match`/`fullmatch` there do one anchored attempt at time proportional to that attempt; `search` tries each candidate start position as a separate attempt but, unlike a naive backtracking search, does not redo the same work at every one, since `run_memo` caches every proven failure at the (instruction, position) level for a backreference-free pattern on this engine, and `OP_REPEAT1` additionally uses precomputed skip-ahead tables, together making `search` linear rather than quadratic for a pattern built from simple repeated atoms; a repeat over a *compound* body only gets the failure-memoization, and a pattern with a backreference disables memoization entirely (unsound there, Section 5) and can still be worst-case exponential, exactly as in CPython. The Pike VM has a literal prefilter (a pattern beginning with a mandatory literal or class skips injecting a new search attempt at a position that atom cannot match, `concept.md` 7.9) but still no lazy DFA state caching or allocation pooling, so an ordinary, non-adversarial pattern that happens to be Pike VM eligible can still measure slower in absolute terms than the same pattern would have on the backtracking engine, a real, documented trade-off (`concept.md` 7.8/7.9), narrower than before the prefilter but not closed, not an oversight either way. See README.md "Implementation status" and "The Pike VM" for the measured numbers on both engines and citations to the published techniques each uses.
**Memory:** the tables above cost real, measured memory, not just time complexity: `search`/`finditer`/`split`/`sub` on a pattern using `OP_REPEAT1` peak at roughly 8.4x the input length (measured with Valgrind/Massif; `match`/`fullmatch` do not allocate these tables at all and use proportionally less). README.md "Memory footprint" has the full measured breakdown, including a fix that was tried, measured, and deliberately reverted because it traded a fast allocation failure for an effectively-unbounded hang, which is a worse failure mode, not a better one.
+19 -15
View File
@@ -429,21 +429,25 @@ int main(void) {
printf("\nWhat this measures and does not measure: regexx is Python `re`\n"
"compatible and dispatches every pattern here (all backreference-,\n"
"lookaround-, and atomic-group-free) to its Pike VM (README.md \"The\n"
"Pike VM\", concept.md 7.2/7.7), a Thompson-NFA simulation with no\n"
"literal prefilter, lazy DFA state caching, or allocation pooling yet\n"
"(concept.md 7.7's explicitly deferred performance axis); glibc's\n"
"<regex.h> is a mature, DFA-backed engine with decades of that exact\n"
"kind of optimization behind it and a much smaller feature set (no\n"
"named groups, no lazy quantifiers, no lookaround, no atomic groups,\n"
"and, structurally, no way to search past an embedded NUL byte at all,\n"
"docs/API.md Section 1.3 vs examples/binary_scan.c). Scenarios A-C\n"
"measure that missing optimization layer directly: this build was\n"
"previously measured on the same three scenarios, on this engine's\n"
"backtracking predecessor plus its own precomputed skip-ahead tables\n"
"(README.md \"Implementation status\"), at roughly 2-3x less of a gap\n"
"to POSIX than shown here, a real, expected regression for ordinary\n"
"patterns from moving to a v1 engine still missing that layer, not\n"
"a defect introduced by the new engine's correctness.\n\n"
"Pike VM\", concept.md 7.2/7.7/7.9), a Thompson-NFA simulation with a\n"
"literal prefilter (added after this benchmark first found scenario A\n"
"regressed the most, concept.md 7.9) but still no lazy DFA state\n"
"caching or allocation pooling; glibc's <regex.h> is a mature,\n"
"DFA-backed engine with decades of that exact kind of optimization\n"
"behind it and a much smaller feature set (no named groups, no lazy\n"
"quantifiers, no lookaround, no atomic groups, and, structurally, no\n"
"way to search past an embedded NUL byte at all, docs/API.md Section\n"
"1.3 vs examples/binary_scan.c). Scenarios A-C measure what remains of\n"
"that missing optimization layer: this build was previously measured\n"
"on this engine's backtracking predecessor plus its own precomputed\n"
"skip-ahead tables (README.md \"Implementation status\") at roughly\n"
"2x-25x against POSIX on these same three scenarios; the Pike VM\n"
"without a prefilter regressed that to roughly 8x-70x, and the\n"
"prefilter closed most of that back up for A (a plain literal, its\n"
"best case) and part of it for B (starts with a class, not a single\n"
"literal); C's a*b starts with a nullable loop no prefilter can help,\n"
"and is unchanged. What remains (roughly 7x-27x here) is genuinely no\n"
"lazy DFA caching or allocation pooling yet, not a defect.\n\n"
"Scenario D inverts entirely, and is the actual point of building this\n"
"engine at all: (a+)+b has no backreference, so it is Pike VM eligible,\n"
"and a Thompson-NFA simulation has no notion of \"try one split, then\n"
+36 -7
View File
@@ -1392,12 +1392,40 @@ static int pike_find(Prog *nfa_pr, const uint32_t *text, const uint8_t *text8, i
int matched = 0, oom = 0;
int64_t sp = pos;
int64_t *seed = malloc((size_t)ncaps * sizeof(int64_t));
if (!seed) oom = 1;
else {
for (int k = 0; k < ncaps; k++) seed[k] = -1;
seed[0] = pos;
if (!pike_addthread(&clist, &sctx, 0, seed, sp)) oom = 1;
/* Literal prefilter, added after v1 (concept.md 7.9): when the very
* first instruction is a mandatory OP_CHAR or OP_CLASS (a pattern
* that begins with a required literal or class, not a nullable loop
* or a leading assertion), a freshly injected start thread at a
* position that instruction does not accept is certain to die on
* the very next pike_step call, exactly the way an un-prefiltered
* one already does, just after paying for a malloc and a full
* pike_addthread call first. Checking the identical condition
* pike_step's own OP_CHAR/OP_CLASS case checks, before injecting
* rather than after, changes nothing about which threads ever exist
* (a position this skips would have produced a thread that dies
* unobserved on the next step regardless), only how much work is
* spent finding that out; this is the literal prefilter technique
* concept.md 7.5/README.md "Implementation status" already cite for
* the backtracking engine's own compute_next_prevmatch, applied
* here in its simplest form (per position, not precomputed). A
* pattern starting with `^`, a group, an alternation, or a nullable
* repeat gets no benefit and none of the risk: should_inject stays
* 1 and every position is injected exactly as in v1. */
int first_op = nfa_pr->n > 0 ? nfa_pr->insts[0].op : -1;
Node *first_cls = (first_op == OP_CLASS) ? ((Node **)nfa_pr->classnodes.data)[nfa_pr->insts[0].data] : NULL;
#define PIKE_SHOULD_INJECT(atpos) \
(first_op == OP_CHAR ? ((atpos) < endpos && char_eq(text_at(&sctx, (atpos)), nfa_pr->insts[0].data, nfa_pr->mode, nfa_pr->flags)) : \
first_op == OP_CLASS ? ((atpos) < endpos && class_match(first_cls, text_at(&sctx, (atpos)), nfa_pr->mode, nfa_pr->flags)) : \
1)
if (PIKE_SHOULD_INJECT(pos)) {
int64_t *seed = malloc((size_t)ncaps * sizeof(int64_t));
if (!seed) oom = 1;
else {
for (int k = 0; k < ncaps; k++) seed[k] = -1;
seed[0] = pos;
if (!pike_addthread(&clist, &sctx, 0, seed, sp)) oom = 1;
}
}
while (!oom) {
@@ -1406,7 +1434,7 @@ static int pike_find(Prog *nfa_pr, const uint32_t *text, const uint8_t *text8, i
if (!step_ok) { oom = 1; break; }
if (sp >= endpos) break;
sp++;
if (!anchored && !matched) {
if (!anchored && !matched && PIKE_SHOULD_INJECT(sp)) {
int64_t *s2 = malloc((size_t)ncaps * sizeof(int64_t));
if (!s2) { oom = 1; break; }
for (int k = 0; k < ncaps; k++) s2[k] = -1;
@@ -1431,6 +1459,7 @@ static int pike_find(Prog *nfa_pr, const uint32_t *text, const uint8_t *text8, i
* position that would have matched. */
if (clist.count == 0 && (anchored || matched)) break;
}
#undef PIKE_SHOULD_INJECT
pikelist_free(&clist);
pikelist_free(&nlist);