Add a literal prefilter to the Pike VM, profile its memory with Massif
Closed two of the three gaps the previous commit's honest self-review left open (the third, full streaming/bounded-memory input, remains out of scope for this pass and is still documented as such). Literal prefilter (regexx.c, pike_find): when the NFA-only Prog's first instruction is a mandatory OP_CHAR or OP_CLASS, a pattern beginning with a required literal or class rather than a nullable loop or a leading assertion, injecting a fresh unanchored start thread at a position that instruction would reject is certain to die on the very next pike_step call regardless; checking that identical condition before injecting rather than after changes nothing about which threads ever exist, only how much wasted work is done finding out. This targeted exactly examples/bench_vs_posix.c's worst regression: the literal-search scenario went from roughly 70x slower than POSIX <regex.h> (up from roughly 22x before the Pike VM existed) down to roughly 11x-13x, better than the original pre-Pike-VM number; number extraction (starts with a class) improved more modestly; a*b (starts with a nullable loop, structurally unhelped) is unchanged, as expected. Verified with the full 3,252-case suite, three clean AddressSanitizer/UndefinedBehaviorSanitizer passes, and a rerun of the whitebox dual-engine cross-check (24,000 match/fullmatch/search plus ~2,700 finditer comparisons between the two engines on the same compiled patterns, zero mismatches). Memory profiling (concept.md 7.9): Valgrind/Massif on the same adversarial, prefilter-proof pattern shape (a*b, nullable leading loop) used for the backtracking engine's own worst case, for a fair comparison. Peak heap was almost entirely the 10MB input buffer itself; the Pike VM's own contribution was roughly 12KB, confirming the design's O(instruction count x group count), input-length- independent memory bound actually holds for the v1 implementation, not only on paper. concept.md Section 10's table, which had only a "not yet profiled" caveat for this row before, is updated with the measured result. README.md, docs/API.md, USAGE.md, and bench_vs_posix.c's own printed summary are updated throughout with the corrected numbers, rather than left describing the pre-prefilter regression as current. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
This commit is contained in:
@@ -184,9 +184,35 @@ result worth restating plainly here: `(a+)+b`, this document's own running
|
||||
example of the backtracking engine's remaining weak spot, is Pike VM
|
||||
eligible and now measures as linear, not quadratic, from *n* = 10,000 to
|
||||
*n* = 160,000 (0.0018s to 0.0308s, roughly 2x per doubling of *n*
|
||||
throughout). Relative performance between the two engines on ordinary,
|
||||
non-adversarial patterns is not yet measured (`concept.md` 7.7/7.8 both say
|
||||
so plainly); this section will be updated once it is.
|
||||
throughout).
|
||||
|
||||
`concept.md` 7.9 records three further findings, closing gaps 7.8 had left
|
||||
open rather than repeating its numbers unchanged. First, the dual-engine
|
||||
debug mode 7.8 said had not been built was built (an ad hoc whitebox
|
||||
harness, not a permanent one, 7.9 explains why): 24,000 `match`/
|
||||
`fullmatch`/`search` comparisons and roughly 2,700 `finditer` comparisons
|
||||
between the two engines on the same compiled patterns, zero mismatches.
|
||||
Second, a literal prefilter was added to the Pike VM's unanchored search
|
||||
(when a pattern must begin with a specific literal or class, a fresh start
|
||||
thread that position cannot satisfy is now skipped before being created,
|
||||
not after), fixing the sharpest instance of the "ordinary patterns got
|
||||
slower" regression "Benchmarks" below used to report: the literal-search
|
||||
scenario there went from roughly 70x slower than POSIX `<regex.h>` to
|
||||
roughly 11x-13x, better than this build's own numbers from before the Pike
|
||||
VM existed at all (roughly 22x-35x). Third, the Pike VM's own memory cost,
|
||||
which the design only asserted a bound for, was profiled directly with
|
||||
Valgrind/Massif on the same adversarial pattern shape used for the
|
||||
backtracking engine's own worst case (`a*b`, nullable leading loop, no
|
||||
benefit from the new prefilter): the engine's own contribution beyond
|
||||
holding the 10MB input itself was roughly 12KB, confirming the design's
|
||||
`O(instruction count x group count)`, input-length-independent memory bound
|
||||
actually holds for the v1 implementation, not only on paper.
|
||||
|
||||
Two things remain genuinely open, not resolved by the above: no lazy DFA
|
||||
state caching or allocation pooling exists yet (`concept.md` 7.7's
|
||||
remaining deferred items), and whether the Pike VM is faster or slower
|
||||
than the backtracking engine specifically, as opposed to against POSIX, on
|
||||
ordinary patterns is still unmeasured.
|
||||
|
||||
### Memory footprint
|
||||
|
||||
@@ -194,12 +220,13 @@ Measured directly with Valgrind/Massif (a real 10MB search) and by watching
|
||||
`VmRSS`/`VmHWM` on real 200MB-1GB files, not estimated from reading the
|
||||
code, for the **backtracking engine**; the Pike VM's own memory cost has a
|
||||
different shape (a bounded number of threads, each carrying its own small
|
||||
capture array, `concept.md` 7.7) and has not yet been measured the same
|
||||
detailed way, only checked at a 20MB scale in passing (a peak of roughly
|
||||
6x the input length across a combined `finditer`/`split`/`sub` run,
|
||||
`concept.md` 7.8), so the multiplier below is specific to patterns still
|
||||
running on the backtracking engine (a backreference, lookaround, or atomic
|
||||
group present), not a claim about every pattern. `Pattern_search`/`finditer`/
|
||||
capture array, `concept.md` 7.7) and was profiled the same rigorous way
|
||||
separately (`concept.md` 7.9, "The Pike VM" above): on the same 10MB
|
||||
adversarial pattern shape used below, its own contribution beyond the input
|
||||
itself was roughly 12KB, not a multiplier of the input length at all, so
|
||||
the multiplier below is specific to patterns still running on the
|
||||
backtracking engine (a backreference, lookaround, or atomic group present),
|
||||
not a claim about every pattern. `Pattern_search`/`finditer`/
|
||||
`split`/`sub` (the operations that try more than one start position) on
|
||||
that engine currently use, at peak, about **8.4x** the input length in
|
||||
memory for a pattern using `OP_REPEAT1`
|
||||
@@ -437,17 +464,21 @@ atomic group, so every one of them runs on the Pike VM ("The Pike VM"
|
||||
above), not the backtracking engine.
|
||||
|
||||
Reported without adjustment in either direction, on this machine: glibc's
|
||||
engine is roughly 8x to 70x faster on the three ordinary scenarios
|
||||
(a literal search, extracting every number from a text, `a*b`), a real,
|
||||
expected regression from where this build measured previously on those
|
||||
same three scenarios (then roughly 2x to 25x, running on the backtracking
|
||||
engine plus its own precomputed skip-ahead tables, before the Pike VM
|
||||
existed at all): the Pike VM has none of that engine's constant-factor
|
||||
optimizations yet (no literal prefilter, no lazy DFA state caching, no
|
||||
allocation pooling, `concept.md` 7.7's explicitly deferred performance
|
||||
axis; correctness came first, per the instruction this engine was built
|
||||
under), and glibc's decades of exactly that kind of optimization show up
|
||||
directly in the gap widening rather than narrowing. The fourth scenario
|
||||
engine is roughly 7x to 27x faster on the three ordinary scenarios (a
|
||||
literal search, extracting every number from a text, `a*b`). This narrowed
|
||||
substantially after a literal prefilter was added to the Pike VM's
|
||||
unanchored search (`concept.md` 7.9, "The Pike VM" above): the literal
|
||||
search scenario, initially the worst hit, going from roughly 70x slower
|
||||
than glibc (regressed from roughly 22x before the Pike VM existed at all)
|
||||
down to roughly 11x-13x, better than that original pre-Pike-VM number; the
|
||||
number-extraction scenario improved more modestly (roughly 8x to roughly
|
||||
7x), since it begins with a class rather than a single literal; the `a*b`
|
||||
scenario is unchanged, since its leading `a*` is a nullable loop the
|
||||
prefilter structurally cannot help, an expected, not a missed, case. No
|
||||
lazy DFA state caching or allocation pooling exists yet
|
||||
(`concept.md` 7.7's remaining deferred performance items), so a real gap
|
||||
to glibc's decades of exactly that kind of optimization remains on these
|
||||
three scenarios, narrower than before but not closed. The fourth scenario
|
||||
inverts entirely: the textbook ReDoS shape `(a+)+b` now runs *faster than
|
||||
glibc*, with no atomic group needed at all, because a Thompson-NFA
|
||||
simulation has no notion of "try one split, then backtrack and try
|
||||
|
||||
Reference in New Issue
Block a user