Add a literal prefilter to the Pike VM, profile its memory with Massif

Closed two of the three gaps the previous commit's honest self-review
left open (the third, full streaming/bounded-memory input, remains
out of scope for this pass and is still documented as such).

Literal prefilter (regexx.c, pike_find): when the NFA-only Prog's
first instruction is a mandatory OP_CHAR or OP_CLASS, a pattern
beginning with a required literal or class rather than a nullable
loop or a leading assertion, injecting a fresh unanchored start
thread at a position that instruction would reject is certain to die
on the very next pike_step call regardless; checking that identical
condition before injecting rather than after changes nothing about
which threads ever exist, only how much wasted work is done finding
out. This targeted exactly examples/bench_vs_posix.c's worst
regression: the literal-search scenario went from roughly 70x slower
than POSIX <regex.h> (up from roughly 22x before the Pike VM existed)
down to roughly 11x-13x, better than the original pre-Pike-VM number;
number extraction (starts with a class) improved more modestly; a*b
(starts with a nullable loop, structurally unhelped) is unchanged, as
expected. Verified with the full 3,252-case suite, three clean
AddressSanitizer/UndefinedBehaviorSanitizer passes, and a rerun of
the whitebox dual-engine cross-check (24,000 match/fullmatch/search
plus ~2,700 finditer comparisons between the two engines on the same
compiled patterns, zero mismatches).

Memory profiling (concept.md 7.9): Valgrind/Massif on the same
adversarial, prefilter-proof pattern shape (a*b, nullable leading
loop) used for the backtracking engine's own worst case, for a fair
comparison. Peak heap was almost entirely the 10MB input buffer
itself; the Pike VM's own contribution was roughly 12KB, confirming
the design's O(instruction count x group count), input-length-
independent memory bound actually holds for the v1 implementation,
not only on paper. concept.md Section 10's table, which had only a
"not yet profiled" caveat for this row before, is updated with the
measured result.

README.md, docs/API.md, USAGE.md, and bench_vs_posix.c's own printed
summary are updated throughout with the corrected numbers, rather
than left describing the pre-prefilter regression as current.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
This commit is contained in:
2026-09-14 12:21:15 +00:00
co-authored by Claude Sonnet 5
parent f54cda3311
commit 5ae68d4ea6
6 changed files with 123 additions and 48 deletions
+51 -20
View File
@@ -184,9 +184,35 @@ result worth restating plainly here: `(a+)+b`, this document's own running
example of the backtracking engine's remaining weak spot, is Pike VM
eligible and now measures as linear, not quadratic, from *n* = 10,000 to
*n* = 160,000 (0.0018s to 0.0308s, roughly 2x per doubling of *n*
throughout). Relative performance between the two engines on ordinary,
non-adversarial patterns is not yet measured (`concept.md` 7.7/7.8 both say
so plainly); this section will be updated once it is.
throughout).
`concept.md` 7.9 records three further findings, closing gaps 7.8 had left
open rather than repeating its numbers unchanged. First, the dual-engine
debug mode 7.8 said had not been built was built (an ad hoc whitebox
harness, not a permanent one, 7.9 explains why): 24,000 `match`/
`fullmatch`/`search` comparisons and roughly 2,700 `finditer` comparisons
between the two engines on the same compiled patterns, zero mismatches.
Second, a literal prefilter was added to the Pike VM's unanchored search
(when a pattern must begin with a specific literal or class, a fresh start
thread that position cannot satisfy is now skipped before being created,
not after), fixing the sharpest instance of the "ordinary patterns got
slower" regression "Benchmarks" below used to report: the literal-search
scenario there went from roughly 70x slower than POSIX `<regex.h>` to
roughly 11x-13x, better than this build's own numbers from before the Pike
VM existed at all (roughly 22x-35x). Third, the Pike VM's own memory cost,
which the design only asserted a bound for, was profiled directly with
Valgrind/Massif on the same adversarial pattern shape used for the
backtracking engine's own worst case (`a*b`, nullable leading loop, no
benefit from the new prefilter): the engine's own contribution beyond
holding the 10MB input itself was roughly 12KB, confirming the design's
`O(instruction count x group count)`, input-length-independent memory bound
actually holds for the v1 implementation, not only on paper.
Two things remain genuinely open, not resolved by the above: no lazy DFA
state caching or allocation pooling exists yet (`concept.md` 7.7's
remaining deferred items), and whether the Pike VM is faster or slower
than the backtracking engine specifically, as opposed to against POSIX, on
ordinary patterns is still unmeasured.
### Memory footprint
@@ -194,12 +220,13 @@ Measured directly with Valgrind/Massif (a real 10MB search) and by watching
`VmRSS`/`VmHWM` on real 200MB-1GB files, not estimated from reading the
code, for the **backtracking engine**; the Pike VM's own memory cost has a
different shape (a bounded number of threads, each carrying its own small
capture array, `concept.md` 7.7) and has not yet been measured the same
detailed way, only checked at a 20MB scale in passing (a peak of roughly
6x the input length across a combined `finditer`/`split`/`sub` run,
`concept.md` 7.8), so the multiplier below is specific to patterns still
running on the backtracking engine (a backreference, lookaround, or atomic
group present), not a claim about every pattern. `Pattern_search`/`finditer`/
capture array, `concept.md` 7.7) and was profiled the same rigorous way
separately (`concept.md` 7.9, "The Pike VM" above): on the same 10MB
adversarial pattern shape used below, its own contribution beyond the input
itself was roughly 12KB, not a multiplier of the input length at all, so
the multiplier below is specific to patterns still running on the
backtracking engine (a backreference, lookaround, or atomic group present),
not a claim about every pattern. `Pattern_search`/`finditer`/
`split`/`sub` (the operations that try more than one start position) on
that engine currently use, at peak, about **8.4x** the input length in
memory for a pattern using `OP_REPEAT1`
@@ -437,17 +464,21 @@ atomic group, so every one of them runs on the Pike VM ("The Pike VM"
above), not the backtracking engine.
Reported without adjustment in either direction, on this machine: glibc's
engine is roughly 8x to 70x faster on the three ordinary scenarios
(a literal search, extracting every number from a text, `a*b`), a real,
expected regression from where this build measured previously on those
same three scenarios (then roughly 2x to 25x, running on the backtracking
engine plus its own precomputed skip-ahead tables, before the Pike VM
existed at all): the Pike VM has none of that engine's constant-factor
optimizations yet (no literal prefilter, no lazy DFA state caching, no
allocation pooling, `concept.md` 7.7's explicitly deferred performance
axis; correctness came first, per the instruction this engine was built
under), and glibc's decades of exactly that kind of optimization show up
directly in the gap widening rather than narrowing. The fourth scenario
engine is roughly 7x to 27x faster on the three ordinary scenarios (a
literal search, extracting every number from a text, `a*b`). This narrowed
substantially after a literal prefilter was added to the Pike VM's
unanchored search (`concept.md` 7.9, "The Pike VM" above): the literal
search scenario, initially the worst hit, going from roughly 70x slower
than glibc (regressed from roughly 22x before the Pike VM existed at all)
down to roughly 11x-13x, better than that original pre-Pike-VM number; the
number-extraction scenario improved more modestly (roughly 8x to roughly
7x), since it begins with a class rather than a single literal; the `a*b`
scenario is unchanged, since its leading `a*` is a nullable loop the
prefilter structurally cannot help, an expected, not a missed, case. No
lazy DFA state caching or allocation pooling exists yet
(`concept.md` 7.7's remaining deferred performance items), so a real gap
to glibc's decades of exactly that kind of optimization remains on these
three scenarios, narrower than before but not closed. The fourth scenario
inverts entirely: the textbook ReDoS shape `(a+)+b` now runs *faster than
glibc*, with no atomic group needed at all, because a Thompson-NFA
simulation has no notion of "try one split, then backtrack and try