Add a literal prefilter to the Pike VM, profile its memory with Massif

Closed two of the three gaps the previous commit's honest self-review
left open (the third, full streaming/bounded-memory input, remains
out of scope for this pass and is still documented as such).

Literal prefilter (regexx.c, pike_find): when the NFA-only Prog's
first instruction is a mandatory OP_CHAR or OP_CLASS, a pattern
beginning with a required literal or class rather than a nullable
loop or a leading assertion, injecting a fresh unanchored start
thread at a position that instruction would reject is certain to die
on the very next pike_step call regardless; checking that identical
condition before injecting rather than after changes nothing about
which threads ever exist, only how much wasted work is done finding
out. This targeted exactly examples/bench_vs_posix.c's worst
regression: the literal-search scenario went from roughly 70x slower
than POSIX <regex.h> (up from roughly 22x before the Pike VM existed)
down to roughly 11x-13x, better than the original pre-Pike-VM number;
number extraction (starts with a class) improved more modestly; a*b
(starts with a nullable loop, structurally unhelped) is unchanged, as
expected. Verified with the full 3,252-case suite, three clean
AddressSanitizer/UndefinedBehaviorSanitizer passes, and a rerun of
the whitebox dual-engine cross-check (24,000 match/fullmatch/search
plus ~2,700 finditer comparisons between the two engines on the same
compiled patterns, zero mismatches).

Memory profiling (concept.md 7.9): Valgrind/Massif on the same
adversarial, prefilter-proof pattern shape (a*b, nullable leading
loop) used for the backtracking engine's own worst case, for a fair
comparison. Peak heap was almost entirely the 10MB input buffer
itself; the Pike VM's own contribution was roughly 12KB, confirming
the design's O(instruction count x group count), input-length-
independent memory bound actually holds for the v1 implementation,
not only on paper. concept.md Section 10's table, which had only a
"not yet profiled" caveat for this row before, is updated with the
measured result.

README.md, docs/API.md, USAGE.md, and bench_vs_posix.c's own printed
summary are updated throughout with the corrected numbers, rather
than left describing the pre-prefilter regression as current.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
This commit is contained in:
2026-09-14 12:21:15 +00:00
co-authored by Claude Sonnet 5
parent f54cda3311
commit 5ae68d4ea6
6 changed files with 123 additions and 48 deletions
+19 -15
View File
@@ -429,21 +429,25 @@ int main(void) {
printf("\nWhat this measures and does not measure: regexx is Python `re`\n"
"compatible and dispatches every pattern here (all backreference-,\n"
"lookaround-, and atomic-group-free) to its Pike VM (README.md \"The\n"
"Pike VM\", concept.md 7.2/7.7), a Thompson-NFA simulation with no\n"
"literal prefilter, lazy DFA state caching, or allocation pooling yet\n"
"(concept.md 7.7's explicitly deferred performance axis); glibc's\n"
"<regex.h> is a mature, DFA-backed engine with decades of that exact\n"
"kind of optimization behind it and a much smaller feature set (no\n"
"named groups, no lazy quantifiers, no lookaround, no atomic groups,\n"
"and, structurally, no way to search past an embedded NUL byte at all,\n"
"docs/API.md Section 1.3 vs examples/binary_scan.c). Scenarios A-C\n"
"measure that missing optimization layer directly: this build was\n"
"previously measured on the same three scenarios, on this engine's\n"
"backtracking predecessor plus its own precomputed skip-ahead tables\n"
"(README.md \"Implementation status\"), at roughly 2-3x less of a gap\n"
"to POSIX than shown here, a real, expected regression for ordinary\n"
"patterns from moving to a v1 engine still missing that layer, not\n"
"a defect introduced by the new engine's correctness.\n\n"
"Pike VM\", concept.md 7.2/7.7/7.9), a Thompson-NFA simulation with a\n"
"literal prefilter (added after this benchmark first found scenario A\n"
"regressed the most, concept.md 7.9) but still no lazy DFA state\n"
"caching or allocation pooling; glibc's <regex.h> is a mature,\n"
"DFA-backed engine with decades of that exact kind of optimization\n"
"behind it and a much smaller feature set (no named groups, no lazy\n"
"quantifiers, no lookaround, no atomic groups, and, structurally, no\n"
"way to search past an embedded NUL byte at all, docs/API.md Section\n"
"1.3 vs examples/binary_scan.c). Scenarios A-C measure what remains of\n"
"that missing optimization layer: this build was previously measured\n"
"on this engine's backtracking predecessor plus its own precomputed\n"
"skip-ahead tables (README.md \"Implementation status\") at roughly\n"
"2x-25x against POSIX on these same three scenarios; the Pike VM\n"
"without a prefilter regressed that to roughly 8x-70x, and the\n"
"prefilter closed most of that back up for A (a plain literal, its\n"
"best case) and part of it for B (starts with a class, not a single\n"
"literal); C's a*b starts with a nullable loop no prefilter can help,\n"
"and is unchanged. What remains (roughly 7x-27x here) is genuinely no\n"
"lazy DFA caching or allocation pooling yet, not a defect.\n\n"
"Scenario D inverts entirely, and is the actual point of building this\n"
"engine at all: (a+)+b has no backreference, so it is Pike VM eligible,\n"
"and a Thompson-NFA simulation has no notion of \"try one split, then\n"