Files
regexx/examples
retoorandClaude Sonnet 5 5ae68d4ea6 Add a literal prefilter to the Pike VM, profile its memory with Massif
Closed two of the three gaps the previous commit's honest self-review
left open (the third, full streaming/bounded-memory input, remains
out of scope for this pass and is still documented as such).

Literal prefilter (regexx.c, pike_find): when the NFA-only Prog's
first instruction is a mandatory OP_CHAR or OP_CLASS, a pattern
beginning with a required literal or class rather than a nullable
loop or a leading assertion, injecting a fresh unanchored start
thread at a position that instruction would reject is certain to die
on the very next pike_step call regardless; checking that identical
condition before injecting rather than after changes nothing about
which threads ever exist, only how much wasted work is done finding
out. This targeted exactly examples/bench_vs_posix.c's worst
regression: the literal-search scenario went from roughly 70x slower
than POSIX <regex.h> (up from roughly 22x before the Pike VM existed)
down to roughly 11x-13x, better than the original pre-Pike-VM number;
number extraction (starts with a class) improved more modestly; a*b
(starts with a nullable loop, structurally unhelped) is unchanged, as
expected. Verified with the full 3,252-case suite, three clean
AddressSanitizer/UndefinedBehaviorSanitizer passes, and a rerun of
the whitebox dual-engine cross-check (24,000 match/fullmatch/search
plus ~2,700 finditer comparisons between the two engines on the same
compiled patterns, zero mismatches).

Memory profiling (concept.md 7.9): Valgrind/Massif on the same
adversarial, prefilter-proof pattern shape (a*b, nullable leading
loop) used for the backtracking engine's own worst case, for a fair
comparison. Peak heap was almost entirely the 10MB input buffer
itself; the Pike VM's own contribution was roughly 12KB, confirming
the design's O(instruction count x group count), input-length-
independent memory bound actually holds for the v1 implementation,
not only on paper. concept.md Section 10's table, which had only a
"not yet profiled" caveat for this row before, is updated with the
measured result.

README.md, docs/API.md, USAGE.md, and bench_vs_posix.c's own printed
summary are updated throughout with the corrected numbers, rather
than left describing the pre-prefilter regression as current.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
2026-09-14 12:21:15 +00:00
..

Examples

Each program here is self-contained (main, no shared helper code) and isolates one distinct feature rather than being a general purpose tool; the one exception is rxgrep.c, a small grep-like program that ties several of the same features together into something actually usable from a shell. ../USAGE.md walks through the same API surface function by function, with smaller inline snippets; these programs are complete, runnable, and closer to what real code doing this specific thing looks like.

Build all of them with make examples from the repository root (or make all / make example for just rxgrep, the default). Each is also a single cc -I. file.c libregexx.a -o file invocation away from being built directly, no build system required.

Program Demonstrates
rxgrep.c A complete grep-like CLI: all three data modes via -m, Pattern_search/finditer for line and per-match output, Pattern_subn for --sub, reading from a real file or a pipe. Run ./rxgrep --help.
binary_scan.c BINARY mode: a byte-range character class ([\x00-\xff]) matching raw bytes 0-255, including an embedded NUL and an embedded 0x0A, neither of which is a terminator or a line break in this mode the way it would be to a C string function or a line-oriented reader. This is the one thing ASCII/UTF8 mode cannot do, since both introduce either a restricted classification or a decoding step.
utf8_scripts.c UTF8 mode: \w recognizing letters across Latin, Greek, Cyrillic, and CJK text (not only ASCII), and the resulting difference between Match_span (code point units) and Match_span_byte (byte units) once a match spans characters that are more than one byte wide.
ascii_logparse.c ASCII mode doing what it is ordinarily used for: parsing structured, line-oriented text with named groups ((?P<name>...)) and reading fields back out by name via Match_group/Pattern_groupindex_lookup, not by a numeric position the caller has to remember.
redos_atomic.c Two-part: first, that the textbook (a+)+b ReDoS shape is now fixed automatically, with no atomic group, because it has no backreference and so runs on the Pike VM (README.md "The Pike VM"); second, a variant with a backreference added specifically to force it back onto the backtracking engine, where the same shape is exponential again and an atomic group ((?>...)) is still the pattern author's own necessary fix.
empty_match_rule.c Pattern_finditer/Pattern_split reproducing a real CPython interpreter's undocumented empty-match retry rule exactly (\d*? against "123abc456" yields 16 matches, not 9), found by probing a real interpreter directly rather than by reading its documentation.
large_file_search.c Input_from_file's mmap-backed reading on a generated 100MB file, with the elapsed time and peak resident memory printed directly so they can be checked against the measured figures in ../README.md "Memory footprint" rather than taken on faith.
bench_vs_posix.c A direct, timed comparison against the C standard library's own <regex.h> (POSIX regcomp/regexec) on six scenarios at multi-megabyte/multi-hundred-thousand-line scale, using only pattern syntax valid in both engines, every one of them Pike VM eligible. Reported honestly, including a real regression: glibc wins the three ordinary scenarios by a wider margin than this build measured before the Pike VM existed (it has none of the backtracking engine's constant-factor optimizations yet), while the adversarial (a+)+b scenario inverts entirely and now beats glibc outright, with no atomic group needed.

Every program below was compiled and actually run while writing it; the claims in its top-of-file comment (what a real CPython interpreter does, what timing difference an atomic group makes, and so on) are checked against that program's own printed output, not written by hand and left unverified.