Researched first (Cox's regexp2 article, the submatch/tagged-NFA follow-up crediting Laurikari, and rust-lang/regex's PikeVM source for the exact leftmost-first mid-search Match handling), then planned in concept.md 7.7 before writing any code, per the explicit instruction to plan before implementing. The existing bytecode compiler turned out to already produce valid Thompson-construction NFA bytecode (OP_SPLIT/OP_JMP/OP_SAVE match Cox's instruction set almost exactly), so no new compiler was needed: Prog gained one field (no_repeat1) and compile_repeat one condition, letting the same compile_node produce a second, NFA-only Prog from the same AST for any pattern containing none of OP_BACKREF/OP_LOOKAHEAD/ OP_LOOKBEHIND/OP_ATOMIC (a possessive quantifier already desugars to the last of these at parse time). That second Prog is matched by a new Pike VM (pike_addthread/pike_step/pike_find): a breadth-first thread list simulation with per-thread capture arrays, epsilon closure implemented with an explicit heap stack rather than C recursion so a pattern with many alternations cannot recurse the C stack, every allocation checked and failing through the existing -1 error convention rather than a NULL dereference. do_one, Pattern_finditer, and Pattern_split each gained an impl->has_nfa branch to this engine, sharing one small helper (find_next) for the "unanchored scan from a position" versus "single anchored attempt at a position" distinction finditer/split's empty-match retry needs. Two real bugs, both found and precisely localized by the existing 3,252-case CPython-derived suite without writing a single test specifically for this engine: OP_MATCH not writing group 0's end position (it has no OP_SAVE; run()'s own OP_MATCH handler sets it directly, and this engine's first version missed replicating that), and an unconditional "thread list empty -> stop" early exit that is wrong for unanchored search specifically, since a freshly injected start thread can die immediately in its own epsilon closure (a leading \b failing outright, repeatedly, inside a longer word like "catalog" for \bcat\b) without that meaning every later position would too. Both fixed; full suite passes, three clean AddressSanitizer/ UndefinedBehaviorSanitizer runs, plus hand-written whitebox checks (a 20,000-branch alternation, UTF8 named groups, BINARY matching across an embedded NUL, greedy/lazy and alternation priority). Measured result: (a+)+b, this project's own running example of the backtracking engine's remaining weak spot, is Pike VM eligible and now measures as genuinely linear (0.0018s to 0.0308s, n=10,000 to 160,000), not merely improved. Measured cost: re-running bench_vs_posix.c's three ordinary scenarios (all now Pike VM eligible too) found the gap to POSIX <regex.h> widened, from roughly 2x-25x before this engine existed to roughly 8x-70x now, the direct, expected cost of this engine's performance axis being explicitly deferred (no literal prefilter, no lazy DFA state caching, no allocation pooling) in favor of correctness first, per the instruction this was built under. Both results, and the reasoning behind deferring the second, are recorded in concept.md 7.7/7.8. README.md, docs/API.md, USAGE.md, and examples/redos_atomic.c and bench_vs_posix.c are updated throughout to describe the new two-engine dispatch accurately, including this real trade-off, rather than leaving the previous single-engine description in place; redos_atomic.c specifically now shows both that (a+)+b no longer needs an atomic group at all and a backreference-forced variant where one still does. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
33 lines
4.0 KiB
Markdown
33 lines
4.0 KiB
Markdown
# Examples
|
|
|
|
Each program here is self-contained (`main`, no shared helper code) and
|
|
isolates one distinct feature rather than being a general purpose tool; the
|
|
one exception is `rxgrep.c`, a small grep-like program that ties several of
|
|
the same features together into something actually usable from a shell.
|
|
[`../USAGE.md`](../USAGE.md) walks through the same API surface function by
|
|
function, with smaller inline snippets; these programs are complete,
|
|
runnable, and closer to what real code doing this specific thing looks
|
|
like.
|
|
|
|
Build all of them with `make examples` from the repository root (or
|
|
`make all` / `make example` for just `rxgrep`, the default). Each is also a
|
|
single `cc -I. file.c libregexx.a -o file` invocation away from being built
|
|
directly, no build system required.
|
|
|
|
| Program | Demonstrates |
|
|
|---|---|
|
|
| [`rxgrep.c`](rxgrep.c) | A complete grep-like CLI: all three data modes via `-m`, `Pattern_search`/`finditer` for line and per-match output, `Pattern_subn` for `--sub`, reading from a real file or a pipe. Run `./rxgrep --help`. |
|
|
| [`binary_scan.c`](binary_scan.c) | `BINARY` mode: a byte-range character class (`[\x00-\xff]`) matching raw bytes 0-255, including an embedded `NUL` and an embedded `0x0A`, neither of which is a terminator or a line break in this mode the way it would be to a C string function or a line-oriented reader. This is the one thing `ASCII`/`UTF8` mode cannot do, since both introduce either a restricted classification or a decoding step. |
|
|
| [`utf8_scripts.c`](utf8_scripts.c) | `UTF8` mode: `\w` recognizing letters across Latin, Greek, Cyrillic, and CJK text (not only ASCII), and the resulting difference between `Match_span` (code point units) and `Match_span_byte` (byte units) once a match spans characters that are more than one byte wide. |
|
|
| [`ascii_logparse.c`](ascii_logparse.c) | `ASCII` mode doing what it is ordinarily used for: parsing structured, line-oriented text with named groups (`(?P<name>...)`) and reading fields back out by name via `Match_group`/`Pattern_groupindex_lookup`, not by a numeric position the caller has to remember. |
|
|
| [`redos_atomic.c`](redos_atomic.c) | Two-part: first, that the textbook `(a+)+b` ReDoS shape is now fixed automatically, with no atomic group, because it has no backreference and so runs on the Pike VM (README.md "The Pike VM"); second, a variant with a backreference added specifically to force it back onto the backtracking engine, where the same shape is exponential again and an atomic group (`(?>...)`) is still the pattern author's own necessary fix. |
|
|
| [`empty_match_rule.c`](empty_match_rule.c) | `Pattern_finditer`/`Pattern_split` reproducing a real CPython interpreter's undocumented empty-match retry rule exactly (`\d*?` against `"123abc456"` yields 16 matches, not 9), found by probing a real interpreter directly rather than by reading its documentation. |
|
|
| [`large_file_search.c`](large_file_search.c) | `Input_from_file`'s `mmap`-backed reading on a generated 100MB file, with the elapsed time and peak resident memory printed directly so they can be checked against the measured figures in `../README.md` "Memory footprint" rather than taken on faith. |
|
|
| [`bench_vs_posix.c`](bench_vs_posix.c) | A direct, timed comparison against the C standard library's own `<regex.h>` (POSIX `regcomp`/`regexec`) on six scenarios at multi-megabyte/multi-hundred-thousand-line scale, using only pattern syntax valid in both engines, every one of them Pike VM eligible. Reported honestly, including a real regression: glibc wins the three ordinary scenarios by a wider margin than this build measured before the Pike VM existed (it has none of the backtracking engine's constant-factor optimizations yet), while the adversarial `(a+)+b` scenario inverts entirely and now beats glibc outright, with no atomic group needed. |
|
|
|
|
Every program below was compiled and actually run while writing it; the
|
|
claims in its top-of-file comment (what a real CPython interpreter does,
|
|
what timing difference an atomic group makes, and so on) are checked
|
|
against that program's own printed output, not written by hand and left
|
|
unverified.
|