Files
regexx/examples
retoorandClaude Sonnet 5 02743cd959 Add six feature examples and a direct benchmark against POSIX <regex.h>
Each new program under examples/ isolates one distinct feature rather
than being a general purpose tool like the existing rxgrep.c:
binary_scan.c (raw byte-range classes including an embedded NUL and
an embedded 0x0A), utf8_scripts.c (\w across Latin/Greek/Cyrillic/CJK
text, code point versus byte offsets), ascii_logparse.c (named groups
against structured log text), redos_atomic.c (atomic groups and
possessive quantifiers timed directly against the unprotected form of
the textbook (a+)+b ReDoS shape), empty_match_rule.c (CPython's
undocumented empty-match retry rule, verified: \d*? against
"123abc456" gives 16 matches, not 9), and large_file_search.c
(Input_from_file's mmap-backed reading on a generated 100MB file,
with elapsed time and peak RSS printed). Every example was compiled
and run while writing it; the claims in each file's top comment are
checked against its own output, not written by hand and left
unverified.

Also adds examples/bench_vs_posix.c, a direct, honestly reported
comparison against the C standard library's own <regex.h>
(regcomp/regexec) on six scenarios at multi-megabyte or
multi-hundred-thousand-line scale, using only pattern syntax valid
for both engines so they run the identical pattern text. glibc's
DFA-backed engine wins five of six scenarios by 2x-35x, which is the
expected outcome of a roughly 2000-line backtracking interpreter
built for Python `re` compatibility competing against a mature,
heavily optimized engine with a much smaller feature set; the sixth
scenario has no POSIX equivalent at all (an atomic group). Every
scenario's match count is cross-checked between the two engines as an
independent correctness signal beyond the existing CPython-derived
test suite.

Two real issues were found and fixed while building this benchmark,
not left in: iterating regexec() over an advancing string pointer is
quadratic in practice (no way to bound the search without an implicit
NUL-scan on every call), fixed by using REG_STARTEND instead; and a
signed integer overflow (undefined behavior, caught by UBSan) in the
benchmark's own pseudo-random text generator, fixed by using an
unsigned accumulator.

README.md and USAGE.md gain pointers to examples/README.md (the new
per-example index) and a "Benchmarks" section summarizing the
POSIX comparison honestly, including where it loses. The Makefile
gains a `make examples` target building all seven programs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
2026-09-14 11:04:10 +00:00
..

Examples

Each program here is self-contained (main, no shared helper code) and isolates one distinct feature rather than being a general purpose tool; the one exception is rxgrep.c, a small grep-like program that ties several of the same features together into something actually usable from a shell. ../USAGE.md walks through the same API surface function by function, with smaller inline snippets; these programs are complete, runnable, and closer to what real code doing this specific thing looks like.

Build all of them with make examples from the repository root (or make all / make example for just rxgrep, the default). Each is also a single cc -I. file.c libregexx.a -o file invocation away from being built directly, no build system required.

Program Demonstrates
rxgrep.c A complete grep-like CLI: all three data modes via -m, Pattern_search/finditer for line and per-match output, Pattern_subn for --sub, reading from a real file or a pipe. Run ./rxgrep --help.
binary_scan.c BINARY mode: a byte-range character class ([\x00-\xff]) matching raw bytes 0-255, including an embedded NUL and an embedded 0x0A, neither of which is a terminator or a line break in this mode the way it would be to a C string function or a line-oriented reader. This is the one thing ASCII/UTF8 mode cannot do, since both introduce either a restricted classification or a decoding step.
utf8_scripts.c UTF8 mode: \w recognizing letters across Latin, Greek, Cyrillic, and CJK text (not only ASCII), and the resulting difference between Match_span (code point units) and Match_span_byte (byte units) once a match spans characters that are more than one byte wide.
ascii_logparse.c ASCII mode doing what it is ordinarily used for: parsing structured, line-oriented text with named groups ((?P<name>...)) and reading fields back out by name via Match_group/Pattern_groupindex_lookup, not by a numeric position the caller has to remember.
redos_atomic.c Atomic groups ((?>...)) and possessive quantifiers (a++) as a real, pattern-author-controlled defense against catastrophic backtracking, timed directly against the plain, unprotected form of the textbook (a+)+b shape on a non-matching adversarial input.
empty_match_rule.c Pattern_finditer/Pattern_split reproducing a real CPython interpreter's undocumented empty-match retry rule exactly (\d*? against "123abc456" yields 16 matches, not 9), found by probing a real interpreter directly rather than by reading its documentation.
large_file_search.c Input_from_file's mmap-backed reading on a generated 100MB file, with the elapsed time and peak resident memory printed directly so they can be checked against the measured figures in ../README.md "Memory footprint" rather than taken on faith.
bench_vs_posix.c A direct, timed comparison against the C standard library's own <regex.h> (POSIX regcomp/regexec) on six scenarios at multi-megabyte/multi-hundred-thousand-line scale, using only pattern syntax valid in both engines. Reported honestly: glibc's DFA-backed engine wins most scenarios by a wide margin, and the one adversarial pattern shape ((a+)+b) where this library's backtracking cost is most visible is shown both in its plain form and fixed with an atomic group, the latter having no POSIX ERE equivalent to compare against at all.

Every program below was compiled and actually run while writing it; the claims in its top-of-file comment (what a real CPython interpreter does, what timing difference an atomic group makes, and so on) are checked against that program's own printed output, not written by hand and left unverified.