Each new program under examples/ isolates one distinct feature rather than being a general purpose tool like the existing rxgrep.c: binary_scan.c (raw byte-range classes including an embedded NUL and an embedded 0x0A), utf8_scripts.c (\w across Latin/Greek/Cyrillic/CJK text, code point versus byte offsets), ascii_logparse.c (named groups against structured log text), redos_atomic.c (atomic groups and possessive quantifiers timed directly against the unprotected form of the textbook (a+)+b ReDoS shape), empty_match_rule.c (CPython's undocumented empty-match retry rule, verified: \d*? against "123abc456" gives 16 matches, not 9), and large_file_search.c (Input_from_file's mmap-backed reading on a generated 100MB file, with elapsed time and peak RSS printed). Every example was compiled and run while writing it; the claims in each file's top comment are checked against its own output, not written by hand and left unverified. Also adds examples/bench_vs_posix.c, a direct, honestly reported comparison against the C standard library's own <regex.h> (regcomp/regexec) on six scenarios at multi-megabyte or multi-hundred-thousand-line scale, using only pattern syntax valid for both engines so they run the identical pattern text. glibc's DFA-backed engine wins five of six scenarios by 2x-35x, which is the expected outcome of a roughly 2000-line backtracking interpreter built for Python `re` compatibility competing against a mature, heavily optimized engine with a much smaller feature set; the sixth scenario has no POSIX equivalent at all (an atomic group). Every scenario's match count is cross-checked between the two engines as an independent correctness signal beyond the existing CPython-derived test suite. Two real issues were found and fixed while building this benchmark, not left in: iterating regexec() over an advancing string pointer is quadratic in practice (no way to bound the search without an implicit NUL-scan on every call), fixed by using REG_STARTEND instead; and a signed integer overflow (undefined behavior, caught by UBSan) in the benchmark's own pseudo-random text generator, fixed by using an unsigned accumulator. README.md and USAGE.md gain pointers to examples/README.md (the new per-example index) and a "Benchmarks" section summarizing the POSIX comparison honestly, including where it loses. The Makefile gains a `make examples` target building all seven programs. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
14 lines
160 B
Plaintext
14 lines
160 B
Plaintext
*.o
|
|
*.a
|
|
rxgrep
|
|
binary_scan
|
|
utf8_scripts
|
|
ascii_logparse
|
|
redos_atomic
|
|
empty_match_rule
|
|
large_file_search
|
|
bench_vs_posix
|
|
tests/generated_tests.c
|
|
/.claude/
|
|
*.dSYM/
|