Files
regexx/Makefile
T
retoorandClaude Sonnet 5 02743cd959 Add six feature examples and a direct benchmark against POSIX <regex.h>
Each new program under examples/ isolates one distinct feature rather
than being a general purpose tool like the existing rxgrep.c:
binary_scan.c (raw byte-range classes including an embedded NUL and
an embedded 0x0A), utf8_scripts.c (\w across Latin/Greek/Cyrillic/CJK
text, code point versus byte offsets), ascii_logparse.c (named groups
against structured log text), redos_atomic.c (atomic groups and
possessive quantifiers timed directly against the unprotected form of
the textbook (a+)+b ReDoS shape), empty_match_rule.c (CPython's
undocumented empty-match retry rule, verified: \d*? against
"123abc456" gives 16 matches, not 9), and large_file_search.c
(Input_from_file's mmap-backed reading on a generated 100MB file,
with elapsed time and peak RSS printed). Every example was compiled
and run while writing it; the claims in each file's top comment are
checked against its own output, not written by hand and left
unverified.

Also adds examples/bench_vs_posix.c, a direct, honestly reported
comparison against the C standard library's own <regex.h>
(regcomp/regexec) on six scenarios at multi-megabyte or
multi-hundred-thousand-line scale, using only pattern syntax valid
for both engines so they run the identical pattern text. glibc's
DFA-backed engine wins five of six scenarios by 2x-35x, which is the
expected outcome of a roughly 2000-line backtracking interpreter
built for Python `re` compatibility competing against a mature,
heavily optimized engine with a much smaller feature set; the sixth
scenario has no POSIX equivalent at all (an atomic group). Every
scenario's match count is cross-checked between the two engines as an
independent correctness signal beyond the existing CPython-derived
test suite.

Two real issues were found and fixed while building this benchmark,
not left in: iterating regexec() over an advancing string pointer is
quadratic in practice (no way to bound the search without an implicit
NUL-scan on every call), fixed by using REG_STARTEND instead; and a
signed integer overflow (undefined behavior, caught by UBSan) in the
benchmark's own pseudo-random text generator, fixed by using an
unsigned accumulator.

README.md and USAGE.md gain pointers to examples/README.md (the new
per-example index) and a "Benchmarks" section summarizing the
POSIX comparison honestly, including where it loses. The Makefile
gains a `make examples` target building all seven programs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
2026-09-14 11:04:10 +00:00

88 lines
3.3 KiB
Makefile

# Makefile for regexx: a single-file C regex interpreter with Python
# `re` semantics. See concept.md for the design and README.md for the
# implementation status.
CC ?= cc
CSTD ?= -std=c11
WARN ?= -Wall -Wextra
OPT ?= -O2
CFLAGS ?= $(CSTD) $(WARN) $(OPT) -g
LDLIBS ?=
PREFIX ?= /usr/local
AR ?= ar
.PHONY: all lib example examples test check clean install fuzz-smoke
all: lib example
# ---- library -----------------------------------------------------------
libregexx.a: regexx.o
$(AR) rcs $@ $^
regexx.o: regexx.c regexx.h
$(CC) $(CFLAGS) -c regexx.c -o $@
lib: libregexx.a
# ---- example -------------------------------------------------------------
example: rxgrep
rxgrep: examples/rxgrep.c libregexx.a regexx.h
$(CC) $(CFLAGS) -I. -o $@ examples/rxgrep.c libregexx.a $(LDLIBS)
# ---- feature examples ------------------------------------------------------
# Each program under examples/ (besides rxgrep) isolates one distinct,
# less-obvious feature (a data mode, an empty-match rule, a ReDoS
# mitigation, mmap-backed large file input) rather than being a general
# purpose tool; see examples/README.md for what each one demonstrates.
EXAMPLE_BINS = binary_scan utf8_scripts ascii_logparse redos_atomic empty_match_rule large_file_search bench_vs_posix
examples: rxgrep $(EXAMPLE_BINS)
binary_scan: examples/binary_scan.c libregexx.a regexx.h
$(CC) $(CFLAGS) -I. -o $@ examples/binary_scan.c libregexx.a $(LDLIBS)
utf8_scripts: examples/utf8_scripts.c libregexx.a regexx.h
$(CC) $(CFLAGS) -I. -o $@ examples/utf8_scripts.c libregexx.a $(LDLIBS)
ascii_logparse: examples/ascii_logparse.c libregexx.a regexx.h
$(CC) $(CFLAGS) -I. -o $@ examples/ascii_logparse.c libregexx.a $(LDLIBS)
redos_atomic: examples/redos_atomic.c libregexx.a regexx.h
$(CC) $(CFLAGS) -I. -o $@ examples/redos_atomic.c libregexx.a $(LDLIBS)
empty_match_rule: examples/empty_match_rule.c libregexx.a regexx.h
$(CC) $(CFLAGS) -I. -o $@ examples/empty_match_rule.c libregexx.a $(LDLIBS)
large_file_search: examples/large_file_search.c libregexx.a regexx.h
$(CC) $(CFLAGS) -I. -o $@ examples/large_file_search.c libregexx.a $(LDLIBS)
bench_vs_posix: examples/bench_vs_posix.c libregexx.a regexx.h
$(CC) $(CFLAGS) -I. -o $@ examples/bench_vs_posix.c libregexx.a $(LDLIBS)
# ---- tests -----------------------------------------------------------------
# tests/generated_tests.c is generated from tests/cases.py using
# CPython's own `re` module as ground truth (concept.md Section 11).
tests/generated_tests.c: tests/cases.py tests/gen.py
python3 tests/gen.py
test: tests/generated_tests.c regexx.c regexx.h tests/harness.c tests/harness.h tests/main.c
$(CC) $(CFLAGS) -I. -o /tmp/regexx_test regexx.c tests/harness.c tests/generated_tests.c tests/main.c $(LDLIBS)
/tmp/regexx_test
# check also runs the suite under AddressSanitizer + UBSan.
check: tests/generated_tests.c
$(CC) $(CSTD) $(WARN) -O0 -g -fsanitize=address,undefined -I. -o /tmp/regexx_test_san \
regexx.c tests/harness.c tests/generated_tests.c tests/main.c $(LDLIBS)
/tmp/regexx_test_san
install: libregexx.a regexx.h
install -d $(DESTDIR)$(PREFIX)/lib $(DESTDIR)$(PREFIX)/include
install -m644 libregexx.a $(DESTDIR)$(PREFIX)/lib/
install -m644 regexx.h $(DESTDIR)$(PREFIX)/include/
clean:
rm -f regexx.o libregexx.a rxgrep tests/generated_tests.c $(EXAMPLE_BINS)