Files
regexx/README.md
T
retoorandClaude Sonnet 5 b0991814cb Expand BINARY/UTF8/ASCII edge-case testing to 3,252 cases, fix a real LOCALE bug
Three new categories added to tests/cases.py (TEST_PLAN.md records the
detail): Category H exercises BINARY mode with genuinely arbitrary raw
bytes (embedded NUL, high bytes, non-UTF-8 sequences) via real Python
`bytes` subjects, not just UTF-8-encoded text; Category I exercises
UTF8 mode edge cases (4-byte/astral code points, combining marks,
Arabic, Hebrew, CJK, offset correctness across multi-byte characters);
Category J is a fixed-seed (reproducible, not flaky) random generator
combining the existing atom/quantifier/grouping vocabulary across all
three modes.

This found and fixed a real bug, not just a test-generation one:
LOCALE, in BINARY/ASCII mode, treated bytes 0x80-0xFF as word
characters, based on an unverified assumption about what the "C"
locale does. Checked directly against both the C standard's own
guarantee for isalnum() under "C" and a real CPython interpreter with
re.LOCALE and the "C" locale explicitly set, neither treats anything
above 0x7f as a word character. Fixed in regexx.c's cls_is_word;
LOCALE is now documented as an accepted no-op in non-UTF8 mode,
matching verified reality instead of a prior assumption (concept.md
13.4, docs/API.md, README.md "Known deviations").

Two more findings were test-generation bugs, not regexx bugs: gen.py's
own ASCII-mode ground truth used Python str + re.ASCII (code-point
space) instead of a bytes pattern against a bytes subject (what
regexx's byte-oriented ASCII mode actually is), and LOCALE combined
with the (now removed as redundant) auto-added re.ASCII flag raised
ValueError in Python for being an incompatible combination. Both
fixed in gen.py.

Two further findings were concrete instances of an already-documented
category (glibc's wctype.h Unicode tables not matching CPython's own
exactly): U+00A0 and fullwidth digits U+FF10-FF19 are recognized by
CPython's \s/\d but not by glibc's iswspace()/iswdigit() under C.utf8.
Recorded in README.md, not patched, for the reason already given for
the first such instance (NBSP) in the previous commit.

After these fixes: all 3,252 committed cases pass, clean under
AddressSanitizer/UndefinedBehaviorSanitizer. The Category J generator
was additionally run against 5 more seeds at 3,000 iterations each
(26,760 further checks) as exploratory validation, all passing; not
committed, to keep the regular suite's size proportionate.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
2026-09-14 09:25:50 +00:00

312 lines
17 KiB
Markdown

# regexx
A single-file C regular expression interpreter that reproduces the observable
behavior of Python's `re` module, including its exact identifier names
(`Pattern`, `Match`, `re_compile`, `re_sub`, `IGNORECASE`, and so on; see
Section 9 of `concept.md`), and that additionally supports binary data
(arbitrary byte streams, including embedded NUL) and UTF-8 text alongside
plain ASCII.
The design rationale, the algorithmic trade-offs, and a full accounting of
what is and is not carried over from Python's `re` and from POSIX's native
`regex.h` are recorded in [`concept.md`](concept.md). This file documents the
implementation that exists today at a glance; [`docs/API.md`](docs/API.md) is
the exhaustive reference (every type, every flag, every function's exact
return-value and memory-ownership convention, checked against a real CPython
interpreter, not against memory).
## Implementation status
This is a v1 implementation. It is a complete, tested engine for the pattern
syntax and operations listed below, executed by a single recursive
backtracking engine (`concept.md` Section 7.3) over a fully materialized copy
of the input.
**It does not yet implement the streaming, bounded-memory regular engine of
`concept.md` Section 7.2.** `Input` (the abstraction over "a source of
chunks", `concept.md` 9.2) is implemented, and `Input_from_file` reads a
whole file into memory before matching. Every public function signature is
already exactly what the streaming design in `concept.md` specifies, so the
non-streaming implementation underneath a given call can be replaced later
without changing any caller. Concretely, today:
- `Pattern_match`/`Pattern_fullmatch` (a single anchored attempt at a fixed
position) run in time proportional to the length of that attempt, and, for
the common case of a single character, class, or `.` repeated by a
quantifier, in *O(1)* recursion depth regardless of input size (the
`OP_REPEAT1` fast path). A repeated *compound* sub-pattern (for example
`(ab)*`) still recurses once per repetition, bounded by a configurable
depth limit (`MAX_DEPTH` in `regexx.c`, currently 60000, not exposed
through the public API, so changing it means editing `regexx.c` and
rebuilding); past that limit, matching fails with a reported error rather
than a stack overflow or a wrong answer.
- **`Pattern_search`/`Pattern_finditer`/`Pattern_split`/`Pattern_sub` run in
linear time for the common case: a pattern built only from simple,
non-backreference constructs where every quantifier's body is a single
character, class, or `.`** (`a*b`, `\s*\d+`, `.*"`, and the great majority
of patterns actually written by hand). This was not always true of this
build; measuring it directly is what caught that it previously was not
(an earlier revision of this file claimed linear time without having
re-verified it, then had to correct that claim, then fixed the underlying
defect; the corrected numbers are below). Two independent techniques make
it true now, neither of them the streaming engine of `concept.md` Section
7.2, which remains unimplemented:
1. **Memoized backtracking.** For any pattern with no `OP_BACKREF`
anywhere, whether execution starting at a given (instruction, text
position) pair can ever reach a match is a fact that never changes
once computed, so `run_memo` (`regexx.c`) caches every *failure* (never
a success, so it cannot change which match is found, only skip
re-deriving failures already known) and consults the cache before
redoing that work. This is a published technique, not a house
invention: a memoization table that records only failures, because
"a matching success ... immediately propagates to the success of the
whole problem," is exactly the scheme described in recent work on
backtracking regex matchers ("Selective Memoization for Efficient
Backtracking Regular Expression Matching", and, for the
lookaround/atomic case specifically, "Efficient Matching with
Memoization for Regexes with Look-around and Atomic Grouping", both
linked below).
2. **Precomputed run lengths and skip-ahead tables for `OP_REPEAT1`.**
Memoizing individual `(instruction, position)` pairs does not help
when each one is already `O(1)` work, which is exactly `a*b`'s
situation: counting how many `a`s follow a position, and then trying
every possible split point against the trailing `b` one at a time,
are each individually cheap but happen `O(remaining length)` times per
start position tried. `compute_maxrun` precomputes, once per
`Pattern_search`/`finditer`/`split` call, how many characters each
repeat can consume from every position in one backward pass, and
`compute_next_prevmatch` precomputes, for a repeat immediately
followed by a single literal/class/`.`, the rightmost position at or
before any given point where that next atom can match, so the
backtrack loop jumps directly to candidates worth trying instead of
visiting every position in between. This is the same idea production
engines call a literal prefilter (RE2's and Rust's `regex` crate's
`memchr`/`memmem`/Teddy prefilters skip positions that provably cannot
match before ever invoking the full engine); those use SIMD-accelerated
library primitives operating directly on bytes, this build uses a
precomputed array, which is slower per lookup but the same
algorithmic idea, and appropriate for `concept.md`'s stated priority of
convenience over performance.
Measured directly: `a*b` searched over *n* bytes of `a` with no `b`
anywhere took 19.2s at *n* = 80,000 before these two techniques and
0.0016s after, with time now scaling linearly (2x per doubling of *n*)
rather than quadratically (4x per doubling).
Cost: these tables use `O(k * n)` memory, `k` being the number of
`OP_REPEAT1` instructions actually present in the compiled pattern,
`n` the input length; `Pattern_match`/`Pattern_fullmatch` never
allocate them (a single anchored attempt gets no benefit from them).
- **A pattern with a backreference can, like CPython's own `_sre`, still
take worst-case exponential time on an adversarial input** (`concept.md`
Section 5, 13.2; memoization above is unsound and therefore disabled
whenever a pattern contains `OP_BACKREF` anywhere, since a backreference's
outcome depends on capture history, not on position alone). This is the
same catastrophic backtracking (ReDoS) behavior CPython itself exhibits on
such patterns, not a regression specific to this engine; a pattern author
who needs to rule it out for a specific pattern can use an atomic group
(`(?>...)`) or a possessive quantifier around the ambiguous repetition,
exactly as they would for CPython's `re`, and exactly as general ReDoS
mitigation guidance recommends (linked below).
- **A backreference-free pattern shaped like nested, overlapping
quantifiers (`(a+)+b`, the textbook ReDoS shape) is no longer
exponential, because it has no backreference and so is memoized, but is
not necessarily linear either.** Measured directly: `(a+)+b` against *n*
characters of `a` with no trailing `b` took over a minute already at
*n* = 40 before memoization (exponential; this build's own stress test
keeps that case at *n* = 24 for exactly this reason) and 3.2s at
*n* = 32,000 after (empirically quadratic: roughly 4x per doubling),
because the inner `a+`'s `OP_REPEAT1` is followed by the group's closing
save, not by a simple atom, so the skip-ahead technique above does not
apply to it, only the memoization does. Exponential to quadratic is a
qualitative change (the *n* = 40 case alone went from longer than this
project could wait for, to instant), not a complete fix; a pattern
actually meant to run unattended against adversarial input should still
avoid this shape, or wrap the inner repetition in an atomic group,
regardless of the improvement.
Further reading on the techniques above: Russ Cox, ["Regular Expression
Matching: the Virtual Machine
Approach"](https://swtch.com/~rsc/regexp/regexp2.html) (why prepending
`.*?` gives linear-time unanchored search only in a Thompson/Pike VM, not in
a backtracking engine, which is why this build needed a different fix);
["Selective Memoization for Efficient Backtracking Regular Expression
Matching"](https://arxiv.org/html/2606.26678) and ["Efficient Matching with
Memoization for Regexes with Look-around and Atomic
Grouping"](https://arxiv.org/pdf/2401.12639) (the failure-only memoization
scheme this build's `run_memo` implements, and a more memory-efficient
selective variant, memoizing only at loop "feedback nodes" rather than every
instruction, that this build does not implement but could); the [Snyk
writeup on ReDoS and catastrophic
backtracking](https://snyk.io/blog/redos-and-catastrophic-backtracking/) for
the general phenomenon and mitigation guidance.
### Pattern syntax supported
Literals; `.` (with `DOTALL`); character classes with ranges, negation, and
`\d \D \w \W \s \S`; `\b \B`; anchors `^ $ \A \Z` (with `MULTILINE`);
quantifiers `* + ? {m,n} {m,} {,n} {m}`, greedy and lazy; possessive
quantifiers `*+ ++ ?+ {m,n}+`; groups `(...) (?:...) (?P<name>...)`;
alternation `|`; backreferences `\1`-`\99`, `(?P=name)`, `\g<name>`,
`\g<N>`; lookahead `(?=...) (?!...)`; fixed-width lookbehind
`(?<=...) (?<!...)`; atomic groups `(?>...)`; comments `(?#...)`; global
inline flags `(?aiLmsux)` at the start of a pattern; escapes
`\n \r \t \f \v \a`, octal `\0`-prefixed escapes, `\xhh`, `\uxxxx`,
`\Uxxxxxxxx`; flags `IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE`, `ASCII`
(all with observable effect; see `docs/API.md` Section 2 for exactly what
each one does), plus `UNICODE`, `LOCALE`, and `DEBUG` (accepted for source
compatibility with Python, currently no-ops in this build: `LOCALE` because
this build commits to the "C" locale only, under which `\w`/`\b`/`\B`
classify only ASCII letters and digits regardless of the flag, verified
directly against both the C standard and a real CPython interpreter).
Rejected at compile time with a clear `PatternError`, rather than
mis-parsed: conditional groups `(?(id)yes|no)`, scoped inline flags
`(?flags:...)`, `\N{NAME}` named code points, and POSIX bracket classes
`[:alpha:]` (which are not part of Python `re` at all, `concept.md` 14.3).
Variable-width lookbehind is also rejected at compile time, matching
CPython.
### Operations supported
`Pattern_match/fullmatch/search/finditer/findall/split/sub/subn/free`,
`Pattern_groupindex_lookup`,
`Match_group/start/end/span/start_byte/end_byte/span_byte/free`,
`re_compile/match/fullmatch/search/finditer/findall/split/sub/subn/escape/purge`,
`PatternError_free`. See `regexx.h` for exact signatures, `docs/API.md` for
the full reference (return values, memory ownership, exact Python
correspondence for each one), and `concept.md` Section 9 for the naming
convention they follow.
### Known deviations from `concept.md` and from CPython, beyond the items above
- `\w`, `\s`, `IGNORECASE` case folding, and `\d` in `UTF8` mode are backed
by glibc's `wctype.h` functions under the `C.utf8` locale, not by a
hand-generated Unicode table (`concept.md` 13.3 anticipated a reduced
static table; using the C library's own tables turned out to be simpler
and more complete, at the cost of depending on the platform's Unicode
version rather than a pinned one). Concretely verified, not just
theoretical: U+00A0 (NO-BREAK SPACE) is in Unicode's `White_Space`
property, so CPython's `\s` matches it, but glibc's `iswspace()` under
`C.utf8` does not, so this build's `\s` does not either. Found by the
large combinatorial test expansion (`tests/cases.py` Category F,
`tests/TEST_PLAN.md`), not anticipated in advance; recorded here rather
than patched, since hand-patching individual code points would start
down the path of maintaining an ad hoc table this design deliberately
avoided by delegating to `wctype.h` in the first place. A second,
same-class instance was found by the later Category I expansion:
fullwidth digits (`U+FF10`-`U+FF19`, Unicode category `Nd`) match
CPython's `\d` but not glibc's `iswdigit()` under `C.utf8` either.
- `lastindex`/`lastgroup` report the highest-numbered capturing group that
participated in the match, which coincides with CPython's "most recently
closed group" rule for straightforward patterns but can differ from it in
pathological cases (nested alternation re-executing a lower-numbered group
after a higher one). Not exercised by the test suite; documented here
rather than silently accepted.
- `Match_free`, `PatternError_free`, `Input_from_buffer`, `Input_from_file`,
and `Input_free` have no Python counterpart and are not mentioned in
`concept.md`'s API surface; they exist because C has no garbage collector.
`Pattern_sub`/`Pattern_subn` take the replacement template and the
callback as two separate parameters rather than one polymorphic argument,
for the same reason (`concept.md` 9.4 already anticipates and justifies
this one).
- Python's `Match.start(group)`/`.end(group)` raise `IndexError` for an
invalid group number and return `-1` only for a valid group that did not
participate; `Match_start`/`Match_end` return `-1` for both cases, since C
has no exception to raise. `Match_group` does distinguish them (`-1` for
no such group, `0` for an unparticipated one), see `docs/API.md` Section
3.23.
- `Pattern.groupindex` has no enumeration function in this build, only
`Pattern_groupindex_lookup(pattern, name)`; there is no way to list every
name a compiled pattern defines without already knowing what to look for.
## Building
Requires a C11 compiler and, for the test suite, Python 3 (used only to
generate ground truth from CPython's own `re` module, `concept.md` Section
11; the library itself has no runtime dependency beyond the C standard
library and `libc`'s `wctype.h`/`locale.h`).
```sh
make # builds libregexx.a and the rxgrep example
make test # regenerates tests/generated_tests.c from Python `re`
# ground truth and runs the full suite
make check # same, under AddressSanitizer + UndefinedBehaviorSanitizer
make clean
```
`make install` installs `libregexx.a` and `regexx.h` under `PREFIX`
(default `/usr/local`).
## Using the library
```c
#include "regexx.h"
#include <string.h>
const char *pattern = "(\\w+)@(\\w+)";
Pattern *pat = re_compile(pattern, strlen(pattern), UTF8, NULL);
Input *in = Input_from_buffer((const uint8_t *)"user@host", strlen("user@host"));
Match m;
if (Pattern_search(pat, in, 0, -1, &m) == 1) {
const char *g; size_t glen;
Match_group(&m, NULL, 1, &g, &glen); /* g/glen -> "user" */
Match_free(&m);
}
char *out; size_t outlen;
Pattern_sub(pat, in, "\\2@\\1", NULL, NULL, 0, &out, &outlen); /* "host@user" */
free(out);
Pattern_free(pat);
Input_free(in);
```
`flags` to `re_compile` combine a data mode, exactly one of `BINARY`,
`ASCII`, or `UTF8` (`concept.md` 9.3), with any of the Python-named flags
(`IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE`, `ASCII` as a flag also
forces ASCII-only `\w`/`\s`/`\d` inside `UTF8` mode, `LOCALE`, `DEBUG`).
## Example: rxgrep
`examples/rxgrep.c` is a small grep-like program built on the library,
demonstrating all three data modes and both the matching and substitution
API:
```sh
./rxgrep -in 'hello' file.txt # case-insensitive, line numbers
./rxgrep -m utf8 -o '\w+' file.txt # print every UTF-8 word, one per line
./rxgrep -c 'error' log.txt # count matching lines
./rxgrep -m binary 'a.c' data.bin # match raw bytes, embedded NUL included
./rxgrep -m utf8 --sub 'REDACTED' '\d{3}-\d{4}' file.txt
```
Run `./rxgrep --help` for the full option list.
## Testing
`tests/cases.py` lists pattern/subject/operation triples, both hand-written
and, for most of the file, generated programmatically from combinations of
quantifiers, groups, backreferences, lookaround, flags, and encoding modes
(`tests/TEST_PLAN.md` records the exact category breakdown and why each one
exists). `tests/gen.py` computes every case's expected result with
CPython's own `re` module and writes `tests/generated_tests.c`, which is
then compiled against `regexx.c` and checked. This is a direct
implementation of the strategy `concept.md` Section 11 describes:
conformance is measured against what CPython actually does, not against a
re-derived reading of its documentation, at a scale (3,252 cases as of this
writing, `make test` reports the current exact count) large enough that it
has already found real defects this way, not only confirmed the absence of
ones anyone thought to write by hand (`tests/TEST_PLAN.md` "Result" names
all four: a C trigraph bug in the test generator itself, a real, previously
undocumented `Pattern_finditer`/`Pattern_split` empty-match defect, a wrong
ground truth for `ASCII` mode in the generator itself, and a real,
verified-wrong implementation of the `LOCALE` flag). `make check`
additionally runs the suite under AddressSanitizer and
UndefinedBehaviorSanitizer.
## License
MIT. See [`LICENSE`](LICENSE).