Auditing the documentation against the actual code (not against what I remembered writing) surfaced a real, previously undocumented defect: README.md claimed a backreference-free, unbounded-lookahead-free pattern "runs in linear time", but that is only true of Pattern_match/ Pattern_fullmatch (one anchored attempt). Pattern_search tries every candidate start position as an independent from-scratch attempt, so it is quadratic in the worst case even for the simplest pattern, since concept.md Section 7.2's engine (which shares work across start positions in one linear pass) is not implemented yet. Measured directly with a*b over a run of plain a's: 0.32s at n=10,000, 19.2s at n=80,000, an ~4x slowdown per doubling. Pattern_finditer/Pattern_split/ Pattern_sub all build on the same search loop and inherit it. README.md and docs/API.md now state this precisely, with the measured numbers, wherever the affected functions are documented, rather than repeating the incorrect blanket "linear time" claim. No code changed; this is a documentation correction only. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
209 lines
10 KiB
Markdown
209 lines
10 KiB
Markdown
# regexx
|
|
|
|
A single-file C regular expression interpreter that reproduces the observable
|
|
behavior of Python's `re` module, including its exact identifier names
|
|
(`Pattern`, `Match`, `re_compile`, `re_sub`, `IGNORECASE`, and so on; see
|
|
Section 9 of `concept.md`), and that additionally supports binary data
|
|
(arbitrary byte streams, including embedded NUL) and UTF-8 text alongside
|
|
plain ASCII.
|
|
|
|
The design rationale, the algorithmic trade-offs, and a full accounting of
|
|
what is and is not carried over from Python's `re` and from POSIX's native
|
|
`regex.h` are recorded in [`concept.md`](concept.md). This file documents the
|
|
implementation that exists today at a glance; [`docs/API.md`](docs/API.md) is
|
|
the exhaustive reference (every type, every flag, every function's exact
|
|
return-value and memory-ownership convention, checked against a real CPython
|
|
interpreter, not against memory).
|
|
|
|
## Implementation status
|
|
|
|
This is a v1 implementation. It is a complete, tested engine for the pattern
|
|
syntax and operations listed below, executed by a single recursive
|
|
backtracking engine (`concept.md` Section 7.3) over a fully materialized copy
|
|
of the input.
|
|
|
|
**It does not yet implement the streaming, bounded-memory regular engine of
|
|
`concept.md` Section 7.2.** `Input` (the abstraction over "a source of
|
|
chunks", `concept.md` 9.2) is implemented, and `Input_from_file` reads a
|
|
whole file into memory before matching. Every public function signature is
|
|
already exactly what the streaming design in `concept.md` specifies, so the
|
|
non-streaming implementation underneath a given call can be replaced later
|
|
without changing any caller. Concretely, today:
|
|
|
|
- `Pattern_match`/`Pattern_fullmatch` (a single anchored attempt at a fixed
|
|
position) run in time proportional to the length of that attempt, and, for
|
|
the common case of a single character, class, or `.` repeated by a
|
|
quantifier, in *O(1)* recursion depth regardless of input size (the
|
|
`OP_REPEAT1` fast path). A repeated *compound* sub-pattern (for example
|
|
`(ab)*`) still recurses once per repetition, bounded by a configurable
|
|
depth limit (`MAX_DEPTH` in `regexx.c`, currently 60000, not exposed
|
|
through the public API, so changing it means editing `regexx.c` and
|
|
rebuilding); past that limit, matching fails with a reported error rather
|
|
than a stack overflow or a wrong answer.
|
|
- **`Pattern_search`/`Pattern_finditer`/`Pattern_split`/`Pattern_sub` are
|
|
quadratic, not linear, in the worst case, even for a pattern with no
|
|
backreference and no unbounded-width lookahead.** v1's search loop tries
|
|
every start position in turn as an independent match attempt (`concept.md`
|
|
Section 7.2's design shares work across start positions in one linear
|
|
pass; that engine is not implemented yet, see above), so a pattern that
|
|
ultimately does not match, or matches only very late, redoes the bulk of
|
|
its work at every position it tries. Measured directly: `a*b` searched
|
|
over *n* bytes of `a` with no `b` anywhere took 0.32s at *n* = 10,000 and
|
|
19.2s at *n* = 80,000, an approximately 4x slowdown per doubling, the
|
|
signature of `O(n^2)`. This is unrelated to, and in addition to, the
|
|
backreference/lookahead ReDoS caveat below; it applies even to the
|
|
simplest possible non-matching pattern, and is the most significant open
|
|
gap between this build and the "gigabyte scale" objective in `concept.md`
|
|
Section 1 for the search-family operations specifically (`Pattern_match`/
|
|
`Pattern_fullmatch`, which never try more than one position, are not
|
|
affected).
|
|
- A pattern with a backreference or an unbounded-width lookahead can, like
|
|
CPython's own `_sre`, take worst-case exponential time on an adversarial
|
|
input (`concept.md` Section 5, 13.2); this is the same catastrophic
|
|
backtracking (ReDoS) behavior CPython itself exhibits on such patterns, not
|
|
a regression specific to this engine.
|
|
|
|
### Pattern syntax supported
|
|
|
|
Literals; `.` (with `DOTALL`); character classes with ranges, negation, and
|
|
`\d \D \w \W \s \S`; `\b \B`; anchors `^ $ \A \Z` (with `MULTILINE`);
|
|
quantifiers `* + ? {m,n} {m,} {,n} {m}`, greedy and lazy; possessive
|
|
quantifiers `*+ ++ ?+ {m,n}+`; groups `(...) (?:...) (?P<name>...)`;
|
|
alternation `|`; backreferences `\1`-`\99`, `(?P=name)`, `\g<name>`,
|
|
`\g<N>`; lookahead `(?=...) (?!...)`; fixed-width lookbehind
|
|
`(?<=...) (?<!...)`; atomic groups `(?>...)`; comments `(?#...)`; global
|
|
inline flags `(?aiLmsux)` at the start of a pattern; escapes
|
|
`\n \r \t \f \v \a`, octal `\0`-prefixed escapes, `\xhh`, `\uxxxx`,
|
|
`\Uxxxxxxxx`; flags `IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE`, `ASCII`,
|
|
`LOCALE` (all with observable effect; see `docs/API.md` Section 2 for exactly
|
|
what each one does), plus `UNICODE` and `DEBUG` (accepted for source
|
|
compatibility with Python, currently no-ops in this build).
|
|
|
|
Rejected at compile time with a clear `PatternError`, rather than
|
|
mis-parsed: conditional groups `(?(id)yes|no)`, scoped inline flags
|
|
`(?flags:...)`, `\N{NAME}` named code points, and POSIX bracket classes
|
|
`[:alpha:]` (which are not part of Python `re` at all, `concept.md` 14.3).
|
|
Variable-width lookbehind is also rejected at compile time, matching
|
|
CPython.
|
|
|
|
### Operations supported
|
|
|
|
`Pattern_match/fullmatch/search/finditer/findall/split/sub/subn/free`,
|
|
`Pattern_groupindex_lookup`,
|
|
`Match_group/start/end/span/start_byte/end_byte/span_byte/free`,
|
|
`re_compile/match/fullmatch/search/finditer/findall/split/sub/subn/escape/purge`,
|
|
`PatternError_free`. See `regexx.h` for exact signatures, `docs/API.md` for
|
|
the full reference (return values, memory ownership, exact Python
|
|
correspondence for each one), and `concept.md` Section 9 for the naming
|
|
convention they follow.
|
|
|
|
### Known deviations from `concept.md` and from CPython, beyond the items above
|
|
|
|
- `\w`, `\s`, `IGNORECASE` case folding, and `\d` in `UTF8` mode are backed
|
|
by glibc's `wctype.h` functions under the `C.utf8` locale, not by a
|
|
hand-generated Unicode table (`concept.md` 13.3 anticipated a reduced
|
|
static table; using the C library's own tables turned out to be simpler
|
|
and more complete, at the cost of depending on the platform's Unicode
|
|
version rather than a pinned one).
|
|
- `lastindex`/`lastgroup` report the highest-numbered capturing group that
|
|
participated in the match, which coincides with CPython's "most recently
|
|
closed group" rule for straightforward patterns but can differ from it in
|
|
pathological cases (nested alternation re-executing a lower-numbered group
|
|
after a higher one). Not exercised by the test suite; documented here
|
|
rather than silently accepted.
|
|
- `Match_free`, `PatternError_free`, `Input_from_buffer`, `Input_from_file`,
|
|
and `Input_free` have no Python counterpart and are not mentioned in
|
|
`concept.md`'s API surface; they exist because C has no garbage collector.
|
|
`Pattern_sub`/`Pattern_subn` take the replacement template and the
|
|
callback as two separate parameters rather than one polymorphic argument,
|
|
for the same reason (`concept.md` 9.4 already anticipates and justifies
|
|
this one).
|
|
- Python's `Match.start(group)`/`.end(group)` raise `IndexError` for an
|
|
invalid group number and return `-1` only for a valid group that did not
|
|
participate; `Match_start`/`Match_end` return `-1` for both cases, since C
|
|
has no exception to raise. `Match_group` does distinguish them (`-1` for
|
|
no such group, `0` for an unparticipated one), see `docs/API.md` Section
|
|
3.23.
|
|
- `Pattern.groupindex` has no enumeration function in this build, only
|
|
`Pattern_groupindex_lookup(pattern, name)`; there is no way to list every
|
|
name a compiled pattern defines without already knowing what to look for.
|
|
|
|
## Building
|
|
|
|
Requires a C11 compiler and, for the test suite, Python 3 (used only to
|
|
generate ground truth from CPython's own `re` module, `concept.md` Section
|
|
11; the library itself has no runtime dependency beyond the C standard
|
|
library and `libc`'s `wctype.h`/`locale.h`).
|
|
|
|
```sh
|
|
make # builds libregexx.a and the rxgrep example
|
|
make test # regenerates tests/generated_tests.c from Python `re`
|
|
# ground truth and runs the full suite
|
|
make check # same, under AddressSanitizer + UndefinedBehaviorSanitizer
|
|
make clean
|
|
```
|
|
|
|
`make install` installs `libregexx.a` and `regexx.h` under `PREFIX`
|
|
(default `/usr/local`).
|
|
|
|
## Using the library
|
|
|
|
```c
|
|
#include "regexx.h"
|
|
#include <string.h>
|
|
|
|
const char *pattern = "(\\w+)@(\\w+)";
|
|
Pattern *pat = re_compile(pattern, strlen(pattern), UTF8, NULL);
|
|
Input *in = Input_from_buffer((const uint8_t *)"user@host", strlen("user@host"));
|
|
|
|
Match m;
|
|
if (Pattern_search(pat, in, 0, -1, &m) == 1) {
|
|
const char *g; size_t glen;
|
|
Match_group(&m, NULL, 1, &g, &glen); /* g/glen -> "user" */
|
|
Match_free(&m);
|
|
}
|
|
|
|
char *out; size_t outlen;
|
|
Pattern_sub(pat, in, "\\2@\\1", NULL, NULL, 0, &out, &outlen); /* "host@user" */
|
|
free(out);
|
|
|
|
Pattern_free(pat);
|
|
Input_free(in);
|
|
```
|
|
|
|
`flags` to `re_compile` combine a data mode, exactly one of `BINARY`,
|
|
`ASCII`, or `UTF8` (`concept.md` 9.3), with any of the Python-named flags
|
|
(`IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE`, `ASCII` as a flag also
|
|
forces ASCII-only `\w`/`\s`/`\d` inside `UTF8` mode, `LOCALE`, `DEBUG`).
|
|
|
|
## Example: rxgrep
|
|
|
|
`examples/rxgrep.c` is a small grep-like program built on the library,
|
|
demonstrating all three data modes and both the matching and substitution
|
|
API:
|
|
|
|
```sh
|
|
./rxgrep -in 'hello' file.txt # case-insensitive, line numbers
|
|
./rxgrep -m utf8 -o '\w+' file.txt # print every UTF-8 word, one per line
|
|
./rxgrep -c 'error' log.txt # count matching lines
|
|
./rxgrep -m binary 'a.c' data.bin # match raw bytes, embedded NUL included
|
|
./rxgrep -m utf8 --sub 'REDACTED' '\d{3}-\d{4}' file.txt
|
|
```
|
|
|
|
Run `./rxgrep --help` for the full option list.
|
|
|
|
## Testing
|
|
|
|
`tests/cases.py` lists pattern/subject/operation triples. `tests/gen.py`
|
|
computes each one's expected result with CPython's own `re` module and
|
|
writes `tests/generated_tests.c`, which is then compiled against `regexx.c`
|
|
and checked. This is a direct implementation of the strategy `concept.md`
|
|
Section 11 describes: conformance is measured against what CPython actually
|
|
does, not against a re-derived reading of its documentation. `make check`
|
|
additionally runs the suite under AddressSanitizer and
|
|
UndefinedBehaviorSanitizer.
|
|
|
|
## License
|
|
|
|
MIT. See [`LICENSE`](LICENSE).
|