Files
regexx/README.md
T

190 lines
8.8 KiB
Markdown
Raw Normal View History

# regexx
A single-file C regular expression interpreter that reproduces the observable
behavior of Python's `re` module, including its exact identifier names
(`Pattern`, `Match`, `re_compile`, `re_sub`, `IGNORECASE`, and so on; see
Section 9 of `concept.md`), and that additionally supports binary data
(arbitrary byte streams, including embedded NUL) and UTF-8 text alongside
plain ASCII.
The design rationale, the algorithmic trade-offs, and a full accounting of
what is and is not carried over from Python's `re` and from POSIX's native
`regex.h` are recorded in [`concept.md`](concept.md). This file documents the
implementation that exists today at a glance; [`docs/API.md`](docs/API.md) is
the exhaustive reference (every type, every flag, every function's exact
return-value and memory-ownership convention, checked against a real CPython
interpreter, not against memory).
## Implementation status
This is a v1 implementation. It is a complete, tested engine for the pattern
syntax and operations listed below, executed by a single recursive
backtracking engine (`concept.md` Section 7.3) over a fully materialized copy
of the input.
**It does not yet implement the streaming, bounded-memory regular engine of
`concept.md` Section 7.2.** `Input` (the abstraction over "a source of
chunks", `concept.md` 9.2) is implemented, and `Input_from_file` reads a
whole file into memory before matching. Every public function signature is
already exactly what the streaming design in `concept.md` specifies, so the
non-streaming implementation underneath a given call can be replaced later
without changing any caller. Concretely, today:
- A pattern with no backreference and no unbounded-width lookahead runs in
linear time and, for the common case of a single character, class, or `.`
repeated by a quantifier, in *O(1)* recursion depth regardless of input
size (the `OP_REPEAT1` fast path). A repeated *compound* sub-pattern (for
example `(ab)*`) still recurses once per repetition, bounded by a
configurable depth limit (`MAX_DEPTH` in `regexx.c`, currently 60000); past
that limit, matching fails with a reported error rather than a stack
overflow or a wrong answer.
- A pattern with a backreference or an unbounded-width lookahead can, like
CPython's own `_sre`, take worst-case exponential time on an adversarial
input (`concept.md` Section 5, 13.2); this is the same catastrophic
backtracking (ReDoS) behavior CPython itself exhibits on such patterns, not
a regression specific to this engine.
### Pattern syntax supported
Literals; `.` (with `DOTALL`); character classes with ranges, negation, and
`\d \D \w \W \s \S`; `\b \B`; anchors `^ $ \A \Z` (with `MULTILINE`);
quantifiers `* + ? {m,n} {m,} {,n} {m}`, greedy and lazy; possessive
quantifiers `*+ ++ ?+ {m,n}+`; groups `(...) (?:...) (?P<name>...)`;
alternation `|`; backreferences `\1`-`\99`, `(?P=name)`, `\g<name>`,
`\g<N>`; lookahead `(?=...) (?!...)`; fixed-width lookbehind
`(?<=...) (?<!...)`; atomic groups `(?>...)`; comments `(?#...)`; global
inline flags `(?aiLmsux)` at the start of a pattern; escapes
`\n \r \t \f \v \a`, octal `\0`-prefixed escapes, `\xhh`, `\uxxxx`,
`\Uxxxxxxxx`; flags `IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE`, `ASCII`,
`LOCALE` (all with observable effect; see `docs/API.md` Section 2 for exactly
what each one does), plus `UNICODE` and `DEBUG` (accepted for source
compatibility with Python, currently no-ops in this build).
Rejected at compile time with a clear `PatternError`, rather than
mis-parsed: conditional groups `(?(id)yes|no)`, scoped inline flags
`(?flags:...)`, `\N{NAME}` named code points, and POSIX bracket classes
`[:alpha:]` (which are not part of Python `re` at all, `concept.md` 14.3).
Variable-width lookbehind is also rejected at compile time, matching
CPython.
### Operations supported
`Pattern_match/fullmatch/search/finditer/findall/split/sub/subn/free`,
`Pattern_groupindex_lookup`,
`Match_group/start/end/span/start_byte/end_byte/span_byte/free`,
`re_compile/match/fullmatch/search/finditer/findall/split/sub/subn/escape/purge`,
`PatternError_free`. See `regexx.h` for exact signatures, `docs/API.md` for
the full reference (return values, memory ownership, exact Python
correspondence for each one), and `concept.md` Section 9 for the naming
convention they follow.
### Known deviations from `concept.md` and from CPython, beyond the items above
- `\w`, `\s`, `IGNORECASE` case folding, and `\d` in `UTF8` mode are backed
by glibc's `wctype.h` functions under the `C.utf8` locale, not by a
hand-generated Unicode table (`concept.md` 13.3 anticipated a reduced
static table; using the C library's own tables turned out to be simpler
and more complete, at the cost of depending on the platform's Unicode
version rather than a pinned one).
- `lastindex`/`lastgroup` report the highest-numbered capturing group that
participated in the match, which coincides with CPython's "most recently
closed group" rule for straightforward patterns but can differ from it in
pathological cases (nested alternation re-executing a lower-numbered group
after a higher one). Not exercised by the test suite; documented here
rather than silently accepted.
- `Match_free`, `PatternError_free`, `Input_from_buffer`, `Input_from_file`,
and `Input_free` have no Python counterpart and are not mentioned in
`concept.md`'s API surface; they exist because C has no garbage collector.
`Pattern_sub`/`Pattern_subn` take the replacement template and the
callback as two separate parameters rather than one polymorphic argument,
for the same reason (`concept.md` 9.4 already anticipates and justifies
this one).
- Python's `Match.start(group)`/`.end(group)` raise `IndexError` for an
invalid group number and return `-1` only for a valid group that did not
participate; `Match_start`/`Match_end` return `-1` for both cases, since C
has no exception to raise. `Match_group` does distinguish them (`-1` for
no such group, `0` for an unparticipated one), see `docs/API.md` Section
3.23.
- `Pattern.groupindex` has no enumeration function in this build, only
`Pattern_groupindex_lookup(pattern, name)`; there is no way to list every
name a compiled pattern defines without already knowing what to look for.
## Building
Requires a C11 compiler and, for the test suite, Python 3 (used only to
generate ground truth from CPython's own `re` module, `concept.md` Section
11; the library itself has no runtime dependency beyond the C standard
library and `libc`'s `wctype.h`/`locale.h`).
```sh
make # builds libregexx.a and the rxgrep example
make test # regenerates tests/generated_tests.c from Python `re`
# ground truth and runs the full suite
make check # same, under AddressSanitizer + UndefinedBehaviorSanitizer
make clean
```
`make install` installs `libregexx.a` and `regexx.h` under `PREFIX`
(default `/usr/local`).
## Using the library
```c
#include "regexx.h"
#include <string.h>
const char *pattern = "(\\w+)@(\\w+)";
Pattern *pat = re_compile(pattern, strlen(pattern), UTF8, NULL);
Input *in = Input_from_buffer((const uint8_t *)"user@host", strlen("user@host"));
Match m;
if (Pattern_search(pat, in, 0, -1, &m) == 1) {
const char *g; size_t glen;
Match_group(&m, NULL, 1, &g, &glen); /* g/glen -> "user" */
Match_free(&m);
}
char *out; size_t outlen;
Pattern_sub(pat, in, "\\2@\\1", NULL, NULL, 0, &out, &outlen); /* "host@user" */
free(out);
Pattern_free(pat);
Input_free(in);
```
`flags` to `re_compile` combine a data mode, exactly one of `BINARY`,
`ASCII`, or `UTF8` (`concept.md` 9.3), with any of the Python-named flags
(`IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE`, `ASCII` as a flag also
forces ASCII-only `\w`/`\s`/`\d` inside `UTF8` mode, `LOCALE`, `DEBUG`).
## Example: rxgrep
`examples/rxgrep.c` is a small grep-like program built on the library,
demonstrating all three data modes and both the matching and substitution
API:
```sh
./rxgrep -in 'hello' file.txt # case-insensitive, line numbers
./rxgrep -m utf8 -o '\w+' file.txt # print every UTF-8 word, one per line
./rxgrep -c 'error' log.txt # count matching lines
./rxgrep -m binary 'a.c' data.bin # match raw bytes, embedded NUL included
./rxgrep -m utf8 --sub 'REDACTED' '\d{3}-\d{4}' file.txt
```
Run `./rxgrep --help` for the full option list.
## Testing
`tests/cases.py` lists pattern/subject/operation triples. `tests/gen.py`
computes each one's expected result with CPython's own `re` module and
writes `tests/generated_tests.c`, which is then compiled against `regexx.c`
and checked. This is a direct implementation of the strategy `concept.md`
Section 11 describes: conformance is measured against what CPython actually
does, not against a re-derived reading of its documentation. `make check`
additionally runs the suite under AddressSanitizer and
UndefinedBehaviorSanitizer.
## License
MIT. See [`LICENSE`](LICENSE).