Profiled a real search with Massif and found the memory-per-input-byte multiplier at 13.4x, dominated by two avoidable costs: - build_matbuf widened every BINARY/ASCII byte into a uint32_t before matching, a 4x copy that mode never needed (a byte never exceeds 255). Removed it: MCtx/MatBuf now carry an optional text8 (borrowed, unwidened) alongside the existing UTF8 text (owned, decoded code points), reconciled per-read through text_at()/buf_at(). BINARY/ ASCII mode now points directly into the Input's own buffer. - Input_from_file always copied the whole file into a malloc'd buffer. It now mmaps regular files read-only (MAP_PRIVATE) instead, so pages stay clean and reclaimable under memory pressure and the file is never copied. Non-seekable sources (pipes, FIFOs, process substitution, stdin) and mmap failures fall back to the previous incremental-read behavior via a separate read_fd_incrementally. Together these bring the multiplier to 8.4x and let a 200MB file that previously OOM-crashed complete a full non-matching search in about 6 seconds at roughly 1.69GB peak RSS. Also tried, measured, and reverted: capping compute_maxrun/ compute_next_prevmatch's table size with a plain-scan fallback above the cap. A real 200MB non-matching search against this fallback hung for minutes instead of failing fast, because disabling either table reintroduces the O(n^2) behavior they exist to prevent, and O(n^2) at n in the hundreds of millions is not practically finite. A fast, diagnosable allocation failure is a better failure mode than a silent, unbounded hang, so the tables are allocated unconditionally again; the finding is recorded in code comments, concept.md 7.6, and README's "Memory footprint" section so it is not retried blindly later. Separately audited every allocation on an input-proportional path (da_push, build_matbuf's UTF-8 decode loop, all three Input_from_file sites) and made each fail cleanly through PatternError instead of crashing on an unchecked NULL dereference. A 1GB file still exceeds available memory in the current environment; this is a property of the machine it was measured on, not a defect, and is documented as such (practical ceiling: available memory / 8.4 for search-family operations, pending the streaming automaton design in concept.md 7.2). Verified with four clean `make test` passes (3252/3252) and a clean ASan/UBSan pass after the change; rxgrep's mmap-backed paths (--sub with a regular file, with a non-seekable process-substitution source, and BINARY-mode embedded NUL handling) re-checked directly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
393 lines
22 KiB
Markdown
393 lines
22 KiB
Markdown
# regexx
|
|
|
|
A single-file C regular expression interpreter that reproduces the observable
|
|
behavior of Python's `re` module, including its exact identifier names
|
|
(`Pattern`, `Match`, `re_compile`, `re_sub`, `IGNORECASE`, and so on; see
|
|
Section 9 of `concept.md`), and that additionally supports binary data
|
|
(arbitrary byte streams, including embedded NUL) and UTF-8 text alongside
|
|
plain ASCII.
|
|
|
|
The design rationale, the algorithmic trade-offs, and a full accounting of
|
|
what is and is not carried over from Python's `re` and from POSIX's native
|
|
`regex.h` are recorded in [`concept.md`](concept.md). This file documents the
|
|
implementation that exists today at a glance; [`docs/API.md`](docs/API.md) is
|
|
the exhaustive reference (every type, every flag, every function's exact
|
|
return-value and memory-ownership convention, checked against a real CPython
|
|
interpreter, not against memory).
|
|
|
|
## Implementation status
|
|
|
|
This is a v1 implementation. It is a complete, tested engine for the pattern
|
|
syntax and operations listed below, executed by a single recursive
|
|
backtracking engine (`concept.md` Section 7.3) over a fully materialized copy
|
|
of the input.
|
|
|
|
**It does not yet implement the streaming, bounded-memory regular engine of
|
|
`concept.md` Section 7.2.** `Input` (the abstraction over "a source of
|
|
chunks", `concept.md` 9.2) is implemented, and `Input_from_file` maps or
|
|
reads a whole file into memory before matching (see "Memory footprint"
|
|
below for exactly how, and for the measured numbers this and other fixes
|
|
were checked against). Every public function signature is already exactly
|
|
what the streaming design in `concept.md` specifies, so the non-streaming
|
|
implementation underneath a given call can be replaced later without
|
|
changing any caller. Concretely, today:
|
|
|
|
- `Pattern_match`/`Pattern_fullmatch` (a single anchored attempt at a fixed
|
|
position) run in time proportional to the length of that attempt, and, for
|
|
the common case of a single character, class, or `.` repeated by a
|
|
quantifier, in *O(1)* recursion depth regardless of input size (the
|
|
`OP_REPEAT1` fast path). A repeated *compound* sub-pattern (for example
|
|
`(ab)*`) still recurses once per repetition, bounded by a configurable
|
|
depth limit (`MAX_DEPTH` in `regexx.c`, currently 60000, not exposed
|
|
through the public API, so changing it means editing `regexx.c` and
|
|
rebuilding); past that limit, matching fails with a reported error rather
|
|
than a stack overflow or a wrong answer.
|
|
- **`Pattern_search`/`Pattern_finditer`/`Pattern_split`/`Pattern_sub` run in
|
|
linear time for the common case: a pattern built only from simple,
|
|
non-backreference constructs where every quantifier's body is a single
|
|
character, class, or `.`** (`a*b`, `\s*\d+`, `.*"`, and the great majority
|
|
of patterns actually written by hand). This was not always true of this
|
|
build; measuring it directly is what caught that it previously was not
|
|
(an earlier revision of this file claimed linear time without having
|
|
re-verified it, then had to correct that claim, then fixed the underlying
|
|
defect; the corrected numbers are below). Two independent techniques make
|
|
it true now, neither of them the streaming engine of `concept.md` Section
|
|
7.2, which remains unimplemented:
|
|
1. **Memoized backtracking.** For any pattern with no `OP_BACKREF`
|
|
anywhere, whether execution starting at a given (instruction, text
|
|
position) pair can ever reach a match is a fact that never changes
|
|
once computed, so `run_memo` (`regexx.c`) caches every *failure* (never
|
|
a success, so it cannot change which match is found, only skip
|
|
re-deriving failures already known) and consults the cache before
|
|
redoing that work. This is a published technique, not a house
|
|
invention: a memoization table that records only failures, because
|
|
"a matching success ... immediately propagates to the success of the
|
|
whole problem," is exactly the scheme described in recent work on
|
|
backtracking regex matchers ("Selective Memoization for Efficient
|
|
Backtracking Regular Expression Matching", and, for the
|
|
lookaround/atomic case specifically, "Efficient Matching with
|
|
Memoization for Regexes with Look-around and Atomic Grouping", both
|
|
linked below).
|
|
2. **Precomputed run lengths and skip-ahead tables for `OP_REPEAT1`.**
|
|
Memoizing individual `(instruction, position)` pairs does not help
|
|
when each one is already `O(1)` work, which is exactly `a*b`'s
|
|
situation: counting how many `a`s follow a position, and then trying
|
|
every possible split point against the trailing `b` one at a time,
|
|
are each individually cheap but happen `O(remaining length)` times per
|
|
start position tried. `compute_maxrun` precomputes, once per
|
|
`Pattern_search`/`finditer`/`split` call, how many characters each
|
|
repeat can consume from every position in one backward pass, and
|
|
`compute_next_prevmatch` precomputes, for a repeat immediately
|
|
followed by a single literal/class/`.`, the rightmost position at or
|
|
before any given point where that next atom can match, so the
|
|
backtrack loop jumps directly to candidates worth trying instead of
|
|
visiting every position in between. This is the same idea production
|
|
engines call a literal prefilter (RE2's and Rust's `regex` crate's
|
|
`memchr`/`memmem`/Teddy prefilters skip positions that provably cannot
|
|
match before ever invoking the full engine); those use SIMD-accelerated
|
|
library primitives operating directly on bytes, this build uses a
|
|
precomputed array, which is slower per lookup but the same
|
|
algorithmic idea, and appropriate for `concept.md`'s stated priority of
|
|
convenience over performance.
|
|
|
|
Measured directly: `a*b` searched over *n* bytes of `a` with no `b`
|
|
anywhere took 19.2s at *n* = 80,000 before these two techniques and
|
|
0.0016s after, with time now scaling linearly (2x per doubling of *n*)
|
|
rather than quadratically (4x per doubling).
|
|
|
|
Cost: these tables use `O(k * n)` memory, `k` being the number of
|
|
`OP_REPEAT1` instructions actually present in the compiled pattern,
|
|
`n` the input length; `Pattern_match`/`Pattern_fullmatch` never
|
|
allocate them (a single anchored attempt gets no benefit from them).
|
|
- **A pattern with a backreference can, like CPython's own `_sre`, still
|
|
take worst-case exponential time on an adversarial input** (`concept.md`
|
|
Section 5, 13.2; memoization above is unsound and therefore disabled
|
|
whenever a pattern contains `OP_BACKREF` anywhere, since a backreference's
|
|
outcome depends on capture history, not on position alone). This is the
|
|
same catastrophic backtracking (ReDoS) behavior CPython itself exhibits on
|
|
such patterns, not a regression specific to this engine; a pattern author
|
|
who needs to rule it out for a specific pattern can use an atomic group
|
|
(`(?>...)`) or a possessive quantifier around the ambiguous repetition,
|
|
exactly as they would for CPython's `re`, and exactly as general ReDoS
|
|
mitigation guidance recommends (linked below).
|
|
- **A backreference-free pattern shaped like nested, overlapping
|
|
quantifiers (`(a+)+b`, the textbook ReDoS shape) is no longer
|
|
exponential, because it has no backreference and so is memoized, but is
|
|
not necessarily linear either.** Measured directly: `(a+)+b` against *n*
|
|
characters of `a` with no trailing `b` took over a minute already at
|
|
*n* = 40 before memoization (exponential; this build's own stress test
|
|
keeps that case at *n* = 24 for exactly this reason) and 3.2s at
|
|
*n* = 32,000 after (empirically quadratic: roughly 4x per doubling),
|
|
because the inner `a+`'s `OP_REPEAT1` is followed by the group's closing
|
|
save, not by a simple atom, so the skip-ahead technique above does not
|
|
apply to it, only the memoization does. Exponential to quadratic is a
|
|
qualitative change (the *n* = 40 case alone went from longer than this
|
|
project could wait for, to instant), not a complete fix; a pattern
|
|
actually meant to run unattended against adversarial input should still
|
|
avoid this shape, or wrap the inner repetition in an atomic group,
|
|
regardless of the improvement.
|
|
|
|
Further reading on the techniques above: Russ Cox, ["Regular Expression
|
|
Matching: the Virtual Machine
|
|
Approach"](https://swtch.com/~rsc/regexp/regexp2.html) (why prepending
|
|
`.*?` gives linear-time unanchored search only in a Thompson/Pike VM, not in
|
|
a backtracking engine, which is why this build needed a different fix);
|
|
["Selective Memoization for Efficient Backtracking Regular Expression
|
|
Matching"](https://arxiv.org/html/2606.26678) and ["Efficient Matching with
|
|
Memoization for Regexes with Look-around and Atomic
|
|
Grouping"](https://arxiv.org/pdf/2401.12639) (the failure-only memoization
|
|
scheme this build's `run_memo` implements, and a more memory-efficient
|
|
selective variant, memoizing only at loop "feedback nodes" rather than every
|
|
instruction, that this build does not implement but could); the [Snyk
|
|
writeup on ReDoS and catastrophic
|
|
backtracking](https://snyk.io/blog/redos-and-catastrophic-backtracking/) for
|
|
the general phenomenon and mitigation guidance.
|
|
|
|
### Memory footprint
|
|
|
|
Measured directly with Valgrind/Massif (a real 10MB search) and by watching
|
|
`VmRSS`/`VmHWM` on real 200MB-1GB files, not estimated from reading the
|
|
code. `Pattern_search`/`finditer`/`split`/`sub` (the operations that try
|
|
more than one start position) currently use, at peak, about **8.4x** the
|
|
input length in memory for a pattern using `OP_REPEAT1`
|
|
(`concept.md` 7.5's `compute_maxrun` and `compute_next_prevmatch` tables,
|
|
`int32_t`-per-input-position each, are the entire remaining cost: 47.75%
|
|
each in the Massif profile, `alloc_memo`'s bitset a further 4.5%).
|
|
`Pattern_match`/`fullmatch` (a single attempt, no search tables) use
|
|
proportionally less.
|
|
|
|
Two real issues were found and fixed getting to that number, in order:
|
|
|
|
1. **`build_matbuf` widened every byte to a 4-byte `uint32_t`, even in
|
|
`BINARY`/`ASCII` mode, where a byte never exceeds 255 and the widening
|
|
bought nothing.** This cost as much extra memory as the input itself,
|
|
four times over, unconditionally, on top of the search tables above.
|
|
Fixed: `BINARY`/`ASCII` mode now reads the input's own bytes directly
|
|
(`MatBuf`/`MCtx`'s `text8` field, `text_at()`/`buf_at()` in `regexx.c`);
|
|
only `UTF8` mode still widens, because it actually needs code points up
|
|
to `0x10FFFF`, which do not fit in a byte. This dropped the measured
|
|
10MB-search peak from 140.3MB (13.4x) to 87.8MB (8.4x).
|
|
2. **`Input_from_file` read every file into a fresh, private, `malloc`'d
|
|
copy, even though the OS's page cache already holds the file's bytes.**
|
|
For a regular, seekable, non-empty file this now uses `mmap()`
|
|
(`PROT_READ`, `MAP_PRIVATE`) instead: the mapped pages are backed
|
|
directly by the file and stay clean (never written), so the kernel can
|
|
reclaim them under memory pressure and re-fault them in from disk later,
|
|
rather than them being pinned for the whole match attempt the way a
|
|
`malloc`'d copy is; it also removes one whole redundant copy of the
|
|
file's bytes. Falls back to the previous `read()`-based incremental
|
|
copy for anything `mmap` does not apply to (a pipe, a FIFO, process
|
|
substitution, stdin, an empty file, or an `mmap()` call that itself
|
|
fails).
|
|
|
|
A **third fix attempt was tried, measured, and reverted** specifically
|
|
because "prevent OOM, keep the footprint small" turned out to have a
|
|
sharp edge worth recording: capping `compute_maxrun`/`compute_next_prevmatch`
|
|
above a size budget and falling back to the plain scan already used when
|
|
either table is `NULL` seemed like an obvious bounded-memory safety valve.
|
|
Measured directly against a real 200MB non-matching search, it was worse
|
|
than doing nothing: both tables are needed together to keep this pattern
|
|
shape (`x*y`-style, unbounded quantifier followed by a required literal
|
|
that never occurs) at linear time; disabling either one alone reintroduces
|
|
the `O(n^2)` behavior they exist to fix, and `O(n^2)` at `n` in the hundreds
|
|
of millions does not finish in any practical amount of time. A fast,
|
|
diagnosable allocation failure (see below) is a better failure mode than a
|
|
silent, effectively-unbounded hang, so the cap was removed; these two
|
|
tables are allocated unconditionally again. There is no way to get both
|
|
bounded memory and linear time out of this technique for this pattern
|
|
shape; only `concept.md` Section 7.2's actual streaming automaton (still
|
|
unimplemented) gets both at once, by construction, which is why it remains
|
|
the correct long-term fix for this axis specifically.
|
|
|
|
**Failing safely.** Every allocation on the input-proportional paths above
|
|
(`build_matbuf`, `compute_maxrun`, `compute_next_prevmatch`,
|
|
`Input_from_file`, the UTF-8 decode arrays) is now checked; a failure
|
|
returns a `PatternError`/`-1` through the ordinary error path instead of
|
|
crashing on a `NULL` dereference, which several of them did before this
|
|
was audited (found by deliberately reasoning through "what happens when
|
|
this specific `malloc` fails on a huge request", not by a tool). This does
|
|
not prevent an out-of-memory condition on a genuinely memory-constrained
|
|
machine; the operating system's OOM killer can still end the process for
|
|
an allocation this library made in good faith (`malloc`/`mmap` returning
|
|
`NULL`/`MAP_FAILED` is the case this library can catch; being killed by
|
|
the kernel before that happens is not something a userspace library can
|
|
intercept). Measured concretely on the machine this was developed on: a
|
|
200MB file search that previously crashed via the OOM killer now completes
|
|
successfully in about 6 seconds at roughly 1.7GB peak RSS; a 1GB file on
|
|
the same machine still exceeded what was available at the time. Both
|
|
numbers are specific to that machine's available memory at the time, not
|
|
a hard property of the library; the 8.4x multiplier above is what actually
|
|
determines the practical ceiling on a given machine (roughly
|
|
`available memory / 8.4` for search-family operations on a pattern using
|
|
`OP_REPEAT1`, more forgiving for `match`/`fullmatch` or for patterns
|
|
without a simple-atom quantifier at all).
|
|
|
|
### Pattern syntax supported
|
|
|
|
Literals; `.` (with `DOTALL`); character classes with ranges, negation, and
|
|
`\d \D \w \W \s \S`; `\b \B`; anchors `^ $ \A \Z` (with `MULTILINE`);
|
|
quantifiers `* + ? {m,n} {m,} {,n} {m}`, greedy and lazy; possessive
|
|
quantifiers `*+ ++ ?+ {m,n}+`; groups `(...) (?:...) (?P<name>...)`;
|
|
alternation `|`; backreferences `\1`-`\99`, `(?P=name)`, `\g<name>`,
|
|
`\g<N>`; lookahead `(?=...) (?!...)`; fixed-width lookbehind
|
|
`(?<=...) (?<!...)`; atomic groups `(?>...)`; comments `(?#...)`; global
|
|
inline flags `(?aiLmsux)` at the start of a pattern; escapes
|
|
`\n \r \t \f \v \a`, octal `\0`-prefixed escapes, `\xhh`, `\uxxxx`,
|
|
`\Uxxxxxxxx`; flags `IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE`, `ASCII`
|
|
(all with observable effect; see `docs/API.md` Section 2 for exactly what
|
|
each one does), plus `UNICODE`, `LOCALE`, and `DEBUG` (accepted for source
|
|
compatibility with Python, currently no-ops in this build: `LOCALE` because
|
|
this build commits to the "C" locale only, under which `\w`/`\b`/`\B`
|
|
classify only ASCII letters and digits regardless of the flag, verified
|
|
directly against both the C standard and a real CPython interpreter).
|
|
|
|
Rejected at compile time with a clear `PatternError`, rather than
|
|
mis-parsed: conditional groups `(?(id)yes|no)`, scoped inline flags
|
|
`(?flags:...)`, `\N{NAME}` named code points, and POSIX bracket classes
|
|
`[:alpha:]` (which are not part of Python `re` at all, `concept.md` 14.3).
|
|
Variable-width lookbehind is also rejected at compile time, matching
|
|
CPython.
|
|
|
|
### Operations supported
|
|
|
|
`Pattern_match/fullmatch/search/finditer/findall/split/sub/subn/free`,
|
|
`Pattern_groupindex_lookup`,
|
|
`Match_group/start/end/span/start_byte/end_byte/span_byte/free`,
|
|
`re_compile/match/fullmatch/search/finditer/findall/split/sub/subn/escape/purge`,
|
|
`PatternError_free`. See `regexx.h` for exact signatures, `docs/API.md` for
|
|
the full reference (return values, memory ownership, exact Python
|
|
correspondence for each one), and `concept.md` Section 9 for the naming
|
|
convention they follow.
|
|
|
|
### Known deviations from `concept.md` and from CPython, beyond the items above
|
|
|
|
- `\w`, `\s`, `IGNORECASE` case folding, and `\d` in `UTF8` mode are backed
|
|
by glibc's `wctype.h` functions under the `C.utf8` locale, not by a
|
|
hand-generated Unicode table (`concept.md` 13.3 anticipated a reduced
|
|
static table; using the C library's own tables turned out to be simpler
|
|
and more complete, at the cost of depending on the platform's Unicode
|
|
version rather than a pinned one). Concretely verified, not just
|
|
theoretical: U+00A0 (NO-BREAK SPACE) is in Unicode's `White_Space`
|
|
property, so CPython's `\s` matches it, but glibc's `iswspace()` under
|
|
`C.utf8` does not, so this build's `\s` does not either. Found by the
|
|
large combinatorial test expansion (`tests/cases.py` Category F,
|
|
`tests/TEST_PLAN.md`), not anticipated in advance; recorded here rather
|
|
than patched, since hand-patching individual code points would start
|
|
down the path of maintaining an ad hoc table this design deliberately
|
|
avoided by delegating to `wctype.h` in the first place. A second,
|
|
same-class instance was found by the later Category I expansion:
|
|
fullwidth digits (`U+FF10`-`U+FF19`, Unicode category `Nd`) match
|
|
CPython's `\d` but not glibc's `iswdigit()` under `C.utf8` either.
|
|
- `lastindex`/`lastgroup` report the highest-numbered capturing group that
|
|
participated in the match, which coincides with CPython's "most recently
|
|
closed group" rule for straightforward patterns but can differ from it in
|
|
pathological cases (nested alternation re-executing a lower-numbered group
|
|
after a higher one). Not exercised by the test suite; documented here
|
|
rather than silently accepted.
|
|
- `Match_free`, `PatternError_free`, `Input_from_buffer`, `Input_from_file`,
|
|
and `Input_free` have no Python counterpart and are not mentioned in
|
|
`concept.md`'s API surface; they exist because C has no garbage collector.
|
|
`Pattern_sub`/`Pattern_subn` take the replacement template and the
|
|
callback as two separate parameters rather than one polymorphic argument,
|
|
for the same reason (`concept.md` 9.4 already anticipates and justifies
|
|
this one).
|
|
- Python's `Match.start(group)`/`.end(group)` raise `IndexError` for an
|
|
invalid group number and return `-1` only for a valid group that did not
|
|
participate; `Match_start`/`Match_end` return `-1` for both cases, since C
|
|
has no exception to raise. `Match_group` does distinguish them (`-1` for
|
|
no such group, `0` for an unparticipated one), see `docs/API.md` Section
|
|
3.23.
|
|
- `Pattern.groupindex` has no enumeration function in this build, only
|
|
`Pattern_groupindex_lookup(pattern, name)`; there is no way to list every
|
|
name a compiled pattern defines without already knowing what to look for.
|
|
|
|
## Building
|
|
|
|
Requires a C11 compiler and, for the test suite, Python 3 (used only to
|
|
generate ground truth from CPython's own `re` module, `concept.md` Section
|
|
11; the library itself has no runtime dependency beyond the C standard
|
|
library and `libc`'s `wctype.h`/`locale.h`).
|
|
|
|
```sh
|
|
make # builds libregexx.a and the rxgrep example
|
|
make test # regenerates tests/generated_tests.c from Python `re`
|
|
# ground truth and runs the full suite
|
|
make check # same, under AddressSanitizer + UndefinedBehaviorSanitizer
|
|
make clean
|
|
```
|
|
|
|
`make install` installs `libregexx.a` and `regexx.h` under `PREFIX`
|
|
(default `/usr/local`).
|
|
|
|
## Using the library
|
|
|
|
```c
|
|
#include "regexx.h"
|
|
#include <string.h>
|
|
|
|
const char *pattern = "(\\w+)@(\\w+)";
|
|
Pattern *pat = re_compile(pattern, strlen(pattern), UTF8, NULL);
|
|
Input *in = Input_from_buffer((const uint8_t *)"user@host", strlen("user@host"));
|
|
|
|
Match m;
|
|
if (Pattern_search(pat, in, 0, -1, &m) == 1) {
|
|
const char *g; size_t glen;
|
|
Match_group(&m, NULL, 1, &g, &glen); /* g/glen -> "user" */
|
|
Match_free(&m);
|
|
}
|
|
|
|
char *out; size_t outlen;
|
|
Pattern_sub(pat, in, "\\2@\\1", NULL, NULL, 0, &out, &outlen); /* "host@user" */
|
|
free(out);
|
|
|
|
Pattern_free(pat);
|
|
Input_free(in);
|
|
```
|
|
|
|
`flags` to `re_compile` combine a data mode, exactly one of `BINARY`,
|
|
`ASCII`, or `UTF8` (`concept.md` 9.3), with any of the Python-named flags
|
|
(`IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE`, `ASCII` as a flag also
|
|
forces ASCII-only `\w`/`\s`/`\d` inside `UTF8` mode, `LOCALE`, `DEBUG`).
|
|
|
|
## Example: rxgrep
|
|
|
|
`examples/rxgrep.c` is a small grep-like program built on the library,
|
|
demonstrating all three data modes and both the matching and substitution
|
|
API:
|
|
|
|
```sh
|
|
./rxgrep -in 'hello' file.txt # case-insensitive, line numbers
|
|
./rxgrep -m utf8 -o '\w+' file.txt # print every UTF-8 word, one per line
|
|
./rxgrep -c 'error' log.txt # count matching lines
|
|
./rxgrep -m binary 'a.c' data.bin # match raw bytes, embedded NUL included
|
|
./rxgrep -m utf8 --sub 'REDACTED' '\d{3}-\d{4}' file.txt
|
|
```
|
|
|
|
Run `./rxgrep --help` for the full option list.
|
|
|
|
## Testing
|
|
|
|
`tests/cases.py` lists pattern/subject/operation triples, both hand-written
|
|
and, for most of the file, generated programmatically from combinations of
|
|
quantifiers, groups, backreferences, lookaround, flags, and encoding modes
|
|
(`tests/TEST_PLAN.md` records the exact category breakdown and why each one
|
|
exists). `tests/gen.py` computes every case's expected result with
|
|
CPython's own `re` module and writes `tests/generated_tests.c`, which is
|
|
then compiled against `regexx.c` and checked. This is a direct
|
|
implementation of the strategy `concept.md` Section 11 describes:
|
|
conformance is measured against what CPython actually does, not against a
|
|
re-derived reading of its documentation, at a scale (3,252 cases as of this
|
|
writing, `make test` reports the current exact count) large enough that it
|
|
has already found real defects this way, not only confirmed the absence of
|
|
ones anyone thought to write by hand (`tests/TEST_PLAN.md` "Result" names
|
|
all four: a C trigraph bug in the test generator itself, a real, previously
|
|
undocumented `Pattern_finditer`/`Pattern_split` empty-match defect, a wrong
|
|
ground truth for `ASCII` mode in the generator itself, and a real,
|
|
verified-wrong implementation of the `LOCALE` flag). `make check`
|
|
additionally runs the suite under AddressSanitizer and
|
|
UndefinedBehaviorSanitizer.
|
|
|
|
## License
|
|
|
|
MIT. See [`LICENSE`](LICENSE).
|