README.md and docs/API.md both previously documented "no way to list every name a pattern defines without already knowing what to look for" as an accepted limitation of Pattern_groupindex_lookup being the only access to groupindex. It was not actually a hard constraint: the underlying GroupIndex struct (regexx.c) already stores every name and its group number in two parallel arrays, populated once at compile time; only a public accessor was missing. Pattern_groupindex_count returns the number of named groups; Pattern_groupindex_at(self, i, &name) for 0 <= i < count writes the i-th name and returns its 1-based group number, or returns -1 for an out-of-range i. Enumeration order is declaration order, verified against a real CPython 3.11 interpreter to match groupindex's own practical (insertion-order) iteration order, not just assumed. Verified directly (count/name/group-number correctness, matching the exact snippet now in USAGE.md's own output), full 3,252-case suite unaffected (3252/3252, this is a pure accessor addition touching no matching logic), clean AddressSanitizer/UndefinedBehaviorSanitizer. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
515 lines
30 KiB
Markdown
515 lines
30 KiB
Markdown
# regexx
|
|
|
|
A single-file C regular expression interpreter that reproduces the observable
|
|
behavior of Python's `re` module, including its exact identifier names
|
|
(`Pattern`, `Match`, `re_compile`, `re_sub`, `IGNORECASE`, and so on; see
|
|
Section 9 of `concept.md`), and that additionally supports binary data
|
|
(arbitrary byte streams, including embedded NUL) and UTF-8 text alongside
|
|
plain ASCII.
|
|
|
|
The design rationale, the algorithmic trade-offs, and a full accounting of
|
|
what is and is not carried over from Python's `re` and from POSIX's native
|
|
`regex.h` are recorded in [`concept.md`](concept.md). This file documents the
|
|
implementation that exists today at a glance; [`docs/API.md`](docs/API.md) is
|
|
the exhaustive reference (every type, every flag, every function's exact
|
|
return-value and memory-ownership convention, checked against a real CPython
|
|
interpreter, not against memory); [`USAGE.md`](USAGE.md) is the task-oriented
|
|
guide (compiled, run, and verified examples covering every operation across
|
|
all three data modes, from "nothing" to "working code").
|
|
|
|
## Implementation status
|
|
|
|
This is a v1 implementation. It is a complete, tested engine for the pattern
|
|
syntax and operations listed below, executed by one of two engines over a
|
|
fully materialized copy of the input, chosen automatically and transparently
|
|
per compiled pattern, never something a caller selects:
|
|
|
|
- **A Pike VM** (`concept.md` 7.2/7.7, a Thompson-NFA simulation with
|
|
per-thread capture tracking), used whenever a pattern contains none of a
|
|
backreference, a lookahead, a lookbehind, or an atomic group/possessive
|
|
quantifier (the last of which is already desugared to an atomic group at
|
|
parse time, `concept.md` Section 4). This is the majority of patterns
|
|
people actually write by hand, and this engine runs every one of them in
|
|
genuinely linear time, not merely the practically-linear-for-the-common-
|
|
case behavior described below for the other engine, because it does not
|
|
backtrack at all: it holds a bounded set of live parse states and steps
|
|
them all forward together, one input position at a time.
|
|
- **A recursive backtracking engine** (`concept.md` Section 7.3), used only
|
|
for the minority of patterns that need a construct the Pike VM cannot
|
|
execute at all: a backreference (not a regular-language construct; no
|
|
Thompson NFA, however built, can execute one), or a lookahead/lookbehind/
|
|
atomic group (each compiles to a self-contained sub-program executed to a
|
|
single yes/no/where-it-ended answer, which needs recursive call structure
|
|
the flat thread-priority simulation does not have, `concept.md` 7.7).
|
|
|
|
**Neither engine yet implements the streaming, bounded-memory version of
|
|
`concept.md` Section 7.2.** `Input` (the abstraction over "a source of
|
|
chunks", `concept.md` 9.2) is implemented, and `Input_from_file` maps or
|
|
reads a whole file into memory before matching (see "Memory footprint"
|
|
below for exactly how, and for the measured numbers this and other fixes
|
|
were checked against). Every public function signature is already exactly
|
|
what the streaming design in `concept.md` specifies, so the non-streaming
|
|
implementation underneath a given call can be replaced later without
|
|
changing any caller. Concretely, today, for the backtracking engine (see
|
|
"The Pike VM" below for the other engine's own, separate numbers):
|
|
|
|
- `Pattern_match`/`Pattern_fullmatch` (a single anchored attempt at a fixed
|
|
position) run in time proportional to the length of that attempt, and, for
|
|
the common case of a single character, class, or `.` repeated by a
|
|
quantifier, in *O(1)* recursion depth regardless of input size (the
|
|
`OP_REPEAT1` fast path). A repeated *compound* sub-pattern (for example
|
|
`(ab)*`) still recurses once per repetition, bounded by a configurable
|
|
depth limit (`MAX_DEPTH` in `regexx.c`, currently 60000, not exposed
|
|
through the public API, so changing it means editing `regexx.c` and
|
|
rebuilding); past that limit, matching fails with a reported error rather
|
|
than a stack overflow or a wrong answer.
|
|
- **`Pattern_search`/`Pattern_finditer`/`Pattern_split`/`Pattern_sub` run in
|
|
linear time for the common case: a pattern built only from simple,
|
|
non-backreference constructs where every quantifier's body is a single
|
|
character, class, or `.`** (`a*b`, `\s*\d+`, `.*"`, and the great majority
|
|
of patterns actually written by hand). This was not always true of this
|
|
build; measuring it directly is what caught that it previously was not
|
|
(an earlier revision of this file claimed linear time without having
|
|
re-verified it, then had to correct that claim, then fixed the underlying
|
|
defect; the corrected numbers are below). Two independent techniques make
|
|
it true now, neither of them the streaming engine of `concept.md` Section
|
|
7.2, which remains unimplemented:
|
|
1. **Memoized backtracking.** For any pattern with no `OP_BACKREF`
|
|
anywhere, whether execution starting at a given (instruction, text
|
|
position) pair can ever reach a match is a fact that never changes
|
|
once computed, so `run_memo` (`regexx.c`) caches every *failure* (never
|
|
a success, so it cannot change which match is found, only skip
|
|
re-deriving failures already known) and consults the cache before
|
|
redoing that work. This is a published technique, not a house
|
|
invention: a memoization table that records only failures, because
|
|
"a matching success ... immediately propagates to the success of the
|
|
whole problem," is exactly the scheme described in recent work on
|
|
backtracking regex matchers ("Selective Memoization for Efficient
|
|
Backtracking Regular Expression Matching", and, for the
|
|
lookaround/atomic case specifically, "Efficient Matching with
|
|
Memoization for Regexes with Look-around and Atomic Grouping", both
|
|
linked below).
|
|
2. **Precomputed run lengths and skip-ahead tables for `OP_REPEAT1`.**
|
|
Memoizing individual `(instruction, position)` pairs does not help
|
|
when each one is already `O(1)` work, which is exactly `a*b`'s
|
|
situation: counting how many `a`s follow a position, and then trying
|
|
every possible split point against the trailing `b` one at a time,
|
|
are each individually cheap but happen `O(remaining length)` times per
|
|
start position tried. `compute_maxrun` precomputes, once per
|
|
`Pattern_search`/`finditer`/`split` call, how many characters each
|
|
repeat can consume from every position in one backward pass, and
|
|
`compute_next_prevmatch` precomputes, for a repeat immediately
|
|
followed by a single literal/class/`.`, the rightmost position at or
|
|
before any given point where that next atom can match, so the
|
|
backtrack loop jumps directly to candidates worth trying instead of
|
|
visiting every position in between. This is the same idea production
|
|
engines call a literal prefilter (RE2's and Rust's `regex` crate's
|
|
`memchr`/`memmem`/Teddy prefilters skip positions that provably cannot
|
|
match before ever invoking the full engine); those use SIMD-accelerated
|
|
library primitives operating directly on bytes, this build uses a
|
|
precomputed array, which is slower per lookup but the same
|
|
algorithmic idea, and appropriate for `concept.md`'s stated priority of
|
|
convenience over performance.
|
|
|
|
Measured directly: `a*b` searched over *n* bytes of `a` with no `b`
|
|
anywhere took 19.2s at *n* = 80,000 before these two techniques and
|
|
0.0016s after, with time now scaling linearly (2x per doubling of *n*)
|
|
rather than quadratically (4x per doubling).
|
|
|
|
Cost: these tables use `O(k * n)` memory, `k` being the number of
|
|
`OP_REPEAT1` instructions actually present in the compiled pattern,
|
|
`n` the input length; `Pattern_match`/`Pattern_fullmatch` never
|
|
allocate them (a single anchored attempt gets no benefit from them).
|
|
- **A pattern with a backreference can, like CPython's own `_sre`, still
|
|
take worst-case exponential time on an adversarial input** (`concept.md`
|
|
Section 5, 13.2; memoization above is unsound and therefore disabled
|
|
whenever a pattern contains `OP_BACKREF` anywhere, since a backreference's
|
|
outcome depends on capture history, not on position alone). This is the
|
|
same catastrophic backtracking (ReDoS) behavior CPython itself exhibits on
|
|
such patterns, not a regression specific to this engine; a pattern author
|
|
who needs to rule it out for a specific pattern can use an atomic group
|
|
(`(?>...)`) or a possessive quantifier around the ambiguous repetition,
|
|
exactly as they would for CPython's `re`, and exactly as general ReDoS
|
|
mitigation guidance recommends (linked below).
|
|
- **A backreference-free pattern shaped like nested, overlapping
|
|
quantifiers (`(a+)+b`, the textbook ReDoS shape) no longer reaches this
|
|
engine at all: it has no backreference, lookaround, or atomic group, so
|
|
it is Pike VM eligible and runs there instead, in genuinely linear time
|
|
(see "The Pike VM" below).** The figures this bullet used to report for
|
|
the backtracking engine specifically (empirically quadratic: 3.2s at
|
|
*n* = 32,000, down from over a minute already at *n* = 40 before
|
|
memoization) remain accurate for what they actually measure, and still
|
|
apply to any pattern shaped like this one that *also* contains a
|
|
backreference or another construct that keeps it on this engine (the
|
|
inner `a+`'s `OP_REPEAT1` is followed by the group's closing save, not a
|
|
simple atom, so the skip-ahead technique above does not apply to it,
|
|
only the memoization does, which is why quadratic, not linear, was and
|
|
remains this engine's own ceiling for a compound repeat). A pattern
|
|
actually meant to run unattended against adversarial input on this
|
|
engine specifically should still avoid this shape, or wrap the inner
|
|
repetition in an atomic group, exactly as before.
|
|
|
|
Further reading on the techniques above: Russ Cox, ["Regular Expression
|
|
Matching: the Virtual Machine
|
|
Approach"](https://swtch.com/~rsc/regexp/regexp2.html) (why prepending
|
|
`.*?` gives linear-time unanchored search only in a Thompson/Pike VM, not in
|
|
a backtracking engine, which is why this build needed a different fix);
|
|
["Selective Memoization for Efficient Backtracking Regular Expression
|
|
Matching"](https://arxiv.org/html/2606.26678) and ["Efficient Matching with
|
|
Memoization for Regexes with Look-around and Atomic
|
|
Grouping"](https://arxiv.org/pdf/2401.12639) (the failure-only memoization
|
|
scheme this build's `run_memo` implements, and a more memory-efficient
|
|
selective variant, memoizing only at loop "feedback nodes" rather than every
|
|
instruction, that this build does not implement but could); the [Snyk
|
|
writeup on ReDoS and catastrophic
|
|
backtracking](https://snyk.io/blog/redos-and-catastrophic-backtracking/) for
|
|
the general phenomenon and mitigation guidance.
|
|
|
|
### The Pike VM
|
|
|
|
`concept.md` 7.7 is the implementation plan, written and researched before
|
|
any code, for the engine described above; 7.8 records what building it
|
|
against the existing test suite actually found, in the same "plan, then
|
|
measured finding" structure this document already uses for the backtracking
|
|
engine's own two sections (7.5, 7.6). In short: two real bugs, both found
|
|
and precisely localized by the existing 3,252-case suite without writing a
|
|
single test specifically for this engine (a missing write of group 0's end
|
|
position, and an unanchored-search early exit that was correct for "nothing
|
|
left to run" but wrong for "nothing left to run *yet*", both `concept.md`
|
|
7.8), a clean AddressSanitizer/UndefinedBehaviorSanitizer pass over the full
|
|
suite plus a further set of hand-written whitebox checks (a 20,000-branch
|
|
alternation, `UTF8`-mode named groups, `BINARY`-mode matching across an
|
|
embedded `NUL`, greedy/lazy and alternation priority), and one measured
|
|
result worth restating plainly here: `(a+)+b`, this document's own running
|
|
example of the backtracking engine's remaining weak spot, is Pike VM
|
|
eligible and now measures as linear, not quadratic, from *n* = 10,000 to
|
|
*n* = 160,000 (0.0018s to 0.0308s, roughly 2x per doubling of *n*
|
|
throughout).
|
|
|
|
`concept.md` 7.9 records three further findings, closing gaps 7.8 had left
|
|
open rather than repeating its numbers unchanged. First, the dual-engine
|
|
debug mode 7.8 said had not been built was built (an ad hoc whitebox
|
|
harness, not a permanent one, 7.9 explains why): 24,000 `match`/
|
|
`fullmatch`/`search` comparisons and roughly 2,700 `finditer` comparisons
|
|
between the two engines on the same compiled patterns, zero mismatches.
|
|
Second, a literal prefilter was added to the Pike VM's unanchored search
|
|
(when a pattern must begin with a specific literal or class, a fresh start
|
|
thread that position cannot satisfy is now skipped before being created,
|
|
not after), fixing the sharpest instance of the "ordinary patterns got
|
|
slower" regression "Benchmarks" below used to report: the literal-search
|
|
scenario there went from roughly 70x slower than POSIX `<regex.h>` to
|
|
roughly 11x-13x, better than this build's own numbers from before the Pike
|
|
VM existed at all (roughly 22x-35x). Third, the Pike VM's own memory cost,
|
|
which the design only asserted a bound for, was profiled directly with
|
|
Valgrind/Massif on the same adversarial pattern shape used for the
|
|
backtracking engine's own worst case (`a*b`, nullable leading loop, no
|
|
benefit from the new prefilter): the engine's own contribution beyond
|
|
holding the 10MB input itself was roughly 12KB, confirming the design's
|
|
`O(instruction count x group count)`, input-length-independent memory bound
|
|
actually holds for the v1 implementation, not only on paper.
|
|
|
|
Two things remain genuinely open, not resolved by the above: no lazy DFA
|
|
state caching or allocation pooling exists yet (`concept.md` 7.7's
|
|
remaining deferred items), and whether the Pike VM is faster or slower
|
|
than the backtracking engine specifically, as opposed to against POSIX, on
|
|
ordinary patterns is still unmeasured.
|
|
|
|
### Memory footprint
|
|
|
|
Measured directly with Valgrind/Massif (a real 10MB search) and by watching
|
|
`VmRSS`/`VmHWM` on real 200MB-1GB files, not estimated from reading the
|
|
code, for the **backtracking engine**; the Pike VM's own memory cost has a
|
|
different shape (a bounded number of threads, each carrying its own small
|
|
capture array, `concept.md` 7.7) and was profiled the same rigorous way
|
|
separately (`concept.md` 7.9, "The Pike VM" above): on the same 10MB
|
|
adversarial pattern shape used below, its own contribution beyond the input
|
|
itself was roughly 12KB, not a multiplier of the input length at all, so
|
|
the multiplier below is specific to patterns still running on the
|
|
backtracking engine (a backreference, lookaround, or atomic group present),
|
|
not a claim about every pattern. `Pattern_search`/`finditer`/
|
|
`split`/`sub` (the operations that try more than one start position) on
|
|
that engine currently use, at peak, about **8.4x** the input length in
|
|
memory for a pattern using `OP_REPEAT1`
|
|
(`concept.md` 7.5's `compute_maxrun` and `compute_next_prevmatch` tables,
|
|
`int32_t`-per-input-position each, are the entire remaining cost: 47.75%
|
|
each in the Massif profile, `alloc_memo`'s bitset a further 4.5%).
|
|
`Pattern_match`/`fullmatch` (a single attempt, no search tables) use
|
|
proportionally less.
|
|
|
|
Two real issues were found and fixed getting to that number, in order:
|
|
|
|
1. **`build_matbuf` widened every byte to a 4-byte `uint32_t`, even in
|
|
`BINARY`/`ASCII` mode, where a byte never exceeds 255 and the widening
|
|
bought nothing.** This cost as much extra memory as the input itself,
|
|
four times over, unconditionally, on top of the search tables above.
|
|
Fixed: `BINARY`/`ASCII` mode now reads the input's own bytes directly
|
|
(`MatBuf`/`MCtx`'s `text8` field, `text_at()`/`buf_at()` in `regexx.c`);
|
|
only `UTF8` mode still widens, because it actually needs code points up
|
|
to `0x10FFFF`, which do not fit in a byte. This dropped the measured
|
|
10MB-search peak from 140.3MB (13.4x) to 87.8MB (8.4x).
|
|
2. **`Input_from_file` read every file into a fresh, private, `malloc`'d
|
|
copy, even though the OS's page cache already holds the file's bytes.**
|
|
For a regular, seekable, non-empty file this now uses `mmap()`
|
|
(`PROT_READ`, `MAP_PRIVATE`) instead: the mapped pages are backed
|
|
directly by the file and stay clean (never written), so the kernel can
|
|
reclaim them under memory pressure and re-fault them in from disk later,
|
|
rather than them being pinned for the whole match attempt the way a
|
|
`malloc`'d copy is; it also removes one whole redundant copy of the
|
|
file's bytes. Falls back to the previous `read()`-based incremental
|
|
copy for anything `mmap` does not apply to (a pipe, a FIFO, process
|
|
substitution, stdin, an empty file, or an `mmap()` call that itself
|
|
fails).
|
|
|
|
A **third fix attempt was tried, measured, and reverted** specifically
|
|
because "prevent OOM, keep the footprint small" turned out to have a
|
|
sharp edge worth recording: capping `compute_maxrun`/`compute_next_prevmatch`
|
|
above a size budget and falling back to the plain scan already used when
|
|
either table is `NULL` seemed like an obvious bounded-memory safety valve.
|
|
Measured directly against a real 200MB non-matching search, it was worse
|
|
than doing nothing: both tables are needed together to keep this pattern
|
|
shape (`x*y`-style, unbounded quantifier followed by a required literal
|
|
that never occurs) at linear time; disabling either one alone reintroduces
|
|
the `O(n^2)` behavior they exist to fix, and `O(n^2)` at `n` in the hundreds
|
|
of millions does not finish in any practical amount of time. A fast,
|
|
diagnosable allocation failure (see below) is a better failure mode than a
|
|
silent, effectively-unbounded hang, so the cap was removed; these two
|
|
tables are allocated unconditionally again. There is no way to get both
|
|
bounded memory and linear time out of this technique for this pattern
|
|
shape; only `concept.md` Section 7.2's actual streaming automaton (still
|
|
unimplemented) gets both at once, by construction, which is why it remains
|
|
the correct long-term fix for this axis specifically.
|
|
|
|
**Failing safely.** Every allocation on the input-proportional paths above
|
|
(`build_matbuf`, `compute_maxrun`, `compute_next_prevmatch`,
|
|
`Input_from_file`, the UTF-8 decode arrays) is now checked; a failure
|
|
returns a `PatternError`/`-1` through the ordinary error path instead of
|
|
crashing on a `NULL` dereference, which several of them did before this
|
|
was audited (found by deliberately reasoning through "what happens when
|
|
this specific `malloc` fails on a huge request", not by a tool). This does
|
|
not prevent an out-of-memory condition on a genuinely memory-constrained
|
|
machine; the operating system's OOM killer can still end the process for
|
|
an allocation this library made in good faith (`malloc`/`mmap` returning
|
|
`NULL`/`MAP_FAILED` is the case this library can catch; being killed by
|
|
the kernel before that happens is not something a userspace library can
|
|
intercept). Measured concretely on the machine this was developed on: a
|
|
200MB file search that previously crashed via the OOM killer now completes
|
|
successfully in about 6 seconds at roughly 1.7GB peak RSS; a 1GB file on
|
|
the same machine still exceeded what was available at the time. Both
|
|
numbers are specific to that machine's available memory at the time, not
|
|
a hard property of the library; the 8.4x multiplier above is what actually
|
|
determines the practical ceiling on a given machine (roughly
|
|
`available memory / 8.4` for search-family operations on a pattern using
|
|
`OP_REPEAT1`, more forgiving for `match`/`fullmatch` or for patterns
|
|
without a simple-atom quantifier at all).
|
|
|
|
### Pattern syntax supported
|
|
|
|
Literals; `.` (with `DOTALL`); character classes with ranges, negation, and
|
|
`\d \D \w \W \s \S`; `\b \B`; anchors `^ $ \A \Z` (with `MULTILINE`);
|
|
quantifiers `* + ? {m,n} {m,} {,n} {m}`, greedy and lazy; possessive
|
|
quantifiers `*+ ++ ?+ {m,n}+`; groups `(...) (?:...) (?P<name>...)`;
|
|
alternation `|`; backreferences `\1`-`\99`, `(?P=name)`, `\g<name>`,
|
|
`\g<N>`; lookahead `(?=...) (?!...)`; fixed-width lookbehind
|
|
`(?<=...) (?<!...)`; atomic groups `(?>...)`; comments `(?#...)`; global
|
|
inline flags `(?aiLmsux)` at the start of a pattern; escapes
|
|
`\n \r \t \f \v \a`, octal `\0`-prefixed escapes, `\xhh`, `\uxxxx`,
|
|
`\Uxxxxxxxx`; flags `IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE`, `ASCII`
|
|
(all with observable effect; see `docs/API.md` Section 2 for exactly what
|
|
each one does), plus `UNICODE`, `LOCALE`, and `DEBUG` (accepted for source
|
|
compatibility with Python, currently no-ops in this build: `LOCALE` because
|
|
this build commits to the "C" locale only, under which `\w`/`\b`/`\B`
|
|
classify only ASCII letters and digits regardless of the flag, verified
|
|
directly against both the C standard and a real CPython interpreter).
|
|
|
|
Rejected at compile time with a clear `PatternError`, rather than
|
|
mis-parsed: conditional groups `(?(id)yes|no)`, scoped inline flags
|
|
`(?flags:...)`, `\N{NAME}` named code points, and POSIX bracket classes
|
|
`[:alpha:]` (which are not part of Python `re` at all, `concept.md` 14.3).
|
|
Variable-width lookbehind is also rejected at compile time, matching
|
|
CPython.
|
|
|
|
### Operations supported
|
|
|
|
`Pattern_match/fullmatch/search/finditer/findall/split/sub/subn/free`,
|
|
`Pattern_groupindex_lookup`,
|
|
`Match_group/start/end/span/start_byte/end_byte/span_byte/free`,
|
|
`re_compile/match/fullmatch/search/finditer/findall/split/sub/subn/escape/purge`,
|
|
`PatternError_free`. See `regexx.h` for exact signatures, `docs/API.md` for
|
|
the full reference (return values, memory ownership, exact Python
|
|
correspondence for each one), and `concept.md` Section 9 for the naming
|
|
convention they follow.
|
|
|
|
### Known deviations from `concept.md` and from CPython, beyond the items above
|
|
|
|
- `\w`, `\s`, `IGNORECASE` case folding, and `\d` in `UTF8` mode are backed
|
|
by glibc's `wctype.h` functions under the `C.utf8` locale, not by a
|
|
hand-generated Unicode table (`concept.md` 13.3 anticipated a reduced
|
|
static table; using the C library's own tables turned out to be simpler
|
|
and more complete, at the cost of depending on the platform's Unicode
|
|
version rather than a pinned one). Concretely verified, not just
|
|
theoretical: U+00A0 (NO-BREAK SPACE) is in Unicode's `White_Space`
|
|
property, so CPython's `\s` matches it, but glibc's `iswspace()` under
|
|
`C.utf8` does not, so this build's `\s` does not either. Found by the
|
|
large combinatorial test expansion (`tests/cases.py` Category F,
|
|
`tests/TEST_PLAN.md`), not anticipated in advance; recorded here rather
|
|
than patched, since hand-patching individual code points would start
|
|
down the path of maintaining an ad hoc table this design deliberately
|
|
avoided by delegating to `wctype.h` in the first place. A second,
|
|
same-class instance was found by the later Category I expansion:
|
|
fullwidth digits (`U+FF10`-`U+FF19`, Unicode category `Nd`) match
|
|
CPython's `\d` but not glibc's `iswdigit()` under `C.utf8` either.
|
|
- `lastindex`/`lastgroup` report the highest-numbered capturing group that
|
|
participated in the match, which coincides with CPython's "most recently
|
|
closed group" rule for straightforward patterns but can differ from it in
|
|
pathological cases (nested alternation re-executing a lower-numbered group
|
|
after a higher one). Not exercised by the test suite; documented here
|
|
rather than silently accepted.
|
|
- `Match_free`, `PatternError_free`, `Input_from_buffer`, `Input_from_file`,
|
|
and `Input_free` have no Python counterpart and are not mentioned in
|
|
`concept.md`'s API surface; they exist because C has no garbage collector.
|
|
`Pattern_sub`/`Pattern_subn` take the replacement template and the
|
|
callback as two separate parameters rather than one polymorphic argument,
|
|
for the same reason (`concept.md` 9.4 already anticipates and justifies
|
|
this one).
|
|
- Python's `Match.start(group)`/`.end(group)` raise `IndexError` for an
|
|
invalid group number and return `-1` only for a valid group that did not
|
|
participate; `Match_start`/`Match_end` return `-1` for both cases, since C
|
|
has no exception to raise. `Match_group` does distinguish them (`-1` for
|
|
no such group, `0` for an unparticipated one), see `docs/API.md` Section
|
|
3.23.
|
|
## Building
|
|
|
|
Requires a C11 compiler and, for the test suite, Python 3 (used only to
|
|
generate ground truth from CPython's own `re` module, `concept.md` Section
|
|
11; the library itself has no runtime dependency beyond the C standard
|
|
library and `libc`'s `wctype.h`/`locale.h`).
|
|
|
|
```sh
|
|
make # builds libregexx.a and the rxgrep example
|
|
make test # regenerates tests/generated_tests.c from Python `re`
|
|
# ground truth and runs the full suite
|
|
make check # same, under AddressSanitizer + UndefinedBehaviorSanitizer
|
|
make clean
|
|
```
|
|
|
|
`make install` installs `libregexx.a` and `regexx.h` under `PREFIX`
|
|
(default `/usr/local`).
|
|
|
|
## Using the library
|
|
|
|
```c
|
|
#include "regexx.h"
|
|
#include <string.h>
|
|
|
|
const char *pattern = "(\\w+)@(\\w+)";
|
|
Pattern *pat = re_compile(pattern, strlen(pattern), UTF8, NULL);
|
|
Input *in = Input_from_buffer((const uint8_t *)"user@host", strlen("user@host"));
|
|
|
|
Match m;
|
|
if (Pattern_search(pat, in, 0, -1, &m) == 1) {
|
|
const char *g; size_t glen;
|
|
Match_group(&m, NULL, 1, &g, &glen); /* g/glen -> "user" */
|
|
Match_free(&m);
|
|
}
|
|
|
|
char *out; size_t outlen;
|
|
Pattern_sub(pat, in, "\\2@\\1", NULL, NULL, 0, &out, &outlen); /* "host@user" */
|
|
free(out);
|
|
|
|
Pattern_free(pat);
|
|
Input_free(in);
|
|
```
|
|
|
|
`flags` to `re_compile` combine a data mode, exactly one of `BINARY`,
|
|
`ASCII`, or `UTF8` (`concept.md` 9.3), with any of the Python-named flags
|
|
(`IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE`, `ASCII` as a flag also
|
|
forces ASCII-only `\w`/`\s`/`\d` inside `UTF8` mode, `LOCALE`, `DEBUG`).
|
|
|
|
## Example: rxgrep
|
|
|
|
`examples/rxgrep.c` is a small grep-like program built on the library,
|
|
demonstrating all three data modes and both the matching and substitution
|
|
API:
|
|
|
|
```sh
|
|
./rxgrep -in 'hello' file.txt # case-insensitive, line numbers
|
|
./rxgrep -m utf8 -o '\w+' file.txt # print every UTF-8 word, one per line
|
|
./rxgrep -c 'error' log.txt # count matching lines
|
|
./rxgrep -m binary 'a.c' data.bin # match raw bytes, embedded NUL included
|
|
./rxgrep -m utf8 --sub 'REDACTED' '\d{3}-\d{4}' file.txt
|
|
```
|
|
|
|
Run `./rxgrep --help` for the full option list.
|
|
|
|
`examples/` has six further programs, each isolating one distinct feature
|
|
(a data mode, the undocumented CPython empty-match rule, atomic-group ReDoS
|
|
mitigation, `mmap`-backed large file input) rather than being a general
|
|
purpose tool; `make examples` builds all of them, and
|
|
[`examples/README.md`](examples/README.md) lists what each one demonstrates.
|
|
|
|
## Benchmarks
|
|
|
|
`examples/bench_vs_posix.c` measures this library directly against the C
|
|
standard library's own `<regex.h>` (POSIX `regcomp`/`regexec`, glibc's
|
|
DFA-backed implementation), on six scenarios at multi-megabyte or
|
|
multi-hundred-thousand-line scale, using only ERE pattern syntax that
|
|
regexx also accepts (no `\d`/`\w`/`\s`, no POSIX bracket classes), so both
|
|
engines run the identical pattern text against the identical subject.
|
|
Every scenario here uses a pattern with no backreference, lookaround, or
|
|
atomic group, so every one of them runs on the Pike VM ("The Pike VM"
|
|
above), not the backtracking engine.
|
|
|
|
Reported without adjustment in either direction, on this machine: glibc's
|
|
engine is roughly 7x to 27x faster on the three ordinary scenarios (a
|
|
literal search, extracting every number from a text, `a*b`). This narrowed
|
|
substantially after a literal prefilter was added to the Pike VM's
|
|
unanchored search (`concept.md` 7.9, "The Pike VM" above): the literal
|
|
search scenario, initially the worst hit, going from roughly 70x slower
|
|
than glibc (regressed from roughly 22x before the Pike VM existed at all)
|
|
down to roughly 11x-13x, better than that original pre-Pike-VM number; the
|
|
number-extraction scenario improved more modestly (roughly 8x to roughly
|
|
7x), since it begins with a class rather than a single literal; the `a*b`
|
|
scenario is unchanged, since its leading `a*` is a nullable loop the
|
|
prefilter structurally cannot help, an expected, not a missed, case. No
|
|
lazy DFA state caching or allocation pooling exists yet
|
|
(`concept.md` 7.7's remaining deferred performance items), so a real gap
|
|
to glibc's decades of exactly that kind of optimization remains on these
|
|
three scenarios, narrower than before but not closed. The fourth scenario
|
|
inverts entirely: the textbook ReDoS shape `(a+)+b` now runs *faster than
|
|
glibc*, with no atomic group needed at all, because a Thompson-NFA
|
|
simulation has no notion of "try one split, then backtrack and try
|
|
another" for the classic nested-quantifier ambiguity to exploit in the
|
|
first place; `examples/redos_atomic.c`'s Part 2 shows the one case this
|
|
specific fix does not reach (a pattern with a backreference forces the
|
|
older, backtracking engine regardless of shape), where an atomic group
|
|
remains the pattern author's own necessary tool, not automatic. Every
|
|
scenario's match count is cross-checked between the two engines and
|
|
reported as agreeing or differing, an independent correctness check beyond
|
|
the CPython-derived test suite below.
|
|
|
|
## Testing
|
|
|
|
`tests/cases.py` lists pattern/subject/operation triples, both hand-written
|
|
and, for most of the file, generated programmatically from combinations of
|
|
quantifiers, groups, backreferences, lookaround, flags, and encoding modes
|
|
(`tests/TEST_PLAN.md` records the exact category breakdown and why each one
|
|
exists). `tests/gen.py` computes every case's expected result with
|
|
CPython's own `re` module and writes `tests/generated_tests.c`, which is
|
|
then compiled against `regexx.c` and checked. This is a direct
|
|
implementation of the strategy `concept.md` Section 11 describes:
|
|
conformance is measured against what CPython actually does, not against a
|
|
re-derived reading of its documentation, at a scale (3,252 cases as of this
|
|
writing, `make test` reports the current exact count) large enough that it
|
|
has already found real defects this way, not only confirmed the absence of
|
|
ones anyone thought to write by hand (`tests/TEST_PLAN.md` "Result" names
|
|
all four: a C trigraph bug in the test generator itself, a real, previously
|
|
undocumented `Pattern_finditer`/`Pattern_split` empty-match defect, a wrong
|
|
ground truth for `ASCII` mode in the generator itself, and a real,
|
|
verified-wrong implementation of the `LOCALE` flag). `make check`
|
|
additionally runs the suite under AddressSanitizer and
|
|
UndefinedBehaviorSanitizer.
|
|
|
|
## License
|
|
|
|
MIT. See [`LICENSE`](LICENSE).
|