Profiled a real search with Massif and found the memory-per-input-byte multiplier at 13.4x, dominated by two avoidable costs: - build_matbuf widened every BINARY/ASCII byte into a uint32_t before matching, a 4x copy that mode never needed (a byte never exceeds 255). Removed it: MCtx/MatBuf now carry an optional text8 (borrowed, unwidened) alongside the existing UTF8 text (owned, decoded code points), reconciled per-read through text_at()/buf_at(). BINARY/ ASCII mode now points directly into the Input's own buffer. - Input_from_file always copied the whole file into a malloc'd buffer. It now mmaps regular files read-only (MAP_PRIVATE) instead, so pages stay clean and reclaimable under memory pressure and the file is never copied. Non-seekable sources (pipes, FIFOs, process substitution, stdin) and mmap failures fall back to the previous incremental-read behavior via a separate read_fd_incrementally. Together these bring the multiplier to 8.4x and let a 200MB file that previously OOM-crashed complete a full non-matching search in about 6 seconds at roughly 1.69GB peak RSS. Also tried, measured, and reverted: capping compute_maxrun/ compute_next_prevmatch's table size with a plain-scan fallback above the cap. A real 200MB non-matching search against this fallback hung for minutes instead of failing fast, because disabling either table reintroduces the O(n^2) behavior they exist to prevent, and O(n^2) at n in the hundreds of millions is not practically finite. A fast, diagnosable allocation failure is a better failure mode than a silent, unbounded hang, so the tables are allocated unconditionally again; the finding is recorded in code comments, concept.md 7.6, and README's "Memory footprint" section so it is not retried blindly later. Separately audited every allocation on an input-proportional path (da_push, build_matbuf's UTF-8 decode loop, all three Input_from_file sites) and made each fail cleanly through PatternError instead of crashing on an unchecked NULL dereference. A 1GB file still exceeds available memory in the current environment; this is a property of the machine it was measured on, not a defect, and is documented as such (practical ceiling: available memory / 8.4 for search-family operations, pending the streaming automaton design in concept.md 7.2). Verified with four clean `make test` passes (3252/3252) and a clean ASan/UBSan pass after the change; rxgrep's mmap-backed paths (--sub with a regular file, with a non-seekable process-substitution source, and BINARY-mode embedded NUL handling) re-checked directly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
regexx
A single-file C regular expression interpreter that reproduces the observable
behavior of Python's re module, including its exact identifier names
(Pattern, Match, re_compile, re_sub, IGNORECASE, and so on; see
Section 9 of concept.md), and that additionally supports binary data
(arbitrary byte streams, including embedded NUL) and UTF-8 text alongside
plain ASCII.
The design rationale, the algorithmic trade-offs, and a full accounting of
what is and is not carried over from Python's re and from POSIX's native
regex.h are recorded in concept.md. This file documents the
implementation that exists today at a glance; docs/API.md is
the exhaustive reference (every type, every flag, every function's exact
return-value and memory-ownership convention, checked against a real CPython
interpreter, not against memory).
Implementation status
This is a v1 implementation. It is a complete, tested engine for the pattern
syntax and operations listed below, executed by a single recursive
backtracking engine (concept.md Section 7.3) over a fully materialized copy
of the input.
It does not yet implement the streaming, bounded-memory regular engine of
concept.md Section 7.2. Input (the abstraction over "a source of
chunks", concept.md 9.2) is implemented, and Input_from_file maps or
reads a whole file into memory before matching (see "Memory footprint"
below for exactly how, and for the measured numbers this and other fixes
were checked against). Every public function signature is already exactly
what the streaming design in concept.md specifies, so the non-streaming
implementation underneath a given call can be replaced later without
changing any caller. Concretely, today:
Pattern_match/Pattern_fullmatch(a single anchored attempt at a fixed position) run in time proportional to the length of that attempt, and, for the common case of a single character, class, or.repeated by a quantifier, in O(1) recursion depth regardless of input size (theOP_REPEAT1fast path). A repeated compound sub-pattern (for example(ab)*) still recurses once per repetition, bounded by a configurable depth limit (MAX_DEPTHinregexx.c, currently 60000, not exposed through the public API, so changing it means editingregexx.cand rebuilding); past that limit, matching fails with a reported error rather than a stack overflow or a wrong answer.Pattern_search/Pattern_finditer/Pattern_split/Pattern_subrun in linear time for the common case: a pattern built only from simple, non-backreference constructs where every quantifier's body is a single character, class, or.(a*b,\s*\d+,.*", and the great majority of patterns actually written by hand). This was not always true of this build; measuring it directly is what caught that it previously was not (an earlier revision of this file claimed linear time without having re-verified it, then had to correct that claim, then fixed the underlying defect; the corrected numbers are below). Two independent techniques make it true now, neither of them the streaming engine ofconcept.mdSection 7.2, which remains unimplemented:-
Memoized backtracking. For any pattern with no
OP_BACKREFanywhere, whether execution starting at a given (instruction, text position) pair can ever reach a match is a fact that never changes once computed, sorun_memo(regexx.c) caches every failure (never a success, so it cannot change which match is found, only skip re-deriving failures already known) and consults the cache before redoing that work. This is a published technique, not a house invention: a memoization table that records only failures, because "a matching success ... immediately propagates to the success of the whole problem," is exactly the scheme described in recent work on backtracking regex matchers ("Selective Memoization for Efficient Backtracking Regular Expression Matching", and, for the lookaround/atomic case specifically, "Efficient Matching with Memoization for Regexes with Look-around and Atomic Grouping", both linked below). -
Precomputed run lengths and skip-ahead tables for
OP_REPEAT1. Memoizing individual(instruction, position)pairs does not help when each one is alreadyO(1)work, which is exactlya*b's situation: counting how manyas follow a position, and then trying every possible split point against the trailingbone at a time, are each individually cheap but happenO(remaining length)times per start position tried.compute_maxrunprecomputes, once perPattern_search/finditer/splitcall, how many characters each repeat can consume from every position in one backward pass, andcompute_next_prevmatchprecomputes, for a repeat immediately followed by a single literal/class/., the rightmost position at or before any given point where that next atom can match, so the backtrack loop jumps directly to candidates worth trying instead of visiting every position in between. This is the same idea production engines call a literal prefilter (RE2's and Rust'sregexcrate'smemchr/memmem/Teddy prefilters skip positions that provably cannot match before ever invoking the full engine); those use SIMD-accelerated library primitives operating directly on bytes, this build uses a precomputed array, which is slower per lookup but the same algorithmic idea, and appropriate forconcept.md's stated priority of convenience over performance.Measured directly:
a*bsearched over n bytes ofawith nobanywhere took 19.2s at n = 80,000 before these two techniques and 0.0016s after, with time now scaling linearly (2x per doubling of n) rather than quadratically (4x per doubling).Cost: these tables use
O(k * n)memory,kbeing the number ofOP_REPEAT1instructions actually present in the compiled pattern,nthe input length;Pattern_match/Pattern_fullmatchnever allocate them (a single anchored attempt gets no benefit from them).
-
- A pattern with a backreference can, like CPython's own
_sre, still take worst-case exponential time on an adversarial input (concept.mdSection 5, 13.2; memoization above is unsound and therefore disabled whenever a pattern containsOP_BACKREFanywhere, since a backreference's outcome depends on capture history, not on position alone). This is the same catastrophic backtracking (ReDoS) behavior CPython itself exhibits on such patterns, not a regression specific to this engine; a pattern author who needs to rule it out for a specific pattern can use an atomic group ((?>...)) or a possessive quantifier around the ambiguous repetition, exactly as they would for CPython'sre, and exactly as general ReDoS mitigation guidance recommends (linked below). - A backreference-free pattern shaped like nested, overlapping
quantifiers (
(a+)+b, the textbook ReDoS shape) is no longer exponential, because it has no backreference and so is memoized, but is not necessarily linear either. Measured directly:(a+)+bagainst n characters ofawith no trailingbtook over a minute already at n = 40 before memoization (exponential; this build's own stress test keeps that case at n = 24 for exactly this reason) and 3.2s at n = 32,000 after (empirically quadratic: roughly 4x per doubling), because the innera+'sOP_REPEAT1is followed by the group's closing save, not by a simple atom, so the skip-ahead technique above does not apply to it, only the memoization does. Exponential to quadratic is a qualitative change (the n = 40 case alone went from longer than this project could wait for, to instant), not a complete fix; a pattern actually meant to run unattended against adversarial input should still avoid this shape, or wrap the inner repetition in an atomic group, regardless of the improvement.
Further reading on the techniques above: Russ Cox, "Regular Expression
Matching: the Virtual Machine
Approach" (why prepending
.*? gives linear-time unanchored search only in a Thompson/Pike VM, not in
a backtracking engine, which is why this build needed a different fix);
"Selective Memoization for Efficient Backtracking Regular Expression
Matching" and "Efficient Matching with
Memoization for Regexes with Look-around and Atomic
Grouping" (the failure-only memoization
scheme this build's run_memo implements, and a more memory-efficient
selective variant, memoizing only at loop "feedback nodes" rather than every
instruction, that this build does not implement but could); the Snyk
writeup on ReDoS and catastrophic
backtracking for
the general phenomenon and mitigation guidance.
Memory footprint
Measured directly with Valgrind/Massif (a real 10MB search) and by watching
VmRSS/VmHWM on real 200MB-1GB files, not estimated from reading the
code. Pattern_search/finditer/split/sub (the operations that try
more than one start position) currently use, at peak, about 8.4x the
input length in memory for a pattern using OP_REPEAT1
(concept.md 7.5's compute_maxrun and compute_next_prevmatch tables,
int32_t-per-input-position each, are the entire remaining cost: 47.75%
each in the Massif profile, alloc_memo's bitset a further 4.5%).
Pattern_match/fullmatch (a single attempt, no search tables) use
proportionally less.
Two real issues were found and fixed getting to that number, in order:
build_matbufwidened every byte to a 4-byteuint32_t, even inBINARY/ASCIImode, where a byte never exceeds 255 and the widening bought nothing. This cost as much extra memory as the input itself, four times over, unconditionally, on top of the search tables above. Fixed:BINARY/ASCIImode now reads the input's own bytes directly (MatBuf/MCtx'stext8field,text_at()/buf_at()inregexx.c); onlyUTF8mode still widens, because it actually needs code points up to0x10FFFF, which do not fit in a byte. This dropped the measured 10MB-search peak from 140.3MB (13.4x) to 87.8MB (8.4x).Input_from_fileread every file into a fresh, private,malloc'd copy, even though the OS's page cache already holds the file's bytes. For a regular, seekable, non-empty file this now usesmmap()(PROT_READ,MAP_PRIVATE) instead: the mapped pages are backed directly by the file and stay clean (never written), so the kernel can reclaim them under memory pressure and re-fault them in from disk later, rather than them being pinned for the whole match attempt the way amalloc'd copy is; it also removes one whole redundant copy of the file's bytes. Falls back to the previousread()-based incremental copy for anythingmmapdoes not apply to (a pipe, a FIFO, process substitution, stdin, an empty file, or anmmap()call that itself fails).
A third fix attempt was tried, measured, and reverted specifically
because "prevent OOM, keep the footprint small" turned out to have a
sharp edge worth recording: capping compute_maxrun/compute_next_prevmatch
above a size budget and falling back to the plain scan already used when
either table is NULL seemed like an obvious bounded-memory safety valve.
Measured directly against a real 200MB non-matching search, it was worse
than doing nothing: both tables are needed together to keep this pattern
shape (x*y-style, unbounded quantifier followed by a required literal
that never occurs) at linear time; disabling either one alone reintroduces
the O(n^2) behavior they exist to fix, and O(n^2) at n in the hundreds
of millions does not finish in any practical amount of time. A fast,
diagnosable allocation failure (see below) is a better failure mode than a
silent, effectively-unbounded hang, so the cap was removed; these two
tables are allocated unconditionally again. There is no way to get both
bounded memory and linear time out of this technique for this pattern
shape; only concept.md Section 7.2's actual streaming automaton (still
unimplemented) gets both at once, by construction, which is why it remains
the correct long-term fix for this axis specifically.
Failing safely. Every allocation on the input-proportional paths above
(build_matbuf, compute_maxrun, compute_next_prevmatch,
Input_from_file, the UTF-8 decode arrays) is now checked; a failure
returns a PatternError/-1 through the ordinary error path instead of
crashing on a NULL dereference, which several of them did before this
was audited (found by deliberately reasoning through "what happens when
this specific malloc fails on a huge request", not by a tool). This does
not prevent an out-of-memory condition on a genuinely memory-constrained
machine; the operating system's OOM killer can still end the process for
an allocation this library made in good faith (malloc/mmap returning
NULL/MAP_FAILED is the case this library can catch; being killed by
the kernel before that happens is not something a userspace library can
intercept). Measured concretely on the machine this was developed on: a
200MB file search that previously crashed via the OOM killer now completes
successfully in about 6 seconds at roughly 1.7GB peak RSS; a 1GB file on
the same machine still exceeded what was available at the time. Both
numbers are specific to that machine's available memory at the time, not
a hard property of the library; the 8.4x multiplier above is what actually
determines the practical ceiling on a given machine (roughly
available memory / 8.4 for search-family operations on a pattern using
OP_REPEAT1, more forgiving for match/fullmatch or for patterns
without a simple-atom quantifier at all).
Pattern syntax supported
Literals; . (with DOTALL); character classes with ranges, negation, and
\d \D \w \W \s \S; \b \B; anchors ^ $ \A \Z (with MULTILINE);
quantifiers * + ? {m,n} {m,} {,n} {m}, greedy and lazy; possessive
quantifiers *+ ++ ?+ {m,n}+; groups (...) (?:...) (?P<name>...);
alternation |; backreferences \1-\99, (?P=name), \g<name>,
\g<N>; lookahead (?=...) (?!...); fixed-width lookbehind
(?<=...) (?<!...); atomic groups (?>...); comments (?#...); global
inline flags (?aiLmsux) at the start of a pattern; escapes
\n \r \t \f \v \a, octal \0-prefixed escapes, \xhh, \uxxxx,
\Uxxxxxxxx; flags IGNORECASE, MULTILINE, DOTALL, VERBOSE, ASCII
(all with observable effect; see docs/API.md Section 2 for exactly what
each one does), plus UNICODE, LOCALE, and DEBUG (accepted for source
compatibility with Python, currently no-ops in this build: LOCALE because
this build commits to the "C" locale only, under which \w/\b/\B
classify only ASCII letters and digits regardless of the flag, verified
directly against both the C standard and a real CPython interpreter).
Rejected at compile time with a clear PatternError, rather than
mis-parsed: conditional groups (?(id)yes|no), scoped inline flags
(?flags:...), \N{NAME} named code points, and POSIX bracket classes
[:alpha:] (which are not part of Python re at all, concept.md 14.3).
Variable-width lookbehind is also rejected at compile time, matching
CPython.
Operations supported
Pattern_match/fullmatch/search/finditer/findall/split/sub/subn/free,
Pattern_groupindex_lookup,
Match_group/start/end/span/start_byte/end_byte/span_byte/free,
re_compile/match/fullmatch/search/finditer/findall/split/sub/subn/escape/purge,
PatternError_free. See regexx.h for exact signatures, docs/API.md for
the full reference (return values, memory ownership, exact Python
correspondence for each one), and concept.md Section 9 for the naming
convention they follow.
Known deviations from concept.md and from CPython, beyond the items above
\w,\s,IGNORECASEcase folding, and\dinUTF8mode are backed by glibc'swctype.hfunctions under theC.utf8locale, not by a hand-generated Unicode table (concept.md13.3 anticipated a reduced static table; using the C library's own tables turned out to be simpler and more complete, at the cost of depending on the platform's Unicode version rather than a pinned one). Concretely verified, not just theoretical: U+00A0 (NO-BREAK SPACE) is in Unicode'sWhite_Spaceproperty, so CPython's\smatches it, but glibc'siswspace()underC.utf8does not, so this build's\sdoes not either. Found by the large combinatorial test expansion (tests/cases.pyCategory F,tests/TEST_PLAN.md), not anticipated in advance; recorded here rather than patched, since hand-patching individual code points would start down the path of maintaining an ad hoc table this design deliberately avoided by delegating towctype.hin the first place. A second, same-class instance was found by the later Category I expansion: fullwidth digits (U+FF10-U+FF19, Unicode categoryNd) match CPython's\dbut not glibc'siswdigit()underC.utf8either.lastindex/lastgroupreport the highest-numbered capturing group that participated in the match, which coincides with CPython's "most recently closed group" rule for straightforward patterns but can differ from it in pathological cases (nested alternation re-executing a lower-numbered group after a higher one). Not exercised by the test suite; documented here rather than silently accepted.Match_free,PatternError_free,Input_from_buffer,Input_from_file, andInput_freehave no Python counterpart and are not mentioned inconcept.md's API surface; they exist because C has no garbage collector.Pattern_sub/Pattern_subntake the replacement template and the callback as two separate parameters rather than one polymorphic argument, for the same reason (concept.md9.4 already anticipates and justifies this one).- Python's
Match.start(group)/.end(group)raiseIndexErrorfor an invalid group number and return-1only for a valid group that did not participate;Match_start/Match_endreturn-1for both cases, since C has no exception to raise.Match_groupdoes distinguish them (-1for no such group,0for an unparticipated one), seedocs/API.mdSection 3.23. Pattern.groupindexhas no enumeration function in this build, onlyPattern_groupindex_lookup(pattern, name); there is no way to list every name a compiled pattern defines without already knowing what to look for.
Building
Requires a C11 compiler and, for the test suite, Python 3 (used only to
generate ground truth from CPython's own re module, concept.md Section
11; the library itself has no runtime dependency beyond the C standard
library and libc's wctype.h/locale.h).
make # builds libregexx.a and the rxgrep example
make test # regenerates tests/generated_tests.c from Python `re`
# ground truth and runs the full suite
make check # same, under AddressSanitizer + UndefinedBehaviorSanitizer
make clean
make install installs libregexx.a and regexx.h under PREFIX
(default /usr/local).
Using the library
#include "regexx.h"
#include <string.h>
const char *pattern = "(\\w+)@(\\w+)";
Pattern *pat = re_compile(pattern, strlen(pattern), UTF8, NULL);
Input *in = Input_from_buffer((const uint8_t *)"user@host", strlen("user@host"));
Match m;
if (Pattern_search(pat, in, 0, -1, &m) == 1) {
const char *g; size_t glen;
Match_group(&m, NULL, 1, &g, &glen); /* g/glen -> "user" */
Match_free(&m);
}
char *out; size_t outlen;
Pattern_sub(pat, in, "\\2@\\1", NULL, NULL, 0, &out, &outlen); /* "host@user" */
free(out);
Pattern_free(pat);
Input_free(in);
flags to re_compile combine a data mode, exactly one of BINARY,
ASCII, or UTF8 (concept.md 9.3), with any of the Python-named flags
(IGNORECASE, MULTILINE, DOTALL, VERBOSE, ASCII as a flag also
forces ASCII-only \w/\s/\d inside UTF8 mode, LOCALE, DEBUG).
Example: rxgrep
examples/rxgrep.c is a small grep-like program built on the library,
demonstrating all three data modes and both the matching and substitution
API:
./rxgrep -in 'hello' file.txt # case-insensitive, line numbers
./rxgrep -m utf8 -o '\w+' file.txt # print every UTF-8 word, one per line
./rxgrep -c 'error' log.txt # count matching lines
./rxgrep -m binary 'a.c' data.bin # match raw bytes, embedded NUL included
./rxgrep -m utf8 --sub 'REDACTED' '\d{3}-\d{4}' file.txt
Run ./rxgrep --help for the full option list.
Testing
tests/cases.py lists pattern/subject/operation triples, both hand-written
and, for most of the file, generated programmatically from combinations of
quantifiers, groups, backreferences, lookaround, flags, and encoding modes
(tests/TEST_PLAN.md records the exact category breakdown and why each one
exists). tests/gen.py computes every case's expected result with
CPython's own re module and writes tests/generated_tests.c, which is
then compiled against regexx.c and checked. This is a direct
implementation of the strategy concept.md Section 11 describes:
conformance is measured against what CPython actually does, not against a
re-derived reading of its documentation, at a scale (3,252 cases as of this
writing, make test reports the current exact count) large enough that it
has already found real defects this way, not only confirmed the absence of
ones anyone thought to write by hand (tests/TEST_PLAN.md "Result" names
all four: a C trigraph bug in the test generator itself, a real, previously
undocumented Pattern_finditer/Pattern_split empty-match defect, a wrong
ground truth for ASCII mode in the generator itself, and a real,
verified-wrong implementation of the LOCALE flag). make check
additionally runs the suite under AddressSanitizer and
UndefinedBehaviorSanitizer.
License
MIT. See LICENSE.