retoorandClaude Sonnet 5 a9ca517981 Bring concept.md and file-header comments in line with the memoization fix
concept.md described the bounded backtracking engine (7.3) as plain
recursive backtracking, with the Regular engine's linear-time guarantee
reserved for patterns the streaming Pike VM (7.2) handles. v1 has no
Pike VM yet, so every pattern actually runs on 7.3's engine; the
previous commit added memoization and precomputed skip-ahead tables to
that engine specifically because plain backtracking search was
quadratic even for ordinary patterns, not just exponential for
pathological ones.

Added Section 7.5 recording that finding precisely: what the two
techniques buy (linear time for the backreference-free majority,
quadratic instead of exponential for the textbook (a+)+b shape), what
they do not buy (7.2's bounded, input-length-independent memory
guarantee, still the correct target for that axis), and citations to
the published techniques they match. Added a short note to the Section
10 complexity table pointing at 7.5 so the table is read as the two
engines' designed guarantees, not conflated with what a specific
implementation measures. Updated regexx.c/regexx.h's top-of-file
comments, which still described a plain backtracking engine.

No code changes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
2026-09-14 08:53:19 +00:00

regexx

A single-file C regular expression interpreter that reproduces the observable behavior of Python's re module, including its exact identifier names (Pattern, Match, re_compile, re_sub, IGNORECASE, and so on; see Section 9 of concept.md), and that additionally supports binary data (arbitrary byte streams, including embedded NUL) and UTF-8 text alongside plain ASCII.

The design rationale, the algorithmic trade-offs, and a full accounting of what is and is not carried over from Python's re and from POSIX's native regex.h are recorded in concept.md. This file documents the implementation that exists today at a glance; docs/API.md is the exhaustive reference (every type, every flag, every function's exact return-value and memory-ownership convention, checked against a real CPython interpreter, not against memory).

Implementation status

This is a v1 implementation. It is a complete, tested engine for the pattern syntax and operations listed below, executed by a single recursive backtracking engine (concept.md Section 7.3) over a fully materialized copy of the input.

It does not yet implement the streaming, bounded-memory regular engine of concept.md Section 7.2. Input (the abstraction over "a source of chunks", concept.md 9.2) is implemented, and Input_from_file reads a whole file into memory before matching. Every public function signature is already exactly what the streaming design in concept.md specifies, so the non-streaming implementation underneath a given call can be replaced later without changing any caller. Concretely, today:

  • Pattern_match/Pattern_fullmatch (a single anchored attempt at a fixed position) run in time proportional to the length of that attempt, and, for the common case of a single character, class, or . repeated by a quantifier, in O(1) recursion depth regardless of input size (the OP_REPEAT1 fast path). A repeated compound sub-pattern (for example (ab)*) still recurses once per repetition, bounded by a configurable depth limit (MAX_DEPTH in regexx.c, currently 60000, not exposed through the public API, so changing it means editing regexx.c and rebuilding); past that limit, matching fails with a reported error rather than a stack overflow or a wrong answer.
  • Pattern_search/Pattern_finditer/Pattern_split/Pattern_sub run in linear time for the common case: a pattern built only from simple, non-backreference constructs where every quantifier's body is a single character, class, or . (a*b, \s*\d+, .*", and the great majority of patterns actually written by hand). This was not always true of this build; measuring it directly is what caught that it previously was not (an earlier revision of this file claimed linear time without having re-verified it, then had to correct that claim, then fixed the underlying defect; the corrected numbers are below). Two independent techniques make it true now, neither of them the streaming engine of concept.md Section 7.2, which remains unimplemented:
    1. Memoized backtracking. For any pattern with no OP_BACKREF anywhere, whether execution starting at a given (instruction, text position) pair can ever reach a match is a fact that never changes once computed, so run_memo (regexx.c) caches every failure (never a success, so it cannot change which match is found, only skip re-deriving failures already known) and consults the cache before redoing that work. This is a published technique, not a house invention: a memoization table that records only failures, because "a matching success ... immediately propagates to the success of the whole problem," is exactly the scheme described in recent work on backtracking regex matchers ("Selective Memoization for Efficient Backtracking Regular Expression Matching", and, for the lookaround/atomic case specifically, "Efficient Matching with Memoization for Regexes with Look-around and Atomic Grouping", both linked below).

    2. Precomputed run lengths and skip-ahead tables for OP_REPEAT1. Memoizing individual (instruction, position) pairs does not help when each one is already O(1) work, which is exactly a*b's situation: counting how many as follow a position, and then trying every possible split point against the trailing b one at a time, are each individually cheap but happen O(remaining length) times per start position tried. compute_maxrun precomputes, once per Pattern_search/finditer/split call, how many characters each repeat can consume from every position in one backward pass, and compute_next_prevmatch precomputes, for a repeat immediately followed by a single literal/class/., the rightmost position at or before any given point where that next atom can match, so the backtrack loop jumps directly to candidates worth trying instead of visiting every position in between. This is the same idea production engines call a literal prefilter (RE2's and Rust's regex crate's memchr/memmem/Teddy prefilters skip positions that provably cannot match before ever invoking the full engine); those use SIMD-accelerated library primitives operating directly on bytes, this build uses a precomputed array, which is slower per lookup but the same algorithmic idea, and appropriate for concept.md's stated priority of convenience over performance.

      Measured directly: a*b searched over n bytes of a with no b anywhere took 19.2s at n = 80,000 before these two techniques and 0.0016s after, with time now scaling linearly (2x per doubling of n) rather than quadratically (4x per doubling).

      Cost: these tables use O(k * n) memory, k being the number of OP_REPEAT1 instructions actually present in the compiled pattern, n the input length; Pattern_match/Pattern_fullmatch never allocate them (a single anchored attempt gets no benefit from them).

  • A pattern with a backreference can, like CPython's own _sre, still take worst-case exponential time on an adversarial input (concept.md Section 5, 13.2; memoization above is unsound and therefore disabled whenever a pattern contains OP_BACKREF anywhere, since a backreference's outcome depends on capture history, not on position alone). This is the same catastrophic backtracking (ReDoS) behavior CPython itself exhibits on such patterns, not a regression specific to this engine; a pattern author who needs to rule it out for a specific pattern can use an atomic group ((?>...)) or a possessive quantifier around the ambiguous repetition, exactly as they would for CPython's re, and exactly as general ReDoS mitigation guidance recommends (linked below).
  • A backreference-free pattern shaped like nested, overlapping quantifiers ((a+)+b, the textbook ReDoS shape) is no longer exponential, because it has no backreference and so is memoized, but is not necessarily linear either. Measured directly: (a+)+b against n characters of a with no trailing b took over a minute already at n = 40 before memoization (exponential; this build's own stress test keeps that case at n = 24 for exactly this reason) and 3.2s at n = 32,000 after (empirically quadratic: roughly 4x per doubling), because the inner a+'s OP_REPEAT1 is followed by the group's closing save, not by a simple atom, so the skip-ahead technique above does not apply to it, only the memoization does. Exponential to quadratic is a qualitative change (the n = 40 case alone went from longer than this project could wait for, to instant), not a complete fix; a pattern actually meant to run unattended against adversarial input should still avoid this shape, or wrap the inner repetition in an atomic group, regardless of the improvement.

Further reading on the techniques above: Russ Cox, "Regular Expression Matching: the Virtual Machine Approach" (why prepending .*? gives linear-time unanchored search only in a Thompson/Pike VM, not in a backtracking engine, which is why this build needed a different fix); "Selective Memoization for Efficient Backtracking Regular Expression Matching" and "Efficient Matching with Memoization for Regexes with Look-around and Atomic Grouping" (the failure-only memoization scheme this build's run_memo implements, and a more memory-efficient selective variant, memoizing only at loop "feedback nodes" rather than every instruction, that this build does not implement but could); the Snyk writeup on ReDoS and catastrophic backtracking for the general phenomenon and mitigation guidance.

Pattern syntax supported

Literals; . (with DOTALL); character classes with ranges, negation, and \d \D \w \W \s \S; \b \B; anchors ^ $ \A \Z (with MULTILINE); quantifiers * + ? {m,n} {m,} {,n} {m}, greedy and lazy; possessive quantifiers *+ ++ ?+ {m,n}+; groups (...) (?:...) (?P<name>...); alternation |; backreferences \1-\99, (?P=name), \g<name>, \g<N>; lookahead (?=...) (?!...); fixed-width lookbehind (?<=...) (?<!...); atomic groups (?>...); comments (?#...); global inline flags (?aiLmsux) at the start of a pattern; escapes \n \r \t \f \v \a, octal \0-prefixed escapes, \xhh, \uxxxx, \Uxxxxxxxx; flags IGNORECASE, MULTILINE, DOTALL, VERBOSE, ASCII, LOCALE (all with observable effect; see docs/API.md Section 2 for exactly what each one does), plus UNICODE and DEBUG (accepted for source compatibility with Python, currently no-ops in this build).

Rejected at compile time with a clear PatternError, rather than mis-parsed: conditional groups (?(id)yes|no), scoped inline flags (?flags:...), \N{NAME} named code points, and POSIX bracket classes [:alpha:] (which are not part of Python re at all, concept.md 14.3). Variable-width lookbehind is also rejected at compile time, matching CPython.

Operations supported

Pattern_match/fullmatch/search/finditer/findall/split/sub/subn/free, Pattern_groupindex_lookup, Match_group/start/end/span/start_byte/end_byte/span_byte/free, re_compile/match/fullmatch/search/finditer/findall/split/sub/subn/escape/purge, PatternError_free. See regexx.h for exact signatures, docs/API.md for the full reference (return values, memory ownership, exact Python correspondence for each one), and concept.md Section 9 for the naming convention they follow.

Known deviations from concept.md and from CPython, beyond the items above

  • \w, \s, IGNORECASE case folding, and \d in UTF8 mode are backed by glibc's wctype.h functions under the C.utf8 locale, not by a hand-generated Unicode table (concept.md 13.3 anticipated a reduced static table; using the C library's own tables turned out to be simpler and more complete, at the cost of depending on the platform's Unicode version rather than a pinned one).
  • lastindex/lastgroup report the highest-numbered capturing group that participated in the match, which coincides with CPython's "most recently closed group" rule for straightforward patterns but can differ from it in pathological cases (nested alternation re-executing a lower-numbered group after a higher one). Not exercised by the test suite; documented here rather than silently accepted.
  • Match_free, PatternError_free, Input_from_buffer, Input_from_file, and Input_free have no Python counterpart and are not mentioned in concept.md's API surface; they exist because C has no garbage collector. Pattern_sub/Pattern_subn take the replacement template and the callback as two separate parameters rather than one polymorphic argument, for the same reason (concept.md 9.4 already anticipates and justifies this one).
  • Python's Match.start(group)/.end(group) raise IndexError for an invalid group number and return -1 only for a valid group that did not participate; Match_start/Match_end return -1 for both cases, since C has no exception to raise. Match_group does distinguish them (-1 for no such group, 0 for an unparticipated one), see docs/API.md Section 3.23.
  • Pattern.groupindex has no enumeration function in this build, only Pattern_groupindex_lookup(pattern, name); there is no way to list every name a compiled pattern defines without already knowing what to look for.

Building

Requires a C11 compiler and, for the test suite, Python 3 (used only to generate ground truth from CPython's own re module, concept.md Section 11; the library itself has no runtime dependency beyond the C standard library and libc's wctype.h/locale.h).

make            # builds libregexx.a and the rxgrep example
make test       # regenerates tests/generated_tests.c from Python `re`
                # ground truth and runs the full suite
make check      # same, under AddressSanitizer + UndefinedBehaviorSanitizer
make clean

make install installs libregexx.a and regexx.h under PREFIX (default /usr/local).

Using the library

#include "regexx.h"
#include <string.h>

const char *pattern = "(\\w+)@(\\w+)";
Pattern *pat = re_compile(pattern, strlen(pattern), UTF8, NULL);
Input *in = Input_from_buffer((const uint8_t *)"user@host", strlen("user@host"));

Match m;
if (Pattern_search(pat, in, 0, -1, &m) == 1) {
    const char *g; size_t glen;
    Match_group(&m, NULL, 1, &g, &glen);   /* g/glen -> "user" */
    Match_free(&m);
}

char *out; size_t outlen;
Pattern_sub(pat, in, "\\2@\\1", NULL, NULL, 0, &out, &outlen); /* "host@user" */
free(out);

Pattern_free(pat);
Input_free(in);

flags to re_compile combine a data mode, exactly one of BINARY, ASCII, or UTF8 (concept.md 9.3), with any of the Python-named flags (IGNORECASE, MULTILINE, DOTALL, VERBOSE, ASCII as a flag also forces ASCII-only \w/\s/\d inside UTF8 mode, LOCALE, DEBUG).

Example: rxgrep

examples/rxgrep.c is a small grep-like program built on the library, demonstrating all three data modes and both the matching and substitution API:

./rxgrep -in 'hello' file.txt        # case-insensitive, line numbers
./rxgrep -m utf8 -o '\w+' file.txt   # print every UTF-8 word, one per line
./rxgrep -c 'error' log.txt          # count matching lines
./rxgrep -m binary 'a.c' data.bin    # match raw bytes, embedded NUL included
./rxgrep -m utf8 --sub 'REDACTED' '\d{3}-\d{4}' file.txt

Run ./rxgrep --help for the full option list.

Testing

tests/cases.py lists pattern/subject/operation triples. tests/gen.py computes each one's expected result with CPython's own re module and writes tests/generated_tests.c, which is then compiled against regexx.c and checked. This is a direct implementation of the strategy concept.md Section 11 describes: conformance is measured against what CPython actually does, not against a re-derived reading of its documentation. make check additionally runs the suite under AddressSanitizer and UndefinedBehaviorSanitizer.

License

MIT. See LICENSE.

S
Description
A single-file C regex interpreter with Python re semantics: BINARY, ASCII, and UTF-8 modes, a Pike VM and a memoized backtracking engine, MIT licensed.
Readme MIT
212 KiB
Languages
C 80.5%
Python 17.6%
Makefile 1.9%