concept.md is the full design document: the objective (Python re parity plus binary/ASCII/UTF-8 modes, gigabyte-scale input, single C file), the automata-theory argument for why unrestricted backreferences/lookaround are incompatible with strict single-pass constant memory, the resulting two-engine architecture, the exact Python-mirroring naming convention, and a full comparison against POSIX regex.h for C-background readers. regexx.c/regexx.h are the v1 implementation: parser, compiler to a Pike/backtracking-style bytecode, and a single recursive backtracking engine covering the pattern syntax and operations listed in README.md, validated against CPython's own re module output (tests/), clean under AddressSanitizer/UBSan, and stress-tested (50MB simple-quantifier match, graceful failure rather than a crash on complex repeats over large input, clean rejection of every intentionally unsupported construct). Also included: examples/rxgrep.c (a small grep-like program exercising all three data modes and the substitution API), the Makefile, the MIT LICENSE, and docs/API.md, an exhaustive reference for every type, flag, and function's exact return-value and memory-ownership convention, checked against the current source and against a real CPython interpreter rather than against memory. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
46 KiB
Concept: A Single File C Regex Interpreter with Python re Semantics
1. Objective
The objective is a regex interpreter, implemented as a single C source file, that reproduces the observable behavior of Python's re module (pattern syntax, flags, and the match, search, fullmatch, findall, finditer, split, sub, and subn operations) while being usable on inputs of arbitrary size, including multi-gigabyte files, without holding the entire input in memory. The implementation is required to operate in three data modes: binary (arbitrary byte streams), ASCII text, and UTF-8 text. Simplicity of the C code takes priority over raw execution speed.
Sections 2 through 4 record what "full Python re support" concretely means. Section 5 records why an unrestricted single pass, constant memory implementation of that full feature set is not mathematically possible, and states the boundary precisely. Sections 6 through 12 describe the architecture chosen to get as close to the objective as that boundary allows, in the simplest C design found. Section 13 lists the residual gaps against CPython's re. Section 14 records, for a reader coming from C rather than Python, the respects in which the native C regular expression facility (POSIX regex.h) differs from Python re, and therefore from this design, none of which this design adopts beyond what Section 1 already asks for.
2. Reference Semantics: Python re Pattern Syntax
The engine parses the following syntax, matching CPython's documented and observed behavior.
2.1 Atoms and literals
- Literal characters (bytes in binary/ASCII mode, decoded code points in UTF-8 mode).
.matches any character except\n, or any character at all underDOTALL.\followed by a non-alphanumeric character is that literal character.- Escapes:
\n \r \t \f \v \a \0, octal\ooo, hexadecimal\xhh,\uxxxx,\Uxxxxxxxx, and named code points\N{NAME}(requires a Unicode name table; see 13.3).
2.2 Character classes
[...]with ranges (a-z), negation ([^...]), and literal],-,^when escaped or positionally safe, matching CPython's class parser exactly (including that]as the first class member is literal).- Shorthand classes
\d \D \w \W \s \S, each with an ASCII definition and a Unicode definition, selected by mode and by theASCII/UNICODEflags exactly as CPython selects them forstrversusbytespatterns. \band\B(word boundary and non boundary), defined with the same "word character" set as\win the active mode.
2.3 Anchors
^and$: string boundaries by default; line boundaries underMULTILINE.\Aand\Z: string boundaries, unaffected byMULTILINE.
2.4 Quantifiers
- Greedy:
* + ? {m,n} {m,} {,n} {m}. - Lazy:
*? +? ?? {m,n}?. - Possessive (CPython 3.11 and later):
*+ ++ ?+ {m,n}+. - Atomic groups (CPython 3.11 and later):
(?>...).
2.5 Groups and grouping constructs
(...)capturing group, numbered left to right by opening parenthesis.(?:...)non capturing group.(?P<name>...)named capturing group;(?P=name)named backreference;\1..\99numbered backreference;\g<name>and\g<1>backreference forms usable inside a pattern as well as in a replacement string.(?#...)comment, discarded at parse time.(?=...),(?!...): lookahead, positive and negative, of unrestricted width.(?<=...),(?<!...): lookbehind, positive and negative. CPython requires the lookbehind body to be fixed width (a fixed number of characters, alternation of equal width branches permitted); the engine enforces the same restriction at compile time and rejects variable width lookbehind with a compile error, matching CPython'serror: look-behind requires fixed-width pattern.(?(id)yes|no)and(?(name)yes|no): conditional branch on whether a numbered or named group has already matched; the|nobranch is optional ((?(id)yes)is valid, matching CPython, with an implicit empty "no" branch).(?aiLmsux)global inline flags, valid only at the start of the pattern, and(?aiLmsux-imsx:...)scoped inline flags valid anywhere, matching CPython's restriction that onlyi, m, s, xare removable anda, L, uare global only.|alternation, ordered, first match wins (not longest match), matching backtracking semantics rather than POSIX leftmost longest.
2.6 Flags
IGNORECASE, MULTILINE, DOTALL, VERBOSE (whitespace and # comments outside classes and outside escapes are ignored), ASCII, UNICODE (default for text mode), LOCALE (accepted for compatibility, implemented as byte range [\x80-\xff] word characters under the "C" locale only; see 13.4), DEBUG (accepted, prints the compiled program instead of executing it).
3. Reference Semantics: Python re Operations
match(pattern, string): anchored at position 0, not required to consume the whole string.fullmatch(pattern, string): anchored at position 0 and at the end of the string.search(pattern, string): first match anywhere.finditer/findall: all non overlapping matches, left to right, with the CPython 3.7+ rule that an empty match advances one position and is followed by a search starting immediately after it rather than being merged with an adjacent non empty match at the same start position.split(pattern, string, maxsplit=0): text of capturing groups is interleaved into the result list, matching CPython; splitting on a pattern that can match an empty string is permitted, matching the CPython 3.7+ behavior change.sub(pattern, repl, string, count=0)/subn:replis either a template string honoring\g<name>,\g<1>,\1, and literal backslash escapes, or a callback invoked once per match with a match record and expected to return replacement bytes/text (the C equivalent of a Python callable, see 9.4).escape(string): backslash escaping of all characters outside[A-Za-z0-9_]in the same wayre.escapedoes since Python 3.7 (that version narrowed the escaped set relative to earlier Python releases; the engine follows the narrowed, current set).- Match record fields:
group(n),group(name),groups(),groupdict(),start(n),end(n),span(n),lastindex,lastgroup.
4. Feature Compatibility Table
| Feature | Status |
|---|---|
| Literals, classes, anchors, quantifiers (greedy/lazy) | Full |
| Alternation, grouping, named groups | Full |
| Backreferences (pattern and replacement) | Full, bounded (Section 5) |
| Lookahead, fixed width | Full |
| Lookahead, unbounded width | Full, bounded (Section 5), same window as backreferences |
| Lookbehind, fixed width | Full |
| Lookbehind, variable width | Rejected at compile time, as in CPython |
| Conditional groups `(?(id)yes | no)` |
| Possessive quantifiers, atomic groups | Full (OP_ATOMIC, Section 7.1) |
| Inline and scoped flags | Full |
\N{NAME} named code points |
Partial (Section 13.3) |
Full Unicode \w/\s/IGNORECASE case folding |
Partial (Section 13.3) |
LOCALE flag beyond the "C" locale |
Not supported (Section 13.4) |
| Streaming over unbounded input | Full for the regular subset, bounded window for backreferences/lookaround (Section 5) |
5. The Single Pass, Bounded Memory Constraint
This section states a limit that shapes the rest of the design, so it is recorded before the architecture.
A pattern language restricted to literals, classes, anchors, quantifiers, grouping, and alternation is a regular language. Regular languages are recognized by a finite automaton, and a finite automaton processes an input stream in one pass, in time linear in the input length, using memory bounded by the automaton's state count, independent of input length. Thompson's construction (converting a pattern to a nondeterministic finite automaton) and its simulation without backtracking (the approach used by grep -E, awk, RE2, and Rob Pike's regular expression virtual machine) achieve exactly this, including for findall style capture extraction.
Backreferences (\1, (?P=name)) break this property. No fixed size finite automaton recognizes a language defined with a backreference in general, because such an automaton would need to remember an arbitrarily long previously matched substring verbatim and compare it later, and a finite automaton has, by definition, only finitely many states with which to do so. The precise formal result here is a combined complexity result: deciding whether a string matches a pattern is NP-hard when both the pattern and the string are counted as part of the problem input. That result does not, by itself, say anything about a fixed pattern, compiled once, matched against a growing string, which is this design's actual situation; for a fixed pattern the practically relevant obstacle is different and better known by name, worst case exponential backtracking time in the length of the string, the mechanism behind catastrophic backtracking (commonly called ReDoS) in every production backtracking engine, including CPython's own _sre. Unbounded width lookahead and lookbehind create the same obstacle for a related reason: resolving them can require holding an unbounded span of the input, forward or backward, before the assertion's truth value is known. This is a property of the language class and of the evaluation strategy required to decide it, not of any particular implementation choice, and it means a literal reading of "full Python re support" and "does not remain in memory or hold the input" are mutually exclusive whenever a pattern actually uses a backreference or an unbounded width lookaround on unbounded input.
The engine resolves this by splitting execution into two engines sharing one bytecode format:
-
Regular engine. Any compiled pattern that contains no backreference and no lookaround whose body has unbounded width is executed by a Thompson/Pike style simulation: single pass, one buffered chunk of input at a time, memory bounded by the number of program instructions multiplied by the number of capture groups, independent of input length. This covers the large majority of patterns used in practice, including nested quantifiers, alternation, and fixed width lookaround.
-
Bounded backtracking engine. Any compiled pattern that contains a backreference or an unbounded width lookahead (fixed width lookaround of either polarity always stays in the regular engine, per 7.1) is executed by a backtracking simulation over a sliding window of the input, of a fixed configurable size (default 1 MiB, see 7.3). This engine gives exact CPython semantics as long as the text a backreference or lookahead needs to inspect fits inside the window relative to the current match attempt. If it does not, the engine reports a recoverable error (
ERANGE-style status) identifying the offending construct and offset, rather than silently returning a wrong answer or reading the whole file into memory. This is the same trade every production streaming text tool with backreference support makes; the engine documents the bound instead of hiding it.
This split is decided once, at compile time, from the parsed pattern, before any input is read. A caller who needs a hard guarantee of bounded memory on arbitrary input can inspect the compiled pattern's engine selection before running it.
6. Architecture Overview
Single C file, four sections in this order, each independent of the ones after it:
- Parser: pattern text to abstract syntax tree (AST). Recursive descent, one function per grammar production (
parse_alt,parse_concat,parse_repeat,parse_atom), matching the structure of the grammar in Section 2 directly, so the parser can be read as an executable grammar. - Compiler: AST to bytecode, by direct recursive translation (Thompson's construction), one code generation function per AST node kind. Same bytecode format is emitted regardless of which of the two engines (5.1/5.2) will run it; the engine choice is a separate flag computed from the AST (does it contain
OP_BACKREF, or anOP_LOOKAHEADwhose body width is unbounded; CPython already forcesOP_LOOKBEHINDto be fixed width, Section 2.5, so lookbehind never contributes to this flag). - Engines: the regular engine (Pike VM) and the bounded backtracking engine (recursive backtracking over the sliding window), described in Section 7.
- Public API: the Python
reequivalent entry points, described in Section 9.
Keeping parser, compiler, and both engines as pure functions over explicit structs (no hidden global state except one user supplied allocator, see 8.1) is what keeps the single file simple to read despite covering the full grammar.
7. Execution Engines
7.1 Bytecode
One flat instruction set, an array of tagged structs, used by both engines:
enum opcode {
OP_CHAR, /* match one literal byte/codepoint */
OP_CLASS, /* match one byte/codepoint against a class */
OP_ANY, /* match one byte/codepoint, DOTALL-sensitive */
OP_SPLIT, /* two continuations (alternation, quantifiers) */
OP_JMP,
OP_SAVE, /* record current offset into capture slot N */
OP_MATCH,
OP_ASSERT, /* zero-width: ^ $ \b \B \A \Z */
OP_BACKREF, /* forces bounded backtracking engine */
OP_LOOKAHEAD, /* sub-program, zero-width, polarity flag */
OP_LOOKBEHIND, /* sub-program, fixed width, zero-width, polarity*/
OP_ATOMIC, /* sub-program, consuming, discards choice points */
OP_COND, /* branch on whether group N has matched */
};
struct inst { enum opcode op; int32_t x, y; uint32_t data; };
This is the same instruction shape used by Pike's virtual machine and by RE2's bytecode; reusing it rather than inventing a new one is what keeps the compiler small (roughly one case per AST node).
OP_LOOKAHEAD and OP_LOOKBEHIND each carry, alongside the sub-program pointer and polarity bit, a compile time computed width: a concrete integer for OP_LOOKBEHIND (CPython requires this to exist and be fixed, Section 2.5) and either a concrete integer or an explicit "unbounded" marker for OP_LOOKAHEAD. Only an unbounded width OP_LOOKAHEAD, together with OP_BACKREF, forces engine selection to the bounded backtracking engine (7.3); a fixed width OP_LOOKAHEAD or OP_LOOKBEHIND, of either polarity, is executed by the regular engine (7.2) using a peek buffer sized to exactly that width, never the full window W.
Atomic groups ((?>...)) and possessive quantifiers (*+ ++ ?+ {m,n}+) both compile to OP_ATOMIC, which is the only new opcode either needs: a possessive quantifier is first desugared, at compile time, into the atomic group wrapping its ordinary greedy form (X*+ becomes (?>X*), X{m,n}+ becomes (?>X{m,n}), and so on), so the compiler and both engines only ever have to implement OP_ATOMIC once. OP_ATOMIC runs its sub-program to its single highest priority success (one priority ordered thread simulation restricted to the sub-program in the regular engine, 7.2; one recursive match attempt in the bounded engine, 7.3), advances the current position past whatever it consumed, and then permanently discards every choice point created while matching the sub-program, so that if matching fails later in the overall pattern, the engine never backtracks into the atomic group looking for a different internal match, which is the defining behavior of both constructs in CPython. Unlike OP_LOOKAHEAD, OP_ATOMIC never forces the bounded backtracking engine, regardless of the width of its body: it advances the stream position as it matches, so, unlike a zero-width assertion, it never needs to hold matched text in memory for a decision made later. Only OP_BACKREF and an unbounded width OP_LOOKAHEAD trigger the bounded engine (Section 5, 6).
7.2 Regular engine: Pike VM over a chunk stream
Standard Thompson NFA simulation extended with capture slots, run breadth first ("all current threads advance over the same input character, in priority order, duplicate states are merged"). Per input character the engine holds at most N threads, N being the instruction count, each thread owning only its capture slot array (2 * ngroups offsets), so per character memory is O(N * ngroups), not O(input length).
Streaming adaptation: input arrives as a sequence of chunks (see 7.3) rather than one buffer. Thread capture slots store absolute stream offsets (a 64 bit counter incremented across chunk boundaries), not pointers into the chunk buffer, so a thread survives a chunk boundary without copying. Once every live thread's earliest referenced offset has advanced past a chunk boundary, that chunk is released back to the caller supplied allocator. This is the entire mechanism that lets the regular engine run over an arbitrarily large file in bounded memory: it never needs to look backward, so it never needs to keep anything but the current chunk and the small thread list.
Two small, constant size pieces of state cross a chunk boundary alongside the thread list. First, in UTF8 mode, a partially read multi-byte sequence: at most 3 pending lead bytes (the longest UTF-8 sequence is 4 bytes), carried into the next chunk before character classification resumes; a chunk is only eligible for release once any sequence straddling its end has been completed by the following chunk. Second, under MULTILINE, one bit recording whether the byte immediately before the current chunk was \n, needed to classify ^ at the very first position of a new chunk without rereading the previous one; \n (0x0A) cannot appear as a non-initial byte of any valid multi-byte UTF-8 sequence, so this bit and the UTF-8 continuation state never interact with each other. Neither addition affects the O(instruction count x group count) memory bound of Section 10, since both are O(1) regardless of chunk size or input length.
7.3 Bounded backtracking engine: sliding window
Used only for the minority of patterns containing a backreference or an unbounded width lookahead (7.1). Maintains an explicit ring buffer window of W bytes (default W = 1 MiB, a run time parameter). The window always contains the current match attempt's start position and everything from there forward that has been read so far, up to W bytes. A recursive backtracking matcher, structurally the direct translation of the AST (one function per node kind, exactly as _sre and most textbook backtracking matchers are structured), walks the bytecode against the window. If a match attempt's required span would exceed W, the call returns the documented bounded-window error described in Section 5 instead of growing the window past its configured limit.
Because this engine is only invoked for patterns that need it, ordinary patterns (the large majority) never pay for the ring buffer or for backtracking, and get the linear time guarantee of 7.2 instead.
7.4 Anchoring across chunks
Both engines expose the same chunk boundary contract: a match cannot be reported as final until either (a) OP_MATCH is reached, or (b) enough trailing context has been seen to prove that no continuation of the current input would change the answer for the leftmost still-open match attempt. For quantifiers this is decided directly by the bytecode's OP_SPLIT/OP_JMP shape; for anchors ($, \Z) the final chunk is distinguished by an explicit "end of stream" sentinel token, so $ and \Z behave identically whether or not MULTILINE is set, matching CPython.
8. Core Data Structures
Kept intentionally minimal, all defined in the single file, no dependency beyond the C standard library (stdint.h, stddef.h, string.h):
8.1 Allocator
One struct of three function pointers (alloc, realloc, free) passed once at engine creation, defaulting to the libc equivalents. This is the only piece of "infrastructure" abstraction in the file, and it exists so the sliding window (7.3) and chunk buffers (7.2) can be sized and released under caller control, which is a prerequisite for the gigabyte scale requirement.
8.2 Dynamic array
One generic growable array ({ void *data; size_t len, cap, elemsize; }) with push/get, used for the instruction array, the capture slot array, and the AST node pool. Deliberately not a macro-heavy generic container; three functions (da_init, da_push, da_free) cover every use site in the file.
8.3 Byte class table
A 256 bit set (uint32_t bits[8]) per compiled [...] class or shorthand class, precomputed at compile time. In UTF-8 mode, code points above 127 are matched against a small number of precompiled Unicode range tables (Section 13.3) instead of the 256 bit set.
8.4 Capture slots
A flat array of 2 * groups capture records per thread (regular engine) or per backtracking call frame (bounded engine), -1 meaning unset. In BINARY and ASCII mode a record is a single int64_t byte offset. In UTF8 mode a record is a pair, { int64_t byte_offset; int64_t codepoint_index; }: the byte offset is what is needed to read the matched bytes back out of Input, and the code point index is what is needed to report Match_start/Match_end/Match_span in the same unit Python uses for str subjects (Section 9.1). The code point index is a running counter incremented once per decoded scalar value as the engine advances, so recording it at a SAVE costs one extra integer copy, not a second pass over the input; it changes the constant factor of the O(instruction count x group count) memory bound of Section 10 in UTF8 mode, not its asymptotic class. This single representation backs Match_group, Match_start, Match_end, Match_span, Match_start_byte, Match_end_byte, Match_span_byte, and Match_groupdict in the public API.
9. Public API (Python re equivalents)
9.0 Naming convention
Every public identifier that has a direct counterpart in Python's re module uses that counterpart's exact spelling, not a transliterated or prefixed variant. Concretely:
- The two data types Python's
reexposes,PatternandMatch, are C structs namedPatternandMatch, notregex_torregexx_pattern. - A method Python calls as
pattern_obj.search(...)is writtenPattern_search(Pattern *self, ...); a method called asmatch_obj.group(...)is writtenMatch_group(Match *self, ...). TheType_methodshape is the direct C rendering oftype.method, needed only because C has no bound methods; the two name fragments either side of the underscore are otherwise exactly the Python names. - A module level function such as
re.compile(...)orre.sub(...)is writtenre_compile(...),re_sub(...), and so on: there_prefix stands for the module the function lives in in Python (re.compile), the same relationshipPattern_andMatch_have to their types. - Flag constants (
IGNORECASE,MULTILINE,DOTALL,VERBOSE,ASCII,UNICODE,LOCALE,DEBUG) are#defineorenumconstants with exactly those names, noRE_orREGEXX_prefix, combined with bitwise|exactly asre.IGNORECASE | re.MULTILINEis combined with Python's|. Patternfields are namedpattern,flags,groups,groupindex, matchingre.Pattern.pattern,.flags,.groups,.groupindexexactly.Matchfields are namedstring,pos,endpos,lastindex,lastgroup,re(a pointer back to the owningPattern, matchingre.Match.re), matchingre.Match's attributes exactly.- The exception is
Input(9.1), which has no Python counterpart because Python'srenever streams: it always operates on an in-memorystrorbytesobject.Inputis the one C-only type the design adds, and it is named descriptively rather than after a nonexistent Python name, precisely so it stands out as the one addition a reader should not go looking for in theredocumentation. - The error type is named
PatternError, matching the alias CPython itself introduced forre.error(re.PatternError), rather than a plainerror(which would collide too easily witherrno.h-style conventions) or an inventedregexx_error.
9.1 Declarations
typedef struct Pattern Pattern; /* re.Pattern */
typedef struct Match Match; /* re.Match */
typedef struct Input Input; /* no Python counterpart, see 9.0; left without a
* struct body here because its concrete layout is
* one of the two variants described in 9.2 and no
* code outside the Input implementation itself
* needs to see inside it, unlike Pattern and Match
* whose fields are part of the public, Python-
* mirroring surface. */
typedef void (*MatchIterCb)(void *ctx, const Match *m);
typedef void (*MatchSubCb)(void *ctx, const Match *m, char **out, size_t *outlen);
typedef struct PatternError PatternError; /* re.error / re.PatternError */
struct PatternError {
const char *msg; /* re.error.msg */
const char *pattern; /* re.error.pattern */
int64_t pos; /* re.error.pos */
int64_t lineno; /* re.error.lineno */
int64_t colno; /* re.error.colno */
};
struct Pattern {
const char *pattern; /* re.Pattern.pattern */
int flags; /* re.Pattern.flags */
int groups; /* re.Pattern.groups */
void *groupindex; /* re.Pattern.groupindex, name -> group number */
void *program; /* compiled bytecode, private (7.1) */
};
struct Match {
Pattern *re; /* re.Match.re */
Input *string; /* re.Match.string */
int64_t pos, endpos; /* re.Match.pos, re.Match.endpos; code point indices in UTF8 mode, byte offsets otherwise, see 8.4 */
int lastindex; /* re.Match.lastindex */
const char *lastgroup; /* re.Match.lastgroup */
void *slots; /* private, see 8.4 */
};
/* Pattern methods: the primitives, bound to an already compiled Pattern,
* mirroring re.Pattern.match / .search / .fullmatch / .finditer / .findall /
* .split / .sub / .subn exactly, including their pos/endpos parameters. */
int Pattern_match(Pattern *self, Input *string, int64_t pos, int64_t endpos, Match *out);
int Pattern_fullmatch(Pattern *self, Input *string, int64_t pos, int64_t endpos, Match *out);
int Pattern_search(Pattern *self, Input *string, int64_t pos, int64_t endpos, Match *out);
int Pattern_finditer(Pattern *self, Input *string, int64_t pos, int64_t endpos, MatchIterCb cb, void *ctx);
int Pattern_findall(Pattern *self, Input *string, int64_t pos, int64_t endpos, MatchIterCb cb, void *ctx);
int Pattern_split(Pattern *self, Input *string, int maxsplit, MatchIterCb cb, void *ctx);
int Pattern_sub(Pattern *self, Input *string, const char *repl, MatchSubCb cb, void *ctx, int count, char **out, size_t *outlen);
int Pattern_subn(Pattern *self, Input *string, const char *repl, MatchSubCb cb, void *ctx, int count, char **out, size_t *outlen, int *n);
void Pattern_free(Pattern *self);
/* Match accessors, matching re.Match's bound methods and attributes */
int Match_group(Match *self, const char *name_or_null, int index, const char **out, size_t *outlen);
void Match_groups(Match *self, /* out array of (ptr,len) */ void *out);
void Match_groupdict(Match *self, /* out name -> (ptr,len) map */ void *out);
int64_t Match_start(Match *self, int group); /* code point index in UTF8 mode, byte offset otherwise */
int64_t Match_end(Match *self, int group);
void Match_span(Match *self, int group, int64_t *start, int64_t *end);
int64_t Match_start_byte(Match *self, int group); /* always a byte offset into Input, see 8.4 */
int64_t Match_end_byte(Match *self, int group);
void Match_span_byte(Match *self, int group, int64_t *start, int64_t *end);
void Match_expand(Match *self, const char *template, char **out, size_t *outlen);
/* Module level functions, mirroring re.compile / re.match / re.search / ...
* exactly: each of the search-family functions compiles pattern through an
* internal bounded cache and then calls the matching Pattern_ function,
* exactly as CPython's re/__init__.py implements re.match as
* _compile(pattern, flags).match(string). */
Pattern *re_compile(const char *pattern, size_t len, int flags, PatternError *err);
int re_match(const char *pattern, size_t len, int flags, Input *string, Match *out);
int re_fullmatch(const char *pattern, size_t len, int flags, Input *string, Match *out);
int re_search(const char *pattern, size_t len, int flags, Input *string, Match *out);
int re_finditer(const char *pattern, size_t len, int flags, Input *string, MatchIterCb cb, void *ctx);
int re_findall(const char *pattern, size_t len, int flags, Input *string, MatchIterCb cb, void *ctx);
int re_split(const char *pattern, size_t len, int flags, Input *string, int maxsplit, MatchIterCb cb, void *ctx);
int re_sub(const char *pattern, size_t len, int flags, Input *string, const char *repl, MatchSubCb cb, void *ctx, int count, char **out, size_t *outlen);
int re_subn(const char *pattern, size_t len, int flags, Input *string, const char *repl, MatchSubCb cb, void *ctx, int count, char **out, size_t *outlen, int *n);
void re_escape(const char *in, size_t len, char **out, size_t *outlen);
void re_purge(void);
The dependency runs from module level to Pattern, not the other way around, which is the same direction CPython itself uses: re.py defines match, search, and the rest as thin wrappers that call _compile(pattern, flags) and then the corresponding Pattern method. Each module level function above holds an internal cache keyed by (pattern, len, flags), bounded to a fixed capacity and cleared entirely on overflow rather than evicting individual entries, again mirroring the strategy CPython's own re module cache uses. re_purge() clears that cache on demand, matching re.purge() exactly, including that it has no effect on any Pattern * a caller is still holding a direct reference to. re_compile bypasses the cache and always produces a fresh Pattern, matching re.compile. A caller working against a large or streaming Input and applying the same pattern repeatedly should call re_compile once and use the Pattern_ functions directly, exactly as idiomatic Python precompiles a pattern that is reused in a loop rather than calling the module level function repeatedly; the module level functions exist for parity with re.match/re.search/etc., not as the recommended entry point for the gigabyte scale case this design targets.
9.2 Input
An abstraction over "a source of chunks": either a fixed buffer (small strings, a drop in replacement for the CPython str/bytes case) or a caller supplied read(void *ctx, uint8_t *buf, size_t cap) -> size_t callback (files, pipes, sockets), which is how gigabyte scale input is supplied without ever requiring the caller to load it fully into memory. See 9.0 for why this type does not carry a Python name.
pos and endpos on the Pattern_ functions are expressed in the same unit as Match_start/Match_end for the pattern's mode (code point index in UTF8 mode, byte offset otherwise, Section 8.4). Honoring a nonzero pos against a streaming Input that only exposes sequential read requires decoding forward from the start of the stream until that position is reached, an O(pos) cost paid once per call, not a departure from the per-character bound of Section 10, which is stated per byte of input actually scanned. An Input that also exposes an optional seek(void *ctx, int64_t byte_offset) -> int callback lets the engine skip that decode pass in ASCII/BINARY mode, or in UTF8 mode whenever the caller already knows the target byte offset, for example one returned earlier by Match_start_byte against the same Input.
9.3 Encoding mode
A field of the compile-time flags alongside IGNORECASE, MULTILINE, and the rest: BINARY, ASCII, UTF8 (no RE_ prefix, per 9.0; CPython has no equivalent constant because the choice between binary and text mode is implicit in whether a bytes or str pattern was compiled, so Pattern_compile here makes that same choice explicit through a flag instead). This flag selects, at compile time, which byte class tables (8.3), which ./\w/\s definitions (2.2), and which decoder (raw byte, or the UTF-8 decoder producing code points for classification while recording both a code point index and a byte offset per capture, Section 8.4) the compiled program uses. Binary mode never decodes: every byte value 0 to 255, including embedded NUL, is a valid atom, matching the behavior Python gets by compiling a bytes pattern against a bytes subject.
A byte sequence in UTF8 mode that is not valid UTF-8 at the point the decoder reaches it is handled the way bytes.decode('utf-8', errors=...) is in Python: the default policy, strict, surfaces a PatternError identifying the byte offset of the first invalid byte, matching the fact that CPython can never hand re a str that was not already validly decoded in the first place. An opt-in replace policy substitutes the Unicode replacement character U+FFFD for the offending bytes and continues, for callers that must process untrusted or partially corrupt streams without aborting. BINARY and ASCII mode have no decode step and so have no analogous failure mode; a byte outside 0-127 in ASCII mode is simply a byte no ASCII-mode class matches, not an error.
9.4 Replacement callback
Pattern_sub/Pattern_subn (and the module level re_sub/re_subn that wrap them, 9.1) take both a repl template string and a cb callback of type MatchSubCb as separate parameters, of which exactly one is non-NULL on any given call: a non-NULL repl is parsed once, at compile time, into a small list of literal/backreference segments, mirroring Section 3's replacement syntax; a non-NULL cb is invoked once per match with the match record and a caller supplied ctx, and is expected to write the replacement text through out/outlen, which is the C shape of Python's callable repl argument to re.sub. Two separate parameters, rather than one parameter serving both roles, is the direct C consequence of Python's single repl argument being polymorphic (string or callable) in a way C's static type system cannot express in one slot, the same kind of unavoidable, minimal departure from a one to one name and shape mapping that 9.0 already accepts for Type_method.
10. Complexity Summary
| Engine | Time | Memory | Applies to |
|---|---|---|---|
| Regular (Pike VM, 7.2) | O(input length x instruction count) | O(instruction count x group count), independent of input length | Patterns with no backreference and no unbounded lookaround |
| Bounded backtracking (7.3) | Worst case exponential in window size, as in CPython | O(window size W), independent of input length beyond W |
Patterns with a backreference or an unbounded width lookahead |
Both figures are stated relative to input length specifically because that is the axis the gigabyte scale requirement constrains; instruction count and group count are properties of the pattern, not the input, and are expected to stay small (tens to low hundreds) for realistically written patterns.
11. Testing Strategy
CPython ships its own re test suite (Lib/test/test_re.py / re_tests.py) as executable pattern, string, expected-result triples. The plan is to translate that suite mechanically into a table of C test cases run against re_match/re_search/re_sub, so the engine's conformance is measured against CPython's own stated behavior rather than against a re-derived interpretation of the documentation. Cases that exercise the explicitly out of scope items in Section 13 are recorded as known deviations rather than deleted, so the gap stays visible.
12. Worked Example: Why This Is the Simplest Design That Reaches the Goal
A single unified backtracking engine (matching CPython's own _sre design most closely) would be simpler to write than the two-engine split in Section 5 and Section 7, but it cannot satisfy the gigabyte scale, bounded memory requirement for the common case, because a naive backtracking matcher's stack depth and re-scan behavior scale with input length for ordinary patterns, not only for pattern using backreferences. Conversely, a single unified automaton engine (no backtracking at all) is simpler still, but cannot express backreferences, or the same unbounded width lookaround built from the widening this design already restricts to the bounded engine (Section 5, 7.1), which the objective in Section 1 requires. The two-engine split is the smallest design found that keeps the automaton engine's linear-time, bounded-memory property for the patterns that admit it, while still offering exact backreference and lookaround semantics for the patterns that need them, at an explicit and configurable memory cost.
13. Known Gaps Against CPython re
13.1 POSIX leftmost-longest matching
Not applicable; CPython's re itself uses ordered, first-alternative-wins backtracking semantics, and this engine matches that, not POSIX grep -E semantics. See Section 14.1 for the full comparison against native C regex.h behavior, which this section only touched on briefly before that comparison existed.
13.2 Recursion limit parity
CPython raises RecursionError past a configurable backtracking depth. The bounded backtracking engine (7.3) instead bounds by window size and an explicit call depth counter with a similar default; exact error message parity is not a goal, only the presence of a safe failure mode.
13.3 Full Unicode tables
\N{NAME} lookup, full IGNORECASE case folding (including special casing such as German ß), and complete \w/\s Unicode category coverage require the Unicode Character Database. The concept ships a reduced set of range tables covering the common categories (letters, digits, marks, common whitespace) rather than the full database, to keep the single file small; the compiled tables are generated from the Unicode Character Database offline and checked in as static arrays, with the generation script kept outside the single interpreter file.
13.4 LOCALE flag
Only the "C" locale behavior is implemented (byte range \x80-\xff treated as word characters); full locale.h integration is out of scope because it reintroduces global, environment dependent state into an otherwise pure, single file design.
13.5 regex third party module extensions
Constructs from the third party regex package (set operations inside classes such as --/&&, fuzzy matching, recursive patterns (?R), variable width lookbehind) are not part of CPython's re and are out of scope by Section 1's own definition of "what Python supports."
13.6 Concurrency of the module level pattern cache
The internal cache backing the module level functions of Section 9.1 (re_match, re_search, and the rest) is shared, mutable state. CPython's own equivalent cache is implicitly protected by the GIL; this design has no equivalent, so a caller invoking the module level functions from more than one thread concurrently must serialize access to the cache itself (a mutex around lookup and insertion, sized independently of the allocator in 8.1) or avoid the module level functions entirely and call re_compile once per pattern up front, sharing the resulting read only Pattern * across threads. The latter is already the recommended pattern for the gigabyte scale case (9.1's closing paragraph), so this limitation is expected to be inactive on the path the design is optimized for.
14. POSIX / Native C Regex (regex.h) Capabilities Absent From Python re
Because this project is delivered as a C module, a reader coming from C rather than from Python may reasonably expect it to behave like the regular expression facility native to the C standard library, POSIX regex.h (regcomp, regexec, regfree, regerror, specified by IEEE Std 1003.1). This section records, completely and for that reader specifically, the respects in which POSIX's native facility does something Python re does not do at all. Every item is either omitted by design, meaning it is incompatible with matching Python re's own behavior and is therefore excluded by Section 1's own definition of the target, or out of scope, meaning it would not conflict with Python parity but was not requested and is not free to add under Section 1's simplicity priority. Section 14.6 closes with the one respect in which this design already exceeds POSIX regex.h, included so the comparison is not one sided.
14.1 Leftmost-longest ("POSIX") matching
POSIX regex.h is specified to find the leftmost match and, among matches starting at that leftmost position, the longest one, applied recursively to subexpressions as well as to the overall match. Python re, like every Perl-derived engine, instead uses leftmost-first, ordered-alternation, backtracking semantics: the first alternative that leads to any overall match wins, even when a later alternative would consume more text, and quantifier greediness is resolved by backtracking order rather than by a global longest-match search. These are two different, mutually incompatible definitions of "the match" for the same pattern and string; a|ab against "ab" matches "a" under Python/Perl semantics and "ab" under POSIX semantics. This design follows Python's ordered semantics throughout (Section 2.5, Section 13.1), by the objective in Section 1. Omitted by design: Pike's priority ordered thread simulation (7.2), used here specifically to reproduce Perl style ordered semantics, is a different algorithm from the one POSIX-longest resolution requires, and running both simultaneously would cost the single, simple engine design Section 12 argues for, for a mode nothing in Section 1 asks for.
14.2 POSIX bracket-expression collating symbols and equivalence classes
A POSIX bracket expression may contain [.collating-symbol.] (a named, possibly multi-character collating element treated as one unit, useful for ranges) and [=equivalence-class=] (every character the active locale's collation treats as primary equivalent to the given one). Python's [...] syntax has no counterpart to either: [[.ch.]] and [[=e=]] in a Python pattern parse as plain sets of the literal characters [, ., c, h, ] and [, =, e, ], never as collating constructs, because Python bracket expressions are defined purely over literal characters and ranges, never over locale collation data. Out of scope: these constructs only have observable effect in locales with genuine multi-character collating elements, which is rare in practice, Python's re has never implemented them for str or bytes patterns, and adding them would require linking the compiled tables of Section 8.3 to the system's LC_COLLATE data, in direct tension with the dependency-free, pure-function design of Section 8.
14.3 POSIX named character classes inside bracket expressions
POSIX bracket expressions accept the twelve standard named classes, alpha, digit, alnum, upper, lower, space, blank, cntrl, graph, print, punct, xdigit, written as [:name:] inside a bracket expression, for example [[:alpha:][:digit:]]. Python's re has no equivalent syntax; the same intent is expressed with ranges and the shorthand classes of Section 2.2 instead ([a-zA-Z], \d, \s), and [[:alpha:]] in Python parses as a literal set containing :, a, l, p, h, [, ]. Out of scope by direct consequence of Section 1: Python re genuinely has no such syntax, so parity with Python re already excludes it; it is recorded here only because a C-background reader is likely to look for it and be surprised to find it silently absent rather than documented.
14.4 Locale collating sequence for bracket-expression ranges
POSIX specifies that a bracket-expression range such as [a-z] is resolved according to the current locale's collating sequence (LC_COLLATE), not according to raw code point or byte value order; in a locale whose collation is not a simple ascending code point order, [a-z] can therefore include, exclude, or reorder characters relative to what it means in the "C" locale, a well known source of surprising results in POSIX tools run under a non-"C" locale. Python's re never does this: a range in a Python pattern is always defined by code point value (str patterns) or byte value (bytes patterns), unconditionally, in every locale. This design follows Python exactly, ranges are always resolved by code point or byte value (Section 2.2), independent of the LOCALE flag (Section 13.4). Omitted by design, for the same reason LOCALE itself is restricted to the "C" locale in Section 13.4: honoring arbitrary system collation would reintroduce global, environment dependent state into a design that is otherwise a pure function of its inputs, and would make the meaning of a compiled Pattern depend on a process-wide setting instead of on the flags given to re_compile.
14.5 REG_NOSUB: compiling a pattern that reports no subexpression positions
regcomp(..., REG_NOSUB) compiles a pattern that reports only whether it matched, not where its subexpressions matched, letting an implementation skip the capture bookkeeping of Section 8.4 entirely. Python's re has no equivalent compile-time flag; a compiled Pattern always reports full match and group data through Match. Out of scope: this is a pure performance optimization with no observable behavior difference, and Section 1 places convenience above performance throughout, so there is nothing here worth the added compile-time flag and the second code path it would require in both engines.
14.6 Where this design already exceeds plain POSIX regex.h
regexec takes a nul-terminated C string, so plain POSIX regex.h, on most implementations, cannot search a subject containing an embedded NUL byte at all; the widely available but non-standard REG_STARTEND extension (present in glibc and the BSDs, not part of IEEE Std 1003.1 itself) works around this only on the platforms that provide it. Python's re has never had this limitation, since str and bytes are always explicit-length, never nul-terminated, and this design follows Python and inherits the same freedom from it directly: Input (Section 9.2) is always an explicit-length byte source, and BINARY mode (Section 9.3) explicitly allows an embedded NUL as an ordinary byte value, matching the gigabyte scale, arbitrary-binary-data objective of Section 1. This item is placed last, and out of sequence with the rest of Section 14's "absent from Python re" framing, specifically so the comparison in this section is accurate in both directions rather than reading as one sided.