This document records, exhaustively, the complete public surface of
`regexx.h`/`regexx.c`: every type, every flag, every function, its exact
return-value convention, its memory ownership rule, and its relationship to
the corresponding Python `re` name. `concept.md` records the design
rationale; `README.md` is the short entry point (build, usage, implementation
status at a glance). This document is the complete reference the other two
point to when a precise answer is needed.
Every fact below was checked against the current source (`regexx.c`,
`regexx.h`) and, where a Python behavior is cited, against a real CPython 3
interpreter, not against memory or documentation alone.
## 1. Types
### 1.1 `Pattern` (`re.Pattern`)
```c
structPattern{
constchar*pattern;/* re.Pattern.pattern */
intflags;/* re.Pattern.flags */
intgroups;/* re.Pattern.groups */
void*groupindex;/* re.Pattern.groupindex, opaque, see 3.11 */
void*program;/* private: compiled bytecode */
};
```
-`pattern`: the exact source text passed to `re_compile`, NUL-terminated, owned by the `Pattern` (freed by `Pattern_free`). Read-only for callers.
-`flags`: the exact `flags` value passed to `re_compile`, including the encoding-mode bits (`BINARY`/`ASCII`/`UTF8`) and any leading global inline flags folded in during parsing (Section 2.6).
-`groups`: the number of capturing groups, matching Python's `re.Pattern.groups` exactly (group 0, the whole match, is not counted).
-`groupindex`: opaque; see `Pattern_groupindex_lookup` (3.11) for the only supported access to it.
-`program`: private, never dereference directly.
A `Pattern` is created only by `re_compile` (directly, or indirectly through the module level cache in `re_match`/`re_search`/etc.) and is freed only by `Pattern_free`.
-`re`: the `Pattern` that produced this match. Borrowed reference; do not free it through this pointer, and do not call `Pattern_free` on it while any `Match` from it is still in use.
-`string`: the `Input` the match was found in. Borrowed reference, same lifetime rule as `re`.
-`pos`, `endpos`: the effective search bounds used to produce this match (the `pos`/`endpos` arguments to whichever `Pattern_` function created it, clamped to `[0, length]`), in the same unit as `Match_start`/`Match_end` (code point index in `UTF8` mode, byte offset otherwise). Matches `re.Match.pos`/`.endpos` exactly.
-`lastindex`: the highest-numbered capturing group that participated in the match, or `-1` if none did. **Deviation from CPython:** Python's `lastindex` is "the index of the last group to match", which for a pattern with re-entrant alternation can differ from "the highest-numbered group that participated"; this build uses the latter, simpler rule. They coincide for every straightforward pattern (any pattern without a capturing group inside a repeated alternative that can also match via a different, lower-numbered branch later in the same attempt).
-`lastgroup`: the name of that group, or `NULL` if it is unnamed or `lastindex` is `-1`. Borrowed pointer into the `Pattern`'s group name table; valid as long as the `Pattern` is.
-`slots`: private.
A `Match` is filled in by one of the `Pattern_`/`re_` matching functions (never allocated separately by the caller: pass the address of a stack or heap `Match` struct as `out`, zero-initialize it first). Its private `slots` are freed by `Match_free`, which does **not** free the `Match` struct itself (the caller owns that memory, stack or heap).
### 1.3 `Input` (no Python counterpart)
Opaque. Python's `re` never streams; it always operates on an in-memory `str`/`bytes` object already held by the caller. `Input` is the type this design adds so a C caller has an explicit thing to construct from a buffer or a file (`concept.md` 9.0/9.2).
-`Input_from_buffer`: wraps an existing buffer. **Does not copy it and does not take ownership.** The buffer must outlive the `Input` and every `Match` produced from it (`Match_group` returns pointers directly into it, Section 3.9). Freeing an `Input_from_buffer``Input` never frees the underlying buffer; the caller is responsible for that.
-`Input_from_file`: for a regular file, `mmap`s it read-only (`MAP_PRIVATE`) instead of copying it into the heap, so the pages are backed by the file and the kernel can reclaim them under memory pressure (`README.md` "Memory footprint"). For a non-seekable source (pipe, FIFO, process substitution, `/dev/stdin`) or when `mmap` itself fails, it falls back to reading incrementally into a growable, owned heap buffer. Either way, the entire input is addressable before any matching happens (`README.md` "Implementation status"); only the second path actually copies it. On failure (cannot open, cannot allocate) returns `NULL` and, if `err` is non-`NULL`, fills it with `strerror(errno)` as `msg`.
-`Input_free`: frees the `Input` struct itself always, and additionally releases the underlying buffer unless it came from `Input_from_buffer`: `munmap`s it if it was mapped, or `free`s it if it was read into an owned heap buffer (both cases of `Input_from_file`).
Named after the alias CPython itself introduced for `re.error` (`re.PatternError`, `concept.md` 9.0). Fields match CPython's `re.error` attributes (added in CPython 3.5) exactly: `msg` is the human-readable message, `pattern` is the offending pattern text (or `NULL` when the error is not about pattern syntax, for example `Input_from_file` failing to open a file), `pos` is the byte offset into `pattern` the error was detected at, `lineno`/`colno` are computed from `pos` the same way CPython computes them (1-based, counting `\n` bytes in `pattern` up to `pos`).
**Ownership:**`msg` and `pattern` are heap allocated (`strdup`) by whichever call filled the struct in. Zero-initialize a `PatternError` before passing its address in, and call `PatternError_free` on it once done reading it, whether or not the call that filled it in succeeded (a `NULL``err` argument to any function is always safe to pass and simply skips error reporting). `PatternError_free` is safe to call on an all-zero or already-freed `PatternError`.
Every function that can fail accepts an optional `PatternError *err` (pass `NULL` to ignore); `re_compile` and `Input_from_file` are the two that actually produce one today. `re_match`/`re_search`/etc. (the module level convenience functions, 3.12) do not expose a `PatternError` parameter at all, matching the fact that CPython's own `re.match`/`re.search`/etc. do not return one either (a syntax error there raises, which has no C equivalent; here it instead causes the call to return `-1` with no further diagnostic, which is why `Pattern`-based usage, not the module level convenience functions, is recommended whenever a compile error needs to be reported, `concept.md`/README "recommended entry point" note).
`MatchIterCb` is invoked once per result by `Pattern_finditer`, `Pattern_findall` (an alias of `finditer` in this build, 3.7), and `Pattern_split` (3.8, with a different per-call meaning, documented there). The `Match` passed in is only valid for the duration of the call; it is freed immediately after the callback returns, so do not retain the pointer.
`MatchSubCb` is the C shape of Python's callable `repl` argument to `re.sub`/`re.subn`: invoked once per match, expected to `malloc` a replacement buffer, write its address into `*out` and its length into `*outlen`. The callback's `*out` becomes owned by `Pattern_sub`/`Pattern_subn`, which frees it after copying its content into the final result buffer.
## 2. Flags
Passed as `flags` to `re_compile` or any `re_` module level function, combined with bitwise `|`, using the exact spelling CPython uses (`concept.md` 9.0).
| Flag | Value | Effect | Status |
|---|---|---|---|
| `IGNORECASE` | `0x0001` | Case-insensitive literal, class, and backreference comparison. Class ranges are matched under both a character's original case and its swapped case (README "Known deviations": an approximation, not full Unicode case folding). | Implemented |
| `MULTILINE` | `0x0002` | `^`/`$` also match at the start/end of each line, not only the start/end of the string. | Implemented |
| `VERBOSE` | `0x0008` | Unescaped whitespace and `#`-to-end-of-line comments outside character classes are stripped from the pattern before parsing. | Implemented |
| `ASCII` | `0x0010` | Two roles: (a) with neither `BINARY` nor `UTF8` also given, selects the `ASCII`**encoding mode** (Section 2, below); (b) in `UTF8` mode, forces `\d`/`\w`/`\s` and `IGNORECASE` swapcase to their ASCII-only definitions instead of consulting `wctype.h`. | Implemented |
| `UNICODE` | `0x0020` | Accepted for source compatibility with Python. **No effect**: `UTF8` mode already behaves as CPython's default (Unicode) `str` matching, so there is no separate "unicode" flag needed the way `ASCII` needs one to opt out. | Accepted, no-op |
| `LOCALE` | `0x0040` | Accepted for source compatibility with Python. **No observable effect**: this build commits to the "C" locale only (`concept.md` 13.4), and under the "C" locale, `\w`/`\b`/`\B` classify only the ASCII letters and digits regardless of this flag, matching both the C standard's own guarantee for `isalnum()` under "C" and a real CPython interpreter's `re.LOCALE` explicitly set to "C" (verified directly; an earlier revision of this build instead treated bytes `0x80`-`0xFF` as word characters under this flag, based on an unverified assumption about "C" locale behavior that turned out to be false, caught by the combinatorial test expansion, corrected in `regexx.c`'s `cls_is_word`). | Accepted, no-op |
| `DEBUG` | `0x0080` | Accepted for source compatibility with Python. **No effect** in this build: nothing is printed and matching is unaffected. | Accepted, no-op |
| `BINARY` | `0x0100` | Encoding mode: raw bytes, `0`-`255` all valid, no decoding, embedded `NUL` is an ordinary byte. | Implemented |
| `UTF8` | `0x0200` | Encoding mode: input is decoded as UTF-8 into code points before matching; `Match_start`/`Match_end` report code point indices (`Match_start_byte`/`Match_end_byte` report byte offsets, Section 3.9). | Implemented |
Exactly one encoding mode is active for any compiled `Pattern`: `BINARY` and `UTF8` are checked first, in that order, and if neither is present the mode is `ASCII` regardless of whether the `ASCII` flag bit itself was set (so `re_compile(p, n, 0, &err)` and `re_compile(p, n, ASCII, &err)` compile to the same encoding mode; only combining `ASCII` with `UTF8` changes anything, per row (b) above).
## 3. Functions
Return value convention used throughout, unless noted otherwise for a specific function: `1` success/matched, `0` no match (not an error), `-1` an error occurred (invalid UTF-8 in the subject for a `UTF8`-mode `Pattern`, or the backtracking depth limit was reached, `README.md` "Implementation status"; no further detail is available through the return value itself in this build, only through the fact that it is negative rather than `0`).
Compiles `pattern` (`len` bytes, need not be `NUL`-terminated) under `flags` into a fresh, independent `Pattern`, matching `re.compile`. Returns `NULL` and fills `err` (if non-`NULL`) on a syntax error, an unsupported construct (Section 5 of this document lists all of them), or a lookbehind that is not fixed-width. Never consults or populates the module level cache (3.12); always allocates a new `Pattern`, freed only by `Pattern_free`.
Mirror `re.Pattern.match`/`.fullmatch`/`.search` exactly, including the `pos`/`endpos` parameters (pass `0` and `-1` for CPython's own defaults, "search the whole string"). `pos`/`endpos` are in the pattern's native unit (code point index for `UTF8` mode, byte offset otherwise, Section 1.2); a negative `endpos` means "to the end". `out` must point to a zero-initialized `Match` (or one already released with `Match_free`); on a `0` or `-1` return it is left untouched.
-`match`: anchored at `pos`, need not reach `endpos`.
-`fullmatch`: anchored at `pos`, must also reach exactly `endpos`.
-`search`: tries every start position from `pos` to `endpos` inclusive, left to right, and reports the first that admits any match (ordinary backtracking priority decides which match that is at that position, `concept.md` 2.5).
**Performance:** which of two engines runs a given `Pattern` is decided once, at `re_compile` time, and is never a caller's choice (README.md "Implementation status", `concept.md` 7.2/7.7). A pattern with no backreference, lookahead, lookbehind, or atomic group (a possessive quantifier already desugars to the last of these) runs on the Pike VM, a Thompson-NFA simulation with no backtracking at all: every operation in this section is genuinely linear in input length on that engine, including `search` against an adversarial pattern shape like `(a+)+b` that would otherwise invite catastrophic backtracking, with no atomic group needed (README.md "The Pike VM"). Everything else (a backreference anywhere, or a lookaround/atomic construct) runs on the recursive backtracking engine instead: `match`/`fullmatch` there do one anchored attempt at time proportional to that attempt; `search` tries each candidate start position as a separate attempt but, unlike a naive backtracking search, does not redo the same work at every one, since `run_memo` caches every proven failure at the (instruction, position) level for a backreference-free pattern on this engine, and `OP_REPEAT1` additionally uses precomputed skip-ahead tables, together making `search` linear rather than quadratic for a pattern built from simple repeated atoms; a repeat over a *compound* body only gets the failure-memoization, and a pattern with a backreference disables memoization entirely (unsound there, Section 5) and can still be worst-case exponential, exactly as in CPython. The Pike VM has a literal prefilter (a pattern beginning with a mandatory literal or class skips injecting a new search attempt at a position that atom cannot match, `concept.md` 7.9) but still no lazy DFA state caching or allocation pooling, so an ordinary, non-adversarial pattern that happens to be Pike VM eligible can still measure slower in absolute terms than the same pattern would have on the backtracking engine, a real, documented trade-off (`concept.md` 7.8/7.9), narrower than before the prefilter but not closed, not an oversight either way. See README.md "Implementation status" and "The Pike VM" for the measured numbers on both engines and citations to the published techniques each uses.
**Memory:** the tables above cost real, measured memory, not just time complexity: `search`/`finditer`/`split`/`sub` on a pattern using `OP_REPEAT1` peak at roughly 8.4x the input length (measured with Valgrind/Massif; `match`/`fullmatch` do not allocate these tables at all and use proportionally less). README.md "Memory footprint" has the full measured breakdown, including a fix that was tried, measured, and deliberately reverted because it traded a fast allocation failure for an effectively-unbounded hang, which is a worse failure mode, not a better one.
`Pattern_findall` is defined as a call to `Pattern_finditer` with the same arguments; both invoke `cb(ctx, m)` once per non-overlapping match, left to right, applying CPython's own empty-match rule exactly, including the part of it CPython does not document (`concept.md` 3, `regexx.c`'s `Pattern_finditer` comment): if the match found at a given start position is empty, a second, non-empty match is additionally searched for and reported at that same start before the scan moves on, so `\d*?` against `"123abc456"` yields 16 matches, not 9, matching a real CPython interpreter exactly (verified by generating this case's expectation from one, `tests/cases.py`). Returns the number of matches found, or `-1` on error. Shares one memoization table and one set of `OP_REPEAT1` precomputed tables across the whole scan (built once, not once per match), so it inherits `search`'s performance characteristics in Section 3.2-3.4 exactly, not a worse case from repeating the scan. This build does not collapse a no-groups match down to "just the matched string" or a multi-group match to a Python tuple the way `re.findall` does at the Python level; the callback always receives a full `Match`, from which the caller reads whatever it needs via `Match_group`. This is a deliberate simplification: `findall`'s string/tuple collapsing is a Python-object-model convenience with no C equivalent to collapse into, so this build gives the caller the same, uniform `Match`-based access `finditer` does, and the two functions exist separately only for name-for-name parity with `re.findall`/`re.finditer`.
`cb` is invoked once per element of the list `re.split()` would return, **in order**: this is the entire contract, and it is unambiguous by construction (earlier drafts of this function called `cb` twice per match with the caller left to infer which call meant what; that design was replaced before release specifically because it was ambiguous). Each element's text is read via `Match_group(m, NULL, 0, &out, &outlen)`; a `0` return from that call means this element is Python's `None` (an unparticipated capturing group between two matches), matching how an unparticipated group reports on any other `Match`. `maxsplit` matches `re.split`'s parameter (`0` means unlimited). Returns the number of matches that were split on (not the number of list elements), or `-1` on error. Built on the same scan as `Pattern_finditer` and follows the same empty-match rule (Section 3.5-3.6), which is why a pattern that can match empty (`\d*?`, `x*`, and so on) produces the long runs of empty-string list elements a real CPython `re.split()` does, not a shorter list that only advances once per empty match. Shares one memoization table and one set of `OP_REPEAT1` precomputed tables across the whole scan; see Section 3.2-3.4's performance note.
Mirror `re.Pattern.sub`/`.subn`. Exactly one of `repl` (a template string) or `cb` (a callback) must be non-`NULL`; passing both or neither is a caller error with unspecified behavior. This two-parameter shape is the direct C consequence of Python's single `repl` argument being polymorphic (string or callable) in a way C's static type system cannot express in one slot (`concept.md` 9.4, 9.0).
`repl` template syntax: `\g<name>`, `\g<N>`, `\N` (one or two digits), `\n`, `\t`, `\\`, and any other `\X` as the literal character `X`, matching `concept.md` Section 3.
`count` matches `re.sub`'s `count` parameter (`0` means unlimited; a positive `count` stops substituting after that many matches, leaving the rest of the subject, including any further matches within it, untouched, exactly as CPython leaves it).
`Pattern_sub` and `Pattern_subn` differ only in whether the number of substitutions actually made is reported back through `n` (mirroring `re.sub` returning just the string versus `re.subn` returning `(string, count)`); `Pattern_sub` is implemented as a call to `Pattern_subn` with a throwaway `n`. Both are implemented on top of `Pattern_finditer` and so share its performance characteristics, Section 3.2-3.4's performance note.
`*out` is a freshly `malloc`'d, `NUL`-terminated buffer of length `*outlen`; the caller must `free` it. On `-1` (error), `*out`/`*outlen` are left untouched.
### 3.10 `Pattern_free`
```c
voidPattern_free(Pattern*self);
```
Frees a `Pattern` and everything it owns (the compiled program, the retained parse tree, `groupindex`, the copy of the pattern text). Do not call this while any `Match` produced from this `Pattern` is still in use (`Match.re`/`Match.lastgroup` borrow from it, Section 1.2); free every such `Match` with `Match_free` first, or simply free them in the reverse order they were created, which is always safe.
Looks up a named group in `re.Pattern.groupindex`; returns its 1-based group number, or `-1` if no group by that name exists in this pattern. This is the only supported access to `groupindex`'s content in this build: there is no enumeration function (no way to list every name a pattern defines without already knowing what to look for), unlike Python's `groupindex`, which is a full mapping object supporting iteration and `len()`. A caller that needs every name a pattern uses must track the names it compiled the pattern with itself.
Each compiles `pattern` through an internal cache and then calls the matching `Pattern_` function with `pos=0`, `endpos=-1` (module level `re.match`/`re.search`/etc. do not expose `pos`/`endpos` either, only the `Pattern` methods do, matching CPython exactly), exactly as CPython's own `re/__init__.py` implements `re.match` as `_compile(pattern, flags).match(string)`. The cache is keyed by `(pattern, len, flags)`, holds up to 512 entries, and is cleared entirely on overflow rather than evicting individual entries (mirroring the strategy CPython's own `re` module cache uses). A compile error inside these functions is silently reported as a `-1` return, with no `PatternError` available (Section 1.4); use `re_compile` plus a `Pattern_` function directly whenever a compile error needs to be diagnosed, or whenever the same pattern is applied more than once (idiomatic Python precompiles a pattern reused in a loop rather than calling the module level function repeatedly, and so should idiomatic use of this API, `concept.md` 9.1).
**Concurrency:** the cache is shared, mutable, process-wide state with no internal locking (`concept.md` 13.6). Do not call any `re_`-prefixed module level function from more than one thread without external synchronization; `Pattern_`-prefixed functions on a `Pattern` no thread is concurrently modifying (which is all of them, since nothing here mutates a compiled `Pattern`) have no such restriction.
### 3.21 `re_purge`
```c
voidre_purge(void);
```
Clears the module level cache, matching `re.purge()` exactly, including that it has no effect on any `Pattern *` a caller already holds a direct reference to (only the cache entry is dropped; already-returned pointers remain valid until `Pattern_free`d).
Matches `re.escape` exactly, including the narrowed escaped-character set CPython adopted in 3.7 (only characters that are actually special in a regex, plus non-ASCII bytes are passed through unescaped rather than every non-alphanumeric character as in pre-3.7 Python). `*out` is a freshly `malloc`'d buffer the caller must `free`.
`Match_group`: if `name_or_null` is non-`NULL`, `index` is ignored and the group is looked up by name (via the same table `Pattern_groupindex_lookup` uses); otherwise `index` (`0` for the whole match) selects the group directly. Returns `1` and sets `*out`/`*outlen` to a borrowed pointer into the underlying `Input`'s buffer (valid as long as both the `Match` and the `Input` are) when the group matched; returns `0` and sets `*out = NULL, *outlen = 0` when the group exists but did not participate (Python's `None`); returns `-1`, leaving `*out`/`*outlen` untouched, when no such group exists at all (by index or by name).
`Match_start`/`Match_end`/`Match_span`: report the group's span in the pattern's native unit (code point index in `UTF8` mode, byte offset otherwise). **Deviation from CPython:** Python's `Match.start(group)`/`.end(group)` raise `IndexError` for an out-of-range group number and return `-1` only for a valid, unparticipated group; this build returns `-1` for both cases uniformly, since C has no exception to raise. A caller that must tell "no such group" apart from "this group did not participate" should use `Match_group` instead, which does distinguish them (`-1` versus `0` above).
`Match_start_byte`/`Match_end_byte`/`Match_span_byte`: always report a byte offset into the `Input`, regardless of encoding mode; identical to the non-`_byte` accessors in `BINARY`/`ASCII` mode, and the byte-offset translation of the same span in `UTF8` mode. Use these, not the code-point ones, whenever the result will be used to slice or seek into the raw `Input` buffer or file (`Match_group` already does this translation internally, so most callers only need these directly for cases `Match_group` does not cover, such as reporting a byte offset to an external tool).
`Match_free`: frees the private per-match state (capture slots and, in `UTF8` mode, the code-point-to-byte-offset table). Does **not** free the `Match` struct itself, and does not affect `Match.re`/`Match.string` (borrowed, Section 1.2). Safe to call more than once on the same `Match` (the second call is a no-op, since the first sets `slots` to `NULL`).
## 4. Memory ownership summary
| Object | Created by | Freed by | Notes |
|---|---|---|---|
| `Pattern *` | `re_compile` | `Pattern_free` | Never returned by the `re_`-prefixed module level functions; those only take a pattern string, they do not hand back the `Pattern` they compiled internally. |
| `Input *` | `Input_from_buffer` / `Input_from_file` | `Input_free` | `Input_from_buffer` never owns the wrapped buffer; `Input_from_file` always owns the buffer it read. |
| `Match` (the struct) | The caller (stack or heap) | The caller | Never allocated by the library; only its private `slots` are, released by `Match_free`. |
| `PatternError` (the struct) | The caller (stack or heap) | The caller | Only `msg`/`pattern` are heap allocated; release them with `PatternError_free`. |
| `*out` from `Pattern_sub`/`subn`, `re_escape`, and a `MatchSubCb`'s own `*out` | The library (or, for the callback, the callback itself) | The caller (or, for the callback's `*out`, `Pattern_sub`/`subn`, immediately after copying it) | Plain `malloc`'d buffers; `free()` them normally. |
| Text returned by `Match_group` | Borrowed from the `Input`'s buffer | Nobody (not a separate allocation) | Valid exactly as long as the `Input` and the `Match` both are. |
## 5. Rejected constructs
These fail `re_compile` with a `PatternError` naming the construct, rather than being silently mis-parsed. See `concept.md` Section 4/13 for why each is out of scope for this build specifically (as opposed to out of scope for Python `re`, Section 6 below).
- Scoped inline flags: `(?i:...)`, `(?imsx-imsx:...)` (global inline flags at the very start of the pattern, `(?aiLmsux)`, are supported).
- Named code points: `\N{NAME}`.
- Variable-width lookbehind: `(?<=...)`/`(?<!...)` whose body is not a single, statically known width (matches CPython's own restriction, not an additional one this build adds).
## 6. What is not Python `re` syntax at all
POSIX bracket-expression syntax (`[[:alpha:]]`, `[.collating-symbol.]`, `[=equivalence-class=]`) is rejected the same way, but for a different reason: it is not part of Python `re`'s grammar in the first place (`concept.md` Section 14 is the complete comparison against POSIX `regex.h`, for a reader coming from C who might otherwise expect it).
## See also
- [`../concept.md`](../concept.md): the full design document (why a two-engine split was planned, the automata-theory argument for it, the POSIX comparison).
- [`../README.md`](../README.md): build instructions, the `rxgrep` example, and the "Implementation status" summary this document expands on.