Files
retoorandClaude Sonnet 5 e4fc067aa6 Eliminate the Pike VM's per-thread allocation, add a real BMH literal
prefilter, research and reject a lazy DFA cache

Researched two production Pike-VM-family engines directly (RE2's
nfa.cc, rust-lang/regex's regex-automata), fetched and read, not
recalled, specifically for what explains the two weaknesses the last
benchmark found (a*b's nullable loop, dense finditer): both avoid
per-thread malloc/free, RE2 via a Thread free list, regex-automata via
a flat SlotTable indexed directly by NFA state. Also researched
Hyperscan's Teddy (SIMD literal matching) and set it aside: it needs
platform-specific intrinsics, in tension with this project's
single-file, portable-C, simplicity-over-performance priority: a
portable Boyer-Moore-Horspool skip was judged proportionate where
Teddy was not.

Rewrote the Pike VM's memory model accordingly. PikeThread's owned
int64_t* is gone; a committed thread's capture row now lives at a
fixed offset in PikeList.table, indexed directly by instruction PC
(Cox's one-thread-per-PC invariant already made this index unique, so
no allocation is needed to store or discard one). Transient rows
needed mid-closure, before the eventual terminal PC is known, come
from ScratchPool, a fixed block with an explicit free list; a plain
bump/decrement counter was tried first and proven incorrect by hand
before being written into the file (an OP_SPLIT's second branch's row
can outlive several non-branching pushes that reuse an *earlier* row
without reallocating it, which only a real free list handles safely
regardless of release order). The closure stack is pre-sized once
instead of grown by realloc on demand. All three, plus a reusable
best-match buffer and the prefilter below, bundle into one PikeEngine,
built once per top level Pattern_/re_ call and reused across every
match found within it, never cached on PatternImpl itself (that would
make concurrent Pattern_search calls on the same compiled Pattern from
different threads race on shared state, breaking the existing
no-synchronization-needed guarantee for a Pattern nothing mutates).

Extended the literal prefilter from a single leading character to the
full mandatory literal prefix, with a real Boyer-Moore-Horspool
bad-character skip table for BINARY/ASCII mode (UTF8 keeps a
without-skip fallback: a byte-indexed table cannot cover code points
past 0x10FFFF). The skip-ahead only ever applies to where a new
unanchored start is injected, never to advancing sp itself while a
thread from an earlier start position is still alive.

Researched and did not build a lazy DFA state cache (memoizing a
live-instruction-set-plus-byte transition). Not an omission: RE2's own
lazy DFA cannot track submatch boundaries, the same structural reason
applies here, since every call wants at least group 0's span, and a
cached transition only answers whether a match is possible, not which
path was taken; using one would need a two-phase architecture deserving
its own research-and-plan pass. Checked empirically too: a*b, the case
such a cache would help most, has at most two live instructions at any
position for its whole run, so there is no repeated state worth
caching in the first place.

Verified: full 3,252-case suite (four clean runs), a whitebox
dual-engine cross-check extended to also cover finditer/split/BINARY
mode (30,000 + 3,334 + 2,500 comparisons, zero mismatches), a targeted
suite for the new prefilter machinery including the classic
Boyer-Moore-Horspool overlapping-suffix correctness trap (12/12), two
clean AddressSanitizer/UndefinedBehaviorSanitizer passes on each (a
third run of each hit the same pre-existing sandbox flake already
documented, confirmed unrelated by retrying clean).

Measured: literal search went from 0.64x of the backtracking engine's
time (already ahead) to 0.03x (~33x faster), and against POSIX
<regex.h> from roughly 11x slower to roughly 2x *faster* than glibc
outright; a*b (no prefilter benefit at all) improved from 3.45x slower
than backtracking to 1.97x, from the allocation fix alone; dense
finditer over [0-9]+/\w+ improved from 3.2x/5.6x slower to 1.7x/2.8x.
Full tables and citations in concept.md 7.10; README.md, docs/API.md,
USAGE.md, and bench_vs_posix.c's own printed summary updated
throughout with the corrected numbers and the full history, not just
the final ones.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
2026-09-14 18:49:21 +00:00

29 KiB

regexx usage guide

This document is task oriented: how to build a pattern, run it against data in each of the three supported modes, and read the result back out, with complete, compiled, and run examples. docs/API.md is the exhaustive reference for every function's exact return value and memory ownership rule; this document exists to get from "nothing" to "working code" without reading that whole reference first. concept.md records the design rationale for anyone who wants to know why a given choice was made rather than only what it is.

Every example below was compiled against the current regexx.c and its actual printed output is shown, not a hand-written guess at what it would print.

1. Building and linking

make                    # builds libregexx.a and the rxgrep example

A program using the library needs regexx.h on its include path and libregexx.a (or regexx.c compiled directly into the program, since the library is a single file) linked in:

cc -std=c11 -I/path/to/regexx -o myprog myprog.c /path/to/regexx/libregexx.a

No third-party dependency is required; regexx.c uses only the C standard library plus, for UTF8 mode's Unicode classification and Input_from_file, POSIX (wctype.h, locale.h, mmap/open/fstat, concept.md Section 8).

2. The three data modes

Exactly one of BINARY, ASCII, or UTF8 is passed in flags to re_compile (or any re_-prefixed module level function); it selects how the subject's bytes are interpreted, not how the pattern text itself is encoded (the pattern is always read as ASCII/UTF-8 source text regardless of this choice).

Mode Subject interpretation Match_start/Match_end unit
BINARY Raw bytes, 0-255 all valid, embedded NUL is an ordinary byte, no decoding at all. Byte offset
ASCII Bytes, but \d/\w/\s and IGNORECASE restrict themselves to the ASCII subset (the default when neither BINARY nor UTF8 is given, docs/API.md Section 2). Byte offset
UTF8 Decoded as UTF-8 into Unicode code points before matching; \d/\w/\s and case folding consult the platform's Unicode tables via wctype.h. Code point index (use Match_start_byte/_end_byte for a byte offset into the raw buffer, Section 6 below)

Passing invalid UTF-8 to a UTF8-mode pattern's matching call returns -1 (the general error return, docs/API.md Section 3); BINARY and ASCII mode never reject a subject on this basis, since neither one decodes it.

Choosing a mode is a decision made once per Pattern (it is baked into flags at re_compile time), not per matching call: a Pattern compiled with UTF8 decodes every Input it is later run against as UTF-8, and a Pattern compiled with BINARY never does, regardless of what the actual bytes happen to look like.

3. Compiling a pattern

#include "regexx.h"
#include <string.h>

PatternError err; memset(&err, 0, sizeof err);
Pattern *pat = re_compile("(\\w+)@(\\w+)\\.(\\w+)", strlen("(\\w+)@(\\w+)\\.(\\w+)"), ASCII, &err);
if (!pat) {
    fprintf(stderr, "pattern error at %lld: %s\n", (long long)err.pos, err.msg);
    PatternError_free(&err);
    /* handle failure */
}
  • pattern/len: the pattern source; it need not be NUL-terminated, len is authoritative.
  • flags: the data mode (Section 2) combined with any of IGNORECASE, MULTILINE, DOTALL, VERBOSE, ASCII (also usable as a flag inside UTF8 mode to force ASCII-only \d/\w/\s), UNICODE, LOCALE, DEBUG (the last three accepted for source compatibility with Python, currently no-ops, docs/API.md Section 2).
  • err: optional (NULL is safe); filled in on a syntax error or an unsupported construct (docs/API.md Section 5 lists every construct this build rejects at compile time, such as conditional groups and scoped inline flags). Call PatternError_free on it once read, whether or not compilation succeeded.

A compiled Pattern is freed with Pattern_free, once, after every Match produced from it has been released (Match.re borrows from the Pattern, so freeing the Pattern first and a Match from it afterward is a use after free, docs/API.md Section 1.2).

4. Building an Input

Every matching call takes an Input, not a raw pointer, so the library has one place to resolve "where are the bytes" regardless of whether they came from memory the caller already has or a file on disk.

Input *Input_from_buffer(const uint8_t *buf, size_t len); /* wraps, does not copy or own */
Input *Input_from_file(const char *path, PatternError *err); /* mmaps or reads a whole file */
void   Input_free(Input *in);

Input_from_buffer wraps an existing buffer (a string literal, a malloc'd region, a memory-mapped region the caller manages itself) without copying it; the buffer must outlive the Input and every Match produced from it, since Match_group returns pointers directly into it (docs/API.md Section 1.3). Input_free on an Input_from_buffer result never frees the wrapped buffer.

Input_from_file maps a regular file read-only with mmap() instead of copying it, so the pages stay reclaimable under memory pressure (README.md "Memory footprint" has the measured numbers); a non-seekable source (a pipe, a FIFO, process substitution, /dev/stdin) or an mmap() failure falls back to reading it incrementally into an owned buffer instead. Either way Input_free releases what it acquired (munmap or free, as appropriate) automatically; the caller never needs to know which path was taken.

PatternError err; memset(&err, 0, sizeof err);
Input *in = Input_from_file("data.txt", &err);
if (!in) { fprintf(stderr, "cannot read file: %s\n", err.msg); PatternError_free(&err); }

Output, verified:

=== error_file ===
in=(nil) msg='No such file or directory'
int Pattern_match(Pattern *self, Input *string, int64_t pos, int64_t endpos, Match *out);
int Pattern_fullmatch(Pattern *self, Input *string, int64_t pos, int64_t endpos, Match *out);
int Pattern_search(Pattern *self, Input *string, int64_t pos, int64_t endpos, Match *out);

pos/endpos bound the attempt (pass 0/-1 for "the whole input", the same defaults re.Pattern.match/.search/etc. use); match and fullmatch are anchored at pos (fullmatch additionally requires reaching exactly endpos), search tries every start position from pos to endpos and reports the first that admits a match. All three return 1 matched, 0 no match, -1 error; out must point at a zero-initialized Match (never allocated by the library itself).

const char *pattern = "(\\w+)@(\\w+)\\.(\\w+)";
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
const char *subject = "contact: user@example.com today";
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));

Match m; memset(&m, 0, sizeof m);
if (Pattern_search(pat, in, 0, -1, &m) == 1) {
    const char *g; size_t glen;
    Match_group(&m, NULL, 0, &g, &glen);   /* whole match: group 0 */
    Match_group(&m, NULL, 1, &g, &glen);   /* first capturing group */
    int64_t s, e;
    Match_span(&m, 0, &s, &e);
    Match_free(&m);
}
Input_free(in);
Pattern_free(pat);

Output, verified:

=== ascii_basic ===
search rc=1
group0='user@example.com'
group1='user'
group2='example'
group3='com'
span0=[9,25)

Named groups

const char *pattern = "(?P<user>\\w+)@(?P<host>\\w+)";
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
const char *subject = "user@host";
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
Match m; memset(&m, 0, sizeof m);
if (Pattern_search(pat, in, 0, -1, &m) == 1) {
    const char *g; size_t glen;
    Match_group(&m, "user", 0, &g, &glen);  /* name_or_null non-NULL: index is ignored */
    Match_group(&m, "host", 0, &g, &glen);
    int idx = Pattern_groupindex_lookup(pat, "host");  /* 1-based group number, or -1 */
    Match_free(&m);
}

Output, verified:

=== ascii_named ===
user='user'
host='host'
groupindex(host)=2

Match_group returns 1 and a borrowed pointer into the Input's buffer when the group participated, 0 and *out = NULL when the group exists in the pattern but did not participate (Python's None), and -1 when no such group exists at all, by index or by name; this three-way distinction is the only place in the API that tells "no such group" apart from "this group matched nothing" (Match_start/Match_end collapse both cases to -1, docs/API.md Section 3.23-3.29).

Every name a Pattern defines can also be enumerated, without already knowing what to look for, via Pattern_groupindex_count/_at (Python: len(pattern.groupindex), dict(pattern.groupindex).items()):

int n = Pattern_groupindex_count(pat);
for (int i = 0; i < n; i++) {
    const char *name;
    int gnum = Pattern_groupindex_at(pat, i, &name);
    printf("  [%d] name=%s group=%d\n", i, name, gnum);
}

Output, verified, for the same pat above:

count=2
  [0] name=user group=1
  [1] name=host group=2

6. Code points versus bytes in UTF8 mode

In UTF8 mode, Match_start/Match_end/Match_span report a code point index (so, for example, a 2-byte UTF-8 character like é counts as one unit, the same way Python's str indexing does), while Match_start_byte/Match_end_byte/Match_span_byte always report a byte offset into the raw Input buffer, in every mode. Use the _byte variants whenever the result is used to slice or seek into the raw buffer or file directly (Match_group already does this translation internally, so it needs neither variant directly; a caller that needs to, for example, report a byte offset to an external tool that has no notion of code points is the usual reason to reach for _byte directly).

const char *pattern = "\\w+";
Pattern *pat = re_compile(pattern, strlen(pattern), UTF8, &err);
const char *subject = "caf\xc3\xa9 today"; /* "café today", é is 2 UTF-8 bytes */
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
Match m; memset(&m, 0, sizeof m);
if (Pattern_search(pat, in, 0, -1, &m) == 1) {
    int64_t cs, ce, bs, be;
    Match_span(&m, 0, &cs, &ce);          /* code point units */
    Match_span_byte(&m, 0, &bs, &be);     /* byte units */
    Match_free(&m);
}

Output, verified:

=== utf8_offsets ===
search rc=1
codepoint span=[0,4) byte span=[0,5)
group bytes: 'café' (5 bytes)

café is 4 code points (c, a, f, é) but 5 bytes (é alone is 2 bytes in UTF-8), which is exactly what the two spans above report.

Iterating UTF-8 text

const char *pattern = "\\w+";
Pattern *pat = re_compile(pattern, strlen(pattern), UTF8, &err);
const char *subject = "na\xc3\xafve caf\xc3\xa9 r\xc3\xa9sum\xc3\xa9"; /* naïve café résumé */
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
Pattern_finditer(pat, in, 0, -1, print_match_cb, &ctx);

Output, verified (\w under UTF8 mode matches letters outside ASCII too, consulting the platform's Unicode tables, docs/API.md Section 2):

=== utf8_finditer ===
  match 0: 'naïve'
  match 1: 'café'
  match 2: 'résumé'
total matches=3

7. Binary data

BINARY mode never decodes the subject; every byte 0-255 is a valid input value, including 0x00, and a ./character class/literal matches by raw byte value.

const char *pattern = "a.c";
Pattern *pat = re_compile(pattern, strlen(pattern), BINARY, &err);
uint8_t subject[] = { 'x', 'a', 0x00, 'c', 'y' };
Input *in = Input_from_buffer(subject, sizeof subject);
Match m; memset(&m, 0, sizeof m);
Pattern_search(pat, in, 0, -1, &m);

Output, verified (. matches the embedded NUL byte at offset 2, since DOTALL is irrelevant here: BINARY mode has no notion of "newline" ending .'s match, byte 0x0a is simply a byte like any other unless the pattern excludes it explicitly):

=== binary_basic ===
search rc=1
matched 3 bytes: 61 00 63

Byte ranges in a character class work the same way, by raw value:

const char *pattern = "[\\x00-\\x1f]+";
Pattern *pat = re_compile(pattern, strlen(pattern), BINARY, &err);
uint8_t subject[] = { 'A', 0x01, 0x02, 0x1f, 'B' };

Output, verified:

=== binary_class ===
search rc=1
span=[1,4)

\d/\w/\s in BINARY mode classify the same way they do in ASCII mode (ASCII letters/digits/whitespace only; there is no "Unicode binary" notion to fall back to, since BINARY mode has no decoding step at all).

8. Finding all matches: finditer/findall

int Pattern_finditer(Pattern *self, Input *string, int64_t pos, int64_t endpos, MatchIterCb cb, void *ctx);

cb(ctx, m) fires once per non-overlapping match, left to right; the Match passed in is only valid for the duration of the call (freed immediately after cb returns, so do not retain the pointer past it). Pattern_findall is the same call under a second name, for parity with re.findall; this build always hands the callback a full Match rather than collapsing it to "just the string" or a tuple the way CPython's findall does at the Python-object level, since there is no C object model to collapse into (docs/API.md Section 3.5-3.6).

typedef struct { int n; } IterCtx;
static void print_match_cb(void *ctx, const Match *m_const) {
    IterCtx *c = ctx;
    Match *m = (Match *)m_const;
    const char *g; size_t glen;
    Match_group(m, NULL, 0, &g, &glen);
    printf("  match %d: '%.*s'\n", c->n, (int)glen, g);
    c->n++;
}

const char *pattern = "\\d+";
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
const char *subject = "order 12 has 345 items, batch 6";
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
IterCtx ctx = { 0 };
int n = Pattern_finditer(pat, in, 0, -1, print_match_cb, &ctx);

Output, verified:

=== ascii_finditer ===
  match 0: '12'
  match 1: '345'
  match 2: '6'
total matches=3

finditer/findall/split/sub all share one memoization table and one set of precomputed skip-ahead tables across the whole scan (built once per call, not once per match), which is what makes them run in linear time for the common case rather than redoing quadratic work; see README.md "Implementation status" and "Memory footprint" for the measured numbers and what that costs in memory.

9. Splitting: Pattern_split

int Pattern_split(Pattern *self, Input *string, int maxsplit, MatchIterCb cb, void *ctx);

cb fires once per element of the list re.split() would return, in order; each element's text is read the same way as any other match, via Match_group(m, NULL, 0, &out, &outlen), and a 0 return from that call means the element is Python's None (an unparticipated capturing group between two matches, docs/API.md Section 3.7). maxsplit matches re.split's parameter (0 unlimited).

const char *pattern = "\\s*,\\s*";
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
const char *subject = "red, green,blue ,  yellow";
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
IterCtx ctx = { 0 };
int n = Pattern_split(pat, in, 0, print_split_cb, &ctx);

Output, verified:

=== ascii_split ===
  elem 0: 'red'
  elem 1: 'green'
  elem 2: 'blue'
  elem 3: 'yellow'
splits=3

10. Substitution: Pattern_sub/Pattern_subn

int Pattern_sub(Pattern *self, Input *string, const char *repl, MatchSubCb cb, void *ctx, int count, char **out, size_t *outlen);
int Pattern_subn(Pattern *self, Input *string, const char *repl, MatchSubCb cb, void *ctx, int count, char **out, size_t *outlen, int *n);

Exactly one of repl (a template string) or cb (a callback) is non-NULL; this two-parameter shape exists because C cannot express Python's single polymorphic repl argument (string or callable) in one slot (docs/API.md Section 3.8-3.9). Pattern_subn additionally reports the number of substitutions made through n; Pattern_sub is the same operation with that count discarded. *out is a freshly malloc'd, NUL-terminated buffer the caller must free.

Template substitution

Template syntax: \g<name>, \g<N>, \N (one or two digits), \n, \t, \\, and any other \X as the literal character X.

const char *pattern = "(\\w+)@(\\w+)";
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
const char *subject = "user@host";
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
char *out; size_t outlen; int n;
Pattern_subn(pat, in, "\\2@\\1", NULL, NULL, 0, &out, &outlen, &n);
/* use out/outlen, then: */
free(out);

Output, verified:

=== ascii_sub_template ===
result='host@user' n=1

Callback substitution

static void upper_cb(void *ctx, const Match *m_const, char **out, size_t *outlen) {
    Match *m = (Match *)m_const;
    const char *g; size_t glen;
    Match_group(m, NULL, 0, &g, &glen);
    char *buf = malloc(glen);
    for (size_t i = 0; i < glen; i++) {
        char c = g[i];
        buf[i] = (c >= 'a' && c <= 'z') ? (char)(c - 32) : c;
    }
    *out = buf; *outlen = glen;   /* ownership passes to Pattern_sub */
}

const char *pattern = "\\w+";
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
const char *subject = "shout this loudly";
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
char *out; size_t outlen;
Pattern_sub(pat, in, NULL, upper_cb, NULL, 0, &out, &outlen);
free(out);

Output, verified:

=== ascii_sub_callback ===
result='SHOUT THIS LOUDLY'

The callback is expected to malloc its replacement and hand ownership of it off through *out/*outlen; Pattern_sub/Pattern_subn frees it immediately after copying its content into the final result buffer (docs/API.md Section 1.5).

11. Flags

Pattern *pat = re_compile("hello", 5, ASCII | IGNORECASE, &err);
const char *pattern = "hello";
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII | IGNORECASE, &err);
const char *subject = "HELLO world";
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
Match m; memset(&m, 0, sizeof m);
Pattern_search(pat, in, 0, -1, &m);   /* rc=1 */
const char *pattern = "^line";
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII | MULTILINE, &err);
const char *subject = "first\nline two\nline three";
/* Pattern_finditer finds 2 matches: MULTILINE makes ^ match after every \n too */
const char *pattern = "a.b";
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII | DOTALL, &err);
const char *subject = "a\nb";
/* rc=1: DOTALL makes . match \n too */
const char *pattern = "\\d+  # a number\n\\s+ \\w+  # then a word";
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII | VERBOSE, &err);
const char *subject = "42 answer";
/* rc=1: VERBOSE strips the unescaped whitespace and # comments before parsing */

Output, verified for all four:

=== flags_example ===
IGNORECASE rc=1
  match 0: 'line'
  match 1: 'line'
MULTILINE matches=2
DOTALL rc=1
VERBOSE rc=1

UNICODE, LOCALE, and DEBUG are accepted (for source compatibility with Python) and combine with any of the above via |, but currently have no observable effect (docs/API.md Section 2 records exactly why for each one, in particular why LOCALE is a no-op rather than the bug it used to be, caught and fixed by this project's test suite).

12. Matching against a file

const char *pattern = "line \\w+";
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
Input *in = Input_from_file("data.txt", &err);
IterCtx ctx = { 0 };
int n = Pattern_finditer(pat, in, 0, -1, print_match_cb, &ctx);
Input_free(in);

Given a file containing line one\nline two has data\nline three\n, output, verified:

=== file_input ===
Input_from_file rc=0x5b30fbc44e20
  match 0: 'line one'
  match 1: 'line two'
  match 2: 'line three'
total matches=3

As Section 4 explains, Input_from_file mmaps a regular file rather than copying it; nothing about matching against it differs from matching against an Input_from_buffer-wrapped in-memory buffer, the mode (BINARY/ASCII/ UTF8) and every function above work identically regardless of which Input_from_* constructor produced the Input. See README.md "Memory footprint" for the measured memory cost of matching a large file this way, and for a documented, deliberately reverted attempt at bounding it that traded a fast allocation failure for an effectively unbounded hang (worth reading before assuming a size cap is a safe thing to add here).

13. Module level convenience functions

int re_search(const char *pattern, size_t len, int flags, Input *string, Match *out);

re_match/re_fullmatch/re_search/re_finditer/re_findall/re_split/ re_sub/re_subn compile pattern through an internal cache (up to 512 entries, keyed by (pattern, len, flags), cleared entirely on overflow) and then call the matching Pattern_ function with pos=0, endpos=-1, mirroring how CPython itself implements re.match as _compile(pattern, flags).match(string).

const char *pattern = "\\d+";
const char *subject = "abc123def";
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
Match m; memset(&m, 0, sizeof m);
int r = re_search(pattern, strlen(pattern), ASCII, in, &m);

Output, verified:

=== module_level ===
re_search rc=1
matched='123'

A compile error inside these functions is reported only as a -1 return, with no PatternError available; use re_compile plus a Pattern_ function directly whenever a compile error needs to be diagnosed, or whenever the same pattern is used more than once (precompiling once and reusing the Pattern avoids repeated cache lookups, matching idiomatic Python's own preference for re.compile in a loop over calling the module level function repeatedly). re_purge() clears the cache, matching re.purge(). This cache is shared, mutable, process-wide state with no internal locking; do not call a re_-prefixed function from more than one thread without external synchronization (Pattern_-prefixed functions on an already-compiled Pattern have no such restriction, since nothing here mutates a compiled Pattern).

14. Escaping literal text: re_escape

void re_escape(const char *in, size_t len, char **out, size_t *outlen);
const char *s = "3.14 (pi)";
char *out; size_t outlen;
re_escape(s, strlen(s), &out, &outlen);
free(out);

Output, verified:

=== escape_example ===
escaped='3\.14\ \(pi\)'

Matches re.escape exactly, including the narrowed escaped-character set CPython adopted in 3.7 (only characters that are actually special in a regex are escaped; other non-alphanumeric bytes are passed through unescaped).

15. Error handling

Two kinds of failure exist in this API: a compile-time PatternError (re_compile, Input_from_file) and a matching-time -1 return with no further detail (docs/API.md Section 3). Always check the return value of a compile call before using the Pattern it was supposed to produce.

const char *pattern = "(unclosed";
PatternError err; memset(&err, 0, sizeof err);
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
if (!pat) {
    fprintf(stderr, "pattern error at %lld: %s\n", (long long)err.pos, err.msg);
}
PatternError_free(&err);

Output, verified:

=== error_compile ===
pat=(nil) msg='missing ), unterminated subpattern' pos=9 lineno=1 colno=10

err.msg/err.pattern are heap allocated by whichever call filled the struct in; zero-initialize a PatternError before use and call PatternError_free on it once done reading it, whether or not the call that filled it in succeeded. A NULL err argument is always safe to pass to any function that accepts one, and simply skips error reporting.

A matching call (Pattern_search, and so on) returning -1 means either invalid UTF-8 in the subject against a UTF8-mode Pattern, or the backtracking depth limit was reached (MAX_DEPTH in regexx.c, currently 60000, not adjustable without editing regexx.c and rebuilding); no further detail is available through the return value alone in this build.

16. Pattern syntax reference

Literals; . (with DOTALL for it to match \n); character classes with ranges and negation ([abc], [^abc], [a-z]); \d \D \w \W \s \S; \b \B; anchors ^ $ \A \Z (^/$ also match at line boundaries with MULTILINE); quantifiers * + ? {m,n} {m,} {,n} {m}, both greedy and lazy (*?, +?, and so on) forms; possessive quantifiers *+ ++ ?+ {m,n}+; groups (...), non-capturing (?:...), named (?P<name>...); alternation |; backreferences \1-\99, (?P=name), \g<name>, \g<N>; lookahead (?=...)/(?!...); fixed-width lookbehind (?<=...)/(?<!...); atomic groups (?>...); comments (?#...); global inline flags (?aiLmsux) at the very start of a pattern; escapes \n \r \t \f \v \a, \0-prefixed octal escapes, \xhh, \uxxxx, \Uxxxxxxxx.

Rejected at compile time with a PatternError naming the construct, rather than silently mis-parsed: conditional groups (?(id)yes|no), scoped inline flags (?flags:...) (only the global, start-of-pattern form is supported), \N{NAME} named code points, variable-width lookbehind (matching CPython's own restriction), and POSIX bracket-expression syntax [[:alpha:]] (not part of Python re's grammar at all, concept.md Section 14). See docs/API.md Section 5-6 for the complete, exact list.

17. Performance and memory, in one paragraph

A pattern with no backreference, lookaround, or atomic group runs on the Pike VM (README.md "The Pike VM"), which is genuinely linear time, including on adversarial nested-quantifier shapes like (a+)+b that would invite catastrophic backtracking elsewhere, with no atomic group needed; it has no per-thread allocation (a flat, PC-indexed table instead, concept.md 7.10) and a real Boyer-Moore-Horspool literal prefilter, and a pattern beginning with a literal string now typically runs faster than POSIX <regex.h> itself, not merely faster than before. A pattern that cannot be prefiltered at all (a nullable leading loop like a*) still runs somewhat slower in absolute terms than the backtracking engine would on the same pattern, a real, documented, and now much narrower trade-off (no lazy DFA state caching exists, a choice concept.md 7.10 explains rather than an unexamined gap). Everything else (a backreference, lookaround, or atomic group anywhere) runs on the recursive backtracking engine instead: match/fullmatch (a single anchored attempt) cost time proportional to that attempt; search/ finditer/split/sub are linear time, not quadratic, for the common case there (no backreference, every quantifier's body a single character/class/.), at a measured memory cost of roughly 8.4x the input length for a pattern using a simple-atom quantifier, and a pattern with a backreference can still be worse than linear, the same way CPython's own re can be on the same patterns. README.md "Implementation status", "The Pike VM", and "Memory footprint" have the full measured account, including one fix that was tried, measured, and deliberately reverted; read it before assuming a given pattern's cost, rather than assuming linear time and bounded memory apply unconditionally.

18. Complete example: rxgrep

examples/rxgrep.c is a complete, working grep-like program built on this API, exercising all three data modes and both the matching and substitution API against real files and pipes:

./rxgrep -in 'hello' file.txt        # case-insensitive, line numbers
./rxgrep -m utf8 -o '\w+' file.txt   # print every UTF-8 word, one per line
./rxgrep -c 'error' log.txt          # count matching lines
./rxgrep -m binary 'a.c' data.bin    # match raw bytes, embedded NUL included
./rxgrep -m utf8 --sub 'REDACTED' '\d{3}-\d{4}' file.txt

Reading its source alongside this document is a reasonable next step once the examples above are familiar: it shows every function here used together in one program, including the parts this document simplified away for clarity (option parsing, line splitting, output formatting).

examples/ also has six smaller, single-purpose programs, each isolating one distinct feature this section-by-section walkthrough only touches briefly: binary_scan.c (raw byte-range classes including embedded NUL), utf8_scripts.c (\w across non-Latin scripts, code point versus byte offsets), ascii_logparse.c (named groups against structured log text), redos_atomic.c (the textbook (a+)+b ReDoS shape, now fixed automatically by the Pike VM with no atomic group needed, contrasted with a backreference-forced variant where an atomic group is still necessary), empty_match_rule.c (the undocumented CPython empty-match retry rule, verified against a real interpreter), large_file_search.c (mmap-backed file input at 100MB, with elapsed time and peak memory printed), and bench_vs_posix.c (a direct, honestly-reported timing comparison against the C standard library's own <regex.h> on six scenarios: glibc wins the three ordinary ones by a wide margin, and the one adversarial pattern shape inverts entirely, now beating glibc outright with no atomic group). See examples/README.md for the complete list with what each one demonstrates; make examples builds all of them.