Add USAGE.md, a task-oriented guide with compiled and verified examples

docs/API.md is the exhaustive reference; concept.md is the design
rationale. Neither is the right document for "I have never used this
library before, show me working code," so this adds one that is:
compiling a pattern, building an Input from a buffer or a file,
match/fullmatch/search, named groups, finditer/findall, split,
sub/subn with both template and callback replacement, all six flags,
error handling, the module level convenience functions and their
cache, re_escape, and a pattern syntax quick reference, each with
BINARY, ASCII, or UTF8 examples as appropriate (including the
code-point-versus-byte-offset distinction UTF8 mode introduces).

Every example was written as a real, compiled program linked against
the current regexx.c and run; the output shown in the document is
that program's actual output, not a hand-written guess, the same
verification standard already used for README.md's own measured
numbers.

README.md gets a one-line pointer to it alongside the existing
pointer to docs/API.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
This commit is contained in:
2026-09-14 10:33:51 +00:00
co-authored by Claude Sonnet 5
parent b24d63e5fd
commit 71e8bdaaa2
2 changed files with 724 additions and 1 deletions
+3 -1
View File
@@ -13,7 +13,9 @@ what is and is not carried over from Python's `re` and from POSIX's native
implementation that exists today at a glance; [`docs/API.md`](docs/API.md) is
the exhaustive reference (every type, every flag, every function's exact
return-value and memory-ownership convention, checked against a real CPython
interpreter, not against memory).
interpreter, not against memory); [`USAGE.md`](USAGE.md) is the task-oriented
guide (compiled, run, and verified examples covering every operation across
all three data modes, from "nothing" to "working code").
## Implementation status
+721
View File
@@ -0,0 +1,721 @@
# regexx usage guide
This document is task oriented: how to build a pattern, run it against data
in each of the three supported modes, and read the result back out, with
complete, compiled, and run examples. [`docs/API.md`](docs/API.md) is the
exhaustive reference for every function's exact return value and memory
ownership rule; this document exists to get from "nothing" to "working code"
without reading that whole reference first. [`concept.md`](concept.md)
records the design rationale for anyone who wants to know why a given
choice was made rather than only what it is.
Every example below was compiled against the current `regexx.c` and its
actual printed output is shown, not a hand-written guess at what it would
print.
## 1. Building and linking
```sh
make # builds libregexx.a and the rxgrep example
```
A program using the library needs `regexx.h` on its include path and
`libregexx.a` (or `regexx.c` compiled directly into the program, since the
library is a single file) linked in:
```sh
cc -std=c11 -I/path/to/regexx -o myprog myprog.c /path/to/regexx/libregexx.a
```
No third-party dependency is required; `regexx.c` uses only the C standard
library plus, for `UTF8` mode's Unicode classification and `Input_from_file`,
POSIX (`wctype.h`, `locale.h`, `mmap`/`open`/`fstat`, `concept.md` Section 8).
## 2. The three data modes
Exactly one of `BINARY`, `ASCII`, or `UTF8` is passed in `flags` to
`re_compile` (or any `re_`-prefixed module level function); it selects how
the subject's bytes are interpreted, not how the pattern text itself is
encoded (the pattern is always read as ASCII/UTF-8 source text regardless of
this choice).
| Mode | Subject interpretation | `Match_start`/`Match_end` unit |
|---|---|---|
| `BINARY` | Raw bytes, `0`-`255` all valid, embedded `NUL` is an ordinary byte, no decoding at all. | Byte offset |
| `ASCII` | Bytes, but `\d`/`\w`/`\s` and `IGNORECASE` restrict themselves to the ASCII subset (the default when neither `BINARY` nor `UTF8` is given, `docs/API.md` Section 2). | Byte offset |
| `UTF8` | Decoded as UTF-8 into Unicode code points before matching; `\d`/`\w`/`\s` and case folding consult the platform's Unicode tables via `wctype.h`. | Code point index (use `Match_start_byte`/`_end_byte` for a byte offset into the raw buffer, Section 6 below) |
Passing invalid UTF-8 to a `UTF8`-mode pattern's matching call returns `-1`
(the general error return, `docs/API.md` Section 3); `BINARY` and `ASCII`
mode never reject a subject on this basis, since neither one decodes it.
Choosing a mode is a decision made once per `Pattern` (it is baked into
`flags` at `re_compile` time), not per matching call: a `Pattern` compiled
with `UTF8` decodes every `Input` it is later run against as UTF-8, and a
`Pattern` compiled with `BINARY` never does, regardless of what the actual
bytes happen to look like.
## 3. Compiling a pattern
```c
#include "regexx.h"
#include <string.h>
PatternError err; memset(&err, 0, sizeof err);
Pattern *pat = re_compile("(\\w+)@(\\w+)\\.(\\w+)", strlen("(\\w+)@(\\w+)\\.(\\w+)"), ASCII, &err);
if (!pat) {
fprintf(stderr, "pattern error at %lld: %s\n", (long long)err.pos, err.msg);
PatternError_free(&err);
/* handle failure */
}
```
- `pattern`/`len`: the pattern source; it need not be `NUL`-terminated, `len`
is authoritative.
- `flags`: the data mode (Section 2) combined with any of `IGNORECASE`,
`MULTILINE`, `DOTALL`, `VERBOSE`, `ASCII` (also usable as a flag inside
`UTF8` mode to force ASCII-only `\d`/`\w`/`\s`), `UNICODE`, `LOCALE`,
`DEBUG` (the last three accepted for source compatibility with Python,
currently no-ops, `docs/API.md` Section 2).
- `err`: optional (`NULL` is safe); filled in on a syntax error or an
unsupported construct (`docs/API.md` Section 5 lists every construct this
build rejects at compile time, such as conditional groups and scoped
inline flags). Call `PatternError_free` on it once read, whether or not
compilation succeeded.
A compiled `Pattern` is freed with `Pattern_free`, once, after every `Match`
produced from it has been released (`Match.re` borrows from the `Pattern`,
so freeing the `Pattern` first and a `Match` from it afterward is a use
after free, `docs/API.md` Section 1.2).
## 4. Building an `Input`
Every matching call takes an `Input`, not a raw pointer, so the library has
one place to resolve "where are the bytes" regardless of whether they came
from memory the caller already has or a file on disk.
```c
Input *Input_from_buffer(const uint8_t *buf, size_t len); /* wraps, does not copy or own */
Input *Input_from_file(const char *path, PatternError *err); /* mmaps or reads a whole file */
void Input_free(Input *in);
```
`Input_from_buffer` wraps an existing buffer (a string literal, a `malloc`'d
region, a memory-mapped region the caller manages itself) without copying
it; the buffer must outlive the `Input` and every `Match` produced from it,
since `Match_group` returns pointers directly into it (`docs/API.md` Section
1.3). `Input_free` on an `Input_from_buffer` result never frees the wrapped
buffer.
`Input_from_file` maps a regular file read-only with `mmap()` instead of
copying it, so the pages stay reclaimable under memory pressure
(`README.md` "Memory footprint" has the measured numbers); a non-seekable
source (a pipe, a FIFO, process substitution, `/dev/stdin`) or an `mmap()`
failure falls back to reading it incrementally into an owned buffer instead.
Either way `Input_free` releases what it acquired (`munmap` or `free`, as
appropriate) automatically; the caller never needs to know which path was
taken.
```c
PatternError err; memset(&err, 0, sizeof err);
Input *in = Input_from_file("data.txt", &err);
if (!in) { fprintf(stderr, "cannot read file: %s\n", err.msg); PatternError_free(&err); }
```
Output, verified:
```
=== error_file ===
in=(nil) msg='No such file or directory'
```
## 5. Matching: `match`, `fullmatch`, `search`
```c
int Pattern_match(Pattern *self, Input *string, int64_t pos, int64_t endpos, Match *out);
int Pattern_fullmatch(Pattern *self, Input *string, int64_t pos, int64_t endpos, Match *out);
int Pattern_search(Pattern *self, Input *string, int64_t pos, int64_t endpos, Match *out);
```
`pos`/`endpos` bound the attempt (pass `0`/`-1` for "the whole input", the
same defaults `re.Pattern.match`/`.search`/etc. use); `match` and
`fullmatch` are anchored at `pos` (`fullmatch` additionally requires
reaching exactly `endpos`), `search` tries every start position from `pos`
to `endpos` and reports the first that admits a match. All three return `1`
matched, `0` no match, `-1` error; `out` must point at a zero-initialized
`Match` (never allocated by the library itself).
```c
const char *pattern = "(\\w+)@(\\w+)\\.(\\w+)";
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
const char *subject = "contact: user@example.com today";
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
Match m; memset(&m, 0, sizeof m);
if (Pattern_search(pat, in, 0, -1, &m) == 1) {
const char *g; size_t glen;
Match_group(&m, NULL, 0, &g, &glen); /* whole match: group 0 */
Match_group(&m, NULL, 1, &g, &glen); /* first capturing group */
int64_t s, e;
Match_span(&m, 0, &s, &e);
Match_free(&m);
}
Input_free(in);
Pattern_free(pat);
```
Output, verified:
```
=== ascii_basic ===
search rc=1
group0='user@example.com'
group1='user'
group2='example'
group3='com'
span0=[9,25)
```
### Named groups
```c
const char *pattern = "(?P<user>\\w+)@(?P<host>\\w+)";
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
const char *subject = "user@host";
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
Match m; memset(&m, 0, sizeof m);
if (Pattern_search(pat, in, 0, -1, &m) == 1) {
const char *g; size_t glen;
Match_group(&m, "user", 0, &g, &glen); /* name_or_null non-NULL: index is ignored */
Match_group(&m, "host", 0, &g, &glen);
int idx = Pattern_groupindex_lookup(pat, "host"); /* 1-based group number, or -1 */
Match_free(&m);
}
```
Output, verified:
```
=== ascii_named ===
user='user'
host='host'
groupindex(host)=2
```
`Match_group` returns `1` and a borrowed pointer into the `Input`'s buffer
when the group participated, `0` and `*out = NULL` when the group exists in
the pattern but did not participate (Python's `None`), and `-1` when no
such group exists at all, by index or by name; this three-way distinction
is the only place in the API that tells "no such group" apart from "this
group matched nothing" (`Match_start`/`Match_end` collapse both cases to
`-1`, `docs/API.md` Section 3.23-3.29).
## 6. Code points versus bytes in `UTF8` mode
In `UTF8` mode, `Match_start`/`Match_end`/`Match_span` report a code point
index (so, for example, a 2-byte UTF-8 character like `é` counts as one
unit, the same way Python's `str` indexing does), while
`Match_start_byte`/`Match_end_byte`/`Match_span_byte` always report a byte
offset into the raw `Input` buffer, in every mode. Use the `_byte` variants
whenever the result is used to slice or seek into the raw buffer or file
directly (`Match_group` already does this translation internally, so it
needs neither variant directly; a caller that needs to, for example, report
a byte offset to an external tool that has no notion of code points is the
usual reason to reach for `_byte` directly).
```c
const char *pattern = "\\w+";
Pattern *pat = re_compile(pattern, strlen(pattern), UTF8, &err);
const char *subject = "caf\xc3\xa9 today"; /* "café today", é is 2 UTF-8 bytes */
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
Match m; memset(&m, 0, sizeof m);
if (Pattern_search(pat, in, 0, -1, &m) == 1) {
int64_t cs, ce, bs, be;
Match_span(&m, 0, &cs, &ce); /* code point units */
Match_span_byte(&m, 0, &bs, &be); /* byte units */
Match_free(&m);
}
```
Output, verified:
```
=== utf8_offsets ===
search rc=1
codepoint span=[0,4) byte span=[0,5)
group bytes: 'café' (5 bytes)
```
`café` is 4 code points (`c`, `a`, `f`, `é`) but 5 bytes (`é` alone is 2
bytes in UTF-8), which is exactly what the two spans above report.
### Iterating UTF-8 text
```c
const char *pattern = "\\w+";
Pattern *pat = re_compile(pattern, strlen(pattern), UTF8, &err);
const char *subject = "na\xc3\xafve caf\xc3\xa9 r\xc3\xa9sum\xc3\xa9"; /* naïve café résumé */
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
Pattern_finditer(pat, in, 0, -1, print_match_cb, &ctx);
```
Output, verified (`\w` under `UTF8` mode matches letters outside ASCII too,
consulting the platform's Unicode tables, `docs/API.md` Section 2):
```
=== utf8_finditer ===
match 0: 'naïve'
match 1: 'café'
match 2: 'résumé'
total matches=3
```
## 7. Binary data
`BINARY` mode never decodes the subject; every byte `0`-`255` is a valid
input value, including `0x00`, and a `.`/character class/literal matches by
raw byte value.
```c
const char *pattern = "a.c";
Pattern *pat = re_compile(pattern, strlen(pattern), BINARY, &err);
uint8_t subject[] = { 'x', 'a', 0x00, 'c', 'y' };
Input *in = Input_from_buffer(subject, sizeof subject);
Match m; memset(&m, 0, sizeof m);
Pattern_search(pat, in, 0, -1, &m);
```
Output, verified (`.` matches the embedded `NUL` byte at offset 2, since
`DOTALL` is irrelevant here: `BINARY` mode has no notion of "newline" ending
`.`'s match, byte `0x0a` is simply a byte like any other unless the pattern
excludes it explicitly):
```
=== binary_basic ===
search rc=1
matched 3 bytes: 61 00 63
```
Byte ranges in a character class work the same way, by raw value:
```c
const char *pattern = "[\\x00-\\x1f]+";
Pattern *pat = re_compile(pattern, strlen(pattern), BINARY, &err);
uint8_t subject[] = { 'A', 0x01, 0x02, 0x1f, 'B' };
```
Output, verified:
```
=== binary_class ===
search rc=1
span=[1,4)
```
`\d`/`\w`/`\s` in `BINARY` mode classify the same way they do in `ASCII`
mode (ASCII letters/digits/whitespace only; there is no "Unicode binary"
notion to fall back to, since `BINARY` mode has no decoding step at all).
## 8. Finding all matches: `finditer`/`findall`
```c
int Pattern_finditer(Pattern *self, Input *string, int64_t pos, int64_t endpos, MatchIterCb cb, void *ctx);
```
`cb(ctx, m)` fires once per non-overlapping match, left to right; the
`Match` passed in is only valid for the duration of the call (freed
immediately after `cb` returns, so do not retain the pointer past it).
`Pattern_findall` is the same call under a second name, for parity with
`re.findall`; this build always hands the callback a full `Match` rather
than collapsing it to "just the string" or a tuple the way CPython's
`findall` does at the Python-object level, since there is no C object model
to collapse into (`docs/API.md` Section 3.5-3.6).
```c
typedef struct { int n; } IterCtx;
static void print_match_cb(void *ctx, const Match *m_const) {
IterCtx *c = ctx;
Match *m = (Match *)m_const;
const char *g; size_t glen;
Match_group(m, NULL, 0, &g, &glen);
printf(" match %d: '%.*s'\n", c->n, (int)glen, g);
c->n++;
}
const char *pattern = "\\d+";
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
const char *subject = "order 12 has 345 items, batch 6";
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
IterCtx ctx = { 0 };
int n = Pattern_finditer(pat, in, 0, -1, print_match_cb, &ctx);
```
Output, verified:
```
=== ascii_finditer ===
match 0: '12'
match 1: '345'
match 2: '6'
total matches=3
```
`finditer`/`findall`/`split`/`sub` all share one memoization table and one
set of precomputed skip-ahead tables across the whole scan (built once per
call, not once per match), which is what makes them run in linear time for
the common case rather than redoing quadratic work; see README.md
"Implementation status" and "Memory footprint" for the measured numbers and
what that costs in memory.
## 9. Splitting: `Pattern_split`
```c
int Pattern_split(Pattern *self, Input *string, int maxsplit, MatchIterCb cb, void *ctx);
```
`cb` fires once per element of the list `re.split()` would return, in
order; each element's text is read the same way as any other match, via
`Match_group(m, NULL, 0, &out, &outlen)`, and a `0` return from that call
means the element is Python's `None` (an unparticipated capturing group
between two matches, `docs/API.md` Section 3.7). `maxsplit` matches
`re.split`'s parameter (`0` unlimited).
```c
const char *pattern = "\\s*,\\s*";
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
const char *subject = "red, green,blue , yellow";
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
IterCtx ctx = { 0 };
int n = Pattern_split(pat, in, 0, print_split_cb, &ctx);
```
Output, verified:
```
=== ascii_split ===
elem 0: 'red'
elem 1: 'green'
elem 2: 'blue'
elem 3: 'yellow'
splits=3
```
## 10. Substitution: `Pattern_sub`/`Pattern_subn`
```c
int Pattern_sub(Pattern *self, Input *string, const char *repl, MatchSubCb cb, void *ctx, int count, char **out, size_t *outlen);
int Pattern_subn(Pattern *self, Input *string, const char *repl, MatchSubCb cb, void *ctx, int count, char **out, size_t *outlen, int *n);
```
Exactly one of `repl` (a template string) or `cb` (a callback) is
non-`NULL`; this two-parameter shape exists because C cannot express
Python's single polymorphic `repl` argument (string or callable) in one
slot (`docs/API.md` Section 3.8-3.9). `Pattern_subn` additionally reports
the number of substitutions made through `n`; `Pattern_sub` is the same
operation with that count discarded. `*out` is a freshly `malloc`'d,
`NUL`-terminated buffer the caller must `free`.
### Template substitution
Template syntax: `\g<name>`, `\g<N>`, `\N` (one or two digits), `\n`, `\t`,
`\\`, and any other `\X` as the literal character `X`.
```c
const char *pattern = "(\\w+)@(\\w+)";
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
const char *subject = "user@host";
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
char *out; size_t outlen; int n;
Pattern_subn(pat, in, "\\2@\\1", NULL, NULL, 0, &out, &outlen, &n);
/* use out/outlen, then: */
free(out);
```
Output, verified:
```
=== ascii_sub_template ===
result='host@user' n=1
```
### Callback substitution
```c
static void upper_cb(void *ctx, const Match *m_const, char **out, size_t *outlen) {
Match *m = (Match *)m_const;
const char *g; size_t glen;
Match_group(m, NULL, 0, &g, &glen);
char *buf = malloc(glen);
for (size_t i = 0; i < glen; i++) {
char c = g[i];
buf[i] = (c >= 'a' && c <= 'z') ? (char)(c - 32) : c;
}
*out = buf; *outlen = glen; /* ownership passes to Pattern_sub */
}
const char *pattern = "\\w+";
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
const char *subject = "shout this loudly";
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
char *out; size_t outlen;
Pattern_sub(pat, in, NULL, upper_cb, NULL, 0, &out, &outlen);
free(out);
```
Output, verified:
```
=== ascii_sub_callback ===
result='SHOUT THIS LOUDLY'
```
The callback is expected to `malloc` its replacement and hand ownership of
it off through `*out`/`*outlen`; `Pattern_sub`/`Pattern_subn` frees it
immediately after copying its content into the final result buffer
(`docs/API.md` Section 1.5).
## 11. Flags
```c
Pattern *pat = re_compile("hello", 5, ASCII | IGNORECASE, &err);
```
```c
const char *pattern = "hello";
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII | IGNORECASE, &err);
const char *subject = "HELLO world";
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
Match m; memset(&m, 0, sizeof m);
Pattern_search(pat, in, 0, -1, &m); /* rc=1 */
```
```c
const char *pattern = "^line";
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII | MULTILINE, &err);
const char *subject = "first\nline two\nline three";
/* Pattern_finditer finds 2 matches: MULTILINE makes ^ match after every \n too */
```
```c
const char *pattern = "a.b";
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII | DOTALL, &err);
const char *subject = "a\nb";
/* rc=1: DOTALL makes . match \n too */
```
```c
const char *pattern = "\\d+ # a number\n\\s+ \\w+ # then a word";
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII | VERBOSE, &err);
const char *subject = "42 answer";
/* rc=1: VERBOSE strips the unescaped whitespace and # comments before parsing */
```
Output, verified for all four:
```
=== flags_example ===
IGNORECASE rc=1
match 0: 'line'
match 1: 'line'
MULTILINE matches=2
DOTALL rc=1
VERBOSE rc=1
```
`UNICODE`, `LOCALE`, and `DEBUG` are accepted (for source compatibility with
Python) and combine with any of the above via `|`, but currently have no
observable effect (`docs/API.md` Section 2 records exactly why for each
one, in particular why `LOCALE` is a no-op rather than the bug it used to
be, caught and fixed by this project's test suite).
## 12. Matching against a file
```c
const char *pattern = "line \\w+";
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
Input *in = Input_from_file("data.txt", &err);
IterCtx ctx = { 0 };
int n = Pattern_finditer(pat, in, 0, -1, print_match_cb, &ctx);
Input_free(in);
```
Given a file containing `line one\nline two has data\nline three\n`, output,
verified:
```
=== file_input ===
Input_from_file rc=0x5b30fbc44e20
match 0: 'line one'
match 1: 'line two'
match 2: 'line three'
total matches=3
```
As Section 4 explains, `Input_from_file` mmaps a regular file rather than
copying it; nothing about matching against it differs from matching against
an `Input_from_buffer`-wrapped in-memory buffer, the mode (`BINARY`/`ASCII`/
`UTF8`) and every function above work identically regardless of which
`Input_from_*` constructor produced the `Input`. See README.md "Memory
footprint" for the measured memory cost of matching a large file this way,
and for a documented, deliberately reverted attempt at bounding it that
traded a fast allocation failure for an effectively unbounded hang (worth
reading before assuming a size cap is a safe thing to add here).
## 13. Module level convenience functions
```c
int re_search(const char *pattern, size_t len, int flags, Input *string, Match *out);
```
`re_match`/`re_fullmatch`/`re_search`/`re_finditer`/`re_findall`/`re_split`/
`re_sub`/`re_subn` compile `pattern` through an internal cache (up to 512
entries, keyed by `(pattern, len, flags)`, cleared entirely on overflow)
and then call the matching `Pattern_` function with `pos=0`, `endpos=-1`,
mirroring how CPython itself implements `re.match` as
`_compile(pattern, flags).match(string)`.
```c
const char *pattern = "\\d+";
const char *subject = "abc123def";
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
Match m; memset(&m, 0, sizeof m);
int r = re_search(pattern, strlen(pattern), ASCII, in, &m);
```
Output, verified:
```
=== module_level ===
re_search rc=1
matched='123'
```
A compile error inside these functions is reported only as a `-1` return,
with no `PatternError` available; use `re_compile` plus a `Pattern_`
function directly whenever a compile error needs to be diagnosed, or
whenever the same pattern is used more than once (precompiling once and
reusing the `Pattern` avoids repeated cache lookups, matching idiomatic
Python's own preference for `re.compile` in a loop over calling the module
level function repeatedly). `re_purge()` clears the cache, matching
`re.purge()`. This cache is shared, mutable, process-wide state with no
internal locking; do not call a `re_`-prefixed function from more than one
thread without external synchronization (`Pattern_`-prefixed functions on
an already-compiled `Pattern` have no such restriction, since nothing here
mutates a compiled `Pattern`).
## 14. Escaping literal text: `re_escape`
```c
void re_escape(const char *in, size_t len, char **out, size_t *outlen);
```
```c
const char *s = "3.14 (pi)";
char *out; size_t outlen;
re_escape(s, strlen(s), &out, &outlen);
free(out);
```
Output, verified:
```
=== escape_example ===
escaped='3\.14\ \(pi\)'
```
Matches `re.escape` exactly, including the narrowed escaped-character set
CPython adopted in 3.7 (only characters that are actually special in a
regex are escaped; other non-alphanumeric bytes are passed through
unescaped).
## 15. Error handling
Two kinds of failure exist in this API: a compile-time `PatternError`
(`re_compile`, `Input_from_file`) and a matching-time `-1` return with no
further detail (`docs/API.md` Section 3). Always check the return value of
a compile call before using the `Pattern` it was supposed to produce.
```c
const char *pattern = "(unclosed";
PatternError err; memset(&err, 0, sizeof err);
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
if (!pat) {
fprintf(stderr, "pattern error at %lld: %s\n", (long long)err.pos, err.msg);
}
PatternError_free(&err);
```
Output, verified:
```
=== error_compile ===
pat=(nil) msg='missing ), unterminated subpattern' pos=9 lineno=1 colno=10
```
`err.msg`/`err.pattern` are heap allocated by whichever call filled the
struct in; zero-initialize a `PatternError` before use and call
`PatternError_free` on it once done reading it, whether or not the call
that filled it in succeeded. A `NULL` `err` argument is always safe to pass
to any function that accepts one, and simply skips error reporting.
A matching call (`Pattern_search`, and so on) returning `-1` means either
invalid UTF-8 in the subject against a `UTF8`-mode `Pattern`, or the
backtracking depth limit was reached (`MAX_DEPTH` in `regexx.c`, currently
60000, not adjustable without editing `regexx.c` and rebuilding); no
further detail is available through the return value alone in this build.
## 16. Pattern syntax reference
Literals; `.` (with `DOTALL` for it to match `\n`); character classes with
ranges and negation (`[abc]`, `[^abc]`, `[a-z]`); `\d \D \w \W \s \S`;
`\b \B`; anchors `^ $ \A \Z` (`^`/`$` also match at line boundaries with
`MULTILINE`); quantifiers `* + ? {m,n} {m,} {,n} {m}`, both greedy and lazy
(`*?`, `+?`, and so on) forms; possessive quantifiers `*+ ++ ?+ {m,n}+`;
groups `(...)`, non-capturing `(?:...)`, named `(?P<name>...)`; alternation
`|`; backreferences `\1`-`\99`, `(?P=name)`, `\g<name>`, `\g<N>`; lookahead
`(?=...)`/`(?!...)`; fixed-width lookbehind `(?<=...)`/`(?<!...)`; atomic
groups `(?>...)`; comments `(?#...)`; global inline flags `(?aiLmsux)` at
the very start of a pattern; escapes `\n \r \t \f \v \a`, `\0`-prefixed
octal escapes, `\xhh`, `\uxxxx`, `\Uxxxxxxxx`.
Rejected at compile time with a `PatternError` naming the construct, rather
than silently mis-parsed: conditional groups `(?(id)yes|no)`, scoped inline
flags `(?flags:...)` (only the global, start-of-pattern form is supported),
`\N{NAME}` named code points, variable-width lookbehind (matching CPython's
own restriction), and POSIX bracket-expression syntax `[[:alpha:]]` (not
part of Python `re`'s grammar at all, `concept.md` Section 14). See
`docs/API.md` Section 5-6 for the complete, exact list.
## 17. Performance and memory, in one paragraph
`match`/`fullmatch` (a single anchored attempt) cost time proportional to
that attempt. `search`/`finditer`/`split`/`sub` are linear time, not
quadratic, for the common case (no backreference, every quantifier's body a
single character/class/`.`), at a measured memory cost of roughly 8.4x the
input length for a pattern using a simple-atom quantifier. A pattern with a
backreference, or shaped like nested overlapping quantifiers
(`(a+)+b`-style), can still be worse than linear, the same way CPython's own
`re` can be on the same patterns. `README.md` "Implementation status" and
"Memory footprint" have the full measured account, including one fix that
was tried, measured, and deliberately reverted; read it before assuming a
given pattern's cost, rather than assuming linear time and bounded memory
apply unconditionally.
## 18. Complete example: `rxgrep`
`examples/rxgrep.c` is a complete, working grep-like program built on this
API, exercising all three data modes and both the matching and substitution
API against real files and pipes:
```sh
./rxgrep -in 'hello' file.txt # case-insensitive, line numbers
./rxgrep -m utf8 -o '\w+' file.txt # print every UTF-8 word, one per line
./rxgrep -c 'error' log.txt # count matching lines
./rxgrep -m binary 'a.c' data.bin # match raw bytes, embedded NUL included
./rxgrep -m utf8 --sub 'REDACTED' '\d{3}-\d{4}' file.txt
```
Reading its source alongside this document is a reasonable next step once
the examples above are familiar: it shows every function here used together
in one program, including the parts this document simplified away for
clarity (option parsing, line splitting, output formatting).