diff --git a/README.md b/README.md index 60b3b00..1aec56d 100644 --- a/README.md +++ b/README.md @@ -13,7 +13,9 @@ what is and is not carried over from Python's `re` and from POSIX's native implementation that exists today at a glance; [`docs/API.md`](docs/API.md) is the exhaustive reference (every type, every flag, every function's exact return-value and memory-ownership convention, checked against a real CPython -interpreter, not against memory). +interpreter, not against memory); [`USAGE.md`](USAGE.md) is the task-oriented +guide (compiled, run, and verified examples covering every operation across +all three data modes, from "nothing" to "working code"). ## Implementation status diff --git a/USAGE.md b/USAGE.md new file mode 100644 index 0000000..ada1b3e --- /dev/null +++ b/USAGE.md @@ -0,0 +1,721 @@ +# regexx usage guide + +This document is task oriented: how to build a pattern, run it against data +in each of the three supported modes, and read the result back out, with +complete, compiled, and run examples. [`docs/API.md`](docs/API.md) is the +exhaustive reference for every function's exact return value and memory +ownership rule; this document exists to get from "nothing" to "working code" +without reading that whole reference first. [`concept.md`](concept.md) +records the design rationale for anyone who wants to know why a given +choice was made rather than only what it is. + +Every example below was compiled against the current `regexx.c` and its +actual printed output is shown, not a hand-written guess at what it would +print. + +## 1. Building and linking + +```sh +make # builds libregexx.a and the rxgrep example +``` + +A program using the library needs `regexx.h` on its include path and +`libregexx.a` (or `regexx.c` compiled directly into the program, since the +library is a single file) linked in: + +```sh +cc -std=c11 -I/path/to/regexx -o myprog myprog.c /path/to/regexx/libregexx.a +``` + +No third-party dependency is required; `regexx.c` uses only the C standard +library plus, for `UTF8` mode's Unicode classification and `Input_from_file`, +POSIX (`wctype.h`, `locale.h`, `mmap`/`open`/`fstat`, `concept.md` Section 8). + +## 2. The three data modes + +Exactly one of `BINARY`, `ASCII`, or `UTF8` is passed in `flags` to +`re_compile` (or any `re_`-prefixed module level function); it selects how +the subject's bytes are interpreted, not how the pattern text itself is +encoded (the pattern is always read as ASCII/UTF-8 source text regardless of +this choice). + +| Mode | Subject interpretation | `Match_start`/`Match_end` unit | +|---|---|---| +| `BINARY` | Raw bytes, `0`-`255` all valid, embedded `NUL` is an ordinary byte, no decoding at all. | Byte offset | +| `ASCII` | Bytes, but `\d`/`\w`/`\s` and `IGNORECASE` restrict themselves to the ASCII subset (the default when neither `BINARY` nor `UTF8` is given, `docs/API.md` Section 2). | Byte offset | +| `UTF8` | Decoded as UTF-8 into Unicode code points before matching; `\d`/`\w`/`\s` and case folding consult the platform's Unicode tables via `wctype.h`. | Code point index (use `Match_start_byte`/`_end_byte` for a byte offset into the raw buffer, Section 6 below) | + +Passing invalid UTF-8 to a `UTF8`-mode pattern's matching call returns `-1` +(the general error return, `docs/API.md` Section 3); `BINARY` and `ASCII` +mode never reject a subject on this basis, since neither one decodes it. + +Choosing a mode is a decision made once per `Pattern` (it is baked into +`flags` at `re_compile` time), not per matching call: a `Pattern` compiled +with `UTF8` decodes every `Input` it is later run against as UTF-8, and a +`Pattern` compiled with `BINARY` never does, regardless of what the actual +bytes happen to look like. + +## 3. Compiling a pattern + +```c +#include "regexx.h" +#include + +PatternError err; memset(&err, 0, sizeof err); +Pattern *pat = re_compile("(\\w+)@(\\w+)\\.(\\w+)", strlen("(\\w+)@(\\w+)\\.(\\w+)"), ASCII, &err); +if (!pat) { + fprintf(stderr, "pattern error at %lld: %s\n", (long long)err.pos, err.msg); + PatternError_free(&err); + /* handle failure */ +} +``` + +- `pattern`/`len`: the pattern source; it need not be `NUL`-terminated, `len` + is authoritative. +- `flags`: the data mode (Section 2) combined with any of `IGNORECASE`, + `MULTILINE`, `DOTALL`, `VERBOSE`, `ASCII` (also usable as a flag inside + `UTF8` mode to force ASCII-only `\d`/`\w`/`\s`), `UNICODE`, `LOCALE`, + `DEBUG` (the last three accepted for source compatibility with Python, + currently no-ops, `docs/API.md` Section 2). +- `err`: optional (`NULL` is safe); filled in on a syntax error or an + unsupported construct (`docs/API.md` Section 5 lists every construct this + build rejects at compile time, such as conditional groups and scoped + inline flags). Call `PatternError_free` on it once read, whether or not + compilation succeeded. + +A compiled `Pattern` is freed with `Pattern_free`, once, after every `Match` +produced from it has been released (`Match.re` borrows from the `Pattern`, +so freeing the `Pattern` first and a `Match` from it afterward is a use +after free, `docs/API.md` Section 1.2). + +## 4. Building an `Input` + +Every matching call takes an `Input`, not a raw pointer, so the library has +one place to resolve "where are the bytes" regardless of whether they came +from memory the caller already has or a file on disk. + +```c +Input *Input_from_buffer(const uint8_t *buf, size_t len); /* wraps, does not copy or own */ +Input *Input_from_file(const char *path, PatternError *err); /* mmaps or reads a whole file */ +void Input_free(Input *in); +``` + +`Input_from_buffer` wraps an existing buffer (a string literal, a `malloc`'d +region, a memory-mapped region the caller manages itself) without copying +it; the buffer must outlive the `Input` and every `Match` produced from it, +since `Match_group` returns pointers directly into it (`docs/API.md` Section +1.3). `Input_free` on an `Input_from_buffer` result never frees the wrapped +buffer. + +`Input_from_file` maps a regular file read-only with `mmap()` instead of +copying it, so the pages stay reclaimable under memory pressure +(`README.md` "Memory footprint" has the measured numbers); a non-seekable +source (a pipe, a FIFO, process substitution, `/dev/stdin`) or an `mmap()` +failure falls back to reading it incrementally into an owned buffer instead. +Either way `Input_free` releases what it acquired (`munmap` or `free`, as +appropriate) automatically; the caller never needs to know which path was +taken. + +```c +PatternError err; memset(&err, 0, sizeof err); +Input *in = Input_from_file("data.txt", &err); +if (!in) { fprintf(stderr, "cannot read file: %s\n", err.msg); PatternError_free(&err); } +``` + +Output, verified: + +``` +=== error_file === +in=(nil) msg='No such file or directory' +``` + +## 5. Matching: `match`, `fullmatch`, `search` + +```c +int Pattern_match(Pattern *self, Input *string, int64_t pos, int64_t endpos, Match *out); +int Pattern_fullmatch(Pattern *self, Input *string, int64_t pos, int64_t endpos, Match *out); +int Pattern_search(Pattern *self, Input *string, int64_t pos, int64_t endpos, Match *out); +``` + +`pos`/`endpos` bound the attempt (pass `0`/`-1` for "the whole input", the +same defaults `re.Pattern.match`/`.search`/etc. use); `match` and +`fullmatch` are anchored at `pos` (`fullmatch` additionally requires +reaching exactly `endpos`), `search` tries every start position from `pos` +to `endpos` and reports the first that admits a match. All three return `1` +matched, `0` no match, `-1` error; `out` must point at a zero-initialized +`Match` (never allocated by the library itself). + +```c +const char *pattern = "(\\w+)@(\\w+)\\.(\\w+)"; +Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err); +const char *subject = "contact: user@example.com today"; +Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject)); + +Match m; memset(&m, 0, sizeof m); +if (Pattern_search(pat, in, 0, -1, &m) == 1) { + const char *g; size_t glen; + Match_group(&m, NULL, 0, &g, &glen); /* whole match: group 0 */ + Match_group(&m, NULL, 1, &g, &glen); /* first capturing group */ + int64_t s, e; + Match_span(&m, 0, &s, &e); + Match_free(&m); +} +Input_free(in); +Pattern_free(pat); +``` + +Output, verified: + +``` +=== ascii_basic === +search rc=1 +group0='user@example.com' +group1='user' +group2='example' +group3='com' +span0=[9,25) +``` + +### Named groups + +```c +const char *pattern = "(?P\\w+)@(?P\\w+)"; +Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err); +const char *subject = "user@host"; +Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject)); +Match m; memset(&m, 0, sizeof m); +if (Pattern_search(pat, in, 0, -1, &m) == 1) { + const char *g; size_t glen; + Match_group(&m, "user", 0, &g, &glen); /* name_or_null non-NULL: index is ignored */ + Match_group(&m, "host", 0, &g, &glen); + int idx = Pattern_groupindex_lookup(pat, "host"); /* 1-based group number, or -1 */ + Match_free(&m); +} +``` + +Output, verified: + +``` +=== ascii_named === +user='user' +host='host' +groupindex(host)=2 +``` + +`Match_group` returns `1` and a borrowed pointer into the `Input`'s buffer +when the group participated, `0` and `*out = NULL` when the group exists in +the pattern but did not participate (Python's `None`), and `-1` when no +such group exists at all, by index or by name; this three-way distinction +is the only place in the API that tells "no such group" apart from "this +group matched nothing" (`Match_start`/`Match_end` collapse both cases to +`-1`, `docs/API.md` Section 3.23-3.29). + +## 6. Code points versus bytes in `UTF8` mode + +In `UTF8` mode, `Match_start`/`Match_end`/`Match_span` report a code point +index (so, for example, a 2-byte UTF-8 character like `é` counts as one +unit, the same way Python's `str` indexing does), while +`Match_start_byte`/`Match_end_byte`/`Match_span_byte` always report a byte +offset into the raw `Input` buffer, in every mode. Use the `_byte` variants +whenever the result is used to slice or seek into the raw buffer or file +directly (`Match_group` already does this translation internally, so it +needs neither variant directly; a caller that needs to, for example, report +a byte offset to an external tool that has no notion of code points is the +usual reason to reach for `_byte` directly). + +```c +const char *pattern = "\\w+"; +Pattern *pat = re_compile(pattern, strlen(pattern), UTF8, &err); +const char *subject = "caf\xc3\xa9 today"; /* "café today", é is 2 UTF-8 bytes */ +Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject)); +Match m; memset(&m, 0, sizeof m); +if (Pattern_search(pat, in, 0, -1, &m) == 1) { + int64_t cs, ce, bs, be; + Match_span(&m, 0, &cs, &ce); /* code point units */ + Match_span_byte(&m, 0, &bs, &be); /* byte units */ + Match_free(&m); +} +``` + +Output, verified: + +``` +=== utf8_offsets === +search rc=1 +codepoint span=[0,4) byte span=[0,5) +group bytes: 'café' (5 bytes) +``` + +`café` is 4 code points (`c`, `a`, `f`, `é`) but 5 bytes (`é` alone is 2 +bytes in UTF-8), which is exactly what the two spans above report. + +### Iterating UTF-8 text + +```c +const char *pattern = "\\w+"; +Pattern *pat = re_compile(pattern, strlen(pattern), UTF8, &err); +const char *subject = "na\xc3\xafve caf\xc3\xa9 r\xc3\xa9sum\xc3\xa9"; /* naïve café résumé */ +Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject)); +Pattern_finditer(pat, in, 0, -1, print_match_cb, &ctx); +``` + +Output, verified (`\w` under `UTF8` mode matches letters outside ASCII too, +consulting the platform's Unicode tables, `docs/API.md` Section 2): + +``` +=== utf8_finditer === + match 0: 'naïve' + match 1: 'café' + match 2: 'résumé' +total matches=3 +``` + +## 7. Binary data + +`BINARY` mode never decodes the subject; every byte `0`-`255` is a valid +input value, including `0x00`, and a `.`/character class/literal matches by +raw byte value. + +```c +const char *pattern = "a.c"; +Pattern *pat = re_compile(pattern, strlen(pattern), BINARY, &err); +uint8_t subject[] = { 'x', 'a', 0x00, 'c', 'y' }; +Input *in = Input_from_buffer(subject, sizeof subject); +Match m; memset(&m, 0, sizeof m); +Pattern_search(pat, in, 0, -1, &m); +``` + +Output, verified (`.` matches the embedded `NUL` byte at offset 2, since +`DOTALL` is irrelevant here: `BINARY` mode has no notion of "newline" ending +`.`'s match, byte `0x0a` is simply a byte like any other unless the pattern +excludes it explicitly): + +``` +=== binary_basic === +search rc=1 +matched 3 bytes: 61 00 63 +``` + +Byte ranges in a character class work the same way, by raw value: + +```c +const char *pattern = "[\\x00-\\x1f]+"; +Pattern *pat = re_compile(pattern, strlen(pattern), BINARY, &err); +uint8_t subject[] = { 'A', 0x01, 0x02, 0x1f, 'B' }; +``` + +Output, verified: + +``` +=== binary_class === +search rc=1 +span=[1,4) +``` + +`\d`/`\w`/`\s` in `BINARY` mode classify the same way they do in `ASCII` +mode (ASCII letters/digits/whitespace only; there is no "Unicode binary" +notion to fall back to, since `BINARY` mode has no decoding step at all). + +## 8. Finding all matches: `finditer`/`findall` + +```c +int Pattern_finditer(Pattern *self, Input *string, int64_t pos, int64_t endpos, MatchIterCb cb, void *ctx); +``` + +`cb(ctx, m)` fires once per non-overlapping match, left to right; the +`Match` passed in is only valid for the duration of the call (freed +immediately after `cb` returns, so do not retain the pointer past it). +`Pattern_findall` is the same call under a second name, for parity with +`re.findall`; this build always hands the callback a full `Match` rather +than collapsing it to "just the string" or a tuple the way CPython's +`findall` does at the Python-object level, since there is no C object model +to collapse into (`docs/API.md` Section 3.5-3.6). + +```c +typedef struct { int n; } IterCtx; +static void print_match_cb(void *ctx, const Match *m_const) { + IterCtx *c = ctx; + Match *m = (Match *)m_const; + const char *g; size_t glen; + Match_group(m, NULL, 0, &g, &glen); + printf(" match %d: '%.*s'\n", c->n, (int)glen, g); + c->n++; +} + +const char *pattern = "\\d+"; +Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err); +const char *subject = "order 12 has 345 items, batch 6"; +Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject)); +IterCtx ctx = { 0 }; +int n = Pattern_finditer(pat, in, 0, -1, print_match_cb, &ctx); +``` + +Output, verified: + +``` +=== ascii_finditer === + match 0: '12' + match 1: '345' + match 2: '6' +total matches=3 +``` + +`finditer`/`findall`/`split`/`sub` all share one memoization table and one +set of precomputed skip-ahead tables across the whole scan (built once per +call, not once per match), which is what makes them run in linear time for +the common case rather than redoing quadratic work; see README.md +"Implementation status" and "Memory footprint" for the measured numbers and +what that costs in memory. + +## 9. Splitting: `Pattern_split` + +```c +int Pattern_split(Pattern *self, Input *string, int maxsplit, MatchIterCb cb, void *ctx); +``` + +`cb` fires once per element of the list `re.split()` would return, in +order; each element's text is read the same way as any other match, via +`Match_group(m, NULL, 0, &out, &outlen)`, and a `0` return from that call +means the element is Python's `None` (an unparticipated capturing group +between two matches, `docs/API.md` Section 3.7). `maxsplit` matches +`re.split`'s parameter (`0` unlimited). + +```c +const char *pattern = "\\s*,\\s*"; +Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err); +const char *subject = "red, green,blue , yellow"; +Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject)); +IterCtx ctx = { 0 }; +int n = Pattern_split(pat, in, 0, print_split_cb, &ctx); +``` + +Output, verified: + +``` +=== ascii_split === + elem 0: 'red' + elem 1: 'green' + elem 2: 'blue' + elem 3: 'yellow' +splits=3 +``` + +## 10. Substitution: `Pattern_sub`/`Pattern_subn` + +```c +int Pattern_sub(Pattern *self, Input *string, const char *repl, MatchSubCb cb, void *ctx, int count, char **out, size_t *outlen); +int Pattern_subn(Pattern *self, Input *string, const char *repl, MatchSubCb cb, void *ctx, int count, char **out, size_t *outlen, int *n); +``` + +Exactly one of `repl` (a template string) or `cb` (a callback) is +non-`NULL`; this two-parameter shape exists because C cannot express +Python's single polymorphic `repl` argument (string or callable) in one +slot (`docs/API.md` Section 3.8-3.9). `Pattern_subn` additionally reports +the number of substitutions made through `n`; `Pattern_sub` is the same +operation with that count discarded. `*out` is a freshly `malloc`'d, +`NUL`-terminated buffer the caller must `free`. + +### Template substitution + +Template syntax: `\g`, `\g`, `\N` (one or two digits), `\n`, `\t`, +`\\`, and any other `\X` as the literal character `X`. + +```c +const char *pattern = "(\\w+)@(\\w+)"; +Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err); +const char *subject = "user@host"; +Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject)); +char *out; size_t outlen; int n; +Pattern_subn(pat, in, "\\2@\\1", NULL, NULL, 0, &out, &outlen, &n); +/* use out/outlen, then: */ +free(out); +``` + +Output, verified: + +``` +=== ascii_sub_template === +result='host@user' n=1 +``` + +### Callback substitution + +```c +static void upper_cb(void *ctx, const Match *m_const, char **out, size_t *outlen) { + Match *m = (Match *)m_const; + const char *g; size_t glen; + Match_group(m, NULL, 0, &g, &glen); + char *buf = malloc(glen); + for (size_t i = 0; i < glen; i++) { + char c = g[i]; + buf[i] = (c >= 'a' && c <= 'z') ? (char)(c - 32) : c; + } + *out = buf; *outlen = glen; /* ownership passes to Pattern_sub */ +} + +const char *pattern = "\\w+"; +Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err); +const char *subject = "shout this loudly"; +Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject)); +char *out; size_t outlen; +Pattern_sub(pat, in, NULL, upper_cb, NULL, 0, &out, &outlen); +free(out); +``` + +Output, verified: + +``` +=== ascii_sub_callback === +result='SHOUT THIS LOUDLY' +``` + +The callback is expected to `malloc` its replacement and hand ownership of +it off through `*out`/`*outlen`; `Pattern_sub`/`Pattern_subn` frees it +immediately after copying its content into the final result buffer +(`docs/API.md` Section 1.5). + +## 11. Flags + +```c +Pattern *pat = re_compile("hello", 5, ASCII | IGNORECASE, &err); +``` + +```c +const char *pattern = "hello"; +Pattern *pat = re_compile(pattern, strlen(pattern), ASCII | IGNORECASE, &err); +const char *subject = "HELLO world"; +Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject)); +Match m; memset(&m, 0, sizeof m); +Pattern_search(pat, in, 0, -1, &m); /* rc=1 */ +``` + +```c +const char *pattern = "^line"; +Pattern *pat = re_compile(pattern, strlen(pattern), ASCII | MULTILINE, &err); +const char *subject = "first\nline two\nline three"; +/* Pattern_finditer finds 2 matches: MULTILINE makes ^ match after every \n too */ +``` + +```c +const char *pattern = "a.b"; +Pattern *pat = re_compile(pattern, strlen(pattern), ASCII | DOTALL, &err); +const char *subject = "a\nb"; +/* rc=1: DOTALL makes . match \n too */ +``` + +```c +const char *pattern = "\\d+ # a number\n\\s+ \\w+ # then a word"; +Pattern *pat = re_compile(pattern, strlen(pattern), ASCII | VERBOSE, &err); +const char *subject = "42 answer"; +/* rc=1: VERBOSE strips the unescaped whitespace and # comments before parsing */ +``` + +Output, verified for all four: + +``` +=== flags_example === +IGNORECASE rc=1 + match 0: 'line' + match 1: 'line' +MULTILINE matches=2 +DOTALL rc=1 +VERBOSE rc=1 +``` + +`UNICODE`, `LOCALE`, and `DEBUG` are accepted (for source compatibility with +Python) and combine with any of the above via `|`, but currently have no +observable effect (`docs/API.md` Section 2 records exactly why for each +one, in particular why `LOCALE` is a no-op rather than the bug it used to +be, caught and fixed by this project's test suite). + +## 12. Matching against a file + +```c +const char *pattern = "line \\w+"; +Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err); +Input *in = Input_from_file("data.txt", &err); +IterCtx ctx = { 0 }; +int n = Pattern_finditer(pat, in, 0, -1, print_match_cb, &ctx); +Input_free(in); +``` + +Given a file containing `line one\nline two has data\nline three\n`, output, +verified: + +``` +=== file_input === +Input_from_file rc=0x5b30fbc44e20 + match 0: 'line one' + match 1: 'line two' + match 2: 'line three' +total matches=3 +``` + +As Section 4 explains, `Input_from_file` mmaps a regular file rather than +copying it; nothing about matching against it differs from matching against +an `Input_from_buffer`-wrapped in-memory buffer, the mode (`BINARY`/`ASCII`/ +`UTF8`) and every function above work identically regardless of which +`Input_from_*` constructor produced the `Input`. See README.md "Memory +footprint" for the measured memory cost of matching a large file this way, +and for a documented, deliberately reverted attempt at bounding it that +traded a fast allocation failure for an effectively unbounded hang (worth +reading before assuming a size cap is a safe thing to add here). + +## 13. Module level convenience functions + +```c +int re_search(const char *pattern, size_t len, int flags, Input *string, Match *out); +``` + +`re_match`/`re_fullmatch`/`re_search`/`re_finditer`/`re_findall`/`re_split`/ +`re_sub`/`re_subn` compile `pattern` through an internal cache (up to 512 +entries, keyed by `(pattern, len, flags)`, cleared entirely on overflow) +and then call the matching `Pattern_` function with `pos=0`, `endpos=-1`, +mirroring how CPython itself implements `re.match` as +`_compile(pattern, flags).match(string)`. + +```c +const char *pattern = "\\d+"; +const char *subject = "abc123def"; +Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject)); +Match m; memset(&m, 0, sizeof m); +int r = re_search(pattern, strlen(pattern), ASCII, in, &m); +``` + +Output, verified: + +``` +=== module_level === +re_search rc=1 +matched='123' +``` + +A compile error inside these functions is reported only as a `-1` return, +with no `PatternError` available; use `re_compile` plus a `Pattern_` +function directly whenever a compile error needs to be diagnosed, or +whenever the same pattern is used more than once (precompiling once and +reusing the `Pattern` avoids repeated cache lookups, matching idiomatic +Python's own preference for `re.compile` in a loop over calling the module +level function repeatedly). `re_purge()` clears the cache, matching +`re.purge()`. This cache is shared, mutable, process-wide state with no +internal locking; do not call a `re_`-prefixed function from more than one +thread without external synchronization (`Pattern_`-prefixed functions on +an already-compiled `Pattern` have no such restriction, since nothing here +mutates a compiled `Pattern`). + +## 14. Escaping literal text: `re_escape` + +```c +void re_escape(const char *in, size_t len, char **out, size_t *outlen); +``` + +```c +const char *s = "3.14 (pi)"; +char *out; size_t outlen; +re_escape(s, strlen(s), &out, &outlen); +free(out); +``` + +Output, verified: + +``` +=== escape_example === +escaped='3\.14\ \(pi\)' +``` + +Matches `re.escape` exactly, including the narrowed escaped-character set +CPython adopted in 3.7 (only characters that are actually special in a +regex are escaped; other non-alphanumeric bytes are passed through +unescaped). + +## 15. Error handling + +Two kinds of failure exist in this API: a compile-time `PatternError` +(`re_compile`, `Input_from_file`) and a matching-time `-1` return with no +further detail (`docs/API.md` Section 3). Always check the return value of +a compile call before using the `Pattern` it was supposed to produce. + +```c +const char *pattern = "(unclosed"; +PatternError err; memset(&err, 0, sizeof err); +Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err); +if (!pat) { + fprintf(stderr, "pattern error at %lld: %s\n", (long long)err.pos, err.msg); +} +PatternError_free(&err); +``` + +Output, verified: + +``` +=== error_compile === +pat=(nil) msg='missing ), unterminated subpattern' pos=9 lineno=1 colno=10 +``` + +`err.msg`/`err.pattern` are heap allocated by whichever call filled the +struct in; zero-initialize a `PatternError` before use and call +`PatternError_free` on it once done reading it, whether or not the call +that filled it in succeeded. A `NULL` `err` argument is always safe to pass +to any function that accepts one, and simply skips error reporting. + +A matching call (`Pattern_search`, and so on) returning `-1` means either +invalid UTF-8 in the subject against a `UTF8`-mode `Pattern`, or the +backtracking depth limit was reached (`MAX_DEPTH` in `regexx.c`, currently +60000, not adjustable without editing `regexx.c` and rebuilding); no +further detail is available through the return value alone in this build. + +## 16. Pattern syntax reference + +Literals; `.` (with `DOTALL` for it to match `\n`); character classes with +ranges and negation (`[abc]`, `[^abc]`, `[a-z]`); `\d \D \w \W \s \S`; +`\b \B`; anchors `^ $ \A \Z` (`^`/`$` also match at line boundaries with +`MULTILINE`); quantifiers `* + ? {m,n} {m,} {,n} {m}`, both greedy and lazy +(`*?`, `+?`, and so on) forms; possessive quantifiers `*+ ++ ?+ {m,n}+`; +groups `(...)`, non-capturing `(?:...)`, named `(?P...)`; alternation +`|`; backreferences `\1`-`\99`, `(?P=name)`, `\g`, `\g`; lookahead +`(?=...)`/`(?!...)`; fixed-width lookbehind `(?<=...)`/`(?...)`; comments `(?#...)`; global inline flags `(?aiLmsux)` at +the very start of a pattern; escapes `\n \r \t \f \v \a`, `\0`-prefixed +octal escapes, `\xhh`, `\uxxxx`, `\Uxxxxxxxx`. + +Rejected at compile time with a `PatternError` naming the construct, rather +than silently mis-parsed: conditional groups `(?(id)yes|no)`, scoped inline +flags `(?flags:...)` (only the global, start-of-pattern form is supported), +`\N{NAME}` named code points, variable-width lookbehind (matching CPython's +own restriction), and POSIX bracket-expression syntax `[[:alpha:]]` (not +part of Python `re`'s grammar at all, `concept.md` Section 14). See +`docs/API.md` Section 5-6 for the complete, exact list. + +## 17. Performance and memory, in one paragraph + +`match`/`fullmatch` (a single anchored attempt) cost time proportional to +that attempt. `search`/`finditer`/`split`/`sub` are linear time, not +quadratic, for the common case (no backreference, every quantifier's body a +single character/class/`.`), at a measured memory cost of roughly 8.4x the +input length for a pattern using a simple-atom quantifier. A pattern with a +backreference, or shaped like nested overlapping quantifiers +(`(a+)+b`-style), can still be worse than linear, the same way CPython's own +`re` can be on the same patterns. `README.md` "Implementation status" and +"Memory footprint" have the full measured account, including one fix that +was tried, measured, and deliberately reverted; read it before assuming a +given pattern's cost, rather than assuming linear time and bounded memory +apply unconditionally. + +## 18. Complete example: `rxgrep` + +`examples/rxgrep.c` is a complete, working grep-like program built on this +API, exercising all three data modes and both the matching and substitution +API against real files and pipes: + +```sh +./rxgrep -in 'hello' file.txt # case-insensitive, line numbers +./rxgrep -m utf8 -o '\w+' file.txt # print every UTF-8 word, one per line +./rxgrep -c 'error' log.txt # count matching lines +./rxgrep -m binary 'a.c' data.bin # match raw bytes, embedded NUL included +./rxgrep -m utf8 --sub 'REDACTED' '\d{3}-\d{4}' file.txt +``` + +Reading its source alongside this document is a reasonable next step once +the examples above are familiar: it shows every function here used together +in one program, including the parts this document simplified away for +clarity (option parsing, line splitting, output formatting).