README.md and docs/API.md both previously documented "no way to list every name a pattern defines without already knowing what to look for" as an accepted limitation of Pattern_groupindex_lookup being the only access to groupindex. It was not actually a hard constraint: the underlying GroupIndex struct (regexx.c) already stores every name and its group number in two parallel arrays, populated once at compile time; only a public accessor was missing. Pattern_groupindex_count returns the number of named groups; Pattern_groupindex_at(self, i, &name) for 0 <= i < count writes the i-th name and returns its 1-based group number, or returns -1 for an out-of-range i. Enumeration order is declaration order, verified against a real CPython 3.11 interpreter to match groupindex's own practical (insertion-order) iteration order, not just assumed. Verified directly (count/name/group-number correctness, matching the exact snippet now in USAGE.md's own output), full 3,252-case suite unaffected (3252/3252, this is a pure accessor addition touching no matching logic), clean AddressSanitizer/UndefinedBehaviorSanitizer. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
772 lines
29 KiB
Markdown
772 lines
29 KiB
Markdown
# regexx usage guide
|
|
|
|
This document is task oriented: how to build a pattern, run it against data
|
|
in each of the three supported modes, and read the result back out, with
|
|
complete, compiled, and run examples. [`docs/API.md`](docs/API.md) is the
|
|
exhaustive reference for every function's exact return value and memory
|
|
ownership rule; this document exists to get from "nothing" to "working code"
|
|
without reading that whole reference first. [`concept.md`](concept.md)
|
|
records the design rationale for anyone who wants to know why a given
|
|
choice was made rather than only what it is.
|
|
|
|
Every example below was compiled against the current `regexx.c` and its
|
|
actual printed output is shown, not a hand-written guess at what it would
|
|
print.
|
|
|
|
## 1. Building and linking
|
|
|
|
```sh
|
|
make # builds libregexx.a and the rxgrep example
|
|
```
|
|
|
|
A program using the library needs `regexx.h` on its include path and
|
|
`libregexx.a` (or `regexx.c` compiled directly into the program, since the
|
|
library is a single file) linked in:
|
|
|
|
```sh
|
|
cc -std=c11 -I/path/to/regexx -o myprog myprog.c /path/to/regexx/libregexx.a
|
|
```
|
|
|
|
No third-party dependency is required; `regexx.c` uses only the C standard
|
|
library plus, for `UTF8` mode's Unicode classification and `Input_from_file`,
|
|
POSIX (`wctype.h`, `locale.h`, `mmap`/`open`/`fstat`, `concept.md` Section 8).
|
|
|
|
## 2. The three data modes
|
|
|
|
Exactly one of `BINARY`, `ASCII`, or `UTF8` is passed in `flags` to
|
|
`re_compile` (or any `re_`-prefixed module level function); it selects how
|
|
the subject's bytes are interpreted, not how the pattern text itself is
|
|
encoded (the pattern is always read as ASCII/UTF-8 source text regardless of
|
|
this choice).
|
|
|
|
| Mode | Subject interpretation | `Match_start`/`Match_end` unit |
|
|
|---|---|---|
|
|
| `BINARY` | Raw bytes, `0`-`255` all valid, embedded `NUL` is an ordinary byte, no decoding at all. | Byte offset |
|
|
| `ASCII` | Bytes, but `\d`/`\w`/`\s` and `IGNORECASE` restrict themselves to the ASCII subset (the default when neither `BINARY` nor `UTF8` is given, `docs/API.md` Section 2). | Byte offset |
|
|
| `UTF8` | Decoded as UTF-8 into Unicode code points before matching; `\d`/`\w`/`\s` and case folding consult the platform's Unicode tables via `wctype.h`. | Code point index (use `Match_start_byte`/`_end_byte` for a byte offset into the raw buffer, Section 6 below) |
|
|
|
|
Passing invalid UTF-8 to a `UTF8`-mode pattern's matching call returns `-1`
|
|
(the general error return, `docs/API.md` Section 3); `BINARY` and `ASCII`
|
|
mode never reject a subject on this basis, since neither one decodes it.
|
|
|
|
Choosing a mode is a decision made once per `Pattern` (it is baked into
|
|
`flags` at `re_compile` time), not per matching call: a `Pattern` compiled
|
|
with `UTF8` decodes every `Input` it is later run against as UTF-8, and a
|
|
`Pattern` compiled with `BINARY` never does, regardless of what the actual
|
|
bytes happen to look like.
|
|
|
|
## 3. Compiling a pattern
|
|
|
|
```c
|
|
#include "regexx.h"
|
|
#include <string.h>
|
|
|
|
PatternError err; memset(&err, 0, sizeof err);
|
|
Pattern *pat = re_compile("(\\w+)@(\\w+)\\.(\\w+)", strlen("(\\w+)@(\\w+)\\.(\\w+)"), ASCII, &err);
|
|
if (!pat) {
|
|
fprintf(stderr, "pattern error at %lld: %s\n", (long long)err.pos, err.msg);
|
|
PatternError_free(&err);
|
|
/* handle failure */
|
|
}
|
|
```
|
|
|
|
- `pattern`/`len`: the pattern source; it need not be `NUL`-terminated, `len`
|
|
is authoritative.
|
|
- `flags`: the data mode (Section 2) combined with any of `IGNORECASE`,
|
|
`MULTILINE`, `DOTALL`, `VERBOSE`, `ASCII` (also usable as a flag inside
|
|
`UTF8` mode to force ASCII-only `\d`/`\w`/`\s`), `UNICODE`, `LOCALE`,
|
|
`DEBUG` (the last three accepted for source compatibility with Python,
|
|
currently no-ops, `docs/API.md` Section 2).
|
|
- `err`: optional (`NULL` is safe); filled in on a syntax error or an
|
|
unsupported construct (`docs/API.md` Section 5 lists every construct this
|
|
build rejects at compile time, such as conditional groups and scoped
|
|
inline flags). Call `PatternError_free` on it once read, whether or not
|
|
compilation succeeded.
|
|
|
|
A compiled `Pattern` is freed with `Pattern_free`, once, after every `Match`
|
|
produced from it has been released (`Match.re` borrows from the `Pattern`,
|
|
so freeing the `Pattern` first and a `Match` from it afterward is a use
|
|
after free, `docs/API.md` Section 1.2).
|
|
|
|
## 4. Building an `Input`
|
|
|
|
Every matching call takes an `Input`, not a raw pointer, so the library has
|
|
one place to resolve "where are the bytes" regardless of whether they came
|
|
from memory the caller already has or a file on disk.
|
|
|
|
```c
|
|
Input *Input_from_buffer(const uint8_t *buf, size_t len); /* wraps, does not copy or own */
|
|
Input *Input_from_file(const char *path, PatternError *err); /* mmaps or reads a whole file */
|
|
void Input_free(Input *in);
|
|
```
|
|
|
|
`Input_from_buffer` wraps an existing buffer (a string literal, a `malloc`'d
|
|
region, a memory-mapped region the caller manages itself) without copying
|
|
it; the buffer must outlive the `Input` and every `Match` produced from it,
|
|
since `Match_group` returns pointers directly into it (`docs/API.md` Section
|
|
1.3). `Input_free` on an `Input_from_buffer` result never frees the wrapped
|
|
buffer.
|
|
|
|
`Input_from_file` maps a regular file read-only with `mmap()` instead of
|
|
copying it, so the pages stay reclaimable under memory pressure
|
|
(`README.md` "Memory footprint" has the measured numbers); a non-seekable
|
|
source (a pipe, a FIFO, process substitution, `/dev/stdin`) or an `mmap()`
|
|
failure falls back to reading it incrementally into an owned buffer instead.
|
|
Either way `Input_free` releases what it acquired (`munmap` or `free`, as
|
|
appropriate) automatically; the caller never needs to know which path was
|
|
taken.
|
|
|
|
```c
|
|
PatternError err; memset(&err, 0, sizeof err);
|
|
Input *in = Input_from_file("data.txt", &err);
|
|
if (!in) { fprintf(stderr, "cannot read file: %s\n", err.msg); PatternError_free(&err); }
|
|
```
|
|
|
|
Output, verified:
|
|
|
|
```
|
|
=== error_file ===
|
|
in=(nil) msg='No such file or directory'
|
|
```
|
|
|
|
## 5. Matching: `match`, `fullmatch`, `search`
|
|
|
|
```c
|
|
int Pattern_match(Pattern *self, Input *string, int64_t pos, int64_t endpos, Match *out);
|
|
int Pattern_fullmatch(Pattern *self, Input *string, int64_t pos, int64_t endpos, Match *out);
|
|
int Pattern_search(Pattern *self, Input *string, int64_t pos, int64_t endpos, Match *out);
|
|
```
|
|
|
|
`pos`/`endpos` bound the attempt (pass `0`/`-1` for "the whole input", the
|
|
same defaults `re.Pattern.match`/`.search`/etc. use); `match` and
|
|
`fullmatch` are anchored at `pos` (`fullmatch` additionally requires
|
|
reaching exactly `endpos`), `search` tries every start position from `pos`
|
|
to `endpos` and reports the first that admits a match. All three return `1`
|
|
matched, `0` no match, `-1` error; `out` must point at a zero-initialized
|
|
`Match` (never allocated by the library itself).
|
|
|
|
```c
|
|
const char *pattern = "(\\w+)@(\\w+)\\.(\\w+)";
|
|
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
|
|
const char *subject = "contact: user@example.com today";
|
|
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
|
|
|
|
Match m; memset(&m, 0, sizeof m);
|
|
if (Pattern_search(pat, in, 0, -1, &m) == 1) {
|
|
const char *g; size_t glen;
|
|
Match_group(&m, NULL, 0, &g, &glen); /* whole match: group 0 */
|
|
Match_group(&m, NULL, 1, &g, &glen); /* first capturing group */
|
|
int64_t s, e;
|
|
Match_span(&m, 0, &s, &e);
|
|
Match_free(&m);
|
|
}
|
|
Input_free(in);
|
|
Pattern_free(pat);
|
|
```
|
|
|
|
Output, verified:
|
|
|
|
```
|
|
=== ascii_basic ===
|
|
search rc=1
|
|
group0='user@example.com'
|
|
group1='user'
|
|
group2='example'
|
|
group3='com'
|
|
span0=[9,25)
|
|
```
|
|
|
|
### Named groups
|
|
|
|
```c
|
|
const char *pattern = "(?P<user>\\w+)@(?P<host>\\w+)";
|
|
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
|
|
const char *subject = "user@host";
|
|
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
|
|
Match m; memset(&m, 0, sizeof m);
|
|
if (Pattern_search(pat, in, 0, -1, &m) == 1) {
|
|
const char *g; size_t glen;
|
|
Match_group(&m, "user", 0, &g, &glen); /* name_or_null non-NULL: index is ignored */
|
|
Match_group(&m, "host", 0, &g, &glen);
|
|
int idx = Pattern_groupindex_lookup(pat, "host"); /* 1-based group number, or -1 */
|
|
Match_free(&m);
|
|
}
|
|
```
|
|
|
|
Output, verified:
|
|
|
|
```
|
|
=== ascii_named ===
|
|
user='user'
|
|
host='host'
|
|
groupindex(host)=2
|
|
```
|
|
|
|
`Match_group` returns `1` and a borrowed pointer into the `Input`'s buffer
|
|
when the group participated, `0` and `*out = NULL` when the group exists in
|
|
the pattern but did not participate (Python's `None`), and `-1` when no
|
|
such group exists at all, by index or by name; this three-way distinction
|
|
is the only place in the API that tells "no such group" apart from "this
|
|
group matched nothing" (`Match_start`/`Match_end` collapse both cases to
|
|
`-1`, `docs/API.md` Section 3.23-3.29).
|
|
|
|
Every name a `Pattern` defines can also be enumerated, without already
|
|
knowing what to look for, via `Pattern_groupindex_count`/`_at` (Python:
|
|
`len(pattern.groupindex)`, `dict(pattern.groupindex).items()`):
|
|
|
|
```c
|
|
int n = Pattern_groupindex_count(pat);
|
|
for (int i = 0; i < n; i++) {
|
|
const char *name;
|
|
int gnum = Pattern_groupindex_at(pat, i, &name);
|
|
printf(" [%d] name=%s group=%d\n", i, name, gnum);
|
|
}
|
|
```
|
|
|
|
Output, verified, for the same `pat` above:
|
|
|
|
```
|
|
count=2
|
|
[0] name=user group=1
|
|
[1] name=host group=2
|
|
```
|
|
|
|
## 6. Code points versus bytes in `UTF8` mode
|
|
|
|
In `UTF8` mode, `Match_start`/`Match_end`/`Match_span` report a code point
|
|
index (so, for example, a 2-byte UTF-8 character like `é` counts as one
|
|
unit, the same way Python's `str` indexing does), while
|
|
`Match_start_byte`/`Match_end_byte`/`Match_span_byte` always report a byte
|
|
offset into the raw `Input` buffer, in every mode. Use the `_byte` variants
|
|
whenever the result is used to slice or seek into the raw buffer or file
|
|
directly (`Match_group` already does this translation internally, so it
|
|
needs neither variant directly; a caller that needs to, for example, report
|
|
a byte offset to an external tool that has no notion of code points is the
|
|
usual reason to reach for `_byte` directly).
|
|
|
|
```c
|
|
const char *pattern = "\\w+";
|
|
Pattern *pat = re_compile(pattern, strlen(pattern), UTF8, &err);
|
|
const char *subject = "caf\xc3\xa9 today"; /* "café today", é is 2 UTF-8 bytes */
|
|
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
|
|
Match m; memset(&m, 0, sizeof m);
|
|
if (Pattern_search(pat, in, 0, -1, &m) == 1) {
|
|
int64_t cs, ce, bs, be;
|
|
Match_span(&m, 0, &cs, &ce); /* code point units */
|
|
Match_span_byte(&m, 0, &bs, &be); /* byte units */
|
|
Match_free(&m);
|
|
}
|
|
```
|
|
|
|
Output, verified:
|
|
|
|
```
|
|
=== utf8_offsets ===
|
|
search rc=1
|
|
codepoint span=[0,4) byte span=[0,5)
|
|
group bytes: 'café' (5 bytes)
|
|
```
|
|
|
|
`café` is 4 code points (`c`, `a`, `f`, `é`) but 5 bytes (`é` alone is 2
|
|
bytes in UTF-8), which is exactly what the two spans above report.
|
|
|
|
### Iterating UTF-8 text
|
|
|
|
```c
|
|
const char *pattern = "\\w+";
|
|
Pattern *pat = re_compile(pattern, strlen(pattern), UTF8, &err);
|
|
const char *subject = "na\xc3\xafve caf\xc3\xa9 r\xc3\xa9sum\xc3\xa9"; /* naïve café résumé */
|
|
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
|
|
Pattern_finditer(pat, in, 0, -1, print_match_cb, &ctx);
|
|
```
|
|
|
|
Output, verified (`\w` under `UTF8` mode matches letters outside ASCII too,
|
|
consulting the platform's Unicode tables, `docs/API.md` Section 2):
|
|
|
|
```
|
|
=== utf8_finditer ===
|
|
match 0: 'naïve'
|
|
match 1: 'café'
|
|
match 2: 'résumé'
|
|
total matches=3
|
|
```
|
|
|
|
## 7. Binary data
|
|
|
|
`BINARY` mode never decodes the subject; every byte `0`-`255` is a valid
|
|
input value, including `0x00`, and a `.`/character class/literal matches by
|
|
raw byte value.
|
|
|
|
```c
|
|
const char *pattern = "a.c";
|
|
Pattern *pat = re_compile(pattern, strlen(pattern), BINARY, &err);
|
|
uint8_t subject[] = { 'x', 'a', 0x00, 'c', 'y' };
|
|
Input *in = Input_from_buffer(subject, sizeof subject);
|
|
Match m; memset(&m, 0, sizeof m);
|
|
Pattern_search(pat, in, 0, -1, &m);
|
|
```
|
|
|
|
Output, verified (`.` matches the embedded `NUL` byte at offset 2, since
|
|
`DOTALL` is irrelevant here: `BINARY` mode has no notion of "newline" ending
|
|
`.`'s match, byte `0x0a` is simply a byte like any other unless the pattern
|
|
excludes it explicitly):
|
|
|
|
```
|
|
=== binary_basic ===
|
|
search rc=1
|
|
matched 3 bytes: 61 00 63
|
|
```
|
|
|
|
Byte ranges in a character class work the same way, by raw value:
|
|
|
|
```c
|
|
const char *pattern = "[\\x00-\\x1f]+";
|
|
Pattern *pat = re_compile(pattern, strlen(pattern), BINARY, &err);
|
|
uint8_t subject[] = { 'A', 0x01, 0x02, 0x1f, 'B' };
|
|
```
|
|
|
|
Output, verified:
|
|
|
|
```
|
|
=== binary_class ===
|
|
search rc=1
|
|
span=[1,4)
|
|
```
|
|
|
|
`\d`/`\w`/`\s` in `BINARY` mode classify the same way they do in `ASCII`
|
|
mode (ASCII letters/digits/whitespace only; there is no "Unicode binary"
|
|
notion to fall back to, since `BINARY` mode has no decoding step at all).
|
|
|
|
## 8. Finding all matches: `finditer`/`findall`
|
|
|
|
```c
|
|
int Pattern_finditer(Pattern *self, Input *string, int64_t pos, int64_t endpos, MatchIterCb cb, void *ctx);
|
|
```
|
|
|
|
`cb(ctx, m)` fires once per non-overlapping match, left to right; the
|
|
`Match` passed in is only valid for the duration of the call (freed
|
|
immediately after `cb` returns, so do not retain the pointer past it).
|
|
`Pattern_findall` is the same call under a second name, for parity with
|
|
`re.findall`; this build always hands the callback a full `Match` rather
|
|
than collapsing it to "just the string" or a tuple the way CPython's
|
|
`findall` does at the Python-object level, since there is no C object model
|
|
to collapse into (`docs/API.md` Section 3.5-3.6).
|
|
|
|
```c
|
|
typedef struct { int n; } IterCtx;
|
|
static void print_match_cb(void *ctx, const Match *m_const) {
|
|
IterCtx *c = ctx;
|
|
Match *m = (Match *)m_const;
|
|
const char *g; size_t glen;
|
|
Match_group(m, NULL, 0, &g, &glen);
|
|
printf(" match %d: '%.*s'\n", c->n, (int)glen, g);
|
|
c->n++;
|
|
}
|
|
|
|
const char *pattern = "\\d+";
|
|
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
|
|
const char *subject = "order 12 has 345 items, batch 6";
|
|
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
|
|
IterCtx ctx = { 0 };
|
|
int n = Pattern_finditer(pat, in, 0, -1, print_match_cb, &ctx);
|
|
```
|
|
|
|
Output, verified:
|
|
|
|
```
|
|
=== ascii_finditer ===
|
|
match 0: '12'
|
|
match 1: '345'
|
|
match 2: '6'
|
|
total matches=3
|
|
```
|
|
|
|
`finditer`/`findall`/`split`/`sub` all share one memoization table and one
|
|
set of precomputed skip-ahead tables across the whole scan (built once per
|
|
call, not once per match), which is what makes them run in linear time for
|
|
the common case rather than redoing quadratic work; see README.md
|
|
"Implementation status" and "Memory footprint" for the measured numbers and
|
|
what that costs in memory.
|
|
|
|
## 9. Splitting: `Pattern_split`
|
|
|
|
```c
|
|
int Pattern_split(Pattern *self, Input *string, int maxsplit, MatchIterCb cb, void *ctx);
|
|
```
|
|
|
|
`cb` fires once per element of the list `re.split()` would return, in
|
|
order; each element's text is read the same way as any other match, via
|
|
`Match_group(m, NULL, 0, &out, &outlen)`, and a `0` return from that call
|
|
means the element is Python's `None` (an unparticipated capturing group
|
|
between two matches, `docs/API.md` Section 3.7). `maxsplit` matches
|
|
`re.split`'s parameter (`0` unlimited).
|
|
|
|
```c
|
|
const char *pattern = "\\s*,\\s*";
|
|
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
|
|
const char *subject = "red, green,blue , yellow";
|
|
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
|
|
IterCtx ctx = { 0 };
|
|
int n = Pattern_split(pat, in, 0, print_split_cb, &ctx);
|
|
```
|
|
|
|
Output, verified:
|
|
|
|
```
|
|
=== ascii_split ===
|
|
elem 0: 'red'
|
|
elem 1: 'green'
|
|
elem 2: 'blue'
|
|
elem 3: 'yellow'
|
|
splits=3
|
|
```
|
|
|
|
## 10. Substitution: `Pattern_sub`/`Pattern_subn`
|
|
|
|
```c
|
|
int Pattern_sub(Pattern *self, Input *string, const char *repl, MatchSubCb cb, void *ctx, int count, char **out, size_t *outlen);
|
|
int Pattern_subn(Pattern *self, Input *string, const char *repl, MatchSubCb cb, void *ctx, int count, char **out, size_t *outlen, int *n);
|
|
```
|
|
|
|
Exactly one of `repl` (a template string) or `cb` (a callback) is
|
|
non-`NULL`; this two-parameter shape exists because C cannot express
|
|
Python's single polymorphic `repl` argument (string or callable) in one
|
|
slot (`docs/API.md` Section 3.8-3.9). `Pattern_subn` additionally reports
|
|
the number of substitutions made through `n`; `Pattern_sub` is the same
|
|
operation with that count discarded. `*out` is a freshly `malloc`'d,
|
|
`NUL`-terminated buffer the caller must `free`.
|
|
|
|
### Template substitution
|
|
|
|
Template syntax: `\g<name>`, `\g<N>`, `\N` (one or two digits), `\n`, `\t`,
|
|
`\\`, and any other `\X` as the literal character `X`.
|
|
|
|
```c
|
|
const char *pattern = "(\\w+)@(\\w+)";
|
|
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
|
|
const char *subject = "user@host";
|
|
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
|
|
char *out; size_t outlen; int n;
|
|
Pattern_subn(pat, in, "\\2@\\1", NULL, NULL, 0, &out, &outlen, &n);
|
|
/* use out/outlen, then: */
|
|
free(out);
|
|
```
|
|
|
|
Output, verified:
|
|
|
|
```
|
|
=== ascii_sub_template ===
|
|
result='host@user' n=1
|
|
```
|
|
|
|
### Callback substitution
|
|
|
|
```c
|
|
static void upper_cb(void *ctx, const Match *m_const, char **out, size_t *outlen) {
|
|
Match *m = (Match *)m_const;
|
|
const char *g; size_t glen;
|
|
Match_group(m, NULL, 0, &g, &glen);
|
|
char *buf = malloc(glen);
|
|
for (size_t i = 0; i < glen; i++) {
|
|
char c = g[i];
|
|
buf[i] = (c >= 'a' && c <= 'z') ? (char)(c - 32) : c;
|
|
}
|
|
*out = buf; *outlen = glen; /* ownership passes to Pattern_sub */
|
|
}
|
|
|
|
const char *pattern = "\\w+";
|
|
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
|
|
const char *subject = "shout this loudly";
|
|
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
|
|
char *out; size_t outlen;
|
|
Pattern_sub(pat, in, NULL, upper_cb, NULL, 0, &out, &outlen);
|
|
free(out);
|
|
```
|
|
|
|
Output, verified:
|
|
|
|
```
|
|
=== ascii_sub_callback ===
|
|
result='SHOUT THIS LOUDLY'
|
|
```
|
|
|
|
The callback is expected to `malloc` its replacement and hand ownership of
|
|
it off through `*out`/`*outlen`; `Pattern_sub`/`Pattern_subn` frees it
|
|
immediately after copying its content into the final result buffer
|
|
(`docs/API.md` Section 1.5).
|
|
|
|
## 11. Flags
|
|
|
|
```c
|
|
Pattern *pat = re_compile("hello", 5, ASCII | IGNORECASE, &err);
|
|
```
|
|
|
|
```c
|
|
const char *pattern = "hello";
|
|
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII | IGNORECASE, &err);
|
|
const char *subject = "HELLO world";
|
|
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
|
|
Match m; memset(&m, 0, sizeof m);
|
|
Pattern_search(pat, in, 0, -1, &m); /* rc=1 */
|
|
```
|
|
|
|
```c
|
|
const char *pattern = "^line";
|
|
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII | MULTILINE, &err);
|
|
const char *subject = "first\nline two\nline three";
|
|
/* Pattern_finditer finds 2 matches: MULTILINE makes ^ match after every \n too */
|
|
```
|
|
|
|
```c
|
|
const char *pattern = "a.b";
|
|
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII | DOTALL, &err);
|
|
const char *subject = "a\nb";
|
|
/* rc=1: DOTALL makes . match \n too */
|
|
```
|
|
|
|
```c
|
|
const char *pattern = "\\d+ # a number\n\\s+ \\w+ # then a word";
|
|
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII | VERBOSE, &err);
|
|
const char *subject = "42 answer";
|
|
/* rc=1: VERBOSE strips the unescaped whitespace and # comments before parsing */
|
|
```
|
|
|
|
Output, verified for all four:
|
|
|
|
```
|
|
=== flags_example ===
|
|
IGNORECASE rc=1
|
|
match 0: 'line'
|
|
match 1: 'line'
|
|
MULTILINE matches=2
|
|
DOTALL rc=1
|
|
VERBOSE rc=1
|
|
```
|
|
|
|
`UNICODE`, `LOCALE`, and `DEBUG` are accepted (for source compatibility with
|
|
Python) and combine with any of the above via `|`, but currently have no
|
|
observable effect (`docs/API.md` Section 2 records exactly why for each
|
|
one, in particular why `LOCALE` is a no-op rather than the bug it used to
|
|
be, caught and fixed by this project's test suite).
|
|
|
|
## 12. Matching against a file
|
|
|
|
```c
|
|
const char *pattern = "line \\w+";
|
|
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
|
|
Input *in = Input_from_file("data.txt", &err);
|
|
IterCtx ctx = { 0 };
|
|
int n = Pattern_finditer(pat, in, 0, -1, print_match_cb, &ctx);
|
|
Input_free(in);
|
|
```
|
|
|
|
Given a file containing `line one\nline two has data\nline three\n`, output,
|
|
verified:
|
|
|
|
```
|
|
=== file_input ===
|
|
Input_from_file rc=0x5b30fbc44e20
|
|
match 0: 'line one'
|
|
match 1: 'line two'
|
|
match 2: 'line three'
|
|
total matches=3
|
|
```
|
|
|
|
As Section 4 explains, `Input_from_file` mmaps a regular file rather than
|
|
copying it; nothing about matching against it differs from matching against
|
|
an `Input_from_buffer`-wrapped in-memory buffer, the mode (`BINARY`/`ASCII`/
|
|
`UTF8`) and every function above work identically regardless of which
|
|
`Input_from_*` constructor produced the `Input`. See README.md "Memory
|
|
footprint" for the measured memory cost of matching a large file this way,
|
|
and for a documented, deliberately reverted attempt at bounding it that
|
|
traded a fast allocation failure for an effectively unbounded hang (worth
|
|
reading before assuming a size cap is a safe thing to add here).
|
|
|
|
## 13. Module level convenience functions
|
|
|
|
```c
|
|
int re_search(const char *pattern, size_t len, int flags, Input *string, Match *out);
|
|
```
|
|
|
|
`re_match`/`re_fullmatch`/`re_search`/`re_finditer`/`re_findall`/`re_split`/
|
|
`re_sub`/`re_subn` compile `pattern` through an internal cache (up to 512
|
|
entries, keyed by `(pattern, len, flags)`, cleared entirely on overflow)
|
|
and then call the matching `Pattern_` function with `pos=0`, `endpos=-1`,
|
|
mirroring how CPython itself implements `re.match` as
|
|
`_compile(pattern, flags).match(string)`.
|
|
|
|
```c
|
|
const char *pattern = "\\d+";
|
|
const char *subject = "abc123def";
|
|
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
|
|
Match m; memset(&m, 0, sizeof m);
|
|
int r = re_search(pattern, strlen(pattern), ASCII, in, &m);
|
|
```
|
|
|
|
Output, verified:
|
|
|
|
```
|
|
=== module_level ===
|
|
re_search rc=1
|
|
matched='123'
|
|
```
|
|
|
|
A compile error inside these functions is reported only as a `-1` return,
|
|
with no `PatternError` available; use `re_compile` plus a `Pattern_`
|
|
function directly whenever a compile error needs to be diagnosed, or
|
|
whenever the same pattern is used more than once (precompiling once and
|
|
reusing the `Pattern` avoids repeated cache lookups, matching idiomatic
|
|
Python's own preference for `re.compile` in a loop over calling the module
|
|
level function repeatedly). `re_purge()` clears the cache, matching
|
|
`re.purge()`. This cache is shared, mutable, process-wide state with no
|
|
internal locking; do not call a `re_`-prefixed function from more than one
|
|
thread without external synchronization (`Pattern_`-prefixed functions on
|
|
an already-compiled `Pattern` have no such restriction, since nothing here
|
|
mutates a compiled `Pattern`).
|
|
|
|
## 14. Escaping literal text: `re_escape`
|
|
|
|
```c
|
|
void re_escape(const char *in, size_t len, char **out, size_t *outlen);
|
|
```
|
|
|
|
```c
|
|
const char *s = "3.14 (pi)";
|
|
char *out; size_t outlen;
|
|
re_escape(s, strlen(s), &out, &outlen);
|
|
free(out);
|
|
```
|
|
|
|
Output, verified:
|
|
|
|
```
|
|
=== escape_example ===
|
|
escaped='3\.14\ \(pi\)'
|
|
```
|
|
|
|
Matches `re.escape` exactly, including the narrowed escaped-character set
|
|
CPython adopted in 3.7 (only characters that are actually special in a
|
|
regex are escaped; other non-alphanumeric bytes are passed through
|
|
unescaped).
|
|
|
|
## 15. Error handling
|
|
|
|
Two kinds of failure exist in this API: a compile-time `PatternError`
|
|
(`re_compile`, `Input_from_file`) and a matching-time `-1` return with no
|
|
further detail (`docs/API.md` Section 3). Always check the return value of
|
|
a compile call before using the `Pattern` it was supposed to produce.
|
|
|
|
```c
|
|
const char *pattern = "(unclosed";
|
|
PatternError err; memset(&err, 0, sizeof err);
|
|
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
|
|
if (!pat) {
|
|
fprintf(stderr, "pattern error at %lld: %s\n", (long long)err.pos, err.msg);
|
|
}
|
|
PatternError_free(&err);
|
|
```
|
|
|
|
Output, verified:
|
|
|
|
```
|
|
=== error_compile ===
|
|
pat=(nil) msg='missing ), unterminated subpattern' pos=9 lineno=1 colno=10
|
|
```
|
|
|
|
`err.msg`/`err.pattern` are heap allocated by whichever call filled the
|
|
struct in; zero-initialize a `PatternError` before use and call
|
|
`PatternError_free` on it once done reading it, whether or not the call
|
|
that filled it in succeeded. A `NULL` `err` argument is always safe to pass
|
|
to any function that accepts one, and simply skips error reporting.
|
|
|
|
A matching call (`Pattern_search`, and so on) returning `-1` means either
|
|
invalid UTF-8 in the subject against a `UTF8`-mode `Pattern`, or the
|
|
backtracking depth limit was reached (`MAX_DEPTH` in `regexx.c`, currently
|
|
60000, not adjustable without editing `regexx.c` and rebuilding); no
|
|
further detail is available through the return value alone in this build.
|
|
|
|
## 16. Pattern syntax reference
|
|
|
|
Literals; `.` (with `DOTALL` for it to match `\n`); character classes with
|
|
ranges and negation (`[abc]`, `[^abc]`, `[a-z]`); `\d \D \w \W \s \S`;
|
|
`\b \B`; anchors `^ $ \A \Z` (`^`/`$` also match at line boundaries with
|
|
`MULTILINE`); quantifiers `* + ? {m,n} {m,} {,n} {m}`, both greedy and lazy
|
|
(`*?`, `+?`, and so on) forms; possessive quantifiers `*+ ++ ?+ {m,n}+`;
|
|
groups `(...)`, non-capturing `(?:...)`, named `(?P<name>...)`; alternation
|
|
`|`; backreferences `\1`-`\99`, `(?P=name)`, `\g<name>`, `\g<N>`; lookahead
|
|
`(?=...)`/`(?!...)`; fixed-width lookbehind `(?<=...)`/`(?<!...)`; atomic
|
|
groups `(?>...)`; comments `(?#...)`; global inline flags `(?aiLmsux)` at
|
|
the very start of a pattern; escapes `\n \r \t \f \v \a`, `\0`-prefixed
|
|
octal escapes, `\xhh`, `\uxxxx`, `\Uxxxxxxxx`.
|
|
|
|
Rejected at compile time with a `PatternError` naming the construct, rather
|
|
than silently mis-parsed: conditional groups `(?(id)yes|no)`, scoped inline
|
|
flags `(?flags:...)` (only the global, start-of-pattern form is supported),
|
|
`\N{NAME}` named code points, variable-width lookbehind (matching CPython's
|
|
own restriction), and POSIX bracket-expression syntax `[[:alpha:]]` (not
|
|
part of Python `re`'s grammar at all, `concept.md` Section 14). See
|
|
`docs/API.md` Section 5-6 for the complete, exact list.
|
|
|
|
## 17. Performance and memory, in one paragraph
|
|
|
|
A pattern with no backreference, lookaround, or atomic group runs on the
|
|
Pike VM (`README.md` "The Pike VM"), which is genuinely linear time,
|
|
including on adversarial nested-quantifier shapes like `(a+)+b` that would
|
|
invite catastrophic backtracking elsewhere, with no atomic group needed;
|
|
it has a literal prefilter (`README.md` "The Pike VM") but still no lazy
|
|
DFA state caching or allocation pooling, so an ordinary pattern can
|
|
currently measure slower in absolute terms than before this engine
|
|
existed, narrower than before the prefilter but a real, documented
|
|
trade-off still, not an oversight. Everything else
|
|
(a backreference, lookaround, or atomic group anywhere) runs on the
|
|
recursive backtracking engine instead: `match`/`fullmatch` (a single
|
|
anchored attempt) cost time proportional to that attempt; `search`/
|
|
`finditer`/`split`/`sub` are linear time, not quadratic, for the common
|
|
case there (no backreference, every quantifier's body a single
|
|
character/class/`.`), at a measured memory cost of roughly 8.4x the input
|
|
length for a pattern using a simple-atom quantifier, and a pattern with a
|
|
backreference can still be worse than linear, the same way CPython's own
|
|
`re` can be on the same patterns. `README.md` "Implementation status",
|
|
"The Pike VM", and "Memory footprint" have the full measured account,
|
|
including one fix that was tried, measured, and deliberately reverted;
|
|
read it before assuming a
|
|
given pattern's cost, rather than assuming linear time and bounded memory
|
|
apply unconditionally.
|
|
|
|
## 18. Complete example: `rxgrep`
|
|
|
|
`examples/rxgrep.c` is a complete, working grep-like program built on this
|
|
API, exercising all three data modes and both the matching and substitution
|
|
API against real files and pipes:
|
|
|
|
```sh
|
|
./rxgrep -in 'hello' file.txt # case-insensitive, line numbers
|
|
./rxgrep -m utf8 -o '\w+' file.txt # print every UTF-8 word, one per line
|
|
./rxgrep -c 'error' log.txt # count matching lines
|
|
./rxgrep -m binary 'a.c' data.bin # match raw bytes, embedded NUL included
|
|
./rxgrep -m utf8 --sub 'REDACTED' '\d{3}-\d{4}' file.txt
|
|
```
|
|
|
|
Reading its source alongside this document is a reasonable next step once
|
|
the examples above are familiar: it shows every function here used together
|
|
in one program, including the parts this document simplified away for
|
|
clarity (option parsing, line splitting, output formatting).
|
|
|
|
`examples/` also has six smaller, single-purpose programs, each isolating
|
|
one distinct feature this section-by-section walkthrough only touches
|
|
briefly: `binary_scan.c` (raw byte-range classes including embedded `NUL`),
|
|
`utf8_scripts.c` (`\w` across non-Latin scripts, code point versus byte
|
|
offsets), `ascii_logparse.c` (named groups against structured log text),
|
|
`redos_atomic.c` (the textbook `(a+)+b` ReDoS shape, now fixed
|
|
automatically by the Pike VM with no atomic group needed, contrasted with
|
|
a backreference-forced variant where an atomic group is still necessary),
|
|
`empty_match_rule.c` (the undocumented CPython empty-match retry rule,
|
|
verified against a real interpreter), `large_file_search.c` (`mmap`-backed
|
|
file input at 100MB, with elapsed time and peak memory printed), and
|
|
`bench_vs_posix.c` (a direct, honestly-reported timing comparison against
|
|
the C standard library's own `<regex.h>` on six scenarios: glibc wins the
|
|
three ordinary ones by a wide margin, and the one adversarial pattern
|
|
shape inverts entirely, now beating glibc outright with no atomic group).
|
|
See [`examples/README.md`](examples/README.md) for the complete list with
|
|
what each one demonstrates; `make examples` builds all of them.
|