Add USAGE.md, a task-oriented guide with compiled and verified examples
docs/API.md is the exhaustive reference; concept.md is the design rationale. Neither is the right document for "I have never used this library before, show me working code," so this adds one that is: compiling a pattern, building an Input from a buffer or a file, match/fullmatch/search, named groups, finditer/findall, split, sub/subn with both template and callback replacement, all six flags, error handling, the module level convenience functions and their cache, re_escape, and a pattern syntax quick reference, each with BINARY, ASCII, or UTF8 examples as appropriate (including the code-point-versus-byte-offset distinction UTF8 mode introduces). Every example was written as a real, compiled program linked against the current regexx.c and run; the output shown in the document is that program's actual output, not a hand-written guess, the same verification standard already used for README.md's own measured numbers. README.md gets a one-line pointer to it alongside the existing pointer to docs/API.md. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
This commit is contained in:
@@ -13,7 +13,9 @@ what is and is not carried over from Python's `re` and from POSIX's native
|
||||
implementation that exists today at a glance; [`docs/API.md`](docs/API.md) is
|
||||
the exhaustive reference (every type, every flag, every function's exact
|
||||
return-value and memory-ownership convention, checked against a real CPython
|
||||
interpreter, not against memory).
|
||||
interpreter, not against memory); [`USAGE.md`](USAGE.md) is the task-oriented
|
||||
guide (compiled, run, and verified examples covering every operation across
|
||||
all three data modes, from "nothing" to "working code").
|
||||
|
||||
## Implementation status
|
||||
|
||||
|
||||
@@ -0,0 +1,721 @@
|
||||
# regexx usage guide
|
||||
|
||||
This document is task oriented: how to build a pattern, run it against data
|
||||
in each of the three supported modes, and read the result back out, with
|
||||
complete, compiled, and run examples. [`docs/API.md`](docs/API.md) is the
|
||||
exhaustive reference for every function's exact return value and memory
|
||||
ownership rule; this document exists to get from "nothing" to "working code"
|
||||
without reading that whole reference first. [`concept.md`](concept.md)
|
||||
records the design rationale for anyone who wants to know why a given
|
||||
choice was made rather than only what it is.
|
||||
|
||||
Every example below was compiled against the current `regexx.c` and its
|
||||
actual printed output is shown, not a hand-written guess at what it would
|
||||
print.
|
||||
|
||||
## 1. Building and linking
|
||||
|
||||
```sh
|
||||
make # builds libregexx.a and the rxgrep example
|
||||
```
|
||||
|
||||
A program using the library needs `regexx.h` on its include path and
|
||||
`libregexx.a` (or `regexx.c` compiled directly into the program, since the
|
||||
library is a single file) linked in:
|
||||
|
||||
```sh
|
||||
cc -std=c11 -I/path/to/regexx -o myprog myprog.c /path/to/regexx/libregexx.a
|
||||
```
|
||||
|
||||
No third-party dependency is required; `regexx.c` uses only the C standard
|
||||
library plus, for `UTF8` mode's Unicode classification and `Input_from_file`,
|
||||
POSIX (`wctype.h`, `locale.h`, `mmap`/`open`/`fstat`, `concept.md` Section 8).
|
||||
|
||||
## 2. The three data modes
|
||||
|
||||
Exactly one of `BINARY`, `ASCII`, or `UTF8` is passed in `flags` to
|
||||
`re_compile` (or any `re_`-prefixed module level function); it selects how
|
||||
the subject's bytes are interpreted, not how the pattern text itself is
|
||||
encoded (the pattern is always read as ASCII/UTF-8 source text regardless of
|
||||
this choice).
|
||||
|
||||
| Mode | Subject interpretation | `Match_start`/`Match_end` unit |
|
||||
|---|---|---|
|
||||
| `BINARY` | Raw bytes, `0`-`255` all valid, embedded `NUL` is an ordinary byte, no decoding at all. | Byte offset |
|
||||
| `ASCII` | Bytes, but `\d`/`\w`/`\s` and `IGNORECASE` restrict themselves to the ASCII subset (the default when neither `BINARY` nor `UTF8` is given, `docs/API.md` Section 2). | Byte offset |
|
||||
| `UTF8` | Decoded as UTF-8 into Unicode code points before matching; `\d`/`\w`/`\s` and case folding consult the platform's Unicode tables via `wctype.h`. | Code point index (use `Match_start_byte`/`_end_byte` for a byte offset into the raw buffer, Section 6 below) |
|
||||
|
||||
Passing invalid UTF-8 to a `UTF8`-mode pattern's matching call returns `-1`
|
||||
(the general error return, `docs/API.md` Section 3); `BINARY` and `ASCII`
|
||||
mode never reject a subject on this basis, since neither one decodes it.
|
||||
|
||||
Choosing a mode is a decision made once per `Pattern` (it is baked into
|
||||
`flags` at `re_compile` time), not per matching call: a `Pattern` compiled
|
||||
with `UTF8` decodes every `Input` it is later run against as UTF-8, and a
|
||||
`Pattern` compiled with `BINARY` never does, regardless of what the actual
|
||||
bytes happen to look like.
|
||||
|
||||
## 3. Compiling a pattern
|
||||
|
||||
```c
|
||||
#include "regexx.h"
|
||||
#include <string.h>
|
||||
|
||||
PatternError err; memset(&err, 0, sizeof err);
|
||||
Pattern *pat = re_compile("(\\w+)@(\\w+)\\.(\\w+)", strlen("(\\w+)@(\\w+)\\.(\\w+)"), ASCII, &err);
|
||||
if (!pat) {
|
||||
fprintf(stderr, "pattern error at %lld: %s\n", (long long)err.pos, err.msg);
|
||||
PatternError_free(&err);
|
||||
/* handle failure */
|
||||
}
|
||||
```
|
||||
|
||||
- `pattern`/`len`: the pattern source; it need not be `NUL`-terminated, `len`
|
||||
is authoritative.
|
||||
- `flags`: the data mode (Section 2) combined with any of `IGNORECASE`,
|
||||
`MULTILINE`, `DOTALL`, `VERBOSE`, `ASCII` (also usable as a flag inside
|
||||
`UTF8` mode to force ASCII-only `\d`/`\w`/`\s`), `UNICODE`, `LOCALE`,
|
||||
`DEBUG` (the last three accepted for source compatibility with Python,
|
||||
currently no-ops, `docs/API.md` Section 2).
|
||||
- `err`: optional (`NULL` is safe); filled in on a syntax error or an
|
||||
unsupported construct (`docs/API.md` Section 5 lists every construct this
|
||||
build rejects at compile time, such as conditional groups and scoped
|
||||
inline flags). Call `PatternError_free` on it once read, whether or not
|
||||
compilation succeeded.
|
||||
|
||||
A compiled `Pattern` is freed with `Pattern_free`, once, after every `Match`
|
||||
produced from it has been released (`Match.re` borrows from the `Pattern`,
|
||||
so freeing the `Pattern` first and a `Match` from it afterward is a use
|
||||
after free, `docs/API.md` Section 1.2).
|
||||
|
||||
## 4. Building an `Input`
|
||||
|
||||
Every matching call takes an `Input`, not a raw pointer, so the library has
|
||||
one place to resolve "where are the bytes" regardless of whether they came
|
||||
from memory the caller already has or a file on disk.
|
||||
|
||||
```c
|
||||
Input *Input_from_buffer(const uint8_t *buf, size_t len); /* wraps, does not copy or own */
|
||||
Input *Input_from_file(const char *path, PatternError *err); /* mmaps or reads a whole file */
|
||||
void Input_free(Input *in);
|
||||
```
|
||||
|
||||
`Input_from_buffer` wraps an existing buffer (a string literal, a `malloc`'d
|
||||
region, a memory-mapped region the caller manages itself) without copying
|
||||
it; the buffer must outlive the `Input` and every `Match` produced from it,
|
||||
since `Match_group` returns pointers directly into it (`docs/API.md` Section
|
||||
1.3). `Input_free` on an `Input_from_buffer` result never frees the wrapped
|
||||
buffer.
|
||||
|
||||
`Input_from_file` maps a regular file read-only with `mmap()` instead of
|
||||
copying it, so the pages stay reclaimable under memory pressure
|
||||
(`README.md` "Memory footprint" has the measured numbers); a non-seekable
|
||||
source (a pipe, a FIFO, process substitution, `/dev/stdin`) or an `mmap()`
|
||||
failure falls back to reading it incrementally into an owned buffer instead.
|
||||
Either way `Input_free` releases what it acquired (`munmap` or `free`, as
|
||||
appropriate) automatically; the caller never needs to know which path was
|
||||
taken.
|
||||
|
||||
```c
|
||||
PatternError err; memset(&err, 0, sizeof err);
|
||||
Input *in = Input_from_file("data.txt", &err);
|
||||
if (!in) { fprintf(stderr, "cannot read file: %s\n", err.msg); PatternError_free(&err); }
|
||||
```
|
||||
|
||||
Output, verified:
|
||||
|
||||
```
|
||||
=== error_file ===
|
||||
in=(nil) msg='No such file or directory'
|
||||
```
|
||||
|
||||
## 5. Matching: `match`, `fullmatch`, `search`
|
||||
|
||||
```c
|
||||
int Pattern_match(Pattern *self, Input *string, int64_t pos, int64_t endpos, Match *out);
|
||||
int Pattern_fullmatch(Pattern *self, Input *string, int64_t pos, int64_t endpos, Match *out);
|
||||
int Pattern_search(Pattern *self, Input *string, int64_t pos, int64_t endpos, Match *out);
|
||||
```
|
||||
|
||||
`pos`/`endpos` bound the attempt (pass `0`/`-1` for "the whole input", the
|
||||
same defaults `re.Pattern.match`/`.search`/etc. use); `match` and
|
||||
`fullmatch` are anchored at `pos` (`fullmatch` additionally requires
|
||||
reaching exactly `endpos`), `search` tries every start position from `pos`
|
||||
to `endpos` and reports the first that admits a match. All three return `1`
|
||||
matched, `0` no match, `-1` error; `out` must point at a zero-initialized
|
||||
`Match` (never allocated by the library itself).
|
||||
|
||||
```c
|
||||
const char *pattern = "(\\w+)@(\\w+)\\.(\\w+)";
|
||||
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
|
||||
const char *subject = "contact: user@example.com today";
|
||||
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
|
||||
|
||||
Match m; memset(&m, 0, sizeof m);
|
||||
if (Pattern_search(pat, in, 0, -1, &m) == 1) {
|
||||
const char *g; size_t glen;
|
||||
Match_group(&m, NULL, 0, &g, &glen); /* whole match: group 0 */
|
||||
Match_group(&m, NULL, 1, &g, &glen); /* first capturing group */
|
||||
int64_t s, e;
|
||||
Match_span(&m, 0, &s, &e);
|
||||
Match_free(&m);
|
||||
}
|
||||
Input_free(in);
|
||||
Pattern_free(pat);
|
||||
```
|
||||
|
||||
Output, verified:
|
||||
|
||||
```
|
||||
=== ascii_basic ===
|
||||
search rc=1
|
||||
group0='user@example.com'
|
||||
group1='user'
|
||||
group2='example'
|
||||
group3='com'
|
||||
span0=[9,25)
|
||||
```
|
||||
|
||||
### Named groups
|
||||
|
||||
```c
|
||||
const char *pattern = "(?P<user>\\w+)@(?P<host>\\w+)";
|
||||
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
|
||||
const char *subject = "user@host";
|
||||
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
|
||||
Match m; memset(&m, 0, sizeof m);
|
||||
if (Pattern_search(pat, in, 0, -1, &m) == 1) {
|
||||
const char *g; size_t glen;
|
||||
Match_group(&m, "user", 0, &g, &glen); /* name_or_null non-NULL: index is ignored */
|
||||
Match_group(&m, "host", 0, &g, &glen);
|
||||
int idx = Pattern_groupindex_lookup(pat, "host"); /* 1-based group number, or -1 */
|
||||
Match_free(&m);
|
||||
}
|
||||
```
|
||||
|
||||
Output, verified:
|
||||
|
||||
```
|
||||
=== ascii_named ===
|
||||
user='user'
|
||||
host='host'
|
||||
groupindex(host)=2
|
||||
```
|
||||
|
||||
`Match_group` returns `1` and a borrowed pointer into the `Input`'s buffer
|
||||
when the group participated, `0` and `*out = NULL` when the group exists in
|
||||
the pattern but did not participate (Python's `None`), and `-1` when no
|
||||
such group exists at all, by index or by name; this three-way distinction
|
||||
is the only place in the API that tells "no such group" apart from "this
|
||||
group matched nothing" (`Match_start`/`Match_end` collapse both cases to
|
||||
`-1`, `docs/API.md` Section 3.23-3.29).
|
||||
|
||||
## 6. Code points versus bytes in `UTF8` mode
|
||||
|
||||
In `UTF8` mode, `Match_start`/`Match_end`/`Match_span` report a code point
|
||||
index (so, for example, a 2-byte UTF-8 character like `é` counts as one
|
||||
unit, the same way Python's `str` indexing does), while
|
||||
`Match_start_byte`/`Match_end_byte`/`Match_span_byte` always report a byte
|
||||
offset into the raw `Input` buffer, in every mode. Use the `_byte` variants
|
||||
whenever the result is used to slice or seek into the raw buffer or file
|
||||
directly (`Match_group` already does this translation internally, so it
|
||||
needs neither variant directly; a caller that needs to, for example, report
|
||||
a byte offset to an external tool that has no notion of code points is the
|
||||
usual reason to reach for `_byte` directly).
|
||||
|
||||
```c
|
||||
const char *pattern = "\\w+";
|
||||
Pattern *pat = re_compile(pattern, strlen(pattern), UTF8, &err);
|
||||
const char *subject = "caf\xc3\xa9 today"; /* "café today", é is 2 UTF-8 bytes */
|
||||
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
|
||||
Match m; memset(&m, 0, sizeof m);
|
||||
if (Pattern_search(pat, in, 0, -1, &m) == 1) {
|
||||
int64_t cs, ce, bs, be;
|
||||
Match_span(&m, 0, &cs, &ce); /* code point units */
|
||||
Match_span_byte(&m, 0, &bs, &be); /* byte units */
|
||||
Match_free(&m);
|
||||
}
|
||||
```
|
||||
|
||||
Output, verified:
|
||||
|
||||
```
|
||||
=== utf8_offsets ===
|
||||
search rc=1
|
||||
codepoint span=[0,4) byte span=[0,5)
|
||||
group bytes: 'café' (5 bytes)
|
||||
```
|
||||
|
||||
`café` is 4 code points (`c`, `a`, `f`, `é`) but 5 bytes (`é` alone is 2
|
||||
bytes in UTF-8), which is exactly what the two spans above report.
|
||||
|
||||
### Iterating UTF-8 text
|
||||
|
||||
```c
|
||||
const char *pattern = "\\w+";
|
||||
Pattern *pat = re_compile(pattern, strlen(pattern), UTF8, &err);
|
||||
const char *subject = "na\xc3\xafve caf\xc3\xa9 r\xc3\xa9sum\xc3\xa9"; /* naïve café résumé */
|
||||
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
|
||||
Pattern_finditer(pat, in, 0, -1, print_match_cb, &ctx);
|
||||
```
|
||||
|
||||
Output, verified (`\w` under `UTF8` mode matches letters outside ASCII too,
|
||||
consulting the platform's Unicode tables, `docs/API.md` Section 2):
|
||||
|
||||
```
|
||||
=== utf8_finditer ===
|
||||
match 0: 'naïve'
|
||||
match 1: 'café'
|
||||
match 2: 'résumé'
|
||||
total matches=3
|
||||
```
|
||||
|
||||
## 7. Binary data
|
||||
|
||||
`BINARY` mode never decodes the subject; every byte `0`-`255` is a valid
|
||||
input value, including `0x00`, and a `.`/character class/literal matches by
|
||||
raw byte value.
|
||||
|
||||
```c
|
||||
const char *pattern = "a.c";
|
||||
Pattern *pat = re_compile(pattern, strlen(pattern), BINARY, &err);
|
||||
uint8_t subject[] = { 'x', 'a', 0x00, 'c', 'y' };
|
||||
Input *in = Input_from_buffer(subject, sizeof subject);
|
||||
Match m; memset(&m, 0, sizeof m);
|
||||
Pattern_search(pat, in, 0, -1, &m);
|
||||
```
|
||||
|
||||
Output, verified (`.` matches the embedded `NUL` byte at offset 2, since
|
||||
`DOTALL` is irrelevant here: `BINARY` mode has no notion of "newline" ending
|
||||
`.`'s match, byte `0x0a` is simply a byte like any other unless the pattern
|
||||
excludes it explicitly):
|
||||
|
||||
```
|
||||
=== binary_basic ===
|
||||
search rc=1
|
||||
matched 3 bytes: 61 00 63
|
||||
```
|
||||
|
||||
Byte ranges in a character class work the same way, by raw value:
|
||||
|
||||
```c
|
||||
const char *pattern = "[\\x00-\\x1f]+";
|
||||
Pattern *pat = re_compile(pattern, strlen(pattern), BINARY, &err);
|
||||
uint8_t subject[] = { 'A', 0x01, 0x02, 0x1f, 'B' };
|
||||
```
|
||||
|
||||
Output, verified:
|
||||
|
||||
```
|
||||
=== binary_class ===
|
||||
search rc=1
|
||||
span=[1,4)
|
||||
```
|
||||
|
||||
`\d`/`\w`/`\s` in `BINARY` mode classify the same way they do in `ASCII`
|
||||
mode (ASCII letters/digits/whitespace only; there is no "Unicode binary"
|
||||
notion to fall back to, since `BINARY` mode has no decoding step at all).
|
||||
|
||||
## 8. Finding all matches: `finditer`/`findall`
|
||||
|
||||
```c
|
||||
int Pattern_finditer(Pattern *self, Input *string, int64_t pos, int64_t endpos, MatchIterCb cb, void *ctx);
|
||||
```
|
||||
|
||||
`cb(ctx, m)` fires once per non-overlapping match, left to right; the
|
||||
`Match` passed in is only valid for the duration of the call (freed
|
||||
immediately after `cb` returns, so do not retain the pointer past it).
|
||||
`Pattern_findall` is the same call under a second name, for parity with
|
||||
`re.findall`; this build always hands the callback a full `Match` rather
|
||||
than collapsing it to "just the string" or a tuple the way CPython's
|
||||
`findall` does at the Python-object level, since there is no C object model
|
||||
to collapse into (`docs/API.md` Section 3.5-3.6).
|
||||
|
||||
```c
|
||||
typedef struct { int n; } IterCtx;
|
||||
static void print_match_cb(void *ctx, const Match *m_const) {
|
||||
IterCtx *c = ctx;
|
||||
Match *m = (Match *)m_const;
|
||||
const char *g; size_t glen;
|
||||
Match_group(m, NULL, 0, &g, &glen);
|
||||
printf(" match %d: '%.*s'\n", c->n, (int)glen, g);
|
||||
c->n++;
|
||||
}
|
||||
|
||||
const char *pattern = "\\d+";
|
||||
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
|
||||
const char *subject = "order 12 has 345 items, batch 6";
|
||||
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
|
||||
IterCtx ctx = { 0 };
|
||||
int n = Pattern_finditer(pat, in, 0, -1, print_match_cb, &ctx);
|
||||
```
|
||||
|
||||
Output, verified:
|
||||
|
||||
```
|
||||
=== ascii_finditer ===
|
||||
match 0: '12'
|
||||
match 1: '345'
|
||||
match 2: '6'
|
||||
total matches=3
|
||||
```
|
||||
|
||||
`finditer`/`findall`/`split`/`sub` all share one memoization table and one
|
||||
set of precomputed skip-ahead tables across the whole scan (built once per
|
||||
call, not once per match), which is what makes them run in linear time for
|
||||
the common case rather than redoing quadratic work; see README.md
|
||||
"Implementation status" and "Memory footprint" for the measured numbers and
|
||||
what that costs in memory.
|
||||
|
||||
## 9. Splitting: `Pattern_split`
|
||||
|
||||
```c
|
||||
int Pattern_split(Pattern *self, Input *string, int maxsplit, MatchIterCb cb, void *ctx);
|
||||
```
|
||||
|
||||
`cb` fires once per element of the list `re.split()` would return, in
|
||||
order; each element's text is read the same way as any other match, via
|
||||
`Match_group(m, NULL, 0, &out, &outlen)`, and a `0` return from that call
|
||||
means the element is Python's `None` (an unparticipated capturing group
|
||||
between two matches, `docs/API.md` Section 3.7). `maxsplit` matches
|
||||
`re.split`'s parameter (`0` unlimited).
|
||||
|
||||
```c
|
||||
const char *pattern = "\\s*,\\s*";
|
||||
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
|
||||
const char *subject = "red, green,blue , yellow";
|
||||
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
|
||||
IterCtx ctx = { 0 };
|
||||
int n = Pattern_split(pat, in, 0, print_split_cb, &ctx);
|
||||
```
|
||||
|
||||
Output, verified:
|
||||
|
||||
```
|
||||
=== ascii_split ===
|
||||
elem 0: 'red'
|
||||
elem 1: 'green'
|
||||
elem 2: 'blue'
|
||||
elem 3: 'yellow'
|
||||
splits=3
|
||||
```
|
||||
|
||||
## 10. Substitution: `Pattern_sub`/`Pattern_subn`
|
||||
|
||||
```c
|
||||
int Pattern_sub(Pattern *self, Input *string, const char *repl, MatchSubCb cb, void *ctx, int count, char **out, size_t *outlen);
|
||||
int Pattern_subn(Pattern *self, Input *string, const char *repl, MatchSubCb cb, void *ctx, int count, char **out, size_t *outlen, int *n);
|
||||
```
|
||||
|
||||
Exactly one of `repl` (a template string) or `cb` (a callback) is
|
||||
non-`NULL`; this two-parameter shape exists because C cannot express
|
||||
Python's single polymorphic `repl` argument (string or callable) in one
|
||||
slot (`docs/API.md` Section 3.8-3.9). `Pattern_subn` additionally reports
|
||||
the number of substitutions made through `n`; `Pattern_sub` is the same
|
||||
operation with that count discarded. `*out` is a freshly `malloc`'d,
|
||||
`NUL`-terminated buffer the caller must `free`.
|
||||
|
||||
### Template substitution
|
||||
|
||||
Template syntax: `\g<name>`, `\g<N>`, `\N` (one or two digits), `\n`, `\t`,
|
||||
`\\`, and any other `\X` as the literal character `X`.
|
||||
|
||||
```c
|
||||
const char *pattern = "(\\w+)@(\\w+)";
|
||||
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
|
||||
const char *subject = "user@host";
|
||||
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
|
||||
char *out; size_t outlen; int n;
|
||||
Pattern_subn(pat, in, "\\2@\\1", NULL, NULL, 0, &out, &outlen, &n);
|
||||
/* use out/outlen, then: */
|
||||
free(out);
|
||||
```
|
||||
|
||||
Output, verified:
|
||||
|
||||
```
|
||||
=== ascii_sub_template ===
|
||||
result='host@user' n=1
|
||||
```
|
||||
|
||||
### Callback substitution
|
||||
|
||||
```c
|
||||
static void upper_cb(void *ctx, const Match *m_const, char **out, size_t *outlen) {
|
||||
Match *m = (Match *)m_const;
|
||||
const char *g; size_t glen;
|
||||
Match_group(m, NULL, 0, &g, &glen);
|
||||
char *buf = malloc(glen);
|
||||
for (size_t i = 0; i < glen; i++) {
|
||||
char c = g[i];
|
||||
buf[i] = (c >= 'a' && c <= 'z') ? (char)(c - 32) : c;
|
||||
}
|
||||
*out = buf; *outlen = glen; /* ownership passes to Pattern_sub */
|
||||
}
|
||||
|
||||
const char *pattern = "\\w+";
|
||||
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
|
||||
const char *subject = "shout this loudly";
|
||||
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
|
||||
char *out; size_t outlen;
|
||||
Pattern_sub(pat, in, NULL, upper_cb, NULL, 0, &out, &outlen);
|
||||
free(out);
|
||||
```
|
||||
|
||||
Output, verified:
|
||||
|
||||
```
|
||||
=== ascii_sub_callback ===
|
||||
result='SHOUT THIS LOUDLY'
|
||||
```
|
||||
|
||||
The callback is expected to `malloc` its replacement and hand ownership of
|
||||
it off through `*out`/`*outlen`; `Pattern_sub`/`Pattern_subn` frees it
|
||||
immediately after copying its content into the final result buffer
|
||||
(`docs/API.md` Section 1.5).
|
||||
|
||||
## 11. Flags
|
||||
|
||||
```c
|
||||
Pattern *pat = re_compile("hello", 5, ASCII | IGNORECASE, &err);
|
||||
```
|
||||
|
||||
```c
|
||||
const char *pattern = "hello";
|
||||
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII | IGNORECASE, &err);
|
||||
const char *subject = "HELLO world";
|
||||
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
|
||||
Match m; memset(&m, 0, sizeof m);
|
||||
Pattern_search(pat, in, 0, -1, &m); /* rc=1 */
|
||||
```
|
||||
|
||||
```c
|
||||
const char *pattern = "^line";
|
||||
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII | MULTILINE, &err);
|
||||
const char *subject = "first\nline two\nline three";
|
||||
/* Pattern_finditer finds 2 matches: MULTILINE makes ^ match after every \n too */
|
||||
```
|
||||
|
||||
```c
|
||||
const char *pattern = "a.b";
|
||||
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII | DOTALL, &err);
|
||||
const char *subject = "a\nb";
|
||||
/* rc=1: DOTALL makes . match \n too */
|
||||
```
|
||||
|
||||
```c
|
||||
const char *pattern = "\\d+ # a number\n\\s+ \\w+ # then a word";
|
||||
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII | VERBOSE, &err);
|
||||
const char *subject = "42 answer";
|
||||
/* rc=1: VERBOSE strips the unescaped whitespace and # comments before parsing */
|
||||
```
|
||||
|
||||
Output, verified for all four:
|
||||
|
||||
```
|
||||
=== flags_example ===
|
||||
IGNORECASE rc=1
|
||||
match 0: 'line'
|
||||
match 1: 'line'
|
||||
MULTILINE matches=2
|
||||
DOTALL rc=1
|
||||
VERBOSE rc=1
|
||||
```
|
||||
|
||||
`UNICODE`, `LOCALE`, and `DEBUG` are accepted (for source compatibility with
|
||||
Python) and combine with any of the above via `|`, but currently have no
|
||||
observable effect (`docs/API.md` Section 2 records exactly why for each
|
||||
one, in particular why `LOCALE` is a no-op rather than the bug it used to
|
||||
be, caught and fixed by this project's test suite).
|
||||
|
||||
## 12. Matching against a file
|
||||
|
||||
```c
|
||||
const char *pattern = "line \\w+";
|
||||
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
|
||||
Input *in = Input_from_file("data.txt", &err);
|
||||
IterCtx ctx = { 0 };
|
||||
int n = Pattern_finditer(pat, in, 0, -1, print_match_cb, &ctx);
|
||||
Input_free(in);
|
||||
```
|
||||
|
||||
Given a file containing `line one\nline two has data\nline three\n`, output,
|
||||
verified:
|
||||
|
||||
```
|
||||
=== file_input ===
|
||||
Input_from_file rc=0x5b30fbc44e20
|
||||
match 0: 'line one'
|
||||
match 1: 'line two'
|
||||
match 2: 'line three'
|
||||
total matches=3
|
||||
```
|
||||
|
||||
As Section 4 explains, `Input_from_file` mmaps a regular file rather than
|
||||
copying it; nothing about matching against it differs from matching against
|
||||
an `Input_from_buffer`-wrapped in-memory buffer, the mode (`BINARY`/`ASCII`/
|
||||
`UTF8`) and every function above work identically regardless of which
|
||||
`Input_from_*` constructor produced the `Input`. See README.md "Memory
|
||||
footprint" for the measured memory cost of matching a large file this way,
|
||||
and for a documented, deliberately reverted attempt at bounding it that
|
||||
traded a fast allocation failure for an effectively unbounded hang (worth
|
||||
reading before assuming a size cap is a safe thing to add here).
|
||||
|
||||
## 13. Module level convenience functions
|
||||
|
||||
```c
|
||||
int re_search(const char *pattern, size_t len, int flags, Input *string, Match *out);
|
||||
```
|
||||
|
||||
`re_match`/`re_fullmatch`/`re_search`/`re_finditer`/`re_findall`/`re_split`/
|
||||
`re_sub`/`re_subn` compile `pattern` through an internal cache (up to 512
|
||||
entries, keyed by `(pattern, len, flags)`, cleared entirely on overflow)
|
||||
and then call the matching `Pattern_` function with `pos=0`, `endpos=-1`,
|
||||
mirroring how CPython itself implements `re.match` as
|
||||
`_compile(pattern, flags).match(string)`.
|
||||
|
||||
```c
|
||||
const char *pattern = "\\d+";
|
||||
const char *subject = "abc123def";
|
||||
Input *in = Input_from_buffer((const uint8_t *)subject, strlen(subject));
|
||||
Match m; memset(&m, 0, sizeof m);
|
||||
int r = re_search(pattern, strlen(pattern), ASCII, in, &m);
|
||||
```
|
||||
|
||||
Output, verified:
|
||||
|
||||
```
|
||||
=== module_level ===
|
||||
re_search rc=1
|
||||
matched='123'
|
||||
```
|
||||
|
||||
A compile error inside these functions is reported only as a `-1` return,
|
||||
with no `PatternError` available; use `re_compile` plus a `Pattern_`
|
||||
function directly whenever a compile error needs to be diagnosed, or
|
||||
whenever the same pattern is used more than once (precompiling once and
|
||||
reusing the `Pattern` avoids repeated cache lookups, matching idiomatic
|
||||
Python's own preference for `re.compile` in a loop over calling the module
|
||||
level function repeatedly). `re_purge()` clears the cache, matching
|
||||
`re.purge()`. This cache is shared, mutable, process-wide state with no
|
||||
internal locking; do not call a `re_`-prefixed function from more than one
|
||||
thread without external synchronization (`Pattern_`-prefixed functions on
|
||||
an already-compiled `Pattern` have no such restriction, since nothing here
|
||||
mutates a compiled `Pattern`).
|
||||
|
||||
## 14. Escaping literal text: `re_escape`
|
||||
|
||||
```c
|
||||
void re_escape(const char *in, size_t len, char **out, size_t *outlen);
|
||||
```
|
||||
|
||||
```c
|
||||
const char *s = "3.14 (pi)";
|
||||
char *out; size_t outlen;
|
||||
re_escape(s, strlen(s), &out, &outlen);
|
||||
free(out);
|
||||
```
|
||||
|
||||
Output, verified:
|
||||
|
||||
```
|
||||
=== escape_example ===
|
||||
escaped='3\.14\ \(pi\)'
|
||||
```
|
||||
|
||||
Matches `re.escape` exactly, including the narrowed escaped-character set
|
||||
CPython adopted in 3.7 (only characters that are actually special in a
|
||||
regex are escaped; other non-alphanumeric bytes are passed through
|
||||
unescaped).
|
||||
|
||||
## 15. Error handling
|
||||
|
||||
Two kinds of failure exist in this API: a compile-time `PatternError`
|
||||
(`re_compile`, `Input_from_file`) and a matching-time `-1` return with no
|
||||
further detail (`docs/API.md` Section 3). Always check the return value of
|
||||
a compile call before using the `Pattern` it was supposed to produce.
|
||||
|
||||
```c
|
||||
const char *pattern = "(unclosed";
|
||||
PatternError err; memset(&err, 0, sizeof err);
|
||||
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
|
||||
if (!pat) {
|
||||
fprintf(stderr, "pattern error at %lld: %s\n", (long long)err.pos, err.msg);
|
||||
}
|
||||
PatternError_free(&err);
|
||||
```
|
||||
|
||||
Output, verified:
|
||||
|
||||
```
|
||||
=== error_compile ===
|
||||
pat=(nil) msg='missing ), unterminated subpattern' pos=9 lineno=1 colno=10
|
||||
```
|
||||
|
||||
`err.msg`/`err.pattern` are heap allocated by whichever call filled the
|
||||
struct in; zero-initialize a `PatternError` before use and call
|
||||
`PatternError_free` on it once done reading it, whether or not the call
|
||||
that filled it in succeeded. A `NULL` `err` argument is always safe to pass
|
||||
to any function that accepts one, and simply skips error reporting.
|
||||
|
||||
A matching call (`Pattern_search`, and so on) returning `-1` means either
|
||||
invalid UTF-8 in the subject against a `UTF8`-mode `Pattern`, or the
|
||||
backtracking depth limit was reached (`MAX_DEPTH` in `regexx.c`, currently
|
||||
60000, not adjustable without editing `regexx.c` and rebuilding); no
|
||||
further detail is available through the return value alone in this build.
|
||||
|
||||
## 16. Pattern syntax reference
|
||||
|
||||
Literals; `.` (with `DOTALL` for it to match `\n`); character classes with
|
||||
ranges and negation (`[abc]`, `[^abc]`, `[a-z]`); `\d \D \w \W \s \S`;
|
||||
`\b \B`; anchors `^ $ \A \Z` (`^`/`$` also match at line boundaries with
|
||||
`MULTILINE`); quantifiers `* + ? {m,n} {m,} {,n} {m}`, both greedy and lazy
|
||||
(`*?`, `+?`, and so on) forms; possessive quantifiers `*+ ++ ?+ {m,n}+`;
|
||||
groups `(...)`, non-capturing `(?:...)`, named `(?P<name>...)`; alternation
|
||||
`|`; backreferences `\1`-`\99`, `(?P=name)`, `\g<name>`, `\g<N>`; lookahead
|
||||
`(?=...)`/`(?!...)`; fixed-width lookbehind `(?<=...)`/`(?<!...)`; atomic
|
||||
groups `(?>...)`; comments `(?#...)`; global inline flags `(?aiLmsux)` at
|
||||
the very start of a pattern; escapes `\n \r \t \f \v \a`, `\0`-prefixed
|
||||
octal escapes, `\xhh`, `\uxxxx`, `\Uxxxxxxxx`.
|
||||
|
||||
Rejected at compile time with a `PatternError` naming the construct, rather
|
||||
than silently mis-parsed: conditional groups `(?(id)yes|no)`, scoped inline
|
||||
flags `(?flags:...)` (only the global, start-of-pattern form is supported),
|
||||
`\N{NAME}` named code points, variable-width lookbehind (matching CPython's
|
||||
own restriction), and POSIX bracket-expression syntax `[[:alpha:]]` (not
|
||||
part of Python `re`'s grammar at all, `concept.md` Section 14). See
|
||||
`docs/API.md` Section 5-6 for the complete, exact list.
|
||||
|
||||
## 17. Performance and memory, in one paragraph
|
||||
|
||||
`match`/`fullmatch` (a single anchored attempt) cost time proportional to
|
||||
that attempt. `search`/`finditer`/`split`/`sub` are linear time, not
|
||||
quadratic, for the common case (no backreference, every quantifier's body a
|
||||
single character/class/`.`), at a measured memory cost of roughly 8.4x the
|
||||
input length for a pattern using a simple-atom quantifier. A pattern with a
|
||||
backreference, or shaped like nested overlapping quantifiers
|
||||
(`(a+)+b`-style), can still be worse than linear, the same way CPython's own
|
||||
`re` can be on the same patterns. `README.md` "Implementation status" and
|
||||
"Memory footprint" have the full measured account, including one fix that
|
||||
was tried, measured, and deliberately reverted; read it before assuming a
|
||||
given pattern's cost, rather than assuming linear time and bounded memory
|
||||
apply unconditionally.
|
||||
|
||||
## 18. Complete example: `rxgrep`
|
||||
|
||||
`examples/rxgrep.c` is a complete, working grep-like program built on this
|
||||
API, exercising all three data modes and both the matching and substitution
|
||||
API against real files and pipes:
|
||||
|
||||
```sh
|
||||
./rxgrep -in 'hello' file.txt # case-insensitive, line numbers
|
||||
./rxgrep -m utf8 -o '\w+' file.txt # print every UTF-8 word, one per line
|
||||
./rxgrep -c 'error' log.txt # count matching lines
|
||||
./rxgrep -m binary 'a.c' data.bin # match raw bytes, embedded NUL included
|
||||
./rxgrep -m utf8 --sub 'REDACTED' '\d{3}-\d{4}' file.txt
|
||||
```
|
||||
|
||||
Reading its source alongside this document is a reasonable next step once
|
||||
the examples above are familiar: it shows every function here used together
|
||||
in one program, including the parts this document simplified away for
|
||||
clarity (option parsing, line splitting, output formatting).
|
||||
Reference in New Issue
Block a user