Files
regexx/examples/README.md
T

33 lines
4.0 KiB
Markdown
Raw Normal View History

# Examples
Each program here is self-contained (`main`, no shared helper code) and
isolates one distinct feature rather than being a general purpose tool; the
one exception is `rxgrep.c`, a small grep-like program that ties several of
the same features together into something actually usable from a shell.
[`../USAGE.md`](../USAGE.md) walks through the same API surface function by
function, with smaller inline snippets; these programs are complete,
runnable, and closer to what real code doing this specific thing looks
like.
Build all of them with `make examples` from the repository root (or
`make all` / `make example` for just `rxgrep`, the default). Each is also a
single `cc -I. file.c libregexx.a -o file` invocation away from being built
directly, no build system required.
| Program | Demonstrates |
|---|---|
| [`rxgrep.c`](rxgrep.c) | A complete grep-like CLI: all three data modes via `-m`, `Pattern_search`/`finditer` for line and per-match output, `Pattern_subn` for `--sub`, reading from a real file or a pipe. Run `./rxgrep --help`. |
| [`binary_scan.c`](binary_scan.c) | `BINARY` mode: a byte-range character class (`[\x00-\xff]`) matching raw bytes 0-255, including an embedded `NUL` and an embedded `0x0A`, neither of which is a terminator or a line break in this mode the way it would be to a C string function or a line-oriented reader. This is the one thing `ASCII`/`UTF8` mode cannot do, since both introduce either a restricted classification or a decoding step. |
| [`utf8_scripts.c`](utf8_scripts.c) | `UTF8` mode: `\w` recognizing letters across Latin, Greek, Cyrillic, and CJK text (not only ASCII), and the resulting difference between `Match_span` (code point units) and `Match_span_byte` (byte units) once a match spans characters that are more than one byte wide. |
| [`ascii_logparse.c`](ascii_logparse.c) | `ASCII` mode doing what it is ordinarily used for: parsing structured, line-oriented text with named groups (`(?P<name>...)`) and reading fields back out by name via `Match_group`/`Pattern_groupindex_lookup`, not by a numeric position the caller has to remember. |
| [`redos_atomic.c`](redos_atomic.c) | Two-part: first, that the textbook `(a+)+b` ReDoS shape is now fixed automatically, with no atomic group, because it has no backreference and so runs on the Pike VM (README.md "The Pike VM"); second, a variant with a backreference added specifically to force it back onto the backtracking engine, where the same shape is exponential again and an atomic group (`(?>...)`) is still the pattern author's own necessary fix. |
| [`empty_match_rule.c`](empty_match_rule.c) | `Pattern_finditer`/`Pattern_split` reproducing a real CPython interpreter's undocumented empty-match retry rule exactly (`\d*?` against `"123abc456"` yields 16 matches, not 9), found by probing a real interpreter directly rather than by reading its documentation. |
| [`large_file_search.c`](large_file_search.c) | `Input_from_file`'s `mmap`-backed reading on a generated 100MB file, with the elapsed time and peak resident memory printed directly so they can be checked against the measured figures in `../README.md` "Memory footprint" rather than taken on faith. |
| [`bench_vs_posix.c`](bench_vs_posix.c) | A direct, timed comparison against the C standard library's own `<regex.h>` (POSIX `regcomp`/`regexec`) on six scenarios at multi-megabyte/multi-hundred-thousand-line scale, using only pattern syntax valid in both engines, every one of them Pike VM eligible. Reported honestly, including a real regression: glibc wins the three ordinary scenarios by a wider margin than this build measured before the Pike VM existed (it has none of the backtracking engine's constant-factor optimizations yet), while the adversarial `(a+)+b` scenario inverts entirely and now beats glibc outright, with no atomic group needed. |
Every program below was compiled and actually run while writing it; the
claims in its top-of-file comment (what a real CPython interpreter does,
what timing difference an atomic group makes, and so on) are checked
against that program's own printed output, not written by hand and left
unverified.