Files
regexx/examples/binary_scan.c
T
retoorandClaude Sonnet 5 02743cd959 Add six feature examples and a direct benchmark against POSIX <regex.h>
Each new program under examples/ isolates one distinct feature rather
than being a general purpose tool like the existing rxgrep.c:
binary_scan.c (raw byte-range classes including an embedded NUL and
an embedded 0x0A), utf8_scripts.c (\w across Latin/Greek/Cyrillic/CJK
text, code point versus byte offsets), ascii_logparse.c (named groups
against structured log text), redos_atomic.c (atomic groups and
possessive quantifiers timed directly against the unprotected form of
the textbook (a+)+b ReDoS shape), empty_match_rule.c (CPython's
undocumented empty-match retry rule, verified: \d*? against
"123abc456" gives 16 matches, not 9), and large_file_search.c
(Input_from_file's mmap-backed reading on a generated 100MB file,
with elapsed time and peak RSS printed). Every example was compiled
and run while writing it; the claims in each file's top comment are
checked against its own output, not written by hand and left
unverified.

Also adds examples/bench_vs_posix.c, a direct, honestly reported
comparison against the C standard library's own <regex.h>
(regcomp/regexec) on six scenarios at multi-megabyte or
multi-hundred-thousand-line scale, using only pattern syntax valid
for both engines so they run the identical pattern text. glibc's
DFA-backed engine wins five of six scenarios by 2x-35x, which is the
expected outcome of a roughly 2000-line backtracking interpreter
built for Python `re` compatibility competing against a mature,
heavily optimized engine with a much smaller feature set; the sixth
scenario has no POSIX equivalent at all (an atomic group). Every
scenario's match count is cross-checked between the two engines as an
independent correctness signal beyond the existing CPython-derived
test suite.

Two real issues were found and fixed while building this benchmark,
not left in: iterating regexec() over an advancing string pointer is
quadratic in practice (no way to bound the search without an implicit
NUL-scan on every call), fixed by using REG_STARTEND instead; and a
signed integer overflow (undefined behavior, caught by UBSan) in the
benchmark's own pseudo-random text generator, fixed by using an
unsigned accumulator.

README.md and USAGE.md gain pointers to examples/README.md (the new
per-example index) and a "Benchmarks" section summarizing the
POSIX comparison honestly, including where it loses. The Makefile
gains a `make examples` target building all seven programs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
2026-09-14 11:04:10 +00:00

65 lines
2.7 KiB
C

/* binary_scan - demonstrates BINARY mode: matching against raw bytes with
* no decoding step at all, where every value 0-255 is an ordinary input
* character, including 0x00 (NUL) and 0x0A (the byte '\n' means in text).
* This is the one thing ASCII and UTF8 mode structurally cannot do: ASCII
* mode still treats the subject as bytes but restricts \d/\w/\s to the
* ASCII subset, and UTF8 mode decodes the subject before matching at all,
* so an arbitrary byte class spanning 0x00-0xFF describes decoded code
* points there, not raw bytes on the wire.
*
* The scenario: a small binary record format concatenated in one buffer,
* each record a 2-byte magic number followed by 4 raw payload bytes that
* may be anything, including bytes a C string function would treat as a
* terminator or a line break. A pattern built from re.escape'd magic bytes
* and a [\x00-\xff]{4} class extracts every record correctly regardless.
*/
#include "regexx.h"
#include <stdio.h>
#include <string.h>
typedef struct { int n; } Ctx;
static void print_record(void *ctx, const Match *m_const) {
Ctx *c = ctx;
Match *m = (Match *)m_const;
const char *g; size_t glen;
Match_group(m, NULL, 0, &g, &glen);
printf("record %d: magic=%02x%02x payload=", c->n, (unsigned char)g[0], (unsigned char)g[1]);
for (size_t i = 2; i < glen; i++) printf("%02x ", (unsigned char)g[i]);
printf("\n");
c->n++;
}
int main(void) {
/* Three records, back to back, with no separator: a real length-value
* binary format would carry its own framing; this example only needs
* a fixed-size payload to keep the pattern simple. Payload bytes
* deliberately include 0x00 and 0x0a, the two values that would break
* a NUL-terminated-string or line-oriented approach to scanning this
* same buffer. */
uint8_t data[] = {
0xca, 0xfe, 0x01, 0x00, 0x00, 0x00, /* record 0: payload has an embedded NUL */
0xca, 0xfe, 0x0a, 0x0a, 0xff, 0x7f, /* record 1: payload has embedded 0x0a bytes */
0xca, 0xfe, 0xde, 0xad, 0xbe, 0xef, /* record 2: ordinary payload */
};
const char *pattern = "\\xca\\xfe[\\x00-\\xff]{4}";
PatternError err; memset(&err, 0, sizeof err);
Pattern *pat = re_compile(pattern, strlen(pattern), BINARY, &err);
if (!pat) {
fprintf(stderr, "compile error: %s\n", err.msg);
PatternError_free(&err);
return 1;
}
Input *in = Input_from_buffer(data, sizeof data);
Ctx ctx = { 0 };
int n = Pattern_finditer(pat, in, 0, -1, print_record, &ctx);
printf("total records found: %d (buffer is %zu bytes with %d embedded NUL, %d embedded 0x0a)\n",
n, sizeof data, 1, 2);
Input_free(in);
Pattern_free(pat);
return 0;
}