Researched first (Cox's regexp2 article, the submatch/tagged-NFA follow-up crediting Laurikari, and rust-lang/regex's PikeVM source for the exact leftmost-first mid-search Match handling), then planned in concept.md 7.7 before writing any code, per the explicit instruction to plan before implementing. The existing bytecode compiler turned out to already produce valid Thompson-construction NFA bytecode (OP_SPLIT/OP_JMP/OP_SAVE match Cox's instruction set almost exactly), so no new compiler was needed: Prog gained one field (no_repeat1) and compile_repeat one condition, letting the same compile_node produce a second, NFA-only Prog from the same AST for any pattern containing none of OP_BACKREF/OP_LOOKAHEAD/ OP_LOOKBEHIND/OP_ATOMIC (a possessive quantifier already desugars to the last of these at parse time). That second Prog is matched by a new Pike VM (pike_addthread/pike_step/pike_find): a breadth-first thread list simulation with per-thread capture arrays, epsilon closure implemented with an explicit heap stack rather than C recursion so a pattern with many alternations cannot recurse the C stack, every allocation checked and failing through the existing -1 error convention rather than a NULL dereference. do_one, Pattern_finditer, and Pattern_split each gained an impl->has_nfa branch to this engine, sharing one small helper (find_next) for the "unanchored scan from a position" versus "single anchored attempt at a position" distinction finditer/split's empty-match retry needs. Two real bugs, both found and precisely localized by the existing 3,252-case CPython-derived suite without writing a single test specifically for this engine: OP_MATCH not writing group 0's end position (it has no OP_SAVE; run()'s own OP_MATCH handler sets it directly, and this engine's first version missed replicating that), and an unconditional "thread list empty -> stop" early exit that is wrong for unanchored search specifically, since a freshly injected start thread can die immediately in its own epsilon closure (a leading \b failing outright, repeatedly, inside a longer word like "catalog" for \bcat\b) without that meaning every later position would too. Both fixed; full suite passes, three clean AddressSanitizer/ UndefinedBehaviorSanitizer runs, plus hand-written whitebox checks (a 20,000-branch alternation, UTF8 named groups, BINARY matching across an embedded NUL, greedy/lazy and alternation priority). Measured result: (a+)+b, this project's own running example of the backtracking engine's remaining weak spot, is Pike VM eligible and now measures as genuinely linear (0.0018s to 0.0308s, n=10,000 to 160,000), not merely improved. Measured cost: re-running bench_vs_posix.c's three ordinary scenarios (all now Pike VM eligible too) found the gap to POSIX <regex.h> widened, from roughly 2x-25x before this engine existed to roughly 8x-70x now, the direct, expected cost of this engine's performance axis being explicitly deferred (no literal prefilter, no lazy DFA state caching, no allocation pooling) in favor of correctness first, per the instruction this was built under. Both results, and the reasoning behind deferring the second, are recorded in concept.md 7.7/7.8. README.md, docs/API.md, USAGE.md, and examples/redos_atomic.c and bench_vs_posix.c are updated throughout to describe the new two-engine dispatch accurately, including this real trade-off, rather than leaving the previous single-engine description in place; redos_atomic.c specifically now shows both that (a+)+b no longer needs an atomic group at all and a backreference-forced variant where one still does. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
103 lines
3.8 KiB
C
103 lines
3.8 KiB
C
/* large_file_search - demonstrates Input_from_file's mmap-backed reading
|
|
* (README.md "Memory footprint") on a file large enough that the
|
|
* difference between mapping it and copying it into a fresh malloc'd
|
|
* buffer is not theoretical. The example generates its own input (so the
|
|
* program is self-contained and does not depend on a file existing on the
|
|
* machine it runs on), searches it for a pattern that does not occur
|
|
* anywhere (the worst case for a search: every position is tried, since
|
|
* none of them succeeds early), and reports both the elapsed time and,
|
|
* where /proc is available, the process's peak resident memory, so the
|
|
* two numbers can be compared directly against the measured figures in
|
|
* README.md rather than taken on faith.
|
|
*/
|
|
#define _POSIX_C_SOURCE 199309L
|
|
#include "regexx.h"
|
|
#include <stdio.h>
|
|
#include <string.h>
|
|
#include <stdlib.h>
|
|
#include <time.h>
|
|
|
|
static double now_seconds(void) {
|
|
struct timespec ts;
|
|
clock_gettime(CLOCK_MONOTONIC, &ts);
|
|
return (double)ts.tv_sec + (double)ts.tv_nsec / 1e9;
|
|
}
|
|
|
|
static void print_peak_rss(void) {
|
|
FILE *f = fopen("/proc/self/status", "r");
|
|
if (!f) { printf("(peak RSS: /proc not available on this platform)\n"); return; }
|
|
char line[256];
|
|
while (fgets(line, sizeof line, f)) {
|
|
if (strncmp(line, "VmHWM:", 6) == 0) {
|
|
printf("peak RSS: %s", line + 6);
|
|
break;
|
|
}
|
|
}
|
|
fclose(f);
|
|
}
|
|
|
|
int main(void) {
|
|
const char *path = "/tmp/regexx_large_file_search_example.dat";
|
|
size_t size = 100u * 1024 * 1024; /* 100MB: large enough to be a real
|
|
* measurement, small enough to run
|
|
* in a few seconds on ordinary
|
|
* hardware. */
|
|
|
|
printf("Generating a %zu MB file of non-matching text at %s ...\n", size / (1024 * 1024), path);
|
|
{
|
|
FILE *f = fopen(path, "wb");
|
|
if (!f) { perror("fopen"); return 1; }
|
|
/* Repeating, harmless filler text: never contains the digit
|
|
* sequence the search pattern below looks for. */
|
|
static const char chunk[] =
|
|
"the quick brown fox jumps over the lazy dog. ";
|
|
size_t chunklen = sizeof(chunk) - 1;
|
|
size_t written = 0;
|
|
while (written < size) {
|
|
size_t n = fwrite(chunk, 1, chunklen, f);
|
|
if (n == 0) break;
|
|
written += n;
|
|
}
|
|
fclose(f);
|
|
}
|
|
|
|
const char *pattern = "\\d{10,}"; /* ten or more consecutive digits: not present */
|
|
PatternError err; memset(&err, 0, sizeof err);
|
|
Pattern *pat = re_compile(pattern, strlen(pattern), ASCII, &err);
|
|
if (!pat) {
|
|
fprintf(stderr, "compile error: %s\n", err.msg);
|
|
PatternError_free(&err);
|
|
remove(path);
|
|
return 1;
|
|
}
|
|
|
|
Input *in = Input_from_file(path, &err);
|
|
if (!in) {
|
|
fprintf(stderr, "cannot read file: %s\n", err.msg);
|
|
PatternError_free(&err);
|
|
Pattern_free(pat);
|
|
remove(path);
|
|
return 1;
|
|
}
|
|
|
|
printf("Searching (Input_from_file mmaps this file rather than copying it,\n"
|
|
"so this process's own heap never holds a private copy of the 100MB\n"
|
|
"file content itself; this pattern has no backreference, lookaround,\n"
|
|
"or atomic group, so it also runs on the Pike VM, README.md \"The\n"
|
|
"Pike VM\", which needs none of the backtracking engine's own search\n"
|
|
"tables, README.md \"Memory footprint\", at all)...\n");
|
|
Match m; memset(&m, 0, sizeof m);
|
|
double t0 = now_seconds();
|
|
int rc = Pattern_search(pat, in, 0, -1, &m);
|
|
double t1 = now_seconds();
|
|
printf("result: rc=%d (0 = no match, as expected) in %.3fs\n", rc, t1 - t0);
|
|
print_peak_rss();
|
|
|
|
if (rc == 1) Match_free(&m);
|
|
Input_free(in);
|
|
Pattern_free(pat);
|
|
remove(path);
|
|
printf("(temporary file removed)\n");
|
|
return 0;
|
|
}
|