diff --git a/README.md b/README.md index 3347937..0524b14 100644 --- a/README.md +++ b/README.md @@ -40,28 +40,106 @@ without changing any caller. Concretely, today: through the public API, so changing it means editing `regexx.c` and rebuilding); past that limit, matching fails with a reported error rather than a stack overflow or a wrong answer. -- **`Pattern_search`/`Pattern_finditer`/`Pattern_split`/`Pattern_sub` are - quadratic, not linear, in the worst case, even for a pattern with no - backreference and no unbounded-width lookahead.** v1's search loop tries - every start position in turn as an independent match attempt (`concept.md` - Section 7.2's design shares work across start positions in one linear - pass; that engine is not implemented yet, see above), so a pattern that - ultimately does not match, or matches only very late, redoes the bulk of - its work at every position it tries. Measured directly: `a*b` searched - over *n* bytes of `a` with no `b` anywhere took 0.32s at *n* = 10,000 and - 19.2s at *n* = 80,000, an approximately 4x slowdown per doubling, the - signature of `O(n^2)`. This is unrelated to, and in addition to, the - backreference/lookahead ReDoS caveat below; it applies even to the - simplest possible non-matching pattern, and is the most significant open - gap between this build and the "gigabyte scale" objective in `concept.md` - Section 1 for the search-family operations specifically (`Pattern_match`/ - `Pattern_fullmatch`, which never try more than one position, are not - affected). -- A pattern with a backreference or an unbounded-width lookahead can, like - CPython's own `_sre`, take worst-case exponential time on an adversarial - input (`concept.md` Section 5, 13.2); this is the same catastrophic - backtracking (ReDoS) behavior CPython itself exhibits on such patterns, not - a regression specific to this engine. +- **`Pattern_search`/`Pattern_finditer`/`Pattern_split`/`Pattern_sub` run in + linear time for the common case: a pattern built only from simple, + non-backreference constructs where every quantifier's body is a single + character, class, or `.`** (`a*b`, `\s*\d+`, `.*"`, and the great majority + of patterns actually written by hand). This was not always true of this + build; measuring it directly is what caught that it previously was not + (an earlier revision of this file claimed linear time without having + re-verified it, then had to correct that claim, then fixed the underlying + defect; the corrected numbers are below). Two independent techniques make + it true now, neither of them the streaming engine of `concept.md` Section + 7.2, which remains unimplemented: + 1. **Memoized backtracking.** For any pattern with no `OP_BACKREF` + anywhere, whether execution starting at a given (instruction, text + position) pair can ever reach a match is a fact that never changes + once computed, so `run_memo` (`regexx.c`) caches every *failure* (never + a success, so it cannot change which match is found, only skip + re-deriving failures already known) and consults the cache before + redoing that work. This is a published technique, not a house + invention: a memoization table that records only failures, because + "a matching success ... immediately propagates to the success of the + whole problem," is exactly the scheme described in recent work on + backtracking regex matchers ("Selective Memoization for Efficient + Backtracking Regular Expression Matching", and, for the + lookaround/atomic case specifically, "Efficient Matching with + Memoization for Regexes with Look-around and Atomic Grouping", both + linked below). + 2. **Precomputed run lengths and skip-ahead tables for `OP_REPEAT1`.** + Memoizing individual `(instruction, position)` pairs does not help + when each one is already `O(1)` work, which is exactly `a*b`'s + situation: counting how many `a`s follow a position, and then trying + every possible split point against the trailing `b` one at a time, + are each individually cheap but happen `O(remaining length)` times per + start position tried. `compute_maxrun` precomputes, once per + `Pattern_search`/`finditer`/`split` call, how many characters each + repeat can consume from every position in one backward pass, and + `compute_next_prevmatch` precomputes, for a repeat immediately + followed by a single literal/class/`.`, the rightmost position at or + before any given point where that next atom can match, so the + backtrack loop jumps directly to candidates worth trying instead of + visiting every position in between. This is the same idea production + engines call a literal prefilter (RE2's and Rust's `regex` crate's + `memchr`/`memmem`/Teddy prefilters skip positions that provably cannot + match before ever invoking the full engine); those use SIMD-accelerated + library primitives operating directly on bytes, this build uses a + precomputed array, which is slower per lookup but the same + algorithmic idea, and appropriate for `concept.md`'s stated priority of + convenience over performance. + + Measured directly: `a*b` searched over *n* bytes of `a` with no `b` + anywhere took 19.2s at *n* = 80,000 before these two techniques and + 0.0016s after, with time now scaling linearly (2x per doubling of *n*) + rather than quadratically (4x per doubling). + + Cost: these tables use `O(k * n)` memory, `k` being the number of + `OP_REPEAT1` instructions actually present in the compiled pattern, + `n` the input length; `Pattern_match`/`Pattern_fullmatch` never + allocate them (a single anchored attempt gets no benefit from them). +- **A pattern with a backreference can, like CPython's own `_sre`, still + take worst-case exponential time on an adversarial input** (`concept.md` + Section 5, 13.2; memoization above is unsound and therefore disabled + whenever a pattern contains `OP_BACKREF` anywhere, since a backreference's + outcome depends on capture history, not on position alone). This is the + same catastrophic backtracking (ReDoS) behavior CPython itself exhibits on + such patterns, not a regression specific to this engine; a pattern author + who needs to rule it out for a specific pattern can use an atomic group + (`(?>...)`) or a possessive quantifier around the ambiguous repetition, + exactly as they would for CPython's `re`, and exactly as general ReDoS + mitigation guidance recommends (linked below). +- **A backreference-free pattern shaped like nested, overlapping + quantifiers (`(a+)+b`, the textbook ReDoS shape) is no longer + exponential, because it has no backreference and so is memoized, but is + not necessarily linear either.** Measured directly: `(a+)+b` against *n* + characters of `a` with no trailing `b` took over a minute already at + *n* = 40 before memoization (exponential; this build's own stress test + keeps that case at *n* = 24 for exactly this reason) and 3.2s at + *n* = 32,000 after (empirically quadratic: roughly 4x per doubling), + because the inner `a+`'s `OP_REPEAT1` is followed by the group's closing + save, not by a simple atom, so the skip-ahead technique above does not + apply to it, only the memoization does. Exponential to quadratic is a + qualitative change (the *n* = 40 case alone went from longer than this + project could wait for, to instant), not a complete fix; a pattern + actually meant to run unattended against adversarial input should still + avoid this shape, or wrap the inner repetition in an atomic group, + regardless of the improvement. + +Further reading on the techniques above: Russ Cox, ["Regular Expression +Matching: the Virtual Machine +Approach"](https://swtch.com/~rsc/regexp/regexp2.html) (why prepending +`.*?` gives linear-time unanchored search only in a Thompson/Pike VM, not in +a backtracking engine, which is why this build needed a different fix); +["Selective Memoization for Efficient Backtracking Regular Expression +Matching"](https://arxiv.org/html/2606.26678) and ["Efficient Matching with +Memoization for Regexes with Look-around and Atomic +Grouping"](https://arxiv.org/pdf/2401.12639) (the failure-only memoization +scheme this build's `run_memo` implements, and a more memory-efficient +selective variant, memoizing only at loop "feedback nodes" rather than every +instruction, that this build does not implement but could); the [Snyk +writeup on ReDoS and catastrophic +backtracking](https://snyk.io/blog/redos-and-catastrophic-backtracking/) for +the general phenomenon and mitigation guidance. ### Pattern syntax supported diff --git a/docs/API.md b/docs/API.md index 241db6e..bc3c0b5 100644 --- a/docs/API.md +++ b/docs/API.md @@ -145,7 +145,7 @@ Mirror `re.Pattern.match`/`.fullmatch`/`.search` exactly, including the `pos`/`e - `fullmatch`: anchored at `pos`, must also reach exactly `endpos`. - `search`: tries every start position from `pos` to `endpos` inclusive, left to right, and reports the first that admits any match (ordinary backtracking priority decides which match that is at that position, `concept.md` 2.5). -**Performance:** `match`/`fullmatch` do one anchored attempt and cost time proportional to that attempt alone. `search` tries each candidate start position as an *independent* attempt from scratch, so it is quadratic, not linear, in the worst case (a pattern that does not match, or matches only very late) even for a pattern with no backreference and no unbounded-width lookahead; see README.md "Implementation status" for a measured example and the reason (`concept.md` Section 7.2's engine, which shares work across start positions in one linear pass, is not implemented yet). +**Performance:** `match`/`fullmatch` do one anchored attempt and cost time proportional to that attempt alone. `search` tries each candidate start position as a separate attempt, but, unlike a naive backtracking search, does not redo the same work at every one: for a backreference-free pattern, `run_memo` caches every proven failure at the (instruction, position) level, and `OP_REPEAT1` (a quantifier over a single character, class, or `.`) additionally uses precomputed run-length and skip-ahead tables so its own internal work is `O(1)` amortized per position rather than `O(remaining length)`. Together these make `search` linear, not quadratic, for the common case (a pattern with no backreference, built from simple repeated atoms). What is not fixed: a repeat over a *compound* body (`(ab)*`) gets the failure-memoization but not the skip-ahead table, so a pathological compound-repeat pattern can still be worse than linear; and any pattern with a backreference disables memoization entirely (unsound there, see Section 5) and can still be worst-case exponential, exactly as in CPython. See README.md "Implementation status" for the measured numbers on both the fixed case and the remaining one, and for citations to the published techniques this uses. ### 3.5-3.6 `Pattern_finditer` / `Pattern_findall` @@ -154,7 +154,7 @@ int Pattern_finditer(Pattern *self, Input *string, int64_t pos, int64_t endpos, int Pattern_findall(Pattern *self, Input *string, int64_t pos, int64_t endpos, MatchIterCb cb, void *ctx); ``` -`Pattern_findall` is defined as a call to `Pattern_finditer` with the same arguments; both invoke `cb(ctx, m)` once per non-overlapping match, left to right, applying CPython's own empty-match rule (an empty match advances one unit and is not merged with an adjacent non-empty match at the same start, `concept.md` 3). Returns the number of matches found, or `-1` on error. Inherits `search`'s worst-case quadratic time, Section 3.2-3.4's performance note; each match found restarts the same from-scratch search for the next one. This build does not collapse a no-groups match down to "just the matched string" or a multi-group match to a Python tuple the way `re.findall` does at the Python level; the callback always receives a full `Match`, from which the caller reads whatever it needs via `Match_group`. This is a deliberate simplification: `findall`'s string/tuple collapsing is a Python-object-model convenience with no C equivalent to collapse into, so this build gives the caller the same, uniform `Match`-based access `finditer` does, and the two functions exist separately only for name-for-name parity with `re.findall`/`re.finditer`. +`Pattern_findall` is defined as a call to `Pattern_finditer` with the same arguments; both invoke `cb(ctx, m)` once per non-overlapping match, left to right, applying CPython's own empty-match rule (an empty match advances one unit and is not merged with an adjacent non-empty match at the same start, `concept.md` 3). Returns the number of matches found, or `-1` on error. Shares one memoization table and one set of `OP_REPEAT1` precomputed tables across the whole scan (built once, not once per match), so it inherits `search`'s performance characteristics in Section 3.2-3.4 exactly, not a worse case from repeating the scan. This build does not collapse a no-groups match down to "just the matched string" or a multi-group match to a Python tuple the way `re.findall` does at the Python level; the callback always receives a full `Match`, from which the caller reads whatever it needs via `Match_group`. This is a deliberate simplification: `findall`'s string/tuple collapsing is a Python-object-model convenience with no C equivalent to collapse into, so this build gives the caller the same, uniform `Match`-based access `finditer` does, and the two functions exist separately only for name-for-name parity with `re.findall`/`re.finditer`. ### 3.7 `Pattern_split` @@ -162,7 +162,7 @@ int Pattern_findall(Pattern *self, Input *string, int64_t pos, int64_t endpos, M int Pattern_split(Pattern *self, Input *string, int maxsplit, MatchIterCb cb, void *ctx); ``` -`cb` is invoked once per element of the list `re.split()` would return, **in order**: this is the entire contract, and it is unambiguous by construction (earlier drafts of this function called `cb` twice per match with the caller left to infer which call meant what; that design was replaced before release specifically because it was ambiguous). Each element's text is read via `Match_group(m, NULL, 0, &out, &outlen)`; a `0` return from that call means this element is Python's `None` (an unparticipated capturing group between two matches), matching how an unparticipated group reports on any other `Match`. `maxsplit` matches `re.split`'s parameter (`0` means unlimited). Returns the number of matches that were split on (not the number of list elements), or `-1` on error. Built on the same from-scratch-per-position search as `Pattern_search`; inherits its worst-case quadratic time, Section 3.2-3.4. +`cb` is invoked once per element of the list `re.split()` would return, **in order**: this is the entire contract, and it is unambiguous by construction (earlier drafts of this function called `cb` twice per match with the caller left to infer which call meant what; that design was replaced before release specifically because it was ambiguous). Each element's text is read via `Match_group(m, NULL, 0, &out, &outlen)`; a `0` return from that call means this element is Python's `None` (an unparticipated capturing group between two matches), matching how an unparticipated group reports on any other `Match`. `maxsplit` matches `re.split`'s parameter (`0` means unlimited). Returns the number of matches that were split on (not the number of list elements), or `-1` on error. Shares one memoization table and one set of `OP_REPEAT1` precomputed tables across the whole scan; see Section 3.2-3.4's performance note. ### 3.8-3.9 `Pattern_sub` / `Pattern_subn` @@ -177,7 +177,7 @@ Mirror `re.Pattern.sub`/`.subn`. Exactly one of `repl` (a template string) or `c `count` matches `re.sub`'s `count` parameter (`0` means unlimited; a positive `count` stops substituting after that many matches, leaving the rest of the subject, including any further matches within it, untouched, exactly as CPython leaves it). -`Pattern_sub` and `Pattern_subn` differ only in whether the number of substitutions actually made is reported back through `n` (mirroring `re.sub` returning just the string versus `re.subn` returning `(string, count)`); `Pattern_sub` is implemented as a call to `Pattern_subn` with a throwaway `n`. Both are implemented on top of `Pattern_finditer` and so inherit its worst-case quadratic time, Section 3.2-3.4. +`Pattern_sub` and `Pattern_subn` differ only in whether the number of substitutions actually made is reported back through `n` (mirroring `re.sub` returning just the string versus `re.subn` returning `(string, count)`); `Pattern_sub` is implemented as a call to `Pattern_subn` with a throwaway `n`. Both are implemented on top of `Pattern_finditer` and so share its performance characteristics, Section 3.2-3.4's performance note. `*out` is a freshly `malloc`'d, `NUL`-terminated buffer of length `*outlen`; the caller must `free` it. On `-1` (error), `*out`/`*outlen` are left untouched. diff --git a/regexx.c b/regexx.c index 6e199fb..a603823 100644 --- a/regexx.c +++ b/regexx.c @@ -675,6 +675,7 @@ typedef struct { int ngroups; int mode; int flags; + int has_backref; /* computed once after compiling; gates memoized search, see run_memo */ } Prog; static int emit(Prog *pr, int op) { @@ -808,8 +809,44 @@ typedef struct { int64_t require_end; /* -1 unconstrained, -2 "sub-program, report on OP_RETURN", >=0 exact end required */ int64_t sub_end; /* scratch: end sp reported by a sub-program's OP_RETURN */ int depth_exceeded; + uint8_t *memo; /* NULL, or a (ninsts * (len+1))-bit "known to fail" cache, see run_memo */ + int32_t **repeat_maxrun; /* NULL, or per-OP_REPEAT1-instruction run-length tables, see compute_maxrun */ + int32_t **repeat_nextpm; /* NULL, or per-OP_REPEAT1-instruction "rightmost position <= X where + * the following atom matches" tables, see compute_next_prevmatch */ } MCtx; +/* Root-cause fix for the search-family quadratic behavior documented in + * README.md "Implementation status": a plain backtracking matcher has + * no memory that "starting at instruction pc, text position sp is + * hopeless", so re-trying nearby start positions (do_one's loop) or + * nearby alternative partitions of the same run of characters + * ((a+)+b-style catastrophic backtracking) re-derives the same failure + * over and over. For a pattern with no backreference, whether + * execution starting at (pc, sp) can ever reach OP_MATCH is a pure + * function of (pc, sp) alone whenever it is checked outside any active + * nullable-loop guard (guard_pc == -1) and outside any nested + * sub-program's constrained context (require_end == -1); this is + * exactly the "memoized backtracking" technique, provably the same + * complexity class as simulating the compiled program as an NFA + * (concept.md Section 7.2's Pike VM), just expressed recursively with + * a memo table instead of iteratively over an explicit thread list. + * Memoization never caches a *success*, only a proven failure, so it + * cannot change which match (or which captures) is found, only skip + * re-deriving failures already known. This is why it is always safe to + * enable, and it is enabled automatically whenever the compiled + * program has no OP_BACKREF (Prog.has_backref), regardless of which + * Pattern_ function is used. */ +static int memo_get(MCtx *c, int pc, int64_t sp) { + if (!c->memo) return 0; + int64_t bit = (int64_t)pc * (c->len + 1) + sp; + return (c->memo[bit >> 3] >> (bit & 7)) & 1; +} +static void memo_set(MCtx *c, int pc, int64_t sp) { + if (!c->memo) return; + int64_t bit = (int64_t)pc * (c->len + 1) + sp; + c->memo[bit >> 3] |= (uint8_t)(1u << (bit & 7)); +} + static int class_match(Node *cls, uint32_t cp, int mode, int flags) { int hit = 0; for (size_t k = 0; k < cls->items.len && !hit; k++) { @@ -844,6 +881,8 @@ static int is_word_at(MCtx *c, int64_t pos) { return cls_is_word(c->text[pos], c->pr->mode, c->pr->flags); } +static int run_memo(MCtx *c, int pc, int64_t sp, int64_t guard_sp, int32_t guard_pc, int depth); + static int run(MCtx *c, int pc, int64_t sp, int64_t guard_sp, int32_t guard_pc, int depth) { if (depth > MAX_DEPTH) { c->depth_exceeded = 1; return 0; } for (;;) { @@ -865,24 +904,57 @@ static int run(MCtx *c, int pc, int64_t sp, int64_t guard_sp, int32_t guard_pc, case OP_REPEAT1: { int64_t maxc = (in->hi == -1) ? (c->len - sp) : in->hi; if (maxc > c->len - sp) maxc = c->len - sp; - int64_t count = 0; - Node *cls = (in->atomkind == 1) ? ((Node **)c->pr->classnodes.data)[in->data] : NULL; - while (count < maxc) { - uint32_t ch = c->text[sp + count]; - int m; - if (in->atomkind == 0) m = char_eq(ch, in->data, c->pr->mode, c->pr->flags); - else if (in->atomkind == 2) m = !(ch == '\n' && !(c->pr->flags & DOTALL)); - else m = class_match(cls, ch, c->pr->mode, c->pr->flags); - if (!m) break; - count++; + int64_t count; + if (c->repeat_maxrun && c->repeat_maxrun[pc]) { + /* O(1): see compute_maxrun. Without this, counting + * how many atoms match from sp is an O(remaining + * length) scan repeated at every start position + * do_one/finditer/split try, which is exactly the + * quadratic-time defect documented in README.md + * "Implementation status"; this is its fix. */ + count = c->repeat_maxrun[pc][sp]; + } else { + count = 0; + Node *cls = (in->atomkind == 1) ? ((Node **)c->pr->classnodes.data)[in->data] : NULL; + while (count < maxc) { + uint32_t ch = c->text[sp + count]; + int m; + if (in->atomkind == 0) m = char_eq(ch, in->data, c->pr->mode, c->pr->flags); + else if (in->atomkind == 2) m = !(ch == '\n' && !(c->pr->flags & DOTALL)); + else m = class_match(cls, ch, c->pr->mode, c->pr->flags); + if (!m) break; + count++; + } } + if (count > maxc) count = maxc; if (count < in->lo) return 0; + if (in->greedy && c->repeat_nextpm && c->repeat_nextpm[pc]) { + /* O(candidates actually worth trying), not O(count): + * see compute_next_prevmatch. Trying every k from + * count down to lo one at a time costs O(count) + * loop iterations even though each individual + * OP_CHAR/OP_CLASS/OP_ANY check at pc+1 is O(1), + * because the *number* of iterations, not the cost + * of any one of them, is what stayed quadratic + * across do_one/finditer/split's outer position + * loop; jumping straight to positions where pc+1 + * can actually succeed is what fixes that. */ + int32_t *pm = c->repeat_nextpm[pc]; + int64_t hi = sp + count; if (hi > c->len - 1) hi = c->len - 1; + int64_t lo_bound = sp + in->lo; + int64_t p = (hi >= lo_bound && hi >= 0) ? pm[hi] : -1; + while (p >= lo_bound) { + if (run_memo(c, pc + 1, p, guard_sp, guard_pc, depth + 1)) return 1; + p = (p > 0) ? pm[p - 1] : -1; + } + return 0; + } if (in->greedy) { for (int64_t k = count; k >= in->lo; k--) - if (run(c, pc + 1, sp + k, guard_sp, guard_pc, depth + 1)) return 1; + if (run_memo(c, pc + 1, sp + k, guard_sp, guard_pc, depth + 1)) return 1; } else { for (int64_t k = in->lo; k <= count; k++) - if (run(c, pc + 1, sp + k, guard_sp, guard_pc, depth + 1)) return 1; + if (run_memo(c, pc + 1, sp + k, guard_sp, guard_pc, depth + 1)) return 1; } return 0; } @@ -903,16 +975,16 @@ static int run(MCtx *c, int pc, int64_t sp, int64_t guard_sp, int32_t guard_pc, case OP_SAVE: { int64_t old = c->caps[in->data]; c->caps[in->data] = sp; - if (run(c, pc + 1, sp, guard_sp, guard_pc, depth + 1)) return 1; + if (run_memo(c, pc + 1, sp, guard_sp, guard_pc, depth + 1)) return 1; c->caps[in->data] = old; return 0; } case OP_SPLIT: if (in->is_loop) { - if (run(c, in->x, sp, sp, pc, depth + 1)) return 1; + if (run_memo(c, in->x, sp, sp, pc, depth + 1)) return 1; pc = in->y; continue; } else { - if (run(c, in->x, sp, guard_sp, guard_pc, depth + 1)) return 1; + if (run_memo(c, in->x, sp, guard_sp, guard_pc, depth + 1)) return 1; pc = in->y; continue; } case OP_JMP: @@ -934,7 +1006,7 @@ static int run(MCtx *c, int pc, int64_t sp, int64_t guard_sp, int32_t guard_pc, int64_t *snap = malloc(ncaps * sizeof(int64_t)); memcpy(snap, c->caps, ncaps * sizeof(int64_t)); int64_t save_req = c->require_end; c->require_end = -2; - int ok = run(c, in->x, sp, -1, -1, depth + 1); + int ok = run_memo(c, in->x, sp, -1, -1, depth + 1); c->require_end = save_req; int accept = in->neg ? !ok : ok; if (!accept) { memcpy(c->caps, snap, ncaps * sizeof(int64_t)); free(snap); return 0; } @@ -949,7 +1021,7 @@ static int run(MCtx *c, int pc, int64_t sp, int64_t guard_sp, int32_t guard_pc, int ok = 0; if (start >= 0) { int64_t save_req = c->require_end; c->require_end = sp; - ok = run(c, in->x, start, -1, -1, depth + 1); + ok = run_memo(c, in->x, start, -1, -1, depth + 1); c->require_end = save_req; } int accept = in->neg ? !ok : ok; @@ -959,7 +1031,7 @@ static int run(MCtx *c, int pc, int64_t sp, int64_t guard_sp, int32_t guard_pc, } case OP_ATOMIC: { int64_t save_req = c->require_end; c->require_end = -2; - int ok = run(c, in->x, sp, -1, -1, depth + 1); + int ok = run_memo(c, in->x, sp, -1, -1, depth + 1); c->require_end = save_req; if (!ok) return 0; sp = c->sub_end; @@ -979,6 +1051,25 @@ static int run(MCtx *c, int pc, int64_t sp, int64_t guard_sp, int32_t guard_pc, } } +/* The only call site allowed to consult/populate the memo (see the + * comment on MCtx.memo above): every recursive call inside run(), and + * every external entry point, goes through here instead of run() + * directly. Caching is restricted to guard_pc == -1 (not inside an + * active nullable-loop guard) and c->require_end == -1 (not inside a + * fullmatch's exact-end constraint, nor inside a nested lookaround/ + * atomic sub-program, which always sets require_end elsewhere first); + * outside that combination, whether (pc, sp) succeeds can depend on + * context beyond (pc, sp) itself, so it is never cached or consulted + * there, only computed directly by run(), exactly as before this + * optimization existed. */ +static int run_memo(MCtx *c, int pc, int64_t sp, int64_t guard_sp, int32_t guard_pc, int depth) { + int cacheable = c->memo && guard_pc == -1 && c->require_end == -1; + if (cacheable && memo_get(c, pc, sp)) return 0; + int result = run(c, pc, sp, guard_sp, guard_pc, depth); + if (cacheable && !result && !c->depth_exceeded) memo_set(c, pc, sp); + return result; +} + /* ================================================================ * 7. Pattern / Match / Input structures * ================================================================ */ @@ -1115,6 +1206,7 @@ Pattern *re_compile(const char *pattern, size_t len, int flags, PatternError *er da_init(&prog.classnodes, sizeof(Node *)); compile_node(&prog, ast); emit(&prog, OP_MATCH); + for (int k = 0; k < prog.n; k++) if (prog.insts[k].op == OP_BACKREF) { prog.has_backref = 1; break; } PatternImpl *impl = calloc(1, sizeof(PatternImpl)); impl->prog = prog; @@ -1279,6 +1371,112 @@ static void fill_match_from_caps(Match *out, Pattern *self, Input *string, int64 static int64_t clamp(int64_t v, int64_t lo, int64_t hi) { return v < lo ? lo : (v > hi ? hi : v); } +/* Allocated once per top-level Pattern_/re_ call (not once per start + * position tried, and, for Pattern_finditer/Pattern_split, not once + * per match found either): see the comment on MCtx.memo. NULL, with no + * behavior change beyond the missing optimization, whenever the + * pattern contains a backreference anywhere. */ +static uint8_t *alloc_memo(Prog *pr, int64_t textlen) { + if (pr->has_backref) return NULL; + int64_t bits = (int64_t)pr->n * (textlen + 1); + int64_t bytes = (bits + 7) / 8; + return calloc((size_t)(bytes > 0 ? bytes : 1), 1); +} + +/* mr[i] = how many consecutive positions starting at text[i] satisfy + * in's atom predicate (0 if text[i] itself does not), computed in one + * backward O(textlen) pass instead of redone forward, from scratch, at + * every position a caller asks about it. This is what makes + * OP_REPEAT1 O(1) per position instead of O(remaining run length), + * which otherwise stays quadratic across do_one/finditer/split's outer + * position loop even with run_memo's (pc, sp) memoization, because the + * counting scan is an internal C loop, not expressed as (pc, sp) + * recursive calls at all. int32_t bounds a single repeat's run length + * to ~2 billion, far past any input this build's fully-materializing + * Input can hold in memory in the first place. */ +static int32_t *compute_maxrun(Prog *pr, Inst *in, const uint32_t *text, int64_t len) { + int32_t *mr = malloc((size_t)(len + 1) * sizeof(int32_t)); + if (!mr) return NULL; + mr[len] = 0; + Node *cls = (in->atomkind == 1) ? ((Node **)pr->classnodes.data)[in->data] : NULL; + for (int64_t i = len - 1; i >= 0; i--) { + uint32_t ch = text[i]; + int m; + if (in->atomkind == 0) m = char_eq(ch, in->data, pr->mode, pr->flags); + else if (in->atomkind == 2) m = !(ch == '\n' && !(pr->flags & DOTALL)); + else m = class_match(cls, ch, pr->mode, pr->flags); + int64_t next = mr[i + 1]; + mr[i] = m ? (int32_t)(next < INT32_MAX ? next + 1 : INT32_MAX) : 0; + } + return mr; +} + +/* One table per OP_REPEAT1 instruction actually present in the + * program, indexed by instruction number; only worth the O(text + * length) memory per instruction for the multi-position operations + * (search/finditer/split) that would otherwise redo the scan at every + * position, so match/fullmatch (a single attempt) do not allocate it. */ +static int32_t **alloc_repeat_maxrun(Prog *pr, const uint32_t *text, int64_t len) { + int32_t **arr = calloc((size_t)pr->n, sizeof(int32_t *)); + if (!arr) return NULL; + for (int i = 0; i < pr->n; i++) + if (pr->insts[i].op == OP_REPEAT1) + arr[i] = compute_maxrun(pr, &pr->insts[i], text, len); + return arr; +} +static void free_repeat_maxrun(Prog *pr, int32_t **arr) { + if (!arr) return; + for (int i = 0; i < pr->n; i++) free(arr[i]); + free(arr); +} + +static int inst_atom_match(Prog *pr, Inst *in, uint32_t ch) { + switch (in->op) { + case OP_CHAR: return char_eq(ch, in->data, pr->mode, pr->flags); + case OP_ANY: return !(ch == '\n' && !(pr->flags & DOTALL)); + case OP_CLASS: return class_match(((Node **)pr->classnodes.data)[in->data], ch, pr->mode, pr->flags); + default: return 0; + } +} + +/* pm[i] = the largest position <= i where the instruction right after + * an OP_REPEAT1 matches, or -1 if none exists in [0, i]. Only computed + * when that next instruction is itself a single simple atom (OP_CHAR/ + * OP_CLASS/OP_ANY); this is what lets OP_REPEAT1's greedy backtrack + * jump straight from "the largest k worth trying" to "the next + * smaller one worth trying" instead of visiting every k in between + * (see the comment at its one call site). Built in one forward + * O(textlen) pass instead of walked freshly, backward, from every + * position a caller asks about it. */ +static int32_t *compute_next_prevmatch(Prog *pr, Inst *next, const uint32_t *text, int64_t len) { + if (len <= 0) return NULL; + int32_t *pm = malloc((size_t)len * sizeof(int32_t)); + if (!pm) return NULL; + int32_t last = -1; + for (int64_t i = 0; i < len; i++) { + if (inst_atom_match(pr, next, text[i])) last = (int32_t)i; + pm[i] = last; + } + return pm; +} +static int32_t **alloc_repeat_nextpm(Prog *pr, const uint32_t *text, int64_t len) { + int32_t **arr = calloc((size_t)pr->n, sizeof(int32_t *)); + if (!arr) return NULL; + for (int i = 0; i < pr->n; i++) { + if (pr->insts[i].op != OP_REPEAT1 || !pr->insts[i].greedy) continue; + if (i + 1 >= pr->n) continue; + int nextop = pr->insts[i + 1].op; + if (nextop == OP_CHAR || nextop == OP_ANY || nextop == OP_CLASS) + arr[i] = compute_next_prevmatch(pr, &pr->insts[i + 1], text, len); + } + return arr; +} +static void free_repeat_nextpm(Prog *pr, int32_t **arr) { + if (!arr) return; + for (int i = 0; i < pr->n; i++) free(arr[i]); + free(arr); +} + static int do_one(Pattern *self, Input *string, int64_t pos, int64_t endpos, int anchored, int fullmatch, Match *out) { MatBuf mb; if (!build_matbuf(self, string, &mb, NULL)) return -1; @@ -1287,16 +1485,27 @@ static int do_one(Pattern *self, Input *string, int64_t pos, int64_t endpos, int PatternImpl *impl = (PatternImpl *)self->program; int ng = impl->prog.ngroups; int64_t *caps; alloc_caps(&caps, ng); + uint8_t *memo = alloc_memo(&impl->prog, ep); + /* Only search (anchored == 0) tries more than one position, so + * only search pays for the run-length precompute (see + * compute_maxrun); match/fullmatch's single attempt gets no + * benefit from it and skips the O(text length) memory. */ + int32_t **maxrun = anchored ? NULL : alloc_repeat_maxrun(&impl->prog, mb.text, ep); + int32_t **nextpm = anchored ? NULL : alloc_repeat_nextpm(&impl->prog, mb.text, ep); int found = 0; int64_t last_start = anchored ? p0 : ep; for (int64_t start = p0; start <= last_start && !found; start++) { reset_caps(caps, ng); caps[0] = start; MCtx c; c.pr = &impl->prog; c.text = mb.text; c.len = ep; c.caps = caps; - c.depth_exceeded = 0; c.require_end = fullmatch ? ep : -1; c.sub_end = -1; - found = run(&c, 0, start, -1, -1, 0); - if (!found && c.depth_exceeded) { free(caps); free_matbuf(&mb); return -1; } + c.depth_exceeded = 0; c.require_end = fullmatch ? ep : -1; c.sub_end = -1; c.memo = memo; + c.repeat_maxrun = maxrun; c.repeat_nextpm = nextpm; + found = run_memo(&c, 0, start, -1, -1, 0); + if (!found && c.depth_exceeded) { free(caps); free(memo); free_repeat_maxrun(&impl->prog, maxrun); free_repeat_nextpm(&impl->prog, nextpm); free_matbuf(&mb); return -1; } } + free(memo); + free_repeat_maxrun(&impl->prog, maxrun); + free_repeat_nextpm(&impl->prog, nextpm); if (!found) { free(caps); free_matbuf(&mb); return 0; } fill_match_from_caps(out, self, string, p0, ep, caps, ng, &mb); free_matbuf(&mb); @@ -1320,6 +1529,14 @@ int Pattern_finditer(Pattern *self, Input *string, int64_t pos, int64_t endpos, int64_t start = clamp(pos, 0, mb.len); PatternImpl *impl = (PatternImpl *)self->program; int ng = impl->prog.ngroups; + /* One memo table for the whole scan: every position tried, across + * every match found, shares it (see the comment on MCtx.memo). A + * fact it records ("(pc, sp) cannot reach OP_MATCH") never becomes + * false later in the same scan, so nothing here ever needs to + * invalidate or reset it between matches. */ + uint8_t *memo = alloc_memo(&impl->prog, ep); + int32_t **maxrun = alloc_repeat_maxrun(&impl->prog, mb.text, ep); /* see compute_maxrun */ + int32_t **nextpm = alloc_repeat_nextpm(&impl->prog, mb.text, ep); /* see compute_next_prevmatch */ int count = 0; while (start <= ep) { int64_t *caps; alloc_caps(&caps, ng); @@ -1329,9 +1546,10 @@ int Pattern_finditer(Pattern *self, Input *string, int64_t pos, int64_t endpos, reset_caps(caps, ng); caps[0] = s; MCtx c; c.pr = &impl->prog; c.text = mb.text; c.len = ep; c.caps = caps; - c.depth_exceeded = 0; c.require_end = -1; c.sub_end = -1; - found = run(&c, 0, s, -1, -1, 0); - if (!found && c.depth_exceeded) { free(caps); free_matbuf(&mb); return -1; } + c.depth_exceeded = 0; c.require_end = -1; c.sub_end = -1; c.memo = memo; + c.repeat_maxrun = maxrun; c.repeat_nextpm = nextpm; + found = run_memo(&c, 0, s, -1, -1, 0); + if (!found && c.depth_exceeded) { free(caps); free(memo); free_repeat_maxrun(&impl->prog, maxrun); free_repeat_nextpm(&impl->prog, nextpm); free_matbuf(&mb); return -1; } } if (!found) { free(caps); break; } Match m; memset(&m, 0, sizeof m); @@ -1343,6 +1561,9 @@ int Pattern_finditer(Pattern *self, Input *string, int64_t pos, int64_t endpos, start = (mend > caps[0]) ? mend : caps[0] + 1; Match_free(&m); } + free(memo); + free_repeat_maxrun(&impl->prog, maxrun); + free_repeat_nextpm(&impl->prog, nextpm); free_matbuf(&mb); return count; } @@ -1372,6 +1593,9 @@ int Pattern_split(Pattern *self, Input *string, int maxsplit, MatchIterCb cb, vo int64_t start = 0, seg_start = 0; PatternImpl *impl = (PatternImpl *)self->program; int ng = impl->prog.ngroups; + uint8_t *memo = alloc_memo(&impl->prog, ep); /* one table for the whole scan, see MCtx.memo */ + int32_t **maxrun = alloc_repeat_maxrun(&impl->prog, mb.text, ep); /* see compute_maxrun */ + int32_t **nextpm = alloc_repeat_nextpm(&impl->prog, mb.text, ep); /* see compute_next_prevmatch */ int n = 0; while (start <= ep) { if (maxsplit > 0 && n >= maxsplit) break; @@ -1381,9 +1605,10 @@ int Pattern_split(Pattern *self, Input *string, int maxsplit, MatchIterCb cb, vo reset_caps(caps, ng); caps[0] = s; MCtx c; c.pr = &impl->prog; c.text = mb.text; c.len = ep; c.caps = caps; - c.depth_exceeded = 0; c.require_end = -1; c.sub_end = -1; - found = run(&c, 0, s, -1, -1, 0); - if (!found && c.depth_exceeded) { free(caps); free_matbuf(&mb); return -1; } + c.depth_exceeded = 0; c.require_end = -1; c.sub_end = -1; c.memo = memo; + c.repeat_maxrun = maxrun; c.repeat_nextpm = nextpm; + found = run_memo(&c, 0, s, -1, -1, 0); + if (!found && c.depth_exceeded) { free(caps); free(memo); free_repeat_maxrun(&impl->prog, maxrun); free_repeat_nextpm(&impl->prog, nextpm); free_matbuf(&mb); return -1; } } if (!found) { free(caps); break; } emit_split_item(cb, ctx, self, string, &mb, seg_start, caps[0]); @@ -1393,6 +1618,9 @@ int Pattern_split(Pattern *self, Input *string, int maxsplit, MatchIterCb cb, vo start = (caps[1] > caps[0]) ? caps[1] : caps[0] + 1; free(caps); } + free(memo); + free_repeat_maxrun(&impl->prog, maxrun); + free_repeat_nextpm(&impl->prog, nextpm); emit_split_item(cb, ctx, self, string, &mb, seg_start, ep); free_matbuf(&mb); return n;