Fix the search-family quadratic-time defect with memoized backtracking

Root cause: do_one/Pattern_finditer/Pattern_split tried every start
position as an independent from-scratch backtracking attempt, so a
pattern that ultimately fails, or matches only very late, redid the
bulk of its work at every position. Measured: a*b over a run of a's
with no b took 19.2s at n=80,000 (quadratic, confirmed by the ~4x
slowdown per doubling).

Two techniques fix it, researched against and matching published
prior art rather than invented ad hoc:

- run_memo caches every proven match failure at the (instruction, text
  position) level, never a success (so it cannot change which match is
  found), and is only enabled when the compiled program has no
  OP_BACKREF anywhere, since a backreference's outcome depends on
  capture history, not on position alone. This is the same "memoize
  only failures" scheme described in recent work on backtracking regex
  matchers (Selective Memoization for Efficient Backtracking Regular
  Expression Matching; Efficient Matching with Memoization for Regexes
  with Look-around and Atomic Grouping).
- compute_maxrun and compute_next_prevmatch precompute, once per
  Pattern_search/finditer/split call, how far a simple-atom quantifier
  (OP_REPEAT1) can run from any position and, when it is immediately
  followed by a single literal/class/., the rightmost position where
  that next atom can match. This lets the backtrack loop jump straight
  to candidates worth trying instead of visiting every position in
  between, the same idea RE2 and Rust's regex crate call a literal
  prefilter, implemented here with a precomputed array instead of a
  SIMD memchr/memmem call, in keeping with concept.md's convenience
  over performance priority.

Result: a*b at n=80,000 dropped from 19.2s to 0.0016s, now scaling
linearly. As a side effect, since it needs no backreference, the
textbook ReDoS pattern (a+)+b also went from exponential (already
impractical past n=40) to empirically quadratic (3.2s at n=32,000),
though not linear: the inner a+'s OP_REPEAT1 is followed by the
group's closing save rather than a simple atom, so the skip-ahead
table does not apply to it, only the failure memoization does.
README.md and docs/API.md are updated with the measured numbers, the
precise remaining limitations, and citations to the sources this was
checked against.

No test behavior changed: all 81 cases (checked against CPython's own
re module output) still pass, clean under AddressSanitizer/UBSan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
This commit is contained in:
2026-09-14 08:04:29 +00:00
co-authored by Claude Sonnet 5
parent 44f0b3eb39
commit 19a7f678a9
3 changed files with 359 additions and 53 deletions
+4 -4
View File
@@ -145,7 +145,7 @@ Mirror `re.Pattern.match`/`.fullmatch`/`.search` exactly, including the `pos`/`e
- `fullmatch`: anchored at `pos`, must also reach exactly `endpos`.
- `search`: tries every start position from `pos` to `endpos` inclusive, left to right, and reports the first that admits any match (ordinary backtracking priority decides which match that is at that position, `concept.md` 2.5).
**Performance:** `match`/`fullmatch` do one anchored attempt and cost time proportional to that attempt alone. `search` tries each candidate start position as an *independent* attempt from scratch, so it is quadratic, not linear, in the worst case (a pattern that does not match, or matches only very late) even for a pattern with no backreference and no unbounded-width lookahead; see README.md "Implementation status" for a measured example and the reason (`concept.md` Section 7.2's engine, which shares work across start positions in one linear pass, is not implemented yet).
**Performance:** `match`/`fullmatch` do one anchored attempt and cost time proportional to that attempt alone. `search` tries each candidate start position as a separate attempt, but, unlike a naive backtracking search, does not redo the same work at every one: for a backreference-free pattern, `run_memo` caches every proven failure at the (instruction, position) level, and `OP_REPEAT1` (a quantifier over a single character, class, or `.`) additionally uses precomputed run-length and skip-ahead tables so its own internal work is `O(1)` amortized per position rather than `O(remaining length)`. Together these make `search` linear, not quadratic, for the common case (a pattern with no backreference, built from simple repeated atoms). What is not fixed: a repeat over a *compound* body (`(ab)*`) gets the failure-memoization but not the skip-ahead table, so a pathological compound-repeat pattern can still be worse than linear; and any pattern with a backreference disables memoization entirely (unsound there, see Section 5) and can still be worst-case exponential, exactly as in CPython. See README.md "Implementation status" for the measured numbers on both the fixed case and the remaining one, and for citations to the published techniques this uses.
### 3.5-3.6 `Pattern_finditer` / `Pattern_findall`
@@ -154,7 +154,7 @@ int Pattern_finditer(Pattern *self, Input *string, int64_t pos, int64_t endpos,
int Pattern_findall(Pattern *self, Input *string, int64_t pos, int64_t endpos, MatchIterCb cb, void *ctx);
```
`Pattern_findall` is defined as a call to `Pattern_finditer` with the same arguments; both invoke `cb(ctx, m)` once per non-overlapping match, left to right, applying CPython's own empty-match rule (an empty match advances one unit and is not merged with an adjacent non-empty match at the same start, `concept.md` 3). Returns the number of matches found, or `-1` on error. Inherits `search`'s worst-case quadratic time, Section 3.2-3.4's performance note; each match found restarts the same from-scratch search for the next one. This build does not collapse a no-groups match down to "just the matched string" or a multi-group match to a Python tuple the way `re.findall` does at the Python level; the callback always receives a full `Match`, from which the caller reads whatever it needs via `Match_group`. This is a deliberate simplification: `findall`'s string/tuple collapsing is a Python-object-model convenience with no C equivalent to collapse into, so this build gives the caller the same, uniform `Match`-based access `finditer` does, and the two functions exist separately only for name-for-name parity with `re.findall`/`re.finditer`.
`Pattern_findall` is defined as a call to `Pattern_finditer` with the same arguments; both invoke `cb(ctx, m)` once per non-overlapping match, left to right, applying CPython's own empty-match rule (an empty match advances one unit and is not merged with an adjacent non-empty match at the same start, `concept.md` 3). Returns the number of matches found, or `-1` on error. Shares one memoization table and one set of `OP_REPEAT1` precomputed tables across the whole scan (built once, not once per match), so it inherits `search`'s performance characteristics in Section 3.2-3.4 exactly, not a worse case from repeating the scan. This build does not collapse a no-groups match down to "just the matched string" or a multi-group match to a Python tuple the way `re.findall` does at the Python level; the callback always receives a full `Match`, from which the caller reads whatever it needs via `Match_group`. This is a deliberate simplification: `findall`'s string/tuple collapsing is a Python-object-model convenience with no C equivalent to collapse into, so this build gives the caller the same, uniform `Match`-based access `finditer` does, and the two functions exist separately only for name-for-name parity with `re.findall`/`re.finditer`.
### 3.7 `Pattern_split`
@@ -162,7 +162,7 @@ int Pattern_findall(Pattern *self, Input *string, int64_t pos, int64_t endpos, M
int Pattern_split(Pattern *self, Input *string, int maxsplit, MatchIterCb cb, void *ctx);
```
`cb` is invoked once per element of the list `re.split()` would return, **in order**: this is the entire contract, and it is unambiguous by construction (earlier drafts of this function called `cb` twice per match with the caller left to infer which call meant what; that design was replaced before release specifically because it was ambiguous). Each element's text is read via `Match_group(m, NULL, 0, &out, &outlen)`; a `0` return from that call means this element is Python's `None` (an unparticipated capturing group between two matches), matching how an unparticipated group reports on any other `Match`. `maxsplit` matches `re.split`'s parameter (`0` means unlimited). Returns the number of matches that were split on (not the number of list elements), or `-1` on error. Built on the same from-scratch-per-position search as `Pattern_search`; inherits its worst-case quadratic time, Section 3.2-3.4.
`cb` is invoked once per element of the list `re.split()` would return, **in order**: this is the entire contract, and it is unambiguous by construction (earlier drafts of this function called `cb` twice per match with the caller left to infer which call meant what; that design was replaced before release specifically because it was ambiguous). Each element's text is read via `Match_group(m, NULL, 0, &out, &outlen)`; a `0` return from that call means this element is Python's `None` (an unparticipated capturing group between two matches), matching how an unparticipated group reports on any other `Match`. `maxsplit` matches `re.split`'s parameter (`0` means unlimited). Returns the number of matches that were split on (not the number of list elements), or `-1` on error. Shares one memoization table and one set of `OP_REPEAT1` precomputed tables across the whole scan; see Section 3.2-3.4's performance note.
### 3.8-3.9 `Pattern_sub` / `Pattern_subn`
@@ -177,7 +177,7 @@ Mirror `re.Pattern.sub`/`.subn`. Exactly one of `repl` (a template string) or `c
`count` matches `re.sub`'s `count` parameter (`0` means unlimited; a positive `count` stops substituting after that many matches, leaving the rest of the subject, including any further matches within it, untouched, exactly as CPython leaves it).
`Pattern_sub` and `Pattern_subn` differ only in whether the number of substitutions actually made is reported back through `n` (mirroring `re.sub` returning just the string versus `re.subn` returning `(string, count)`); `Pattern_sub` is implemented as a call to `Pattern_subn` with a throwaway `n`. Both are implemented on top of `Pattern_finditer` and so inherit its worst-case quadratic time, Section 3.2-3.4.
`Pattern_sub` and `Pattern_subn` differ only in whether the number of substitutions actually made is reported back through `n` (mirroring `re.sub` returning just the string versus `re.subn` returning `(string, count)`); `Pattern_sub` is implemented as a call to `Pattern_subn` with a throwaway `n`. Both are implemented on top of `Pattern_finditer` and so share its performance characteristics, Section 3.2-3.4's performance note.
`*out` is a freshly `malloc`'d, `NUL`-terminated buffer of length `*outlen`; the caller must `free` it. On `-1` (error), `*out`/`*outlen` are left untouched.