Root cause: do_one/Pattern_finditer/Pattern_split tried every start
position as an independent from-scratch backtracking attempt, so a
pattern that ultimately fails, or matches only very late, redid the
bulk of its work at every position. Measured: a*b over a run of a's
with no b took 19.2s at n=80,000 (quadratic, confirmed by the ~4x
slowdown per doubling).
Two techniques fix it, researched against and matching published
prior art rather than invented ad hoc:
- run_memo caches every proven match failure at the (instruction, text
position) level, never a success (so it cannot change which match is
found), and is only enabled when the compiled program has no
OP_BACKREF anywhere, since a backreference's outcome depends on
capture history, not on position alone. This is the same "memoize
only failures" scheme described in recent work on backtracking regex
matchers (Selective Memoization for Efficient Backtracking Regular
Expression Matching; Efficient Matching with Memoization for Regexes
with Look-around and Atomic Grouping).
- compute_maxrun and compute_next_prevmatch precompute, once per
Pattern_search/finditer/split call, how far a simple-atom quantifier
(OP_REPEAT1) can run from any position and, when it is immediately
followed by a single literal/class/., the rightmost position where
that next atom can match. This lets the backtrack loop jump straight
to candidates worth trying instead of visiting every position in
between, the same idea RE2 and Rust's regex crate call a literal
prefilter, implemented here with a precomputed array instead of a
SIMD memchr/memmem call, in keeping with concept.md's convenience
over performance priority.
Result: a*b at n=80,000 dropped from 19.2s to 0.0016s, now scaling
linearly. As a side effect, since it needs no backreference, the
textbook ReDoS pattern (a+)+b also went from exponential (already
impractical past n=40) to empirically quadratic (3.2s at n=32,000),
though not linear: the inner a+'s OP_REPEAT1 is followed by the
group's closing save rather than a simple atom, so the skip-ahead
table does not apply to it, only the failure memoization does.
README.md and docs/API.md are updated with the measured numbers, the
precise remaining limitations, and citations to the sources this was
checked against.
No test behavior changed: all 81 cases (checked against CPython's own
re module output) still pass, clean under AddressSanitizer/UBSan.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY