Commit Graph
3 Commits
Author SHA1 Message Date
retoorandClaude Sonnet 5 19a7f678a9 Fix the search-family quadratic-time defect with memoized backtracking
Root cause: do_one/Pattern_finditer/Pattern_split tried every start
position as an independent from-scratch backtracking attempt, so a
pattern that ultimately fails, or matches only very late, redid the
bulk of its work at every position. Measured: a*b over a run of a's
with no b took 19.2s at n=80,000 (quadratic, confirmed by the ~4x
slowdown per doubling).

Two techniques fix it, researched against and matching published
prior art rather than invented ad hoc:

- run_memo caches every proven match failure at the (instruction, text
  position) level, never a success (so it cannot change which match is
  found), and is only enabled when the compiled program has no
  OP_BACKREF anywhere, since a backreference's outcome depends on
  capture history, not on position alone. This is the same "memoize
  only failures" scheme described in recent work on backtracking regex
  matchers (Selective Memoization for Efficient Backtracking Regular
  Expression Matching; Efficient Matching with Memoization for Regexes
  with Look-around and Atomic Grouping).
- compute_maxrun and compute_next_prevmatch precompute, once per
  Pattern_search/finditer/split call, how far a simple-atom quantifier
  (OP_REPEAT1) can run from any position and, when it is immediately
  followed by a single literal/class/., the rightmost position where
  that next atom can match. This lets the backtrack loop jump straight
  to candidates worth trying instead of visiting every position in
  between, the same idea RE2 and Rust's regex crate call a literal
  prefilter, implemented here with a precomputed array instead of a
  SIMD memchr/memmem call, in keeping with concept.md's convenience
  over performance priority.

Result: a*b at n=80,000 dropped from 19.2s to 0.0016s, now scaling
linearly. As a side effect, since it needs no backreference, the
textbook ReDoS pattern (a+)+b also went from exponential (already
impractical past n=40) to empirically quadratic (3.2s at n=32,000),
though not linear: the inner a+'s OP_REPEAT1 is followed by the
group's closing save rather than a simple atom, so the skip-ahead
table does not apply to it, only the failure memoization does.
README.md and docs/API.md are updated with the measured numbers, the
precise remaining limitations, and citations to the sources this was
checked against.

No test behavior changed: all 81 cases (checked against CPython's own
re module output) still pass, clean under AddressSanitizer/UBSan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
2026-09-14 08:04:29 +00:00
retoorandClaude Sonnet 5 44f0b3eb39 Correct a false linear-time claim: search-family ops are quadratic
Auditing the documentation against the actual code (not against what I
remembered writing) surfaced a real, previously undocumented defect:
README.md claimed a backreference-free, unbounded-lookahead-free
pattern "runs in linear time", but that is only true of Pattern_match/
Pattern_fullmatch (one anchored attempt). Pattern_search tries every
candidate start position as an independent from-scratch attempt, so it
is quadratic in the worst case even for the simplest pattern, since
concept.md Section 7.2's engine (which shares work across start
positions in one linear pass) is not implemented yet. Measured directly
with a*b over a run of plain a's: 0.32s at n=10,000, 19.2s at n=80,000,
an ~4x slowdown per doubling. Pattern_finditer/Pattern_split/
Pattern_sub all build on the same search loop and inherit it.

README.md and docs/API.md now state this precisely, with the measured
numbers, wherever the affected functions are documented, rather than
repeating the incorrect blanket "linear time" claim. No code changed;
this is a documentation correction only.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
2026-09-14 07:02:26 +00:00
retoorandClaude Sonnet 5 8f6afd6cd4 Add regexx: a single-file C regex interpreter with Python re semantics
concept.md is the full design document: the objective (Python re parity
plus binary/ASCII/UTF-8 modes, gigabyte-scale input, single C file), the
automata-theory argument for why unrestricted backreferences/lookaround
are incompatible with strict single-pass constant memory, the resulting
two-engine architecture, the exact Python-mirroring naming convention,
and a full comparison against POSIX regex.h for C-background readers.

regexx.c/regexx.h are the v1 implementation: parser, compiler to a
Pike/backtracking-style bytecode, and a single recursive backtracking
engine covering the pattern syntax and operations listed in README.md,
validated against CPython's own re module output (tests/), clean under
AddressSanitizer/UBSan, and stress-tested (50MB simple-quantifier match,
graceful failure rather than a crash on complex repeats over large
input, clean rejection of every intentionally unsupported construct).

Also included: examples/rxgrep.c (a small grep-like program exercising
all three data modes and the substitution API), the Makefile, the MIT
LICENSE, and docs/API.md, an exhaustive reference for every type, flag,
and function's exact return-value and memory-ownership convention,
checked against the current source and against a real CPython
interpreter rather than against memory.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
2026-09-14 05:56:16 +00:00