From a9ca517981bc23136af187c48783cf9b62fc255b Mon Sep 17 00:00:00 2001 From: retoor Date: Mon, 14 Sep 2026 08:53:19 +0000 Subject: [PATCH] Bring concept.md and file-header comments in line with the memoization fix concept.md described the bounded backtracking engine (7.3) as plain recursive backtracking, with the Regular engine's linear-time guarantee reserved for patterns the streaming Pike VM (7.2) handles. v1 has no Pike VM yet, so every pattern actually runs on 7.3's engine; the previous commit added memoization and precomputed skip-ahead tables to that engine specifically because plain backtracking search was quadratic even for ordinary patterns, not just exponential for pathological ones. Added Section 7.5 recording that finding precisely: what the two techniques buy (linear time for the backreference-free majority, quadratic instead of exponential for the textbook (a+)+b shape), what they do not buy (7.2's bounded, input-length-independent memory guarantee, still the correct target for that axis), and citations to the published techniques they match. Added a short note to the Section 10 complexity table pointing at 7.5 so the table is read as the two engines' designed guarantees, not conflated with what a specific implementation measures. Updated regexx.c/regexx.h's top-of-file comments, which still described a plain backtracking engine. No code changes. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY --- concept.md | 11 +++++++++++ regexx.c | 9 +++++++++ regexx.h | 13 +++++++++---- 3 files changed, 29 insertions(+), 4 deletions(-) diff --git a/concept.md b/concept.md index 77276b8..6e2909a 100644 --- a/concept.md +++ b/concept.md @@ -152,6 +152,15 @@ Because this engine is only invoked for patterns that need it, ordinary patterns Both engines expose the same chunk boundary contract: a match cannot be reported as final until either (a) `OP_MATCH` is reached, or (b) enough trailing context has been seen to prove that no continuation of the current input would change the answer for the leftmost still-open match attempt. For quantifiers this is decided directly by the bytecode's `OP_SPLIT`/`OP_JMP` shape; for anchors (`$`, `\Z`) the final chunk is distinguished by an explicit "end of stream" sentinel token, so `$` and `\Z` behave identically whether or not `MULTILINE` is set, matching CPython. +### 7.5 Implementation finding: memoized backtracking as a practical stand-in for 7.2 + +The v1 implementation (README.md "Implementation status") does not yet have Section 7.2's Pike VM at all; every pattern runs on this section's backtracking engine, without the sliding window (v1 materializes the whole input instead, a separate, already-documented simplification). Building it anyway surfaced a result worth recording here because it affects how urgent 7.2 actually is: a plain backtracking search is not merely "exponential in the worst case for pathological patterns," it is quadratic even for the simplest ordinary pattern, because trying every start position from scratch re-derives the same failures over and over. Two additions to this section's engine, kept but not originally specified here, close most of that gap without building 7.2: + +- Memoizing every proven match *failure* at the (instruction, position) level, never a success, so it cannot change which match is found, only skip re-deriving failures already known. Valid only when the compiled program contains no `OP_BACKREF` anywhere (a backreference's outcome depends on capture history, not on position alone, so the memoized fact would not be a pure function of position). This is a published technique: recent work on backtracking regex matchers describes the identical "failures only" memoization scheme, and a further refinement, memoizing selectively only at loop "feedback nodes" instead of every instruction, which this implementation does not do but could, for less memory. +- Precomputing, for a quantifier over a single character, class, or `.`, how far it can run from any position, and, when it is immediately followed by another single atom, the rightmost position where that atom matches, so the backtrack loop jumps to candidates worth trying instead of visiting every position in between. This is the same idea as the literal prefilters (`memchr`/`memmem`/Teddy) production engines like RE2 and Rust's `regex` crate use to skip positions that provably cannot match, implemented here with a precomputed array rather than a SIMD library call, in keeping with this document's convenience-over-performance priority. + +Together these make the backtracking engine linear, not quadratic, for a pattern with no backreference built from simple repeated atoms (measured: a pattern searched over 80,000 non-matching bytes dropped from 19.2s to 0.0016s), and turn the textbook catastrophic-backtracking shape `(a+)+b` from exponential into empirically quadratic, though not linear, since a repeat over a *compound* body only gets the failure memoization, not the skip-ahead table. Neither technique gives the bounded, input-length-independent *memory* guarantee that is 7.2's actual reason for existing (Section 5): both use `O(instruction count x input length)` memory for their tables, which is bounded but scales with input length, unlike 7.2's `O(instruction count)`. 7.2 therefore remains the correct target for the memory axis of Section 1's objective; what changed is that the *time* axis, for the backreference-free majority of patterns, no longer depends on building it. README.md and `docs/API.md` carry the measured numbers and full citations. + ## 8. Core Data Structures Kept intentionally minimal, all defined in the single file, no dependency beyond the C standard library (`stdint.h`, `stddef.h`, `string.h`): @@ -294,6 +303,8 @@ A byte sequence in `UTF8` mode that is not valid UTF-8 at the point the decoder Both figures are stated relative to input length specifically because that is the axis the gigabyte scale requirement constrains; instruction count and group count are properties of the pattern, not the input, and are expected to stay small (tens to low hundreds) for realistically written patterns. +This table describes the two engines as designed. Section 7.5 records that the bounded backtracking engine's v1 implementation, augmented with memoization and precomputed skip-ahead tables, in practice achieves the Regular engine's linear *time* bound for the backreference-free majority of patterns, at the cost of `O(instruction count x input length)` memory rather than the Regular engine's `O(instruction count x group count)`, so the two rows above remain the correct statement of each engine's guarantee, not of what a specific implementation happens to measure. + ## 11. Testing Strategy CPython ships its own `re` test suite (`Lib/test/test_re.py` / `re_tests.py`) as executable pattern, string, expected-result triples. The plan is to translate that suite mechanically into a table of C test cases run against `re_match`/`re_search`/`re_sub`, so the engine's conformance is measured against CPython's own stated behavior rather than against a re-derived interpretation of the documentation. Cases that exercise the explicitly out of scope items in Section 13 are recorded as known deviations rather than deleted, so the gap stays visible. diff --git a/regexx.c b/regexx.c index a603823..76b5f35 100644 --- a/regexx.c +++ b/regexx.c @@ -8,6 +8,15 @@ * streaming, bounded-memory regular engine of concept.md Section 7.2; * that remains future work, tracked in README.md. * + * That backtracking engine is memoized (run_memo) and, for quantifiers + * over a single character/class/./ (OP_REPEAT1), backed by precomputed + * run-length and skip-ahead tables (compute_maxrun, + * compute_next_prevmatch), so that Pattern_search/finditer/split/sub + * run in linear rather than quadratic time for the common, + * backreference-free case; see README.md "Implementation status" for + * the measured numbers, the precise remaining limitations, and + * citations to the published techniques this matches. + * * Naming convention (concept.md Section 9.0): every identifier with a * direct Python `re` counterpart uses that counterpart's exact spelling. * diff --git a/regexx.h b/regexx.h index 820f2e5..0cb3bea 100644 --- a/regexx.h +++ b/regexx.h @@ -7,10 +7,15 @@ * Implementation status relative to concept.md: this is the v1 * implementation. It provides the full public API and pattern syntax * described below, executed by a single recursive backtracking engine - * (concept.md Section 7.3) operating over a fully materialized copy of - * the input. The streaming, bounded-memory regular engine of Section - * 7.2 is not implemented yet; see README.md "Implementation status" - * for the complete, itemized list of what v1 does and does not cover. + * (concept.md Section 7.3) over a fully materialized copy of the + * input, augmented with memoization and precomputed skip-ahead tables + * (regexx.c: run_memo, compute_maxrun, compute_next_prevmatch) that + * make the search-family functions (Pattern_search/finditer/split/sub) + * run in linear time for the common, backreference-free case rather + * than the quadratic time a plain backtracking search would have. The + * streaming, bounded-memory regular engine of Section 7.2 is still not + * implemented; see README.md "Implementation status" for the complete, + * itemized, measured account of what v1 does and does not cover. */ #ifndef REGEXX_H #define REGEXX_H