Expand the test suite to 2,247 combinatorial cases, fix two real bugs it found
tests/TEST_PLAN.md records the plan before the result: seven categories (quantifier x grouping x flags, alternation x backreference x groups, lookaround combinatorics, sub/split specifics, real-world combination patterns, mode cross-checks, flag combination stress), each generated programmatically in tests/cases.py rather than hand-typed, still every case checked against a real CPython re result computed by tests/gen.py, per concept.md Section 11's existing strategy, just at 2,247 cases instead of 81. Running it immediately found two real, previously latent bugs, not just confirmed correctness of what was already covered by hand: - tests/gen.py's own C string-literal encoder had a trigraph bug: any generated pattern containing `??)` (the lazy quantifier next to a closing paren) was silently rewritten by the C compiler from 7 bytes to 5 before the suite ever ran, confirmed directly by compiling and printing the corrupted string. Fixed by escaping '?' as '\?', which is always safe and makes trigraph formation impossible. - Pattern_finditer/Pattern_split had a real, previously undocumented correctness defect: CPython's empty-match handling additionally searches for, and reports, a second, non-empty match at the same start position whenever the natural match found there was empty (reverse engineered against a real interpreter, since this is not written down in CPython's own documentation; \d*? against "123abc456" yields 16 matches, not 9). Fixed with a new MCtx forbid_empty flag that forces exactly that second search by rejecting the empty solution at OP_MATCH and letting ordinary backtracking find the next alternative, wired into both functions (they have independent scan loops). concept.md Section 3 and docs/API.md now state the rule precisely instead of the previous, incomplete description. A third, more mundane finding: the combinatorial mode cross-check category surfaced that ASCII-mode ground truth was being computed wrong in tests/gen.py itself (Python str + re.ASCII, which stays in code-point space, instead of a bytes pattern against a bytes subject, which is what regexx's byte-oriented ASCII mode actually is), and separately surfaced a genuine, now precisely documented Unicode-table gap already anticipated in principle by README.md's "Known deviations" (glibc's iswspace() under the C.utf8 locale does not classify U+00A0 NO-BREAK SPACE as whitespace; CPython's \s does). All 2,247 cases pass, clean under AddressSanitizer/UndefinedBehavior- Sanitizer; the 50MB/quadratic-time and ReDoS-scaling benchmarks from the previous two commits are unaffected. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
This commit is contained in:
+2
-2
@@ -154,7 +154,7 @@ int Pattern_finditer(Pattern *self, Input *string, int64_t pos, int64_t endpos,
|
||||
int Pattern_findall(Pattern *self, Input *string, int64_t pos, int64_t endpos, MatchIterCb cb, void *ctx);
|
||||
```
|
||||
|
||||
`Pattern_findall` is defined as a call to `Pattern_finditer` with the same arguments; both invoke `cb(ctx, m)` once per non-overlapping match, left to right, applying CPython's own empty-match rule (an empty match advances one unit and is not merged with an adjacent non-empty match at the same start, `concept.md` 3). Returns the number of matches found, or `-1` on error. Shares one memoization table and one set of `OP_REPEAT1` precomputed tables across the whole scan (built once, not once per match), so it inherits `search`'s performance characteristics in Section 3.2-3.4 exactly, not a worse case from repeating the scan. This build does not collapse a no-groups match down to "just the matched string" or a multi-group match to a Python tuple the way `re.findall` does at the Python level; the callback always receives a full `Match`, from which the caller reads whatever it needs via `Match_group`. This is a deliberate simplification: `findall`'s string/tuple collapsing is a Python-object-model convenience with no C equivalent to collapse into, so this build gives the caller the same, uniform `Match`-based access `finditer` does, and the two functions exist separately only for name-for-name parity with `re.findall`/`re.finditer`.
|
||||
`Pattern_findall` is defined as a call to `Pattern_finditer` with the same arguments; both invoke `cb(ctx, m)` once per non-overlapping match, left to right, applying CPython's own empty-match rule exactly, including the part of it CPython does not document (`concept.md` 3, `regexx.c`'s `Pattern_finditer` comment): if the match found at a given start position is empty, a second, non-empty match is additionally searched for and reported at that same start before the scan moves on, so `\d*?` against `"123abc456"` yields 16 matches, not 9, matching a real CPython interpreter exactly (verified by generating this case's expectation from one, `tests/cases.py`). Returns the number of matches found, or `-1` on error. Shares one memoization table and one set of `OP_REPEAT1` precomputed tables across the whole scan (built once, not once per match), so it inherits `search`'s performance characteristics in Section 3.2-3.4 exactly, not a worse case from repeating the scan. This build does not collapse a no-groups match down to "just the matched string" or a multi-group match to a Python tuple the way `re.findall` does at the Python level; the callback always receives a full `Match`, from which the caller reads whatever it needs via `Match_group`. This is a deliberate simplification: `findall`'s string/tuple collapsing is a Python-object-model convenience with no C equivalent to collapse into, so this build gives the caller the same, uniform `Match`-based access `finditer` does, and the two functions exist separately only for name-for-name parity with `re.findall`/`re.finditer`.
|
||||
|
||||
### 3.7 `Pattern_split`
|
||||
|
||||
@@ -162,7 +162,7 @@ int Pattern_findall(Pattern *self, Input *string, int64_t pos, int64_t endpos, M
|
||||
int Pattern_split(Pattern *self, Input *string, int maxsplit, MatchIterCb cb, void *ctx);
|
||||
```
|
||||
|
||||
`cb` is invoked once per element of the list `re.split()` would return, **in order**: this is the entire contract, and it is unambiguous by construction (earlier drafts of this function called `cb` twice per match with the caller left to infer which call meant what; that design was replaced before release specifically because it was ambiguous). Each element's text is read via `Match_group(m, NULL, 0, &out, &outlen)`; a `0` return from that call means this element is Python's `None` (an unparticipated capturing group between two matches), matching how an unparticipated group reports on any other `Match`. `maxsplit` matches `re.split`'s parameter (`0` means unlimited). Returns the number of matches that were split on (not the number of list elements), or `-1` on error. Shares one memoization table and one set of `OP_REPEAT1` precomputed tables across the whole scan; see Section 3.2-3.4's performance note.
|
||||
`cb` is invoked once per element of the list `re.split()` would return, **in order**: this is the entire contract, and it is unambiguous by construction (earlier drafts of this function called `cb` twice per match with the caller left to infer which call meant what; that design was replaced before release specifically because it was ambiguous). Each element's text is read via `Match_group(m, NULL, 0, &out, &outlen)`; a `0` return from that call means this element is Python's `None` (an unparticipated capturing group between two matches), matching how an unparticipated group reports on any other `Match`. `maxsplit` matches `re.split`'s parameter (`0` means unlimited). Returns the number of matches that were split on (not the number of list elements), or `-1` on error. Built on the same scan as `Pattern_finditer` and follows the same empty-match rule (Section 3.5-3.6), which is why a pattern that can match empty (`\d*?`, `x*`, and so on) produces the long runs of empty-string list elements a real CPython `re.split()` does, not a shorter list that only advances once per empty match. Shares one memoization table and one set of `OP_REPEAT1` precomputed tables across the whole scan; see Section 3.2-3.4's performance note.
|
||||
|
||||
### 3.8-3.9 `Pattern_sub` / `Pattern_subn`
|
||||
|
||||
|
||||
Reference in New Issue
Block a user