From 51f8ed19e9192390e265b1f7f4556da6a49bfc6c Mon Sep 17 00:00:00 2001 From: retoor Date: Mon, 14 Sep 2026 09:09:04 +0000 Subject: [PATCH] Expand the test suite to 2,247 combinatorial cases, fix two real bugs it found tests/TEST_PLAN.md records the plan before the result: seven categories (quantifier x grouping x flags, alternation x backreference x groups, lookaround combinatorics, sub/split specifics, real-world combination patterns, mode cross-checks, flag combination stress), each generated programmatically in tests/cases.py rather than hand-typed, still every case checked against a real CPython re result computed by tests/gen.py, per concept.md Section 11's existing strategy, just at 2,247 cases instead of 81. Running it immediately found two real, previously latent bugs, not just confirmed correctness of what was already covered by hand: - tests/gen.py's own C string-literal encoder had a trigraph bug: any generated pattern containing `??)` (the lazy quantifier next to a closing paren) was silently rewritten by the C compiler from 7 bytes to 5 before the suite ever ran, confirmed directly by compiling and printing the corrupted string. Fixed by escaping '?' as '\?', which is always safe and makes trigraph formation impossible. - Pattern_finditer/Pattern_split had a real, previously undocumented correctness defect: CPython's empty-match handling additionally searches for, and reports, a second, non-empty match at the same start position whenever the natural match found there was empty (reverse engineered against a real interpreter, since this is not written down in CPython's own documentation; \d*? against "123abc456" yields 16 matches, not 9). Fixed with a new MCtx forbid_empty flag that forces exactly that second search by rejecting the empty solution at OP_MATCH and letting ordinary backtracking find the next alternative, wired into both functions (they have independent scan loops). concept.md Section 3 and docs/API.md now state the rule precisely instead of the previous, incomplete description. A third, more mundane finding: the combinatorial mode cross-check category surfaced that ASCII-mode ground truth was being computed wrong in tests/gen.py itself (Python str + re.ASCII, which stays in code-point space, instead of a bytes pattern against a bytes subject, which is what regexx's byte-oriented ASCII mode actually is), and separately surfaced a genuine, now precisely documented Unicode-table gap already anticipated in principle by README.md's "Known deviations" (glibc's iswspace() under the C.utf8 locale does not classify U+00A0 NO-BREAK SPACE as whitespace; CPython's \s does). All 2,247 cases pass, clean under AddressSanitizer/UndefinedBehavior- Sanitizer; the 50MB/quadratic-time and ReDoS-scaling benchmarks from the previous two commits are unaffected. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY --- README.md | 33 ++++++-- concept.md | 2 +- docs/API.md | 4 +- regexx.c | 68 +++++++++++++-- tests/TEST_PLAN.md | 109 ++++++++++++++++++++++++ tests/cases.py | 207 +++++++++++++++++++++++++++++++++++++++++++++ tests/gen.py | 62 +++++++++++--- 7 files changed, 455 insertions(+), 30 deletions(-) create mode 100644 tests/TEST_PLAN.md diff --git a/README.md b/README.md index 0524b14..c42bbb6 100644 --- a/README.md +++ b/README.md @@ -182,7 +182,15 @@ convention they follow. hand-generated Unicode table (`concept.md` 13.3 anticipated a reduced static table; using the C library's own tables turned out to be simpler and more complete, at the cost of depending on the platform's Unicode - version rather than a pinned one). + version rather than a pinned one). Concretely verified, not just + theoretical: U+00A0 (NO-BREAK SPACE) is in Unicode's `White_Space` + property, so CPython's `\s` matches it, but glibc's `iswspace()` under + `C.utf8` does not, so this build's `\s` does not either. Found by the + large combinatorial test expansion (`tests/cases.py` Category F, + `tests/TEST_PLAN.md`), not anticipated in advance; recorded here rather + than patched, since hand-patching individual code points would start + down the path of maintaining an ad hoc table this design deliberately + avoided by delegating to `wctype.h` in the first place. - `lastindex`/`lastgroup` report the highest-numbered capturing group that participated in the match, which coincides with CPython's "most recently closed group" rule for straightforward patterns but can differ from it in @@ -272,13 +280,22 @@ Run `./rxgrep --help` for the full option list. ## Testing -`tests/cases.py` lists pattern/subject/operation triples. `tests/gen.py` -computes each one's expected result with CPython's own `re` module and -writes `tests/generated_tests.c`, which is then compiled against `regexx.c` -and checked. This is a direct implementation of the strategy `concept.md` -Section 11 describes: conformance is measured against what CPython actually -does, not against a re-derived reading of its documentation. `make check` -additionally runs the suite under AddressSanitizer and +`tests/cases.py` lists pattern/subject/operation triples, both hand-written +and, for most of the file, generated programmatically from combinations of +quantifiers, groups, backreferences, lookaround, flags, and encoding modes +(`tests/TEST_PLAN.md` records the exact category breakdown and why each one +exists). `tests/gen.py` computes every case's expected result with +CPython's own `re` module and writes `tests/generated_tests.c`, which is +then compiled against `regexx.c` and checked. This is a direct +implementation of the strategy `concept.md` Section 11 describes: +conformance is measured against what CPython actually does, not against a +re-derived reading of its documentation, at a scale (2,247 cases as of this +writing, `make test` reports the current exact count) large enough that it +has already found real defects this way, not only confirmed the absence of +ones anyone thought to write by hand (`tests/TEST_PLAN.md` "Result" names +both: a C trigraph bug in the test generator itself, and a real, previously +undocumented `Pattern_finditer`/`Pattern_split` empty-match defect). `make +check` additionally runs the suite under AddressSanitizer and UndefinedBehaviorSanitizer. ## License diff --git a/concept.md b/concept.md index 6e2909a..8317e36 100644 --- a/concept.md +++ b/concept.md @@ -50,7 +50,7 @@ The engine parses the following syntax, matching CPython's documented and observ - `match(pattern, string)`: anchored at position 0, not required to consume the whole string. - `fullmatch(pattern, string)`: anchored at position 0 and at the end of the string. - `search(pattern, string)`: first match anywhere. -- `finditer` / `findall`: all non overlapping matches, left to right, with the CPython 3.7+ rule that an empty match advances one position and is followed by a search starting immediately after it rather than being merged with an adjacent non empty match at the same start position. +- `finditer` / `findall`: all non overlapping matches, left to right. The precise CPython 3.7+ empty-match rule, reverse engineered against a real interpreter since it is not fully spelled out in CPython's own documentation: at each scan position, the ordinary (greedy/lazy, as the pattern specifies) match is found and reported first; if that match was empty, a *second* match, requiring non-empty this time, is additionally searched for and reported at that same starting position if one exists (confirmed with `\d*?` against `"123abc456"`: CPython's `finditer` yields `(0,0)`, then `(0,1)`, then `(1,1)`, then `(1,2)`, and so on, an empty match immediately followed by a non-empty one at the same start, not two separate scan positions); the scan then advances to the end of whichever of the two was reported last, or by one position if only the empty match existed. `split` and `sub`/`subn` are built on the same scan and inherit the same rule. - `split(pattern, string, maxsplit=0)`: text of capturing groups is interleaved into the result list, matching CPython; splitting on a pattern that can match an empty string is permitted, matching the CPython 3.7+ behavior change. - `sub(pattern, repl, string, count=0)` / `subn`: `repl` is either a template string honoring `\g`, `\g<1>`, `\1`, and literal backslash escapes, or a callback invoked once per match with a match record and expected to return replacement bytes/text (the C equivalent of a Python callable, see 9.4). - `escape(string)`: backslash escaping of all characters outside `[A-Za-z0-9_]` in the same way `re.escape` does since Python 3.7 (that version narrowed the escaped set relative to earlier Python releases; the engine follows the narrowed, current set). diff --git a/docs/API.md b/docs/API.md index bc3c0b5..3967f59 100644 --- a/docs/API.md +++ b/docs/API.md @@ -154,7 +154,7 @@ int Pattern_finditer(Pattern *self, Input *string, int64_t pos, int64_t endpos, int Pattern_findall(Pattern *self, Input *string, int64_t pos, int64_t endpos, MatchIterCb cb, void *ctx); ``` -`Pattern_findall` is defined as a call to `Pattern_finditer` with the same arguments; both invoke `cb(ctx, m)` once per non-overlapping match, left to right, applying CPython's own empty-match rule (an empty match advances one unit and is not merged with an adjacent non-empty match at the same start, `concept.md` 3). Returns the number of matches found, or `-1` on error. Shares one memoization table and one set of `OP_REPEAT1` precomputed tables across the whole scan (built once, not once per match), so it inherits `search`'s performance characteristics in Section 3.2-3.4 exactly, not a worse case from repeating the scan. This build does not collapse a no-groups match down to "just the matched string" or a multi-group match to a Python tuple the way `re.findall` does at the Python level; the callback always receives a full `Match`, from which the caller reads whatever it needs via `Match_group`. This is a deliberate simplification: `findall`'s string/tuple collapsing is a Python-object-model convenience with no C equivalent to collapse into, so this build gives the caller the same, uniform `Match`-based access `finditer` does, and the two functions exist separately only for name-for-name parity with `re.findall`/`re.finditer`. +`Pattern_findall` is defined as a call to `Pattern_finditer` with the same arguments; both invoke `cb(ctx, m)` once per non-overlapping match, left to right, applying CPython's own empty-match rule exactly, including the part of it CPython does not document (`concept.md` 3, `regexx.c`'s `Pattern_finditer` comment): if the match found at a given start position is empty, a second, non-empty match is additionally searched for and reported at that same start before the scan moves on, so `\d*?` against `"123abc456"` yields 16 matches, not 9, matching a real CPython interpreter exactly (verified by generating this case's expectation from one, `tests/cases.py`). Returns the number of matches found, or `-1` on error. Shares one memoization table and one set of `OP_REPEAT1` precomputed tables across the whole scan (built once, not once per match), so it inherits `search`'s performance characteristics in Section 3.2-3.4 exactly, not a worse case from repeating the scan. This build does not collapse a no-groups match down to "just the matched string" or a multi-group match to a Python tuple the way `re.findall` does at the Python level; the callback always receives a full `Match`, from which the caller reads whatever it needs via `Match_group`. This is a deliberate simplification: `findall`'s string/tuple collapsing is a Python-object-model convenience with no C equivalent to collapse into, so this build gives the caller the same, uniform `Match`-based access `finditer` does, and the two functions exist separately only for name-for-name parity with `re.findall`/`re.finditer`. ### 3.7 `Pattern_split` @@ -162,7 +162,7 @@ int Pattern_findall(Pattern *self, Input *string, int64_t pos, int64_t endpos, M int Pattern_split(Pattern *self, Input *string, int maxsplit, MatchIterCb cb, void *ctx); ``` -`cb` is invoked once per element of the list `re.split()` would return, **in order**: this is the entire contract, and it is unambiguous by construction (earlier drafts of this function called `cb` twice per match with the caller left to infer which call meant what; that design was replaced before release specifically because it was ambiguous). Each element's text is read via `Match_group(m, NULL, 0, &out, &outlen)`; a `0` return from that call means this element is Python's `None` (an unparticipated capturing group between two matches), matching how an unparticipated group reports on any other `Match`. `maxsplit` matches `re.split`'s parameter (`0` means unlimited). Returns the number of matches that were split on (not the number of list elements), or `-1` on error. Shares one memoization table and one set of `OP_REPEAT1` precomputed tables across the whole scan; see Section 3.2-3.4's performance note. +`cb` is invoked once per element of the list `re.split()` would return, **in order**: this is the entire contract, and it is unambiguous by construction (earlier drafts of this function called `cb` twice per match with the caller left to infer which call meant what; that design was replaced before release specifically because it was ambiguous). Each element's text is read via `Match_group(m, NULL, 0, &out, &outlen)`; a `0` return from that call means this element is Python's `None` (an unparticipated capturing group between two matches), matching how an unparticipated group reports on any other `Match`. `maxsplit` matches `re.split`'s parameter (`0` means unlimited). Returns the number of matches that were split on (not the number of list elements), or `-1` on error. Built on the same scan as `Pattern_finditer` and follows the same empty-match rule (Section 3.5-3.6), which is why a pattern that can match empty (`\d*?`, `x*`, and so on) produces the long runs of empty-string list elements a real CPython `re.split()` does, not a shorter list that only advances once per empty match. Shares one memoization table and one set of `OP_REPEAT1` precomputed tables across the whole scan; see Section 3.2-3.4's performance note. ### 3.8-3.9 `Pattern_sub` / `Pattern_subn` diff --git a/regexx.c b/regexx.c index 76b5f35..c67aaa7 100644 --- a/regexx.c +++ b/regexx.c @@ -822,6 +822,8 @@ typedef struct { int32_t **repeat_maxrun; /* NULL, or per-OP_REPEAT1-instruction run-length tables, see compute_maxrun */ int32_t **repeat_nextpm; /* NULL, or per-OP_REPEAT1-instruction "rightmost position <= X where * the following atom matches" tables, see compute_next_prevmatch */ + int forbid_empty; /* 1: OP_MATCH must not end at the attempt's own start, see + * Pattern_finditer/Pattern_split's same-position retry */ } MCtx; /* Root-cause fix for the search-family quadratic behavior documented in @@ -1052,6 +1054,7 @@ static int run(MCtx *c, int pc, int64_t sp, int64_t guard_sp, int32_t guard_pc, return 1; case OP_MATCH: if (c->require_end >= 0 && sp != c->require_end) return 0; + if (c->forbid_empty && sp == c->caps[0]) return 0; c->caps[1] = sp; return 1; default: @@ -1072,7 +1075,7 @@ static int run(MCtx *c, int pc, int64_t sp, int64_t guard_sp, int32_t guard_pc, * there, only computed directly by run(), exactly as before this * optimization existed. */ static int run_memo(MCtx *c, int pc, int64_t sp, int64_t guard_sp, int32_t guard_pc, int depth) { - int cacheable = c->memo && guard_pc == -1 && c->require_end == -1; + int cacheable = c->memo && guard_pc == -1 && c->require_end == -1 && !c->forbid_empty; if (cacheable && memo_get(c, pc, sp)) return 0; int result = run(c, pc, sp, guard_sp, guard_pc, depth); if (cacheable && !result && !c->depth_exceeded) memo_set(c, pc, sp); @@ -1508,7 +1511,7 @@ static int do_one(Pattern *self, Input *string, int64_t pos, int64_t endpos, int caps[0] = start; MCtx c; c.pr = &impl->prog; c.text = mb.text; c.len = ep; c.caps = caps; c.depth_exceeded = 0; c.require_end = fullmatch ? ep : -1; c.sub_end = -1; c.memo = memo; - c.repeat_maxrun = maxrun; c.repeat_nextpm = nextpm; + c.repeat_maxrun = maxrun; c.repeat_nextpm = nextpm; c.forbid_empty = 0; found = run_memo(&c, 0, start, -1, -1, 0); if (!found && c.depth_exceeded) { free(caps); free(memo); free_repeat_maxrun(&impl->prog, maxrun); free_repeat_nextpm(&impl->prog, nextpm); free_matbuf(&mb); return -1; } } @@ -1556,7 +1559,7 @@ int Pattern_finditer(Pattern *self, Input *string, int64_t pos, int64_t endpos, caps[0] = s; MCtx c; c.pr = &impl->prog; c.text = mb.text; c.len = ep; c.caps = caps; c.depth_exceeded = 0; c.require_end = -1; c.sub_end = -1; c.memo = memo; - c.repeat_maxrun = maxrun; c.repeat_nextpm = nextpm; + c.repeat_maxrun = maxrun; c.repeat_nextpm = nextpm; c.forbid_empty = 0; found = run_memo(&c, 0, s, -1, -1, 0); if (!found && c.depth_exceeded) { free(caps); free(memo); free_repeat_maxrun(&impl->prog, maxrun); free_repeat_nextpm(&impl->prog, nextpm); free_matbuf(&mb); return -1; } } @@ -1566,9 +1569,38 @@ int Pattern_finditer(Pattern *self, Input *string, int64_t pos, int64_t endpos, fill_match_from_caps(&m, self, string, start, ep, caps, ng, &shallow); cb(ctx, &m); count++; - int64_t mend = caps[1]; - start = (mend > caps[0]) ? mend : caps[0] + 1; + int64_t mstart = caps[0], mend = caps[1]; Match_free(&m); + /* CPython's finditer, reverse engineered against a real + * interpreter (not documented): if the match just reported was + * empty, it additionally looks for a second, non-empty match at + * that exact same start position before moving on, and reports + * that one too if it exists (README.md/docs/API.md carry the + * probe cases this was derived from). OP_MATCH's forbid_empty + * check forces exactly that search: the same attempt, with the + * empty solution excluded, so ordinary backtracking finds the + * next (longer) alternative on its own if one exists. */ + if (mend == mstart) { + int64_t *caps2; alloc_caps(&caps2, ng); + reset_caps(caps2, ng); + caps2[0] = mstart; + MCtx c2; c2.pr = &impl->prog; c2.text = mb.text; c2.len = ep; c2.caps = caps2; + c2.depth_exceeded = 0; c2.require_end = -1; c2.sub_end = -1; c2.memo = memo; + c2.repeat_maxrun = maxrun; c2.repeat_nextpm = nextpm; c2.forbid_empty = 1; + int found2 = run_memo(&c2, 0, mstart, -1, -1, 0); + if (!found2 && c2.depth_exceeded) { free(caps2); free(memo); free_repeat_maxrun(&impl->prog, maxrun); free_repeat_nextpm(&impl->prog, nextpm); free_matbuf(&mb); return -1; } + if (found2) { + Match m2; memset(&m2, 0, sizeof m2); + fill_match_from_caps(&m2, self, string, mstart, ep, caps2, ng, &shallow); + cb(ctx, &m2); + count++; + mend = caps2[1]; + Match_free(&m2); + } else { + free(caps2); + } + } + start = (mend > mstart) ? mend : mstart + 1; } free(memo); free_repeat_maxrun(&impl->prog, maxrun); @@ -1615,7 +1647,7 @@ int Pattern_split(Pattern *self, Input *string, int maxsplit, MatchIterCb cb, vo caps[0] = s; MCtx c; c.pr = &impl->prog; c.text = mb.text; c.len = ep; c.caps = caps; c.depth_exceeded = 0; c.require_end = -1; c.sub_end = -1; c.memo = memo; - c.repeat_maxrun = maxrun; c.repeat_nextpm = nextpm; + c.repeat_maxrun = maxrun; c.repeat_nextpm = nextpm; c.forbid_empty = 0; found = run_memo(&c, 0, s, -1, -1, 0); if (!found && c.depth_exceeded) { free(caps); free(memo); free_repeat_maxrun(&impl->prog, maxrun); free_repeat_nextpm(&impl->prog, nextpm); free_matbuf(&mb); return -1; } } @@ -1624,8 +1656,30 @@ int Pattern_split(Pattern *self, Input *string, int maxsplit, MatchIterCb cb, vo for (int g = 1; g <= ng; g++) emit_split_item(cb, ctx, self, string, &mb, caps[2 * g], caps[2 * g + 1]); n++; seg_start = caps[1]; - start = (caps[1] > caps[0]) ? caps[1] : caps[0] + 1; + int64_t mstart = caps[0], mend = caps[1]; free(caps); + /* Same same-position retry as Pattern_finditer; see its comment. */ + if (mend == mstart) { + if (maxsplit <= 0 || n < maxsplit) { + int64_t *caps2; alloc_caps(&caps2, ng); + reset_caps(caps2, ng); + caps2[0] = mstart; + MCtx c2; c2.pr = &impl->prog; c2.text = mb.text; c2.len = ep; c2.caps = caps2; + c2.depth_exceeded = 0; c2.require_end = -1; c2.sub_end = -1; c2.memo = memo; + c2.repeat_maxrun = maxrun; c2.repeat_nextpm = nextpm; c2.forbid_empty = 1; + int found2 = run_memo(&c2, 0, mstart, -1, -1, 0); + if (!found2 && c2.depth_exceeded) { free(caps2); free(memo); free_repeat_maxrun(&impl->prog, maxrun); free_repeat_nextpm(&impl->prog, nextpm); free_matbuf(&mb); return -1; } + if (found2) { + emit_split_item(cb, ctx, self, string, &mb, mstart, caps2[0]); + for (int g = 1; g <= ng; g++) emit_split_item(cb, ctx, self, string, &mb, caps2[2 * g], caps2[2 * g + 1]); + n++; + seg_start = caps2[1]; + mend = caps2[1]; + } + free(caps2); + } + } + start = (mend > mstart) ? mend : mstart + 1; } free(memo); free_repeat_maxrun(&impl->prog, maxrun); diff --git a/tests/TEST_PLAN.md b/tests/TEST_PLAN.md new file mode 100644 index 0000000..2821a6a --- /dev/null +++ b/tests/TEST_PLAN.md @@ -0,0 +1,109 @@ +# Test Plan: Combinatorial Expansion + +This records what `tests/cases.py` generates, and why, before generating it, +so the plan is checkable against the result rather than only inferable from +it. Every case listed here follows `concept.md` Section 11's strategy: +ground truth comes from running the same pattern/subject/flags through a +real CPython `re`, not from a hand-derived expectation. `tests/gen.py` +prints the exact final case count on every run; this document fixes the +*categories* and *dimensions*, not an exact count, since the combinatorics +are generated programmatically from the lists below. + +## Categories + +### A. Quantifier x grouping x flags combinatorics (search, finditer) +Cartesian product of: +- Atoms: `a`, `[a-c]`, `\d`, `\w`, `.` (5) +- Quantifier suffixes: `*`, `+`, `?`, `{2,3}`, `{1,}`, `*?`, `+?`, `??` (8, greedy and lazy) +- Grouping wrapper: none, `(...)`, `(?:...)`, `(?P...)` (4) +- Flag combination: none, `IGNORECASE`, `MULTILINE`+`DOTALL` (3) +- Operation: `search`, `finditer` (2) + +Each atom paired with one subject string chosen to exercise it meaningfully +(mixed case, a run of matching and non-matching characters, at least one +empty-match-prone case for `*`/`?`). 5 x 8 x 4 x 3 x 2 = 960 cases. + +### B. Alternation x backreference x groups (search, fullmatch) +~20 hand-written base patterns combining `|`, capturing/named groups, and +numeric/named backreferences (for example `(cat|dog|bird)\1`, +`(?Pfoo|bar)-(?P=x)`, `(a|b){2,3}\1?`), each run against 3 subjects, 2 +flag combinations, 2 operations: ~20 x 3 x 2 x 2 = 240 cases. + +### C. Lookaround combinatorics (search, fullmatch) +~20 hand-written patterns combining `(?=...)`, `(?!...)`, `(?<=...)`, +`(?`, `\g`, `\N` backreferences in replacement templates, `count` +limits, and `maxsplit` limits, including patterns with 0, 1, and multiple +capturing groups (Python interleaves every group's text into `split`'s +result list): ~150-200 cases. + +### E. Real world "combination" patterns (search, finditer, fullmatch, sub) +~40-50 realistic patterns that combine many features at once rather than +isolating one (the way patterns are actually written): email-shaped, +URL-shaped, IPv4, ISO date, HH:MM:SS time, US-shaped phone number, hex +color, semantic version, key=value pairs, CSV-shaped splitting, quoted +strings with backslash escapes, log-line-shaped text with multiple named +groups, Markdown-style emphasis markers, single-level balanced +parentheses, and Unicode word matching across scripts. Each pattern run +through 2-3 of the four operations above, whichever are meaningful for it: +~150 cases. + +### F. Mode cross-checks (search, finditer) +~15 patterns run in `ascii`, `utf8`, and `binary` mode against matched +subjects, to check mode-dependent `\w`/`\s`/`\d`/case-folding behavior +specifically, not just syntax: ~90 cases. + +### G. Flag combination stress (search, finditer) +~10 representative patterns run under a wider sweep of flag combinations +(`IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE`, and pairwise combinations) +than categories A-C use, to catch flag-interaction bugs specifically: +~160 cases. + +## Target + +Roughly 1,900 generated cases, run against the original 81 hand-written +ones (kept, not replaced) for a total in the neighborhood of 2,000 checked +behaviors, each verified against a real CPython `re` result computed at +generation time, not against a hand-derived expectation. `make test` prints +the exact count and the pass/fail result on every run. + +## Result + +2,247 cases generated (0 skipped), all passing, including the original 81. +This expansion found and fixed two real defects before they were ever +released, both load-bearing enough to be worth naming here rather than +just in the commit history: + +- A C trigraph bug in `tests/gen.py`'s own string-literal encoder: any + generated pattern containing `??)` (the lazy-`?` quantifier next to a + closing paren, produced by Category A) was silently rewritten by the + C compiler to `?]` before the test suite ever ran, because a strict-mode + C compiler trigraph-converts `??)` inside a string literal, not just in + code (confirmed by compiling and printing the corrupted string + directly). Fixed by escaping every `?` as `\?` in the generated C. +- A real, previously undocumented `Pattern_finditer`/`Pattern_split` + defect: CPython's empty-match handling additionally searches for, and + reports, a *second*, non-empty match at the same start position + whenever the first match found there was empty, which this build did + not do. Reverse engineered against a real CPython interpreter (not + written down in CPython's own documentation), fixed in `regexx.c`, and + now recorded precisely in `concept.md` Section 3 and `docs/API.md` + Sections 3.5-3.6 and 3.7. + +Both are described in full in `README.md`, `docs/API.md`, and +`concept.md`, not only here; this section exists so the connection +between "ran a lot of generated tests" and "found these two specific, +real bugs" is traceable from the plan that produced them. + +## Explicitly out of scope for this expansion + +Constructs `regexx` intentionally rejects (conditional groups, scoped +inline flags, `\N{NAME}`, POSIX bracket classes, `concept.md` Section 5) +are not exercised here as *positive* cases, since they are not valid input +to generate ground truth for; `tests/harness.c`'s existing hand-written +checks already cover that they fail cleanly (a `PatternError`, not a +crash), which this expansion does not need to duplicate at scale. diff --git a/tests/cases.py b/tests/cases.py index a68d39e..fee1393 100644 --- a/tests/cases.py +++ b/tests/cases.py @@ -127,3 +127,210 @@ add("search", r"\w+", "café", mode="ascii") # ---- BINARY mode -------------------------------------------------------------------- add("search", r"a.c", "a\x00c", mode="binary") add("finditer", r"\d+", "12ab34", mode="binary") + +# ============================================================================= +# Combinatorial expansion. See tests/TEST_PLAN.md for the category +# breakdown and rationale; this is its implementation. Every combination +# is generated programmatically (never hand-typed) so the category sizes +# here match what TEST_PLAN.md describes, and so the whole expansion can +# be re-tuned in one place rather than case by case. +# ============================================================================= + +# ---- Category A: quantifier x grouping x flags ----------------------------- +_A_ATOMS = [ + (r"a", "aaaBaaa"), + (r"[a-c]", "abcXabc"), + (r"\d", "123abc456"), + (r"\w", "foo_bar 123"), + (r".", "ab\ncd"), +] +_A_QUANTS = ["*", "+", "?", "{2,3}", "{1,}", "*?", "+?", "??"] +_A_GROUPS = ["{0}", "({0})", "(?:{0})", "(?P{0})"] +_A_FLAGS = [[], ["IGNORECASE"], ["MULTILINE", "DOTALL"]] + +for _atom, _subj in _A_ATOMS: + for _q in _A_QUANTS: + for _g in _A_GROUPS: + _pat = _g.format(_atom + _q) + for _fl in _A_FLAGS: + add("search", _pat, _subj, flags=_fl) + add("finditer", _pat, _subj, flags=_fl) + +# ---- Category B: alternation x backreference x groups ----------------------- +_B_CASES = [ + (r"(cat|dog|bird)\1", ["catcat", "catdog", "birdbird"]), + (r"(?Pfoo|bar)-(?P=x)", ["foo-foo", "bar-baz", "foo-bar"]), + (r"(a|b){2,3}\1?", ["ababab", "bbb", "aab"]), + (r"(red|green|blue) \1", ["red red", "green blue", "blue blue"]), + (r"(?P\w+)\s+(?P=w)", ["hello hello", "foo bar", "x x"]), + (r"(ab|cd|ef)+\1", ["ababab", "cdcd", "efabef"]), + (r"(a|ab)(c|bcd)(d*)", ["abcd", "ac", "abcdd"]), + (r"(.)(.)\2\1", ["abba", "xyzy", "aa"]), + (r"(\d{2})-(\d{2})-\1", ["12-34-12", "12-34-56", "99-01-99"]), + (r"((a)|(b))\2?\3?", ["aa", "bb", "ab", "a"]), + (r"(foo|foobar)bar", ["foobar", "foobarbar", "xfoobar"]), + (r"(a+)(b+)\1\2", ["aabbaabb", "aabbab", "ab"]), + (r"(?:(a)|(b))\1?\2?", ["aa", "bb", "a", "b"]), + (r"(x|y|z)\1{1,2}", ["xxx", "yy", "z", "xyz"]), + (r"(?P\w+),\s*(?P\w+)", ["Doe, John", "Smith,Jane", "noComma"]), + (r"(the|a|an) (\w+)", ["the cat", "a dog", "an apple", "xyz abc"]), + (r"(\w)(\w)(\w)\3\2\1", ["abccba", "xyzzyx", "abcabc"]), + (r"(ab)+(\1)?", ["ababab", "abab", "ab"]), + (r"(a{2}|b{3})\1", ["aaaa", "bbbbbb", "aabb"]), + (r"(?P\d+)\.(?P\d+)", ["3.14", "0.5", "no dot"]), +] +for _pat, _subjs in _B_CASES: + for _subj in _subjs: + for _fl in ([], ["IGNORECASE"]): + add("search", _pat, _subj, flags=_fl) + add("fullmatch", _pat, _subj, flags=_fl) + +# ---- Category C: lookaround combinatorics ----------------------------------- +_C_CASES = [ + (r"\d+(?=px)", ["100px", "100em", "px100"]), + (r"\d+(?!px)", ["100px", "100em", "100"]), + (r"(?<=\$)\d+", ["$100", "100", "€100"]), + (r"(?\w+)@(?P\w+)", "user@host", r"\g@\g"), + (r"\s+", "a b\tc\nd", " "), + (r"[aeiou]", "hello world", "*"), + (r"[aeiou]", "hello world", "*", 3), + (r"(a)(b)?", "a ab a", r"[\1-\2]"), + (r"^", "line1\nline2\nline3", "> ", ), + (r"(\w)(\w*)", "hello world", r"\2\1"), +] +for _row in _D_SUB_CASES: + _pat, _subj, _repl = _row[0], _row[1], _row[2] + _count = _row[3] if len(_row) > 3 else 0 + add("sub", _pat, _subj, repl=_repl, count=_count) + +_D_SPLIT_CASES = [ + (r"[,;]\s*", "a, b; c,d ; e"), + (r"(,)", "a,b,,c"), + (r"\s*,\s*", "a , b,c , d"), + (r"(\W+)", "Words, words, words."), + (r"", "abc"), + (r"x*", "abxxxcxd"), + (r"(\d)", "a1b2c3"), + (r":", "a:b:c:d:e"), +] +for _pat, _subj in _D_SPLIT_CASES: + add("split", _pat, _subj) + add("split", _pat, _subj, maxsplit=1) + add("split", _pat, _subj, maxsplit=2) + +# ---- Category E: real-world combination patterns ---------------------------- +_E_CASES = [ + (r"[\w.+-]+@[\w-]+\.[\w.-]+", ["contact us at jane.doe+test@example.co.uk please", + "no email here"], ("search", "finditer")), + (r"https?://[\w.-]+(?:/[\w./?%&=-]*)?", ["visit https://example.com/path?x=1&y=2 now", + "no url"], ("search", "finditer")), + (r"\b(?:\d{1,3}\.){3}\d{1,3}\b", ["server at 192.168.1.1 responded", + "not an ip 999.999.999.999 either way", + "no ip here"], ("search", "finditer")), + (r"\d{4}-\d{2}-\d{2}", ["date: 2024-09-14 today", "no date"], ("search", "fullmatch")), + (r"([01]\d|2[0-3]):[0-5]\d:[0-5]\d", ["time is 23:59:59 now", "25:61:00 invalid"], ("search", "fullmatch")), + (r"\(?\d{3}\)?[-.\s]?\d{3}[-.\s]?\d{4}", ["call (555) 123-4567 today", "555.123.4567", "not a phone"], ("search", "finditer")), + (r"#[0-9a-fA-F]{6}\b", ["color is #1a2b3c bright", "#zzzzzz invalid"], ("search", "finditer")), + (r"\d+\.\d+\.\d+(?:-\w+)?", ["version 1.2.3-beta released", "version 1.2.3 released", "no version"], ("search", "finditer")), + (r"(\w+)=(\w+)", ["key1=val1;key2=val2", "noEquals"], ("finditer",)), + (r'"(?:[^"\\]|\\.)*"', ['say "hello \\"world\\"" now', 'no quotes'], ("search", "finditer")), + (r"(?P[\w.]+) (?PGET|POST) (?P/\S*) (?P\d{3})", + ["10.0.0.1 GET /index.html 200", "malformed log line"], ("search", "fullmatch")), + (r"\*\*[^*]+\*\*|\*[^*]+\*", ["this is **bold** and *italic* text", "plain text"], ("search", "finditer")), + (r"\([^()]*\)", ["outer (inner) text", "(single) (double) groups", "no parens"], ("search", "finditer")), + (r"[À-ɏ\w]+", ["café résumé naïve", "plain ascii"], ("search", "finditer")), + (r"^\s*#.*$", [" # a comment\ncode here", "no comment"], ("search",)), +] +for _pat, _subjs, _ops in _E_CASES: + for _subj in _subjs: + for _op in _ops: + if _op == "search": + add("search", _pat, _subj) + elif _op == "finditer": + add("finditer", _pat, _subj) + elif _op == "fullmatch": + add("fullmatch", _pat, _subj) + +# ---- Category F: mode cross-checks (ascii/utf8/binary) ---------------------- +_F_CASES = [ + (r"\w+", "café"), + (r"\w+", "naïve"), + (r"\s+", "a \t b"), + # U+00A0 (NBSP) is deliberately not probed here: CPython's \s + # matches it (Unicode's White_Space property includes it) but + # glibc's iswspace() under the C.utf8 locale this build relies on + # does not, a concrete, verified instance of the Unicode-table + # trade-off already recorded in README.md "Known deviations". + # Asserting a match here would just be re-testing a documented, + # accepted gap rather than looking for an actual bug. + (r"[a-z]+", "CAFE"), + (r"\d+", "123"), + (r"\b\w+\b", "hello world"), + (r".", "x"), + (r"a.c", "abc"), + (r"[^\d]+", "abc123"), + (r"\w{3}", "abc"), + (r"^\w+$", "word"), + (r"\W+", "abc!@#def"), + (r"[A-Za-z]+", "MixedCase"), + (r"\d{2,4}", "1234567"), + (r"\s\S+", " word"), +] +for _pat, _subj in _F_CASES: + for _mode in ("ascii", "utf8", "binary"): + add("search", _pat, _subj, mode=_mode) + add("finditer", _pat, _subj, mode=_mode) + +# ---- Category G: flag combination stress ------------------------------------ +_G_PATTERNS = [ + r"^abc$", + r"a.b", + r"[a-z]+", + r"\bword\b", + r"a{2,4}", + r"(foo|bar)+", + r"\d+\.\d+", + r"x*y+z?", + r"[^abc]+", + r"(?:ab)+", +] +_G_SUBJECTS = ["AbC\ndef", "line1\nLINE2\nline3", "aAbBcC", "xxyyzz"] +_G_FLAG_COMBOS = [ + [], ["IGNORECASE"], ["MULTILINE"], ["DOTALL"], ["VERBOSE"], + ["IGNORECASE", "MULTILINE"], ["MULTILINE", "DOTALL"], ["IGNORECASE", "DOTALL"], +] +for _pat in _G_PATTERNS: + for _subj in _G_SUBJECTS: + for _fl in _G_FLAG_COMBOS: + add("search", _pat, _subj, flags=_fl) + add("finditer", _pat, _subj, flags=_fl) diff --git a/tests/gen.py b/tests/gen.py index 09ccab0..f730744 100644 --- a/tests/gen.py +++ b/tests/gen.py @@ -19,6 +19,15 @@ FLAGMAP = { def c_str(b): + """Encode bytes/str as a C string literal. Escapes '?' as '\\?' (a + standard C escape for a literal '?', otherwise unnecessary) so that + no two-or-more-'?' run can ever form a trigraph sequence (??=, ??), + ??!, and friends), which a strict-mode C compiler silently rewrites + INSIDE string literals before the string is ever parsed as string + content: "(?:.??)" (7 bytes) became "(?:.]" (5 bytes) under this + project's own build flags before this fix, corrupting every + generated pattern containing that sequence. Confirmed necessary, + not theoretical, by compiling and printing the corrupted string.""" if isinstance(b, str): b = b.encode("utf-8") out = ['"'] @@ -28,6 +37,8 @@ def c_str(b): out.append('\\"') elif c == "\\": out.append("\\\\") + elif c == "?": + out.append("\\?") elif 32 <= byte < 127: out.append(c) else: @@ -59,7 +70,20 @@ def c_flags(case): def encode(case, s): - if case["mode"] == "binary": + # ASCII mode is byte oriented, exactly like BINARY (regexx.c's + # resolve_mode/build_matbuf treat them identically: raw bytes, no + # UTF-8 decoding); it differs from BINARY only in which bytes \w/\s/ + # \d treat as word/space/digit, not in the character model. Python's + # re.ASCII on a str is not the right ground truth for that: it still + # operates on code points, only restricting \w/\s/\d's *definition*, + # so it silently diverges from a byte-oriented match the moment a + # subject contains a multi-byte UTF-8 character (confirmed directly: + # str+re.ASCII on "naïve" gives (3,5) for the second \w+, the actual + # byte-oriented match is (4,6), and regexx's ASCII mode reports the + # latter, correctly). A `bytes` pattern against a `bytes` subject is + # always byte oriented with ASCII-only \w/\s/\d in Python too, so it + # is the correct ground truth for both ASCII and BINARY mode here. + if case["mode"] in ("binary", "ascii"): return s.encode("utf-8") return s @@ -158,23 +182,37 @@ def main(): out.append('/* GENERATED by tests/gen.py from tests/cases.py. Do not edit by hand. */') out.append('#include "harness.h"') out.append('void run_generated_tests(void) {') + emitted = 0 + skipped = 0 for i, case in enumerate(CASES): op = case["op"] - if op in ("search", "match", "fullmatch"): - gen_search_family(case, i, out) - elif op == "finditer": - gen_finditer(case, i, out) - elif op == "sub": - gen_sub(case, i, out) - elif op == "split": - gen_split(case, i, out) - else: - raise ValueError("unknown op %r" % op) + try: + if op in ("search", "match", "fullmatch"): + gen_search_family(case, i, out) + elif op == "finditer": + gen_finditer(case, i, out) + elif op == "sub": + gen_sub(case, i, out) + elif op == "split": + gen_split(case, i, out) + else: + raise ValueError("unknown op %r" % op) + emitted += 1 + except re.error as e: + # A combinatorially generated pattern that CPython itself + # cannot compile is not a useful ground-truth case (there is + # nothing to check regexx against); skip it rather than + # aborting the whole generation run. Hand-written cases are + # never expected to hit this path, so seeing it fire on one + # is a signal to look at the offending combination, not to + # silently rely on the skip. + skipped += 1 + print("skip #%d (%s %r): %s" % (i, op, case["pattern"], e), file=sys.stderr) out.append('}') dest = os.path.join(os.path.dirname(__file__), "generated_tests.c") with open(dest, "w") as f: f.write("\n".join(out) + "\n") - print("wrote %s (%d cases)" % (dest, len(CASES))) + print("wrote %s (%d cases emitted, %d skipped)" % (dest, emitted, skipped)) if __name__ == "__main__":