diff --git a/README.md b/README.md index 0524b14..c42bbb6 100644 --- a/README.md +++ b/README.md @@ -182,7 +182,15 @@ convention they follow. hand-generated Unicode table (`concept.md` 13.3 anticipated a reduced static table; using the C library's own tables turned out to be simpler and more complete, at the cost of depending on the platform's Unicode - version rather than a pinned one). + version rather than a pinned one). Concretely verified, not just + theoretical: U+00A0 (NO-BREAK SPACE) is in Unicode's `White_Space` + property, so CPython's `\s` matches it, but glibc's `iswspace()` under + `C.utf8` does not, so this build's `\s` does not either. Found by the + large combinatorial test expansion (`tests/cases.py` Category F, + `tests/TEST_PLAN.md`), not anticipated in advance; recorded here rather + than patched, since hand-patching individual code points would start + down the path of maintaining an ad hoc table this design deliberately + avoided by delegating to `wctype.h` in the first place. - `lastindex`/`lastgroup` report the highest-numbered capturing group that participated in the match, which coincides with CPython's "most recently closed group" rule for straightforward patterns but can differ from it in @@ -272,13 +280,22 @@ Run `./rxgrep --help` for the full option list. ## Testing -`tests/cases.py` lists pattern/subject/operation triples. `tests/gen.py` -computes each one's expected result with CPython's own `re` module and -writes `tests/generated_tests.c`, which is then compiled against `regexx.c` -and checked. This is a direct implementation of the strategy `concept.md` -Section 11 describes: conformance is measured against what CPython actually -does, not against a re-derived reading of its documentation. `make check` -additionally runs the suite under AddressSanitizer and +`tests/cases.py` lists pattern/subject/operation triples, both hand-written +and, for most of the file, generated programmatically from combinations of +quantifiers, groups, backreferences, lookaround, flags, and encoding modes +(`tests/TEST_PLAN.md` records the exact category breakdown and why each one +exists). `tests/gen.py` computes every case's expected result with +CPython's own `re` module and writes `tests/generated_tests.c`, which is +then compiled against `regexx.c` and checked. This is a direct +implementation of the strategy `concept.md` Section 11 describes: +conformance is measured against what CPython actually does, not against a +re-derived reading of its documentation, at a scale (2,247 cases as of this +writing, `make test` reports the current exact count) large enough that it +has already found real defects this way, not only confirmed the absence of +ones anyone thought to write by hand (`tests/TEST_PLAN.md` "Result" names +both: a C trigraph bug in the test generator itself, and a real, previously +undocumented `Pattern_finditer`/`Pattern_split` empty-match defect). `make +check` additionally runs the suite under AddressSanitizer and UndefinedBehaviorSanitizer. ## License diff --git a/concept.md b/concept.md index 6e2909a..8317e36 100644 --- a/concept.md +++ b/concept.md @@ -50,7 +50,7 @@ The engine parses the following syntax, matching CPython's documented and observ - `match(pattern, string)`: anchored at position 0, not required to consume the whole string. - `fullmatch(pattern, string)`: anchored at position 0 and at the end of the string. - `search(pattern, string)`: first match anywhere. -- `finditer` / `findall`: all non overlapping matches, left to right, with the CPython 3.7+ rule that an empty match advances one position and is followed by a search starting immediately after it rather than being merged with an adjacent non empty match at the same start position. +- `finditer` / `findall`: all non overlapping matches, left to right. The precise CPython 3.7+ empty-match rule, reverse engineered against a real interpreter since it is not fully spelled out in CPython's own documentation: at each scan position, the ordinary (greedy/lazy, as the pattern specifies) match is found and reported first; if that match was empty, a *second* match, requiring non-empty this time, is additionally searched for and reported at that same starting position if one exists (confirmed with `\d*?` against `"123abc456"`: CPython's `finditer` yields `(0,0)`, then `(0,1)`, then `(1,1)`, then `(1,2)`, and so on, an empty match immediately followed by a non-empty one at the same start, not two separate scan positions); the scan then advances to the end of whichever of the two was reported last, or by one position if only the empty match existed. `split` and `sub`/`subn` are built on the same scan and inherit the same rule. - `split(pattern, string, maxsplit=0)`: text of capturing groups is interleaved into the result list, matching CPython; splitting on a pattern that can match an empty string is permitted, matching the CPython 3.7+ behavior change. - `sub(pattern, repl, string, count=0)` / `subn`: `repl` is either a template string honoring `\g`, `\g<1>`, `\1`, and literal backslash escapes, or a callback invoked once per match with a match record and expected to return replacement bytes/text (the C equivalent of a Python callable, see 9.4). - `escape(string)`: backslash escaping of all characters outside `[A-Za-z0-9_]` in the same way `re.escape` does since Python 3.7 (that version narrowed the escaped set relative to earlier Python releases; the engine follows the narrowed, current set). diff --git a/docs/API.md b/docs/API.md index bc3c0b5..3967f59 100644 --- a/docs/API.md +++ b/docs/API.md @@ -154,7 +154,7 @@ int Pattern_finditer(Pattern *self, Input *string, int64_t pos, int64_t endpos, int Pattern_findall(Pattern *self, Input *string, int64_t pos, int64_t endpos, MatchIterCb cb, void *ctx); ``` -`Pattern_findall` is defined as a call to `Pattern_finditer` with the same arguments; both invoke `cb(ctx, m)` once per non-overlapping match, left to right, applying CPython's own empty-match rule (an empty match advances one unit and is not merged with an adjacent non-empty match at the same start, `concept.md` 3). Returns the number of matches found, or `-1` on error. Shares one memoization table and one set of `OP_REPEAT1` precomputed tables across the whole scan (built once, not once per match), so it inherits `search`'s performance characteristics in Section 3.2-3.4 exactly, not a worse case from repeating the scan. This build does not collapse a no-groups match down to "just the matched string" or a multi-group match to a Python tuple the way `re.findall` does at the Python level; the callback always receives a full `Match`, from which the caller reads whatever it needs via `Match_group`. This is a deliberate simplification: `findall`'s string/tuple collapsing is a Python-object-model convenience with no C equivalent to collapse into, so this build gives the caller the same, uniform `Match`-based access `finditer` does, and the two functions exist separately only for name-for-name parity with `re.findall`/`re.finditer`. +`Pattern_findall` is defined as a call to `Pattern_finditer` with the same arguments; both invoke `cb(ctx, m)` once per non-overlapping match, left to right, applying CPython's own empty-match rule exactly, including the part of it CPython does not document (`concept.md` 3, `regexx.c`'s `Pattern_finditer` comment): if the match found at a given start position is empty, a second, non-empty match is additionally searched for and reported at that same start before the scan moves on, so `\d*?` against `"123abc456"` yields 16 matches, not 9, matching a real CPython interpreter exactly (verified by generating this case's expectation from one, `tests/cases.py`). Returns the number of matches found, or `-1` on error. Shares one memoization table and one set of `OP_REPEAT1` precomputed tables across the whole scan (built once, not once per match), so it inherits `search`'s performance characteristics in Section 3.2-3.4 exactly, not a worse case from repeating the scan. This build does not collapse a no-groups match down to "just the matched string" or a multi-group match to a Python tuple the way `re.findall` does at the Python level; the callback always receives a full `Match`, from which the caller reads whatever it needs via `Match_group`. This is a deliberate simplification: `findall`'s string/tuple collapsing is a Python-object-model convenience with no C equivalent to collapse into, so this build gives the caller the same, uniform `Match`-based access `finditer` does, and the two functions exist separately only for name-for-name parity with `re.findall`/`re.finditer`. ### 3.7 `Pattern_split` @@ -162,7 +162,7 @@ int Pattern_findall(Pattern *self, Input *string, int64_t pos, int64_t endpos, M int Pattern_split(Pattern *self, Input *string, int maxsplit, MatchIterCb cb, void *ctx); ``` -`cb` is invoked once per element of the list `re.split()` would return, **in order**: this is the entire contract, and it is unambiguous by construction (earlier drafts of this function called `cb` twice per match with the caller left to infer which call meant what; that design was replaced before release specifically because it was ambiguous). Each element's text is read via `Match_group(m, NULL, 0, &out, &outlen)`; a `0` return from that call means this element is Python's `None` (an unparticipated capturing group between two matches), matching how an unparticipated group reports on any other `Match`. `maxsplit` matches `re.split`'s parameter (`0` means unlimited). Returns the number of matches that were split on (not the number of list elements), or `-1` on error. Shares one memoization table and one set of `OP_REPEAT1` precomputed tables across the whole scan; see Section 3.2-3.4's performance note. +`cb` is invoked once per element of the list `re.split()` would return, **in order**: this is the entire contract, and it is unambiguous by construction (earlier drafts of this function called `cb` twice per match with the caller left to infer which call meant what; that design was replaced before release specifically because it was ambiguous). Each element's text is read via `Match_group(m, NULL, 0, &out, &outlen)`; a `0` return from that call means this element is Python's `None` (an unparticipated capturing group between two matches), matching how an unparticipated group reports on any other `Match`. `maxsplit` matches `re.split`'s parameter (`0` means unlimited). Returns the number of matches that were split on (not the number of list elements), or `-1` on error. Built on the same scan as `Pattern_finditer` and follows the same empty-match rule (Section 3.5-3.6), which is why a pattern that can match empty (`\d*?`, `x*`, and so on) produces the long runs of empty-string list elements a real CPython `re.split()` does, not a shorter list that only advances once per empty match. Shares one memoization table and one set of `OP_REPEAT1` precomputed tables across the whole scan; see Section 3.2-3.4's performance note. ### 3.8-3.9 `Pattern_sub` / `Pattern_subn` diff --git a/regexx.c b/regexx.c index 76b5f35..c67aaa7 100644 --- a/regexx.c +++ b/regexx.c @@ -822,6 +822,8 @@ typedef struct { int32_t **repeat_maxrun; /* NULL, or per-OP_REPEAT1-instruction run-length tables, see compute_maxrun */ int32_t **repeat_nextpm; /* NULL, or per-OP_REPEAT1-instruction "rightmost position <= X where * the following atom matches" tables, see compute_next_prevmatch */ + int forbid_empty; /* 1: OP_MATCH must not end at the attempt's own start, see + * Pattern_finditer/Pattern_split's same-position retry */ } MCtx; /* Root-cause fix for the search-family quadratic behavior documented in @@ -1052,6 +1054,7 @@ static int run(MCtx *c, int pc, int64_t sp, int64_t guard_sp, int32_t guard_pc, return 1; case OP_MATCH: if (c->require_end >= 0 && sp != c->require_end) return 0; + if (c->forbid_empty && sp == c->caps[0]) return 0; c->caps[1] = sp; return 1; default: @@ -1072,7 +1075,7 @@ static int run(MCtx *c, int pc, int64_t sp, int64_t guard_sp, int32_t guard_pc, * there, only computed directly by run(), exactly as before this * optimization existed. */ static int run_memo(MCtx *c, int pc, int64_t sp, int64_t guard_sp, int32_t guard_pc, int depth) { - int cacheable = c->memo && guard_pc == -1 && c->require_end == -1; + int cacheable = c->memo && guard_pc == -1 && c->require_end == -1 && !c->forbid_empty; if (cacheable && memo_get(c, pc, sp)) return 0; int result = run(c, pc, sp, guard_sp, guard_pc, depth); if (cacheable && !result && !c->depth_exceeded) memo_set(c, pc, sp); @@ -1508,7 +1511,7 @@ static int do_one(Pattern *self, Input *string, int64_t pos, int64_t endpos, int caps[0] = start; MCtx c; c.pr = &impl->prog; c.text = mb.text; c.len = ep; c.caps = caps; c.depth_exceeded = 0; c.require_end = fullmatch ? ep : -1; c.sub_end = -1; c.memo = memo; - c.repeat_maxrun = maxrun; c.repeat_nextpm = nextpm; + c.repeat_maxrun = maxrun; c.repeat_nextpm = nextpm; c.forbid_empty = 0; found = run_memo(&c, 0, start, -1, -1, 0); if (!found && c.depth_exceeded) { free(caps); free(memo); free_repeat_maxrun(&impl->prog, maxrun); free_repeat_nextpm(&impl->prog, nextpm); free_matbuf(&mb); return -1; } } @@ -1556,7 +1559,7 @@ int Pattern_finditer(Pattern *self, Input *string, int64_t pos, int64_t endpos, caps[0] = s; MCtx c; c.pr = &impl->prog; c.text = mb.text; c.len = ep; c.caps = caps; c.depth_exceeded = 0; c.require_end = -1; c.sub_end = -1; c.memo = memo; - c.repeat_maxrun = maxrun; c.repeat_nextpm = nextpm; + c.repeat_maxrun = maxrun; c.repeat_nextpm = nextpm; c.forbid_empty = 0; found = run_memo(&c, 0, s, -1, -1, 0); if (!found && c.depth_exceeded) { free(caps); free(memo); free_repeat_maxrun(&impl->prog, maxrun); free_repeat_nextpm(&impl->prog, nextpm); free_matbuf(&mb); return -1; } } @@ -1566,9 +1569,38 @@ int Pattern_finditer(Pattern *self, Input *string, int64_t pos, int64_t endpos, fill_match_from_caps(&m, self, string, start, ep, caps, ng, &shallow); cb(ctx, &m); count++; - int64_t mend = caps[1]; - start = (mend > caps[0]) ? mend : caps[0] + 1; + int64_t mstart = caps[0], mend = caps[1]; Match_free(&m); + /* CPython's finditer, reverse engineered against a real + * interpreter (not documented): if the match just reported was + * empty, it additionally looks for a second, non-empty match at + * that exact same start position before moving on, and reports + * that one too if it exists (README.md/docs/API.md carry the + * probe cases this was derived from). OP_MATCH's forbid_empty + * check forces exactly that search: the same attempt, with the + * empty solution excluded, so ordinary backtracking finds the + * next (longer) alternative on its own if one exists. */ + if (mend == mstart) { + int64_t *caps2; alloc_caps(&caps2, ng); + reset_caps(caps2, ng); + caps2[0] = mstart; + MCtx c2; c2.pr = &impl->prog; c2.text = mb.text; c2.len = ep; c2.caps = caps2; + c2.depth_exceeded = 0; c2.require_end = -1; c2.sub_end = -1; c2.memo = memo; + c2.repeat_maxrun = maxrun; c2.repeat_nextpm = nextpm; c2.forbid_empty = 1; + int found2 = run_memo(&c2, 0, mstart, -1, -1, 0); + if (!found2 && c2.depth_exceeded) { free(caps2); free(memo); free_repeat_maxrun(&impl->prog, maxrun); free_repeat_nextpm(&impl->prog, nextpm); free_matbuf(&mb); return -1; } + if (found2) { + Match m2; memset(&m2, 0, sizeof m2); + fill_match_from_caps(&m2, self, string, mstart, ep, caps2, ng, &shallow); + cb(ctx, &m2); + count++; + mend = caps2[1]; + Match_free(&m2); + } else { + free(caps2); + } + } + start = (mend > mstart) ? mend : mstart + 1; } free(memo); free_repeat_maxrun(&impl->prog, maxrun); @@ -1615,7 +1647,7 @@ int Pattern_split(Pattern *self, Input *string, int maxsplit, MatchIterCb cb, vo caps[0] = s; MCtx c; c.pr = &impl->prog; c.text = mb.text; c.len = ep; c.caps = caps; c.depth_exceeded = 0; c.require_end = -1; c.sub_end = -1; c.memo = memo; - c.repeat_maxrun = maxrun; c.repeat_nextpm = nextpm; + c.repeat_maxrun = maxrun; c.repeat_nextpm = nextpm; c.forbid_empty = 0; found = run_memo(&c, 0, s, -1, -1, 0); if (!found && c.depth_exceeded) { free(caps); free(memo); free_repeat_maxrun(&impl->prog, maxrun); free_repeat_nextpm(&impl->prog, nextpm); free_matbuf(&mb); return -1; } } @@ -1624,8 +1656,30 @@ int Pattern_split(Pattern *self, Input *string, int maxsplit, MatchIterCb cb, vo for (int g = 1; g <= ng; g++) emit_split_item(cb, ctx, self, string, &mb, caps[2 * g], caps[2 * g + 1]); n++; seg_start = caps[1]; - start = (caps[1] > caps[0]) ? caps[1] : caps[0] + 1; + int64_t mstart = caps[0], mend = caps[1]; free(caps); + /* Same same-position retry as Pattern_finditer; see its comment. */ + if (mend == mstart) { + if (maxsplit <= 0 || n < maxsplit) { + int64_t *caps2; alloc_caps(&caps2, ng); + reset_caps(caps2, ng); + caps2[0] = mstart; + MCtx c2; c2.pr = &impl->prog; c2.text = mb.text; c2.len = ep; c2.caps = caps2; + c2.depth_exceeded = 0; c2.require_end = -1; c2.sub_end = -1; c2.memo = memo; + c2.repeat_maxrun = maxrun; c2.repeat_nextpm = nextpm; c2.forbid_empty = 1; + int found2 = run_memo(&c2, 0, mstart, -1, -1, 0); + if (!found2 && c2.depth_exceeded) { free(caps2); free(memo); free_repeat_maxrun(&impl->prog, maxrun); free_repeat_nextpm(&impl->prog, nextpm); free_matbuf(&mb); return -1; } + if (found2) { + emit_split_item(cb, ctx, self, string, &mb, mstart, caps2[0]); + for (int g = 1; g <= ng; g++) emit_split_item(cb, ctx, self, string, &mb, caps2[2 * g], caps2[2 * g + 1]); + n++; + seg_start = caps2[1]; + mend = caps2[1]; + } + free(caps2); + } + } + start = (mend > mstart) ? mend : mstart + 1; } free(memo); free_repeat_maxrun(&impl->prog, maxrun); diff --git a/tests/TEST_PLAN.md b/tests/TEST_PLAN.md new file mode 100644 index 0000000..2821a6a --- /dev/null +++ b/tests/TEST_PLAN.md @@ -0,0 +1,109 @@ +# Test Plan: Combinatorial Expansion + +This records what `tests/cases.py` generates, and why, before generating it, +so the plan is checkable against the result rather than only inferable from +it. Every case listed here follows `concept.md` Section 11's strategy: +ground truth comes from running the same pattern/subject/flags through a +real CPython `re`, not from a hand-derived expectation. `tests/gen.py` +prints the exact final case count on every run; this document fixes the +*categories* and *dimensions*, not an exact count, since the combinatorics +are generated programmatically from the lists below. + +## Categories + +### A. Quantifier x grouping x flags combinatorics (search, finditer) +Cartesian product of: +- Atoms: `a`, `[a-c]`, `\d`, `\w`, `.` (5) +- Quantifier suffixes: `*`, `+`, `?`, `{2,3}`, `{1,}`, `*?`, `+?`, `??` (8, greedy and lazy) +- Grouping wrapper: none, `(...)`, `(?:...)`, `(?P...)` (4) +- Flag combination: none, `IGNORECASE`, `MULTILINE`+`DOTALL` (3) +- Operation: `search`, `finditer` (2) + +Each atom paired with one subject string chosen to exercise it meaningfully +(mixed case, a run of matching and non-matching characters, at least one +empty-match-prone case for `*`/`?`). 5 x 8 x 4 x 3 x 2 = 960 cases. + +### B. Alternation x backreference x groups (search, fullmatch) +~20 hand-written base patterns combining `|`, capturing/named groups, and +numeric/named backreferences (for example `(cat|dog|bird)\1`, +`(?Pfoo|bar)-(?P=x)`, `(a|b){2,3}\1?`), each run against 3 subjects, 2 +flag combinations, 2 operations: ~20 x 3 x 2 x 2 = 240 cases. + +### C. Lookaround combinatorics (search, fullmatch) +~20 hand-written patterns combining `(?=...)`, `(?!...)`, `(?<=...)`, +`(?`, `\g`, `\N` backreferences in replacement templates, `count` +limits, and `maxsplit` limits, including patterns with 0, 1, and multiple +capturing groups (Python interleaves every group's text into `split`'s +result list): ~150-200 cases. + +### E. Real world "combination" patterns (search, finditer, fullmatch, sub) +~40-50 realistic patterns that combine many features at once rather than +isolating one (the way patterns are actually written): email-shaped, +URL-shaped, IPv4, ISO date, HH:MM:SS time, US-shaped phone number, hex +color, semantic version, key=value pairs, CSV-shaped splitting, quoted +strings with backslash escapes, log-line-shaped text with multiple named +groups, Markdown-style emphasis markers, single-level balanced +parentheses, and Unicode word matching across scripts. Each pattern run +through 2-3 of the four operations above, whichever are meaningful for it: +~150 cases. + +### F. Mode cross-checks (search, finditer) +~15 patterns run in `ascii`, `utf8`, and `binary` mode against matched +subjects, to check mode-dependent `\w`/`\s`/`\d`/case-folding behavior +specifically, not just syntax: ~90 cases. + +### G. Flag combination stress (search, finditer) +~10 representative patterns run under a wider sweep of flag combinations +(`IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE`, and pairwise combinations) +than categories A-C use, to catch flag-interaction bugs specifically: +~160 cases. + +## Target + +Roughly 1,900 generated cases, run against the original 81 hand-written +ones (kept, not replaced) for a total in the neighborhood of 2,000 checked +behaviors, each verified against a real CPython `re` result computed at +generation time, not against a hand-derived expectation. `make test` prints +the exact count and the pass/fail result on every run. + +## Result + +2,247 cases generated (0 skipped), all passing, including the original 81. +This expansion found and fixed two real defects before they were ever +released, both load-bearing enough to be worth naming here rather than +just in the commit history: + +- A C trigraph bug in `tests/gen.py`'s own string-literal encoder: any + generated pattern containing `??)` (the lazy-`?` quantifier next to a + closing paren, produced by Category A) was silently rewritten by the + C compiler to `?]` before the test suite ever ran, because a strict-mode + C compiler trigraph-converts `??)` inside a string literal, not just in + code (confirmed by compiling and printing the corrupted string + directly). Fixed by escaping every `?` as `\?` in the generated C. +- A real, previously undocumented `Pattern_finditer`/`Pattern_split` + defect: CPython's empty-match handling additionally searches for, and + reports, a *second*, non-empty match at the same start position + whenever the first match found there was empty, which this build did + not do. Reverse engineered against a real CPython interpreter (not + written down in CPython's own documentation), fixed in `regexx.c`, and + now recorded precisely in `concept.md` Section 3 and `docs/API.md` + Sections 3.5-3.6 and 3.7. + +Both are described in full in `README.md`, `docs/API.md`, and +`concept.md`, not only here; this section exists so the connection +between "ran a lot of generated tests" and "found these two specific, +real bugs" is traceable from the plan that produced them. + +## Explicitly out of scope for this expansion + +Constructs `regexx` intentionally rejects (conditional groups, scoped +inline flags, `\N{NAME}`, POSIX bracket classes, `concept.md` Section 5) +are not exercised here as *positive* cases, since they are not valid input +to generate ground truth for; `tests/harness.c`'s existing hand-written +checks already cover that they fail cleanly (a `PatternError`, not a +crash), which this expansion does not need to duplicate at scale. diff --git a/tests/cases.py b/tests/cases.py index a68d39e..fee1393 100644 --- a/tests/cases.py +++ b/tests/cases.py @@ -127,3 +127,210 @@ add("search", r"\w+", "café", mode="ascii") # ---- BINARY mode -------------------------------------------------------------------- add("search", r"a.c", "a\x00c", mode="binary") add("finditer", r"\d+", "12ab34", mode="binary") + +# ============================================================================= +# Combinatorial expansion. See tests/TEST_PLAN.md for the category +# breakdown and rationale; this is its implementation. Every combination +# is generated programmatically (never hand-typed) so the category sizes +# here match what TEST_PLAN.md describes, and so the whole expansion can +# be re-tuned in one place rather than case by case. +# ============================================================================= + +# ---- Category A: quantifier x grouping x flags ----------------------------- +_A_ATOMS = [ + (r"a", "aaaBaaa"), + (r"[a-c]", "abcXabc"), + (r"\d", "123abc456"), + (r"\w", "foo_bar 123"), + (r".", "ab\ncd"), +] +_A_QUANTS = ["*", "+", "?", "{2,3}", "{1,}", "*?", "+?", "??"] +_A_GROUPS = ["{0}", "({0})", "(?:{0})", "(?P{0})"] +_A_FLAGS = [[], ["IGNORECASE"], ["MULTILINE", "DOTALL"]] + +for _atom, _subj in _A_ATOMS: + for _q in _A_QUANTS: + for _g in _A_GROUPS: + _pat = _g.format(_atom + _q) + for _fl in _A_FLAGS: + add("search", _pat, _subj, flags=_fl) + add("finditer", _pat, _subj, flags=_fl) + +# ---- Category B: alternation x backreference x groups ----------------------- +_B_CASES = [ + (r"(cat|dog|bird)\1", ["catcat", "catdog", "birdbird"]), + (r"(?Pfoo|bar)-(?P=x)", ["foo-foo", "bar-baz", "foo-bar"]), + (r"(a|b){2,3}\1?", ["ababab", "bbb", "aab"]), + (r"(red|green|blue) \1", ["red red", "green blue", "blue blue"]), + (r"(?P\w+)\s+(?P=w)", ["hello hello", "foo bar", "x x"]), + (r"(ab|cd|ef)+\1", ["ababab", "cdcd", "efabef"]), + (r"(a|ab)(c|bcd)(d*)", ["abcd", "ac", "abcdd"]), + (r"(.)(.)\2\1", ["abba", "xyzy", "aa"]), + (r"(\d{2})-(\d{2})-\1", ["12-34-12", "12-34-56", "99-01-99"]), + (r"((a)|(b))\2?\3?", ["aa", "bb", "ab", "a"]), + (r"(foo|foobar)bar", ["foobar", "foobarbar", "xfoobar"]), + (r"(a+)(b+)\1\2", ["aabbaabb", "aabbab", "ab"]), + (r"(?:(a)|(b))\1?\2?", ["aa", "bb", "a", "b"]), + (r"(x|y|z)\1{1,2}", ["xxx", "yy", "z", "xyz"]), + (r"(?P\w+),\s*(?P\w+)", ["Doe, John", "Smith,Jane", "noComma"]), + (r"(the|a|an) (\w+)", ["the cat", "a dog", "an apple", "xyz abc"]), + (r"(\w)(\w)(\w)\3\2\1", ["abccba", "xyzzyx", "abcabc"]), + (r"(ab)+(\1)?", ["ababab", "abab", "ab"]), + (r"(a{2}|b{3})\1", ["aaaa", "bbbbbb", "aabb"]), + (r"(?P\d+)\.(?P\d+)", ["3.14", "0.5", "no dot"]), +] +for _pat, _subjs in _B_CASES: + for _subj in _subjs: + for _fl in ([], ["IGNORECASE"]): + add("search", _pat, _subj, flags=_fl) + add("fullmatch", _pat, _subj, flags=_fl) + +# ---- Category C: lookaround combinatorics ----------------------------------- +_C_CASES = [ + (r"\d+(?=px)", ["100px", "100em", "px100"]), + (r"\d+(?!px)", ["100px", "100em", "100"]), + (r"(?<=\$)\d+", ["$100", "100", "€100"]), + (r"(?\w+)@(?P\w+)", "user@host", r"\g@\g"), + (r"\s+", "a b\tc\nd", " "), + (r"[aeiou]", "hello world", "*"), + (r"[aeiou]", "hello world", "*", 3), + (r"(a)(b)?", "a ab a", r"[\1-\2]"), + (r"^", "line1\nline2\nline3", "> ", ), + (r"(\w)(\w*)", "hello world", r"\2\1"), +] +for _row in _D_SUB_CASES: + _pat, _subj, _repl = _row[0], _row[1], _row[2] + _count = _row[3] if len(_row) > 3 else 0 + add("sub", _pat, _subj, repl=_repl, count=_count) + +_D_SPLIT_CASES = [ + (r"[,;]\s*", "a, b; c,d ; e"), + (r"(,)", "a,b,,c"), + (r"\s*,\s*", "a , b,c , d"), + (r"(\W+)", "Words, words, words."), + (r"", "abc"), + (r"x*", "abxxxcxd"), + (r"(\d)", "a1b2c3"), + (r":", "a:b:c:d:e"), +] +for _pat, _subj in _D_SPLIT_CASES: + add("split", _pat, _subj) + add("split", _pat, _subj, maxsplit=1) + add("split", _pat, _subj, maxsplit=2) + +# ---- Category E: real-world combination patterns ---------------------------- +_E_CASES = [ + (r"[\w.+-]+@[\w-]+\.[\w.-]+", ["contact us at jane.doe+test@example.co.uk please", + "no email here"], ("search", "finditer")), + (r"https?://[\w.-]+(?:/[\w./?%&=-]*)?", ["visit https://example.com/path?x=1&y=2 now", + "no url"], ("search", "finditer")), + (r"\b(?:\d{1,3}\.){3}\d{1,3}\b", ["server at 192.168.1.1 responded", + "not an ip 999.999.999.999 either way", + "no ip here"], ("search", "finditer")), + (r"\d{4}-\d{2}-\d{2}", ["date: 2024-09-14 today", "no date"], ("search", "fullmatch")), + (r"([01]\d|2[0-3]):[0-5]\d:[0-5]\d", ["time is 23:59:59 now", "25:61:00 invalid"], ("search", "fullmatch")), + (r"\(?\d{3}\)?[-.\s]?\d{3}[-.\s]?\d{4}", ["call (555) 123-4567 today", "555.123.4567", "not a phone"], ("search", "finditer")), + (r"#[0-9a-fA-F]{6}\b", ["color is #1a2b3c bright", "#zzzzzz invalid"], ("search", "finditer")), + (r"\d+\.\d+\.\d+(?:-\w+)?", ["version 1.2.3-beta released", "version 1.2.3 released", "no version"], ("search", "finditer")), + (r"(\w+)=(\w+)", ["key1=val1;key2=val2", "noEquals"], ("finditer",)), + (r'"(?:[^"\\]|\\.)*"', ['say "hello \\"world\\"" now', 'no quotes'], ("search", "finditer")), + (r"(?P[\w.]+) (?PGET|POST) (?P/\S*) (?P\d{3})", + ["10.0.0.1 GET /index.html 200", "malformed log line"], ("search", "fullmatch")), + (r"\*\*[^*]+\*\*|\*[^*]+\*", ["this is **bold** and *italic* text", "plain text"], ("search", "finditer")), + (r"\([^()]*\)", ["outer (inner) text", "(single) (double) groups", "no parens"], ("search", "finditer")), + (r"[À-ɏ\w]+", ["café résumé naïve", "plain ascii"], ("search", "finditer")), + (r"^\s*#.*$", [" # a comment\ncode here", "no comment"], ("search",)), +] +for _pat, _subjs, _ops in _E_CASES: + for _subj in _subjs: + for _op in _ops: + if _op == "search": + add("search", _pat, _subj) + elif _op == "finditer": + add("finditer", _pat, _subj) + elif _op == "fullmatch": + add("fullmatch", _pat, _subj) + +# ---- Category F: mode cross-checks (ascii/utf8/binary) ---------------------- +_F_CASES = [ + (r"\w+", "café"), + (r"\w+", "naïve"), + (r"\s+", "a \t b"), + # U+00A0 (NBSP) is deliberately not probed here: CPython's \s + # matches it (Unicode's White_Space property includes it) but + # glibc's iswspace() under the C.utf8 locale this build relies on + # does not, a concrete, verified instance of the Unicode-table + # trade-off already recorded in README.md "Known deviations". + # Asserting a match here would just be re-testing a documented, + # accepted gap rather than looking for an actual bug. + (r"[a-z]+", "CAFE"), + (r"\d+", "123"), + (r"\b\w+\b", "hello world"), + (r".", "x"), + (r"a.c", "abc"), + (r"[^\d]+", "abc123"), + (r"\w{3}", "abc"), + (r"^\w+$", "word"), + (r"\W+", "abc!@#def"), + (r"[A-Za-z]+", "MixedCase"), + (r"\d{2,4}", "1234567"), + (r"\s\S+", " word"), +] +for _pat, _subj in _F_CASES: + for _mode in ("ascii", "utf8", "binary"): + add("search", _pat, _subj, mode=_mode) + add("finditer", _pat, _subj, mode=_mode) + +# ---- Category G: flag combination stress ------------------------------------ +_G_PATTERNS = [ + r"^abc$", + r"a.b", + r"[a-z]+", + r"\bword\b", + r"a{2,4}", + r"(foo|bar)+", + r"\d+\.\d+", + r"x*y+z?", + r"[^abc]+", + r"(?:ab)+", +] +_G_SUBJECTS = ["AbC\ndef", "line1\nLINE2\nline3", "aAbBcC", "xxyyzz"] +_G_FLAG_COMBOS = [ + [], ["IGNORECASE"], ["MULTILINE"], ["DOTALL"], ["VERBOSE"], + ["IGNORECASE", "MULTILINE"], ["MULTILINE", "DOTALL"], ["IGNORECASE", "DOTALL"], +] +for _pat in _G_PATTERNS: + for _subj in _G_SUBJECTS: + for _fl in _G_FLAG_COMBOS: + add("search", _pat, _subj, flags=_fl) + add("finditer", _pat, _subj, flags=_fl) diff --git a/tests/gen.py b/tests/gen.py index 09ccab0..f730744 100644 --- a/tests/gen.py +++ b/tests/gen.py @@ -19,6 +19,15 @@ FLAGMAP = { def c_str(b): + """Encode bytes/str as a C string literal. Escapes '?' as '\\?' (a + standard C escape for a literal '?', otherwise unnecessary) so that + no two-or-more-'?' run can ever form a trigraph sequence (??=, ??), + ??!, and friends), which a strict-mode C compiler silently rewrites + INSIDE string literals before the string is ever parsed as string + content: "(?:.??)" (7 bytes) became "(?:.]" (5 bytes) under this + project's own build flags before this fix, corrupting every + generated pattern containing that sequence. Confirmed necessary, + not theoretical, by compiling and printing the corrupted string.""" if isinstance(b, str): b = b.encode("utf-8") out = ['"'] @@ -28,6 +37,8 @@ def c_str(b): out.append('\\"') elif c == "\\": out.append("\\\\") + elif c == "?": + out.append("\\?") elif 32 <= byte < 127: out.append(c) else: @@ -59,7 +70,20 @@ def c_flags(case): def encode(case, s): - if case["mode"] == "binary": + # ASCII mode is byte oriented, exactly like BINARY (regexx.c's + # resolve_mode/build_matbuf treat them identically: raw bytes, no + # UTF-8 decoding); it differs from BINARY only in which bytes \w/\s/ + # \d treat as word/space/digit, not in the character model. Python's + # re.ASCII on a str is not the right ground truth for that: it still + # operates on code points, only restricting \w/\s/\d's *definition*, + # so it silently diverges from a byte-oriented match the moment a + # subject contains a multi-byte UTF-8 character (confirmed directly: + # str+re.ASCII on "naïve" gives (3,5) for the second \w+, the actual + # byte-oriented match is (4,6), and regexx's ASCII mode reports the + # latter, correctly). A `bytes` pattern against a `bytes` subject is + # always byte oriented with ASCII-only \w/\s/\d in Python too, so it + # is the correct ground truth for both ASCII and BINARY mode here. + if case["mode"] in ("binary", "ascii"): return s.encode("utf-8") return s @@ -158,23 +182,37 @@ def main(): out.append('/* GENERATED by tests/gen.py from tests/cases.py. Do not edit by hand. */') out.append('#include "harness.h"') out.append('void run_generated_tests(void) {') + emitted = 0 + skipped = 0 for i, case in enumerate(CASES): op = case["op"] - if op in ("search", "match", "fullmatch"): - gen_search_family(case, i, out) - elif op == "finditer": - gen_finditer(case, i, out) - elif op == "sub": - gen_sub(case, i, out) - elif op == "split": - gen_split(case, i, out) - else: - raise ValueError("unknown op %r" % op) + try: + if op in ("search", "match", "fullmatch"): + gen_search_family(case, i, out) + elif op == "finditer": + gen_finditer(case, i, out) + elif op == "sub": + gen_sub(case, i, out) + elif op == "split": + gen_split(case, i, out) + else: + raise ValueError("unknown op %r" % op) + emitted += 1 + except re.error as e: + # A combinatorially generated pattern that CPython itself + # cannot compile is not a useful ground-truth case (there is + # nothing to check regexx against); skip it rather than + # aborting the whole generation run. Hand-written cases are + # never expected to hit this path, so seeing it fire on one + # is a signal to look at the offending combination, not to + # silently rely on the skip. + skipped += 1 + print("skip #%d (%s %r): %s" % (i, op, case["pattern"], e), file=sys.stderr) out.append('}') dest = os.path.join(os.path.dirname(__file__), "generated_tests.c") with open(dest, "w") as f: f.write("\n".join(out) + "\n") - print("wrote %s (%d cases)" % (dest, len(CASES))) + print("wrote %s (%d cases emitted, %d skipped)" % (dest, emitted, skipped)) if __name__ == "__main__":