tests/TEST_PLAN.md records the plan before the result: seven categories
(quantifier x grouping x flags, alternation x backreference x groups,
lookaround combinatorics, sub/split specifics, real-world combination
patterns, mode cross-checks, flag combination stress), each generated
programmatically in tests/cases.py rather than hand-typed, still every
case checked against a real CPython re result computed by tests/gen.py,
per concept.md Section 11's existing strategy, just at 2,247 cases
instead of 81.
Running it immediately found two real, previously latent bugs, not
just confirmed correctness of what was already covered by hand:
- tests/gen.py's own C string-literal encoder had a trigraph bug: any
generated pattern containing `??)` (the lazy quantifier next to a
closing paren) was silently rewritten by the C compiler from 7 bytes
to 5 before the suite ever ran, confirmed directly by compiling and
printing the corrupted string. Fixed by escaping '?' as '\?', which
is always safe and makes trigraph formation impossible.
- Pattern_finditer/Pattern_split had a real, previously undocumented
correctness defect: CPython's empty-match handling additionally
searches for, and reports, a second, non-empty match at the same
start position whenever the natural match found there was empty
(reverse engineered against a real interpreter, since this is not
written down in CPython's own documentation; \d*? against
"123abc456" yields 16 matches, not 9). Fixed with a new MCtx
forbid_empty flag that forces exactly that second search by
rejecting the empty solution at OP_MATCH and letting ordinary
backtracking find the next alternative, wired into both functions
(they have independent scan loops). concept.md Section 3 and
docs/API.md now state the rule precisely instead of the previous,
incomplete description.
A third, more mundane finding: the combinatorial mode cross-check
category surfaced that ASCII-mode ground truth was being computed
wrong in tests/gen.py itself (Python str + re.ASCII, which stays in
code-point space, instead of a bytes pattern against a bytes subject,
which is what regexx's byte-oriented ASCII mode actually is), and
separately surfaced a genuine, now precisely documented Unicode-table
gap already anticipated in principle by README.md's "Known deviations"
(glibc's iswspace() under the C.utf8 locale does not classify U+00A0
NO-BREAK SPACE as whitespace; CPython's \s does).
All 2,247 cases pass, clean under AddressSanitizer/UndefinedBehavior-
Sanitizer; the 50MB/quadratic-time and ReDoS-scaling benchmarks from
the previous two commits are unaffected.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY