Three new categories added to tests/cases.py (TEST_PLAN.md records the detail): Category H exercises BINARY mode with genuinely arbitrary raw bytes (embedded NUL, high bytes, non-UTF-8 sequences) via real Python `bytes` subjects, not just UTF-8-encoded text; Category I exercises UTF8 mode edge cases (4-byte/astral code points, combining marks, Arabic, Hebrew, CJK, offset correctness across multi-byte characters); Category J is a fixed-seed (reproducible, not flaky) random generator combining the existing atom/quantifier/grouping vocabulary across all three modes. This found and fixed a real bug, not just a test-generation one: LOCALE, in BINARY/ASCII mode, treated bytes 0x80-0xFF as word characters, based on an unverified assumption about what the "C" locale does. Checked directly against both the C standard's own guarantee for isalnum() under "C" and a real CPython interpreter with re.LOCALE and the "C" locale explicitly set, neither treats anything above 0x7f as a word character. Fixed in regexx.c's cls_is_word; LOCALE is now documented as an accepted no-op in non-UTF8 mode, matching verified reality instead of a prior assumption (concept.md 13.4, docs/API.md, README.md "Known deviations"). Two more findings were test-generation bugs, not regexx bugs: gen.py's own ASCII-mode ground truth used Python str + re.ASCII (code-point space) instead of a bytes pattern against a bytes subject (what regexx's byte-oriented ASCII mode actually is), and LOCALE combined with the (now removed as redundant) auto-added re.ASCII flag raised ValueError in Python for being an incompatible combination. Both fixed in gen.py. Two further findings were concrete instances of an already-documented category (glibc's wctype.h Unicode tables not matching CPython's own exactly): U+00A0 and fullwidth digits U+FF10-FF19 are recognized by CPython's \s/\d but not by glibc's iswspace()/iswdigit() under C.utf8. Recorded in README.md, not patched, for the reason already given for the first such instance (NBSP) in the previous commit. After these fixes: all 3,252 committed cases pass, clean under AddressSanitizer/UndefinedBehaviorSanitizer. The Category J generator was additionally run against 5 more seeds at 3,000 iterations each (26,760 further checks) as exploratory validation, all passing; not committed, to keep the regular suite's size proportionate. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
9.0 KiB
Test Plan: Combinatorial Expansion
This records what tests/cases.py generates, and why, before generating it,
so the plan is checkable against the result rather than only inferable from
it. Every case listed here follows concept.md Section 11's strategy:
ground truth comes from running the same pattern/subject/flags through a
real CPython re, not from a hand-derived expectation. tests/gen.py
prints the exact final case count on every run; this document fixes the
categories and dimensions, not an exact count, since the combinatorics
are generated programmatically from the lists below.
Categories
A. Quantifier x grouping x flags combinatorics (search, finditer)
Cartesian product of:
- Atoms:
a,[a-c],\d,\w,.(5) - Quantifier suffixes:
*,+,?,{2,3},{1,},*?,+?,??(8, greedy and lazy) - Grouping wrapper: none,
(...),(?:...),(?P<g>...)(4) - Flag combination: none,
IGNORECASE,MULTILINE+DOTALL(3) - Operation:
search,finditer(2)
Each atom paired with one subject string chosen to exercise it meaningfully
(mixed case, a run of matching and non-matching characters, at least one
empty-match-prone case for */?). 5 x 8 x 4 x 3 x 2 = 960 cases.
B. Alternation x backreference x groups (search, fullmatch)
~20 hand-written base patterns combining |, capturing/named groups, and
numeric/named backreferences (for example (cat|dog|bird)\1,
(?P<x>foo|bar)-(?P=x), (a|b){2,3}\1?), each run against 3 subjects, 2
flag combinations, 2 operations: ~20 x 3 x 2 x 2 = 240 cases.
C. Lookaround combinatorics (search, fullmatch)
~20 hand-written patterns combining (?=...), (?!...), (?<=...),
(?<!...) with quantifiers, groups, and each other (nested lookarounds),
each run against 3 subjects, 2 operations: ~20 x 3 x 2 = 120 cases.
D. sub/split specific combinatorics
~30 pattern/replacement/count and pattern/maxsplit combinations exercising
\g<name>, \g<N>, \N backreferences in replacement templates, count
limits, and maxsplit limits, including patterns with 0, 1, and multiple
capturing groups (Python interleaves every group's text into split's
result list): ~150-200 cases.
E. Real world "combination" patterns (search, finditer, fullmatch, sub)
~40-50 realistic patterns that combine many features at once rather than isolating one (the way patterns are actually written): email-shaped, URL-shaped, IPv4, ISO date, HH:MM:SS time, US-shaped phone number, hex color, semantic version, key=value pairs, CSV-shaped splitting, quoted strings with backslash escapes, log-line-shaped text with multiple named groups, Markdown-style emphasis markers, single-level balanced parentheses, and Unicode word matching across scripts. Each pattern run through 2-3 of the four operations above, whichever are meaningful for it: ~150 cases.
F. Mode cross-checks (search, finditer)
~15 patterns run in ascii, utf8, and binary mode against matched
subjects, to check mode-dependent \w/\s/\d/case-folding behavior
specifically, not just syntax: ~90 cases.
G. Flag combination stress (search, finditer)
~10 representative patterns run under a wider sweep of flag combinations
(IGNORECASE, MULTILINE, DOTALL, VERBOSE, and pairwise combinations)
than categories A-C use, to catch flag-interaction bugs specifically:
~160 cases.
H. BINARY mode with genuinely arbitrary raw bytes
Unlike Category F (which exercises BINARY/ASCII/UTF8 mode with
ordinary UTF-8-encoded text), these cases use real Python bytes objects
containing embedded NUL, high bytes (0x80-0xFF), and byte sequences
that are not, and are not meant to be, valid UTF-8 at all: ~20 hand-written
cases across search, finditer, sub, and split.
I. UTF8 mode edge cases
Real multi-byte content beyond simple accented Latin: 4-byte (astral
plane) code points such as emoji, combining marks, right-to-left scripts
(Arabic, Hebrew), CJK, mixed-width strings, and offset correctness for
backreferences/groups/lookaround spanning multi-byte characters: ~20
hand-written cases across search, fullmatch, finditer, sub, and
split.
J. Seeded random combinatorics
A fixed-seed (0xC0FFEE) random generator building 900 pattern/subject
pairs from the same atom/quantifier/grouping vocabulary as Category A, but
combined randomly rather than exhaustively, across all three modes and a
sweep of flag combinations. The fixed seed makes this a deterministic,
reproducible regression test, not a flaky fuzzer: the same 900 cases
generate every run, so a failure here is exactly as reproducible and
reportable as a hand-written one, while still exploring combinations no
one sat down and thought to write by hand.
Target
Roughly 1,900 generated cases for Categories A-G, run against the original
81 hand-written ones (kept, not replaced); Categories H-J added later
brought the total to just over 3,200. make test prints the exact current
count and the pass/fail result on every run.
Result
3,252 cases generated (0 skipped), all passing, including the original 81. This expansion found and fixed four real defects before they were ever released, load-bearing enough to be worth naming here rather than only in the commit history:
- A C trigraph bug in
tests/gen.py's own string-literal encoder: any generated pattern containing??)(the lazy-?quantifier next to a closing paren, produced by Category A) was silently rewritten by the C compiler to?]before the test suite ever ran, because a strict-mode C compiler trigraph-converts??)inside a string literal, not just in code (confirmed by compiling and printing the corrupted string directly). Fixed by escaping every?as\?in the generated C. - A real, previously undocumented
Pattern_finditer/Pattern_splitdefect: CPython's empty-match handling additionally searches for, and reports, a second, non-empty match at the same start position whenever the first match found there was empty, which this build did not do. Reverse engineered against a real CPython interpreter (not written down in CPython's own documentation), fixed inregexx.c, and now recorded precisely inconcept.mdSection 3 anddocs/API.mdSections 3.5-3.6 and 3.7. tests/gen.py's ownASCII-mode ground truth was wrong (Pythonstr+re.ASCII, which stays in code-point space, instead of abytespattern against abytessubject, which is what regexx's byte-orientedASCIImode actually is), caught by Category F's mode cross-checks against a subject containing a multi-byte character.- A real, verified-wrong implementation of the
LOCALEflag: an earlier revision treated bytes0x80-0xFFas word characters underLOCALEinBINARY/ASCIImode, based on an unverified assumption about "C" locale behavior. Caught by Category H, checked directly against both the C standard's guarantee forisalnum()under the "C" locale and a real CPython interpreter withre.LOCALEand the "C" locale explicitly set, both of which classify only ASCII letters and digits.LOCALEis now documented as an accepted no-op inBINARY/ASCIImode, matching verified reality instead of a prior assumption (concept.md13.4,docs/API.mdSection 2,README.md"Known deviations").
Two further, smaller findings were concrete instances of an
already-documented category (glibc's wctype.h Unicode tables not
matching CPython's own bundled ones exactly, README.md "Known
deviations"), not new defects: U+00A0 (NO-BREAK SPACE, Category F) and
fullwidth digits U+FF10-U+FF19 (Category I) are both recognized by
CPython's \s/\d but not by glibc's iswspace()/iswdigit() under the
C.utf8 locale this build relies on. Recorded, not patched, for the same
reason already given for the first instance: hand-patching individual
code points would start down the path of maintaining an ad hoc table this
design deliberately avoided.
After these fixes, the random Category J generator was additionally run
against five more seeds (1, 42, 777, 99999, 123456) at 3,000
iterations each, 26,760 further checks beyond the 900 committed here, all
passing; only the committed seed's 900 cases are part of the regular test
suite, the rest was exploratory validation, not committed as a permanent
fixture, to avoid the build slowing down for marginal additional coverage
over what is already committed.
All of the above is described in full in README.md, docs/API.md, and
concept.md, not only here; this section exists so the connection between
"ran a lot of generated tests" and "found these specific, real bugs" is
traceable from the plan that produced them.
Explicitly out of scope for this expansion
Constructs regexx intentionally rejects (conditional groups, scoped
inline flags, \N{NAME}, POSIX bracket classes, concept.md Section 5)
are not exercised here as positive cases, since they are not valid input
to generate ground truth for; tests/harness.c's existing hand-written
checks already cover that they fail cleanly (a PatternError, not a
crash), which this expansion does not need to duplicate at scale.