Files
regexx/tests/TEST_PLAN.md
T
retoorandClaude Sonnet 5 b0991814cb Expand BINARY/UTF8/ASCII edge-case testing to 3,252 cases, fix a real LOCALE bug
Three new categories added to tests/cases.py (TEST_PLAN.md records the
detail): Category H exercises BINARY mode with genuinely arbitrary raw
bytes (embedded NUL, high bytes, non-UTF-8 sequences) via real Python
`bytes` subjects, not just UTF-8-encoded text; Category I exercises
UTF8 mode edge cases (4-byte/astral code points, combining marks,
Arabic, Hebrew, CJK, offset correctness across multi-byte characters);
Category J is a fixed-seed (reproducible, not flaky) random generator
combining the existing atom/quantifier/grouping vocabulary across all
three modes.

This found and fixed a real bug, not just a test-generation one:
LOCALE, in BINARY/ASCII mode, treated bytes 0x80-0xFF as word
characters, based on an unverified assumption about what the "C"
locale does. Checked directly against both the C standard's own
guarantee for isalnum() under "C" and a real CPython interpreter with
re.LOCALE and the "C" locale explicitly set, neither treats anything
above 0x7f as a word character. Fixed in regexx.c's cls_is_word;
LOCALE is now documented as an accepted no-op in non-UTF8 mode,
matching verified reality instead of a prior assumption (concept.md
13.4, docs/API.md, README.md "Known deviations").

Two more findings were test-generation bugs, not regexx bugs: gen.py's
own ASCII-mode ground truth used Python str + re.ASCII (code-point
space) instead of a bytes pattern against a bytes subject (what
regexx's byte-oriented ASCII mode actually is), and LOCALE combined
with the (now removed as redundant) auto-added re.ASCII flag raised
ValueError in Python for being an incompatible combination. Both
fixed in gen.py.

Two further findings were concrete instances of an already-documented
category (glibc's wctype.h Unicode tables not matching CPython's own
exactly): U+00A0 and fullwidth digits U+FF10-FF19 are recognized by
CPython's \s/\d but not by glibc's iswspace()/iswdigit() under C.utf8.
Recorded in README.md, not patched, for the reason already given for
the first such instance (NBSP) in the previous commit.

After these fixes: all 3,252 committed cases pass, clean under
AddressSanitizer/UndefinedBehaviorSanitizer. The Category J generator
was additionally run against 5 more seeds at 3,000 iterations each
(26,760 further checks) as exploratory validation, all passing; not
committed, to keep the regular suite's size proportionate.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
2026-09-14 09:25:50 +00:00

168 lines
9.0 KiB
Markdown

# Test Plan: Combinatorial Expansion
This records what `tests/cases.py` generates, and why, before generating it,
so the plan is checkable against the result rather than only inferable from
it. Every case listed here follows `concept.md` Section 11's strategy:
ground truth comes from running the same pattern/subject/flags through a
real CPython `re`, not from a hand-derived expectation. `tests/gen.py`
prints the exact final case count on every run; this document fixes the
*categories* and *dimensions*, not an exact count, since the combinatorics
are generated programmatically from the lists below.
## Categories
### A. Quantifier x grouping x flags combinatorics (search, finditer)
Cartesian product of:
- Atoms: `a`, `[a-c]`, `\d`, `\w`, `.` (5)
- Quantifier suffixes: `*`, `+`, `?`, `{2,3}`, `{1,}`, `*?`, `+?`, `??` (8, greedy and lazy)
- Grouping wrapper: none, `(...)`, `(?:...)`, `(?P<g>...)` (4)
- Flag combination: none, `IGNORECASE`, `MULTILINE`+`DOTALL` (3)
- Operation: `search`, `finditer` (2)
Each atom paired with one subject string chosen to exercise it meaningfully
(mixed case, a run of matching and non-matching characters, at least one
empty-match-prone case for `*`/`?`). 5 x 8 x 4 x 3 x 2 = 960 cases.
### B. Alternation x backreference x groups (search, fullmatch)
~20 hand-written base patterns combining `|`, capturing/named groups, and
numeric/named backreferences (for example `(cat|dog|bird)\1`,
`(?P<x>foo|bar)-(?P=x)`, `(a|b){2,3}\1?`), each run against 3 subjects, 2
flag combinations, 2 operations: ~20 x 3 x 2 x 2 = 240 cases.
### C. Lookaround combinatorics (search, fullmatch)
~20 hand-written patterns combining `(?=...)`, `(?!...)`, `(?<=...)`,
`(?<!...)` with quantifiers, groups, and each other (nested lookarounds),
each run against 3 subjects, 2 operations: ~20 x 3 x 2 = 120 cases.
### D. `sub`/`split` specific combinatorics
~30 pattern/replacement/count and pattern/maxsplit combinations exercising
`\g<name>`, `\g<N>`, `\N` backreferences in replacement templates, `count`
limits, and `maxsplit` limits, including patterns with 0, 1, and multiple
capturing groups (Python interleaves every group's text into `split`'s
result list): ~150-200 cases.
### E. Real world "combination" patterns (search, finditer, fullmatch, sub)
~40-50 realistic patterns that combine many features at once rather than
isolating one (the way patterns are actually written): email-shaped,
URL-shaped, IPv4, ISO date, HH:MM:SS time, US-shaped phone number, hex
color, semantic version, key=value pairs, CSV-shaped splitting, quoted
strings with backslash escapes, log-line-shaped text with multiple named
groups, Markdown-style emphasis markers, single-level balanced
parentheses, and Unicode word matching across scripts. Each pattern run
through 2-3 of the four operations above, whichever are meaningful for it:
~150 cases.
### F. Mode cross-checks (search, finditer)
~15 patterns run in `ascii`, `utf8`, and `binary` mode against matched
subjects, to check mode-dependent `\w`/`\s`/`\d`/case-folding behavior
specifically, not just syntax: ~90 cases.
### G. Flag combination stress (search, finditer)
~10 representative patterns run under a wider sweep of flag combinations
(`IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE`, and pairwise combinations)
than categories A-C use, to catch flag-interaction bugs specifically:
~160 cases.
### H. `BINARY` mode with genuinely arbitrary raw bytes
Unlike Category F (which exercises `BINARY`/`ASCII`/`UTF8` mode with
ordinary UTF-8-encoded text), these cases use real Python `bytes` objects
containing embedded `NUL`, high bytes (`0x80`-`0xFF`), and byte sequences
that are not, and are not meant to be, valid UTF-8 at all: ~20 hand-written
cases across `search`, `finditer`, `sub`, and `split`.
### I. `UTF8` mode edge cases
Real multi-byte content beyond simple accented Latin: 4-byte (astral
plane) code points such as emoji, combining marks, right-to-left scripts
(Arabic, Hebrew), CJK, mixed-width strings, and offset correctness for
backreferences/groups/lookaround spanning multi-byte characters: ~20
hand-written cases across `search`, `fullmatch`, `finditer`, `sub`, and
`split`.
### J. Seeded random combinatorics
A fixed-seed (`0xC0FFEE`) random generator building 900 pattern/subject
pairs from the same atom/quantifier/grouping vocabulary as Category A, but
combined randomly rather than exhaustively, across all three modes and a
sweep of flag combinations. The fixed seed makes this a deterministic,
reproducible regression test, not a flaky fuzzer: the same 900 cases
generate every run, so a failure here is exactly as reproducible and
reportable as a hand-written one, while still exploring combinations no
one sat down and thought to write by hand.
## Target
Roughly 1,900 generated cases for Categories A-G, run against the original
81 hand-written ones (kept, not replaced); Categories H-J added later
brought the total to just over 3,200. `make test` prints the exact current
count and the pass/fail result on every run.
## Result
3,252 cases generated (0 skipped), all passing, including the original 81.
This expansion found and fixed four real defects before they were ever
released, load-bearing enough to be worth naming here rather than only in
the commit history:
- A C trigraph bug in `tests/gen.py`'s own string-literal encoder: any
generated pattern containing `??)` (the lazy-`?` quantifier next to a
closing paren, produced by Category A) was silently rewritten by the
C compiler to `?]` before the test suite ever ran, because a strict-mode
C compiler trigraph-converts `??)` inside a string literal, not just in
code (confirmed by compiling and printing the corrupted string
directly). Fixed by escaping every `?` as `\?` in the generated C.
- A real, previously undocumented `Pattern_finditer`/`Pattern_split`
defect: CPython's empty-match handling additionally searches for, and
reports, a *second*, non-empty match at the same start position
whenever the first match found there was empty, which this build did
not do. Reverse engineered against a real CPython interpreter (not
written down in CPython's own documentation), fixed in `regexx.c`, and
now recorded precisely in `concept.md` Section 3 and `docs/API.md`
Sections 3.5-3.6 and 3.7.
- `tests/gen.py`'s own `ASCII`-mode ground truth was wrong (Python
`str` + `re.ASCII`, which stays in code-point space, instead of a
`bytes` pattern against a `bytes` subject, which is what regexx's
byte-oriented `ASCII` mode actually is), caught by Category F's mode
cross-checks against a subject containing a multi-byte character.
- A real, verified-wrong implementation of the `LOCALE` flag: an earlier
revision treated bytes `0x80`-`0xFF` as word characters under `LOCALE`
in `BINARY`/`ASCII` mode, based on an unverified assumption about "C"
locale behavior. Caught by Category H, checked directly against both
the C standard's guarantee for `isalnum()` under the "C" locale and a
real CPython interpreter with `re.LOCALE` and the "C" locale explicitly
set, both of which classify only ASCII letters and digits. `LOCALE` is
now documented as an accepted no-op in `BINARY`/`ASCII` mode, matching
verified reality instead of a prior assumption (`concept.md` 13.4,
`docs/API.md` Section 2, `README.md` "Known deviations").
Two further, smaller findings were concrete instances of an
already-documented category (glibc's `wctype.h` Unicode tables not
matching CPython's own bundled ones exactly, `README.md` "Known
deviations"), not new defects: U+00A0 (NO-BREAK SPACE, Category F) and
fullwidth digits `U+FF10`-`U+FF19` (Category I) are both recognized by
CPython's `\s`/`\d` but not by glibc's `iswspace()`/`iswdigit()` under the
`C.utf8` locale this build relies on. Recorded, not patched, for the same
reason already given for the first instance: hand-patching individual
code points would start down the path of maintaining an ad hoc table this
design deliberately avoided.
After these fixes, the random Category J generator was additionally run
against five more seeds (`1`, `42`, `777`, `99999`, `123456`) at 3,000
iterations each, 26,760 further checks beyond the 900 committed here, all
passing; only the committed seed's 900 cases are part of the regular test
suite, the rest was exploratory validation, not committed as a permanent
fixture, to avoid the build slowing down for marginal additional coverage
over what is already committed.
All of the above is described in full in `README.md`, `docs/API.md`, and
`concept.md`, not only here; this section exists so the connection between
"ran a lot of generated tests" and "found these specific, real bugs" is
traceable from the plan that produced them.
## Explicitly out of scope for this expansion
Constructs `regexx` intentionally rejects (conditional groups, scoped
inline flags, `\N{NAME}`, POSIX bracket classes, `concept.md` Section 5)
are not exercised here as *positive* cases, since they are not valid input
to generate ground truth for; `tests/harness.c`'s existing hand-written
checks already cover that they fail cleanly (a `PatternError`, not a
crash), which this expansion does not need to duplicate at scale.