Files
regexx/tests/TEST_PLAN.md
T
retoorandClaude Sonnet 5 b0991814cb Expand BINARY/UTF8/ASCII edge-case testing to 3,252 cases, fix a real LOCALE bug
Three new categories added to tests/cases.py (TEST_PLAN.md records the
detail): Category H exercises BINARY mode with genuinely arbitrary raw
bytes (embedded NUL, high bytes, non-UTF-8 sequences) via real Python
`bytes` subjects, not just UTF-8-encoded text; Category I exercises
UTF8 mode edge cases (4-byte/astral code points, combining marks,
Arabic, Hebrew, CJK, offset correctness across multi-byte characters);
Category J is a fixed-seed (reproducible, not flaky) random generator
combining the existing atom/quantifier/grouping vocabulary across all
three modes.

This found and fixed a real bug, not just a test-generation one:
LOCALE, in BINARY/ASCII mode, treated bytes 0x80-0xFF as word
characters, based on an unverified assumption about what the "C"
locale does. Checked directly against both the C standard's own
guarantee for isalnum() under "C" and a real CPython interpreter with
re.LOCALE and the "C" locale explicitly set, neither treats anything
above 0x7f as a word character. Fixed in regexx.c's cls_is_word;
LOCALE is now documented as an accepted no-op in non-UTF8 mode,
matching verified reality instead of a prior assumption (concept.md
13.4, docs/API.md, README.md "Known deviations").

Two more findings were test-generation bugs, not regexx bugs: gen.py's
own ASCII-mode ground truth used Python str + re.ASCII (code-point
space) instead of a bytes pattern against a bytes subject (what
regexx's byte-oriented ASCII mode actually is), and LOCALE combined
with the (now removed as redundant) auto-added re.ASCII flag raised
ValueError in Python for being an incompatible combination. Both
fixed in gen.py.

Two further findings were concrete instances of an already-documented
category (glibc's wctype.h Unicode tables not matching CPython's own
exactly): U+00A0 and fullwidth digits U+FF10-FF19 are recognized by
CPython's \s/\d but not by glibc's iswspace()/iswdigit() under C.utf8.
Recorded in README.md, not patched, for the reason already given for
the first such instance (NBSP) in the previous commit.

After these fixes: all 3,252 committed cases pass, clean under
AddressSanitizer/UndefinedBehaviorSanitizer. The Category J generator
was additionally run against 5 more seeds at 3,000 iterations each
(26,760 further checks) as exploratory validation, all passing; not
committed, to keep the regular suite's size proportionate.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
2026-09-14 09:25:50 +00:00

9.0 KiB

Test Plan: Combinatorial Expansion

This records what tests/cases.py generates, and why, before generating it, so the plan is checkable against the result rather than only inferable from it. Every case listed here follows concept.md Section 11's strategy: ground truth comes from running the same pattern/subject/flags through a real CPython re, not from a hand-derived expectation. tests/gen.py prints the exact final case count on every run; this document fixes the categories and dimensions, not an exact count, since the combinatorics are generated programmatically from the lists below.

Categories

A. Quantifier x grouping x flags combinatorics (search, finditer)

Cartesian product of:

  • Atoms: a, [a-c], \d, \w, . (5)
  • Quantifier suffixes: *, +, ?, {2,3}, {1,}, *?, +?, ?? (8, greedy and lazy)
  • Grouping wrapper: none, (...), (?:...), (?P<g>...) (4)
  • Flag combination: none, IGNORECASE, MULTILINE+DOTALL (3)
  • Operation: search, finditer (2)

Each atom paired with one subject string chosen to exercise it meaningfully (mixed case, a run of matching and non-matching characters, at least one empty-match-prone case for */?). 5 x 8 x 4 x 3 x 2 = 960 cases.

B. Alternation x backreference x groups (search, fullmatch)

~20 hand-written base patterns combining |, capturing/named groups, and numeric/named backreferences (for example (cat|dog|bird)\1, (?P<x>foo|bar)-(?P=x), (a|b){2,3}\1?), each run against 3 subjects, 2 flag combinations, 2 operations: ~20 x 3 x 2 x 2 = 240 cases.

C. Lookaround combinatorics (search, fullmatch)

~20 hand-written patterns combining (?=...), (?!...), (?<=...), (?<!...) with quantifiers, groups, and each other (nested lookarounds), each run against 3 subjects, 2 operations: ~20 x 3 x 2 = 120 cases.

D. sub/split specific combinatorics

~30 pattern/replacement/count and pattern/maxsplit combinations exercising \g<name>, \g<N>, \N backreferences in replacement templates, count limits, and maxsplit limits, including patterns with 0, 1, and multiple capturing groups (Python interleaves every group's text into split's result list): ~150-200 cases.

E. Real world "combination" patterns (search, finditer, fullmatch, sub)

~40-50 realistic patterns that combine many features at once rather than isolating one (the way patterns are actually written): email-shaped, URL-shaped, IPv4, ISO date, HH:MM:SS time, US-shaped phone number, hex color, semantic version, key=value pairs, CSV-shaped splitting, quoted strings with backslash escapes, log-line-shaped text with multiple named groups, Markdown-style emphasis markers, single-level balanced parentheses, and Unicode word matching across scripts. Each pattern run through 2-3 of the four operations above, whichever are meaningful for it: ~150 cases.

F. Mode cross-checks (search, finditer)

~15 patterns run in ascii, utf8, and binary mode against matched subjects, to check mode-dependent \w/\s/\d/case-folding behavior specifically, not just syntax: ~90 cases.

G. Flag combination stress (search, finditer)

~10 representative patterns run under a wider sweep of flag combinations (IGNORECASE, MULTILINE, DOTALL, VERBOSE, and pairwise combinations) than categories A-C use, to catch flag-interaction bugs specifically: ~160 cases.

H. BINARY mode with genuinely arbitrary raw bytes

Unlike Category F (which exercises BINARY/ASCII/UTF8 mode with ordinary UTF-8-encoded text), these cases use real Python bytes objects containing embedded NUL, high bytes (0x80-0xFF), and byte sequences that are not, and are not meant to be, valid UTF-8 at all: ~20 hand-written cases across search, finditer, sub, and split.

I. UTF8 mode edge cases

Real multi-byte content beyond simple accented Latin: 4-byte (astral plane) code points such as emoji, combining marks, right-to-left scripts (Arabic, Hebrew), CJK, mixed-width strings, and offset correctness for backreferences/groups/lookaround spanning multi-byte characters: ~20 hand-written cases across search, fullmatch, finditer, sub, and split.

J. Seeded random combinatorics

A fixed-seed (0xC0FFEE) random generator building 900 pattern/subject pairs from the same atom/quantifier/grouping vocabulary as Category A, but combined randomly rather than exhaustively, across all three modes and a sweep of flag combinations. The fixed seed makes this a deterministic, reproducible regression test, not a flaky fuzzer: the same 900 cases generate every run, so a failure here is exactly as reproducible and reportable as a hand-written one, while still exploring combinations no one sat down and thought to write by hand.

Target

Roughly 1,900 generated cases for Categories A-G, run against the original 81 hand-written ones (kept, not replaced); Categories H-J added later brought the total to just over 3,200. make test prints the exact current count and the pass/fail result on every run.

Result

3,252 cases generated (0 skipped), all passing, including the original 81. This expansion found and fixed four real defects before they were ever released, load-bearing enough to be worth naming here rather than only in the commit history:

  • A C trigraph bug in tests/gen.py's own string-literal encoder: any generated pattern containing ??) (the lazy-? quantifier next to a closing paren, produced by Category A) was silently rewritten by the C compiler to ?] before the test suite ever ran, because a strict-mode C compiler trigraph-converts ??) inside a string literal, not just in code (confirmed by compiling and printing the corrupted string directly). Fixed by escaping every ? as \? in the generated C.
  • A real, previously undocumented Pattern_finditer/Pattern_split defect: CPython's empty-match handling additionally searches for, and reports, a second, non-empty match at the same start position whenever the first match found there was empty, which this build did not do. Reverse engineered against a real CPython interpreter (not written down in CPython's own documentation), fixed in regexx.c, and now recorded precisely in concept.md Section 3 and docs/API.md Sections 3.5-3.6 and 3.7.
  • tests/gen.py's own ASCII-mode ground truth was wrong (Python str + re.ASCII, which stays in code-point space, instead of a bytes pattern against a bytes subject, which is what regexx's byte-oriented ASCII mode actually is), caught by Category F's mode cross-checks against a subject containing a multi-byte character.
  • A real, verified-wrong implementation of the LOCALE flag: an earlier revision treated bytes 0x80-0xFF as word characters under LOCALE in BINARY/ASCII mode, based on an unverified assumption about "C" locale behavior. Caught by Category H, checked directly against both the C standard's guarantee for isalnum() under the "C" locale and a real CPython interpreter with re.LOCALE and the "C" locale explicitly set, both of which classify only ASCII letters and digits. LOCALE is now documented as an accepted no-op in BINARY/ASCII mode, matching verified reality instead of a prior assumption (concept.md 13.4, docs/API.md Section 2, README.md "Known deviations").

Two further, smaller findings were concrete instances of an already-documented category (glibc's wctype.h Unicode tables not matching CPython's own bundled ones exactly, README.md "Known deviations"), not new defects: U+00A0 (NO-BREAK SPACE, Category F) and fullwidth digits U+FF10-U+FF19 (Category I) are both recognized by CPython's \s/\d but not by glibc's iswspace()/iswdigit() under the C.utf8 locale this build relies on. Recorded, not patched, for the same reason already given for the first instance: hand-patching individual code points would start down the path of maintaining an ad hoc table this design deliberately avoided.

After these fixes, the random Category J generator was additionally run against five more seeds (1, 42, 777, 99999, 123456) at 3,000 iterations each, 26,760 further checks beyond the 900 committed here, all passing; only the committed seed's 900 cases are part of the regular test suite, the rest was exploratory validation, not committed as a permanent fixture, to avoid the build slowing down for marginal additional coverage over what is already committed.

All of the above is described in full in README.md, docs/API.md, and concept.md, not only here; this section exists so the connection between "ran a lot of generated tests" and "found these specific, real bugs" is traceable from the plan that produced them.

Explicitly out of scope for this expansion

Constructs regexx intentionally rejects (conditional groups, scoped inline flags, \N{NAME}, POSIX bracket classes, concept.md Section 5) are not exercised here as positive cases, since they are not valid input to generate ground truth for; tests/harness.c's existing hand-written checks already cover that they fail cleanly (a PatternError, not a crash), which this expansion does not need to duplicate at scale.