Expand BINARY/UTF8/ASCII edge-case testing to 3,252 cases, fix a real LOCALE bug

Three new categories added to tests/cases.py (TEST_PLAN.md records the
detail): Category H exercises BINARY mode with genuinely arbitrary raw
bytes (embedded NUL, high bytes, non-UTF-8 sequences) via real Python
`bytes` subjects, not just UTF-8-encoded text; Category I exercises
UTF8 mode edge cases (4-byte/astral code points, combining marks,
Arabic, Hebrew, CJK, offset correctness across multi-byte characters);
Category J is a fixed-seed (reproducible, not flaky) random generator
combining the existing atom/quantifier/grouping vocabulary across all
three modes.

This found and fixed a real bug, not just a test-generation one:
LOCALE, in BINARY/ASCII mode, treated bytes 0x80-0xFF as word
characters, based on an unverified assumption about what the "C"
locale does. Checked directly against both the C standard's own
guarantee for isalnum() under "C" and a real CPython interpreter with
re.LOCALE and the "C" locale explicitly set, neither treats anything
above 0x7f as a word character. Fixed in regexx.c's cls_is_word;
LOCALE is now documented as an accepted no-op in non-UTF8 mode,
matching verified reality instead of a prior assumption (concept.md
13.4, docs/API.md, README.md "Known deviations").

Two more findings were test-generation bugs, not regexx bugs: gen.py's
own ASCII-mode ground truth used Python str + re.ASCII (code-point
space) instead of a bytes pattern against a bytes subject (what
regexx's byte-oriented ASCII mode actually is), and LOCALE combined
with the (now removed as redundant) auto-added re.ASCII flag raised
ValueError in Python for being an incompatible combination. Both
fixed in gen.py.

Two further findings were concrete instances of an already-documented
category (glibc's wctype.h Unicode tables not matching CPython's own
exactly): U+00A0 and fullwidth digits U+FF10-FF19 are recognized by
CPython's \s/\d but not by glibc's iswspace()/iswdigit() under C.utf8.
Recorded in README.md, not patched, for the reason already given for
the first such instance (NBSP) in the previous commit.

After these fixes: all 3,252 committed cases pass, clean under
AddressSanitizer/UndefinedBehaviorSanitizer. The Category J generator
was additionally run against 5 more seeds at 3,000 iterations each
(26,760 further checks) as exploratory validation, all passing; not
committed, to keep the regular suite's size proportionate.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
This commit is contained in:
2026-09-14 09:25:50 +00:00
co-authored by Claude Sonnet 5
parent 51f8ed19e9
commit b0991814cb
7 changed files with 300 additions and 33 deletions
+71 -13
View File
@@ -63,20 +63,44 @@ specifically, not just syntax: ~90 cases.
than categories A-C use, to catch flag-interaction bugs specifically:
~160 cases.
### H. `BINARY` mode with genuinely arbitrary raw bytes
Unlike Category F (which exercises `BINARY`/`ASCII`/`UTF8` mode with
ordinary UTF-8-encoded text), these cases use real Python `bytes` objects
containing embedded `NUL`, high bytes (`0x80`-`0xFF`), and byte sequences
that are not, and are not meant to be, valid UTF-8 at all: ~20 hand-written
cases across `search`, `finditer`, `sub`, and `split`.
### I. `UTF8` mode edge cases
Real multi-byte content beyond simple accented Latin: 4-byte (astral
plane) code points such as emoji, combining marks, right-to-left scripts
(Arabic, Hebrew), CJK, mixed-width strings, and offset correctness for
backreferences/groups/lookaround spanning multi-byte characters: ~20
hand-written cases across `search`, `fullmatch`, `finditer`, `sub`, and
`split`.
### J. Seeded random combinatorics
A fixed-seed (`0xC0FFEE`) random generator building 900 pattern/subject
pairs from the same atom/quantifier/grouping vocabulary as Category A, but
combined randomly rather than exhaustively, across all three modes and a
sweep of flag combinations. The fixed seed makes this a deterministic,
reproducible regression test, not a flaky fuzzer: the same 900 cases
generate every run, so a failure here is exactly as reproducible and
reportable as a hand-written one, while still exploring combinations no
one sat down and thought to write by hand.
## Target
Roughly 1,900 generated cases, run against the original 81 hand-written
ones (kept, not replaced) for a total in the neighborhood of 2,000 checked
behaviors, each verified against a real CPython `re` result computed at
generation time, not against a hand-derived expectation. `make test` prints
the exact count and the pass/fail result on every run.
Roughly 1,900 generated cases for Categories A-G, run against the original
81 hand-written ones (kept, not replaced); Categories H-J added later
brought the total to just over 3,200. `make test` prints the exact current
count and the pass/fail result on every run.
## Result
2,247 cases generated (0 skipped), all passing, including the original 81.
This expansion found and fixed two real defects before they were ever
released, both load-bearing enough to be worth naming here rather than
just in the commit history:
3,252 cases generated (0 skipped), all passing, including the original 81.
This expansion found and fixed four real defects before they were ever
released, load-bearing enough to be worth naming here rather than only in
the commit history:
- A C trigraph bug in `tests/gen.py`'s own string-literal encoder: any
generated pattern containing `??)` (the lazy-`?` quantifier next to a
@@ -93,11 +117,45 @@ just in the commit history:
written down in CPython's own documentation), fixed in `regexx.c`, and
now recorded precisely in `concept.md` Section 3 and `docs/API.md`
Sections 3.5-3.6 and 3.7.
- `tests/gen.py`'s own `ASCII`-mode ground truth was wrong (Python
`str` + `re.ASCII`, which stays in code-point space, instead of a
`bytes` pattern against a `bytes` subject, which is what regexx's
byte-oriented `ASCII` mode actually is), caught by Category F's mode
cross-checks against a subject containing a multi-byte character.
- A real, verified-wrong implementation of the `LOCALE` flag: an earlier
revision treated bytes `0x80`-`0xFF` as word characters under `LOCALE`
in `BINARY`/`ASCII` mode, based on an unverified assumption about "C"
locale behavior. Caught by Category H, checked directly against both
the C standard's guarantee for `isalnum()` under the "C" locale and a
real CPython interpreter with `re.LOCALE` and the "C" locale explicitly
set, both of which classify only ASCII letters and digits. `LOCALE` is
now documented as an accepted no-op in `BINARY`/`ASCII` mode, matching
verified reality instead of a prior assumption (`concept.md` 13.4,
`docs/API.md` Section 2, `README.md` "Known deviations").
Both are described in full in `README.md`, `docs/API.md`, and
`concept.md`, not only here; this section exists so the connection
between "ran a lot of generated tests" and "found these two specific,
real bugs" is traceable from the plan that produced them.
Two further, smaller findings were concrete instances of an
already-documented category (glibc's `wctype.h` Unicode tables not
matching CPython's own bundled ones exactly, `README.md` "Known
deviations"), not new defects: U+00A0 (NO-BREAK SPACE, Category F) and
fullwidth digits `U+FF10`-`U+FF19` (Category I) are both recognized by
CPython's `\s`/`\d` but not by glibc's `iswspace()`/`iswdigit()` under the
`C.utf8` locale this build relies on. Recorded, not patched, for the same
reason already given for the first instance: hand-patching individual
code points would start down the path of maintaining an ad hoc table this
design deliberately avoided.
After these fixes, the random Category J generator was additionally run
against five more seeds (`1`, `42`, `777`, `99999`, `123456`) at 3,000
iterations each, 26,760 further checks beyond the 900 committed here, all
passing; only the committed seed's 900 cases are part of the regular test
suite, the rest was exploratory validation, not committed as a permanent
fixture, to avoid the build slowing down for marginal additional coverage
over what is already committed.
All of the above is described in full in `README.md`, `docs/API.md`, and
`concept.md`, not only here; this section exists so the connection between
"ran a lot of generated tests" and "found these specific, real bugs" is
traceable from the plan that produced them.
## Explicitly out of scope for this expansion
+183
View File
@@ -3,6 +3,7 @@
an operation; gen.py computes the expected result with Python's `re`
and emits a C assertion that checks regexx against it.
"""
import re
CASES = []
@@ -334,3 +335,185 @@ for _pat in _G_PATTERNS:
for _fl in _G_FLAG_COMBOS:
add("search", _pat, _subj, flags=_fl)
add("finditer", _pat, _subj, flags=_fl)
# =============================================================================
# Category H: BINARY mode with genuinely arbitrary raw bytes (not just UTF-8
# encoded text). `subject` here is a real Python `bytes` object; gen.py's
# encode() passes it through unchanged rather than UTF-8-encoding it, so
# these exercise byte values and byte sequences that are not, and are not
# meant to be, valid UTF-8 at all.
# =============================================================================
_H_CASES = [
(r"a.c", b"a\x00c"),
(r"a.c", b"a\xffc"),
(r"[\x00-\x1f]+", b"hello\x01\x02\x03world"),
(r"\x00", b"abc\x00def"),
(r".", b"\xff"),
(r".+", b"\x00\x01\x02\xfd\xfe\xff"),
(r"a\x00+b", b"a\x00\x00\x00b"),
(r"[\x80-\xff]+", b"abc\x80\x90\xa0\xffdef"),
(r"[\x80-\xff]+", b"abcdef"),
(r"\w+", b"abc\x80\x90def", "ascii"),
(r"\w+", b"abc\x80\x90def", "ascii", ("LOCALE",)),
(r"(.)\1", b"\x00\x00"),
(r"(.)\1", b"\x00\x01"),
(r"(\xff+)", b"\xff\xff\xff"),
(r"[^\x00]+", b"abc\x00def"),
(r"^\xff", b"\xffabc"),
(r"\xff$", b"abc\xff"),
(r"a", b"A", "binary", ("IGNORECASE",)),
(r"[a-z]+", b"ABC\x80abc", "binary", ("IGNORECASE",)),
]
for _row in _H_CASES:
_pat, _subj = _row[0], _row[1]
_mode = _row[2] if len(_row) > 2 else "binary"
_fl = _row[3] if len(_row) > 3 else ()
add("search", _pat, _subj, mode=_mode, flags=_fl)
add("finditer", _pat, _subj, mode=_mode, flags=_fl)
# Raw-byte sub/split: replacement templates and split still need to work
# correctly with embedded NUL and high bytes on both sides.
_H_SUB_CASES = [
(r"\x00", b"a\x00b\x00c", "-"),
(r"[\x80-\xff]", b"a\x80b\x90c", "?"),
(r"(.)\x00", b"a\x00b\x00", r"[\1]"),
]
for _pat, _subj, _repl in _H_SUB_CASES:
add("sub", _pat, _subj, mode="binary", repl=_repl)
_H_SPLIT_CASES = [
(r"\x00+", b"a\x00\x00b\x00c"),
(r"[\x80-\xff]", b"a\x80b\x90c\xa0d"),
]
for _pat, _subj in _H_SPLIT_CASES:
add("split", _pat, _subj, mode="binary")
# =============================================================================
# Category I: UTF-8 edge cases -- real multi-byte content beyond simple
# accented Latin: 4-byte (astral plane) code points, combining marks,
# right-to-left scripts, CJK, mixed-width strings, and offset correctness
# for backreferences/groups/lookaround spanning multi-byte characters.
# =============================================================================
_I_CASES = [
(r"\w+", "hello \U0001F600 world"), # emoji, a 4-byte code point
(r".", "\U0001F600"),
(r"\W", "\U0001F600"),
(r"(.)", "\U0001F600\U0001F601"),
(r"é", "é"), # 'e' + combining acute accent
(r"\w+", "éclair"),
(r"[؀-ۿ]+", "السلام"), # Arabic
(r"\w+", "שלום"), # Hebrew
(r"[一-鿿]+", "中文字符"), # CJK unified ideographs
(r"\w+", "日本語"), # Japanese
(r"(\w)(\w)\2\1", "文字字文"),
(r"(.)\1+", "\U0001F600\U0001F600\U0001F600"),
(r"^.{3}$", "aé\U0001F600"), # mixed 1/2/4-byte codepoints, fixed count
(r"\bworld\b", "café world café"),
(r"(?<=é)\w+", "cafémonde"),
(r"(?=\U0001F600)", "x\U0001F600y"),
(r"[^\x00-\x7f]+", "abcéèêdef"),
# Fullwidth digits (U+FF10-FF19, Unicode category Nd) are deliberately
# not probed here: CPython's \d matches them (it matches the Nd
# category), but glibc's iswdigit() under the C.utf8 locale this
# build relies on does not, the same class of Unicode-table gap as
# the NBSP case in README.md "Known deviations", which now also
# records this specific instance.
(r"café", "CAFÉ", "utf8", ("IGNORECASE",)),
(r"ß", "ß", "utf8", ("IGNORECASE",)), # German sharp s
]
for _row in _I_CASES:
_pat, _subj = _row[0], _row[1]
_mode = _row[2] if len(_row) > 2 else "utf8"
_fl = _row[3] if len(_row) > 3 else ()
add("search", _pat, _subj, mode=_mode, flags=_fl)
add("fullmatch", _pat, _subj, mode=_mode, flags=_fl)
add("finditer", _pat, _subj, mode=_mode, flags=_fl)
_I_SUB_CASES = [
(r"\U0001F600", "hi \U0001F600 there", ":)"),
(r"(\w)", "café", r"[\1]"),
(r"[一-鿿]", "中文abc", "?"),
]
for _pat, _subj, _repl in _I_SUB_CASES:
add("sub", _pat, _subj, mode="utf8", repl=_repl)
_I_SPLIT_CASES = [
(r"\s+", "café au lait 中文"),
(r"(\U0001F600)", "a\U0001F600b\U0001F600c"),
]
for _pat, _subj in _I_SPLIT_CASES:
add("split", _pat, _subj, mode="utf8")
# =============================================================================
# Category J: seeded random combinatorics. A fixed seed makes this a
# deterministic, reproducible regression test, not a flaky fuzzer: the same
# cases are generated every run, so a failure here is exactly as
# reproducible and reportable as a hand-written one. Explores combinations
# no one sat down and thought to write by hand, across all three modes.
# =============================================================================
import random as _random # noqa: E402
_rng = _random.Random(0xC0FFEE)
_J_ASCII_ATOMS = ["a", "b", "c", "1", "2", "_", " ", "[a-c]", "[0-9]", r"\d", r"\w", r"\s", "."]
_J_UTF8_ATOMS = _J_ASCII_ATOMS + ["é", "è", "中", "ا", "\U0001F600"]
_J_QUANTS = ["", "?", "*", "+", "{1,2}", "{0,2}", "*?", "+?", "??"]
_J_WRAPS = ["{0}", "({0})", "(?:{0})", "(?P<g%d>{0})"]
_J_ASCII_SUBJ_CHARS = list("abc123_ ")
_J_UTF8_SUBJ_CHARS = _J_ASCII_SUBJ_CHARS + ["é", "è", "中", "文", "ا", "\U0001F600", "\U0001F601"]
_J_SPECIAL_ATOMS = {r"\d", r"\w", r"\s", ".", "[a-c]", "[0-9]"}
def _j_random_fragment(mode, group_no):
atoms = _J_UTF8_ATOMS if mode == "utf8" else _J_ASCII_ATOMS
atom = _rng.choice(atoms)
if atom not in _J_SPECIAL_ATOMS:
atom = re.escape(atom)
q = _rng.choice(_J_QUANTS)
wrap = _rng.choice(_J_WRAPS)
if "%d" in wrap:
wrap = wrap % group_no
return wrap.format(atom + q)
def _j_random_pattern(mode):
n = _rng.randint(1, 4)
group_no = 1
parts = []
for _ in range(n):
frag = _j_random_fragment(mode, group_no)
if frag.startswith("(") and not frag.startswith("(?:"):
group_no += 1
parts.append(frag)
joiner = "|" if _rng.random() < 0.25 else ""
return joiner.join(parts)
def _j_random_subject(mode, length):
chars = _J_UTF8_SUBJ_CHARS if mode == "utf8" else _J_ASCII_SUBJ_CHARS
return "".join(_rng.choice(chars) for _ in range(length))
_J_FLAG_CHOICES = [[], ["IGNORECASE"], ["MULTILINE"], ["DOTALL"], ["IGNORECASE", "MULTILINE"]]
_J_MODES = ["ascii", "utf8", "binary"]
_J_OPS = ["search", "finditer", "fullmatch"]
# Patterns/subjects for "binary" mode are still built from the ASCII-safe
# atom pool (Category H already covers genuinely arbitrary, non-UTF-8
# binary content deliberately and by hand); here binary mode exercises
# byte-for-byte matching of ordinary ASCII text through the BINARY
# encoding path specifically, distinct from the same text through ASCII
# or UTF8 mode, which Category F already probes with a fixed pattern set.
_j_generated = 0
for _ in range(900):
_mode = _rng.choice(_J_MODES)
_atom_mode = "utf8" if _mode == "utf8" else "ascii"
_pat = _j_random_pattern(_atom_mode)
_subj = _j_random_subject(_atom_mode, _rng.randint(0, 10))
_fl = _rng.choice(_J_FLAG_CHOICES)
_op = _rng.choice(_J_OPS)
add(_op, _pat, _subj, mode=_mode, flags=_fl)
_j_generated += 1
+12 -5
View File
@@ -48,9 +48,14 @@ def c_str(b):
def py_flags(case):
# re.ASCII is not added here for mode == "ascii": ground truth for
# that mode is already a `bytes` pattern against a `bytes` subject
# (encode(), above), and a bytes pattern is inherently ASCII-only
# for \w/\s/\d regardless of re.ASCII, so adding it would be purely
# redundant, and actively wrong whenever a case also uses LOCALE,
# since Python rejects ASCII and LOCALE combined (confirmed
# directly: ValueError: ASCII and LOCALE flags are incompatible).
f = 0
if case["mode"] == "ascii":
f |= re.ASCII
for name in case["flags"]:
f |= FLAGMAP[name]
return f
@@ -83,6 +88,8 @@ def encode(case, s):
# latter, correctly). A `bytes` pattern against a `bytes` subject is
# always byte oriented with ASCII-only \w/\s/\d in Python too, so it
# is the correct ground truth for both ASCII and BINARY mode here.
if isinstance(s, bytes):
return s # already raw bytes (arbitrary, not necessarily valid UTF-8): pass through
if case["mode"] in ("binary", "ascii"):
return s.encode("utf-8")
return s
@@ -186,6 +193,8 @@ def main():
skipped = 0
for i, case in enumerate(CASES):
op = case["op"]
if op not in ("search", "match", "fullmatch", "finditer", "sub", "split"):
raise ValueError("unknown op %r" % op) # a typo in a hand-written case, not a skippable combination
try:
if op in ("search", "match", "fullmatch"):
gen_search_family(case, i, out)
@@ -195,10 +204,8 @@ def main():
gen_sub(case, i, out)
elif op == "split":
gen_split(case, i, out)
else:
raise ValueError("unknown op %r" % op)
emitted += 1
except re.error as e:
except (re.error, ValueError) as e:
# A combinatorially generated pattern that CPython itself
# cannot compile is not a useful ground-truth case (there is
# nothing to check regexx against); skip it rather than