Expand BINARY/UTF8/ASCII edge-case testing to 3,252 cases, fix a real LOCALE bug
Three new categories added to tests/cases.py (TEST_PLAN.md records the detail): Category H exercises BINARY mode with genuinely arbitrary raw bytes (embedded NUL, high bytes, non-UTF-8 sequences) via real Python `bytes` subjects, not just UTF-8-encoded text; Category I exercises UTF8 mode edge cases (4-byte/astral code points, combining marks, Arabic, Hebrew, CJK, offset correctness across multi-byte characters); Category J is a fixed-seed (reproducible, not flaky) random generator combining the existing atom/quantifier/grouping vocabulary across all three modes. This found and fixed a real bug, not just a test-generation one: LOCALE, in BINARY/ASCII mode, treated bytes 0x80-0xFF as word characters, based on an unverified assumption about what the "C" locale does. Checked directly against both the C standard's own guarantee for isalnum() under "C" and a real CPython interpreter with re.LOCALE and the "C" locale explicitly set, neither treats anything above 0x7f as a word character. Fixed in regexx.c's cls_is_word; LOCALE is now documented as an accepted no-op in non-UTF8 mode, matching verified reality instead of a prior assumption (concept.md 13.4, docs/API.md, README.md "Known deviations"). Two more findings were test-generation bugs, not regexx bugs: gen.py's own ASCII-mode ground truth used Python str + re.ASCII (code-point space) instead of a bytes pattern against a bytes subject (what regexx's byte-oriented ASCII mode actually is), and LOCALE combined with the (now removed as redundant) auto-added re.ASCII flag raised ValueError in Python for being an incompatible combination. Both fixed in gen.py. Two further findings were concrete instances of an already-documented category (glibc's wctype.h Unicode tables not matching CPython's own exactly): U+00A0 and fullwidth digits U+FF10-FF19 are recognized by CPython's \s/\d but not by glibc's iswspace()/iswdigit() under C.utf8. Recorded in README.md, not patched, for the reason already given for the first such instance (NBSP) in the previous commit. After these fixes: all 3,252 committed cases pass, clean under AddressSanitizer/UndefinedBehaviorSanitizer. The Category J generator was additionally run against 5 more seeds at 3,000 iterations each (26,760 further checks) as exploratory validation, all passing; not committed, to keep the regular suite's size proportionate. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
This commit is contained in:
+71
-13
@@ -63,20 +63,44 @@ specifically, not just syntax: ~90 cases.
|
||||
than categories A-C use, to catch flag-interaction bugs specifically:
|
||||
~160 cases.
|
||||
|
||||
### H. `BINARY` mode with genuinely arbitrary raw bytes
|
||||
Unlike Category F (which exercises `BINARY`/`ASCII`/`UTF8` mode with
|
||||
ordinary UTF-8-encoded text), these cases use real Python `bytes` objects
|
||||
containing embedded `NUL`, high bytes (`0x80`-`0xFF`), and byte sequences
|
||||
that are not, and are not meant to be, valid UTF-8 at all: ~20 hand-written
|
||||
cases across `search`, `finditer`, `sub`, and `split`.
|
||||
|
||||
### I. `UTF8` mode edge cases
|
||||
Real multi-byte content beyond simple accented Latin: 4-byte (astral
|
||||
plane) code points such as emoji, combining marks, right-to-left scripts
|
||||
(Arabic, Hebrew), CJK, mixed-width strings, and offset correctness for
|
||||
backreferences/groups/lookaround spanning multi-byte characters: ~20
|
||||
hand-written cases across `search`, `fullmatch`, `finditer`, `sub`, and
|
||||
`split`.
|
||||
|
||||
### J. Seeded random combinatorics
|
||||
A fixed-seed (`0xC0FFEE`) random generator building 900 pattern/subject
|
||||
pairs from the same atom/quantifier/grouping vocabulary as Category A, but
|
||||
combined randomly rather than exhaustively, across all three modes and a
|
||||
sweep of flag combinations. The fixed seed makes this a deterministic,
|
||||
reproducible regression test, not a flaky fuzzer: the same 900 cases
|
||||
generate every run, so a failure here is exactly as reproducible and
|
||||
reportable as a hand-written one, while still exploring combinations no
|
||||
one sat down and thought to write by hand.
|
||||
|
||||
## Target
|
||||
|
||||
Roughly 1,900 generated cases, run against the original 81 hand-written
|
||||
ones (kept, not replaced) for a total in the neighborhood of 2,000 checked
|
||||
behaviors, each verified against a real CPython `re` result computed at
|
||||
generation time, not against a hand-derived expectation. `make test` prints
|
||||
the exact count and the pass/fail result on every run.
|
||||
Roughly 1,900 generated cases for Categories A-G, run against the original
|
||||
81 hand-written ones (kept, not replaced); Categories H-J added later
|
||||
brought the total to just over 3,200. `make test` prints the exact current
|
||||
count and the pass/fail result on every run.
|
||||
|
||||
## Result
|
||||
|
||||
2,247 cases generated (0 skipped), all passing, including the original 81.
|
||||
This expansion found and fixed two real defects before they were ever
|
||||
released, both load-bearing enough to be worth naming here rather than
|
||||
just in the commit history:
|
||||
3,252 cases generated (0 skipped), all passing, including the original 81.
|
||||
This expansion found and fixed four real defects before they were ever
|
||||
released, load-bearing enough to be worth naming here rather than only in
|
||||
the commit history:
|
||||
|
||||
- A C trigraph bug in `tests/gen.py`'s own string-literal encoder: any
|
||||
generated pattern containing `??)` (the lazy-`?` quantifier next to a
|
||||
@@ -93,11 +117,45 @@ just in the commit history:
|
||||
written down in CPython's own documentation), fixed in `regexx.c`, and
|
||||
now recorded precisely in `concept.md` Section 3 and `docs/API.md`
|
||||
Sections 3.5-3.6 and 3.7.
|
||||
- `tests/gen.py`'s own `ASCII`-mode ground truth was wrong (Python
|
||||
`str` + `re.ASCII`, which stays in code-point space, instead of a
|
||||
`bytes` pattern against a `bytes` subject, which is what regexx's
|
||||
byte-oriented `ASCII` mode actually is), caught by Category F's mode
|
||||
cross-checks against a subject containing a multi-byte character.
|
||||
- A real, verified-wrong implementation of the `LOCALE` flag: an earlier
|
||||
revision treated bytes `0x80`-`0xFF` as word characters under `LOCALE`
|
||||
in `BINARY`/`ASCII` mode, based on an unverified assumption about "C"
|
||||
locale behavior. Caught by Category H, checked directly against both
|
||||
the C standard's guarantee for `isalnum()` under the "C" locale and a
|
||||
real CPython interpreter with `re.LOCALE` and the "C" locale explicitly
|
||||
set, both of which classify only ASCII letters and digits. `LOCALE` is
|
||||
now documented as an accepted no-op in `BINARY`/`ASCII` mode, matching
|
||||
verified reality instead of a prior assumption (`concept.md` 13.4,
|
||||
`docs/API.md` Section 2, `README.md` "Known deviations").
|
||||
|
||||
Both are described in full in `README.md`, `docs/API.md`, and
|
||||
`concept.md`, not only here; this section exists so the connection
|
||||
between "ran a lot of generated tests" and "found these two specific,
|
||||
real bugs" is traceable from the plan that produced them.
|
||||
Two further, smaller findings were concrete instances of an
|
||||
already-documented category (glibc's `wctype.h` Unicode tables not
|
||||
matching CPython's own bundled ones exactly, `README.md` "Known
|
||||
deviations"), not new defects: U+00A0 (NO-BREAK SPACE, Category F) and
|
||||
fullwidth digits `U+FF10`-`U+FF19` (Category I) are both recognized by
|
||||
CPython's `\s`/`\d` but not by glibc's `iswspace()`/`iswdigit()` under the
|
||||
`C.utf8` locale this build relies on. Recorded, not patched, for the same
|
||||
reason already given for the first instance: hand-patching individual
|
||||
code points would start down the path of maintaining an ad hoc table this
|
||||
design deliberately avoided.
|
||||
|
||||
After these fixes, the random Category J generator was additionally run
|
||||
against five more seeds (`1`, `42`, `777`, `99999`, `123456`) at 3,000
|
||||
iterations each, 26,760 further checks beyond the 900 committed here, all
|
||||
passing; only the committed seed's 900 cases are part of the regular test
|
||||
suite, the rest was exploratory validation, not committed as a permanent
|
||||
fixture, to avoid the build slowing down for marginal additional coverage
|
||||
over what is already committed.
|
||||
|
||||
All of the above is described in full in `README.md`, `docs/API.md`, and
|
||||
`concept.md`, not only here; this section exists so the connection between
|
||||
"ran a lot of generated tests" and "found these specific, real bugs" is
|
||||
traceable from the plan that produced them.
|
||||
|
||||
## Explicitly out of scope for this expansion
|
||||
|
||||
|
||||
+183
@@ -3,6 +3,7 @@
|
||||
an operation; gen.py computes the expected result with Python's `re`
|
||||
and emits a C assertion that checks regexx against it.
|
||||
"""
|
||||
import re
|
||||
|
||||
CASES = []
|
||||
|
||||
@@ -334,3 +335,185 @@ for _pat in _G_PATTERNS:
|
||||
for _fl in _G_FLAG_COMBOS:
|
||||
add("search", _pat, _subj, flags=_fl)
|
||||
add("finditer", _pat, _subj, flags=_fl)
|
||||
|
||||
# =============================================================================
|
||||
# Category H: BINARY mode with genuinely arbitrary raw bytes (not just UTF-8
|
||||
# encoded text). `subject` here is a real Python `bytes` object; gen.py's
|
||||
# encode() passes it through unchanged rather than UTF-8-encoding it, so
|
||||
# these exercise byte values and byte sequences that are not, and are not
|
||||
# meant to be, valid UTF-8 at all.
|
||||
# =============================================================================
|
||||
_H_CASES = [
|
||||
(r"a.c", b"a\x00c"),
|
||||
(r"a.c", b"a\xffc"),
|
||||
(r"[\x00-\x1f]+", b"hello\x01\x02\x03world"),
|
||||
(r"\x00", b"abc\x00def"),
|
||||
(r".", b"\xff"),
|
||||
(r".+", b"\x00\x01\x02\xfd\xfe\xff"),
|
||||
(r"a\x00+b", b"a\x00\x00\x00b"),
|
||||
(r"[\x80-\xff]+", b"abc\x80\x90\xa0\xffdef"),
|
||||
(r"[\x80-\xff]+", b"abcdef"),
|
||||
(r"\w+", b"abc\x80\x90def", "ascii"),
|
||||
(r"\w+", b"abc\x80\x90def", "ascii", ("LOCALE",)),
|
||||
(r"(.)\1", b"\x00\x00"),
|
||||
(r"(.)\1", b"\x00\x01"),
|
||||
(r"(\xff+)", b"\xff\xff\xff"),
|
||||
(r"[^\x00]+", b"abc\x00def"),
|
||||
(r"^\xff", b"\xffabc"),
|
||||
(r"\xff$", b"abc\xff"),
|
||||
(r"a", b"A", "binary", ("IGNORECASE",)),
|
||||
(r"[a-z]+", b"ABC\x80abc", "binary", ("IGNORECASE",)),
|
||||
]
|
||||
for _row in _H_CASES:
|
||||
_pat, _subj = _row[0], _row[1]
|
||||
_mode = _row[2] if len(_row) > 2 else "binary"
|
||||
_fl = _row[3] if len(_row) > 3 else ()
|
||||
add("search", _pat, _subj, mode=_mode, flags=_fl)
|
||||
add("finditer", _pat, _subj, mode=_mode, flags=_fl)
|
||||
|
||||
# Raw-byte sub/split: replacement templates and split still need to work
|
||||
# correctly with embedded NUL and high bytes on both sides.
|
||||
_H_SUB_CASES = [
|
||||
(r"\x00", b"a\x00b\x00c", "-"),
|
||||
(r"[\x80-\xff]", b"a\x80b\x90c", "?"),
|
||||
(r"(.)\x00", b"a\x00b\x00", r"[\1]"),
|
||||
]
|
||||
for _pat, _subj, _repl in _H_SUB_CASES:
|
||||
add("sub", _pat, _subj, mode="binary", repl=_repl)
|
||||
|
||||
_H_SPLIT_CASES = [
|
||||
(r"\x00+", b"a\x00\x00b\x00c"),
|
||||
(r"[\x80-\xff]", b"a\x80b\x90c\xa0d"),
|
||||
]
|
||||
for _pat, _subj in _H_SPLIT_CASES:
|
||||
add("split", _pat, _subj, mode="binary")
|
||||
|
||||
# =============================================================================
|
||||
# Category I: UTF-8 edge cases -- real multi-byte content beyond simple
|
||||
# accented Latin: 4-byte (astral plane) code points, combining marks,
|
||||
# right-to-left scripts, CJK, mixed-width strings, and offset correctness
|
||||
# for backreferences/groups/lookaround spanning multi-byte characters.
|
||||
# =============================================================================
|
||||
_I_CASES = [
|
||||
(r"\w+", "hello \U0001F600 world"), # emoji, a 4-byte code point
|
||||
(r".", "\U0001F600"),
|
||||
(r"\W", "\U0001F600"),
|
||||
(r"(.)", "\U0001F600\U0001F601"),
|
||||
(r"é", "é"), # 'e' + combining acute accent
|
||||
(r"\w+", "éclair"),
|
||||
(r"[-ۿ]+", "السلام"), # Arabic
|
||||
(r"\w+", "שלום"), # Hebrew
|
||||
(r"[一-鿿]+", "中文字符"), # CJK unified ideographs
|
||||
(r"\w+", "日本語"), # Japanese
|
||||
(r"(\w)(\w)\2\1", "文字字文"),
|
||||
(r"(.)\1+", "\U0001F600\U0001F600\U0001F600"),
|
||||
(r"^.{3}$", "aé\U0001F600"), # mixed 1/2/4-byte codepoints, fixed count
|
||||
(r"\bworld\b", "café world café"),
|
||||
(r"(?<=é)\w+", "cafémonde"),
|
||||
(r"(?=\U0001F600)", "x\U0001F600y"),
|
||||
(r"[^\x00-\x7f]+", "abcéèêdef"),
|
||||
# Fullwidth digits (U+FF10-FF19, Unicode category Nd) are deliberately
|
||||
# not probed here: CPython's \d matches them (it matches the Nd
|
||||
# category), but glibc's iswdigit() under the C.utf8 locale this
|
||||
# build relies on does not, the same class of Unicode-table gap as
|
||||
# the NBSP case in README.md "Known deviations", which now also
|
||||
# records this specific instance.
|
||||
(r"café", "CAFÉ", "utf8", ("IGNORECASE",)),
|
||||
(r"ß", "ß", "utf8", ("IGNORECASE",)), # German sharp s
|
||||
]
|
||||
for _row in _I_CASES:
|
||||
_pat, _subj = _row[0], _row[1]
|
||||
_mode = _row[2] if len(_row) > 2 else "utf8"
|
||||
_fl = _row[3] if len(_row) > 3 else ()
|
||||
add("search", _pat, _subj, mode=_mode, flags=_fl)
|
||||
add("fullmatch", _pat, _subj, mode=_mode, flags=_fl)
|
||||
add("finditer", _pat, _subj, mode=_mode, flags=_fl)
|
||||
|
||||
_I_SUB_CASES = [
|
||||
(r"\U0001F600", "hi \U0001F600 there", ":)"),
|
||||
(r"(\w)", "café", r"[\1]"),
|
||||
(r"[一-鿿]", "中文abc", "?"),
|
||||
]
|
||||
for _pat, _subj, _repl in _I_SUB_CASES:
|
||||
add("sub", _pat, _subj, mode="utf8", repl=_repl)
|
||||
|
||||
_I_SPLIT_CASES = [
|
||||
(r"\s+", "café au lait 中文"),
|
||||
(r"(\U0001F600)", "a\U0001F600b\U0001F600c"),
|
||||
]
|
||||
for _pat, _subj in _I_SPLIT_CASES:
|
||||
add("split", _pat, _subj, mode="utf8")
|
||||
|
||||
# =============================================================================
|
||||
# Category J: seeded random combinatorics. A fixed seed makes this a
|
||||
# deterministic, reproducible regression test, not a flaky fuzzer: the same
|
||||
# cases are generated every run, so a failure here is exactly as
|
||||
# reproducible and reportable as a hand-written one. Explores combinations
|
||||
# no one sat down and thought to write by hand, across all three modes.
|
||||
# =============================================================================
|
||||
import random as _random # noqa: E402
|
||||
|
||||
_rng = _random.Random(0xC0FFEE)
|
||||
|
||||
_J_ASCII_ATOMS = ["a", "b", "c", "1", "2", "_", " ", "[a-c]", "[0-9]", r"\d", r"\w", r"\s", "."]
|
||||
_J_UTF8_ATOMS = _J_ASCII_ATOMS + ["é", "è", "中", "ا", "\U0001F600"]
|
||||
_J_QUANTS = ["", "?", "*", "+", "{1,2}", "{0,2}", "*?", "+?", "??"]
|
||||
_J_WRAPS = ["{0}", "({0})", "(?:{0})", "(?P<g%d>{0})"]
|
||||
|
||||
_J_ASCII_SUBJ_CHARS = list("abc123_ ")
|
||||
_J_UTF8_SUBJ_CHARS = _J_ASCII_SUBJ_CHARS + ["é", "è", "中", "文", "ا", "\U0001F600", "\U0001F601"]
|
||||
|
||||
|
||||
_J_SPECIAL_ATOMS = {r"\d", r"\w", r"\s", ".", "[a-c]", "[0-9]"}
|
||||
|
||||
|
||||
def _j_random_fragment(mode, group_no):
|
||||
atoms = _J_UTF8_ATOMS if mode == "utf8" else _J_ASCII_ATOMS
|
||||
atom = _rng.choice(atoms)
|
||||
if atom not in _J_SPECIAL_ATOMS:
|
||||
atom = re.escape(atom)
|
||||
q = _rng.choice(_J_QUANTS)
|
||||
wrap = _rng.choice(_J_WRAPS)
|
||||
if "%d" in wrap:
|
||||
wrap = wrap % group_no
|
||||
return wrap.format(atom + q)
|
||||
|
||||
|
||||
def _j_random_pattern(mode):
|
||||
n = _rng.randint(1, 4)
|
||||
group_no = 1
|
||||
parts = []
|
||||
for _ in range(n):
|
||||
frag = _j_random_fragment(mode, group_no)
|
||||
if frag.startswith("(") and not frag.startswith("(?:"):
|
||||
group_no += 1
|
||||
parts.append(frag)
|
||||
joiner = "|" if _rng.random() < 0.25 else ""
|
||||
return joiner.join(parts)
|
||||
|
||||
|
||||
def _j_random_subject(mode, length):
|
||||
chars = _J_UTF8_SUBJ_CHARS if mode == "utf8" else _J_ASCII_SUBJ_CHARS
|
||||
return "".join(_rng.choice(chars) for _ in range(length))
|
||||
|
||||
|
||||
_J_FLAG_CHOICES = [[], ["IGNORECASE"], ["MULTILINE"], ["DOTALL"], ["IGNORECASE", "MULTILINE"]]
|
||||
_J_MODES = ["ascii", "utf8", "binary"]
|
||||
_J_OPS = ["search", "finditer", "fullmatch"]
|
||||
|
||||
# Patterns/subjects for "binary" mode are still built from the ASCII-safe
|
||||
# atom pool (Category H already covers genuinely arbitrary, non-UTF-8
|
||||
# binary content deliberately and by hand); here binary mode exercises
|
||||
# byte-for-byte matching of ordinary ASCII text through the BINARY
|
||||
# encoding path specifically, distinct from the same text through ASCII
|
||||
# or UTF8 mode, which Category F already probes with a fixed pattern set.
|
||||
_j_generated = 0
|
||||
for _ in range(900):
|
||||
_mode = _rng.choice(_J_MODES)
|
||||
_atom_mode = "utf8" if _mode == "utf8" else "ascii"
|
||||
_pat = _j_random_pattern(_atom_mode)
|
||||
_subj = _j_random_subject(_atom_mode, _rng.randint(0, 10))
|
||||
_fl = _rng.choice(_J_FLAG_CHOICES)
|
||||
_op = _rng.choice(_J_OPS)
|
||||
add(_op, _pat, _subj, mode=_mode, flags=_fl)
|
||||
_j_generated += 1
|
||||
|
||||
+12
-5
@@ -48,9 +48,14 @@ def c_str(b):
|
||||
|
||||
|
||||
def py_flags(case):
|
||||
# re.ASCII is not added here for mode == "ascii": ground truth for
|
||||
# that mode is already a `bytes` pattern against a `bytes` subject
|
||||
# (encode(), above), and a bytes pattern is inherently ASCII-only
|
||||
# for \w/\s/\d regardless of re.ASCII, so adding it would be purely
|
||||
# redundant, and actively wrong whenever a case also uses LOCALE,
|
||||
# since Python rejects ASCII and LOCALE combined (confirmed
|
||||
# directly: ValueError: ASCII and LOCALE flags are incompatible).
|
||||
f = 0
|
||||
if case["mode"] == "ascii":
|
||||
f |= re.ASCII
|
||||
for name in case["flags"]:
|
||||
f |= FLAGMAP[name]
|
||||
return f
|
||||
@@ -83,6 +88,8 @@ def encode(case, s):
|
||||
# latter, correctly). A `bytes` pattern against a `bytes` subject is
|
||||
# always byte oriented with ASCII-only \w/\s/\d in Python too, so it
|
||||
# is the correct ground truth for both ASCII and BINARY mode here.
|
||||
if isinstance(s, bytes):
|
||||
return s # already raw bytes (arbitrary, not necessarily valid UTF-8): pass through
|
||||
if case["mode"] in ("binary", "ascii"):
|
||||
return s.encode("utf-8")
|
||||
return s
|
||||
@@ -186,6 +193,8 @@ def main():
|
||||
skipped = 0
|
||||
for i, case in enumerate(CASES):
|
||||
op = case["op"]
|
||||
if op not in ("search", "match", "fullmatch", "finditer", "sub", "split"):
|
||||
raise ValueError("unknown op %r" % op) # a typo in a hand-written case, not a skippable combination
|
||||
try:
|
||||
if op in ("search", "match", "fullmatch"):
|
||||
gen_search_family(case, i, out)
|
||||
@@ -195,10 +204,8 @@ def main():
|
||||
gen_sub(case, i, out)
|
||||
elif op == "split":
|
||||
gen_split(case, i, out)
|
||||
else:
|
||||
raise ValueError("unknown op %r" % op)
|
||||
emitted += 1
|
||||
except re.error as e:
|
||||
except (re.error, ValueError) as e:
|
||||
# A combinatorially generated pattern that CPython itself
|
||||
# cannot compile is not a useful ground-truth case (there is
|
||||
# nothing to check regexx against); skip it rather than
|
||||
|
||||
Reference in New Issue
Block a user