Files
regexx/tests
retoorandClaude Sonnet 5 b0991814cb Expand BINARY/UTF8/ASCII edge-case testing to 3,252 cases, fix a real LOCALE bug
Three new categories added to tests/cases.py (TEST_PLAN.md records the
detail): Category H exercises BINARY mode with genuinely arbitrary raw
bytes (embedded NUL, high bytes, non-UTF-8 sequences) via real Python
`bytes` subjects, not just UTF-8-encoded text; Category I exercises
UTF8 mode edge cases (4-byte/astral code points, combining marks,
Arabic, Hebrew, CJK, offset correctness across multi-byte characters);
Category J is a fixed-seed (reproducible, not flaky) random generator
combining the existing atom/quantifier/grouping vocabulary across all
three modes.

This found and fixed a real bug, not just a test-generation one:
LOCALE, in BINARY/ASCII mode, treated bytes 0x80-0xFF as word
characters, based on an unverified assumption about what the "C"
locale does. Checked directly against both the C standard's own
guarantee for isalnum() under "C" and a real CPython interpreter with
re.LOCALE and the "C" locale explicitly set, neither treats anything
above 0x7f as a word character. Fixed in regexx.c's cls_is_word;
LOCALE is now documented as an accepted no-op in non-UTF8 mode,
matching verified reality instead of a prior assumption (concept.md
13.4, docs/API.md, README.md "Known deviations").

Two more findings were test-generation bugs, not regexx bugs: gen.py's
own ASCII-mode ground truth used Python str + re.ASCII (code-point
space) instead of a bytes pattern against a bytes subject (what
regexx's byte-oriented ASCII mode actually is), and LOCALE combined
with the (now removed as redundant) auto-added re.ASCII flag raised
ValueError in Python for being an incompatible combination. Both
fixed in gen.py.

Two further findings were concrete instances of an already-documented
category (glibc's wctype.h Unicode tables not matching CPython's own
exactly): U+00A0 and fullwidth digits U+FF10-FF19 are recognized by
CPython's \s/\d but not by glibc's iswspace()/iswdigit() under C.utf8.
Recorded in README.md, not patched, for the reason already given for
the first such instance (NBSP) in the previous commit.

After these fixes: all 3,252 committed cases pass, clean under
AddressSanitizer/UndefinedBehaviorSanitizer. The Category J generator
was additionally run against 5 more seeds at 3,000 iterations each
(26,760 further checks) as exploratory validation, all passing; not
committed, to keep the regular suite's size proportionate.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
2026-09-14 09:25:50 +00:00
..