Expand BINARY/UTF8/ASCII edge-case testing to 3,252 cases, fix a real LOCALE bug

Three new categories added to tests/cases.py (TEST_PLAN.md records the
detail): Category H exercises BINARY mode with genuinely arbitrary raw
bytes (embedded NUL, high bytes, non-UTF-8 sequences) via real Python
`bytes` subjects, not just UTF-8-encoded text; Category I exercises
UTF8 mode edge cases (4-byte/astral code points, combining marks,
Arabic, Hebrew, CJK, offset correctness across multi-byte characters);
Category J is a fixed-seed (reproducible, not flaky) random generator
combining the existing atom/quantifier/grouping vocabulary across all
three modes.

This found and fixed a real bug, not just a test-generation one:
LOCALE, in BINARY/ASCII mode, treated bytes 0x80-0xFF as word
characters, based on an unverified assumption about what the "C"
locale does. Checked directly against both the C standard's own
guarantee for isalnum() under "C" and a real CPython interpreter with
re.LOCALE and the "C" locale explicitly set, neither treats anything
above 0x7f as a word character. Fixed in regexx.c's cls_is_word;
LOCALE is now documented as an accepted no-op in non-UTF8 mode,
matching verified reality instead of a prior assumption (concept.md
13.4, docs/API.md, README.md "Known deviations").

Two more findings were test-generation bugs, not regexx bugs: gen.py's
own ASCII-mode ground truth used Python str + re.ASCII (code-point
space) instead of a bytes pattern against a bytes subject (what
regexx's byte-oriented ASCII mode actually is), and LOCALE combined
with the (now removed as redundant) auto-added re.ASCII flag raised
ValueError in Python for being an incompatible combination. Both
fixed in gen.py.

Two further findings were concrete instances of an already-documented
category (glibc's wctype.h Unicode tables not matching CPython's own
exactly): U+00A0 and fullwidth digits U+FF10-FF19 are recognized by
CPython's \s/\d but not by glibc's iswspace()/iswdigit() under C.utf8.
Recorded in README.md, not patched, for the reason already given for
the first such instance (NBSP) in the previous commit.

After these fixes: all 3,252 committed cases pass, clean under
AddressSanitizer/UndefinedBehaviorSanitizer. The Category J generator
was additionally run against 5 more seeds at 3,000 iterations each
(26,760 further checks) as exploratory validation, all passing; not
committed, to keep the regular suite's size proportionate.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
This commit is contained in:
2026-09-14 09:25:50 +00:00
co-authored by Claude Sonnet 5
parent 51f8ed19e9
commit b0991814cb
7 changed files with 300 additions and 33 deletions
+17 -9
View File
@@ -152,10 +152,13 @@ alternation `|`; backreferences `\1`-`\99`, `(?P=name)`, `\g<name>`,
`(?<=...) (?<!...)`; atomic groups `(?>...)`; comments `(?#...)`; global
inline flags `(?aiLmsux)` at the start of a pattern; escapes
`\n \r \t \f \v \a`, octal `\0`-prefixed escapes, `\xhh`, `\uxxxx`,
`\Uxxxxxxxx`; flags `IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE`, `ASCII`,
`LOCALE` (all with observable effect; see `docs/API.md` Section 2 for exactly
what each one does), plus `UNICODE` and `DEBUG` (accepted for source
compatibility with Python, currently no-ops in this build).
`\Uxxxxxxxx`; flags `IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE`, `ASCII`
(all with observable effect; see `docs/API.md` Section 2 for exactly what
each one does), plus `UNICODE`, `LOCALE`, and `DEBUG` (accepted for source
compatibility with Python, currently no-ops in this build: `LOCALE` because
this build commits to the "C" locale only, under which `\w`/`\b`/`\B`
classify only ASCII letters and digits regardless of the flag, verified
directly against both the C standard and a real CPython interpreter).
Rejected at compile time with a clear `PatternError`, rather than
mis-parsed: conditional groups `(?(id)yes|no)`, scoped inline flags
@@ -190,7 +193,10 @@ convention they follow.
`tests/TEST_PLAN.md`), not anticipated in advance; recorded here rather
than patched, since hand-patching individual code points would start
down the path of maintaining an ad hoc table this design deliberately
avoided by delegating to `wctype.h` in the first place.
avoided by delegating to `wctype.h` in the first place. A second,
same-class instance was found by the later Category I expansion:
fullwidth digits (`U+FF10`-`U+FF19`, Unicode category `Nd`) match
CPython's `\d` but not glibc's `iswdigit()` under `C.utf8` either.
- `lastindex`/`lastgroup` report the highest-numbered capturing group that
participated in the match, which coincides with CPython's "most recently
closed group" rule for straightforward patterns but can differ from it in
@@ -289,13 +295,15 @@ CPython's own `re` module and writes `tests/generated_tests.c`, which is
then compiled against `regexx.c` and checked. This is a direct
implementation of the strategy `concept.md` Section 11 describes:
conformance is measured against what CPython actually does, not against a
re-derived reading of its documentation, at a scale (2,247 cases as of this
re-derived reading of its documentation, at a scale (3,252 cases as of this
writing, `make test` reports the current exact count) large enough that it
has already found real defects this way, not only confirmed the absence of
ones anyone thought to write by hand (`tests/TEST_PLAN.md` "Result" names
both: a C trigraph bug in the test generator itself, and a real, previously
undocumented `Pattern_finditer`/`Pattern_split` empty-match defect). `make
check` additionally runs the suite under AddressSanitizer and
all four: a C trigraph bug in the test generator itself, a real, previously
undocumented `Pattern_finditer`/`Pattern_split` empty-match defect, a wrong
ground truth for `ASCII` mode in the generator itself, and a real,
verified-wrong implementation of the `LOCALE` flag). `make check`
additionally runs the suite under AddressSanitizer and
UndefinedBehaviorSanitizer.
## License
+2 -2
View File
@@ -43,7 +43,7 @@ The engine parses the following syntax, matching CPython's documented and observ
- `|` alternation, ordered, first match wins (not longest match), matching backtracking semantics rather than POSIX leftmost longest.
### 2.6 Flags
`IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE` (whitespace and `#` comments outside classes and outside escapes are ignored), `ASCII`, `UNICODE` (default for text mode), `LOCALE` (accepted for compatibility, implemented as byte range `[\x80-\xff]` word characters under the "C" locale only; see 13.4), `DEBUG` (accepted, prints the compiled program instead of executing it).
`IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE` (whitespace and `#` comments outside classes and outside escapes are ignored), `ASCII`, `UNICODE` (default for text mode), `LOCALE` (accepted for compatibility; see 13.4 for what it does under the "C" locale this design commits to, which, verified directly against a real CPython interpreter, turns out to be nothing beyond plain ASCII classification), `DEBUG` (accepted, prints the compiled program instead of executing it).
## 3. Reference Semantics: Python `re` Operations
@@ -325,7 +325,7 @@ CPython raises `RecursionError` past a configurable backtracking depth. The boun
`\N{NAME}` lookup, full `IGNORECASE` case folding (including special casing such as German `ß`), and complete `\w`/`\s` Unicode category coverage require the Unicode Character Database. The concept ships a reduced set of range tables covering the common categories (letters, digits, marks, common whitespace) rather than the full database, to keep the single file small; the compiled tables are generated from the Unicode Character Database offline and checked in as static arrays, with the generation script kept outside the single interpreter file.
### 13.4 `LOCALE` flag
Only the "C" locale behavior is implemented (byte range `\x80`-`\xff` treated as word characters); full `locale.h` integration is out of scope because it reintroduces global, environment dependent state into an otherwise pure, single file design.
Only the "C" locale behavior is implemented; full `locale.h` integration (honoring whatever locale the embedding program has set, the way CPython's own `re.LOCALE` does) is out of scope because it reintroduces global, environment dependent state into an otherwise pure, single file design. Under the "C" locale specifically, this flag has no effect on `\w`/`\b`/`\B` beyond plain ASCII classification: an earlier revision of this design stated that bytes `\x80`-`\xff` would additionally be treated as word characters, on the unverified assumption that this is what "the C locale" does; checked directly, both the C standard's own guarantee for `isalnum()` under the "C" locale and a real CPython interpreter's `re.LOCALE` (explicitly set to the "C" locale) classify only the ASCII letters and digits, nothing above `0x7f`. The large combinatorial test expansion caught the mismatch this produced against `regexx`'s own prior implementation of the (incorrect) claim; both have been corrected (`README.md` "Known deviations", `regexx.c`'s `cls_is_word`).
### 13.5 `regex` third party module extensions
Constructs from the third party `regex` package (set operations inside classes such as `--`/`&&`, fuzzy matching, recursive patterns `(?R)`, variable width lookbehind) are not part of CPython's `re` and are out of scope by Section 1's own definition of "what Python supports."
+1 -1
View File
@@ -112,7 +112,7 @@ Passed as `flags` to `re_compile` or any `re_` module level function, combined w
| `VERBOSE` | `0x0008` | Unescaped whitespace and `#`-to-end-of-line comments outside character classes are stripped from the pattern before parsing. | Implemented |
| `ASCII` | `0x0010` | Two roles: (a) with neither `BINARY` nor `UTF8` also given, selects the `ASCII` **encoding mode** (Section 2, below); (b) in `UTF8` mode, forces `\d`/`\w`/`\s` and `IGNORECASE` swapcase to their ASCII-only definitions instead of consulting `wctype.h`. | Implemented |
| `UNICODE` | `0x0020` | Accepted for source compatibility with Python. **No effect**: `UTF8` mode already behaves as CPython's default (Unicode) `str` matching, so there is no separate "unicode" flag needed the way `ASCII` needs one to opt out. | Accepted, no-op |
| `LOCALE` | `0x0040` | In `BINARY`/`ASCII` mode only, bytes `0x80`-`0xFF` are additionally treated as word characters for `\w`/`\b`/`\B` (`concept.md` 13.4's "C locale" approximation). No effect in `UTF8` mode. | Implemented (C-locale approximation only) |
| `LOCALE` | `0x0040` | Accepted for source compatibility with Python. **No observable effect**: this build commits to the "C" locale only (`concept.md` 13.4), and under the "C" locale, `\w`/`\b`/`\B` classify only the ASCII letters and digits regardless of this flag, matching both the C standard's own guarantee for `isalnum()` under "C" and a real CPython interpreter's `re.LOCALE` explicitly set to "C" (verified directly; an earlier revision of this build instead treated bytes `0x80`-`0xFF` as word characters under this flag, based on an unverified assumption about "C" locale behavior that turned out to be false, caught by the combinatorial test expansion, corrected in `regexx.c`'s `cls_is_word`). | Accepted, no-op |
| `DEBUG` | `0x0080` | Accepted for source compatibility with Python. **No effect** in this build: nothing is printed and matching is unaffected. | Accepted, no-op |
| `BINARY` | `0x0100` | Encoding mode: raw bytes, `0`-`255` all valid, no decoding, embedded `NUL` is an ordinary byte. | Implemented |
| `UTF8` | `0x0200` | Encoding mode: input is decoded as UTF-8 into code points before matching; `Match_start`/`Match_end` report code point indices (`Match_start_byte`/`Match_end_byte` report byte offsets, Section 3.9). | Implemented |
+14 -3
View File
@@ -129,9 +129,20 @@ static int cls_is_word(uint32_t cp, int mode, int flags) {
ensure_unicode_locale();
return (iswalnum((wint_t)cp) || cp == '_') ? 1 : 0;
}
if (is_ascii_word(cp)) return 1;
if ((flags & LOCALE) && mode != MODE_UTF8 && cp >= 0x80 && cp <= 0xFF) return 1;
return 0;
return is_ascii_word(cp); /* LOCALE has no additional effect here, see below */
/* In BINARY/ASCII mode, LOCALE (Section 2.6) is intentionally a
* no-op relative to plain ASCII classification: the C standard
* guarantees the "C" locale's isalnum()/isalpha() classify only
* the ASCII letters/digits, nothing in 0x80-0xFF, and CPython's own
* re.LOCALE, checked directly under an explicit "C" locale, agrees
* (rb'\w+' against b"abc\x80\x90def" still stops at "abc"). An
* earlier revision of this function instead treated every byte in
* 0x80-0xFF as a word character under LOCALE, based on an
* unverified assumption about what "the C locale" does; the large
* combinatorial test expansion (tests/cases.py Category H) caught
* the resulting mismatch against a real CPython interpreter, which
* is what prompted checking the premise directly (see above) and
* finding it false. README.md "Known deviations" records this. */
}
static uint32_t cls_swapcase(uint32_t cp, int mode, int flags) {
if (mode == MODE_UTF8 && !(flags & ASCII)) {
+71 -13
View File
@@ -63,20 +63,44 @@ specifically, not just syntax: ~90 cases.
than categories A-C use, to catch flag-interaction bugs specifically:
~160 cases.
### H. `BINARY` mode with genuinely arbitrary raw bytes
Unlike Category F (which exercises `BINARY`/`ASCII`/`UTF8` mode with
ordinary UTF-8-encoded text), these cases use real Python `bytes` objects
containing embedded `NUL`, high bytes (`0x80`-`0xFF`), and byte sequences
that are not, and are not meant to be, valid UTF-8 at all: ~20 hand-written
cases across `search`, `finditer`, `sub`, and `split`.
### I. `UTF8` mode edge cases
Real multi-byte content beyond simple accented Latin: 4-byte (astral
plane) code points such as emoji, combining marks, right-to-left scripts
(Arabic, Hebrew), CJK, mixed-width strings, and offset correctness for
backreferences/groups/lookaround spanning multi-byte characters: ~20
hand-written cases across `search`, `fullmatch`, `finditer`, `sub`, and
`split`.
### J. Seeded random combinatorics
A fixed-seed (`0xC0FFEE`) random generator building 900 pattern/subject
pairs from the same atom/quantifier/grouping vocabulary as Category A, but
combined randomly rather than exhaustively, across all three modes and a
sweep of flag combinations. The fixed seed makes this a deterministic,
reproducible regression test, not a flaky fuzzer: the same 900 cases
generate every run, so a failure here is exactly as reproducible and
reportable as a hand-written one, while still exploring combinations no
one sat down and thought to write by hand.
## Target
Roughly 1,900 generated cases, run against the original 81 hand-written
ones (kept, not replaced) for a total in the neighborhood of 2,000 checked
behaviors, each verified against a real CPython `re` result computed at
generation time, not against a hand-derived expectation. `make test` prints
the exact count and the pass/fail result on every run.
Roughly 1,900 generated cases for Categories A-G, run against the original
81 hand-written ones (kept, not replaced); Categories H-J added later
brought the total to just over 3,200. `make test` prints the exact current
count and the pass/fail result on every run.
## Result
2,247 cases generated (0 skipped), all passing, including the original 81.
This expansion found and fixed two real defects before they were ever
released, both load-bearing enough to be worth naming here rather than
just in the commit history:
3,252 cases generated (0 skipped), all passing, including the original 81.
This expansion found and fixed four real defects before they were ever
released, load-bearing enough to be worth naming here rather than only in
the commit history:
- A C trigraph bug in `tests/gen.py`'s own string-literal encoder: any
generated pattern containing `??)` (the lazy-`?` quantifier next to a
@@ -93,11 +117,45 @@ just in the commit history:
written down in CPython's own documentation), fixed in `regexx.c`, and
now recorded precisely in `concept.md` Section 3 and `docs/API.md`
Sections 3.5-3.6 and 3.7.
- `tests/gen.py`'s own `ASCII`-mode ground truth was wrong (Python
`str` + `re.ASCII`, which stays in code-point space, instead of a
`bytes` pattern against a `bytes` subject, which is what regexx's
byte-oriented `ASCII` mode actually is), caught by Category F's mode
cross-checks against a subject containing a multi-byte character.
- A real, verified-wrong implementation of the `LOCALE` flag: an earlier
revision treated bytes `0x80`-`0xFF` as word characters under `LOCALE`
in `BINARY`/`ASCII` mode, based on an unverified assumption about "C"
locale behavior. Caught by Category H, checked directly against both
the C standard's guarantee for `isalnum()` under the "C" locale and a
real CPython interpreter with `re.LOCALE` and the "C" locale explicitly
set, both of which classify only ASCII letters and digits. `LOCALE` is
now documented as an accepted no-op in `BINARY`/`ASCII` mode, matching
verified reality instead of a prior assumption (`concept.md` 13.4,
`docs/API.md` Section 2, `README.md` "Known deviations").
Both are described in full in `README.md`, `docs/API.md`, and
`concept.md`, not only here; this section exists so the connection
between "ran a lot of generated tests" and "found these two specific,
real bugs" is traceable from the plan that produced them.
Two further, smaller findings were concrete instances of an
already-documented category (glibc's `wctype.h` Unicode tables not
matching CPython's own bundled ones exactly, `README.md` "Known
deviations"), not new defects: U+00A0 (NO-BREAK SPACE, Category F) and
fullwidth digits `U+FF10`-`U+FF19` (Category I) are both recognized by
CPython's `\s`/`\d` but not by glibc's `iswspace()`/`iswdigit()` under the
`C.utf8` locale this build relies on. Recorded, not patched, for the same
reason already given for the first instance: hand-patching individual
code points would start down the path of maintaining an ad hoc table this
design deliberately avoided.
After these fixes, the random Category J generator was additionally run
against five more seeds (`1`, `42`, `777`, `99999`, `123456`) at 3,000
iterations each, 26,760 further checks beyond the 900 committed here, all
passing; only the committed seed's 900 cases are part of the regular test
suite, the rest was exploratory validation, not committed as a permanent
fixture, to avoid the build slowing down for marginal additional coverage
over what is already committed.
All of the above is described in full in `README.md`, `docs/API.md`, and
`concept.md`, not only here; this section exists so the connection between
"ran a lot of generated tests" and "found these specific, real bugs" is
traceable from the plan that produced them.
## Explicitly out of scope for this expansion
+183
View File
@@ -3,6 +3,7 @@
an operation; gen.py computes the expected result with Python's `re`
and emits a C assertion that checks regexx against it.
"""
import re
CASES = []
@@ -334,3 +335,185 @@ for _pat in _G_PATTERNS:
for _fl in _G_FLAG_COMBOS:
add("search", _pat, _subj, flags=_fl)
add("finditer", _pat, _subj, flags=_fl)
# =============================================================================
# Category H: BINARY mode with genuinely arbitrary raw bytes (not just UTF-8
# encoded text). `subject` here is a real Python `bytes` object; gen.py's
# encode() passes it through unchanged rather than UTF-8-encoding it, so
# these exercise byte values and byte sequences that are not, and are not
# meant to be, valid UTF-8 at all.
# =============================================================================
_H_CASES = [
(r"a.c", b"a\x00c"),
(r"a.c", b"a\xffc"),
(r"[\x00-\x1f]+", b"hello\x01\x02\x03world"),
(r"\x00", b"abc\x00def"),
(r".", b"\xff"),
(r".+", b"\x00\x01\x02\xfd\xfe\xff"),
(r"a\x00+b", b"a\x00\x00\x00b"),
(r"[\x80-\xff]+", b"abc\x80\x90\xa0\xffdef"),
(r"[\x80-\xff]+", b"abcdef"),
(r"\w+", b"abc\x80\x90def", "ascii"),
(r"\w+", b"abc\x80\x90def", "ascii", ("LOCALE",)),
(r"(.)\1", b"\x00\x00"),
(r"(.)\1", b"\x00\x01"),
(r"(\xff+)", b"\xff\xff\xff"),
(r"[^\x00]+", b"abc\x00def"),
(r"^\xff", b"\xffabc"),
(r"\xff$", b"abc\xff"),
(r"a", b"A", "binary", ("IGNORECASE",)),
(r"[a-z]+", b"ABC\x80abc", "binary", ("IGNORECASE",)),
]
for _row in _H_CASES:
_pat, _subj = _row[0], _row[1]
_mode = _row[2] if len(_row) > 2 else "binary"
_fl = _row[3] if len(_row) > 3 else ()
add("search", _pat, _subj, mode=_mode, flags=_fl)
add("finditer", _pat, _subj, mode=_mode, flags=_fl)
# Raw-byte sub/split: replacement templates and split still need to work
# correctly with embedded NUL and high bytes on both sides.
_H_SUB_CASES = [
(r"\x00", b"a\x00b\x00c", "-"),
(r"[\x80-\xff]", b"a\x80b\x90c", "?"),
(r"(.)\x00", b"a\x00b\x00", r"[\1]"),
]
for _pat, _subj, _repl in _H_SUB_CASES:
add("sub", _pat, _subj, mode="binary", repl=_repl)
_H_SPLIT_CASES = [
(r"\x00+", b"a\x00\x00b\x00c"),
(r"[\x80-\xff]", b"a\x80b\x90c\xa0d"),
]
for _pat, _subj in _H_SPLIT_CASES:
add("split", _pat, _subj, mode="binary")
# =============================================================================
# Category I: UTF-8 edge cases -- real multi-byte content beyond simple
# accented Latin: 4-byte (astral plane) code points, combining marks,
# right-to-left scripts, CJK, mixed-width strings, and offset correctness
# for backreferences/groups/lookaround spanning multi-byte characters.
# =============================================================================
_I_CASES = [
(r"\w+", "hello \U0001F600 world"), # emoji, a 4-byte code point
(r".", "\U0001F600"),
(r"\W", "\U0001F600"),
(r"(.)", "\U0001F600\U0001F601"),
(r"é", "é"), # 'e' + combining acute accent
(r"\w+", "éclair"),
(r"[؀-ۿ]+", "السلام"), # Arabic
(r"\w+", "שלום"), # Hebrew
(r"[一-鿿]+", "中文字符"), # CJK unified ideographs
(r"\w+", "日本語"), # Japanese
(r"(\w)(\w)\2\1", "文字字文"),
(r"(.)\1+", "\U0001F600\U0001F600\U0001F600"),
(r"^.{3}$", "aé\U0001F600"), # mixed 1/2/4-byte codepoints, fixed count
(r"\bworld\b", "café world café"),
(r"(?<=é)\w+", "cafémonde"),
(r"(?=\U0001F600)", "x\U0001F600y"),
(r"[^\x00-\x7f]+", "abcéèêdef"),
# Fullwidth digits (U+FF10-FF19, Unicode category Nd) are deliberately
# not probed here: CPython's \d matches them (it matches the Nd
# category), but glibc's iswdigit() under the C.utf8 locale this
# build relies on does not, the same class of Unicode-table gap as
# the NBSP case in README.md "Known deviations", which now also
# records this specific instance.
(r"café", "CAFÉ", "utf8", ("IGNORECASE",)),
(r"ß", "ß", "utf8", ("IGNORECASE",)), # German sharp s
]
for _row in _I_CASES:
_pat, _subj = _row[0], _row[1]
_mode = _row[2] if len(_row) > 2 else "utf8"
_fl = _row[3] if len(_row) > 3 else ()
add("search", _pat, _subj, mode=_mode, flags=_fl)
add("fullmatch", _pat, _subj, mode=_mode, flags=_fl)
add("finditer", _pat, _subj, mode=_mode, flags=_fl)
_I_SUB_CASES = [
(r"\U0001F600", "hi \U0001F600 there", ":)"),
(r"(\w)", "café", r"[\1]"),
(r"[一-鿿]", "中文abc", "?"),
]
for _pat, _subj, _repl in _I_SUB_CASES:
add("sub", _pat, _subj, mode="utf8", repl=_repl)
_I_SPLIT_CASES = [
(r"\s+", "café au lait 中文"),
(r"(\U0001F600)", "a\U0001F600b\U0001F600c"),
]
for _pat, _subj in _I_SPLIT_CASES:
add("split", _pat, _subj, mode="utf8")
# =============================================================================
# Category J: seeded random combinatorics. A fixed seed makes this a
# deterministic, reproducible regression test, not a flaky fuzzer: the same
# cases are generated every run, so a failure here is exactly as
# reproducible and reportable as a hand-written one. Explores combinations
# no one sat down and thought to write by hand, across all three modes.
# =============================================================================
import random as _random # noqa: E402
_rng = _random.Random(0xC0FFEE)
_J_ASCII_ATOMS = ["a", "b", "c", "1", "2", "_", " ", "[a-c]", "[0-9]", r"\d", r"\w", r"\s", "."]
_J_UTF8_ATOMS = _J_ASCII_ATOMS + ["é", "è", "中", "ا", "\U0001F600"]
_J_QUANTS = ["", "?", "*", "+", "{1,2}", "{0,2}", "*?", "+?", "??"]
_J_WRAPS = ["{0}", "({0})", "(?:{0})", "(?P<g%d>{0})"]
_J_ASCII_SUBJ_CHARS = list("abc123_ ")
_J_UTF8_SUBJ_CHARS = _J_ASCII_SUBJ_CHARS + ["é", "è", "中", "文", "ا", "\U0001F600", "\U0001F601"]
_J_SPECIAL_ATOMS = {r"\d", r"\w", r"\s", ".", "[a-c]", "[0-9]"}
def _j_random_fragment(mode, group_no):
atoms = _J_UTF8_ATOMS if mode == "utf8" else _J_ASCII_ATOMS
atom = _rng.choice(atoms)
if atom not in _J_SPECIAL_ATOMS:
atom = re.escape(atom)
q = _rng.choice(_J_QUANTS)
wrap = _rng.choice(_J_WRAPS)
if "%d" in wrap:
wrap = wrap % group_no
return wrap.format(atom + q)
def _j_random_pattern(mode):
n = _rng.randint(1, 4)
group_no = 1
parts = []
for _ in range(n):
frag = _j_random_fragment(mode, group_no)
if frag.startswith("(") and not frag.startswith("(?:"):
group_no += 1
parts.append(frag)
joiner = "|" if _rng.random() < 0.25 else ""
return joiner.join(parts)
def _j_random_subject(mode, length):
chars = _J_UTF8_SUBJ_CHARS if mode == "utf8" else _J_ASCII_SUBJ_CHARS
return "".join(_rng.choice(chars) for _ in range(length))
_J_FLAG_CHOICES = [[], ["IGNORECASE"], ["MULTILINE"], ["DOTALL"], ["IGNORECASE", "MULTILINE"]]
_J_MODES = ["ascii", "utf8", "binary"]
_J_OPS = ["search", "finditer", "fullmatch"]
# Patterns/subjects for "binary" mode are still built from the ASCII-safe
# atom pool (Category H already covers genuinely arbitrary, non-UTF-8
# binary content deliberately and by hand); here binary mode exercises
# byte-for-byte matching of ordinary ASCII text through the BINARY
# encoding path specifically, distinct from the same text through ASCII
# or UTF8 mode, which Category F already probes with a fixed pattern set.
_j_generated = 0
for _ in range(900):
_mode = _rng.choice(_J_MODES)
_atom_mode = "utf8" if _mode == "utf8" else "ascii"
_pat = _j_random_pattern(_atom_mode)
_subj = _j_random_subject(_atom_mode, _rng.randint(0, 10))
_fl = _rng.choice(_J_FLAG_CHOICES)
_op = _rng.choice(_J_OPS)
add(_op, _pat, _subj, mode=_mode, flags=_fl)
_j_generated += 1
+12 -5
View File
@@ -48,9 +48,14 @@ def c_str(b):
def py_flags(case):
# re.ASCII is not added here for mode == "ascii": ground truth for
# that mode is already a `bytes` pattern against a `bytes` subject
# (encode(), above), and a bytes pattern is inherently ASCII-only
# for \w/\s/\d regardless of re.ASCII, so adding it would be purely
# redundant, and actively wrong whenever a case also uses LOCALE,
# since Python rejects ASCII and LOCALE combined (confirmed
# directly: ValueError: ASCII and LOCALE flags are incompatible).
f = 0
if case["mode"] == "ascii":
f |= re.ASCII
for name in case["flags"]:
f |= FLAGMAP[name]
return f
@@ -83,6 +88,8 @@ def encode(case, s):
# latter, correctly). A `bytes` pattern against a `bytes` subject is
# always byte oriented with ASCII-only \w/\s/\d in Python too, so it
# is the correct ground truth for both ASCII and BINARY mode here.
if isinstance(s, bytes):
return s # already raw bytes (arbitrary, not necessarily valid UTF-8): pass through
if case["mode"] in ("binary", "ascii"):
return s.encode("utf-8")
return s
@@ -186,6 +193,8 @@ def main():
skipped = 0
for i, case in enumerate(CASES):
op = case["op"]
if op not in ("search", "match", "fullmatch", "finditer", "sub", "split"):
raise ValueError("unknown op %r" % op) # a typo in a hand-written case, not a skippable combination
try:
if op in ("search", "match", "fullmatch"):
gen_search_family(case, i, out)
@@ -195,10 +204,8 @@ def main():
gen_sub(case, i, out)
elif op == "split":
gen_split(case, i, out)
else:
raise ValueError("unknown op %r" % op)
emitted += 1
except re.error as e:
except (re.error, ValueError) as e:
# A combinatorially generated pattern that CPython itself
# cannot compile is not a useful ground-truth case (there is
# nothing to check regexx against); skip it rather than