Expand BINARY/UTF8/ASCII edge-case testing to 3,252 cases, fix a real LOCALE bug
Three new categories added to tests/cases.py (TEST_PLAN.md records the detail): Category H exercises BINARY mode with genuinely arbitrary raw bytes (embedded NUL, high bytes, non-UTF-8 sequences) via real Python `bytes` subjects, not just UTF-8-encoded text; Category I exercises UTF8 mode edge cases (4-byte/astral code points, combining marks, Arabic, Hebrew, CJK, offset correctness across multi-byte characters); Category J is a fixed-seed (reproducible, not flaky) random generator combining the existing atom/quantifier/grouping vocabulary across all three modes. This found and fixed a real bug, not just a test-generation one: LOCALE, in BINARY/ASCII mode, treated bytes 0x80-0xFF as word characters, based on an unverified assumption about what the "C" locale does. Checked directly against both the C standard's own guarantee for isalnum() under "C" and a real CPython interpreter with re.LOCALE and the "C" locale explicitly set, neither treats anything above 0x7f as a word character. Fixed in regexx.c's cls_is_word; LOCALE is now documented as an accepted no-op in non-UTF8 mode, matching verified reality instead of a prior assumption (concept.md 13.4, docs/API.md, README.md "Known deviations"). Two more findings were test-generation bugs, not regexx bugs: gen.py's own ASCII-mode ground truth used Python str + re.ASCII (code-point space) instead of a bytes pattern against a bytes subject (what regexx's byte-oriented ASCII mode actually is), and LOCALE combined with the (now removed as redundant) auto-added re.ASCII flag raised ValueError in Python for being an incompatible combination. Both fixed in gen.py. Two further findings were concrete instances of an already-documented category (glibc's wctype.h Unicode tables not matching CPython's own exactly): U+00A0 and fullwidth digits U+FF10-FF19 are recognized by CPython's \s/\d but not by glibc's iswspace()/iswdigit() under C.utf8. Recorded in README.md, not patched, for the reason already given for the first such instance (NBSP) in the previous commit. After these fixes: all 3,252 committed cases pass, clean under AddressSanitizer/UndefinedBehaviorSanitizer. The Category J generator was additionally run against 5 more seeds at 3,000 iterations each (26,760 further checks) as exploratory validation, all passing; not committed, to keep the regular suite's size proportionate. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
This commit is contained in:
@@ -152,10 +152,13 @@ alternation `|`; backreferences `\1`-`\99`, `(?P=name)`, `\g<name>`,
|
||||
`(?<=...) (?<!...)`; atomic groups `(?>...)`; comments `(?#...)`; global
|
||||
inline flags `(?aiLmsux)` at the start of a pattern; escapes
|
||||
`\n \r \t \f \v \a`, octal `\0`-prefixed escapes, `\xhh`, `\uxxxx`,
|
||||
`\Uxxxxxxxx`; flags `IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE`, `ASCII`,
|
||||
`LOCALE` (all with observable effect; see `docs/API.md` Section 2 for exactly
|
||||
what each one does), plus `UNICODE` and `DEBUG` (accepted for source
|
||||
compatibility with Python, currently no-ops in this build).
|
||||
`\Uxxxxxxxx`; flags `IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE`, `ASCII`
|
||||
(all with observable effect; see `docs/API.md` Section 2 for exactly what
|
||||
each one does), plus `UNICODE`, `LOCALE`, and `DEBUG` (accepted for source
|
||||
compatibility with Python, currently no-ops in this build: `LOCALE` because
|
||||
this build commits to the "C" locale only, under which `\w`/`\b`/`\B`
|
||||
classify only ASCII letters and digits regardless of the flag, verified
|
||||
directly against both the C standard and a real CPython interpreter).
|
||||
|
||||
Rejected at compile time with a clear `PatternError`, rather than
|
||||
mis-parsed: conditional groups `(?(id)yes|no)`, scoped inline flags
|
||||
@@ -190,7 +193,10 @@ convention they follow.
|
||||
`tests/TEST_PLAN.md`), not anticipated in advance; recorded here rather
|
||||
than patched, since hand-patching individual code points would start
|
||||
down the path of maintaining an ad hoc table this design deliberately
|
||||
avoided by delegating to `wctype.h` in the first place.
|
||||
avoided by delegating to `wctype.h` in the first place. A second,
|
||||
same-class instance was found by the later Category I expansion:
|
||||
fullwidth digits (`U+FF10`-`U+FF19`, Unicode category `Nd`) match
|
||||
CPython's `\d` but not glibc's `iswdigit()` under `C.utf8` either.
|
||||
- `lastindex`/`lastgroup` report the highest-numbered capturing group that
|
||||
participated in the match, which coincides with CPython's "most recently
|
||||
closed group" rule for straightforward patterns but can differ from it in
|
||||
@@ -289,13 +295,15 @@ CPython's own `re` module and writes `tests/generated_tests.c`, which is
|
||||
then compiled against `regexx.c` and checked. This is a direct
|
||||
implementation of the strategy `concept.md` Section 11 describes:
|
||||
conformance is measured against what CPython actually does, not against a
|
||||
re-derived reading of its documentation, at a scale (2,247 cases as of this
|
||||
re-derived reading of its documentation, at a scale (3,252 cases as of this
|
||||
writing, `make test` reports the current exact count) large enough that it
|
||||
has already found real defects this way, not only confirmed the absence of
|
||||
ones anyone thought to write by hand (`tests/TEST_PLAN.md` "Result" names
|
||||
both: a C trigraph bug in the test generator itself, and a real, previously
|
||||
undocumented `Pattern_finditer`/`Pattern_split` empty-match defect). `make
|
||||
check` additionally runs the suite under AddressSanitizer and
|
||||
all four: a C trigraph bug in the test generator itself, a real, previously
|
||||
undocumented `Pattern_finditer`/`Pattern_split` empty-match defect, a wrong
|
||||
ground truth for `ASCII` mode in the generator itself, and a real,
|
||||
verified-wrong implementation of the `LOCALE` flag). `make check`
|
||||
additionally runs the suite under AddressSanitizer and
|
||||
UndefinedBehaviorSanitizer.
|
||||
|
||||
## License
|
||||
|
||||
+2
-2
@@ -43,7 +43,7 @@ The engine parses the following syntax, matching CPython's documented and observ
|
||||
- `|` alternation, ordered, first match wins (not longest match), matching backtracking semantics rather than POSIX leftmost longest.
|
||||
|
||||
### 2.6 Flags
|
||||
`IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE` (whitespace and `#` comments outside classes and outside escapes are ignored), `ASCII`, `UNICODE` (default for text mode), `LOCALE` (accepted for compatibility, implemented as byte range `[\x80-\xff]` word characters under the "C" locale only; see 13.4), `DEBUG` (accepted, prints the compiled program instead of executing it).
|
||||
`IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE` (whitespace and `#` comments outside classes and outside escapes are ignored), `ASCII`, `UNICODE` (default for text mode), `LOCALE` (accepted for compatibility; see 13.4 for what it does under the "C" locale this design commits to, which, verified directly against a real CPython interpreter, turns out to be nothing beyond plain ASCII classification), `DEBUG` (accepted, prints the compiled program instead of executing it).
|
||||
|
||||
## 3. Reference Semantics: Python `re` Operations
|
||||
|
||||
@@ -325,7 +325,7 @@ CPython raises `RecursionError` past a configurable backtracking depth. The boun
|
||||
`\N{NAME}` lookup, full `IGNORECASE` case folding (including special casing such as German `ß`), and complete `\w`/`\s` Unicode category coverage require the Unicode Character Database. The concept ships a reduced set of range tables covering the common categories (letters, digits, marks, common whitespace) rather than the full database, to keep the single file small; the compiled tables are generated from the Unicode Character Database offline and checked in as static arrays, with the generation script kept outside the single interpreter file.
|
||||
|
||||
### 13.4 `LOCALE` flag
|
||||
Only the "C" locale behavior is implemented (byte range `\x80`-`\xff` treated as word characters); full `locale.h` integration is out of scope because it reintroduces global, environment dependent state into an otherwise pure, single file design.
|
||||
Only the "C" locale behavior is implemented; full `locale.h` integration (honoring whatever locale the embedding program has set, the way CPython's own `re.LOCALE` does) is out of scope because it reintroduces global, environment dependent state into an otherwise pure, single file design. Under the "C" locale specifically, this flag has no effect on `\w`/`\b`/`\B` beyond plain ASCII classification: an earlier revision of this design stated that bytes `\x80`-`\xff` would additionally be treated as word characters, on the unverified assumption that this is what "the C locale" does; checked directly, both the C standard's own guarantee for `isalnum()` under the "C" locale and a real CPython interpreter's `re.LOCALE` (explicitly set to the "C" locale) classify only the ASCII letters and digits, nothing above `0x7f`. The large combinatorial test expansion caught the mismatch this produced against `regexx`'s own prior implementation of the (incorrect) claim; both have been corrected (`README.md` "Known deviations", `regexx.c`'s `cls_is_word`).
|
||||
|
||||
### 13.5 `regex` third party module extensions
|
||||
Constructs from the third party `regex` package (set operations inside classes such as `--`/`&&`, fuzzy matching, recursive patterns `(?R)`, variable width lookbehind) are not part of CPython's `re` and are out of scope by Section 1's own definition of "what Python supports."
|
||||
|
||||
+1
-1
@@ -112,7 +112,7 @@ Passed as `flags` to `re_compile` or any `re_` module level function, combined w
|
||||
| `VERBOSE` | `0x0008` | Unescaped whitespace and `#`-to-end-of-line comments outside character classes are stripped from the pattern before parsing. | Implemented |
|
||||
| `ASCII` | `0x0010` | Two roles: (a) with neither `BINARY` nor `UTF8` also given, selects the `ASCII` **encoding mode** (Section 2, below); (b) in `UTF8` mode, forces `\d`/`\w`/`\s` and `IGNORECASE` swapcase to their ASCII-only definitions instead of consulting `wctype.h`. | Implemented |
|
||||
| `UNICODE` | `0x0020` | Accepted for source compatibility with Python. **No effect**: `UTF8` mode already behaves as CPython's default (Unicode) `str` matching, so there is no separate "unicode" flag needed the way `ASCII` needs one to opt out. | Accepted, no-op |
|
||||
| `LOCALE` | `0x0040` | In `BINARY`/`ASCII` mode only, bytes `0x80`-`0xFF` are additionally treated as word characters for `\w`/`\b`/`\B` (`concept.md` 13.4's "C locale" approximation). No effect in `UTF8` mode. | Implemented (C-locale approximation only) |
|
||||
| `LOCALE` | `0x0040` | Accepted for source compatibility with Python. **No observable effect**: this build commits to the "C" locale only (`concept.md` 13.4), and under the "C" locale, `\w`/`\b`/`\B` classify only the ASCII letters and digits regardless of this flag, matching both the C standard's own guarantee for `isalnum()` under "C" and a real CPython interpreter's `re.LOCALE` explicitly set to "C" (verified directly; an earlier revision of this build instead treated bytes `0x80`-`0xFF` as word characters under this flag, based on an unverified assumption about "C" locale behavior that turned out to be false, caught by the combinatorial test expansion, corrected in `regexx.c`'s `cls_is_word`). | Accepted, no-op |
|
||||
| `DEBUG` | `0x0080` | Accepted for source compatibility with Python. **No effect** in this build: nothing is printed and matching is unaffected. | Accepted, no-op |
|
||||
| `BINARY` | `0x0100` | Encoding mode: raw bytes, `0`-`255` all valid, no decoding, embedded `NUL` is an ordinary byte. | Implemented |
|
||||
| `UTF8` | `0x0200` | Encoding mode: input is decoded as UTF-8 into code points before matching; `Match_start`/`Match_end` report code point indices (`Match_start_byte`/`Match_end_byte` report byte offsets, Section 3.9). | Implemented |
|
||||
|
||||
@@ -129,9 +129,20 @@ static int cls_is_word(uint32_t cp, int mode, int flags) {
|
||||
ensure_unicode_locale();
|
||||
return (iswalnum((wint_t)cp) || cp == '_') ? 1 : 0;
|
||||
}
|
||||
if (is_ascii_word(cp)) return 1;
|
||||
if ((flags & LOCALE) && mode != MODE_UTF8 && cp >= 0x80 && cp <= 0xFF) return 1;
|
||||
return 0;
|
||||
return is_ascii_word(cp); /* LOCALE has no additional effect here, see below */
|
||||
/* In BINARY/ASCII mode, LOCALE (Section 2.6) is intentionally a
|
||||
* no-op relative to plain ASCII classification: the C standard
|
||||
* guarantees the "C" locale's isalnum()/isalpha() classify only
|
||||
* the ASCII letters/digits, nothing in 0x80-0xFF, and CPython's own
|
||||
* re.LOCALE, checked directly under an explicit "C" locale, agrees
|
||||
* (rb'\w+' against b"abc\x80\x90def" still stops at "abc"). An
|
||||
* earlier revision of this function instead treated every byte in
|
||||
* 0x80-0xFF as a word character under LOCALE, based on an
|
||||
* unverified assumption about what "the C locale" does; the large
|
||||
* combinatorial test expansion (tests/cases.py Category H) caught
|
||||
* the resulting mismatch against a real CPython interpreter, which
|
||||
* is what prompted checking the premise directly (see above) and
|
||||
* finding it false. README.md "Known deviations" records this. */
|
||||
}
|
||||
static uint32_t cls_swapcase(uint32_t cp, int mode, int flags) {
|
||||
if (mode == MODE_UTF8 && !(flags & ASCII)) {
|
||||
|
||||
+71
-13
@@ -63,20 +63,44 @@ specifically, not just syntax: ~90 cases.
|
||||
than categories A-C use, to catch flag-interaction bugs specifically:
|
||||
~160 cases.
|
||||
|
||||
### H. `BINARY` mode with genuinely arbitrary raw bytes
|
||||
Unlike Category F (which exercises `BINARY`/`ASCII`/`UTF8` mode with
|
||||
ordinary UTF-8-encoded text), these cases use real Python `bytes` objects
|
||||
containing embedded `NUL`, high bytes (`0x80`-`0xFF`), and byte sequences
|
||||
that are not, and are not meant to be, valid UTF-8 at all: ~20 hand-written
|
||||
cases across `search`, `finditer`, `sub`, and `split`.
|
||||
|
||||
### I. `UTF8` mode edge cases
|
||||
Real multi-byte content beyond simple accented Latin: 4-byte (astral
|
||||
plane) code points such as emoji, combining marks, right-to-left scripts
|
||||
(Arabic, Hebrew), CJK, mixed-width strings, and offset correctness for
|
||||
backreferences/groups/lookaround spanning multi-byte characters: ~20
|
||||
hand-written cases across `search`, `fullmatch`, `finditer`, `sub`, and
|
||||
`split`.
|
||||
|
||||
### J. Seeded random combinatorics
|
||||
A fixed-seed (`0xC0FFEE`) random generator building 900 pattern/subject
|
||||
pairs from the same atom/quantifier/grouping vocabulary as Category A, but
|
||||
combined randomly rather than exhaustively, across all three modes and a
|
||||
sweep of flag combinations. The fixed seed makes this a deterministic,
|
||||
reproducible regression test, not a flaky fuzzer: the same 900 cases
|
||||
generate every run, so a failure here is exactly as reproducible and
|
||||
reportable as a hand-written one, while still exploring combinations no
|
||||
one sat down and thought to write by hand.
|
||||
|
||||
## Target
|
||||
|
||||
Roughly 1,900 generated cases, run against the original 81 hand-written
|
||||
ones (kept, not replaced) for a total in the neighborhood of 2,000 checked
|
||||
behaviors, each verified against a real CPython `re` result computed at
|
||||
generation time, not against a hand-derived expectation. `make test` prints
|
||||
the exact count and the pass/fail result on every run.
|
||||
Roughly 1,900 generated cases for Categories A-G, run against the original
|
||||
81 hand-written ones (kept, not replaced); Categories H-J added later
|
||||
brought the total to just over 3,200. `make test` prints the exact current
|
||||
count and the pass/fail result on every run.
|
||||
|
||||
## Result
|
||||
|
||||
2,247 cases generated (0 skipped), all passing, including the original 81.
|
||||
This expansion found and fixed two real defects before they were ever
|
||||
released, both load-bearing enough to be worth naming here rather than
|
||||
just in the commit history:
|
||||
3,252 cases generated (0 skipped), all passing, including the original 81.
|
||||
This expansion found and fixed four real defects before they were ever
|
||||
released, load-bearing enough to be worth naming here rather than only in
|
||||
the commit history:
|
||||
|
||||
- A C trigraph bug in `tests/gen.py`'s own string-literal encoder: any
|
||||
generated pattern containing `??)` (the lazy-`?` quantifier next to a
|
||||
@@ -93,11 +117,45 @@ just in the commit history:
|
||||
written down in CPython's own documentation), fixed in `regexx.c`, and
|
||||
now recorded precisely in `concept.md` Section 3 and `docs/API.md`
|
||||
Sections 3.5-3.6 and 3.7.
|
||||
- `tests/gen.py`'s own `ASCII`-mode ground truth was wrong (Python
|
||||
`str` + `re.ASCII`, which stays in code-point space, instead of a
|
||||
`bytes` pattern against a `bytes` subject, which is what regexx's
|
||||
byte-oriented `ASCII` mode actually is), caught by Category F's mode
|
||||
cross-checks against a subject containing a multi-byte character.
|
||||
- A real, verified-wrong implementation of the `LOCALE` flag: an earlier
|
||||
revision treated bytes `0x80`-`0xFF` as word characters under `LOCALE`
|
||||
in `BINARY`/`ASCII` mode, based on an unverified assumption about "C"
|
||||
locale behavior. Caught by Category H, checked directly against both
|
||||
the C standard's guarantee for `isalnum()` under the "C" locale and a
|
||||
real CPython interpreter with `re.LOCALE` and the "C" locale explicitly
|
||||
set, both of which classify only ASCII letters and digits. `LOCALE` is
|
||||
now documented as an accepted no-op in `BINARY`/`ASCII` mode, matching
|
||||
verified reality instead of a prior assumption (`concept.md` 13.4,
|
||||
`docs/API.md` Section 2, `README.md` "Known deviations").
|
||||
|
||||
Both are described in full in `README.md`, `docs/API.md`, and
|
||||
`concept.md`, not only here; this section exists so the connection
|
||||
between "ran a lot of generated tests" and "found these two specific,
|
||||
real bugs" is traceable from the plan that produced them.
|
||||
Two further, smaller findings were concrete instances of an
|
||||
already-documented category (glibc's `wctype.h` Unicode tables not
|
||||
matching CPython's own bundled ones exactly, `README.md` "Known
|
||||
deviations"), not new defects: U+00A0 (NO-BREAK SPACE, Category F) and
|
||||
fullwidth digits `U+FF10`-`U+FF19` (Category I) are both recognized by
|
||||
CPython's `\s`/`\d` but not by glibc's `iswspace()`/`iswdigit()` under the
|
||||
`C.utf8` locale this build relies on. Recorded, not patched, for the same
|
||||
reason already given for the first instance: hand-patching individual
|
||||
code points would start down the path of maintaining an ad hoc table this
|
||||
design deliberately avoided.
|
||||
|
||||
After these fixes, the random Category J generator was additionally run
|
||||
against five more seeds (`1`, `42`, `777`, `99999`, `123456`) at 3,000
|
||||
iterations each, 26,760 further checks beyond the 900 committed here, all
|
||||
passing; only the committed seed's 900 cases are part of the regular test
|
||||
suite, the rest was exploratory validation, not committed as a permanent
|
||||
fixture, to avoid the build slowing down for marginal additional coverage
|
||||
over what is already committed.
|
||||
|
||||
All of the above is described in full in `README.md`, `docs/API.md`, and
|
||||
`concept.md`, not only here; this section exists so the connection between
|
||||
"ran a lot of generated tests" and "found these specific, real bugs" is
|
||||
traceable from the plan that produced them.
|
||||
|
||||
## Explicitly out of scope for this expansion
|
||||
|
||||
|
||||
+183
@@ -3,6 +3,7 @@
|
||||
an operation; gen.py computes the expected result with Python's `re`
|
||||
and emits a C assertion that checks regexx against it.
|
||||
"""
|
||||
import re
|
||||
|
||||
CASES = []
|
||||
|
||||
@@ -334,3 +335,185 @@ for _pat in _G_PATTERNS:
|
||||
for _fl in _G_FLAG_COMBOS:
|
||||
add("search", _pat, _subj, flags=_fl)
|
||||
add("finditer", _pat, _subj, flags=_fl)
|
||||
|
||||
# =============================================================================
|
||||
# Category H: BINARY mode with genuinely arbitrary raw bytes (not just UTF-8
|
||||
# encoded text). `subject` here is a real Python `bytes` object; gen.py's
|
||||
# encode() passes it through unchanged rather than UTF-8-encoding it, so
|
||||
# these exercise byte values and byte sequences that are not, and are not
|
||||
# meant to be, valid UTF-8 at all.
|
||||
# =============================================================================
|
||||
_H_CASES = [
|
||||
(r"a.c", b"a\x00c"),
|
||||
(r"a.c", b"a\xffc"),
|
||||
(r"[\x00-\x1f]+", b"hello\x01\x02\x03world"),
|
||||
(r"\x00", b"abc\x00def"),
|
||||
(r".", b"\xff"),
|
||||
(r".+", b"\x00\x01\x02\xfd\xfe\xff"),
|
||||
(r"a\x00+b", b"a\x00\x00\x00b"),
|
||||
(r"[\x80-\xff]+", b"abc\x80\x90\xa0\xffdef"),
|
||||
(r"[\x80-\xff]+", b"abcdef"),
|
||||
(r"\w+", b"abc\x80\x90def", "ascii"),
|
||||
(r"\w+", b"abc\x80\x90def", "ascii", ("LOCALE",)),
|
||||
(r"(.)\1", b"\x00\x00"),
|
||||
(r"(.)\1", b"\x00\x01"),
|
||||
(r"(\xff+)", b"\xff\xff\xff"),
|
||||
(r"[^\x00]+", b"abc\x00def"),
|
||||
(r"^\xff", b"\xffabc"),
|
||||
(r"\xff$", b"abc\xff"),
|
||||
(r"a", b"A", "binary", ("IGNORECASE",)),
|
||||
(r"[a-z]+", b"ABC\x80abc", "binary", ("IGNORECASE",)),
|
||||
]
|
||||
for _row in _H_CASES:
|
||||
_pat, _subj = _row[0], _row[1]
|
||||
_mode = _row[2] if len(_row) > 2 else "binary"
|
||||
_fl = _row[3] if len(_row) > 3 else ()
|
||||
add("search", _pat, _subj, mode=_mode, flags=_fl)
|
||||
add("finditer", _pat, _subj, mode=_mode, flags=_fl)
|
||||
|
||||
# Raw-byte sub/split: replacement templates and split still need to work
|
||||
# correctly with embedded NUL and high bytes on both sides.
|
||||
_H_SUB_CASES = [
|
||||
(r"\x00", b"a\x00b\x00c", "-"),
|
||||
(r"[\x80-\xff]", b"a\x80b\x90c", "?"),
|
||||
(r"(.)\x00", b"a\x00b\x00", r"[\1]"),
|
||||
]
|
||||
for _pat, _subj, _repl in _H_SUB_CASES:
|
||||
add("sub", _pat, _subj, mode="binary", repl=_repl)
|
||||
|
||||
_H_SPLIT_CASES = [
|
||||
(r"\x00+", b"a\x00\x00b\x00c"),
|
||||
(r"[\x80-\xff]", b"a\x80b\x90c\xa0d"),
|
||||
]
|
||||
for _pat, _subj in _H_SPLIT_CASES:
|
||||
add("split", _pat, _subj, mode="binary")
|
||||
|
||||
# =============================================================================
|
||||
# Category I: UTF-8 edge cases -- real multi-byte content beyond simple
|
||||
# accented Latin: 4-byte (astral plane) code points, combining marks,
|
||||
# right-to-left scripts, CJK, mixed-width strings, and offset correctness
|
||||
# for backreferences/groups/lookaround spanning multi-byte characters.
|
||||
# =============================================================================
|
||||
_I_CASES = [
|
||||
(r"\w+", "hello \U0001F600 world"), # emoji, a 4-byte code point
|
||||
(r".", "\U0001F600"),
|
||||
(r"\W", "\U0001F600"),
|
||||
(r"(.)", "\U0001F600\U0001F601"),
|
||||
(r"é", "é"), # 'e' + combining acute accent
|
||||
(r"\w+", "éclair"),
|
||||
(r"[-ۿ]+", "السلام"), # Arabic
|
||||
(r"\w+", "שלום"), # Hebrew
|
||||
(r"[一-鿿]+", "中文字符"), # CJK unified ideographs
|
||||
(r"\w+", "日本語"), # Japanese
|
||||
(r"(\w)(\w)\2\1", "文字字文"),
|
||||
(r"(.)\1+", "\U0001F600\U0001F600\U0001F600"),
|
||||
(r"^.{3}$", "aé\U0001F600"), # mixed 1/2/4-byte codepoints, fixed count
|
||||
(r"\bworld\b", "café world café"),
|
||||
(r"(?<=é)\w+", "cafémonde"),
|
||||
(r"(?=\U0001F600)", "x\U0001F600y"),
|
||||
(r"[^\x00-\x7f]+", "abcéèêdef"),
|
||||
# Fullwidth digits (U+FF10-FF19, Unicode category Nd) are deliberately
|
||||
# not probed here: CPython's \d matches them (it matches the Nd
|
||||
# category), but glibc's iswdigit() under the C.utf8 locale this
|
||||
# build relies on does not, the same class of Unicode-table gap as
|
||||
# the NBSP case in README.md "Known deviations", which now also
|
||||
# records this specific instance.
|
||||
(r"café", "CAFÉ", "utf8", ("IGNORECASE",)),
|
||||
(r"ß", "ß", "utf8", ("IGNORECASE",)), # German sharp s
|
||||
]
|
||||
for _row in _I_CASES:
|
||||
_pat, _subj = _row[0], _row[1]
|
||||
_mode = _row[2] if len(_row) > 2 else "utf8"
|
||||
_fl = _row[3] if len(_row) > 3 else ()
|
||||
add("search", _pat, _subj, mode=_mode, flags=_fl)
|
||||
add("fullmatch", _pat, _subj, mode=_mode, flags=_fl)
|
||||
add("finditer", _pat, _subj, mode=_mode, flags=_fl)
|
||||
|
||||
_I_SUB_CASES = [
|
||||
(r"\U0001F600", "hi \U0001F600 there", ":)"),
|
||||
(r"(\w)", "café", r"[\1]"),
|
||||
(r"[一-鿿]", "中文abc", "?"),
|
||||
]
|
||||
for _pat, _subj, _repl in _I_SUB_CASES:
|
||||
add("sub", _pat, _subj, mode="utf8", repl=_repl)
|
||||
|
||||
_I_SPLIT_CASES = [
|
||||
(r"\s+", "café au lait 中文"),
|
||||
(r"(\U0001F600)", "a\U0001F600b\U0001F600c"),
|
||||
]
|
||||
for _pat, _subj in _I_SPLIT_CASES:
|
||||
add("split", _pat, _subj, mode="utf8")
|
||||
|
||||
# =============================================================================
|
||||
# Category J: seeded random combinatorics. A fixed seed makes this a
|
||||
# deterministic, reproducible regression test, not a flaky fuzzer: the same
|
||||
# cases are generated every run, so a failure here is exactly as
|
||||
# reproducible and reportable as a hand-written one. Explores combinations
|
||||
# no one sat down and thought to write by hand, across all three modes.
|
||||
# =============================================================================
|
||||
import random as _random # noqa: E402
|
||||
|
||||
_rng = _random.Random(0xC0FFEE)
|
||||
|
||||
_J_ASCII_ATOMS = ["a", "b", "c", "1", "2", "_", " ", "[a-c]", "[0-9]", r"\d", r"\w", r"\s", "."]
|
||||
_J_UTF8_ATOMS = _J_ASCII_ATOMS + ["é", "è", "中", "ا", "\U0001F600"]
|
||||
_J_QUANTS = ["", "?", "*", "+", "{1,2}", "{0,2}", "*?", "+?", "??"]
|
||||
_J_WRAPS = ["{0}", "({0})", "(?:{0})", "(?P<g%d>{0})"]
|
||||
|
||||
_J_ASCII_SUBJ_CHARS = list("abc123_ ")
|
||||
_J_UTF8_SUBJ_CHARS = _J_ASCII_SUBJ_CHARS + ["é", "è", "中", "文", "ا", "\U0001F600", "\U0001F601"]
|
||||
|
||||
|
||||
_J_SPECIAL_ATOMS = {r"\d", r"\w", r"\s", ".", "[a-c]", "[0-9]"}
|
||||
|
||||
|
||||
def _j_random_fragment(mode, group_no):
|
||||
atoms = _J_UTF8_ATOMS if mode == "utf8" else _J_ASCII_ATOMS
|
||||
atom = _rng.choice(atoms)
|
||||
if atom not in _J_SPECIAL_ATOMS:
|
||||
atom = re.escape(atom)
|
||||
q = _rng.choice(_J_QUANTS)
|
||||
wrap = _rng.choice(_J_WRAPS)
|
||||
if "%d" in wrap:
|
||||
wrap = wrap % group_no
|
||||
return wrap.format(atom + q)
|
||||
|
||||
|
||||
def _j_random_pattern(mode):
|
||||
n = _rng.randint(1, 4)
|
||||
group_no = 1
|
||||
parts = []
|
||||
for _ in range(n):
|
||||
frag = _j_random_fragment(mode, group_no)
|
||||
if frag.startswith("(") and not frag.startswith("(?:"):
|
||||
group_no += 1
|
||||
parts.append(frag)
|
||||
joiner = "|" if _rng.random() < 0.25 else ""
|
||||
return joiner.join(parts)
|
||||
|
||||
|
||||
def _j_random_subject(mode, length):
|
||||
chars = _J_UTF8_SUBJ_CHARS if mode == "utf8" else _J_ASCII_SUBJ_CHARS
|
||||
return "".join(_rng.choice(chars) for _ in range(length))
|
||||
|
||||
|
||||
_J_FLAG_CHOICES = [[], ["IGNORECASE"], ["MULTILINE"], ["DOTALL"], ["IGNORECASE", "MULTILINE"]]
|
||||
_J_MODES = ["ascii", "utf8", "binary"]
|
||||
_J_OPS = ["search", "finditer", "fullmatch"]
|
||||
|
||||
# Patterns/subjects for "binary" mode are still built from the ASCII-safe
|
||||
# atom pool (Category H already covers genuinely arbitrary, non-UTF-8
|
||||
# binary content deliberately and by hand); here binary mode exercises
|
||||
# byte-for-byte matching of ordinary ASCII text through the BINARY
|
||||
# encoding path specifically, distinct from the same text through ASCII
|
||||
# or UTF8 mode, which Category F already probes with a fixed pattern set.
|
||||
_j_generated = 0
|
||||
for _ in range(900):
|
||||
_mode = _rng.choice(_J_MODES)
|
||||
_atom_mode = "utf8" if _mode == "utf8" else "ascii"
|
||||
_pat = _j_random_pattern(_atom_mode)
|
||||
_subj = _j_random_subject(_atom_mode, _rng.randint(0, 10))
|
||||
_fl = _rng.choice(_J_FLAG_CHOICES)
|
||||
_op = _rng.choice(_J_OPS)
|
||||
add(_op, _pat, _subj, mode=_mode, flags=_fl)
|
||||
_j_generated += 1
|
||||
|
||||
+12
-5
@@ -48,9 +48,14 @@ def c_str(b):
|
||||
|
||||
|
||||
def py_flags(case):
|
||||
# re.ASCII is not added here for mode == "ascii": ground truth for
|
||||
# that mode is already a `bytes` pattern against a `bytes` subject
|
||||
# (encode(), above), and a bytes pattern is inherently ASCII-only
|
||||
# for \w/\s/\d regardless of re.ASCII, so adding it would be purely
|
||||
# redundant, and actively wrong whenever a case also uses LOCALE,
|
||||
# since Python rejects ASCII and LOCALE combined (confirmed
|
||||
# directly: ValueError: ASCII and LOCALE flags are incompatible).
|
||||
f = 0
|
||||
if case["mode"] == "ascii":
|
||||
f |= re.ASCII
|
||||
for name in case["flags"]:
|
||||
f |= FLAGMAP[name]
|
||||
return f
|
||||
@@ -83,6 +88,8 @@ def encode(case, s):
|
||||
# latter, correctly). A `bytes` pattern against a `bytes` subject is
|
||||
# always byte oriented with ASCII-only \w/\s/\d in Python too, so it
|
||||
# is the correct ground truth for both ASCII and BINARY mode here.
|
||||
if isinstance(s, bytes):
|
||||
return s # already raw bytes (arbitrary, not necessarily valid UTF-8): pass through
|
||||
if case["mode"] in ("binary", "ascii"):
|
||||
return s.encode("utf-8")
|
||||
return s
|
||||
@@ -186,6 +193,8 @@ def main():
|
||||
skipped = 0
|
||||
for i, case in enumerate(CASES):
|
||||
op = case["op"]
|
||||
if op not in ("search", "match", "fullmatch", "finditer", "sub", "split"):
|
||||
raise ValueError("unknown op %r" % op) # a typo in a hand-written case, not a skippable combination
|
||||
try:
|
||||
if op in ("search", "match", "fullmatch"):
|
||||
gen_search_family(case, i, out)
|
||||
@@ -195,10 +204,8 @@ def main():
|
||||
gen_sub(case, i, out)
|
||||
elif op == "split":
|
||||
gen_split(case, i, out)
|
||||
else:
|
||||
raise ValueError("unknown op %r" % op)
|
||||
emitted += 1
|
||||
except re.error as e:
|
||||
except (re.error, ValueError) as e:
|
||||
# A combinatorially generated pattern that CPython itself
|
||||
# cannot compile is not a useful ground-truth case (there is
|
||||
# nothing to check regexx against); skip it rather than
|
||||
|
||||
Reference in New Issue
Block a user