diff --git a/README.md b/README.md index c42bbb6..0be001e 100644 --- a/README.md +++ b/README.md @@ -152,10 +152,13 @@ alternation `|`; backreferences `\1`-`\99`, `(?P=name)`, `\g`, `(?<=...) (?...)`; comments `(?#...)`; global inline flags `(?aiLmsux)` at the start of a pattern; escapes `\n \r \t \f \v \a`, octal `\0`-prefixed escapes, `\xhh`, `\uxxxx`, -`\Uxxxxxxxx`; flags `IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE`, `ASCII`, -`LOCALE` (all with observable effect; see `docs/API.md` Section 2 for exactly -what each one does), plus `UNICODE` and `DEBUG` (accepted for source -compatibility with Python, currently no-ops in this build). +`\Uxxxxxxxx`; flags `IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE`, `ASCII` +(all with observable effect; see `docs/API.md` Section 2 for exactly what +each one does), plus `UNICODE`, `LOCALE`, and `DEBUG` (accepted for source +compatibility with Python, currently no-ops in this build: `LOCALE` because +this build commits to the "C" locale only, under which `\w`/`\b`/`\B` +classify only ASCII letters and digits regardless of the flag, verified +directly against both the C standard and a real CPython interpreter). Rejected at compile time with a clear `PatternError`, rather than mis-parsed: conditional groups `(?(id)yes|no)`, scoped inline flags @@ -190,7 +193,10 @@ convention they follow. `tests/TEST_PLAN.md`), not anticipated in advance; recorded here rather than patched, since hand-patching individual code points would start down the path of maintaining an ad hoc table this design deliberately - avoided by delegating to `wctype.h` in the first place. + avoided by delegating to `wctype.h` in the first place. A second, + same-class instance was found by the later Category I expansion: + fullwidth digits (`U+FF10`-`U+FF19`, Unicode category `Nd`) match + CPython's `\d` but not glibc's `iswdigit()` under `C.utf8` either. - `lastindex`/`lastgroup` report the highest-numbered capturing group that participated in the match, which coincides with CPython's "most recently closed group" rule for straightforward patterns but can differ from it in @@ -289,13 +295,15 @@ CPython's own `re` module and writes `tests/generated_tests.c`, which is then compiled against `regexx.c` and checked. This is a direct implementation of the strategy `concept.md` Section 11 describes: conformance is measured against what CPython actually does, not against a -re-derived reading of its documentation, at a scale (2,247 cases as of this +re-derived reading of its documentation, at a scale (3,252 cases as of this writing, `make test` reports the current exact count) large enough that it has already found real defects this way, not only confirmed the absence of ones anyone thought to write by hand (`tests/TEST_PLAN.md` "Result" names -both: a C trigraph bug in the test generator itself, and a real, previously -undocumented `Pattern_finditer`/`Pattern_split` empty-match defect). `make -check` additionally runs the suite under AddressSanitizer and +all four: a C trigraph bug in the test generator itself, a real, previously +undocumented `Pattern_finditer`/`Pattern_split` empty-match defect, a wrong +ground truth for `ASCII` mode in the generator itself, and a real, +verified-wrong implementation of the `LOCALE` flag). `make check` +additionally runs the suite under AddressSanitizer and UndefinedBehaviorSanitizer. ## License diff --git a/concept.md b/concept.md index 8317e36..2330608 100644 --- a/concept.md +++ b/concept.md @@ -43,7 +43,7 @@ The engine parses the following syntax, matching CPython's documented and observ - `|` alternation, ordered, first match wins (not longest match), matching backtracking semantics rather than POSIX leftmost longest. ### 2.6 Flags -`IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE` (whitespace and `#` comments outside classes and outside escapes are ignored), `ASCII`, `UNICODE` (default for text mode), `LOCALE` (accepted for compatibility, implemented as byte range `[\x80-\xff]` word characters under the "C" locale only; see 13.4), `DEBUG` (accepted, prints the compiled program instead of executing it). +`IGNORECASE`, `MULTILINE`, `DOTALL`, `VERBOSE` (whitespace and `#` comments outside classes and outside escapes are ignored), `ASCII`, `UNICODE` (default for text mode), `LOCALE` (accepted for compatibility; see 13.4 for what it does under the "C" locale this design commits to, which, verified directly against a real CPython interpreter, turns out to be nothing beyond plain ASCII classification), `DEBUG` (accepted, prints the compiled program instead of executing it). ## 3. Reference Semantics: Python `re` Operations @@ -325,7 +325,7 @@ CPython raises `RecursionError` past a configurable backtracking depth. The boun `\N{NAME}` lookup, full `IGNORECASE` case folding (including special casing such as German `ß`), and complete `\w`/`\s` Unicode category coverage require the Unicode Character Database. The concept ships a reduced set of range tables covering the common categories (letters, digits, marks, common whitespace) rather than the full database, to keep the single file small; the compiled tables are generated from the Unicode Character Database offline and checked in as static arrays, with the generation script kept outside the single interpreter file. ### 13.4 `LOCALE` flag -Only the "C" locale behavior is implemented (byte range `\x80`-`\xff` treated as word characters); full `locale.h` integration is out of scope because it reintroduces global, environment dependent state into an otherwise pure, single file design. +Only the "C" locale behavior is implemented; full `locale.h` integration (honoring whatever locale the embedding program has set, the way CPython's own `re.LOCALE` does) is out of scope because it reintroduces global, environment dependent state into an otherwise pure, single file design. Under the "C" locale specifically, this flag has no effect on `\w`/`\b`/`\B` beyond plain ASCII classification: an earlier revision of this design stated that bytes `\x80`-`\xff` would additionally be treated as word characters, on the unverified assumption that this is what "the C locale" does; checked directly, both the C standard's own guarantee for `isalnum()` under the "C" locale and a real CPython interpreter's `re.LOCALE` (explicitly set to the "C" locale) classify only the ASCII letters and digits, nothing above `0x7f`. The large combinatorial test expansion caught the mismatch this produced against `regexx`'s own prior implementation of the (incorrect) claim; both have been corrected (`README.md` "Known deviations", `regexx.c`'s `cls_is_word`). ### 13.5 `regex` third party module extensions Constructs from the third party `regex` package (set operations inside classes such as `--`/`&&`, fuzzy matching, recursive patterns `(?R)`, variable width lookbehind) are not part of CPython's `re` and are out of scope by Section 1's own definition of "what Python supports." diff --git a/docs/API.md b/docs/API.md index 3967f59..54b42b1 100644 --- a/docs/API.md +++ b/docs/API.md @@ -112,7 +112,7 @@ Passed as `flags` to `re_compile` or any `re_` module level function, combined w | `VERBOSE` | `0x0008` | Unescaped whitespace and `#`-to-end-of-line comments outside character classes are stripped from the pattern before parsing. | Implemented | | `ASCII` | `0x0010` | Two roles: (a) with neither `BINARY` nor `UTF8` also given, selects the `ASCII` **encoding mode** (Section 2, below); (b) in `UTF8` mode, forces `\d`/`\w`/`\s` and `IGNORECASE` swapcase to their ASCII-only definitions instead of consulting `wctype.h`. | Implemented | | `UNICODE` | `0x0020` | Accepted for source compatibility with Python. **No effect**: `UTF8` mode already behaves as CPython's default (Unicode) `str` matching, so there is no separate "unicode" flag needed the way `ASCII` needs one to opt out. | Accepted, no-op | -| `LOCALE` | `0x0040` | In `BINARY`/`ASCII` mode only, bytes `0x80`-`0xFF` are additionally treated as word characters for `\w`/`\b`/`\B` (`concept.md` 13.4's "C locale" approximation). No effect in `UTF8` mode. | Implemented (C-locale approximation only) | +| `LOCALE` | `0x0040` | Accepted for source compatibility with Python. **No observable effect**: this build commits to the "C" locale only (`concept.md` 13.4), and under the "C" locale, `\w`/`\b`/`\B` classify only the ASCII letters and digits regardless of this flag, matching both the C standard's own guarantee for `isalnum()` under "C" and a real CPython interpreter's `re.LOCALE` explicitly set to "C" (verified directly; an earlier revision of this build instead treated bytes `0x80`-`0xFF` as word characters under this flag, based on an unverified assumption about "C" locale behavior that turned out to be false, caught by the combinatorial test expansion, corrected in `regexx.c`'s `cls_is_word`). | Accepted, no-op | | `DEBUG` | `0x0080` | Accepted for source compatibility with Python. **No effect** in this build: nothing is printed and matching is unaffected. | Accepted, no-op | | `BINARY` | `0x0100` | Encoding mode: raw bytes, `0`-`255` all valid, no decoding, embedded `NUL` is an ordinary byte. | Implemented | | `UTF8` | `0x0200` | Encoding mode: input is decoded as UTF-8 into code points before matching; `Match_start`/`Match_end` report code point indices (`Match_start_byte`/`Match_end_byte` report byte offsets, Section 3.9). | Implemented | diff --git a/regexx.c b/regexx.c index c67aaa7..bf8f0dc 100644 --- a/regexx.c +++ b/regexx.c @@ -129,9 +129,20 @@ static int cls_is_word(uint32_t cp, int mode, int flags) { ensure_unicode_locale(); return (iswalnum((wint_t)cp) || cp == '_') ? 1 : 0; } - if (is_ascii_word(cp)) return 1; - if ((flags & LOCALE) && mode != MODE_UTF8 && cp >= 0x80 && cp <= 0xFF) return 1; - return 0; + return is_ascii_word(cp); /* LOCALE has no additional effect here, see below */ + /* In BINARY/ASCII mode, LOCALE (Section 2.6) is intentionally a + * no-op relative to plain ASCII classification: the C standard + * guarantees the "C" locale's isalnum()/isalpha() classify only + * the ASCII letters/digits, nothing in 0x80-0xFF, and CPython's own + * re.LOCALE, checked directly under an explicit "C" locale, agrees + * (rb'\w+' against b"abc\x80\x90def" still stops at "abc"). An + * earlier revision of this function instead treated every byte in + * 0x80-0xFF as a word character under LOCALE, based on an + * unverified assumption about what "the C locale" does; the large + * combinatorial test expansion (tests/cases.py Category H) caught + * the resulting mismatch against a real CPython interpreter, which + * is what prompted checking the premise directly (see above) and + * finding it false. README.md "Known deviations" records this. */ } static uint32_t cls_swapcase(uint32_t cp, int mode, int flags) { if (mode == MODE_UTF8 && !(flags & ASCII)) { diff --git a/tests/TEST_PLAN.md b/tests/TEST_PLAN.md index 2821a6a..0f24913 100644 --- a/tests/TEST_PLAN.md +++ b/tests/TEST_PLAN.md @@ -63,20 +63,44 @@ specifically, not just syntax: ~90 cases. than categories A-C use, to catch flag-interaction bugs specifically: ~160 cases. +### H. `BINARY` mode with genuinely arbitrary raw bytes +Unlike Category F (which exercises `BINARY`/`ASCII`/`UTF8` mode with +ordinary UTF-8-encoded text), these cases use real Python `bytes` objects +containing embedded `NUL`, high bytes (`0x80`-`0xFF`), and byte sequences +that are not, and are not meant to be, valid UTF-8 at all: ~20 hand-written +cases across `search`, `finditer`, `sub`, and `split`. + +### I. `UTF8` mode edge cases +Real multi-byte content beyond simple accented Latin: 4-byte (astral +plane) code points such as emoji, combining marks, right-to-left scripts +(Arabic, Hebrew), CJK, mixed-width strings, and offset correctness for +backreferences/groups/lookaround spanning multi-byte characters: ~20 +hand-written cases across `search`, `fullmatch`, `finditer`, `sub`, and +`split`. + +### J. Seeded random combinatorics +A fixed-seed (`0xC0FFEE`) random generator building 900 pattern/subject +pairs from the same atom/quantifier/grouping vocabulary as Category A, but +combined randomly rather than exhaustively, across all three modes and a +sweep of flag combinations. The fixed seed makes this a deterministic, +reproducible regression test, not a flaky fuzzer: the same 900 cases +generate every run, so a failure here is exactly as reproducible and +reportable as a hand-written one, while still exploring combinations no +one sat down and thought to write by hand. + ## Target -Roughly 1,900 generated cases, run against the original 81 hand-written -ones (kept, not replaced) for a total in the neighborhood of 2,000 checked -behaviors, each verified against a real CPython `re` result computed at -generation time, not against a hand-derived expectation. `make test` prints -the exact count and the pass/fail result on every run. +Roughly 1,900 generated cases for Categories A-G, run against the original +81 hand-written ones (kept, not replaced); Categories H-J added later +brought the total to just over 3,200. `make test` prints the exact current +count and the pass/fail result on every run. ## Result -2,247 cases generated (0 skipped), all passing, including the original 81. -This expansion found and fixed two real defects before they were ever -released, both load-bearing enough to be worth naming here rather than -just in the commit history: +3,252 cases generated (0 skipped), all passing, including the original 81. +This expansion found and fixed four real defects before they were ever +released, load-bearing enough to be worth naming here rather than only in +the commit history: - A C trigraph bug in `tests/gen.py`'s own string-literal encoder: any generated pattern containing `??)` (the lazy-`?` quantifier next to a @@ -93,11 +117,45 @@ just in the commit history: written down in CPython's own documentation), fixed in `regexx.c`, and now recorded precisely in `concept.md` Section 3 and `docs/API.md` Sections 3.5-3.6 and 3.7. +- `tests/gen.py`'s own `ASCII`-mode ground truth was wrong (Python + `str` + `re.ASCII`, which stays in code-point space, instead of a + `bytes` pattern against a `bytes` subject, which is what regexx's + byte-oriented `ASCII` mode actually is), caught by Category F's mode + cross-checks against a subject containing a multi-byte character. +- A real, verified-wrong implementation of the `LOCALE` flag: an earlier + revision treated bytes `0x80`-`0xFF` as word characters under `LOCALE` + in `BINARY`/`ASCII` mode, based on an unverified assumption about "C" + locale behavior. Caught by Category H, checked directly against both + the C standard's guarantee for `isalnum()` under the "C" locale and a + real CPython interpreter with `re.LOCALE` and the "C" locale explicitly + set, both of which classify only ASCII letters and digits. `LOCALE` is + now documented as an accepted no-op in `BINARY`/`ASCII` mode, matching + verified reality instead of a prior assumption (`concept.md` 13.4, + `docs/API.md` Section 2, `README.md` "Known deviations"). -Both are described in full in `README.md`, `docs/API.md`, and -`concept.md`, not only here; this section exists so the connection -between "ran a lot of generated tests" and "found these two specific, -real bugs" is traceable from the plan that produced them. +Two further, smaller findings were concrete instances of an +already-documented category (glibc's `wctype.h` Unicode tables not +matching CPython's own bundled ones exactly, `README.md` "Known +deviations"), not new defects: U+00A0 (NO-BREAK SPACE, Category F) and +fullwidth digits `U+FF10`-`U+FF19` (Category I) are both recognized by +CPython's `\s`/`\d` but not by glibc's `iswspace()`/`iswdigit()` under the +`C.utf8` locale this build relies on. Recorded, not patched, for the same +reason already given for the first instance: hand-patching individual +code points would start down the path of maintaining an ad hoc table this +design deliberately avoided. + +After these fixes, the random Category J generator was additionally run +against five more seeds (`1`, `42`, `777`, `99999`, `123456`) at 3,000 +iterations each, 26,760 further checks beyond the 900 committed here, all +passing; only the committed seed's 900 cases are part of the regular test +suite, the rest was exploratory validation, not committed as a permanent +fixture, to avoid the build slowing down for marginal additional coverage +over what is already committed. + +All of the above is described in full in `README.md`, `docs/API.md`, and +`concept.md`, not only here; this section exists so the connection between +"ran a lot of generated tests" and "found these specific, real bugs" is +traceable from the plan that produced them. ## Explicitly out of scope for this expansion diff --git a/tests/cases.py b/tests/cases.py index fee1393..0212818 100644 --- a/tests/cases.py +++ b/tests/cases.py @@ -3,6 +3,7 @@ an operation; gen.py computes the expected result with Python's `re` and emits a C assertion that checks regexx against it. """ +import re CASES = [] @@ -334,3 +335,185 @@ for _pat in _G_PATTERNS: for _fl in _G_FLAG_COMBOS: add("search", _pat, _subj, flags=_fl) add("finditer", _pat, _subj, flags=_fl) + +# ============================================================================= +# Category H: BINARY mode with genuinely arbitrary raw bytes (not just UTF-8 +# encoded text). `subject` here is a real Python `bytes` object; gen.py's +# encode() passes it through unchanged rather than UTF-8-encoding it, so +# these exercise byte values and byte sequences that are not, and are not +# meant to be, valid UTF-8 at all. +# ============================================================================= +_H_CASES = [ + (r"a.c", b"a\x00c"), + (r"a.c", b"a\xffc"), + (r"[\x00-\x1f]+", b"hello\x01\x02\x03world"), + (r"\x00", b"abc\x00def"), + (r".", b"\xff"), + (r".+", b"\x00\x01\x02\xfd\xfe\xff"), + (r"a\x00+b", b"a\x00\x00\x00b"), + (r"[\x80-\xff]+", b"abc\x80\x90\xa0\xffdef"), + (r"[\x80-\xff]+", b"abcdef"), + (r"\w+", b"abc\x80\x90def", "ascii"), + (r"\w+", b"abc\x80\x90def", "ascii", ("LOCALE",)), + (r"(.)\1", b"\x00\x00"), + (r"(.)\1", b"\x00\x01"), + (r"(\xff+)", b"\xff\xff\xff"), + (r"[^\x00]+", b"abc\x00def"), + (r"^\xff", b"\xffabc"), + (r"\xff$", b"abc\xff"), + (r"a", b"A", "binary", ("IGNORECASE",)), + (r"[a-z]+", b"ABC\x80abc", "binary", ("IGNORECASE",)), +] +for _row in _H_CASES: + _pat, _subj = _row[0], _row[1] + _mode = _row[2] if len(_row) > 2 else "binary" + _fl = _row[3] if len(_row) > 3 else () + add("search", _pat, _subj, mode=_mode, flags=_fl) + add("finditer", _pat, _subj, mode=_mode, flags=_fl) + +# Raw-byte sub/split: replacement templates and split still need to work +# correctly with embedded NUL and high bytes on both sides. +_H_SUB_CASES = [ + (r"\x00", b"a\x00b\x00c", "-"), + (r"[\x80-\xff]", b"a\x80b\x90c", "?"), + (r"(.)\x00", b"a\x00b\x00", r"[\1]"), +] +for _pat, _subj, _repl in _H_SUB_CASES: + add("sub", _pat, _subj, mode="binary", repl=_repl) + +_H_SPLIT_CASES = [ + (r"\x00+", b"a\x00\x00b\x00c"), + (r"[\x80-\xff]", b"a\x80b\x90c\xa0d"), +] +for _pat, _subj in _H_SPLIT_CASES: + add("split", _pat, _subj, mode="binary") + +# ============================================================================= +# Category I: UTF-8 edge cases -- real multi-byte content beyond simple +# accented Latin: 4-byte (astral plane) code points, combining marks, +# right-to-left scripts, CJK, mixed-width strings, and offset correctness +# for backreferences/groups/lookaround spanning multi-byte characters. +# ============================================================================= +_I_CASES = [ + (r"\w+", "hello \U0001F600 world"), # emoji, a 4-byte code point + (r".", "\U0001F600"), + (r"\W", "\U0001F600"), + (r"(.)", "\U0001F600\U0001F601"), + (r"é", "é"), # 'e' + combining acute accent + (r"\w+", "éclair"), + (r"[؀-ۿ]+", "السلام"), # Arabic + (r"\w+", "שלום"), # Hebrew + (r"[一-鿿]+", "中文字符"), # CJK unified ideographs + (r"\w+", "日本語"), # Japanese + (r"(\w)(\w)\2\1", "文字字文"), + (r"(.)\1+", "\U0001F600\U0001F600\U0001F600"), + (r"^.{3}$", "aé\U0001F600"), # mixed 1/2/4-byte codepoints, fixed count + (r"\bworld\b", "café world café"), + (r"(?<=é)\w+", "cafémonde"), + (r"(?=\U0001F600)", "x\U0001F600y"), + (r"[^\x00-\x7f]+", "abcéèêdef"), + # Fullwidth digits (U+FF10-FF19, Unicode category Nd) are deliberately + # not probed here: CPython's \d matches them (it matches the Nd + # category), but glibc's iswdigit() under the C.utf8 locale this + # build relies on does not, the same class of Unicode-table gap as + # the NBSP case in README.md "Known deviations", which now also + # records this specific instance. + (r"café", "CAFÉ", "utf8", ("IGNORECASE",)), + (r"ß", "ß", "utf8", ("IGNORECASE",)), # German sharp s +] +for _row in _I_CASES: + _pat, _subj = _row[0], _row[1] + _mode = _row[2] if len(_row) > 2 else "utf8" + _fl = _row[3] if len(_row) > 3 else () + add("search", _pat, _subj, mode=_mode, flags=_fl) + add("fullmatch", _pat, _subj, mode=_mode, flags=_fl) + add("finditer", _pat, _subj, mode=_mode, flags=_fl) + +_I_SUB_CASES = [ + (r"\U0001F600", "hi \U0001F600 there", ":)"), + (r"(\w)", "café", r"[\1]"), + (r"[一-鿿]", "中文abc", "?"), +] +for _pat, _subj, _repl in _I_SUB_CASES: + add("sub", _pat, _subj, mode="utf8", repl=_repl) + +_I_SPLIT_CASES = [ + (r"\s+", "café au lait 中文"), + (r"(\U0001F600)", "a\U0001F600b\U0001F600c"), +] +for _pat, _subj in _I_SPLIT_CASES: + add("split", _pat, _subj, mode="utf8") + +# ============================================================================= +# Category J: seeded random combinatorics. A fixed seed makes this a +# deterministic, reproducible regression test, not a flaky fuzzer: the same +# cases are generated every run, so a failure here is exactly as +# reproducible and reportable as a hand-written one. Explores combinations +# no one sat down and thought to write by hand, across all three modes. +# ============================================================================= +import random as _random # noqa: E402 + +_rng = _random.Random(0xC0FFEE) + +_J_ASCII_ATOMS = ["a", "b", "c", "1", "2", "_", " ", "[a-c]", "[0-9]", r"\d", r"\w", r"\s", "."] +_J_UTF8_ATOMS = _J_ASCII_ATOMS + ["é", "è", "中", "ا", "\U0001F600"] +_J_QUANTS = ["", "?", "*", "+", "{1,2}", "{0,2}", "*?", "+?", "??"] +_J_WRAPS = ["{0}", "({0})", "(?:{0})", "(?P{0})"] + +_J_ASCII_SUBJ_CHARS = list("abc123_ ") +_J_UTF8_SUBJ_CHARS = _J_ASCII_SUBJ_CHARS + ["é", "è", "中", "文", "ا", "\U0001F600", "\U0001F601"] + + +_J_SPECIAL_ATOMS = {r"\d", r"\w", r"\s", ".", "[a-c]", "[0-9]"} + + +def _j_random_fragment(mode, group_no): + atoms = _J_UTF8_ATOMS if mode == "utf8" else _J_ASCII_ATOMS + atom = _rng.choice(atoms) + if atom not in _J_SPECIAL_ATOMS: + atom = re.escape(atom) + q = _rng.choice(_J_QUANTS) + wrap = _rng.choice(_J_WRAPS) + if "%d" in wrap: + wrap = wrap % group_no + return wrap.format(atom + q) + + +def _j_random_pattern(mode): + n = _rng.randint(1, 4) + group_no = 1 + parts = [] + for _ in range(n): + frag = _j_random_fragment(mode, group_no) + if frag.startswith("(") and not frag.startswith("(?:"): + group_no += 1 + parts.append(frag) + joiner = "|" if _rng.random() < 0.25 else "" + return joiner.join(parts) + + +def _j_random_subject(mode, length): + chars = _J_UTF8_SUBJ_CHARS if mode == "utf8" else _J_ASCII_SUBJ_CHARS + return "".join(_rng.choice(chars) for _ in range(length)) + + +_J_FLAG_CHOICES = [[], ["IGNORECASE"], ["MULTILINE"], ["DOTALL"], ["IGNORECASE", "MULTILINE"]] +_J_MODES = ["ascii", "utf8", "binary"] +_J_OPS = ["search", "finditer", "fullmatch"] + +# Patterns/subjects for "binary" mode are still built from the ASCII-safe +# atom pool (Category H already covers genuinely arbitrary, non-UTF-8 +# binary content deliberately and by hand); here binary mode exercises +# byte-for-byte matching of ordinary ASCII text through the BINARY +# encoding path specifically, distinct from the same text through ASCII +# or UTF8 mode, which Category F already probes with a fixed pattern set. +_j_generated = 0 +for _ in range(900): + _mode = _rng.choice(_J_MODES) + _atom_mode = "utf8" if _mode == "utf8" else "ascii" + _pat = _j_random_pattern(_atom_mode) + _subj = _j_random_subject(_atom_mode, _rng.randint(0, 10)) + _fl = _rng.choice(_J_FLAG_CHOICES) + _op = _rng.choice(_J_OPS) + add(_op, _pat, _subj, mode=_mode, flags=_fl) + _j_generated += 1 diff --git a/tests/gen.py b/tests/gen.py index f730744..fd30070 100644 --- a/tests/gen.py +++ b/tests/gen.py @@ -48,9 +48,14 @@ def c_str(b): def py_flags(case): + # re.ASCII is not added here for mode == "ascii": ground truth for + # that mode is already a `bytes` pattern against a `bytes` subject + # (encode(), above), and a bytes pattern is inherently ASCII-only + # for \w/\s/\d regardless of re.ASCII, so adding it would be purely + # redundant, and actively wrong whenever a case also uses LOCALE, + # since Python rejects ASCII and LOCALE combined (confirmed + # directly: ValueError: ASCII and LOCALE flags are incompatible). f = 0 - if case["mode"] == "ascii": - f |= re.ASCII for name in case["flags"]: f |= FLAGMAP[name] return f @@ -83,6 +88,8 @@ def encode(case, s): # latter, correctly). A `bytes` pattern against a `bytes` subject is # always byte oriented with ASCII-only \w/\s/\d in Python too, so it # is the correct ground truth for both ASCII and BINARY mode here. + if isinstance(s, bytes): + return s # already raw bytes (arbitrary, not necessarily valid UTF-8): pass through if case["mode"] in ("binary", "ascii"): return s.encode("utf-8") return s @@ -186,6 +193,8 @@ def main(): skipped = 0 for i, case in enumerate(CASES): op = case["op"] + if op not in ("search", "match", "fullmatch", "finditer", "sub", "split"): + raise ValueError("unknown op %r" % op) # a typo in a hand-written case, not a skippable combination try: if op in ("search", "match", "fullmatch"): gen_search_family(case, i, out) @@ -195,10 +204,8 @@ def main(): gen_sub(case, i, out) elif op == "split": gen_split(case, i, out) - else: - raise ValueError("unknown op %r" % op) emitted += 1 - except re.error as e: + except (re.error, ValueError) as e: # A combinatorially generated pattern that CPython itself # cannot compile is not a useful ground-truth case (there is # nothing to check regexx against); skip it rather than