Commit Graph
3 Commits
Author SHA1 Message Date
retoorandClaude Sonnet 5 b0991814cb Expand BINARY/UTF8/ASCII edge-case testing to 3,252 cases, fix a real LOCALE bug
Three new categories added to tests/cases.py (TEST_PLAN.md records the
detail): Category H exercises BINARY mode with genuinely arbitrary raw
bytes (embedded NUL, high bytes, non-UTF-8 sequences) via real Python
`bytes` subjects, not just UTF-8-encoded text; Category I exercises
UTF8 mode edge cases (4-byte/astral code points, combining marks,
Arabic, Hebrew, CJK, offset correctness across multi-byte characters);
Category J is a fixed-seed (reproducible, not flaky) random generator
combining the existing atom/quantifier/grouping vocabulary across all
three modes.

This found and fixed a real bug, not just a test-generation one:
LOCALE, in BINARY/ASCII mode, treated bytes 0x80-0xFF as word
characters, based on an unverified assumption about what the "C"
locale does. Checked directly against both the C standard's own
guarantee for isalnum() under "C" and a real CPython interpreter with
re.LOCALE and the "C" locale explicitly set, neither treats anything
above 0x7f as a word character. Fixed in regexx.c's cls_is_word;
LOCALE is now documented as an accepted no-op in non-UTF8 mode,
matching verified reality instead of a prior assumption (concept.md
13.4, docs/API.md, README.md "Known deviations").

Two more findings were test-generation bugs, not regexx bugs: gen.py's
own ASCII-mode ground truth used Python str + re.ASCII (code-point
space) instead of a bytes pattern against a bytes subject (what
regexx's byte-oriented ASCII mode actually is), and LOCALE combined
with the (now removed as redundant) auto-added re.ASCII flag raised
ValueError in Python for being an incompatible combination. Both
fixed in gen.py.

Two further findings were concrete instances of an already-documented
category (glibc's wctype.h Unicode tables not matching CPython's own
exactly): U+00A0 and fullwidth digits U+FF10-FF19 are recognized by
CPython's \s/\d but not by glibc's iswspace()/iswdigit() under C.utf8.
Recorded in README.md, not patched, for the reason already given for
the first such instance (NBSP) in the previous commit.

After these fixes: all 3,252 committed cases pass, clean under
AddressSanitizer/UndefinedBehaviorSanitizer. The Category J generator
was additionally run against 5 more seeds at 3,000 iterations each
(26,760 further checks) as exploratory validation, all passing; not
committed, to keep the regular suite's size proportionate.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
2026-09-14 09:25:50 +00:00
retoorandClaude Sonnet 5 51f8ed19e9 Expand the test suite to 2,247 combinatorial cases, fix two real bugs it found
tests/TEST_PLAN.md records the plan before the result: seven categories
(quantifier x grouping x flags, alternation x backreference x groups,
lookaround combinatorics, sub/split specifics, real-world combination
patterns, mode cross-checks, flag combination stress), each generated
programmatically in tests/cases.py rather than hand-typed, still every
case checked against a real CPython re result computed by tests/gen.py,
per concept.md Section 11's existing strategy, just at 2,247 cases
instead of 81.

Running it immediately found two real, previously latent bugs, not
just confirmed correctness of what was already covered by hand:

- tests/gen.py's own C string-literal encoder had a trigraph bug: any
  generated pattern containing `??)` (the lazy quantifier next to a
  closing paren) was silently rewritten by the C compiler from 7 bytes
  to 5 before the suite ever ran, confirmed directly by compiling and
  printing the corrupted string. Fixed by escaping '?' as '\?', which
  is always safe and makes trigraph formation impossible.
- Pattern_finditer/Pattern_split had a real, previously undocumented
  correctness defect: CPython's empty-match handling additionally
  searches for, and reports, a second, non-empty match at the same
  start position whenever the natural match found there was empty
  (reverse engineered against a real interpreter, since this is not
  written down in CPython's own documentation; \d*? against
  "123abc456" yields 16 matches, not 9). Fixed with a new MCtx
  forbid_empty flag that forces exactly that second search by
  rejecting the empty solution at OP_MATCH and letting ordinary
  backtracking find the next alternative, wired into both functions
  (they have independent scan loops). concept.md Section 3 and
  docs/API.md now state the rule precisely instead of the previous,
  incomplete description.

A third, more mundane finding: the combinatorial mode cross-check
category surfaced that ASCII-mode ground truth was being computed
wrong in tests/gen.py itself (Python str + re.ASCII, which stays in
code-point space, instead of a bytes pattern against a bytes subject,
which is what regexx's byte-oriented ASCII mode actually is), and
separately surfaced a genuine, now precisely documented Unicode-table
gap already anticipated in principle by README.md's "Known deviations"
(glibc's iswspace() under the C.utf8 locale does not classify U+00A0
NO-BREAK SPACE as whitespace; CPython's \s does).

All 2,247 cases pass, clean under AddressSanitizer/UndefinedBehavior-
Sanitizer; the 50MB/quadratic-time and ReDoS-scaling benchmarks from
the previous two commits are unaffected.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
2026-09-14 09:09:04 +00:00
retoorandClaude Sonnet 5 8f6afd6cd4 Add regexx: a single-file C regex interpreter with Python re semantics
concept.md is the full design document: the objective (Python re parity
plus binary/ASCII/UTF-8 modes, gigabyte-scale input, single C file), the
automata-theory argument for why unrestricted backreferences/lookaround
are incompatible with strict single-pass constant memory, the resulting
two-engine architecture, the exact Python-mirroring naming convention,
and a full comparison against POSIX regex.h for C-background readers.

regexx.c/regexx.h are the v1 implementation: parser, compiler to a
Pike/backtracking-style bytecode, and a single recursive backtracking
engine covering the pattern syntax and operations listed in README.md,
validated against CPython's own re module output (tests/), clean under
AddressSanitizer/UBSan, and stress-tested (50MB simple-quantifier match,
graceful failure rather than a crash on complex repeats over large
input, clean rejection of every intentionally unsupported construct).

Also included: examples/rxgrep.c (a small grep-like program exercising
all three data modes and the substitution API), the Makefile, the MIT
LICENSE, and docs/API.md, an exhaustive reference for every type, flag,
and function's exact return-value and memory-ownership convention,
checked against the current source and against a real CPython
interpreter rather than against memory.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
2026-09-14 05:56:16 +00:00