Files
regexx/regexx.h
T
retoorandClaude Sonnet 5 fa07622671 Add a Pike VM (regular engine, concept.md 7.2), dispatched automatically
Researched first (Cox's regexp2 article, the submatch/tagged-NFA
follow-up crediting Laurikari, and rust-lang/regex's PikeVM source for
the exact leftmost-first mid-search Match handling), then planned in
concept.md 7.7 before writing any code, per the explicit instruction to
plan before implementing.

The existing bytecode compiler turned out to already produce valid
Thompson-construction NFA bytecode (OP_SPLIT/OP_JMP/OP_SAVE match
Cox's instruction set almost exactly), so no new compiler was needed:
Prog gained one field (no_repeat1) and compile_repeat one condition,
letting the same compile_node produce a second, NFA-only Prog from the
same AST for any pattern containing none of OP_BACKREF/OP_LOOKAHEAD/
OP_LOOKBEHIND/OP_ATOMIC (a possessive quantifier already desugars to
the last of these at parse time). That second Prog is matched by a new
Pike VM (pike_addthread/pike_step/pike_find): a breadth-first thread
list simulation with per-thread capture arrays, epsilon closure
implemented with an explicit heap stack rather than C recursion so a
pattern with many alternations cannot recurse the C stack, every
allocation checked and failing through the existing -1 error
convention rather than a NULL dereference. do_one, Pattern_finditer,
and Pattern_split each gained an impl->has_nfa branch to this engine,
sharing one small helper (find_next) for the "unanchored scan from a
position" versus "single anchored attempt at a position" distinction
finditer/split's empty-match retry needs.

Two real bugs, both found and precisely localized by the existing
3,252-case CPython-derived suite without writing a single test
specifically for this engine: OP_MATCH not writing group 0's end
position (it has no OP_SAVE; run()'s own OP_MATCH handler sets it
directly, and this engine's first version missed replicating that),
and an unconditional "thread list empty -> stop" early exit that is
wrong for unanchored search specifically, since a freshly injected
start thread can die immediately in its own epsilon closure (a
leading \b failing outright, repeatedly, inside a longer word like
"catalog" for \bcat\b) without that meaning every later position
would too. Both fixed; full suite passes, three clean AddressSanitizer/
UndefinedBehaviorSanitizer runs, plus hand-written whitebox checks
(a 20,000-branch alternation, UTF8 named groups, BINARY matching
across an embedded NUL, greedy/lazy and alternation priority).

Measured result: (a+)+b, this project's own running example of the
backtracking engine's remaining weak spot, is Pike VM eligible and now
measures as genuinely linear (0.0018s to 0.0308s, n=10,000 to
160,000), not merely improved. Measured cost: re-running
bench_vs_posix.c's three ordinary scenarios (all now Pike VM eligible
too) found the gap to POSIX <regex.h> widened, from roughly 2x-25x
before this engine existed to roughly 8x-70x now, the direct,
expected cost of this engine's performance axis being explicitly
deferred (no literal prefilter, no lazy DFA state caching, no
allocation pooling) in favor of correctness first, per the instruction
this was built under. Both results, and the reasoning behind
deferring the second, are recorded in concept.md 7.7/7.8.

README.md, docs/API.md, USAGE.md, and examples/redos_atomic.c and
bench_vs_posix.c are updated throughout to describe the new two-engine
dispatch accurately, including this real trade-off, rather than
leaving the previous single-engine description in place; redos_atomic.c
specifically now shows both that (a+)+b no longer needs an atomic group
at all and a backreference-forced variant where one still does.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
2026-09-14 11:38:44 +00:00

144 lines
7.4 KiB
C

/* regexx.h - public API for regexx, a single-file C regex interpreter
* with Python `re` semantics. See concept.md for the design rationale
* and Section 9 in particular for the naming convention this header
* follows: every identifier with a direct Python `re` counterpart uses
* that counterpart's exact spelling.
*
* Implementation status relative to concept.md: this is the v1
* implementation. It provides the full public API and pattern syntax
* described below, executed by one of two engines over a fully
* materialized copy of the input, chosen automatically per compiled
* pattern (concept.md Section 7.2/7.7): a Pike VM (a Thompson-NFA
* simulation, genuinely linear time, no backtracking at all) for any
* pattern with no backreference, lookaround, or atomic group; a
* recursive backtracking engine (concept.md Section 7.3), augmented
* with memoization and precomputed skip-ahead tables (regexx.c:
* run_memo, compute_maxrun, compute_next_prevmatch) that make the
* search-family functions (Pattern_search/finditer/split/sub) run in
* linear time for the common, backreference-free case there too,
* otherwise. Neither engine yet implements the streaming, bounded-
* memory version of Section 7.2; see README.md "Implementation status"
* and "The Pike VM" for the complete, itemized, measured account of
* what v1 does and does not cover, including a real, documented
* performance trade-off between the two engines that is not yet
* resolved (concept.md 7.8).
*/
#ifndef REGEXX_H
#define REGEXX_H
#include <stddef.h>
#include <stdint.h>
#ifdef __cplusplus
extern "C" {
#endif
/* ---- Flags (re.compile flags; bitwise OR, exact Python names) ---- */
#define IGNORECASE 0x0001
#define MULTILINE 0x0002
#define DOTALL 0x0004
#define VERBOSE 0x0008
#define ASCII 0x0010
#define UNICODE 0x0020 /* default for UTF8 mode; explicit for clarity */
#define LOCALE 0x0040
#define DEBUG 0x0080
/* Encoding mode, not a Python flag (Section 9.3): exactly one required. */
#define BINARY 0x0100
#define UTF8 0x0200
/* ASCII (0x0010) doubles as the ASCII *encoding* mode when neither
* BINARY nor UTF8 is given; see regexx.c encoding-mode resolution. */
/* ---- Types ---- */
typedef struct Pattern Pattern; /* re.Pattern */
typedef struct Match Match; /* re.Match */
typedef struct Input Input; /* no Python counterpart, see concept.md 9.0/9.2 */
typedef void (*MatchIterCb)(void *ctx, const Match *m);
typedef void (*MatchSubCb)(void *ctx, const Match *m, char **out, size_t *outlen);
typedef struct PatternError PatternError; /* re.error / re.PatternError */
struct PatternError {
const char *msg; /* re.error.msg */
const char *pattern; /* re.error.pattern */
int64_t pos; /* re.error.pos */
int64_t lineno; /* re.error.lineno */
int64_t colno; /* re.error.colno */
};
/* msg and pattern are heap allocated by whichever call filled this
* struct in; call PatternError_free once you are done reading it
* (safe to call on a zero-initialized or already-freed PatternError). */
void PatternError_free(PatternError *err);
struct Pattern {
const char *pattern; /* re.Pattern.pattern */
int flags; /* re.Pattern.flags */
int groups; /* re.Pattern.groups */
void *groupindex; /* re.Pattern.groupindex, name -> group number */
void *program; /* compiled bytecode, private */
};
struct Match {
Pattern *re; /* re.Match.re */
Input *string; /* re.Match.string */
int64_t pos, endpos; /* re.Match.pos, re.Match.endpos */
int lastindex; /* re.Match.lastindex */
const char *lastgroup; /* re.Match.lastgroup */
void *slots; /* private */
};
/* ---- Input construction (no Python counterpart) ---- */
Input *Input_from_buffer(const uint8_t *buf, size_t len); /* does not copy or take ownership */
Input *Input_from_file(const char *path, PatternError *err); /* seekable or not (pipes, "-" via /dev/stdin, etc. all work) */
void Input_free(Input *in);
/* ---- Pattern methods (bound to an already compiled Pattern) ---- */
int Pattern_match(Pattern *self, Input *string, int64_t pos, int64_t endpos, Match *out);
int Pattern_fullmatch(Pattern *self, Input *string, int64_t pos, int64_t endpos, Match *out);
int Pattern_search(Pattern *self, Input *string, int64_t pos, int64_t endpos, Match *out);
int Pattern_finditer(Pattern *self, Input *string, int64_t pos, int64_t endpos, MatchIterCb cb, void *ctx);
int Pattern_findall(Pattern *self, Input *string, int64_t pos, int64_t endpos, MatchIterCb cb, void *ctx);
/* cb fires once per element of the list re.split() would return, in
* order; read each element's text via Match_group(m, NULL, 0, ...). */
int Pattern_split(Pattern *self, Input *string, int maxsplit, MatchIterCb cb, void *ctx);
int Pattern_sub(Pattern *self, Input *string, const char *repl, MatchSubCb cb, void *ctx, int count, char **out, size_t *outlen);
int Pattern_subn(Pattern *self, Input *string, const char *repl, MatchSubCb cb, void *ctx, int count, char **out, size_t *outlen, int *n);
void Pattern_free(Pattern *self);
/* Looks up a named group in re.Pattern.groupindex; returns its 1-based
* group number, or -1 if no group by that name exists. There is no
* enumeration function for groupindex as a whole in this build: a
* caller that needs every name must track the names it used to build
* the pattern itself, since Pattern->groupindex is otherwise opaque. */
int Pattern_groupindex_lookup(Pattern *self, const char *name);
/* ---- Match accessors ---- */
int Match_group(Match *self, const char *name_or_null, int index, const char **out, size_t *outlen);
int64_t Match_start(Match *self, int group);
int64_t Match_end(Match *self, int group);
void Match_span(Match *self, int group, int64_t *start, int64_t *end);
int64_t Match_start_byte(Match *self, int group);
int64_t Match_end_byte(Match *self, int group);
void Match_span_byte(Match *self, int group, int64_t *start, int64_t *end);
void Match_free(Match *self);
/* ---- Module level functions ---- */
Pattern *re_compile(const char *pattern, size_t len, int flags, PatternError *err);
int re_match(const char *pattern, size_t len, int flags, Input *string, Match *out);
int re_fullmatch(const char *pattern, size_t len, int flags, Input *string, Match *out);
int re_search(const char *pattern, size_t len, int flags, Input *string, Match *out);
int re_finditer(const char *pattern, size_t len, int flags, Input *string, MatchIterCb cb, void *ctx);
int re_findall(const char *pattern, size_t len, int flags, Input *string, MatchIterCb cb, void *ctx);
int re_split(const char *pattern, size_t len, int flags, Input *string, int maxsplit, MatchIterCb cb, void *ctx);
int re_sub(const char *pattern, size_t len, int flags, Input *string, const char *repl, MatchSubCb cb, void *ctx, int count, char **out, size_t *outlen);
int re_subn(const char *pattern, size_t len, int flags, Input *string, const char *repl, MatchSubCb cb, void *ctx, int count, char **out, size_t *outlen, int *n);
void re_escape(const char *in, size_t len, char **out, size_t *outlen);
void re_purge(void);
#ifdef __cplusplus
}
#endif
#endif /* REGEXX_H */