Reduce memory footprint for large inputs, harden allocation failure paths
Profiled a real search with Massif and found the memory-per-input-byte multiplier at 13.4x, dominated by two avoidable costs: - build_matbuf widened every BINARY/ASCII byte into a uint32_t before matching, a 4x copy that mode never needed (a byte never exceeds 255). Removed it: MCtx/MatBuf now carry an optional text8 (borrowed, unwidened) alongside the existing UTF8 text (owned, decoded code points), reconciled per-read through text_at()/buf_at(). BINARY/ ASCII mode now points directly into the Input's own buffer. - Input_from_file always copied the whole file into a malloc'd buffer. It now mmaps regular files read-only (MAP_PRIVATE) instead, so pages stay clean and reclaimable under memory pressure and the file is never copied. Non-seekable sources (pipes, FIFOs, process substitution, stdin) and mmap failures fall back to the previous incremental-read behavior via a separate read_fd_incrementally. Together these bring the multiplier to 8.4x and let a 200MB file that previously OOM-crashed complete a full non-matching search in about 6 seconds at roughly 1.69GB peak RSS. Also tried, measured, and reverted: capping compute_maxrun/ compute_next_prevmatch's table size with a plain-scan fallback above the cap. A real 200MB non-matching search against this fallback hung for minutes instead of failing fast, because disabling either table reintroduces the O(n^2) behavior they exist to prevent, and O(n^2) at n in the hundreds of millions is not practically finite. A fast, diagnosable allocation failure is a better failure mode than a silent, unbounded hang, so the tables are allocated unconditionally again; the finding is recorded in code comments, concept.md 7.6, and README's "Memory footprint" section so it is not retried blindly later. Separately audited every allocation on an input-proportional path (da_push, build_matbuf's UTF-8 decode loop, all three Input_from_file sites) and made each fail cleanly through PatternError instead of crashing on an unchecked NULL dereference. A 1GB file still exceeds available memory in the current environment; this is a property of the machine it was measured on, not a defect, and is documented as such (practical ceiling: available memory / 8.4 for search-family operations, pending the streaming automaton design in concept.md 7.2). Verified with four clean `make test` passes (3252/3252) and a clean ASan/UBSan pass after the change; rxgrep's mmap-backed paths (--sub with a regular file, with a non-seekable process-substitution source, and BINARY-mode embedded NUL handling) re-checked directly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EjuMk8kY9SDus1wWe2K9xY
This commit is contained in:
+4
-2
@@ -67,8 +67,8 @@ void Input_free(Input *in);
|
||||
```
|
||||
|
||||
- `Input_from_buffer`: wraps an existing buffer. **Does not copy it and does not take ownership.** The buffer must outlive the `Input` and every `Match` produced from it (`Match_group` returns pointers directly into it, Section 3.9). Freeing an `Input_from_buffer` `Input` never frees the underlying buffer; the caller is responsible for that.
|
||||
- `Input_from_file`: reads the whole file into a freshly allocated, owned buffer. Works on both seekable files and non-seekable sources (pipes, FIFOs, process substitution, `/dev/stdin`): a seekable source is read in one `fread` after `fseek`/`ftell` sizing it; a non-seekable source is read incrementally into a growable buffer. Either way, the entire input ends up in memory before any matching happens (`README.md` "Implementation status"). On failure (cannot open, cannot allocate) returns `NULL` and, if `err` is non-`NULL`, fills it with `strerror(errno)` as `msg`.
|
||||
- `Input_free`: frees the `Input` and, only if it owns its buffer (true for `Input_from_file`, false for `Input_from_buffer`), the buffer too.
|
||||
- `Input_from_file`: for a regular file, `mmap`s it read-only (`MAP_PRIVATE`) instead of copying it into the heap, so the pages are backed by the file and the kernel can reclaim them under memory pressure (`README.md` "Memory footprint"). For a non-seekable source (pipe, FIFO, process substitution, `/dev/stdin`) or when `mmap` itself fails, it falls back to reading incrementally into a growable, owned heap buffer. Either way, the entire input is addressable before any matching happens (`README.md` "Implementation status"); only the second path actually copies it. On failure (cannot open, cannot allocate) returns `NULL` and, if `err` is non-`NULL`, fills it with `strerror(errno)` as `msg`.
|
||||
- `Input_free`: frees the `Input` struct itself always, and additionally releases the underlying buffer unless it came from `Input_from_buffer`: `munmap`s it if it was mapped, or `free`s it if it was read into an owned heap buffer (both cases of `Input_from_file`).
|
||||
|
||||
### 1.4 `PatternError` (`re.error` / `re.PatternError`)
|
||||
|
||||
@@ -147,6 +147,8 @@ Mirror `re.Pattern.match`/`.fullmatch`/`.search` exactly, including the `pos`/`e
|
||||
|
||||
**Performance:** `match`/`fullmatch` do one anchored attempt and cost time proportional to that attempt alone. `search` tries each candidate start position as a separate attempt, but, unlike a naive backtracking search, does not redo the same work at every one: for a backreference-free pattern, `run_memo` caches every proven failure at the (instruction, position) level, and `OP_REPEAT1` (a quantifier over a single character, class, or `.`) additionally uses precomputed run-length and skip-ahead tables so its own internal work is `O(1)` amortized per position rather than `O(remaining length)`. Together these make `search` linear, not quadratic, for the common case (a pattern with no backreference, built from simple repeated atoms). What is not fixed: a repeat over a *compound* body (`(ab)*`) gets the failure-memoization but not the skip-ahead table, so a pathological compound-repeat pattern can still be worse than linear; and any pattern with a backreference disables memoization entirely (unsound there, see Section 5) and can still be worst-case exponential, exactly as in CPython. See README.md "Implementation status" for the measured numbers on both the fixed case and the remaining one, and for citations to the published techniques this uses.
|
||||
|
||||
**Memory:** the tables above cost real, measured memory, not just time complexity: `search`/`finditer`/`split`/`sub` on a pattern using `OP_REPEAT1` peak at roughly 8.4x the input length (measured with Valgrind/Massif; `match`/`fullmatch` do not allocate these tables at all and use proportionally less). README.md "Memory footprint" has the full measured breakdown, including a fix that was tried, measured, and deliberately reverted because it traded a fast allocation failure for an effectively-unbounded hang, which is a worse failure mode, not a better one.
|
||||
|
||||
### 3.5-3.6 `Pattern_finditer` / `Pattern_findall`
|
||||
|
||||
```c
|
||||
|
||||
Reference in New Issue
Block a user