Fix the O(n^2) bulk create/unlink: flat array -> persistent treap

BENCH.md's earlier finding was real: every structural write (create/
unlink/mkdir) copied the entire sorted UpperSnapshot entry array before
publishing the next snapshot, making bulk sequential creation O(n^2).
concept.md Section 5.3 names the exact condition for reconsidering this
("a persistent structurally-shared tree structure is not required
until this assumption is empirically violated") -- that condition was
measured, not hypothesized, so this closes it rather than leaving it
as a documented-but-open limitation.

The index is now a persistent treap (src/upper.c): a structural write
copies only the O(log n) nodes on the path to the change, sharing
every other node (its own refcount, cascading like MutCell's and
UpperSnapshot's) with whichever snapshot(s) it was built from. Chosen
over a persistent AVL/red-black/weight-balanced tree because deletion
in those can need O(log n) rebalancing rotations -- each a real
allocation in a persistent setting -- where a treap needs only O(1)
amortized rotations for insert and delete (Seidel & Aragon 1996), with
expected O(log n) height regardless of insertion order, including the
sorted-by-creation-order pattern that made the flat array quadratic in
the first place. Liljenzin's "Confluently Persistent Sets and Maps"
(arXiv:1301.3388) documents persistent treaps giving O(1) snapshots
for MVCC specifically, which is this exact use case. Full reasoning
and the ownership convention (functions consume one ref of their tree
arguments, return one owned ref) are in upper.c's comment above
struct TreapNode.

Blast radius kept deliberately small: snapshot_upsert/snapshot_remove
keep their exact original signatures, so upper_create/upper_mkdir/
upper_remove/upper_rename/upper_copy_up needed zero changes.
upper_lookup/upper_has_children keep their exact contracts. Only
upstd_readdir and overlay_readdir's manual array scans became calls to
a new upper_visit_range (O(log n + r) range query, replacing an O(n)
scan in both, a bonus fix beyond what was strictly necessary) since
there's no flat array left to scan.

Added tests/test_index_stress.c: thousands of randomized (not
sequential) creates/deletes/renames across nested directories,
cross-checked against an independent reference model after every
round, not just "did it not crash" -- exercises exactly the code path
the O(n^2) bug and this fix live in, at a scale the other tests don't
reach. Verified under -fsanitize=undefined (150+ runs across this
change's lifetime, 0 failures) and -fsanitize=address (100+ runs, 0
real findings; known sandbox ASan-startup flakes excluded, see prior
commits) per CLAUDE.md's sanitizer rule for upper.c/overlay.c changes.

Measured result (make bench, same environment as the original
finding): mem create 257x faster, unlink 330x faster, mkdir 65x
faster, concurrent mixed workload 131x faster -- and, the comparison
that matters, mem now beats raw fs at every one of these (was losing
by 6-22x before). The complexity-class change is confirmed the same
way the O(n^2) was found: mkdir at N=4,000 vs create at N=20,000 now
shows a 7.15x slowdown for a 5x increase in N, matching the O(n log n)
prediction (5.97x) rather than the old O(n^2) one (25x). Full
before/after tables in BENCH.md's new "Resolution" section, which
keeps the original run as the historical record rather than
overwriting it, per this project's own documentation standard.

Recorded the fix in CLAUDE.md's "Known performance characteristics"
(marked RESOLVED, not silently removed) and its architecture map.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
This commit is contained in:
2026-09-14 08:09:49 +00:00
co-authored by Claude Sonnet 5
parent 0b3b207bc0
commit 64283d587b
8 changed files with 753 additions and 257 deletions
+11 -7
View File
@@ -69,13 +69,17 @@ process's lifetime.
`bench/bench.c` (`make bench`) measures PackFS against the host filesystem
across metadata operations (create/read/stat/readdir/unlink/mkdir), large
sequential I/O, random-access pack reads, and concurrent mixed workloads.
Full results from one run, with honest analysis of both the wins and the
losses — including a real, quantitatively-confirmed O(n²) cost in bulk
sequential file creation that `concept.md` Section 5.3 explicitly
anticipated and named the trigger condition for — are in
[`BENCH.md`](BENCH.md). Read it before quoting a number from it: what
`raw fs` vs `raw+fsync` vs `dir` each actually measure is not
interchangeable, and the file explains why.
[`BENCH.md`](BENCH.md) has the full results and honest analysis of both the
wins and the losses, including the story of a real bug this benchmark
found: an early run turned up a quantitatively-confirmed O(n²) cost in bulk
sequential file creation, which `concept.md` Section 5.3 had explicitly
anticipated and named the trigger condition for closing. That's since been
fixed (`src/upper.c`'s index is a persistent treap now, not a flat array),
re-benchmarked, and re-confirmed as an actual complexity-class change, not
just a faster constant factor — see `BENCH.md`'s "Resolution" section for
the before/after numbers. Read the whole file before quoting a number from
it: what `raw fs` vs `raw+fsync` vs `dir` each actually measure is not
interchangeable, and it explains why.
`openat2`/Landlock support (Section 6) is detected automatically at compile
time via `<sys/syscall.h>`; on kernels or platforms without them, `dir`