Fix pack_write's O(n^2) dedup + correctness bug; document mount table O(n^2)
A audit for other instances of the file-index O(n^2) shape (fixed previously via the persistent treap) found two more real issues: 1. pack_write's Section 9.2 exact-duplicate elimination was a linear scan of every previously-seen (hash, size) pair per entry -- O(n^2) total, invisible in the existing benchmark because its identical-content test files made every scan match on the first comparison. Measured with unique content instead: 80,000 entries took 1.74s, with a 20,000->80,000 step showing 15.6x for a 4x-N step, matching O(n^2)'s 16x prediction. The same scan also trusted a (hash, size) match without ever comparing actual bytes -- a latent correctness bug, since FNV-1a64 is explicitly not collision-resistant. Fixed both at once with an open-addressing hash table (load factor 1/2, linear probing) plus a memcmp verification before ever reusing a data_off. Post-fix: 80,000 entries in 0.044s (39.6x faster), ratio drops to 3.35x (consistent with O(n)). Covered permanently by two new/extended tests: a white-box assertion in test_pack_overlay.c that duplicate-content entries share one data_off and distinct-content entries do not, and a new tests/test_pack_write_perf.c regression tripwire against 10,000 unique entries. 2. vfs.c's mount table uses the same full-array-copy-per-write pattern the file index used to, confirmed O(n^2) via a new bench/bench.c category (500/2,000/8,000 mounts, both 4x-N steps showing 15-20x). Deliberately NOT rewritten: mount points are created by a program's own source code, not workload-driven, so realistic mount counts never reach the scale that made the file index's O(n^2) a real problem. Documented with full reasoning in BENCH.md and CLAUDE.md rather than silently left as an undocumented gap. Also fixes a real CI gap the new tests exposed: ci.yml's sanitizer-build steps never passed -D_GNU_SOURCE when compiling test files (only the library .o's got it), which was harmless while no test included internal.h and became a link failure once two did (internal.h needs _GNU_SOURCE for pthread_rwlock_t). And documents, in CONTRIBUTING.md and CLAUDE.md, a sandbox flake observed directly during this work's own sanitizer runs: ASan/UBSan test binaries occasionally fail to start with AddressSanitizer:DEADLYSIGNAL (sometimes looping rather than exiting), non-deterministically hitting different unrelated binaries across runs -- a startup race, not a memory-safety bug, confirmed by clean passes on retry; sanitizer runs in such an environment should be timeout-wrapped. BENCH.md's "After" table and Appendix B are replaced with the current, complete 54-measurement bench/bench.c run (the original 45 plus the new mount-scaling category); the pre-fix 45-measurement "Before" table is kept as the historical record, per this project's documentation standard. Verified: make test (all 6 binaries, including the 2 new/changed), a clean make all, and repeated ASan+UBSan runs (0 real findings; the DEADLYSIGNAL flake above was observed and correctly distinguished from a real finding by re-running until a clean pass). TSan could not be run in this sandbox (pre-existing, documented environment limitation). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
This commit is contained in:
@@ -68,18 +68,21 @@ process's lifetime.
|
||||
|
||||
`bench/bench.c` (`make bench`) measures PackFS against the host filesystem
|
||||
across metadata operations (create/read/stat/readdir/unlink/mkdir), large
|
||||
sequential I/O, random-access pack reads, and concurrent mixed workloads.
|
||||
[`BENCH.md`](BENCH.md) has the full results and honest analysis of both the
|
||||
wins and the losses, including the story of a real bug this benchmark
|
||||
found: an early run turned up a quantitatively-confirmed O(n²) cost in bulk
|
||||
sequential file creation, which `concept.md` Section 5.3 had explicitly
|
||||
anticipated and named the trigger condition for closing. That's since been
|
||||
fixed (`src/upper.c`'s index is a persistent treap now, not a flat array),
|
||||
re-benchmarked, and re-confirmed as an actual complexity-class change, not
|
||||
just a faster constant factor — see `BENCH.md`'s "Resolution" section for
|
||||
the before/after numbers. Read the whole file before quoting a number from
|
||||
it: what `raw fs` vs `raw+fsync` vs `dir` each actually measure is not
|
||||
interchangeable, and it explains why.
|
||||
sequential I/O, random-access pack reads, mount-table scaling, and
|
||||
concurrent mixed workloads. [`BENCH.md`](BENCH.md) has the full results and
|
||||
honest analysis of both the wins and the losses, including the story of
|
||||
three real O(n²) findings this project's own benchmarking turned up — not
|
||||
just the wins. Two are fixed: bulk sequential file creation (`src/upper.c`'s
|
||||
index is a persistent treap now, not a flat array — see "Resolution") and
|
||||
`pack_write`'s compaction-time duplicate-content elimination, which also had
|
||||
a latent correctness bug now closed alongside it (see "Resolution #2"). One
|
||||
is confirmed and *deliberately not* fixed: the mount table scales O(n²) in
|
||||
mount count, the same way the file index used to, but mount counts are
|
||||
bounded by a program's own source code rather than workload-driven, so it
|
||||
isn't worth the added complexity — see "Finding: mount table scaling" for
|
||||
the reasoning and the numbers behind that call. Read the whole file before
|
||||
quoting a number from it: what `raw fs` vs `raw+fsync` vs `dir` each
|
||||
actually measure is not interchangeable, and it explains why.
|
||||
|
||||
`openat2`/Landlock support (Section 6) is detected automatically at compile
|
||||
time via `<sys/syscall.h>`; on kernels or platforms without them, `dir`
|
||||
|
||||
Reference in New Issue
Block a user