Add a benchmark suite comparing PackFS against the host filesystem

bench/bench.c (`make bench`) measures create/read/stat/readdir/unlink/
mkdir on mem, dir, and raw fs; large sequential I/O; random-access pack
reads via mmap vs raw fs; compaction throughput; and 8-thread
concurrent mixed workloads. Results from one full run, with honest
analysis (what each "raw fs" vs "raw+fsync" vs "dir" label actually
measures, so they aren't misread as interchangeable), are in BENCH.md.

The benchmark surfaced a real, quantitatively-confirmed finding, not
just favorable numbers: bulk sequential create/unlink on mem/dir is
O(n^2) in file count (~9-22x slower than raw fs at N=20,000), because
every structural write copies the entire snapshot entry array before
publishing it (Section 5.3). concept.md itself names the exact trigger
condition for reconsidering this ("a persistent structurally-shared
tree structure is not required until this assumption is empirically
violated") — this benchmark is that violation, measured rather than
hypothesized: mkdir at N=4,000 vs create at N=20,000 (same mechanism,
5x the N) shows a 28.4x slowdown, matching the O(n^2) prediction (25x)
far better than O(n) (5x).

Also found: dir-backend stat() costs ~2x raw stat() (open+fstat vs one
syscall, the direct cost of openat2 containment on the metadata path);
mem-backed large writes lose to raw fs above ~16MB (Section 5.7's
buffer-growth discipline re-copies prior writes on every capacity
doubling, the price of never exposing a reader to a freed buffer).

Recorded the O(n^2) finding in CLAUDE.md's new "Known performance
characteristics" section, per the same pattern used for the
reclaim_gate use-after-free discovery, since it's exactly the kind of
fact that would otherwise have to be rediscovered by benchmarking
again from scratch. Linked from README and CONTRIBUTING.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
This commit is contained in:
2026-09-14 07:40:42 +00:00
co-authored by Claude Sonnet 5
parent dbbbadf025
commit 0b3b207bc0
6 changed files with 790 additions and 2 deletions
+14
View File
@@ -48,6 +48,7 @@ preference).
make # builds libpackfs.a and libpackfs.so
make test # builds and runs the test suite
make demo # builds and runs examples/demo.c — see "Try it" below
make bench # builds and runs bench/bench.c — see "Benchmarks" below
make install # installs to $PREFIX (default /usr/local)
```
@@ -63,6 +64,19 @@ first `readdir` shows the first run's files, proving that compaction and
reload actually persist data through the pack file, not just within one
process's lifetime.
## Benchmarks
`bench/bench.c` (`make bench`) measures PackFS against the host filesystem
across metadata operations (create/read/stat/readdir/unlink/mkdir), large
sequential I/O, random-access pack reads, and concurrent mixed workloads.
Full results from one run, with honest analysis of both the wins and the
losses — including a real, quantitatively-confirmed O(n²) cost in bulk
sequential file creation that `concept.md` Section 5.3 explicitly
anticipated and named the trigger condition for — are in
[`BENCH.md`](BENCH.md). Read it before quoting a number from it: what
`raw fs` vs `raw+fsync` vs `dir` each actually measure is not
interchangeable, and the file explains why.
`openat2`/Landlock support (Section 6) is detected automatically at compile
time via `<sys/syscall.h>`; on kernels or platforms without them, `dir`
mounts fall back to the weaker, documented residual-risk posture described