# PackFS vs. the host filesystem: benchmark results **Status: two O(n²) findings this document originally reported have been fixed; one further O(n²) finding is documented below and deliberately left unfixed, with justification.** The file index (`UpperSnapshot`) is now a persistent treap instead of a flat sorted array (`src/upper.c`; see `CLAUDE.md`'s "Known performance characteristics") — see "Resolution" below. `pack_write`'s compaction-time duplicate-content elimination (Section 9.2) was also a linear scan with the same O(n²) shape, plus a latent correctness bug (a hash match was trusted without a `memcmp` verification); both are fixed — see "Resolution #2" below. The mount table (`vfs.c`'s `MountSnapshot`) uses the same full-array-copy pattern the file index used to, is confirmed O(n²) in mount count, and is *not* rewritten — see "Finding: mount table scaling" below for why that is the correct call, not an oversight. This document keeps every prior run's numbers as the historical record of each problem — per this project's own documentation standard, a correction extends the record rather than erasing it — and adds each fix's numbers alongside them, in full, not a curated subset of either. **Read "Resolution," "Resolution #2," and "Finding: mount table scaling" below before drawing conclusions from the "Before" tables**, which describe a version of the index this codebase no longer has. This document reports the results of `bench/bench.c` (`make bench`), each run once on the environment described below. These are single runs, not a statistically-averaged series — treat the numbers as illustrative of *shape* (which operations are faster, by roughly what factor, and why) more than as precise absolute figures for a different machine. The methodology, including exactly what each backend label measures, is documented in `bench/bench.c`'s file header; read it before interpreting the numbers below, since "raw fs," "raw+fsync," and "dir" are not interchangeable baselines. The "Before" table below reproduces every one of the 45 measurements that run's `=== summary (45 measurements) ===` block reported; the "After" table reproduces every one of the 54 measurements the current `bench/bench.c` reports (the original 45, unchanged in category, plus 9 new mount-table- scaling measurements added after the original run — see "Finding: mount table scaling"). None omitted, none combined in either table, and the complete verbatim program output for both runs, exactly as captured, is preserved in the two appendices at the end as the primary source record. The `pack_write` dedup fix ("Resolution #2") is measured separately, by calling `pack_write` directly rather than through `bench/bench.c` — see that section for why. ## Environment - Host: containerized, 12 logical CPUs (AMD Ryzen 5 3600), running under an `overlay` root filesystem (`overlay` on `overlay`, per `mount`) — there is no separate tmpfs at `/tmp` in this environment; every "raw fs" and `dir` number below is against that container overlay filesystem, not a bare disk or tmpfs. This matters most for the `fsync` numbers (see below). - `N_SMALL` = 20,000 files × 128 bytes for the metadata-heavy suite; `N_DIRS` = 4,000; concurrency = 8 threads × 4,000 ops; large files up to 64 MB; `N_MOUNTS` = 2,000 (run at N/4, N, N×4 — 500/2,000/8,000 — for the mount-table-scaling category added after the original run). Full parameters are in `bench/bench.c`. - Wall-clock time: 5m49s (before), 4m4s (after the treap fix, before the mount-scaling category existed), 5m24s (current, with the mount-scaling category added). The dominant cost throughout is the single `raw+fsync` 20,000-file create test (~116s in every run — this cost is on the host's storage, not PackFS, and is within measurement noise across all three runs, as expected, since that test never touches PackFS's index at all); the added mount-scaling category itself accounts for only ~3.7s of the increase from 4m4s to 5m24s, so the remainder is run-to-run variance in untimed setup/teardown, consistent with this document's stated single-run, illustrative-of-shape methodology, not a regression. ## Resolution The original run found bulk sequential `create`/`unlink`/`mkdir` on `mem`/`dir` to be O(n²) in file count, because every structural write copied the entire sorted entry array before publishing the next snapshot (Section 5.3). `concept.md` itself names the trigger condition for reconsidering that design — "a persistent (structurally shared) tree structure is not required until this assumption is empirically violated" — and the original run was exactly that violation, measured rather than hypothesized. The index was rewritten as a **persistent treap** (`src/upper.c`): a structural write now builds only the O(log n) nodes on the path to the change, sharing every other node (refcounted) with the snapshot it was built from, instead of copying all n entries. The choice of treap over a persistent AVL/red-black/weight-balanced tree, and the reasoning behind it, is documented in `src/upper.c`'s comment above `struct TreapNode` and cited there: Seidel & Aragon, "Randomized Search Trees" (1996), for the O(1) amortized rotation count on both insert and delete (persistent balanced trees that need O(log n) rotations pay a real, repeated allocation cost per rotation, which a treap avoids); Liljenzin, "Confluently Persistent Sets and Maps" (arXiv:1301.3388), for persistent treaps specifically as an MVCC snapshot mechanism, which is this exact use case. **Headline result, measured, same methodology as the original finding.** The "After" figures below are drawn from the "After: complete results" table further down this document (the current, canonical post-fix run), not from an intermediate run kept only in this section — every number here can be found again in that table: | Category | Before | After | Speedup | |---|---|---|---| | create 20,000 files (mem) | 9.5565s / 2,093 ops/s | 0.0400s / 499,809 ops/s | **239x** | | unlink 20,000 files (mem) | 10.5575s / 1,894 ops/s | 0.0179s / 1,117,742 ops/s | **590x** | | mkdir 4,000 dirs (mem) | 0.3363s / 11,896 ops/s | 0.0056s / 719,696 ops/s | **60x** | | rmdir 4,000 dirs (mem) | 0.3395s / 11,782 ops/s | 0.0040s / 993,724 ops/s | **85x** | | create 20,000 files (dir) | 12.0367s / 1,662 ops/s | 1.1418s / 17,517 ops/s | **11x** | | unlink 20,000 files (dir) | 10.1851s / 1,964 ops/s | 0.4911s / 40,724 ops/s | **21x** | | concurrent create+read+unlink (mem) | 36.7261s / 2,614 ops/s | 0.2750s / 349,127 ops/s | **134x** | (An earlier revision of this table cited a different post-fix run — same fixed code, an earlier `make bench` invocation — whose numbers were 65–330x rather than 60–590x; the two runs agree on every qualitative conclusion below, but `unlink 20,000 files (mem)` in particular differs by nearly 2x between them, 0.0320s vs 0.0179s, both measuring an operation fast enough (sub-20ms for 20,000 entries) that ordinary system jitter in this containerized environment has real proportional effect. This is exactly the kind of run-to-run variance this document's stated methodology warns about — "illustrative of shape... more than precise absolute figures" — and the reason the table now points at one single canonical run rather than mixing figures from two.) `mem` now *beats* raw fs (not merely "no longer loses to it") at every one of these: create is 25x faster than raw fs (was 9.5x slower), unlink is 27x faster (was 22x slower), and the concurrent mixed workload is 21x faster (was 6x slower) — the single-writer design's "optimizes for read-heavy concurrent workloads" caveat from the original analysis (below) no longer needs to exclude bulk-write workloads, because the write side is no longer the bottleneck it was. **The complexity-class change is confirmed quantitatively, the same way the original O(n²) finding was** — not just "it got faster," which an unrelated constant-factor optimization could also produce: `mkdir` at N=4,000 vs. `create` at N=20,000 (same structural-write mechanism, 5x the N) now shows a **7.14x** slowdown (0.0400s / 0.0056s). O(n) predicts 5.0x; O(n log n) predicts 5.97x; the old O(n²) regime predicted, and measured, 25–28.4x. 7.14x lands close to the O(n log n) prediction (some excess is expected — `create` writes bytes into a fresh `MutCell` and `mkdir` does not, a constant-factor difference the pure index-complexity comparison doesn't account for) and nowhere near quadratic — and is itself nearly identical to the 7.15x this same check gave against the earlier run referenced above, which is the reassurance that matters here: the *ratio* this check depends on is stable across runs even where the fastest individual measurements are not. The index's asymptotic behavior changed, not merely its constant factor. Full output of both runs is reproduced verbatim in the appendices below and is reproducible again with `make bench`; the correctness of the new index under heavy randomized structural churn (non-sequential creates/deletes/ renames, nested directories, thousands of entries, cross-checked against an independent reference model — not just "the benchmark ran fast without crashing") is `tests/test_index_stress.c`, part of `make test`. ## Resolution #2: pack_write compaction dedup A second, independent O(n²) finding, found while auditing the codebase for other instances of the same shape of bug after "Resolution" above: `src/ pack.c`'s `pack_write` implements Section 9.2's exact-duplicate elimination (two entries with byte-identical content share one `data_off` in the compacted pack) as a linear scan, for each entry, of every previously-seen `(hash, size)` pair. That is O(n) per entry, O(n²) total across n distinct blobs — the same complexity-class bug "Resolution" fixed in the file index, here in the compaction writer instead. It went unnoticed by the original benchmark for a specific, identifiable reason: `bench/bench.c`'s compaction test (category 8, "compact N entries to pack") writes byte-identical content into every file, so every entry's linear scan matched on the very first comparison (against entry 0) and never actually scanned — the benchmark's own test data happened to be the linear scan's *best* case, not a representative one. Measuring with unique per-file content instead (each entry's actual worst case) surfaced it: | N (unique entries) | Time | |---|---| | 5,000 | 0.0138s | | 20,000 | 0.1114s | | 80,000 | 1.7389s | 5,000 → 20,000 is a 4x step in N; time increased 8.07x. 20,000 → 80,000 is also a 4x step; time increased 15.61x, converging toward O(n²)'s 16x prediction as N grows (the smaller first ratio reflects fixed per-call overhead not yet dominated by the quadratic term) — the same convergence pattern "Resolution" observed for the file index, and the same standard used there to distinguish a real complexity-class finding from noise. This scan had a second, latent problem beyond speed: it compared only `(hash, size)`, never the actual bytes, before deciding two entries were duplicates and sharing their `data_off`. FNV-1a64 (`src/hash.c`) is explicitly not collision-resistant — it is a fast checksum, not a cryptographic hash — so two different-content entries that happened to collide on `(hash, size)` would have been silently merged into one blob, corrupting one of them. This was never observed in practice (no test's data happened to collide), but was true of the code as written, not a hypothetical: the fix required for the O(n²) issue also closes it. **Fix:** the linear scan was replaced with an open-addressing hash table (load factor 1/2, linear probing) over `(hash, size)`, plus — critically — an actual `memcmp` against the candidate's stored data pointer before ever reusing a `data_off`, so a hash match is only ever a candidate, never accepted as proof of equality. Comment and full listing: `src/pack.c`'s `DedupSlot` struct and the rewritten first pass of `pack_write`. **After the fix, same methodology:** | N (unique entries) | Time | vs. before | |---|---|---| | 5,000 | 0.0079s | 1.75x faster | | 20,000 | 0.0131s | 8.50x faster | | 80,000 | 0.0439s | 39.6x faster | 5,000 → 20,000 now shows 1.66x (O(n) predicts 4x, but a hash table's dominant cost at these sizes is close to constant per insert, so sub-linear growth here is expected, not suspicious); 20,000 → 80,000 shows 3.35x, close to O(n)'s 4x prediction and nowhere near O(n²)'s 16x. The complexity class changed, exactly as with the file index. This fix is not reflected in the "After" table/Appendix B below, because that `make bench` run predates it — but it does not need to be: category 8's "compact N entries to pack" measurement uses identical content (as explained above), which was already the linear scan's O(1)-per-entry best case before this fix and is unaffected by the fix (a hash table lookup that matches on the first probe is also O(1) per entry); the number in the "After" table (0.0267s, 20,000 identical-content entries) remains accurate and does not need to be re-measured. The fix matters for compacting many files with *distinct* content, which `bench/bench.c`'s existing compaction test does not exercise — hence the separate, direct `pack_write` measurement above rather than a `bench/bench.c` change. Correctness is covered permanently by `tests/test_pack_overlay.c` (asserts `/a.txt` and `/dup.txt`, identical content, share one `data_off`, while `/b.txt`, different content, shares neither's — catching both "dedup stopped happening" and "dedup over-matched"); the performance fix is covered permanently by `tests/test_pack_write_perf.c` (10,000 unique entries under a generous time bound, a regression tripwire rather than a tight assertion, run as part of `make test`). ## Finding: mount table scaling (O(n²), documented, deliberately not fixed) `src/vfs.c`'s mount table (`MountSnapshot`) uses the exact same pattern the file index used before "Resolution" above: every `vfs_mount`/`vfs_unmount` copies the entire array of mounted backends before publishing the next snapshot (Section 5.3's single-writer model applies to the mount table too, per `concept.md`'s explicit statement that "the mount table is not separate global state exempt from this model"). This is confirmed O(n²) in mount count by the `bench/bench.c` category 10 measurements (full numbers in the "After" table below and Appendix B): | N mounts | mount | resolve (stat through N mounts) | unmount | |---|---|---|---| | 500 | 0.0054s | 0.0020s | 0.0046s | | 2,000 | 0.0831s | 0.0296s | 0.0758s | | 8,000 | 1.4739s | 0.4938s | 1.5172s | 500 → 2,000 is a 4x step in N: mount increased 15.4x, resolve 14.8x, unmount 16.5x. 2,000 → 8,000 is also a 4x step: mount increased 17.7x, resolve 16.7x, unmount 20.0x. All six ratios cluster around O(n²)'s 16x prediction for a 4x-N step, and nowhere near O(n)'s 4x — the same complexity-class-confirmation methodology used for both findings above, applied here and yielding the same verdict. Notably, `resolve` (a single `vfs_stat` through an already-built N-mount table) is *also* O(n²) across N resolves, not just O(n) per resolve as might be assumed — consistent with longest-prefix-match dispatch being a linear scan of the mount table per call (Section 3's dispatch design), which is itself O(n) per call and O(n²) across N calls, independent of the snapshot-copy cost on the write side. **This is deliberately not fixed, and the reasoning is the same trigger condition `concept.md` Section 5.3 names for the file index — inverted:** "a persistent (structurally shared) tree structure is not required until this assumption is empirically violated." The file index's assumption was violated because file counts are workload-driven and can plausibly reach tens of thousands (or more) at runtime, chosen by whatever program embeds PackFS and whatever data it manages. Mount counts are different in kind, not just degree: a mount point is created by a call to `vfs_mount` written into a program's own source code, at a location the program's author chose — it is bounded by how many lines of "mount this backend at this path" a person or a build script is willing to write, not by user data, input size, or any other externally-driven quantity. A program with 8,000 mount points, structured that way on purpose, is not a realistic workload this system needs to serve well; a program with 8,000 *files* is an ordinary one. Converting the mount table to a persistent treap would fix this at the cost of real, permanent complexity (an additional generic persistent-tree instantiation, or a second copy of the treap logic specialized to prefix matching) for a case that has not been and is not expected to be hit. Should a future use case genuinely need thousands of mount points at runtime — for instance, one mount per user in a multi-tenant embedding — this finding, and the numbers above, are the starting point for that decision; until then, this is recorded as a known, measured, and consciously accepted limitation, not an unexamined one. ## Before: complete results (pre-fix; kept as the historical record — see "Resolution" above) All 45 measurements from that run's summary block, none omitted or combined. | Category | Backend | Time | Throughput | MB/s | |---|---|---|---|---| | create 20,000 files | mem | 9.5565s | 2,093 ops/s | 0.3 | | create 20,000 files | raw fs | 1.0083s | 19,836 ops/s | 2.4 | | create 20,000 files | raw+fsync | 116.1030s | 172 ops/s | 0.0 | | read 20,000 files | mem | 0.0071s | 2,798,731 ops/s | 341.6 | | read 20,000 files | raw fs | 0.1611s | 124,122 ops/s | 15.2 | | stat 20,000 files | mem | 0.0062s | 3,223,599 ops/s | — | | stat 20,000 files | raw fs | 0.0703s | 284,465 ops/s | — | | readdir (20,000 entries) | mem | 0.0037s | 5,397,828 /s | — | | readdir (20,000 entries) | raw fs | 0.0071s | 2,831,904 /s | — | | create 20,000 files | dir | 12.0367s | 1,662 ops/s | 0.2 | | read 20,000 files | dir | 0.1674s | 119,468 ops/s | 14.6 | | stat 20,000 files | dir | 0.1516s | 131,937 ops/s | — | | readdir (20,000 entries) | dir | 0.0028s | 7,058,003 /s | — | | unlink 20,000 files | mem | 10.5575s | 1,894 ops/s | — | | unlink 20,000 files | dir | 10.1851s | 1,964 ops/s | — | | unlink 20,000 files | raw fs | 0.4782s | 41,823 ops/s | — | | mkdir 4,000 dirs | mem | 0.3363s | 11,896 ops/s | — | | rmdir 4,000 dirs | mem | 0.3395s | 11,782 ops/s | — | | mkdir 4,000 dirs | dir | 0.5109s | 7,829 ops/s | — | | rmdir 4,000 dirs | dir | 0.4892s | 8,177 ops/s | — | | mkdir 4,000 dirs | raw fs | 0.1591s | 25,143 ops/s | — | | rmdir 4,000 dirs | raw fs | 0.2119s | 18,876 ops/s | — | | write 1MB | mem | 0.0001s | 6,687 ops/s | 6,687.1 | | read 1MB | mem | 0.0000s | 50,941 ops/s | 50,941.4 | | write 16MB | mem | 0.0138s | 72 ops/s | 1,157.8 | | read 16MB | mem | 0.0010s | 1,015 ops/s | 16,240.8 | | write 64MB | mem | 0.0566s | 18 ops/s | 1,129.8 | | read 64MB | mem | 0.0034s | 292 ops/s | 18,705.5 | | write 1MB | raw fs | 0.0004s | 2,685 ops/s | 2,685.4 | | read 1MB | raw fs | 0.0001s | 16,584 ops/s | 16,583.7 | | write 16MB | raw fs | 0.0047s | 213 ops/s | 3,412.8 | | read 16MB | raw fs | 0.0010s | 960 ops/s | 15,363.8 | | write 64MB | raw fs | 0.0186s | 54 ops/s | 3,435.6 | | read 64MB | raw fs | 0.0047s | 213 ops/s | 13,635.3 | | write 1MB | raw+fsync | 0.0087s | 116 ops/s | 115.5 | | read 1MB | raw+fsync | 0.0001s | 8,857 ops/s | 8,857.4 | | write 16MB | raw+fsync | 0.0840s | 12 ops/s | 190.5 | | read 16MB | raw+fsync | 0.0013s | 796 ops/s | 12,737.8 | | write 64MB | raw+fsync | 0.0891s | 11 ops/s | 718.3 | | read 64MB | raw+fsync | 0.0045s | 220 ops/s | 14,096.3 | | compact 20,000 entries to pack | pack | 0.0291s | 687,097 ops/s | 83.9 | | random-read 20,000 entries | pack (mmap'd) | 0.0085s | 2,350,224 ops/s | 286.9 | | random-read 20,000 entries | raw fs | 0.1767s | 113,198 ops/s | 13.8 | | concurrent create+read+unlink (8×4,000×3) | mem | 36.7261s | 2,614 ops/s | — | | concurrent create+read+unlink (8×4,000×3) | raw fs | 6.1202s | 15,686 ops/s | — | ## After: complete results (post-fix, current) All 54 measurements from the current `bench/bench.c`'s summary block, none omitted or combined: the original 45 (re-run, post-treap-fix) plus 9 new mount-table-scaling measurements (category 10, added after the original run — see "Finding: mount table scaling" above). The `pack_write` dedup fix ("Resolution #2" above) postdates this run but does not change any number in it — see that section for why the compaction row below is unaffected. | Category | Backend | Time | Throughput | MB/s | |---|---|---|---|---| | create 20,000 files | mem | 0.0400s | 499,809 ops/s | 61.0 | | create 20,000 files | raw fs | 0.9943s | 20,114 ops/s | 2.5 | | create 20,000 files | raw+fsync | 116.2147s | 172 ops/s | 0.0 | | read 20,000 files | mem | 0.0099s | 2,011,901 ops/s | 245.6 | | read 20,000 files | raw fs | 0.1606s | 124,543 ops/s | 15.2 | | stat 20,000 files | mem | 0.0077s | 2,598,241 ops/s | — | | stat 20,000 files | raw fs | 0.0699s | 286,222 ops/s | — | | readdir (20,000 entries) | mem | 0.0041s | 4,845,826 /s | — | | readdir (20,000 entries) | raw fs | 0.0069s | 2,894,815 /s | — | | create 20,000 files | dir | 1.1418s | 17,517 ops/s | 2.1 | | read 20,000 files | dir | 0.1530s | 130,705 ops/s | 16.0 | | stat 20,000 files | dir | 0.1379s | 145,056 ops/s | — | | readdir (20,000 entries) | dir | 0.0037s | 5,465,162 /s | — | | unlink 20,000 files | mem | 0.0179s | 1,117,742 ops/s | — | | unlink 20,000 files | dir | 0.4911s | 40,724 ops/s | — | | unlink 20,000 files | raw fs | 0.4746s | 42,141 ops/s | — | | mkdir 4,000 dirs | mem | 0.0056s | 719,696 ops/s | — | | rmdir 4,000 dirs | mem | 0.0040s | 993,724 ops/s | — | | mkdir 4,000 dirs | dir | 0.1710s | 23,386 ops/s | — | | rmdir 4,000 dirs | dir | 0.1257s | 31,827 ops/s | — | | mkdir 4,000 dirs | raw fs | 0.1374s | 29,122 ops/s | — | | rmdir 4,000 dirs | raw fs | 0.0951s | 42,043 ops/s | — | | write 1MB | mem | 0.0001s | 7,058 ops/s | 7,058.1 | | read 1MB | mem | 0.0000s | 50,051 ops/s | 50,050.9 | | write 16MB | mem | 0.0110s | 91 ops/s | 1,458.2 | | read 16MB | mem | 0.0009s | 1,058 ops/s | 16,920.1 | | write 64MB | mem | 0.0586s | 17 ops/s | 1,091.3 | | read 64MB | mem | 0.0032s | 313 ops/s | 20,056.5 | | write 1MB | raw fs | 0.0004s | 2,759 ops/s | 2,759.4 | | read 1MB | raw fs | 0.0001s | 16,812 ops/s | 16,812.4 | | write 16MB | raw fs | 0.0044s | 229 ops/s | 3,656.4 | | read 16MB | raw fs | 0.0010s | 989 ops/s | 15,827.9 | | write 64MB | raw fs | 0.0186s | 54 ops/s | 3,449.7 | | read 64MB | raw fs | 0.0047s | 213 ops/s | 13,633.1 | | write 1MB | raw+fsync | 0.0180s | 56 ops/s | 55.7 | | read 1MB | raw+fsync | 0.0001s | 9,059 ops/s | 9,058.8 | | write 16MB | raw+fsync | 0.0231s | 43 ops/s | 691.7 | | read 16MB | raw+fsync | 0.0012s | 855 ops/s | 13,672.2 | | write 64MB | raw+fsync | 0.0803s | 12 ops/s | 797.2 | | read 64MB | raw+fsync | 0.0048s | 209 ops/s | 13,393.5 | | compact 20,000 entries to pack | pack | 0.0267s | 747,839 ops/s | 91.3 | | random-read 20,000 entries | pack (mmap'd) | 0.0082s | 2,435,930 ops/s | 297.4 | | random-read 20,000 entries | raw fs | 0.1629s | 122,749 ops/s | 15.0 | | concurrent create+read+unlink (8×4,000×3) | mem | 0.2750s | 349,127 ops/s | — | | concurrent create+read+unlink (8×4,000×3) | raw fs | 5.6498s | 16,992 ops/s | — | | mount 500 backends | vfs | 0.0054s | 92,833 ops/s | — | | resolve, 500 mounts | vfs | 0.0020s | 246,064 ops/s | — | | unmount 500 backends | vfs | 0.0046s | 109,207 ops/s | — | | mount 2,000 backends | vfs | 0.0831s | 24,081 ops/s | — | | resolve, 2,000 mounts | vfs | 0.0296s | 67,458 ops/s | — | | unmount 2,000 backends | vfs | 0.0758s | 26,397 ops/s | — | | mount 8,000 backends | vfs | 1.4739s | 5,428 ops/s | — | | resolve, 8,000 mounts | vfs | 0.4938s | 16,200 ops/s | — | | unmount 8,000 backends | vfs | 1.5172s | 5,273 ops/s | — | Everything not involving bulk `mem`/`dir` structural writes (reads, `stat`, `readdir`, large sequential I/O, pack random-access, compaction) is unchanged within normal run-to-run noise, as the row-by-row comparison against the "Before" table shows — the treap rewrite only touches the structural-write path (Section 5.3), not content reads, content writes, or the pack format. The `raw+fsync` create row is unchanged within noise (116.1030s before, 116.2147s here) precisely because it never touches PackFS at all. The nine new mount-table rows have no "Before" counterpart — the mount table was never rewritten, so there is no pre/post comparison to make for them; they are current-state measurements only, discussed in "Finding: mount table scaling" above. ## Analysis The bullets below are checked against the current "After" table (the 54-measurement run) rather than preserved unchanged from an earlier draft — an earlier revision of this section cited exact figures from the prior 45-measurement after-run, which drifted out of sync with the "After" table once that table was replaced with the newer run; the numbers below were recomputed from the table actually in this document, not carried over. ### Where PackFS wins clearly - **Reads of small files: 15–16x faster than both raw fs and the `dir` backend** (2.0M ops/s vs ~125–131K ops/s). No syscall per read — `mem` reads are an expected-O(log n) treap lookup plus a `memcpy` out of an already-resident buffer (Section 5.3, 9.1). - **`stat`: ~9x faster than raw fs** (2.6M ops/s vs 286K ops/s). Same reason — no syscall, and the index is sorted for expected-O(log n) lookup rather than requiring a directory entry scan. - **Random-access reads against a compacted pack: ~20x faster than raw fs** (2.4M ops/s vs 123K ops/s), because the pack is one `mmap`'d file with a binary-searchable index (Section 9.1), against 20,000 individual `open`/`read`/`close` syscall triples on the raw-fs side. This is the single result that most directly validates the architecture's stated purpose — Section 3.2's claim that "random access is a flat table lookup, not a linear scan or a directory-parse-then-seek" — under an actual measured workload, not just by construction. - **Compaction throughput**: ~91 MB/s / ~748K entries/s to serialize the live tree into a fresh pack (Section 4.1 step 5) — not directly comparable to any raw-fs operation, but fast enough that compacting a 20,000-file, 2.5 MB tree is not a practically-felt pause (27ms). - **Bulk create/unlink/mkdir/rmdir: RESOLVED, now 11–590x faster than the pre-fix `mem`/`dir` numbers, and (for `mem`) 21–27x faster than raw fs** — see "Resolution" above; this used to be PackFS's clearest loss and is now among its clearest wins. - **Concurrent mixed create+read+unlink: RESOLVED, now 21x faster than raw fs** (was 6x *slower* before the fix) — see "Resolution." - **`pack_write` compaction dedup on distinct-content files: RESOLVED, 1.75– 39.6x faster depending on N, plus a latent hash-collision correctness bug closed alongside it** — see "Resolution #2." Not visible in this benchmark's own compaction test (category 8), which uses identical content and so never exercised either problem; measured separately by calling `pack_write` directly. ### Where raw fs still wins, and why that's expected, not a bug - **`dir` backend `stat` is ~2.0x *slower* than raw fs `stat`** (145K ops/s vs 286K ops/s) — because `pfs_dir_statat` (Section 6.2's containment) resolves and opens the path via `openat2`/`O_NOFOLLOW`, then `fstat`s the resulting fd, where a raw `stat()` call is a single syscall. This is the direct, measured cost of path containment on the metadata path, separate from and smaller than the containment cost paid on `open` itself (which raw fs pays an equivalent single-syscall cost for anyway). Unaffected by the index rewrite. - **`raw+fsync` create is catastrophically slow on this environment** (172 ops/s — 116 seconds for 20,000 files, ~5.8ms per `fsync`, within measurement noise across the "Before" and "After" runs — 116.1030s vs 116.2147s, a 0.1% difference), because this container's overlay filesystem has poor per-call `fsync` latency. This is a property of the container, not of PackFS or of a bare disk — flagged in the environment section above precisely so it isn't misread as "PackFS's journal must be this slow too." The journal (Section 4.4) does call `fsync` once per journaled write for the same durability reason raw `fsync` is slow here, which is a real, inherited cost on this kind of storage — not a PackFS-specific one. Unaffected by the index rewrite (the journal's cost is per-write `fsync` latency, not index maintenance). - **Large writes above ~16 MB: raw fs is faster than `mem`** (3,450– 3,656 MB/s vs 1,091–1,458 MB/s). Section 5.7's buffer-growth discipline (allocate new, copy old + new, publish, retire old — never realloc in place) means every capacity doubling re-copies everything written so far; a plain unsynced `write()` to a real file only ever appends new pages to the page cache, never re-copying prior ones. This is entirely independent of the index structure (Section 5.7's buffer-growth discipline governs a `MutCell`'s content buffer, not the index the treap rewrite replaced) and is unaffected by this fix — the safety property Section 5.7 requires (no reader can ever see a freed buffer) still has a real, quantifiable cost for very large sequential writes, traded for correctness under concurrent access that a plain in-place realloc would not have. ## Appendix A: complete verbatim output, before the fix Captured exactly as produced by `make bench` against the pre-fix (flat array) index; reformatted into the tables above, but reproduced here unedited as the primary source record. ```text === PackFS vs. host filesystem: benchmark === environment: raw fs test root: /tmp/packfs_bench_raw_unzuwh dir backend root: /tmp/packfs_bench_dir_MXB4CT small files (N): 20000, 128 bytes each directories (N): 4000 concurrency: 8 threads x 4000 ops (create+read+unlink) large-file sizes: 1MB 16MB 64MB NOTE: see the file header for what "raw fs" vs "raw+fsync" vs "dir" actually measure — they are not interchangeable. == 1. create N small files == create N small files mem 9.5565s 2093 ops/s 0.3 MB/s create N small files raw fs 1.0083s 19836 ops/s 2.4 MB/s create N small files raw+fsync 116.1030s 172 ops/s 0.0 MB/s == 2. read N small files == read N small files mem 0.0071s 2798731 ops/s 341.6 MB/s read N small files raw fs 0.1611s 124122 ops/s 15.2 MB/s == 3. stat N files == stat N files mem 0.0062s 3223599 ops/s stat N files raw fs 0.0703s 284465 ops/s == 4. readdir == readdir (N entries) mem 0.0037s 5397828 ops/s readdir (N entries) raw fs 0.0071s 2831904 ops/s == 1b. create N small files (dir backend vs. what it wraps) == create N small files dir 12.0367s 1662 ops/s 0.2 MB/s == 2b. read N small files (dir backend) == read N small files dir 0.1674s 119468 ops/s 14.6 MB/s == 3b. stat N files (dir backend) == stat N files dir 0.1516s 131937 ops/s == 4b. readdir (dir backend) == readdir (N entries) dir 0.0028s 7058003 ops/s == 5. unlink N files == unlink N files mem 10.5575s 1894 ops/s unlink N files dir 10.1851s 1964 ops/s unlink N files raw fs 0.4782s 41823 ops/s == 6. mkdir/rmdir N directories == mkdir N dirs mem 0.3363s 11896 ops/s rmdir N dirs mem 0.3395s 11782 ops/s mkdir N dirs dir 0.5109s 7829 ops/s rmdir N dirs dir 0.4892s 8177 ops/s mkdir N dirs raw fs 0.1591s 25143 ops/s rmdir N dirs raw fs 0.2119s 18876 ops/s == 7. large sequential write/read == write 1MB mem 0.0001s 6687 ops/s 6687.1 MB/s read 1MB mem 0.0000s 50941 ops/s 50941.4 MB/s write 16MB mem 0.0138s 72 ops/s 1157.8 MB/s read 16MB mem 0.0010s 1015 ops/s 16240.8 MB/s write 64MB mem 0.0566s 18 ops/s 1129.8 MB/s read 64MB mem 0.0034s 292 ops/s 18705.5 MB/s write 1MB raw fs 0.0004s 2685 ops/s 2685.4 MB/s read 1MB raw fs 0.0001s 16584 ops/s 16583.7 MB/s write 16MB raw fs 0.0047s 213 ops/s 3412.8 MB/s read 16MB raw fs 0.0010s 960 ops/s 15363.8 MB/s write 64MB raw fs 0.0186s 54 ops/s 3435.6 MB/s read 64MB raw fs 0.0047s 213 ops/s 13635.3 MB/s write 1MB raw+fsync 0.0087s 116 ops/s 115.5 MB/s read 1MB raw+fsync 0.0001s 8857 ops/s 8857.4 MB/s write 16MB raw+fsync 0.0840s 12 ops/s 190.5 MB/s read 16MB raw+fsync 0.0013s 796 ops/s 12737.8 MB/s write 64MB raw+fsync 0.0891s 11 ops/s 718.3 MB/s read 64MB raw+fsync 0.0045s 220 ops/s 14096.3 MB/s == 8. random-access read: mmap'd pack vs. raw fs == compact N entries to pack pack 0.0291s 687097 ops/s 83.9 MB/s random-read N entries pack 0.0085s 2350224 ops/s 286.9 MB/s random-read N entries raw fs 0.1767s 113198 ops/s 13.8 MB/s == 9. concurrency (create+read+unlink) == concurrent create+read+unlink mem 36.7261s 2614 ops/s concurrent create+read+unlink raw fs 6.1202s 15686 ops/s === summary (45 measurements) === category backend seconds throughput MB/s create N small files mem 9.5565s 2093 ops/s 0.3 MB/s create N small files raw fs 1.0083s 19836 ops/s 2.4 MB/s create N small files raw+fsync 116.1030s 172 ops/s 0.0 MB/s read N small files mem 0.0071s 2798731 ops/s 341.6 MB/s read N small files raw fs 0.1611s 124122 ops/s 15.2 MB/s stat N files mem 0.0062s 3223599 ops/s stat N files raw fs 0.0703s 284465 ops/s readdir (N entries) mem 0.0037s 5397828 ops/s readdir (N entries) raw fs 0.0071s 2831904 ops/s create N small files dir 12.0367s 1662 ops/s 0.2 MB/s read N small files dir 0.1674s 119468 ops/s 14.6 MB/s stat N files dir 0.1516s 131937 ops/s readdir (N entries) dir 0.0028s 7058003 ops/s unlink N files mem 10.5575s 1894 ops/s unlink N files dir 10.1851s 1964 ops/s unlink N files raw fs 0.4782s 41823 ops/s mkdir N dirs mem 0.3363s 11896 ops/s rmdir N dirs mem 0.3395s 11782 ops/s mkdir N dirs dir 0.5109s 7829 ops/s rmdir N dirs dir 0.4892s 8177 ops/s mkdir N dirs raw fs 0.1591s 25143 ops/s rmdir N dirs raw fs 0.2119s 18876 ops/s write 1MB mem 0.0001s 6687 ops/s 6687.1 MB/s read 1MB mem 0.0000s 50941 ops/s 50941.4 MB/s write 16MB mem 0.0138s 72 ops/s 1157.8 MB/s read 16MB mem 0.0010s 1015 ops/s 16240.8 MB/s write 64MB mem 0.0566s 18 ops/s 1129.8 MB/s read 64MB mem 0.0034s 292 ops/s 18705.5 MB/s write 1MB raw fs 0.0004s 2685 ops/s 2685.4 MB/s read 1MB raw fs 0.0001s 16584 ops/s 16583.7 MB/s write 16MB raw fs 0.0047s 213 ops/s 3412.8 MB/s read 16MB raw fs 0.0010s 960 ops/s 15363.8 MB/s write 64MB raw fs 0.0186s 54 ops/s 3435.6 MB/s read 64MB raw fs 0.0047s 213 ops/s 13635.3 MB/s write 1MB raw+fsync 0.0087s 116 ops/s 115.5 MB/s read 1MB raw+fsync 0.0001s 8857 ops/s 8857.4 MB/s write 16MB raw+fsync 0.0840s 12 ops/s 190.5 MB/s read 16MB raw+fsync 0.0013s 796 ops/s 12737.8 MB/s write 64MB raw+fsync 0.0891s 11 ops/s 718.3 MB/s read 64MB raw+fsync 0.0045s 220 ops/s 14096.3 MB/s compact N entries to pack pack 0.0291s 687097 ops/s 83.9 MB/s random-read N entries pack 0.0085s 2350224 ops/s 286.9 MB/s random-read N entries raw fs 0.1767s 113198 ops/s 13.8 MB/s concurrent create+read+unlink mem 36.7261s 2614 ops/s concurrent create+read+unlink raw fs 6.1202s 15686 ops/s real 5m49.507s user 2m9.600s sys 0m20.941s ``` ## Appendix B: complete verbatim output, after the fix (current, 54 measurements) Captured exactly as produced by the current `make bench` (post-treap-fix index, with the mount-table-scaling category added); reformatted into the tables above, but reproduced here unedited as the primary source record. This run predates the `pack_write` dedup fix ("Resolution #2"), which does not change any number in it — see that section for why. ```text === PackFS vs. host filesystem: benchmark === environment: raw fs test root: /tmp/packfs_bench_raw_PnMAT7 dir backend root: /tmp/packfs_bench_dir_U8Vgxi small files (N): 20000, 128 bytes each directories (N): 4000 concurrency: 8 threads x 4000 ops (create+read+unlink) mounts (N): 2000 large-file sizes: 1MB 16MB 64MB NOTE: see the file header for what "raw fs" vs "raw+fsync" vs "dir" actually measure — they are not interchangeable. == 1. create N small files == create N small files mem 0.0400s 499809 ops/s 61.0 MB/s create N small files raw fs 0.9943s 20114 ops/s 2.5 MB/s create N small files raw+fsync 116.2147s 172 ops/s 0.0 MB/s == 2. read N small files == read N small files mem 0.0099s 2011901 ops/s 245.6 MB/s read N small files raw fs 0.1606s 124543 ops/s 15.2 MB/s == 3. stat N files == stat N files mem 0.0077s 2598241 ops/s stat N files raw fs 0.0699s 286222 ops/s == 4. readdir == readdir (N entries) mem 0.0041s 4845826 ops/s readdir (N entries) raw fs 0.0069s 2894815 ops/s == 1b. create N small files (dir backend vs. what it wraps) == create N small files dir 1.1418s 17517 ops/s 2.1 MB/s == 2b. read N small files (dir backend) == read N small files dir 0.1530s 130705 ops/s 16.0 MB/s == 3b. stat N files (dir backend) == stat N files dir 0.1379s 145056 ops/s == 4b. readdir (dir backend) == readdir (N entries) dir 0.0037s 5465162 ops/s == 5. unlink N files == unlink N files mem 0.0179s 1117742 ops/s unlink N files dir 0.4911s 40724 ops/s unlink N files raw fs 0.4746s 42141 ops/s == 6. mkdir/rmdir N directories == mkdir N dirs mem 0.0056s 719696 ops/s rmdir N dirs mem 0.0040s 993724 ops/s mkdir N dirs dir 0.1710s 23386 ops/s rmdir N dirs dir 0.1257s 31827 ops/s mkdir N dirs raw fs 0.1374s 29122 ops/s rmdir N dirs raw fs 0.0951s 42043 ops/s == 7. large sequential write/read == write 1MB mem 0.0001s 7058 ops/s 7058.1 MB/s read 1MB mem 0.0000s 50051 ops/s 50050.9 MB/s write 16MB mem 0.0110s 91 ops/s 1458.2 MB/s read 16MB mem 0.0009s 1058 ops/s 16920.1 MB/s write 64MB mem 0.0586s 17 ops/s 1091.3 MB/s read 64MB mem 0.0032s 313 ops/s 20056.5 MB/s write 1MB raw fs 0.0004s 2759 ops/s 2759.4 MB/s read 1MB raw fs 0.0001s 16812 ops/s 16812.4 MB/s write 16MB raw fs 0.0044s 229 ops/s 3656.4 MB/s read 16MB raw fs 0.0010s 989 ops/s 15827.9 MB/s write 64MB raw fs 0.0186s 54 ops/s 3449.7 MB/s read 64MB raw fs 0.0047s 213 ops/s 13633.1 MB/s write 1MB raw+fsync 0.0180s 56 ops/s 55.7 MB/s read 1MB raw+fsync 0.0001s 9059 ops/s 9058.8 MB/s write 16MB raw+fsync 0.0231s 43 ops/s 691.7 MB/s read 16MB raw+fsync 0.0012s 855 ops/s 13672.2 MB/s write 64MB raw+fsync 0.0803s 12 ops/s 797.2 MB/s read 64MB raw+fsync 0.0048s 209 ops/s 13393.5 MB/s == 8. random-access read: mmap'd pack vs. raw fs == compact N entries to pack pack 0.0267s 747839 ops/s 91.3 MB/s random-read N entries pack 0.0082s 2435930 ops/s 297.4 MB/s random-read N entries raw fs 0.1629s 122749 ops/s 15.0 MB/s == 9. concurrency (create+read+unlink) == concurrent create+read+unlink mem 0.2750s 349127 ops/s concurrent create+read+unlink raw fs 5.6498s 16992 ops/s == 10. mount table scaling (the mount table uses the same full-snapshot-copy pattern the file index used to) == mount 500 backends vfs 0.0054s 92833 ops/s resolve, 500 mounts vfs 0.0020s 246064 ops/s unmount 500 backends vfs 0.0046s 109207 ops/s mount 2000 backends vfs 0.0831s 24081 ops/s resolve, 2000 mounts vfs 0.0296s 67458 ops/s unmount 2000 backends vfs 0.0758s 26397 ops/s mount 8000 backends vfs 1.4739s 5428 ops/s resolve, 8000 mounts vfs 0.4938s 16200 ops/s unmount 8000 backends vfs 1.5172s 5273 ops/s === summary (54 measurements) === category backend seconds throughput MB/s create N small files mem 0.0400s 499809 ops/s 61.0 MB/s create N small files raw fs 0.9943s 20114 ops/s 2.5 MB/s create N small files raw+fsync 116.2147s 172 ops/s 0.0 MB/s read N small files mem 0.0099s 2011901 ops/s 245.6 MB/s read N small files raw fs 0.1606s 124543 ops/s 15.2 MB/s stat N files mem 0.0077s 2598241 ops/s stat N files raw fs 0.0699s 286222 ops/s readdir (N entries) mem 0.0041s 4845826 ops/s readdir (N entries) raw fs 0.0069s 2894815 ops/s create N small files dir 1.1418s 17517 ops/s 2.1 MB/s read N small files dir 0.1530s 130705 ops/s 16.0 MB/s stat N files dir 0.1379s 145056 ops/s readdir (N entries) dir 0.0037s 5465162 ops/s unlink N files mem 0.0179s 1117742 ops/s unlink N files dir 0.4911s 40724 ops/s unlink N files raw fs 0.4746s 42141 ops/s mkdir N dirs mem 0.0056s 719696 ops/s rmdir N dirs mem 0.0040s 993724 ops/s mkdir N dirs dir 0.1710s 23386 ops/s rmdir N dirs dir 0.1257s 31827 ops/s mkdir N dirs raw fs 0.1374s 29122 ops/s rmdir N dirs raw fs 0.0951s 42043 ops/s write 1MB mem 0.0001s 7058 ops/s 7058.1 MB/s read 1MB mem 0.0000s 50051 ops/s 50050.9 MB/s write 16MB mem 0.0110s 91 ops/s 1458.2 MB/s read 16MB mem 0.0009s 1058 ops/s 16920.1 MB/s write 64MB mem 0.0586s 17 ops/s 1091.3 MB/s read 64MB mem 0.0032s 313 ops/s 20056.5 MB/s write 1MB raw fs 0.0004s 2759 ops/s 2759.4 MB/s read 1MB raw fs 0.0001s 16812 ops/s 16812.4 MB/s write 16MB raw fs 0.0044s 229 ops/s 3656.4 MB/s read 16MB raw fs 0.0010s 989 ops/s 15827.9 MB/s write 64MB raw fs 0.0186s 54 ops/s 3449.7 MB/s read 64MB raw fs 0.0047s 213 ops/s 13633.1 MB/s write 1MB raw+fsync 0.0180s 56 ops/s 55.7 MB/s read 1MB raw+fsync 0.0001s 9059 ops/s 9058.8 MB/s write 16MB raw+fsync 0.0231s 43 ops/s 691.7 MB/s read 16MB raw+fsync 0.0012s 855 ops/s 13672.2 MB/s write 64MB raw+fsync 0.0803s 12 ops/s 797.2 MB/s read 64MB raw+fsync 0.0048s 209 ops/s 13393.5 MB/s compact N entries to pack pack 0.0267s 747839 ops/s 91.3 MB/s random-read N entries pack 0.0082s 2435930 ops/s 297.4 MB/s random-read N entries raw fs 0.1629s 122749 ops/s 15.0 MB/s concurrent create+read+unlink mem 0.2750s 349127 ops/s concurrent create+read+unlink raw fs 5.6498s 16992 ops/s mount 500 backends vfs 0.0054s 92833 ops/s resolve, 500 mounts vfs 0.0020s 246064 ops/s unmount 500 backends vfs 0.0046s 109207 ops/s mount 2000 backends vfs 0.0831s 24081 ops/s resolve, 2000 mounts vfs 0.0296s 67458 ops/s unmount 2000 backends vfs 0.0758s 26397 ops/s mount 8000 backends vfs 1.4739s 5428 ops/s resolve, 8000 mounts vfs 0.4938s 16200 ops/s unmount 8000 backends vfs 1.5172s 5273 ops/s real 5m23.896s user 0m4.732s sys 0m19.047s ``` ## Reproducing ```sh make bench ``` Takes several minutes on a similarly slow-`fsync` environment, dominated by the `raw+fsync` create test; expect well under a minute on a host with normal disk or tmpfs `fsync` latency.