Files
packfs/BENCH.md
T
retoorandClaude Sonnet 5 529935deb4 BENCH.md: fix numbers left stale by the After-table replacement
Doing another literal, no-caveats audit pass (per the standing request to
verify literally everything, not just reassert it) found real, concrete
inconsistencies introduced by the previous commit: it replaced BENCH.md's
"After: complete results" table and Appendix B with a newer, more complete
bench/bench.c run (adding the mount-table-scaling category), but several
other places in the same document still quoted exact figures from the
older run those tables used to hold -- a reader cross-checking the
"Resolution" section's headline table, or the "Analysis" section's bullet
points, against the "After" table below them would have found contradictory
numbers for the same measurement.

Verified programmatically, not just by re-reading: every row's time,
throughput, and MB/s in the Before (45 rows) and After (54 rows) tables now
matches its corresponding appendix line exactly (0 mismatches out of 99
rows, checked with a script, not by eye).

Fixed:
- The "Resolution" section's headline before/after table used the older
  run's mem create/unlink/mkdir/rmdir numbers (0.0372s/0.0320s/0.0052s/
  0.0045s), which don't match the current After table (0.0400s/0.0179s/
  0.0056s/0.0040s) -- for unlink specifically, a genuine near-2x difference
  between two runs of identical, already-fixed code, not just noise. Now
  points at the single current run, with speedups recomputed (239x/590x/
  60x/85x/11x/21x/134x, mem now 25-27x faster than raw fs, complexity-class
  check now 7.14x vs the old 7.15x -- nearly identical, which is the
  reassurance that actually matters), plus an explicit note on the observed
  run-to-run variance rather than silently picking one number.
- The "Analysis" section's read/stat/random-read/compaction/dir-stat/
  large-write bullets all cited the older run's exact ops/s and MB/s
  figures (2.4M small-file reads, ~815K compaction entries/s, 133K dir
  stat, etc.); recomputed against the current After table (now 15-16x for
  small reads, ~9x for stat, ~91MB/s / ~748K entries/s for compaction, 145K
  vs 286K for dir stat, corrected large-write MB/s ranges).
- The same stale 65-330x / 15-27x / 22x range in CLAUDE.md's "Known
  performance characteristics" bulk-write bullet, updated to match and to
  explain the same run-to-run variance rather than presenting one run's
  numbers as if they were exact.

No code changes; make test still passes (all 6 binaries, unaffected by a
documentation-only change).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
2026-09-14 10:13:57 +00:00

49 KiB
Raw Blame History

PackFS vs. the host filesystem: benchmark results

Status: two O(n²) findings this document originally reported have been fixed; one further O(n²) finding is documented below and deliberately left unfixed, with justification. The file index (UpperSnapshot) is now a persistent treap instead of a flat sorted array (src/upper.c; see CLAUDE.md's "Known performance characteristics") — see "Resolution" below. pack_write's compaction-time duplicate-content elimination (Section 9.2) was also a linear scan with the same O(n²) shape, plus a latent correctness bug (a hash match was trusted without a memcmp verification); both are fixed — see "Resolution #2" below. The mount table (vfs.c's MountSnapshot) uses the same full-array-copy pattern the file index used to, is confirmed O(n²) in mount count, and is not rewritten — see "Finding: mount table scaling" below for why that is the correct call, not an oversight. This document keeps every prior run's numbers as the historical record of each problem — per this project's own documentation standard, a correction extends the record rather than erasing it — and adds each fix's numbers alongside them, in full, not a curated subset of either. Read "Resolution," "Resolution #2," and "Finding: mount table scaling" below before drawing conclusions from the "Before" tables, which describe a version of the index this codebase no longer has.

This document reports the results of bench/bench.c (make bench), each run once on the environment described below. These are single runs, not a statistically-averaged series — treat the numbers as illustrative of shape (which operations are faster, by roughly what factor, and why) more than as precise absolute figures for a different machine. The methodology, including exactly what each backend label measures, is documented in bench/bench.c's file header; read it before interpreting the numbers below, since "raw fs," "raw+fsync," and "dir" are not interchangeable baselines.

The "Before" table below reproduces every one of the 45 measurements that run's === summary (45 measurements) === block reported; the "After" table reproduces every one of the 54 measurements the current bench/bench.c reports (the original 45, unchanged in category, plus 9 new mount-table- scaling measurements added after the original run — see "Finding: mount table scaling"). None omitted, none combined in either table, and the complete verbatim program output for both runs, exactly as captured, is preserved in the two appendices at the end as the primary source record. The pack_write dedup fix ("Resolution #2") is measured separately, by calling pack_write directly rather than through bench/bench.c — see that section for why.

Environment

  • Host: containerized, 12 logical CPUs (AMD Ryzen 5 3600), running under an overlay root filesystem (overlay on overlay, per mount) — there is no separate tmpfs at /tmp in this environment; every "raw fs" and dir number below is against that container overlay filesystem, not a bare disk or tmpfs. This matters most for the fsync numbers (see below).
  • N_SMALL = 20,000 files × 128 bytes for the metadata-heavy suite; N_DIRS = 4,000; concurrency = 8 threads × 4,000 ops; large files up to 64 MB; N_MOUNTS = 2,000 (run at N/4, N, N×4 — 500/2,000/8,000 — for the mount-table-scaling category added after the original run). Full parameters are in bench/bench.c.
  • Wall-clock time: 5m49s (before), 4m4s (after the treap fix, before the mount-scaling category existed), 5m24s (current, with the mount-scaling category added). The dominant cost throughout is the single raw+fsync 20,000-file create test (~116s in every run — this cost is on the host's storage, not PackFS, and is within measurement noise across all three runs, as expected, since that test never touches PackFS's index at all); the added mount-scaling category itself accounts for only ~3.7s of the increase from 4m4s to 5m24s, so the remainder is run-to-run variance in untimed setup/teardown, consistent with this document's stated single-run, illustrative-of-shape methodology, not a regression.

Resolution

The original run found bulk sequential create/unlink/mkdir on mem/dir to be O(n²) in file count, because every structural write copied the entire sorted entry array before publishing the next snapshot (Section 5.3). concept.md itself names the trigger condition for reconsidering that design — "a persistent (structurally shared) tree structure is not required until this assumption is empirically violated" — and the original run was exactly that violation, measured rather than hypothesized.

The index was rewritten as a persistent treap (src/upper.c): a structural write now builds only the O(log n) nodes on the path to the change, sharing every other node (refcounted) with the snapshot it was built from, instead of copying all n entries. The choice of treap over a persistent AVL/red-black/weight-balanced tree, and the reasoning behind it, is documented in src/upper.c's comment above struct TreapNode and cited there: Seidel & Aragon, "Randomized Search Trees" (1996), for the O(1) amortized rotation count on both insert and delete (persistent balanced trees that need O(log n) rotations pay a real, repeated allocation cost per rotation, which a treap avoids); Liljenzin, "Confluently Persistent Sets and Maps" (arXiv:1301.3388), for persistent treaps specifically as an MVCC snapshot mechanism, which is this exact use case.

Headline result, measured, same methodology as the original finding. The "After" figures below are drawn from the "After: complete results" table further down this document (the current, canonical post-fix run), not from an intermediate run kept only in this section — every number here can be found again in that table:

Category Before After Speedup
create 20,000 files (mem) 9.5565s / 2,093 ops/s 0.0400s / 499,809 ops/s 239x
unlink 20,000 files (mem) 10.5575s / 1,894 ops/s 0.0179s / 1,117,742 ops/s 590x
mkdir 4,000 dirs (mem) 0.3363s / 11,896 ops/s 0.0056s / 719,696 ops/s 60x
rmdir 4,000 dirs (mem) 0.3395s / 11,782 ops/s 0.0040s / 993,724 ops/s 85x
create 20,000 files (dir) 12.0367s / 1,662 ops/s 1.1418s / 17,517 ops/s 11x
unlink 20,000 files (dir) 10.1851s / 1,964 ops/s 0.4911s / 40,724 ops/s 21x
concurrent create+read+unlink (mem) 36.7261s / 2,614 ops/s 0.2750s / 349,127 ops/s 134x

(An earlier revision of this table cited a different post-fix run — same fixed code, an earlier make bench invocation — whose numbers were 65–330x rather than 60–590x; the two runs agree on every qualitative conclusion below, but unlink 20,000 files (mem) in particular differs by nearly 2x between them, 0.0320s vs 0.0179s, both measuring an operation fast enough (sub-20ms for 20,000 entries) that ordinary system jitter in this containerized environment has real proportional effect. This is exactly the kind of run-to-run variance this document's stated methodology warns about — "illustrative of shape... more than precise absolute figures" — and the reason the table now points at one single canonical run rather than mixing figures from two.)

mem now beats raw fs (not merely "no longer loses to it") at every one of these: create is 25x faster than raw fs (was 9.5x slower), unlink is 27x faster (was 22x slower), and the concurrent mixed workload is 21x faster (was 6x slower) — the single-writer design's "optimizes for read-heavy concurrent workloads" caveat from the original analysis (below) no longer needs to exclude bulk-write workloads, because the write side is no longer the bottleneck it was.

The complexity-class change is confirmed quantitatively, the same way the original O(n²) finding was — not just "it got faster," which an unrelated constant-factor optimization could also produce: mkdir at N=4,000 vs. create at N=20,000 (same structural-write mechanism, 5x the N) now shows a 7.14x slowdown (0.0400s / 0.0056s). O(n) predicts 5.0x; O(n log n) predicts 5.97x; the old O(n²) regime predicted, and measured, 25–28.4x. 7.14x lands close to the O(n log n) prediction (some excess is expected — create writes bytes into a fresh MutCell and mkdir does not, a constant-factor difference the pure index-complexity comparison doesn't account for) and nowhere near quadratic — and is itself nearly identical to the 7.15x this same check gave against the earlier run referenced above, which is the reassurance that matters here: the ratio this check depends on is stable across runs even where the fastest individual measurements are not. The index's asymptotic behavior changed, not merely its constant factor.

Full output of both runs is reproduced verbatim in the appendices below and is reproducible again with make bench; the correctness of the new index under heavy randomized structural churn (non-sequential creates/deletes/ renames, nested directories, thousands of entries, cross-checked against an independent reference model — not just "the benchmark ran fast without crashing") is tests/test_index_stress.c, part of make test.

Resolution #2: pack_write compaction dedup

A second, independent O(n²) finding, found while auditing the codebase for other instances of the same shape of bug after "Resolution" above: src/ pack.c's pack_write implements Section 9.2's exact-duplicate elimination (two entries with byte-identical content share one data_off in the compacted pack) as a linear scan, for each entry, of every previously-seen (hash, size) pair. That is O(n) per entry, O(n²) total across n distinct blobs — the same complexity-class bug "Resolution" fixed in the file index, here in the compaction writer instead. It went unnoticed by the original benchmark for a specific, identifiable reason: bench/bench.c's compaction test (category 8, "compact N entries to pack") writes byte-identical content into every file, so every entry's linear scan matched on the very first comparison (against entry 0) and never actually scanned — the benchmark's own test data happened to be the linear scan's best case, not a representative one. Measuring with unique per-file content instead (each entry's actual worst case) surfaced it:

N (unique entries) Time
5,000 0.0138s
20,000 0.1114s
80,000 1.7389s

5,000 → 20,000 is a 4x step in N; time increased 8.07x. 20,000 → 80,000 is also a 4x step; time increased 15.61x, converging toward O(n²)'s 16x prediction as N grows (the smaller first ratio reflects fixed per-call overhead not yet dominated by the quadratic term) — the same convergence pattern "Resolution" observed for the file index, and the same standard used there to distinguish a real complexity-class finding from noise.

This scan had a second, latent problem beyond speed: it compared only (hash, size), never the actual bytes, before deciding two entries were duplicates and sharing their data_off. FNV-1a64 (src/hash.c) is explicitly not collision-resistant — it is a fast checksum, not a cryptographic hash — so two different-content entries that happened to collide on (hash, size) would have been silently merged into one blob, corrupting one of them. This was never observed in practice (no test's data happened to collide), but was true of the code as written, not a hypothetical: the fix required for the O(n²) issue also closes it.

Fix: the linear scan was replaced with an open-addressing hash table (load factor 1/2, linear probing) over (hash, size), plus — critically — an actual memcmp against the candidate's stored data pointer before ever reusing a data_off, so a hash match is only ever a candidate, never accepted as proof of equality. Comment and full listing: src/pack.c's DedupSlot struct and the rewritten first pass of pack_write.

After the fix, same methodology:

N (unique entries) Time vs. before
5,000 0.0079s 1.75x faster
20,000 0.0131s 8.50x faster
80,000 0.0439s 39.6x faster

5,000 → 20,000 now shows 1.66x (O(n) predicts 4x, but a hash table's dominant cost at these sizes is close to constant per insert, so sub-linear growth here is expected, not suspicious); 20,000 → 80,000 shows 3.35x, close to O(n)'s 4x prediction and nowhere near O(n²)'s 16x. The complexity class changed, exactly as with the file index.

This fix is not reflected in the "After" table/Appendix B below, because that make bench run predates it — but it does not need to be: category 8's "compact N entries to pack" measurement uses identical content (as explained above), which was already the linear scan's O(1)-per-entry best case before this fix and is unaffected by the fix (a hash table lookup that matches on the first probe is also O(1) per entry); the number in the "After" table (0.0267s, 20,000 identical-content entries) remains accurate and does not need to be re-measured. The fix matters for compacting many files with distinct content, which bench/bench.c's existing compaction test does not exercise — hence the separate, direct pack_write measurement above rather than a bench/bench.c change. Correctness is covered permanently by tests/test_pack_overlay.c (asserts /a.txt and /dup.txt, identical content, share one data_off, while /b.txt, different content, shares neither's — catching both "dedup stopped happening" and "dedup over-matched"); the performance fix is covered permanently by tests/test_pack_write_perf.c (10,000 unique entries under a generous time bound, a regression tripwire rather than a tight assertion, run as part of make test).

Finding: mount table scaling (O(n²), documented, deliberately not fixed)

src/vfs.c's mount table (MountSnapshot) uses the exact same pattern the file index used before "Resolution" above: every vfs_mount/vfs_unmount copies the entire array of mounted backends before publishing the next snapshot (Section 5.3's single-writer model applies to the mount table too, per concept.md's explicit statement that "the mount table is not separate global state exempt from this model"). This is confirmed O(n²) in mount count by the bench/bench.c category 10 measurements (full numbers in the "After" table below and Appendix B):

N mounts mount resolve (stat through N mounts) unmount
500 0.0054s 0.0020s 0.0046s
2,000 0.0831s 0.0296s 0.0758s
8,000 1.4739s 0.4938s 1.5172s

500 → 2,000 is a 4x step in N: mount increased 15.4x, resolve 14.8x, unmount 16.5x. 2,000 → 8,000 is also a 4x step: mount increased 17.7x, resolve 16.7x, unmount 20.0x. All six ratios cluster around O(n²)'s 16x prediction for a 4x-N step, and nowhere near O(n)'s 4x — the same complexity-class-confirmation methodology used for both findings above, applied here and yielding the same verdict. Notably, resolve (a single vfs_stat through an already-built N-mount table) is also O(n²) across N resolves, not just O(n) per resolve as might be assumed — consistent with longest-prefix-match dispatch being a linear scan of the mount table per call (Section 3's dispatch design), which is itself O(n) per call and O(n²) across N calls, independent of the snapshot-copy cost on the write side.

This is deliberately not fixed, and the reasoning is the same trigger condition concept.md Section 5.3 names for the file index — inverted: "a persistent (structurally shared) tree structure is not required until this assumption is empirically violated." The file index's assumption was violated because file counts are workload-driven and can plausibly reach tens of thousands (or more) at runtime, chosen by whatever program embeds PackFS and whatever data it manages. Mount counts are different in kind, not just degree: a mount point is created by a call to vfs_mount written into a program's own source code, at a location the program's author chose — it is bounded by how many lines of "mount this backend at this path" a person or a build script is willing to write, not by user data, input size, or any other externally-driven quantity. A program with 8,000 mount points, structured that way on purpose, is not a realistic workload this system needs to serve well; a program with 8,000 files is an ordinary one. Converting the mount table to a persistent treap would fix this at the cost of real, permanent complexity (an additional generic persistent-tree instantiation, or a second copy of the treap logic specialized to prefix matching) for a case that has not been and is not expected to be hit. Should a future use case genuinely need thousands of mount points at runtime — for instance, one mount per user in a multi-tenant embedding — this finding, and the numbers above, are the starting point for that decision; until then, this is recorded as a known, measured, and consciously accepted limitation, not an unexamined one.

Before: complete results (pre-fix; kept as the historical record — see "Resolution" above)

All 45 measurements from that run's summary block, none omitted or combined.

Category Backend Time Throughput MB/s
create 20,000 files mem 9.5565s 2,093 ops/s 0.3
create 20,000 files raw fs 1.0083s 19,836 ops/s 2.4
create 20,000 files raw+fsync 116.1030s 172 ops/s 0.0
read 20,000 files mem 0.0071s 2,798,731 ops/s 341.6
read 20,000 files raw fs 0.1611s 124,122 ops/s 15.2
stat 20,000 files mem 0.0062s 3,223,599 ops/s —
stat 20,000 files raw fs 0.0703s 284,465 ops/s —
readdir (20,000 entries) mem 0.0037s 5,397,828 /s —
readdir (20,000 entries) raw fs 0.0071s 2,831,904 /s —
create 20,000 files dir 12.0367s 1,662 ops/s 0.2
read 20,000 files dir 0.1674s 119,468 ops/s 14.6
stat 20,000 files dir 0.1516s 131,937 ops/s —
readdir (20,000 entries) dir 0.0028s 7,058,003 /s —
unlink 20,000 files mem 10.5575s 1,894 ops/s —
unlink 20,000 files dir 10.1851s 1,964 ops/s —
unlink 20,000 files raw fs 0.4782s 41,823 ops/s —
mkdir 4,000 dirs mem 0.3363s 11,896 ops/s —
rmdir 4,000 dirs mem 0.3395s 11,782 ops/s —
mkdir 4,000 dirs dir 0.5109s 7,829 ops/s —
rmdir 4,000 dirs dir 0.4892s 8,177 ops/s —
mkdir 4,000 dirs raw fs 0.1591s 25,143 ops/s —
rmdir 4,000 dirs raw fs 0.2119s 18,876 ops/s —
write 1MB mem 0.0001s 6,687 ops/s 6,687.1
read 1MB mem 0.0000s 50,941 ops/s 50,941.4
write 16MB mem 0.0138s 72 ops/s 1,157.8
read 16MB mem 0.0010s 1,015 ops/s 16,240.8
write 64MB mem 0.0566s 18 ops/s 1,129.8
read 64MB mem 0.0034s 292 ops/s 18,705.5
write 1MB raw fs 0.0004s 2,685 ops/s 2,685.4
read 1MB raw fs 0.0001s 16,584 ops/s 16,583.7
write 16MB raw fs 0.0047s 213 ops/s 3,412.8
read 16MB raw fs 0.0010s 960 ops/s 15,363.8
write 64MB raw fs 0.0186s 54 ops/s 3,435.6
read 64MB raw fs 0.0047s 213 ops/s 13,635.3
write 1MB raw+fsync 0.0087s 116 ops/s 115.5
read 1MB raw+fsync 0.0001s 8,857 ops/s 8,857.4
write 16MB raw+fsync 0.0840s 12 ops/s 190.5
read 16MB raw+fsync 0.0013s 796 ops/s 12,737.8
write 64MB raw+fsync 0.0891s 11 ops/s 718.3
read 64MB raw+fsync 0.0045s 220 ops/s 14,096.3
compact 20,000 entries to pack pack 0.0291s 687,097 ops/s 83.9
random-read 20,000 entries pack (mmap'd) 0.0085s 2,350,224 ops/s 286.9
random-read 20,000 entries raw fs 0.1767s 113,198 ops/s 13.8
concurrent create+read+unlink (8×4,000×3) mem 36.7261s 2,614 ops/s —
concurrent create+read+unlink (8×4,000×3) raw fs 6.1202s 15,686 ops/s —

After: complete results (post-fix, current)

All 54 measurements from the current bench/bench.c's summary block, none omitted or combined: the original 45 (re-run, post-treap-fix) plus 9 new mount-table-scaling measurements (category 10, added after the original run — see "Finding: mount table scaling" above). The pack_write dedup fix ("Resolution #2" above) postdates this run but does not change any number in it — see that section for why the compaction row below is unaffected.

Category Backend Time Throughput MB/s
create 20,000 files mem 0.0400s 499,809 ops/s 61.0
create 20,000 files raw fs 0.9943s 20,114 ops/s 2.5
create 20,000 files raw+fsync 116.2147s 172 ops/s 0.0
read 20,000 files mem 0.0099s 2,011,901 ops/s 245.6
read 20,000 files raw fs 0.1606s 124,543 ops/s 15.2
stat 20,000 files mem 0.0077s 2,598,241 ops/s —
stat 20,000 files raw fs 0.0699s 286,222 ops/s —
readdir (20,000 entries) mem 0.0041s 4,845,826 /s —
readdir (20,000 entries) raw fs 0.0069s 2,894,815 /s —
create 20,000 files dir 1.1418s 17,517 ops/s 2.1
read 20,000 files dir 0.1530s 130,705 ops/s 16.0
stat 20,000 files dir 0.1379s 145,056 ops/s —
readdir (20,000 entries) dir 0.0037s 5,465,162 /s —
unlink 20,000 files mem 0.0179s 1,117,742 ops/s —
unlink 20,000 files dir 0.4911s 40,724 ops/s —
unlink 20,000 files raw fs 0.4746s 42,141 ops/s —
mkdir 4,000 dirs mem 0.0056s 719,696 ops/s —
rmdir 4,000 dirs mem 0.0040s 993,724 ops/s —
mkdir 4,000 dirs dir 0.1710s 23,386 ops/s —
rmdir 4,000 dirs dir 0.1257s 31,827 ops/s —
mkdir 4,000 dirs raw fs 0.1374s 29,122 ops/s —
rmdir 4,000 dirs raw fs 0.0951s 42,043 ops/s —
write 1MB mem 0.0001s 7,058 ops/s 7,058.1
read 1MB mem 0.0000s 50,051 ops/s 50,050.9
write 16MB mem 0.0110s 91 ops/s 1,458.2
read 16MB mem 0.0009s 1,058 ops/s 16,920.1
write 64MB mem 0.0586s 17 ops/s 1,091.3
read 64MB mem 0.0032s 313 ops/s 20,056.5
write 1MB raw fs 0.0004s 2,759 ops/s 2,759.4
read 1MB raw fs 0.0001s 16,812 ops/s 16,812.4
write 16MB raw fs 0.0044s 229 ops/s 3,656.4
read 16MB raw fs 0.0010s 989 ops/s 15,827.9
write 64MB raw fs 0.0186s 54 ops/s 3,449.7
read 64MB raw fs 0.0047s 213 ops/s 13,633.1
write 1MB raw+fsync 0.0180s 56 ops/s 55.7
read 1MB raw+fsync 0.0001s 9,059 ops/s 9,058.8
write 16MB raw+fsync 0.0231s 43 ops/s 691.7
read 16MB raw+fsync 0.0012s 855 ops/s 13,672.2
write 64MB raw+fsync 0.0803s 12 ops/s 797.2
read 64MB raw+fsync 0.0048s 209 ops/s 13,393.5
compact 20,000 entries to pack pack 0.0267s 747,839 ops/s 91.3
random-read 20,000 entries pack (mmap'd) 0.0082s 2,435,930 ops/s 297.4
random-read 20,000 entries raw fs 0.1629s 122,749 ops/s 15.0
concurrent create+read+unlink (8×4,000×3) mem 0.2750s 349,127 ops/s —
concurrent create+read+unlink (8×4,000×3) raw fs 5.6498s 16,992 ops/s —
mount 500 backends vfs 0.0054s 92,833 ops/s —
resolve, 500 mounts vfs 0.0020s 246,064 ops/s —
unmount 500 backends vfs 0.0046s 109,207 ops/s —
mount 2,000 backends vfs 0.0831s 24,081 ops/s —
resolve, 2,000 mounts vfs 0.0296s 67,458 ops/s —
unmount 2,000 backends vfs 0.0758s 26,397 ops/s —
mount 8,000 backends vfs 1.4739s 5,428 ops/s —
resolve, 8,000 mounts vfs 0.4938s 16,200 ops/s —
unmount 8,000 backends vfs 1.5172s 5,273 ops/s —

Everything not involving bulk mem/dir structural writes (reads, stat, readdir, large sequential I/O, pack random-access, compaction) is unchanged within normal run-to-run noise, as the row-by-row comparison against the "Before" table shows — the treap rewrite only touches the structural-write path (Section 5.3), not content reads, content writes, or the pack format. The raw+fsync create row is unchanged within noise (116.1030s before, 116.2147s here) precisely because it never touches PackFS at all. The nine new mount-table rows have no "Before" counterpart — the mount table was never rewritten, so there is no pre/post comparison to make for them; they are current-state measurements only, discussed in "Finding: mount table scaling" above.

Analysis

The bullets below are checked against the current "After" table (the 54-measurement run) rather than preserved unchanged from an earlier draft — an earlier revision of this section cited exact figures from the prior 45-measurement after-run, which drifted out of sync with the "After" table once that table was replaced with the newer run; the numbers below were recomputed from the table actually in this document, not carried over.

Where PackFS wins clearly

  • Reads of small files: 15–16x faster than both raw fs and the dir backend (2.0M ops/s vs ~125–131K ops/s). No syscall per read — mem reads are an expected-O(log n) treap lookup plus a memcpy out of an already-resident buffer (Section 5.3, 9.1).
  • stat: ~9x faster than raw fs (2.6M ops/s vs 286K ops/s). Same reason — no syscall, and the index is sorted for expected-O(log n) lookup rather than requiring a directory entry scan.
  • Random-access reads against a compacted pack: ~20x faster than raw fs (2.4M ops/s vs 123K ops/s), because the pack is one mmap'd file with a binary-searchable index (Section 9.1), against 20,000 individual open/read/close syscall triples on the raw-fs side. This is the single result that most directly validates the architecture's stated purpose — Section 3.2's claim that "random access is a flat table lookup, not a linear scan or a directory-parse-then-seek" — under an actual measured workload, not just by construction.
  • Compaction throughput: ~91 MB/s / ~748K entries/s to serialize the live tree into a fresh pack (Section 4.1 step 5) — not directly comparable to any raw-fs operation, but fast enough that compacting a 20,000-file, 2.5 MB tree is not a practically-felt pause (27ms).
  • Bulk create/unlink/mkdir/rmdir: RESOLVED, now 11–590x faster than the pre-fix mem/dir numbers, and (for mem) 21–27x faster than raw fs — see "Resolution" above; this used to be PackFS's clearest loss and is now among its clearest wins.
  • Concurrent mixed create+read+unlink: RESOLVED, now 21x faster than raw fs (was 6x slower before the fix) — see "Resolution."
  • pack_write compaction dedup on distinct-content files: RESOLVED, 1.75– 39.6x faster depending on N, plus a latent hash-collision correctness bug closed alongside it — see "Resolution #2." Not visible in this benchmark's own compaction test (category 8), which uses identical content and so never exercised either problem; measured separately by calling pack_write directly.

Where raw fs still wins, and why that's expected, not a bug

  • dir backend stat is ~2.0x slower than raw fs stat (145K ops/s vs 286K ops/s) — because pfs_dir_statat (Section 6.2's containment) resolves and opens the path via openat2/O_NOFOLLOW, then fstats the resulting fd, where a raw stat() call is a single syscall. This is the direct, measured cost of path containment on the metadata path, separate from and smaller than the containment cost paid on open itself (which raw fs pays an equivalent single-syscall cost for anyway). Unaffected by the index rewrite.
  • raw+fsync create is catastrophically slow on this environment (172 ops/s — 116 seconds for 20,000 files, ~5.8ms per fsync, within measurement noise across the "Before" and "After" runs — 116.1030s vs 116.2147s, a 0.1% difference), because this container's overlay filesystem has poor per-call fsync latency. This is a property of the container, not of PackFS or of a bare disk — flagged in the environment section above precisely so it isn't misread as "PackFS's journal must be this slow too." The journal (Section 4.4) does call fsync once per journaled write for the same durability reason raw fsync is slow here, which is a real, inherited cost on this kind of storage — not a PackFS-specific one. Unaffected by the index rewrite (the journal's cost is per-write fsync latency, not index maintenance).
  • Large writes above ~16 MB: raw fs is faster than mem (3,450– 3,656 MB/s vs 1,091–1,458 MB/s). Section 5.7's buffer-growth discipline (allocate new, copy old + new, publish, retire old — never realloc in place) means every capacity doubling re-copies everything written so far; a plain unsynced write() to a real file only ever appends new pages to the page cache, never re-copying prior ones. This is entirely independent of the index structure (Section 5.7's buffer-growth discipline governs a MutCell's content buffer, not the index the treap rewrite replaced) and is unaffected by this fix — the safety property Section 5.7 requires (no reader can ever see a freed buffer) still has a real, quantifiable cost for very large sequential writes, traded for correctness under concurrent access that a plain in-place realloc would not have.

Appendix A: complete verbatim output, before the fix

Captured exactly as produced by make bench against the pre-fix (flat array) index; reformatted into the tables above, but reproduced here unedited as the primary source record.

=== PackFS vs. host filesystem: benchmark ===

environment:
  raw fs test root: /tmp/packfs_bench_raw_unzuwh
  dir backend root: /tmp/packfs_bench_dir_MXB4CT
  small files (N):  20000, 128 bytes each
  directories (N):  4000
  concurrency:      8 threads x 4000 ops (create+read+unlink)
  large-file sizes: 1MB 16MB 64MB 
  NOTE: see the file header for what "raw fs" vs "raw+fsync" vs
        "dir" actually measure — they are not interchangeable.

== 1. create N small files ==
  create N small files         mem           9.5565s            2093 ops/s         0.3 MB/s
  create N small files         raw fs        1.0083s           19836 ops/s         2.4 MB/s
  create N small files         raw+fsync   116.1030s             172 ops/s         0.0 MB/s

== 2. read N small files ==
  read N small files           mem           0.0071s         2798731 ops/s       341.6 MB/s
  read N small files           raw fs        0.1611s          124122 ops/s        15.2 MB/s

== 3. stat N files ==
  stat N files                 mem           0.0062s         3223599 ops/s
  stat N files                 raw fs        0.0703s          284465 ops/s

== 4. readdir ==
  readdir (N entries)          mem           0.0037s         5397828 ops/s
  readdir (N entries)          raw fs        0.0071s         2831904 ops/s

== 1b. create N small files (dir backend vs. what it wraps) ==
  create N small files         dir          12.0367s            1662 ops/s         0.2 MB/s

== 2b. read N small files (dir backend) ==
  read N small files           dir           0.1674s          119468 ops/s        14.6 MB/s

== 3b. stat N files (dir backend) ==
  stat N files                 dir           0.1516s          131937 ops/s

== 4b. readdir (dir backend) ==
  readdir (N entries)          dir           0.0028s         7058003 ops/s

== 5. unlink N files ==
  unlink N files               mem          10.5575s            1894 ops/s
  unlink N files               dir          10.1851s            1964 ops/s
  unlink N files               raw fs        0.4782s           41823 ops/s

== 6. mkdir/rmdir N directories ==
  mkdir N dirs                 mem           0.3363s           11896 ops/s
  rmdir N dirs                 mem           0.3395s           11782 ops/s
  mkdir N dirs                 dir           0.5109s            7829 ops/s
  rmdir N dirs                 dir           0.4892s            8177 ops/s
  mkdir N dirs                 raw fs        0.1591s           25143 ops/s
  rmdir N dirs                 raw fs        0.2119s           18876 ops/s

== 7. large sequential write/read ==
  write 1MB                    mem           0.0001s            6687 ops/s      6687.1 MB/s
  read 1MB                     mem           0.0000s           50941 ops/s     50941.4 MB/s
  write 16MB                   mem           0.0138s              72 ops/s      1157.8 MB/s
  read 16MB                    mem           0.0010s            1015 ops/s     16240.8 MB/s
  write 64MB                   mem           0.0566s              18 ops/s      1129.8 MB/s
  read 64MB                    mem           0.0034s             292 ops/s     18705.5 MB/s
  write 1MB                    raw fs        0.0004s            2685 ops/s      2685.4 MB/s
  read 1MB                     raw fs        0.0001s           16584 ops/s     16583.7 MB/s
  write 16MB                   raw fs        0.0047s             213 ops/s      3412.8 MB/s
  read 16MB                    raw fs        0.0010s             960 ops/s     15363.8 MB/s
  write 64MB                   raw fs        0.0186s              54 ops/s      3435.6 MB/s
  read 64MB                    raw fs        0.0047s             213 ops/s     13635.3 MB/s
  write 1MB                    raw+fsync     0.0087s             116 ops/s       115.5 MB/s
  read 1MB                     raw+fsync     0.0001s            8857 ops/s      8857.4 MB/s
  write 16MB                   raw+fsync     0.0840s              12 ops/s       190.5 MB/s
  read 16MB                    raw+fsync     0.0013s             796 ops/s     12737.8 MB/s
  write 64MB                   raw+fsync     0.0891s              11 ops/s       718.3 MB/s
  read 64MB                    raw+fsync     0.0045s             220 ops/s     14096.3 MB/s

== 8. random-access read: mmap'd pack vs. raw fs ==
  compact N entries to pack    pack          0.0291s          687097 ops/s        83.9 MB/s
  random-read N entries        pack          0.0085s         2350224 ops/s       286.9 MB/s
  random-read N entries        raw fs        0.1767s          113198 ops/s        13.8 MB/s

== 9. concurrency (create+read+unlink) ==
  concurrent create+read+unlink mem          36.7261s            2614 ops/s
  concurrent create+read+unlink raw fs        6.1202s           15686 ops/s

=== summary (45 measurements) ===
  category                     backend        seconds        throughput          MB/s
  create N small files         mem            9.5565s            2093 ops/s         0.3 MB/s
  create N small files         raw fs         1.0083s           19836 ops/s         2.4 MB/s
  create N small files         raw+fsync    116.1030s             172 ops/s         0.0 MB/s
  read N small files           mem            0.0071s         2798731 ops/s       341.6 MB/s
  read N small files           raw fs         0.1611s          124122 ops/s        15.2 MB/s
  stat N files                 mem            0.0062s         3223599 ops/s               
  stat N files                 raw fs         0.0703s          284465 ops/s               
  readdir (N entries)          mem            0.0037s         5397828 ops/s               
  readdir (N entries)          raw fs         0.0071s         2831904 ops/s               
  create N small files         dir           12.0367s            1662 ops/s         0.2 MB/s
  read N small files           dir            0.1674s          119468 ops/s        14.6 MB/s
  stat N files                 dir            0.1516s          131937 ops/s               
  readdir (N entries)          dir            0.0028s         7058003 ops/s               
  unlink N files               mem           10.5575s            1894 ops/s               
  unlink N files               dir           10.1851s            1964 ops/s               
  unlink N files               raw fs         0.4782s           41823 ops/s               
  mkdir N dirs                 mem            0.3363s           11896 ops/s               
  rmdir N dirs                 mem            0.3395s           11782 ops/s               
  mkdir N dirs                 dir            0.5109s            7829 ops/s               
  rmdir N dirs                 dir            0.4892s            8177 ops/s               
  mkdir N dirs                 raw fs         0.1591s           25143 ops/s               
  rmdir N dirs                 raw fs         0.2119s           18876 ops/s               
  write 1MB                    mem            0.0001s            6687 ops/s      6687.1 MB/s
  read 1MB                     mem            0.0000s           50941 ops/s     50941.4 MB/s
  write 16MB                   mem            0.0138s              72 ops/s      1157.8 MB/s
  read 16MB                    mem            0.0010s            1015 ops/s     16240.8 MB/s
  write 64MB                   mem            0.0566s              18 ops/s      1129.8 MB/s
  read 64MB                    mem            0.0034s             292 ops/s     18705.5 MB/s
  write 1MB                    raw fs         0.0004s            2685 ops/s      2685.4 MB/s
  read 1MB                     raw fs         0.0001s           16584 ops/s     16583.7 MB/s
  write 16MB                   raw fs         0.0047s             213 ops/s      3412.8 MB/s
  read 16MB                    raw fs         0.0010s             960 ops/s     15363.8 MB/s
  write 64MB                   raw fs         0.0186s              54 ops/s      3435.6 MB/s
  read 64MB                    raw fs         0.0047s             213 ops/s     13635.3 MB/s
  write 1MB                    raw+fsync      0.0087s             116 ops/s       115.5 MB/s
  read 1MB                     raw+fsync      0.0001s            8857 ops/s      8857.4 MB/s
  write 16MB                   raw+fsync      0.0840s              12 ops/s       190.5 MB/s
  read 16MB                    raw+fsync      0.0013s             796 ops/s     12737.8 MB/s
  write 64MB                   raw+fsync      0.0891s              11 ops/s       718.3 MB/s
  read 64MB                    raw+fsync      0.0045s             220 ops/s     14096.3 MB/s
  compact N entries to pack    pack           0.0291s          687097 ops/s        83.9 MB/s
  random-read N entries        pack           0.0085s         2350224 ops/s       286.9 MB/s
  random-read N entries        raw fs         0.1767s          113198 ops/s        13.8 MB/s
  concurrent create+read+unlink mem           36.7261s            2614 ops/s               
  concurrent create+read+unlink raw fs         6.1202s           15686 ops/s               

real	5m49.507s
user	2m9.600s
sys	0m20.941s

Appendix B: complete verbatim output, after the fix (current, 54 measurements)

Captured exactly as produced by the current make bench (post-treap-fix index, with the mount-table-scaling category added); reformatted into the tables above, but reproduced here unedited as the primary source record. This run predates the pack_write dedup fix ("Resolution #2"), which does not change any number in it — see that section for why.

=== PackFS vs. host filesystem: benchmark ===

environment:
  raw fs test root: /tmp/packfs_bench_raw_PnMAT7
  dir backend root: /tmp/packfs_bench_dir_U8Vgxi
  small files (N):  20000, 128 bytes each
  directories (N):  4000
  concurrency:      8 threads x 4000 ops (create+read+unlink)
  mounts (N):       2000
  large-file sizes: 1MB 16MB 64MB 
  NOTE: see the file header for what "raw fs" vs "raw+fsync" vs
        "dir" actually measure — they are not interchangeable.

== 1. create N small files ==
  create N small files         mem           0.0400s          499809 ops/s        61.0 MB/s
  create N small files         raw fs        0.9943s           20114 ops/s         2.5 MB/s
  create N small files         raw+fsync   116.2147s             172 ops/s         0.0 MB/s

== 2. read N small files ==
  read N small files           mem           0.0099s         2011901 ops/s       245.6 MB/s
  read N small files           raw fs        0.1606s          124543 ops/s        15.2 MB/s

== 3. stat N files ==
  stat N files                 mem           0.0077s         2598241 ops/s
  stat N files                 raw fs        0.0699s          286222 ops/s

== 4. readdir ==
  readdir (N entries)          mem           0.0041s         4845826 ops/s
  readdir (N entries)          raw fs        0.0069s         2894815 ops/s

== 1b. create N small files (dir backend vs. what it wraps) ==
  create N small files         dir           1.1418s           17517 ops/s         2.1 MB/s

== 2b. read N small files (dir backend) ==
  read N small files           dir           0.1530s          130705 ops/s        16.0 MB/s

== 3b. stat N files (dir backend) ==
  stat N files                 dir           0.1379s          145056 ops/s

== 4b. readdir (dir backend) ==
  readdir (N entries)          dir           0.0037s         5465162 ops/s

== 5. unlink N files ==
  unlink N files               mem           0.0179s         1117742 ops/s
  unlink N files               dir           0.4911s           40724 ops/s
  unlink N files               raw fs        0.4746s           42141 ops/s

== 6. mkdir/rmdir N directories ==
  mkdir N dirs                 mem           0.0056s          719696 ops/s
  rmdir N dirs                 mem           0.0040s          993724 ops/s
  mkdir N dirs                 dir           0.1710s           23386 ops/s
  rmdir N dirs                 dir           0.1257s           31827 ops/s
  mkdir N dirs                 raw fs        0.1374s           29122 ops/s
  rmdir N dirs                 raw fs        0.0951s           42043 ops/s

== 7. large sequential write/read ==
  write 1MB                    mem           0.0001s            7058 ops/s      7058.1 MB/s
  read 1MB                     mem           0.0000s           50051 ops/s     50050.9 MB/s
  write 16MB                   mem           0.0110s              91 ops/s      1458.2 MB/s
  read 16MB                    mem           0.0009s            1058 ops/s     16920.1 MB/s
  write 64MB                   mem           0.0586s              17 ops/s      1091.3 MB/s
  read 64MB                    mem           0.0032s             313 ops/s     20056.5 MB/s
  write 1MB                    raw fs        0.0004s            2759 ops/s      2759.4 MB/s
  read 1MB                     raw fs        0.0001s           16812 ops/s     16812.4 MB/s
  write 16MB                   raw fs        0.0044s             229 ops/s      3656.4 MB/s
  read 16MB                    raw fs        0.0010s             989 ops/s     15827.9 MB/s
  write 64MB                   raw fs        0.0186s              54 ops/s      3449.7 MB/s
  read 64MB                    raw fs        0.0047s             213 ops/s     13633.1 MB/s
  write 1MB                    raw+fsync     0.0180s              56 ops/s        55.7 MB/s
  read 1MB                     raw+fsync     0.0001s            9059 ops/s      9058.8 MB/s
  write 16MB                   raw+fsync     0.0231s              43 ops/s       691.7 MB/s
  read 16MB                    raw+fsync     0.0012s             855 ops/s     13672.2 MB/s
  write 64MB                   raw+fsync     0.0803s              12 ops/s       797.2 MB/s
  read 64MB                    raw+fsync     0.0048s             209 ops/s     13393.5 MB/s

== 8. random-access read: mmap'd pack vs. raw fs ==
  compact N entries to pack    pack          0.0267s          747839 ops/s        91.3 MB/s
  random-read N entries        pack          0.0082s         2435930 ops/s       297.4 MB/s
  random-read N entries        raw fs        0.1629s          122749 ops/s        15.0 MB/s

== 9. concurrency (create+read+unlink) ==
  concurrent create+read+unlink mem           0.2750s          349127 ops/s
  concurrent create+read+unlink raw fs        5.6498s           16992 ops/s

== 10. mount table scaling (the mount table uses the same full-snapshot-copy pattern the file index used to) ==
  mount 500 backends           vfs           0.0054s           92833 ops/s
  resolve, 500 mounts          vfs           0.0020s          246064 ops/s
  unmount 500 backends         vfs           0.0046s          109207 ops/s
  mount 2000 backends          vfs           0.0831s           24081 ops/s
  resolve, 2000 mounts         vfs           0.0296s           67458 ops/s
  unmount 2000 backends        vfs           0.0758s           26397 ops/s
  mount 8000 backends          vfs           1.4739s            5428 ops/s
  resolve, 8000 mounts         vfs           0.4938s           16200 ops/s
  unmount 8000 backends        vfs           1.5172s            5273 ops/s

=== summary (54 measurements) ===
  category                     backend        seconds        throughput          MB/s
  create N small files         mem            0.0400s          499809 ops/s        61.0 MB/s
  create N small files         raw fs         0.9943s           20114 ops/s         2.5 MB/s
  create N small files         raw+fsync    116.2147s             172 ops/s         0.0 MB/s
  read N small files           mem            0.0099s         2011901 ops/s       245.6 MB/s
  read N small files           raw fs         0.1606s          124543 ops/s        15.2 MB/s
  stat N files                 mem            0.0077s         2598241 ops/s               
  stat N files                 raw fs         0.0699s          286222 ops/s               
  readdir (N entries)          mem            0.0041s         4845826 ops/s               
  readdir (N entries)          raw fs         0.0069s         2894815 ops/s               
  create N small files         dir            1.1418s           17517 ops/s         2.1 MB/s
  read N small files           dir            0.1530s          130705 ops/s        16.0 MB/s
  stat N files                 dir            0.1379s          145056 ops/s               
  readdir (N entries)          dir            0.0037s         5465162 ops/s               
  unlink N files               mem            0.0179s         1117742 ops/s               
  unlink N files               dir            0.4911s           40724 ops/s               
  unlink N files               raw fs         0.4746s           42141 ops/s               
  mkdir N dirs                 mem            0.0056s          719696 ops/s               
  rmdir N dirs                 mem            0.0040s          993724 ops/s               
  mkdir N dirs                 dir            0.1710s           23386 ops/s               
  rmdir N dirs                 dir            0.1257s           31827 ops/s               
  mkdir N dirs                 raw fs         0.1374s           29122 ops/s               
  rmdir N dirs                 raw fs         0.0951s           42043 ops/s               
  write 1MB                    mem            0.0001s            7058 ops/s      7058.1 MB/s
  read 1MB                     mem            0.0000s           50051 ops/s     50050.9 MB/s
  write 16MB                   mem            0.0110s              91 ops/s      1458.2 MB/s
  read 16MB                    mem            0.0009s            1058 ops/s     16920.1 MB/s
  write 64MB                   mem            0.0586s              17 ops/s      1091.3 MB/s
  read 64MB                    mem            0.0032s             313 ops/s     20056.5 MB/s
  write 1MB                    raw fs         0.0004s            2759 ops/s      2759.4 MB/s
  read 1MB                     raw fs         0.0001s           16812 ops/s     16812.4 MB/s
  write 16MB                   raw fs         0.0044s             229 ops/s      3656.4 MB/s
  read 16MB                    raw fs         0.0010s             989 ops/s     15827.9 MB/s
  write 64MB                   raw fs         0.0186s              54 ops/s      3449.7 MB/s
  read 64MB                    raw fs         0.0047s             213 ops/s     13633.1 MB/s
  write 1MB                    raw+fsync      0.0180s              56 ops/s        55.7 MB/s
  read 1MB                     raw+fsync      0.0001s            9059 ops/s      9058.8 MB/s
  write 16MB                   raw+fsync      0.0231s              43 ops/s       691.7 MB/s
  read 16MB                    raw+fsync      0.0012s             855 ops/s     13672.2 MB/s
  write 64MB                   raw+fsync      0.0803s              12 ops/s       797.2 MB/s
  read 64MB                    raw+fsync      0.0048s             209 ops/s     13393.5 MB/s
  compact N entries to pack    pack           0.0267s          747839 ops/s        91.3 MB/s
  random-read N entries        pack           0.0082s         2435930 ops/s       297.4 MB/s
  random-read N entries        raw fs         0.1629s          122749 ops/s        15.0 MB/s
  concurrent create+read+unlink mem            0.2750s          349127 ops/s               
  concurrent create+read+unlink raw fs         5.6498s           16992 ops/s               
  mount 500 backends           vfs            0.0054s           92833 ops/s               
  resolve, 500 mounts          vfs            0.0020s          246064 ops/s               
  unmount 500 backends         vfs            0.0046s          109207 ops/s               
  mount 2000 backends          vfs            0.0831s           24081 ops/s               
  resolve, 2000 mounts         vfs            0.0296s           67458 ops/s               
  unmount 2000 backends        vfs            0.0758s           26397 ops/s               
  mount 8000 backends          vfs            1.4739s            5428 ops/s               
  resolve, 8000 mounts         vfs            0.4938s           16200 ops/s               
  unmount 8000 backends        vfs            1.5172s            5273 ops/s               

real	5m23.896s
user	0m4.732s
sys	0m19.047s

Reproducing

make bench

Takes several minutes on a similarly slow-fsync environment, dominated by the raw+fsync create test; expect well under a minute on a host with normal disk or tmpfs fsync latency.