Fix pack_write's O(n^2) dedup + correctness bug; document mount table O(n^2)

A audit for other instances of the file-index O(n^2) shape (fixed
previously via the persistent treap) found two more real issues:

1. pack_write's Section 9.2 exact-duplicate elimination was a linear scan
   of every previously-seen (hash, size) pair per entry -- O(n^2) total,
   invisible in the existing benchmark because its identical-content test
   files made every scan match on the first comparison. Measured with
   unique content instead: 80,000 entries took 1.74s, with a 20,000->80,000
   step showing 15.6x for a 4x-N step, matching O(n^2)'s 16x prediction.
   The same scan also trusted a (hash, size) match without ever comparing
   actual bytes -- a latent correctness bug, since FNV-1a64 is explicitly
   not collision-resistant. Fixed both at once with an open-addressing hash
   table (load factor 1/2, linear probing) plus a memcmp verification
   before ever reusing a data_off. Post-fix: 80,000 entries in 0.044s
   (39.6x faster), ratio drops to 3.35x (consistent with O(n)).

   Covered permanently by two new/extended tests: a white-box assertion in
   test_pack_overlay.c that duplicate-content entries share one data_off
   and distinct-content entries do not, and a new
   tests/test_pack_write_perf.c regression tripwire against 10,000 unique
   entries.

2. vfs.c's mount table uses the same full-array-copy-per-write pattern the
   file index used to, confirmed O(n^2) via a new bench/bench.c category
   (500/2,000/8,000 mounts, both 4x-N steps showing 15-20x). Deliberately
   NOT rewritten: mount points are created by a program's own source code,
   not workload-driven, so realistic mount counts never reach the scale
   that made the file index's O(n^2) a real problem. Documented with full
   reasoning in BENCH.md and CLAUDE.md rather than silently left as an
   undocumented gap.

Also fixes a real CI gap the new tests exposed: ci.yml's sanitizer-build
steps never passed -D_GNU_SOURCE when compiling test files (only the
library .o's got it), which was harmless while no test included
internal.h and became a link failure once two did (internal.h needs
_GNU_SOURCE for pthread_rwlock_t). And documents, in CONTRIBUTING.md and
CLAUDE.md, a sandbox flake observed directly during this work's own
sanitizer runs: ASan/UBSan test binaries occasionally fail to start with
AddressSanitizer:DEADLYSIGNAL (sometimes looping rather than exiting),
non-deterministically hitting different unrelated binaries across runs --
a startup race, not a memory-safety bug, confirmed by clean passes on
retry; sanitizer runs in such an environment should be timeout-wrapped.

BENCH.md's "After" table and Appendix B are replaced with the current,
complete 54-measurement bench/bench.c run (the original 45 plus the new
mount-scaling category); the pre-fix 45-measurement "Before" table is kept
as the historical record, per this project's documentation standard.

Verified: make test (all 6 binaries, including the 2 new/changed), a clean
make all, and repeated ASan+UBSan runs (0 real findings; the DEADLYSIGNAL
flake above was observed and correctly distinguished from a real finding
by re-running until a clean pass). TSan could not be run in this sandbox
(pre-existing, documented environment limitation).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
This commit is contained in:
2026-09-14 09:55:47 +00:00
co-authored by Claude Sonnet 5
parent edbf97b3f1
commit 684c6d95de
9 changed files with 693 additions and 205 deletions
+376 -172
View File
@@ -1,15 +1,24 @@
# PackFS vs. the host filesystem: benchmark results
**Status: the O(n²) bulk create/unlink finding this document originally
reported has been fixed.** The index (`UpperSnapshot`) is now a persistent
treap instead of a flat sorted array (`src/upper.c`; see `CLAUDE.md`'s
"Known performance characteristics"). This document keeps the original
run's numbers as the historical record of the problem — per this project's
own documentation standard, a correction extends the record rather than
erasing it — and adds the post-fix run's numbers alongside them, in full,
not a curated subset of either. **Read "Resolution" below before drawing
conclusions from the "Before" tables**, which describe a version of the
index this codebase no longer has.
**Status: two O(n²) findings this document originally reported have been
fixed; one further O(n²) finding is documented below and deliberately left
unfixed, with justification.** The file index (`UpperSnapshot`) is now a
persistent treap instead of a flat sorted array (`src/upper.c`; see
`CLAUDE.md`'s "Known performance characteristics") — see "Resolution"
below. `pack_write`'s compaction-time duplicate-content elimination
(Section 9.2) was also a linear scan with the same O(n²) shape, plus a
latent correctness bug (a hash match was trusted without a `memcmp`
verification); both are fixed — see "Resolution #2" below. The mount
table (`vfs.c`'s `MountSnapshot`) uses the same full-array-copy pattern the
file index used to, is confirmed O(n²) in mount count, and is *not*
rewritten — see "Finding: mount table scaling" below for why that is the
correct call, not an oversight. This document keeps every prior run's
numbers as the historical record of each problem — per this project's own
documentation standard, a correction extends the record rather than erasing
it — and adds each fix's numbers alongside them, in full, not a curated
subset of either. **Read "Resolution," "Resolution #2," and "Finding: mount
table scaling" below before drawing conclusions from the "Before" tables**,
which describe a version of the index this codebase no longer has.
This document reports the results of `bench/bench.c` (`make bench`), each
run once on the environment described below. These are single runs, not a
@@ -20,11 +29,17 @@ exactly what each backend label measures, is documented in `bench/bench.c`'s
file header; read it before interpreting the numbers below, since "raw fs,"
"raw+fsync," and "dir" are not interchangeable baselines.
Every table below reproduces every one of the 45 measurements each run's
`=== summary (45 measurements) ===` block reported — none omitted, none
combined — and the complete verbatim program output for both runs, exactly
as captured, is preserved in the two appendices at the end as the primary
source record.
The "Before" table below reproduces every one of the 45 measurements that
run's `=== summary (45 measurements) ===` block reported; the "After" table
reproduces every one of the 54 measurements the current `bench/bench.c`
reports (the original 45, unchanged in category, plus 9 new mount-table-
scaling measurements added after the original run — see "Finding: mount
table scaling"). None omitted, none combined in either table, and the
complete verbatim program output for both runs, exactly as captured, is
preserved in the two appendices at the end as the primary source record.
The `pack_write` dedup fix ("Resolution #2") is measured separately, by
calling `pack_write` directly rather than through `bench/bench.c` — see
that section for why.
## Environment
@@ -35,12 +50,19 @@ source record.
disk or tmpfs. This matters most for the `fsync` numbers (see below).
- `N_SMALL` = 20,000 files × 128 bytes for the metadata-heavy suite;
`N_DIRS` = 4,000; concurrency = 8 threads × 4,000 ops; large files up to
64 MB. Full parameters are in `bench/bench.c`.
- Wall-clock time: 5m49s (before), 4m4s (after), almost entirely spent in
the single `raw+fsync` 20,000-file create test (~116s in both runs — this
cost is on the host's storage, not PackFS, and is within measurement noise
of identical before and after the fix, as expected, since that test never
touches PackFS's index at all).
64 MB; `N_MOUNTS` = 2,000 (run at N/4, N, N×4 — 500/2,000/8,000 — for the
mount-table-scaling category added after the original run). Full
parameters are in `bench/bench.c`.
- Wall-clock time: 5m49s (before), 4m4s (after the treap fix, before the
mount-scaling category existed), 5m24s (current, with the mount-scaling
category added). The dominant cost throughout is the single `raw+fsync`
20,000-file create test (~116s in every run — this cost is on the host's
storage, not PackFS, and is within measurement noise across all three
runs, as expected, since that test never touches PackFS's index at all);
the added mount-scaling category itself accounts for only ~3.7s of the
increase from 4m4s to 5m24s, so the remainder is run-to-run variance in
untimed setup/teardown, consistent with this document's stated single-run,
illustrative-of-shape methodology, not a regression.
## Resolution
@@ -105,6 +127,141 @@ renames, nested directories, thousands of entries, cross-checked against an
independent reference model — not just "the benchmark ran fast without
crashing") is `tests/test_index_stress.c`, part of `make test`.
## Resolution #2: pack_write compaction dedup
A second, independent O(n²) finding, found while auditing the codebase for
other instances of the same shape of bug after "Resolution" above: `src/
pack.c`'s `pack_write` implements Section 9.2's exact-duplicate elimination
(two entries with byte-identical content share one `data_off` in the
compacted pack) as a linear scan, for each entry, of every previously-seen
`(hash, size)` pair. That is O(n) per entry, O(n²) total across n distinct
blobs — the same complexity-class bug "Resolution" fixed in the file index,
here in the compaction writer instead. It went unnoticed by the original
benchmark for a specific, identifiable reason: `bench/bench.c`'s compaction
test (category 8, "compact N entries to pack") writes byte-identical content
into every file, so every entry's linear scan matched on the very first
comparison (against entry 0) and never actually scanned — the benchmark's
own test data happened to be the linear scan's *best* case, not a
representative one. Measuring with unique per-file content instead (each
entry's actual worst case) surfaced it:
| N (unique entries) | Time |
|---|---|
| 5,000 | 0.0138s |
| 20,000 | 0.1114s |
| 80,000 | 1.7389s |
5,000 → 20,000 is a 4x step in N; time increased 8.07x. 20,000 → 80,000 is
also a 4x step; time increased 15.61x, converging toward O(n²)'s 16x
prediction as N grows (the smaller first ratio reflects fixed per-call
overhead not yet dominated by the quadratic term) — the same convergence
pattern "Resolution" observed for the file index, and the same standard used
there to distinguish a real complexity-class finding from noise.
This scan had a second, latent problem beyond speed: it compared only
`(hash, size)`, never the actual bytes, before deciding two entries were
duplicates and sharing their `data_off`. FNV-1a64 (`src/hash.c`) is
explicitly not collision-resistant — it is a fast checksum, not a
cryptographic hash — so two different-content entries that happened to
collide on `(hash, size)` would have been silently merged into one blob,
corrupting one of them. This was never observed in practice (no test's data
happened to collide), but was true of the code as written, not a
hypothetical: the fix required for the O(n²) issue also closes it.
**Fix:** the linear scan was replaced with an open-addressing hash table
(load factor 1/2, linear probing) over `(hash, size)`, plus — critically —
an actual `memcmp` against the candidate's stored data pointer before ever
reusing a `data_off`, so a hash match is only ever a candidate, never
accepted as proof of equality. Comment and full listing:
`src/pack.c`'s `DedupSlot` struct and the rewritten first pass of
`pack_write`.
**After the fix, same methodology:**
| N (unique entries) | Time | vs. before |
|---|---|---|
| 5,000 | 0.0079s | 1.75x faster |
| 20,000 | 0.0131s | 8.50x faster |
| 80,000 | 0.0439s | 39.6x faster |
5,000 → 20,000 now shows 1.66x (O(n) predicts 4x, but a hash table's
dominant cost at these sizes is close to constant per insert, so sub-linear
growth here is expected, not suspicious); 20,000 → 80,000 shows 3.35x,
close to O(n)'s 4x prediction and nowhere near O(n²)'s 16x. The complexity
class changed, exactly as with the file index.
This fix is not reflected in the "After" table/Appendix B below, because
that `make bench` run predates it — but it does not need to be: category
8's "compact N entries to pack" measurement uses identical content (as
explained above), which was already the linear scan's O(1)-per-entry best
case before this fix and is unaffected by the fix (a hash table lookup that
matches on the first probe is also O(1) per entry); the number in the
"After" table (0.0267s, 20,000 identical-content entries) remains accurate
and does not need to be re-measured. The fix matters for compacting many
files with *distinct* content, which `bench/bench.c`'s existing compaction
test does not exercise — hence the separate, direct `pack_write`
measurement above rather than a `bench/bench.c` change. Correctness is
covered permanently by `tests/test_pack_overlay.c` (asserts `/a.txt` and
`/dup.txt`, identical content, share one `data_off`, while `/b.txt`,
different content, shares neither's — catching both "dedup stopped
happening" and "dedup over-matched"); the performance fix is covered
permanently by `tests/test_pack_write_perf.c` (10,000 unique entries under
a generous time bound, a regression tripwire rather than a tight
assertion, run as part of `make test`).
## Finding: mount table scaling (O(n²), documented, deliberately not fixed)
`src/vfs.c`'s mount table (`MountSnapshot`) uses the exact same pattern the
file index used before "Resolution" above: every `vfs_mount`/`vfs_unmount`
copies the entire array of mounted backends before publishing the next
snapshot (Section 5.3's single-writer model applies to the mount table too,
per `concept.md`'s explicit statement that "the mount table is not separate
global state exempt from this model"). This is confirmed O(n²) in mount
count by the `bench/bench.c` category 10 measurements (full numbers in the
"After" table below and Appendix B):
| N mounts | mount | resolve (stat through N mounts) | unmount |
|---|---|---|---|
| 500 | 0.0054s | 0.0020s | 0.0046s |
| 2,000 | 0.0831s | 0.0296s | 0.0758s |
| 8,000 | 1.4739s | 0.4938s | 1.5172s |
500 → 2,000 is a 4x step in N: mount increased 15.4x, resolve 14.8x,
unmount 16.5x. 2,000 → 8,000 is also a 4x step: mount increased 17.7x,
resolve 16.7x, unmount 20.0x. All six ratios cluster around O(n²)'s 16x
prediction for a 4x-N step, and nowhere near O(n)'s 4x — the same
complexity-class-confirmation methodology used for both findings above,
applied here and yielding the same verdict. Notably, `resolve` (a single
`vfs_stat` through an already-built N-mount table) is *also* O(n²) across N
resolves, not just O(n) per resolve as might be assumed — consistent with
longest-prefix-match dispatch being a linear scan of the mount table per
call (Section 3's dispatch design), which is itself O(n) per call and O(n²)
across N calls, independent of the snapshot-copy cost on the write side.
**This is deliberately not fixed, and the reasoning is the same trigger
condition `concept.md` Section 5.3 names for the file index — inverted:**
"a persistent (structurally shared) tree structure is not required until
this assumption is empirically violated." The file index's assumption was
violated because file counts are workload-driven and can plausibly reach
tens of thousands (or more) at runtime, chosen by whatever program embeds
PackFS and whatever data it manages. Mount counts are different in kind,
not just degree: a mount point is created by a call to `vfs_mount` written
into a program's own source code, at a location the program's author
chose — it is bounded by how many lines of "mount this backend at this
path" a person or a build script is willing to write, not by user data,
input size, or any other externally-driven quantity. A program with 8,000
mount points, structured that way on purpose, is not a realistic workload
this system needs to serve well; a program with 8,000 *files* is an
ordinary one. Converting the mount table to a persistent treap would fix
this at the cost of real, permanent complexity (an additional generic
persistent-tree instantiation, or a second copy of the treap logic
specialized to prefix matching) for a case that has not been and is not
expected to be hit. Should a future use case genuinely need thousands of
mount points at runtime — for instance, one mount per user in a
multi-tenant embedding — this finding, and the numbers above, are the
starting point for that decision; until then, this is recorded as a known,
measured, and consciously accepted limitation, not an unexamined one.
## Before: complete results (pre-fix; kept as the historical record — see "Resolution" above)
All 45 measurements from that run's summary block, none omitted or combined.
@@ -157,65 +314,83 @@ All 45 measurements from that run's summary block, none omitted or combined.
| concurrent create+read+unlink (8×4,000×3) | mem | 36.7261s | 2,614 ops/s | — |
| concurrent create+read+unlink (8×4,000×3) | raw fs | 6.1202s | 15,686 ops/s | — |
## After: complete results (post-fix)
## After: complete results (post-fix, current)
All 45 measurements from that run's summary block, none omitted or combined.
All 54 measurements from the current `bench/bench.c`'s summary block, none
omitted or combined: the original 45 (re-run, post-treap-fix) plus 9 new
mount-table-scaling measurements (category 10, added after the original
run — see "Finding: mount table scaling" above). The `pack_write` dedup fix
("Resolution #2" above) postdates this run but does not change any number
in it — see that section for why the compaction row below is unaffected.
| Category | Backend | Time | Throughput | MB/s |
|---|---|---|---|---|
| create 20,000 files | mem | 0.0372s | 537,103 ops/s | 65.6 |
| create 20,000 files | raw fs | 1.0098s | 19,805 ops/s | 2.4 |
| create 20,000 files | raw+fsync | 116.0980s | 172 ops/s | 0.0 |
| read 20,000 files | mem | 0.0085s | 2,364,410 ops/s | 288.6 |
| read 20,000 files | raw fs | 0.1594s | 125,459 ops/s | 15.3 |
| stat 20,000 files | mem | 0.0073s | 2,723,001 ops/s | — |
| stat 20,000 files | raw fs | 0.0697s | 287,063 ops/s | — |
| readdir (20,000 entries) | mem | 0.0042s | 4,816,418 /s | — |
| readdir (20,000 entries) | raw fs | 0.0070s | 2,840,890 /s | — |
| create 20,000 files | dir | 1.1531s | 17,345 ops/s | 2.1 |
| read 20,000 files | dir | 0.1601s | 124,933 ops/s | 15.3 |
| stat 20,000 files | dir | 0.1509s | 132,517 ops/s | — |
| readdir (20,000 entries) | dir | 0.0039s | 5,128,581 /s | — |
| unlink 20,000 files | mem | 0.0320s | 625,437 ops/s | — |
| unlink 20,000 files | dir | 0.5146s | 38,864 ops/s | — |
| unlink 20,000 files | raw fs | 0.4923s | 40,628 ops/s | — |
| mkdir 4,000 dirs | mem | 0.0052s | 771,085 ops/s | — |
| rmdir 4,000 dirs | mem | 0.0045s | 898,169 ops/s | — |
| mkdir 4,000 dirs | dir | 0.1774s | 22,551 ops/s | — |
| rmdir 4,000 dirs | dir | 0.1118s | 35,778 ops/s | — |
| mkdir 4,000 dirs | raw fs | 0.1367s | 29,251 ops/s | — |
| rmdir 4,000 dirs | raw fs | 0.0901s | 44,388 ops/s | — |
| write 1MB | mem | 0.0003s | 2,937 ops/s | 2,937.3 |
| read 1MB | mem | 0.0000s | 49,752 ops/s | 49,751.7 |
| write 16MB | mem | 0.0108s | 93 ops/s | 1,483.0 |
| read 16MB | mem | 0.0010s | 1,050 ops/s | 16,796.7 |
| write 64MB | mem | 0.0589s | 17 ops/s | 1,086.8 |
| read 64MB | mem | 0.0033s | 304 ops/s | 19,475.8 |
| write 1MB | raw fs | 0.0004s | 2,463 ops/s | 2,462.6 |
| read 1MB | raw fs | 0.0001s | 15,738 ops/s | 15,738.2 |
| write 16MB | raw fs | 0.0046s | 217 ops/s | 3,477.6 |
| read 16MB | raw fs | 0.0011s | 945 ops/s | 15,121.6 |
| write 64MB | raw fs | 0.0183s | 55 ops/s | 3,495.7 |
| read 64MB | raw fs | 0.0050s | 198 ops/s | 12,702.5 |
| write 1MB | raw+fsync | 0.0235s | 43 ops/s | 42.5 |
| read 1MB | raw+fsync | 0.0001s | 9,816 ops/s | 9,816.4 |
| write 16MB | raw+fsync | 0.0228s | 44 ops/s | 701.5 |
| read 16MB | raw+fsync | 0.0012s | 832 ops/s | 13,316.1 |
| write 64MB | raw+fsync | 0.0825s | 12 ops/s | 775.9 |
| read 64MB | raw+fsync | 0.0057s | 176 ops/s | 11,254.5 |
| compact 20,000 entries to pack | pack | 0.0245s | 814,864 ops/s | 99.5 |
| random-read 20,000 entries | pack (mmap'd) | 0.0082s | 2,445,144 ops/s | 298.5 |
| random-read 20,000 entries | raw fs | 0.1657s | 120,709 ops/s | 14.7 |
| concurrent create+read+unlink (8×4,000×3) | mem | 0.2807s | 341,972 ops/s | — |
| concurrent create+read+unlink (8×4,000×3) | raw fs | 6.2980s | 15,243 ops/s | — |
| create 20,000 files | mem | 0.0400s | 499,809 ops/s | 61.0 |
| create 20,000 files | raw fs | 0.9943s | 20,114 ops/s | 2.5 |
| create 20,000 files | raw+fsync | 116.2147s | 172 ops/s | 0.0 |
| read 20,000 files | mem | 0.0099s | 2,011,901 ops/s | 245.6 |
| read 20,000 files | raw fs | 0.1606s | 124,543 ops/s | 15.2 |
| stat 20,000 files | mem | 0.0077s | 2,598,241 ops/s | — |
| stat 20,000 files | raw fs | 0.0699s | 286,222 ops/s | — |
| readdir (20,000 entries) | mem | 0.0041s | 4,845,826 /s | — |
| readdir (20,000 entries) | raw fs | 0.0069s | 2,894,815 /s | — |
| create 20,000 files | dir | 1.1418s | 17,517 ops/s | 2.1 |
| read 20,000 files | dir | 0.1530s | 130,705 ops/s | 16.0 |
| stat 20,000 files | dir | 0.1379s | 145,056 ops/s | — |
| readdir (20,000 entries) | dir | 0.0037s | 5,465,162 /s | — |
| unlink 20,000 files | mem | 0.0179s | 1,117,742 ops/s | — |
| unlink 20,000 files | dir | 0.4911s | 40,724 ops/s | — |
| unlink 20,000 files | raw fs | 0.4746s | 42,141 ops/s | — |
| mkdir 4,000 dirs | mem | 0.0056s | 719,696 ops/s | — |
| rmdir 4,000 dirs | mem | 0.0040s | 993,724 ops/s | — |
| mkdir 4,000 dirs | dir | 0.1710s | 23,386 ops/s | — |
| rmdir 4,000 dirs | dir | 0.1257s | 31,827 ops/s | — |
| mkdir 4,000 dirs | raw fs | 0.1374s | 29,122 ops/s | — |
| rmdir 4,000 dirs | raw fs | 0.0951s | 42,043 ops/s | — |
| write 1MB | mem | 0.0001s | 7,058 ops/s | 7,058.1 |
| read 1MB | mem | 0.0000s | 50,051 ops/s | 50,050.9 |
| write 16MB | mem | 0.0110s | 91 ops/s | 1,458.2 |
| read 16MB | mem | 0.0009s | 1,058 ops/s | 16,920.1 |
| write 64MB | mem | 0.0586s | 17 ops/s | 1,091.3 |
| read 64MB | mem | 0.0032s | 313 ops/s | 20,056.5 |
| write 1MB | raw fs | 0.0004s | 2,759 ops/s | 2,759.4 |
| read 1MB | raw fs | 0.0001s | 16,812 ops/s | 16,812.4 |
| write 16MB | raw fs | 0.0044s | 229 ops/s | 3,656.4 |
| read 16MB | raw fs | 0.0010s | 989 ops/s | 15,827.9 |
| write 64MB | raw fs | 0.0186s | 54 ops/s | 3,449.7 |
| read 64MB | raw fs | 0.0047s | 213 ops/s | 13,633.1 |
| write 1MB | raw+fsync | 0.0180s | 56 ops/s | 55.7 |
| read 1MB | raw+fsync | 0.0001s | 9,059 ops/s | 9,058.8 |
| write 16MB | raw+fsync | 0.0231s | 43 ops/s | 691.7 |
| read 16MB | raw+fsync | 0.0012s | 855 ops/s | 13,672.2 |
| write 64MB | raw+fsync | 0.0803s | 12 ops/s | 797.2 |
| read 64MB | raw+fsync | 0.0048s | 209 ops/s | 13,393.5 |
| compact 20,000 entries to pack | pack | 0.0267s | 747,839 ops/s | 91.3 |
| random-read 20,000 entries | pack (mmap'd) | 0.0082s | 2,435,930 ops/s | 297.4 |
| random-read 20,000 entries | raw fs | 0.1629s | 122,749 ops/s | 15.0 |
| concurrent create+read+unlink (8×4,000×3) | mem | 0.2750s | 349,127 ops/s | — |
| concurrent create+read+unlink (8×4,000×3) | raw fs | 5.6498s | 16,992 ops/s | — |
| mount 500 backends | vfs | 0.0054s | 92,833 ops/s | — |
| resolve, 500 mounts | vfs | 0.0020s | 246,064 ops/s | — |
| unmount 500 backends | vfs | 0.0046s | 109,207 ops/s | — |
| mount 2,000 backends | vfs | 0.0831s | 24,081 ops/s | — |
| resolve, 2,000 mounts | vfs | 0.0296s | 67,458 ops/s | — |
| unmount 2,000 backends | vfs | 0.0758s | 26,397 ops/s | — |
| mount 8,000 backends | vfs | 1.4739s | 5,428 ops/s | — |
| resolve, 8,000 mounts | vfs | 0.4938s | 16,200 ops/s | — |
| unmount 8,000 backends | vfs | 1.5172s | 5,273 ops/s | — |
Everything not involving bulk `mem`/`dir` structural writes (reads, `stat`,
`readdir`, large sequential I/O, pack random-access, compaction) is
unchanged within normal run-to-run noise, as the row-by-row comparison
above shows — the treap rewrite only touches the structural-write path
(Section 5.3), not content reads, content writes, or the pack format. The
`raw+fsync` create row is identical to three decimal places (116.1030s vs
116.0980s) precisely because it never touches PackFS at all.
against the "Before" table shows — the treap rewrite only touches the
structural-write path (Section 5.3), not content reads, content writes, or
the pack format. The `raw+fsync` create row is unchanged within noise
(116.1030s before, 116.2147s here) precisely because it never touches
PackFS at all. The nine new mount-table rows have no "Before" counterpart —
the mount table was never rewritten, so there is no pre/post comparison to
make for them; they are current-state measurements only, discussed in
"Finding: mount table scaling" above.
## Analysis
@@ -251,6 +426,12 @@ now marked where "Resolution" above supersedes it.
among its clearest wins.
- **Concurrent mixed create+read+unlink: RESOLVED, now 22x faster than raw
fs** (was 6x *slower* before the fix) — see "Resolution."
- **`pack_write` compaction dedup on distinct-content files: RESOLVED, 1.75–
39.6x faster depending on N, plus a latent hash-collision correctness bug
closed alongside it** — see "Resolution #2." Not visible in this
benchmark's own compaction test (category 8), which uses identical
content and so never exercised either problem; measured separately by
calling `pack_write` directly.
### Where raw fs still wins, and why that's expected, not a bug
@@ -430,147 +611,170 @@ user 2m9.600s
sys 0m20.941s
```
## Appendix B: complete verbatim output, after the fix
## Appendix B: complete verbatim output, after the fix (current, 54 measurements)
Captured exactly as produced by `make bench` against the post-fix
(persistent treap) index; reformatted into the tables above, but reproduced
here unedited as the primary source record.
Captured exactly as produced by the current `make bench` (post-treap-fix
index, with the mount-table-scaling category added); reformatted into the
tables above, but reproduced here unedited as the primary source record.
This run predates the `pack_write` dedup fix ("Resolution #2"), which does
not change any number in it — see that section for why.
```text
=== PackFS vs. host filesystem: benchmark ===
environment:
raw fs test root: /tmp/packfs_bench_raw_fbOZ1b
dir backend root: /tmp/packfs_bench_dir_7tEOsC
raw fs test root: /tmp/packfs_bench_raw_PnMAT7
dir backend root: /tmp/packfs_bench_dir_U8Vgxi
small files (N): 20000, 128 bytes each
directories (N): 4000
concurrency: 8 threads x 4000 ops (create+read+unlink)
mounts (N): 2000
large-file sizes: 1MB 16MB 64MB
NOTE: see the file header for what "raw fs" vs "raw+fsync" vs
"dir" actually measure — they are not interchangeable.
== 1. create N small files ==
create N small files mem 0.0372s 537103 ops/s 65.6 MB/s
create N small files raw fs 1.0098s 19805 ops/s 2.4 MB/s
create N small files raw+fsync 116.0980s 172 ops/s 0.0 MB/s
create N small files mem 0.0400s 499809 ops/s 61.0 MB/s
create N small files raw fs 0.9943s 20114 ops/s 2.5 MB/s
create N small files raw+fsync 116.2147s 172 ops/s 0.0 MB/s
== 2. read N small files ==
read N small files mem 0.0085s 2364410 ops/s 288.6 MB/s
read N small files raw fs 0.1594s 125459 ops/s 15.3 MB/s
read N small files mem 0.0099s 2011901 ops/s 245.6 MB/s
read N small files raw fs 0.1606s 124543 ops/s 15.2 MB/s
== 3. stat N files ==
stat N files mem 0.0073s 2723001 ops/s
stat N files raw fs 0.0697s 287063 ops/s
stat N files mem 0.0077s 2598241 ops/s
stat N files raw fs 0.0699s 286222 ops/s
== 4. readdir ==
readdir (N entries) mem 0.0042s 4816418 ops/s
readdir (N entries) raw fs 0.0070s 2840890 ops/s
readdir (N entries) mem 0.0041s 4845826 ops/s
readdir (N entries) raw fs 0.0069s 2894815 ops/s
== 1b. create N small files (dir backend vs. what it wraps) ==
create N small files dir 1.1531s 17345 ops/s 2.1 MB/s
create N small files dir 1.1418s 17517 ops/s 2.1 MB/s
== 2b. read N small files (dir backend) ==
read N small files dir 0.1601s 124933 ops/s 15.3 MB/s
read N small files dir 0.1530s 130705 ops/s 16.0 MB/s
== 3b. stat N files (dir backend) ==
stat N files dir 0.1509s 132517 ops/s
stat N files dir 0.1379s 145056 ops/s
== 4b. readdir (dir backend) ==
readdir (N entries) dir 0.0039s 5128581 ops/s
readdir (N entries) dir 0.0037s 5465162 ops/s
== 5. unlink N files ==
unlink N files mem 0.0320s 625437 ops/s
unlink N files dir 0.5146s 38864 ops/s
unlink N files raw fs 0.4923s 40628 ops/s
unlink N files mem 0.0179s 1117742 ops/s
unlink N files dir 0.4911s 40724 ops/s
unlink N files raw fs 0.4746s 42141 ops/s
== 6. mkdir/rmdir N directories ==
mkdir N dirs mem 0.0052s 771085 ops/s
rmdir N dirs mem 0.0045s 898169 ops/s
mkdir N dirs dir 0.1774s 22551 ops/s
rmdir N dirs dir 0.1118s 35778 ops/s
mkdir N dirs raw fs 0.1367s 29251 ops/s
rmdir N dirs raw fs 0.0901s 44388 ops/s
mkdir N dirs mem 0.0056s 719696 ops/s
rmdir N dirs mem 0.0040s 993724 ops/s
mkdir N dirs dir 0.1710s 23386 ops/s
rmdir N dirs dir 0.1257s 31827 ops/s
mkdir N dirs raw fs 0.1374s 29122 ops/s
rmdir N dirs raw fs 0.0951s 42043 ops/s
== 7. large sequential write/read ==
write 1MB mem 0.0003s 2937 ops/s 2937.3 MB/s
read 1MB mem 0.0000s 49752 ops/s 49751.7 MB/s
write 16MB mem 0.0108s 93 ops/s 1483.0 MB/s
read 16MB mem 0.0010s 1050 ops/s 16796.7 MB/s
write 64MB mem 0.0589s 17 ops/s 1086.8 MB/s
read 64MB mem 0.0033s 304 ops/s 19475.8 MB/s
write 1MB raw fs 0.0004s 2463 ops/s 2462.6 MB/s
read 1MB raw fs 0.0001s 15738 ops/s 15738.2 MB/s
write 16MB raw fs 0.0046s 217 ops/s 3477.6 MB/s
read 16MB raw fs 0.0011s 945 ops/s 15121.6 MB/s
write 64MB raw fs 0.0183s 55 ops/s 3495.7 MB/s
read 64MB raw fs 0.0050s 198 ops/s 12702.5 MB/s
write 1MB raw+fsync 0.0235s 43 ops/s 42.5 MB/s
read 1MB raw+fsync 0.0001s 9816 ops/s 9816.4 MB/s
write 16MB raw+fsync 0.0228s 44 ops/s 701.5 MB/s
read 16MB raw+fsync 0.0012s 832 ops/s 13316.1 MB/s
write 64MB raw+fsync 0.0825s 12 ops/s 775.9 MB/s
read 64MB raw+fsync 0.0057s 176 ops/s 11254.5 MB/s
write 1MB mem 0.0001s 7058 ops/s 7058.1 MB/s
read 1MB mem 0.0000s 50051 ops/s 50050.9 MB/s
write 16MB mem 0.0110s 91 ops/s 1458.2 MB/s
read 16MB mem 0.0009s 1058 ops/s 16920.1 MB/s
write 64MB mem 0.0586s 17 ops/s 1091.3 MB/s
read 64MB mem 0.0032s 313 ops/s 20056.5 MB/s
write 1MB raw fs 0.0004s 2759 ops/s 2759.4 MB/s
read 1MB raw fs 0.0001s 16812 ops/s 16812.4 MB/s
write 16MB raw fs 0.0044s 229 ops/s 3656.4 MB/s
read 16MB raw fs 0.0010s 989 ops/s 15827.9 MB/s
write 64MB raw fs 0.0186s 54 ops/s 3449.7 MB/s
read 64MB raw fs 0.0047s 213 ops/s 13633.1 MB/s
write 1MB raw+fsync 0.0180s 56 ops/s 55.7 MB/s
read 1MB raw+fsync 0.0001s 9059 ops/s 9058.8 MB/s
write 16MB raw+fsync 0.0231s 43 ops/s 691.7 MB/s
read 16MB raw+fsync 0.0012s 855 ops/s 13672.2 MB/s
write 64MB raw+fsync 0.0803s 12 ops/s 797.2 MB/s
read 64MB raw+fsync 0.0048s 209 ops/s 13393.5 MB/s
== 8. random-access read: mmap'd pack vs. raw fs ==
compact N entries to pack pack 0.0245s 814864 ops/s 99.5 MB/s
random-read N entries pack 0.0082s 2445144 ops/s 298.5 MB/s
random-read N entries raw fs 0.1657s 120709 ops/s 14.7 MB/s
compact N entries to pack pack 0.0267s 747839 ops/s 91.3 MB/s
random-read N entries pack 0.0082s 2435930 ops/s 297.4 MB/s
random-read N entries raw fs 0.1629s 122749 ops/s 15.0 MB/s
== 9. concurrency (create+read+unlink) ==
concurrent create+read+unlink mem 0.2807s 341972 ops/s
concurrent create+read+unlink raw fs 6.2980s 15243 ops/s
concurrent create+read+unlink mem 0.2750s 349127 ops/s
concurrent create+read+unlink raw fs 5.6498s 16992 ops/s
=== summary (45 measurements) ===
== 10. mount table scaling (the mount table uses the same full-snapshot-copy pattern the file index used to) ==
mount 500 backends vfs 0.0054s 92833 ops/s
resolve, 500 mounts vfs 0.0020s 246064 ops/s
unmount 500 backends vfs 0.0046s 109207 ops/s
mount 2000 backends vfs 0.0831s 24081 ops/s
resolve, 2000 mounts vfs 0.0296s 67458 ops/s
unmount 2000 backends vfs 0.0758s 26397 ops/s
mount 8000 backends vfs 1.4739s 5428 ops/s
resolve, 8000 mounts vfs 0.4938s 16200 ops/s
unmount 8000 backends vfs 1.5172s 5273 ops/s
=== summary (54 measurements) ===
category backend seconds throughput MB/s
create N small files mem 0.0372s 537103 ops/s 65.6 MB/s
create N small files raw fs 1.0098s 19805 ops/s 2.4 MB/s
create N small files raw+fsync 116.0980s 172 ops/s 0.0 MB/s
read N small files mem 0.0085s 2364410 ops/s 288.6 MB/s
read N small files raw fs 0.1594s 125459 ops/s 15.3 MB/s
stat N files mem 0.0073s 2723001 ops/s
stat N files raw fs 0.0697s 287063 ops/s
readdir (N entries) mem 0.0042s 4816418 ops/s
readdir (N entries) raw fs 0.0070s 2840890 ops/s
create N small files dir 1.1531s 17345 ops/s 2.1 MB/s
read N small files dir 0.1601s 124933 ops/s 15.3 MB/s
stat N files dir 0.1509s 132517 ops/s
readdir (N entries) dir 0.0039s 5128581 ops/s
unlink N files mem 0.0320s 625437 ops/s
unlink N files dir 0.5146s 38864 ops/s
unlink N files raw fs 0.4923s 40628 ops/s
mkdir N dirs mem 0.0052s 771085 ops/s
rmdir N dirs mem 0.0045s 898169 ops/s
mkdir N dirs dir 0.1774s 22551 ops/s
rmdir N dirs dir 0.1118s 35778 ops/s
mkdir N dirs raw fs 0.1367s 29251 ops/s
rmdir N dirs raw fs 0.0901s 44388 ops/s
write 1MB mem 0.0003s 2937 ops/s 2937.3 MB/s
read 1MB mem 0.0000s 49752 ops/s 49751.7 MB/s
write 16MB mem 0.0108s 93 ops/s 1483.0 MB/s
read 16MB mem 0.0010s 1050 ops/s 16796.7 MB/s
write 64MB mem 0.0589s 17 ops/s 1086.8 MB/s
read 64MB mem 0.0033s 304 ops/s 19475.8 MB/s
write 1MB raw fs 0.0004s 2463 ops/s 2462.6 MB/s
read 1MB raw fs 0.0001s 15738 ops/s 15738.2 MB/s
write 16MB raw fs 0.0046s 217 ops/s 3477.6 MB/s
read 16MB raw fs 0.0011s 945 ops/s 15121.6 MB/s
write 64MB raw fs 0.0183s 55 ops/s 3495.7 MB/s
read 64MB raw fs 0.0050s 198 ops/s 12702.5 MB/s
write 1MB raw+fsync 0.0235s 43 ops/s 42.5 MB/s
read 1MB raw+fsync 0.0001s 9816 ops/s 9816.4 MB/s
write 16MB raw+fsync 0.0228s 44 ops/s 701.5 MB/s
read 16MB raw+fsync 0.0012s 832 ops/s 13316.1 MB/s
write 64MB raw+fsync 0.0825s 12 ops/s 775.9 MB/s
read 64MB raw+fsync 0.0057s 176 ops/s 11254.5 MB/s
compact N entries to pack pack 0.0245s 814864 ops/s 99.5 MB/s
random-read N entries pack 0.0082s 2445144 ops/s 298.5 MB/s
random-read N entries raw fs 0.1657s 120709 ops/s 14.7 MB/s
concurrent create+read+unlink mem 0.2807s 341972 ops/s
concurrent create+read+unlink raw fs 6.2980s 15243 ops/s
create N small files mem 0.0400s 499809 ops/s 61.0 MB/s
create N small files raw fs 0.9943s 20114 ops/s 2.5 MB/s
create N small files raw+fsync 116.2147s 172 ops/s 0.0 MB/s
read N small files mem 0.0099s 2011901 ops/s 245.6 MB/s
read N small files raw fs 0.1606s 124543 ops/s 15.2 MB/s
stat N files mem 0.0077s 2598241 ops/s
stat N files raw fs 0.0699s 286222 ops/s
readdir (N entries) mem 0.0041s 4845826 ops/s
readdir (N entries) raw fs 0.0069s 2894815 ops/s
create N small files dir 1.1418s 17517 ops/s 2.1 MB/s
read N small files dir 0.1530s 130705 ops/s 16.0 MB/s
stat N files dir 0.1379s 145056 ops/s
readdir (N entries) dir 0.0037s 5465162 ops/s
unlink N files mem 0.0179s 1117742 ops/s
unlink N files dir 0.4911s 40724 ops/s
unlink N files raw fs 0.4746s 42141 ops/s
mkdir N dirs mem 0.0056s 719696 ops/s
rmdir N dirs mem 0.0040s 993724 ops/s
mkdir N dirs dir 0.1710s 23386 ops/s
rmdir N dirs dir 0.1257s 31827 ops/s
mkdir N dirs raw fs 0.1374s 29122 ops/s
rmdir N dirs raw fs 0.0951s 42043 ops/s
write 1MB mem 0.0001s 7058 ops/s 7058.1 MB/s
read 1MB mem 0.0000s 50051 ops/s 50050.9 MB/s
write 16MB mem 0.0110s 91 ops/s 1458.2 MB/s
read 16MB mem 0.0009s 1058 ops/s 16920.1 MB/s
write 64MB mem 0.0586s 17 ops/s 1091.3 MB/s
read 64MB mem 0.0032s 313 ops/s 20056.5 MB/s
write 1MB raw fs 0.0004s 2759 ops/s 2759.4 MB/s
read 1MB raw fs 0.0001s 16812 ops/s 16812.4 MB/s
write 16MB raw fs 0.0044s 229 ops/s 3656.4 MB/s
read 16MB raw fs 0.0010s 989 ops/s 15827.9 MB/s
write 64MB raw fs 0.0186s 54 ops/s 3449.7 MB/s
read 64MB raw fs 0.0047s 213 ops/s 13633.1 MB/s
write 1MB raw+fsync 0.0180s 56 ops/s 55.7 MB/s
read 1MB raw+fsync 0.0001s 9059 ops/s 9058.8 MB/s
write 16MB raw+fsync 0.0231s 43 ops/s 691.7 MB/s
read 16MB raw+fsync 0.0012s 855 ops/s 13672.2 MB/s
write 64MB raw+fsync 0.0803s 12 ops/s 797.2 MB/s
read 64MB raw+fsync 0.0048s 209 ops/s 13393.5 MB/s
compact N entries to pack pack 0.0267s 747839 ops/s 91.3 MB/s
random-read N entries pack 0.0082s 2435930 ops/s 297.4 MB/s
random-read N entries raw fs 0.1629s 122749 ops/s 15.0 MB/s
concurrent create+read+unlink mem 0.2750s 349127 ops/s
concurrent create+read+unlink raw fs 5.6498s 16992 ops/s
mount 500 backends vfs 0.0054s 92833 ops/s
resolve, 500 mounts vfs 0.0020s 246064 ops/s
unmount 500 backends vfs 0.0046s 109207 ops/s
mount 2000 backends vfs 0.0831s 24081 ops/s
resolve, 2000 mounts vfs 0.0296s 67458 ops/s
unmount 2000 backends vfs 0.0758s 26397 ops/s
mount 8000 backends vfs 1.4739s 5428 ops/s
resolve, 8000 mounts vfs 0.4938s 16200 ops/s
unmount 8000 backends vfs 1.5172s 5273 ops/s
real 4m3.827s
user 0m1.039s
sys 0m19.841s
real 5m23.896s
user 0m4.732s
sys 0m19.047s
```
## Reproducing