Doing another literal, no-caveats audit pass (per the standing request to verify literally everything, not just reassert it) found real, concrete inconsistencies introduced by the previous commit: it replaced BENCH.md's "After: complete results" table and Appendix B with a newer, more complete bench/bench.c run (adding the mount-table-scaling category), but several other places in the same document still quoted exact figures from the older run those tables used to hold -- a reader cross-checking the "Resolution" section's headline table, or the "Analysis" section's bullet points, against the "After" table below them would have found contradictory numbers for the same measurement. Verified programmatically, not just by re-reading: every row's time, throughput, and MB/s in the Before (45 rows) and After (54 rows) tables now matches its corresponding appendix line exactly (0 mismatches out of 99 rows, checked with a script, not by eye). Fixed: - The "Resolution" section's headline before/after table used the older run's mem create/unlink/mkdir/rmdir numbers (0.0372s/0.0320s/0.0052s/ 0.0045s), which don't match the current After table (0.0400s/0.0179s/ 0.0056s/0.0040s) -- for unlink specifically, a genuine near-2x difference between two runs of identical, already-fixed code, not just noise. Now points at the single current run, with speedups recomputed (239x/590x/ 60x/85x/11x/21x/134x, mem now 25-27x faster than raw fs, complexity-class check now 7.14x vs the old 7.15x -- nearly identical, which is the reassurance that actually matters), plus an explicit note on the observed run-to-run variance rather than silently picking one number. - The "Analysis" section's read/stat/random-read/compaction/dir-stat/ large-write bullets all cited the older run's exact ops/s and MB/s figures (2.4M small-file reads, ~815K compaction entries/s, 133K dir stat, etc.); recomputed against the current After table (now 15-16x for small reads, ~9x for stat, ~91MB/s / ~748K entries/s for compaction, 145K vs 286K for dir stat, corrected large-write MB/s ranges). - The same stale 65-330x / 15-27x / 22x range in CLAUDE.md's "Known performance characteristics" bulk-write bullet, updated to match and to explain the same run-to-run variance rather than presenting one run's numbers as if they were exact. No code changes; make test still passes (all 6 binaries, unaffected by a documentation-only change). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
49 KiB
PackFS vs. the host filesystem: benchmark results
Status: two O(n²) findings this document originally reported have been
fixed; one further O(n²) finding is documented below and deliberately left
unfixed, with justification. The file index (UpperSnapshot) is now a
persistent treap instead of a flat sorted array (src/upper.c; see
CLAUDE.md's "Known performance characteristics") — see "Resolution"
below. pack_write's compaction-time duplicate-content elimination
(Section 9.2) was also a linear scan with the same O(n²) shape, plus a
latent correctness bug (a hash match was trusted without a memcmp
verification); both are fixed — see "Resolution #2" below. The mount
table (vfs.c's MountSnapshot) uses the same full-array-copy pattern the
file index used to, is confirmed O(n²) in mount count, and is not
rewritten — see "Finding: mount table scaling" below for why that is the
correct call, not an oversight. This document keeps every prior run's
numbers as the historical record of each problem — per this project's own
documentation standard, a correction extends the record rather than erasing
it — and adds each fix's numbers alongside them, in full, not a curated
subset of either. Read "Resolution," "Resolution #2," and "Finding: mount
table scaling" below before drawing conclusions from the "Before" tables,
which describe a version of the index this codebase no longer has.
This document reports the results of bench/bench.c (make bench), each
run once on the environment described below. These are single runs, not a
statistically-averaged series — treat the numbers as illustrative of shape
(which operations are faster, by roughly what factor, and why) more than as
precise absolute figures for a different machine. The methodology, including
exactly what each backend label measures, is documented in bench/bench.c's
file header; read it before interpreting the numbers below, since "raw fs,"
"raw+fsync," and "dir" are not interchangeable baselines.
The "Before" table below reproduces every one of the 45 measurements that
run's === summary (45 measurements) === block reported; the "After" table
reproduces every one of the 54 measurements the current bench/bench.c
reports (the original 45, unchanged in category, plus 9 new mount-table-
scaling measurements added after the original run — see "Finding: mount
table scaling"). None omitted, none combined in either table, and the
complete verbatim program output for both runs, exactly as captured, is
preserved in the two appendices at the end as the primary source record.
The pack_write dedup fix ("Resolution #2") is measured separately, by
calling pack_write directly rather than through bench/bench.c — see
that section for why.
Environment
- Host: containerized, 12 logical CPUs (AMD Ryzen 5 3600), running under an
overlayroot filesystem (overlayonoverlay, permount) — there is no separate tmpfs at/tmpin this environment; every "raw fs" anddirnumber below is against that container overlay filesystem, not a bare disk or tmpfs. This matters most for thefsyncnumbers (see below). N_SMALL= 20,000 files × 128 bytes for the metadata-heavy suite;N_DIRS= 4,000; concurrency = 8 threads × 4,000 ops; large files up to 64 MB;N_MOUNTS= 2,000 (run at N/4, N, N×4 — 500/2,000/8,000 — for the mount-table-scaling category added after the original run). Full parameters are inbench/bench.c.- Wall-clock time: 5m49s (before), 4m4s (after the treap fix, before the
mount-scaling category existed), 5m24s (current, with the mount-scaling
category added). The dominant cost throughout is the single
raw+fsync20,000-file create test (~116s in every run — this cost is on the host's storage, not PackFS, and is within measurement noise across all three runs, as expected, since that test never touches PackFS's index at all); the added mount-scaling category itself accounts for only ~3.7s of the increase from 4m4s to 5m24s, so the remainder is run-to-run variance in untimed setup/teardown, consistent with this document's stated single-run, illustrative-of-shape methodology, not a regression.
Resolution
The original run found bulk sequential create/unlink/mkdir on
mem/dir to be O(n²) in file count, because every structural write copied
the entire sorted entry array before publishing the next snapshot
(Section 5.3). concept.md itself names the trigger condition for
reconsidering that design — "a persistent (structurally shared) tree
structure is not required until this assumption is empirically violated" —
and the original run was exactly that violation, measured rather than
hypothesized.
The index was rewritten as a persistent treap (src/upper.c): a
structural write now builds only the O(log n) nodes on the path to the
change, sharing every other node (refcounted) with the snapshot it was
built from, instead of copying all n entries. The choice of treap over a
persistent AVL/red-black/weight-balanced tree, and the reasoning behind it,
is documented in src/upper.c's comment above struct TreapNode and cited
there: Seidel & Aragon, "Randomized Search Trees" (1996), for the O(1)
amortized rotation count on both insert and delete (persistent balanced
trees that need O(log n) rotations pay a real, repeated allocation cost per
rotation, which a treap avoids); Liljenzin, "Confluently Persistent Sets and
Maps" (arXiv:1301.3388), for persistent treaps specifically as an MVCC
snapshot mechanism, which is this exact use case.
Headline result, measured, same methodology as the original finding. The "After" figures below are drawn from the "After: complete results" table further down this document (the current, canonical post-fix run), not from an intermediate run kept only in this section — every number here can be found again in that table:
| Category | Before | After | Speedup |
|---|---|---|---|
| create 20,000 files (mem) | 9.5565s / 2,093 ops/s | 0.0400s / 499,809 ops/s | 239x |
| unlink 20,000 files (mem) | 10.5575s / 1,894 ops/s | 0.0179s / 1,117,742 ops/s | 590x |
| mkdir 4,000 dirs (mem) | 0.3363s / 11,896 ops/s | 0.0056s / 719,696 ops/s | 60x |
| rmdir 4,000 dirs (mem) | 0.3395s / 11,782 ops/s | 0.0040s / 993,724 ops/s | 85x |
| create 20,000 files (dir) | 12.0367s / 1,662 ops/s | 1.1418s / 17,517 ops/s | 11x |
| unlink 20,000 files (dir) | 10.1851s / 1,964 ops/s | 0.4911s / 40,724 ops/s | 21x |
| concurrent create+read+unlink (mem) | 36.7261s / 2,614 ops/s | 0.2750s / 349,127 ops/s | 134x |
(An earlier revision of this table cited a different post-fix run — same
fixed code, an earlier make bench invocation — whose numbers were 65–330x
rather than 60–590x; the two runs agree on every qualitative conclusion
below, but unlink 20,000 files (mem) in particular differs by nearly 2x
between them, 0.0320s vs 0.0179s, both measuring an operation fast enough
(sub-20ms for 20,000 entries) that ordinary system jitter in this
containerized environment has real proportional effect. This is exactly
the kind of run-to-run variance this document's stated methodology warns
about — "illustrative of shape... more than precise absolute figures" —
and the reason the table now points at one single canonical run rather
than mixing figures from two.)
mem now beats raw fs (not merely "no longer loses to it") at every one
of these: create is 25x faster than raw fs (was 9.5x slower), unlink is 27x
faster (was 22x slower), and the concurrent mixed workload is 21x faster
(was 6x slower) — the single-writer design's "optimizes for read-heavy
concurrent workloads" caveat from the original analysis (below) no longer
needs to exclude bulk-write workloads, because the write side is no longer
the bottleneck it was.
The complexity-class change is confirmed quantitatively, the same way the
original O(n²) finding was — not just "it got faster," which an unrelated
constant-factor optimization could also produce: mkdir at N=4,000 vs.
create at N=20,000 (same structural-write mechanism, 5x the N) now shows a
7.14x slowdown (0.0400s / 0.0056s). O(n) predicts 5.0x; O(n log n)
predicts 5.97x; the old O(n²) regime predicted, and measured, 25–28.4x.
7.14x lands close to the O(n log n) prediction (some excess is expected —
create writes bytes into a fresh MutCell and mkdir does not, a
constant-factor difference the pure index-complexity comparison doesn't
account for) and nowhere near quadratic — and is itself nearly identical to
the 7.15x this same check gave against the earlier run referenced above,
which is the reassurance that matters here: the ratio this check depends
on is stable across runs even where the fastest individual measurements
are not. The index's asymptotic behavior changed, not merely its constant
factor.
Full output of both runs is reproduced verbatim in the appendices below and
is reproducible again with make bench; the correctness of the new index
under heavy randomized structural churn (non-sequential creates/deletes/
renames, nested directories, thousands of entries, cross-checked against an
independent reference model — not just "the benchmark ran fast without
crashing") is tests/test_index_stress.c, part of make test.
Resolution #2: pack_write compaction dedup
A second, independent O(n²) finding, found while auditing the codebase for
other instances of the same shape of bug after "Resolution" above: src/ pack.c's pack_write implements Section 9.2's exact-duplicate elimination
(two entries with byte-identical content share one data_off in the
compacted pack) as a linear scan, for each entry, of every previously-seen
(hash, size) pair. That is O(n) per entry, O(n²) total across n distinct
blobs — the same complexity-class bug "Resolution" fixed in the file index,
here in the compaction writer instead. It went unnoticed by the original
benchmark for a specific, identifiable reason: bench/bench.c's compaction
test (category 8, "compact N entries to pack") writes byte-identical content
into every file, so every entry's linear scan matched on the very first
comparison (against entry 0) and never actually scanned — the benchmark's
own test data happened to be the linear scan's best case, not a
representative one. Measuring with unique per-file content instead (each
entry's actual worst case) surfaced it:
| N (unique entries) | Time |
|---|---|
| 5,000 | 0.0138s |
| 20,000 | 0.1114s |
| 80,000 | 1.7389s |
5,000 → 20,000 is a 4x step in N; time increased 8.07x. 20,000 → 80,000 is also a 4x step; time increased 15.61x, converging toward O(n²)'s 16x prediction as N grows (the smaller first ratio reflects fixed per-call overhead not yet dominated by the quadratic term) — the same convergence pattern "Resolution" observed for the file index, and the same standard used there to distinguish a real complexity-class finding from noise.
This scan had a second, latent problem beyond speed: it compared only
(hash, size), never the actual bytes, before deciding two entries were
duplicates and sharing their data_off. FNV-1a64 (src/hash.c) is
explicitly not collision-resistant — it is a fast checksum, not a
cryptographic hash — so two different-content entries that happened to
collide on (hash, size) would have been silently merged into one blob,
corrupting one of them. This was never observed in practice (no test's data
happened to collide), but was true of the code as written, not a
hypothetical: the fix required for the O(n²) issue also closes it.
Fix: the linear scan was replaced with an open-addressing hash table
(load factor 1/2, linear probing) over (hash, size), plus — critically —
an actual memcmp against the candidate's stored data pointer before ever
reusing a data_off, so a hash match is only ever a candidate, never
accepted as proof of equality. Comment and full listing:
src/pack.c's DedupSlot struct and the rewritten first pass of
pack_write.
After the fix, same methodology:
| N (unique entries) | Time | vs. before |
|---|---|---|
| 5,000 | 0.0079s | 1.75x faster |
| 20,000 | 0.0131s | 8.50x faster |
| 80,000 | 0.0439s | 39.6x faster |
5,000 → 20,000 now shows 1.66x (O(n) predicts 4x, but a hash table's dominant cost at these sizes is close to constant per insert, so sub-linear growth here is expected, not suspicious); 20,000 → 80,000 shows 3.35x, close to O(n)'s 4x prediction and nowhere near O(n²)'s 16x. The complexity class changed, exactly as with the file index.
This fix is not reflected in the "After" table/Appendix B below, because
that make bench run predates it — but it does not need to be: category
8's "compact N entries to pack" measurement uses identical content (as
explained above), which was already the linear scan's O(1)-per-entry best
case before this fix and is unaffected by the fix (a hash table lookup that
matches on the first probe is also O(1) per entry); the number in the
"After" table (0.0267s, 20,000 identical-content entries) remains accurate
and does not need to be re-measured. The fix matters for compacting many
files with distinct content, which bench/bench.c's existing compaction
test does not exercise — hence the separate, direct pack_write
measurement above rather than a bench/bench.c change. Correctness is
covered permanently by tests/test_pack_overlay.c (asserts /a.txt and
/dup.txt, identical content, share one data_off, while /b.txt,
different content, shares neither's — catching both "dedup stopped
happening" and "dedup over-matched"); the performance fix is covered
permanently by tests/test_pack_write_perf.c (10,000 unique entries under
a generous time bound, a regression tripwire rather than a tight
assertion, run as part of make test).
Finding: mount table scaling (O(n²), documented, deliberately not fixed)
src/vfs.c's mount table (MountSnapshot) uses the exact same pattern the
file index used before "Resolution" above: every vfs_mount/vfs_unmount
copies the entire array of mounted backends before publishing the next
snapshot (Section 5.3's single-writer model applies to the mount table too,
per concept.md's explicit statement that "the mount table is not separate
global state exempt from this model"). This is confirmed O(n²) in mount
count by the bench/bench.c category 10 measurements (full numbers in the
"After" table below and Appendix B):
| N mounts | mount | resolve (stat through N mounts) | unmount |
|---|---|---|---|
| 500 | 0.0054s | 0.0020s | 0.0046s |
| 2,000 | 0.0831s | 0.0296s | 0.0758s |
| 8,000 | 1.4739s | 0.4938s | 1.5172s |
500 → 2,000 is a 4x step in N: mount increased 15.4x, resolve 14.8x,
unmount 16.5x. 2,000 → 8,000 is also a 4x step: mount increased 17.7x,
resolve 16.7x, unmount 20.0x. All six ratios cluster around O(n²)'s 16x
prediction for a 4x-N step, and nowhere near O(n)'s 4x — the same
complexity-class-confirmation methodology used for both findings above,
applied here and yielding the same verdict. Notably, resolve (a single
vfs_stat through an already-built N-mount table) is also O(n²) across N
resolves, not just O(n) per resolve as might be assumed — consistent with
longest-prefix-match dispatch being a linear scan of the mount table per
call (Section 3's dispatch design), which is itself O(n) per call and O(n²)
across N calls, independent of the snapshot-copy cost on the write side.
This is deliberately not fixed, and the reasoning is the same trigger
condition concept.md Section 5.3 names for the file index — inverted:
"a persistent (structurally shared) tree structure is not required until
this assumption is empirically violated." The file index's assumption was
violated because file counts are workload-driven and can plausibly reach
tens of thousands (or more) at runtime, chosen by whatever program embeds
PackFS and whatever data it manages. Mount counts are different in kind,
not just degree: a mount point is created by a call to vfs_mount written
into a program's own source code, at a location the program's author
chose — it is bounded by how many lines of "mount this backend at this
path" a person or a build script is willing to write, not by user data,
input size, or any other externally-driven quantity. A program with 8,000
mount points, structured that way on purpose, is not a realistic workload
this system needs to serve well; a program with 8,000 files is an
ordinary one. Converting the mount table to a persistent treap would fix
this at the cost of real, permanent complexity (an additional generic
persistent-tree instantiation, or a second copy of the treap logic
specialized to prefix matching) for a case that has not been and is not
expected to be hit. Should a future use case genuinely need thousands of
mount points at runtime — for instance, one mount per user in a
multi-tenant embedding — this finding, and the numbers above, are the
starting point for that decision; until then, this is recorded as a known,
measured, and consciously accepted limitation, not an unexamined one.
Before: complete results (pre-fix; kept as the historical record — see "Resolution" above)
All 45 measurements from that run's summary block, none omitted or combined.
| Category | Backend | Time | Throughput | MB/s |
|---|---|---|---|---|
| create 20,000 files | mem | 9.5565s | 2,093 ops/s | 0.3 |
| create 20,000 files | raw fs | 1.0083s | 19,836 ops/s | 2.4 |
| create 20,000 files | raw+fsync | 116.1030s | 172 ops/s | 0.0 |
| read 20,000 files | mem | 0.0071s | 2,798,731 ops/s | 341.6 |
| read 20,000 files | raw fs | 0.1611s | 124,122 ops/s | 15.2 |
| stat 20,000 files | mem | 0.0062s | 3,223,599 ops/s | — |
| stat 20,000 files | raw fs | 0.0703s | 284,465 ops/s | — |
| readdir (20,000 entries) | mem | 0.0037s | 5,397,828 /s | — |
| readdir (20,000 entries) | raw fs | 0.0071s | 2,831,904 /s | — |
| create 20,000 files | dir | 12.0367s | 1,662 ops/s | 0.2 |
| read 20,000 files | dir | 0.1674s | 119,468 ops/s | 14.6 |
| stat 20,000 files | dir | 0.1516s | 131,937 ops/s | — |
| readdir (20,000 entries) | dir | 0.0028s | 7,058,003 /s | — |
| unlink 20,000 files | mem | 10.5575s | 1,894 ops/s | — |
| unlink 20,000 files | dir | 10.1851s | 1,964 ops/s | — |
| unlink 20,000 files | raw fs | 0.4782s | 41,823 ops/s | — |
| mkdir 4,000 dirs | mem | 0.3363s | 11,896 ops/s | — |
| rmdir 4,000 dirs | mem | 0.3395s | 11,782 ops/s | — |
| mkdir 4,000 dirs | dir | 0.5109s | 7,829 ops/s | — |
| rmdir 4,000 dirs | dir | 0.4892s | 8,177 ops/s | — |
| mkdir 4,000 dirs | raw fs | 0.1591s | 25,143 ops/s | — |
| rmdir 4,000 dirs | raw fs | 0.2119s | 18,876 ops/s | — |
| write 1MB | mem | 0.0001s | 6,687 ops/s | 6,687.1 |
| read 1MB | mem | 0.0000s | 50,941 ops/s | 50,941.4 |
| write 16MB | mem | 0.0138s | 72 ops/s | 1,157.8 |
| read 16MB | mem | 0.0010s | 1,015 ops/s | 16,240.8 |
| write 64MB | mem | 0.0566s | 18 ops/s | 1,129.8 |
| read 64MB | mem | 0.0034s | 292 ops/s | 18,705.5 |
| write 1MB | raw fs | 0.0004s | 2,685 ops/s | 2,685.4 |
| read 1MB | raw fs | 0.0001s | 16,584 ops/s | 16,583.7 |
| write 16MB | raw fs | 0.0047s | 213 ops/s | 3,412.8 |
| read 16MB | raw fs | 0.0010s | 960 ops/s | 15,363.8 |
| write 64MB | raw fs | 0.0186s | 54 ops/s | 3,435.6 |
| read 64MB | raw fs | 0.0047s | 213 ops/s | 13,635.3 |
| write 1MB | raw+fsync | 0.0087s | 116 ops/s | 115.5 |
| read 1MB | raw+fsync | 0.0001s | 8,857 ops/s | 8,857.4 |
| write 16MB | raw+fsync | 0.0840s | 12 ops/s | 190.5 |
| read 16MB | raw+fsync | 0.0013s | 796 ops/s | 12,737.8 |
| write 64MB | raw+fsync | 0.0891s | 11 ops/s | 718.3 |
| read 64MB | raw+fsync | 0.0045s | 220 ops/s | 14,096.3 |
| compact 20,000 entries to pack | pack | 0.0291s | 687,097 ops/s | 83.9 |
| random-read 20,000 entries | pack (mmap'd) | 0.0085s | 2,350,224 ops/s | 286.9 |
| random-read 20,000 entries | raw fs | 0.1767s | 113,198 ops/s | 13.8 |
| concurrent create+read+unlink (8×4,000×3) | mem | 36.7261s | 2,614 ops/s | — |
| concurrent create+read+unlink (8×4,000×3) | raw fs | 6.1202s | 15,686 ops/s | — |
After: complete results (post-fix, current)
All 54 measurements from the current bench/bench.c's summary block, none
omitted or combined: the original 45 (re-run, post-treap-fix) plus 9 new
mount-table-scaling measurements (category 10, added after the original
run — see "Finding: mount table scaling" above). The pack_write dedup fix
("Resolution #2" above) postdates this run but does not change any number
in it — see that section for why the compaction row below is unaffected.
| Category | Backend | Time | Throughput | MB/s |
|---|---|---|---|---|
| create 20,000 files | mem | 0.0400s | 499,809 ops/s | 61.0 |
| create 20,000 files | raw fs | 0.9943s | 20,114 ops/s | 2.5 |
| create 20,000 files | raw+fsync | 116.2147s | 172 ops/s | 0.0 |
| read 20,000 files | mem | 0.0099s | 2,011,901 ops/s | 245.6 |
| read 20,000 files | raw fs | 0.1606s | 124,543 ops/s | 15.2 |
| stat 20,000 files | mem | 0.0077s | 2,598,241 ops/s | — |
| stat 20,000 files | raw fs | 0.0699s | 286,222 ops/s | — |
| readdir (20,000 entries) | mem | 0.0041s | 4,845,826 /s | — |
| readdir (20,000 entries) | raw fs | 0.0069s | 2,894,815 /s | — |
| create 20,000 files | dir | 1.1418s | 17,517 ops/s | 2.1 |
| read 20,000 files | dir | 0.1530s | 130,705 ops/s | 16.0 |
| stat 20,000 files | dir | 0.1379s | 145,056 ops/s | — |
| readdir (20,000 entries) | dir | 0.0037s | 5,465,162 /s | — |
| unlink 20,000 files | mem | 0.0179s | 1,117,742 ops/s | — |
| unlink 20,000 files | dir | 0.4911s | 40,724 ops/s | — |
| unlink 20,000 files | raw fs | 0.4746s | 42,141 ops/s | — |
| mkdir 4,000 dirs | mem | 0.0056s | 719,696 ops/s | — |
| rmdir 4,000 dirs | mem | 0.0040s | 993,724 ops/s | — |
| mkdir 4,000 dirs | dir | 0.1710s | 23,386 ops/s | — |
| rmdir 4,000 dirs | dir | 0.1257s | 31,827 ops/s | — |
| mkdir 4,000 dirs | raw fs | 0.1374s | 29,122 ops/s | — |
| rmdir 4,000 dirs | raw fs | 0.0951s | 42,043 ops/s | — |
| write 1MB | mem | 0.0001s | 7,058 ops/s | 7,058.1 |
| read 1MB | mem | 0.0000s | 50,051 ops/s | 50,050.9 |
| write 16MB | mem | 0.0110s | 91 ops/s | 1,458.2 |
| read 16MB | mem | 0.0009s | 1,058 ops/s | 16,920.1 |
| write 64MB | mem | 0.0586s | 17 ops/s | 1,091.3 |
| read 64MB | mem | 0.0032s | 313 ops/s | 20,056.5 |
| write 1MB | raw fs | 0.0004s | 2,759 ops/s | 2,759.4 |
| read 1MB | raw fs | 0.0001s | 16,812 ops/s | 16,812.4 |
| write 16MB | raw fs | 0.0044s | 229 ops/s | 3,656.4 |
| read 16MB | raw fs | 0.0010s | 989 ops/s | 15,827.9 |
| write 64MB | raw fs | 0.0186s | 54 ops/s | 3,449.7 |
| read 64MB | raw fs | 0.0047s | 213 ops/s | 13,633.1 |
| write 1MB | raw+fsync | 0.0180s | 56 ops/s | 55.7 |
| read 1MB | raw+fsync | 0.0001s | 9,059 ops/s | 9,058.8 |
| write 16MB | raw+fsync | 0.0231s | 43 ops/s | 691.7 |
| read 16MB | raw+fsync | 0.0012s | 855 ops/s | 13,672.2 |
| write 64MB | raw+fsync | 0.0803s | 12 ops/s | 797.2 |
| read 64MB | raw+fsync | 0.0048s | 209 ops/s | 13,393.5 |
| compact 20,000 entries to pack | pack | 0.0267s | 747,839 ops/s | 91.3 |
| random-read 20,000 entries | pack (mmap'd) | 0.0082s | 2,435,930 ops/s | 297.4 |
| random-read 20,000 entries | raw fs | 0.1629s | 122,749 ops/s | 15.0 |
| concurrent create+read+unlink (8×4,000×3) | mem | 0.2750s | 349,127 ops/s | — |
| concurrent create+read+unlink (8×4,000×3) | raw fs | 5.6498s | 16,992 ops/s | — |
| mount 500 backends | vfs | 0.0054s | 92,833 ops/s | — |
| resolve, 500 mounts | vfs | 0.0020s | 246,064 ops/s | — |
| unmount 500 backends | vfs | 0.0046s | 109,207 ops/s | — |
| mount 2,000 backends | vfs | 0.0831s | 24,081 ops/s | — |
| resolve, 2,000 mounts | vfs | 0.0296s | 67,458 ops/s | — |
| unmount 2,000 backends | vfs | 0.0758s | 26,397 ops/s | — |
| mount 8,000 backends | vfs | 1.4739s | 5,428 ops/s | — |
| resolve, 8,000 mounts | vfs | 0.4938s | 16,200 ops/s | — |
| unmount 8,000 backends | vfs | 1.5172s | 5,273 ops/s | — |
Everything not involving bulk mem/dir structural writes (reads, stat,
readdir, large sequential I/O, pack random-access, compaction) is
unchanged within normal run-to-run noise, as the row-by-row comparison
against the "Before" table shows — the treap rewrite only touches the
structural-write path (Section 5.3), not content reads, content writes, or
the pack format. The raw+fsync create row is unchanged within noise
(116.1030s before, 116.2147s here) precisely because it never touches
PackFS at all. The nine new mount-table rows have no "Before" counterpart —
the mount table was never rewritten, so there is no pre/post comparison to
make for them; they are current-state measurements only, discussed in
"Finding: mount table scaling" above.
Analysis
The bullets below are checked against the current "After" table (the 54-measurement run) rather than preserved unchanged from an earlier draft — an earlier revision of this section cited exact figures from the prior 45-measurement after-run, which drifted out of sync with the "After" table once that table was replaced with the newer run; the numbers below were recomputed from the table actually in this document, not carried over.
Where PackFS wins clearly
- Reads of small files: 15–16x faster than both raw fs and the
dirbackend (2.0M ops/s vs ~125–131K ops/s). No syscall per read —memreads are an expected-O(log n) treap lookup plus amemcpyout of an already-resident buffer (Section 5.3, 9.1). stat: ~9x faster than raw fs (2.6M ops/s vs 286K ops/s). Same reason — no syscall, and the index is sorted for expected-O(log n) lookup rather than requiring a directory entry scan.- Random-access reads against a compacted pack: ~20x faster than raw fs
(2.4M ops/s vs 123K ops/s), because the pack is one
mmap'd file with a binary-searchable index (Section 9.1), against 20,000 individualopen/read/closesyscall triples on the raw-fs side. This is the single result that most directly validates the architecture's stated purpose — Section 3.2's claim that "random access is a flat table lookup, not a linear scan or a directory-parse-then-seek" — under an actual measured workload, not just by construction. - Compaction throughput: ~91 MB/s / ~748K entries/s to serialize the live tree into a fresh pack (Section 4.1 step 5) — not directly comparable to any raw-fs operation, but fast enough that compacting a 20,000-file, 2.5 MB tree is not a practically-felt pause (27ms).
- Bulk create/unlink/mkdir/rmdir: RESOLVED, now 11–590x faster than the
pre-fix
mem/dirnumbers, and (formem) 21–27x faster than raw fs — see "Resolution" above; this used to be PackFS's clearest loss and is now among its clearest wins. - Concurrent mixed create+read+unlink: RESOLVED, now 21x faster than raw fs (was 6x slower before the fix) — see "Resolution."
pack_writecompaction dedup on distinct-content files: RESOLVED, 1.75– 39.6x faster depending on N, plus a latent hash-collision correctness bug closed alongside it — see "Resolution #2." Not visible in this benchmark's own compaction test (category 8), which uses identical content and so never exercised either problem; measured separately by callingpack_writedirectly.
Where raw fs still wins, and why that's expected, not a bug
dirbackendstatis ~2.0x slower than raw fsstat(145K ops/s vs 286K ops/s) — becausepfs_dir_statat(Section 6.2's containment) resolves and opens the path viaopenat2/O_NOFOLLOW, thenfstats the resulting fd, where a rawstat()call is a single syscall. This is the direct, measured cost of path containment on the metadata path, separate from and smaller than the containment cost paid onopenitself (which raw fs pays an equivalent single-syscall cost for anyway). Unaffected by the index rewrite.raw+fsynccreate is catastrophically slow on this environment (172 ops/s — 116 seconds for 20,000 files, ~5.8ms perfsync, within measurement noise across the "Before" and "After" runs — 116.1030s vs 116.2147s, a 0.1% difference), because this container's overlay filesystem has poor per-callfsynclatency. This is a property of the container, not of PackFS or of a bare disk — flagged in the environment section above precisely so it isn't misread as "PackFS's journal must be this slow too." The journal (Section 4.4) does callfsynconce per journaled write for the same durability reason rawfsyncis slow here, which is a real, inherited cost on this kind of storage — not a PackFS-specific one. Unaffected by the index rewrite (the journal's cost is per-writefsynclatency, not index maintenance).- Large writes above ~16 MB: raw fs is faster than
mem(3,450– 3,656 MB/s vs 1,091–1,458 MB/s). Section 5.7's buffer-growth discipline (allocate new, copy old + new, publish, retire old — never realloc in place) means every capacity doubling re-copies everything written so far; a plain unsyncedwrite()to a real file only ever appends new pages to the page cache, never re-copying prior ones. This is entirely independent of the index structure (Section 5.7's buffer-growth discipline governs aMutCell's content buffer, not the index the treap rewrite replaced) and is unaffected by this fix — the safety property Section 5.7 requires (no reader can ever see a freed buffer) still has a real, quantifiable cost for very large sequential writes, traded for correctness under concurrent access that a plain in-place realloc would not have.
Appendix A: complete verbatim output, before the fix
Captured exactly as produced by make bench against the pre-fix (flat
array) index; reformatted into the tables above, but reproduced here
unedited as the primary source record.
=== PackFS vs. host filesystem: benchmark ===
environment:
raw fs test root: /tmp/packfs_bench_raw_unzuwh
dir backend root: /tmp/packfs_bench_dir_MXB4CT
small files (N): 20000, 128 bytes each
directories (N): 4000
concurrency: 8 threads x 4000 ops (create+read+unlink)
large-file sizes: 1MB 16MB 64MB
NOTE: see the file header for what "raw fs" vs "raw+fsync" vs
"dir" actually measure — they are not interchangeable.
== 1. create N small files ==
create N small files mem 9.5565s 2093 ops/s 0.3 MB/s
create N small files raw fs 1.0083s 19836 ops/s 2.4 MB/s
create N small files raw+fsync 116.1030s 172 ops/s 0.0 MB/s
== 2. read N small files ==
read N small files mem 0.0071s 2798731 ops/s 341.6 MB/s
read N small files raw fs 0.1611s 124122 ops/s 15.2 MB/s
== 3. stat N files ==
stat N files mem 0.0062s 3223599 ops/s
stat N files raw fs 0.0703s 284465 ops/s
== 4. readdir ==
readdir (N entries) mem 0.0037s 5397828 ops/s
readdir (N entries) raw fs 0.0071s 2831904 ops/s
== 1b. create N small files (dir backend vs. what it wraps) ==
create N small files dir 12.0367s 1662 ops/s 0.2 MB/s
== 2b. read N small files (dir backend) ==
read N small files dir 0.1674s 119468 ops/s 14.6 MB/s
== 3b. stat N files (dir backend) ==
stat N files dir 0.1516s 131937 ops/s
== 4b. readdir (dir backend) ==
readdir (N entries) dir 0.0028s 7058003 ops/s
== 5. unlink N files ==
unlink N files mem 10.5575s 1894 ops/s
unlink N files dir 10.1851s 1964 ops/s
unlink N files raw fs 0.4782s 41823 ops/s
== 6. mkdir/rmdir N directories ==
mkdir N dirs mem 0.3363s 11896 ops/s
rmdir N dirs mem 0.3395s 11782 ops/s
mkdir N dirs dir 0.5109s 7829 ops/s
rmdir N dirs dir 0.4892s 8177 ops/s
mkdir N dirs raw fs 0.1591s 25143 ops/s
rmdir N dirs raw fs 0.2119s 18876 ops/s
== 7. large sequential write/read ==
write 1MB mem 0.0001s 6687 ops/s 6687.1 MB/s
read 1MB mem 0.0000s 50941 ops/s 50941.4 MB/s
write 16MB mem 0.0138s 72 ops/s 1157.8 MB/s
read 16MB mem 0.0010s 1015 ops/s 16240.8 MB/s
write 64MB mem 0.0566s 18 ops/s 1129.8 MB/s
read 64MB mem 0.0034s 292 ops/s 18705.5 MB/s
write 1MB raw fs 0.0004s 2685 ops/s 2685.4 MB/s
read 1MB raw fs 0.0001s 16584 ops/s 16583.7 MB/s
write 16MB raw fs 0.0047s 213 ops/s 3412.8 MB/s
read 16MB raw fs 0.0010s 960 ops/s 15363.8 MB/s
write 64MB raw fs 0.0186s 54 ops/s 3435.6 MB/s
read 64MB raw fs 0.0047s 213 ops/s 13635.3 MB/s
write 1MB raw+fsync 0.0087s 116 ops/s 115.5 MB/s
read 1MB raw+fsync 0.0001s 8857 ops/s 8857.4 MB/s
write 16MB raw+fsync 0.0840s 12 ops/s 190.5 MB/s
read 16MB raw+fsync 0.0013s 796 ops/s 12737.8 MB/s
write 64MB raw+fsync 0.0891s 11 ops/s 718.3 MB/s
read 64MB raw+fsync 0.0045s 220 ops/s 14096.3 MB/s
== 8. random-access read: mmap'd pack vs. raw fs ==
compact N entries to pack pack 0.0291s 687097 ops/s 83.9 MB/s
random-read N entries pack 0.0085s 2350224 ops/s 286.9 MB/s
random-read N entries raw fs 0.1767s 113198 ops/s 13.8 MB/s
== 9. concurrency (create+read+unlink) ==
concurrent create+read+unlink mem 36.7261s 2614 ops/s
concurrent create+read+unlink raw fs 6.1202s 15686 ops/s
=== summary (45 measurements) ===
category backend seconds throughput MB/s
create N small files mem 9.5565s 2093 ops/s 0.3 MB/s
create N small files raw fs 1.0083s 19836 ops/s 2.4 MB/s
create N small files raw+fsync 116.1030s 172 ops/s 0.0 MB/s
read N small files mem 0.0071s 2798731 ops/s 341.6 MB/s
read N small files raw fs 0.1611s 124122 ops/s 15.2 MB/s
stat N files mem 0.0062s 3223599 ops/s
stat N files raw fs 0.0703s 284465 ops/s
readdir (N entries) mem 0.0037s 5397828 ops/s
readdir (N entries) raw fs 0.0071s 2831904 ops/s
create N small files dir 12.0367s 1662 ops/s 0.2 MB/s
read N small files dir 0.1674s 119468 ops/s 14.6 MB/s
stat N files dir 0.1516s 131937 ops/s
readdir (N entries) dir 0.0028s 7058003 ops/s
unlink N files mem 10.5575s 1894 ops/s
unlink N files dir 10.1851s 1964 ops/s
unlink N files raw fs 0.4782s 41823 ops/s
mkdir N dirs mem 0.3363s 11896 ops/s
rmdir N dirs mem 0.3395s 11782 ops/s
mkdir N dirs dir 0.5109s 7829 ops/s
rmdir N dirs dir 0.4892s 8177 ops/s
mkdir N dirs raw fs 0.1591s 25143 ops/s
rmdir N dirs raw fs 0.2119s 18876 ops/s
write 1MB mem 0.0001s 6687 ops/s 6687.1 MB/s
read 1MB mem 0.0000s 50941 ops/s 50941.4 MB/s
write 16MB mem 0.0138s 72 ops/s 1157.8 MB/s
read 16MB mem 0.0010s 1015 ops/s 16240.8 MB/s
write 64MB mem 0.0566s 18 ops/s 1129.8 MB/s
read 64MB mem 0.0034s 292 ops/s 18705.5 MB/s
write 1MB raw fs 0.0004s 2685 ops/s 2685.4 MB/s
read 1MB raw fs 0.0001s 16584 ops/s 16583.7 MB/s
write 16MB raw fs 0.0047s 213 ops/s 3412.8 MB/s
read 16MB raw fs 0.0010s 960 ops/s 15363.8 MB/s
write 64MB raw fs 0.0186s 54 ops/s 3435.6 MB/s
read 64MB raw fs 0.0047s 213 ops/s 13635.3 MB/s
write 1MB raw+fsync 0.0087s 116 ops/s 115.5 MB/s
read 1MB raw+fsync 0.0001s 8857 ops/s 8857.4 MB/s
write 16MB raw+fsync 0.0840s 12 ops/s 190.5 MB/s
read 16MB raw+fsync 0.0013s 796 ops/s 12737.8 MB/s
write 64MB raw+fsync 0.0891s 11 ops/s 718.3 MB/s
read 64MB raw+fsync 0.0045s 220 ops/s 14096.3 MB/s
compact N entries to pack pack 0.0291s 687097 ops/s 83.9 MB/s
random-read N entries pack 0.0085s 2350224 ops/s 286.9 MB/s
random-read N entries raw fs 0.1767s 113198 ops/s 13.8 MB/s
concurrent create+read+unlink mem 36.7261s 2614 ops/s
concurrent create+read+unlink raw fs 6.1202s 15686 ops/s
real 5m49.507s
user 2m9.600s
sys 0m20.941s
Appendix B: complete verbatim output, after the fix (current, 54 measurements)
Captured exactly as produced by the current make bench (post-treap-fix
index, with the mount-table-scaling category added); reformatted into the
tables above, but reproduced here unedited as the primary source record.
This run predates the pack_write dedup fix ("Resolution #2"), which does
not change any number in it — see that section for why.
=== PackFS vs. host filesystem: benchmark ===
environment:
raw fs test root: /tmp/packfs_bench_raw_PnMAT7
dir backend root: /tmp/packfs_bench_dir_U8Vgxi
small files (N): 20000, 128 bytes each
directories (N): 4000
concurrency: 8 threads x 4000 ops (create+read+unlink)
mounts (N): 2000
large-file sizes: 1MB 16MB 64MB
NOTE: see the file header for what "raw fs" vs "raw+fsync" vs
"dir" actually measure — they are not interchangeable.
== 1. create N small files ==
create N small files mem 0.0400s 499809 ops/s 61.0 MB/s
create N small files raw fs 0.9943s 20114 ops/s 2.5 MB/s
create N small files raw+fsync 116.2147s 172 ops/s 0.0 MB/s
== 2. read N small files ==
read N small files mem 0.0099s 2011901 ops/s 245.6 MB/s
read N small files raw fs 0.1606s 124543 ops/s 15.2 MB/s
== 3. stat N files ==
stat N files mem 0.0077s 2598241 ops/s
stat N files raw fs 0.0699s 286222 ops/s
== 4. readdir ==
readdir (N entries) mem 0.0041s 4845826 ops/s
readdir (N entries) raw fs 0.0069s 2894815 ops/s
== 1b. create N small files (dir backend vs. what it wraps) ==
create N small files dir 1.1418s 17517 ops/s 2.1 MB/s
== 2b. read N small files (dir backend) ==
read N small files dir 0.1530s 130705 ops/s 16.0 MB/s
== 3b. stat N files (dir backend) ==
stat N files dir 0.1379s 145056 ops/s
== 4b. readdir (dir backend) ==
readdir (N entries) dir 0.0037s 5465162 ops/s
== 5. unlink N files ==
unlink N files mem 0.0179s 1117742 ops/s
unlink N files dir 0.4911s 40724 ops/s
unlink N files raw fs 0.4746s 42141 ops/s
== 6. mkdir/rmdir N directories ==
mkdir N dirs mem 0.0056s 719696 ops/s
rmdir N dirs mem 0.0040s 993724 ops/s
mkdir N dirs dir 0.1710s 23386 ops/s
rmdir N dirs dir 0.1257s 31827 ops/s
mkdir N dirs raw fs 0.1374s 29122 ops/s
rmdir N dirs raw fs 0.0951s 42043 ops/s
== 7. large sequential write/read ==
write 1MB mem 0.0001s 7058 ops/s 7058.1 MB/s
read 1MB mem 0.0000s 50051 ops/s 50050.9 MB/s
write 16MB mem 0.0110s 91 ops/s 1458.2 MB/s
read 16MB mem 0.0009s 1058 ops/s 16920.1 MB/s
write 64MB mem 0.0586s 17 ops/s 1091.3 MB/s
read 64MB mem 0.0032s 313 ops/s 20056.5 MB/s
write 1MB raw fs 0.0004s 2759 ops/s 2759.4 MB/s
read 1MB raw fs 0.0001s 16812 ops/s 16812.4 MB/s
write 16MB raw fs 0.0044s 229 ops/s 3656.4 MB/s
read 16MB raw fs 0.0010s 989 ops/s 15827.9 MB/s
write 64MB raw fs 0.0186s 54 ops/s 3449.7 MB/s
read 64MB raw fs 0.0047s 213 ops/s 13633.1 MB/s
write 1MB raw+fsync 0.0180s 56 ops/s 55.7 MB/s
read 1MB raw+fsync 0.0001s 9059 ops/s 9058.8 MB/s
write 16MB raw+fsync 0.0231s 43 ops/s 691.7 MB/s
read 16MB raw+fsync 0.0012s 855 ops/s 13672.2 MB/s
write 64MB raw+fsync 0.0803s 12 ops/s 797.2 MB/s
read 64MB raw+fsync 0.0048s 209 ops/s 13393.5 MB/s
== 8. random-access read: mmap'd pack vs. raw fs ==
compact N entries to pack pack 0.0267s 747839 ops/s 91.3 MB/s
random-read N entries pack 0.0082s 2435930 ops/s 297.4 MB/s
random-read N entries raw fs 0.1629s 122749 ops/s 15.0 MB/s
== 9. concurrency (create+read+unlink) ==
concurrent create+read+unlink mem 0.2750s 349127 ops/s
concurrent create+read+unlink raw fs 5.6498s 16992 ops/s
== 10. mount table scaling (the mount table uses the same full-snapshot-copy pattern the file index used to) ==
mount 500 backends vfs 0.0054s 92833 ops/s
resolve, 500 mounts vfs 0.0020s 246064 ops/s
unmount 500 backends vfs 0.0046s 109207 ops/s
mount 2000 backends vfs 0.0831s 24081 ops/s
resolve, 2000 mounts vfs 0.0296s 67458 ops/s
unmount 2000 backends vfs 0.0758s 26397 ops/s
mount 8000 backends vfs 1.4739s 5428 ops/s
resolve, 8000 mounts vfs 0.4938s 16200 ops/s
unmount 8000 backends vfs 1.5172s 5273 ops/s
=== summary (54 measurements) ===
category backend seconds throughput MB/s
create N small files mem 0.0400s 499809 ops/s 61.0 MB/s
create N small files raw fs 0.9943s 20114 ops/s 2.5 MB/s
create N small files raw+fsync 116.2147s 172 ops/s 0.0 MB/s
read N small files mem 0.0099s 2011901 ops/s 245.6 MB/s
read N small files raw fs 0.1606s 124543 ops/s 15.2 MB/s
stat N files mem 0.0077s 2598241 ops/s
stat N files raw fs 0.0699s 286222 ops/s
readdir (N entries) mem 0.0041s 4845826 ops/s
readdir (N entries) raw fs 0.0069s 2894815 ops/s
create N small files dir 1.1418s 17517 ops/s 2.1 MB/s
read N small files dir 0.1530s 130705 ops/s 16.0 MB/s
stat N files dir 0.1379s 145056 ops/s
readdir (N entries) dir 0.0037s 5465162 ops/s
unlink N files mem 0.0179s 1117742 ops/s
unlink N files dir 0.4911s 40724 ops/s
unlink N files raw fs 0.4746s 42141 ops/s
mkdir N dirs mem 0.0056s 719696 ops/s
rmdir N dirs mem 0.0040s 993724 ops/s
mkdir N dirs dir 0.1710s 23386 ops/s
rmdir N dirs dir 0.1257s 31827 ops/s
mkdir N dirs raw fs 0.1374s 29122 ops/s
rmdir N dirs raw fs 0.0951s 42043 ops/s
write 1MB mem 0.0001s 7058 ops/s 7058.1 MB/s
read 1MB mem 0.0000s 50051 ops/s 50050.9 MB/s
write 16MB mem 0.0110s 91 ops/s 1458.2 MB/s
read 16MB mem 0.0009s 1058 ops/s 16920.1 MB/s
write 64MB mem 0.0586s 17 ops/s 1091.3 MB/s
read 64MB mem 0.0032s 313 ops/s 20056.5 MB/s
write 1MB raw fs 0.0004s 2759 ops/s 2759.4 MB/s
read 1MB raw fs 0.0001s 16812 ops/s 16812.4 MB/s
write 16MB raw fs 0.0044s 229 ops/s 3656.4 MB/s
read 16MB raw fs 0.0010s 989 ops/s 15827.9 MB/s
write 64MB raw fs 0.0186s 54 ops/s 3449.7 MB/s
read 64MB raw fs 0.0047s 213 ops/s 13633.1 MB/s
write 1MB raw+fsync 0.0180s 56 ops/s 55.7 MB/s
read 1MB raw+fsync 0.0001s 9059 ops/s 9058.8 MB/s
write 16MB raw+fsync 0.0231s 43 ops/s 691.7 MB/s
read 16MB raw+fsync 0.0012s 855 ops/s 13672.2 MB/s
write 64MB raw+fsync 0.0803s 12 ops/s 797.2 MB/s
read 64MB raw+fsync 0.0048s 209 ops/s 13393.5 MB/s
compact N entries to pack pack 0.0267s 747839 ops/s 91.3 MB/s
random-read N entries pack 0.0082s 2435930 ops/s 297.4 MB/s
random-read N entries raw fs 0.1629s 122749 ops/s 15.0 MB/s
concurrent create+read+unlink mem 0.2750s 349127 ops/s
concurrent create+read+unlink raw fs 5.6498s 16992 ops/s
mount 500 backends vfs 0.0054s 92833 ops/s
resolve, 500 mounts vfs 0.0020s 246064 ops/s
unmount 500 backends vfs 0.0046s 109207 ops/s
mount 2000 backends vfs 0.0831s 24081 ops/s
resolve, 2000 mounts vfs 0.0296s 67458 ops/s
unmount 2000 backends vfs 0.0758s 26397 ops/s
mount 8000 backends vfs 1.4739s 5428 ops/s
resolve, 8000 mounts vfs 0.4938s 16200 ops/s
unmount 8000 backends vfs 1.5172s 5273 ops/s
real 5m23.896s
user 0m4.732s
sys 0m19.047s
Reproducing
make bench
Takes several minutes on a similarly slow-fsync environment, dominated
by the raw+fsync create test; expect well under a minute on a host with
normal disk or tmpfs fsync latency.