Files
packfs/BENCH.md
T
retoorandClaude Sonnet 5 64283d587b Fix the O(n^2) bulk create/unlink: flat array -> persistent treap
BENCH.md's earlier finding was real: every structural write (create/
unlink/mkdir) copied the entire sorted UpperSnapshot entry array before
publishing the next snapshot, making bulk sequential creation O(n^2).
concept.md Section 5.3 names the exact condition for reconsidering this
("a persistent structurally-shared tree structure is not required
until this assumption is empirically violated") -- that condition was
measured, not hypothesized, so this closes it rather than leaving it
as a documented-but-open limitation.

The index is now a persistent treap (src/upper.c): a structural write
copies only the O(log n) nodes on the path to the change, sharing
every other node (its own refcount, cascading like MutCell's and
UpperSnapshot's) with whichever snapshot(s) it was built from. Chosen
over a persistent AVL/red-black/weight-balanced tree because deletion
in those can need O(log n) rebalancing rotations -- each a real
allocation in a persistent setting -- where a treap needs only O(1)
amortized rotations for insert and delete (Seidel & Aragon 1996), with
expected O(log n) height regardless of insertion order, including the
sorted-by-creation-order pattern that made the flat array quadratic in
the first place. Liljenzin's "Confluently Persistent Sets and Maps"
(arXiv:1301.3388) documents persistent treaps giving O(1) snapshots
for MVCC specifically, which is this exact use case. Full reasoning
and the ownership convention (functions consume one ref of their tree
arguments, return one owned ref) are in upper.c's comment above
struct TreapNode.

Blast radius kept deliberately small: snapshot_upsert/snapshot_remove
keep their exact original signatures, so upper_create/upper_mkdir/
upper_remove/upper_rename/upper_copy_up needed zero changes.
upper_lookup/upper_has_children keep their exact contracts. Only
upstd_readdir and overlay_readdir's manual array scans became calls to
a new upper_visit_range (O(log n + r) range query, replacing an O(n)
scan in both, a bonus fix beyond what was strictly necessary) since
there's no flat array left to scan.

Added tests/test_index_stress.c: thousands of randomized (not
sequential) creates/deletes/renames across nested directories,
cross-checked against an independent reference model after every
round, not just "did it not crash" -- exercises exactly the code path
the O(n^2) bug and this fix live in, at a scale the other tests don't
reach. Verified under -fsanitize=undefined (150+ runs across this
change's lifetime, 0 failures) and -fsanitize=address (100+ runs, 0
real findings; known sandbox ASan-startup flakes excluded, see prior
commits) per CLAUDE.md's sanitizer rule for upper.c/overlay.c changes.

Measured result (make bench, same environment as the original
finding): mem create 257x faster, unlink 330x faster, mkdir 65x
faster, concurrent mixed workload 131x faster -- and, the comparison
that matters, mem now beats raw fs at every one of these (was losing
by 6-22x before). The complexity-class change is confirmed the same
way the O(n^2) was found: mkdir at N=4,000 vs create at N=20,000 now
shows a 7.15x slowdown for a 5x increase in N, matching the O(n log n)
prediction (5.97x) rather than the old O(n^2) one (25x). Full
before/after tables in BENCH.md's new "Resolution" section, which
keeps the original run as the historical record rather than
overwriting it, per this project's own documentation standard.

Recorded the fix in CLAUDE.md's "Known performance characteristics"
(marked RESOLVED, not silently removed) and its architecture map.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
2026-09-14 08:09:49 +00:00

14 KiB
Raw Blame History

PackFS vs. the host filesystem: benchmark results

Status: the O(n²) bulk create/unlink finding this document originally reported has been fixed. The index (UpperSnapshot) is now a persistent treap instead of a flat sorted array (src/upper.c; see CLAUDE.md's "Known performance characteristics"). This document keeps the original run's numbers as the historical record of the problem — per this project's own documentation standard, a correction extends the record rather than erasing it — and adds the post-fix run's numbers alongside them. Read "Resolution" below before drawing conclusions from the "Original results" section, which describes a version of the index this codebase no longer has.

This document reports the results of bench/bench.c (make bench), each run once on the environment described below. These are single runs, not a statistically-averaged series — treat the numbers as illustrative of shape (which operations are faster, by roughly what factor, and why) more than as precise absolute figures for a different machine. The methodology, including exactly what each backend label measures, is documented in bench/bench.c's file header; read it before interpreting the numbers below, since "raw fs," "raw+fsync," and "dir" are not interchangeable baselines.

Environment

  • Host: containerized, 12 logical CPUs (AMD Ryzen 5 3600), running under an overlay root filesystem (overlay on overlay, per mount) — there is no separate tmpfs at /tmp in this environment; every "raw fs" and dir number below is against that container overlay filesystem, not a bare disk or tmpfs. This matters most for the fsync numbers (see below).
  • N_SMALL = 20,000 files × 128 bytes for the metadata-heavy suite; N_DIRS = 4,000; concurrency = 8 threads × 4,000 ops; large files up to 64 MB. Full parameters are in bench/bench.c.
  • Both runs took several minutes of wall-clock time, almost entirely spent in the single raw+fsync 20,000-file create test (~116s in both runs — this cost is on the host's storage, not PackFS, and is identical before and after the fix, as expected).

Resolution

The original run below found bulk sequential create/unlink/mkdir on mem/dir to be O(n²) in file count, because every structural write copied the entire sorted entry array before publishing the next snapshot (Section 5.3). concept.md itself names the trigger condition for reconsidering that design — "a persistent (structurally shared) tree structure is not required until this assumption is empirically violated" — and the original run was exactly that violation, measured rather than hypothesized.

The index was rewritten as a persistent treap (src/upper.c): a structural write now builds only the O(log n) nodes on the path to the change, sharing every other node (refcounted) with the snapshot it was built from, instead of copying all n entries. The choice of treap over a persistent AVL/red-black/weight-balanced tree, and the reasoning behind it, is documented in src/upper.c's comment above struct TreapNode and cited there: Seidel & Aragon, "Randomized Search Trees" (1996), for the O(1) amortized rotation count on both insert and delete (persistent balanced trees that need O(log n) rotations pay a real, repeated allocation cost per rotation, which a treap avoids); Liljenzin, "Confluently Persistent Sets and Maps" (arXiv:1301.3388), for persistent treaps specifically as an MVCC snapshot mechanism, which is this exact use case.

Result, measured, same methodology as the original finding:

Category Before After Speedup
create 20,000 files (mem) 9.56s / 2,093 ops/s 0.037s / 537,103 ops/s 257x
unlink 20,000 files (mem) 10.56s / 1,894 ops/s 0.032s / 625,437 ops/s 330x
mkdir 4,000 dirs (mem) 0.336s / 11,896 ops/s 0.0052s / 771,085 ops/s 65x
create 20,000 files (dir) 12.04s / 1,662 ops/s 1.15s / 17,345 ops/s 10x
concurrent create+read+unlink (mem) 36.73s / 2,614 ops/s 0.28s / 341,972 ops/s 131x

mem now beats raw fs (not merely "no longer loses to it") at every one of these: create is 27x faster than raw fs (was 9.5x slower), unlink is 15x faster (was 22x slower), and the concurrent mixed workload is 22x faster (was 6x slower) — the single-writer design's "optimizes for read-heavy concurrent workloads" caveat from the original analysis (below) no longer needs to exclude bulk-write workloads, because the write side is no longer the bottleneck it was.

The complexity-class change is confirmed quantitatively, the same way the original O(n²) finding was — not just "it got faster," which an unrelated constant-factor optimization could also produce: mkdir at N=4,000 vs. create at N=20,000 (same structural-write mechanism, 5x the N) now shows a 7.15x slowdown. O(n) predicts 5.0x; O(n log n) predicts 5.97x; the old O(n²) regime predicted, and measured, 25–28.4x. 7.15x lands close to the O(n log n) prediction (some excess is expected — content writes' buffer-allocation and per-file host effects differ between create, which writes bytes, and mkdir, which does not) and nowhere near quadratic. The index's asymptotic behavior changed, not merely its constant factor.

Full output of the post-fix run is reproducible with make bench; the correctness of the new index under heavy randomized structural churn (non-sequential creates/deletes/renames, nested directories, thousands of entries, cross-checked against an independent reference model — not just "the benchmark ran fast without crashing") is tests/test_index_stress.c, part of make test.

Original results (pre-fix; kept as the historical record — see "Resolution" above)

Category Backend Time Throughput MB/s
create 20,000 files mem 9.56s 2,093 ops/s 0.3
create 20,000 files dir 12.04s 1,662 ops/s 0.2
create 20,000 files raw fs 1.01s 19,836 ops/s 2.4
create 20,000 files raw+fsync 116.10s 172 ops/s 0.0
read 20,000 files mem 0.007s 2,798,731 ops/s 341.6
read 20,000 files dir 0.167s 119,468 ops/s 14.6
read 20,000 files raw fs 0.161s 124,122 ops/s 15.2
stat 20,000 files mem 0.006s 3,223,599 ops/s —
stat 20,000 files dir 0.152s 131,937 ops/s —
stat 20,000 files raw fs 0.070s 284,465 ops/s —
readdir (20,000 entries) mem 0.0037s 5,397,828 /s —
readdir (20,000 entries) dir 0.0028s 7,058,003 /s —
readdir (20,000 entries) raw fs 0.0071s 2,831,904 /s —
unlink 20,000 files mem 10.56s 1,894 ops/s —
unlink 20,000 files dir 10.19s 1,964 ops/s —
unlink 20,000 files raw fs 0.48s 41,823 ops/s —
mkdir 4,000 dirs mem 0.336s 11,896 ops/s —
mkdir 4,000 dirs dir 0.511s 7,829 ops/s —
mkdir 4,000 dirs raw fs 0.159s 25,143 ops/s —
write 1 / 16 / 64 MB mem — — 6,687 / 1,158 / 1,130
write 1 / 16 / 64 MB raw fs (no fsync) — — 2,685 / 3,413 / 3,436
write 1 / 16 / 64 MB raw+fsync — — 116 / 191 / 718
random-read 20,000 entries pack (mmap'd) 0.0085s 2,350,224 ops/s 286.9
random-read 20,000 entries raw fs 0.177s 113,198 ops/s 13.8
compact 20,000 entries to pack pack 0.029s 687,097 ops/s 83.9
concurrent create+read+unlink (8×4,000×3) mem 36.73s 2,614 ops/s —
concurrent create+read+unlink (8×4,000×3) raw fs 6.12s 15,686 ops/s —

Post-fix results

Category Backend Time Throughput MB/s
create 20,000 files mem 0.037s 537,103 ops/s 65.6
create 20,000 files dir 1.153s 17,345 ops/s 2.1
create 20,000 files raw fs 1.010s 19,805 ops/s 2.4
create 20,000 files raw+fsync 116.098s 172 ops/s 0.0
read 20,000 files mem 0.0085s 2,364,410 ops/s 288.6
read 20,000 files dir 0.160s 124,933 ops/s 15.3
read 20,000 files raw fs 0.159s 125,459 ops/s 15.3
stat 20,000 files mem 0.0073s 2,723,001 ops/s —
stat 20,000 files dir 0.151s 132,517 ops/s —
stat 20,000 files raw fs 0.070s 287,063 ops/s —
readdir (20,000 entries) mem 0.0042s 4,816,418 /s —
readdir (20,000 entries) dir 0.0039s 5,128,581 /s —
readdir (20,000 entries) raw fs 0.0070s 2,840,890 /s —
unlink 20,000 files mem 0.032s 625,437 ops/s —
unlink 20,000 files dir 0.515s 38,864 ops/s —
unlink 20,000 files raw fs 0.492s 40,628 ops/s —
mkdir 4,000 dirs mem 0.0052s 771,085 ops/s —
mkdir 4,000 dirs dir 0.177s 22,551 ops/s —
mkdir 4,000 dirs raw fs 0.137s 29,251 ops/s —
write 1 / 16 / 64 MB mem — — 2,937 / 1,483 / 1,087
write 1 / 16 / 64 MB raw fs (no fsync) — — 2,463 / 3,478 / 3,496
write 1 / 16 / 64 MB raw+fsync — — 42 / 702 / 776
random-read 20,000 entries pack (mmap'd) 0.0082s 2,445,144 ops/s 298.5
random-read 20,000 entries raw fs 0.166s 120,709 ops/s 14.7
compact 20,000 entries to pack pack 0.025s 814,864 ops/s 99.5
concurrent create+read+unlink (8×4,000×3) mem 0.281s 341,972 ops/s —
concurrent create+read+unlink (8×4,000×3) raw fs 6.298s 15,243 ops/s —

Everything not involving bulk mem/dir structural writes (reads, stat, readdir, large sequential I/O, pack random-access, compaction) is unchanged within normal run-to-run noise, as expected — the treap rewrite only touches the structural-write path (Section 5.3), not content reads, content writes, or the pack format.

Analysis

The bullets below are preserved from the original run for the categories the fix did not change (the "wins" and the non-structural-write "losses"), with the historical O(n²) explanation kept as the record of what was found and superseded by "Resolution" above where it no longer applies.

Where PackFS wins clearly

  • Reads of small files: 20–23x faster than both raw fs and the dir backend (2.4M ops/s vs ~125K ops/s). No syscall per read — mem reads are an expected-O(log n) treap lookup plus a memcpy out of an already-resident buffer (Section 5.3, 9.1).
  • stat: ~10x faster than raw fs. Same reason — no syscall, and the index is sorted for expected-O(log n) lookup rather than requiring a directory entry scan.
  • Random-access reads against a compacted pack: ~20x faster than raw fs (2.4M ops/s vs 121K ops/s), because the pack is one mmap'd file with a binary-searchable index (Section 9.1), against 20,000 individual open/read/close syscall triples on the raw-fs side. This is the single result that most directly validates the architecture's stated purpose — Section 3.2's claim that "random access is a flat table lookup, not a linear scan or a directory-parse-then-seek" — under an actual measured workload, not just by construction.
  • Compaction throughput: ~100 MB/s / ~815K entries/s to serialize the live tree into a fresh pack (Section 4.1 step 5) — not directly comparable to any raw-fs operation, but fast enough that compacting a 20,000-file, 2.5 MB tree is not a practically-felt pause (25ms).
  • Bulk create/unlink/mkdir: now 10–330x faster than the pre-fix mem/ dir numbers, and (for mem) 15–27x faster than raw fs — see "Resolution" above; this used to be PackFS's clearest loss and is now among its clearest wins.
  • Concurrent mixed create+read+unlink: now 22x faster than raw fs (was 6x slower before the fix) — see "Resolution."

Where raw fs still wins, and why that's expected, not a bug

  • dir backend stat is ~2.2x slower than raw fs stat (133K ops/s vs 287K ops/s) — because pfs_dir_statat (Section 6.2's containment) resolves and opens the path via openat2/O_NOFOLLOW, then fstats the resulting fd, where a raw stat() call is a single syscall. This is the direct, measured cost of path containment on the metadata path, separate from and smaller than the containment cost paid on open itself (which raw fs pays an equivalent single-syscall cost for anyway). Unaffected by the index rewrite.
  • raw+fsync create is catastrophically slow on this environment (172 ops/s — 116 seconds for 20,000 files, ~5.8ms per fsync), because this container's overlay filesystem has poor per-call fsync latency. This is a property of the container, not of PackFS or of a bare disk — flagged in the environment section above precisely so it isn't misread as "PackFS's journal must be this slow too." The journal (Section 4.4) does call fsync once per journaled write for the same durability reason raw fsync is slow here, which is a real, inherited cost on this kind of storage — not a PackFS-specific one. Unaffected by the index rewrite (the journal's cost is per-write fsync latency, not index maintenance).
  • Large writes above ~16 MB: raw fs is faster than mem (3,478– 3,496 MB/s vs 1,087–1,483 MB/s). Section 5.7's buffer-growth discipline (allocate new, copy old + new, publish, retire old — never realloc in place) means every capacity doubling re-copies everything written so far; a plain unsynced write() to a real file only ever appends new pages to the page cache, never re-copying prior ones. This is entirely independent of the index structure (Section 5.7's buffer-growth discipline governs a MutCell's content buffer, not the index the treap rewrite replaced) and is unaffected by this fix — the safety property Section 5.7 requires (no reader can ever see a freed buffer) still has a real, quantifiable cost for very large sequential writes, traded for correctness under concurrent access that a plain in-place realloc would not have.

Reproducing

make bench

Takes several minutes on a similarly slow-fsync environment, dominated by the raw+fsync create test; expect well under a minute on a host with normal disk or tmpfs fsync latency.