BENCH.md's earlier finding was real: every structural write (create/
unlink/mkdir) copied the entire sorted UpperSnapshot entry array before
publishing the next snapshot, making bulk sequential creation O(n^2).
concept.md Section 5.3 names the exact condition for reconsidering this
("a persistent structurally-shared tree structure is not required
until this assumption is empirically violated") -- that condition was
measured, not hypothesized, so this closes it rather than leaving it
as a documented-but-open limitation.
The index is now a persistent treap (src/upper.c): a structural write
copies only the O(log n) nodes on the path to the change, sharing
every other node (its own refcount, cascading like MutCell's and
UpperSnapshot's) with whichever snapshot(s) it was built from. Chosen
over a persistent AVL/red-black/weight-balanced tree because deletion
in those can need O(log n) rebalancing rotations -- each a real
allocation in a persistent setting -- where a treap needs only O(1)
amortized rotations for insert and delete (Seidel & Aragon 1996), with
expected O(log n) height regardless of insertion order, including the
sorted-by-creation-order pattern that made the flat array quadratic in
the first place. Liljenzin's "Confluently Persistent Sets and Maps"
(arXiv:1301.3388) documents persistent treaps giving O(1) snapshots
for MVCC specifically, which is this exact use case. Full reasoning
and the ownership convention (functions consume one ref of their tree
arguments, return one owned ref) are in upper.c's comment above
struct TreapNode.
Blast radius kept deliberately small: snapshot_upsert/snapshot_remove
keep their exact original signatures, so upper_create/upper_mkdir/
upper_remove/upper_rename/upper_copy_up needed zero changes.
upper_lookup/upper_has_children keep their exact contracts. Only
upstd_readdir and overlay_readdir's manual array scans became calls to
a new upper_visit_range (O(log n + r) range query, replacing an O(n)
scan in both, a bonus fix beyond what was strictly necessary) since
there's no flat array left to scan.
Added tests/test_index_stress.c: thousands of randomized (not
sequential) creates/deletes/renames across nested directories,
cross-checked against an independent reference model after every
round, not just "did it not crash" -- exercises exactly the code path
the O(n^2) bug and this fix live in, at a scale the other tests don't
reach. Verified under -fsanitize=undefined (150+ runs across this
change's lifetime, 0 failures) and -fsanitize=address (100+ runs, 0
real findings; known sandbox ASan-startup flakes excluded, see prior
commits) per CLAUDE.md's sanitizer rule for upper.c/overlay.c changes.
Measured result (make bench, same environment as the original
finding): mem create 257x faster, unlink 330x faster, mkdir 65x
faster, concurrent mixed workload 131x faster -- and, the comparison
that matters, mem now beats raw fs at every one of these (was losing
by 6-22x before). The complexity-class change is confirmed the same
way the O(n^2) was found: mkdir at N=4,000 vs create at N=20,000 now
shows a 7.15x slowdown for a 5x increase in N, matching the O(n log n)
prediction (5.97x) rather than the old O(n^2) one (25x). Full
before/after tables in BENCH.md's new "Resolution" section, which
keeps the original run as the historical record rather than
overwriting it, per this project's own documentation standard.
Recorded the fix in CLAUDE.md's "Known performance characteristics"
(marked RESOLVED, not silently removed) and its architecture map.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
14 KiB
PackFS vs. the host filesystem: benchmark results
Status: the O(n²) bulk create/unlink finding this document originally
reported has been fixed. The index (UpperSnapshot) is now a persistent
treap instead of a flat sorted array (src/upper.c; see CLAUDE.md's "Known
performance characteristics"). This document keeps the original run's
numbers as the historical record of the problem — per this project's own
documentation standard, a correction extends the record rather than erasing
it — and adds the post-fix run's numbers alongside them. Read "Resolution"
below before drawing conclusions from the "Original results" section,
which describes a version of the index this codebase no longer has.
This document reports the results of bench/bench.c (make bench), each
run once on the environment described below. These are single runs, not a
statistically-averaged series — treat the numbers as illustrative of shape
(which operations are faster, by roughly what factor, and why) more than as
precise absolute figures for a different machine. The methodology, including
exactly what each backend label measures, is documented in bench/bench.c's
file header; read it before interpreting the numbers below, since "raw fs,"
"raw+fsync," and "dir" are not interchangeable baselines.
Environment
- Host: containerized, 12 logical CPUs (AMD Ryzen 5 3600), running under an
overlayroot filesystem (overlayonoverlay, permount) — there is no separate tmpfs at/tmpin this environment; every "raw fs" anddirnumber below is against that container overlay filesystem, not a bare disk or tmpfs. This matters most for thefsyncnumbers (see below). N_SMALL= 20,000 files × 128 bytes for the metadata-heavy suite;N_DIRS= 4,000; concurrency = 8 threads × 4,000 ops; large files up to 64 MB. Full parameters are inbench/bench.c.- Both runs took several minutes of wall-clock time, almost entirely spent
in the single
raw+fsync20,000-file create test (~116s in both runs — this cost is on the host's storage, not PackFS, and is identical before and after the fix, as expected).
Resolution
The original run below found bulk sequential create/unlink/mkdir on
mem/dir to be O(n²) in file count, because every structural write copied
the entire sorted entry array before publishing the next snapshot
(Section 5.3). concept.md itself names the trigger condition for
reconsidering that design — "a persistent (structurally shared) tree
structure is not required until this assumption is empirically violated" —
and the original run was exactly that violation, measured rather than
hypothesized.
The index was rewritten as a persistent treap (src/upper.c): a
structural write now builds only the O(log n) nodes on the path to the
change, sharing every other node (refcounted) with the snapshot it was
built from, instead of copying all n entries. The choice of treap over a
persistent AVL/red-black/weight-balanced tree, and the reasoning behind it,
is documented in src/upper.c's comment above struct TreapNode and cited
there: Seidel & Aragon, "Randomized Search Trees" (1996), for the O(1)
amortized rotation count on both insert and delete (persistent balanced
trees that need O(log n) rotations pay a real, repeated allocation cost per
rotation, which a treap avoids); Liljenzin, "Confluently Persistent Sets and
Maps" (arXiv:1301.3388), for persistent treaps specifically as an MVCC
snapshot mechanism, which is this exact use case.
Result, measured, same methodology as the original finding:
| Category | Before | After | Speedup |
|---|---|---|---|
| create 20,000 files (mem) | 9.56s / 2,093 ops/s | 0.037s / 537,103 ops/s | 257x |
| unlink 20,000 files (mem) | 10.56s / 1,894 ops/s | 0.032s / 625,437 ops/s | 330x |
| mkdir 4,000 dirs (mem) | 0.336s / 11,896 ops/s | 0.0052s / 771,085 ops/s | 65x |
| create 20,000 files (dir) | 12.04s / 1,662 ops/s | 1.15s / 17,345 ops/s | 10x |
| concurrent create+read+unlink (mem) | 36.73s / 2,614 ops/s | 0.28s / 341,972 ops/s | 131x |
mem now beats raw fs (not merely "no longer loses to it") at every one
of these: create is 27x faster than raw fs (was 9.5x slower), unlink is 15x
faster (was 22x slower), and the concurrent mixed workload is 22x faster
(was 6x slower) — the single-writer design's "optimizes for read-heavy
concurrent workloads" caveat from the original analysis (below) no longer
needs to exclude bulk-write workloads, because the write side is no longer
the bottleneck it was.
The complexity-class change is confirmed quantitatively, the same way the
original O(n²) finding was — not just "it got faster," which an unrelated
constant-factor optimization could also produce: mkdir at N=4,000 vs.
create at N=20,000 (same structural-write mechanism, 5x the N) now shows a
7.15x slowdown. O(n) predicts 5.0x; O(n log n) predicts 5.97x; the old
O(n²) regime predicted, and measured, 25–28.4x. 7.15x lands close to the
O(n log n) prediction (some excess is expected — content writes'
buffer-allocation and per-file host effects differ between create, which
writes bytes, and mkdir, which does not) and nowhere near quadratic. The
index's asymptotic behavior changed, not merely its constant factor.
Full output of the post-fix run is reproducible with make bench; the
correctness of the new index under heavy randomized structural churn
(non-sequential creates/deletes/renames, nested directories, thousands of
entries, cross-checked against an independent reference model — not just
"the benchmark ran fast without crashing") is tests/test_index_stress.c,
part of make test.
Original results (pre-fix; kept as the historical record — see "Resolution" above)
| Category | Backend | Time | Throughput | MB/s |
|---|---|---|---|---|
| create 20,000 files | mem | 9.56s | 2,093 ops/s | 0.3 |
| create 20,000 files | dir | 12.04s | 1,662 ops/s | 0.2 |
| create 20,000 files | raw fs | 1.01s | 19,836 ops/s | 2.4 |
| create 20,000 files | raw+fsync | 116.10s | 172 ops/s | 0.0 |
| read 20,000 files | mem | 0.007s | 2,798,731 ops/s | 341.6 |
| read 20,000 files | dir | 0.167s | 119,468 ops/s | 14.6 |
| read 20,000 files | raw fs | 0.161s | 124,122 ops/s | 15.2 |
| stat 20,000 files | mem | 0.006s | 3,223,599 ops/s | — |
| stat 20,000 files | dir | 0.152s | 131,937 ops/s | — |
| stat 20,000 files | raw fs | 0.070s | 284,465 ops/s | — |
| readdir (20,000 entries) | mem | 0.0037s | 5,397,828 /s | — |
| readdir (20,000 entries) | dir | 0.0028s | 7,058,003 /s | — |
| readdir (20,000 entries) | raw fs | 0.0071s | 2,831,904 /s | — |
| unlink 20,000 files | mem | 10.56s | 1,894 ops/s | — |
| unlink 20,000 files | dir | 10.19s | 1,964 ops/s | — |
| unlink 20,000 files | raw fs | 0.48s | 41,823 ops/s | — |
| mkdir 4,000 dirs | mem | 0.336s | 11,896 ops/s | — |
| mkdir 4,000 dirs | dir | 0.511s | 7,829 ops/s | — |
| mkdir 4,000 dirs | raw fs | 0.159s | 25,143 ops/s | — |
| write 1 / 16 / 64 MB | mem | — | — | 6,687 / 1,158 / 1,130 |
| write 1 / 16 / 64 MB | raw fs (no fsync) | — | — | 2,685 / 3,413 / 3,436 |
| write 1 / 16 / 64 MB | raw+fsync | — | — | 116 / 191 / 718 |
| random-read 20,000 entries | pack (mmap'd) | 0.0085s | 2,350,224 ops/s | 286.9 |
| random-read 20,000 entries | raw fs | 0.177s | 113,198 ops/s | 13.8 |
| compact 20,000 entries to pack | pack | 0.029s | 687,097 ops/s | 83.9 |
| concurrent create+read+unlink (8×4,000×3) | mem | 36.73s | 2,614 ops/s | — |
| concurrent create+read+unlink (8×4,000×3) | raw fs | 6.12s | 15,686 ops/s | — |
Post-fix results
| Category | Backend | Time | Throughput | MB/s |
|---|---|---|---|---|
| create 20,000 files | mem | 0.037s | 537,103 ops/s | 65.6 |
| create 20,000 files | dir | 1.153s | 17,345 ops/s | 2.1 |
| create 20,000 files | raw fs | 1.010s | 19,805 ops/s | 2.4 |
| create 20,000 files | raw+fsync | 116.098s | 172 ops/s | 0.0 |
| read 20,000 files | mem | 0.0085s | 2,364,410 ops/s | 288.6 |
| read 20,000 files | dir | 0.160s | 124,933 ops/s | 15.3 |
| read 20,000 files | raw fs | 0.159s | 125,459 ops/s | 15.3 |
| stat 20,000 files | mem | 0.0073s | 2,723,001 ops/s | — |
| stat 20,000 files | dir | 0.151s | 132,517 ops/s | — |
| stat 20,000 files | raw fs | 0.070s | 287,063 ops/s | — |
| readdir (20,000 entries) | mem | 0.0042s | 4,816,418 /s | — |
| readdir (20,000 entries) | dir | 0.0039s | 5,128,581 /s | — |
| readdir (20,000 entries) | raw fs | 0.0070s | 2,840,890 /s | — |
| unlink 20,000 files | mem | 0.032s | 625,437 ops/s | — |
| unlink 20,000 files | dir | 0.515s | 38,864 ops/s | — |
| unlink 20,000 files | raw fs | 0.492s | 40,628 ops/s | — |
| mkdir 4,000 dirs | mem | 0.0052s | 771,085 ops/s | — |
| mkdir 4,000 dirs | dir | 0.177s | 22,551 ops/s | — |
| mkdir 4,000 dirs | raw fs | 0.137s | 29,251 ops/s | — |
| write 1 / 16 / 64 MB | mem | — | — | 2,937 / 1,483 / 1,087 |
| write 1 / 16 / 64 MB | raw fs (no fsync) | — | — | 2,463 / 3,478 / 3,496 |
| write 1 / 16 / 64 MB | raw+fsync | — | — | 42 / 702 / 776 |
| random-read 20,000 entries | pack (mmap'd) | 0.0082s | 2,445,144 ops/s | 298.5 |
| random-read 20,000 entries | raw fs | 0.166s | 120,709 ops/s | 14.7 |
| compact 20,000 entries to pack | pack | 0.025s | 814,864 ops/s | 99.5 |
| concurrent create+read+unlink (8×4,000×3) | mem | 0.281s | 341,972 ops/s | — |
| concurrent create+read+unlink (8×4,000×3) | raw fs | 6.298s | 15,243 ops/s | — |
Everything not involving bulk mem/dir structural writes (reads, stat,
readdir, large sequential I/O, pack random-access, compaction) is
unchanged within normal run-to-run noise, as expected — the treap rewrite
only touches the structural-write path (Section 5.3), not content reads,
content writes, or the pack format.
Analysis
The bullets below are preserved from the original run for the categories the fix did not change (the "wins" and the non-structural-write "losses"), with the historical O(n²) explanation kept as the record of what was found and superseded by "Resolution" above where it no longer applies.
Where PackFS wins clearly
- Reads of small files: 20–23x faster than both raw fs and the
dirbackend (2.4M ops/s vs ~125K ops/s). No syscall per read —memreads are an expected-O(log n) treap lookup plus amemcpyout of an already-resident buffer (Section 5.3, 9.1). stat: ~10x faster than raw fs. Same reason — no syscall, and the index is sorted for expected-O(log n) lookup rather than requiring a directory entry scan.- Random-access reads against a compacted pack: ~20x faster than raw fs
(2.4M ops/s vs 121K ops/s), because the pack is one
mmap'd file with a binary-searchable index (Section 9.1), against 20,000 individualopen/read/closesyscall triples on the raw-fs side. This is the single result that most directly validates the architecture's stated purpose — Section 3.2's claim that "random access is a flat table lookup, not a linear scan or a directory-parse-then-seek" — under an actual measured workload, not just by construction. - Compaction throughput: ~100 MB/s / ~815K entries/s to serialize the live tree into a fresh pack (Section 4.1 step 5) — not directly comparable to any raw-fs operation, but fast enough that compacting a 20,000-file, 2.5 MB tree is not a practically-felt pause (25ms).
- Bulk create/unlink/mkdir: now 10–330x faster than the pre-fix
mem/dirnumbers, and (formem) 15–27x faster than raw fs — see "Resolution" above; this used to be PackFS's clearest loss and is now among its clearest wins. - Concurrent mixed create+read+unlink: now 22x faster than raw fs (was 6x slower before the fix) — see "Resolution."
Where raw fs still wins, and why that's expected, not a bug
dirbackendstatis ~2.2x slower than raw fsstat(133K ops/s vs 287K ops/s) — becausepfs_dir_statat(Section 6.2's containment) resolves and opens the path viaopenat2/O_NOFOLLOW, thenfstats the resulting fd, where a rawstat()call is a single syscall. This is the direct, measured cost of path containment on the metadata path, separate from and smaller than the containment cost paid onopenitself (which raw fs pays an equivalent single-syscall cost for anyway). Unaffected by the index rewrite.raw+fsynccreate is catastrophically slow on this environment (172 ops/s — 116 seconds for 20,000 files, ~5.8ms perfsync), because this container's overlay filesystem has poor per-callfsynclatency. This is a property of the container, not of PackFS or of a bare disk — flagged in the environment section above precisely so it isn't misread as "PackFS's journal must be this slow too." The journal (Section 4.4) does callfsynconce per journaled write for the same durability reason rawfsyncis slow here, which is a real, inherited cost on this kind of storage — not a PackFS-specific one. Unaffected by the index rewrite (the journal's cost is per-writefsynclatency, not index maintenance).- Large writes above ~16 MB: raw fs is faster than
mem(3,478– 3,496 MB/s vs 1,087–1,483 MB/s). Section 5.7's buffer-growth discipline (allocate new, copy old + new, publish, retire old — never realloc in place) means every capacity doubling re-copies everything written so far; a plain unsyncedwrite()to a real file only ever appends new pages to the page cache, never re-copying prior ones. This is entirely independent of the index structure (Section 5.7's buffer-growth discipline governs aMutCell's content buffer, not the index the treap rewrite replaced) and is unaffected by this fix — the safety property Section 5.7 requires (no reader can ever see a freed buffer) still has a real, quantifiable cost for very large sequential writes, traded for correctness under concurrent access that a plain in-place realloc would not have.
Reproducing
make bench
Takes several minutes on a similarly slow-fsync environment, dominated
by the raw+fsync create test; expect well under a minute on a host with
normal disk or tmpfs fsync latency.