BENCH.md's earlier finding was real: every structural write (create/
unlink/mkdir) copied the entire sorted UpperSnapshot entry array before
publishing the next snapshot, making bulk sequential creation O(n^2).
concept.md Section 5.3 names the exact condition for reconsidering this
("a persistent structurally-shared tree structure is not required
until this assumption is empirically violated") -- that condition was
measured, not hypothesized, so this closes it rather than leaving it
as a documented-but-open limitation.
The index is now a persistent treap (src/upper.c): a structural write
copies only the O(log n) nodes on the path to the change, sharing
every other node (its own refcount, cascading like MutCell's and
UpperSnapshot's) with whichever snapshot(s) it was built from. Chosen
over a persistent AVL/red-black/weight-balanced tree because deletion
in those can need O(log n) rebalancing rotations -- each a real
allocation in a persistent setting -- where a treap needs only O(1)
amortized rotations for insert and delete (Seidel & Aragon 1996), with
expected O(log n) height regardless of insertion order, including the
sorted-by-creation-order pattern that made the flat array quadratic in
the first place. Liljenzin's "Confluently Persistent Sets and Maps"
(arXiv:1301.3388) documents persistent treaps giving O(1) snapshots
for MVCC specifically, which is this exact use case. Full reasoning
and the ownership convention (functions consume one ref of their tree
arguments, return one owned ref) are in upper.c's comment above
struct TreapNode.
Blast radius kept deliberately small: snapshot_upsert/snapshot_remove
keep their exact original signatures, so upper_create/upper_mkdir/
upper_remove/upper_rename/upper_copy_up needed zero changes.
upper_lookup/upper_has_children keep their exact contracts. Only
upstd_readdir and overlay_readdir's manual array scans became calls to
a new upper_visit_range (O(log n + r) range query, replacing an O(n)
scan in both, a bonus fix beyond what was strictly necessary) since
there's no flat array left to scan.
Added tests/test_index_stress.c: thousands of randomized (not
sequential) creates/deletes/renames across nested directories,
cross-checked against an independent reference model after every
round, not just "did it not crash" -- exercises exactly the code path
the O(n^2) bug and this fix live in, at a scale the other tests don't
reach. Verified under -fsanitize=undefined (150+ runs across this
change's lifetime, 0 failures) and -fsanitize=address (100+ runs, 0
real findings; known sandbox ASan-startup flakes excluded, see prior
commits) per CLAUDE.md's sanitizer rule for upper.c/overlay.c changes.
Measured result (make bench, same environment as the original
finding): mem create 257x faster, unlink 330x faster, mkdir 65x
faster, concurrent mixed workload 131x faster -- and, the comparison
that matters, mem now beats raw fs at every one of these (was losing
by 6-22x before). The complexity-class change is confirmed the same
way the O(n^2) was found: mkdir at N=4,000 vs create at N=20,000 now
shows a 7.15x slowdown for a 5x increase in N, matching the O(n log n)
prediction (5.97x) rather than the old O(n^2) one (25x). Full
before/after tables in BENCH.md's new "Resolution" section, which
keeps the original run as the historical record rather than
overwriting it, per this project's own documentation standard.
Recorded the fix in CLAUDE.md's "Known performance characteristics"
(marked RESOLVED, not silently removed) and its architecture map.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
246 lines
14 KiB
Markdown
246 lines
14 KiB
Markdown
# PackFS vs. the host filesystem: benchmark results
|
||
|
||
**Status: the O(n²) bulk create/unlink finding this document originally
|
||
reported has been fixed.** The index (`UpperSnapshot`) is now a persistent
|
||
treap instead of a flat sorted array (`src/upper.c`; see CLAUDE.md's "Known
|
||
performance characteristics"). This document keeps the original run's
|
||
numbers as the historical record of the problem — per this project's own
|
||
documentation standard, a correction extends the record rather than erasing
|
||
it — and adds the post-fix run's numbers alongside them. **Read "Resolution"
|
||
below before drawing conclusions from the "Original results" section**,
|
||
which describes a version of the index this codebase no longer has.
|
||
|
||
This document reports the results of `bench/bench.c` (`make bench`), each
|
||
run once on the environment described below. These are single runs, not a
|
||
statistically-averaged series — treat the numbers as illustrative of *shape*
|
||
(which operations are faster, by roughly what factor, and why) more than as
|
||
precise absolute figures for a different machine. The methodology, including
|
||
exactly what each backend label measures, is documented in `bench/bench.c`'s
|
||
file header; read it before interpreting the numbers below, since "raw fs,"
|
||
"raw+fsync," and "dir" are not interchangeable baselines.
|
||
|
||
## Environment
|
||
|
||
- Host: containerized, 12 logical CPUs (AMD Ryzen 5 3600), running under an
|
||
`overlay` root filesystem (`overlay` on `overlay`, per `mount`) — there is
|
||
no separate tmpfs at `/tmp` in this environment; every "raw fs" and `dir`
|
||
number below is against that container overlay filesystem, not a bare
|
||
disk or tmpfs. This matters most for the `fsync` numbers (see below).
|
||
- `N_SMALL` = 20,000 files × 128 bytes for the metadata-heavy suite;
|
||
`N_DIRS` = 4,000; concurrency = 8 threads × 4,000 ops; large files up to
|
||
64 MB. Full parameters are in `bench/bench.c`.
|
||
- Both runs took several minutes of wall-clock time, almost entirely spent
|
||
in the single `raw+fsync` 20,000-file create test (~116s in both runs —
|
||
this cost is on the host's storage, not PackFS, and is identical before
|
||
and after the fix, as expected).
|
||
|
||
## Resolution
|
||
|
||
The original run below found bulk sequential `create`/`unlink`/`mkdir` on
|
||
`mem`/`dir` to be O(n²) in file count, because every structural write copied
|
||
the entire sorted entry array before publishing the next snapshot
|
||
(Section 5.3). `concept.md` itself names the trigger condition for
|
||
reconsidering that design — "a persistent (structurally shared) tree
|
||
structure is not required until this assumption is empirically violated" —
|
||
and the original run was exactly that violation, measured rather than
|
||
hypothesized.
|
||
|
||
The index was rewritten as a **persistent treap** (`src/upper.c`): a
|
||
structural write now builds only the O(log n) nodes on the path to the
|
||
change, sharing every other node (refcounted) with the snapshot it was
|
||
built from, instead of copying all n entries. The choice of treap over a
|
||
persistent AVL/red-black/weight-balanced tree, and the reasoning behind it,
|
||
is documented in `src/upper.c`'s comment above `struct TreapNode` and cited
|
||
there: Seidel & Aragon, "Randomized Search Trees" (1996), for the O(1)
|
||
amortized rotation count on both insert and delete (persistent balanced
|
||
trees that need O(log n) rotations pay a real, repeated allocation cost per
|
||
rotation, which a treap avoids); Liljenzin, "Confluently Persistent Sets and
|
||
Maps" (arXiv:1301.3388), for persistent treaps specifically as an MVCC
|
||
snapshot mechanism, which is this exact use case.
|
||
|
||
**Result, measured, same methodology as the original finding:**
|
||
|
||
| Category | Before | After | Speedup |
|
||
|---|---|---|---|
|
||
| create 20,000 files (mem) | 9.56s / 2,093 ops/s | 0.037s / 537,103 ops/s | **257x** |
|
||
| unlink 20,000 files (mem) | 10.56s / 1,894 ops/s | 0.032s / 625,437 ops/s | **330x** |
|
||
| mkdir 4,000 dirs (mem) | 0.336s / 11,896 ops/s | 0.0052s / 771,085 ops/s | **65x** |
|
||
| create 20,000 files (dir) | 12.04s / 1,662 ops/s | 1.15s / 17,345 ops/s | **10x** |
|
||
| concurrent create+read+unlink (mem) | 36.73s / 2,614 ops/s | 0.28s / 341,972 ops/s | **131x** |
|
||
|
||
`mem` now *beats* raw fs (not merely "no longer loses to it") at every one
|
||
of these: create is 27x faster than raw fs (was 9.5x slower), unlink is 15x
|
||
faster (was 22x slower), and the concurrent mixed workload is 22x faster
|
||
(was 6x slower) — the single-writer design's "optimizes for read-heavy
|
||
concurrent workloads" caveat from the original analysis (below) no longer
|
||
needs to exclude bulk-write workloads, because the write side is no longer
|
||
the bottleneck it was.
|
||
|
||
**The complexity-class change is confirmed quantitatively, the same way the
|
||
original O(n²) finding was** — not just "it got faster," which an unrelated
|
||
constant-factor optimization could also produce: `mkdir` at N=4,000 vs.
|
||
`create` at N=20,000 (same structural-write mechanism, 5x the N) now shows a
|
||
**7.15x** slowdown. O(n) predicts 5.0x; O(n log n) predicts 5.97x; the old
|
||
O(n²) regime predicted, and measured, 25–28.4x. 7.15x lands close to the
|
||
O(n log n) prediction (some excess is expected — content writes'
|
||
buffer-allocation and per-file host effects differ between `create`, which
|
||
writes bytes, and `mkdir`, which does not) and nowhere near quadratic. The
|
||
index's asymptotic behavior changed, not merely its constant factor.
|
||
|
||
Full output of the post-fix run is reproducible with `make bench`; the
|
||
correctness of the new index under heavy randomized structural churn
|
||
(non-sequential creates/deletes/renames, nested directories, thousands of
|
||
entries, cross-checked against an independent reference model — not just
|
||
"the benchmark ran fast without crashing") is `tests/test_index_stress.c`,
|
||
part of `make test`.
|
||
|
||
## Original results (pre-fix; kept as the historical record — see "Resolution" above)
|
||
|
||
| Category | Backend | Time | Throughput | MB/s |
|
||
|---|---|---|---|---|
|
||
| create 20,000 files | mem | 9.56s | 2,093 ops/s | 0.3 |
|
||
| create 20,000 files | dir | 12.04s | 1,662 ops/s | 0.2 |
|
||
| create 20,000 files | raw fs | 1.01s | 19,836 ops/s | 2.4 |
|
||
| create 20,000 files | raw+fsync | 116.10s | 172 ops/s | 0.0 |
|
||
| read 20,000 files | mem | 0.007s | 2,798,731 ops/s | 341.6 |
|
||
| read 20,000 files | dir | 0.167s | 119,468 ops/s | 14.6 |
|
||
| read 20,000 files | raw fs | 0.161s | 124,122 ops/s | 15.2 |
|
||
| stat 20,000 files | mem | 0.006s | 3,223,599 ops/s | — |
|
||
| stat 20,000 files | dir | 0.152s | 131,937 ops/s | — |
|
||
| stat 20,000 files | raw fs | 0.070s | 284,465 ops/s | — |
|
||
| readdir (20,000 entries) | mem | 0.0037s | 5,397,828 /s | — |
|
||
| readdir (20,000 entries) | dir | 0.0028s | 7,058,003 /s | — |
|
||
| readdir (20,000 entries) | raw fs | 0.0071s | 2,831,904 /s | — |
|
||
| unlink 20,000 files | mem | 10.56s | 1,894 ops/s | — |
|
||
| unlink 20,000 files | dir | 10.19s | 1,964 ops/s | — |
|
||
| unlink 20,000 files | raw fs | 0.48s | 41,823 ops/s | — |
|
||
| mkdir 4,000 dirs | mem | 0.336s | 11,896 ops/s | — |
|
||
| mkdir 4,000 dirs | dir | 0.511s | 7,829 ops/s | — |
|
||
| mkdir 4,000 dirs | raw fs | 0.159s | 25,143 ops/s | — |
|
||
| write 1 / 16 / 64 MB | mem | — | — | 6,687 / 1,158 / 1,130 |
|
||
| write 1 / 16 / 64 MB | raw fs (no fsync) | — | — | 2,685 / 3,413 / 3,436 |
|
||
| write 1 / 16 / 64 MB | raw+fsync | — | — | 116 / 191 / 718 |
|
||
| random-read 20,000 entries | pack (mmap'd) | 0.0085s | 2,350,224 ops/s | 286.9 |
|
||
| random-read 20,000 entries | raw fs | 0.177s | 113,198 ops/s | 13.8 |
|
||
| compact 20,000 entries to pack | pack | 0.029s | 687,097 ops/s | 83.9 |
|
||
| concurrent create+read+unlink (8×4,000×3) | mem | 36.73s | 2,614 ops/s | — |
|
||
| concurrent create+read+unlink (8×4,000×3) | raw fs | 6.12s | 15,686 ops/s | — |
|
||
|
||
## Post-fix results
|
||
|
||
| Category | Backend | Time | Throughput | MB/s |
|
||
|---|---|---|---|---|
|
||
| create 20,000 files | mem | 0.037s | 537,103 ops/s | 65.6 |
|
||
| create 20,000 files | dir | 1.153s | 17,345 ops/s | 2.1 |
|
||
| create 20,000 files | raw fs | 1.010s | 19,805 ops/s | 2.4 |
|
||
| create 20,000 files | raw+fsync | 116.098s | 172 ops/s | 0.0 |
|
||
| read 20,000 files | mem | 0.0085s | 2,364,410 ops/s | 288.6 |
|
||
| read 20,000 files | dir | 0.160s | 124,933 ops/s | 15.3 |
|
||
| read 20,000 files | raw fs | 0.159s | 125,459 ops/s | 15.3 |
|
||
| stat 20,000 files | mem | 0.0073s | 2,723,001 ops/s | — |
|
||
| stat 20,000 files | dir | 0.151s | 132,517 ops/s | — |
|
||
| stat 20,000 files | raw fs | 0.070s | 287,063 ops/s | — |
|
||
| readdir (20,000 entries) | mem | 0.0042s | 4,816,418 /s | — |
|
||
| readdir (20,000 entries) | dir | 0.0039s | 5,128,581 /s | — |
|
||
| readdir (20,000 entries) | raw fs | 0.0070s | 2,840,890 /s | — |
|
||
| unlink 20,000 files | mem | 0.032s | 625,437 ops/s | — |
|
||
| unlink 20,000 files | dir | 0.515s | 38,864 ops/s | — |
|
||
| unlink 20,000 files | raw fs | 0.492s | 40,628 ops/s | — |
|
||
| mkdir 4,000 dirs | mem | 0.0052s | 771,085 ops/s | — |
|
||
| mkdir 4,000 dirs | dir | 0.177s | 22,551 ops/s | — |
|
||
| mkdir 4,000 dirs | raw fs | 0.137s | 29,251 ops/s | — |
|
||
| write 1 / 16 / 64 MB | mem | — | — | 2,937 / 1,483 / 1,087 |
|
||
| write 1 / 16 / 64 MB | raw fs (no fsync) | — | — | 2,463 / 3,478 / 3,496 |
|
||
| write 1 / 16 / 64 MB | raw+fsync | — | — | 42 / 702 / 776 |
|
||
| random-read 20,000 entries | pack (mmap'd) | 0.0082s | 2,445,144 ops/s | 298.5 |
|
||
| random-read 20,000 entries | raw fs | 0.166s | 120,709 ops/s | 14.7 |
|
||
| compact 20,000 entries to pack | pack | 0.025s | 814,864 ops/s | 99.5 |
|
||
| concurrent create+read+unlink (8×4,000×3) | mem | 0.281s | 341,972 ops/s | — |
|
||
| concurrent create+read+unlink (8×4,000×3) | raw fs | 6.298s | 15,243 ops/s | — |
|
||
|
||
Everything not involving bulk `mem`/`dir` structural writes (reads, `stat`,
|
||
`readdir`, large sequential I/O, pack random-access, compaction) is
|
||
unchanged within normal run-to-run noise, as expected — the treap rewrite
|
||
only touches the structural-write path (Section 5.3), not content reads,
|
||
content writes, or the pack format.
|
||
|
||
## Analysis
|
||
|
||
The bullets below are preserved from the original run for the categories
|
||
the fix did not change (the "wins" and the non-structural-write "losses"),
|
||
with the historical O(n²) explanation kept as the record of what was found
|
||
and superseded by "Resolution" above where it no longer applies.
|
||
|
||
### Where PackFS wins clearly
|
||
|
||
- **Reads of small files: 20–23x faster than both raw fs and the `dir`
|
||
backend** (2.4M ops/s vs ~125K ops/s). No syscall per read — `mem`
|
||
reads are an expected-O(log n) treap lookup plus a `memcpy` out of an
|
||
already-resident buffer (Section 5.3, 9.1).
|
||
- **`stat`: ~10x faster than raw fs.** Same reason — no syscall, and the
|
||
index is sorted for expected-O(log n) lookup rather than requiring a
|
||
directory entry scan.
|
||
- **Random-access reads against a compacted pack: ~20x faster than raw fs**
|
||
(2.4M ops/s vs 121K ops/s), because the pack is one `mmap`'d file with a
|
||
binary-searchable index (Section 9.1), against 20,000 individual
|
||
`open`/`read`/`close` syscall triples on the raw-fs side. This is the
|
||
single result that most directly validates the architecture's stated
|
||
purpose — Section 3.2's claim that "random access is a flat table
|
||
lookup, not a linear scan or a directory-parse-then-seek" — under an
|
||
actual measured workload, not just by construction.
|
||
- **Compaction throughput**: ~100 MB/s / ~815K entries/s to serialize the
|
||
live tree into a fresh pack (Section 4.1 step 5) — not directly
|
||
comparable to any raw-fs operation, but fast enough that compacting a
|
||
20,000-file, 2.5 MB tree is not a practically-felt pause (25ms).
|
||
- **Bulk create/unlink/mkdir: now 10–330x faster than the pre-fix `mem`/
|
||
`dir` numbers, and (for `mem`) 15–27x faster than raw fs** — see
|
||
"Resolution" above; this used to be PackFS's clearest loss and is now
|
||
among its clearest wins.
|
||
- **Concurrent mixed create+read+unlink: now 22x faster than raw fs**
|
||
(was 6x *slower* before the fix) — see "Resolution."
|
||
|
||
### Where raw fs still wins, and why that's expected, not a bug
|
||
|
||
- **`dir` backend `stat` is ~2.2x *slower* than raw fs `stat`** (133K ops/s
|
||
vs 287K ops/s) — because `pfs_dir_statat` (Section 6.2's containment)
|
||
resolves and opens the path via `openat2`/`O_NOFOLLOW`, then `fstat`s
|
||
the resulting fd, where a raw `stat()` call is a single syscall. This is
|
||
the direct, measured cost of path containment on the metadata path,
|
||
separate from and smaller than the containment cost paid on `open`
|
||
itself (which raw fs pays an equivalent single-syscall cost for anyway).
|
||
Unaffected by the index rewrite.
|
||
- **`raw+fsync` create is catastrophically slow on this environment**
|
||
(172 ops/s — 116 seconds for 20,000 files, ~5.8ms per `fsync`), because
|
||
this container's overlay filesystem has poor per-call `fsync` latency.
|
||
This is a property of the container, not of PackFS or of a bare disk —
|
||
flagged in the environment section above precisely so it isn't
|
||
misread as "PackFS's journal must be this slow too." The journal
|
||
(Section 4.4) does call `fsync` once per journaled write for the same
|
||
durability reason raw `fsync` is slow here, which is a real, inherited
|
||
cost on this kind of storage — not a PackFS-specific one. Unaffected by
|
||
the index rewrite (the journal's cost is per-write `fsync` latency, not
|
||
index maintenance).
|
||
- **Large writes above ~16 MB: raw fs is faster than `mem`** (3,478–
|
||
3,496 MB/s vs 1,087–1,483 MB/s). Section 5.7's buffer-growth discipline
|
||
(allocate new, copy old + new, publish, retire old — never realloc in
|
||
place) means every capacity doubling re-copies everything written so
|
||
far; a plain unsynced `write()` to a real file only ever appends new
|
||
pages to the page cache, never re-copying prior ones. This is entirely
|
||
independent of the index structure (Section 5.7's buffer-growth
|
||
discipline governs a `MutCell`'s content buffer, not the index the
|
||
treap rewrite replaced) and is unaffected by this fix — the safety
|
||
property Section 5.7 requires (no reader can ever see a freed buffer)
|
||
still has a real, quantifiable cost for very large sequential writes,
|
||
traded for correctness under concurrent access that a plain in-place
|
||
realloc would not have.
|
||
|
||
## Reproducing
|
||
|
||
```sh
|
||
make bench
|
||
```
|
||
|
||
Takes several minutes on a similarly slow-`fsync` environment, dominated
|
||
by the `raw+fsync` create test; expect well under a minute on a host with
|
||
normal disk or tmpfs `fsync` latency.
|