Files
packfs/BENCH.md
T
retoorandClaude Sonnet 5 0b3b207bc0 Add a benchmark suite comparing PackFS against the host filesystem
bench/bench.c (`make bench`) measures create/read/stat/readdir/unlink/
mkdir on mem, dir, and raw fs; large sequential I/O; random-access pack
reads via mmap vs raw fs; compaction throughput; and 8-thread
concurrent mixed workloads. Results from one full run, with honest
analysis (what each "raw fs" vs "raw+fsync" vs "dir" label actually
measures, so they aren't misread as interchangeable), are in BENCH.md.

The benchmark surfaced a real, quantitatively-confirmed finding, not
just favorable numbers: bulk sequential create/unlink on mem/dir is
O(n^2) in file count (~9-22x slower than raw fs at N=20,000), because
every structural write copies the entire snapshot entry array before
publishing it (Section 5.3). concept.md itself names the exact trigger
condition for reconsidering this ("a persistent structurally-shared
tree structure is not required until this assumption is empirically
violated") — this benchmark is that violation, measured rather than
hypothesized: mkdir at N=4,000 vs create at N=20,000 (same mechanism,
5x the N) shows a 28.4x slowdown, matching the O(n^2) prediction (25x)
far better than O(n) (5x).

Also found: dir-backend stat() costs ~2x raw stat() (open+fstat vs one
syscall, the direct cost of openat2 containment on the metadata path);
mem-backed large writes lose to raw fs above ~16MB (Section 5.7's
buffer-growth discipline re-copies prior writes on every capacity
doubling, the price of never exposing a reader to a freed buffer).

Recorded the O(n^2) finding in CLAUDE.md's new "Known performance
characteristics" section, per the same pattern used for the
reclaim_gate use-after-free discovery, since it's exactly the kind of
fact that would otherwise have to be rediscovered by benchmarking
again from scratch. Linked from README and CONTRIBUTING.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
2026-09-14 07:40:42 +00:00

164 lines
9.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# PackFS vs. the host filesystem: benchmark results
This document reports the results of `bench/bench.c` (`make bench`), run to
completion once on the environment described below. It is a single run, not
a statistically-averaged series — treat the numbers as illustrative of
*shape* (which operations are faster, by roughly what factor, and why) more
than as precise absolute figures for a different machine. The methodology,
including exactly what each backend label measures, is documented in
`bench/bench.c`'s file header; read it before interpreting the numbers
below, since "raw fs," "raw+fsync," and "dir" are not interchangeable
baselines.
## Environment
- Host: containerized, 12 logical CPUs (AMD Ryzen 5 3600), running under an
`overlay` root filesystem (`overlay` on `overlay`, per `mount`) — there is
no separate tmpfs at `/tmp` in this environment; every "raw fs" and `dir`
number below is against that container overlay filesystem, not a bare
disk or tmpfs. This matters most for the `fsync` numbers (see below).
- `N_SMALL` = 20,000 files × 128 bytes for the metadata-heavy suite;
`N_DIRS` = 4,000; concurrency = 8 threads × 4,000 ops; large files up to
64 MB. Full parameters are in `bench/bench.c`.
- Wall-clock time for the full run: 5m49s, almost entirely spent in the
single `raw+fsync` 20,000-file create test (116s) — see below.
## Results
| Category | Backend | Time | Throughput | MB/s |
|---|---|---|---|---|
| create 20,000 files | mem | 9.56s | 2,093 ops/s | 0.3 |
| create 20,000 files | dir | 12.04s | 1,662 ops/s | 0.2 |
| create 20,000 files | raw fs | 1.01s | 19,836 ops/s | 2.4 |
| create 20,000 files | raw+fsync | 116.10s | 172 ops/s | 0.0 |
| read 20,000 files | mem | 0.007s | 2,798,731 ops/s | 341.6 |
| read 20,000 files | dir | 0.167s | 119,468 ops/s | 14.6 |
| read 20,000 files | raw fs | 0.161s | 124,122 ops/s | 15.2 |
| stat 20,000 files | mem | 0.006s | 3,223,599 ops/s | — |
| stat 20,000 files | dir | 0.152s | 131,937 ops/s | — |
| stat 20,000 files | raw fs | 0.070s | 284,465 ops/s | — |
| readdir (20,000 entries) | mem | 0.0037s | 5,397,828 /s | — |
| readdir (20,000 entries) | dir | 0.0028s | 7,058,003 /s | — |
| readdir (20,000 entries) | raw fs | 0.0071s | 2,831,904 /s | — |
| unlink 20,000 files | mem | 10.56s | 1,894 ops/s | — |
| unlink 20,000 files | dir | 10.19s | 1,964 ops/s | — |
| unlink 20,000 files | raw fs | 0.48s | 41,823 ops/s | — |
| mkdir 4,000 dirs | mem | 0.336s | 11,896 ops/s | — |
| mkdir 4,000 dirs | dir | 0.511s | 7,829 ops/s | — |
| mkdir 4,000 dirs | raw fs | 0.159s | 25,143 ops/s | — |
| write 1 / 16 / 64 MB | mem | — | — | 6,687 / 1,158 / 1,130 |
| write 1 / 16 / 64 MB | raw fs (no fsync) | — | — | 2,685 / 3,413 / 3,436 |
| write 1 / 16 / 64 MB | raw+fsync | — | — | 116 / 191 / 718 |
| random-read 20,000 entries | pack (mmap'd) | 0.0085s | 2,350,224 ops/s | 286.9 |
| random-read 20,000 entries | raw fs | 0.177s | 113,198 ops/s | 13.8 |
| compact 20,000 entries to pack | pack | 0.029s | 687,097 ops/s | 83.9 |
| concurrent create+read+unlink (8×4,000×3) | mem | 36.73s | 2,614 ops/s | — |
| concurrent create+read+unlink (8×4,000×3) | raw fs | 6.12s | 15,686 ops/s | — |
Full per-run output, including the read-throughput MB/s columns omitted
above for brevity, is reproducible with `make bench`.
## Analysis
### Where PackFS wins clearly
- **Reads of small files: 20–23x faster than both raw fs and the `dir`
backend** (2.8M ops/s vs ~120K ops/s). No syscall per read — `mem`
reads are a binary-search lookup plus a `memcpy` out of an
already-resident buffer (Section 5.3, 9.1).
- **`stat`: ~11x faster than raw fs.** Same reason — no syscall, and the
index is sorted for O(log n) lookup rather than requiring a directory
entry scan.
- **Random-access reads against a compacted pack: ~21x faster than raw fs**
(2.35M ops/s vs 113K ops/s), because the pack is one `mmap`'d file with a
binary-searchable index (Section 9.1), against 20,000 individual
`open`/`read`/`close` syscall triples on the raw-fs side. This is the
single result that most directly validates the architecture's stated
purpose — Section 3.2's claim that "random access is a flat table
lookup, not a linear scan or a directory-parse-then-seek" — under an
actual measured workload, not just by construction.
- **Compaction throughput**: 84 MB/s / 687K entries/s to serialize the
live tree into a fresh pack (Section 4.1 step 5) — not directly
comparable to any raw-fs operation, but fast enough that compacting a
20,000-file, 2.5 MB tree is not a practically-felt pause (29ms).
- **Large sequential writes below roughly 4–8 MB total: 2.5x faster than
page-cache-buffered raw fs** (6,687 MB/s vs 2,685 MB/s at 1 MB) — pure
`malloc`+`memcpy` beats even an unsynced `write()` syscall at this size.
### Where raw fs wins clearly, and why that's expected, not a bug
- **Bulk create/unlink of many files is 9–22x *slower* on `mem` and `dir`
than on raw fs** (2,093 and 1,662 ops/s vs 19,836 ops/s for create;
1,894 and 1,964 vs 41,823 for unlink). This is not incidental overhead —
it is the direct, predictable cost of Section 5.3's design: every
structural write (`create`, `unlink`, `mkdir`) builds a **complete copy**
of the snapshot's sorted entry array before publishing it. `concept.md`
states this tradeoff explicitly and names its own limit: "at
agent/sandbox scale the index is small enough that a full copy... is
acceptable; a persistent (structurally shared) tree structure is not
required until this assumption is empirically violated" (Section 5.3).
**This benchmark is that empirical violation.** At N=20,000 sequential
creates, the total cost of copying an array that grows from 0 to 20,000
entries is quadratic in the file count, not linear — confirmed
quantitatively, not just asserted: `mkdir` at N=4,000 (a 5x smaller N,
same structural-write mechanism) takes 0.336s, while `create` at
N=20,000 takes 9.56s — a 28.4x slowdown for a 5x increase in N.
O(n) scaling predicts a 5.0x slowdown; O(n²) predicts 25.0x. The
observed 28.4x is close to the quadratic prediction and nowhere near
the linear one. **If bulk sequential creation/deletion of tens of
thousands of files is a workload this project needs to support well,
the index needs to stop being "copy the whole array per write" — this
is not a proposal to do that rewrite, only a record that the
spec's own stated trigger condition for reconsidering it has now been
measured, not merely hypothesized.**
- **`dir` backend `stat` is ~2.2x *slower* than raw fs `stat`** (132K ops/s
vs 284K ops/s) — because `pfs_dir_statat` (Section 6.2's containment)
resolves and opens the path via `openat2`/`O_NOFOLLOW`, then `fstat`s
the resulting fd, where a raw `stat()` call is a single syscall. This is
the direct, measured cost of path containment on the metadata path,
separate from and smaller than the containment cost paid on `open`
itself (which raw fs pays an equivalent single-syscall cost for anyway).
- **`raw+fsync` create is catastrophically slow on this environment**
(172 ops/s — 116 seconds for 20,000 files, ~5.8ms per `fsync`), because
this container's overlay filesystem has poor per-call `fsync` latency.
This is a property of the container, not of PackFS or of a bare disk —
flagged in the environment section above precisely so it isn't
misread as "PackFS's journal must be this slow too." The journal
(Section 4.4) does call `fsync` once per journaled write for the same
durability reason raw `fsync` is slow here, which is a real, inherited
cost on this kind of storage — not a PackFS-specific one.
- **Large writes above ~16 MB: raw fs is ~3x faster than `mem`** (3,413–
3,436 MB/s vs 1,130–1,158 MB/s). Section 5.7's buffer-growth discipline
(allocate new, copy old + new, publish, retire old — never realloc in
place) means every capacity doubling re-copies everything written so
far; a plain unsynced `write()` to a real file only ever appends new
pages to the page cache, never re-copying prior ones. The crossover is
visible in the data: `mem` beats raw fs at 1 MB (6,687 vs 2,685 MB/s)
and loses to it by 16 MB (1,158 vs 3,413 MB/s) — the safety property
Section 5.7 requires (no reader can ever see a freed buffer) has a real,
quantifiable cost for very large sequential writes, traded for
correctness under concurrent access that a plain in-place realloc
would not have.
- **Concurrent mixed create+read+unlink: raw fs is ~6x faster** (15,686
vs 2,614 ops/s). This workload is create/unlink-dominated (two-thirds
of each thread's three operations per iteration are structural writes),
which is exactly the case just shown to be `mem`'s weakest point,
further serialized through the single-writer lock (Section 5.3) across
all 8 threads. This is not a counterexample to "wait-free readers" —
reads within this same run are still wait-free — it is a demonstration
that the single-writer design optimizes for read-heavy concurrent
workloads specifically, not for concurrent bulk metadata churn, exactly
as Section 5.1 frames the goal ("a documented, well-understood
concurrency pattern," modeled on LMDB, which makes the identical
tradeoff for the identical reason).
## Reproducing
```sh
make bench
```
Takes several minutes on a similarly slow-`fsync` environment, dominated
by the `raw+fsync` create test; expect well under a minute on a host with
normal disk or tmpfs `fsync` latency.