# PackFS vs. the host filesystem: benchmark results This document reports the results of `bench/bench.c` (`make bench`), run to completion once on the environment described below. It is a single run, not a statistically-averaged series — treat the numbers as illustrative of *shape* (which operations are faster, by roughly what factor, and why) more than as precise absolute figures for a different machine. The methodology, including exactly what each backend label measures, is documented in `bench/bench.c`'s file header; read it before interpreting the numbers below, since "raw fs," "raw+fsync," and "dir" are not interchangeable baselines. ## Environment - Host: containerized, 12 logical CPUs (AMD Ryzen 5 3600), running under an `overlay` root filesystem (`overlay` on `overlay`, per `mount`) — there is no separate tmpfs at `/tmp` in this environment; every "raw fs" and `dir` number below is against that container overlay filesystem, not a bare disk or tmpfs. This matters most for the `fsync` numbers (see below). - `N_SMALL` = 20,000 files × 128 bytes for the metadata-heavy suite; `N_DIRS` = 4,000; concurrency = 8 threads × 4,000 ops; large files up to 64 MB. Full parameters are in `bench/bench.c`. - Wall-clock time for the full run: 5m49s, almost entirely spent in the single `raw+fsync` 20,000-file create test (116s) — see below. ## Results | Category | Backend | Time | Throughput | MB/s | |---|---|---|---|---| | create 20,000 files | mem | 9.56s | 2,093 ops/s | 0.3 | | create 20,000 files | dir | 12.04s | 1,662 ops/s | 0.2 | | create 20,000 files | raw fs | 1.01s | 19,836 ops/s | 2.4 | | create 20,000 files | raw+fsync | 116.10s | 172 ops/s | 0.0 | | read 20,000 files | mem | 0.007s | 2,798,731 ops/s | 341.6 | | read 20,000 files | dir | 0.167s | 119,468 ops/s | 14.6 | | read 20,000 files | raw fs | 0.161s | 124,122 ops/s | 15.2 | | stat 20,000 files | mem | 0.006s | 3,223,599 ops/s | — | | stat 20,000 files | dir | 0.152s | 131,937 ops/s | — | | stat 20,000 files | raw fs | 0.070s | 284,465 ops/s | — | | readdir (20,000 entries) | mem | 0.0037s | 5,397,828 /s | — | | readdir (20,000 entries) | dir | 0.0028s | 7,058,003 /s | — | | readdir (20,000 entries) | raw fs | 0.0071s | 2,831,904 /s | — | | unlink 20,000 files | mem | 10.56s | 1,894 ops/s | — | | unlink 20,000 files | dir | 10.19s | 1,964 ops/s | — | | unlink 20,000 files | raw fs | 0.48s | 41,823 ops/s | — | | mkdir 4,000 dirs | mem | 0.336s | 11,896 ops/s | — | | mkdir 4,000 dirs | dir | 0.511s | 7,829 ops/s | — | | mkdir 4,000 dirs | raw fs | 0.159s | 25,143 ops/s | — | | write 1 / 16 / 64 MB | mem | — | — | 6,687 / 1,158 / 1,130 | | write 1 / 16 / 64 MB | raw fs (no fsync) | — | — | 2,685 / 3,413 / 3,436 | | write 1 / 16 / 64 MB | raw+fsync | — | — | 116 / 191 / 718 | | random-read 20,000 entries | pack (mmap'd) | 0.0085s | 2,350,224 ops/s | 286.9 | | random-read 20,000 entries | raw fs | 0.177s | 113,198 ops/s | 13.8 | | compact 20,000 entries to pack | pack | 0.029s | 687,097 ops/s | 83.9 | | concurrent create+read+unlink (8×4,000×3) | mem | 36.73s | 2,614 ops/s | — | | concurrent create+read+unlink (8×4,000×3) | raw fs | 6.12s | 15,686 ops/s | — | Full per-run output, including the read-throughput MB/s columns omitted above for brevity, is reproducible with `make bench`. ## Analysis ### Where PackFS wins clearly - **Reads of small files: 20–23x faster than both raw fs and the `dir` backend** (2.8M ops/s vs ~120K ops/s). No syscall per read — `mem` reads are a binary-search lookup plus a `memcpy` out of an already-resident buffer (Section 5.3, 9.1). - **`stat`: ~11x faster than raw fs.** Same reason — no syscall, and the index is sorted for O(log n) lookup rather than requiring a directory entry scan. - **Random-access reads against a compacted pack: ~21x faster than raw fs** (2.35M ops/s vs 113K ops/s), because the pack is one `mmap`'d file with a binary-searchable index (Section 9.1), against 20,000 individual `open`/`read`/`close` syscall triples on the raw-fs side. This is the single result that most directly validates the architecture's stated purpose — Section 3.2's claim that "random access is a flat table lookup, not a linear scan or a directory-parse-then-seek" — under an actual measured workload, not just by construction. - **Compaction throughput**: 84 MB/s / 687K entries/s to serialize the live tree into a fresh pack (Section 4.1 step 5) — not directly comparable to any raw-fs operation, but fast enough that compacting a 20,000-file, 2.5 MB tree is not a practically-felt pause (29ms). - **Large sequential writes below roughly 4–8 MB total: 2.5x faster than page-cache-buffered raw fs** (6,687 MB/s vs 2,685 MB/s at 1 MB) — pure `malloc`+`memcpy` beats even an unsynced `write()` syscall at this size. ### Where raw fs wins clearly, and why that's expected, not a bug - **Bulk create/unlink of many files is 9–22x *slower* on `mem` and `dir` than on raw fs** (2,093 and 1,662 ops/s vs 19,836 ops/s for create; 1,894 and 1,964 vs 41,823 for unlink). This is not incidental overhead — it is the direct, predictable cost of Section 5.3's design: every structural write (`create`, `unlink`, `mkdir`) builds a **complete copy** of the snapshot's sorted entry array before publishing it. `concept.md` states this tradeoff explicitly and names its own limit: "at agent/sandbox scale the index is small enough that a full copy... is acceptable; a persistent (structurally shared) tree structure is not required until this assumption is empirically violated" (Section 5.3). **This benchmark is that empirical violation.** At N=20,000 sequential creates, the total cost of copying an array that grows from 0 to 20,000 entries is quadratic in the file count, not linear — confirmed quantitatively, not just asserted: `mkdir` at N=4,000 (a 5x smaller N, same structural-write mechanism) takes 0.336s, while `create` at N=20,000 takes 9.56s — a 28.4x slowdown for a 5x increase in N. O(n) scaling predicts a 5.0x slowdown; O(n²) predicts 25.0x. The observed 28.4x is close to the quadratic prediction and nowhere near the linear one. **If bulk sequential creation/deletion of tens of thousands of files is a workload this project needs to support well, the index needs to stop being "copy the whole array per write" — this is not a proposal to do that rewrite, only a record that the spec's own stated trigger condition for reconsidering it has now been measured, not merely hypothesized.** - **`dir` backend `stat` is ~2.2x *slower* than raw fs `stat`** (132K ops/s vs 284K ops/s) — because `pfs_dir_statat` (Section 6.2's containment) resolves and opens the path via `openat2`/`O_NOFOLLOW`, then `fstat`s the resulting fd, where a raw `stat()` call is a single syscall. This is the direct, measured cost of path containment on the metadata path, separate from and smaller than the containment cost paid on `open` itself (which raw fs pays an equivalent single-syscall cost for anyway). - **`raw+fsync` create is catastrophically slow on this environment** (172 ops/s — 116 seconds for 20,000 files, ~5.8ms per `fsync`), because this container's overlay filesystem has poor per-call `fsync` latency. This is a property of the container, not of PackFS or of a bare disk — flagged in the environment section above precisely so it isn't misread as "PackFS's journal must be this slow too." The journal (Section 4.4) does call `fsync` once per journaled write for the same durability reason raw `fsync` is slow here, which is a real, inherited cost on this kind of storage — not a PackFS-specific one. - **Large writes above ~16 MB: raw fs is ~3x faster than `mem`** (3,413– 3,436 MB/s vs 1,130–1,158 MB/s). Section 5.7's buffer-growth discipline (allocate new, copy old + new, publish, retire old — never realloc in place) means every capacity doubling re-copies everything written so far; a plain unsynced `write()` to a real file only ever appends new pages to the page cache, never re-copying prior ones. The crossover is visible in the data: `mem` beats raw fs at 1 MB (6,687 vs 2,685 MB/s) and loses to it by 16 MB (1,158 vs 3,413 MB/s) — the safety property Section 5.7 requires (no reader can ever see a freed buffer) has a real, quantifiable cost for very large sequential writes, traded for correctness under concurrent access that a plain in-place realloc would not have. - **Concurrent mixed create+read+unlink: raw fs is ~6x faster** (15,686 vs 2,614 ops/s). This workload is create/unlink-dominated (two-thirds of each thread's three operations per iteration are structural writes), which is exactly the case just shown to be `mem`'s weakest point, further serialized through the single-writer lock (Section 5.3) across all 8 threads. This is not a counterexample to "wait-free readers" — reads within this same run are still wait-free — it is a demonstration that the single-writer design optimizes for read-heavy concurrent workloads specifically, not for concurrent bulk metadata churn, exactly as Section 5.1 frames the goal ("a documented, well-understood concurrency pattern," modeled on LMDB, which makes the identical tradeoff for the identical reason). ## Reproducing ```sh make bench ``` Takes several minutes on a similarly slow-`fsync` environment, dominated by the `raw+fsync` create test; expect well under a minute on a host with normal disk or tmpfs `fsync` latency.