Files
packfs/BENCH.md
T

164 lines
9.4 KiB
Markdown
Raw Normal View History

# PackFS vs. the host filesystem: benchmark results
This document reports the results of `bench/bench.c` (`make bench`), run to
completion once on the environment described below. It is a single run, not
a statistically-averaged series — treat the numbers as illustrative of
*shape* (which operations are faster, by roughly what factor, and why) more
than as precise absolute figures for a different machine. The methodology,
including exactly what each backend label measures, is documented in
`bench/bench.c`'s file header; read it before interpreting the numbers
below, since "raw fs," "raw+fsync," and "dir" are not interchangeable
baselines.
## Environment
- Host: containerized, 12 logical CPUs (AMD Ryzen 5 3600), running under an
`overlay` root filesystem (`overlay` on `overlay`, per `mount`) — there is
no separate tmpfs at `/tmp` in this environment; every "raw fs" and `dir`
number below is against that container overlay filesystem, not a bare
disk or tmpfs. This matters most for the `fsync` numbers (see below).
- `N_SMALL` = 20,000 files × 128 bytes for the metadata-heavy suite;
`N_DIRS` = 4,000; concurrency = 8 threads × 4,000 ops; large files up to
64 MB. Full parameters are in `bench/bench.c`.
- Wall-clock time for the full run: 5m49s, almost entirely spent in the
single `raw+fsync` 20,000-file create test (116s) — see below.
## Results
| Category | Backend | Time | Throughput | MB/s |
|---|---|---|---|---|
| create 20,000 files | mem | 9.56s | 2,093 ops/s | 0.3 |
| create 20,000 files | dir | 12.04s | 1,662 ops/s | 0.2 |
| create 20,000 files | raw fs | 1.01s | 19,836 ops/s | 2.4 |
| create 20,000 files | raw+fsync | 116.10s | 172 ops/s | 0.0 |
| read 20,000 files | mem | 0.007s | 2,798,731 ops/s | 341.6 |
| read 20,000 files | dir | 0.167s | 119,468 ops/s | 14.6 |
| read 20,000 files | raw fs | 0.161s | 124,122 ops/s | 15.2 |
| stat 20,000 files | mem | 0.006s | 3,223,599 ops/s | — |
| stat 20,000 files | dir | 0.152s | 131,937 ops/s | — |
| stat 20,000 files | raw fs | 0.070s | 284,465 ops/s | — |
| readdir (20,000 entries) | mem | 0.0037s | 5,397,828 /s | — |
| readdir (20,000 entries) | dir | 0.0028s | 7,058,003 /s | — |
| readdir (20,000 entries) | raw fs | 0.0071s | 2,831,904 /s | — |
| unlink 20,000 files | mem | 10.56s | 1,894 ops/s | — |
| unlink 20,000 files | dir | 10.19s | 1,964 ops/s | — |
| unlink 20,000 files | raw fs | 0.48s | 41,823 ops/s | — |
| mkdir 4,000 dirs | mem | 0.336s | 11,896 ops/s | — |
| mkdir 4,000 dirs | dir | 0.511s | 7,829 ops/s | — |
| mkdir 4,000 dirs | raw fs | 0.159s | 25,143 ops/s | — |
| write 1 / 16 / 64 MB | mem | — | — | 6,687 / 1,158 / 1,130 |
| write 1 / 16 / 64 MB | raw fs (no fsync) | — | — | 2,685 / 3,413 / 3,436 |
| write 1 / 16 / 64 MB | raw+fsync | — | — | 116 / 191 / 718 |
| random-read 20,000 entries | pack (mmap'd) | 0.0085s | 2,350,224 ops/s | 286.9 |
| random-read 20,000 entries | raw fs | 0.177s | 113,198 ops/s | 13.8 |
| compact 20,000 entries to pack | pack | 0.029s | 687,097 ops/s | 83.9 |
| concurrent create+read+unlink (8×4,000×3) | mem | 36.73s | 2,614 ops/s | — |
| concurrent create+read+unlink (8×4,000×3) | raw fs | 6.12s | 15,686 ops/s | — |
Full per-run output, including the read-throughput MB/s columns omitted
above for brevity, is reproducible with `make bench`.
## Analysis
### Where PackFS wins clearly
- **Reads of small files: 20–23x faster than both raw fs and the `dir`
backend** (2.8M ops/s vs ~120K ops/s). No syscall per read — `mem`
reads are a binary-search lookup plus a `memcpy` out of an
already-resident buffer (Section 5.3, 9.1).
- **`stat`: ~11x faster than raw fs.** Same reason — no syscall, and the
index is sorted for O(log n) lookup rather than requiring a directory
entry scan.
- **Random-access reads against a compacted pack: ~21x faster than raw fs**
(2.35M ops/s vs 113K ops/s), because the pack is one `mmap`'d file with a
binary-searchable index (Section 9.1), against 20,000 individual
`open`/`read`/`close` syscall triples on the raw-fs side. This is the
single result that most directly validates the architecture's stated
purpose — Section 3.2's claim that "random access is a flat table
lookup, not a linear scan or a directory-parse-then-seek" — under an
actual measured workload, not just by construction.
- **Compaction throughput**: 84 MB/s / 687K entries/s to serialize the
live tree into a fresh pack (Section 4.1 step 5) — not directly
comparable to any raw-fs operation, but fast enough that compacting a
20,000-file, 2.5 MB tree is not a practically-felt pause (29ms).
- **Large sequential writes below roughly 4–8 MB total: 2.5x faster than
page-cache-buffered raw fs** (6,687 MB/s vs 2,685 MB/s at 1 MB) — pure
`malloc`+`memcpy` beats even an unsynced `write()` syscall at this size.
### Where raw fs wins clearly, and why that's expected, not a bug
- **Bulk create/unlink of many files is 9–22x *slower* on `mem` and `dir`
than on raw fs** (2,093 and 1,662 ops/s vs 19,836 ops/s for create;
1,894 and 1,964 vs 41,823 for unlink). This is not incidental overhead —
it is the direct, predictable cost of Section 5.3's design: every
structural write (`create`, `unlink`, `mkdir`) builds a **complete copy**
of the snapshot's sorted entry array before publishing it. `concept.md`
states this tradeoff explicitly and names its own limit: "at
agent/sandbox scale the index is small enough that a full copy... is
acceptable; a persistent (structurally shared) tree structure is not
required until this assumption is empirically violated" (Section 5.3).
**This benchmark is that empirical violation.** At N=20,000 sequential
creates, the total cost of copying an array that grows from 0 to 20,000
entries is quadratic in the file count, not linear — confirmed
quantitatively, not just asserted: `mkdir` at N=4,000 (a 5x smaller N,
same structural-write mechanism) takes 0.336s, while `create` at
N=20,000 takes 9.56s — a 28.4x slowdown for a 5x increase in N.
O(n) scaling predicts a 5.0x slowdown; O(n²) predicts 25.0x. The
observed 28.4x is close to the quadratic prediction and nowhere near
the linear one. **If bulk sequential creation/deletion of tens of
thousands of files is a workload this project needs to support well,
the index needs to stop being "copy the whole array per write" — this
is not a proposal to do that rewrite, only a record that the
spec's own stated trigger condition for reconsidering it has now been
measured, not merely hypothesized.**
- **`dir` backend `stat` is ~2.2x *slower* than raw fs `stat`** (132K ops/s
vs 284K ops/s) — because `pfs_dir_statat` (Section 6.2's containment)
resolves and opens the path via `openat2`/`O_NOFOLLOW`, then `fstat`s
the resulting fd, where a raw `stat()` call is a single syscall. This is
the direct, measured cost of path containment on the metadata path,
separate from and smaller than the containment cost paid on `open`
itself (which raw fs pays an equivalent single-syscall cost for anyway).
- **`raw+fsync` create is catastrophically slow on this environment**
(172 ops/s — 116 seconds for 20,000 files, ~5.8ms per `fsync`), because
this container's overlay filesystem has poor per-call `fsync` latency.
This is a property of the container, not of PackFS or of a bare disk —
flagged in the environment section above precisely so it isn't
misread as "PackFS's journal must be this slow too." The journal
(Section 4.4) does call `fsync` once per journaled write for the same
durability reason raw `fsync` is slow here, which is a real, inherited
cost on this kind of storage — not a PackFS-specific one.
- **Large writes above ~16 MB: raw fs is ~3x faster than `mem`** (3,413–
3,436 MB/s vs 1,130–1,158 MB/s). Section 5.7's buffer-growth discipline
(allocate new, copy old + new, publish, retire old — never realloc in
place) means every capacity doubling re-copies everything written so
far; a plain unsynced `write()` to a real file only ever appends new
pages to the page cache, never re-copying prior ones. The crossover is
visible in the data: `mem` beats raw fs at 1 MB (6,687 vs 2,685 MB/s)
and loses to it by 16 MB (1,158 vs 3,413 MB/s) — the safety property
Section 5.7 requires (no reader can ever see a freed buffer) has a real,
quantifiable cost for very large sequential writes, traded for
correctness under concurrent access that a plain in-place realloc
would not have.
- **Concurrent mixed create+read+unlink: raw fs is ~6x faster** (15,686
vs 2,614 ops/s). This workload is create/unlink-dominated (two-thirds
of each thread's three operations per iteration are structural writes),
which is exactly the case just shown to be `mem`'s weakest point,
further serialized through the single-writer lock (Section 5.3) across
all 8 threads. This is not a counterexample to "wait-free readers" —
reads within this same run are still wait-free — it is a demonstration
that the single-writer design optimizes for read-heavy concurrent
workloads specifically, not for concurrent bulk metadata churn, exactly
as Section 5.1 frames the goal ("a documented, well-understood
concurrency pattern," modeled on LMDB, which makes the identical
tradeoff for the identical reason).
## Reproducing
```sh
make bench
```
Takes several minutes on a similarly slow-`fsync` environment, dominated
by the `raw+fsync` create test; expect well under a minute on a host with
normal disk or tmpfs `fsync` latency.