164 lines
9.4 KiB
Markdown
164 lines
9.4 KiB
Markdown
# PackFS vs. the host filesystem: benchmark results
|
||||
|
|
|
|||
|
|
This document reports the results of `bench/bench.c` (`make bench`), run to
|
|||
|
|
completion once on the environment described below. It is a single run, not
|
|||
|
|
a statistically-averaged series — treat the numbers as illustrative of
|
|||
|
|
*shape* (which operations are faster, by roughly what factor, and why) more
|
|||
|
|
than as precise absolute figures for a different machine. The methodology,
|
|||
|
|
including exactly what each backend label measures, is documented in
|
|||
|
|
`bench/bench.c`'s file header; read it before interpreting the numbers
|
|||
|
|
below, since "raw fs," "raw+fsync," and "dir" are not interchangeable
|
|||
|
|
baselines.
|
|||
|
|
|
|||
|
|
## Environment
|
|||
|
|
|
|||
|
|
- Host: containerized, 12 logical CPUs (AMD Ryzen 5 3600), running under an
|
|||
|
|
`overlay` root filesystem (`overlay` on `overlay`, per `mount`) — there is
|
|||
|
|
no separate tmpfs at `/tmp` in this environment; every "raw fs" and `dir`
|
|||
|
|
number below is against that container overlay filesystem, not a bare
|
|||
|
|
disk or tmpfs. This matters most for the `fsync` numbers (see below).
|
|||
|
|
- `N_SMALL` = 20,000 files × 128 bytes for the metadata-heavy suite;
|
|||
|
|
`N_DIRS` = 4,000; concurrency = 8 threads × 4,000 ops; large files up to
|
|||
|
|
64 MB. Full parameters are in `bench/bench.c`.
|
|||
|
|
- Wall-clock time for the full run: 5m49s, almost entirely spent in the
|
|||
|
|
single `raw+fsync` 20,000-file create test (116s) — see below.
|
|||
|
|
|
|||
|
|
## Results
|
|||
|
|
|
|||
|
|
| Category | Backend | Time | Throughput | MB/s |
|
|||
|
|
|---|---|---|---|---|
|
|||
|
|
| create 20,000 files | mem | 9.56s | 2,093 ops/s | 0.3 |
|
|||
|
|
| create 20,000 files | dir | 12.04s | 1,662 ops/s | 0.2 |
|
|||
|
|
| create 20,000 files | raw fs | 1.01s | 19,836 ops/s | 2.4 |
|
|||
|
|
| create 20,000 files | raw+fsync | 116.10s | 172 ops/s | 0.0 |
|
|||
|
|
| read 20,000 files | mem | 0.007s | 2,798,731 ops/s | 341.6 |
|
|||
|
|
| read 20,000 files | dir | 0.167s | 119,468 ops/s | 14.6 |
|
|||
|
|
| read 20,000 files | raw fs | 0.161s | 124,122 ops/s | 15.2 |
|
|||
|
|
| stat 20,000 files | mem | 0.006s | 3,223,599 ops/s | — |
|
|||
|
|
| stat 20,000 files | dir | 0.152s | 131,937 ops/s | — |
|
|||
|
|
| stat 20,000 files | raw fs | 0.070s | 284,465 ops/s | — |
|
|||
|
|
| readdir (20,000 entries) | mem | 0.0037s | 5,397,828 /s | — |
|
|||
|
|
| readdir (20,000 entries) | dir | 0.0028s | 7,058,003 /s | — |
|
|||
|
|
| readdir (20,000 entries) | raw fs | 0.0071s | 2,831,904 /s | — |
|
|||
|
|
| unlink 20,000 files | mem | 10.56s | 1,894 ops/s | — |
|
|||
|
|
| unlink 20,000 files | dir | 10.19s | 1,964 ops/s | — |
|
|||
|
|
| unlink 20,000 files | raw fs | 0.48s | 41,823 ops/s | — |
|
|||
|
|
| mkdir 4,000 dirs | mem | 0.336s | 11,896 ops/s | — |
|
|||
|
|
| mkdir 4,000 dirs | dir | 0.511s | 7,829 ops/s | — |
|
|||
|
|
| mkdir 4,000 dirs | raw fs | 0.159s | 25,143 ops/s | — |
|
|||
|
|
| write 1 / 16 / 64 MB | mem | — | — | 6,687 / 1,158 / 1,130 |
|
|||
|
|
| write 1 / 16 / 64 MB | raw fs (no fsync) | — | — | 2,685 / 3,413 / 3,436 |
|
|||
|
|
| write 1 / 16 / 64 MB | raw+fsync | — | — | 116 / 191 / 718 |
|
|||
|
|
| random-read 20,000 entries | pack (mmap'd) | 0.0085s | 2,350,224 ops/s | 286.9 |
|
|||
|
|
| random-read 20,000 entries | raw fs | 0.177s | 113,198 ops/s | 13.8 |
|
|||
|
|
| compact 20,000 entries to pack | pack | 0.029s | 687,097 ops/s | 83.9 |
|
|||
|
|
| concurrent create+read+unlink (8×4,000×3) | mem | 36.73s | 2,614 ops/s | — |
|
|||
|
|
| concurrent create+read+unlink (8×4,000×3) | raw fs | 6.12s | 15,686 ops/s | — |
|
|||
|
|
|
|||
|
|
Full per-run output, including the read-throughput MB/s columns omitted
|
|||
|
|
above for brevity, is reproducible with `make bench`.
|
|||
|
|
|
|||
|
|
## Analysis
|
|||
|
|
|
|||
|
|
### Where PackFS wins clearly
|
|||
|
|
|
|||
|
|
- **Reads of small files: 20–23x faster than both raw fs and the `dir`
|
|||
|
|
backend** (2.8M ops/s vs ~120K ops/s). No syscall per read — `mem`
|
|||
|
|
reads are a binary-search lookup plus a `memcpy` out of an
|
|||
|
|
already-resident buffer (Section 5.3, 9.1).
|
|||
|
|
- **`stat`: ~11x faster than raw fs.** Same reason — no syscall, and the
|
|||
|
|
index is sorted for O(log n) lookup rather than requiring a directory
|
|||
|
|
entry scan.
|
|||
|
|
- **Random-access reads against a compacted pack: ~21x faster than raw fs**
|
|||
|
|
(2.35M ops/s vs 113K ops/s), because the pack is one `mmap`'d file with a
|
|||
|
|
binary-searchable index (Section 9.1), against 20,000 individual
|
|||
|
|
`open`/`read`/`close` syscall triples on the raw-fs side. This is the
|
|||
|
|
single result that most directly validates the architecture's stated
|
|||
|
|
purpose — Section 3.2's claim that "random access is a flat table
|
|||
|
|
lookup, not a linear scan or a directory-parse-then-seek" — under an
|
|||
|
|
actual measured workload, not just by construction.
|
|||
|
|
- **Compaction throughput**: 84 MB/s / 687K entries/s to serialize the
|
|||
|
|
live tree into a fresh pack (Section 4.1 step 5) — not directly
|
|||
|
|
comparable to any raw-fs operation, but fast enough that compacting a
|
|||
|
|
20,000-file, 2.5 MB tree is not a practically-felt pause (29ms).
|
|||
|
|
- **Large sequential writes below roughly 4–8 MB total: 2.5x faster than
|
|||
|
|
page-cache-buffered raw fs** (6,687 MB/s vs 2,685 MB/s at 1 MB) — pure
|
|||
|
|
`malloc`+`memcpy` beats even an unsynced `write()` syscall at this size.
|
|||
|
|
|
|||
|
|
### Where raw fs wins clearly, and why that's expected, not a bug
|
|||
|
|
|
|||
|
|
- **Bulk create/unlink of many files is 9–22x *slower* on `mem` and `dir`
|
|||
|
|
than on raw fs** (2,093 and 1,662 ops/s vs 19,836 ops/s for create;
|
|||
|
|
1,894 and 1,964 vs 41,823 for unlink). This is not incidental overhead —
|
|||
|
|
it is the direct, predictable cost of Section 5.3's design: every
|
|||
|
|
structural write (`create`, `unlink`, `mkdir`) builds a **complete copy**
|
|||
|
|
of the snapshot's sorted entry array before publishing it. `concept.md`
|
|||
|
|
states this tradeoff explicitly and names its own limit: "at
|
|||
|
|
agent/sandbox scale the index is small enough that a full copy... is
|
|||
|
|
acceptable; a persistent (structurally shared) tree structure is not
|
|||
|
|
required until this assumption is empirically violated" (Section 5.3).
|
|||
|
|
**This benchmark is that empirical violation.** At N=20,000 sequential
|
|||
|
|
creates, the total cost of copying an array that grows from 0 to 20,000
|
|||
|
|
entries is quadratic in the file count, not linear — confirmed
|
|||
|
|
quantitatively, not just asserted: `mkdir` at N=4,000 (a 5x smaller N,
|
|||
|
|
same structural-write mechanism) takes 0.336s, while `create` at
|
|||
|
|
N=20,000 takes 9.56s — a 28.4x slowdown for a 5x increase in N.
|
|||
|
|
O(n) scaling predicts a 5.0x slowdown; O(n²) predicts 25.0x. The
|
|||
|
|
observed 28.4x is close to the quadratic prediction and nowhere near
|
|||
|
|
the linear one. **If bulk sequential creation/deletion of tens of
|
|||
|
|
thousands of files is a workload this project needs to support well,
|
|||
|
|
the index needs to stop being "copy the whole array per write" — this
|
|||
|
|
is not a proposal to do that rewrite, only a record that the
|
|||
|
|
spec's own stated trigger condition for reconsidering it has now been
|
|||
|
|
measured, not merely hypothesized.**
|
|||
|
|
- **`dir` backend `stat` is ~2.2x *slower* than raw fs `stat`** (132K ops/s
|
|||
|
|
vs 284K ops/s) — because `pfs_dir_statat` (Section 6.2's containment)
|
|||
|
|
resolves and opens the path via `openat2`/`O_NOFOLLOW`, then `fstat`s
|
|||
|
|
the resulting fd, where a raw `stat()` call is a single syscall. This is
|
|||
|
|
the direct, measured cost of path containment on the metadata path,
|
|||
|
|
separate from and smaller than the containment cost paid on `open`
|
|||
|
|
itself (which raw fs pays an equivalent single-syscall cost for anyway).
|
|||
|
|
- **`raw+fsync` create is catastrophically slow on this environment**
|
|||
|
|
(172 ops/s — 116 seconds for 20,000 files, ~5.8ms per `fsync`), because
|
|||
|
|
this container's overlay filesystem has poor per-call `fsync` latency.
|
|||
|
|
This is a property of the container, not of PackFS or of a bare disk —
|
|||
|
|
flagged in the environment section above precisely so it isn't
|
|||
|
|
misread as "PackFS's journal must be this slow too." The journal
|
|||
|
|
(Section 4.4) does call `fsync` once per journaled write for the same
|
|||
|
|
durability reason raw `fsync` is slow here, which is a real, inherited
|
|||
|
|
cost on this kind of storage — not a PackFS-specific one.
|
|||
|
|
- **Large writes above ~16 MB: raw fs is ~3x faster than `mem`** (3,413–
|
|||
|
|
3,436 MB/s vs 1,130–1,158 MB/s). Section 5.7's buffer-growth discipline
|
|||
|
|
(allocate new, copy old + new, publish, retire old — never realloc in
|
|||
|
|
place) means every capacity doubling re-copies everything written so
|
|||
|
|
far; a plain unsynced `write()` to a real file only ever appends new
|
|||
|
|
pages to the page cache, never re-copying prior ones. The crossover is
|
|||
|
|
visible in the data: `mem` beats raw fs at 1 MB (6,687 vs 2,685 MB/s)
|
|||
|
|
and loses to it by 16 MB (1,158 vs 3,413 MB/s) — the safety property
|
|||
|
|
Section 5.7 requires (no reader can ever see a freed buffer) has a real,
|
|||
|
|
quantifiable cost for very large sequential writes, traded for
|
|||
|
|
correctness under concurrent access that a plain in-place realloc
|
|||
|
|
would not have.
|
|||
|
|
- **Concurrent mixed create+read+unlink: raw fs is ~6x faster** (15,686
|
|||
|
|
vs 2,614 ops/s). This workload is create/unlink-dominated (two-thirds
|
|||
|
|
of each thread's three operations per iteration are structural writes),
|
|||
|
|
which is exactly the case just shown to be `mem`'s weakest point,
|
|||
|
|
further serialized through the single-writer lock (Section 5.3) across
|
|||
|
|
all 8 threads. This is not a counterexample to "wait-free readers" —
|
|||
|
|
reads within this same run are still wait-free — it is a demonstration
|
|||
|
|
that the single-writer design optimizes for read-heavy concurrent
|
|||
|
|
workloads specifically, not for concurrent bulk metadata churn, exactly
|
|||
|
|
as Section 5.1 frames the goal ("a documented, well-understood
|
|||
|
|
concurrency pattern," modeled on LMDB, which makes the identical
|
|||
|
|
tradeoff for the identical reason).
|
|||
|
|
|
|||
|
|
## Reproducing
|
|||
|
|
|
|||
|
|
```sh
|
|||
|
|
make bench
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Takes several minutes on a similarly slow-`fsync` environment, dominated
|
|||
|
|
by the `raw+fsync` create test; expect well under a minute on a host with
|
|||
|
|
normal disk or tmpfs `fsync` latency.
|