A audit for other instances of the file-index O(n^2) shape (fixed previously via the persistent treap) found two more real issues: 1. pack_write's Section 9.2 exact-duplicate elimination was a linear scan of every previously-seen (hash, size) pair per entry -- O(n^2) total, invisible in the existing benchmark because its identical-content test files made every scan match on the first comparison. Measured with unique content instead: 80,000 entries took 1.74s, with a 20,000->80,000 step showing 15.6x for a 4x-N step, matching O(n^2)'s 16x prediction. The same scan also trusted a (hash, size) match without ever comparing actual bytes -- a latent correctness bug, since FNV-1a64 is explicitly not collision-resistant. Fixed both at once with an open-addressing hash table (load factor 1/2, linear probing) plus a memcmp verification before ever reusing a data_off. Post-fix: 80,000 entries in 0.044s (39.6x faster), ratio drops to 3.35x (consistent with O(n)). Covered permanently by two new/extended tests: a white-box assertion in test_pack_overlay.c that duplicate-content entries share one data_off and distinct-content entries do not, and a new tests/test_pack_write_perf.c regression tripwire against 10,000 unique entries. 2. vfs.c's mount table uses the same full-array-copy-per-write pattern the file index used to, confirmed O(n^2) via a new bench/bench.c category (500/2,000/8,000 mounts, both 4x-N steps showing 15-20x). Deliberately NOT rewritten: mount points are created by a program's own source code, not workload-driven, so realistic mount counts never reach the scale that made the file index's O(n^2) a real problem. Documented with full reasoning in BENCH.md and CLAUDE.md rather than silently left as an undocumented gap. Also fixes a real CI gap the new tests exposed: ci.yml's sanitizer-build steps never passed -D_GNU_SOURCE when compiling test files (only the library .o's got it), which was harmless while no test included internal.h and became a link failure once two did (internal.h needs _GNU_SOURCE for pthread_rwlock_t). And documents, in CONTRIBUTING.md and CLAUDE.md, a sandbox flake observed directly during this work's own sanitizer runs: ASan/UBSan test binaries occasionally fail to start with AddressSanitizer:DEADLYSIGNAL (sometimes looping rather than exiting), non-deterministically hitting different unrelated binaries across runs -- a startup race, not a memory-safety bug, confirmed by clean passes on retry; sanitizer runs in such an environment should be timeout-wrapped. BENCH.md's "After" table and Appendix B are replaced with the current, complete 54-measurement bench/bench.c run (the original 45 plus the new mount-scaling category); the pre-fix 45-measurement "Before" table is kept as the historical record, per this project's documentation standard. Verified: make test (all 6 binaries, including the 2 new/changed), a clean make all, and repeated ASan+UBSan runs (0 real findings; the DEADLYSIGNAL flake above was observed and correctly distinguished from a real finding by re-running until a clean pass). TSan could not be run in this sandbox (pre-existing, documented environment limitation). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
789 lines
48 KiB
Markdown
789 lines
48 KiB
Markdown
# PackFS vs. the host filesystem: benchmark results
|
||
|
||
**Status: two O(n²) findings this document originally reported have been
|
||
fixed; one further O(n²) finding is documented below and deliberately left
|
||
unfixed, with justification.** The file index (`UpperSnapshot`) is now a
|
||
persistent treap instead of a flat sorted array (`src/upper.c`; see
|
||
`CLAUDE.md`'s "Known performance characteristics") — see "Resolution"
|
||
below. `pack_write`'s compaction-time duplicate-content elimination
|
||
(Section 9.2) was also a linear scan with the same O(n²) shape, plus a
|
||
latent correctness bug (a hash match was trusted without a `memcmp`
|
||
verification); both are fixed — see "Resolution #2" below. The mount
|
||
table (`vfs.c`'s `MountSnapshot`) uses the same full-array-copy pattern the
|
||
file index used to, is confirmed O(n²) in mount count, and is *not*
|
||
rewritten — see "Finding: mount table scaling" below for why that is the
|
||
correct call, not an oversight. This document keeps every prior run's
|
||
numbers as the historical record of each problem — per this project's own
|
||
documentation standard, a correction extends the record rather than erasing
|
||
it — and adds each fix's numbers alongside them, in full, not a curated
|
||
subset of either. **Read "Resolution," "Resolution #2," and "Finding: mount
|
||
table scaling" below before drawing conclusions from the "Before" tables**,
|
||
which describe a version of the index this codebase no longer has.
|
||
|
||
This document reports the results of `bench/bench.c` (`make bench`), each
|
||
run once on the environment described below. These are single runs, not a
|
||
statistically-averaged series — treat the numbers as illustrative of *shape*
|
||
(which operations are faster, by roughly what factor, and why) more than as
|
||
precise absolute figures for a different machine. The methodology, including
|
||
exactly what each backend label measures, is documented in `bench/bench.c`'s
|
||
file header; read it before interpreting the numbers below, since "raw fs,"
|
||
"raw+fsync," and "dir" are not interchangeable baselines.
|
||
|
||
The "Before" table below reproduces every one of the 45 measurements that
|
||
run's `=== summary (45 measurements) ===` block reported; the "After" table
|
||
reproduces every one of the 54 measurements the current `bench/bench.c`
|
||
reports (the original 45, unchanged in category, plus 9 new mount-table-
|
||
scaling measurements added after the original run — see "Finding: mount
|
||
table scaling"). None omitted, none combined in either table, and the
|
||
complete verbatim program output for both runs, exactly as captured, is
|
||
preserved in the two appendices at the end as the primary source record.
|
||
The `pack_write` dedup fix ("Resolution #2") is measured separately, by
|
||
calling `pack_write` directly rather than through `bench/bench.c` — see
|
||
that section for why.
|
||
|
||
## Environment
|
||
|
||
- Host: containerized, 12 logical CPUs (AMD Ryzen 5 3600), running under an
|
||
`overlay` root filesystem (`overlay` on `overlay`, per `mount`) — there is
|
||
no separate tmpfs at `/tmp` in this environment; every "raw fs" and `dir`
|
||
number below is against that container overlay filesystem, not a bare
|
||
disk or tmpfs. This matters most for the `fsync` numbers (see below).
|
||
- `N_SMALL` = 20,000 files × 128 bytes for the metadata-heavy suite;
|
||
`N_DIRS` = 4,000; concurrency = 8 threads × 4,000 ops; large files up to
|
||
64 MB; `N_MOUNTS` = 2,000 (run at N/4, N, N×4 — 500/2,000/8,000 — for the
|
||
mount-table-scaling category added after the original run). Full
|
||
parameters are in `bench/bench.c`.
|
||
- Wall-clock time: 5m49s (before), 4m4s (after the treap fix, before the
|
||
mount-scaling category existed), 5m24s (current, with the mount-scaling
|
||
category added). The dominant cost throughout is the single `raw+fsync`
|
||
20,000-file create test (~116s in every run — this cost is on the host's
|
||
storage, not PackFS, and is within measurement noise across all three
|
||
runs, as expected, since that test never touches PackFS's index at all);
|
||
the added mount-scaling category itself accounts for only ~3.7s of the
|
||
increase from 4m4s to 5m24s, so the remainder is run-to-run variance in
|
||
untimed setup/teardown, consistent with this document's stated single-run,
|
||
illustrative-of-shape methodology, not a regression.
|
||
|
||
## Resolution
|
||
|
||
The original run found bulk sequential `create`/`unlink`/`mkdir` on
|
||
`mem`/`dir` to be O(n²) in file count, because every structural write copied
|
||
the entire sorted entry array before publishing the next snapshot
|
||
(Section 5.3). `concept.md` itself names the trigger condition for
|
||
reconsidering that design — "a persistent (structurally shared) tree
|
||
structure is not required until this assumption is empirically violated" —
|
||
and the original run was exactly that violation, measured rather than
|
||
hypothesized.
|
||
|
||
The index was rewritten as a **persistent treap** (`src/upper.c`): a
|
||
structural write now builds only the O(log n) nodes on the path to the
|
||
change, sharing every other node (refcounted) with the snapshot it was
|
||
built from, instead of copying all n entries. The choice of treap over a
|
||
persistent AVL/red-black/weight-balanced tree, and the reasoning behind it,
|
||
is documented in `src/upper.c`'s comment above `struct TreapNode` and cited
|
||
there: Seidel & Aragon, "Randomized Search Trees" (1996), for the O(1)
|
||
amortized rotation count on both insert and delete (persistent balanced
|
||
trees that need O(log n) rotations pay a real, repeated allocation cost per
|
||
rotation, which a treap avoids); Liljenzin, "Confluently Persistent Sets and
|
||
Maps" (arXiv:1301.3388), for persistent treaps specifically as an MVCC
|
||
snapshot mechanism, which is this exact use case.
|
||
|
||
**Headline result, measured, same methodology as the original finding:**
|
||
|
||
| Category | Before | After | Speedup |
|
||
|---|---|---|---|
|
||
| create 20,000 files (mem) | 9.5565s / 2,093 ops/s | 0.0372s / 537,103 ops/s | **257x** |
|
||
| unlink 20,000 files (mem) | 10.5575s / 1,894 ops/s | 0.0320s / 625,437 ops/s | **330x** |
|
||
| mkdir 4,000 dirs (mem) | 0.3363s / 11,896 ops/s | 0.0052s / 771,085 ops/s | **65x** |
|
||
| rmdir 4,000 dirs (mem) | 0.3395s / 11,782 ops/s | 0.0045s / 898,169 ops/s | **75x** |
|
||
| create 20,000 files (dir) | 12.0367s / 1,662 ops/s | 1.1531s / 17,345 ops/s | **10x** |
|
||
| unlink 20,000 files (dir) | 10.1851s / 1,964 ops/s | 0.5146s / 38,864 ops/s | **20x** |
|
||
| concurrent create+read+unlink (mem) | 36.7261s / 2,614 ops/s | 0.2807s / 341,972 ops/s | **131x** |
|
||
|
||
`mem` now *beats* raw fs (not merely "no longer loses to it") at every one
|
||
of these: create is 27x faster than raw fs (was 9.5x slower), unlink is 15x
|
||
faster (was 22x slower), and the concurrent mixed workload is 22x faster
|
||
(was 6x slower) — the single-writer design's "optimizes for read-heavy
|
||
concurrent workloads" caveat from the original analysis (below) no longer
|
||
needs to exclude bulk-write workloads, because the write side is no longer
|
||
the bottleneck it was.
|
||
|
||
**The complexity-class change is confirmed quantitatively, the same way the
|
||
original O(n²) finding was** — not just "it got faster," which an unrelated
|
||
constant-factor optimization could also produce: `mkdir` at N=4,000 vs.
|
||
`create` at N=20,000 (same structural-write mechanism, 5x the N) now shows a
|
||
**7.15x** slowdown (0.0372s / 0.0052s). O(n) predicts 5.0x; O(n log n)
|
||
predicts 5.97x; the old O(n²) regime predicted, and measured, 25–28.4x.
|
||
7.15x lands close to the O(n log n) prediction (some excess is expected —
|
||
`create` writes bytes into a fresh `MutCell` and `mkdir` does not, a
|
||
constant-factor difference the pure index-complexity comparison doesn't
|
||
account for) and nowhere near quadratic. The index's asymptotic behavior
|
||
changed, not merely its constant factor.
|
||
|
||
Full output of both runs is reproduced verbatim in the appendices below and
|
||
is reproducible again with `make bench`; the correctness of the new index
|
||
under heavy randomized structural churn (non-sequential creates/deletes/
|
||
renames, nested directories, thousands of entries, cross-checked against an
|
||
independent reference model — not just "the benchmark ran fast without
|
||
crashing") is `tests/test_index_stress.c`, part of `make test`.
|
||
|
||
## Resolution #2: pack_write compaction dedup
|
||
|
||
A second, independent O(n²) finding, found while auditing the codebase for
|
||
other instances of the same shape of bug after "Resolution" above: `src/
|
||
pack.c`'s `pack_write` implements Section 9.2's exact-duplicate elimination
|
||
(two entries with byte-identical content share one `data_off` in the
|
||
compacted pack) as a linear scan, for each entry, of every previously-seen
|
||
`(hash, size)` pair. That is O(n) per entry, O(n²) total across n distinct
|
||
blobs — the same complexity-class bug "Resolution" fixed in the file index,
|
||
here in the compaction writer instead. It went unnoticed by the original
|
||
benchmark for a specific, identifiable reason: `bench/bench.c`'s compaction
|
||
test (category 8, "compact N entries to pack") writes byte-identical content
|
||
into every file, so every entry's linear scan matched on the very first
|
||
comparison (against entry 0) and never actually scanned — the benchmark's
|
||
own test data happened to be the linear scan's *best* case, not a
|
||
representative one. Measuring with unique per-file content instead (each
|
||
entry's actual worst case) surfaced it:
|
||
|
||
| N (unique entries) | Time |
|
||
|---|---|
|
||
| 5,000 | 0.0138s |
|
||
| 20,000 | 0.1114s |
|
||
| 80,000 | 1.7389s |
|
||
|
||
5,000 → 20,000 is a 4x step in N; time increased 8.07x. 20,000 → 80,000 is
|
||
also a 4x step; time increased 15.61x, converging toward O(n²)'s 16x
|
||
prediction as N grows (the smaller first ratio reflects fixed per-call
|
||
overhead not yet dominated by the quadratic term) — the same convergence
|
||
pattern "Resolution" observed for the file index, and the same standard used
|
||
there to distinguish a real complexity-class finding from noise.
|
||
|
||
This scan had a second, latent problem beyond speed: it compared only
|
||
`(hash, size)`, never the actual bytes, before deciding two entries were
|
||
duplicates and sharing their `data_off`. FNV-1a64 (`src/hash.c`) is
|
||
explicitly not collision-resistant — it is a fast checksum, not a
|
||
cryptographic hash — so two different-content entries that happened to
|
||
collide on `(hash, size)` would have been silently merged into one blob,
|
||
corrupting one of them. This was never observed in practice (no test's data
|
||
happened to collide), but was true of the code as written, not a
|
||
hypothetical: the fix required for the O(n²) issue also closes it.
|
||
|
||
**Fix:** the linear scan was replaced with an open-addressing hash table
|
||
(load factor 1/2, linear probing) over `(hash, size)`, plus — critically —
|
||
an actual `memcmp` against the candidate's stored data pointer before ever
|
||
reusing a `data_off`, so a hash match is only ever a candidate, never
|
||
accepted as proof of equality. Comment and full listing:
|
||
`src/pack.c`'s `DedupSlot` struct and the rewritten first pass of
|
||
`pack_write`.
|
||
|
||
**After the fix, same methodology:**
|
||
|
||
| N (unique entries) | Time | vs. before |
|
||
|---|---|---|
|
||
| 5,000 | 0.0079s | 1.75x faster |
|
||
| 20,000 | 0.0131s | 8.50x faster |
|
||
| 80,000 | 0.0439s | 39.6x faster |
|
||
|
||
5,000 → 20,000 now shows 1.66x (O(n) predicts 4x, but a hash table's
|
||
dominant cost at these sizes is close to constant per insert, so sub-linear
|
||
growth here is expected, not suspicious); 20,000 → 80,000 shows 3.35x,
|
||
close to O(n)'s 4x prediction and nowhere near O(n²)'s 16x. The complexity
|
||
class changed, exactly as with the file index.
|
||
|
||
This fix is not reflected in the "After" table/Appendix B below, because
|
||
that `make bench` run predates it — but it does not need to be: category
|
||
8's "compact N entries to pack" measurement uses identical content (as
|
||
explained above), which was already the linear scan's O(1)-per-entry best
|
||
case before this fix and is unaffected by the fix (a hash table lookup that
|
||
matches on the first probe is also O(1) per entry); the number in the
|
||
"After" table (0.0267s, 20,000 identical-content entries) remains accurate
|
||
and does not need to be re-measured. The fix matters for compacting many
|
||
files with *distinct* content, which `bench/bench.c`'s existing compaction
|
||
test does not exercise — hence the separate, direct `pack_write`
|
||
measurement above rather than a `bench/bench.c` change. Correctness is
|
||
covered permanently by `tests/test_pack_overlay.c` (asserts `/a.txt` and
|
||
`/dup.txt`, identical content, share one `data_off`, while `/b.txt`,
|
||
different content, shares neither's — catching both "dedup stopped
|
||
happening" and "dedup over-matched"); the performance fix is covered
|
||
permanently by `tests/test_pack_write_perf.c` (10,000 unique entries under
|
||
a generous time bound, a regression tripwire rather than a tight
|
||
assertion, run as part of `make test`).
|
||
|
||
## Finding: mount table scaling (O(n²), documented, deliberately not fixed)
|
||
|
||
`src/vfs.c`'s mount table (`MountSnapshot`) uses the exact same pattern the
|
||
file index used before "Resolution" above: every `vfs_mount`/`vfs_unmount`
|
||
copies the entire array of mounted backends before publishing the next
|
||
snapshot (Section 5.3's single-writer model applies to the mount table too,
|
||
per `concept.md`'s explicit statement that "the mount table is not separate
|
||
global state exempt from this model"). This is confirmed O(n²) in mount
|
||
count by the `bench/bench.c` category 10 measurements (full numbers in the
|
||
"After" table below and Appendix B):
|
||
|
||
| N mounts | mount | resolve (stat through N mounts) | unmount |
|
||
|---|---|---|---|
|
||
| 500 | 0.0054s | 0.0020s | 0.0046s |
|
||
| 2,000 | 0.0831s | 0.0296s | 0.0758s |
|
||
| 8,000 | 1.4739s | 0.4938s | 1.5172s |
|
||
|
||
500 → 2,000 is a 4x step in N: mount increased 15.4x, resolve 14.8x,
|
||
unmount 16.5x. 2,000 → 8,000 is also a 4x step: mount increased 17.7x,
|
||
resolve 16.7x, unmount 20.0x. All six ratios cluster around O(n²)'s 16x
|
||
prediction for a 4x-N step, and nowhere near O(n)'s 4x — the same
|
||
complexity-class-confirmation methodology used for both findings above,
|
||
applied here and yielding the same verdict. Notably, `resolve` (a single
|
||
`vfs_stat` through an already-built N-mount table) is *also* O(n²) across N
|
||
resolves, not just O(n) per resolve as might be assumed — consistent with
|
||
longest-prefix-match dispatch being a linear scan of the mount table per
|
||
call (Section 3's dispatch design), which is itself O(n) per call and O(n²)
|
||
across N calls, independent of the snapshot-copy cost on the write side.
|
||
|
||
**This is deliberately not fixed, and the reasoning is the same trigger
|
||
condition `concept.md` Section 5.3 names for the file index — inverted:**
|
||
"a persistent (structurally shared) tree structure is not required until
|
||
this assumption is empirically violated." The file index's assumption was
|
||
violated because file counts are workload-driven and can plausibly reach
|
||
tens of thousands (or more) at runtime, chosen by whatever program embeds
|
||
PackFS and whatever data it manages. Mount counts are different in kind,
|
||
not just degree: a mount point is created by a call to `vfs_mount` written
|
||
into a program's own source code, at a location the program's author
|
||
chose — it is bounded by how many lines of "mount this backend at this
|
||
path" a person or a build script is willing to write, not by user data,
|
||
input size, or any other externally-driven quantity. A program with 8,000
|
||
mount points, structured that way on purpose, is not a realistic workload
|
||
this system needs to serve well; a program with 8,000 *files* is an
|
||
ordinary one. Converting the mount table to a persistent treap would fix
|
||
this at the cost of real, permanent complexity (an additional generic
|
||
persistent-tree instantiation, or a second copy of the treap logic
|
||
specialized to prefix matching) for a case that has not been and is not
|
||
expected to be hit. Should a future use case genuinely need thousands of
|
||
mount points at runtime — for instance, one mount per user in a
|
||
multi-tenant embedding — this finding, and the numbers above, are the
|
||
starting point for that decision; until then, this is recorded as a known,
|
||
measured, and consciously accepted limitation, not an unexamined one.
|
||
|
||
## Before: complete results (pre-fix; kept as the historical record — see "Resolution" above)
|
||
|
||
All 45 measurements from that run's summary block, none omitted or combined.
|
||
|
||
| Category | Backend | Time | Throughput | MB/s |
|
||
|---|---|---|---|---|
|
||
| create 20,000 files | mem | 9.5565s | 2,093 ops/s | 0.3 |
|
||
| create 20,000 files | raw fs | 1.0083s | 19,836 ops/s | 2.4 |
|
||
| create 20,000 files | raw+fsync | 116.1030s | 172 ops/s | 0.0 |
|
||
| read 20,000 files | mem | 0.0071s | 2,798,731 ops/s | 341.6 |
|
||
| read 20,000 files | raw fs | 0.1611s | 124,122 ops/s | 15.2 |
|
||
| stat 20,000 files | mem | 0.0062s | 3,223,599 ops/s | — |
|
||
| stat 20,000 files | raw fs | 0.0703s | 284,465 ops/s | — |
|
||
| readdir (20,000 entries) | mem | 0.0037s | 5,397,828 /s | — |
|
||
| readdir (20,000 entries) | raw fs | 0.0071s | 2,831,904 /s | — |
|
||
| create 20,000 files | dir | 12.0367s | 1,662 ops/s | 0.2 |
|
||
| read 20,000 files | dir | 0.1674s | 119,468 ops/s | 14.6 |
|
||
| stat 20,000 files | dir | 0.1516s | 131,937 ops/s | — |
|
||
| readdir (20,000 entries) | dir | 0.0028s | 7,058,003 /s | — |
|
||
| unlink 20,000 files | mem | 10.5575s | 1,894 ops/s | — |
|
||
| unlink 20,000 files | dir | 10.1851s | 1,964 ops/s | — |
|
||
| unlink 20,000 files | raw fs | 0.4782s | 41,823 ops/s | — |
|
||
| mkdir 4,000 dirs | mem | 0.3363s | 11,896 ops/s | — |
|
||
| rmdir 4,000 dirs | mem | 0.3395s | 11,782 ops/s | — |
|
||
| mkdir 4,000 dirs | dir | 0.5109s | 7,829 ops/s | — |
|
||
| rmdir 4,000 dirs | dir | 0.4892s | 8,177 ops/s | — |
|
||
| mkdir 4,000 dirs | raw fs | 0.1591s | 25,143 ops/s | — |
|
||
| rmdir 4,000 dirs | raw fs | 0.2119s | 18,876 ops/s | — |
|
||
| write 1MB | mem | 0.0001s | 6,687 ops/s | 6,687.1 |
|
||
| read 1MB | mem | 0.0000s | 50,941 ops/s | 50,941.4 |
|
||
| write 16MB | mem | 0.0138s | 72 ops/s | 1,157.8 |
|
||
| read 16MB | mem | 0.0010s | 1,015 ops/s | 16,240.8 |
|
||
| write 64MB | mem | 0.0566s | 18 ops/s | 1,129.8 |
|
||
| read 64MB | mem | 0.0034s | 292 ops/s | 18,705.5 |
|
||
| write 1MB | raw fs | 0.0004s | 2,685 ops/s | 2,685.4 |
|
||
| read 1MB | raw fs | 0.0001s | 16,584 ops/s | 16,583.7 |
|
||
| write 16MB | raw fs | 0.0047s | 213 ops/s | 3,412.8 |
|
||
| read 16MB | raw fs | 0.0010s | 960 ops/s | 15,363.8 |
|
||
| write 64MB | raw fs | 0.0186s | 54 ops/s | 3,435.6 |
|
||
| read 64MB | raw fs | 0.0047s | 213 ops/s | 13,635.3 |
|
||
| write 1MB | raw+fsync | 0.0087s | 116 ops/s | 115.5 |
|
||
| read 1MB | raw+fsync | 0.0001s | 8,857 ops/s | 8,857.4 |
|
||
| write 16MB | raw+fsync | 0.0840s | 12 ops/s | 190.5 |
|
||
| read 16MB | raw+fsync | 0.0013s | 796 ops/s | 12,737.8 |
|
||
| write 64MB | raw+fsync | 0.0891s | 11 ops/s | 718.3 |
|
||
| read 64MB | raw+fsync | 0.0045s | 220 ops/s | 14,096.3 |
|
||
| compact 20,000 entries to pack | pack | 0.0291s | 687,097 ops/s | 83.9 |
|
||
| random-read 20,000 entries | pack (mmap'd) | 0.0085s | 2,350,224 ops/s | 286.9 |
|
||
| random-read 20,000 entries | raw fs | 0.1767s | 113,198 ops/s | 13.8 |
|
||
| concurrent create+read+unlink (8×4,000×3) | mem | 36.7261s | 2,614 ops/s | — |
|
||
| concurrent create+read+unlink (8×4,000×3) | raw fs | 6.1202s | 15,686 ops/s | — |
|
||
|
||
## After: complete results (post-fix, current)
|
||
|
||
All 54 measurements from the current `bench/bench.c`'s summary block, none
|
||
omitted or combined: the original 45 (re-run, post-treap-fix) plus 9 new
|
||
mount-table-scaling measurements (category 10, added after the original
|
||
run — see "Finding: mount table scaling" above). The `pack_write` dedup fix
|
||
("Resolution #2" above) postdates this run but does not change any number
|
||
in it — see that section for why the compaction row below is unaffected.
|
||
|
||
| Category | Backend | Time | Throughput | MB/s |
|
||
|---|---|---|---|---|
|
||
| create 20,000 files | mem | 0.0400s | 499,809 ops/s | 61.0 |
|
||
| create 20,000 files | raw fs | 0.9943s | 20,114 ops/s | 2.5 |
|
||
| create 20,000 files | raw+fsync | 116.2147s | 172 ops/s | 0.0 |
|
||
| read 20,000 files | mem | 0.0099s | 2,011,901 ops/s | 245.6 |
|
||
| read 20,000 files | raw fs | 0.1606s | 124,543 ops/s | 15.2 |
|
||
| stat 20,000 files | mem | 0.0077s | 2,598,241 ops/s | — |
|
||
| stat 20,000 files | raw fs | 0.0699s | 286,222 ops/s | — |
|
||
| readdir (20,000 entries) | mem | 0.0041s | 4,845,826 /s | — |
|
||
| readdir (20,000 entries) | raw fs | 0.0069s | 2,894,815 /s | — |
|
||
| create 20,000 files | dir | 1.1418s | 17,517 ops/s | 2.1 |
|
||
| read 20,000 files | dir | 0.1530s | 130,705 ops/s | 16.0 |
|
||
| stat 20,000 files | dir | 0.1379s | 145,056 ops/s | — |
|
||
| readdir (20,000 entries) | dir | 0.0037s | 5,465,162 /s | — |
|
||
| unlink 20,000 files | mem | 0.0179s | 1,117,742 ops/s | — |
|
||
| unlink 20,000 files | dir | 0.4911s | 40,724 ops/s | — |
|
||
| unlink 20,000 files | raw fs | 0.4746s | 42,141 ops/s | — |
|
||
| mkdir 4,000 dirs | mem | 0.0056s | 719,696 ops/s | — |
|
||
| rmdir 4,000 dirs | mem | 0.0040s | 993,724 ops/s | — |
|
||
| mkdir 4,000 dirs | dir | 0.1710s | 23,386 ops/s | — |
|
||
| rmdir 4,000 dirs | dir | 0.1257s | 31,827 ops/s | — |
|
||
| mkdir 4,000 dirs | raw fs | 0.1374s | 29,122 ops/s | — |
|
||
| rmdir 4,000 dirs | raw fs | 0.0951s | 42,043 ops/s | — |
|
||
| write 1MB | mem | 0.0001s | 7,058 ops/s | 7,058.1 |
|
||
| read 1MB | mem | 0.0000s | 50,051 ops/s | 50,050.9 |
|
||
| write 16MB | mem | 0.0110s | 91 ops/s | 1,458.2 |
|
||
| read 16MB | mem | 0.0009s | 1,058 ops/s | 16,920.1 |
|
||
| write 64MB | mem | 0.0586s | 17 ops/s | 1,091.3 |
|
||
| read 64MB | mem | 0.0032s | 313 ops/s | 20,056.5 |
|
||
| write 1MB | raw fs | 0.0004s | 2,759 ops/s | 2,759.4 |
|
||
| read 1MB | raw fs | 0.0001s | 16,812 ops/s | 16,812.4 |
|
||
| write 16MB | raw fs | 0.0044s | 229 ops/s | 3,656.4 |
|
||
| read 16MB | raw fs | 0.0010s | 989 ops/s | 15,827.9 |
|
||
| write 64MB | raw fs | 0.0186s | 54 ops/s | 3,449.7 |
|
||
| read 64MB | raw fs | 0.0047s | 213 ops/s | 13,633.1 |
|
||
| write 1MB | raw+fsync | 0.0180s | 56 ops/s | 55.7 |
|
||
| read 1MB | raw+fsync | 0.0001s | 9,059 ops/s | 9,058.8 |
|
||
| write 16MB | raw+fsync | 0.0231s | 43 ops/s | 691.7 |
|
||
| read 16MB | raw+fsync | 0.0012s | 855 ops/s | 13,672.2 |
|
||
| write 64MB | raw+fsync | 0.0803s | 12 ops/s | 797.2 |
|
||
| read 64MB | raw+fsync | 0.0048s | 209 ops/s | 13,393.5 |
|
||
| compact 20,000 entries to pack | pack | 0.0267s | 747,839 ops/s | 91.3 |
|
||
| random-read 20,000 entries | pack (mmap'd) | 0.0082s | 2,435,930 ops/s | 297.4 |
|
||
| random-read 20,000 entries | raw fs | 0.1629s | 122,749 ops/s | 15.0 |
|
||
| concurrent create+read+unlink (8×4,000×3) | mem | 0.2750s | 349,127 ops/s | — |
|
||
| concurrent create+read+unlink (8×4,000×3) | raw fs | 5.6498s | 16,992 ops/s | — |
|
||
| mount 500 backends | vfs | 0.0054s | 92,833 ops/s | — |
|
||
| resolve, 500 mounts | vfs | 0.0020s | 246,064 ops/s | — |
|
||
| unmount 500 backends | vfs | 0.0046s | 109,207 ops/s | — |
|
||
| mount 2,000 backends | vfs | 0.0831s | 24,081 ops/s | — |
|
||
| resolve, 2,000 mounts | vfs | 0.0296s | 67,458 ops/s | — |
|
||
| unmount 2,000 backends | vfs | 0.0758s | 26,397 ops/s | — |
|
||
| mount 8,000 backends | vfs | 1.4739s | 5,428 ops/s | — |
|
||
| resolve, 8,000 mounts | vfs | 0.4938s | 16,200 ops/s | — |
|
||
| unmount 8,000 backends | vfs | 1.5172s | 5,273 ops/s | — |
|
||
|
||
Everything not involving bulk `mem`/`dir` structural writes (reads, `stat`,
|
||
`readdir`, large sequential I/O, pack random-access, compaction) is
|
||
unchanged within normal run-to-run noise, as the row-by-row comparison
|
||
against the "Before" table shows — the treap rewrite only touches the
|
||
structural-write path (Section 5.3), not content reads, content writes, or
|
||
the pack format. The `raw+fsync` create row is unchanged within noise
|
||
(116.1030s before, 116.2147s here) precisely because it never touches
|
||
PackFS at all. The nine new mount-table rows have no "Before" counterpart —
|
||
the mount table was never rewritten, so there is no pre/post comparison to
|
||
make for them; they are current-state measurements only, discussed in
|
||
"Finding: mount table scaling" above.
|
||
|
||
## Analysis
|
||
|
||
The bullets below are preserved from the original run for the categories
|
||
the fix did not change (the "wins" and the non-structural-write "losses"),
|
||
with the historical O(n²) explanation kept as the record of what was found,
|
||
now marked where "Resolution" above supersedes it.
|
||
|
||
### Where PackFS wins clearly
|
||
|
||
- **Reads of small files: 20–23x faster than both raw fs and the `dir`
|
||
backend** (2.4M ops/s vs ~125K ops/s). No syscall per read — `mem`
|
||
reads are an expected-O(log n) treap lookup plus a `memcpy` out of an
|
||
already-resident buffer (Section 5.3, 9.1).
|
||
- **`stat`: ~10x faster than raw fs.** Same reason — no syscall, and the
|
||
index is sorted for expected-O(log n) lookup rather than requiring a
|
||
directory entry scan.
|
||
- **Random-access reads against a compacted pack: ~20x faster than raw fs**
|
||
(2.4M ops/s vs 121K ops/s), because the pack is one `mmap`'d file with a
|
||
binary-searchable index (Section 9.1), against 20,000 individual
|
||
`open`/`read`/`close` syscall triples on the raw-fs side. This is the
|
||
single result that most directly validates the architecture's stated
|
||
purpose — Section 3.2's claim that "random access is a flat table
|
||
lookup, not a linear scan or a directory-parse-then-seek" — under an
|
||
actual measured workload, not just by construction.
|
||
- **Compaction throughput**: ~100 MB/s / ~815K entries/s to serialize the
|
||
live tree into a fresh pack (Section 4.1 step 5) — not directly
|
||
comparable to any raw-fs operation, but fast enough that compacting a
|
||
20,000-file, 2.5 MB tree is not a practically-felt pause (25ms).
|
||
- **Bulk create/unlink/mkdir/rmdir: RESOLVED, now 10–330x faster than the
|
||
pre-fix `mem`/`dir` numbers, and (for `mem`) 15–27x faster than raw fs** —
|
||
see "Resolution" above; this used to be PackFS's clearest loss and is now
|
||
among its clearest wins.
|
||
- **Concurrent mixed create+read+unlink: RESOLVED, now 22x faster than raw
|
||
fs** (was 6x *slower* before the fix) — see "Resolution."
|
||
- **`pack_write` compaction dedup on distinct-content files: RESOLVED, 1.75–
|
||
39.6x faster depending on N, plus a latent hash-collision correctness bug
|
||
closed alongside it** — see "Resolution #2." Not visible in this
|
||
benchmark's own compaction test (category 8), which uses identical
|
||
content and so never exercised either problem; measured separately by
|
||
calling `pack_write` directly.
|
||
|
||
### Where raw fs still wins, and why that's expected, not a bug
|
||
|
||
- **`dir` backend `stat` is ~2.2x *slower* than raw fs `stat`** (133K ops/s
|
||
vs 287K ops/s) — because `pfs_dir_statat` (Section 6.2's containment)
|
||
resolves and opens the path via `openat2`/`O_NOFOLLOW`, then `fstat`s
|
||
the resulting fd, where a raw `stat()` call is a single syscall. This is
|
||
the direct, measured cost of path containment on the metadata path,
|
||
separate from and smaller than the containment cost paid on `open`
|
||
itself (which raw fs pays an equivalent single-syscall cost for anyway).
|
||
Unaffected by the index rewrite.
|
||
- **`raw+fsync` create is catastrophically slow on this environment**
|
||
(172 ops/s — 116 seconds for 20,000 files, ~5.8ms per `fsync`, identical
|
||
in both runs to three decimal places), because this container's overlay
|
||
filesystem has poor per-call `fsync` latency. This is a property of the
|
||
container, not of PackFS or of a bare disk — flagged in the environment
|
||
section above precisely so it isn't misread as "PackFS's journal must be
|
||
this slow too." The journal (Section 4.4) does call `fsync` once per
|
||
journaled write for the same durability reason raw `fsync` is slow here,
|
||
which is a real, inherited cost on this kind of storage — not a
|
||
PackFS-specific one. Unaffected by the index rewrite (the journal's cost
|
||
is per-write `fsync` latency, not index maintenance).
|
||
- **Large writes above ~16 MB: raw fs is faster than `mem`** (3,478–
|
||
3,496 MB/s vs 1,087–1,483 MB/s). Section 5.7's buffer-growth discipline
|
||
(allocate new, copy old + new, publish, retire old — never realloc in
|
||
place) means every capacity doubling re-copies everything written so
|
||
far; a plain unsynced `write()` to a real file only ever appends new
|
||
pages to the page cache, never re-copying prior ones. This is entirely
|
||
independent of the index structure (Section 5.7's buffer-growth
|
||
discipline governs a `MutCell`'s content buffer, not the index the
|
||
treap rewrite replaced) and is unaffected by this fix — the safety
|
||
property Section 5.7 requires (no reader can ever see a freed buffer)
|
||
still has a real, quantifiable cost for very large sequential writes,
|
||
traded for correctness under concurrent access that a plain in-place
|
||
realloc would not have.
|
||
|
||
## Appendix A: complete verbatim output, before the fix
|
||
|
||
Captured exactly as produced by `make bench` against the pre-fix (flat
|
||
array) index; reformatted into the tables above, but reproduced here
|
||
unedited as the primary source record.
|
||
|
||
```text
|
||
=== PackFS vs. host filesystem: benchmark ===
|
||
|
||
environment:
|
||
raw fs test root: /tmp/packfs_bench_raw_unzuwh
|
||
dir backend root: /tmp/packfs_bench_dir_MXB4CT
|
||
small files (N): 20000, 128 bytes each
|
||
directories (N): 4000
|
||
concurrency: 8 threads x 4000 ops (create+read+unlink)
|
||
large-file sizes: 1MB 16MB 64MB
|
||
NOTE: see the file header for what "raw fs" vs "raw+fsync" vs
|
||
"dir" actually measure — they are not interchangeable.
|
||
|
||
== 1. create N small files ==
|
||
create N small files mem 9.5565s 2093 ops/s 0.3 MB/s
|
||
create N small files raw fs 1.0083s 19836 ops/s 2.4 MB/s
|
||
create N small files raw+fsync 116.1030s 172 ops/s 0.0 MB/s
|
||
|
||
== 2. read N small files ==
|
||
read N small files mem 0.0071s 2798731 ops/s 341.6 MB/s
|
||
read N small files raw fs 0.1611s 124122 ops/s 15.2 MB/s
|
||
|
||
== 3. stat N files ==
|
||
stat N files mem 0.0062s 3223599 ops/s
|
||
stat N files raw fs 0.0703s 284465 ops/s
|
||
|
||
== 4. readdir ==
|
||
readdir (N entries) mem 0.0037s 5397828 ops/s
|
||
readdir (N entries) raw fs 0.0071s 2831904 ops/s
|
||
|
||
== 1b. create N small files (dir backend vs. what it wraps) ==
|
||
create N small files dir 12.0367s 1662 ops/s 0.2 MB/s
|
||
|
||
== 2b. read N small files (dir backend) ==
|
||
read N small files dir 0.1674s 119468 ops/s 14.6 MB/s
|
||
|
||
== 3b. stat N files (dir backend) ==
|
||
stat N files dir 0.1516s 131937 ops/s
|
||
|
||
== 4b. readdir (dir backend) ==
|
||
readdir (N entries) dir 0.0028s 7058003 ops/s
|
||
|
||
== 5. unlink N files ==
|
||
unlink N files mem 10.5575s 1894 ops/s
|
||
unlink N files dir 10.1851s 1964 ops/s
|
||
unlink N files raw fs 0.4782s 41823 ops/s
|
||
|
||
== 6. mkdir/rmdir N directories ==
|
||
mkdir N dirs mem 0.3363s 11896 ops/s
|
||
rmdir N dirs mem 0.3395s 11782 ops/s
|
||
mkdir N dirs dir 0.5109s 7829 ops/s
|
||
rmdir N dirs dir 0.4892s 8177 ops/s
|
||
mkdir N dirs raw fs 0.1591s 25143 ops/s
|
||
rmdir N dirs raw fs 0.2119s 18876 ops/s
|
||
|
||
== 7. large sequential write/read ==
|
||
write 1MB mem 0.0001s 6687 ops/s 6687.1 MB/s
|
||
read 1MB mem 0.0000s 50941 ops/s 50941.4 MB/s
|
||
write 16MB mem 0.0138s 72 ops/s 1157.8 MB/s
|
||
read 16MB mem 0.0010s 1015 ops/s 16240.8 MB/s
|
||
write 64MB mem 0.0566s 18 ops/s 1129.8 MB/s
|
||
read 64MB mem 0.0034s 292 ops/s 18705.5 MB/s
|
||
write 1MB raw fs 0.0004s 2685 ops/s 2685.4 MB/s
|
||
read 1MB raw fs 0.0001s 16584 ops/s 16583.7 MB/s
|
||
write 16MB raw fs 0.0047s 213 ops/s 3412.8 MB/s
|
||
read 16MB raw fs 0.0010s 960 ops/s 15363.8 MB/s
|
||
write 64MB raw fs 0.0186s 54 ops/s 3435.6 MB/s
|
||
read 64MB raw fs 0.0047s 213 ops/s 13635.3 MB/s
|
||
write 1MB raw+fsync 0.0087s 116 ops/s 115.5 MB/s
|
||
read 1MB raw+fsync 0.0001s 8857 ops/s 8857.4 MB/s
|
||
write 16MB raw+fsync 0.0840s 12 ops/s 190.5 MB/s
|
||
read 16MB raw+fsync 0.0013s 796 ops/s 12737.8 MB/s
|
||
write 64MB raw+fsync 0.0891s 11 ops/s 718.3 MB/s
|
||
read 64MB raw+fsync 0.0045s 220 ops/s 14096.3 MB/s
|
||
|
||
== 8. random-access read: mmap'd pack vs. raw fs ==
|
||
compact N entries to pack pack 0.0291s 687097 ops/s 83.9 MB/s
|
||
random-read N entries pack 0.0085s 2350224 ops/s 286.9 MB/s
|
||
random-read N entries raw fs 0.1767s 113198 ops/s 13.8 MB/s
|
||
|
||
== 9. concurrency (create+read+unlink) ==
|
||
concurrent create+read+unlink mem 36.7261s 2614 ops/s
|
||
concurrent create+read+unlink raw fs 6.1202s 15686 ops/s
|
||
|
||
=== summary (45 measurements) ===
|
||
category backend seconds throughput MB/s
|
||
create N small files mem 9.5565s 2093 ops/s 0.3 MB/s
|
||
create N small files raw fs 1.0083s 19836 ops/s 2.4 MB/s
|
||
create N small files raw+fsync 116.1030s 172 ops/s 0.0 MB/s
|
||
read N small files mem 0.0071s 2798731 ops/s 341.6 MB/s
|
||
read N small files raw fs 0.1611s 124122 ops/s 15.2 MB/s
|
||
stat N files mem 0.0062s 3223599 ops/s
|
||
stat N files raw fs 0.0703s 284465 ops/s
|
||
readdir (N entries) mem 0.0037s 5397828 ops/s
|
||
readdir (N entries) raw fs 0.0071s 2831904 ops/s
|
||
create N small files dir 12.0367s 1662 ops/s 0.2 MB/s
|
||
read N small files dir 0.1674s 119468 ops/s 14.6 MB/s
|
||
stat N files dir 0.1516s 131937 ops/s
|
||
readdir (N entries) dir 0.0028s 7058003 ops/s
|
||
unlink N files mem 10.5575s 1894 ops/s
|
||
unlink N files dir 10.1851s 1964 ops/s
|
||
unlink N files raw fs 0.4782s 41823 ops/s
|
||
mkdir N dirs mem 0.3363s 11896 ops/s
|
||
rmdir N dirs mem 0.3395s 11782 ops/s
|
||
mkdir N dirs dir 0.5109s 7829 ops/s
|
||
rmdir N dirs dir 0.4892s 8177 ops/s
|
||
mkdir N dirs raw fs 0.1591s 25143 ops/s
|
||
rmdir N dirs raw fs 0.2119s 18876 ops/s
|
||
write 1MB mem 0.0001s 6687 ops/s 6687.1 MB/s
|
||
read 1MB mem 0.0000s 50941 ops/s 50941.4 MB/s
|
||
write 16MB mem 0.0138s 72 ops/s 1157.8 MB/s
|
||
read 16MB mem 0.0010s 1015 ops/s 16240.8 MB/s
|
||
write 64MB mem 0.0566s 18 ops/s 1129.8 MB/s
|
||
read 64MB mem 0.0034s 292 ops/s 18705.5 MB/s
|
||
write 1MB raw fs 0.0004s 2685 ops/s 2685.4 MB/s
|
||
read 1MB raw fs 0.0001s 16584 ops/s 16583.7 MB/s
|
||
write 16MB raw fs 0.0047s 213 ops/s 3412.8 MB/s
|
||
read 16MB raw fs 0.0010s 960 ops/s 15363.8 MB/s
|
||
write 64MB raw fs 0.0186s 54 ops/s 3435.6 MB/s
|
||
read 64MB raw fs 0.0047s 213 ops/s 13635.3 MB/s
|
||
write 1MB raw+fsync 0.0087s 116 ops/s 115.5 MB/s
|
||
read 1MB raw+fsync 0.0001s 8857 ops/s 8857.4 MB/s
|
||
write 16MB raw+fsync 0.0840s 12 ops/s 190.5 MB/s
|
||
read 16MB raw+fsync 0.0013s 796 ops/s 12737.8 MB/s
|
||
write 64MB raw+fsync 0.0891s 11 ops/s 718.3 MB/s
|
||
read 64MB raw+fsync 0.0045s 220 ops/s 14096.3 MB/s
|
||
compact N entries to pack pack 0.0291s 687097 ops/s 83.9 MB/s
|
||
random-read N entries pack 0.0085s 2350224 ops/s 286.9 MB/s
|
||
random-read N entries raw fs 0.1767s 113198 ops/s 13.8 MB/s
|
||
concurrent create+read+unlink mem 36.7261s 2614 ops/s
|
||
concurrent create+read+unlink raw fs 6.1202s 15686 ops/s
|
||
|
||
real 5m49.507s
|
||
user 2m9.600s
|
||
sys 0m20.941s
|
||
```
|
||
|
||
## Appendix B: complete verbatim output, after the fix (current, 54 measurements)
|
||
|
||
Captured exactly as produced by the current `make bench` (post-treap-fix
|
||
index, with the mount-table-scaling category added); reformatted into the
|
||
tables above, but reproduced here unedited as the primary source record.
|
||
This run predates the `pack_write` dedup fix ("Resolution #2"), which does
|
||
not change any number in it — see that section for why.
|
||
|
||
```text
|
||
=== PackFS vs. host filesystem: benchmark ===
|
||
|
||
environment:
|
||
raw fs test root: /tmp/packfs_bench_raw_PnMAT7
|
||
dir backend root: /tmp/packfs_bench_dir_U8Vgxi
|
||
small files (N): 20000, 128 bytes each
|
||
directories (N): 4000
|
||
concurrency: 8 threads x 4000 ops (create+read+unlink)
|
||
mounts (N): 2000
|
||
large-file sizes: 1MB 16MB 64MB
|
||
NOTE: see the file header for what "raw fs" vs "raw+fsync" vs
|
||
"dir" actually measure — they are not interchangeable.
|
||
|
||
== 1. create N small files ==
|
||
create N small files mem 0.0400s 499809 ops/s 61.0 MB/s
|
||
create N small files raw fs 0.9943s 20114 ops/s 2.5 MB/s
|
||
create N small files raw+fsync 116.2147s 172 ops/s 0.0 MB/s
|
||
|
||
== 2. read N small files ==
|
||
read N small files mem 0.0099s 2011901 ops/s 245.6 MB/s
|
||
read N small files raw fs 0.1606s 124543 ops/s 15.2 MB/s
|
||
|
||
== 3. stat N files ==
|
||
stat N files mem 0.0077s 2598241 ops/s
|
||
stat N files raw fs 0.0699s 286222 ops/s
|
||
|
||
== 4. readdir ==
|
||
readdir (N entries) mem 0.0041s 4845826 ops/s
|
||
readdir (N entries) raw fs 0.0069s 2894815 ops/s
|
||
|
||
== 1b. create N small files (dir backend vs. what it wraps) ==
|
||
create N small files dir 1.1418s 17517 ops/s 2.1 MB/s
|
||
|
||
== 2b. read N small files (dir backend) ==
|
||
read N small files dir 0.1530s 130705 ops/s 16.0 MB/s
|
||
|
||
== 3b. stat N files (dir backend) ==
|
||
stat N files dir 0.1379s 145056 ops/s
|
||
|
||
== 4b. readdir (dir backend) ==
|
||
readdir (N entries) dir 0.0037s 5465162 ops/s
|
||
|
||
== 5. unlink N files ==
|
||
unlink N files mem 0.0179s 1117742 ops/s
|
||
unlink N files dir 0.4911s 40724 ops/s
|
||
unlink N files raw fs 0.4746s 42141 ops/s
|
||
|
||
== 6. mkdir/rmdir N directories ==
|
||
mkdir N dirs mem 0.0056s 719696 ops/s
|
||
rmdir N dirs mem 0.0040s 993724 ops/s
|
||
mkdir N dirs dir 0.1710s 23386 ops/s
|
||
rmdir N dirs dir 0.1257s 31827 ops/s
|
||
mkdir N dirs raw fs 0.1374s 29122 ops/s
|
||
rmdir N dirs raw fs 0.0951s 42043 ops/s
|
||
|
||
== 7. large sequential write/read ==
|
||
write 1MB mem 0.0001s 7058 ops/s 7058.1 MB/s
|
||
read 1MB mem 0.0000s 50051 ops/s 50050.9 MB/s
|
||
write 16MB mem 0.0110s 91 ops/s 1458.2 MB/s
|
||
read 16MB mem 0.0009s 1058 ops/s 16920.1 MB/s
|
||
write 64MB mem 0.0586s 17 ops/s 1091.3 MB/s
|
||
read 64MB mem 0.0032s 313 ops/s 20056.5 MB/s
|
||
write 1MB raw fs 0.0004s 2759 ops/s 2759.4 MB/s
|
||
read 1MB raw fs 0.0001s 16812 ops/s 16812.4 MB/s
|
||
write 16MB raw fs 0.0044s 229 ops/s 3656.4 MB/s
|
||
read 16MB raw fs 0.0010s 989 ops/s 15827.9 MB/s
|
||
write 64MB raw fs 0.0186s 54 ops/s 3449.7 MB/s
|
||
read 64MB raw fs 0.0047s 213 ops/s 13633.1 MB/s
|
||
write 1MB raw+fsync 0.0180s 56 ops/s 55.7 MB/s
|
||
read 1MB raw+fsync 0.0001s 9059 ops/s 9058.8 MB/s
|
||
write 16MB raw+fsync 0.0231s 43 ops/s 691.7 MB/s
|
||
read 16MB raw+fsync 0.0012s 855 ops/s 13672.2 MB/s
|
||
write 64MB raw+fsync 0.0803s 12 ops/s 797.2 MB/s
|
||
read 64MB raw+fsync 0.0048s 209 ops/s 13393.5 MB/s
|
||
|
||
== 8. random-access read: mmap'd pack vs. raw fs ==
|
||
compact N entries to pack pack 0.0267s 747839 ops/s 91.3 MB/s
|
||
random-read N entries pack 0.0082s 2435930 ops/s 297.4 MB/s
|
||
random-read N entries raw fs 0.1629s 122749 ops/s 15.0 MB/s
|
||
|
||
== 9. concurrency (create+read+unlink) ==
|
||
concurrent create+read+unlink mem 0.2750s 349127 ops/s
|
||
concurrent create+read+unlink raw fs 5.6498s 16992 ops/s
|
||
|
||
== 10. mount table scaling (the mount table uses the same full-snapshot-copy pattern the file index used to) ==
|
||
mount 500 backends vfs 0.0054s 92833 ops/s
|
||
resolve, 500 mounts vfs 0.0020s 246064 ops/s
|
||
unmount 500 backends vfs 0.0046s 109207 ops/s
|
||
mount 2000 backends vfs 0.0831s 24081 ops/s
|
||
resolve, 2000 mounts vfs 0.0296s 67458 ops/s
|
||
unmount 2000 backends vfs 0.0758s 26397 ops/s
|
||
mount 8000 backends vfs 1.4739s 5428 ops/s
|
||
resolve, 8000 mounts vfs 0.4938s 16200 ops/s
|
||
unmount 8000 backends vfs 1.5172s 5273 ops/s
|
||
|
||
=== summary (54 measurements) ===
|
||
category backend seconds throughput MB/s
|
||
create N small files mem 0.0400s 499809 ops/s 61.0 MB/s
|
||
create N small files raw fs 0.9943s 20114 ops/s 2.5 MB/s
|
||
create N small files raw+fsync 116.2147s 172 ops/s 0.0 MB/s
|
||
read N small files mem 0.0099s 2011901 ops/s 245.6 MB/s
|
||
read N small files raw fs 0.1606s 124543 ops/s 15.2 MB/s
|
||
stat N files mem 0.0077s 2598241 ops/s
|
||
stat N files raw fs 0.0699s 286222 ops/s
|
||
readdir (N entries) mem 0.0041s 4845826 ops/s
|
||
readdir (N entries) raw fs 0.0069s 2894815 ops/s
|
||
create N small files dir 1.1418s 17517 ops/s 2.1 MB/s
|
||
read N small files dir 0.1530s 130705 ops/s 16.0 MB/s
|
||
stat N files dir 0.1379s 145056 ops/s
|
||
readdir (N entries) dir 0.0037s 5465162 ops/s
|
||
unlink N files mem 0.0179s 1117742 ops/s
|
||
unlink N files dir 0.4911s 40724 ops/s
|
||
unlink N files raw fs 0.4746s 42141 ops/s
|
||
mkdir N dirs mem 0.0056s 719696 ops/s
|
||
rmdir N dirs mem 0.0040s 993724 ops/s
|
||
mkdir N dirs dir 0.1710s 23386 ops/s
|
||
rmdir N dirs dir 0.1257s 31827 ops/s
|
||
mkdir N dirs raw fs 0.1374s 29122 ops/s
|
||
rmdir N dirs raw fs 0.0951s 42043 ops/s
|
||
write 1MB mem 0.0001s 7058 ops/s 7058.1 MB/s
|
||
read 1MB mem 0.0000s 50051 ops/s 50050.9 MB/s
|
||
write 16MB mem 0.0110s 91 ops/s 1458.2 MB/s
|
||
read 16MB mem 0.0009s 1058 ops/s 16920.1 MB/s
|
||
write 64MB mem 0.0586s 17 ops/s 1091.3 MB/s
|
||
read 64MB mem 0.0032s 313 ops/s 20056.5 MB/s
|
||
write 1MB raw fs 0.0004s 2759 ops/s 2759.4 MB/s
|
||
read 1MB raw fs 0.0001s 16812 ops/s 16812.4 MB/s
|
||
write 16MB raw fs 0.0044s 229 ops/s 3656.4 MB/s
|
||
read 16MB raw fs 0.0010s 989 ops/s 15827.9 MB/s
|
||
write 64MB raw fs 0.0186s 54 ops/s 3449.7 MB/s
|
||
read 64MB raw fs 0.0047s 213 ops/s 13633.1 MB/s
|
||
write 1MB raw+fsync 0.0180s 56 ops/s 55.7 MB/s
|
||
read 1MB raw+fsync 0.0001s 9059 ops/s 9058.8 MB/s
|
||
write 16MB raw+fsync 0.0231s 43 ops/s 691.7 MB/s
|
||
read 16MB raw+fsync 0.0012s 855 ops/s 13672.2 MB/s
|
||
write 64MB raw+fsync 0.0803s 12 ops/s 797.2 MB/s
|
||
read 64MB raw+fsync 0.0048s 209 ops/s 13393.5 MB/s
|
||
compact N entries to pack pack 0.0267s 747839 ops/s 91.3 MB/s
|
||
random-read N entries pack 0.0082s 2435930 ops/s 297.4 MB/s
|
||
random-read N entries raw fs 0.1629s 122749 ops/s 15.0 MB/s
|
||
concurrent create+read+unlink mem 0.2750s 349127 ops/s
|
||
concurrent create+read+unlink raw fs 5.6498s 16992 ops/s
|
||
mount 500 backends vfs 0.0054s 92833 ops/s
|
||
resolve, 500 mounts vfs 0.0020s 246064 ops/s
|
||
unmount 500 backends vfs 0.0046s 109207 ops/s
|
||
mount 2000 backends vfs 0.0831s 24081 ops/s
|
||
resolve, 2000 mounts vfs 0.0296s 67458 ops/s
|
||
unmount 2000 backends vfs 0.0758s 26397 ops/s
|
||
mount 8000 backends vfs 1.4739s 5428 ops/s
|
||
resolve, 8000 mounts vfs 0.4938s 16200 ops/s
|
||
unmount 8000 backends vfs 1.5172s 5273 ops/s
|
||
|
||
real 5m23.896s
|
||
user 0m4.732s
|
||
sys 0m19.047s
|
||
```
|
||
|
||
## Reproducing
|
||
|
||
```sh
|
||
make bench
|
||
```
|
||
|
||
Takes several minutes on a similarly slow-`fsync` environment, dominated
|
||
by the `raw+fsync` create test; expect well under a minute on a host with
|
||
normal disk or tmpfs `fsync` latency.
|