Doing another literal, no-caveats audit pass (per the standing request to verify literally everything, not just reassert it) found real, concrete inconsistencies introduced by the previous commit: it replaced BENCH.md's "After: complete results" table and Appendix B with a newer, more complete bench/bench.c run (adding the mount-table-scaling category), but several other places in the same document still quoted exact figures from the older run those tables used to hold -- a reader cross-checking the "Resolution" section's headline table, or the "Analysis" section's bullet points, against the "After" table below them would have found contradictory numbers for the same measurement. Verified programmatically, not just by re-reading: every row's time, throughput, and MB/s in the Before (45 rows) and After (54 rows) tables now matches its corresponding appendix line exactly (0 mismatches out of 99 rows, checked with a script, not by eye). Fixed: - The "Resolution" section's headline before/after table used the older run's mem create/unlink/mkdir/rmdir numbers (0.0372s/0.0320s/0.0052s/ 0.0045s), which don't match the current After table (0.0400s/0.0179s/ 0.0056s/0.0040s) -- for unlink specifically, a genuine near-2x difference between two runs of identical, already-fixed code, not just noise. Now points at the single current run, with speedups recomputed (239x/590x/ 60x/85x/11x/21x/134x, mem now 25-27x faster than raw fs, complexity-class check now 7.14x vs the old 7.15x -- nearly identical, which is the reassurance that actually matters), plus an explicit note on the observed run-to-run variance rather than silently picking one number. - The "Analysis" section's read/stat/random-read/compaction/dir-stat/ large-write bullets all cited the older run's exact ops/s and MB/s figures (2.4M small-file reads, ~815K compaction entries/s, 133K dir stat, etc.); recomputed against the current After table (now 15-16x for small reads, ~9x for stat, ~91MB/s / ~748K entries/s for compaction, 145K vs 286K for dir stat, corrected large-write MB/s ranges). - The same stale 65-330x / 15-27x / 22x range in CLAUDE.md's "Known performance characteristics" bulk-write bullet, updated to match and to explain the same run-to-run variance rather than presenting one run's numbers as if they were exact. No code changes; make test still passes (all 6 binaries, unaffected by a documentation-only change). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
PackFS
A statically linked, in-process virtual file system for C. PackFS treats a
shipped file tree as an immutable pack image and derives writability
from a mem or dir upper layer through a copy-on-write overlay. zip and
tar are not part of this project and never will be — see "Status" below.
The full design rationale — why this shape, what alternatives were rejected,
and the concurrency, path-containment, and integrity models this
implementation follows — is specified in concept.md, which is
frozen (see CLAUDE.md) and is the authoritative source for
every design decision below. This README documents the implementation that
followed from it, not a restatement of the rationale.
Status
This is an initial, partial implementation of the spec, not a complete one.
Implemented and tested: mount table, mem/dir/pack/overlay backends
(including the standalone read-only pack backend, backend_pack_new —
every mutating call against it returns VFS_ERR_PERM), copy-up, whiteouts,
compaction, an append journal, path containment, and pack integrity
validation.
Will never be built, by explicit project decision: the zip (miniz) and
tar (USTAR) import/export backends. concept.md Section 11 recommends
them, but that recommendation is superseded — see CLAUDE.md, "Project
decisions that supersede concept.md." backend_overlay_new reads and writes
this project's own pack format exclusively; there is no zip/tar support and
none is planned. Do not open an issue or PR adding one.
Deliberately out of scope for v0 (concept.md Section 10, not gaps):
full POSIX semantics, enforced permissions/symlinks/hard links,
cross-process concurrency (design specified in Section 5.6, unimplemented),
and content-defined chunking/delta compression.
Tested on Linux only, in one environment. The dir-mount containment
fallback path for kernels without openat2 (Section 6.3) is implemented but
has not been exercised on such a kernel, nor on macOS or Windows.
Building
Zero required third-party dependencies — only a C11 compiler, make, and
pthread (Section 11.1 of concept.md makes this a hard constraint, not a
preference).
make # builds libpackfs.a and libpackfs.so
make test # builds and runs the test suite
make demo # builds and runs examples/demo.c — see "Try it" below
make bench # builds and runs bench/bench.c — see "Benchmarks" below
make install # installs to $PREFIX (default /usr/local)
Try it
examples/demo.c is a small, runnable, human-readable program — not another
automated test — that exercises the library end to end and prints what it
did at each step: a pack-backed overlay (write, read, mkdir, readdir,
stat, a copy-up-then-whiteout delete, compaction via vfs_sync), and a
sandboxed dir mount that demonstrates a ../../../etc/passwd escape
attempt being rejected. Run make demo twice in a row: the second run's
first readdir shows the first run's files, proving that compaction and
reload actually persist data through the pack file, not just within one
process's lifetime.
Benchmarks
bench/bench.c (make bench) measures PackFS against the host filesystem
across metadata operations (create/read/stat/readdir/unlink/mkdir), large
sequential I/O, random-access pack reads, mount-table scaling, and
concurrent mixed workloads. BENCH.md has the full results and
honest analysis of both the wins and the losses, including the story of
three real O(n²) findings this project's own benchmarking turned up — not
just the wins. Two are fixed: bulk sequential file creation (src/upper.c's
index is a persistent treap now, not a flat array — see "Resolution") and
pack_write's compaction-time duplicate-content elimination, which also had
a latent correctness bug now closed alongside it (see "Resolution #2"). One
is confirmed and deliberately not fixed: the mount table scales O(n²) in
mount count, the same way the file index used to, but mount counts are
bounded by a program's own source code rather than workload-driven, so it
isn't worth the added complexity — see "Finding: mount table scaling" for
the reasoning and the numbers behind that call. Read the whole file before
quoting a number from it: what raw fs vs raw+fsync vs dir each
actually measure is not interchangeable, and it explains why.
openat2/Landlock support (Section 6) is detected automatically at compile
time via <sys/syscall.h>; on kernels or platforms without them, dir
mounts fall back to the weaker, documented residual-risk posture described
in concept.md Section 6.3 rather than failing to build.
Quick example
#include <packfs.h>
Vfs *v = vfs_new();
/* a plain in-memory writable tree */
Backend *mem = backend_mem_new();
vfs_mount(v, "/", mem);
int err = 0;
VfsFile *f = vfs_open(v, "/hello.txt", VFS_O_WRONLY | VFS_O_CREAT, &err);
vfs_write(f, "hello", 5);
vfs_close(f);
vfs_free(v);
backend_free(mem);
A shipped pack with a writable overlay on top:
Backend *mem = backend_mem_new();
int err = 0;
Backend *ov = backend_overlay_new("assets.pack", mem, &err); /* loads assets.pack if it exists */
vfs_mount(v, "/", ov);
/* ... reads served straight from the pack; writes copy-up into mem ... */
vfs_sync(v, "/"); /* compacts the overlay into a fresh assets.pack (Section 4.1, 5.3) */
A sandboxed host directory:
int derr = 0;
Backend *dir = backend_dir_new("/var/lib/myapp/data", &derr);
vfs_mount(v, "/data", dir);
/* every lookup under /data is contained to that directory (Section 6.2);
* ".." and symlink escapes are rejected, not merely discouraged. */
A shipped pack mounted read-only, no writable layer at all:
int perr = 0;
Backend *ro = backend_pack_new("assets.pack", &perr);
vfs_mount(v, "/assets", ro);
/* vfs_write, vfs_mkdir, vfs_unlink, vfs_rename, vfs_sync against anything
* under /assets all return VFS_ERR_PERM; there is no upper layer to
* absorb a write into. */
API
The public API is include/packfs.h; every function and
struct is documented there with a pointer to the concept.md section that
specifies its behavior. In outline:
vfs_new/vfs_free— aVfsowns a mount table, nothing else.backend_mem_new/backend_dir_new/backend_pack_new/backend_overlay_new— construct a backend;vfs_mount/vfs_unmountattach or detach it at a path prefix.backend_pack_newmounts a pack read-only, with no writable upper layer at all — every mutating call against it returnsVFS_ERR_PERM; wrap the same pack inbackend_overlay_newinstead when writability is wanted.vfs_open/vfs_read/vfs_write/vfs_close— file I/O.vfs_stat/vfs_readdir/vfs_mkdir/vfs_unlink/vfs_rename— metadata and namespace operations.vfs_sync— compacts an overlay mount into a fresh pack.vfs_harden_process_with_landlock— optional, opt-in, process-wide Landlock confinement to the process's currentdirmounts. Deliberately not applied automatically byvfs_mount(Section 6.2 explains why: Landlock restrictions are irreversible and process-wide, which would be a surprising side effect for an embeddable library to trigger on its own).
Concurrency
Single-writer, wait-free-reader (Section 5): readers never take a lock and
never observe a write in progress; structural changes (create/unlink/rename/
mkdir/mount/unmount) publish a new immutable snapshot via one atomic pointer
swap; ordinary content writes to an already-existing file update a per-entry
cell directly and never touch the snapshot. mem-backed buffer growth
never mutates a buffer address a reader might be reading (Section 5.7):
growth always allocates a new buffer and publishes it, never reallocates in
place. tests/test_concurrency.c exercises this under concurrent reader and
writer threads, and the suite is regularly run under ThreadSanitizer and
AddressSanitizer (see .github/workflows/ci.yml).
Security
dir mounts are capability-scoped (Section 6.2): a mount holds an already-open
directory file descriptor, not a path string, and every lookup beneath it is
resolved with openat2(RESOLVE_BENEATH | RESOLVE_NO_SYMLINKS) on Linux 5.6+,
which atomically rejects .. and symlink escapes in one kernel call. Where
that syscall is unavailable, containment falls back to per-component
O_NOFOLLOW resolution — weaker, and documented as such (Section 6.3), not
silently assumed equivalent. Pack files are treated as untrusted input unless
they came from this process's own compaction: every on-disk offset is
bounds-checked and the pack's checksum is verified before any of it is
trusted (Section 7). See tests/test_dir.c for a containment regression test
and tests/test_pack_overlay.c for a corrupted-pack rejection test.
Contributing
See CONTRIBUTING.md. Read concept.md and CLAUDE.md
first — they are the project's actual specification and its enforced
documentation standard, respectively, and every design decision in the code
traces back to one of them.
License
MIT — see LICENSE.