# PackFS A statically linked, in-process virtual file system for C. PackFS treats a shipped file tree as an immutable **pack** image and derives writability from a `mem` or `dir` upper layer through a copy-on-write overlay. `zip` and `tar` are not part of this project and never will be — see "Status" below. The full design rationale — why this shape, what alternatives were rejected, and the concurrency, path-containment, and integrity models this implementation follows — is specified in [`concept.md`](concept.md), which is frozen (see [`CLAUDE.md`](CLAUDE.md)) and is the authoritative source for every design decision below. This README documents the implementation that followed from it, not a restatement of the rationale. ## Status This is an initial, partial implementation of the spec, not a complete one. **Implemented and tested:** mount table, `mem`/`dir`/`pack`/overlay backends (including the standalone read-only `pack` backend, `backend_pack_new` — every mutating call against it returns `VFS_ERR_PERM`), copy-up, whiteouts, compaction, an append journal, path containment, and pack integrity validation. **Will never be built, by explicit project decision:** the `zip` (miniz) and `tar` (USTAR) import/export backends. `concept.md` Section 11 recommends them, but that recommendation is superseded — see `CLAUDE.md`, "Project decisions that supersede concept.md." `backend_overlay_new` reads and writes this project's own pack format exclusively; there is no zip/tar support and none is planned. Do not open an issue or PR adding one. **Deliberately out of scope for v0** (`concept.md` Section 10, not gaps): full POSIX semantics, enforced permissions/symlinks/hard links, cross-process concurrency (design specified in Section 5.6, unimplemented), and content-defined chunking/delta compression. **Tested on Linux only**, in one environment. The `dir`-mount containment fallback path for kernels without `openat2` (Section 6.3) is implemented but has not been exercised on such a kernel, nor on macOS or Windows. ## Building Zero required third-party dependencies — only a C11 compiler, `make`, and `pthread` (Section 11.1 of `concept.md` makes this a hard constraint, not a preference). ```sh make # builds libpackfs.a and libpackfs.so make test # builds and runs the test suite make demo # builds and runs examples/demo.c — see "Try it" below make bench # builds and runs bench/bench.c — see "Benchmarks" below make install # installs to $PREFIX (default /usr/local), including a pkg-config file ``` `make install` also generates and installs `packfs.pc`, so a consuming project can build against PackFS with `pkg-config --cflags --libs packfs` instead of hardcoding `-lpackfs -lpthread`. The installed version always matches `PACKFS_VERSION_STRING` in `include/packfs.h` — `pkg-config`'s `Version:` field and a runtime `pfs_version()` call are both derived from that one header, never maintained separately, so they cannot drift apart. ## Try it `examples/demo.c` is a small, runnable, human-readable program — not another automated test — that exercises the library end to end and prints what it did at each step: a pack-backed overlay (write, read, `mkdir`, `readdir`, `stat`, a copy-up-then-whiteout delete, compaction via `vfs_sync`), and a sandboxed `dir` mount that demonstrates a `../../../etc/passwd` escape attempt being rejected. Run `make demo` twice in a row: the second run's first `readdir` shows the first run's files, proving that compaction and reload actually persist data through the pack file, not just within one process's lifetime. ## Benchmarks `bench/bench.c` (`make bench`) measures PackFS against the host filesystem across metadata operations (create/read/stat/readdir/unlink/mkdir), large sequential I/O, random-access pack reads, mount-table scaling, and concurrent mixed workloads. [`BENCH.md`](BENCH.md) has the full results and honest analysis of both the wins and the losses, including the story of three real O(n²) findings this project's own benchmarking turned up — not just the wins. Two are fixed: bulk sequential file creation (`src/upper.c`'s index is a persistent treap now, not a flat array — see "Resolution") and `pack_write`'s compaction-time duplicate-content elimination, which also had a latent correctness bug now closed alongside it (see "Resolution #2"). One is confirmed and *deliberately not* fixed: the mount table scales O(n²) in mount count, the same way the file index used to, but mount counts are bounded by a program's own source code rather than workload-driven, so it isn't worth the added complexity — see "Finding: mount table scaling" for the reasoning and the numbers behind that call. Read the whole file before quoting a number from it: what `raw fs` vs `raw+fsync` vs `dir` each actually measure is not interchangeable, and it explains why. `openat2`/Landlock support (Section 6) is detected automatically at compile time via ``; on kernels or platforms without them, `dir` mounts fall back to the weaker, documented residual-risk posture described in `concept.md` Section 6.3 rather than failing to build. ### Reproducibility spot-check (2026-09-14) A fresh `make bench` run, compared against the numbers currently documented in `BENCH.md`'s "After" table, on the same environment described there. Not a replacement for `BENCH.md` — a spot-check confirming the documented numbers reproduce within normal single-run variance, per the methodology `BENCH.md` itself states ("illustrative of shape... not precise absolute figures"). | Category | Backend | Documented (BENCH.md) | New run | Delta | |---|---|---|---|---| | create 20,000 files | mem | 0.0400s | 0.0394s | -1.5% | | create 20,000 files | raw fs | 0.9943s | 1.0171s | +2.3% | | create 20,000 files | raw+fsync | 116.2147s | 116.7567s | +0.5% | | read 20,000 files | mem | 0.0099s | 0.0091s | -8.1% | | read 20,000 files | raw fs | 0.1606s | 0.1634s | +1.7% | | stat 20,000 files | mem | 0.0077s | 0.0078s | +1.3% | | stat 20,000 files | raw fs | 0.0699s | 0.0727s | +4.0% | | readdir (20,000 entries) | mem | 0.0041s | 0.0043s | +4.9% | | readdir (20,000 entries) | raw fs | 0.0069s | 0.0072s | +4.3% | | create 20,000 files | dir | 1.1418s | 1.2030s | +5.4% | | read 20,000 files | dir | 0.1530s | 0.1715s | +12.1% | | stat 20,000 files | dir | 0.1379s | 0.1546s | +12.1% | | readdir (20,000 entries) | dir | 0.0037s | 0.0041s | +10.8% | | unlink 20,000 files | mem | 0.0179s | 0.0182s | +1.7% | | unlink 20,000 files | dir | 0.4911s | 0.5978s | +21.7% | | unlink 20,000 files | raw fs | 0.4746s | 0.5062s | +6.7% | | mkdir 4,000 dirs | mem | 0.0056s | 0.0050s | -10.7% | | rmdir 4,000 dirs | mem | 0.0040s | 0.0039s | -2.5% | | mkdir 4,000 dirs | dir | 0.1710s | 0.1864s | +9.0% | | rmdir 4,000 dirs | dir | 0.1257s | 0.1169s | -7.0% | | mkdir 4,000 dirs | raw fs | 0.1374s | 0.1444s | +5.1% | | rmdir 4,000 dirs | raw fs | 0.0951s | 0.1005s | +5.7% | | write 1MB | mem | 0.0001s | 0.0001s | +0.0% | | read 1MB | mem | 0.0000s | 0.0000s | +0.0% | | write 16MB | mem | 0.0110s | 0.0110s | +0.0% | | read 16MB | mem | 0.0009s | 0.0010s | +11.1% | | write 64MB | mem | 0.0586s | 0.0621s | +6.0% | | read 64MB | mem | 0.0032s | 0.0036s | +12.5% | | write 1MB | raw fs | 0.0004s | 0.0004s | +0.0% | | read 1MB | raw fs | 0.0001s | 0.0001s | +0.0% | | write 16MB | raw fs | 0.0044s | 0.0047s | +6.8% | | read 16MB | raw fs | 0.0010s | 0.0012s | +20.0% | | write 64MB | raw fs | 0.0186s | 0.0189s | +1.6% | | read 64MB | raw fs | 0.0047s | 0.0050s | +6.4% | | write 1MB | raw+fsync | 0.0180s | 0.0340s | +88.9% | | read 1MB | raw+fsync | 0.0001s | 0.0001s | +0.0% | | write 16MB | raw+fsync | 0.0231s | 0.0242s | +4.8% | | read 16MB | raw+fsync | 0.0012s | 0.0012s | +0.0% | | write 64MB | raw+fsync | 0.0803s | 0.0835s | +4.0% | | read 64MB | raw+fsync | 0.0048s | 0.0049s | +2.1% | | compact 20,000 entries to pack | pack | 0.0267s | 0.0280s | +4.9% | | random-read 20,000 entries | pack (mmap'd) | 0.0082s | 0.0085s | +3.7% | | random-read 20,000 entries | raw fs | 0.1629s | 0.1912s | +17.4% | | concurrent create+read+unlink (8×4,000×3) | mem | 0.2750s | 0.2791s | +1.5% | | concurrent create+read+unlink (8×4,000×3) | raw fs | 5.6498s | 6.1583s | +9.0% | | mount 500 backends | vfs | 0.0054s | 0.0060s | +11.1% | | resolve, 500 mounts | vfs | 0.0020s | 0.0022s | +10.0% | | unmount 500 backends | vfs | 0.0046s | 0.0046s | +0.0% | | mount 2,000 backends | vfs | 0.0831s | 0.0863s | +3.9% | | resolve, 2,000 mounts | vfs | 0.0296s | 0.0329s | +11.1% | | unmount 2,000 backends | vfs | 0.0758s | 0.0780s | +2.9% | | mount 8,000 backends | vfs | 1.4739s | 1.5380s | +4.3% | | resolve, 8,000 mounts | vfs | 0.4938s | 0.5298s | +7.3% | | unmount 8,000 backends | vfs | 1.5172s | 1.5191s | +0.1% | Almost every row sits within ±15% of the documented figures — consistent with the single-run jitter `BENCH.md` already warns about, not a regression. Two rows exceed that: `unlink 20,000 files (dir)` (+21.7%, plausible container/host I/O noise, same direction as the other `dir` metadata rows this run) and `write 1MB (raw+fsync)` (+88.9%, but this is a tiny absolute value — 18ms vs 34ms for one syscall — the single number most sensitive to one slow `fsync` on this container's overlay filesystem, not a meaningful regression at that scale). Total wall-clock: 4m7.7s this run vs 5m23.9s documented, itself within the same single-run variance. ## Quick example ```c #include Vfs *v = vfs_new(); /* a plain in-memory writable tree */ Backend *mem = backend_mem_new(); vfs_mount(v, "/", mem); int err = 0; VfsFile *f = vfs_open(v, "/hello.txt", VFS_O_WRONLY | VFS_O_CREAT, &err); vfs_write(f, "hello", 5); vfs_close(f); vfs_free(v); backend_free(mem); ``` A shipped pack with a writable overlay on top: ```c Backend *mem = backend_mem_new(); int err = 0; Backend *ov = backend_overlay_new("assets.pack", mem, &err); /* loads assets.pack if it exists */ vfs_mount(v, "/", ov); /* ... reads served straight from the pack; writes copy-up into mem ... */ vfs_sync(v, "/"); /* compacts the overlay into a fresh assets.pack (Section 4.1, 5.3) */ ``` A sandboxed host directory: ```c int derr = 0; Backend *dir = backend_dir_new("/var/lib/myapp/data", &derr); vfs_mount(v, "/data", dir); /* every lookup under /data is contained to that directory (Section 6.2); * ".." and symlink escapes are rejected, not merely discouraged. */ ``` A shipped pack mounted read-only, no writable layer at all: ```c int perr = 0; Backend *ro = backend_pack_new("assets.pack", &perr); vfs_mount(v, "/assets", ro); /* vfs_write, vfs_mkdir, vfs_unlink, vfs_rename, vfs_sync against anything * under /assets all return VFS_ERR_PERM; there is no upper layer to * absorb a write into. */ ``` ## API The public API is [`include/packfs.h`](include/packfs.h); every function and struct is documented there with a pointer to the `concept.md` section that specifies its behavior. In outline: - `vfs_new` / `vfs_free` — a `Vfs` owns a mount table, nothing else. - `backend_mem_new` / `backend_dir_new` / `backend_pack_new` / `backend_overlay_new` — construct a backend; `vfs_mount`/`vfs_unmount` attach or detach it at a path prefix. `backend_pack_new` mounts a pack read-only, with no writable upper layer at all — every mutating call against it returns `VFS_ERR_PERM`; wrap the same pack in `backend_overlay_new` instead when writability is wanted. - `vfs_open` / `vfs_read` / `vfs_write` / `vfs_close` — file I/O. - `vfs_stat` / `vfs_readdir` / `vfs_mkdir` / `vfs_unlink` / `vfs_rename` — metadata and namespace operations. - `vfs_sync` — compacts an overlay mount into a fresh pack. - `vfs_harden_process_with_landlock` — optional, opt-in, process-wide Landlock confinement to the process's current `dir` mounts. Deliberately **not** applied automatically by `vfs_mount` (Section 6.2 explains why: Landlock restrictions are irreversible and process-wide, which would be a surprising side effect for an embeddable library to trigger on its own). ## Concurrency Single-writer, wait-free-reader (Section 5): readers never take a lock and never observe a write in progress; structural changes (create/unlink/rename/ mkdir/mount/unmount) publish a new immutable snapshot via one atomic pointer swap; ordinary content writes to an already-existing file update a per-entry cell directly and never touch the snapshot. `mem`-backed buffer growth never mutates a buffer address a reader might be reading (Section 5.7): growth always allocates a new buffer and publishes it, never reallocates in place. `tests/test_concurrency.c` exercises this under concurrent reader and writer threads. The suite is regularly run under AddressSanitizer/ UndefinedBehaviorSanitizer, both locally and in CI (see `.gitea/workflows/ci.yml`), and this is not aspirational — ASan caught a real heap-use-after-free in the snapshot-reclamation logic during development (see the `reclaim_gate` note in `src/internal.h`). ThreadSanitizer is configured the same way but, as of this writing, has not actually completed a run in either environment this project has been built and tested in so far — the local development sandbox and this project's own CI runner both block the `personality(ADDR_NO_RANDOMIZE)` syscall TSan needs to start (see `CONTRIBUTING.md` for the confirming tests in each case). CI's TSan step is written to not fail the build over that specific, known-benign failure, which means a green CI run is not evidence TSan actually executed — stated plainly here rather than left to be assumed from CI showing green. ## Security `dir` mounts are capability-scoped (Section 6.2): a mount holds an already-open directory file descriptor, not a path string, and every lookup beneath it is resolved with `openat2(RESOLVE_BENEATH | RESOLVE_NO_SYMLINKS)` on Linux 5.6+, which atomically rejects `..` and symlink escapes in one kernel call. Where that syscall is unavailable, containment falls back to per-component `O_NOFOLLOW` resolution — weaker, and documented as such (Section 6.3), not silently assumed equivalent. Pack files are treated as untrusted input unless they came from this process's own compaction: every on-disk offset is bounds-checked and the pack's checksum is verified before any of it is trusted (Section 7). See `tests/test_dir.c` for a containment regression test and `tests/test_pack_overlay.c` for a corrupted-pack rejection test. See [`SECURITY.md`](SECURITY.md) for exactly what is and isn't claimed as a security boundary, and how to report a vulnerability. ## Contributing See [`CONTRIBUTING.md`](CONTRIBUTING.md). Read `concept.md` and `CLAUDE.md` first — they are the project's actual specification and its enforced documentation standard, respectively, and every design decision in the code traces back to one of them. ## License MIT — see [`LICENSE`](LICENSE).