A audit for other instances of the file-index O(n^2) shape (fixed previously via the persistent treap) found two more real issues: 1. pack_write's Section 9.2 exact-duplicate elimination was a linear scan of every previously-seen (hash, size) pair per entry -- O(n^2) total, invisible in the existing benchmark because its identical-content test files made every scan match on the first comparison. Measured with unique content instead: 80,000 entries took 1.74s, with a 20,000->80,000 step showing 15.6x for a 4x-N step, matching O(n^2)'s 16x prediction. The same scan also trusted a (hash, size) match without ever comparing actual bytes -- a latent correctness bug, since FNV-1a64 is explicitly not collision-resistant. Fixed both at once with an open-addressing hash table (load factor 1/2, linear probing) plus a memcmp verification before ever reusing a data_off. Post-fix: 80,000 entries in 0.044s (39.6x faster), ratio drops to 3.35x (consistent with O(n)). Covered permanently by two new/extended tests: a white-box assertion in test_pack_overlay.c that duplicate-content entries share one data_off and distinct-content entries do not, and a new tests/test_pack_write_perf.c regression tripwire against 10,000 unique entries. 2. vfs.c's mount table uses the same full-array-copy-per-write pattern the file index used to, confirmed O(n^2) via a new bench/bench.c category (500/2,000/8,000 mounts, both 4x-N steps showing 15-20x). Deliberately NOT rewritten: mount points are created by a program's own source code, not workload-driven, so realistic mount counts never reach the scale that made the file index's O(n^2) a real problem. Documented with full reasoning in BENCH.md and CLAUDE.md rather than silently left as an undocumented gap. Also fixes a real CI gap the new tests exposed: ci.yml's sanitizer-build steps never passed -D_GNU_SOURCE when compiling test files (only the library .o's got it), which was harmless while no test included internal.h and became a link failure once two did (internal.h needs _GNU_SOURCE for pthread_rwlock_t). And documents, in CONTRIBUTING.md and CLAUDE.md, a sandbox flake observed directly during this work's own sanitizer runs: ASan/UBSan test binaries occasionally fail to start with AddressSanitizer:DEADLYSIGNAL (sometimes looping rather than exiting), non-deterministically hitting different unrelated binaries across runs -- a startup race, not a memory-safety bug, confirmed by clean passes on retry; sanitizer runs in such an environment should be timeout-wrapped. BENCH.md's "After" table and Appendix B are replaced with the current, complete 54-measurement bench/bench.c run (the original 45 plus the new mount-scaling category); the pre-fix 45-measurement "Before" table is kept as the historical record, per this project's documentation standard. Verified: make test (all 6 binaries, including the 2 new/changed), a clean make all, and repeated ASan+UBSan runs (0 real findings; the DEADLYSIGNAL flake above was observed and correctly distinguished from a real finding by re-running until a clean pass). TSan could not be run in this sandbox (pre-existing, documented environment limitation). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
206 lines
9.0 KiB
Markdown
206 lines
9.0 KiB
Markdown
# PackFS
|
|
|
|
A statically linked, in-process virtual file system for C. PackFS treats a
|
|
shipped file tree as an immutable **pack** image and derives writability
|
|
from a `mem` or `dir` upper layer through a copy-on-write overlay. `zip` and
|
|
`tar` are not part of this project and never will be — see "Status" below.
|
|
|
|
The full design rationale — why this shape, what alternatives were rejected,
|
|
and the concurrency, path-containment, and integrity models this
|
|
implementation follows — is specified in [`concept.md`](concept.md), which is
|
|
frozen (see [`CLAUDE.md`](CLAUDE.md)) and is the authoritative source for
|
|
every design decision below. This README documents the implementation that
|
|
followed from it, not a restatement of the rationale.
|
|
|
|
## Status
|
|
|
|
This is an initial, partial implementation of the spec, not a complete one.
|
|
|
|
**Implemented and tested:** mount table, `mem`/`dir`/`pack`/overlay backends
|
|
(including the standalone read-only `pack` backend, `backend_pack_new` —
|
|
every mutating call against it returns `VFS_ERR_PERM`), copy-up, whiteouts,
|
|
compaction, an append journal, path containment, and pack integrity
|
|
validation.
|
|
|
|
**Will never be built, by explicit project decision:** the `zip` (miniz) and
|
|
`tar` (USTAR) import/export backends. `concept.md` Section 11 recommends
|
|
them, but that recommendation is superseded — see `CLAUDE.md`, "Project
|
|
decisions that supersede concept.md." `backend_overlay_new` reads and writes
|
|
this project's own pack format exclusively; there is no zip/tar support and
|
|
none is planned. Do not open an issue or PR adding one.
|
|
|
|
**Deliberately out of scope for v0** (`concept.md` Section 10, not gaps):
|
|
full POSIX semantics, enforced permissions/symlinks/hard links,
|
|
cross-process concurrency (design specified in Section 5.6, unimplemented),
|
|
and content-defined chunking/delta compression.
|
|
|
|
**Tested on Linux only**, in one environment. The `dir`-mount containment
|
|
fallback path for kernels without `openat2` (Section 6.3) is implemented but
|
|
has not been exercised on such a kernel, nor on macOS or Windows.
|
|
|
|
## Building
|
|
|
|
Zero required third-party dependencies — only a C11 compiler, `make`, and
|
|
`pthread` (Section 11.1 of `concept.md` makes this a hard constraint, not a
|
|
preference).
|
|
|
|
```sh
|
|
make # builds libpackfs.a and libpackfs.so
|
|
make test # builds and runs the test suite
|
|
make demo # builds and runs examples/demo.c — see "Try it" below
|
|
make bench # builds and runs bench/bench.c — see "Benchmarks" below
|
|
make install # installs to $PREFIX (default /usr/local)
|
|
```
|
|
|
|
## Try it
|
|
|
|
`examples/demo.c` is a small, runnable, human-readable program — not another
|
|
automated test — that exercises the library end to end and prints what it
|
|
did at each step: a pack-backed overlay (write, read, `mkdir`, `readdir`,
|
|
`stat`, a copy-up-then-whiteout delete, compaction via `vfs_sync`), and a
|
|
sandboxed `dir` mount that demonstrates a `../../../etc/passwd` escape
|
|
attempt being rejected. Run `make demo` twice in a row: the second run's
|
|
first `readdir` shows the first run's files, proving that compaction and
|
|
reload actually persist data through the pack file, not just within one
|
|
process's lifetime.
|
|
|
|
## Benchmarks
|
|
|
|
`bench/bench.c` (`make bench`) measures PackFS against the host filesystem
|
|
across metadata operations (create/read/stat/readdir/unlink/mkdir), large
|
|
sequential I/O, random-access pack reads, mount-table scaling, and
|
|
concurrent mixed workloads. [`BENCH.md`](BENCH.md) has the full results and
|
|
honest analysis of both the wins and the losses, including the story of
|
|
three real O(n²) findings this project's own benchmarking turned up — not
|
|
just the wins. Two are fixed: bulk sequential file creation (`src/upper.c`'s
|
|
index is a persistent treap now, not a flat array — see "Resolution") and
|
|
`pack_write`'s compaction-time duplicate-content elimination, which also had
|
|
a latent correctness bug now closed alongside it (see "Resolution #2"). One
|
|
is confirmed and *deliberately not* fixed: the mount table scales O(n²) in
|
|
mount count, the same way the file index used to, but mount counts are
|
|
bounded by a program's own source code rather than workload-driven, so it
|
|
isn't worth the added complexity — see "Finding: mount table scaling" for
|
|
the reasoning and the numbers behind that call. Read the whole file before
|
|
quoting a number from it: what `raw fs` vs `raw+fsync` vs `dir` each
|
|
actually measure is not interchangeable, and it explains why.
|
|
|
|
`openat2`/Landlock support (Section 6) is detected automatically at compile
|
|
time via `<sys/syscall.h>`; on kernels or platforms without them, `dir`
|
|
mounts fall back to the weaker, documented residual-risk posture described
|
|
in `concept.md` Section 6.3 rather than failing to build.
|
|
|
|
## Quick example
|
|
|
|
```c
|
|
#include <packfs.h>
|
|
|
|
Vfs *v = vfs_new();
|
|
|
|
/* a plain in-memory writable tree */
|
|
Backend *mem = backend_mem_new();
|
|
vfs_mount(v, "/", mem);
|
|
|
|
int err = 0;
|
|
VfsFile *f = vfs_open(v, "/hello.txt", VFS_O_WRONLY | VFS_O_CREAT, &err);
|
|
vfs_write(f, "hello", 5);
|
|
vfs_close(f);
|
|
|
|
vfs_free(v);
|
|
backend_free(mem);
|
|
```
|
|
|
|
A shipped pack with a writable overlay on top:
|
|
|
|
```c
|
|
Backend *mem = backend_mem_new();
|
|
int err = 0;
|
|
Backend *ov = backend_overlay_new("assets.pack", mem, &err); /* loads assets.pack if it exists */
|
|
vfs_mount(v, "/", ov);
|
|
|
|
/* ... reads served straight from the pack; writes copy-up into mem ... */
|
|
|
|
vfs_sync(v, "/"); /* compacts the overlay into a fresh assets.pack (Section 4.1, 5.3) */
|
|
```
|
|
|
|
A sandboxed host directory:
|
|
|
|
```c
|
|
int derr = 0;
|
|
Backend *dir = backend_dir_new("/var/lib/myapp/data", &derr);
|
|
vfs_mount(v, "/data", dir);
|
|
/* every lookup under /data is contained to that directory (Section 6.2);
|
|
* ".." and symlink escapes are rejected, not merely discouraged. */
|
|
```
|
|
|
|
A shipped pack mounted read-only, no writable layer at all:
|
|
|
|
```c
|
|
int perr = 0;
|
|
Backend *ro = backend_pack_new("assets.pack", &perr);
|
|
vfs_mount(v, "/assets", ro);
|
|
/* vfs_write, vfs_mkdir, vfs_unlink, vfs_rename, vfs_sync against anything
|
|
* under /assets all return VFS_ERR_PERM; there is no upper layer to
|
|
* absorb a write into. */
|
|
```
|
|
|
|
## API
|
|
|
|
The public API is [`include/packfs.h`](include/packfs.h); every function and
|
|
struct is documented there with a pointer to the `concept.md` section that
|
|
specifies its behavior. In outline:
|
|
|
|
- `vfs_new` / `vfs_free` — a `Vfs` owns a mount table, nothing else.
|
|
- `backend_mem_new` / `backend_dir_new` / `backend_pack_new` /
|
|
`backend_overlay_new` — construct a backend; `vfs_mount`/`vfs_unmount`
|
|
attach or detach it at a path prefix. `backend_pack_new` mounts a pack
|
|
read-only, with no writable upper layer at all — every mutating call
|
|
against it returns `VFS_ERR_PERM`; wrap the same pack in
|
|
`backend_overlay_new` instead when writability is wanted.
|
|
- `vfs_open` / `vfs_read` / `vfs_write` / `vfs_close` — file I/O.
|
|
- `vfs_stat` / `vfs_readdir` / `vfs_mkdir` / `vfs_unlink` / `vfs_rename` —
|
|
metadata and namespace operations.
|
|
- `vfs_sync` — compacts an overlay mount into a fresh pack.
|
|
- `vfs_harden_process_with_landlock` — optional, opt-in, process-wide
|
|
Landlock confinement to the process's current `dir` mounts. Deliberately
|
|
**not** applied automatically by `vfs_mount` (Section 6.2 explains why:
|
|
Landlock restrictions are irreversible and process-wide, which would be a
|
|
surprising side effect for an embeddable library to trigger on its own).
|
|
|
|
## Concurrency
|
|
|
|
Single-writer, wait-free-reader (Section 5): readers never take a lock and
|
|
never observe a write in progress; structural changes (create/unlink/rename/
|
|
mkdir/mount/unmount) publish a new immutable snapshot via one atomic pointer
|
|
swap; ordinary content writes to an already-existing file update a per-entry
|
|
cell directly and never touch the snapshot. `mem`-backed buffer growth
|
|
never mutates a buffer address a reader might be reading (Section 5.7):
|
|
growth always allocates a new buffer and publishes it, never reallocates in
|
|
place. `tests/test_concurrency.c` exercises this under concurrent reader and
|
|
writer threads, and the suite is regularly run under ThreadSanitizer and
|
|
AddressSanitizer (see `.github/workflows/ci.yml`).
|
|
|
|
## Security
|
|
|
|
`dir` mounts are capability-scoped (Section 6.2): a mount holds an already-open
|
|
directory file descriptor, not a path string, and every lookup beneath it is
|
|
resolved with `openat2(RESOLVE_BENEATH | RESOLVE_NO_SYMLINKS)` on Linux 5.6+,
|
|
which atomically rejects `..` and symlink escapes in one kernel call. Where
|
|
that syscall is unavailable, containment falls back to per-component
|
|
`O_NOFOLLOW` resolution — weaker, and documented as such (Section 6.3), not
|
|
silently assumed equivalent. Pack files are treated as untrusted input unless
|
|
they came from this process's own compaction: every on-disk offset is
|
|
bounds-checked and the pack's checksum is verified before any of it is
|
|
trusted (Section 7). See `tests/test_dir.c` for a containment regression test
|
|
and `tests/test_pack_overlay.c` for a corrupted-pack rejection test.
|
|
|
|
## Contributing
|
|
|
|
See [`CONTRIBUTING.md`](CONTRIBUTING.md). Read `concept.md` and `CLAUDE.md`
|
|
first — they are the project's actual specification and its enforced
|
|
documentation standard, respectively, and every design decision in the code
|
|
traces back to one of them.
|
|
|
|
## License
|
|
|
|
MIT — see [`LICENSE`](LICENSE).
|