BENCH.md's earlier finding was real: every structural write (create/
unlink/mkdir) copied the entire sorted UpperSnapshot entry array before
publishing the next snapshot, making bulk sequential creation O(n^2).
concept.md Section 5.3 names the exact condition for reconsidering this
("a persistent structurally-shared tree structure is not required
until this assumption is empirically violated") -- that condition was
measured, not hypothesized, so this closes it rather than leaving it
as a documented-but-open limitation.
The index is now a persistent treap (src/upper.c): a structural write
copies only the O(log n) nodes on the path to the change, sharing
every other node (its own refcount, cascading like MutCell's and
UpperSnapshot's) with whichever snapshot(s) it was built from. Chosen
over a persistent AVL/red-black/weight-balanced tree because deletion
in those can need O(log n) rebalancing rotations -- each a real
allocation in a persistent setting -- where a treap needs only O(1)
amortized rotations for insert and delete (Seidel & Aragon 1996), with
expected O(log n) height regardless of insertion order, including the
sorted-by-creation-order pattern that made the flat array quadratic in
the first place. Liljenzin's "Confluently Persistent Sets and Maps"
(arXiv:1301.3388) documents persistent treaps giving O(1) snapshots
for MVCC specifically, which is this exact use case. Full reasoning
and the ownership convention (functions consume one ref of their tree
arguments, return one owned ref) are in upper.c's comment above
struct TreapNode.
Blast radius kept deliberately small: snapshot_upsert/snapshot_remove
keep their exact original signatures, so upper_create/upper_mkdir/
upper_remove/upper_rename/upper_copy_up needed zero changes.
upper_lookup/upper_has_children keep their exact contracts. Only
upstd_readdir and overlay_readdir's manual array scans became calls to
a new upper_visit_range (O(log n + r) range query, replacing an O(n)
scan in both, a bonus fix beyond what was strictly necessary) since
there's no flat array left to scan.
Added tests/test_index_stress.c: thousands of randomized (not
sequential) creates/deletes/renames across nested directories,
cross-checked against an independent reference model after every
round, not just "did it not crash" -- exercises exactly the code path
the O(n^2) bug and this fix live in, at a scale the other tests don't
reach. Verified under -fsanitize=undefined (150+ runs across this
change's lifetime, 0 failures) and -fsanitize=address (100+ runs, 0
real findings; known sandbox ASan-startup flakes excluded, see prior
commits) per CLAUDE.md's sanitizer rule for upper.c/overlay.c changes.
Measured result (make bench, same environment as the original
finding): mem create 257x faster, unlink 330x faster, mkdir 65x
faster, concurrent mixed workload 131x faster -- and, the comparison
that matters, mem now beats raw fs at every one of these (was losing
by 6-22x before). The complexity-class change is confirmed the same
way the O(n^2) was found: mkdir at N=4,000 vs create at N=20,000 now
shows a 7.15x slowdown for a 5x increase in N, matching the O(n log n)
prediction (5.97x) rather than the old O(n^2) one (25x). Full
before/after tables in BENCH.md's new "Resolution" section, which
keeps the original run as the historical record rather than
overwriting it, per this project's own documentation standard.
Recorded the fix in CLAUDE.md's "Known performance characteristics"
(marked RESOLVED, not silently removed) and its architecture map.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
203 lines
8.8 KiB
Markdown
203 lines
8.8 KiB
Markdown
# PackFS
|
|
|
|
A statically linked, in-process virtual file system for C. PackFS treats a
|
|
shipped file tree as an immutable **pack** image and derives writability
|
|
from a `mem` or `dir` upper layer through a copy-on-write overlay. `zip` and
|
|
`tar` are not part of this project and never will be — see "Status" below.
|
|
|
|
The full design rationale — why this shape, what alternatives were rejected,
|
|
and the concurrency, path-containment, and integrity models this
|
|
implementation follows — is specified in [`concept.md`](concept.md), which is
|
|
frozen (see [`CLAUDE.md`](CLAUDE.md)) and is the authoritative source for
|
|
every design decision below. This README documents the implementation that
|
|
followed from it, not a restatement of the rationale.
|
|
|
|
## Status
|
|
|
|
This is an initial, partial implementation of the spec, not a complete one.
|
|
|
|
**Implemented and tested:** mount table, `mem`/`dir`/`pack`/overlay backends
|
|
(including the standalone read-only `pack` backend, `backend_pack_new` —
|
|
every mutating call against it returns `VFS_ERR_PERM`), copy-up, whiteouts,
|
|
compaction, an append journal, path containment, and pack integrity
|
|
validation.
|
|
|
|
**Will never be built, by explicit project decision:** the `zip` (miniz) and
|
|
`tar` (USTAR) import/export backends. `concept.md` Section 11 recommends
|
|
them, but that recommendation is superseded — see `CLAUDE.md`, "Project
|
|
decisions that supersede concept.md." `backend_overlay_new` reads and writes
|
|
this project's own pack format exclusively; there is no zip/tar support and
|
|
none is planned. Do not open an issue or PR adding one.
|
|
|
|
**Deliberately out of scope for v0** (`concept.md` Section 10, not gaps):
|
|
full POSIX semantics, enforced permissions/symlinks/hard links,
|
|
cross-process concurrency (design specified in Section 5.6, unimplemented),
|
|
and content-defined chunking/delta compression.
|
|
|
|
**Tested on Linux only**, in one environment. The `dir`-mount containment
|
|
fallback path for kernels without `openat2` (Section 6.3) is implemented but
|
|
has not been exercised on such a kernel, nor on macOS or Windows.
|
|
|
|
## Building
|
|
|
|
Zero required third-party dependencies — only a C11 compiler, `make`, and
|
|
`pthread` (Section 11.1 of `concept.md` makes this a hard constraint, not a
|
|
preference).
|
|
|
|
```sh
|
|
make # builds libpackfs.a and libpackfs.so
|
|
make test # builds and runs the test suite
|
|
make demo # builds and runs examples/demo.c — see "Try it" below
|
|
make bench # builds and runs bench/bench.c — see "Benchmarks" below
|
|
make install # installs to $PREFIX (default /usr/local)
|
|
```
|
|
|
|
## Try it
|
|
|
|
`examples/demo.c` is a small, runnable, human-readable program — not another
|
|
automated test — that exercises the library end to end and prints what it
|
|
did at each step: a pack-backed overlay (write, read, `mkdir`, `readdir`,
|
|
`stat`, a copy-up-then-whiteout delete, compaction via `vfs_sync`), and a
|
|
sandboxed `dir` mount that demonstrates a `../../../etc/passwd` escape
|
|
attempt being rejected. Run `make demo` twice in a row: the second run's
|
|
first `readdir` shows the first run's files, proving that compaction and
|
|
reload actually persist data through the pack file, not just within one
|
|
process's lifetime.
|
|
|
|
## Benchmarks
|
|
|
|
`bench/bench.c` (`make bench`) measures PackFS against the host filesystem
|
|
across metadata operations (create/read/stat/readdir/unlink/mkdir), large
|
|
sequential I/O, random-access pack reads, and concurrent mixed workloads.
|
|
[`BENCH.md`](BENCH.md) has the full results and honest analysis of both the
|
|
wins and the losses, including the story of a real bug this benchmark
|
|
found: an early run turned up a quantitatively-confirmed O(n²) cost in bulk
|
|
sequential file creation, which `concept.md` Section 5.3 had explicitly
|
|
anticipated and named the trigger condition for closing. That's since been
|
|
fixed (`src/upper.c`'s index is a persistent treap now, not a flat array),
|
|
re-benchmarked, and re-confirmed as an actual complexity-class change, not
|
|
just a faster constant factor — see `BENCH.md`'s "Resolution" section for
|
|
the before/after numbers. Read the whole file before quoting a number from
|
|
it: what `raw fs` vs `raw+fsync` vs `dir` each actually measure is not
|
|
interchangeable, and it explains why.
|
|
|
|
`openat2`/Landlock support (Section 6) is detected automatically at compile
|
|
time via `<sys/syscall.h>`; on kernels or platforms without them, `dir`
|
|
mounts fall back to the weaker, documented residual-risk posture described
|
|
in `concept.md` Section 6.3 rather than failing to build.
|
|
|
|
## Quick example
|
|
|
|
```c
|
|
#include <packfs.h>
|
|
|
|
Vfs *v = vfs_new();
|
|
|
|
/* a plain in-memory writable tree */
|
|
Backend *mem = backend_mem_new();
|
|
vfs_mount(v, "/", mem);
|
|
|
|
int err = 0;
|
|
VfsFile *f = vfs_open(v, "/hello.txt", VFS_O_WRONLY | VFS_O_CREAT, &err);
|
|
vfs_write(f, "hello", 5);
|
|
vfs_close(f);
|
|
|
|
vfs_free(v);
|
|
backend_free(mem);
|
|
```
|
|
|
|
A shipped pack with a writable overlay on top:
|
|
|
|
```c
|
|
Backend *mem = backend_mem_new();
|
|
int err = 0;
|
|
Backend *ov = backend_overlay_new("assets.pack", mem, &err); /* loads assets.pack if it exists */
|
|
vfs_mount(v, "/", ov);
|
|
|
|
/* ... reads served straight from the pack; writes copy-up into mem ... */
|
|
|
|
vfs_sync(v, "/"); /* compacts the overlay into a fresh assets.pack (Section 4.1, 5.3) */
|
|
```
|
|
|
|
A sandboxed host directory:
|
|
|
|
```c
|
|
int derr = 0;
|
|
Backend *dir = backend_dir_new("/var/lib/myapp/data", &derr);
|
|
vfs_mount(v, "/data", dir);
|
|
/* every lookup under /data is contained to that directory (Section 6.2);
|
|
* ".." and symlink escapes are rejected, not merely discouraged. */
|
|
```
|
|
|
|
A shipped pack mounted read-only, no writable layer at all:
|
|
|
|
```c
|
|
int perr = 0;
|
|
Backend *ro = backend_pack_new("assets.pack", &perr);
|
|
vfs_mount(v, "/assets", ro);
|
|
/* vfs_write, vfs_mkdir, vfs_unlink, vfs_rename, vfs_sync against anything
|
|
* under /assets all return VFS_ERR_PERM; there is no upper layer to
|
|
* absorb a write into. */
|
|
```
|
|
|
|
## API
|
|
|
|
The public API is [`include/packfs.h`](include/packfs.h); every function and
|
|
struct is documented there with a pointer to the `concept.md` section that
|
|
specifies its behavior. In outline:
|
|
|
|
- `vfs_new` / `vfs_free` — a `Vfs` owns a mount table, nothing else.
|
|
- `backend_mem_new` / `backend_dir_new` / `backend_pack_new` /
|
|
`backend_overlay_new` — construct a backend; `vfs_mount`/`vfs_unmount`
|
|
attach or detach it at a path prefix. `backend_pack_new` mounts a pack
|
|
read-only, with no writable upper layer at all — every mutating call
|
|
against it returns `VFS_ERR_PERM`; wrap the same pack in
|
|
`backend_overlay_new` instead when writability is wanted.
|
|
- `vfs_open` / `vfs_read` / `vfs_write` / `vfs_close` — file I/O.
|
|
- `vfs_stat` / `vfs_readdir` / `vfs_mkdir` / `vfs_unlink` / `vfs_rename` —
|
|
metadata and namespace operations.
|
|
- `vfs_sync` — compacts an overlay mount into a fresh pack.
|
|
- `vfs_harden_process_with_landlock` — optional, opt-in, process-wide
|
|
Landlock confinement to the process's current `dir` mounts. Deliberately
|
|
**not** applied automatically by `vfs_mount` (Section 6.2 explains why:
|
|
Landlock restrictions are irreversible and process-wide, which would be a
|
|
surprising side effect for an embeddable library to trigger on its own).
|
|
|
|
## Concurrency
|
|
|
|
Single-writer, wait-free-reader (Section 5): readers never take a lock and
|
|
never observe a write in progress; structural changes (create/unlink/rename/
|
|
mkdir/mount/unmount) publish a new immutable snapshot via one atomic pointer
|
|
swap; ordinary content writes to an already-existing file update a per-entry
|
|
cell directly and never touch the snapshot. `mem`-backed buffer growth
|
|
never mutates a buffer address a reader might be reading (Section 5.7):
|
|
growth always allocates a new buffer and publishes it, never reallocates in
|
|
place. `tests/test_concurrency.c` exercises this under concurrent reader and
|
|
writer threads, and the suite is regularly run under ThreadSanitizer and
|
|
AddressSanitizer (see `.github/workflows/ci.yml`).
|
|
|
|
## Security
|
|
|
|
`dir` mounts are capability-scoped (Section 6.2): a mount holds an already-open
|
|
directory file descriptor, not a path string, and every lookup beneath it is
|
|
resolved with `openat2(RESOLVE_BENEATH | RESOLVE_NO_SYMLINKS)` on Linux 5.6+,
|
|
which atomically rejects `..` and symlink escapes in one kernel call. Where
|
|
that syscall is unavailable, containment falls back to per-component
|
|
`O_NOFOLLOW` resolution — weaker, and documented as such (Section 6.3), not
|
|
silently assumed equivalent. Pack files are treated as untrusted input unless
|
|
they came from this process's own compaction: every on-disk offset is
|
|
bounds-checked and the pack's checksum is verified before any of it is
|
|
trusted (Section 7). See `tests/test_dir.c` for a containment regression test
|
|
and `tests/test_pack_overlay.c` for a corrupted-pack rejection test.
|
|
|
|
## Contributing
|
|
|
|
See [`CONTRIBUTING.md`](CONTRIBUTING.md). Read `concept.md` and `CLAUDE.md`
|
|
first — they are the project's actual specification and its enforced
|
|
documentation standard, respectively, and every design decision in the code
|
|
traces back to one of them.
|
|
|
|
## License
|
|
|
|
MIT — see [`LICENSE`](LICENSE).
|