Prompted by "fix literally everything still open" after the previous
session's data-integrity work. Went through each open item in turn:
1. TSan: tried a genuinely different execution environment (a remote
cloud sandbox, via a dedicated agent) rather than re-stating the local
sandbox's limitation. Result: identical block there too --
personality(ADDR_NO_RANDOMIZE) returns EPERM, a trivial pthread
program fails TSan identically, and all 6 PackFS test binaries fail
with the same FATAL: ThreadSanitizer: unexpected memory mapping
signature. This is now confirmed in two independent environments, not
one -- strong evidence it's a real infrastructure restriction, not a
one-off fluke worth chasing further with the tools available here.
2. While investigating the "disk-full mid-write" gap flagged as untested
last session, found two real, previously-unknown bugs by reading the
journal code (not by a test catching them unprompted):
- journal_append_record and everything that called it were void, and
none of the fwrite/fflush/fsync calls inside had their return values
checked. A real write failure (disk full, quota, I/O error) was
silently reported as success to vfs_write/vfs_mkdir/vfs_unlink/
vfs_rename -- directly contradicting Section 4.4's premise that
success means durable.
- Fixing that alone was not enough, confirmed by direct reproduction:
a partial write leaves a torn record in the journal, and
journal_replay correctly stops at the first record it can't fully
read (Section 4.3) -- which means every record appended *after* the
torn one, including ones that themselves wrote perfectly fine later,
became silently unreachable on reopen. Reproduced directly before
fixing: a forced-failed write followed by a genuinely successful one
was unrecoverable. Fixed by rolling the journal file back to its
exact pre-record length on any failed write.
Both closed in src/overlay.c (journal_append_record/_put/_delete/
_mkdir/journal_put_current now return and propagate success/failure;
overlay_write/_mkdir/_unlink/_rename return VFS_ERR_IO on a durability
failure without rolling back the already-applied in-memory change,
the same asymmetry a real write()-then-failed-fsync() has). Covered
permanently by the new tests/test_journal_failure.c, which forces a
real failure via RLIMIT_FSIZE + ignored SIGXFSZ, not a mock.
Also fixed in the same pass, found by inspection while touching this
code: journal_put_current used to pass a NULL buffer into a memcpy of
a nonzero size when malloc(size) failed (an OOM-triggered NULL-pointer
dereference) -- closed with an explicit allocation-failure check.
Not test-triggered (forcing malloc() failure portably isn't practical
here); verified by code inspection instead, stated as such rather than
claimed as tested.
3. The remaining "journal-truncation-specific crash window" gap from last
session was investigated, not silently dropped: reliably targeting
that narrow a window would need real concurrency (a second writer
thread racing the kill) for benefit the existing compaction-crash test
already gets probabilistically -- a poor trade, so left as a stated,
deliberate non-goal (CLAUDE.md) rather than built.
4. Cross-process contention is NOT addressed here and should not be read
as an oversight: it is concept.md's own explicit, permanent "not
implemented in v0" scope boundary (a specified-but-unbuilt LMDB-style
reader-table design), not a bug -- building it would be a large,
unrequested feature addition outside this session's actual scope.
Verified: clean make all + make test (all 8 binaries), make bench and
make demo still build and the demo runs correctly end to end, and a full
ASan/UBSan sweep of all 8 binaries with zero real findings (some retries
needed for the already-documented DEADLYSIGNAL flake, which
test_crash_consistency hits more often than other tests simply because it
forks 60+ subprocesses per run -- noted in CONTRIBUTING.md so this isn't
mistaken for a regression later).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
PackFS
A statically linked, in-process virtual file system for C. PackFS treats a
shipped file tree as an immutable pack image and derives writability
from a mem or dir upper layer through a copy-on-write overlay. zip and
tar are not part of this project and never will be — see "Status" below.
The full design rationale — why this shape, what alternatives were rejected,
and the concurrency, path-containment, and integrity models this
implementation follows — is specified in concept.md, which is
frozen (see CLAUDE.md) and is the authoritative source for
every design decision below. This README documents the implementation that
followed from it, not a restatement of the rationale.
Status
This is an initial, partial implementation of the spec, not a complete one.
Implemented and tested: mount table, mem/dir/pack/overlay backends
(including the standalone read-only pack backend, backend_pack_new —
every mutating call against it returns VFS_ERR_PERM), copy-up, whiteouts,
compaction, an append journal, path containment, and pack integrity
validation.
Will never be built, by explicit project decision: the zip (miniz) and
tar (USTAR) import/export backends. concept.md Section 11 recommends
them, but that recommendation is superseded — see CLAUDE.md, "Project
decisions that supersede concept.md." backend_overlay_new reads and writes
this project's own pack format exclusively; there is no zip/tar support and
none is planned. Do not open an issue or PR adding one.
Deliberately out of scope for v0 (concept.md Section 10, not gaps):
full POSIX semantics, enforced permissions/symlinks/hard links,
cross-process concurrency (design specified in Section 5.6, unimplemented),
and content-defined chunking/delta compression.
Tested on Linux only, in one environment. The dir-mount containment
fallback path for kernels without openat2 (Section 6.3) is implemented but
has not been exercised on such a kernel, nor on macOS or Windows.
Building
Zero required third-party dependencies — only a C11 compiler, make, and
pthread (Section 11.1 of concept.md makes this a hard constraint, not a
preference).
make # builds libpackfs.a and libpackfs.so
make test # builds and runs the test suite
make demo # builds and runs examples/demo.c — see "Try it" below
make bench # builds and runs bench/bench.c — see "Benchmarks" below
make install # installs to $PREFIX (default /usr/local), including a pkg-config file
make install also generates and installs packfs.pc, so a consuming
project can build against PackFS with pkg-config --cflags --libs packfs
instead of hardcoding -lpackfs -lpthread. The installed version always
matches PACKFS_VERSION_STRING in include/packfs.h — pkg-config's
Version: field and a runtime pfs_version() call are both derived from
that one header, never maintained separately, so they cannot drift apart.
Try it
examples/demo.c is a small, runnable, human-readable program — not another
automated test — that exercises the library end to end and prints what it
did at each step: a pack-backed overlay (write, read, mkdir, readdir,
stat, a copy-up-then-whiteout delete, compaction via vfs_sync), and a
sandboxed dir mount that demonstrates a ../../../etc/passwd escape
attempt being rejected. Run make demo twice in a row: the second run's
first readdir shows the first run's files, proving that compaction and
reload actually persist data through the pack file, not just within one
process's lifetime.
Benchmarks
bench/bench.c (make bench) measures PackFS against the host filesystem
across metadata operations (create/read/stat/readdir/unlink/mkdir), large
sequential I/O, random-access pack reads, mount-table scaling, and
concurrent mixed workloads. BENCH.md has the full results and
honest analysis of both the wins and the losses, including the story of
three real O(n²) findings this project's own benchmarking turned up — not
just the wins. Two are fixed: bulk sequential file creation (src/upper.c's
index is a persistent treap now, not a flat array — see "Resolution") and
pack_write's compaction-time duplicate-content elimination, which also had
a latent correctness bug now closed alongside it (see "Resolution #2"). One
is confirmed and deliberately not fixed: the mount table scales O(n²) in
mount count, the same way the file index used to, but mount counts are
bounded by a program's own source code rather than workload-driven, so it
isn't worth the added complexity — see "Finding: mount table scaling" for
the reasoning and the numbers behind that call. Read the whole file before
quoting a number from it: what raw fs vs raw+fsync vs dir each
actually measure is not interchangeable, and it explains why.
openat2/Landlock support (Section 6) is detected automatically at compile
time via <sys/syscall.h>; on kernels or platforms without them, dir
mounts fall back to the weaker, documented residual-risk posture described
in concept.md Section 6.3 rather than failing to build.
Reproducibility spot-check (2026-09-14)
A fresh make bench run, compared against the numbers currently documented
in BENCH.md's "After" table, on the same environment described there. Not
a replacement for BENCH.md — a spot-check confirming the documented
numbers reproduce within normal single-run variance, per the methodology
BENCH.md itself states ("illustrative of shape... not precise absolute
figures"). Every column from the raw make bench output is kept for both
runs — time, throughput, and MB/s where the category reports one — not
just a time delta, so every category (including the pack rows:
compaction and random-access read) can be checked in full, not summarized
away.
| Category | Backend | Doc Time | Doc Throughput | Doc MB/s | New Time | New Throughput | New MB/s | Δ Time |
|---|---|---|---|---|---|---|---|---|
| create 20,000 files | mem | 0.0400s | 499,809 ops/s | 61.0 | 0.0373s | 536,419 ops/s | 65.5 | -6.8% |
| create 20,000 files | raw fs | 0.9943s | 20,114 ops/s | 2.5 | 0.9883s | 20,237 ops/s | 2.5 | -0.6% |
| create 20,000 files | raw+fsync | 116.2147s | 172 ops/s | 0.0 | 116.3080s | 172 ops/s | 0.0 | +0.1% |
| read 20,000 files | mem | 0.0099s | 2,011,901 ops/s | 245.6 | 0.0116s | 1,723,405 ops/s | 210.4 | +17.2% |
| read 20,000 files | raw fs | 0.1606s | 124,543 ops/s | 15.2 | 0.2040s | 98,030 ops/s | 12.0 | +27.0% |
| stat 20,000 files | mem | 0.0077s | 2,598,241 ops/s | — | 0.0079s | 2,519,089 ops/s | — | +2.6% |
| stat 20,000 files | raw fs | 0.0699s | 286,222 ops/s | — | 0.0954s | 209,541 ops/s | — | +36.5% |
| readdir (20,000 entries) | mem | 0.0041s | 4,845,826 ops/s | — | 0.0043s | 4,675,035 ops/s | — | +4.9% |
| readdir (20,000 entries) | raw fs | 0.0069s | 2,894,815 ops/s | — | 0.0071s | 2,808,607 ops/s | — | +2.9% |
| create 20,000 files | dir | 1.1418s | 17,517 ops/s | 2.1 | 1.7006s | 11,760 ops/s | 1.4 | +48.9% |
| read 20,000 files | dir | 0.1530s | 130,705 ops/s | 16.0 | 0.2238s | 89,371 ops/s | 10.9 | +46.3% |
| stat 20,000 files | dir | 0.1379s | 145,056 ops/s | — | 0.1984s | 100,826 ops/s | — | +43.9% |
| readdir (20,000 entries) | dir | 0.0037s | 5,465,162 ops/s | — | 0.0042s | 4,799,105 ops/s | — | +13.5% |
| unlink 20,000 files | mem | 0.0179s | 1,117,742 ops/s | — | 0.0183s | 1,094,715 ops/s | — | +2.2% |
| unlink 20,000 files | dir | 0.4911s | 40,724 ops/s | — | 0.7500s | 26,667 ops/s | — | +52.7% |
| unlink 20,000 files | raw fs | 0.4746s | 42,141 ops/s | — | 0.6404s | 31,228 ops/s | — | +34.9% |
| mkdir 4,000 dirs | mem | 0.0056s | 719,696 ops/s | — | 0.0051s | 791,918 ops/s | — | -8.9% |
| rmdir 4,000 dirs | mem | 0.0040s | 993,724 ops/s | — | 0.0435s | 91,873 ops/s | — | +987.5% |
| mkdir 4,000 dirs | dir | 0.1710s | 23,386 ops/s | — | 0.2123s | 18,838 ops/s | — | +24.2% |
| rmdir 4,000 dirs | dir | 0.1257s | 31,827 ops/s | — | 0.1372s | 29,163 ops/s | — | +9.1% |
| mkdir 4,000 dirs | raw fs | 0.1374s | 29,122 ops/s | — | 0.1837s | 21,772 ops/s | — | +33.7% |
| rmdir 4,000 dirs | raw fs | 0.0951s | 42,043 ops/s | — | 0.1305s | 30,640 ops/s | — | +37.2% |
| write 1MB | mem | 0.0001s | 7,058 ops/s | 7,058.1 | 0.0002s | 6,313 ops/s | 6,312.7 | +100.0% |
| read 1MB | mem | 0.0000s | 50,051 ops/s | 50,050.9 | 0.0000s | 48,616 ops/s | 48,616.4 | +0.0% |
| write 16MB | mem | 0.0110s | 91 ops/s | 1,458.2 | 0.0407s | 25 ops/s | 393.3 | +270.0% |
| read 16MB | mem | 0.0009s | 1,058 ops/s | 16,920.1 | 0.0010s | 1,021 ops/s | 16,338.0 | +11.1% |
| write 64MB | mem | 0.0586s | 17 ops/s | 1,091.3 | 0.0599s | 17 ops/s | 1,068.3 | +2.2% |
| read 64MB | mem | 0.0032s | 313 ops/s | 20,056.5 | 0.0300s | 33 ops/s | 2,131.0 | +837.5% |
| write 1MB | raw fs | 0.0004s | 2,759 ops/s | 2,759.4 | 0.0004s | 2,474 ops/s | 2,474.0 | +0.0% |
| read 1MB | raw fs | 0.0001s | 16,812 ops/s | 16,812.4 | 0.0001s | 16,756 ops/s | 16,756.0 | +0.0% |
| write 16MB | raw fs | 0.0044s | 229 ops/s | 3,656.4 | 0.0044s | 228 ops/s | 3,645.0 | +0.0% |
| read 16MB | raw fs | 0.0010s | 989 ops/s | 15,827.9 | 0.0010s | 965 ops/s | 15,443.5 | +0.0% |
| write 64MB | raw fs | 0.0186s | 54 ops/s | 3,449.7 | 0.0184s | 54 ops/s | 3,478.9 | -1.1% |
| read 64MB | raw fs | 0.0047s | 213 ops/s | 13,633.1 | 0.0049s | 203 ops/s | 13,003.0 | +4.3% |
| write 1MB | raw+fsync | 0.0180s | 56 ops/s | 55.7 | 0.0162s | 62 ops/s | 61.7 | -10.0% |
| read 1MB | raw+fsync | 0.0001s | 9,059 ops/s | 9,058.8 | 0.0001s | 12,025 ops/s | 12,025.1 | +0.0% |
| write 16MB | raw+fsync | 0.0231s | 43 ops/s | 691.7 | 0.0247s | 41 ops/s | 648.1 | +6.9% |
| read 16MB | raw+fsync | 0.0012s | 855 ops/s | 13,672.2 | 0.0014s | 726 ops/s | 11,610.9 | +16.7% |
| write 64MB | raw+fsync | 0.0803s | 12 ops/s | 797.2 | 0.1001s | 10 ops/s | 639.5 | +24.7% |
| read 64MB | raw+fsync | 0.0048s | 209 ops/s | 13,393.5 | 0.0200s | 50 ops/s | 3,193.8 | +316.7% |
| compact 20,000 entries to pack | pack | 0.0267s | 747,839 ops/s | 91.3 | 0.0250s | 800,791 ops/s | 97.8 | -6.4% |
| random-read 20,000 entries | pack (mmap'd) | 0.0082s | 2,435,930 ops/s | 297.4 | 0.0083s | 2,397,682 ops/s | 292.7 | +1.2% |
| random-read 20,000 entries | raw fs | 0.1629s | 122,749 ops/s | 15.0 | 0.1691s | 118,290 ops/s | 14.4 | +3.8% |
| concurrent create+read+unlink (8×4,000×3) | mem | 0.2750s | 349,127 ops/s | — | 0.2981s | 322,041 ops/s | — | +8.4% |
| concurrent create+read+unlink (8×4,000×3) | raw fs | 5.6498s | 16,992 ops/s | — | 6.3837s | 15,038 ops/s | — | +13.0% |
| mount 500 backends | vfs | 0.0054s | 92,833 ops/s | — | 0.0054s | 91,871 ops/s | — | +0.0% |
| resolve, 500 mounts | vfs | 0.0020s | 246,064 ops/s | — | 0.0019s | 262,671 ops/s | — | -5.0% |
| unmount 500 backends | vfs | 0.0046s | 109,207 ops/s | — | 0.0045s | 110,348 ops/s | — | -2.2% |
| mount 2,000 backends | vfs | 0.0831s | 24,081 ops/s | — | 0.0836s | 23,926 ops/s | — | +0.6% |
| resolve, 2,000 mounts | vfs | 0.0296s | 67,458 ops/s | — | 0.0293s | 68,318 ops/s | — | -1.0% |
| unmount 2,000 backends | vfs | 0.0758s | 26,397 ops/s | — | 0.0802s | 24,932 ops/s | — | +5.8% |
| mount 8,000 backends | vfs | 1.4739s | 5,428 ops/s | — | 1.4905s | 5,367 ops/s | — | +1.1% |
| resolve, 8,000 mounts | vfs | 0.4938s | 16,200 ops/s | — | 0.4616s | 17,332 ops/s | — | -6.5% |
| unmount 8,000 backends | vfs | 1.5172s | 5,273 ops/s | — | 1.5176s | 5,271 ops/s | — | +0.0% |
The pack-backend rows are the most stable in the whole table (compaction
-6.4%, random-access read +1.2%/+3.8%), consistent with BENCH.md's point
that pack random access is a flat mmap'd table lookup rather than
anything sensitive to scheduling jitter the way syscall-heavy dir/raw-fs
metadata operations are.
Most rows sit within ±15% of the documented figures — single-run jitter,
not a regression, exactly as BENCH.md's stated methodology warns to
expect. This run was noisier than the previous spot-check, though, with
several rows well outside that band, worth naming rather than averaging
away: every dir-backend metadata row (create/read/stat/unlink/
mkdir) is elevated 24–53% together, pointing at host-side I/O contention
during this run rather than anything PackFS-specific (the mem and pack
rows, touching no host filesystem, mostly did not move); and two mem
rows are outliers even by that standard — rmdir 4,000 dirs (+987.5%,
0.0040s → 0.0435s) and read 64MB (+837.5%, 0.0032s → 0.0300s). Both are
large relative jumps on a small absolute base (tens of milliseconds),
exactly where a single scheduling stall has the most disproportionate
effect on a percentage, and neither is corroborated by a neighboring
row moving the same way (mkdir and unlink on mem, right next to
rmdir, both stayed flat or improved; read 16MB/write 64MB on mem,
right next to read 64MB, both stayed within a few percent) — the
single-run, non-statistical methodology BENCH.md states up front means
this can't be distinguished from noise without repeated runs, so it is
reported here rather than either hidden or asserted as a regression.
Total wall-clock: 4m10.1s this run vs 5m23.9s documented — faster overall
despite the several elevated dir/raw-fs rows above, because the dominant
cost throughout every run of this benchmark is the single raw+fsync
20,000-file create test (~116s here too, unaffected — see BENCH.md's
"Environment" section for why that test in particular is unrelated to
PackFS's own performance).
Quick example
#include <packfs.h>
Vfs *v = vfs_new();
/* a plain in-memory writable tree */
Backend *mem = backend_mem_new();
vfs_mount(v, "/", mem);
int err = 0;
VfsFile *f = vfs_open(v, "/hello.txt", VFS_O_WRONLY | VFS_O_CREAT, &err);
vfs_write(f, "hello", 5);
vfs_close(f);
vfs_free(v);
backend_free(mem);
A shipped pack with a writable overlay on top:
Backend *mem = backend_mem_new();
int err = 0;
Backend *ov = backend_overlay_new("assets.pack", mem, &err); /* loads assets.pack if it exists */
vfs_mount(v, "/", ov);
/* ... reads served straight from the pack; writes copy-up into mem ... */
vfs_sync(v, "/"); /* compacts the overlay into a fresh assets.pack (Section 4.1, 5.3) */
A sandboxed host directory:
int derr = 0;
Backend *dir = backend_dir_new("/var/lib/myapp/data", &derr);
vfs_mount(v, "/data", dir);
/* every lookup under /data is contained to that directory (Section 6.2);
* ".." and symlink escapes are rejected, not merely discouraged. */
A shipped pack mounted read-only, no writable layer at all:
int perr = 0;
Backend *ro = backend_pack_new("assets.pack", &perr);
vfs_mount(v, "/assets", ro);
/* vfs_write, vfs_mkdir, vfs_unlink, vfs_rename, vfs_sync against anything
* under /assets all return VFS_ERR_PERM; there is no upper layer to
* absorb a write into. */
API
The public API is include/packfs.h; every function and
struct is documented there with a pointer to the concept.md section that
specifies its behavior. In outline:
vfs_new/vfs_free— aVfsowns a mount table, nothing else.backend_mem_new/backend_dir_new/backend_pack_new/backend_overlay_new— construct a backend;vfs_mount/vfs_unmountattach or detach it at a path prefix.backend_pack_newmounts a pack read-only, with no writable upper layer at all — every mutating call against it returnsVFS_ERR_PERM; wrap the same pack inbackend_overlay_newinstead when writability is wanted.vfs_open/vfs_read/vfs_write/vfs_close— file I/O.vfs_stat/vfs_readdir/vfs_mkdir/vfs_unlink/vfs_rename— metadata and namespace operations.vfs_sync— compacts an overlay mount into a fresh pack.vfs_harden_process_with_landlock— optional, opt-in, process-wide Landlock confinement to the process's currentdirmounts. Deliberately not applied automatically byvfs_mount(Section 6.2 explains why: Landlock restrictions are irreversible and process-wide, which would be a surprising side effect for an embeddable library to trigger on its own).
Concurrency
Single-writer, wait-free-reader (Section 5): readers never take a lock and
never observe a write in progress; structural changes (create/unlink/rename/
mkdir/mount/unmount) publish a new immutable snapshot via one atomic pointer
swap; ordinary content writes to an already-existing file update a per-entry
cell directly and never touch the snapshot. mem-backed buffer growth
never mutates a buffer address a reader might be reading (Section 5.7):
growth always allocates a new buffer and publishes it, never reallocates in
place. tests/test_concurrency.c exercises this under concurrent reader and
writer threads. The suite is regularly run under AddressSanitizer/
UndefinedBehaviorSanitizer, both locally and in CI (see
.gitea/workflows/ci.yml), and this is not aspirational — ASan caught a
real heap-use-after-free in the snapshot-reclamation logic during
development (see the reclaim_gate note in src/internal.h).
ThreadSanitizer is configured the same way but, as of this writing, has
not actually completed a run in either environment this project has been
built and tested in so far — the local development sandbox and this
project's own CI runner both block the personality(ADDR_NO_RANDOMIZE)
syscall TSan needs to start (see CONTRIBUTING.md for the confirming
tests in each case). CI's TSan step is written to not fail the build over
that specific, known-benign failure, which means a green CI run is not
evidence TSan actually executed — stated plainly here rather than left to
be assumed from CI showing green.
Security
dir mounts are capability-scoped (Section 6.2): a mount holds an already-open
directory file descriptor, not a path string, and every lookup beneath it is
resolved with openat2(RESOLVE_BENEATH | RESOLVE_NO_SYMLINKS) on Linux 5.6+,
which atomically rejects .. and symlink escapes in one kernel call. Where
that syscall is unavailable, containment falls back to per-component
O_NOFOLLOW resolution — weaker, and documented as such (Section 6.3), not
silently assumed equivalent. Pack files are treated as untrusted input unless
they came from this process's own compaction: every on-disk offset is
bounds-checked and the pack's checksum is verified before any of it is
trusted (Section 7). See tests/test_dir.c for a containment regression test
and tests/test_pack_overlay.c for a corrupted-pack rejection test. See
SECURITY.md for exactly what is and isn't claimed as a
security boundary, and how to report a vulnerability.
Contributing
See CONTRIBUTING.md. Read concept.md and CLAUDE.md
first — they are the project's actual specification and its enforced
documentation standard, respectively, and every design decision in the code
traces back to one of them.
License
MIT — see LICENSE.