retoorandClaude Sonnet 5 153d44ee3c
CI / build-and-test (push) Failing after 33s
Fix silent journal-write-failure data loss; confirm TSan blocked twice over
Prompted by "fix literally everything still open" after the previous
session's data-integrity work. Went through each open item in turn:

1. TSan: tried a genuinely different execution environment (a remote
   cloud sandbox, via a dedicated agent) rather than re-stating the local
   sandbox's limitation. Result: identical block there too --
   personality(ADDR_NO_RANDOMIZE) returns EPERM, a trivial pthread
   program fails TSan identically, and all 6 PackFS test binaries fail
   with the same FATAL: ThreadSanitizer: unexpected memory mapping
   signature. This is now confirmed in two independent environments, not
   one -- strong evidence it's a real infrastructure restriction, not a
   one-off fluke worth chasing further with the tools available here.

2. While investigating the "disk-full mid-write" gap flagged as untested
   last session, found two real, previously-unknown bugs by reading the
   journal code (not by a test catching them unprompted):

   - journal_append_record and everything that called it were void, and
     none of the fwrite/fflush/fsync calls inside had their return values
     checked. A real write failure (disk full, quota, I/O error) was
     silently reported as success to vfs_write/vfs_mkdir/vfs_unlink/
     vfs_rename -- directly contradicting Section 4.4's premise that
     success means durable.

   - Fixing that alone was not enough, confirmed by direct reproduction:
     a partial write leaves a torn record in the journal, and
     journal_replay correctly stops at the first record it can't fully
     read (Section 4.3) -- which means every record appended *after* the
     torn one, including ones that themselves wrote perfectly fine later,
     became silently unreachable on reopen. Reproduced directly before
     fixing: a forced-failed write followed by a genuinely successful one
     was unrecoverable. Fixed by rolling the journal file back to its
     exact pre-record length on any failed write.

   Both closed in src/overlay.c (journal_append_record/_put/_delete/
   _mkdir/journal_put_current now return and propagate success/failure;
   overlay_write/_mkdir/_unlink/_rename return VFS_ERR_IO on a durability
   failure without rolling back the already-applied in-memory change,
   the same asymmetry a real write()-then-failed-fsync() has). Covered
   permanently by the new tests/test_journal_failure.c, which forces a
   real failure via RLIMIT_FSIZE + ignored SIGXFSZ, not a mock.

   Also fixed in the same pass, found by inspection while touching this
   code: journal_put_current used to pass a NULL buffer into a memcpy of
   a nonzero size when malloc(size) failed (an OOM-triggered NULL-pointer
   dereference) -- closed with an explicit allocation-failure check.
   Not test-triggered (forcing malloc() failure portably isn't practical
   here); verified by code inspection instead, stated as such rather than
   claimed as tested.

3. The remaining "journal-truncation-specific crash window" gap from last
   session was investigated, not silently dropped: reliably targeting
   that narrow a window would need real concurrency (a second writer
   thread racing the kill) for benefit the existing compaction-crash test
   already gets probabilistically -- a poor trade, so left as a stated,
   deliberate non-goal (CLAUDE.md) rather than built.

4. Cross-process contention is NOT addressed here and should not be read
   as an oversight: it is concept.md's own explicit, permanent "not
   implemented in v0" scope boundary (a specified-but-unbuilt LMDB-style
   reader-table design), not a bug -- building it would be a large,
   unrequested feature addition outside this session's actual scope.

Verified: clean make all + make test (all 8 binaries), make bench and
make demo still build and the demo runs correctly end to end, and a full
ASan/UBSan sweep of all 8 binaries with zero real findings (some retries
needed for the already-documented DEADLYSIGNAL flake, which
test_crash_consistency hits more often than other tests simply because it
forks 60+ subprocesses per run -- noted in CONTRIBUTING.md so this isn't
mistaken for a regression later).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
2026-09-14 19:40:23 +00:00

PackFS

A statically linked, in-process virtual file system for C. PackFS treats a shipped file tree as an immutable pack image and derives writability from a mem or dir upper layer through a copy-on-write overlay. zip and tar are not part of this project and never will be — see "Status" below.

The full design rationale — why this shape, what alternatives were rejected, and the concurrency, path-containment, and integrity models this implementation follows — is specified in concept.md, which is frozen (see CLAUDE.md) and is the authoritative source for every design decision below. This README documents the implementation that followed from it, not a restatement of the rationale.

Status

This is an initial, partial implementation of the spec, not a complete one.

Implemented and tested: mount table, mem/dir/pack/overlay backends (including the standalone read-only pack backend, backend_pack_new — every mutating call against it returns VFS_ERR_PERM), copy-up, whiteouts, compaction, an append journal, path containment, and pack integrity validation.

Will never be built, by explicit project decision: the zip (miniz) and tar (USTAR) import/export backends. concept.md Section 11 recommends them, but that recommendation is superseded — see CLAUDE.md, "Project decisions that supersede concept.md." backend_overlay_new reads and writes this project's own pack format exclusively; there is no zip/tar support and none is planned. Do not open an issue or PR adding one.

Deliberately out of scope for v0 (concept.md Section 10, not gaps): full POSIX semantics, enforced permissions/symlinks/hard links, cross-process concurrency (design specified in Section 5.6, unimplemented), and content-defined chunking/delta compression.

Tested on Linux only, in one environment. The dir-mount containment fallback path for kernels without openat2 (Section 6.3) is implemented but has not been exercised on such a kernel, nor on macOS or Windows.

Building

Zero required third-party dependencies — only a C11 compiler, make, and pthread (Section 11.1 of concept.md makes this a hard constraint, not a preference).

make            # builds libpackfs.a and libpackfs.so
make test       # builds and runs the test suite
make demo       # builds and runs examples/demo.c — see "Try it" below
make bench      # builds and runs bench/bench.c — see "Benchmarks" below
make install    # installs to $PREFIX (default /usr/local), including a pkg-config file

make install also generates and installs packfs.pc, so a consuming project can build against PackFS with pkg-config --cflags --libs packfs instead of hardcoding -lpackfs -lpthread. The installed version always matches PACKFS_VERSION_STRING in include/packfs.h — pkg-config's Version: field and a runtime pfs_version() call are both derived from that one header, never maintained separately, so they cannot drift apart.

Try it

examples/demo.c is a small, runnable, human-readable program — not another automated test — that exercises the library end to end and prints what it did at each step: a pack-backed overlay (write, read, mkdir, readdir, stat, a copy-up-then-whiteout delete, compaction via vfs_sync), and a sandboxed dir mount that demonstrates a ../../../etc/passwd escape attempt being rejected. Run make demo twice in a row: the second run's first readdir shows the first run's files, proving that compaction and reload actually persist data through the pack file, not just within one process's lifetime.

Benchmarks

bench/bench.c (make bench) measures PackFS against the host filesystem across metadata operations (create/read/stat/readdir/unlink/mkdir), large sequential I/O, random-access pack reads, mount-table scaling, and concurrent mixed workloads. BENCH.md has the full results and honest analysis of both the wins and the losses, including the story of three real O(n²) findings this project's own benchmarking turned up — not just the wins. Two are fixed: bulk sequential file creation (src/upper.c's index is a persistent treap now, not a flat array — see "Resolution") and pack_write's compaction-time duplicate-content elimination, which also had a latent correctness bug now closed alongside it (see "Resolution #2"). One is confirmed and deliberately not fixed: the mount table scales O(n²) in mount count, the same way the file index used to, but mount counts are bounded by a program's own source code rather than workload-driven, so it isn't worth the added complexity — see "Finding: mount table scaling" for the reasoning and the numbers behind that call. Read the whole file before quoting a number from it: what raw fs vs raw+fsync vs dir each actually measure is not interchangeable, and it explains why.

openat2/Landlock support (Section 6) is detected automatically at compile time via <sys/syscall.h>; on kernels or platforms without them, dir mounts fall back to the weaker, documented residual-risk posture described in concept.md Section 6.3 rather than failing to build.

Reproducibility spot-check (2026-09-14)

A fresh make bench run, compared against the numbers currently documented in BENCH.md's "After" table, on the same environment described there. Not a replacement for BENCH.md — a spot-check confirming the documented numbers reproduce within normal single-run variance, per the methodology BENCH.md itself states ("illustrative of shape... not precise absolute figures"). Every column from the raw make bench output is kept for both runs — time, throughput, and MB/s where the category reports one — not just a time delta, so every category (including the pack rows: compaction and random-access read) can be checked in full, not summarized away.

Category Backend Doc Time Doc Throughput Doc MB/s New Time New Throughput New MB/s Δ Time
create 20,000 files mem 0.0400s 499,809 ops/s 61.0 0.0373s 536,419 ops/s 65.5 -6.8%
create 20,000 files raw fs 0.9943s 20,114 ops/s 2.5 0.9883s 20,237 ops/s 2.5 -0.6%
create 20,000 files raw+fsync 116.2147s 172 ops/s 0.0 116.3080s 172 ops/s 0.0 +0.1%
read 20,000 files mem 0.0099s 2,011,901 ops/s 245.6 0.0116s 1,723,405 ops/s 210.4 +17.2%
read 20,000 files raw fs 0.1606s 124,543 ops/s 15.2 0.2040s 98,030 ops/s 12.0 +27.0%
stat 20,000 files mem 0.0077s 2,598,241 ops/s — 0.0079s 2,519,089 ops/s — +2.6%
stat 20,000 files raw fs 0.0699s 286,222 ops/s — 0.0954s 209,541 ops/s — +36.5%
readdir (20,000 entries) mem 0.0041s 4,845,826 ops/s — 0.0043s 4,675,035 ops/s — +4.9%
readdir (20,000 entries) raw fs 0.0069s 2,894,815 ops/s — 0.0071s 2,808,607 ops/s — +2.9%
create 20,000 files dir 1.1418s 17,517 ops/s 2.1 1.7006s 11,760 ops/s 1.4 +48.9%
read 20,000 files dir 0.1530s 130,705 ops/s 16.0 0.2238s 89,371 ops/s 10.9 +46.3%
stat 20,000 files dir 0.1379s 145,056 ops/s — 0.1984s 100,826 ops/s — +43.9%
readdir (20,000 entries) dir 0.0037s 5,465,162 ops/s — 0.0042s 4,799,105 ops/s — +13.5%
unlink 20,000 files mem 0.0179s 1,117,742 ops/s — 0.0183s 1,094,715 ops/s — +2.2%
unlink 20,000 files dir 0.4911s 40,724 ops/s — 0.7500s 26,667 ops/s — +52.7%
unlink 20,000 files raw fs 0.4746s 42,141 ops/s — 0.6404s 31,228 ops/s — +34.9%
mkdir 4,000 dirs mem 0.0056s 719,696 ops/s — 0.0051s 791,918 ops/s — -8.9%
rmdir 4,000 dirs mem 0.0040s 993,724 ops/s — 0.0435s 91,873 ops/s — +987.5%
mkdir 4,000 dirs dir 0.1710s 23,386 ops/s — 0.2123s 18,838 ops/s — +24.2%
rmdir 4,000 dirs dir 0.1257s 31,827 ops/s — 0.1372s 29,163 ops/s — +9.1%
mkdir 4,000 dirs raw fs 0.1374s 29,122 ops/s — 0.1837s 21,772 ops/s — +33.7%
rmdir 4,000 dirs raw fs 0.0951s 42,043 ops/s — 0.1305s 30,640 ops/s — +37.2%
write 1MB mem 0.0001s 7,058 ops/s 7,058.1 0.0002s 6,313 ops/s 6,312.7 +100.0%
read 1MB mem 0.0000s 50,051 ops/s 50,050.9 0.0000s 48,616 ops/s 48,616.4 +0.0%
write 16MB mem 0.0110s 91 ops/s 1,458.2 0.0407s 25 ops/s 393.3 +270.0%
read 16MB mem 0.0009s 1,058 ops/s 16,920.1 0.0010s 1,021 ops/s 16,338.0 +11.1%
write 64MB mem 0.0586s 17 ops/s 1,091.3 0.0599s 17 ops/s 1,068.3 +2.2%
read 64MB mem 0.0032s 313 ops/s 20,056.5 0.0300s 33 ops/s 2,131.0 +837.5%
write 1MB raw fs 0.0004s 2,759 ops/s 2,759.4 0.0004s 2,474 ops/s 2,474.0 +0.0%
read 1MB raw fs 0.0001s 16,812 ops/s 16,812.4 0.0001s 16,756 ops/s 16,756.0 +0.0%
write 16MB raw fs 0.0044s 229 ops/s 3,656.4 0.0044s 228 ops/s 3,645.0 +0.0%
read 16MB raw fs 0.0010s 989 ops/s 15,827.9 0.0010s 965 ops/s 15,443.5 +0.0%
write 64MB raw fs 0.0186s 54 ops/s 3,449.7 0.0184s 54 ops/s 3,478.9 -1.1%
read 64MB raw fs 0.0047s 213 ops/s 13,633.1 0.0049s 203 ops/s 13,003.0 +4.3%
write 1MB raw+fsync 0.0180s 56 ops/s 55.7 0.0162s 62 ops/s 61.7 -10.0%
read 1MB raw+fsync 0.0001s 9,059 ops/s 9,058.8 0.0001s 12,025 ops/s 12,025.1 +0.0%
write 16MB raw+fsync 0.0231s 43 ops/s 691.7 0.0247s 41 ops/s 648.1 +6.9%
read 16MB raw+fsync 0.0012s 855 ops/s 13,672.2 0.0014s 726 ops/s 11,610.9 +16.7%
write 64MB raw+fsync 0.0803s 12 ops/s 797.2 0.1001s 10 ops/s 639.5 +24.7%
read 64MB raw+fsync 0.0048s 209 ops/s 13,393.5 0.0200s 50 ops/s 3,193.8 +316.7%
compact 20,000 entries to pack pack 0.0267s 747,839 ops/s 91.3 0.0250s 800,791 ops/s 97.8 -6.4%
random-read 20,000 entries pack (mmap'd) 0.0082s 2,435,930 ops/s 297.4 0.0083s 2,397,682 ops/s 292.7 +1.2%
random-read 20,000 entries raw fs 0.1629s 122,749 ops/s 15.0 0.1691s 118,290 ops/s 14.4 +3.8%
concurrent create+read+unlink (8×4,000×3) mem 0.2750s 349,127 ops/s — 0.2981s 322,041 ops/s — +8.4%
concurrent create+read+unlink (8×4,000×3) raw fs 5.6498s 16,992 ops/s — 6.3837s 15,038 ops/s — +13.0%
mount 500 backends vfs 0.0054s 92,833 ops/s — 0.0054s 91,871 ops/s — +0.0%
resolve, 500 mounts vfs 0.0020s 246,064 ops/s — 0.0019s 262,671 ops/s — -5.0%
unmount 500 backends vfs 0.0046s 109,207 ops/s — 0.0045s 110,348 ops/s — -2.2%
mount 2,000 backends vfs 0.0831s 24,081 ops/s — 0.0836s 23,926 ops/s — +0.6%
resolve, 2,000 mounts vfs 0.0296s 67,458 ops/s — 0.0293s 68,318 ops/s — -1.0%
unmount 2,000 backends vfs 0.0758s 26,397 ops/s — 0.0802s 24,932 ops/s — +5.8%
mount 8,000 backends vfs 1.4739s 5,428 ops/s — 1.4905s 5,367 ops/s — +1.1%
resolve, 8,000 mounts vfs 0.4938s 16,200 ops/s — 0.4616s 17,332 ops/s — -6.5%
unmount 8,000 backends vfs 1.5172s 5,273 ops/s — 1.5176s 5,271 ops/s — +0.0%

The pack-backend rows are the most stable in the whole table (compaction -6.4%, random-access read +1.2%/+3.8%), consistent with BENCH.md's point that pack random access is a flat mmap'd table lookup rather than anything sensitive to scheduling jitter the way syscall-heavy dir/raw-fs metadata operations are.

Most rows sit within ±15% of the documented figures — single-run jitter, not a regression, exactly as BENCH.md's stated methodology warns to expect. This run was noisier than the previous spot-check, though, with several rows well outside that band, worth naming rather than averaging away: every dir-backend metadata row (create/read/stat/unlink/ mkdir) is elevated 24–53% together, pointing at host-side I/O contention during this run rather than anything PackFS-specific (the mem and pack rows, touching no host filesystem, mostly did not move); and two mem rows are outliers even by that standard — rmdir 4,000 dirs (+987.5%, 0.0040s → 0.0435s) and read 64MB (+837.5%, 0.0032s → 0.0300s). Both are large relative jumps on a small absolute base (tens of milliseconds), exactly where a single scheduling stall has the most disproportionate effect on a percentage, and neither is corroborated by a neighboring row moving the same way (mkdir and unlink on mem, right next to rmdir, both stayed flat or improved; read 16MB/write 64MB on mem, right next to read 64MB, both stayed within a few percent) — the single-run, non-statistical methodology BENCH.md states up front means this can't be distinguished from noise without repeated runs, so it is reported here rather than either hidden or asserted as a regression. Total wall-clock: 4m10.1s this run vs 5m23.9s documented — faster overall despite the several elevated dir/raw-fs rows above, because the dominant cost throughout every run of this benchmark is the single raw+fsync 20,000-file create test (~116s here too, unaffected — see BENCH.md's "Environment" section for why that test in particular is unrelated to PackFS's own performance).

Quick example

#include <packfs.h>

Vfs *v = vfs_new();

/* a plain in-memory writable tree */
Backend *mem = backend_mem_new();
vfs_mount(v, "/", mem);

int err = 0;
VfsFile *f = vfs_open(v, "/hello.txt", VFS_O_WRONLY | VFS_O_CREAT, &err);
vfs_write(f, "hello", 5);
vfs_close(f);

vfs_free(v);
backend_free(mem);

A shipped pack with a writable overlay on top:

Backend *mem = backend_mem_new();
int err = 0;
Backend *ov = backend_overlay_new("assets.pack", mem, &err); /* loads assets.pack if it exists */
vfs_mount(v, "/", ov);

/* ... reads served straight from the pack; writes copy-up into mem ... */

vfs_sync(v, "/"); /* compacts the overlay into a fresh assets.pack (Section 4.1, 5.3) */

A sandboxed host directory:

int derr = 0;
Backend *dir = backend_dir_new("/var/lib/myapp/data", &derr);
vfs_mount(v, "/data", dir);
/* every lookup under /data is contained to that directory (Section 6.2);
 * ".." and symlink escapes are rejected, not merely discouraged. */

A shipped pack mounted read-only, no writable layer at all:

int perr = 0;
Backend *ro = backend_pack_new("assets.pack", &perr);
vfs_mount(v, "/assets", ro);
/* vfs_write, vfs_mkdir, vfs_unlink, vfs_rename, vfs_sync against anything
 * under /assets all return VFS_ERR_PERM; there is no upper layer to
 * absorb a write into. */

API

The public API is include/packfs.h; every function and struct is documented there with a pointer to the concept.md section that specifies its behavior. In outline:

  • vfs_new / vfs_free — a Vfs owns a mount table, nothing else.
  • backend_mem_new / backend_dir_new / backend_pack_new / backend_overlay_new — construct a backend; vfs_mount/vfs_unmount attach or detach it at a path prefix. backend_pack_new mounts a pack read-only, with no writable upper layer at all — every mutating call against it returns VFS_ERR_PERM; wrap the same pack in backend_overlay_new instead when writability is wanted.
  • vfs_open / vfs_read / vfs_write / vfs_close — file I/O.
  • vfs_stat / vfs_readdir / vfs_mkdir / vfs_unlink / vfs_rename — metadata and namespace operations.
  • vfs_sync — compacts an overlay mount into a fresh pack.
  • vfs_harden_process_with_landlock — optional, opt-in, process-wide Landlock confinement to the process's current dir mounts. Deliberately not applied automatically by vfs_mount (Section 6.2 explains why: Landlock restrictions are irreversible and process-wide, which would be a surprising side effect for an embeddable library to trigger on its own).

Concurrency

Single-writer, wait-free-reader (Section 5): readers never take a lock and never observe a write in progress; structural changes (create/unlink/rename/ mkdir/mount/unmount) publish a new immutable snapshot via one atomic pointer swap; ordinary content writes to an already-existing file update a per-entry cell directly and never touch the snapshot. mem-backed buffer growth never mutates a buffer address a reader might be reading (Section 5.7): growth always allocates a new buffer and publishes it, never reallocates in place. tests/test_concurrency.c exercises this under concurrent reader and writer threads. The suite is regularly run under AddressSanitizer/ UndefinedBehaviorSanitizer, both locally and in CI (see .gitea/workflows/ci.yml), and this is not aspirational — ASan caught a real heap-use-after-free in the snapshot-reclamation logic during development (see the reclaim_gate note in src/internal.h). ThreadSanitizer is configured the same way but, as of this writing, has not actually completed a run in either environment this project has been built and tested in so far — the local development sandbox and this project's own CI runner both block the personality(ADDR_NO_RANDOMIZE) syscall TSan needs to start (see CONTRIBUTING.md for the confirming tests in each case). CI's TSan step is written to not fail the build over that specific, known-benign failure, which means a green CI run is not evidence TSan actually executed — stated plainly here rather than left to be assumed from CI showing green.

Security

dir mounts are capability-scoped (Section 6.2): a mount holds an already-open directory file descriptor, not a path string, and every lookup beneath it is resolved with openat2(RESOLVE_BENEATH | RESOLVE_NO_SYMLINKS) on Linux 5.6+, which atomically rejects .. and symlink escapes in one kernel call. Where that syscall is unavailable, containment falls back to per-component O_NOFOLLOW resolution — weaker, and documented as such (Section 6.3), not silently assumed equivalent. Pack files are treated as untrusted input unless they came from this process's own compaction: every on-disk offset is bounds-checked and the pack's checksum is verified before any of it is trusted (Section 7). See tests/test_dir.c for a containment regression test and tests/test_pack_overlay.c for a corrupted-pack rejection test. See SECURITY.md for exactly what is and isn't claimed as a security boundary, and how to report a vulnerability.

Contributing

See CONTRIBUTING.md. Read concept.md and CLAUDE.md first — they are the project's actual specification and its enforced documentation standard, respectively, and every design decision in the code traces back to one of them.

License

MIT — see LICENSE.

S
Description
A statically linked, in-process virtual file system for C.
Readme MIT
371 KiB
Languages
C 98.7%
Makefile 1.3%