CI / build-and-test (push) Failing after 23s
The real Gitea Actions run on this project's own registered runner just failed exactly as CONTRIBUTING.md already anticipated it might: FATAL: ThreadSanitizer: unexpected memory mapping, the same signature already documented as a sandbox/container seccomp restriction blocking personality(ADDR_NO_RANDOMIZE), which TSan needs to start at all. Build, test suite, and ASan/UBSan all passed -- only the TSan step failed, on an environment issue, not a code issue. Fixed the CI step itself rather than just noting the failure: it now classifies each binary's TSan run as PASS, a known flake (that exact FATAL signature and nothing indicating an actual race was found), or a real failure (anything else -- a genuine data race, a crash, any other error). Only a real failure fails the build. Verified the classification logic locally against four cases (a synthetic real race report, a plain assertion failure, the known flake signature alone, and a clean pass) -- each classified correctly -- and against this sandbox's own six test binaries, all six of which hit the real flake (this sandbox has never been able to run TSan either) and correctly did not fail the build. This does NOT mean TSan verification is happening in CI -- it means CI no longer conflates "TSan couldn't start" with "the build is broken." Updated CONTRIBUTING.md and CLAUDE.md from "whether the runner can execute TSan is unverified" (a hedge) to the now-confirmed fact that it can't, and corrected an overclaim in README.md that the suite is "regularly run under ThreadSanitizer" -- as far as this project has been able to confirm, TSan has not actually completed a run in any environment it's been built in yet, local sandbox or CI. ASan/UBSan remain the real, run, load-bearing sanitizer coverage. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
303 lines
15 KiB
Markdown
303 lines
15 KiB
Markdown
# PackFS
|
||
|
||
A statically linked, in-process virtual file system for C. PackFS treats a
|
||
shipped file tree as an immutable **pack** image and derives writability
|
||
from a `mem` or `dir` upper layer through a copy-on-write overlay. `zip` and
|
||
`tar` are not part of this project and never will be — see "Status" below.
|
||
|
||
The full design rationale — why this shape, what alternatives were rejected,
|
||
and the concurrency, path-containment, and integrity models this
|
||
implementation follows — is specified in [`concept.md`](concept.md), which is
|
||
frozen (see [`CLAUDE.md`](CLAUDE.md)) and is the authoritative source for
|
||
every design decision below. This README documents the implementation that
|
||
followed from it, not a restatement of the rationale.
|
||
|
||
## Status
|
||
|
||
This is an initial, partial implementation of the spec, not a complete one.
|
||
|
||
**Implemented and tested:** mount table, `mem`/`dir`/`pack`/overlay backends
|
||
(including the standalone read-only `pack` backend, `backend_pack_new` —
|
||
every mutating call against it returns `VFS_ERR_PERM`), copy-up, whiteouts,
|
||
compaction, an append journal, path containment, and pack integrity
|
||
validation.
|
||
|
||
**Will never be built, by explicit project decision:** the `zip` (miniz) and
|
||
`tar` (USTAR) import/export backends. `concept.md` Section 11 recommends
|
||
them, but that recommendation is superseded — see `CLAUDE.md`, "Project
|
||
decisions that supersede concept.md." `backend_overlay_new` reads and writes
|
||
this project's own pack format exclusively; there is no zip/tar support and
|
||
none is planned. Do not open an issue or PR adding one.
|
||
|
||
**Deliberately out of scope for v0** (`concept.md` Section 10, not gaps):
|
||
full POSIX semantics, enforced permissions/symlinks/hard links,
|
||
cross-process concurrency (design specified in Section 5.6, unimplemented),
|
||
and content-defined chunking/delta compression.
|
||
|
||
**Tested on Linux only**, in one environment. The `dir`-mount containment
|
||
fallback path for kernels without `openat2` (Section 6.3) is implemented but
|
||
has not been exercised on such a kernel, nor on macOS or Windows.
|
||
|
||
## Building
|
||
|
||
Zero required third-party dependencies — only a C11 compiler, `make`, and
|
||
`pthread` (Section 11.1 of `concept.md` makes this a hard constraint, not a
|
||
preference).
|
||
|
||
```sh
|
||
make # builds libpackfs.a and libpackfs.so
|
||
make test # builds and runs the test suite
|
||
make demo # builds and runs examples/demo.c — see "Try it" below
|
||
make bench # builds and runs bench/bench.c — see "Benchmarks" below
|
||
make install # installs to $PREFIX (default /usr/local), including a pkg-config file
|
||
```
|
||
|
||
`make install` also generates and installs `packfs.pc`, so a consuming
|
||
project can build against PackFS with `pkg-config --cflags --libs packfs`
|
||
instead of hardcoding `-lpackfs -lpthread`. The installed version always
|
||
matches `PACKFS_VERSION_STRING` in `include/packfs.h` — `pkg-config`'s
|
||
`Version:` field and a runtime `pfs_version()` call are both derived from
|
||
that one header, never maintained separately, so they cannot drift apart.
|
||
|
||
## Try it
|
||
|
||
`examples/demo.c` is a small, runnable, human-readable program — not another
|
||
automated test — that exercises the library end to end and prints what it
|
||
did at each step: a pack-backed overlay (write, read, `mkdir`, `readdir`,
|
||
`stat`, a copy-up-then-whiteout delete, compaction via `vfs_sync`), and a
|
||
sandboxed `dir` mount that demonstrates a `../../../etc/passwd` escape
|
||
attempt being rejected. Run `make demo` twice in a row: the second run's
|
||
first `readdir` shows the first run's files, proving that compaction and
|
||
reload actually persist data through the pack file, not just within one
|
||
process's lifetime.
|
||
|
||
## Benchmarks
|
||
|
||
`bench/bench.c` (`make bench`) measures PackFS against the host filesystem
|
||
across metadata operations (create/read/stat/readdir/unlink/mkdir), large
|
||
sequential I/O, random-access pack reads, mount-table scaling, and
|
||
concurrent mixed workloads. [`BENCH.md`](BENCH.md) has the full results and
|
||
honest analysis of both the wins and the losses, including the story of
|
||
three real O(n²) findings this project's own benchmarking turned up — not
|
||
just the wins. Two are fixed: bulk sequential file creation (`src/upper.c`'s
|
||
index is a persistent treap now, not a flat array — see "Resolution") and
|
||
`pack_write`'s compaction-time duplicate-content elimination, which also had
|
||
a latent correctness bug now closed alongside it (see "Resolution #2"). One
|
||
is confirmed and *deliberately not* fixed: the mount table scales O(n²) in
|
||
mount count, the same way the file index used to, but mount counts are
|
||
bounded by a program's own source code rather than workload-driven, so it
|
||
isn't worth the added complexity — see "Finding: mount table scaling" for
|
||
the reasoning and the numbers behind that call. Read the whole file before
|
||
quoting a number from it: what `raw fs` vs `raw+fsync` vs `dir` each
|
||
actually measure is not interchangeable, and it explains why.
|
||
|
||
`openat2`/Landlock support (Section 6) is detected automatically at compile
|
||
time via `<sys/syscall.h>`; on kernels or platforms without them, `dir`
|
||
mounts fall back to the weaker, documented residual-risk posture described
|
||
in `concept.md` Section 6.3 rather than failing to build.
|
||
|
||
### Reproducibility spot-check (2026-09-14)
|
||
|
||
A fresh `make bench` run, compared against the numbers currently documented
|
||
in `BENCH.md`'s "After" table, on the same environment described there. Not
|
||
a replacement for `BENCH.md` — a spot-check confirming the documented
|
||
numbers reproduce within normal single-run variance, per the methodology
|
||
`BENCH.md` itself states ("illustrative of shape... not precise absolute
|
||
figures").
|
||
|
||
| Category | Backend | Documented (BENCH.md) | New run | Delta |
|
||
|---|---|---|---|---|
|
||
| create 20,000 files | mem | 0.0400s | 0.0394s | -1.5% |
|
||
| create 20,000 files | raw fs | 0.9943s | 1.0171s | +2.3% |
|
||
| create 20,000 files | raw+fsync | 116.2147s | 116.7567s | +0.5% |
|
||
| read 20,000 files | mem | 0.0099s | 0.0091s | -8.1% |
|
||
| read 20,000 files | raw fs | 0.1606s | 0.1634s | +1.7% |
|
||
| stat 20,000 files | mem | 0.0077s | 0.0078s | +1.3% |
|
||
| stat 20,000 files | raw fs | 0.0699s | 0.0727s | +4.0% |
|
||
| readdir (20,000 entries) | mem | 0.0041s | 0.0043s | +4.9% |
|
||
| readdir (20,000 entries) | raw fs | 0.0069s | 0.0072s | +4.3% |
|
||
| create 20,000 files | dir | 1.1418s | 1.2030s | +5.4% |
|
||
| read 20,000 files | dir | 0.1530s | 0.1715s | +12.1% |
|
||
| stat 20,000 files | dir | 0.1379s | 0.1546s | +12.1% |
|
||
| readdir (20,000 entries) | dir | 0.0037s | 0.0041s | +10.8% |
|
||
| unlink 20,000 files | mem | 0.0179s | 0.0182s | +1.7% |
|
||
| unlink 20,000 files | dir | 0.4911s | 0.5978s | +21.7% |
|
||
| unlink 20,000 files | raw fs | 0.4746s | 0.5062s | +6.7% |
|
||
| mkdir 4,000 dirs | mem | 0.0056s | 0.0050s | -10.7% |
|
||
| rmdir 4,000 dirs | mem | 0.0040s | 0.0039s | -2.5% |
|
||
| mkdir 4,000 dirs | dir | 0.1710s | 0.1864s | +9.0% |
|
||
| rmdir 4,000 dirs | dir | 0.1257s | 0.1169s | -7.0% |
|
||
| mkdir 4,000 dirs | raw fs | 0.1374s | 0.1444s | +5.1% |
|
||
| rmdir 4,000 dirs | raw fs | 0.0951s | 0.1005s | +5.7% |
|
||
| write 1MB | mem | 0.0001s | 0.0001s | +0.0% |
|
||
| read 1MB | mem | 0.0000s | 0.0000s | +0.0% |
|
||
| write 16MB | mem | 0.0110s | 0.0110s | +0.0% |
|
||
| read 16MB | mem | 0.0009s | 0.0010s | +11.1% |
|
||
| write 64MB | mem | 0.0586s | 0.0621s | +6.0% |
|
||
| read 64MB | mem | 0.0032s | 0.0036s | +12.5% |
|
||
| write 1MB | raw fs | 0.0004s | 0.0004s | +0.0% |
|
||
| read 1MB | raw fs | 0.0001s | 0.0001s | +0.0% |
|
||
| write 16MB | raw fs | 0.0044s | 0.0047s | +6.8% |
|
||
| read 16MB | raw fs | 0.0010s | 0.0012s | +20.0% |
|
||
| write 64MB | raw fs | 0.0186s | 0.0189s | +1.6% |
|
||
| read 64MB | raw fs | 0.0047s | 0.0050s | +6.4% |
|
||
| write 1MB | raw+fsync | 0.0180s | 0.0340s | +88.9% |
|
||
| read 1MB | raw+fsync | 0.0001s | 0.0001s | +0.0% |
|
||
| write 16MB | raw+fsync | 0.0231s | 0.0242s | +4.8% |
|
||
| read 16MB | raw+fsync | 0.0012s | 0.0012s | +0.0% |
|
||
| write 64MB | raw+fsync | 0.0803s | 0.0835s | +4.0% |
|
||
| read 64MB | raw+fsync | 0.0048s | 0.0049s | +2.1% |
|
||
| compact 20,000 entries to pack | pack | 0.0267s | 0.0280s | +4.9% |
|
||
| random-read 20,000 entries | pack (mmap'd) | 0.0082s | 0.0085s | +3.7% |
|
||
| random-read 20,000 entries | raw fs | 0.1629s | 0.1912s | +17.4% |
|
||
| concurrent create+read+unlink (8×4,000×3) | mem | 0.2750s | 0.2791s | +1.5% |
|
||
| concurrent create+read+unlink (8×4,000×3) | raw fs | 5.6498s | 6.1583s | +9.0% |
|
||
| mount 500 backends | vfs | 0.0054s | 0.0060s | +11.1% |
|
||
| resolve, 500 mounts | vfs | 0.0020s | 0.0022s | +10.0% |
|
||
| unmount 500 backends | vfs | 0.0046s | 0.0046s | +0.0% |
|
||
| mount 2,000 backends | vfs | 0.0831s | 0.0863s | +3.9% |
|
||
| resolve, 2,000 mounts | vfs | 0.0296s | 0.0329s | +11.1% |
|
||
| unmount 2,000 backends | vfs | 0.0758s | 0.0780s | +2.9% |
|
||
| mount 8,000 backends | vfs | 1.4739s | 1.5380s | +4.3% |
|
||
| resolve, 8,000 mounts | vfs | 0.4938s | 0.5298s | +7.3% |
|
||
| unmount 8,000 backends | vfs | 1.5172s | 1.5191s | +0.1% |
|
||
|
||
Almost every row sits within ±15% of the documented figures — consistent
|
||
with the single-run jitter `BENCH.md` already warns about, not a
|
||
regression. Two rows exceed that: `unlink 20,000 files (dir)` (+21.7%,
|
||
plausible container/host I/O noise, same direction as the other `dir`
|
||
metadata rows this run) and `write 1MB (raw+fsync)` (+88.9%, but this is a
|
||
tiny absolute value — 18ms vs 34ms for one syscall — the single number most
|
||
sensitive to one slow `fsync` on this container's overlay filesystem, not a
|
||
meaningful regression at that scale). Total wall-clock: 4m7.7s this run vs
|
||
5m23.9s documented, itself within the same single-run variance.
|
||
|
||
## Quick example
|
||
|
||
```c
|
||
#include <packfs.h>
|
||
|
||
Vfs *v = vfs_new();
|
||
|
||
/* a plain in-memory writable tree */
|
||
Backend *mem = backend_mem_new();
|
||
vfs_mount(v, "/", mem);
|
||
|
||
int err = 0;
|
||
VfsFile *f = vfs_open(v, "/hello.txt", VFS_O_WRONLY | VFS_O_CREAT, &err);
|
||
vfs_write(f, "hello", 5);
|
||
vfs_close(f);
|
||
|
||
vfs_free(v);
|
||
backend_free(mem);
|
||
```
|
||
|
||
A shipped pack with a writable overlay on top:
|
||
|
||
```c
|
||
Backend *mem = backend_mem_new();
|
||
int err = 0;
|
||
Backend *ov = backend_overlay_new("assets.pack", mem, &err); /* loads assets.pack if it exists */
|
||
vfs_mount(v, "/", ov);
|
||
|
||
/* ... reads served straight from the pack; writes copy-up into mem ... */
|
||
|
||
vfs_sync(v, "/"); /* compacts the overlay into a fresh assets.pack (Section 4.1, 5.3) */
|
||
```
|
||
|
||
A sandboxed host directory:
|
||
|
||
```c
|
||
int derr = 0;
|
||
Backend *dir = backend_dir_new("/var/lib/myapp/data", &derr);
|
||
vfs_mount(v, "/data", dir);
|
||
/* every lookup under /data is contained to that directory (Section 6.2);
|
||
* ".." and symlink escapes are rejected, not merely discouraged. */
|
||
```
|
||
|
||
A shipped pack mounted read-only, no writable layer at all:
|
||
|
||
```c
|
||
int perr = 0;
|
||
Backend *ro = backend_pack_new("assets.pack", &perr);
|
||
vfs_mount(v, "/assets", ro);
|
||
/* vfs_write, vfs_mkdir, vfs_unlink, vfs_rename, vfs_sync against anything
|
||
* under /assets all return VFS_ERR_PERM; there is no upper layer to
|
||
* absorb a write into. */
|
||
```
|
||
|
||
## API
|
||
|
||
The public API is [`include/packfs.h`](include/packfs.h); every function and
|
||
struct is documented there with a pointer to the `concept.md` section that
|
||
specifies its behavior. In outline:
|
||
|
||
- `vfs_new` / `vfs_free` — a `Vfs` owns a mount table, nothing else.
|
||
- `backend_mem_new` / `backend_dir_new` / `backend_pack_new` /
|
||
`backend_overlay_new` — construct a backend; `vfs_mount`/`vfs_unmount`
|
||
attach or detach it at a path prefix. `backend_pack_new` mounts a pack
|
||
read-only, with no writable upper layer at all — every mutating call
|
||
against it returns `VFS_ERR_PERM`; wrap the same pack in
|
||
`backend_overlay_new` instead when writability is wanted.
|
||
- `vfs_open` / `vfs_read` / `vfs_write` / `vfs_close` — file I/O.
|
||
- `vfs_stat` / `vfs_readdir` / `vfs_mkdir` / `vfs_unlink` / `vfs_rename` —
|
||
metadata and namespace operations.
|
||
- `vfs_sync` — compacts an overlay mount into a fresh pack.
|
||
- `vfs_harden_process_with_landlock` — optional, opt-in, process-wide
|
||
Landlock confinement to the process's current `dir` mounts. Deliberately
|
||
**not** applied automatically by `vfs_mount` (Section 6.2 explains why:
|
||
Landlock restrictions are irreversible and process-wide, which would be a
|
||
surprising side effect for an embeddable library to trigger on its own).
|
||
|
||
## Concurrency
|
||
|
||
Single-writer, wait-free-reader (Section 5): readers never take a lock and
|
||
never observe a write in progress; structural changes (create/unlink/rename/
|
||
mkdir/mount/unmount) publish a new immutable snapshot via one atomic pointer
|
||
swap; ordinary content writes to an already-existing file update a per-entry
|
||
cell directly and never touch the snapshot. `mem`-backed buffer growth
|
||
never mutates a buffer address a reader might be reading (Section 5.7):
|
||
growth always allocates a new buffer and publishes it, never reallocates in
|
||
place. `tests/test_concurrency.c` exercises this under concurrent reader and
|
||
writer threads. The suite is regularly run under AddressSanitizer/
|
||
UndefinedBehaviorSanitizer, both locally and in CI (see
|
||
`.gitea/workflows/ci.yml`), and this is not aspirational — ASan caught a
|
||
real heap-use-after-free in the snapshot-reclamation logic during
|
||
development (see the `reclaim_gate` note in `src/internal.h`).
|
||
ThreadSanitizer is configured the same way but, as of this writing, has
|
||
not actually completed a run in either environment this project has been
|
||
built and tested in so far — the local development sandbox and this
|
||
project's own CI runner both block the `personality(ADDR_NO_RANDOMIZE)`
|
||
syscall TSan needs to start (see `CONTRIBUTING.md` for the confirming
|
||
tests in each case). CI's TSan step is written to not fail the build over
|
||
that specific, known-benign failure, which means a green CI run is not
|
||
evidence TSan actually executed — stated plainly here rather than left to
|
||
be assumed from CI showing green.
|
||
|
||
## Security
|
||
|
||
`dir` mounts are capability-scoped (Section 6.2): a mount holds an already-open
|
||
directory file descriptor, not a path string, and every lookup beneath it is
|
||
resolved with `openat2(RESOLVE_BENEATH | RESOLVE_NO_SYMLINKS)` on Linux 5.6+,
|
||
which atomically rejects `..` and symlink escapes in one kernel call. Where
|
||
that syscall is unavailable, containment falls back to per-component
|
||
`O_NOFOLLOW` resolution — weaker, and documented as such (Section 6.3), not
|
||
silently assumed equivalent. Pack files are treated as untrusted input unless
|
||
they came from this process's own compaction: every on-disk offset is
|
||
bounds-checked and the pack's checksum is verified before any of it is
|
||
trusted (Section 7). See `tests/test_dir.c` for a containment regression test
|
||
and `tests/test_pack_overlay.c` for a corrupted-pack rejection test. See
|
||
[`SECURITY.md`](SECURITY.md) for exactly what is and isn't claimed as a
|
||
security boundary, and how to report a vulnerability.
|
||
|
||
## Contributing
|
||
|
||
See [`CONTRIBUTING.md`](CONTRIBUTING.md). Read `concept.md` and `CLAUDE.md`
|
||
first — they are the project's actual specification and its enforced
|
||
documentation standard, respectively, and every design decision in the code
|
||
traces back to one of them.
|
||
|
||
## License
|
||
|
||
MIT — see [`LICENSE`](LICENSE).
|