Files
retoorandClaude Sonnet 5 7661c94105
CI / build-and-test (push) Successful in 48s
Add POSTMORTEM.md: every real issue found, how, fix, and regression prevention
A detailed, permanent record covering the whole arc of this project's
development so far -- requested explicitly, as detailed as possible, to
prevent regression for ever. Seven sections:

1. Performance findings: the file-index O(n^2) treap rewrite, pack_write's
   O(n^2) dedup + latent hash-collision correctness bug, and the mount
   table's O(n^2) (confirmed, deliberately not fixed, with the reasoning).
2. Environment/tooling limitations: TSan's categorical block (every
   workaround actually tried and ruled out, not just the ones that
   worked), and both variants of the ASan/UBSan sandbox-startup flake.
3. Project professionalization: version API, SPDX, pkg-config, SECURITY.md,
   CHANGELOG.md, the Gitea-not-GitHub migration, and the CODE_OF_CONDUCT
   decision.
4. Git identity correction: the filter-branch rewrite, the tag-object
   tagger field it missed (found only by a full object-database sweep,
   not by re-reading git log), and the exhaustive re-verification.
5. Data-integrity fault injection: test_crash_consistency.c, the real bug
   in its own oracle (128 false failures before a progress side-channel
   fixed it), and the two real overlay.c bugs found while building the
   write-failure test (silent journal failure, torn-record poisoning of
   later records).
6. CI as a second, independent reviewer: the memory leak detect_leaks=0
   had been masking locally, and the set -e scripting bug -- including
   the mistake made while fixing it (reintroducing the same bug in a new
   shape), the follow-on bash-variable-blowup bug, the newly-discovered
   TSan segfault variant, and the retry-budget gap, each confirmed by
   direct reproduction under bash -eo pipefail, not reasoned about.
7. Distilled lessons: ten patterns extracted from the above, written to
   be applied to a new problem, not just recognized in this one.

Every figure cited was checked against BENCH.md/CLAUDE.md's own numbers
before writing this, not reproduced from memory.

Cross-referenced from CLAUDE.md's "Repository status", README.md's
"Contributing" section, and CHANGELOG.md -- including two entries CHANGELOG
itself was missing (the set -e CI fix from the previous commit, and this
document), the exact class of gap POSTMORTEM.md section 3 already
describes happening once before with the v0.1.0 tag.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
2026-09-14 21:02:58 +00:00

333 lines
19 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# PackFS
A statically linked, in-process virtual file system for C. PackFS treats a
shipped file tree as an immutable **pack** image and derives writability
from a `mem` or `dir` upper layer through a copy-on-write overlay. `zip` and
`tar` are not part of this project and never will be — see "Status" below.
The full design rationale — why this shape, what alternatives were rejected,
and the concurrency, path-containment, and integrity models this
implementation follows — is specified in [`concept.md`](concept.md), which is
frozen (see [`CLAUDE.md`](CLAUDE.md)) and is the authoritative source for
every design decision below. This README documents the implementation that
followed from it, not a restatement of the rationale.
## Status
This is an initial, partial implementation of the spec, not a complete one.
**Implemented and tested:** mount table, `mem`/`dir`/`pack`/overlay backends
(including the standalone read-only `pack` backend, `backend_pack_new` —
every mutating call against it returns `VFS_ERR_PERM`), copy-up, whiteouts,
compaction, an append journal, path containment, and pack integrity
validation.
**Will never be built, by explicit project decision:** the `zip` (miniz) and
`tar` (USTAR) import/export backends. `concept.md` Section 11 recommends
them, but that recommendation is superseded — see `CLAUDE.md`, "Project
decisions that supersede concept.md." `backend_overlay_new` reads and writes
this project's own pack format exclusively; there is no zip/tar support and
none is planned. Do not open an issue or PR adding one.
**Deliberately out of scope for v0** (`concept.md` Section 10, not gaps):
full POSIX semantics, enforced permissions/symlinks/hard links,
cross-process concurrency (design specified in Section 5.6, unimplemented),
and content-defined chunking/delta compression.
**Tested on Linux only**, in one environment. The `dir`-mount containment
fallback path for kernels without `openat2` (Section 6.3) is implemented but
has not been exercised on such a kernel, nor on macOS or Windows.
## Building
Zero required third-party dependencies — only a C11 compiler, `make`, and
`pthread` (Section 11.1 of `concept.md` makes this a hard constraint, not a
preference).
```sh
make # builds libpackfs.a and libpackfs.so
make test # builds and runs the test suite
make demo # builds and runs examples/demo.c — see "Try it" below
make bench # builds and runs bench/bench.c — see "Benchmarks" below
make install # installs to $PREFIX (default /usr/local), including a pkg-config file
```
`make install` also generates and installs `packfs.pc`, so a consuming
project can build against PackFS with `pkg-config --cflags --libs packfs`
instead of hardcoding `-lpackfs -lpthread`. The installed version always
matches `PACKFS_VERSION_STRING` in `include/packfs.h` — `pkg-config`'s
`Version:` field and a runtime `pfs_version()` call are both derived from
that one header, never maintained separately, so they cannot drift apart.
## Try it
`examples/demo.c` is a small, runnable, human-readable program — not another
automated test — that exercises the library end to end and prints what it
did at each step: a pack-backed overlay (write, read, `mkdir`, `readdir`,
`stat`, a copy-up-then-whiteout delete, compaction via `vfs_sync`), and a
sandboxed `dir` mount that demonstrates a `../../../etc/passwd` escape
attempt being rejected. Run `make demo` twice in a row: the second run's
first `readdir` shows the first run's files, proving that compaction and
reload actually persist data through the pack file, not just within one
process's lifetime.
## Benchmarks
`bench/bench.c` (`make bench`) measures PackFS against the host filesystem
across metadata operations (create/read/stat/readdir/unlink/mkdir), large
sequential I/O, random-access pack reads, mount-table scaling, and
concurrent mixed workloads. [`BENCH.md`](BENCH.md) has the full results and
honest analysis of both the wins and the losses, including the story of
three real O(n²) findings this project's own benchmarking turned up — not
just the wins. Two are fixed: bulk sequential file creation (`src/upper.c`'s
index is a persistent treap now, not a flat array — see "Resolution") and
`pack_write`'s compaction-time duplicate-content elimination, which also had
a latent correctness bug now closed alongside it (see "Resolution #2"). One
is confirmed and *deliberately not* fixed: the mount table scales O(n²) in
mount count, the same way the file index used to, but mount counts are
bounded by a program's own source code rather than workload-driven, so it
isn't worth the added complexity — see "Finding: mount table scaling" for
the reasoning and the numbers behind that call. Read the whole file before
quoting a number from it: what `raw fs` vs `raw+fsync` vs `dir` each
actually measure is not interchangeable, and it explains why.
`openat2`/Landlock support (Section 6) is detected automatically at compile
time via `<sys/syscall.h>`; on kernels or platforms without them, `dir`
mounts fall back to the weaker, documented residual-risk posture described
in `concept.md` Section 6.3 rather than failing to build.
### Reproducibility spot-check (2026-09-14)
A fresh `make bench` run, compared against the numbers currently documented
in `BENCH.md`'s "After" table, on the same environment described there. Not
a replacement for `BENCH.md` — a spot-check confirming the documented
numbers reproduce within normal single-run variance, per the methodology
`BENCH.md` itself states ("illustrative of shape... not precise absolute
figures"). Every column from the raw `make bench` output is kept for both
runs — time, throughput, and MB/s where the category reports one — not
just a time delta, so every category (including the `pack` rows:
compaction and random-access read) can be checked in full, not summarized
away.
| Category | Backend | Doc Time | Doc Throughput | Doc MB/s | New Time | New Throughput | New MB/s | Δ Time |
|---|---|---|---|---|---|---|---|---|
| create 20,000 files | mem | 0.0400s | 499,809 ops/s | 61.0 | 0.0373s | 536,419 ops/s | 65.5 | -6.8% |
| create 20,000 files | raw fs | 0.9943s | 20,114 ops/s | 2.5 | 0.9883s | 20,237 ops/s | 2.5 | -0.6% |
| create 20,000 files | raw+fsync | 116.2147s | 172 ops/s | 0.0 | 116.3080s | 172 ops/s | 0.0 | +0.1% |
| read 20,000 files | mem | 0.0099s | 2,011,901 ops/s | 245.6 | 0.0116s | 1,723,405 ops/s | 210.4 | +17.2% |
| read 20,000 files | raw fs | 0.1606s | 124,543 ops/s | 15.2 | 0.2040s | 98,030 ops/s | 12.0 | +27.0% |
| stat 20,000 files | mem | 0.0077s | 2,598,241 ops/s | — | 0.0079s | 2,519,089 ops/s | — | +2.6% |
| stat 20,000 files | raw fs | 0.0699s | 286,222 ops/s | — | 0.0954s | 209,541 ops/s | — | +36.5% |
| readdir (20,000 entries) | mem | 0.0041s | 4,845,826 ops/s | — | 0.0043s | 4,675,035 ops/s | — | +4.9% |
| readdir (20,000 entries) | raw fs | 0.0069s | 2,894,815 ops/s | — | 0.0071s | 2,808,607 ops/s | — | +2.9% |
| create 20,000 files | dir | 1.1418s | 17,517 ops/s | 2.1 | 1.7006s | 11,760 ops/s | 1.4 | +48.9% |
| read 20,000 files | dir | 0.1530s | 130,705 ops/s | 16.0 | 0.2238s | 89,371 ops/s | 10.9 | +46.3% |
| stat 20,000 files | dir | 0.1379s | 145,056 ops/s | — | 0.1984s | 100,826 ops/s | — | +43.9% |
| readdir (20,000 entries) | dir | 0.0037s | 5,465,162 ops/s | — | 0.0042s | 4,799,105 ops/s | — | +13.5% |
| unlink 20,000 files | mem | 0.0179s | 1,117,742 ops/s | — | 0.0183s | 1,094,715 ops/s | — | +2.2% |
| unlink 20,000 files | dir | 0.4911s | 40,724 ops/s | — | 0.7500s | 26,667 ops/s | — | +52.7% |
| unlink 20,000 files | raw fs | 0.4746s | 42,141 ops/s | — | 0.6404s | 31,228 ops/s | — | +34.9% |
| mkdir 4,000 dirs | mem | 0.0056s | 719,696 ops/s | — | 0.0051s | 791,918 ops/s | — | -8.9% |
| rmdir 4,000 dirs | mem | 0.0040s | 993,724 ops/s | — | 0.0435s | 91,873 ops/s | — | +987.5% |
| mkdir 4,000 dirs | dir | 0.1710s | 23,386 ops/s | — | 0.2123s | 18,838 ops/s | — | +24.2% |
| rmdir 4,000 dirs | dir | 0.1257s | 31,827 ops/s | — | 0.1372s | 29,163 ops/s | — | +9.1% |
| mkdir 4,000 dirs | raw fs | 0.1374s | 29,122 ops/s | — | 0.1837s | 21,772 ops/s | — | +33.7% |
| rmdir 4,000 dirs | raw fs | 0.0951s | 42,043 ops/s | — | 0.1305s | 30,640 ops/s | — | +37.2% |
| write 1MB | mem | 0.0001s | 7,058 ops/s | 7,058.1 | 0.0002s | 6,313 ops/s | 6,312.7 | +100.0% |
| read 1MB | mem | 0.0000s | 50,051 ops/s | 50,050.9 | 0.0000s | 48,616 ops/s | 48,616.4 | +0.0% |
| write 16MB | mem | 0.0110s | 91 ops/s | 1,458.2 | 0.0407s | 25 ops/s | 393.3 | +270.0% |
| read 16MB | mem | 0.0009s | 1,058 ops/s | 16,920.1 | 0.0010s | 1,021 ops/s | 16,338.0 | +11.1% |
| write 64MB | mem | 0.0586s | 17 ops/s | 1,091.3 | 0.0599s | 17 ops/s | 1,068.3 | +2.2% |
| read 64MB | mem | 0.0032s | 313 ops/s | 20,056.5 | 0.0300s | 33 ops/s | 2,131.0 | +837.5% |
| write 1MB | raw fs | 0.0004s | 2,759 ops/s | 2,759.4 | 0.0004s | 2,474 ops/s | 2,474.0 | +0.0% |
| read 1MB | raw fs | 0.0001s | 16,812 ops/s | 16,812.4 | 0.0001s | 16,756 ops/s | 16,756.0 | +0.0% |
| write 16MB | raw fs | 0.0044s | 229 ops/s | 3,656.4 | 0.0044s | 228 ops/s | 3,645.0 | +0.0% |
| read 16MB | raw fs | 0.0010s | 989 ops/s | 15,827.9 | 0.0010s | 965 ops/s | 15,443.5 | +0.0% |
| write 64MB | raw fs | 0.0186s | 54 ops/s | 3,449.7 | 0.0184s | 54 ops/s | 3,478.9 | -1.1% |
| read 64MB | raw fs | 0.0047s | 213 ops/s | 13,633.1 | 0.0049s | 203 ops/s | 13,003.0 | +4.3% |
| write 1MB | raw+fsync | 0.0180s | 56 ops/s | 55.7 | 0.0162s | 62 ops/s | 61.7 | -10.0% |
| read 1MB | raw+fsync | 0.0001s | 9,059 ops/s | 9,058.8 | 0.0001s | 12,025 ops/s | 12,025.1 | +0.0% |
| write 16MB | raw+fsync | 0.0231s | 43 ops/s | 691.7 | 0.0247s | 41 ops/s | 648.1 | +6.9% |
| read 16MB | raw+fsync | 0.0012s | 855 ops/s | 13,672.2 | 0.0014s | 726 ops/s | 11,610.9 | +16.7% |
| write 64MB | raw+fsync | 0.0803s | 12 ops/s | 797.2 | 0.1001s | 10 ops/s | 639.5 | +24.7% |
| read 64MB | raw+fsync | 0.0048s | 209 ops/s | 13,393.5 | 0.0200s | 50 ops/s | 3,193.8 | +316.7% |
| compact 20,000 entries to pack | pack | 0.0267s | 747,839 ops/s | 91.3 | 0.0250s | 800,791 ops/s | 97.8 | -6.4% |
| random-read 20,000 entries | pack (mmap'd) | 0.0082s | 2,435,930 ops/s | 297.4 | 0.0083s | 2,397,682 ops/s | 292.7 | +1.2% |
| random-read 20,000 entries | raw fs | 0.1629s | 122,749 ops/s | 15.0 | 0.1691s | 118,290 ops/s | 14.4 | +3.8% |
| concurrent create+read+unlink (8×4,000×3) | mem | 0.2750s | 349,127 ops/s | — | 0.2981s | 322,041 ops/s | — | +8.4% |
| concurrent create+read+unlink (8×4,000×3) | raw fs | 5.6498s | 16,992 ops/s | — | 6.3837s | 15,038 ops/s | — | +13.0% |
| mount 500 backends | vfs | 0.0054s | 92,833 ops/s | — | 0.0054s | 91,871 ops/s | — | +0.0% |
| resolve, 500 mounts | vfs | 0.0020s | 246,064 ops/s | — | 0.0019s | 262,671 ops/s | — | -5.0% |
| unmount 500 backends | vfs | 0.0046s | 109,207 ops/s | — | 0.0045s | 110,348 ops/s | — | -2.2% |
| mount 2,000 backends | vfs | 0.0831s | 24,081 ops/s | — | 0.0836s | 23,926 ops/s | — | +0.6% |
| resolve, 2,000 mounts | vfs | 0.0296s | 67,458 ops/s | — | 0.0293s | 68,318 ops/s | — | -1.0% |
| unmount 2,000 backends | vfs | 0.0758s | 26,397 ops/s | — | 0.0802s | 24,932 ops/s | — | +5.8% |
| mount 8,000 backends | vfs | 1.4739s | 5,428 ops/s | — | 1.4905s | 5,367 ops/s | — | +1.1% |
| resolve, 8,000 mounts | vfs | 0.4938s | 16,200 ops/s | — | 0.4616s | 17,332 ops/s | — | -6.5% |
| unmount 8,000 backends | vfs | 1.5172s | 5,273 ops/s | — | 1.5176s | 5,271 ops/s | — | +0.0% |
The `pack`-backend rows are the most stable in the whole table (compaction
-6.4%, random-access read +1.2%/+3.8%), consistent with `BENCH.md`'s point
that `pack` random access is a flat `mmap`'d table lookup rather than
anything sensitive to scheduling jitter the way syscall-heavy `dir`/raw-fs
metadata operations are.
Most rows sit within ±15% of the documented figures — single-run jitter,
not a regression, exactly as `BENCH.md`'s stated methodology warns to
expect. This run was noisier than the previous spot-check, though, with
several rows well outside that band, worth naming rather than averaging
away: every `dir`-backend metadata row (`create`/`read`/`stat`/`unlink`/
`mkdir`) is elevated 24–53% together, pointing at host-side I/O contention
during this run rather than anything PackFS-specific (the `mem` and `pack`
rows, touching no host filesystem, mostly did not move); and two `mem`
rows are outliers even by that standard — `rmdir 4,000 dirs` (+987.5%,
0.0040s → 0.0435s) and `read 64MB` (+837.5%, 0.0032s → 0.0300s). Both are
large relative jumps on a small absolute base (tens of milliseconds),
exactly where a single scheduling stall has the most disproportionate
effect on a percentage, and neither is corroborated by a neighboring
row moving the same way (`mkdir` and `unlink` on `mem`, right next to
`rmdir`, both stayed flat or improved; `read 16MB`/`write 64MB` on `mem`,
right next to `read 64MB`, both stayed within a few percent) — the
single-run, non-statistical methodology `BENCH.md` states up front means
this can't be distinguished from noise without repeated runs, so it is
reported here rather than either hidden or asserted as a regression.
Total wall-clock: 4m10.1s this run vs 5m23.9s documented — faster overall
despite the several elevated `dir`/raw-fs rows above, because the dominant
cost throughout every run of this benchmark is the single `raw+fsync`
20,000-file create test (~116s here too, unaffected — see `BENCH.md`'s
"Environment" section for why that test in particular is unrelated to
PackFS's own performance).
## Quick example
```c
#include <packfs.h>
Vfs *v = vfs_new();
/* a plain in-memory writable tree */
Backend *mem = backend_mem_new();
vfs_mount(v, "/", mem);
int err = 0;
VfsFile *f = vfs_open(v, "/hello.txt", VFS_O_WRONLY | VFS_O_CREAT, &err);
vfs_write(f, "hello", 5);
vfs_close(f);
vfs_free(v);
backend_free(mem);
```
A shipped pack with a writable overlay on top:
```c
Backend *mem = backend_mem_new();
int err = 0;
Backend *ov = backend_overlay_new("assets.pack", mem, &err); /* loads assets.pack if it exists */
vfs_mount(v, "/", ov);
/* ... reads served straight from the pack; writes copy-up into mem ... */
vfs_sync(v, "/"); /* compacts the overlay into a fresh assets.pack (Section 4.1, 5.3) */
```
A sandboxed host directory:
```c
int derr = 0;
Backend *dir = backend_dir_new("/var/lib/myapp/data", &derr);
vfs_mount(v, "/data", dir);
/* every lookup under /data is contained to that directory (Section 6.2);
* ".." and symlink escapes are rejected, not merely discouraged. */
```
A shipped pack mounted read-only, no writable layer at all:
```c
int perr = 0;
Backend *ro = backend_pack_new("assets.pack", &perr);
vfs_mount(v, "/assets", ro);
/* vfs_write, vfs_mkdir, vfs_unlink, vfs_rename, vfs_sync against anything
* under /assets all return VFS_ERR_PERM; there is no upper layer to
* absorb a write into. */
```
## API
The public API is [`include/packfs.h`](include/packfs.h); every function and
struct is documented there with a pointer to the `concept.md` section that
specifies its behavior. In outline:
- `vfs_new` / `vfs_free` — a `Vfs` owns a mount table, nothing else.
- `backend_mem_new` / `backend_dir_new` / `backend_pack_new` /
`backend_overlay_new` — construct a backend; `vfs_mount`/`vfs_unmount`
attach or detach it at a path prefix. `backend_pack_new` mounts a pack
read-only, with no writable upper layer at all — every mutating call
against it returns `VFS_ERR_PERM`; wrap the same pack in
`backend_overlay_new` instead when writability is wanted.
- `vfs_open` / `vfs_read` / `vfs_write` / `vfs_close` — file I/O.
- `vfs_stat` / `vfs_readdir` / `vfs_mkdir` / `vfs_unlink` / `vfs_rename` —
metadata and namespace operations.
- `vfs_sync` — compacts an overlay mount into a fresh pack.
- `vfs_harden_process_with_landlock` — optional, opt-in, process-wide
Landlock confinement to the process's current `dir` mounts. Deliberately
**not** applied automatically by `vfs_mount` (Section 6.2 explains why:
Landlock restrictions are irreversible and process-wide, which would be a
surprising side effect for an embeddable library to trigger on its own).
## Concurrency
Single-writer, wait-free-reader (Section 5): readers never take a lock and
never observe a write in progress; structural changes (create/unlink/rename/
mkdir/mount/unmount) publish a new immutable snapshot via one atomic pointer
swap; ordinary content writes to an already-existing file update a per-entry
cell directly and never touch the snapshot. `mem`-backed buffer growth
never mutates a buffer address a reader might be reading (Section 5.7):
growth always allocates a new buffer and publishes it, never reallocates in
place. `tests/test_concurrency.c` exercises this under concurrent reader and
writer threads. The suite is regularly run under AddressSanitizer/
UndefinedBehaviorSanitizer, both locally and in CI (see
`.gitea/workflows/ci.yml`), and this is not aspirational — ASan caught a
real heap-use-after-free in the snapshot-reclamation logic during
development (see the `reclaim_gate` note in `src/internal.h`).
ThreadSanitizer is configured the same way but, as of this writing, has
not actually completed a run in either environment this project has been
built and tested in so far — the local development sandbox and this
project's own CI runner both block the `personality(ADDR_NO_RANDOMIZE)`
syscall TSan needs to start (see `CONTRIBUTING.md` for the confirming
tests in each case). CI's TSan step is written to not fail the build over
that specific, known-benign failure, which means a green CI run is not
evidence TSan actually executed — stated plainly here rather than left to
be assumed from CI showing green.
## Security
`dir` mounts are capability-scoped (Section 6.2): a mount holds an already-open
directory file descriptor, not a path string, and every lookup beneath it is
resolved with `openat2(RESOLVE_BENEATH | RESOLVE_NO_SYMLINKS)` on Linux 5.6+,
which atomically rejects `..` and symlink escapes in one kernel call. Where
that syscall is unavailable, containment falls back to per-component
`O_NOFOLLOW` resolution — weaker, and documented as such (Section 6.3), not
silently assumed equivalent. Pack files are treated as untrusted input unless
they came from this process's own compaction: every on-disk offset is
bounds-checked and the pack's checksum is verified before any of it is
trusted (Section 7). See `tests/test_dir.c` for a containment regression test
and `tests/test_pack_overlay.c` for a corrupted-pack rejection test. See
[`SECURITY.md`](SECURITY.md) for exactly what is and isn't claimed as a
security boundary, and how to report a vulnerability.
## Contributing
See [`CONTRIBUTING.md`](CONTRIBUTING.md). Read `concept.md` and `CLAUDE.md`
first — they are the project's actual specification and its enforced
documentation standard, respectively, and every design decision in the code
traces back to one of them. [`POSTMORTEM.md`](POSTMORTEM.md) is a detailed,
permanent record of every real bug and environment issue found during this
project's development — what happened, how it was diagnosed, the fix, and
the regression-prevention artifact for each. Check it before re-diagnosing
something that might already be answered there.
## License
MIT — see [`LICENSE`](LICENSE).