Files
packfs/CLAUDE.md
T

106 lines
24 KiB
Markdown
Raw Normal View History

# CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
## Repository status
This repository contains a working v0 implementation of the `concept.md` specification: `include/packfs.h` (public API), `src/*.c` (implementation), `tests/test_*.c` (test suite), and open-source project scaffolding (`README.md`, `LICENSE`, `CONTRIBUTING.md`, `.github/workflows/ci.yml`), alongside the frozen `concept.md` and this file.
## Build, test, and lint commands
```sh
make # builds libpackfs.a and libpackfs.so (zero required third-party deps, Section 11.1)
make test # builds and runs every tests/test_*.c
make clean
make install PREFIX=/some/prefix
```
A single test: `make build/test_mem && ./build/test_mem` (substitute any `test_*` basename). There is no separate lint step — the build itself uses `-Wall -Wextra -Wpedantic`, and a warning introduced by a change is a build failure, not something to leave in place.
Sanitizer builds are not wired into `make test` (they need per-file compilation with sanitizer flags plus `-D_GNU_SOURCE -Iinclude -Isrc`); see `.github/workflows/ci.yml` for the exact invocation, which CI runs on every push. **Any change to `upper.c`, `overlay.c`, or `vfs.c` must be verified under `-fsanitize=address,undefined` and `-fsanitize=thread` before being considered done** — this is not a formality: exactly this process caught a real use-after-free in the snapshot-reclamation logic during initial development (see the `reclaim_gate` note below), which the plain build and even repeated plain test runs never surfaced.
## Code architecture
- `include/packfs.h` — the entire public API.
- `src/internal.h` — every internal type shared across `.c` files; read this first when touching implementation code.
- `src/vfs.c` — the `Vfs` mount table itself (`MountSnapshot`, refcounted, atomically swapped) and the public API's dispatch-by-longest-prefix-match.
- `src/upper.c` — the writable layer shared by `mem`, `dir`, and the overlay's upper side: `UpperSnapshot`/`UpperEntry`/`MutCell`, the persistent-treap index (`TreapNode` and `treap_*`/`node_*` — see "Known performance characteristics" below for why it's a treap and not a flat array), the structural-write functions (`upper_create`/`upper_mkdir`/`upper_remove`/`upper_rename`/`upper_copy_up`), the content-write fast path (`upper_cell_read`/`upper_cell_write`), the range-query API (`upper_visit_range`, used by `readdir` in both this file and `overlay.c`), and the standalone `mem`/`dir` `Backend` (`backend_mem_new`/`backend_dir_new`).
- `src/overlay.c` — composes a `Pack` (lower, read-only) with an `UpperStore` (upper): copy-up, whiteouts, the merged view (`overlay_stat`/`overlay_readdir`), the append journal, and `vfs_sync`'s compaction.
- `src/pack.c` — the on-disk pack format: load-time integrity validation, binary-search lookup/range-query, the compaction writer (atomic rename, exact-duplicate elimination), and the standalone read-only `pack` `Backend` (`backend_pack_new`) for mounting a pack with no writable upper layer at all.
- `src/containment.c` — `dir`-mount path containment (`openat2`/`O_NOFOLLOW` fallback) and opt-in Landlock hardening.
- `src/path.c` — virtual-namespace `.`/`..` canonicalization.
- `src/hash.c` — FNV-1a64, used for pack checksums and compaction dedup.
- `tests/` — one binary per concern (`test_mem`, `test_dir`, `test_pack_overlay`, `test_concurrency`, `test_index_stress` — the treap's correctness under thousands of randomized structural operations, cross-checked against an independent reference model); `test_harness.h` is a small assertion-macro header, not a framework, consistent with the zero-dependency constraint.
**One architectural fact spans every mutable structure and is easy to miss reading any single file in isolation:** `MountSnapshot` (`vfs.c`), `UpperSnapshot` (`upper.c`), and a `mem`-backed `MutCell`'s buffer (`upper.c`, Section 5.7) are all retired through the same two-step pattern — publish the replacement via `atomic_store_explicit(..., memory_order_release)`, then free the superseded object only under a dedicated `reclaim_gate` rwlock's *write* side, while every acquirer takes that same gate's *read* side around its own load-then-increment. This exists because a plain "load a pointer, then atomically increment its refcount" leaves a real gap between those two steps in which a concurrent writer can free the very object being acquired — confirmed by ASan as an actual heap-use-after-free during development, not a theoretical concern. See the comment on `UpperStore.reclaim_gate` in `internal.h` for the full reasoning. Any new refcounted, concurrently-reclaimed structure added to this codebase needs the same gate, not just an atomic pointer and a naive refcount. The one exception, and the reason it's an exception rather than a hole: `TreapNode`'s own per-node refcounting (`upper.c`) does *not* need its own `reclaim_gate`, because nothing ever acquires a `TreapNode*` independently the way `upper_acquire` acquires an `UpperSnapshot*` — a reader only ever reaches tree nodes by walking `s->root` after already holding a valid `UpperSnapshot` reference, which transitively keeps the whole tree alive for the reader's purposes; `node_ref`/`node_unref` are called only by the single writer (under `writer_lock`) and by `upper_release`'s cascade (already inside `reclaim_gate`'s protection at the snapshot level), never by a reader racing a writer the way the snapshot pointer itself is raced.
## What this project is
PackFS is the design for a **statically linked, in-process virtual file system (VFS) written in C**, intended for embedding in sandboxed, agent-driven, TUI, or game-shaped programs. `concept.md` is the full specification (architecture, write model, concurrency model, path-containment model, pack-integrity validation, C object-model sketch, on-disk pack format, explicit v0 exclusions, and a self-evaluation) — read it in full before doing any implementation work here, rather than relying on the summary below.
## Documentation standard (enforced)
Every document in this repository — this file included — is written and revised under the following rule. It was established after `concept.md` was rewritten into its current form, and that rewrite is the reference example: when in doubt about whether a piece of prose conforms, compare it against `concept.md`.
- **Register.** Formal, precise, third-person or directly-stated-constraint prose. No colloquialisms, no rhetorical asides, no filler sections ("Tips," "Support," "Common Tasks") invented for their own sake.
- **Structure scales to content, not the other way around.** A specification-length document (`concept.md`) earns full academic structure — Abstract, Problem Statement, numbered body sections, Evaluation, Conclusion, References. A short instructions file (this one) states its rules directly without manufacturing sections it doesn't need.
- **Constraints are labeled by strength.** A requirement an implementation must satisfy is stated as such explicitly ("hard constraint," "must," "required"). A recommendation that could reasonably be revisited is stated as such ("recommended," "a legitimate simpler fallback"). The two are never left ambiguous relative to each other.
- **Revisions never delete information.** An error, once caught, is corrected and the correction is explained in place; prior content is extended or superseded, not silently dropped. This applies to every living document in this repository; `concept.md` is the one exception, having been declared frozen rather than living — see "`concept.md` is immutable" below.
- **External claims are cited, not asserted.** Where a decision leans on prior art or an established system (e.g., LMDB, OverlayFS), it is attributed with a references entry rather than presented as self-evident.
- **Every specification-length document carries its own self-evaluation.** A short methodology note, an assessment table scored per relevant category, and one final overall grade with justification — see `concept.md` Section 12 for the pattern. This file's own self-evaluation is at the bottom.
## `concept.md` is immutable
As of this revision, `concept.md` is frozen. It is not to be edited, rewritten, appended to, or otherwise modified by any future session, for any reason — including to fix a typo, to close a newly-found gap, or to keep it "in sync" with a later decision. It stands as the fixed record of the design as reasoned through its own research and review cycles; the self-evaluation and grade inside it describe that fixed text, and would themselves become inaccurate if the text under them kept moving.
This is a deliberate exception to the "revisions never delete information" pattern above, not a stricter version of it: that pattern describes how a living document is corrected in place; a frozen document is not corrected in place at all. The correct response to a newly-found gap, error, or improvement in the design is to open a **new, separate document** (for example, an addendum or a versioned successor) that amends or supersedes `concept.md`, leaving its text untouched. If such a document is ever created, this file should be updated to point to it alongside `concept.md`, and to state plainly which of the two is authoritative where they overlap.
The load-bearing constraints summarized below are therefore also frozen in substance: they describe what `concept.md` says, not a moving target. A future change to the actual design changes what is authoritative — a new document — not what `concept.md` says happened.
## Project decisions that supersede concept.md
`concept.md` cannot be edited (above), but the project's actual scope has diverged from it in one place, by an explicit, permanent decision — recorded here per this file's own protocol for citing an amendment alongside the frozen text it overrides:
- **The `zip` and `tar` import/export backends will never be built.** `concept.md` Section 3.1 lists them as candidate backends and Section 11 recommends them ("Zip via miniz as import/export. Tar as snapshot/export"), with Section 11.1 describing a compile-time switch at the `zip`/`tar` backend boundary for `miniz`/`libarchive`. That recommendation is superseded: this project supports the pack format only. Do not propose, scaffold, stub, or partially implement a `zip` or `tar` backend, and do not add `miniz` or `libarchive` as a dependency for any reason. `backend_overlay_new` reads and writes this project's own pack format exclusively. The rest of Section 11's recommendation (core + `mem` + `dir` + custom pack + overlay + compaction) stands as written and is what this codebase implements.
## Load-bearing constraints from concept.md
These are hard requirements stated in the spec, not stylistic suggestions — any implementation work must respect them:
- **Static linking is enforced, not aspirational.** The core (`vfs_*`, `mem`, `dir`, `pack`, overlay, compaction) must build with zero required third-party libraries and zero dynamic loading. `concept.md` describes `miniz`/`libarchive` sitting behind a compile-time switch at the `zip`/`tar` backend boundary as an *optional* addition that must never be required — but per "Project decisions that supersede concept.md" above, that boundary will never actually be built, so this constraint is satisfied trivially: there is no optional dependency at all, required or otherwise.
- **Archive formats (zip, tar) are, and will remain, entirely absent from this codebase — not merely import/export skins.** The system of record is, and is the only supported format, a custom indexed **pack** file (immutable image: header + index + blobs + strings). See "Project decisions that supersede concept.md" above.
- **Writability comes from a copy-on-write overlay, not from mutating the pack.** Pack mounted read-only at a path; `mem` or `dir` mounted as the writable upper layer; `open(O_RDWR)` triggers copy-up. Deleting a lower-layer (pack) entry requires writing a **whiteout** tombstone in the upper layer — there is no way to delete a pack entry directly, and skipping the whiteout leaves deletion of pack-originated files undefined.
- **Compaction and the journal must be crash-safe by construction:** compaction writes `pack.img.tmp`, `fsync`s, then atomically `rename()`s over `pack.img` (never in place); journal records are length/checksum self-describing so replay stops cleanly at the first torn record; journal "truncation" after compaction means writing the surviving tail to a new file and renaming it over the journal, since a file's front cannot be truncated in place.
- **Concurrency model is single-writer, wait-free-reader (LMDB-style MVCC), not a global lock.** Live state — upper index, whiteout set, **and the mount table** — is an immutable, refcounted `VfsSnapshot` reached through one atomic pointer; readers load it once (wait-free) and never observe an in-progress write; exactly one writer at a time builds the next snapshot copy-on-write and publishes it with a single atomic pointer swap using release/acquire ordering (getting this ordering wrong is a real C11 data race). `vfs_mount`/`vfs_unmount` are structural writes like any other — the mount table is not separate global state exempt from this model. **Structural writes (create/unlink/rename/mkdir/whiteout/mount/unmount) publish a new snapshot; content writes (bytes into an already-existing file) do not** — they update a per-entry size/mtime cell under a per-entry lock instead, so an ordinary `vfs_write` never requires copying and republishing the whole index. A `mem`-backed write that grows a file's buffer past its current allocation must **allocate a new buffer and publish it with a release store**, never realloc the existing address in place — a concurrent reader may otherwise copy out of freed memory, not just a stale value; the old buffer is retired under the same reclamation discipline as a retired snapshot. Compaction is just another writer against a frozen snapshot plus a recorded journal watermark; if writing `pack.img.tmp` or its `fsync` fails, compaction aborts and the existing pack + journal are untouched. Known, explicitly-accepted gaps: the per-entry lock covers `dir`-upper only within this process, not against a second uncoordinated process writing the same host directory; rename-over-a-mapped-pack is safe on POSIX but not guaranteed on Windows; naive refcounting contends under high read concurrency (hazard pointers are the recommended upgrade path, not epoch-based reclamation, because this system's workloads can plausibly stall a reader thread); cross-process concurrency has a specified LMDB-style reader-table design but is not implemented in v0.
- **Path containment for `dir` mounts is a hard requirement, not best-effort.** A `dir` mount is an already-open directory file descriptor (capability-scoped, WASI-style), not a re-resolved path string. On Linux 5.6+, every lookup under it uses `openat2(RESOLVE_BENEATH | RESOLVE_NO_SYMLINKS)` to atomically block `..` escapes, symlink escapes, and the canonicalize-then-open TOCTOU window in one kernel call; Landlock (5.13+) is a recommended additional layer. On other platforms, no equivalent atomic primitive is assumed — the fallback (`O_NOFOLLOW` per component + canonicalize-and-verify) is an accepted, explicitly weaker residual-risk posture, not a claim of parity. The VFS core additionally canonicalizes `.`/`..` in the virtual namespace itself, before any backend is reached, so a crafted path cannot escape a mount's prefix even when no host directory is involved.
- **A pack file is untrusted input unless it came from this run's own compaction.** Every index entry's `data_off + size` and `name_off` must be bounds-checked against the actual file/strings-region length before being trusted by `vfs_open`/`vfs_stat`/`vfs_readdir`; a failed check invalidates the whole pack load, it is not skipped silently. An externally-sourced pack additionally needs a checksum verified at load, the same self-checking-record principle already required of the journal.
- **The on-disk and in-memory indexes are kept sorted by full path**, not insertion order, so exact lookups are a binary search and `vfs_readdir` is a bounded range query (two binary searches), not a linear scan.
- **An empty directory needs an explicit index entry.** `vfs_mkdir` writes a zero-size entry with a reserved directory-marker bit in `mode`, otherwise a directory with no files in it has nothing referencing its path and vanishes across compaction. The marker is superseded (not deleted) once any file exists under that path.
- **`mode` is stored/restored but not enforced.** Permission bits, symlinks, and hard links have no access-control or link semantics in v0 — `mode` round-trips through compaction (including the directory-marker bit above) but nothing in the VFS core interprets it as a security boundary.
- **Explicitly out of scope for v0:** FUSE as the first backend, invoking libarchive per-`read()`, full POSIX semantics/locks/sockets/mmap of virtual files, any single format trying to double as initrd + game pak + user home, implementing the cross-process reader-table design (specified, not built), content-defined chunking/delta compression in the pack format (only exact-duplicate elimination during compaction is in scope), and enforced permissions/symlinks/hard links.
`concept.md` will not be updated again (see "`concept.md` is immutable" above), so the constraints above will not drift out from under it. If a future document amends or supersedes part of the design, update this section to cite that document alongside `concept.md` — without editing `concept.md` itself.
## Known performance characteristics
`bench/bench.c` (`make bench`) measures PackFS against the host filesystem; full results and analysis are in `BENCH.md`. The one finding here that future work on this codebase needs to know without re-running the benchmark:
- **RESOLVED: bulk sequential `create`/`unlink`/`mkdir` on `mem`/`dir` was O(n²) in the number of entries — confirmed empirically (`BENCH.md`), then fixed, not just worked around.** The index (`UpperSnapshot`) was a flat sorted array; every structural write copied it in full before publishing the next snapshot (Section 5.3's single-writer model), and at 20,000 sequential creates this measured ~9–22x slower than the raw host filesystem. `concept.md` Section 5.3 names the exact trigger for reconsidering that design — "a persistent (structurally shared) tree structure is not required until this assumption is empirically violated" — and that trigger was hit. The index is now a **persistent treap** (`src/upper.c`, `struct TreapNode` and the functions around it — see the citations in that comment, Seidel & Aragon 1996 and Liljenzin arXiv:1301.3388, for why a treap and not a persistent AVL/red-black/weight-balanced tree): a structural write now copies only the O(log n) nodes on the path to the change. Post-fix, the same benchmark shows `mem` create/unlink/mkdir 65–330x faster than before, and — the more important comparison — 15–27x *faster* than raw fs where it used to be 9–22x *slower*; the concurrent mixed workload went from 6x slower than raw fs to 22x faster. `BENCH.md`'s "Resolution" section has the full before/after table and the complexity-class re-confirmation (the same `mkdir`-at-N=4,000-vs-`create`-at-N=20,000 methodology that found O(n²) now shows a ratio consistent with O(n log n), not O(n²)). A new test, `tests/test_index_stress.c`, cross-checks the treap's correctness under thousands of randomized (non-sequential) creates/deletes/renames against an independent reference model — read it, not just the benchmark, before touching `snapshot_upsert`/`snapshot_remove`/the treap functions again.
- **`dir`-backend `stat` costs roughly 2x a raw `stat()` call**, because `pfs_dir_statat` (containment, Section 6.2) resolves and opens the path, then `fstat`s the fd, where raw `stat()` is one syscall. Expected, and the same containment cost `open` already pays — recorded so it isn't mistaken for a regression if someone benchmarks it again later.
- **`mem`-backed large sequential writes lose to a plain unsynced host `write()` above roughly 16 MB**, because Section 5.7's buffer-growth discipline (allocate new, copy everything, publish, retire old) re-copies previously-written bytes on every capacity doubling, where the host page cache only ever appends new pages. This is the measured cost of the safety property Section 5.7 requires (no reader ever sees a freed buffer), not an accidental inefficiency.
## Self-evaluation
**Methodology.** This file was checked against the repository's actual state (`make test` passing, including a new `test_index_stress`; `include/`, `src/`, `tests/`, `bench/` present and matching the description below; confirmed by directory listing and a live build), against `concept.md` as frozen (for factual consistency of the constraints and architecture summarized above), and against the documentation standard stated in this file. This revision followed a rewrite of `upper.c`'s index structure (flat array to persistent treap) undertaken specifically to close the O(n²) finding this file's "Known performance characteristics" section previously reported as an open, unresolved limitation — the rewrite was verified under `make test`, `test_index_stress`'s randomized correctness check, ASan/UBSan (per the sanitizer rule below), and a full before/after `make bench` run, not merely believed correct because the code compiled.
| Category | Grade | Notes |
|---|---|---|
| Factual accuracy | A | Build/test commands and the code map were verified against a live `make test` run and the actual file layout, not written from memory of intent; the performance-characteristics section was updated from a live `make bench` re-run, not assumed fixed because the rewrite was intended to fix it. |
| Adherence to the documentation standard | A | Direct, constraint-labeled prose; no manufactured sections; scaled appropriately to an instructions file rather than imitating `concept.md`'s full academic structure. |
| Completeness for its purpose | A | Covers repository status, build/test/lint commands, a per-file code map, every hard constraint from `concept.md` relevant to implementation work, the permanent zip/tar exclusion, the cross-cutting `reclaim_gate` pattern (including why `TreapNode` is a deliberate exception to it, not a hole in it), and the now-resolved O(n²) bulk-write characteristic with a pointer to its fix and verification. |
| Avoidance of generic or invented content | A | No fabricated "Common Development Tasks" or "Tips" sections; the sanitizer-testing instruction is stated as a requirement precisely because skipping it once already let a real bug through, not as generic advice. |
**Overall grade: A.** The file states only what is verifiably true of the repository, the spec, and the implementation; labels constraints by strength; documents the architectural pattern (snapshot reclamation via `reclaim_gate`) that spans multiple files; and — unlike the previous revision, which correctly recorded a real limitation but left it as a known gap — now records that limitation's actual, measured resolution, with the same empirical rigor (a complexity-class re-confirmation, not just "it got faster") that found it in the first place.
**A real gap this audit found and fixed, not just documented:** `backend_pack_new` was declared in `include/packfs.h` and referenced in this file's own architecture map, but was never implemented in `src/pack.c` — any program calling it would fail at link time. It is now implemented (a standalone, read-only `pack` `Backend`; every mutating operation returns `VFS_ERR_PERM`), covered by a new case in `tests/test_pack_overlay.c`, and verified under `-fsanitize=undefined` per this file's own testing rule. This is recorded here because it is exactly the kind of error "document literally all" is supposed to catch: a documentation pass that describes an unimplemented function accurately is still wrong in a way that matters.