Implements the core design: a mount table published as an atomically- swapped snapshot; mem/dir/pack/overlay backends; copy-on-write overlay with copy-up and whiteout deletion; a checksummed append journal; compaction with exact-duplicate elimination; single-writer/wait-free- reader concurrency with a structural/content write split; openat2/ Landlock path containment for dir mounts; and load-time pack integrity validation. Zero required third-party dependencies. Sanitizer testing (ASan/UBSan) caught and led to fixing a genuine heap-use-after-free in the snapshot-reclamation path: the textbook "load pointer, then increment its refcount" pattern left a gap a concurrent writer could free through. Closed with a small reclaim_gate rwlock, documented in internal.h and CLAUDE.md since it's a pattern every refcounted structure in the codebase now follows. zip/tar import/export backends, recommended in concept.md Section 11, will not be built — a permanent project decision recorded in CLAUDE.md since concept.md itself is frozen and cannot be edited to reflect it. Includes a runnable demo (examples/demo.c, `make demo`) exercising the library end to end and proving cross-run persistence through the pack file, plus open-source scaffolding: MIT license, README, CONTRIBUTING, and a CI workflow running the test suite under ASan/UBSan/TSan. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
306 lines
40 KiB
Markdown
306 lines
40 KiB
Markdown
# PackFS: Design Rationale and Evaluation of a Statically-Linked, In-Process Virtual File System
|
|
|
|
## Abstract
|
|
|
|
This document specifies the architecture of an in-process virtual file system (VFS) intended for embedding in sandboxed, agent-driven, TUI, or game-shaped C programs. The central claim is that the design problem is not the selection of an archive format ("tar versus zip"), but the construction of an **in-process, statically linked, C-owned path tree** with inexpensive mount points and a copy-on-write overlay. Archive formats (zip, tar) are relegated to the role of an **image backend** — a serialization detail — rather than serving as the filesystem's primary storage representation. The document defines the system's architecture, backend selection rationale, write model, concurrency model, path-containment model, pack-integrity validation procedure, object model, on-disk format, and explicit exclusions from the initial version (v0). This revision extends an earlier version of the specification by adding a path-containment model for host-backed mounts (Section 6) and an on-load integrity-validation procedure for pack files (Section 7) — both previously unaddressed despite direct relevance to the project's stated sandbox use case — and by refining the concurrency model to distinguish structural from content mutation (Section 5.3) and to specify, rather than merely defer, a cross-process design (Section 5.6). The present revision closes three further gaps found on re-examination: the mount table was previously outside the snapshot the concurrency model protects (Section 5.3); buffer growth on the `mem` backend could race a concurrent reader into a use-after-free rather than merely a stale read (Section 5.7); and the pack format had no representation for an empty directory (Section 9.3). It concludes with a self-evaluation across several categories, informed by comparison against established prior art (LMDB, OverlayFS, WASI, Landlock, and the memory-reclamation literature).
|
|
|
|
## 1. Problem Statement
|
|
|
|
The target use cases (sandboxes, TUIs, agents, games, chroot-shaped programs) share the following constraints: small, explicit C; no surprise runtime; no mandatory dynamic dependencies; predictable behavior under crash and concurrent access. These constraints rule out treating a mutable archive format (zip or tar) as the system of record, since neither format was designed for in-place, random-access mutation. The design instead separates **path resolution and mutability** (the VFS core) from **serialization** (the backend), such that the choice of archive format becomes an implementation detail confined to a single backend rather than a constraint on the whole system.
|
|
|
|
### 1.1 Threat Model
|
|
|
|
Two categories of untrusted input are treated as first-class concerns rather than assumed away:
|
|
|
|
- **A pack file may be corrupted or actively adversarial.** It may have been imported through the zip compatibility path, supplied by another party, or read back after a crash. Its on-disk offsets must not be trusted without validation before use (Section 7).
|
|
- **A `dir`-backed mount exposes a real host directory**, and path resolution into it must not be escapable via `..` components, symlinks, or a race between validation and use — regardless of whether the untrusted input is a virtual path supplied by an agent or a symlink planted by a concurrent host-side process (Section 6).
|
|
|
|
Both properties are treated as hard requirements for any mount reachable by untrusted input, and as reasonable-effort hardening elsewhere.
|
|
|
|
## 2. Architectural Overview
|
|
|
|
### 2.1 System Shape
|
|
|
|
```text
|
|
paths
|
|
→ VFS core (mount table, cwd, overlay, fd table)
|
|
→ backends
|
|
mem live, writable
|
|
dir real OS directory
|
|
pack one file, indexed, read-only
|
|
archive import/export only (zip or tar)
|
|
```
|
|
|
|
### 2.2 Invariants (Core Rules)
|
|
|
|
- A single path namespace is exposed (e.g., `/assets/foo`, `/tmp/bar`, `/proc/...` if process-table emulation is desired).
|
|
- Mounts are resolved by longest-prefix match to a backend.
|
|
- Writes never mutate a pack file in place.
|
|
- A writable mount is either `mem` or `dir`, or a copy-on-write overlay sitting atop a read-only pack.
|
|
- Pack files are treated as **images**: they are loaded and compacted, not treated as block devices.
|
|
|
|
Under this model, the choice between zip and tar is a detail confined to the `pack` backend, not a property of the system as a whole.
|
|
|
|
## 3. Backend Selection
|
|
|
|
### 3.1 Candidate Backends
|
|
|
|
| Backend | Role | Use |
|
|
|---|---|---|
|
|
| **mem** | default writable layer | maps, arenas, explicit `new`/`free` |
|
|
| **dir** | escape hatch to the host | optional, easy to disable in a sandbox |
|
|
| **pack** | shipped tree in one file | custom table + blobs, mmap if desired |
|
|
| **zip** | interchange / tools | miniz, static, import into pack or mount read-only |
|
|
| **tar** | interchange / snapshots | hand-rolled USTAR *or* libarchive for export |
|
|
|
|
### 3.2 Selection Rationale
|
|
|
|
The recommended pack format is a custom indexed structure, not tar and not zip used as the live format, for the following reasons:
|
|
|
|
- The design already models resources as explicit types with constructors and destructors (`type T { fields; new(); free(); }`); a custom format is consistent with this style and keeps the pack builder under the project's own control.
|
|
- Random access is a flat table lookup, not a linear scan or a directory-parse-then-seek. Tar has no index at all — locating a member requires reading headers sequentially until the target is reached. Zip does provide random access via its central directory, but this is a second, more complex format to parse, and it was not designed to support in-place edit or delete; removing or resizing an entry still requires rewriting the archive.
|
|
- Static linkage reduces to compiling the project's own `.c` files, with no dependency on an archive library for the hot path.
|
|
- Compression is optional and can be applied per file, rather than as a single transform over the whole archive.
|
|
- The overlay-and-compact procedure is straightforward: the pack is rewritten from the merged (upper + lower) tree.
|
|
|
|
### 3.3 Role Separation: Compatibility versus Storage
|
|
|
|
Zip functions as a **compatibility skin** (accepting archives produced by other tools). Tar functions as a **snapshot skin** (e.g., `dump /home → blob`). Neither is suitable as the inode layer of the live filesystem.
|
|
|
|
## 4. The Write Model (Copy-on-Write Overlay)
|
|
|
|
Archive formats (zip, tar) must not be made writable in place. Instead, the system adopts a copy-on-write overlay model analogous to OverlayFS, without a kernel component.
|
|
|
|
### 4.1 Procedure
|
|
|
|
1. Mount the pack at `/`, read-only.
|
|
2. Mount `mem` (or a host directory) as the upper layer at the same root.
|
|
3. `open(O_RDWR)` triggers **copy-up**: the pack member is copied into the upper layer before it is modified.
|
|
4. `unlink` on a path that exists only in the upper layer removes it from the upper layer directly. `unlink` on a path that exists in the lower (pack) layer cannot mutate the pack; instead, it writes a **whiteout** — a tombstone marker — into the upper layer. The merge step treats a whiteout as "path does not exist," even though the lower entry remains physically present in the pack. `rename` is defined as copy-up-if-needed, followed by a write at the new path and a whiteout at the old path.
|
|
5. `vfs_sync()`, or an equivalent tool, walks the merged view (upper layer takes precedence; whiteouts hide lower entries; compaction subsequently discards both the whiteout and the entry it hides) and writes a **new pack**.
|
|
|
|
### 4.2 Deletion Semantics: The Whiteout Mechanism
|
|
|
|
The whiteout is the mechanism that makes deletion of a read-only, lower-layer entry well-defined; it is not an optional implementation detail. Without it, "delete a file that originated in the read-only pack" has no coherent semantics, since the pack cannot be mutated and the upper layer initially contains nothing corresponding to that path.
|
|
|
|
### 4.3 Durability Requirements
|
|
|
|
The model above is honest under crash, mmap, and sandboxing only if compaction and the journal are independently made crash-safe:
|
|
|
|
- **Compaction must be atomic.** The new pack is written to `pack.img.tmp`, `fsync`ed, and then moved into place via `rename()` over `pack.img`. Same-filesystem rename is atomic, so a crash during compaction leaves either the old pack or the new pack intact, never a partially written file. The new pack must never be written in place. If writing `pack.img.tmp` or its `fsync` fails for any reason (disk full, I/O error), compaction is aborted before the rename: `pack.img` and the full journal are untouched and remain the valid state, and nothing has been partially applied.
|
|
- **Journal records must be self-checking.** Each record carries a length prefix (or fixed size) and a checksum. On replay, the first record that fails to parse fully or fails its checksum terminates replay; everything after that point is treated as never committed. Without this, a crash during an append can corrupt the interpretation of the entire replay, not merely the last write.
|
|
|
|
### 4.4 Append Journal for Key/Value Workloads
|
|
|
|
For workloads that are primarily key/value in character (saves, agent memory, configuration), a small append-only journal alongside the pack suffices:
|
|
|
|
```text
|
|
pack.img immutable
|
|
pack.img.jnl append records (put/delete)
|
|
```
|
|
|
|
On open: load the pack index, then replay the journal. On compaction: emit a new pack, then replace the journal with only the records written **after** the watermark recorded at the start of compaction (see Section 5) — not "whatever remains at completion time," and not an in-place truncation (justified in Section 5.3). See Section 5.3 for how an ordinary content write (bytes into an already-copied-up file) avoids re-publishing a snapshot on every call.
|
|
|
|
## 5. Concurrency Model
|
|
|
|
### 5.1 Design Goal
|
|
|
|
Concurrent access is a first-class requirement, not an afterthought. Two failure modes are common in ad hoc designs: omitting a concurrency model entirely, or defaulting to a single global lock that serializes all readers against all writers indefinitely. The design below avoids both by adopting a documented, well-understood concurrency pattern rather than an invented one.
|
|
|
|
### 5.2 Prior Art
|
|
|
|
- **LMDB** achieves serializable isolation using a single writer combined with copy-on-write shadow paging: exactly one write transaction is active at a time; the number of concurrent readers is unbounded; readers are wait-free and block neither the writer nor each other, because each reader walks an older, still-valid, immutable page tree while the writer constructs new pages elsewhere [1].
|
|
- **OverlayFS** demonstrates that copy-up is a genuine race condition, not a formality: CVE-2023-0386 was a container-escape vulnerability caused by an unsynchronized copy-up operation; the corresponding kernel fix serializes copy-up under an inode lock [2, 3].
|
|
|
|
### 5.3 Snapshot-Based Single-Writer, Wait-Free-Reader Design
|
|
|
|
The system adopts LMDB's structural pattern rather than a single lock over the entire tree:
|
|
|
|
- The live VFS state (upper-layer index and whiteout set) is represented as an **immutable, reference-counted snapshot** (`VfsSnapshot`). Exactly one atomic pointer, `current_snapshot`, is the sole entry point for all reads.
|
|
- **The mount table is part of the same snapshot, not separate global state.** Path resolution (Section 2.2) depends on the mount table exactly as much as it depends on the upper index, so `vfs_mount` and `vfs_unmount` are structural writes like any other, going through the identical single-writer-and-swap path described below rather than a mechanism of their own. This is a deliberate choice, not an oversight left over from treating mounts as fixed at startup: folding the mount table into the snapshot costs nothing at the size mount tables actually reach in practice, and it closes what would otherwise be an entirely separate, unaddressed race between path resolution and a concurrent `vfs_mount`/`vfs_unmount` call.
|
|
- `vfs_open`, `vfs_stat`, and `vfs_readdir` load this pointer once (a single atomic load plus a reference-count increment — wait-free, lock-free), then operate against that frozen view for the remainder of the call. A write in progress is never observed, because an in-progress write has not yet published a new snapshot.
|
|
- Exactly one writer executes at a time; a plain mutex suffices for the in-process case. The writer constructs the next snapshot as a copy-on-write derivative of the current one. At the scale anticipated for agent- or sandbox-oriented workloads, the index is small enough that a full copy of the index and whiteout set per structural write is acceptable; a persistent (structurally shared) tree structure is not required until this assumption is empirically violated. The writer then performs **one atomic pointer swap** to publish the new snapshot — a release store on publication, paired with an acquire load on read. Incorrect memory ordering here constitutes an actual data race under the C11 memory model, not merely a theoretical concern: a reader could observe the new pointer value before observing the writes that constructed the object it references.
|
|
- A snapshot is freed only once its reference count reaches zero — that is, once every reader that acquired it has released it. This is the same grace-period discipline used in read-copy-update (RCU), implemented here via plain reference counting rather than epoch or quiescent-state tracking; Section 5.4 discusses the scaling cost of that choice and the recommended upgrade path.
|
|
- **Copy-up is rendered race-free as a consequence of this structure**, not by additional locking: the full blob is written into upper storage first (a new `mem` block, or write-then-rename under `dir`), and only afterward is its index entry installed in the *next* snapshot. There is no window in which a reader can observe a partially copied file, because readers never observe uncommitted snapshots. No inode lock is required; the transaction boundary performs the equivalent function.
|
|
- **Structural mutation is distinguished from content mutation to avoid unnecessary copying.** A structural mutation — create, unlink, rename, mkdir, whiteout — changes which paths exist, and therefore requires building and publishing a new snapshot as described above. A content mutation — an ordinary `vfs_write` to a path that already has an index entry in the current snapshot — changes only that entry's bytes, size, and mtime, none of which affects which paths exist. Routing every content mutation through the structural path would mean copying and republishing the entire index on every `vfs_write` call, which is wasteful at any index size worth mentioning. Instead, each index entry owns a small mutable metadata cell (size, mtime), updated with plain atomic stores and guarded by a per-entry lock that serializes concurrent writers to the *same* file; this cell's identity is stable across snapshots that do not structurally touch that entry, so publishing a new snapshot is not required to make a content write visible. `vfs_stat` and `vfs_read` observe the cell with an acquire load, so a size update becomes visible to a reader as soon as the writer's release store completes, without waiting for or triggering a snapshot swap.
|
|
- **Compaction is treated as an ordinary writer.** It acquires the current snapshot (a free, wait-free operation), records the journal offset at that instant (the watermark), walks the frozen snapshot to construct the new pack, writes `pack.img.tmp`, calls `fsync`, and renames it over `pack.img`. The live writer continues to append journal records and publish snapshots throughout this process; neither operation blocks the other. Afterward, all journal entries at or before the watermark are dropped and all entries after it are retained. Because a file cannot be truncated from its front in place, this is implemented by writing the surviving tail to a new file and renaming it over the journal — the same atomic-rename technique used for the pack itself.
|
|
|
|
### 5.4 Known Limitations
|
|
|
|
- **Naive per-snapshot reference counting contends under high read concurrency.** Every `vfs_open`/close pair increments and decrements one shared atomic counter; at high thread counts this cache line becomes a bottleneck, and reference counting is documented in the memory-reclamation literature as the weakest-scaling of the standard schemes precisely because of this shared, contended counter [4]. Two alternatives trade this differently: epoch-based reclamation (the mechanism behind the Linux kernel's own RCU) replaces the shared counter with per-thread epoch counters, giving O(1) overhead per read at the cost of unbounded memory growth if a single reader thread stalls and never advances its epoch [4]; hazard pointers bound memory strictly and remain fully lock-free, at the cost of a per-access memory barrier that is more expensive than an epoch counter under low contention [4]. Given that this system's target workloads (agent and sandbox processes) can plausibly stall or hang a reader thread without that being a defect in the VFS itself, the unbounded-growth failure mode of epoch-based reclamation is judged the worse risk of the two. **Hazard pointers are the recommended upgrade path** if reference-count contention is ever measured to matter; plain reference counting remains the v0 default because it is simplest to implement correctly and adequate at the concurrency levels this system targets.
|
|
- **The per-entry lock in Section 5.3 closes the within-process version of the `dir`-upper gap, not the cross-process version.** It does not extend to a second, independent process or tool writing into the same host directory outside this VFS instance's knowledge, which remains genuinely unsynchronized — the snapshot pointer provides a consistent, wait-free view of the index and whiteout set regardless of upper backend, but the bytes on disk under `dir` upper are still subject to whatever concurrency guarantees the host filesystem itself provides once a second, uncoordinated writer is involved.
|
|
- **Rename-over-a-mapped-file is platform-dependent.** Replacing a memory-mapped pack via `rename()` is safe on POSIX systems, where an existing mapping remains valid until the last reader unmaps it. On Windows, this can fail outright, since a file that is still open or mapped generally cannot be renamed over. Where Windows is a genuine target, the mitigations are either to close and reopen the mapping around compaction, or to use a generation-numbered filename (e.g., `pack.img.3`) together with a small pointer file, rather than renaming onto a fixed name.
|
|
- **Cross-process concurrency has a specified design (Section 5.6) but remains unimplemented in v0.** It is no longer an unspecified gap, only a deferred one.
|
|
|
|
### 5.5 Rejected Alternatives
|
|
|
|
- **A reader/writer lock as the primary mechanism** is a legitimate, simpler v0 fallback if correctness with minimal implementation effort is prioritized over throughput (readers block briefly during writes and compaction, which is acceptable if writes are infrequent). It is presented here as a deliberate downgrade from wait-free reads, not a different tier of correctness, and may be an appropriate starting point.
|
|
- **A seqlock** is well suited to a single fixed-size, frequently read field (e.g., a generation counter in the pack header), but not to a variable-size index or directory tree, where safely detecting a torn read is substantially harder. The snapshot pointer already provides the whole-structure equivalent of what a seqlock provides for a single field, so introducing a second mechanism is unnecessary.
|
|
|
|
### 5.6 Cross-Process Concurrency (Deferred Design)
|
|
|
|
Should a future version need two operating-system processes to share one pack file concurrently, LMDB's reader-table design is the specified template, not an invented one [5]:
|
|
|
|
- A small, fixed-size lock file, memory-mapped and shared across processes, holds a table of reader slots, each cache-line-aligned to avoid false sharing between readers running on different cores [5].
|
|
- A reader acquires a slot — guarded by a short-held mutex used only to find a free slot, not held for the duration of the read — and records the snapshot identifier it is using. The read itself then proceeds without holding any lock, mirroring the in-process design in Section 5.3.
|
|
- Before reclaiming or overwriting an old pack generation, the single cross-process writer scans the slot table for the oldest snapshot identifier still recorded by any live reader, and reclaims nothing newer than that [5].
|
|
- This mechanism is additive to, not a replacement for, the in-process atomic-pointer design: within one process, the atomic `current_snapshot` pointer remains the fast path; the reader table exists only so that a writer in a different process can learn the oldest in-use snapshot across process boundaries.
|
|
|
|
This section specifies the mechanism in enough detail to implement; Section 10 continues to treat the implementation itself as out of scope for v0.
|
|
|
|
### 5.7 Byte-Buffer Growth Under Concurrent Read (`mem` Backend)
|
|
|
|
The content-mutation path in Section 5.3 accounts for updating an entry's size and mtime under a per-entry lock, but a `vfs_write` that extends a `mem`-backed file past its current allocation must grow the underlying byte buffer itself, which is a distinct hazard the metadata cell alone does not cover: if the buffer is grown by reallocating its existing address in place, a concurrent `vfs_read` copying out of that address can be left reading freed memory, not merely a stale value — a use-after-free, not just a torn read.
|
|
|
|
The write path must not mutate the existing buffer's address when growing it. Instead, growth follows the same replace-don't-mutate principle already used for the whole-tree snapshot [1, 9]: a new, larger buffer is allocated, the existing bytes plus the newly written bytes are copied into it, and the entry's data pointer is updated with a single release store; a reader's `vfs_read` loads that pointer with an acquire load before copying out of it, so it either sees the old buffer in full or the new one in full, never a torn transition between the two. The old buffer is retired under the same reclamation discipline as a retired snapshot (Section 5.4) — freed once no in-flight reader can still be using it — rather than freed immediately at the point of growth.
|
|
|
|
A write that does not grow the buffer — an in-place overwrite of already-allocated bytes — is not given a stronger guarantee than POSIX itself provides for concurrent, unsynchronized `write()` calls to the same file: the per-entry lock in Section 5.3 serializes concurrent *writers* against each other, but a reader racing an in-place overwrite may observe any interleaving of old and new bytes, not corruption or a use-after-free. Strengthening this further, to a torn-read-free guarantee on every write regardless of growth, is not undertaken here; it would mean copy-on-write at the granularity of every write's byte range rather than only at the granularity of buffer growth, which is materially more machinery for a guarantee POSIX callers do not otherwise expect.
|
|
|
|
## 6. Path Containment and Capability-Scoped Mounts
|
|
|
|
Path containment was absent from earlier revisions of this document despite being directly relevant to the project's stated sandbox use case. It is addressed here as a first-class design concern rather than an implementation afterthought.
|
|
|
|
### 6.1 Virtual-Namespace Canonicalization
|
|
|
|
Before a path reaches any backend, the VFS core resolves `.` and `..` components purely lexically within the virtual namespace and rejects a path that would resolve above the root of the mount it targets. This applies uniformly to `mem`, `pack`, and `dir` backends: even where no host filesystem is involved, an unvalidated `..` could otherwise be used to address a sibling mount's namespace through a mount point that should not expose it.
|
|
|
|
### 6.2 Capability-Scoped `dir` Mounts
|
|
|
|
A `dir` mount is represented internally as an already-open directory file descriptor — obtained once, at mount time, via `open(path, O_DIRECTORY)` — rather than as a path string re-resolved on every access. This is the same capability-based posture WASI uses for its preopened directories, where a module is handed a directory descriptor at startup and has no ambient authority to resolve paths outside it [6]. Every subsequent lookup under that mount is performed relative to that descriptor using `openat2()` with `RESOLVE_BENEATH | RESOLVE_NO_SYMLINKS` on Linux 5.6 and later: the kernel atomically rejects `..` escapes and symlink escapes (including through `/proc` magic links) as a single lookup operation, rather than as a canonicalize-then-open pair, which removes the time-of-check-to-time-of-use window that a symlink swapped in between validation and use would otherwise open [7]. Where Landlock (Linux 5.13+) is available, applying a ruleset scoped to the mount's directory descriptor is recommended as an additional, independent enforcement layer beneath the VFS's own checks, consistent with Landlock's own stacking model, in which a sandboxed thread can only access a path that all of its enforced policy layers grant [8].
|
|
|
|
### 6.3 Platform Coverage and Residual Risk
|
|
|
|
`openat2` with `RESOLVE_BENEATH` is Linux-specific and requires kernel 5.6 or later. On macOS, older Linux, and Windows, no equivalent single-syscall atomic containment primitive is assumed to exist in this specification; the fallback is `O_NOFOLLOW` applied per path component plus a canonicalize-and-verify-prefix check, which closes most but not all of the same time-of-check-to-time-of-use window. This is recorded as an accepted, platform-dependent residual risk rather than a claim of uniform containment — the same posture this document already takes toward the Windows rename-on-mapped-file caveat in Section 5.4.
|
|
|
|
## 7. Pack Integrity Validation on Load
|
|
|
|
A pack file is untrusted input whenever it did not originate from this process's own compaction step within the current run — for example, one imported through the zip compatibility path, supplied by another party, or read back after a crash. Loading such a file must validate its structure before any offset within it is trusted:
|
|
|
|
- The magic value and version field are checked before anything else is read.
|
|
- For every index entry, `data_off + size` is checked against the file's actual length, and `name_off` against the length of the strings region, before that entry is used to satisfy any `vfs_open`, `vfs_stat`, or `vfs_readdir` call. An entry that fails this check is treated as a load-time error for the whole pack, not skipped silently, since a single bad offset can indicate either corruption or a deliberately crafted file.
|
|
- Entries are checked for overlap with the header and index regions themselves, rejecting a pack that aliases file content onto its own metadata.
|
|
- Where the pack did not originate from this run's own compaction, an overall checksum over the index (and optionally the blob region) is verified at load, applying the same self-checking-record principle already required of the journal in Section 4.3, rather than treating a pack file as trusted simply because the journal already is.
|
|
|
|
## 8. Object Model (C API)
|
|
|
|
The following sketch illustrates the intended scope; it is not a POSIX-complete interface. `isize` and `usize` denote project-specific signed and unsigned size typedefs (e.g., built on `ptrdiff_t`/`size_t`) and are not standard C — concrete types must be chosen before compilation.
|
|
|
|
```c
|
|
typedef struct Vfs Vfs;
|
|
typedef struct VfsFile VfsFile;
|
|
|
|
Vfs *vfs_new(void);
|
|
void vfs_free(Vfs *v);
|
|
|
|
int vfs_mount(Vfs *v, const char *path, Backend *b, int flags);
|
|
int vfs_unmount(Vfs *v, const char *path);
|
|
|
|
VfsFile *vfs_open(Vfs *v, const char *path, int flags);
|
|
isize vfs_read(VfsFile *f, void *buf, usize n);
|
|
isize vfs_write(VfsFile *f, const void *buf, usize n);
|
|
int vfs_close(VfsFile *f);
|
|
|
|
int vfs_stat(Vfs *v, const char *path, VfsStat *out);
|
|
int vfs_readdir(Vfs *v, const char *path, VfsDir *out);
|
|
int vfs_mkdir(Vfs *v, const char *path);
|
|
int vfs_unlink(Vfs *v, const char *path);
|
|
```
|
|
|
|
Backends implement a common, small vtable. Emulating every POSIX flag is not required at this stage.
|
|
|
|
## 9. On-Disk Pack Format
|
|
|
|
```text
|
|
"PKFS" u32 version
|
|
u64 index_offset
|
|
u64 index_count
|
|
... blobs ...
|
|
index: { name_off, data_off, size, mode, mtime }
|
|
strings
|
|
```
|
|
|
|
This layout is sufficient to specify a VFS whose on-disk representation remains tractable to reason about in full.
|
|
|
|
### 9.1 Index Ordering for Directory Enumeration
|
|
|
|
The index is stored sorted lexicographically by full path, not in insertion order. Exact-path lookup (`vfs_open`, `vfs_stat`) uses this order for a binary search rather than a linear scan; enumerating the children of a directory (`vfs_readdir`) uses the same order to compute a contiguous range via two binary searches — the lower and upper bounds of the path prefix — rather than scanning the whole index. This mirrors the reason LMDB and similar systems keep their primary structure ordered by key: a range query over a sorted structure is a bounded number of comparisons, not a full pass. The in-memory upper-layer index (Section 5) is expected to preserve the same ordering property, so that a merged directory listing (upper ranked over lower, per Section 4) does not itself degrade to a linear scan on the mutable side.
|
|
|
|
### 9.2 Content-Addressed Deduplication During Compaction
|
|
|
|
Compaction already reads every live blob to rewrite the pack (Section 4.1); this is a natural point, not an additional pass, at which to hash each blob and reuse a previous `data_off` for any blob whose content already appears earlier in the same pack, rather than storing duplicate bytes. This is an optional space optimization, not a correctness requirement, and is scoped strictly to exact-duplicate detection during compaction — it is not a general delta-compression or content-defined-chunking scheme, which would reintroduce the complexity this document's stated minimalism (Section 10) explicitly argues against.
|
|
|
|
### 9.3 Representing Empty Directories
|
|
|
|
The index as specified (Section 9) stores file entries; a directory's existence is otherwise implicit in the path prefixes of the files beneath it. This leaves no representation for a directory containing no files — the result of `vfs_mkdir` followed by nothing else — which would otherwise vanish across a compaction that only ever walks live file entries, since there would be no entry referencing that path at all.
|
|
|
|
`vfs_mkdir` therefore creates an explicit index entry for its path with `size = 0` and a reserved bit in `mode` marking it as a directory marker rather than a file. Compaction preserves a directory-marker entry exactly as it preserves a file entry (Section 4.1). The marker is superseded, not removed, the moment any file is created beneath that path — the directory's existence is then implied by that file as usual, and the marker becomes redundant without needing to be deleted; `vfs_readdir`'s range query (Section 9.1) does not distinguish a directory implied by a child from one made explicit by a marker, since both produce the same enumerated result.
|
|
|
|
## 10. Explicit Exclusions (Out of Scope for v0)
|
|
|
|
- In-place writable tar or zip as the system of record.
|
|
- FUSE as the first backend (a plausible later addition, not a core dependency).
|
|
- Invoking libarchive on every `read()` call.
|
|
- Full POSIX semantics, locking, sockets, or mmap of virtual files, in v0.
|
|
- A single format intended to simultaneously serve as initrd, game asset pack, and user home directory.
|
|
- Implementing the cross-process concurrency design in Section 5.6 (the design is specified; building it is deferred until a second process actually needs the pack file).
|
|
- Content-defined chunking or delta compression in the pack format (Section 9.2 permits only exact-duplicate elimination during compaction).
|
|
- Enforced permission bits, symlinks, and hard links. The `mode` field (Section 9) is stored and restored across compaction, including the directory-marker bit defined in Section 9.3, but is not interpreted as an access-control mechanism in v0; symlink and hard-link semantics have no representation in the index at all.
|
|
|
|
## 11. Recommendation
|
|
|
|
**VFS core + `mem` + `dir` + custom pack + overlay + compaction.**
|
|
Zip is supported via miniz for import/export. Tar is supported as a snapshot/export path where Unix-style semantics are preferred.
|
|
|
|
### 11.1 Static-Linking Constraint
|
|
|
|
This is a hard constraint, not a preference: the core (`vfs_*`, `mem`, `dir`, `pack`, overlay, compaction) must build and link as plain static C with **zero required third-party libraries and zero dynamic loading**. miniz and libarchive are placed behind a compile-time switch strictly at the `zip`/`tar` backend boundary; neither may be required to obtain a working, writable VFS. `openat2` and Landlock (Section 6) are direct syscalls, not libraries, and do not weaken this constraint. This distinction — between "static" as an enforced property and "static" as an aspiration — is treated as load-bearing throughout the design.
|
|
|
|
Consequences of the overall recommendation:
|
|
|
|
- Static C throughout the core.
|
|
- No mandatory third-party dependency at runtime.
|
|
- Writability without claiming tar or zip can safely serve as a writable source of truth.
|
|
- One-file distribution where desired.
|
|
- A host directory where a one-file distribution is not desired.
|
|
- A specified containment mechanism for a chroot-shaped sandbox (Section 6), not merely a place to attach one later.
|
|
|
|
Where the requirement is specifically "one file that supports both `open()` of existing paths and persistence of new writes," the accurate name for that file is **pack + journal**, not zip and not tar.
|
|
|
|
## 12. Evaluation
|
|
|
|
### 12.1 Methodology
|
|
|
|
The design was reviewed against its own stated claims, checked for internal consistency, and cross-referenced against established prior art for the components with the highest risk of subtle error: archive format semantics, overlay filesystem deletion semantics, concurrent access under copy-on-write, memory-reclamation scaling, path-containment mechanisms for host-backed mounts, and the pack format's coverage of the object model it claims to support. Sources consulted are listed in the References section. Issues identified during review were corrected in the current text rather than merely noted, including: the original overlay deletion procedure having no defined behavior for a lower-layer-only entry (resolved by the whiteout mechanism, Section 4.2); durability and concurrency initially asserted without a supporting mechanism (resolved by Sections 4.3 and 5); every `vfs_write` call requiring a full index copy and republish (resolved by the structural/content mutation split, Section 5.3); reference counting's scaling limits left unexamined (resolved by Section 5.4's citation-backed trade-off); cross-process concurrency and path containment being named as deferred without a design behind them (resolved by Sections 5.6 and 6, respectively); pack files being treated as trusted input regardless of origin (resolved by Section 7); the mount table being outside the scope of the snapshot the concurrency model otherwise protects (resolved by Section 5.3); a `mem`-backed buffer growing under a concurrent reader being able to produce a use-after-free rather than merely a stale read (resolved by Section 5.7); and `vfs_mkdir` having no representable, compaction-surviving result for an empty directory (resolved by Section 9.3).
|
|
|
|
### 12.2 Assessment by Category
|
|
|
|
| Category | Grade | Notes |
|
|
|---|---|---|
|
|
| Correctness / technical soundness | A | No outstanding false claims identified as of the current revision; this pass additionally caught and closed a real use-after-free hazard in the concurrency model itself (Section 5.7) rather than only in the parts of the design added previously. |
|
|
| Concurrency design | A | The single-writer, wait-free-reader model is precisely specified, including the structural/content mutation distinction, a citation-backed reclamation trade-off, a concrete (if deferred) cross-process design, mount-table coverage under the same snapshot, and safe buffer growth on the `mem` backend. |
|
|
| Security / containment model | B+ | Virtual-namespace canonicalization and capability-scoped `dir` mounts via `openat2`/Landlock give kernel-enforced containment on Linux 5.6+. Graded below A because no equivalent atomic primitive is claimed or achieved on macOS or Windows — an honestly stated residual risk, not a closed gap. |
|
|
| Architecture clarity | A | The combination of mount table, backend vtable, immutable pack image, and snapshot-pointer-governed mutable state forms one coherent model without internal contradiction. |
|
|
| Static-linking discipline | A | The constraint is explicit and consistently enforced at backend boundaries; every mechanism added across revisions (atomics, mutexes, `openat2`, Landlock) relies on the standard toolchain and kernel syscalls, not additional third-party dependencies. |
|
|
| Completeness for a concept-level specification | A | Path containment, pack-integrity validation, empty-directory representation, and an explicit statement of what `mode` does and does not mean are now specified. Permissions, symlinks, hard links, and full cross-platform containment parity are stated as scoped exclusions (Section 10) rather than silent omissions. |
|
|
| Writing / usability as a specification | A | Precise on every mechanism most prone to subtle error (memory ordering, truncation direction, platform-specific rename and containment semantics, buffer-growth reclamation); implementable directly from this text. |
|
|
|
|
### 12.3 Overall Grade
|
|
|
|
**A.** Each limitation identified in the prior self-evaluations has been addressed to the extent it admits a static, dependency-free solution: reclamation scaling has a citation-backed recommendation, cross-process concurrency and path containment have concrete designs rather than bare deferrals, pack files are no longer treated as trusted by default, the mount table is no longer outside the concurrency model's protection, buffer growth can no longer race a reader into a use-after-free, and empty directories are representable across compaction. The remaining gaps — full cross-platform containment parity, cross-process implementation, and enforced permissions/symlinks/hard links — are stated as explicit, scoped exclusions rather than silent omissions, which this document's own standard treats as the appropriate closure for a concept-level specification.
|
|
|
|
## 13. Conclusion
|
|
|
|
The core architectural claim — that a virtual file system should treat a pack file as an immutable image, derive writability from a `mem` or `dir` upper layer combined with a copy-on-write overlay, and confine tar/zip to import/export roles — is supported both by the internal analysis in this document and by analogous decisions in established systems (LMDB's shadow paging, OverlayFS's whiteout and copy-up mechanisms, WASI's capability-based preopens, and Landlock's stackable access control). The design is considered suitable as a specification from which an initial implementation may proceed, subject to the exclusions enumerated in Section 10 and the limitations enumerated in Sections 5.4, 5.7, and 6.3.
|
|
|
|
## References
|
|
|
|
[1] "LMDB," Database of Databases, https://dbdb.io/db/lmdb.
|
|
[2] "Overlayfs Copy-on-Write Container Escape: CVE-2023-0386 and Writeback Race Mitigations," Systems Hardening, https://www.systemshardening.com/articles/kubernetes/overlayfs-cow-container-escape/.
|
|
[3] "Overlay Filesystem," The Linux Kernel Documentation, https://www.kernel.org/doc/Documentation/filesystems/overlayfs.txt.
|
|
[4] T. E. Hart, P. E. McKenney, A. D. Brown, J. Walpole, "Making Lockless Synchronization Fast: Performance Implications of Memory Reclamation," https://pdfs.semanticscholar.org/ea37/ace00efe3a22791b270146a911930f088102.pdf.
|
|
[5] "Reader Lock Table," LMDB documentation, http://www.lmdb.tech/doc/group__readers.html.
|
|
[6] Y. Nakata, "WASI's Capability-based Security Model," https://www.chikuwa.it/blog/2023/capability/.
|
|
[7] "openat2(2) — Linux manual page," https://man7.org/linux/man-pages/man2/openat2.2.html; "Restricting path name lookup with openat2()," LWN.net, https://lwn.net/Articles/796868/.
|
|
[8] "Landlock: unprivileged access control," The Linux Kernel Documentation, https://docs.kernel.org/userspace-api/landlock.html.
|
|
[9] P. E. McKenney, "What is RCU? Part 2: Usage," LWN.net, https://lwn.net/Articles/263130/.
|