Prompted by "fix literally everything still open" after the previous
session's data-integrity work. Went through each open item in turn:
1. TSan: tried a genuinely different execution environment (a remote
cloud sandbox, via a dedicated agent) rather than re-stating the local
sandbox's limitation. Result: identical block there too --
personality(ADDR_NO_RANDOMIZE) returns EPERM, a trivial pthread
program fails TSan identically, and all 6 PackFS test binaries fail
with the same FATAL: ThreadSanitizer: unexpected memory mapping
signature. This is now confirmed in two independent environments, not
one -- strong evidence it's a real infrastructure restriction, not a
one-off fluke worth chasing further with the tools available here.
2. While investigating the "disk-full mid-write" gap flagged as untested
last session, found two real, previously-unknown bugs by reading the
journal code (not by a test catching them unprompted):
- journal_append_record and everything that called it were void, and
none of the fwrite/fflush/fsync calls inside had their return values
checked. A real write failure (disk full, quota, I/O error) was
silently reported as success to vfs_write/vfs_mkdir/vfs_unlink/
vfs_rename -- directly contradicting Section 4.4's premise that
success means durable.
- Fixing that alone was not enough, confirmed by direct reproduction:
a partial write leaves a torn record in the journal, and
journal_replay correctly stops at the first record it can't fully
read (Section 4.3) -- which means every record appended *after* the
torn one, including ones that themselves wrote perfectly fine later,
became silently unreachable on reopen. Reproduced directly before
fixing: a forced-failed write followed by a genuinely successful one
was unrecoverable. Fixed by rolling the journal file back to its
exact pre-record length on any failed write.
Both closed in src/overlay.c (journal_append_record/_put/_delete/
_mkdir/journal_put_current now return and propagate success/failure;
overlay_write/_mkdir/_unlink/_rename return VFS_ERR_IO on a durability
failure without rolling back the already-applied in-memory change,
the same asymmetry a real write()-then-failed-fsync() has). Covered
permanently by the new tests/test_journal_failure.c, which forces a
real failure via RLIMIT_FSIZE + ignored SIGXFSZ, not a mock.
Also fixed in the same pass, found by inspection while touching this
code: journal_put_current used to pass a NULL buffer into a memcpy of
a nonzero size when malloc(size) failed (an OOM-triggered NULL-pointer
dereference) -- closed with an explicit allocation-failure check.
Not test-triggered (forcing malloc() failure portably isn't practical
here); verified by code inspection instead, stated as such rather than
claimed as tested.
3. The remaining "journal-truncation-specific crash window" gap from last
session was investigated, not silently dropped: reliably targeting
that narrow a window would need real concurrency (a second writer
thread racing the kill) for benefit the existing compaction-crash test
already gets probabilistically -- a poor trade, so left as a stated,
deliberate non-goal (CLAUDE.md) rather than built.
4. Cross-process contention is NOT addressed here and should not be read
as an oversight: it is concept.md's own explicit, permanent "not
implemented in v0" scope boundary (a specified-but-unbuilt LMDB-style
reader-table design), not a bug -- building it would be a large,
unrequested feature addition outside this session's actual scope.
Verified: clean make all + make test (all 8 binaries), make bench and
make demo still build and the demo runs correctly end to end, and a full
ASan/UBSan sweep of all 8 binaries with zero real findings (some retries
needed for the already-documented DEADLYSIGNAL flake, which
test_crash_consistency hits more often than other tests simply because it
forks 60+ subprocesses per run -- noted in CONTRIBUTING.md so this isn't
mistaken for a regression later).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
6.5 KiB
Security Policy
Supported versions
No version has been tagged yet (see CHANGELOG.md); this project is
pre-1.0 (README.md's "Status"). Only the latest commit on the default
branch is supported — there is no LTS branch and no backport policy while
the project is at 0.x.
What PackFS actually claims as a security boundary
Stated precisely, per this project's documentation standard (CLAUDE.md),
rather than a generic "we take security seriously":
- Path containment for
dirmounts is a hard requirement (concept.mdSection 6,CLAUDE.md's "Load-bearing constraints"). On Linux 5.6+, every lookup beneath adirmount usesopenat2(RESOLVE_BENEATH | RESOLVE_NO_SYMLINKS)to atomically block..escapes, symlink escapes, and the canonicalize-then-open TOCTOU window in one kernel call.vfs_harden_process_with_landlock(Section 6.2) is an additional, opt-in kernel-level layer — not applied automatically, since Landlock restrictions are cumulative, irreversible, and process-wide, which would be a surprising side effect for an embeddable library to trigger on its own. - On platforms or kernels without
openat2, containment falls back to per-componentO_NOFOLLOWplus canonicalize-and-verify (Section 6.3) — an accepted, explicitly weaker residual-risk posture, not a claim of equal strength. Code and documentation that describe this fallback as equivalent to theopenat2path would themselves be a bug worth reporting. - The VFS core canonicalizes
./..in the virtual namespace itself, before any backend is reached, so a crafted path cannot escape a mount's prefix even when no host directory is involved. - A pack file is treated as untrusted input unless it came from this
run's own compaction (Section 7). Every index entry's offsets and
sizes are bounds-checked against the actual file/strings-region length
before being trusted; a failed check invalidates the whole pack load. An
externally-sourced pack additionally needs a checksum verified at load.
Loading an attacker-supplied
.imgfile that PackFS accepts without detecting corruption, or that causes an out-of-bounds read/write, is a security bug. That checksum (FNV-1a64,src/hash.c) is explicitly non-cryptographic and not collision-resistant — stated plainly here because this is the one place in this document where it matters most: it is a real, effective check against accidental corruption (a truncated copy, a bit flip, a bad transfer), but it is not a tamper- evident guarantee against a deliberate adversary who can choose the bytes of a.imgfile — someone with that capability could in principle construct a corrupted pack whose checksum still matches. Do not treat "checksum verified at load" as a substitute for verifying the source of an externally-supplied pack through some other channel if the threat model includes a deliberate adversary, not just accidental corruption. - Journal records are length/checksum self-describing (Section 4.4), so replay stops cleanly at the first torn or corrupted record rather than reading past it.
- The crash-safety design above is verified by actual fault injection,
not only by reasoning about the design.
tests/test_crash_consistency.cforks a real child process andSIGKILLs it at randomized points during journaled writes and during compaction, then checks what a fresh reopen recovers, across dozens of trials per run — not a hand-truncated file standing in for a crash. SeeCLAUDE.md's "Known reliability characteristics" for the methodology and results, including a real bug the test itself had on its first draft (conflating "never attempted before the kill" with "lost after completing") — fixed before treating the crash-safety claim as verified. - A real, previously-unknown data-loss bug in this exact area was found
and fixed: a failed journal write (disk full, quota, an I/O error) used
to be silently reported as success, and even after fixing that, a
partial write left behind a torn record that made every later,
individually-successful write unrecoverable on reopen too — not just
the one that actually failed. Both are fixed (
src/overlay.c,journal_append_record) and covered bytests/test_journal_failure.c, which forces a real write failure viaRLIMIT_FSIZErather than simulating one. SeeCLAUDE.md's "Known reliability characteristics" for the full story, including how each half of the bug was confirmed by direct reproduction before and after the fix, not just inspection.
What PackFS explicitly does not claim
Reporting these as vulnerabilities is not necessary — they are documented, known limitations, not silent gaps:
mode(permission bits, symlinks, hard links) is stored and restored but not enforced. Nothing in the VFS core interprets it as an access control boundary in v0 (concept.mdSection 10).- Cross-process concurrency is specified (an LMDB-style reader-table
design, Section 5.6) but not implemented. Two uncoordinated processes
writing the same host directory through a
dirmount are not synchronized against each other. zip/tarare permanently out of scope (CLAUDE.md, "Project decisions that supersede concept.md") — there is no archive-format parsing code in this project to have a vulnerability in.
Reporting a vulnerability
This project is hosted on Gitea, not GitHub (see CONTRIBUTING.md).
Please do not open a public issue on the issue tracker for a suspected
security vulnerability — report it privately, by email, to the
maintainer:
retoor — retoor@molodetz.nl
When reporting, include: the affected function(s) or file(s), a minimal
reproduction (ideally as a tests/test_*.c-style program using only the
public API in include/packfs.h), and which of the guarantees above (or
which concept.md section) you believe is violated — this project's own
verification discipline (make test, ASan/UBSan/TSan per CONTRIBUTING.md)
depends on being able to turn a report into a permanent regression test.
Verification this project already does
Every change to src/upper.c, src/overlay.c, or src/vfs.c is required
to pass under -fsanitize=address,undefined and -fsanitize=thread
before being considered done (CLAUDE.md) — this process has already
caught and fixed a real heap-use-after-free in the snapshot-reclamation
logic during development (see the reclaim_gate note in src/internal.h).
.gitea/workflows/ci.yml (Gitea Actions) runs the full test suite plus
both sanitizer passes on every push.