Commit Graph
8 Commits
Author SHA1 Message Date
retoorandClaude Sonnet 5 4aad7d6f64 Fix real memory leak Gitea CI caught, that local testing had been masking
CI / build-and-test (push) Failing after 38s
Gitea CI flagged two things on the last push:

1. A -Wunused-result warning on an intentionally-ignored write() return
   value in test_crash_consistency.c's progress side-channel. Fixed with
   an explicit (void) cast and a comment explaining why ignoring it is
   safe (a short write there only makes the progress count more
   conservative, per that function's own existing documented tolerance).

2. A real LeakSanitizer failure -- 256 bytes across 4 allocations from
   vfs_new/vfs_unmount. Root cause: tests/test_journal_failure.c's two
   "reopen after the failure, verify recovery" blocks called
   vfs_unmount/backend_free/backend_free inside their `if (ov2)` branch
   (the normal, expected path) but vfs_free(v2) only on the `else`
   branch, which is never actually reached in practice. Fixed by moving
   vfs_free(v2) to run unconditionally after the if, in both blocks.

This bug was invisible locally across many runs because local sanitizer
verification had been using ASAN_OPTIONS=detect_leaks=0 -- adopted
originally for a real reason (a SIGKILLed forked child in
test_crash_consistency.c never runs its own exit-time leak check, so its
allocations were never the actual concern) but applied to the whole test
run, which also suppressed detection of this real bug in the *parent*
process's own code. CI doesn't set that option, so it caught what local
runs couldn't. Documented in CONTRIBUTING.md as a real process gap, not
just a code bug: a local verification habit that diverges from what CI
actually runs can let a real finding through until it reaches CI.

Verified: confirmed the leak directly first (reproduced locally by
dropping the detect_leaks=0 override, matching CI exactly, before
touching any code), then confirmed the fix by re-running the same
no-override sweep across all 8 test binaries with zero leaks found, plus
a clean make all + make test.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
2026-09-14 20:10:25 +00:00
retoorandClaude Sonnet 5 153d44ee3c Fix silent journal-write-failure data loss; confirm TSan blocked twice over
CI / build-and-test (push) Failing after 33s
Prompted by "fix literally everything still open" after the previous
session's data-integrity work. Went through each open item in turn:

1. TSan: tried a genuinely different execution environment (a remote
   cloud sandbox, via a dedicated agent) rather than re-stating the local
   sandbox's limitation. Result: identical block there too --
   personality(ADDR_NO_RANDOMIZE) returns EPERM, a trivial pthread
   program fails TSan identically, and all 6 PackFS test binaries fail
   with the same FATAL: ThreadSanitizer: unexpected memory mapping
   signature. This is now confirmed in two independent environments, not
   one -- strong evidence it's a real infrastructure restriction, not a
   one-off fluke worth chasing further with the tools available here.

2. While investigating the "disk-full mid-write" gap flagged as untested
   last session, found two real, previously-unknown bugs by reading the
   journal code (not by a test catching them unprompted):

   - journal_append_record and everything that called it were void, and
     none of the fwrite/fflush/fsync calls inside had their return values
     checked. A real write failure (disk full, quota, I/O error) was
     silently reported as success to vfs_write/vfs_mkdir/vfs_unlink/
     vfs_rename -- directly contradicting Section 4.4's premise that
     success means durable.

   - Fixing that alone was not enough, confirmed by direct reproduction:
     a partial write leaves a torn record in the journal, and
     journal_replay correctly stops at the first record it can't fully
     read (Section 4.3) -- which means every record appended *after* the
     torn one, including ones that themselves wrote perfectly fine later,
     became silently unreachable on reopen. Reproduced directly before
     fixing: a forced-failed write followed by a genuinely successful one
     was unrecoverable. Fixed by rolling the journal file back to its
     exact pre-record length on any failed write.

   Both closed in src/overlay.c (journal_append_record/_put/_delete/
   _mkdir/journal_put_current now return and propagate success/failure;
   overlay_write/_mkdir/_unlink/_rename return VFS_ERR_IO on a durability
   failure without rolling back the already-applied in-memory change,
   the same asymmetry a real write()-then-failed-fsync() has). Covered
   permanently by the new tests/test_journal_failure.c, which forces a
   real failure via RLIMIT_FSIZE + ignored SIGXFSZ, not a mock.

   Also fixed in the same pass, found by inspection while touching this
   code: journal_put_current used to pass a NULL buffer into a memcpy of
   a nonzero size when malloc(size) failed (an OOM-triggered NULL-pointer
   dereference) -- closed with an explicit allocation-failure check.
   Not test-triggered (forcing malloc() failure portably isn't practical
   here); verified by code inspection instead, stated as such rather than
   claimed as tested.

3. The remaining "journal-truncation-specific crash window" gap from last
   session was investigated, not silently dropped: reliably targeting
   that narrow a window would need real concurrency (a second writer
   thread racing the kill) for benefit the existing compaction-crash test
   already gets probabilistically -- a poor trade, so left as a stated,
   deliberate non-goal (CLAUDE.md) rather than built.

4. Cross-process contention is NOT addressed here and should not be read
   as an oversight: it is concept.md's own explicit, permanent "not
   implemented in v0" scope boundary (a specified-but-unbuilt LMDB-style
   reader-table design), not a bug -- building it would be a large,
   unrequested feature addition outside this session's actual scope.

Verified: clean make all + make test (all 8 binaries), make bench and
make demo still build and the demo runs correctly end to end, and a full
ASan/UBSan sweep of all 8 binaries with zero real findings (some retries
needed for the already-documented DEADLYSIGNAL flake, which
test_crash_consistency hits more often than other tests simply because it
forks 60+ subprocesses per run -- noted in CONTRIBUTING.md so this isn't
mistaken for a regression later).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
2026-09-14 19:40:23 +00:00
retoorandClaude Sonnet 5 e41bd2334a Add real fault-injection crash-safety testing; investigate TSan for real
CI / build-and-test (push) Failing after 38s
Prompted by a direct question about data-integrity trustworthiness: the
crash-safety design (self-checking journal records, atomic-rename
compaction, Section 4.3/4.4) had never actually been tested against a real
crash -- only reasoned about statically and tested against an
already-corrupted file (test_pack_overlay.c, a different scenario).

Added tests/test_crash_consistency.c: real fork()+SIGKILL fault injection
against an actual child process, not a hand-truncated file standing in
for a crash, with empirically-calibrated kill-delay sampling (measured via
throwaway scripts, not guessed) so trials land genuine mid-operation
interruptions rather than always completing first:

- Journal-write crash injection (25 trials): kills a child mid-burst of
  individually-journaled, individually-fsynced writes; verifies a strict
  clean-prefix recovery (every completed write present and correct, every
  write after the kill cleanly absent, no gaps). 20-22/25 trials per run
  land a genuine interruption; zero corruption found.
- Compaction crash injection (40 trials): kills a child mid-vfs_sync;
  verifies every write durably journaled *before* vfs_sync was called
  survives regardless of whether compaction itself completed. 40/40
  trials per run land a genuine interruption; zero corruption found.

A real finding from building this test, not from the code under test: the
first draft reported 128 failures, every one a bug in the test itself --
it couldn't distinguish "the child was killed before ever attempting this
write" from "the write completed and was then lost," since SIGKILL can't
be caught to report progress. Fixed with an independent progress
side-channel (plain write()+fsync() on a separate file, entirely outside
packfs) as ground truth. Documented in CLAUDE.md's new "Known reliability
characteristics" section, including this finding, because a claim is only
as trustworthy as what's measuring it.

Also investigated whether this sandbox's TSan block is actually
unfixable, rather than re-asserting the known limitation: tried
personality(ADDR_NO_RANDOMIZE) directly (EPERM), setarch -R (fails
identically), searched TSAN_OPTIONS for a bypass (none exists), and
checked capabilities/seccomp/unshare --user (zero effective caps, active
seccomp filter, namespace escape also blocked). Confirmed categorical and
documented as such in CLAUDE.md/CONTRIBUTING.md, rather than left as an
unexamined "it doesn't work here."

Also closed a real documentation gap SECURITY.md had: the pack integrity
checksum's non-cryptographic nature (FNV-1a64, already documented in the
context of the compaction-dedup bug) had never been explicitly connected
to its *other* use, load-time pack integrity validation -- added, since
it's a real, relevant caveat for anyone relying on that checksum against
a deliberate adversary rather than accidental corruption.

Verified: clean make all + make test (all 7 binaries), full ASan/UBSan
sweep of all 7 binaries with zero real findings, and the new test run
standalone 4+ times (this session) with stable, consistent results.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
2026-09-14 19:19:21 +00:00
retoorandClaude Sonnet 5 419182bb05 Add version API, pkg-config, SPDX headers, SECURITY.md, CHANGELOG.md
Project-hygiene pass toward being a properly citable, embeddable,
professionally-packaged C library rather than just working code:

- PACKFS_VERSION_MAJOR/MINOR/PATCH/STRING in include/packfs.h, the single
  source of truth for the project's version, plus a runtime pfs_version()
  (src/vfs.c, next to vfs_new/vfs_free) so a dynamically-linked consumer
  can check ABI/API compatibility without recompiling. Covered by a new
  assertion in tests/test_mem.c that the macro and the runtime function
  never disagree.
- packfs.pc.in + a `make install` rule that generates packfs.pc with its
  Version: field derived from PACKFS_VERSION_STRING via a Makefile-level
  grep/sed, never hand-maintained separately -- verified end-to-end with a
  scratch `make install PREFIX=...` + `pkg-config --cflags --libs packfs`
  + `make uninstall`, not just by reading the rule.
- SPDX-License-Identifier: MIT added to every src/*.c and src/internal.h
  (include/packfs.h already had one); the whole distributed source tree
  now carries consistent machine-readable license metadata.
- SECURITY.md, stating precisely what this project's containment and pack-
  integrity code actually claims as a security boundary (concept.md
  Section 6/7) versus what it explicitly does not (unenforced `mode`, no
  cross-process concurrency) -- not generic boilerplate -- with a real
  reporting contact rather than a placeholder.
- CHANGELOG.md (Keep a Changelog format), summarizing the real history in
  git log to date; explicitly notes no version is tagged yet.

Verified: clean `make all` + `make test` (all 6 binaries, including the
new version-check assertion), and a full ASan/UBSan sweep of all 6
binaries with zero real findings (one run hit the already-documented
DEADLYSIGNAL sandbox flake on test_pack_write_perf across all 3 retries;
re-verified directly afterward with 5/5 additional clean passes and timing
well under any plausible timeout, confirming it was the flake, not a
regression, before treating this as done).

Deliberately not done here, by the user's explicit choice: no
CODE_OF_CONDUCT.md, and no git remote/publishing -- this repository still
has neither, and both are decisions left to the maintainer.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
2026-09-14 10:58:35 +00:00
retoorandClaude Sonnet 5 684c6d95de Fix pack_write's O(n^2) dedup + correctness bug; document mount table O(n^2)
A audit for other instances of the file-index O(n^2) shape (fixed
previously via the persistent treap) found two more real issues:

1. pack_write's Section 9.2 exact-duplicate elimination was a linear scan
   of every previously-seen (hash, size) pair per entry -- O(n^2) total,
   invisible in the existing benchmark because its identical-content test
   files made every scan match on the first comparison. Measured with
   unique content instead: 80,000 entries took 1.74s, with a 20,000->80,000
   step showing 15.6x for a 4x-N step, matching O(n^2)'s 16x prediction.
   The same scan also trusted a (hash, size) match without ever comparing
   actual bytes -- a latent correctness bug, since FNV-1a64 is explicitly
   not collision-resistant. Fixed both at once with an open-addressing hash
   table (load factor 1/2, linear probing) plus a memcmp verification
   before ever reusing a data_off. Post-fix: 80,000 entries in 0.044s
   (39.6x faster), ratio drops to 3.35x (consistent with O(n)).

   Covered permanently by two new/extended tests: a white-box assertion in
   test_pack_overlay.c that duplicate-content entries share one data_off
   and distinct-content entries do not, and a new
   tests/test_pack_write_perf.c regression tripwire against 10,000 unique
   entries.

2. vfs.c's mount table uses the same full-array-copy-per-write pattern the
   file index used to, confirmed O(n^2) via a new bench/bench.c category
   (500/2,000/8,000 mounts, both 4x-N steps showing 15-20x). Deliberately
   NOT rewritten: mount points are created by a program's own source code,
   not workload-driven, so realistic mount counts never reach the scale
   that made the file index's O(n^2) a real problem. Documented with full
   reasoning in BENCH.md and CLAUDE.md rather than silently left as an
   undocumented gap.

Also fixes a real CI gap the new tests exposed: ci.yml's sanitizer-build
steps never passed -D_GNU_SOURCE when compiling test files (only the
library .o's got it), which was harmless while no test included
internal.h and became a link failure once two did (internal.h needs
_GNU_SOURCE for pthread_rwlock_t). And documents, in CONTRIBUTING.md and
CLAUDE.md, a sandbox flake observed directly during this work's own
sanitizer runs: ASan/UBSan test binaries occasionally fail to start with
AddressSanitizer:DEADLYSIGNAL (sometimes looping rather than exiting),
non-deterministically hitting different unrelated binaries across runs --
a startup race, not a memory-safety bug, confirmed by clean passes on
retry; sanitizer runs in such an environment should be timeout-wrapped.

BENCH.md's "After" table and Appendix B are replaced with the current,
complete 54-measurement bench/bench.c run (the original 45 plus the new
mount-scaling category); the pre-fix 45-measurement "Before" table is kept
as the historical record, per this project's documentation standard.

Verified: make test (all 6 binaries, including the 2 new/changed), a clean
make all, and repeated ASan+UBSan runs (0 real findings; the DEADLYSIGNAL
flake above was observed and correctly distinguished from a real finding
by re-running until a clean pass). TSan could not be run in this sandbox
(pre-existing, documented environment limitation).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
2026-09-14 09:55:47 +00:00
retoorandClaude Sonnet 5 64283d587b Fix the O(n^2) bulk create/unlink: flat array -> persistent treap
BENCH.md's earlier finding was real: every structural write (create/
unlink/mkdir) copied the entire sorted UpperSnapshot entry array before
publishing the next snapshot, making bulk sequential creation O(n^2).
concept.md Section 5.3 names the exact condition for reconsidering this
("a persistent structurally-shared tree structure is not required
until this assumption is empirically violated") -- that condition was
measured, not hypothesized, so this closes it rather than leaving it
as a documented-but-open limitation.

The index is now a persistent treap (src/upper.c): a structural write
copies only the O(log n) nodes on the path to the change, sharing
every other node (its own refcount, cascading like MutCell's and
UpperSnapshot's) with whichever snapshot(s) it was built from. Chosen
over a persistent AVL/red-black/weight-balanced tree because deletion
in those can need O(log n) rebalancing rotations -- each a real
allocation in a persistent setting -- where a treap needs only O(1)
amortized rotations for insert and delete (Seidel & Aragon 1996), with
expected O(log n) height regardless of insertion order, including the
sorted-by-creation-order pattern that made the flat array quadratic in
the first place. Liljenzin's "Confluently Persistent Sets and Maps"
(arXiv:1301.3388) documents persistent treaps giving O(1) snapshots
for MVCC specifically, which is this exact use case. Full reasoning
and the ownership convention (functions consume one ref of their tree
arguments, return one owned ref) are in upper.c's comment above
struct TreapNode.

Blast radius kept deliberately small: snapshot_upsert/snapshot_remove
keep their exact original signatures, so upper_create/upper_mkdir/
upper_remove/upper_rename/upper_copy_up needed zero changes.
upper_lookup/upper_has_children keep their exact contracts. Only
upstd_readdir and overlay_readdir's manual array scans became calls to
a new upper_visit_range (O(log n + r) range query, replacing an O(n)
scan in both, a bonus fix beyond what was strictly necessary) since
there's no flat array left to scan.

Added tests/test_index_stress.c: thousands of randomized (not
sequential) creates/deletes/renames across nested directories,
cross-checked against an independent reference model after every
round, not just "did it not crash" -- exercises exactly the code path
the O(n^2) bug and this fix live in, at a scale the other tests don't
reach. Verified under -fsanitize=undefined (150+ runs across this
change's lifetime, 0 failures) and -fsanitize=address (100+ runs, 0
real findings; known sandbox ASan-startup flakes excluded, see prior
commits) per CLAUDE.md's sanitizer rule for upper.c/overlay.c changes.

Measured result (make bench, same environment as the original
finding): mem create 257x faster, unlink 330x faster, mkdir 65x
faster, concurrent mixed workload 131x faster -- and, the comparison
that matters, mem now beats raw fs at every one of these (was losing
by 6-22x before). The complexity-class change is confirmed the same
way the O(n^2) was found: mkdir at N=4,000 vs create at N=20,000 now
shows a 7.15x slowdown for a 5x increase in N, matching the O(n log n)
prediction (5.97x) rather than the old O(n^2) one (25x). Full
before/after tables in BENCH.md's new "Resolution" section, which
keeps the original run as the historical record rather than
overwriting it, per this project's own documentation standard.

Recorded the fix in CLAUDE.md's "Known performance characteristics"
(marked RESOLVED, not silently removed) and its architecture map.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
2026-09-14 08:09:49 +00:00
retoorandClaude Sonnet 5 9fb5cf2d9a Document every public function; fix a real gap the audit found
Cross-checked every function declared in include/packfs.h against
nm -D libpackfs.so.0 as the starting point for a full documentation
pass, per the request to document literally everything rather than
just the parts already covered.

That check found a genuine bug, not just a documentation gap:
backend_pack_new was declared in the public header and named in
CLAUDE.md's architecture map, but never implemented in src/pack.c —
any caller would fail at link time. Implemented it as a standalone,
read-only `pack` Backend (every mutating call returns VFS_ERR_PERM,
consistent with concept.md Section 2.1 listing `pack` as its own
backend kind distinct from the overlay), covered it with a new test
case, and verified it under -fsanitize=undefined per CLAUDE.md's
sanitizer rule.

Added a doc comment to every previously-undocumented function and
struct field in packfs.h and internal.h (vfs_open/read/write/close/
stat/readdir/mkdir/unlink/rename, every upper_* structural/content
function, pfs_dir_*, pfs_fnv1a64, PackIndexEntry/Pack fields). Updated
README and CLAUDE.md to mention backend_pack_new and to stop gesturing
at zip/tar as though import/export exists.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
2026-09-14 05:54:52 +00:00
retoorandClaude Sonnet 5 72e3c900f2 Add PackFS v0: statically linked in-process VFS implementing concept.md
Implements the core design: a mount table published as an atomically-
swapped snapshot; mem/dir/pack/overlay backends; copy-on-write overlay
with copy-up and whiteout deletion; a checksummed append journal;
compaction with exact-duplicate elimination; single-writer/wait-free-
reader concurrency with a structural/content write split; openat2/
Landlock path containment for dir mounts; and load-time pack integrity
validation. Zero required third-party dependencies.

Sanitizer testing (ASan/UBSan) caught and led to fixing a genuine
heap-use-after-free in the snapshot-reclamation path: the textbook
"load pointer, then increment its refcount" pattern left a gap a
concurrent writer could free through. Closed with a small reclaim_gate
rwlock, documented in internal.h and CLAUDE.md since it's a pattern
every refcounted structure in the codebase now follows.

zip/tar import/export backends, recommended in concept.md Section 11,
will not be built — a permanent project decision recorded in CLAUDE.md
since concept.md itself is frozen and cannot be edited to reflect it.

Includes a runnable demo (examples/demo.c, `make demo`) exercising the
library end to end and proving cross-run persistence through the pack
file, plus open-source scaffolding: MIT license, README, CONTRIBUTING,
and a CI workflow running the test suite under ASan/UBSan/TSan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
2026-09-14 05:44:08 +00:00