6 Commits
Author SHA1 Message Date
retoorandClaude Sonnet 5 7661c94105 Add POSTMORTEM.md: every real issue found, how, fix, and regression prevention
CI / build-and-test (push) Successful in 48s
A detailed, permanent record covering the whole arc of this project's
development so far -- requested explicitly, as detailed as possible, to
prevent regression for ever. Seven sections:

1. Performance findings: the file-index O(n^2) treap rewrite, pack_write's
   O(n^2) dedup + latent hash-collision correctness bug, and the mount
   table's O(n^2) (confirmed, deliberately not fixed, with the reasoning).
2. Environment/tooling limitations: TSan's categorical block (every
   workaround actually tried and ruled out, not just the ones that
   worked), and both variants of the ASan/UBSan sandbox-startup flake.
3. Project professionalization: version API, SPDX, pkg-config, SECURITY.md,
   CHANGELOG.md, the Gitea-not-GitHub migration, and the CODE_OF_CONDUCT
   decision.
4. Git identity correction: the filter-branch rewrite, the tag-object
   tagger field it missed (found only by a full object-database sweep,
   not by re-reading git log), and the exhaustive re-verification.
5. Data-integrity fault injection: test_crash_consistency.c, the real bug
   in its own oracle (128 false failures before a progress side-channel
   fixed it), and the two real overlay.c bugs found while building the
   write-failure test (silent journal failure, torn-record poisoning of
   later records).
6. CI as a second, independent reviewer: the memory leak detect_leaks=0
   had been masking locally, and the set -e scripting bug -- including
   the mistake made while fixing it (reintroducing the same bug in a new
   shape), the follow-on bash-variable-blowup bug, the newly-discovered
   TSan segfault variant, and the retry-budget gap, each confirmed by
   direct reproduction under bash -eo pipefail, not reasoned about.
7. Distilled lessons: ten patterns extracted from the above, written to
   be applied to a new problem, not just recognized in this one.

Every figure cited was checked against BENCH.md/CLAUDE.md's own numbers
before writing this, not reproduced from memory.

Cross-referenced from CLAUDE.md's "Repository status", README.md's
"Contributing" section, and CHANGELOG.md -- including two entries CHANGELOG
itself was missing (the set -e CI fix from the previous commit, and this
document), the exact class of gap POSTMORTEM.md section 3 already
describes happening once before with the v0.1.0 tag.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
2026-09-14 21:02:58 +00:00
retoorandClaude Sonnet 5 4aad7d6f64 Fix real memory leak Gitea CI caught, that local testing had been masking
CI / build-and-test (push) Failing after 38s
Gitea CI flagged two things on the last push:

1. A -Wunused-result warning on an intentionally-ignored write() return
   value in test_crash_consistency.c's progress side-channel. Fixed with
   an explicit (void) cast and a comment explaining why ignoring it is
   safe (a short write there only makes the progress count more
   conservative, per that function's own existing documented tolerance).

2. A real LeakSanitizer failure -- 256 bytes across 4 allocations from
   vfs_new/vfs_unmount. Root cause: tests/test_journal_failure.c's two
   "reopen after the failure, verify recovery" blocks called
   vfs_unmount/backend_free/backend_free inside their `if (ov2)` branch
   (the normal, expected path) but vfs_free(v2) only on the `else`
   branch, which is never actually reached in practice. Fixed by moving
   vfs_free(v2) to run unconditionally after the if, in both blocks.

This bug was invisible locally across many runs because local sanitizer
verification had been using ASAN_OPTIONS=detect_leaks=0 -- adopted
originally for a real reason (a SIGKILLed forked child in
test_crash_consistency.c never runs its own exit-time leak check, so its
allocations were never the actual concern) but applied to the whole test
run, which also suppressed detection of this real bug in the *parent*
process's own code. CI doesn't set that option, so it caught what local
runs couldn't. Documented in CONTRIBUTING.md as a real process gap, not
just a code bug: a local verification habit that diverges from what CI
actually runs can let a real finding through until it reaches CI.

Verified: confirmed the leak directly first (reproduced locally by
dropping the detect_leaks=0 override, matching CI exactly, before
touching any code), then confirmed the fix by re-running the same
no-override sweep across all 8 test binaries with zero leaks found, plus
a clean make all + make test.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
2026-09-14 20:10:25 +00:00
retoorandClaude Sonnet 5 153d44ee3c Fix silent journal-write-failure data loss; confirm TSan blocked twice over
CI / build-and-test (push) Failing after 33s
Prompted by "fix literally everything still open" after the previous
session's data-integrity work. Went through each open item in turn:

1. TSan: tried a genuinely different execution environment (a remote
   cloud sandbox, via a dedicated agent) rather than re-stating the local
   sandbox's limitation. Result: identical block there too --
   personality(ADDR_NO_RANDOMIZE) returns EPERM, a trivial pthread
   program fails TSan identically, and all 6 PackFS test binaries fail
   with the same FATAL: ThreadSanitizer: unexpected memory mapping
   signature. This is now confirmed in two independent environments, not
   one -- strong evidence it's a real infrastructure restriction, not a
   one-off fluke worth chasing further with the tools available here.

2. While investigating the "disk-full mid-write" gap flagged as untested
   last session, found two real, previously-unknown bugs by reading the
   journal code (not by a test catching them unprompted):

   - journal_append_record and everything that called it were void, and
     none of the fwrite/fflush/fsync calls inside had their return values
     checked. A real write failure (disk full, quota, I/O error) was
     silently reported as success to vfs_write/vfs_mkdir/vfs_unlink/
     vfs_rename -- directly contradicting Section 4.4's premise that
     success means durable.

   - Fixing that alone was not enough, confirmed by direct reproduction:
     a partial write leaves a torn record in the journal, and
     journal_replay correctly stops at the first record it can't fully
     read (Section 4.3) -- which means every record appended *after* the
     torn one, including ones that themselves wrote perfectly fine later,
     became silently unreachable on reopen. Reproduced directly before
     fixing: a forced-failed write followed by a genuinely successful one
     was unrecoverable. Fixed by rolling the journal file back to its
     exact pre-record length on any failed write.

   Both closed in src/overlay.c (journal_append_record/_put/_delete/
   _mkdir/journal_put_current now return and propagate success/failure;
   overlay_write/_mkdir/_unlink/_rename return VFS_ERR_IO on a durability
   failure without rolling back the already-applied in-memory change,
   the same asymmetry a real write()-then-failed-fsync() has). Covered
   permanently by the new tests/test_journal_failure.c, which forces a
   real failure via RLIMIT_FSIZE + ignored SIGXFSZ, not a mock.

   Also fixed in the same pass, found by inspection while touching this
   code: journal_put_current used to pass a NULL buffer into a memcpy of
   a nonzero size when malloc(size) failed (an OOM-triggered NULL-pointer
   dereference) -- closed with an explicit allocation-failure check.
   Not test-triggered (forcing malloc() failure portably isn't practical
   here); verified by code inspection instead, stated as such rather than
   claimed as tested.

3. The remaining "journal-truncation-specific crash window" gap from last
   session was investigated, not silently dropped: reliably targeting
   that narrow a window would need real concurrency (a second writer
   thread racing the kill) for benefit the existing compaction-crash test
   already gets probabilistically -- a poor trade, so left as a stated,
   deliberate non-goal (CLAUDE.md) rather than built.

4. Cross-process contention is NOT addressed here and should not be read
   as an oversight: it is concept.md's own explicit, permanent "not
   implemented in v0" scope boundary (a specified-but-unbuilt LMDB-style
   reader-table design), not a bug -- building it would be a large,
   unrequested feature addition outside this session's actual scope.

Verified: clean make all + make test (all 8 binaries), make bench and
make demo still build and the demo runs correctly end to end, and a full
ASan/UBSan sweep of all 8 binaries with zero real findings (some retries
needed for the already-documented DEADLYSIGNAL flake, which
test_crash_consistency hits more often than other tests simply because it
forks 60+ subprocesses per run -- noted in CONTRIBUTING.md so this isn't
mistaken for a regression later).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
2026-09-14 19:40:23 +00:00
retoorandClaude Sonnet 5 21bcb4640f Move CI to Gitea Actions; this project is hosted on Gitea, not GitHub
CI / build-and-test (push) Failing after 23s
Per explicit direction: this repository is not going on GitHub. Moved
.github/workflows/ci.yml to .gitea/workflows/ci.yml (Gitea Actions'
convention) and updated every doc that referenced the old path or assumed
GitHub-specific features:

- .gitea/workflows/ci.yml: added a header comment on the two things that
  are genuinely Gitea-specific and instance-dependent, not just a renamed
  file -- `runs-on: ubuntu-latest` must match a label the actual
  registered Gitea runner advertises (there is no GitHub-hosted-runner
  equivalent, this is self-hosted), and `actions/checkout@v4` resolves
  against whatever action source that runner is configured with.
- CONTRIBUTING.md: corrected a claim that no longer holds -- it previously
  said CI "runs on a normal, unrestricted GitHub Actions VM where TSan is
  expected to work"; since this is actually a self-hosted Gitea runner
  whose environment isn't controlled by this repo, that assumption isn't
  something this repo can vouch for, so the text now says so rather than
  carrying the old (GitHub-shaped) assumption forward silently.
- SECURITY.md: removed a claim this project can't back up (that "private
  security advisories" are available once hosted -- that's a GitHub
  feature this repo never had access to); reporting is by direct email to
  the maintainer only.
- README.md: fixed a real gap while in here -- the "Security" section
  never actually linked to SECURITY.md despite it existing since the
  previous commit.
- CLAUDE.md, CHANGELOG.md: updated path references; CLAUDE.md's
  self-evaluation section (a historical record of an audit finding) keeps
  the old .github path where it describes what was literally true at that
  time, with a note explaining the rename, rather than rewriting history.

Verified: clean make all + make test, all 6 binaries pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
2026-09-14 11:03:20 +00:00
retoorandClaude Sonnet 5 a9143308f5 CHANGELOG.md: reflect that v0.1.0 is now tagged, not still Unreleased
The v0.1.0 tag was created in the same session right after this file was
written, leaving it self-inconsistent: a tagged v0.1.0 whose own
CHANGELOG.md still said "no version has been tagged yet" under an
[Unreleased] heading. Fixed by renaming the heading to "[0.1.0] -
2026-09-14" and updating the framing paragraph.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
2026-09-14 10:59:09 +00:00
retoorandClaude Sonnet 5 419182bb05 Add version API, pkg-config, SPDX headers, SECURITY.md, CHANGELOG.md
Project-hygiene pass toward being a properly citable, embeddable,
professionally-packaged C library rather than just working code:

- PACKFS_VERSION_MAJOR/MINOR/PATCH/STRING in include/packfs.h, the single
  source of truth for the project's version, plus a runtime pfs_version()
  (src/vfs.c, next to vfs_new/vfs_free) so a dynamically-linked consumer
  can check ABI/API compatibility without recompiling. Covered by a new
  assertion in tests/test_mem.c that the macro and the runtime function
  never disagree.
- packfs.pc.in + a `make install` rule that generates packfs.pc with its
  Version: field derived from PACKFS_VERSION_STRING via a Makefile-level
  grep/sed, never hand-maintained separately -- verified end-to-end with a
  scratch `make install PREFIX=...` + `pkg-config --cflags --libs packfs`
  + `make uninstall`, not just by reading the rule.
- SPDX-License-Identifier: MIT added to every src/*.c and src/internal.h
  (include/packfs.h already had one); the whole distributed source tree
  now carries consistent machine-readable license metadata.
- SECURITY.md, stating precisely what this project's containment and pack-
  integrity code actually claims as a security boundary (concept.md
  Section 6/7) versus what it explicitly does not (unenforced `mode`, no
  cross-process concurrency) -- not generic boilerplate -- with a real
  reporting contact rather than a placeholder.
- CHANGELOG.md (Keep a Changelog format), summarizing the real history in
  git log to date; explicitly notes no version is tagged yet.

Verified: clean `make all` + `make test` (all 6 binaries, including the
new version-check assertion), and a full ASan/UBSan sweep of all 6
binaries with zero real findings (one run hit the already-documented
DEADLYSIGNAL sandbox flake on test_pack_write_perf across all 3 retries;
re-verified directly afterward with 5/5 additional clean passes and timing
well under any plausible timeout, confirming it was the flake, not a
regression, before treating this as done).

Deliberately not done here, by the user's explicit choice: no
CODE_OF_CONDUCT.md, and no git remote/publishing -- this repository still
has neither, and both are decisions left to the maintainer.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
2026-09-14 10:58:35 +00:00