CI / build-and-test (push) Successful in 48s
A detailed, permanent record covering the whole arc of this project's development so far -- requested explicitly, as detailed as possible, to prevent regression for ever. Seven sections: 1. Performance findings: the file-index O(n^2) treap rewrite, pack_write's O(n^2) dedup + latent hash-collision correctness bug, and the mount table's O(n^2) (confirmed, deliberately not fixed, with the reasoning). 2. Environment/tooling limitations: TSan's categorical block (every workaround actually tried and ruled out, not just the ones that worked), and both variants of the ASan/UBSan sandbox-startup flake. 3. Project professionalization: version API, SPDX, pkg-config, SECURITY.md, CHANGELOG.md, the Gitea-not-GitHub migration, and the CODE_OF_CONDUCT decision. 4. Git identity correction: the filter-branch rewrite, the tag-object tagger field it missed (found only by a full object-database sweep, not by re-reading git log), and the exhaustive re-verification. 5. Data-integrity fault injection: test_crash_consistency.c, the real bug in its own oracle (128 false failures before a progress side-channel fixed it), and the two real overlay.c bugs found while building the write-failure test (silent journal failure, torn-record poisoning of later records). 6. CI as a second, independent reviewer: the memory leak detect_leaks=0 had been masking locally, and the set -e scripting bug -- including the mistake made while fixing it (reintroducing the same bug in a new shape), the follow-on bash-variable-blowup bug, the newly-discovered TSan segfault variant, and the retry-budget gap, each confirmed by direct reproduction under bash -eo pipefail, not reasoned about. 7. Distilled lessons: ten patterns extracted from the above, written to be applied to a new problem, not just recognized in this one. Every figure cited was checked against BENCH.md/CLAUDE.md's own numbers before writing this, not reproduced from memory. Cross-referenced from CLAUDE.md's "Repository status", README.md's "Contributing" section, and CHANGELOG.md -- including two entries CHANGELOG itself was missing (the set -e CI fix from the previous commit, and this document), the exact class of gap POSTMORTEM.md section 3 already describes happening once before with the v0.1.0 tag. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
9.5 KiB
9.5 KiB
Changelog
All notable changes to this project are documented in this file, in the
style of Keep a Changelog. This
project follows Semantic Versioning; see
PACKFS_VERSION_STRING in include/packfs.h for the authoritative current
version (packfs.pc, generated by make install, is derived from it, not
maintained separately).
This project is pre-release (0.x, an initial, partial implementation of
concept.md, not a completed one; see README.md's "Status"). v0.1.0 is
tagged locally in this repository's own git history but has not been
pushed to any remote — there is no public release yet, only a local one.
[Unreleased]
Fixed
- A failed journal write (disk full, quota, an I/O error) used to be
silently reported as success.
journal_append_recordand its callers werevoidand never checkedfwrite/fflush/fsync's return values, sovfs_write/vfs_mkdir/vfs_unlink/vfs_renameon an overlay-backed file all reported success even when the durable journal append actually failed. Fixed by propagating success/failure through the whole call chain; the affected call now returnsVFS_ERR_IO. - A second bug, found only by reproducing the first fix's edge case
directly: a partial write left a torn record in the journal that made
every later, individually-successful write unrecoverable on reopen
too, not just the failed one —
journal_replaystops at the first record it can't fully read, regardless of what valid records follow it. Fixed by rolling the journal file back to its exact pre-record length whenever a record fails partway. Both bugs covered permanently by the newtests/test_journal_failure.c, which forces a real write failure viaRLIMIT_FSIZErather than a mock. - Neither bug was caught by an existing test — both were found by code review while investigating a question about data-integrity trustworthiness, then confirmed by direct reproduction before and after each fix, not merely inspected and assumed correct.
tests/test_journal_failure.citself leaked memory on its normal, expected code path (vfs_free(v2)was only called on the unreachedelsebranch, in two separate reopen blocks) — invisible locally because local sanitizer runs had been usingASAN_OPTIONS=detect_leaks=0, caught by Gitea CI, which does not set that option. Fixed, and the local/CI sanitizer-invocation mismatch that let it go unnoticed is now called out explicitly inCONTRIBUTING.md. Also fixed a-Wunused-resultwarning on an intentionally-ignoredwrite()return value intests/test_crash_consistency.c's progress side-channel.
Added
tests/test_crash_consistency.c: realfork()+SIGKILLfault injection (not a hand-truncated file) against journaled writes and compaction, with an independent progress side-channel so the test's own oracle can distinguish "never attempted before the kill" from "attempted and lost" — a distinction its own first draft got wrong, reporting 128 false failures before that side-channel was added. Zero real corruption found across dozens of genuine interruptions per run, both scenarios, on repeated runs including under ASan/UBSan.
Documented
- The pack integrity checksum's (FNV-1a64) non-cryptographic nature —
already documented for the compaction-dedup bug — was never explicitly
connected to its other job, load-time pack integrity validation,
where it matters more (
SECURITY.md). - TSan's inability to run is now confirmed in a second, independent
sandboxed environment (a separate remote cloud sandbox, tried
specifically to check whether the restriction was one environment's
fluke), not just this project's own — same root cause
(
personality(ADDR_NO_RANDOMIZE)blocked by seccomp) confirmed both generically and against this project's actual test suite. POSTMORTEM.md: a detailed, permanent record of every real bug and environment issue found across this project's development, how each was actually diagnosed (including false starts), the final fix, and the regression-prevention artifact for each — written so a future session doesn't have to re-diagnose something already answered there.
Fixed (CI)
.gitea/workflows/ci.yml's ThreadSanitizer step was failing CI for an environment reason, but the actual bug was in the CI script, not the code under test. Gitea Actions'run:steps execute underset -e; an unwrappedOUT=$("./binary" 2>&1)aborts the whole script on the first nonzero exit before the step's own classification logic (which exists specifically to not fail the build over the already-documented TSan-can't-start-here limitation) ever runs. Fixed by wrapping every such invocation inif cmd; then RC=0; else RC=$?; fi. Fixing this surfaced two more real issues, both also fixed: capturing unbounded flake output into a bash variable (rather than a file with a bounded read-back) caused a genuine misclassification at scale, and a second, previously undocumented segfault variant of the same TSan-can't-start flake (confirmed via hitting multiple unrelated binaries at random, not a per-binary bug). The ASan/UBSan step's retry budget was also raised from 3 to 5 after a real 3-in-a-row flake exhaustion was observed during this same verification work. SeePOSTMORTEM.md§6.2 for the full diagnosis, including the exact reproduction technique used to confirm each fix (not just reasoned about).- A
-Wunused-resultwarning ontests/test_crash_consistency.c's intentionally-ignoredwrite()return value built clean locally (a(void)cast) but still warned on the Gitea runner's own gcc — a real toolchain configuration difference, confirmed by reproducing clean locally with identical flags first. Fixed with an actual conditional branch on the return value instead, which is portably honored.
[0.1.0] - 2026-09-14
Added
- Initial implementation of
concept.md:vfs_*core,mem/dir/packbackends, the copy-on-write overlay (copy-up, whiteouts, compaction, an append journal), path containment (openat2/Landlock on Linux 5.6+/5.13+, a documented weaker fallback elsewhere), and pack integrity validation. tests/test_*.ccovering each backend, concurrency, and randomized structural-index stress testing against an independent reference model.bench/bench.c(make bench) andBENCH.md, comparing PackFS against the host filesystem across metadata operations, large sequential I/O, random-access pack reads, mount-table scaling, and concurrent workloads.PACKFS_VERSION_MAJOR/MINOR/PATCH/STRINGandpfs_version()for compile-time and runtime version/ABI checks.packfs.pc(pkg-config), generated bymake installfrompackfs.pc.in, version always derived frominclude/packfs.h.SPDX-License-Identifier: MITon every file undersrc/andinclude/.
Fixed
- Bulk sequential
create/unlink/mkdironmem/dirwas O(n²) in file count. The index (UpperSnapshot) was a flat sorted array copied in full on every structural write; rewritten as a persistent treap (src/upper.c), reducing a structural write to the O(log n) nodes on the path to the change. SeeBENCH.md's "Resolution" for the full before/after measurement and complexity-class confirmation. pack_write's compaction-time exact-duplicate elimination (Section 9.2) was O(n²) in entry count, and trusted a hash match without comparing actual bytes (a latent correctness bug: FNV-1a64 is explicitly not collision-resistant). Rewritten as an open-addressing hash table with amemcmpverification before ever reusing adata_off, fixing both at once. SeeBENCH.md's "Resolution #2".backend_pack_newwas declared ininclude/packfs.hbut never implemented — any program calling it failed at link time. Implemented as a standalone, read-onlypackBackend.- The CI config's (
.gitea/workflows/ci.yml) sanitizer-build steps never passed-D_GNU_SOURCEwhen compiling test files (only the library object files got it), which was harmless until a test includedinternal.h(needed forpthread_rwlock_t) and became a link failure.
Changed
- CI moved from a GitHub-Actions-convention config (
.github/workflows/) to Gitea Actions (.gitea/workflows/) — this project is hosted on Gitea, not GitHub, and was never actually hosted on GitHub; only the scaffolding's starting convention changed.SECURITY.md,CLAUDE.md,CONTRIBUTING.md, andREADME.mdupdated to match.
Documented
vfs.c's mount table (MountSnapshot) is O(n²) in mount count, the same pattern the file index used to have — confirmed, and deliberately left unfixed. Mount counts are bounded by a program's own source code, not workload-driven, so they do not reach the scale that made the file index's O(n²) a real problem. SeeBENCH.md's "Finding: mount table scaling" for the numbers and reasoning.- A ThreadSanitizer environment limitation (some sandboxes block the
personality(ADDR_NO_RANDOMIZE)syscall TSan needs) and a related ASan/UBSan sandbox startup flake (AddressSanitizer:DEADLYSIGNAL), both inCONTRIBUTING.mdandCLAUDE.md, with the confirming tests and the required mitigation (timeout-wrapped sanitizer runs).