CI / build-and-test (push) Failing after 33s
Prompted by "fix literally everything still open" after the previous
session's data-integrity work. Went through each open item in turn:
1. TSan: tried a genuinely different execution environment (a remote
cloud sandbox, via a dedicated agent) rather than re-stating the local
sandbox's limitation. Result: identical block there too --
personality(ADDR_NO_RANDOMIZE) returns EPERM, a trivial pthread
program fails TSan identically, and all 6 PackFS test binaries fail
with the same FATAL: ThreadSanitizer: unexpected memory mapping
signature. This is now confirmed in two independent environments, not
one -- strong evidence it's a real infrastructure restriction, not a
one-off fluke worth chasing further with the tools available here.
2. While investigating the "disk-full mid-write" gap flagged as untested
last session, found two real, previously-unknown bugs by reading the
journal code (not by a test catching them unprompted):
- journal_append_record and everything that called it were void, and
none of the fwrite/fflush/fsync calls inside had their return values
checked. A real write failure (disk full, quota, I/O error) was
silently reported as success to vfs_write/vfs_mkdir/vfs_unlink/
vfs_rename -- directly contradicting Section 4.4's premise that
success means durable.
- Fixing that alone was not enough, confirmed by direct reproduction:
a partial write leaves a torn record in the journal, and
journal_replay correctly stops at the first record it can't fully
read (Section 4.3) -- which means every record appended *after* the
torn one, including ones that themselves wrote perfectly fine later,
became silently unreachable on reopen. Reproduced directly before
fixing: a forced-failed write followed by a genuinely successful one
was unrecoverable. Fixed by rolling the journal file back to its
exact pre-record length on any failed write.
Both closed in src/overlay.c (journal_append_record/_put/_delete/
_mkdir/journal_put_current now return and propagate success/failure;
overlay_write/_mkdir/_unlink/_rename return VFS_ERR_IO on a durability
failure without rolling back the already-applied in-memory change,
the same asymmetry a real write()-then-failed-fsync() has). Covered
permanently by the new tests/test_journal_failure.c, which forces a
real failure via RLIMIT_FSIZE + ignored SIGXFSZ, not a mock.
Also fixed in the same pass, found by inspection while touching this
code: journal_put_current used to pass a NULL buffer into a memcpy of
a nonzero size when malloc(size) failed (an OOM-triggered NULL-pointer
dereference) -- closed with an explicit allocation-failure check.
Not test-triggered (forcing malloc() failure portably isn't practical
here); verified by code inspection instead, stated as such rather than
claimed as tested.
3. The remaining "journal-truncation-specific crash window" gap from last
session was investigated, not silently dropped: reliably targeting
that narrow a window would need real concurrency (a second writer
thread racing the kill) for benefit the existing compaction-crash test
already gets probabilistically -- a poor trade, so left as a stated,
deliberate non-goal (CLAUDE.md) rather than built.
4. Cross-process contention is NOT addressed here and should not be read
as an oversight: it is concept.md's own explicit, permanent "not
implemented in v0" scope boundary (a specified-but-unbuilt LMDB-style
reader-table design), not a bug -- building it would be a large,
unrequested feature addition outside this session's actual scope.
Verified: clean make all + make test (all 8 binaries), make bench and
make demo still build and the demo runs correctly end to end, and a full
ASan/UBSan sweep of all 8 binaries with zero real findings (some retries
needed for the already-documented DEADLYSIGNAL flake, which
test_crash_consistency hits more often than other tests simply because it
forks 60+ subprocesses per run -- noted in CONTRIBUTING.md so this isn't
mistaken for a regression later).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
125 lines
6.9 KiB
Markdown
125 lines
6.9 KiB
Markdown
# Changelog
|
|
|
|
All notable changes to this project are documented in this file, in the
|
|
style of [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). This
|
|
project follows [Semantic Versioning](https://semver.org/); see
|
|
`PACKFS_VERSION_STRING` in `include/packfs.h` for the authoritative current
|
|
version (`packfs.pc`, generated by `make install`, is derived from it, not
|
|
maintained separately).
|
|
|
|
This project is pre-release (`0.x`, an initial, partial implementation of
|
|
`concept.md`, not a completed one; see `README.md`'s "Status"). `v0.1.0` is
|
|
tagged locally in this repository's own git history but has not been
|
|
pushed to any remote — there is no public release yet, only a local one.
|
|
|
|
## [Unreleased]
|
|
|
|
### Fixed
|
|
- **A failed journal write (disk full, quota, an I/O error) used to be
|
|
silently reported as success.** `journal_append_record` and its callers
|
|
were `void` and never checked `fwrite`/`fflush`/`fsync`'s return
|
|
values, so `vfs_write`/`vfs_mkdir`/`vfs_unlink`/`vfs_rename` on an
|
|
overlay-backed file all reported success even when the durable journal
|
|
append actually failed. Fixed by propagating success/failure through
|
|
the whole call chain; the affected call now returns `VFS_ERR_IO`.
|
|
- **A second bug, found only by reproducing the first fix's edge case
|
|
directly: a partial write left a torn record in the journal that made
|
|
every later, individually-*successful* write unrecoverable on reopen
|
|
too**, not just the failed one — `journal_replay` stops at the first
|
|
record it can't fully read, regardless of what valid records follow it.
|
|
Fixed by rolling the journal file back to its exact pre-record length
|
|
whenever a record fails partway. Both bugs covered permanently by the
|
|
new `tests/test_journal_failure.c`, which forces a real write failure
|
|
via `RLIMIT_FSIZE` rather than a mock.
|
|
- Neither bug was caught by an existing test — both were found by code
|
|
review while investigating a question about data-integrity
|
|
trustworthiness, then confirmed by direct reproduction before and after
|
|
each fix, not merely inspected and assumed correct.
|
|
|
|
### Added
|
|
- `tests/test_crash_consistency.c`: real `fork()`+`SIGKILL` fault
|
|
injection (not a hand-truncated file) against journaled writes and
|
|
compaction, with an independent progress side-channel so the test's own
|
|
oracle can distinguish "never attempted before the kill" from
|
|
"attempted and lost" — a distinction its own first draft got wrong,
|
|
reporting 128 false failures before that side-channel was added. Zero
|
|
real corruption found across dozens of genuine interruptions per run,
|
|
both scenarios, on repeated runs including under ASan/UBSan.
|
|
|
|
### Documented
|
|
- The pack integrity checksum's (FNV-1a64) non-cryptographic nature —
|
|
already documented for the compaction-dedup bug — was never explicitly
|
|
connected to its *other* job, load-time pack integrity validation,
|
|
where it matters more (`SECURITY.md`).
|
|
- TSan's inability to run is now confirmed in a *second*, independent
|
|
sandboxed environment (a separate remote cloud sandbox, tried
|
|
specifically to check whether the restriction was one environment's
|
|
fluke), not just this project's own — same root cause
|
|
(`personality(ADDR_NO_RANDOMIZE)` blocked by seccomp) confirmed both
|
|
generically and against this project's actual test suite.
|
|
|
|
## [0.1.0] - 2026-09-14
|
|
|
|
### Added
|
|
- Initial implementation of `concept.md`: `vfs_*` core, `mem`/`dir`/`pack`
|
|
backends, the copy-on-write overlay (copy-up, whiteouts, compaction, an
|
|
append journal), path containment (`openat2`/Landlock on Linux 5.6+/5.13+,
|
|
a documented weaker fallback elsewhere), and pack integrity validation.
|
|
- `tests/test_*.c` covering each backend, concurrency, and randomized
|
|
structural-index stress testing against an independent reference model.
|
|
- `bench/bench.c` (`make bench`) and `BENCH.md`, comparing PackFS against
|
|
the host filesystem across metadata operations, large sequential I/O,
|
|
random-access pack reads, mount-table scaling, and concurrent workloads.
|
|
- `PACKFS_VERSION_MAJOR`/`MINOR`/`PATCH`/`STRING` and `pfs_version()` for
|
|
compile-time and runtime version/ABI checks.
|
|
- `packfs.pc` (pkg-config), generated by `make install` from
|
|
`packfs.pc.in`, version always derived from `include/packfs.h`.
|
|
- `SPDX-License-Identifier: MIT` on every file under `src/` and
|
|
`include/`.
|
|
|
|
### Fixed
|
|
- **Bulk sequential `create`/`unlink`/`mkdir` on `mem`/`dir` was O(n²) in
|
|
file count.** The index (`UpperSnapshot`) was a flat sorted array copied
|
|
in full on every structural write; rewritten as a persistent treap
|
|
(`src/upper.c`), reducing a structural write to the O(log n) nodes on
|
|
the path to the change. See `BENCH.md`'s "Resolution" for the full
|
|
before/after measurement and complexity-class confirmation.
|
|
- **`pack_write`'s compaction-time exact-duplicate elimination (Section
|
|
9.2) was O(n²) in entry count, and trusted a hash match without
|
|
comparing actual bytes** (a latent correctness bug: FNV-1a64 is
|
|
explicitly not collision-resistant). Rewritten as an open-addressing
|
|
hash table with a `memcmp` verification before ever reusing a
|
|
`data_off`, fixing both at once. See `BENCH.md`'s "Resolution #2".
|
|
- **`backend_pack_new` was declared in `include/packfs.h` but never
|
|
implemented** — any program calling it failed at link time. Implemented
|
|
as a standalone, read-only `pack` `Backend`.
|
|
- The CI config's (`.gitea/workflows/ci.yml`) sanitizer-build steps never
|
|
passed `-D_GNU_SOURCE` when compiling test files (only the library
|
|
object files got it), which was harmless until a test included
|
|
`internal.h` (needed for `pthread_rwlock_t`) and became a link failure.
|
|
|
|
### Changed
|
|
- CI moved from a GitHub-Actions-convention config (`.github/workflows/`)
|
|
to Gitea Actions (`.gitea/workflows/`) — this project is hosted on
|
|
Gitea, not GitHub, and was never actually hosted on GitHub; only the
|
|
scaffolding's starting convention changed. `SECURITY.md`, `CLAUDE.md`,
|
|
`CONTRIBUTING.md`, and `README.md` updated to match.
|
|
|
|
### Documented
|
|
- **`vfs.c`'s mount table (`MountSnapshot`) is O(n²) in mount count, the
|
|
same pattern the file index used to have — confirmed, and deliberately
|
|
left unfixed.** Mount counts are bounded by a program's own source code,
|
|
not workload-driven, so they do not reach the scale that made the file
|
|
index's O(n²) a real problem. See `BENCH.md`'s "Finding: mount table
|
|
scaling" for the numbers and reasoning.
|
|
- A ThreadSanitizer environment limitation (some sandboxes block the
|
|
`personality(ADDR_NO_RANDOMIZE)` syscall TSan needs) and a related
|
|
ASan/UBSan sandbox startup flake (`AddressSanitizer:DEADLYSIGNAL`),
|
|
both in `CONTRIBUTING.md` and `CLAUDE.md`, with the confirming tests and
|
|
the required mitigation (`timeout`-wrapped sanitizer runs).
|
|
|
|
<!-- This repository has no remote configured yet (see `git remote -v`),
|
|
so there is no comparison-link URL to cite here truthfully; add
|
|
"[0.1.0]: <repo-url>/releases/tag/v0.1.0" once one exists, per Keep
|
|
a Changelog's own convention. -->
|