Add real fault-injection crash-safety testing; investigate TSan for real
CI / build-and-test (push) Failing after 38s
CI / build-and-test (push) Failing after 38s
Prompted by a direct question about data-integrity trustworthiness: the crash-safety design (self-checking journal records, atomic-rename compaction, Section 4.3/4.4) had never actually been tested against a real crash -- only reasoned about statically and tested against an already-corrupted file (test_pack_overlay.c, a different scenario). Added tests/test_crash_consistency.c: real fork()+SIGKILL fault injection against an actual child process, not a hand-truncated file standing in for a crash, with empirically-calibrated kill-delay sampling (measured via throwaway scripts, not guessed) so trials land genuine mid-operation interruptions rather than always completing first: - Journal-write crash injection (25 trials): kills a child mid-burst of individually-journaled, individually-fsynced writes; verifies a strict clean-prefix recovery (every completed write present and correct, every write after the kill cleanly absent, no gaps). 20-22/25 trials per run land a genuine interruption; zero corruption found. - Compaction crash injection (40 trials): kills a child mid-vfs_sync; verifies every write durably journaled *before* vfs_sync was called survives regardless of whether compaction itself completed. 40/40 trials per run land a genuine interruption; zero corruption found. A real finding from building this test, not from the code under test: the first draft reported 128 failures, every one a bug in the test itself -- it couldn't distinguish "the child was killed before ever attempting this write" from "the write completed and was then lost," since SIGKILL can't be caught to report progress. Fixed with an independent progress side-channel (plain write()+fsync() on a separate file, entirely outside packfs) as ground truth. Documented in CLAUDE.md's new "Known reliability characteristics" section, including this finding, because a claim is only as trustworthy as what's measuring it. Also investigated whether this sandbox's TSan block is actually unfixable, rather than re-asserting the known limitation: tried personality(ADDR_NO_RANDOMIZE) directly (EPERM), setarch -R (fails identically), searched TSAN_OPTIONS for a bypass (none exists), and checked capabilities/seccomp/unshare --user (zero effective caps, active seccomp filter, namespace escape also blocked). Confirmed categorical and documented as such in CLAUDE.md/CONTRIBUTING.md, rather than left as an unexamined "it doesn't work here." Also closed a real documentation gap SECURITY.md had: the pack integrity checksum's non-cryptographic nature (FNV-1a64, already documented in the context of the compaction-dedup bug) had never been explicitly connected to its *other* use, load-time pack integrity validation -- added, since it's a real, relevant caveat for anyone relying on that checksum against a deliberate adversary rather than accidental corruption. Verified: clean make all + make test (all 7 binaries), full ASan/UBSan sweep of all 7 binaries with zero real findings, and the new test run standalone 4+ times (this session) with stable, consistent results. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
This commit is contained in:
@@ -19,7 +19,7 @@ A single test: `make build/test_mem && ./build/test_mem` (substitute any `test_*
|
||||
|
||||
Sanitizer builds are not wired into `make test` (they need per-file compilation with sanitizer flags plus `-D_GNU_SOURCE -Iinclude -Isrc`); see `.gitea/workflows/ci.yml` (Gitea Actions) for the exact invocation, which CI runs on every push. **Any change to `upper.c`, `overlay.c`, or `vfs.c` must be verified under `-fsanitize=address,undefined` and `-fsanitize=thread` before being considered done** — this is not a formality: exactly this process caught a real use-after-free in the snapshot-reclamation logic during initial development (see the `reclaim_gate` note below), which the plain build and even repeated plain test runs never surfaced.
|
||||
|
||||
**If your environment cannot run ThreadSanitizer at all, say so rather than skipping it silently.** TSan needs `personality(ADDR_NO_RANDOMIZE)` to disable ASLR for itself; some sandboxes block that syscall outright, in which case *every* TSan build fails identically (`FATAL: ThreadSanitizer: unexpected memory mapping`), including a trivial unrelated pthread program — that is the confirming test, not a PackFS-specific symptom. **This project's own CI runner is one of them, confirmed by its first real run** (see `CONTRIBUTING.md` for the details and for exactly how `.gitea/workflows/ci.yml`'s TSan step now tells that specific, known failure apart from a real finding rather than either hiding it or failing the build over it) — a green CI run is therefore not proof TSan executed. In that situation, ASan/UBSan are still run and still matter (they are what actually caught the `reclaim_gate` bug above), but they do not perform TSan's happens-before race analysis and are not a substitute for it; a change is TSan-verified only once it has actually passed on a machine confirmed able to execute it, never merely because ASan/UBSan passed or because a CI step reported success.
|
||||
**If your environment cannot run ThreadSanitizer at all, say so rather than skipping it silently.** TSan needs `personality(ADDR_NO_RANDOMIZE)` to disable ASLR for itself; some sandboxes block that syscall outright, in which case *every* TSan build fails identically (`FATAL: ThreadSanitizer: unexpected memory mapping`), including a trivial unrelated pthread program — that is the confirming test, not a PackFS-specific symptom. **This project's own CI runner is one of them, confirmed by its first real run** (see `CONTRIBUTING.md` for the details and for exactly how `.gitea/workflows/ci.yml`'s TSan step now tells that specific, known failure apart from a real finding rather than either hiding it or failing the build over it) — a green CI run is therefore not proof TSan executed. In that situation, ASan/UBSan are still run and still matter (they are what actually caught the `reclaim_gate` bug above), but they do not perform TSan's happens-before race analysis and are not a substitute for it; a change is TSan-verified only once it has actually passed on a machine confirmed able to execute it, never merely because ASan/UBSan passed or because a CI step reported success. **This has been actively investigated, not just accepted at face value:** in this sandbox, `personality(ADDR_NO_RANDOMIZE)` itself returns `EPERM` when called directly (confirmed via a raw syscall, not just observed via TSan's own error); `setarch -R` (which just calls the same syscall) fails identically; there is no `TSAN_OPTIONS` flag that relaxes the startup memory-mapping check (checked against the runtime's own `help=1` option listing); the process holds zero effective capabilities (`CapEff` all-zero) under an active seccomp filter (`Seccomp_filters: 1`) that also blocks `unshare --user`, ruling out working around it via a nested, less-restricted namespace. This is a categorical, syscall-level restriction with no userspace workaround available in this environment, not a configuration this project has simply failed to find yet.
|
||||
|
||||
**A related but separate flake affects ASan/UBSan themselves in that kind of sandbox, not just TSan:** a sanitizer-built test binary can non-deterministically fail to start with `AddressSanitizer:DEADLYSIGNAL`, occasionally as an unbounded repeating loop rather than a single line — a sandbox startup race, confirmed by it hitting different, unrelated binaries across repeated runs, each of which then passes cleanly on retry. See `CONTRIBUTING.md`'s workflow section for the full description and the required mitigation (`timeout`-wrap sanitizer runs in such an environment; treat `DEADLYSIGNAL` alone, without an actual `ERROR: AddressSanitizer` or `runtime error:` string, as inconclusive and re-run rather than as a finding).
|
||||
|
||||
@@ -34,7 +34,7 @@ Sanitizer builds are not wired into `make test` (they need per-file compilation
|
||||
- `src/containment.c` — `dir`-mount path containment (`openat2`/`O_NOFOLLOW` fallback) and opt-in Landlock hardening.
|
||||
- `src/path.c` — virtual-namespace `.`/`..` canonicalization.
|
||||
- `src/hash.c` — FNV-1a64, used for pack checksums and compaction dedup.
|
||||
- `tests/` — one binary per concern (`test_mem`, `test_dir`, `test_pack_overlay`, `test_concurrency`, `test_index_stress` — the treap's correctness under thousands of randomized structural operations, cross-checked against an independent reference model); `test_harness.h` is a small assertion-macro header, not a framework, consistent with the zero-dependency constraint.
|
||||
- `tests/` — one binary per concern (`test_mem`, `test_dir`, `test_pack_overlay`, `test_concurrency`, `test_index_stress` — the treap's correctness under thousands of randomized structural operations, cross-checked against an independent reference model; `test_pack_write_perf` — the `pack_write` dedup O(n²) regression tripwire; `test_crash_consistency` — real fork()+`SIGKILL` fault injection against a real process, not a hand-truncated file, verifying Section 4.3/4.4's crash-safety claims for both journaled writes and compaction, with an independent progress side-channel so the oracle can tell "never attempted before the kill" apart from "attempted and lost" — see its own file header before touching journal or compaction code); `test_harness.h` is a small assertion-macro header, not a framework, consistent with the zero-dependency constraint.
|
||||
|
||||
**One architectural fact spans every mutable structure and is easy to miss reading any single file in isolation:** `MountSnapshot` (`vfs.c`), `UpperSnapshot` (`upper.c`), and a `mem`-backed `MutCell`'s buffer (`upper.c`, Section 5.7) are all retired through the same two-step pattern — publish the replacement via `atomic_store_explicit(..., memory_order_release)`, then free the superseded object only under a dedicated `reclaim_gate` rwlock's *write* side, while every acquirer takes that same gate's *read* side around its own load-then-increment. This exists because a plain "load a pointer, then atomically increment its refcount" leaves a real gap between those two steps in which a concurrent writer can free the very object being acquired — confirmed by ASan as an actual heap-use-after-free during development, not a theoretical concern. See the comment on `UpperStore.reclaim_gate` in `internal.h` for the full reasoning. Any new refcounted, concurrently-reclaimed structure added to this codebase needs the same gate, not just an atomic pointer and a naive refcount. The one exception, and the reason it's an exception rather than a hole: `TreapNode`'s own per-node refcounting (`upper.c`) does *not* need its own `reclaim_gate`, because nothing ever acquires a `TreapNode*` independently the way `upper_acquire` acquires an `UpperSnapshot*` — a reader only ever reaches tree nodes by walking `s->root` after already holding a valid `UpperSnapshot` reference, which transitively keeps the whole tree alive for the reader's purposes; `node_ref`/`node_unref` are called only by the single writer (under `writer_lock`) and by `upper_release`'s cascade (already inside `reclaim_gate`'s protection at the snapshot level), never by a reader racing a writer the way the snapshot pointer itself is raced.
|
||||
|
||||
@@ -95,6 +95,17 @@ These are hard requirements stated in the spec, not stylistic suggestions — an
|
||||
- **RESOLVED: `pack_write`'s compaction-time exact-duplicate elimination (Section 9.2) was O(n²) in entry count, plus a latent correctness bug, both fixed together.** The dedup check was a linear scan of every previously-seen `(hash, size)` pair per entry, comparing only the hash and size — never the actual bytes — before trusting a match, so two different-content entries that happened to collide on `(hash, size)` under FNV-1a64 (explicitly not collision-resistant, `src/hash.c`) would have been silently merged, corrupting one of them; this was never observed in practice but was true of the code as written. `bench/bench.c`'s own compaction benchmark never surfaced either problem because its test files are byte-identical, which made every scan match on the first comparison (the scan's best case, not its worst case). Measuring with unique content instead (`pack_write` called directly, bypassing the journal/fsync path) showed 80,000 entries taking 1.74s, with a 20,000→80,000 (4x N) step showing 15.6x — matching O(n²)'s 16x prediction. Fixed with an open-addressing hash table (`DedupSlot`, load factor 1/2, linear probing) plus a `memcmp` verification before ever reusing a `data_off`, closing both the complexity issue and the correctness gap at once; post-fix, the same 80,000-entry case takes 0.044s (39.6x faster) and the 20,000→80,000 ratio drops to 3.35x, consistent with O(n). `BENCH.md`'s "Resolution #2" has the full before/after numbers. Correctness is covered permanently by `tests/test_pack_overlay.c` (asserts duplicate-content entries share one `data_off` and distinct-content entries do not); the performance fix is covered permanently by `tests/test_pack_write_perf.c` (a regression tripwire against 10,000 unique entries, run as part of `make test`).
|
||||
- **NOT FIXED, by deliberate decision: `vfs.c`'s mount table (`MountSnapshot`) is O(n²) in mount count, using the same full-array-copy-per-write pattern the file index used to.** Confirmed via `bench/bench.c` category 10 (500/2,000/8,000 mounts, both 4x-N steps showing 15–20x, matching O(n²)'s 16x prediction, for mount, resolve, and unmount alike). Unlike the file index, this is not being converted to a persistent tree: mount points are created by calls written into a program's own source code, not driven by workload or user data, so realistic mount counts do not reach the scale that made the file index's O(n²) an actual problem — `concept.md` Section 5.3's own stated trigger for needing a persistent structure ("not required until this assumption is empirically violated") has not been violated here the way it was for the file index. `BENCH.md`'s "Finding: mount table scaling" has the full numbers and reasoning; revisit only if a real use case needs thousands of runtime mount points (e.g. one mount per tenant).
|
||||
|
||||
## Known reliability characteristics
|
||||
|
||||
Section 4.3/4.4's crash-safety claims (self-checking journal records; atomic-rename compaction that leaves the existing pack + journal untouched on failure) were, until now, verified only by static reasoning about the design and by `test_pack_overlay.c`'s test of loading an *already*-corrupted pack file — never by actually killing a process mid-write and checking what a fresh reopen recovers. `tests/test_crash_consistency.c` closes that gap with real fault injection (`fork()` + `SIGKILL` at randomized, empirically-calibrated delays against an actual child process, not a hand-truncated file standing in for a crash):
|
||||
|
||||
- **Journal-write crash injection** (25 trials, ~6ms/journaled-write on this build, delays sampled across the whole burst): kills a child mid-way through a sequence of individually-journaled, individually-fsynced writes. Verifies a strict "clean prefix" after reopening — every write the child actually completed is present with exact content, every write after the kill is cleanly absent, no gaps. 20–22 of 25 trials land a genuine interruption (`WIFSIGNALED`, not `WIFEXITED`) each run; **zero corruption or gaps found**.
|
||||
- **Compaction crash injection** (40 trials, delays sampled across a calibrated ~0.08s window covering pre-compaction durable writes through post-rename): kills a child mid-`vfs_sync`. Verifies that every write durably journaled *before* `vfs_sync` was even called survives regardless of whether compaction itself completed, and that reopening never fails. Uses an independent progress side-channel (a plain `write()`+`fsync()` on a dedicated file, entirely outside packfs) as ground truth for how many writes the child actually completed before dying — without it, a write the child simply never reached in time is indistinguishable from one that completed and was then lost, which is exactly the bug the test itself had on its first draft (see below). 40/40 trials land a genuine interruption each run; **zero corruption or loss found**.
|
||||
|
||||
**A real finding from building this test, recorded here per this file's own "document literally all" standard: the test's first draft reported 128 failures, and every one of them was a bug in the test, not in PackFS.** It checked "is extra write *i* present after the crash" without distinguishing "the child was killed before it ever attempted write *i*" (expected, not a bug) from "write *i* completed and was then lost" (would be a real bug) — SIGKILL cannot be caught, so the child has no way to report its own progress through the normal API. Fixed by adding the independent progress side-channel described above; the corrected test has passed cleanly across every run so far (multiple full runs, plus 3 clean runs under ASan/UBSan). This is the same category of lesson as `BENCH.md`'s own history of findings: a claim (here, "PackFS's crash safety" — there, "the benchmark's numbers") is only as trustworthy as the thing measuring it, and that has to be checked too, not just the subject under test.
|
||||
|
||||
**Still not covered, stated plainly:** these two scenarios don't exercise cross-process contention (two processes crashing/racing against the same files — not implemented in v0, see "Load-bearing constraints" above), disk-full/`ENOSPC` mid-write, or a crash during journal *truncation* specifically (the second atomic-rename step after a successful compaction, `.jnl.tmp` → `.jnl`) as its own isolated case. TSan remains unable to run in this sandbox for the concurrency side of reliability — see the sanitizer section above, now including the specific workarounds tried (not just the limitation) before concluding it's genuinely unfixable here.
|
||||
|
||||
## Self-evaluation
|
||||
|
||||
**Methodology.** This file was checked against the repository's actual state (`make test` passing, including two new tests, `test_pack_write_perf` and the dedup-verification block added to `test_pack_overlay`; `include/`, `src/`, `tests/`, `bench/` present and matching the description below; confirmed by directory listing and a live build), against `concept.md` as frozen (for factual consistency of the constraints and architecture summarized above), and against the documentation standard stated in this file. This revision followed a deliberate audit for other instances of the file index's O(n²) shape elsewhere in the codebase — prompted by a request to leave no caveats undocumented — which found two more real issues, not zero: `pack_write`'s compaction dedup (O(n²), plus a latent hash-collision correctness bug, both now fixed) and the mount table (O(n²), confirmed and deliberately left unfixed, with the reasoning recorded). The audit also found and fixed a CI gap the new tests exposed — the CI config's (now `.gitea/workflows/ci.yml`; at the time of this finding it lived at `.github/workflows/ci.yml`, before this project's scaffolding was set up for Gitea hosting instead of the GitHub-Actions convention it started from — this repository was never actually hosted on GitHub) sanitizer-build steps never passed `-D_GNU_SOURCE` when compiling test files, which was harmless while no test included `internal.h` and became a real link failure once two did (`internal.h` needs it for `pthread_rwlock_t`) — and documented a sandbox flake (ASan/UBSan's own `DEADLYSIGNAL` startup race, distinct from and in addition to the already-documented TSan limitation) observed directly during this audit's own sanitizer runs, in `CONTRIBUTING.md` and cross-referenced here. Every finding below was confirmed empirically (direct `pack_write` timing, `bench/bench.c`'s new mount-scaling category, repeated sanitizer runs with `timeout`), not asserted.
|
||||
|
||||
+22
-1
@@ -38,10 +38,31 @@ rather than a generic "we take security seriously":
|
||||
externally-sourced pack additionally needs a checksum verified at load.
|
||||
Loading an attacker-supplied `.img` file that PackFS accepts without
|
||||
detecting corruption, or that causes an out-of-bounds read/write, is a
|
||||
security bug.
|
||||
security bug. **That checksum (FNV-1a64, `src/hash.c`) is explicitly
|
||||
non-cryptographic and not collision-resistant** — stated plainly here
|
||||
because this is the one place in this document where it matters most:
|
||||
it is a real, effective check against *accidental* corruption (a
|
||||
truncated copy, a bit flip, a bad transfer), but it is not a tamper-
|
||||
evident guarantee against a deliberate adversary who can choose the
|
||||
bytes of a `.img` file — someone with that capability could in
|
||||
principle construct a corrupted pack whose checksum still matches. Do
|
||||
not treat "checksum verified at load" as a substitute for verifying the
|
||||
*source* of an externally-supplied pack through some other channel if
|
||||
the threat model includes a deliberate adversary, not just accidental
|
||||
corruption.
|
||||
- **Journal records are length/checksum self-describing** (Section 4.4),
|
||||
so replay stops cleanly at the first torn or corrupted record rather
|
||||
than reading past it.
|
||||
- **The crash-safety design above is verified by actual fault injection,
|
||||
not only by reasoning about the design.** `tests/test_crash_consistency.c`
|
||||
forks a real child process and `SIGKILL`s it at randomized points during
|
||||
journaled writes and during compaction, then checks what a fresh reopen
|
||||
recovers, across dozens of trials per run — not a hand-truncated file
|
||||
standing in for a crash. See `CLAUDE.md`'s "Known reliability
|
||||
characteristics" for the methodology and results, including a real bug
|
||||
the *test itself* had on its first draft (conflating "never attempted
|
||||
before the kill" with "lost after completing") — fixed before treating
|
||||
the crash-safety claim as verified.
|
||||
|
||||
## What PackFS explicitly does *not* claim
|
||||
|
||||
|
||||
@@ -0,0 +1,401 @@
|
||||
/*
|
||||
* test_crash_consistency.c — real fault injection for Section 4.3/4.4's
|
||||
* crash-safety claims (self-checking journal records, atomic-rename
|
||||
* compaction), via an actual fork()+SIGKILL of a child mid-operation and
|
||||
* a fresh reopen afterward, not a hand-truncated file standing in for a
|
||||
* crash. The two scenarios below (journal writes, compaction) were the
|
||||
* two concrete claims CLAUDE.md/concept.md make about crash behavior;
|
||||
* before this file, neither had ever actually been exercised by killing
|
||||
* a real process mid-write — only by static reasoning about the design
|
||||
* and by test_pack_overlay.c's separate (unrelated) test of loading an
|
||||
* already-corrupted pack file.
|
||||
*
|
||||
* Timing constants below are empirically calibrated for this project's
|
||||
* own build (see the comment above each scenario), not arbitrary: the
|
||||
* goal is to sample kill delays across the real duration of the
|
||||
* operation under test, confirmed per-run by reporting how many trials
|
||||
* actually landed a genuine mid-operation kill (WIFSIGNALED) versus how
|
||||
* many the child outran (WIFEXITED) — a run reporting zero interrupted
|
||||
* trials would mean this test verified nothing and should be treated as
|
||||
* a test bug, not a pass; see the CHECK on INTERRUPTED_MIN below.
|
||||
*/
|
||||
|
||||
#include <fcntl.h>
|
||||
#include <signal.h>
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <sys/wait.h>
|
||||
#include <time.h>
|
||||
#include <unistd.h>
|
||||
|
||||
#include "packfs.h"
|
||||
#include "test_harness.h"
|
||||
|
||||
static void sleep_sec(double s) {
|
||||
struct timespec ts;
|
||||
ts.tv_sec = (time_t)s;
|
||||
ts.tv_nsec = (long)((s - (double)ts.tv_sec) * 1e9);
|
||||
nanosleep(&ts, NULL);
|
||||
}
|
||||
|
||||
/* fork the given child function, let it run for kill_delay seconds, then
|
||||
* SIGKILL it. Returns 1 if the child was genuinely still running (killed
|
||||
* by the signal), 0 if it had already exited on its own. */
|
||||
static int run_and_kill(void (*child_fn)(void *), void *arg, double kill_delay) {
|
||||
pid_t pid = fork();
|
||||
if (pid < 0) { perror("fork"); exit(1); }
|
||||
if (pid == 0) {
|
||||
child_fn(arg);
|
||||
_exit(99); /* child_fn must _exit itself; this is a safety net */
|
||||
}
|
||||
sleep_sec(kill_delay);
|
||||
kill(pid, SIGKILL);
|
||||
int status = 0;
|
||||
waitpid(pid, &status, 0);
|
||||
return WIFSIGNALED(status) && WTERMSIG(status) == SIGKILL;
|
||||
}
|
||||
|
||||
/* ================= scenario A: kill mid-burst of journaled writes ===== */
|
||||
/*
|
||||
* Calibrated: ~6ms per individual journaled create+write+close on this
|
||||
* build/environment (fopen+fwrite+fflush+fsync per record, Section 4.4).
|
||||
* BURST=60 writes takes ~0.36s; sampling kill delays across [0, 0.40s]
|
||||
* covers the whole burst including its very start and its tail.
|
||||
*/
|
||||
#define BURST_N 60
|
||||
#define BURST_TRIALS 25
|
||||
#define BURST_MAX_DELAY 0.40
|
||||
|
||||
typedef struct { char pack_path[512]; } BurstArg;
|
||||
|
||||
static void mk_burst_content(char *buf, size_t n, int i) {
|
||||
snprintf(buf, n, "burst-content-%06d", i);
|
||||
}
|
||||
|
||||
static void burst_child(void *arg_) {
|
||||
BurstArg *arg = (BurstArg *)arg_;
|
||||
Vfs *v = vfs_new();
|
||||
Backend *mem = backend_mem_new();
|
||||
int oerr = 0;
|
||||
Backend *ov = backend_overlay_new(arg->pack_path, mem, &oerr);
|
||||
if (!ov) _exit(2);
|
||||
if (vfs_mount(v, "/", ov) != VFS_OK) _exit(3);
|
||||
|
||||
char name[64], content[64];
|
||||
for (int i = 0; i < BURST_N; i++) {
|
||||
snprintf(name, sizeof(name), "/f%05d.txt", i);
|
||||
mk_burst_content(content, sizeof(content), i);
|
||||
int err = 0;
|
||||
VfsFile *f = vfs_open(v, name, VFS_O_WRONLY | VFS_O_CREAT, &err);
|
||||
if (!f) _exit(4);
|
||||
if (vfs_write(f, content, strlen(content)) != (pfs_isize)strlen(content)) _exit(5);
|
||||
if (vfs_close(f) != VFS_OK) _exit(6);
|
||||
}
|
||||
_exit(0); /* all BURST_N writes durably committed */
|
||||
}
|
||||
|
||||
static void test_journal_burst_crash(void) {
|
||||
BurstArg arg;
|
||||
snprintf(arg.pack_path, sizeof(arg.pack_path), "/tmp/packfs_test_crash_burst_%d.img", (int)getpid());
|
||||
char jpath[600];
|
||||
snprintf(jpath, sizeof(jpath), "%s.jnl", arg.pack_path);
|
||||
|
||||
int interrupted_count = 0;
|
||||
for (int trial = 0; trial < BURST_TRIALS; trial++) {
|
||||
unlink(arg.pack_path);
|
||||
unlink(jpath);
|
||||
double delay = (BURST_MAX_DELAY * (double)trial) / (double)(BURST_TRIALS - 1);
|
||||
int interrupted = run_and_kill(burst_child, &arg, delay);
|
||||
if (interrupted) interrupted_count++;
|
||||
|
||||
Vfs *v2 = vfs_new();
|
||||
Backend *mem2 = backend_mem_new();
|
||||
int oerr2 = 0;
|
||||
Backend *ov2 = backend_overlay_new(arg.pack_path, mem2, &oerr2);
|
||||
CHECK(ov2 != NULL); /* a crash mid-journal must never make reopening fail */
|
||||
if (!ov2) { vfs_free(v2); backend_free(mem2); continue; }
|
||||
CHECK_EQ_INT(vfs_mount(v2, "/", ov2), VFS_OK);
|
||||
|
||||
/* clean-prefix invariant: every present file has exactly correct
|
||||
* content, and there is no gap (file i missing, file i+1 present) */
|
||||
int last_present = -1;
|
||||
char name[64], expect[64], buf[64];
|
||||
for (int i = 0; i < BURST_N; i++) {
|
||||
snprintf(name, sizeof(name), "/f%05d.txt", i);
|
||||
int err = 0;
|
||||
VfsFile *f = vfs_open(v2, name, VFS_O_RDONLY, &err);
|
||||
if (!f) continue;
|
||||
mk_burst_content(expect, sizeof(expect), i);
|
||||
memset(buf, 0, sizeof(buf));
|
||||
pfs_isize n = vfs_read(f, buf, sizeof(buf));
|
||||
vfs_close(f);
|
||||
if (n != (pfs_isize)strlen(expect) || memcmp(buf, expect, strlen(expect)) != 0) {
|
||||
fprintf(stderr, "FAIL trial %d: %s has wrong/corrupt content after crash (delay=%.4f)\n",
|
||||
trial, name, delay);
|
||||
pfs_test_failures++;
|
||||
}
|
||||
if (i != last_present + 1) {
|
||||
fprintf(stderr, "FAIL trial %d: gap in journal replay prefix at %s (delay=%.4f)\n",
|
||||
trial, name, delay);
|
||||
pfs_test_failures++;
|
||||
}
|
||||
last_present = i;
|
||||
}
|
||||
|
||||
vfs_unmount(v2, "/");
|
||||
backend_free(ov2);
|
||||
backend_free(mem2);
|
||||
vfs_free(v2);
|
||||
}
|
||||
|
||||
unlink(arg.pack_path);
|
||||
unlink(jpath);
|
||||
fprintf(stderr, "journal-burst crash test: %d/%d trials genuinely interrupted mid-burst\n",
|
||||
interrupted_count, BURST_TRIALS);
|
||||
/* if this is ever 0, the delay schedule no longer matches this
|
||||
* machine's write latency and the test isn't exercising the crash
|
||||
* path at all -- that is itself a failure, not a quiet pass. */
|
||||
CHECK(interrupted_count >= BURST_TRIALS / 4);
|
||||
}
|
||||
|
||||
/* ================= scenario B: kill mid-compaction ===================== */
|
||||
/*
|
||||
* Calibrated: a 100-file, 2KB-payload durable baseline (built once,
|
||||
* untimed, then compacted once to get a clean golden pack.img) plus a
|
||||
* per-trial 10-file durable "extras" batch takes ~0.058s to journal and
|
||||
* the following compaction takes ~0.012s on this build/environment.
|
||||
* Sampling kill delays across [0, 0.08s] covers extras-journaling,
|
||||
* compaction start, mid-compaction, and post-rename.
|
||||
*
|
||||
* The invariant under test is not "pre-state XOR post-state" -- it's
|
||||
* simpler and is the actual documented guarantee (concept.md's crash-
|
||||
* safety constraints, CLAUDE.md's "if writing pack.img.tmp or its fsync
|
||||
* fails, compaction aborts and the existing pack + journal are
|
||||
* untouched"): anything durably journaled (fsynced) *before* vfs_sync is
|
||||
* even called must survive a crash during that vfs_sync call, no matter
|
||||
* where in it the crash lands -- either via journal replay (if
|
||||
* compaction didn't finish) or by being included in the fresh pack (if
|
||||
* it did).
|
||||
*/
|
||||
#define GOLDEN_N 100
|
||||
#define EXTRA_N 10
|
||||
#define PAYLOAD_SIZE 2048
|
||||
#define COMPACT_TRIALS 40
|
||||
#define COMPACT_MAX_DELAY 0.08
|
||||
|
||||
typedef struct { char pack_path[512]; char golden_path[512]; char progress_path[512]; } CompactArg;
|
||||
|
||||
static void mk_payload(char *buf, size_t n, const char *tag, int i) {
|
||||
int len = snprintf(buf, n, "%s-%06d-", tag, i);
|
||||
for (size_t j = (size_t)len; j < n; j++) buf[j] = (char)('a' + (int)(j % 26));
|
||||
}
|
||||
|
||||
static int copy_file(const char *from, const char *to) {
|
||||
FILE *in = fopen(from, "rb");
|
||||
if (!in) return -1;
|
||||
FILE *out = fopen(to, "wb");
|
||||
if (!out) { fclose(in); return -1; }
|
||||
char buf[65536];
|
||||
size_t n;
|
||||
while ((n = fread(buf, 1, sizeof(buf), in)) > 0) fwrite(buf, 1, n, out);
|
||||
fclose(in);
|
||||
fclose(out);
|
||||
return 0;
|
||||
}
|
||||
|
||||
/* Records "N extras durably completed" via a plain POSIX write+fsync on a
|
||||
* dedicated file, entirely independent of packfs. This is the test's
|
||||
* ground truth for how far the child actually got before SIGKILL landed
|
||||
* -- without it, an extra that the child simply never reached in time
|
||||
* (expected, not a bug) is indistinguishable from one that completed and
|
||||
* was then lost (a real bug), since SIGKILL cannot be caught to report
|
||||
* progress any other way. The one imprecision this leaves: the tiny gap
|
||||
* between vfs_close() returning (extra i durably in the journal) and
|
||||
* this progress write's own fsync landing can make the progress count
|
||||
* undercount by one extra in the worst case -- which only makes the
|
||||
* check *less* strict at the margin (skips verifying the most recent
|
||||
* extra), never produces a false failure. */
|
||||
static void record_progress(const char *progress_path, int n_done) {
|
||||
int fd = open(progress_path, O_WRONLY | O_CREAT | O_TRUNC, 0644);
|
||||
if (fd < 0) return;
|
||||
char buf[16];
|
||||
int len = snprintf(buf, sizeof(buf), "%d", n_done);
|
||||
write(fd, buf, (size_t)len);
|
||||
fsync(fd);
|
||||
close(fd);
|
||||
}
|
||||
|
||||
static int read_progress(const char *progress_path) {
|
||||
FILE *fp = fopen(progress_path, "r");
|
||||
if (!fp) return 0;
|
||||
int n = 0;
|
||||
if (fscanf(fp, "%d", &n) != 1) n = 0;
|
||||
fclose(fp);
|
||||
return n;
|
||||
}
|
||||
|
||||
static void compact_child(void *arg_) {
|
||||
CompactArg *arg = (CompactArg *)arg_;
|
||||
record_progress(arg->progress_path, 0);
|
||||
Vfs *v = vfs_new();
|
||||
Backend *mem = backend_mem_new();
|
||||
int oerr = 0;
|
||||
Backend *ov = backend_overlay_new(arg->pack_path, mem, &oerr);
|
||||
if (!ov) _exit(2);
|
||||
if (vfs_mount(v, "/", ov) != VFS_OK) _exit(3);
|
||||
|
||||
char payload[PAYLOAD_SIZE];
|
||||
for (int i = 0; i < EXTRA_N; i++) {
|
||||
char name[64];
|
||||
snprintf(name, sizeof(name), "/extra%04d.dat", i);
|
||||
mk_payload(payload, sizeof(payload), "extra", i);
|
||||
int err = 0;
|
||||
VfsFile *f = vfs_open(v, name, VFS_O_WRONLY | VFS_O_CREAT, &err);
|
||||
if (!f) _exit(4);
|
||||
if (vfs_write(f, payload, sizeof(payload)) != (pfs_isize)sizeof(payload)) _exit(5);
|
||||
if (vfs_close(f) != VFS_OK) _exit(6);
|
||||
record_progress(arg->progress_path, i + 1); /* extra i is now durable */
|
||||
}
|
||||
/* every extra above is already durable (each fsynced individually);
|
||||
* this is the operation actually being crash-tested */
|
||||
if (vfs_sync(v, "/") != VFS_OK) _exit(7);
|
||||
_exit(0);
|
||||
}
|
||||
|
||||
static void build_golden_baseline(CompactArg *arg) {
|
||||
unlink(arg->pack_path);
|
||||
char jpath[600];
|
||||
snprintf(jpath, sizeof(jpath), "%s.jnl", arg->pack_path);
|
||||
unlink(jpath);
|
||||
|
||||
Vfs *v = vfs_new();
|
||||
Backend *mem = backend_mem_new();
|
||||
int oerr = 0;
|
||||
Backend *ov = backend_overlay_new(arg->pack_path, mem, &oerr);
|
||||
CHECK(ov != NULL);
|
||||
CHECK_EQ_INT(vfs_mount(v, "/", ov), VFS_OK);
|
||||
|
||||
char payload[PAYLOAD_SIZE];
|
||||
for (int i = 0; i < GOLDEN_N; i++) {
|
||||
char name[64];
|
||||
snprintf(name, sizeof(name), "/base%05d.dat", i);
|
||||
mk_payload(payload, sizeof(payload), "base", i);
|
||||
int err = 0;
|
||||
VfsFile *f = vfs_open(v, name, VFS_O_WRONLY | VFS_O_CREAT, &err);
|
||||
CHECK(f != NULL);
|
||||
CHECK_EQ_INT(vfs_write(f, payload, sizeof(payload)), (pfs_isize)sizeof(payload));
|
||||
CHECK_EQ_INT(vfs_close(f), VFS_OK);
|
||||
}
|
||||
CHECK_EQ_INT(vfs_sync(v, "/"), VFS_OK); /* clean, untimed compaction */
|
||||
|
||||
vfs_unmount(v, "/");
|
||||
backend_free(ov);
|
||||
backend_free(mem);
|
||||
vfs_free(v);
|
||||
unlink(jpath); /* golden state: compacted pack, no pending journal */
|
||||
|
||||
CHECK_EQ_INT(copy_file(arg->pack_path, arg->golden_path), 0);
|
||||
}
|
||||
|
||||
static void test_compaction_crash(void) {
|
||||
CompactArg arg;
|
||||
snprintf(arg.pack_path, sizeof(arg.pack_path), "/tmp/packfs_test_crash_compact_%d.img", (int)getpid());
|
||||
snprintf(arg.golden_path, sizeof(arg.golden_path), "/tmp/packfs_test_crash_golden_%d.img", (int)getpid());
|
||||
snprintf(arg.progress_path, sizeof(arg.progress_path), "/tmp/packfs_test_crash_progress_%d.txt", (int)getpid());
|
||||
char jpath[600];
|
||||
snprintf(jpath, sizeof(jpath), "%s.jnl", arg.pack_path);
|
||||
|
||||
build_golden_baseline(&arg);
|
||||
|
||||
int interrupted_count = 0;
|
||||
for (int trial = 0; trial < COMPACT_TRIALS; trial++) {
|
||||
CHECK_EQ_INT(copy_file(arg.golden_path, arg.pack_path), 0);
|
||||
unlink(jpath);
|
||||
|
||||
double delay = (COMPACT_MAX_DELAY * (double)trial) / (double)(COMPACT_TRIALS - 1);
|
||||
int interrupted = run_and_kill(compact_child, &arg, delay);
|
||||
if (interrupted) interrupted_count++;
|
||||
|
||||
Vfs *v2 = vfs_new();
|
||||
Backend *mem2 = backend_mem_new();
|
||||
int oerr2 = 0;
|
||||
Backend *ov2 = backend_overlay_new(arg.pack_path, mem2, &oerr2);
|
||||
CHECK(ov2 != NULL); /* a crash mid-compaction must never make reopening fail */
|
||||
if (!ov2) { vfs_free(v2); backend_free(mem2); continue; }
|
||||
CHECK_EQ_INT(vfs_mount(v2, "/", ov2), VFS_OK);
|
||||
|
||||
/* the pre-existing golden baseline must always survive */
|
||||
char name[64], expect[PAYLOAD_SIZE], buf[PAYLOAD_SIZE];
|
||||
for (int i = 0; i < GOLDEN_N; i++) {
|
||||
snprintf(name, sizeof(name), "/base%05d.dat", i);
|
||||
mk_payload(expect, sizeof(expect), "base", i);
|
||||
int err = 0;
|
||||
VfsFile *f = vfs_open(v2, name, VFS_O_RDONLY, &err);
|
||||
if (!f) {
|
||||
fprintf(stderr, "FAIL trial %d: baseline file %s LOST after compaction crash (delay=%.4f)\n",
|
||||
trial, name, delay);
|
||||
pfs_test_failures++;
|
||||
continue;
|
||||
}
|
||||
memset(buf, 0, sizeof(buf));
|
||||
pfs_isize n = vfs_read(f, buf, sizeof(buf));
|
||||
vfs_close(f);
|
||||
if (n != (pfs_isize)sizeof(buf) || memcmp(buf, expect, sizeof(buf)) != 0) {
|
||||
fprintf(stderr, "FAIL trial %d: baseline file %s CORRUPTED after compaction crash (delay=%.4f)\n",
|
||||
trial, name, delay);
|
||||
pfs_test_failures++;
|
||||
}
|
||||
}
|
||||
|
||||
/* Every "extra" write the child actually completed (vfs_close
|
||||
* returned VFS_OK) before it was killed was already durably
|
||||
* journaled at that point -- it must survive regardless of
|
||||
* whether the subsequent vfs_sync (compaction) itself completed.
|
||||
* completed_extras, from the independent progress side-channel,
|
||||
* is the ground truth for how many extras the child actually
|
||||
* finished; extras beyond that were never attempted and their
|
||||
* absence is expected, not a bug. */
|
||||
int completed_extras = read_progress(arg.progress_path);
|
||||
CHECK(completed_extras >= 0 && completed_extras <= EXTRA_N);
|
||||
for (int i = 0; i < completed_extras; i++) {
|
||||
snprintf(name, sizeof(name), "/extra%04d.dat", i);
|
||||
mk_payload(expect, sizeof(expect), "extra", i);
|
||||
int err = 0;
|
||||
VfsFile *f = vfs_open(v2, name, VFS_O_RDONLY, &err);
|
||||
if (!f) {
|
||||
fprintf(stderr, "FAIL trial %d: durably-completed %s LOST after compaction crash "
|
||||
"(delay=%.4f, completed_extras=%d)\n", trial, name, delay, completed_extras);
|
||||
pfs_test_failures++;
|
||||
continue;
|
||||
}
|
||||
memset(buf, 0, sizeof(buf));
|
||||
pfs_isize n = vfs_read(f, buf, sizeof(buf));
|
||||
vfs_close(f);
|
||||
if (n != (pfs_isize)sizeof(buf) || memcmp(buf, expect, sizeof(buf)) != 0) {
|
||||
fprintf(stderr, "FAIL trial %d: durably-completed %s CORRUPTED after compaction crash "
|
||||
"(delay=%.4f, completed_extras=%d)\n", trial, name, delay, completed_extras);
|
||||
pfs_test_failures++;
|
||||
}
|
||||
}
|
||||
|
||||
vfs_unmount(v2, "/");
|
||||
backend_free(ov2);
|
||||
backend_free(mem2);
|
||||
vfs_free(v2);
|
||||
}
|
||||
|
||||
unlink(arg.pack_path);
|
||||
unlink(arg.golden_path);
|
||||
unlink(arg.progress_path);
|
||||
unlink(jpath);
|
||||
fprintf(stderr, "compaction crash test: %d/%d trials genuinely interrupted mid-compaction\n",
|
||||
interrupted_count, COMPACT_TRIALS);
|
||||
CHECK(interrupted_count >= COMPACT_TRIALS / 4);
|
||||
}
|
||||
|
||||
int main(void) {
|
||||
test_journal_burst_crash();
|
||||
test_compaction_crash();
|
||||
TEST_MAIN_END();
|
||||
}
|
||||
Reference in New Issue
Block a user