Fix pack_write's O(n^2) dedup + correctness bug; document mount table O(n^2)
A audit for other instances of the file-index O(n^2) shape (fixed previously via the persistent treap) found two more real issues: 1. pack_write's Section 9.2 exact-duplicate elimination was a linear scan of every previously-seen (hash, size) pair per entry -- O(n^2) total, invisible in the existing benchmark because its identical-content test files made every scan match on the first comparison. Measured with unique content instead: 80,000 entries took 1.74s, with a 20,000->80,000 step showing 15.6x for a 4x-N step, matching O(n^2)'s 16x prediction. The same scan also trusted a (hash, size) match without ever comparing actual bytes -- a latent correctness bug, since FNV-1a64 is explicitly not collision-resistant. Fixed both at once with an open-addressing hash table (load factor 1/2, linear probing) plus a memcmp verification before ever reusing a data_off. Post-fix: 80,000 entries in 0.044s (39.6x faster), ratio drops to 3.35x (consistent with O(n)). Covered permanently by two new/extended tests: a white-box assertion in test_pack_overlay.c that duplicate-content entries share one data_off and distinct-content entries do not, and a new tests/test_pack_write_perf.c regression tripwire against 10,000 unique entries. 2. vfs.c's mount table uses the same full-array-copy-per-write pattern the file index used to, confirmed O(n^2) via a new bench/bench.c category (500/2,000/8,000 mounts, both 4x-N steps showing 15-20x). Deliberately NOT rewritten: mount points are created by a program's own source code, not workload-driven, so realistic mount counts never reach the scale that made the file index's O(n^2) a real problem. Documented with full reasoning in BENCH.md and CLAUDE.md rather than silently left as an undocumented gap. Also fixes a real CI gap the new tests exposed: ci.yml's sanitizer-build steps never passed -D_GNU_SOURCE when compiling test files (only the library .o's got it), which was harmless while no test included internal.h and became a link failure once two did (internal.h needs _GNU_SOURCE for pthread_rwlock_t). And documents, in CONTRIBUTING.md and CLAUDE.md, a sandbox flake observed directly during this work's own sanitizer runs: ASan/UBSan test binaries occasionally fail to start with AddressSanitizer:DEADLYSIGNAL (sometimes looping rather than exiting), non-deterministically hitting different unrelated binaries across runs -- a startup race, not a memory-safety bug, confirmed by clean passes on retry; sanitizer runs in such an environment should be timeout-wrapped. BENCH.md's "After" table and Appendix B are replaced with the current, complete 54-measurement bench/bench.c run (the original 45 plus the new mount-scaling category); the pre-fix 45-measurement "Before" table is kept as the historical record, per this project's documentation standard. Verified: make test (all 6 binaries, including the 2 new/changed), a clean make all, and repeated ASan+UBSan runs (0 real findings; the DEADLYSIGNAL flake above was observed and correctly distinguished from a real finding by re-running until a clean pass). TSan could not be run in this sandbox (pre-existing, documented environment limitation). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
This commit is contained in:
@@ -10,6 +10,7 @@
|
||||
#include <unistd.h>
|
||||
|
||||
#include "packfs.h"
|
||||
#include "internal.h" /* white-box: verifies Section 9.2 dedup actually shares data_off, not just "doesn't crash" */
|
||||
#include "test_harness.h"
|
||||
|
||||
static char *tmp_pack_path(void) {
|
||||
@@ -51,6 +52,28 @@ int main(void) {
|
||||
|
||||
CHECK_EQ_INT(vfs_sync(v, "/"), VFS_OK); /* compaction */
|
||||
|
||||
/* Section 9.2 dedup, actually verified (not just exercised): /a.txt
|
||||
* and /dup.txt have identical content and must share one data_off
|
||||
* in the compacted pack; /b.txt has different content and must not
|
||||
* share either's. Catches both "dedup silently stopped happening"
|
||||
* (a real, previously-unasserted gap) and "dedup over-matched two
|
||||
* different files onto the same bytes" (the hash-collision-without-
|
||||
* memcmp-verification bug fixed alongside the O(n^2) compaction
|
||||
* cost — see pack.c's DedupSlot comment). */
|
||||
{
|
||||
Pack *p = NULL; int lerr = 0;
|
||||
CHECK_EQ_INT(pack_load(pack_path, &p, &lerr), 0);
|
||||
if (p) {
|
||||
const PackIndexEntry *ea = NULL, *edup = NULL, *eb = NULL;
|
||||
CHECK(pack_find(p, "/a.txt", &ea));
|
||||
CHECK(pack_find(p, "/dup.txt", &edup));
|
||||
CHECK(pack_find(p, "/b.txt", &eb));
|
||||
if (ea && edup) CHECK_EQ_INT(ea->data_off, edup->data_off);
|
||||
if (ea && eb) CHECK(ea->data_off != eb->data_off);
|
||||
pack_close(p);
|
||||
}
|
||||
}
|
||||
|
||||
/* read-only open straight from the freshly-compacted pack */
|
||||
f = vfs_open(v, "/a.txt", VFS_O_RDONLY, &err);
|
||||
CHECK(f != NULL);
|
||||
|
||||
@@ -0,0 +1,105 @@
|
||||
/*
|
||||
* test_pack_write_perf.c — regression guard for a real bug: pack_write's
|
||||
* Section 9.2 exact-duplicate elimination used to be a linear scan of
|
||||
* every previously-seen (hash, size) pair per entry — O(n) per entry,
|
||||
* O(n^2) total across n distinct blobs, found and fixed alongside the
|
||||
* file index's own O(n^2) (see CLAUDE.md's "Known performance
|
||||
* characteristics" and BENCH.md). It is now a hash table (load factor
|
||||
* 1/2, linear probing), which this asserts stays fast at a scale where
|
||||
* the old code was already measurably slow (BENCH.md: 80,000 unique
|
||||
* entries took 1.74s pre-fix, 0.044s post-fix).
|
||||
*
|
||||
* Deliberately bypasses the VFS/overlay/journal path entirely (calls
|
||||
* pack_write directly via internal.h) rather than creating N files
|
||||
* through an overlay: that path journals and fsyncs every write
|
||||
* (Section 4.4), which is a separate, unrelated cost this test has no
|
||||
* reason to pay, and which is pathologically slow on some sandboxed
|
||||
* environments (see CONTRIBUTING.md's ThreadSanitizer note for another
|
||||
* example of the same class of environment quirk) — conflating the two
|
||||
* costs is exactly the mistake that delayed finding this bug in the
|
||||
* first place, so this test deliberately does not repeat it.
|
||||
*/
|
||||
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <time.h>
|
||||
#include <unistd.h>
|
||||
|
||||
#include "internal.h"
|
||||
#include "test_harness.h"
|
||||
|
||||
#define N 10000
|
||||
|
||||
static double now_sec(void) {
|
||||
struct timespec ts;
|
||||
clock_gettime(CLOCK_MONOTONIC, &ts);
|
||||
return (double)ts.tv_sec + (double)ts.tv_nsec / 1e9;
|
||||
}
|
||||
|
||||
int main(void) {
|
||||
PackBuildEntry *entries = (PackBuildEntry *)malloc(sizeof(PackBuildEntry) * N);
|
||||
char **names = (char **)malloc(sizeof(char *) * N);
|
||||
char **datas = (char **)malloc(sizeof(char *) * N);
|
||||
|
||||
/* every entry's content is unique: this is the linear scan's actual
|
||||
* worst case (every comparison fails until the empty tail), and
|
||||
* exactly the case the original benchmark's identical-content test
|
||||
* files accidentally never exercised. */
|
||||
for (int i = 0; i < N; i++) {
|
||||
names[i] = (char *)malloc(32);
|
||||
snprintf(names[i], 32, "/f%06d.dat", i);
|
||||
datas[i] = (char *)malloc(48);
|
||||
int len = snprintf(datas[i], 48, "unique-content-block-%06d", i);
|
||||
entries[i].name = names[i];
|
||||
entries[i].data = datas[i];
|
||||
entries[i].size = (uint64_t)len;
|
||||
entries[i].mode = 0;
|
||||
entries[i].mtime = 0;
|
||||
}
|
||||
|
||||
char path[128];
|
||||
snprintf(path, sizeof(path), "/tmp/packfs_test_dedup_perf_%d.img", (int)getpid());
|
||||
unlink(path);
|
||||
|
||||
int err = 0;
|
||||
double t0 = now_sec();
|
||||
int rc = pack_write(path, entries, N, &err);
|
||||
double elapsed = now_sec() - t0;
|
||||
|
||||
CHECK_EQ_INT(rc, 0);
|
||||
/* Generous bound: post-fix this takes well under 0.1s for N=10,000
|
||||
* on ordinary hardware (BENCH.md: 0.013s for N=20,000). The old
|
||||
* O(n^2) code took long enough at this N to be a clear, unmissable
|
||||
* fail, not a borderline one — this is a regression tripwire, not a
|
||||
* tight performance assertion. */
|
||||
CHECK(elapsed < 5.0);
|
||||
if (elapsed >= 5.0) {
|
||||
fprintf(stderr, "pack_write(%d unique entries) took %.3fs -- the O(n^2) dedup regression may be back\n", N, elapsed);
|
||||
}
|
||||
|
||||
/* correctness, not just speed: load it back and confirm every entry
|
||||
* is findable with its own distinct content intact. */
|
||||
Pack *p = NULL;
|
||||
int lerr = 0;
|
||||
CHECK_EQ_INT(pack_load(path, &p, &lerr), 0);
|
||||
if (p) {
|
||||
for (int i = 0; i < N; i += 997) { /* sample, not all 10,000, to keep this fast */
|
||||
const PackIndexEntry *e = NULL;
|
||||
CHECK(pack_find(p, names[i], &e));
|
||||
if (e) {
|
||||
CHECK_EQ_INT(e->size, strlen(datas[i]));
|
||||
CHECK_EQ_INT(memcmp(pack_entry_data(p, e), datas[i], e->size), 0);
|
||||
}
|
||||
}
|
||||
pack_close(p);
|
||||
}
|
||||
|
||||
unlink(path);
|
||||
for (int i = 0; i < N; i++) { free(names[i]); free(datas[i]); }
|
||||
free(names);
|
||||
free(datas);
|
||||
free(entries);
|
||||
|
||||
TEST_MAIN_END();
|
||||
}
|
||||
Reference in New Issue
Block a user