A audit for other instances of the file-index O(n^2) shape (fixed previously via the persistent treap) found two more real issues: 1. pack_write's Section 9.2 exact-duplicate elimination was a linear scan of every previously-seen (hash, size) pair per entry -- O(n^2) total, invisible in the existing benchmark because its identical-content test files made every scan match on the first comparison. Measured with unique content instead: 80,000 entries took 1.74s, with a 20,000->80,000 step showing 15.6x for a 4x-N step, matching O(n^2)'s 16x prediction. The same scan also trusted a (hash, size) match without ever comparing actual bytes -- a latent correctness bug, since FNV-1a64 is explicitly not collision-resistant. Fixed both at once with an open-addressing hash table (load factor 1/2, linear probing) plus a memcmp verification before ever reusing a data_off. Post-fix: 80,000 entries in 0.044s (39.6x faster), ratio drops to 3.35x (consistent with O(n)). Covered permanently by two new/extended tests: a white-box assertion in test_pack_overlay.c that duplicate-content entries share one data_off and distinct-content entries do not, and a new tests/test_pack_write_perf.c regression tripwire against 10,000 unique entries. 2. vfs.c's mount table uses the same full-array-copy-per-write pattern the file index used to, confirmed O(n^2) via a new bench/bench.c category (500/2,000/8,000 mounts, both 4x-N steps showing 15-20x). Deliberately NOT rewritten: mount points are created by a program's own source code, not workload-driven, so realistic mount counts never reach the scale that made the file index's O(n^2) a real problem. Documented with full reasoning in BENCH.md and CLAUDE.md rather than silently left as an undocumented gap. Also fixes a real CI gap the new tests exposed: ci.yml's sanitizer-build steps never passed -D_GNU_SOURCE when compiling test files (only the library .o's got it), which was harmless while no test included internal.h and became a link failure once two did (internal.h needs _GNU_SOURCE for pthread_rwlock_t). And documents, in CONTRIBUTING.md and CLAUDE.md, a sandbox flake observed directly during this work's own sanitizer runs: ASan/UBSan test binaries occasionally fail to start with AddressSanitizer:DEADLYSIGNAL (sometimes looping rather than exiting), non-deterministically hitting different unrelated binaries across runs -- a startup race, not a memory-safety bug, confirmed by clean passes on retry; sanitizer runs in such an environment should be timeout-wrapped. BENCH.md's "After" table and Appendix B are replaced with the current, complete 54-measurement bench/bench.c run (the original 45 plus the new mount-scaling category); the pre-fix 45-measurement "Before" table is kept as the historical record, per this project's documentation standard. Verified: make test (all 6 binaries, including the 2 new/changed), a clean make all, and repeated ASan+UBSan runs (0 real findings; the DEADLYSIGNAL flake above was observed and correctly distinguished from a real finding by re-running until a clean pass). TSan could not be run in this sandbox (pre-existing, documented environment limitation). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
106 lines
3.9 KiB
C
106 lines
3.9 KiB
C
/*
|
|
* test_pack_write_perf.c — regression guard for a real bug: pack_write's
|
|
* Section 9.2 exact-duplicate elimination used to be a linear scan of
|
|
* every previously-seen (hash, size) pair per entry — O(n) per entry,
|
|
* O(n^2) total across n distinct blobs, found and fixed alongside the
|
|
* file index's own O(n^2) (see CLAUDE.md's "Known performance
|
|
* characteristics" and BENCH.md). It is now a hash table (load factor
|
|
* 1/2, linear probing), which this asserts stays fast at a scale where
|
|
* the old code was already measurably slow (BENCH.md: 80,000 unique
|
|
* entries took 1.74s pre-fix, 0.044s post-fix).
|
|
*
|
|
* Deliberately bypasses the VFS/overlay/journal path entirely (calls
|
|
* pack_write directly via internal.h) rather than creating N files
|
|
* through an overlay: that path journals and fsyncs every write
|
|
* (Section 4.4), which is a separate, unrelated cost this test has no
|
|
* reason to pay, and which is pathologically slow on some sandboxed
|
|
* environments (see CONTRIBUTING.md's ThreadSanitizer note for another
|
|
* example of the same class of environment quirk) — conflating the two
|
|
* costs is exactly the mistake that delayed finding this bug in the
|
|
* first place, so this test deliberately does not repeat it.
|
|
*/
|
|
|
|
#include <stdio.h>
|
|
#include <stdlib.h>
|
|
#include <string.h>
|
|
#include <time.h>
|
|
#include <unistd.h>
|
|
|
|
#include "internal.h"
|
|
#include "test_harness.h"
|
|
|
|
#define N 10000
|
|
|
|
static double now_sec(void) {
|
|
struct timespec ts;
|
|
clock_gettime(CLOCK_MONOTONIC, &ts);
|
|
return (double)ts.tv_sec + (double)ts.tv_nsec / 1e9;
|
|
}
|
|
|
|
int main(void) {
|
|
PackBuildEntry *entries = (PackBuildEntry *)malloc(sizeof(PackBuildEntry) * N);
|
|
char **names = (char **)malloc(sizeof(char *) * N);
|
|
char **datas = (char **)malloc(sizeof(char *) * N);
|
|
|
|
/* every entry's content is unique: this is the linear scan's actual
|
|
* worst case (every comparison fails until the empty tail), and
|
|
* exactly the case the original benchmark's identical-content test
|
|
* files accidentally never exercised. */
|
|
for (int i = 0; i < N; i++) {
|
|
names[i] = (char *)malloc(32);
|
|
snprintf(names[i], 32, "/f%06d.dat", i);
|
|
datas[i] = (char *)malloc(48);
|
|
int len = snprintf(datas[i], 48, "unique-content-block-%06d", i);
|
|
entries[i].name = names[i];
|
|
entries[i].data = datas[i];
|
|
entries[i].size = (uint64_t)len;
|
|
entries[i].mode = 0;
|
|
entries[i].mtime = 0;
|
|
}
|
|
|
|
char path[128];
|
|
snprintf(path, sizeof(path), "/tmp/packfs_test_dedup_perf_%d.img", (int)getpid());
|
|
unlink(path);
|
|
|
|
int err = 0;
|
|
double t0 = now_sec();
|
|
int rc = pack_write(path, entries, N, &err);
|
|
double elapsed = now_sec() - t0;
|
|
|
|
CHECK_EQ_INT(rc, 0);
|
|
/* Generous bound: post-fix this takes well under 0.1s for N=10,000
|
|
* on ordinary hardware (BENCH.md: 0.013s for N=20,000). The old
|
|
* O(n^2) code took long enough at this N to be a clear, unmissable
|
|
* fail, not a borderline one — this is a regression tripwire, not a
|
|
* tight performance assertion. */
|
|
CHECK(elapsed < 5.0);
|
|
if (elapsed >= 5.0) {
|
|
fprintf(stderr, "pack_write(%d unique entries) took %.3fs -- the O(n^2) dedup regression may be back\n", N, elapsed);
|
|
}
|
|
|
|
/* correctness, not just speed: load it back and confirm every entry
|
|
* is findable with its own distinct content intact. */
|
|
Pack *p = NULL;
|
|
int lerr = 0;
|
|
CHECK_EQ_INT(pack_load(path, &p, &lerr), 0);
|
|
if (p) {
|
|
for (int i = 0; i < N; i += 997) { /* sample, not all 10,000, to keep this fast */
|
|
const PackIndexEntry *e = NULL;
|
|
CHECK(pack_find(p, names[i], &e));
|
|
if (e) {
|
|
CHECK_EQ_INT(e->size, strlen(datas[i]));
|
|
CHECK_EQ_INT(memcmp(pack_entry_data(p, e), datas[i], e->size), 0);
|
|
}
|
|
}
|
|
pack_close(p);
|
|
}
|
|
|
|
unlink(path);
|
|
for (int i = 0; i < N; i++) { free(names[i]); free(datas[i]); }
|
|
free(names);
|
|
free(datas);
|
|
free(entries);
|
|
|
|
TEST_MAIN_END();
|
|
}
|