Files
packfs/tests/test_pack_write_perf.c
retoorandClaude Sonnet 5 684c6d95de Fix pack_write's O(n^2) dedup + correctness bug; document mount table O(n^2)
A audit for other instances of the file-index O(n^2) shape (fixed
previously via the persistent treap) found two more real issues:

1. pack_write's Section 9.2 exact-duplicate elimination was a linear scan
   of every previously-seen (hash, size) pair per entry -- O(n^2) total,
   invisible in the existing benchmark because its identical-content test
   files made every scan match on the first comparison. Measured with
   unique content instead: 80,000 entries took 1.74s, with a 20,000->80,000
   step showing 15.6x for a 4x-N step, matching O(n^2)'s 16x prediction.
   The same scan also trusted a (hash, size) match without ever comparing
   actual bytes -- a latent correctness bug, since FNV-1a64 is explicitly
   not collision-resistant. Fixed both at once with an open-addressing hash
   table (load factor 1/2, linear probing) plus a memcmp verification
   before ever reusing a data_off. Post-fix: 80,000 entries in 0.044s
   (39.6x faster), ratio drops to 3.35x (consistent with O(n)).

   Covered permanently by two new/extended tests: a white-box assertion in
   test_pack_overlay.c that duplicate-content entries share one data_off
   and distinct-content entries do not, and a new
   tests/test_pack_write_perf.c regression tripwire against 10,000 unique
   entries.

2. vfs.c's mount table uses the same full-array-copy-per-write pattern the
   file index used to, confirmed O(n^2) via a new bench/bench.c category
   (500/2,000/8,000 mounts, both 4x-N steps showing 15-20x). Deliberately
   NOT rewritten: mount points are created by a program's own source code,
   not workload-driven, so realistic mount counts never reach the scale
   that made the file index's O(n^2) a real problem. Documented with full
   reasoning in BENCH.md and CLAUDE.md rather than silently left as an
   undocumented gap.

Also fixes a real CI gap the new tests exposed: ci.yml's sanitizer-build
steps never passed -D_GNU_SOURCE when compiling test files (only the
library .o's got it), which was harmless while no test included
internal.h and became a link failure once two did (internal.h needs
_GNU_SOURCE for pthread_rwlock_t). And documents, in CONTRIBUTING.md and
CLAUDE.md, a sandbox flake observed directly during this work's own
sanitizer runs: ASan/UBSan test binaries occasionally fail to start with
AddressSanitizer:DEADLYSIGNAL (sometimes looping rather than exiting),
non-deterministically hitting different unrelated binaries across runs --
a startup race, not a memory-safety bug, confirmed by clean passes on
retry; sanitizer runs in such an environment should be timeout-wrapped.

BENCH.md's "After" table and Appendix B are replaced with the current,
complete 54-measurement bench/bench.c run (the original 45 plus the new
mount-scaling category); the pre-fix 45-measurement "Before" table is kept
as the historical record, per this project's documentation standard.

Verified: make test (all 6 binaries, including the 2 new/changed), a clean
make all, and repeated ASan+UBSan runs (0 real findings; the DEADLYSIGNAL
flake above was observed and correctly distinguished from a real finding
by re-running until a clean pass). TSan could not be run in this sandbox
(pre-existing, documented environment limitation).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UqJpkdJ6Njnt1pw3CbghzB
2026-09-14 09:55:47 +00:00

106 lines
3.9 KiB
C

/*
* test_pack_write_perf.c — regression guard for a real bug: pack_write's
* Section 9.2 exact-duplicate elimination used to be a linear scan of
* every previously-seen (hash, size) pair per entry — O(n) per entry,
* O(n^2) total across n distinct blobs, found and fixed alongside the
* file index's own O(n^2) (see CLAUDE.md's "Known performance
* characteristics" and BENCH.md). It is now a hash table (load factor
* 1/2, linear probing), which this asserts stays fast at a scale where
* the old code was already measurably slow (BENCH.md: 80,000 unique
* entries took 1.74s pre-fix, 0.044s post-fix).
*
* Deliberately bypasses the VFS/overlay/journal path entirely (calls
* pack_write directly via internal.h) rather than creating N files
* through an overlay: that path journals and fsyncs every write
* (Section 4.4), which is a separate, unrelated cost this test has no
* reason to pay, and which is pathologically slow on some sandboxed
* environments (see CONTRIBUTING.md's ThreadSanitizer note for another
* example of the same class of environment quirk) — conflating the two
* costs is exactly the mistake that delayed finding this bug in the
* first place, so this test deliberately does not repeat it.
*/
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <time.h>
#include <unistd.h>
#include "internal.h"
#include "test_harness.h"
#define N 10000
static double now_sec(void) {
struct timespec ts;
clock_gettime(CLOCK_MONOTONIC, &ts);
return (double)ts.tv_sec + (double)ts.tv_nsec / 1e9;
}
int main(void) {
PackBuildEntry *entries = (PackBuildEntry *)malloc(sizeof(PackBuildEntry) * N);
char **names = (char **)malloc(sizeof(char *) * N);
char **datas = (char **)malloc(sizeof(char *) * N);
/* every entry's content is unique: this is the linear scan's actual
* worst case (every comparison fails until the empty tail), and
* exactly the case the original benchmark's identical-content test
* files accidentally never exercised. */
for (int i = 0; i < N; i++) {
names[i] = (char *)malloc(32);
snprintf(names[i], 32, "/f%06d.dat", i);
datas[i] = (char *)malloc(48);
int len = snprintf(datas[i], 48, "unique-content-block-%06d", i);
entries[i].name = names[i];
entries[i].data = datas[i];
entries[i].size = (uint64_t)len;
entries[i].mode = 0;
entries[i].mtime = 0;
}
char path[128];
snprintf(path, sizeof(path), "/tmp/packfs_test_dedup_perf_%d.img", (int)getpid());
unlink(path);
int err = 0;
double t0 = now_sec();
int rc = pack_write(path, entries, N, &err);
double elapsed = now_sec() - t0;
CHECK_EQ_INT(rc, 0);
/* Generous bound: post-fix this takes well under 0.1s for N=10,000
* on ordinary hardware (BENCH.md: 0.013s for N=20,000). The old
* O(n^2) code took long enough at this N to be a clear, unmissable
* fail, not a borderline one — this is a regression tripwire, not a
* tight performance assertion. */
CHECK(elapsed < 5.0);
if (elapsed >= 5.0) {
fprintf(stderr, "pack_write(%d unique entries) took %.3fs -- the O(n^2) dedup regression may be back\n", N, elapsed);
}
/* correctness, not just speed: load it back and confirm every entry
* is findable with its own distinct content intact. */
Pack *p = NULL;
int lerr = 0;
CHECK_EQ_INT(pack_load(path, &p, &lerr), 0);
if (p) {
for (int i = 0; i < N; i += 997) { /* sample, not all 10,000, to keep this fast */
const PackIndexEntry *e = NULL;
CHECK(pack_find(p, names[i], &e));
if (e) {
CHECK_EQ_INT(e->size, strlen(datas[i]));
CHECK_EQ_INT(memcmp(pack_entry_data(p, e), datas[i], e->size), 0);
}
}
pack_close(p);
}
unlink(path);
for (int i = 0; i < N; i++) { free(names[i]); free(datas[i]); }
free(names);
free(datas);
free(entries);
TEST_MAIN_END();
}