Files
versioning/research_notes/Backup include exclude rules/binary-large-files.md
T
retoor 3499f6cda0 Remove encryption; user-data filters; SQLite safety; recovery, retention, scheduler
- No client-side encryption: plaintext content-addressed remote (SHA-256
  names), no backup key, no cryptography dependency (server disk encryption
  is the trust model).
- Filters back user data: images, archives, databases, PDFs accepted; 10 MiB
  cap; binary sniffing removed; temp names hardened (~$, #..#, .temp).
- SQLite zero-error policy: backup-API snapshots + integrity_check, journal
  folding, locked/corrupt loud skips, verified restores.
- Recovery: reindex from manifests, remote adopt, on-demand blob fetch,
  5-day retention + thinning, GC, date-guarded remote purge, metrics.
- Scheduler with WebDAV quota signal and 70% pressure backstop (floor kept).
- 37 tests incl. live-monitor capture safety and DB safety.
2026-10-10 03:41:36 +02:00

15 KiB
Raw Blame History

Binary, Large, Image, Archive and Database Files in Versioning/Backup Tools

What size caps or chunking strategies do these tools use, and what breaks with multi-MB binaries in a per-version system (storage blowup, bandwidth)?

Takeaway

Restic, Borg and Kopia all use content-defined chunking (CDC) with ~0.5–8 MiB variable chunks so small edits to large binaries only store 1–2 new chunks; without CDC (fixed blocks or full-file copies) a 1-byte insert re-stores the whole file, and Git-LFS instead punts large files to pointer+blob storage with host-enforced per-file caps.

Cited Findings

  • Restic splits files with Rabin-fingerprint CDC over a 64-byte sliding window, cutting when low 21 bits are zero; files <512 KiB are not split, blobs are 512 KiB–8 MiB, ~1 MiB average — Source; background — Source
  • Restic chunker defaults aim at ~1 MiB average (splitmask = (1<<20)-1) with configurable Min/MaxSize — Source
  • Borg splits files into deduplicated chunks globally across repo (all machines/archives); chunk id is a strong hash/MAC (hmac-sha256 / keyed blake3), not the rolling-hash value — Source
  • Borg default buzhash params: min 2^19 (512 KiB), max 2^23 (8 MiB), mask 21 bits (~2 MiB target), window 4095 B; fixed chunker option for disk/VM images — Source
  • Borg 2.x chunkers: fastcdc (default, fastest), buzhash64/buzhash, keyed AES variants (toeplitz-aes/rabin-aes strongest), fixed for raw disk images where CDC gains little — Source
  • Borg warns fine-grained --chunker-params=buzhash,10,23,16,4095 creates huge chunk counts and RAM/disk load; coarse default suits large volumes — Source
  • Kopia calls CDC "splitters": FIXED vs DYNAMIC BUZHASH/RABINKARP, sizes 1M–8M (default DYNAMIC-4M-BUZHASH); small files = one content, large files split so metadata-only change to 10 GB video uploads only 1–2 chunks (<10 MB) — Source; details — Source; packing — Source
  • Kopia packs many contents into 20–40 MB pack blobs; splitter choice is set at repo creation — Source; chunk-then-hash-then-compress-then-encrypt pipeline — Source
  • Restic has no hard max file size but offers --exclude-larger-than size (suffixes k/M/G/T) to skip files over a threshold — Source
  • Git-LFS has no inherent file-size limit; limits are host-enforced (GitHub: 2 GB Free/Pro, 4 GB Team, 5 GB Enterprise Cloud; >5 GB rejected); pointer file stores version/oid sha256/size — Source
  • Git-LFS tracks by .gitattributes pattern, not by size; --above in migrate import is one-shot, no automatic by-size tracking (2025 --min-size/autotracksize PR still debated/unmerged) — Source
  • Git on Windows pre-2.34 could not smudge/clean files >4 GiB; workaround GIT_LFS_SKIP_SMUDGE=1 + git lfs pull — Source
  • Every Git-LFS revision counts against remote storage/bandwidth quota, so per-version binaries inflate cost; clones/pulls are slow on large repos — Source
  • Keyed CDC (Borg/Restic/Kopia) is fingerprintable: observing chunk sizes of known data recovers keys (CCS 2025, eprint 2025/558; backup-service attacks eprint 2025/532); Restic 0.18 mitigates by random chunk-to-pack assignment — Source; attacks — Source; analysis — Source

Inferences

  • A small local daemon should copy the Restic/Borg default: CDC ~1–2 MiB average, 512 KiB min, 8 MiB max; offer --exclude-larger-than-style cap (e.g. 50–100 MB default-off) plus extension-based excludes rather than a hard cap.
  • Per-version full copies of multi-MB binaries blow up storage and bandwidth linearly with versions; CDC reduces this to delta-of-chunks but still re-reads/re-hashes the whole file each run.
  • For VM/disk images with fixed internal layout, a fixed-block mode is faster and equally effective.

Gaps

  • No reliable 2026 source found on VSCode/JetBrains Local History size caps or binary handling (searches rate-limited); editor-local-history defaults remain unverified.
  • Exact Kopia default content size bounds (min/max around 4 MB mean) not confirmed from primary docs; forum states 1–8 MB range only.

Takeaway

Git (and libgit2 ports) still use NUL-in-first-8000-bytes as the primary binary signal, tightened in 2008 so CRLF conversion defers to diff; it is fast and recommended as a first heuristic but known-insufficient for UTF-16 and NUL-free binaries, so explicit .gitattributes marking is required.

Cited Findings

  • Git buffer_is_binary() checks for NUL via memchr in first 8000 bytes (FIRST_FEW_BYTES) — Source
  • Git convert.c convert_is_binary(): binary if lonecr or nul or (printable>>7) < nonprintable; CRLF auto-conversion bails on binary — Source
  • libgit2 git_blob_is_binary uses core-git heuristic: NUL scan + printable/nonprintable ratio over first 8000 bytes — Source
  • 2008 patch unified heuristics: any NUL forces binary in convert.c so CRLF handling is stricter than diff (prior convert.c used only <1% nonprintable rule, mis-handling tar/word-processor files diff called binary) — Source
  • Git mailing-list guidance: NUL in first 8000 bytes = binary; UTF-16 must be marked explicitly, Git does not handle it internally; short NUL-free binaries must also be marked explicitly — Source
  • git diff --numstat reports -\t- for binary (practical detector); git check-attr only reflects .gitattributes, not the heuristic — Source

Inferences

  • NUL-sniffing remains the recommended cheap first pass for a small daemon (scan first 8 KiB), matching Git/libgit2 behavior and user expectations.
  • It must be paired with an explicit override list (extensions + .gitattributes-style binary/-text) because UTF-16 text false-positives as binary and small high-entropy binaries without NUL false-negative as text.

Gaps

  • No 2023–2026 primary source found revising or deprecating NUL-sniffing; whether modern editors still use exactly 8000 bytes vs 4000/8000 variants (libgit2 history) is unresolved.

How to safely snapshot live sqlite files (WAL checkpoint, .dump, filesystem snapshot) vs copying the raw file?

Takeaway

Never cp/read the raw sqlite file hot: in WAL mode the consistent image spans main+-wal+-shm and byte copies tear or go stale; use the Online Backup API (sqlite3_backup_* / Connection.backup() / .backup CLI), VACUUM INTO, or a quiesced filesystem snapshot, then PRAGMA integrity_check.

Cited Findings

  • Historical cp-under-shared-lock method is fast but blocks writers, cannot copy to/from memory DBs, and risks corruption on power/OS failure — Source
  • Backup API: source read-locked only during each sqlite3_backup_step(nPage); destination write-locked throughout; incremental stepping lets writers proceed; concurrent write by another connection restarts backup automatically — Source; API contract — Source
  • Python exposes as Connection.backup(target, pages, progress, sleep); pages=-1 copies all at once (holds lock), positive pages + sleep=0.250 yields between steps; restart detected when remaining jumps back toward total — Source
  • WAL-mode live DB is three files (main + -wal committed-not-checkpointed frames + -shm index); cp/rsync/snapshot captures them at different instants → SQLITE_CORRUPT/SQLITE_NOTADB; same hazard for -journal in rollback mode — Source
  • Real-world failure: fs.copyFile() of only .db in WAL mode produced 100% corrupt backups across weeks of 6-hour cron; fix was better-sqlite3 .backup() plus hour-granular filenames — Source
  • Correct one-liners: sqlite3 app.db ".backup './db-backups/app.db'" (general) or VACUUM INTO (same safety + compaction, needs SQLite 3.27+, refuses if target exists so rm -f first); for dedup pipelines prefer .backup because compaction reshuffles pages and defeats chunking — Source
  • Borg docs: Borg just copies file as-is; if DB is written mid-read the archive may be inconsistent — use sqlite-aware method (sqlite3 db.sqlite "VACUUM INTO 'copy.sqlite'") or filesystem snapshot — Source; Borg FAQ repeats the VACUUM INTO advice — Source
  • borgmatic best practice: dump (export) databases rather than backing internal files; streams dump directly to Borg — Source
  • Checkpoint+lock schemes (wal_checkpoint(TRUNCATE) then BEGIN IMMEDIATE then copy main+wal) are fragile: checkpoint can fail to get writer lock, writers can interleave, autocheckpoint is per-process — use the backup API instead — Source
  • Post-backup rules: run PRAGMA integrity_check (expect single ok) on every finished image; pre-size target ~1.2× source; set busy_timeout ≥5000ms; write target off the source I/O queue; delete partials; on restore stop app, copy main file, delete stale -wal/-shm — Source; restore checklist — Source

Inferences

  • Small-daemon default: if file header is SQLite format 3, do not copy directly; shell to sqlite3 "$f" ".backup '$tmp'" or VACUUM INTO (when compaction desired) into a temp file, then ingest the temp file; fall back to copy only when DB is quiesced and wal_checkpoint(TRUNCATE) shows zero log pages.
  • Smaller Borg-style chunks (~64–128 KiB target, e.g. fastcdc,15,19,17) cut sqlite re-storage ~3–8× vs 2 MiB defaults at the cost of chunk-count RAM — see next section.

Gaps

  • .dump (SQL text) vs .backup (page image) size/dedup tradeoff not quantified from primary sources in this pass; anecdotal claim that dumps dedup to KBs but take hours on 12 GB DBs is single-issue-report only.

Do images/archives dedup at all, and is per-version storage of them worth it vs plain mirroring?

Takeaway

Recompressed, encrypted, or gzipped-per-version artifacts barely dedup (a 1-byte change avalanches through compression; encrypted chunks are high-entropy), while uncompressed raw images and plain sqlite files dedup well under CDC — so version verbatim binaries that change little, otherwise mirror single-copy.

Cited Findings

  • Restic maintainer: pass uncompressed data; small input change → large compressed-output change → new blob hashes → repo grows (50 GB gzipped DB dumps → 50 GB repo vs 36 GB for raw); CDC blobs identified by SHA-256 — Source
  • Duplicati skips recompression/dedup for known-compressed extensions by default (saves CPU; whole-file handling) because metadata edits rewrite the stream; identical copies still dedup, moves are free — Source
  • Borg on 5 generations of same sqlite DB: gzipped inputs = 667 unique/667 total chunks (zero dedup, 1.6 GB); uncompressed = strong dedup (590 MB total), zstd-5 beats gzip — Source
  • Borg sqlite tuning: default ~2 MiB target wastes a whole chunk per changed 4 KiB page; fastcdc,15,19,17,2 (~128 KiB target) drastically cuts incrementals; finer 14,18,16 helps more; cost is chunk-index RAM — params apply per-run so back up DBs in a separate run — Source; 12.45 GB Vintage Story sqlite: 2nd backup 518 MB default → 62 MB (15,19,17) → 37 MB (10,23,16) — Source
  • Kopia: 8 MB photo with metadata edit re-uploads ~8 MB under 4 MB splitter (chunk > file); fix is smaller splitter at repo-creation cost of more chunks — Source
  • Fixed-splitter warning: 10 GB video + 1 prepended byte re-uploads 10 GB; content-based (buzhash/rabinkarp) uploads 1–2 chunks — Source
  • Backup pipelines must chunk-then-compress (per-chunk); encrypt-then-chunk/compress is useless since ciphertext is indistinguishable from random; encrypted dedup needs weakened convergent/MLE schemes with leakage tradeoffs — Source; survey — Source
  • Kopia compresses each chunk independently (s2 default, gzip optional); splitting costs little ratio (466 MB → 119 MB standalone s2 vs 133 MB via Kopia-s2) — Source

Inferences

  • Daemon policy: (1) never store .gz/.zip/.jpg/.mp4/.sqlite.gz deltas expecting CDC wins — keep 1 mirror copy or thin retention; (2) ingest DBs/VM images uncompressed and let CDC + per-chunk compression do the work; (3) exclude or separately-schedule >100 MB media with tiny chunkers only if churn is low.
  • .dump/uncompressed SQL text is the most dedup-friendly DB form but slowest to produce; page-image .backup is the balanced default.

Gaps

  • Quantitative JPEG/MP4/ZIP re-storage ratios under Restic/Borg/Kopia CDC not found in primary docs; image-specific guidance (EXIF-only edits) relies on forum anecdotes.
  • Whether per-chunk compression (Kopia/Restic 0.16+) rescues any dedup on already-compressed inputs remains unmeasured in sources reviewed.