- No client-side encryption: plaintext content-addressed remote (SHA-256 names), no backup key, no cryptography dependency (server disk encryption is the trust model). - Filters back user data: images, archives, databases, PDFs accepted; 10 MiB cap; binary sniffing removed; temp names hardened (~$, #..#, .temp). - SQLite zero-error policy: backup-API snapshots + integrity_check, journal folding, locked/corrupt loud skips, verified restores. - Recovery: reindex from manifests, remote adopt, on-demand blob fetch, 5-day retention + thinning, GC, date-guarded remote purge, metrics. - Scheduler with WebDAV quota signal and 70% pressure backstop (floor kept). - 37 tests incl. live-monitor capture safety and DB safety.
15 KiB
Binary, Large, Image, Archive and Database Files in Versioning/Backup Tools
What size caps or chunking strategies do these tools use, and what breaks with multi-MB binaries in a per-version system (storage blowup, bandwidth)?
Takeaway
Restic, Borg and Kopia all use content-defined chunking (CDC) with ~0.5–8 MiB variable chunks so small edits to large binaries only store 1–2 new chunks; without CDC (fixed blocks or full-file copies) a 1-byte insert re-stores the whole file, and Git-LFS instead punts large files to pointer+blob storage with host-enforced per-file caps.
Cited Findings
- Restic splits files with Rabin-fingerprint CDC over a 64-byte sliding window, cutting when low 21 bits are zero; files <512 KiB are not split, blobs are 512 KiB–8 MiB, ~1 MiB average — Source; background — Source
- Restic chunker defaults aim at ~1 MiB average (
splitmask = (1<<20)-1) with configurable Min/MaxSize — Source - Borg splits files into deduplicated chunks globally across repo (all machines/archives); chunk id is a strong hash/MAC (hmac-sha256 / keyed blake3), not the rolling-hash value — Source
- Borg default buzhash params: min 2^19 (512 KiB), max 2^23 (8 MiB), mask 21 bits (~2 MiB target), window 4095 B;
fixedchunker option for disk/VM images — Source - Borg 2.x chunkers:
fastcdc(default, fastest),buzhash64/buzhash, keyed AES variants (toeplitz-aes/rabin-aesstrongest),fixedfor raw disk images where CDC gains little — Source - Borg warns fine-grained
--chunker-params=buzhash,10,23,16,4095creates huge chunk counts and RAM/disk load; coarse default suits large volumes — Source - Kopia calls CDC "splitters": FIXED vs DYNAMIC BUZHASH/RABINKARP, sizes 1M–8M (default DYNAMIC-4M-BUZHASH); small files = one content, large files split so metadata-only change to 10 GB video uploads only 1–2 chunks (<10 MB) — Source; details — Source; packing — Source
- Kopia packs many contents into 20–40 MB pack blobs; splitter choice is set at repo creation — Source; chunk-then-hash-then-compress-then-encrypt pipeline — Source
- Restic has no hard max file size but offers
--exclude-larger-than size(suffixes k/M/G/T) to skip files over a threshold — Source - Git-LFS has no inherent file-size limit; limits are host-enforced (GitHub: 2 GB Free/Pro, 4 GB Team, 5 GB Enterprise Cloud; >5 GB rejected); pointer file stores
version/oid sha256/size— Source - Git-LFS tracks by
.gitattributespattern, not by size;--aboveinmigrate importis one-shot, no automatic by-size tracking (2025--min-size/autotracksizePR still debated/unmerged) — Source - Git on Windows pre-2.34 could not smudge/clean files >4 GiB; workaround
GIT_LFS_SKIP_SMUDGE=1+git lfs pull— Source - Every Git-LFS revision counts against remote storage/bandwidth quota, so per-version binaries inflate cost; clones/pulls are slow on large repos — Source
- Keyed CDC (Borg/Restic/Kopia) is fingerprintable: observing chunk sizes of known data recovers keys (CCS 2025, eprint 2025/558; backup-service attacks eprint 2025/532); Restic 0.18 mitigates by random chunk-to-pack assignment — Source; attacks — Source; analysis — Source
Inferences
- A small local daemon should copy the Restic/Borg default: CDC ~1–2 MiB average, 512 KiB min, 8 MiB max; offer
--exclude-larger-than-style cap (e.g. 50–100 MB default-off) plus extension-based excludes rather than a hard cap. - Per-version full copies of multi-MB binaries blow up storage and bandwidth linearly with versions; CDC reduces this to delta-of-chunks but still re-reads/re-hashes the whole file each run.
- For VM/disk images with fixed internal layout, a
fixed-block mode is faster and equally effective.
Gaps
- No reliable 2026 source found on VSCode/JetBrains Local History size caps or binary handling (searches rate-limited); editor-local-history defaults remain unverified.
- Exact Kopia default content size bounds (min/max around 4 MB mean) not confirmed from primary docs; forum states 1–8 MB range only.
NUL-byte sniffing for binary detection: who does it and is it still recommended?
Takeaway
Git (and libgit2 ports) still use NUL-in-first-8000-bytes as the primary binary signal, tightened in 2008 so CRLF conversion defers to diff; it is fast and recommended as a first heuristic but known-insufficient for UTF-16 and NUL-free binaries, so explicit .gitattributes marking is required.
Cited Findings
- Git
buffer_is_binary()checks for NUL viamemchrin first 8000 bytes (FIRST_FEW_BYTES) — Source - Git
convert.cconvert_is_binary(): binary iflonecrornulor(printable>>7) < nonprintable; CRLF auto-conversion bails on binary — Source - libgit2
git_blob_is_binaryuses core-git heuristic: NUL scan + printable/nonprintable ratio over first 8000 bytes — Source - 2008 patch unified heuristics: any NUL forces binary in
convert.cso CRLF handling is stricter than diff (prior convert.c used only <1% nonprintable rule, mis-handling tar/word-processor files diff called binary) — Source - Git mailing-list guidance: NUL in first 8000 bytes = binary; UTF-16 must be marked explicitly, Git does not handle it internally; short NUL-free binaries must also be marked explicitly — Source
git diff --numstatreports-\t-for binary (practical detector);git check-attronly reflects.gitattributes, not the heuristic — Source
Inferences
- NUL-sniffing remains the recommended cheap first pass for a small daemon (scan first 8 KiB), matching Git/libgit2 behavior and user expectations.
- It must be paired with an explicit override list (extensions +
.gitattributes-stylebinary/-text) because UTF-16 text false-positives as binary and small high-entropy binaries without NUL false-negative as text.
Gaps
- No 2023–2026 primary source found revising or deprecating NUL-sniffing; whether modern editors still use exactly 8000 bytes vs 4000/8000 variants (libgit2 history) is unresolved.
How to safely snapshot live sqlite files (WAL checkpoint, .dump, filesystem snapshot) vs copying the raw file?
Takeaway
Never cp/read the raw sqlite file hot: in WAL mode the consistent image spans main+-wal+-shm and byte copies tear or go stale; use the Online Backup API (sqlite3_backup_* / Connection.backup() / .backup CLI), VACUUM INTO, or a quiesced filesystem snapshot, then PRAGMA integrity_check.
Cited Findings
- Historical
cp-under-shared-lock method is fast but blocks writers, cannot copy to/from memory DBs, and risks corruption on power/OS failure — Source - Backup API: source read-locked only during each
sqlite3_backup_step(nPage); destination write-locked throughout; incremental stepping lets writers proceed; concurrent write by another connection restarts backup automatically — Source; API contract — Source - Python exposes as
Connection.backup(target, pages, progress, sleep);pages=-1copies all at once (holds lock), positive pages +sleep=0.250yields between steps; restart detected whenremainingjumps back towardtotal— Source - WAL-mode live DB is three files (main +
-walcommitted-not-checkpointed frames +-shmindex);cp/rsync/snapshot captures them at different instants →SQLITE_CORRUPT/SQLITE_NOTADB; same hazard for-journalin rollback mode — Source - Real-world failure:
fs.copyFile()of only.dbin WAL mode produced 100% corrupt backups across weeks of 6-hour cron; fix wasbetter-sqlite3 .backup()plus hour-granular filenames — Source - Correct one-liners:
sqlite3 app.db ".backup './db-backups/app.db'"(general) orVACUUM INTO(same safety + compaction, needs SQLite 3.27+, refuses if target exists sorm -ffirst); for dedup pipelines prefer.backupbecause compaction reshuffles pages and defeats chunking — Source - Borg docs: Borg just copies file as-is; if DB is written mid-read the archive may be inconsistent — use sqlite-aware method (
sqlite3 db.sqlite "VACUUM INTO 'copy.sqlite'") or filesystem snapshot — Source; Borg FAQ repeats theVACUUM INTOadvice — Source - borgmatic best practice: dump (export) databases rather than backing internal files; streams dump directly to Borg — Source
- Checkpoint+lock schemes (
wal_checkpoint(TRUNCATE)thenBEGIN IMMEDIATEthen copy main+wal) are fragile: checkpoint can fail to get writer lock, writers can interleave, autocheckpoint is per-process — use the backup API instead — Source - Post-backup rules: run
PRAGMA integrity_check(expect singleok) on every finished image; pre-size target ~1.2× source; setbusy_timeout ≥5000ms; write target off the source I/O queue; delete partials; on restore stop app, copy main file, delete stale-wal/-shm— Source; restore checklist — Source
Inferences
- Small-daemon default: if file header is
SQLite format 3, do not copy directly; shell tosqlite3 "$f" ".backup '$tmp'"orVACUUM INTO(when compaction desired) into a temp file, then ingest the temp file; fall back to copy only when DB is quiesced andwal_checkpoint(TRUNCATE)shows zero log pages. - Smaller Borg-style chunks (~64–128 KiB target, e.g.
fastcdc,15,19,17) cut sqlite re-storage ~3–8× vs 2 MiB defaults at the cost of chunk-count RAM — see next section.
Gaps
.dump(SQL text) vs.backup(page image) size/dedup tradeoff not quantified from primary sources in this pass; anecdotal claim that dumps dedup to KBs but take hours on 12 GB DBs is single-issue-report only.
Do images/archives dedup at all, and is per-version storage of them worth it vs plain mirroring?
Takeaway
Recompressed, encrypted, or gzipped-per-version artifacts barely dedup (a 1-byte change avalanches through compression; encrypted chunks are high-entropy), while uncompressed raw images and plain sqlite files dedup well under CDC — so version verbatim binaries that change little, otherwise mirror single-copy.
Cited Findings
- Restic maintainer: pass uncompressed data; small input change → large compressed-output change → new blob hashes → repo grows (50 GB gzipped DB dumps → 50 GB repo vs 36 GB for raw); CDC blobs identified by SHA-256 — Source
- Duplicati skips recompression/dedup for known-compressed extensions by default (saves CPU; whole-file handling) because metadata edits rewrite the stream; identical copies still dedup, moves are free — Source
- Borg on 5 generations of same sqlite DB: gzipped inputs = 667 unique/667 total chunks (zero dedup, 1.6 GB); uncompressed = strong dedup (590 MB total), zstd-5 beats gzip — Source
- Borg sqlite tuning: default ~2 MiB target wastes a whole chunk per changed 4 KiB page;
fastcdc,15,19,17,2(~128 KiB target) drastically cuts incrementals; finer14,18,16helps more; cost is chunk-index RAM — params apply per-run so back up DBs in a separate run — Source; 12.45 GB Vintage Story sqlite: 2nd backup 518 MB default → 62 MB (15,19,17) → 37 MB (10,23,16) — Source - Kopia: 8 MB photo with metadata edit re-uploads ~8 MB under 4 MB splitter (chunk > file); fix is smaller splitter at repo-creation cost of more chunks — Source
- Fixed-splitter warning: 10 GB video + 1 prepended byte re-uploads 10 GB; content-based (buzhash/rabinkarp) uploads 1–2 chunks — Source
- Backup pipelines must chunk-then-compress (per-chunk); encrypt-then-chunk/compress is useless since ciphertext is indistinguishable from random; encrypted dedup needs weakened convergent/MLE schemes with leakage tradeoffs — Source; survey — Source
- Kopia compresses each chunk independently (s2 default, gzip optional); splitting costs little ratio (466 MB → 119 MB standalone s2 vs 133 MB via Kopia-s2) — Source
Inferences
- Daemon policy: (1) never store
.gz/.zip/.jpg/.mp4/.sqlite.gzdeltas expecting CDC wins — keep 1 mirror copy or thin retention; (2) ingest DBs/VM images uncompressed and let CDC + per-chunk compression do the work; (3) exclude or separately-schedule >100 MB media with tiny chunkers only if churn is low. .dump/uncompressed SQL text is the most dedup-friendly DB form but slowest to produce; page-image.backupis the balanced default.
Gaps
- Quantitative JPEG/MP4/ZIP re-storage ratios under Restic/Borg/Kopia CDC not found in primary docs; image-specific guidance (EXIF-only edits) relies on forum anecdotes.
- Whether per-chunk compression (Kopia/Restic 0.16+) rescues any dedup on already-compressed inputs remains unmeasured in sources reviewed.