# Binary, Large, Image, Archive and Database Files in Versioning/Backup Tools ## What size caps or chunking strategies do these tools use, and what breaks with multi-MB binaries in a per-version system (storage blowup, bandwidth)? ### Takeaway Restic, Borg and Kopia all use content-defined chunking (CDC) with ~0.5–8 MiB variable chunks so small edits to large binaries only store 1–2 new chunks; without CDC (fixed blocks or full-file copies) a 1-byte insert re-stores the whole file, and Git-LFS instead punts large files to pointer+blob storage with host-enforced per-file caps. ### Cited Findings - Restic splits files with Rabin-fingerprint CDC over a 64-byte sliding window, cutting when low 21 bits are zero; files <512 KiB are not split, blobs are 512 KiB–8 MiB, ~1 MiB average - [Source](https://github.com/restic/restic/blob/master/doc/design.rst); background - [Source](https://restic.net/blog/2015-09-12/restic-foundation1-cdc/) - Restic chunker defaults aim at ~1 MiB average (`splitmask = (1<<20)-1`) with configurable Min/MaxSize - [Source](https://github.com/restic/chunker/blob/master/chunker.go) - Borg splits files into deduplicated chunks globally across repo (all machines/archives); chunk id is a strong hash/MAC (hmac-sha256 / keyed blake3), not the rolling-hash value - [Source](https://borgbackup.readthedocs.io/en/master/) - Borg default buzhash params: min 2^19 (512 KiB), max 2^23 (8 MiB), mask 21 bits (~2 MiB target), window 4095 B; `fixed` chunker option for disk/VM images - [Source](https://github.com/borgbackup/borg/blob/86fd77fd/docs/internals/data-structures.rst) - Borg 2.x chunkers: `fastcdc` (default, fastest), `buzhash64`/`buzhash`, keyed AES variants (`toeplitz-aes`/`rabin-aes` strongest), `fixed` for raw disk images where CDC gains little - [Source](https://borgbackup.readthedocs.io/en/latest/internals/chunker.html) - Borg warns fine-grained `--chunker-params=buzhash,10,23,16,4095` creates huge chunk counts and RAM/disk load; coarse default suits large volumes - [Source](https://borgbackup.readthedocs.io/en/stable/usage/notes.html) - Kopia calls CDC "splitters": FIXED vs DYNAMIC BUZHASH/RABINKARP, sizes 1M–8M (default DYNAMIC-4M-BUZHASH); small files = one content, large files split so metadata-only change to 10 GB video uploads only 1–2 chunks (<10 MB) - [Source](https://kopia.discourse.group/t/does-kopia-use-content-defined-chunking-cdc/1417); details - [Source](https://kopia.discourse.group/t/difference-between-the-available-splitters/894); packing - [Source](https://kopia.discourse.group/t/do-hashed-chunks-span-multiple-files/1444) - Kopia packs many contents into 20–40 MB pack blobs; splitter choice is set at repo creation - [Source](https://kopia.io/docs/advanced/architecture/); chunk-then-hash-then-compress-then-encrypt pipeline - [Source](https://kopia.io/docs/advanced/compression/) - Restic has no hard max file size but offers `--exclude-larger-than size` (suffixes k/M/G/T) to skip files over a threshold - [Source](https://restic.readthedocs.io/en/stable/040_backup.html) - Git-LFS has no inherent file-size limit; limits are host-enforced (GitHub: 2 GB Free/Pro, 4 GB Team, 5 GB Enterprise Cloud; >5 GB rejected); pointer file stores `version/oid sha256/size` - [Source](https://docs.github.com/en/repositories/working-with-files/managing-large-files/about-git-large-file-storage) - Git-LFS tracks by `.gitattributes` pattern, not by size; `--above` in `migrate import` is one-shot, no automatic by-size tracking (2025 `--min-size/autotracksize` PR still debated/unmerged) - [Source](https://github.com/git-lfs/git-lfs/blob/main/docs/man/git-lfs-faq.adoc) - Git on Windows pre-2.34 could not smudge/clean files >4 GiB; workaround `GIT_LFS_SKIP_SMUDGE=1` + `git lfs pull` - [Source](https://github.com/git-lfs/git-lfs/blob/main/docs/man/git-lfs-faq.adoc) - Every Git-LFS revision counts against remote storage/bandwidth quota, so per-version binaries inflate cost; clones/pulls are slow on large repos - [Source](https://get.assembla.com/blog/git-lfs/) - Keyed CDC (Borg/Restic/Kopia) is fingerprintable: observing chunk sizes of known data recovers keys (CCS 2025, eprint 2025/558; backup-service attacks eprint 2025/532); Restic 0.18 mitigates by random chunk-to-pack assignment - [Source](https://github.com/restic/restic/blob/master/doc/design.rst); attacks - [Source](https://eprint.iacr.org/2025/558); analysis - [Source](https://eprint.iacr.org/2025/532.pdf) ### Inferences - A small local daemon should copy the Restic/Borg default: CDC ~1–2 MiB average, 512 KiB min, 8 MiB max; offer `--exclude-larger-than`-style cap (e.g. 50–100 MB default-off) plus extension-based excludes rather than a hard cap. - Per-version full copies of multi-MB binaries blow up storage and bandwidth linearly with versions; CDC reduces this to delta-of-chunks but still re-reads/re-hashes the whole file each run. - For VM/disk images with fixed internal layout, a `fixed`-block mode is faster and equally effective. ### Gaps - No reliable 2026 source found on VSCode/JetBrains Local History size caps or binary handling (searches rate-limited); editor-local-history defaults remain unverified. - Exact Kopia default content size bounds (min/max around 4 MB mean) not confirmed from primary docs; forum states 1–8 MB range only. ## NUL-byte sniffing for binary detection: who does it and is it still recommended? ### Takeaway Git (and libgit2 ports) still use NUL-in-first-8000-bytes as the primary binary signal, tightened in 2008 so CRLF conversion defers to diff; it is fast and recommended as a first heuristic but known-insufficient for UTF-16 and NUL-free binaries, so explicit `.gitattributes` marking is required. ### Cited Findings - Git `buffer_is_binary()` checks for NUL via `memchr` in first 8000 bytes (`FIRST_FEW_BYTES`) - [Source](https://stackoverflow.com/questions/6119956/how-to-determine-if-git-handles-a-file-as-binary-or-as-text) - Git `convert.c` `convert_is_binary()`: binary if `lonecr` or `nul` or `(printable>>7) < nonprintable`; CRLF auto-conversion bails on binary - [Source](https://code.googlesource.com/git/+/645cc7a2a7274a92403d2848ef643a96f1589d09/convert.c) - libgit2 `git_blob_is_binary` uses core-git heuristic: NUL scan + printable/nonprintable ratio over first 8000 bytes - [Source](https://libgit2.org/docs/reference/main/blob/git_blob_is_binary.html) - 2008 patch unified heuristics: any NUL forces binary in `convert.c` so CRLF handling is stricter than diff (prior convert.c used only <1% nonprintable rule, mis-handling tar/word-processor files diff called binary) - [Source](https://public-inbox.org/git/20080116011321.GD13984@dpotapov.dyndns.org/t/) - Git mailing-list guidance: NUL in first 8000 bytes = binary; UTF-16 must be marked explicitly, Git does not handle it internally; short NUL-free binaries must also be marked explicitly - [Source](https://public-inbox.org/git/20151202004921.GC28197@sigill.intra.peff.net/T/) - `git diff --numstat` reports `-\t-` for binary (practical detector); `git check-attr` only reflects `.gitattributes`, not the heuristic - [Source](https://stackoverflow.com/questions/6119956/how-to-determine-if-git-handles-a-file-as-binary-or-as-text) ### Inferences - NUL-sniffing remains the recommended cheap first pass for a small daemon (scan first 8 KiB), matching Git/libgit2 behavior and user expectations. - It must be paired with an explicit override list (extensions + `.gitattributes`-style `binary`/`-text`) because UTF-16 text false-positives as binary and small high-entropy binaries without NUL false-negative as text. ### Gaps - No 2023–2026 primary source found revising or deprecating NUL-sniffing; whether modern editors still use exactly 8000 bytes vs 4000/8000 variants (libgit2 history) is unresolved. ## How to safely snapshot live sqlite files (WAL checkpoint, .dump, filesystem snapshot) vs copying the raw file? ### Takeaway Never `cp`/read the raw sqlite file hot: in WAL mode the consistent image spans main+`-wal`+`-shm` and byte copies tear or go stale; use the Online Backup API (`sqlite3_backup_*` / `Connection.backup()` / `.backup` CLI), `VACUUM INTO`, or a quiesced filesystem snapshot, then `PRAGMA integrity_check`. ### Cited Findings - Historical `cp`-under-shared-lock method is fast but blocks writers, cannot copy to/from memory DBs, and risks corruption on power/OS failure - [Source](https://sqlite.org/backup.html) - Backup API: source read-locked only during each `sqlite3_backup_step(nPage)`; destination write-locked throughout; incremental stepping lets writers proceed; concurrent write by another connection restarts backup automatically - [Source](https://sqlite.org/backup.html); API contract - [Source](https://sqlite.org/c3ref/backup_finish.html) - Python exposes as `Connection.backup(target, pages, progress, sleep)`; `pages=-1` copies all at once (holds lock), positive pages + `sleep=0.250` yields between steps; restart detected when `remaining` jumps back toward `total` - [Source](https://www.productionhardening.org/backup-recovery-data-integrity/online-backup-api-hot-copies/) - WAL-mode live DB is three files (main + `-wal` committed-not-checkpointed frames + `-shm` index); `cp`/`rsync`/snapshot captures them at different instants → `SQLITE_CORRUPT`/`SQLITE_NOTADB`; same hazard for `-journal` in rollback mode - [Source](https://www.productionhardening.org/backup-recovery-data-integrity/) - Real-world failure: `fs.copyFile()` of only `.db` in WAL mode produced 100% corrupt backups across weeks of 6-hour cron; fix was `better-sqlite3 .backup()` plus hour-granular filenames - [Source](https://scottspence.com/posts/sqlite-corruption-fs-copyfile-issue) - Correct one-liners: `sqlite3 app.db ".backup './db-backups/app.db'"` (general) or `VACUUM INTO` (same safety + compaction, needs SQLite 3.27+, refuses if target exists so `rm -f` first); for dedup pipelines prefer `.backup` because compaction reshuffles pages and defeats chunking - [Source](https://www.backupdata.io/resources/guides/sqlite-backups-you-can-actually-restore) - Borg docs: Borg just copies file as-is; if DB is written mid-read the archive may be inconsistent - use sqlite-aware method (`sqlite3 db.sqlite "VACUUM INTO 'copy.sqlite'"`) or filesystem snapshot - [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst); Borg FAQ repeats the `VACUUM INTO` advice - [Source](https://borgbackup.readthedocs.io/en/master/faq.html) - borgmatic best practice: dump (export) databases rather than backing internal files; streams dump directly to Borg - [Source](https://torsion.org/borgmatic/how-to/backup-your-databases/) - Checkpoint+lock schemes (`wal_checkpoint(TRUNCATE)` then `BEGIN IMMEDIATE` then copy main+wal) are fragile: checkpoint can fail to get writer lock, writers can interleave, autocheckpoint is per-process - use the backup API instead - [Source](https://sqlite.org/forum/forumpost/2ea989bbe9) - Post-backup rules: run `PRAGMA integrity_check` (expect single `ok`) on every finished image; pre-size target ~1.2× source; set `busy_timeout ≥5000ms`; write target off the source I/O queue; delete partials; on restore stop app, copy main file, delete stale `-wal`/`-shm` - [Source](https://www.productionhardening.org/backup-recovery-data-integrity/online-backup-api-hot-copies/); restore checklist - [Source](https://www.backupdata.io/resources/guides/sqlite-backups-you-can-actually-restore) ### Inferences - Small-daemon default: if file header is `SQLite format 3`, do not copy directly; shell to `sqlite3 "$f" ".backup '$tmp'"` or `VACUUM INTO` (when compaction desired) into a temp file, then ingest the temp file; fall back to copy only when DB is quiesced and `wal_checkpoint(TRUNCATE)` shows zero log pages. - Smaller Borg-style chunks (~64–128 KiB target, e.g. `fastcdc,15,19,17`) cut sqlite re-storage ~3–8× vs 2 MiB defaults at the cost of chunk-count RAM - see next section. ### Gaps - `.dump` (SQL text) vs `.backup` (page image) size/dedup tradeoff not quantified from primary sources in this pass; anecdotal claim that dumps dedup to KBs but take hours on 12 GB DBs is single-issue-report only. ## Do images/archives dedup at all, and is per-version storage of them worth it vs plain mirroring? ### Takeaway Recompressed, encrypted, or gzipped-per-version artifacts barely dedup (a 1-byte change avalanches through compression; encrypted chunks are high-entropy), while uncompressed raw images and plain sqlite files dedup well under CDC - so version verbatim binaries that change little, otherwise mirror single-copy. ### Cited Findings - Restic maintainer: pass uncompressed data; small input change → large compressed-output change → new blob hashes → repo grows (50 GB gzipped DB dumps → 50 GB repo vs 36 GB for raw); CDC blobs identified by SHA-256 - [Source](https://github.com/restic/restic/issues/790) - Duplicati skips recompression/dedup for known-compressed extensions by default (saves CPU; whole-file handling) because metadata edits rewrite the stream; identical copies still dedup, moves are free - [Source](https://forum.duplicati.com/t/deduplication-for-large-files/8394) - Borg on 5 generations of same sqlite DB: gzipped inputs = 667 unique/667 total chunks (zero dedup, 1.6 GB); uncompressed = strong dedup (590 MB total), zstd-5 beats gzip - [Source](https://appsintheopen.com/posts/66-backing-up-sqlite-database-with-borg-and-de-duplication) - Borg sqlite tuning: default ~2 MiB target wastes a whole chunk per changed 4 KiB page; `fastcdc,15,19,17,2` (~128 KiB target) drastically cuts incrementals; finer `14,18,16` helps more; cost is chunk-index RAM - params apply per-run so back up DBs in a separate run - [Source](https://borgbackup.readthedocs.io/en/master/faq.html); 12.45 GB Vintage Story sqlite: 2nd backup 518 MB default → 62 MB (15,19,17) → 37 MB (10,23,16) - [Source](https://github.com/borgbackup/borg/issues/5877) - Kopia: 8 MB photo with metadata edit re-uploads ~8 MB under 4 MB splitter (chunk > file); fix is smaller splitter at repo-creation cost of more chunks - [Source](https://kopia.discourse.group/t/chunk-size-setting/1351) - Fixed-splitter warning: 10 GB video + 1 prepended byte re-uploads 10 GB; content-based (buzhash/rabinkarp) uploads 1–2 chunks - [Source](https://kopia.discourse.group/t/difference-between-the-available-splitters/894) - Backup pipelines must chunk-then-compress (per-chunk); encrypt-then-chunk/compress is useless since ciphertext is indistinguishable from random; encrypted dedup needs weakened convergent/MLE schemes with leakage tradeoffs - [Source](https://eprint.iacr.org/2025/532.pdf); survey - [Source](https://dl.acm.org/doi/10.1145/3685278) - Kopia compresses each chunk independently (s2 default, gzip optional); splitting costs little ratio (466 MB → 119 MB standalone s2 vs 133 MB via Kopia-s2) - [Source](https://kopia.io/docs/advanced/compression/) ### Inferences - Daemon policy: (1) never store `.gz/.zip/.jpg/.mp4/.sqlite.gz` deltas expecting CDC wins - keep 1 mirror copy or thin retention; (2) ingest DBs/VM images uncompressed and let CDC + per-chunk compression do the work; (3) exclude or separately-schedule >100 MB media with tiny chunkers only if churn is low. - `.dump`/uncompressed SQL text is the most dedup-friendly DB form but slowest to produce; page-image `.backup` is the balanced default. ### Gaps - Quantitative JPEG/MP4/ZIP re-storage ratios under Restic/Borg/Kopia CDC not found in primary docs; image-specific guidance (EXIF-only edits) relies on forum anecdotes. - Whether per-chunk compression (Kopia/Restic 0.16+) rescues any dedup on already-compressed inputs remains unmeasured in sources reviewed.