Files
versioning/research_notes/Backup include exclude rules/binary-large-files.md
T

99 lines
15 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Binary, Large, Image, Archive and Database Files in Versioning/Backup Tools
## What size caps or chunking strategies do these tools use, and what breaks with multi-MB binaries in a per-version system (storage blowup, bandwidth)?
### Takeaway
Restic, Borg and Kopia all use content-defined chunking (CDC) with ~0.5–8 MiB variable chunks so small edits to large binaries only store 1–2 new chunks; without CDC (fixed blocks or full-file copies) a 1-byte insert re-stores the whole file, and Git-LFS instead punts large files to pointer+blob storage with host-enforced per-file caps.
### Cited Findings
- Restic splits files with Rabin-fingerprint CDC over a 64-byte sliding window, cutting when low 21 bits are zero; files <512 KiB are not split, blobs are 512 KiB–8 MiB, ~1 MiB average - [Source](https://github.com/restic/restic/blob/master/doc/design.rst); background - [Source](https://restic.net/blog/2015-09-12/restic-foundation1-cdc/)
- Restic chunker defaults aim at ~1 MiB average (`splitmask = (1<<20)-1`) with configurable Min/MaxSize - [Source](https://github.com/restic/chunker/blob/master/chunker.go)
- Borg splits files into deduplicated chunks globally across repo (all machines/archives); chunk id is a strong hash/MAC (hmac-sha256 / keyed blake3), not the rolling-hash value - [Source](https://borgbackup.readthedocs.io/en/master/)
- Borg default buzhash params: min 2^19 (512 KiB), max 2^23 (8 MiB), mask 21 bits (~2 MiB target), window 4095 B; `fixed` chunker option for disk/VM images - [Source](https://github.com/borgbackup/borg/blob/86fd77fd/docs/internals/data-structures.rst)
- Borg 2.x chunkers: `fastcdc` (default, fastest), `buzhash64`/`buzhash`, keyed AES variants (`toeplitz-aes`/`rabin-aes` strongest), `fixed` for raw disk images where CDC gains little - [Source](https://borgbackup.readthedocs.io/en/latest/internals/chunker.html)
- Borg warns fine-grained `--chunker-params=buzhash,10,23,16,4095` creates huge chunk counts and RAM/disk load; coarse default suits large volumes - [Source](https://borgbackup.readthedocs.io/en/stable/usage/notes.html)
- Kopia calls CDC "splitters": FIXED vs DYNAMIC BUZHASH/RABINKARP, sizes 1M–8M (default DYNAMIC-4M-BUZHASH); small files = one content, large files split so metadata-only change to 10 GB video uploads only 1–2 chunks (<10 MB) - [Source](https://kopia.discourse.group/t/does-kopia-use-content-defined-chunking-cdc/1417); details - [Source](https://kopia.discourse.group/t/difference-between-the-available-splitters/894); packing - [Source](https://kopia.discourse.group/t/do-hashed-chunks-span-multiple-files/1444)
- Kopia packs many contents into 20–40 MB pack blobs; splitter choice is set at repo creation - [Source](https://kopia.io/docs/advanced/architecture/); chunk-then-hash-then-compress-then-encrypt pipeline - [Source](https://kopia.io/docs/advanced/compression/)
- Restic has no hard max file size but offers `--exclude-larger-than size` (suffixes k/M/G/T) to skip files over a threshold - [Source](https://restic.readthedocs.io/en/stable/040_backup.html)
- Git-LFS has no inherent file-size limit; limits are host-enforced (GitHub: 2 GB Free/Pro, 4 GB Team, 5 GB Enterprise Cloud; >5 GB rejected); pointer file stores `version/oid sha256/size` - [Source](https://docs.github.com/en/repositories/working-with-files/managing-large-files/about-git-large-file-storage)
- Git-LFS tracks by `.gitattributes` pattern, not by size; `--above` in `migrate import` is one-shot, no automatic by-size tracking (2025 `--min-size/autotracksize` PR still debated/unmerged) - [Source](https://github.com/git-lfs/git-lfs/blob/main/docs/man/git-lfs-faq.adoc)
- Git on Windows pre-2.34 could not smudge/clean files >4 GiB; workaround `GIT_LFS_SKIP_SMUDGE=1` + `git lfs pull` - [Source](https://github.com/git-lfs/git-lfs/blob/main/docs/man/git-lfs-faq.adoc)
- Every Git-LFS revision counts against remote storage/bandwidth quota, so per-version binaries inflate cost; clones/pulls are slow on large repos - [Source](https://get.assembla.com/blog/git-lfs/)
- Keyed CDC (Borg/Restic/Kopia) is fingerprintable: observing chunk sizes of known data recovers keys (CCS 2025, eprint 2025/558; backup-service attacks eprint 2025/532); Restic 0.18 mitigates by random chunk-to-pack assignment - [Source](https://github.com/restic/restic/blob/master/doc/design.rst); attacks - [Source](https://eprint.iacr.org/2025/558); analysis - [Source](https://eprint.iacr.org/2025/532.pdf)
### Inferences
- A small local daemon should copy the Restic/Borg default: CDC ~1–2 MiB average, 512 KiB min, 8 MiB max; offer `--exclude-larger-than`-style cap (e.g. 50–100 MB default-off) plus extension-based excludes rather than a hard cap.
- Per-version full copies of multi-MB binaries blow up storage and bandwidth linearly with versions; CDC reduces this to delta-of-chunks but still re-reads/re-hashes the whole file each run.
- For VM/disk images with fixed internal layout, a `fixed`-block mode is faster and equally effective.
### Gaps
- No reliable 2026 source found on VSCode/JetBrains Local History size caps or binary handling (searches rate-limited); editor-local-history defaults remain unverified.
- Exact Kopia default content size bounds (min/max around 4 MB mean) not confirmed from primary docs; forum states 1–8 MB range only.
## NUL-byte sniffing for binary detection: who does it and is it still recommended?
### Takeaway
Git (and libgit2 ports) still use NUL-in-first-8000-bytes as the primary binary signal, tightened in 2008 so CRLF conversion defers to diff; it is fast and recommended as a first heuristic but known-insufficient for UTF-16 and NUL-free binaries, so explicit `.gitattributes` marking is required.
### Cited Findings
- Git `buffer_is_binary()` checks for NUL via `memchr` in first 8000 bytes (`FIRST_FEW_BYTES`) - [Source](https://stackoverflow.com/questions/6119956/how-to-determine-if-git-handles-a-file-as-binary-or-as-text)
- Git `convert.c` `convert_is_binary()`: binary if `lonecr` or `nul` or `(printable>>7) < nonprintable`; CRLF auto-conversion bails on binary - [Source](https://code.googlesource.com/git/+/645cc7a2a7274a92403d2848ef643a96f1589d09/convert.c)
- libgit2 `git_blob_is_binary` uses core-git heuristic: NUL scan + printable/nonprintable ratio over first 8000 bytes - [Source](https://libgit2.org/docs/reference/main/blob/git_blob_is_binary.html)
- 2008 patch unified heuristics: any NUL forces binary in `convert.c` so CRLF handling is stricter than diff (prior convert.c used only <1% nonprintable rule, mis-handling tar/word-processor files diff called binary) - [Source](https://public-inbox.org/git/20080116011321.GD13984@dpotapov.dyndns.org/t/)
- Git mailing-list guidance: NUL in first 8000 bytes = binary; UTF-16 must be marked explicitly, Git does not handle it internally; short NUL-free binaries must also be marked explicitly - [Source](https://public-inbox.org/git/20151202004921.GC28197@sigill.intra.peff.net/T/)
- `git diff --numstat` reports `-\t-` for binary (practical detector); `git check-attr` only reflects `.gitattributes`, not the heuristic - [Source](https://stackoverflow.com/questions/6119956/how-to-determine-if-git-handles-a-file-as-binary-or-as-text)
### Inferences
- NUL-sniffing remains the recommended cheap first pass for a small daemon (scan first 8 KiB), matching Git/libgit2 behavior and user expectations.
- It must be paired with an explicit override list (extensions + `.gitattributes`-style `binary`/`-text`) because UTF-16 text false-positives as binary and small high-entropy binaries without NUL false-negative as text.
### Gaps
- No 2023–2026 primary source found revising or deprecating NUL-sniffing; whether modern editors still use exactly 8000 bytes vs 4000/8000 variants (libgit2 history) is unresolved.
## How to safely snapshot live sqlite files (WAL checkpoint, .dump, filesystem snapshot) vs copying the raw file?
### Takeaway
Never `cp`/read the raw sqlite file hot: in WAL mode the consistent image spans main+`-wal`+`-shm` and byte copies tear or go stale; use the Online Backup API (`sqlite3_backup_*` / `Connection.backup()` / `.backup` CLI), `VACUUM INTO`, or a quiesced filesystem snapshot, then `PRAGMA integrity_check`.
### Cited Findings
- Historical `cp`-under-shared-lock method is fast but blocks writers, cannot copy to/from memory DBs, and risks corruption on power/OS failure - [Source](https://sqlite.org/backup.html)
- Backup API: source read-locked only during each `sqlite3_backup_step(nPage)`; destination write-locked throughout; incremental stepping lets writers proceed; concurrent write by another connection restarts backup automatically - [Source](https://sqlite.org/backup.html); API contract - [Source](https://sqlite.org/c3ref/backup_finish.html)
- Python exposes as `Connection.backup(target, pages, progress, sleep)`; `pages=-1` copies all at once (holds lock), positive pages + `sleep=0.250` yields between steps; restart detected when `remaining` jumps back toward `total` - [Source](https://www.productionhardening.org/backup-recovery-data-integrity/online-backup-api-hot-copies/)
- WAL-mode live DB is three files (main + `-wal` committed-not-checkpointed frames + `-shm` index); `cp`/`rsync`/snapshot captures them at different instants → `SQLITE_CORRUPT`/`SQLITE_NOTADB`; same hazard for `-journal` in rollback mode - [Source](https://www.productionhardening.org/backup-recovery-data-integrity/)
- Real-world failure: `fs.copyFile()` of only `.db` in WAL mode produced 100% corrupt backups across weeks of 6-hour cron; fix was `better-sqlite3 .backup()` plus hour-granular filenames - [Source](https://scottspence.com/posts/sqlite-corruption-fs-copyfile-issue)
- Correct one-liners: `sqlite3 app.db ".backup './db-backups/app.db'"` (general) or `VACUUM INTO` (same safety + compaction, needs SQLite 3.27+, refuses if target exists so `rm -f` first); for dedup pipelines prefer `.backup` because compaction reshuffles pages and defeats chunking - [Source](https://www.backupdata.io/resources/guides/sqlite-backups-you-can-actually-restore)
- Borg docs: Borg just copies file as-is; if DB is written mid-read the archive may be inconsistent - use sqlite-aware method (`sqlite3 db.sqlite "VACUUM INTO 'copy.sqlite'"`) or filesystem snapshot - [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst); Borg FAQ repeats the `VACUUM INTO` advice - [Source](https://borgbackup.readthedocs.io/en/master/faq.html)
- borgmatic best practice: dump (export) databases rather than backing internal files; streams dump directly to Borg - [Source](https://torsion.org/borgmatic/how-to/backup-your-databases/)
- Checkpoint+lock schemes (`wal_checkpoint(TRUNCATE)` then `BEGIN IMMEDIATE` then copy main+wal) are fragile: checkpoint can fail to get writer lock, writers can interleave, autocheckpoint is per-process - use the backup API instead - [Source](https://sqlite.org/forum/forumpost/2ea989bbe9)
- Post-backup rules: run `PRAGMA integrity_check` (expect single `ok`) on every finished image; pre-size target ~1.2× source; set `busy_timeout ≥5000ms`; write target off the source I/O queue; delete partials; on restore stop app, copy main file, delete stale `-wal`/`-shm` - [Source](https://www.productionhardening.org/backup-recovery-data-integrity/online-backup-api-hot-copies/); restore checklist - [Source](https://www.backupdata.io/resources/guides/sqlite-backups-you-can-actually-restore)
### Inferences
- Small-daemon default: if file header is `SQLite format 3`, do not copy directly; shell to `sqlite3 "$f" ".backup '$tmp'"` or `VACUUM INTO` (when compaction desired) into a temp file, then ingest the temp file; fall back to copy only when DB is quiesced and `wal_checkpoint(TRUNCATE)` shows zero log pages.
- Smaller Borg-style chunks (~64–128 KiB target, e.g. `fastcdc,15,19,17`) cut sqlite re-storage ~3–8× vs 2 MiB defaults at the cost of chunk-count RAM - see next section.
### Gaps
- `.dump` (SQL text) vs `.backup` (page image) size/dedup tradeoff not quantified from primary sources in this pass; anecdotal claim that dumps dedup to KBs but take hours on 12 GB DBs is single-issue-report only.
## Do images/archives dedup at all, and is per-version storage of them worth it vs plain mirroring?
### Takeaway
Recompressed, encrypted, or gzipped-per-version artifacts barely dedup (a 1-byte change avalanches through compression; encrypted chunks are high-entropy), while uncompressed raw images and plain sqlite files dedup well under CDC - so version verbatim binaries that change little, otherwise mirror single-copy.
### Cited Findings
- Restic maintainer: pass uncompressed data; small input change → large compressed-output change → new blob hashes → repo grows (50 GB gzipped DB dumps → 50 GB repo vs 36 GB for raw); CDC blobs identified by SHA-256 - [Source](https://github.com/restic/restic/issues/790)
- Duplicati skips recompression/dedup for known-compressed extensions by default (saves CPU; whole-file handling) because metadata edits rewrite the stream; identical copies still dedup, moves are free - [Source](https://forum.duplicati.com/t/deduplication-for-large-files/8394)
- Borg on 5 generations of same sqlite DB: gzipped inputs = 667 unique/667 total chunks (zero dedup, 1.6 GB); uncompressed = strong dedup (590 MB total), zstd-5 beats gzip - [Source](https://appsintheopen.com/posts/66-backing-up-sqlite-database-with-borg-and-de-duplication)
- Borg sqlite tuning: default ~2 MiB target wastes a whole chunk per changed 4 KiB page; `fastcdc,15,19,17,2` (~128 KiB target) drastically cuts incrementals; finer `14,18,16` helps more; cost is chunk-index RAM - params apply per-run so back up DBs in a separate run - [Source](https://borgbackup.readthedocs.io/en/master/faq.html); 12.45 GB Vintage Story sqlite: 2nd backup 518 MB default → 62 MB (15,19,17) → 37 MB (10,23,16) - [Source](https://github.com/borgbackup/borg/issues/5877)
- Kopia: 8 MB photo with metadata edit re-uploads ~8 MB under 4 MB splitter (chunk > file); fix is smaller splitter at repo-creation cost of more chunks - [Source](https://kopia.discourse.group/t/chunk-size-setting/1351)
- Fixed-splitter warning: 10 GB video + 1 prepended byte re-uploads 10 GB; content-based (buzhash/rabinkarp) uploads 1–2 chunks - [Source](https://kopia.discourse.group/t/difference-between-the-available-splitters/894)
- Backup pipelines must chunk-then-compress (per-chunk); encrypt-then-chunk/compress is useless since ciphertext is indistinguishable from random; encrypted dedup needs weakened convergent/MLE schemes with leakage tradeoffs - [Source](https://eprint.iacr.org/2025/532.pdf); survey - [Source](https://dl.acm.org/doi/10.1145/3685278)
- Kopia compresses each chunk independently (s2 default, gzip optional); splitting costs little ratio (466 MB → 119 MB standalone s2 vs 133 MB via Kopia-s2) - [Source](https://kopia.io/docs/advanced/compression/)
### Inferences
- Daemon policy: (1) never store `.gz/.zip/.jpg/.mp4/.sqlite.gz` deltas expecting CDC wins - keep 1 mirror copy or thin retention; (2) ingest DBs/VM images uncompressed and let CDC + per-chunk compression do the work; (3) exclude or separately-schedule >100 MB media with tiny chunkers only if churn is low.
- `.dump`/uncompressed SQL text is the most dedup-friendly DB form but slowest to produce; page-image `.backup` is the balanced default.
### Gaps
- Quantitative JPEG/MP4/ZIP re-storage ratios under Restic/Borg/Kopia CDC not found in primary docs; image-specific guidance (EXIF-only edits) relies on forum anecdotes.
- Whether per-chunk compression (Kopia/Restic 0.16+) rescues any dedup on already-compressed inputs remains unmeasured in sources reviewed.