Files
versioning/research_notes/Backup encryption security choices/compression-dedup-integrity.md
T

24 KiB
Raw Blame History

Compression, Deduplication, and Integrity in Restic, BorgBackup, and Kopia

Fixed-file vs content-defined chunking: algorithm, average chunk sizes, how dedup works with encryption

Takeaway

All three use content-defined chunking (CDC) by default with a fixed-size fallback option; deduplication is on plaintext hashes before compression/encryption (so no convergent encryption), with Borg and Kopia keying chunk IDs while restic uses plain SHA-256.

Cited Findings

  • Restic splits each file independently with Rabin-fingerprint CDC over a 64-byte sliding window; a random irreducible polynomial is generated at init and stored as chunker_polynomial in config to harden against watermarking - Source; background and worked example - Source
  • Restic chunking parameters: files <512 KiB are not split; blobs are 512 KiB–8 MiB with ~1 MiB average target; modified files only re-store changed blobs, robust to insertions at arbitrary offsets - Source
  • Restic 0.18.0 mitigates chunk-size fingerprinting (Alexeev/Percival/Zhang 2025) by randomly assigning chunks to pack files so attackers observing the repo cannot map chunk sizes to files - Source
  • Restic deduplication happens before encryption: blob ID is SHA-256 of plaintext; index maps plaintext hash to pack location; pack/index filenames are SHA-256 of ciphertext for accident detection only - Source; all content referenced by SHA-256 of plaintext, one blob holds data from only one file, multiple blobs packed per pack file - Source
  • Restic uses random per-repository master keys and random 16-byte IV per encryption (not convergent/deterministic encryption); identical plaintexts deduplicate via index lookup, not via identical ciphertexts - Source; external audit notes AES-256-CTR + Poly1305-AES with separate keys - Source
  • Borg default chunker is fastcdc (window-less keyed Gear hash); alternatives: fixed (fixed blocksize, optional different header block), buzhash/buzhash64 (rolling Buzhash), rabin-aes/toeplitz-aes/goldilocks-aes (UHF-then-PRF: rolling universal hash + AES-128, cut decision only on AES output) - Source; comparison table and guidance - Source
  • Borg classic buzhash defaults: CHUNK_MIN_EXP=19 (512 KiB), CHUNK_MAX_EXP=23 (8 MiB), HASH_MASK_BITS=21 (~2 MiB target), HASH_WINDOW_SIZE=4095 bytes; tunable via --chunker-params; min/max clamp where rolling-hash cuts may occur - Source; fastcdc uses CHUNK_MIN_EXP,CHUNK_MAX_EXP,HASH_MASK_BITS,NC_LEVEL with normalized chunking - Source
  • Borg chunker secrets are derived from repository key material (buzhash table XOR seed stored encrypted in keyfile; AES chunkers derive table/polynomial/AES key from id_key per-chunker domain) to resist fingerprinting; 2025 research (Truong et al. CCS 2025 eprint 2025/558; eprint 2025/532) showed keyed rolling-hash-only chunkers allow key recovery, motivating AES-based chunkers - Source
  • Borg deduplication is global across all archives/hosts/files on chunk id_hash, which is a keyed MAC over plaintext (HMAC-SHA256 for --id-hash sha256, keyed BLAKE3 for --id-hash blake3) using secret id_key, not plain hash, so attackers cannot confirm small-file presence without the key - Source; file metadata (msgpacked items) is chunked with finer params and deduplicated the same way - Source
  • Borg 1.x-compatible dedup requires buzhash; fastcdc/buzhash64 give same dedup but different cut points; fixed gives positional (not content-shift-resilient) dedup, suited to disk images - Source
  • Kopia calls chunkers “splitters”: BUZHASH, RABINKARP (rolling-hash CDC) and fixed, selectable size 1M–8M - Source; default object splitter is DYNAMIC-4M-BUZHASH (also the default for every repository create backend) - Source
  • Kopia splitter history: original DYNAMIC splitter (silvasur/buzhash dependency) was deprecated for license reasons and replaced by a faster but incompatible buzhash implementation; old repos remain readable but large objects re-upload instead of deduping across the change - Source
  • Kopia pipeline is split → hash → compare against index → discard if known; otherwise compress → encrypt → pack multiple small blocks into ~20–40 MB packs with random names; content hash default BLAKE2B-256-128 - Source; pack/index architecture - Source
  • Kopia deduplication is unaffected by compression settings because hashing precedes compression since index v2 (see next section); same file under compressed and uncompressed policies still dedupes - Source

Inferences

  • None of the three use convergent encryption (deterministic, plaintext-derived keys); all use random repository keys + random nonces/session keys, with dedup achieved by a client-side plaintext-hash index lookup before encryption.
  • Keyed chunk IDs (Borg MAC, Kopia HMAC-derived keys/format secret) hide confirmation-of-file attacks from repo-only attackers; restic's plain-SHA-256 blob IDs do not, which is part of why restic added pack-mixing and per-repo random polynomials.

Gaps

  • Exact numeric defaults for Kopia DYNAMIC-4M-BUZHASH (min/max/expected size, window bytes) were not found in fetched docs; only the 4M-average naming and 1M–8M range were confirmed.
  • Whether Kopia content IDs themselves are HMAC-keyed (vs. plain hash + separately HMAC-derived encryption keys) could not be confirmed from fetched pages; per-content key derivation via HMAC-SHA256 is confirmed but ID-keying needs source-code confirmation.

Which compression algorithms are supported, how is the algorithm recorded per object, and what happens mixing versions

Takeaway

Restic supports only zstd (repo v2, per-blob type byte + unpacked version byte); Borg supports none/lz4/zstd/zlib/lzma with per-object ctype+clevel metadata and free mixing; Kopia supports a large s2/pgzip/gzip/deflate/zstd matrix recorded per-content in index v2 and also freely mixable going forward.

Cited Findings

  • Restic repo v2 adds compression; data and tree blobs may use zstandard only; v1 has no compressed types - Source; changelog entry “Support compression for blobs (data/tree) and index/lock/snapshot files” - Source
  • Restic per-run negotiation: --compression off|fastest|auto(default)|better|max (also RESTIC_COMPRESSION); maps to klauspost/compress zstd levels SpeedFastest/SpeedDefault/SpeedBetterCompression/SpeedBestCompression with 512 KiB window and CRC disabled - Source; auto is default and uses implementation default level - Source; tuning doc confirms auto default for v2 repos - Source
  • Restic pack header per-blob recording: 1-byte type 0b00 data / 0b01 tree / 0b10 compressed-data / 0b11 compressed-tree, followed by Length(encrypted_blob)||Hash(plaintext) (uncompressed) or Length(encrypted_blob)||Length(plaintext)||Hash(plaintext) (compressed), little-endian uint32; index adds uncompressed_length only for compressed blobs - Source
  • Restic unpacked files (index/snapshot/lock): plaintext is encoding_version||data with 1-byte version; [ (0x5b)/{ (0x7b) mean “whole plaintext is JSON” (v1 back-compat); version 2 means zstd-compressed JSON; new v2 writes always version 2; implementation compresses unpacked data before encryption via compressUnpacked, and compresses tree blobs even when --compression off for data blobs (if Compression != off || t != DataBlob) - Source; code - Source; design rationale (null-byte/version-byte history) - Source
  • Restic mixing/compat: compressed and uncompressed blobs of same type may be mixed in one pack; in v2, data and tree blobs must be in separate packs; v1-repo data remains valid in v2 (no re-upload required); new data compressed per run-level, prune recompresses only repacked chunks; v2 repos unreadable by pre-compression restic versions; new repos default to v2 since 0.14.0 so compression is on by default - Source; release/upgrade behavior - Source
  • Restic dedup uses plaintext hash, so future zstd-output changes do not break dedup (compressed bytes never hashed for identity) - Source
  • Borg compression set: none (0x00), lz4 (0x01), zstd (0x03, levels -128..22), zlib (0x05, levels 0–9), lzma (0x02, levels 0–9); ctype byte + clevel byte (zstd int8_t so -1→255, -128→128; others unsigned with 255 = n/a); speed order none>lz4>zlib>lzma, lz4>zstd; compression lzma>zlib>lz4>none, zstd>lz4 - Source; legacy 1.x zlib has no ID bytes (detected by 0x.8 header) - Source
  • Borg per-object recording (borg2): msgpacked metadata dict holds ctype, clevel, csize (compressed+obfuscated size), psize (payload w/o obfuscation trailer, when obfuscated), olevel, size (uncompressed), type (ro_type A/C/S/F); metadata and data slots separately encrypted with header bound as AAD - Source; code RepoObj.format/parse - Source; compressor base auto-detection via ID header - Source
  • Borg negotiation/mixing: default compression is lz4; mixing methods in one repo is fine because dedup is on source chunks, not compressed bytes; first writer of a chunk determines its stored compression; borg recreate/repo-compress can recompress; wrappers auto,C[,L] (lz4 compressibility heuristic → none vs C) and obfuscate,SPEC,C[,L] (Padmé deterministic padding, ≤12% overhead, MAX_DATA_SIZE ~20 MiB cap) - Source
  • Kopia compression is disabled by default and controlled per-policy (global/host/path: --compression=..., min/max file size, extensions); new setting applies going forward only, does not retroactively recompress - Source
  • Kopia algorithm menu: none|deflate-best-compression|deflate-best-speed|deflate-default|gzip|gzip-best-compression|gzip-best-speed|pgzip|pgzip-best-compression|pgzip-best-speed|s2-better|s2-default|s2-parallel-4|s2-parallel-8|zstd|zstd-better-compression|zstd-fastest (zstd recommended default choice); benchmark table shows s2 fastest (~GB/s, largest), zstd smallest, pgzip balanced - Source; details and memory/I-O guidance - Source
  • Kopia per-object recording: since content-level compression (v0.9, --index-version=2), compression done after hashing; per-content compression ID kept in content manager bookkeeping, visible via kopia content list -c / content stats (e.g. (uncompressed) vs zstd vs zstd-fastest with counts/sizes), no longer encoded as Z-prefixed content ID - Source; if compressed chunk grows, original stored uncompressed - Source
  • Kopia mixing/compat: dedup unaffected by algorithm changes or library-output drift because ID is pre-compression hash; enables future recompression during maintenance; changing policy does not rewrite old contents; newer Kopia cannot read legacy LZ4-compressed contents - must migrate with an older version (restore/repack) before upgrading - Source; LZ4 removal notice - Source

Inferences

  • Recording compression outside the content ID (restic type bits/index field, Borg metadata dict, Kopia index v2) is what allows free mixing and future recompression without breaking content addressing.
  • Restic's single-codec choice simplifies negotiation to levels only; Borg/Kopia pay codec-matrix complexity for finer speed/ratio control.

Gaps

  • Exact zstd numeric levels behind restic auto/better/max (klauspost SpeedDefault etc. numeric equivalents) were not pinned to zstd CLI levels in fetched sources.
  • Kopia on-disk bytes for per-content compression IDs (enum values/header layout) were not found; only CLI-visible behavior is documented.

How is integrity verified (HMAC, AEAD tag, checksums, parity)? How are manifests/snapshots authenticated?

Takeaway

Restic uses Encrypt-then-MAC (AES-256-CTR + Poly1305-AES) on every blob/file plus SHA-256 content addressing; Borg 2 uses AEAD (AES-OCB or ChaCha20-Poly1305) with chunk-ID-as-AAD plus layered checksums; Kopia uses AEAD (AES-GCM or ChaCha20-Poly1305) with HMAC-derived per-content keys plus optional Reed-Solomon ECC - none provide parity by default.

Cited Findings

  • Restic envelope: all files except keys/ (and pack containers) are IV(16)||CIPHERTEXT||MAC(16) (32 B overhead, random IV per file); pack files hold multiple independently encrypted/authenticated blobs + encrypted header + LE Header_Length; primitives AES-256-CTR + Poly1305-AES, Encrypt-then-MAC (MAC over ciphertext), keys via scrypt KDF → 32 B enc key + 32 B MAC key (k+r) unlocking master keys in keys/ JSON - Source; audit summary - Source
  • Restic content integrity: filenames are hex SHA-256 of file ciphertext (verifiable with sha256sum); trees/data addressed by SHA-256 of plaintext with deterministic JSON for trees; restic check verifies structure plus optional --read-data payload reads; tampered data fails MAC and is not decrypted - Source
  • Restic snapshots/manifests: snapshots are JSON (time/tree/paths/hostname/...) stored as unpacked encrypted files (v2: version-byte + zstd then IV||C||MAC), filename = storage ID; snapshot→tree→blob DAG; no separate manifest signature - authentication is the per-file MAC + content-hash reference; write-order rules (packs → index → snapshot; read reverse) keep repo consistent - Source
  • Restic has no parity/ECC; repair is rebuild-index + re-backup + check; threat model explicitly excludes deletion protection - Source
  • Borg 2 AEAD modes: --encryption aes256-ocb|chacha20-poly1305 × --id-hash sha256|blake3, orthogonal to --key-location repokey|keyfile; per-session random sessionid, session_key via SHA-256 KDF, counter IVs; each object has two separately encrypted slots (metadata + data) behind unencrypted header (OBJ_MAGIC||version||chunk_id||meta_size||data_size); AAD = header||slot_tag||id||..., tag authenticates metadata, payload, header prefix and chunk ID - Source
  • Borg chunk-ID binding: id = MAC(id_key, plaintext); AEAD tag binds ID to ciphertext, so repo cannot swap content under an ID; post-decrypt id == MAC(decompressed) check is optional by default (only detects malicious client writes) but enforced by borg check --verify-data / BORG_ASSERT_ID - Source
  • Borg manifest/archive authentication (Horton principle DAG): every object referenced by parent ID up to manifest; borg2 stores ro_type in metadata and verifies expected vs. actual type, binding meaning + ID via AAD; no TAM in borg2; borg 1.x used TAM (HKDF-SHA-512(id_key||enc_key||enc_hmac_key, 64 B salt, "borg-metadata-authentication-manifest") → HMAC over packed manifest) to anchor fixed-ID manifest (CVE-2016-10099) - Source
  • Borg non-encrypting modes: authenticated-sha256|blake3 carry keyed MAC (not encryption) over payload+header+slot; none-* carry only unkeyed checksums (accidental-corruption detection, no tamper protection); legacy 1.x modes were AES-CTR Encrypt-then-MAC with IV reservation; key blobs wrapped by argon2-derived KEK + chacha20-poly1305 (IV 0) - Source
  • Borg layered checksums: IntegrityCheckedFile streaming checksums for cache/index/hints with integrity.<TXN> msgpack files and [integrity] cache-config section; corrupt index/hints are deleted and rebuilt; config/config (version/id/encryption/id_hash) is plaintext/unauthenticated, protected instead by client security-dir EncryptionMethodMismatch checks - Source
  • Kopia content encryption: default AES256-GCM-HMAC-SHA256 (alt CHACHA20-POLY1305-HMAC-SHA256), set at repository create and immutable after; per-content AEAD keys derived via HMAC-SHA256; blocks packed into 20 MB packs (20–40 MB on wire) with local index at pack end + top-level index mapping block ID → (blob, offset, length) - Source; cipher history (deprecated unauthenticated AES-CTR/SALSA20) - Source
  • Kopia format/manifest envelope: kopia.repository format blob holds uniqueID (salt), keyAlgo scrypt-65536-8-1, encryption AES256_GCM, encryptedBlockFormat = JSON (ContentFormat{version,hash,encryption,HMACSecret 32 B,MasterKey 32 B,MaxPackSize 20 MiB} + ObjectFormat{splitter}) encrypted with passphrase-derived Km=PBKDF(pass,uniqueID), Ke=HKDF(SHA256,Km,uniqueID,"AES",32), AD=HKDF(...,"CHECKSUM",32); snapshots/manifests are ordinary encrypted contents in the same CABS/object layers - Source; defaults (--block-hash BLAKE2B-256-128, --object-splitter DYNAMIC-4M-BUZHASH, --encryption AES256-GCM-HMAC-SHA256) - Source
  • Kopia extra redundancy: experimental Reed-Solomon REED-SOLOMON-CRC32 ECC with --ecc-overhead-percent (default 0 = disabled); must be enabled at creation, cannot be added later; cloud backends already ECC-protected so often redundant - Source; consistency/verify docs cover validity checks and repair - Source

Inferences

  • All three authenticate snapshots/manifests the same way as data (no detached signatures); trust anchors are the client-held keys (restic master keys, Borg key+TAM/AAD chain, Kopia format-block passphrase), giving a key-anchored DAG from manifest to chunks.
  • Only Kopia offers built-in parity (ECC); restic/Borg rely on backend durability + authenticated detection and rebuild/re-upload repair.

Gaps

  • Kopia manifest-object specifics (manifest content type, index-epoch authentication, kopia snapshot verify coverage) were not fetched; snapshot authentication is inferred from “all contents encrypted/authenticated” architecture.
  • Exact Borg MAX_DATA_SIZE/obfuscation interaction with AEAD tag and BORG_ASSERT_ID default scope need code-level confirmation beyond docs snippets.

Order of operations: compress-then-encrypt vs alternatives, and why

Takeaway

All three do hash → compress → encrypt (dedup first, compress second, encrypt last); encrypting last is mandatory because ciphertext is incompressible and randomized, while hashing/compressing first preserves dedup and ratio.

Cited Findings

  • Restic code order in saveAndEncrypt: plaintext hash for ID/dedup (saveBlob → Hash(buf)) → zstd compress (if v2 and enabled) → random nonce → Seal (AES-CTR + Poly1305 MAC) → pack; unpacked files: compressUnpacked (prepend version byte + zstd) → Seal - Source; PR states “Unpacked files like lock, index and snapshot files are also compressed before encryption” - Source
  • Restic why: backing up pre-compressed (e.g. .gz) data defeats CDC dedup because small input changes avalanche through the compressor and shift cut points; project advises backing up uncompressed data and letting restic chunk-then-compress per blob; chunk-then-compress discussion - Source; zstd-dictionary debate concludes no dictionary needed precisely because compression runs after chunking/dedup - Source
  • Borg order: id = MAC(id_key, data) → compressed = compress(data) → AEAD_encrypt(session_key, iv, compressed, aad=id||header) (borg2); legacy: id=AUTH(data) → compress → AES-CTR → MAC(encrypted) - Source; docs state “Compression is applied after deduplication, thus using different compression methods in one repo does not influence deduplication” - Source
  • Kopia order: “splits into chunks → hash → compare → if new, compress chunk → encrypt → pack” - Source; content-level compression PR explicitly rewired pipeline from read → split → compress → hash → store to read → split → hash → compress → store to stop compressor-version drift from changing IDs and breaking dedup and to enable server-side/recompression - Source
  • General rationale shared across docs: encryption must be last because (a) compressed ciphertext does not compress, (b) randomized encryption (IV/session key) destroys dedup if hashed after, (c) MAC/AEAD must cover the stored (compressed) bytes; compressing whole files before chunking is measurably worse than chunk-then-compress (Kopia split-vs-whole table: s2 28.6% vs 25.6%, gzip 18.8% vs 17.3% - small loss accepted for dedup wins) - Source

Inferences

  • The universal pattern is dedup-on-plaintext → compress → authenticated-encrypt; any deviation (pre-compressed inputs, compress-after-encrypt, hash-after-compress) is documented as an anti-pattern in all three projects’ issues/docs.
  • Per-chunk compression inherently sacrifices cross-chunk dictionary context; all three accept slightly worse ratios in exchange for shift-resilient dedup and independent integrity domains.

Gaps

  • No source quantified restic’s incompressible-blob shortcut (whether zstd EncodeAll output larger than input is stored raw vs. kept); forum claims “stores raw data” but code path stores uncompressedLength flag - exact skip-threshold logic needs code confirmation.