24 KiB
Compression, Deduplication, and Integrity in Restic, BorgBackup, and Kopia
Fixed-file vs content-defined chunking: algorithm, average chunk sizes, how dedup works with encryption
Takeaway
All three use content-defined chunking (CDC) by default with a fixed-size fallback option; deduplication is on plaintext hashes before compression/encryption (so no convergent encryption), with Borg and Kopia keying chunk IDs while restic uses plain SHA-256.
Cited Findings
- Restic splits each file independently with Rabin-fingerprint CDC over a 64-byte sliding window; a random irreducible polynomial is generated at
initand stored aschunker_polynomialinconfigto harden against watermarking - Source; background and worked example - Source - Restic chunking parameters: files <512 KiB are not split; blobs are 512 KiB–8 MiB with ~1 MiB average target; modified files only re-store changed blobs, robust to insertions at arbitrary offsets - Source
- Restic 0.18.0 mitigates chunk-size fingerprinting (Alexeev/Percival/Zhang 2025) by randomly assigning chunks to pack files so attackers observing the repo cannot map chunk sizes to files - Source
- Restic deduplication happens before encryption: blob ID is SHA-256 of plaintext; index maps plaintext hash to pack location; pack/index filenames are SHA-256 of ciphertext for accident detection only - Source; all content referenced by SHA-256 of plaintext, one blob holds data from only one file, multiple blobs packed per pack file - Source
- Restic uses random per-repository master keys and random 16-byte IV per encryption (not convergent/deterministic encryption); identical plaintexts deduplicate via index lookup, not via identical ciphertexts - Source; external audit notes AES-256-CTR + Poly1305-AES with separate keys - Source
- Borg default chunker is
fastcdc(window-less keyed Gear hash); alternatives:fixed(fixed blocksize, optional different header block),buzhash/buzhash64(rolling Buzhash),rabin-aes/toeplitz-aes/goldilocks-aes(UHF-then-PRF: rolling universal hash + AES-128, cut decision only on AES output) - Source; comparison table and guidance - Source - Borg classic
buzhashdefaults:CHUNK_MIN_EXP=19(512 KiB),CHUNK_MAX_EXP=23(8 MiB),HASH_MASK_BITS=21(~2 MiB target),HASH_WINDOW_SIZE=4095bytes; tunable via--chunker-params; min/max clamp where rolling-hash cuts may occur - Source;fastcdcusesCHUNK_MIN_EXP,CHUNK_MAX_EXP,HASH_MASK_BITS,NC_LEVELwith normalized chunking - Source - Borg chunker secrets are derived from repository key material (buzhash table XOR seed stored encrypted in keyfile; AES chunkers derive table/polynomial/AES key from
id_keyper-chunker domain) to resist fingerprinting; 2025 research (Truong et al. CCS 2025 eprint 2025/558; eprint 2025/532) showed keyed rolling-hash-only chunkers allow key recovery, motivating AES-based chunkers - Source - Borg deduplication is global across all archives/hosts/files on chunk
id_hash, which is a keyed MAC over plaintext (HMAC-SHA256for--id-hash sha256, keyedBLAKE3for--id-hash blake3) using secretid_key, not plain hash, so attackers cannot confirm small-file presence without the key - Source; file metadata (msgpacked items) is chunked with finer params and deduplicated the same way - Source - Borg 1.x-compatible dedup requires
buzhash;fastcdc/buzhash64give same dedup but different cut points;fixedgives positional (not content-shift-resilient) dedup, suited to disk images - Source - Kopia calls chunkers “splitters”:
BUZHASH,RABINKARP(rolling-hash CDC) andfixed, selectable size 1M–8M - Source; default object splitter isDYNAMIC-4M-BUZHASH(also the default for everyrepository createbackend) - Source - Kopia splitter history: original
DYNAMICsplitter (silvasur/buzhash dependency) was deprecated for license reasons and replaced by a faster but incompatible buzhash implementation; old repos remain readable but large objects re-upload instead of deduping across the change - Source - Kopia pipeline is split → hash → compare against index → discard if known; otherwise compress → encrypt → pack multiple small blocks into ~20–40 MB packs with random names; content hash default
BLAKE2B-256-128- Source; pack/index architecture - Source - Kopia deduplication is unaffected by compression settings because hashing precedes compression since index v2 (see next section); same file under compressed and uncompressed policies still dedupes - Source
Inferences
- None of the three use convergent encryption (deterministic, plaintext-derived keys); all use random repository keys + random nonces/session keys, with dedup achieved by a client-side plaintext-hash index lookup before encryption.
- Keyed chunk IDs (Borg MAC, Kopia HMAC-derived keys/format secret) hide confirmation-of-file attacks from repo-only attackers; restic's plain-SHA-256 blob IDs do not, which is part of why restic added pack-mixing and per-repo random polynomials.
Gaps
- Exact numeric defaults for Kopia
DYNAMIC-4M-BUZHASH(min/max/expected size, window bytes) were not found in fetched docs; only the 4M-average naming and 1M–8M range were confirmed. - Whether Kopia content IDs themselves are HMAC-keyed (vs. plain hash + separately HMAC-derived encryption keys) could not be confirmed from fetched pages; per-content key derivation via HMAC-SHA256 is confirmed but ID-keying needs source-code confirmation.
Which compression algorithms are supported, how is the algorithm recorded per object, and what happens mixing versions
Takeaway
Restic supports only zstd (repo v2, per-blob type byte + unpacked version byte); Borg supports none/lz4/zstd/zlib/lzma with per-object ctype+clevel metadata and free mixing; Kopia supports a large s2/pgzip/gzip/deflate/zstd matrix recorded per-content in index v2 and also freely mixable going forward.
Cited Findings
- Restic repo v2 adds compression; data and tree blobs may use zstandard only; v1 has no compressed types - Source; changelog entry “Support compression for blobs (data/tree) and index/lock/snapshot files” - Source
- Restic per-run negotiation:
--compression off|fastest|auto(default)|better|max(alsoRESTIC_COMPRESSION); maps to klauspost/compress zstd levelsSpeedFastest/SpeedDefault/SpeedBetterCompression/SpeedBestCompressionwith 512 KiB window and CRC disabled - Source;autois default and uses implementation default level - Source; tuning doc confirmsautodefault for v2 repos - Source - Restic pack header per-blob recording: 1-byte type
0b00data /0b01tree /0b10compressed-data /0b11compressed-tree, followed byLength(encrypted_blob)||Hash(plaintext)(uncompressed) orLength(encrypted_blob)||Length(plaintext)||Hash(plaintext)(compressed), little-endian uint32; index addsuncompressed_lengthonly for compressed blobs - Source - Restic unpacked files (index/snapshot/lock): plaintext is
encoding_version||datawith 1-byte version;[(0x5b)/{(0x7b) mean “whole plaintext is JSON” (v1 back-compat); version2means zstd-compressed JSON; new v2 writes always version 2; implementation compresses unpacked data before encryption viacompressUnpacked, and compresses tree blobs even when--compression offfor data blobs (if Compression != off || t != DataBlob) - Source; code - Source; design rationale (null-byte/version-byte history) - Source - Restic mixing/compat: compressed and uncompressed blobs of same type may be mixed in one pack; in v2, data and tree blobs must be in separate packs; v1-repo data remains valid in v2 (no re-upload required); new data compressed per run-level,
prunerecompresses only repacked chunks; v2 repos unreadable by pre-compression restic versions; new repos default to v2 since 0.14.0 so compression is on by default - Source; release/upgrade behavior - Source - Restic dedup uses plaintext hash, so future zstd-output changes do not break dedup (compressed bytes never hashed for identity) - Source
- Borg compression set:
none(0x00),lz4(0x01),zstd(0x03, levels -128..22),zlib(0x05, levels 0–9),lzma(0x02, levels 0–9);ctypebyte +clevelbyte (zstdint8_tso -1→255, -128→128; others unsigned with 255 = n/a); speed order none>lz4>zlib>lzma, lz4>zstd; compression lzma>zlib>lz4>none, zstd>lz4 - Source; legacy 1.x zlib has no ID bytes (detected by0x.8header) - Source - Borg per-object recording (borg2): msgpacked metadata dict holds
ctype,clevel,csize(compressed+obfuscated size),psize(payload w/o obfuscation trailer, when obfuscated),olevel,size(uncompressed),type(ro_typeA/C/S/F); metadata and data slots separately encrypted with header bound as AAD - Source; codeRepoObj.format/parse- Source; compressor base auto-detection via ID header - Source - Borg negotiation/mixing: default compression is
lz4; mixing methods in one repo is fine because dedup is on source chunks, not compressed bytes; first writer of a chunk determines its stored compression;borg recreate/repo-compresscan recompress; wrappersauto,C[,L](lz4 compressibility heuristic →nonevsC) andobfuscate,SPEC,C[,L](Padmé deterministic padding, ≤12% overhead,MAX_DATA_SIZE~20 MiB cap) - Source - Kopia compression is disabled by default and controlled per-policy (global/host/path:
--compression=..., min/max file size, extensions); new setting applies going forward only, does not retroactively recompress - Source - Kopia algorithm menu:
none|deflate-best-compression|deflate-best-speed|deflate-default|gzip|gzip-best-compression|gzip-best-speed|pgzip|pgzip-best-compression|pgzip-best-speed|s2-better|s2-default|s2-parallel-4|s2-parallel-8|zstd|zstd-better-compression|zstd-fastest(zstdrecommended default choice); benchmark table shows s2 fastest (~GB/s, largest), zstd smallest, pgzip balanced - Source; details and memory/I-O guidance - Source - Kopia per-object recording: since content-level compression (v0.9,
--index-version=2), compression done after hashing; per-content compression ID kept in content manager bookkeeping, visible viakopia content list -c/content stats(e.g.(uncompressed)vszstdvszstd-fastestwith counts/sizes), no longer encoded asZ-prefixed content ID - Source; if compressed chunk grows, original stored uncompressed - Source - Kopia mixing/compat: dedup unaffected by algorithm changes or library-output drift because ID is pre-compression hash; enables future recompression during maintenance; changing policy does not rewrite old contents; newer Kopia cannot read legacy LZ4-compressed contents - must migrate with an older version (restore/repack) before upgrading - Source; LZ4 removal notice - Source
Inferences
- Recording compression outside the content ID (restic type bits/index field, Borg metadata dict, Kopia index v2) is what allows free mixing and future recompression without breaking content addressing.
- Restic's single-codec choice simplifies negotiation to levels only; Borg/Kopia pay codec-matrix complexity for finer speed/ratio control.
Gaps
- Exact zstd numeric levels behind restic
auto/better/max(klauspostSpeedDefaultetc. numeric equivalents) were not pinned to zstd CLI levels in fetched sources. - Kopia on-disk bytes for per-content compression IDs (enum values/header layout) were not found; only CLI-visible behavior is documented.
How is integrity verified (HMAC, AEAD tag, checksums, parity)? How are manifests/snapshots authenticated?
Takeaway
Restic uses Encrypt-then-MAC (AES-256-CTR + Poly1305-AES) on every blob/file plus SHA-256 content addressing; Borg 2 uses AEAD (AES-OCB or ChaCha20-Poly1305) with chunk-ID-as-AAD plus layered checksums; Kopia uses AEAD (AES-GCM or ChaCha20-Poly1305) with HMAC-derived per-content keys plus optional Reed-Solomon ECC - none provide parity by default.
Cited Findings
- Restic envelope: all files except
keys/(and pack containers) areIV(16)||CIPHERTEXT||MAC(16)(32 B overhead, random IV per file); pack files hold multiple independently encrypted/authenticated blobs + encrypted header + LEHeader_Length; primitives AES-256-CTR + Poly1305-AES, Encrypt-then-MAC (MAC over ciphertext), keys via scrypt KDF → 32 B enc key + 32 B MAC key (k+r) unlocking master keys inkeys/JSON - Source; audit summary - Source - Restic content integrity: filenames are hex SHA-256 of file ciphertext (verifiable with
sha256sum); trees/data addressed by SHA-256 of plaintext with deterministic JSON for trees;restic checkverifies structure plus optional--read-datapayload reads; tampered data fails MAC and is not decrypted - Source - Restic snapshots/manifests: snapshots are JSON (
time/tree/paths/hostname/...) stored as unpacked encrypted files (v2: version-byte + zstd thenIV||C||MAC), filename = storage ID; snapshot→tree→blob DAG; no separate manifest signature - authentication is the per-file MAC + content-hash reference; write-order rules (packs → index → snapshot; read reverse) keep repo consistent - Source - Restic has no parity/ECC; repair is
rebuild-index+ re-backup +check; threat model explicitly excludes deletion protection - Source - Borg 2 AEAD modes:
--encryption aes256-ocb|chacha20-poly1305×--id-hash sha256|blake3, orthogonal to--key-location repokey|keyfile; per-session randomsessionid,session_keyvia SHA-256 KDF, counter IVs; each object has two separately encrypted slots (metadata + data) behind unencrypted header (OBJ_MAGIC||version||chunk_id||meta_size||data_size);AAD = header||slot_tag||id||..., tag authenticates metadata, payload, header prefix and chunk ID - Source - Borg chunk-ID binding:
id = MAC(id_key, plaintext); AEAD tag binds ID to ciphertext, so repo cannot swap content under an ID; post-decryptid == MAC(decompressed)check is optional by default (only detects malicious client writes) but enforced byborg check --verify-data/BORG_ASSERT_ID- Source - Borg manifest/archive authentication (Horton principle DAG): every object referenced by parent ID up to manifest; borg2 stores
ro_typein metadata and verifies expected vs. actual type, binding meaning + ID via AAD; no TAM in borg2; borg 1.x used TAM (HKDF-SHA-512(id_key||enc_key||enc_hmac_key, 64 B salt, "borg-metadata-authentication-manifest")→HMACover packed manifest) to anchor fixed-ID manifest (CVE-2016-10099) - Source - Borg non-encrypting modes:
authenticated-sha256|blake3carry keyed MAC (not encryption) over payload+header+slot;none-*carry only unkeyed checksums (accidental-corruption detection, no tamper protection); legacy 1.x modes were AES-CTR Encrypt-then-MAC with IV reservation; key blobs wrapped by argon2-derived KEK + chacha20-poly1305 (IV 0) - Source - Borg layered checksums:
IntegrityCheckedFilestreaming checksums for cache/index/hints withintegrity.<TXN>msgpack files and[integrity]cache-config section; corrupt index/hints are deleted and rebuilt;config/config(version/id/encryption/id_hash) is plaintext/unauthenticated, protected instead by client security-dirEncryptionMethodMismatchchecks - Source - Kopia content encryption: default
AES256-GCM-HMAC-SHA256(altCHACHA20-POLY1305-HMAC-SHA256), set atrepository createand immutable after; per-content AEAD keys derived via HMAC-SHA256; blocks packed into 20 MB packs (20–40 MB on wire) with local index at pack end + top-level index mapping block ID → (blob, offset, length) - Source; cipher history (deprecated unauthenticated AES-CTR/SALSA20) - Source - Kopia format/manifest envelope:
kopia.repositoryformat blob holdsuniqueID(salt),keyAlgo scrypt-65536-8-1,encryption AES256_GCM,encryptedBlockFormat= JSON (ContentFormat{version,hash,encryption,HMACSecret 32 B,MasterKey 32 B,MaxPackSize 20 MiB}+ObjectFormat{splitter}) encrypted with passphrase-derivedKm=PBKDF(pass,uniqueID),Ke=HKDF(SHA256,Km,uniqueID,"AES",32),AD=HKDF(...,"CHECKSUM",32); snapshots/manifests are ordinary encrypted contents in the same CABS/object layers - Source; defaults (--block-hash BLAKE2B-256-128,--object-splitter DYNAMIC-4M-BUZHASH,--encryption AES256-GCM-HMAC-SHA256) - Source - Kopia extra redundancy: experimental Reed-Solomon
REED-SOLOMON-CRC32ECC with--ecc-overhead-percent(default 0 = disabled); must be enabled at creation, cannot be added later; cloud backends already ECC-protected so often redundant - Source; consistency/verify docs cover validity checks and repair - Source
Inferences
- All three authenticate snapshots/manifests the same way as data (no detached signatures); trust anchors are the client-held keys (restic master keys, Borg key+TAM/AAD chain, Kopia format-block passphrase), giving a key-anchored DAG from manifest to chunks.
- Only Kopia offers built-in parity (ECC); restic/Borg rely on backend durability + authenticated detection and rebuild/re-upload repair.
Gaps
- Kopia manifest-object specifics (manifest content type, index-epoch authentication,
kopia snapshot verifycoverage) were not fetched; snapshot authentication is inferred from “all contents encrypted/authenticated” architecture. - Exact Borg
MAX_DATA_SIZE/obfuscation interaction with AEAD tag andBORG_ASSERT_IDdefault scope need code-level confirmation beyond docs snippets.
Order of operations: compress-then-encrypt vs alternatives, and why
Takeaway
All three do hash → compress → encrypt (dedup first, compress second, encrypt last); encrypting last is mandatory because ciphertext is incompressible and randomized, while hashing/compressing first preserves dedup and ratio.
Cited Findings
- Restic code order in
saveAndEncrypt: plaintext hash for ID/dedup (saveBlob→Hash(buf)) → zstd compress (if v2 and enabled) → random nonce →Seal(AES-CTR + Poly1305 MAC) → pack; unpacked files:compressUnpacked(prepend version byte + zstd) →Seal- Source; PR states “Unpacked files like lock, index and snapshot files are also compressed before encryption” - Source - Restic why: backing up pre-compressed (e.g.
.gz) data defeats CDC dedup because small input changes avalanche through the compressor and shift cut points; project advises backing up uncompressed data and letting restic chunk-then-compress per blob; chunk-then-compress discussion - Source; zstd-dictionary debate concludes no dictionary needed precisely because compression runs after chunking/dedup - Source - Borg order:
id = MAC(id_key, data)→compressed = compress(data)→AEAD_encrypt(session_key, iv, compressed, aad=id||header)(borg2); legacy:id=AUTH(data)→compress→AES-CTR→MAC(encrypted)- Source; docs state “Compression is applied after deduplication, thus using different compression methods in one repo does not influence deduplication” - Source - Kopia order: “splits into chunks → hash → compare → if new, compress chunk → encrypt → pack” - Source; content-level compression PR explicitly rewired pipeline from
read → split → compress → hash → storetoread → split → hash → compress → storeto stop compressor-version drift from changing IDs and breaking dedup and to enable server-side/recompression - Source - General rationale shared across docs: encryption must be last because (a) compressed ciphertext does not compress, (b) randomized encryption (IV/session key) destroys dedup if hashed after, (c) MAC/AEAD must cover the stored (compressed) bytes; compressing whole files before chunking is measurably worse than chunk-then-compress (Kopia split-vs-whole table: s2 28.6% vs 25.6%, gzip 18.8% vs 17.3% - small loss accepted for dedup wins) - Source
Inferences
- The universal pattern is dedup-on-plaintext → compress → authenticated-encrypt; any deviation (pre-compressed inputs, compress-after-encrypt, hash-after-compress) is documented as an anti-pattern in all three projects’ issues/docs.
- Per-chunk compression inherently sacrifices cross-chunk dictionary context; all three accept slightly worse ratios in exchange for shift-resilient dedup and independent integrity domains.
Gaps
- No source quantified restic’s incompressible-blob shortcut (whether zstd
EncodeAlloutput larger than input is stored raw vs. kept); forum claims “stores raw data” but code path storesuncompressedLengthflag - exact skip-threshold logic needs code confirmation.