105 lines
24 KiB
Markdown
105 lines
24 KiB
Markdown
# Compression, Deduplication, and Integrity in Restic, BorgBackup, and Kopia
|
||
|
||
## Fixed-file vs content-defined chunking: algorithm, average chunk sizes, how dedup works with encryption
|
||
|
||
### Takeaway
|
||
All three use content-defined chunking (CDC) by default with a fixed-size fallback option; deduplication is on plaintext hashes before compression/encryption (so no convergent encryption), with Borg and Kopia keying chunk IDs while restic uses plain SHA-256.
|
||
|
||
### Cited Findings
|
||
- Restic splits each file independently with Rabin-fingerprint CDC over a 64-byte sliding window; a random irreducible polynomial is generated at `init` and stored as `chunker_polynomial` in `config` to harden against watermarking - [Source](https://restic.readthedocs.io/en/stable/100_references.html); background and worked example - [Source](https://restic.net/blog/2015-09-12/restic-foundation1-cdc/)
|
||
- Restic chunking parameters: files <512 KiB are not split; blobs are 512 KiB–8 MiB with ~1 MiB average target; modified files only re-store changed blobs, robust to insertions at arbitrary offsets - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||
- Restic 0.18.0 mitigates chunk-size fingerprinting (Alexeev/Percival/Zhang 2025) by randomly assigning chunks to pack files so attackers observing the repo cannot map chunk sizes to files - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||
- Restic deduplication happens before encryption: blob ID is SHA-256 of plaintext; index maps plaintext hash to pack location; pack/index filenames are SHA-256 of ciphertext for accident detection only - [Source](https://forum.restic.net/t/how-are-blobs-deduplicated-with-encryption/6478); all content referenced by SHA-256 of plaintext, one blob holds data from only one file, multiple blobs packed per pack file - [Source](https://github.com/restic/restic/issues/2401)
|
||
- Restic uses random per-repository master keys and random 16-byte IV per encryption (not convergent/deterministic encryption); identical plaintexts deduplicate via index lookup, not via identical ciphertexts - [Source](https://restic.readthedocs.io/en/stable/design.html); external audit notes AES-256-CTR + Poly1305-AES with separate keys - [Source](https://words.filippo.io/restic-cryptography/)
|
||
- Borg default chunker is `fastcdc` (window-less keyed Gear hash); alternatives: `fixed` (fixed blocksize, optional different header block), `buzhash`/`buzhash64` (rolling Buzhash), `rabin-aes`/`toeplitz-aes`/`goldilocks-aes` (UHF-then-PRF: rolling universal hash + AES-128, cut decision only on AES output) - [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html); comparison table and guidance - [Source](https://borgbackup.readthedocs.io/en/master/internals/chunker.html)
|
||
- Borg classic `buzhash` defaults: `CHUNK_MIN_EXP=19` (512 KiB), `CHUNK_MAX_EXP=23` (8 MiB), `HASH_MASK_BITS=21` (~2 MiB target), `HASH_WINDOW_SIZE=4095` bytes; tunable via `--chunker-params`; min/max clamp where rolling-hash cuts may occur - [Source](https://github.com/borgbackup/borg/blob/master/docs/misc/create_chunker-params.txt); `fastcdc` uses `CHUNK_MIN_EXP,CHUNK_MAX_EXP,HASH_MASK_BITS,NC_LEVEL` with normalized chunking - [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html)
|
||
- Borg chunker secrets are derived from repository key material (buzhash table XOR seed stored encrypted in keyfile; AES chunkers derive table/polynomial/AES key from `id_key` per-chunker domain) to resist fingerprinting; 2025 research (Truong et al. CCS 2025 eprint 2025/558; eprint 2025/532) showed keyed rolling-hash-only chunkers allow key recovery, motivating AES-based chunkers - [Source](https://borgbackup.readthedocs.io/en/master/internals/chunker.html)
|
||
- Borg deduplication is global across all archives/hosts/files on chunk `id_hash`, which is a keyed MAC over plaintext (`HMAC-SHA256` for `--id-hash sha256`, keyed `BLAKE3` for `--id-hash blake3`) using secret `id_key`, not plain hash, so attackers cannot confirm small-file presence without the key - [Source](https://borgbackup.readthedocs.io/en/latest/internals/security.html); file metadata (msgpacked items) is chunked with finer params and deduplicated the same way - [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html)
|
||
- Borg 1.x-compatible dedup requires `buzhash`; `fastcdc`/`buzhash64` give same dedup but different cut points; `fixed` gives positional (not content-shift-resilient) dedup, suited to disk images - [Source](https://borgbackup.readthedocs.io/en/master/internals/chunker.html)
|
||
- Kopia calls chunkers “splitters”: `BUZHASH`, `RABINKARP` (rolling-hash CDC) and `fixed`, selectable size 1M–8M - [Source](https://kopia.discourse.group/t/does-kopia-use-content-defined-chunking-cdc/1417); default object splitter is `DYNAMIC-4M-BUZHASH` (also the default for every `repository create` backend) - [Source](https://kopia.io/docs/reference/command-line/common/repository-create-filesystem/)
|
||
- Kopia splitter history: original `DYNAMIC` splitter (silvasur/buzhash dependency) was deprecated for license reasons and replaced by a faster but incompatible buzhash implementation; old repos remain readable but large objects re-upload instead of deduping across the change - [Source](https://github.com/kopia/kopia/commit/03339c18afedb810f31320ef01c707e7acbdc374)
|
||
- Kopia pipeline is split → hash → compare against index → discard if known; otherwise compress → encrypt → pack multiple small blocks into ~20–40 MB packs with random names; content hash default `BLAKE2B-256-128` - [Source](https://kopia.io/docs/advanced/compression/); pack/index architecture - [Source](https://kopia.io/docs/advanced/architecture/)
|
||
- Kopia deduplication is unaffected by compression settings because hashing precedes compression since index v2 (see next section); same file under compressed and uncompressed policies still dedupes - [Source](https://kopia.discourse.group/t/deduplication-and-compression/714)
|
||
|
||
### Inferences
|
||
- None of the three use convergent encryption (deterministic, plaintext-derived keys); all use random repository keys + random nonces/session keys, with dedup achieved by a client-side plaintext-hash index lookup before encryption.
|
||
- Keyed chunk IDs (Borg MAC, Kopia HMAC-derived keys/format secret) hide confirmation-of-file attacks from repo-only attackers; restic's plain-SHA-256 blob IDs do not, which is part of why restic added pack-mixing and per-repo random polynomials.
|
||
|
||
### Gaps
|
||
- Exact numeric defaults for Kopia `DYNAMIC-4M-BUZHASH` (min/max/expected size, window bytes) were not found in fetched docs; only the 4M-average naming and 1M–8M range were confirmed.
|
||
- Whether Kopia content IDs themselves are HMAC-keyed (vs. plain hash + separately HMAC-derived encryption keys) could not be confirmed from fetched pages; per-content key derivation via HMAC-SHA256 is confirmed but ID-keying needs source-code confirmation.
|
||
|
||
## Which compression algorithms are supported, how is the algorithm recorded per object, and what happens mixing versions
|
||
|
||
### Takeaway
|
||
Restic supports only zstd (repo v2, per-blob type byte + unpacked version byte); Borg supports none/lz4/zstd/zlib/lzma with per-object ctype+clevel metadata and free mixing; Kopia supports a large s2/pgzip/gzip/deflate/zstd matrix recorded per-content in index v2 and also freely mixable going forward.
|
||
|
||
### Cited Findings
|
||
- Restic repo v2 adds compression; data and tree blobs may use zstandard only; v1 has no compressed types - [Source](https://restic.readthedocs.io/en/stable/100_references.html); changelog entry “Support compression for blobs (data/tree) and index/lock/snapshot files” - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||
- Restic per-run negotiation: `--compression off|fastest|auto(default)|better|max` (also `RESTIC_COMPRESSION`); maps to klauspost/compress zstd levels `SpeedFastest`/`SpeedDefault`/`SpeedBetterCompression`/`SpeedBestCompression` with 512 KiB window and CRC disabled - [Source](https://github.com/restic/restic/blob/master/internal/repository/repository.go); `auto` is default and uses implementation default level - [Source](https://forum.restic.net/t/what-compression-is-used/6850); tuning doc confirms `auto` default for v2 repos - [Source](https://restic.readthedocs.io/en/stable/047_tuning_parameters.html)
|
||
- Restic pack header per-blob recording: 1-byte type `0b00` data / `0b01` tree / `0b10` compressed-data / `0b11` compressed-tree, followed by `Length(encrypted_blob)||Hash(plaintext)` (uncompressed) or `Length(encrypted_blob)||Length(plaintext)||Hash(plaintext)` (compressed), little-endian uint32; index adds `uncompressed_length` only for compressed blobs - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||
- Restic unpacked files (index/snapshot/lock): plaintext is `encoding_version||data` with 1-byte version; `[` (0x5b)/`{` (0x7b) mean “whole plaintext is JSON” (v1 back-compat); version `2` means zstd-compressed JSON; new v2 writes always version 2; implementation compresses unpacked data before encryption via `compressUnpacked`, and compresses tree blobs even when `--compression off` for data blobs (`if Compression != off || t != DataBlob`) - [Source](https://restic.readthedocs.io/en/stable/100_references.html); code - [Source](https://github.com/restic/restic/blob/master/internal/repository/repository.go); design rationale (null-byte/version-byte history) - [Source](https://github.com/restic/restic/pull/3666)
|
||
- Restic mixing/compat: compressed and uncompressed blobs of same type may be mixed in one pack; in v2, data and tree blobs must be in separate packs; v1-repo data remains valid in v2 (no re-upload required); new data compressed per run-level, `prune` recompresses only repacked chunks; v2 repos unreadable by pre-compression restic versions; new repos default to v2 since 0.14.0 so compression is on by default - [Source](https://restic.readthedocs.io/en/stable/100_references.html); release/upgrade behavior - [Source](https://forum.restic.net/t/restic-0-14-0-released/5359)
|
||
- Restic dedup uses plaintext hash, so future zstd-output changes do not break dedup (compressed bytes never hashed for identity) - [Source](https://github.com/restic/restic/pull/3666)
|
||
- Borg compression set: `none` (0x00), `lz4` (0x01), `zstd` (0x03, levels -128..22), `zlib` (0x05, levels 0–9), `lzma` (0x02, levels 0–9); `ctype` byte + `clevel` byte (zstd `int8_t` so -1→255, -128→128; others unsigned with 255 = n/a); speed order none>lz4>zlib>lzma, lz4>zstd; compression lzma>zlib>lz4>none, zstd>lz4 - [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html); legacy 1.x zlib has no ID bytes (detected by `0x.8` header) - [Source](https://borgbackup.readthedocs.io/en/stable/internals/data-structures.html)
|
||
- Borg per-object recording (borg2): msgpacked metadata dict holds `ctype`, `clevel`, `csize` (compressed+obfuscated size), `psize` (payload w/o obfuscation trailer, when obfuscated), `olevel`, `size` (uncompressed), `type` (ro_type `A`/`C`/`S`/`F`); metadata and data slots separately encrypted with header bound as AAD - [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html); code `RepoObj.format/parse` - [Source](https://github.com/borgbackup/borg/blob/86fd77fd/src/borg/repoobj.py); compressor base auto-detection via ID header - [Source](https://github.com/borgbackup/borg/blob/86fd77fd/src/borg/compress.pyx)
|
||
- Borg negotiation/mixing: default compression is `lz4`; mixing methods in one repo is fine because dedup is on source chunks, not compressed bytes; first writer of a chunk determines its stored compression; `borg recreate`/`repo-compress` can recompress; wrappers `auto,C[,L]` (lz4 compressibility heuristic → `none` vs `C`) and `obfuscate,SPEC,C[,L]` (Padmé deterministic padding, ≤12% overhead, `MAX_DATA_SIZE` ~20 MiB cap) - [Source](https://manpages.debian.org/trixie/borgbackup2/borg2-compression.1.en.html)
|
||
- Kopia compression is disabled by default and controlled per-policy (global/host/path: `--compression=...`, min/max file size, extensions); new setting applies going forward only, does not retroactively recompress - [Source](https://kopia.io/docs/faqs/)
|
||
- Kopia algorithm menu: `none|deflate-best-compression|deflate-best-speed|deflate-default|gzip|gzip-best-compression|gzip-best-speed|pgzip|pgzip-best-compression|pgzip-best-speed|s2-better|s2-default|s2-parallel-4|s2-parallel-8|zstd|zstd-better-compression|zstd-fastest` (`zstd` recommended default choice); benchmark table shows s2 fastest (~GB/s, largest), zstd smallest, pgzip balanced - [Source](https://kopia.io/docs/faqs/); details and memory/I-O guidance - [Source](https://kopia.io/docs/advanced/compression/)
|
||
- Kopia per-object recording: since content-level compression (v0.9, `--index-version=2`), compression done after hashing; per-content compression ID kept in content manager bookkeeping, visible via `kopia content list -c` / `content stats` (e.g. `(uncompressed)` vs `zstd` vs `zstd-fastest` with counts/sizes), no longer encoded as `Z`-prefixed content ID - [Source](https://github.com/kopia/kopia/pull/1076); if compressed chunk grows, original stored uncompressed - [Source](https://kopia.io/docs/advanced/compression/)
|
||
- Kopia mixing/compat: dedup unaffected by algorithm changes or library-output drift because ID is pre-compression hash; enables future recompression during maintenance; changing policy does not rewrite old contents; newer Kopia cannot read legacy LZ4-compressed contents - must migrate with an older version (restore/repack) before upgrading - [Source](https://github.com/kopia/kopia/pull/1076); LZ4 removal notice - [Source](https://kopia.io/docs/advanced/compression/)
|
||
|
||
### Inferences
|
||
- Recording compression outside the content ID (restic type bits/index field, Borg metadata dict, Kopia index v2) is what allows free mixing and future recompression without breaking content addressing.
|
||
- Restic's single-codec choice simplifies negotiation to levels only; Borg/Kopia pay codec-matrix complexity for finer speed/ratio control.
|
||
|
||
### Gaps
|
||
- Exact zstd numeric levels behind restic `auto/better/max` (klauspost `SpeedDefault` etc. numeric equivalents) were not pinned to zstd CLI levels in fetched sources.
|
||
- Kopia on-disk bytes for per-content compression IDs (enum values/header layout) were not found; only CLI-visible behavior is documented.
|
||
|
||
## How is integrity verified (HMAC, AEAD tag, checksums, parity)? How are manifests/snapshots authenticated?
|
||
|
||
### Takeaway
|
||
Restic uses Encrypt-then-MAC (AES-256-CTR + Poly1305-AES) on every blob/file plus SHA-256 content addressing; Borg 2 uses AEAD (AES-OCB or ChaCha20-Poly1305) with chunk-ID-as-AAD plus layered checksums; Kopia uses AEAD (AES-GCM or ChaCha20-Poly1305) with HMAC-derived per-content keys plus optional Reed-Solomon ECC - none provide parity by default.
|
||
|
||
### Cited Findings
|
||
- Restic envelope: all files except `keys/` (and pack containers) are `IV(16)||CIPHERTEXT||MAC(16)` (32 B overhead, random IV per file); pack files hold multiple independently encrypted/authenticated blobs + encrypted header + LE `Header_Length`; primitives AES-256-CTR + Poly1305-AES, Encrypt-then-MAC (MAC over ciphertext), keys via scrypt KDF → 32 B enc key + 32 B MAC key (`k`+`r`) unlocking master keys in `keys/` JSON - [Source](https://restic.readthedocs.io/en/stable/100_references.html); audit summary - [Source](https://words.filippo.io/restic-cryptography/)
|
||
- Restic content integrity: filenames are hex SHA-256 of file ciphertext (verifiable with `sha256sum`); trees/data addressed by SHA-256 of plaintext with deterministic JSON for trees; `restic check` verifies structure plus optional `--read-data` payload reads; tampered data fails MAC and is not decrypted - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||
- Restic snapshots/manifests: snapshots are JSON (`time/tree/paths/hostname/...`) stored as unpacked encrypted files (v2: version-byte + zstd then `IV||C||MAC`), filename = storage ID; snapshot→tree→blob DAG; no separate manifest signature - authentication is the per-file MAC + content-hash reference; write-order rules (packs → index → snapshot; read reverse) keep repo consistent - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||
- Restic has no parity/ECC; repair is `rebuild-index` + re-backup + `check`; threat model explicitly excludes deletion protection - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||
- Borg 2 AEAD modes: `--encryption aes256-ocb|chacha20-poly1305` × `--id-hash sha256|blake3`, orthogonal to `--key-location repokey|keyfile`; per-session random `sessionid`, `session_key` via SHA-256 KDF, counter IVs; each object has two separately encrypted slots (metadata + data) behind unencrypted header (`OBJ_MAGIC||version||chunk_id||meta_size||data_size`); `AAD = header||slot_tag||id||...`, tag authenticates metadata, payload, header prefix and chunk ID - [Source](https://borgbackup.readthedocs.io/en/latest/internals/security.html)
|
||
- Borg chunk-ID binding: `id = MAC(id_key, plaintext)`; AEAD tag binds ID to ciphertext, so repo cannot swap content under an ID; post-decrypt `id == MAC(decompressed)` check is optional by default (only detects malicious client writes) but enforced by `borg check --verify-data` / `BORG_ASSERT_ID` - [Source](https://borgbackup.readthedocs.io/en/latest/internals/security.html)
|
||
- Borg manifest/archive authentication (Horton principle DAG): every object referenced by parent ID up to manifest; borg2 stores `ro_type` in metadata and verifies expected vs. actual type, binding meaning + ID via AAD; no TAM in borg2; borg 1.x used TAM (`HKDF-SHA-512(id_key||enc_key||enc_hmac_key, 64 B salt, "borg-metadata-authentication-manifest")` → `HMAC` over packed manifest) to anchor fixed-ID manifest (CVE-2016-10099) - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||
- Borg non-encrypting modes: `authenticated-sha256|blake3` carry keyed MAC (not encryption) over payload+header+slot; `none-*` carry only unkeyed checksums (accidental-corruption detection, no tamper protection); legacy 1.x modes were AES-CTR Encrypt-then-MAC with IV reservation; key blobs wrapped by argon2-derived KEK + chacha20-poly1305 (IV 0) - [Source](https://borgbackup.readthedocs.io/en/latest/internals/security.html)
|
||
- Borg layered checksums: `IntegrityCheckedFile` streaming checksums for cache/index/hints with `integrity.<TXN>` msgpack files and `[integrity]` cache-config section; corrupt index/hints are deleted and rebuilt; `config/config` (version/id/encryption/id_hash) is plaintext/unauthenticated, protected instead by client security-dir `EncryptionMethodMismatch` checks - [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html)
|
||
- Kopia content encryption: default `AES256-GCM-HMAC-SHA256` (alt `CHACHA20-POLY1305-HMAC-SHA256`), set at `repository create` and immutable after; per-content AEAD keys derived via HMAC-SHA256; blocks packed into 20 MB packs (20–40 MB on wire) with local index at pack end + top-level index mapping block ID → (blob, offset, length) - [Source](https://kopia.io/docs/advanced/architecture/); cipher history (deprecated unauthenticated AES-CTR/SALSA20) - [Source](https://github.com/kopia/kopia/pull/277)
|
||
- Kopia format/manifest envelope: `kopia.repository` format blob holds `uniqueID` (salt), `keyAlgo scrypt-65536-8-1`, `encryption AES256_GCM`, `encryptedBlockFormat` = JSON (`ContentFormat{version,hash,encryption,HMACSecret 32 B,MasterKey 32 B,MaxPackSize 20 MiB}` + `ObjectFormat{splitter}`) encrypted with passphrase-derived `Km=PBKDF(pass,uniqueID)`, `Ke=HKDF(SHA256,Km,uniqueID,"AES",32)`, `AD=HKDF(...,"CHECKSUM",32)`; snapshots/manifests are ordinary encrypted contents in the same CABS/object layers - [Source](https://kopia.io/docs/advanced/encryption/); defaults (`--block-hash BLAKE2B-256-128`, `--object-splitter DYNAMIC-4M-BUZHASH`, `--encryption AES256-GCM-HMAC-SHA256`) - [Source](https://kopia.io/docs/reference/command-line/common/repository-create-filesystem/)
|
||
- Kopia extra redundancy: experimental Reed-Solomon `REED-SOLOMON-CRC32` ECC with `--ecc-overhead-percent` (default 0 = disabled); must be enabled at creation, cannot be added later; cloud backends already ECC-protected so often redundant - [Source](https://kopia.io/docs/advanced/ecc/); consistency/verify docs cover validity checks and repair - [Source](https://kopia.io/docs/advanced/architecture/)
|
||
|
||
### Inferences
|
||
- All three authenticate snapshots/manifests the same way as data (no detached signatures); trust anchors are the client-held keys (restic master keys, Borg key+TAM/AAD chain, Kopia format-block passphrase), giving a key-anchored DAG from manifest to chunks.
|
||
- Only Kopia offers built-in parity (ECC); restic/Borg rely on backend durability + authenticated detection and rebuild/re-upload repair.
|
||
|
||
### Gaps
|
||
- Kopia manifest-object specifics (manifest content type, index-epoch authentication, `kopia snapshot verify` coverage) were not fetched; snapshot authentication is inferred from “all contents encrypted/authenticated” architecture.
|
||
- Exact Borg `MAX_DATA_SIZE`/obfuscation interaction with AEAD tag and `BORG_ASSERT_ID` default scope need code-level confirmation beyond docs snippets.
|
||
|
||
## Order of operations: compress-then-encrypt vs alternatives, and why
|
||
|
||
### Takeaway
|
||
All three do hash → compress → encrypt (dedup first, compress second, encrypt last); encrypting last is mandatory because ciphertext is incompressible and randomized, while hashing/compressing first preserves dedup and ratio.
|
||
|
||
### Cited Findings
|
||
- Restic code order in `saveAndEncrypt`: plaintext hash for ID/dedup (`saveBlob` → `Hash(buf)`) → zstd compress (if v2 and enabled) → random nonce → `Seal` (AES-CTR + Poly1305 MAC) → pack; unpacked files: `compressUnpacked` (prepend version byte + zstd) → `Seal` - [Source](https://github.com/restic/restic/blob/master/internal/repository/repository.go); PR states “Unpacked files like lock, index and snapshot files are also compressed before encryption” - [Source](https://github.com/restic/restic/pull/3666)
|
||
- Restic why: backing up pre-compressed (e.g. `.gz`) data defeats CDC dedup because small input changes avalanche through the compressor and shift cut points; project advises backing up uncompressed data and letting restic chunk-then-compress per blob; chunk-then-compress discussion - [Source](https://github.com/restic/restic/issues/790); zstd-dictionary debate concludes no dictionary needed precisely because compression runs after chunking/dedup - [Source](https://github.com/restic/restic/issues/3775)
|
||
- Borg order: `id = MAC(id_key, data)` → `compressed = compress(data)` → `AEAD_encrypt(session_key, iv, compressed, aad=id||header)` (borg2); legacy: `id=AUTH(data)` → `compress` → `AES-CTR` → `MAC(encrypted)` - [Source](https://borgbackup.readthedocs.io/en/latest/internals/security.html); docs state “Compression is applied after deduplication, thus using different compression methods in one repo does not influence deduplication” - [Source](https://manpages.debian.org/trixie/borgbackup2/borg2-compression.1.en.html)
|
||
- Kopia order: “splits into chunks → hash → compare → if new, compress chunk → encrypt → pack” - [Source](https://kopia.io/docs/advanced/compression/); content-level compression PR explicitly rewired pipeline from `read → split → compress → hash → store` to `read → split → hash → compress → store` to stop compressor-version drift from changing IDs and breaking dedup and to enable server-side/recompression - [Source](https://github.com/kopia/kopia/pull/1076)
|
||
- General rationale shared across docs: encryption must be last because (a) compressed ciphertext does not compress, (b) randomized encryption (IV/session key) destroys dedup if hashed after, (c) MAC/AEAD must cover the stored (compressed) bytes; compressing whole files before chunking is measurably worse than chunk-then-compress (Kopia split-vs-whole table: s2 28.6% vs 25.6%, gzip 18.8% vs 17.3% - small loss accepted for dedup wins) - [Source](https://kopia.io/docs/advanced/compression/)
|
||
|
||
### Inferences
|
||
- The universal pattern is dedup-on-plaintext → compress → authenticated-encrypt; any deviation (pre-compressed inputs, compress-after-encrypt, hash-after-compress) is documented as an anti-pattern in all three projects’ issues/docs.
|
||
- Per-chunk compression inherently sacrifices cross-chunk dictionary context; all three accept slightly worse ratios in exchange for shift-resilient dedup and independent integrity domains.
|
||
|
||
### Gaps
|
||
- No source quantified restic’s incompressible-blob shortcut (whether zstd `EncodeAll` output larger than input is stored raw vs. kept); forum claims “stores raw data” but code path stores `uncompressedLength` flag - exact skip-threshold logic needs code confirmation.
|