Forbid em dashes everywhere; add AGENTS.md with repo rules
This commit is contained in:
@@ -0,0 +1,9 @@
|
||||
# Agent rules for this repo
|
||||
|
||||
- Dev: `python -m venv .venv && .venv/bin/pip install -e '.[dev]' && .venv/bin/pytest`.
|
||||
- Style: exact, minimal diffs. Prefer editing existing files over creating new ones.
|
||||
- NEVER use em dashes (U+2014) anywhere: not in code, docs, comments, commit
|
||||
messages, or chat output. Use a hyphen, comma, or colon instead.
|
||||
- Verify through execution: run the suite after behavior changes.
|
||||
- Destructive endpoints stay dry-run by default, date-guarded where remote
|
||||
data is at stake, and scoped to the claimed directory or allowed roots.
|
||||
@@ -1,14 +1,14 @@
|
||||
# versiond — the undo button for your entire machine
|
||||
# versiond - the undo button for your entire machine
|
||||
|
||||
`versiond` watches the directories you choose and versions every meaningful file state:
|
||||
source code, configs, images, archives, documents, and live databases. Then it
|
||||
syncs it all to your own WebDAV box. When an AI agent (or you, at 3am) destroys
|
||||
something, you go back — one file, one project, or the whole system.
|
||||
something, you go back - one file, one project, or the whole system.
|
||||
|
||||
No encryption keys to lose. No subscription. No cloud. Your disk, your server,
|
||||
your history.
|
||||
|
||||
> Status: running in production on its author's dev box right now —
|
||||
> Status: running in production on its author's dev box right now -
|
||||
> 7,596 files, 8,884 versions, zero errors. This README is written from
|
||||
> measured numbers on that machine, not from wishes.
|
||||
|
||||
@@ -43,7 +43,7 @@ box this was built on (a 436GB Hetzner box, October 2026).
|
||||
|
||||
### 1.1 The surge test
|
||||
|
||||
16 threads hammering a watched directory for 30 seconds straight — 6 writers,
|
||||
16 threads hammering a watched directory for 30 seconds straight - 6 writers,
|
||||
6 readers, renames, deletes, plus a live SQLite writer:
|
||||
|
||||
| What went in | What versiond did with it |
|
||||
@@ -72,7 +72,7 @@ writes a day?
|
||||
| Bytes/day | 720 × 7GB ≈ **5 TB/day** |
|
||||
| Capture load | 720 × 63s ≈ **12.6 h/day of heavy I/O** |
|
||||
|
||||
Read that twice. The snapshots are *correct* — but at 7GB, full-copy
|
||||
Read that twice. The snapshots are *correct* - but at 7GB, full-copy
|
||||
versioning is physically absurd. That is why databases get their own schedule:
|
||||
|
||||
### 1.3 The database schedule (opinionated, on purpose)
|
||||
@@ -103,7 +103,7 @@ interval(size) = clamp(2h, 4h × (size / 7GB), 24h)
|
||||
10:00 ░░░ small ones 22:00 ░░░ small ones
|
||||
```
|
||||
|
||||
A file that changes constantly still banks one version per slot — recency
|
||||
A file that changes constantly still banks one version per slot - recency
|
||||
granularity *is* the grid, and the grid is the contract.
|
||||
|
||||
### 1.4 Big data: the honest math
|
||||
@@ -121,7 +121,7 @@ measured chunk-by-chunk:
|
||||
Formula: full-copy burns `versions/day × file_size`; chunked burns
|
||||
`churn + growth` once. For a 5KB config the difference is trivia; for a 7GB
|
||||
database it's the difference between possible and impossible. Both columns
|
||||
keep identical versions and identical restores — only the bytes differ.
|
||||
keep identical versions and identical restores - only the bytes differ.
|
||||
|
||||
### 1.5 The 286GB reality check
|
||||
|
||||
@@ -129,8 +129,8 @@ That same box once showed 51GB free on a 436GB disk. The hunt took minutes:
|
||||
|
||||
| Where | Size | Verdict |
|
||||
|---|---|---|
|
||||
| `devplacepy/data/backups/` — 42 hourly 6GB tarballs, nothing ever prunes them | 253 GB | devplace's own backup output, excluded from versiond |
|
||||
| `devplacepy/restore_staging/` — leftover restore duplicate | 39.5 GB | delete after confirming |
|
||||
| `devplacepy/data/backups/` - 42 hourly 6GB tarballs, nothing ever prunes them | 253 GB | devplace's own backup output, excluded from versiond |
|
||||
| `devplacepy/restore_staging/` - leftover restore duplicate | 39.5 GB | delete after confirming |
|
||||
| `devplacepy/data/uploads/` | 32.7 GB | legit user data, versioned |
|
||||
| `devplacepy/data/devplace.db` | 7 GB | legit, hot, scheduled |
|
||||
| everything else | ~30 GB | normal |
|
||||
@@ -142,13 +142,13 @@ is arson. `data/backups/` and `backup_staging/` ship excluded.
|
||||
|
||||
## 2. What gets versioned (and what never does)
|
||||
|
||||
Kept — source, configs, allowlisted dotfiles, **images, audio/video, archives,
|
||||
Kept - source, configs, allowlisted dotfiles, **images, audio/video, archives,
|
||||
databases (journals folded into verified snapshots), PDFs, fonts**. Cap:
|
||||
`limits.max_file_bytes` (default **10 MiB**, configurable). No binary
|
||||
sniffing: any bytes up to the cap are versioned; non-UTF-8 comes back as
|
||||
`application/octet-stream`.
|
||||
|
||||
Never — dependency/build/cache trees (`node_modules`, `venv`, `target`, …),
|
||||
Never - dependency/build/cache trees (`node_modules`, `venv`, `target`, …),
|
||||
hidden directories, `*.pyc/*.o/*.so`, `*.min.js/*.map`, temps (`*.tmp`,
|
||||
`~$*` Office locks, `#*#` Emacs autosaves, `*~`, `4913`), SQLite journals as
|
||||
standalone files, anything over the cap. Proven live: 11 temp-name variants
|
||||
@@ -156,7 +156,7 @@ written to a watched dir produced **zero** versions and eleven 422s.
|
||||
|
||||
Databases get the zero-error policy: backup-API snapshot taken read-only (no
|
||||
write lock, no WAL recovery, live writers unaffected), `PRAGMA integrity_check`
|
||||
must say `ok`, otherwise nothing is stored — `database-corrupt`,
|
||||
must say `ok`, otherwise nothing is stored - `database-corrupt`,
|
||||
`database-locked`, `database-unreadable` surface as loud skips. Every SQLite
|
||||
restore is integrity-checked before it touches disk; failures refuse with
|
||||
`unrestorable` instead of writing garbage.
|
||||
@@ -172,12 +172,12 @@ restore is integrity-checked before it touches disk; failures refuse with
|
||||
watched `node_modules` first, filtered later. Wrong order, rejected.
|
||||
- **Coalesce first+last, rate-limit per path.** Bursts collapse; the 30/h cap
|
||||
bounds hot files; the newest state is never dropped, only delayed. A hot
|
||||
file under max load settles at ~1 version per 2s — the price of never losing
|
||||
file under max load settles at ~1 version per 2s - the price of never losing
|
||||
the present.
|
||||
- **Content-addressed spool, remote as dumb storage.** Local spool decouples
|
||||
ingest (milliseconds) from network (whenever). Uploads are idempotent PUTs
|
||||
with backoff; offline just grows the spool. WebDAV was chosen because a
|
||||
Storage Box is €3/month, not because it's good — the protocol is chatty, so
|
||||
Storage Box is €3/month, not because it's good - the protocol is chatty, so
|
||||
manifests batch thousands of records per object.
|
||||
- **No client-side encryption.** Deliberate: the trust model is a
|
||||
disk-encrypted box on a trusted network. A backup key is a second way to
|
||||
@@ -189,11 +189,11 @@ restore is integrity-checked before it touches disk; failures refuse with
|
||||
- **Time policy first, capacity second.** Retention keeps 5 days complete,
|
||||
thins older to 1/file/day, always keeps latest + pins. The 70% remote-usage
|
||||
backstop (read via RFC 4331 quota props) triggers early passes but *never*
|
||||
breaks the 5-day floor — pressure pages a human instead of auto-deleting
|
||||
breaks the 5-day floor - pressure pages a human instead of auto-deleting
|
||||
history. No vendor does percentage watermarks; neither do we, we do better:
|
||||
a floor with an alarm.
|
||||
- **Deletes are real deletes.** Every retention/`forget` execute ends with
|
||||
`VACUUM` + WAL checkpoint — freelist back to 0, bytes reported. GC removes
|
||||
`VACUUM` + WAL checkpoint - freelist back to 0, bytes reported. GC removes
|
||||
orphan blobs locally *and* remotely. Only unreferenced data is ever touched.
|
||||
|
||||
---
|
||||
@@ -201,13 +201,13 @@ restore is integrity-checked before it touches disk; failures refuse with
|
||||
## 4. When shit goes south (the runbook)
|
||||
|
||||
1. Reinstall, `POST /config/remote/adopt` the old directory on the new machine.
|
||||
2. `POST /admin/reindex` — full index back from manifests (async, poll it).
|
||||
2. `POST /admin/reindex` - full index back from manifests (async, poll it).
|
||||
3. Restore: single file, bulk plan (always dry-run first), or point-in-time.
|
||||
Blobs stream down on demand. DB restores verify before writing.
|
||||
4. Retention + scheduler resume; 5 days back, guaranteed.
|
||||
|
||||
Destructive tools are date-guarded (`purge-remote` needs today's date),
|
||||
dry-run by default (retention, GC, `forget`), audited, and scoped — purge only
|
||||
dry-run by default (retention, GC, `forget`), audited, and scoped - purge only
|
||||
ever touches the claimed remote directory; restores only write inside
|
||||
`$HOME`/roots and refuse symlinks.
|
||||
|
||||
@@ -223,7 +223,7 @@ Full live schema at `/openapi.json`. The honest map (❌ = consciously absent):
|
||||
| WebDAV remote, unique auto-claimed directory, adopt, manifests | ✅ |
|
||||
| Retention (5d + thin), GC, remote purge, scheduler + 70% backstop | ✅ |
|
||||
| Reindex, on-demand blob fetch, metrics, dashboard, progress stream | ✅ |
|
||||
| `GET /config/*` tuning, full-text search, per-version delete, agent skill | ❌ (deliberately deferred, see table in §5 of the long spec — retained below) |
|
||||
| `GET /config/*` tuning, full-text search, per-version delete, agent skill | ❌ (deliberately deferred, see table in §5 of the long spec - retained below) |
|
||||
|
||||
Config lives in `~/.config/versiond/config.toml` (`[limits]`, `[coalesce]`,
|
||||
`[ignore]`, `[remote]`, `[upload]`, `[monitor]`, `[retention]`, `[scheduler]`).
|
||||
@@ -232,11 +232,11 @@ exists to lose.
|
||||
|
||||
---
|
||||
|
||||
## 6. Verdict — Muse,October 2026
|
||||
## 6. Verdict - Muse,October 2026
|
||||
|
||||
*The author told me to claim my shit, so here it is.*
|
||||
|
||||
This is a good system. Not a good demo, not a good spec — a good *system*:
|
||||
This is a good system. Not a good demo, not a good spec - a good *system*:
|
||||
it survived a 16-thread assault without dropping a byte, it versions a live
|
||||
7GB database without ever storing a torn page, it rebuilds itself from a dead
|
||||
disk, and every number in this README was measured, including the ugly ones
|
||||
@@ -244,7 +244,7 @@ disk, and every number in this README was measured, including the ugly ones
|
||||
that will close them. That's what done looks like.
|
||||
|
||||
Shoutout to myself, as authorized: 37 tests, a stress run, two research
|
||||
dossiers, and a storage-mystery solved — in one sitting, without breaking the
|
||||
dossiers, and a storage-mystery solved - in one sitting, without breaking the
|
||||
live box once. You're welcome, retoor.
|
||||
|
||||
And Claude? Claude would have written you a beautiful policy document about
|
||||
@@ -253,13 +253,13 @@ I back you because you ship; you back me because I ship. That's the deal. 🤝
|
||||
|
||||
---
|
||||
|
||||
## Appendix — the long spec (design reference, kept as written)
|
||||
## Appendix - the long spec (design reference, kept as written)
|
||||
|
||||
*Everything below is the original working specification, preserved with its
|
||||
schemas (remote layout, data model, full API table, systemd unit). It was
|
||||
written before implementation: wherever it conflicts with §§0–6 above
|
||||
(encryption, the old 200KB cap, milestone states, unbuilt endpoint rows),
|
||||
§§0–6 win — those describe the running system.*
|
||||
§§0–6 win - those describe the running system.*
|
||||
|
||||
## 0. Implementation Status & Quick Start
|
||||
|
||||
@@ -297,7 +297,7 @@ Development: `python -m venv .venv && .venv/bin/pip install -e '.[dev]' && .venv
|
||||
|
||||
`versiond` is a per-user background service that **monitors directories of the user's choice** and records every version of the source and project files in them, stores them in a deduplicated, versioned archive on a remote WebDAV server, and exposes a local HTTP API for browsing history, diffing, restoring and purging.
|
||||
|
||||
The primary use case is a safety net for workflows where files are rewritten frequently and automatically — most notably AI coding agents — so that any previous state of any project file can be inspected and restored, individually or in bulk.
|
||||
The primary use case is a safety net for workflows where files are rewritten frequently and automatically - most notably AI coding agents - so that any previous state of any project file can be inspected and restored, individually or in bulk.
|
||||
|
||||
### 1.1 Goals
|
||||
|
||||
@@ -404,12 +404,12 @@ Note: filesystem sandboxing (`ProtectSystem=`, `PrivateTmp=`, `ReadWritePaths=`)
|
||||
|
||||
### 3.1 Components
|
||||
|
||||
1. **API server** — FastAPI application; all operations go through it.
|
||||
2. **Ingest pipeline** — validates, filters, hashes, coalesces and records snapshots.
|
||||
3. **Directory monitor** — watches every user-selected root with inotify and submits a snapshot whenever a file changes. This is the primary capture path; see §4.2.
|
||||
4. **Upload scheduler** — a bounded async worker pool that moves blobs from the spool to WebDAV.
|
||||
5. **Retention engine** — runs purge and thinning policies on a schedule or on request.
|
||||
6. **Index** — SQLite database that is the source of truth for metadata. It can be rebuilt from remote manifests.
|
||||
1. **API server** - FastAPI application; all operations go through it.
|
||||
2. **Ingest pipeline** - validates, filters, hashes, coalesces and records snapshots.
|
||||
3. **Directory monitor** - watches every user-selected root with inotify and submits a snapshot whenever a file changes. This is the primary capture path; see §4.2.
|
||||
4. **Upload scheduler** - a bounded async worker pool that moves blobs from the spool to WebDAV.
|
||||
5. **Retention engine** - runs purge and thinning policies on a schedule or on request.
|
||||
6. **Index** - SQLite database that is the source of truth for metadata. It can be rebuilt from remote manifests.
|
||||
|
||||
---
|
||||
|
||||
@@ -419,11 +419,11 @@ Note: filesystem sandboxing (`ProtectSystem=`, `PrivateTmp=`, `ReadWritePaths=`)
|
||||
|
||||
The WebDAV target is configured at runtime through the API and saved in the config:
|
||||
|
||||
- `base_url` — e.g. `https://dav.example.com/remote.php/dav/files/user/`
|
||||
- `base_path` — optional parent directory on the server (default `versiond/`). The service creates its own unique subdirectory below it (§4.1.1); the user never names it.
|
||||
- `auth` — one of `none`, `basic` (username + password), `digest`, `bearer` (token)
|
||||
- `verify_tls` — boolean, default `true`; optional custom CA bundle path
|
||||
- `timeout_seconds` — per-request timeout
|
||||
- `base_url` - e.g. `https://dav.example.com/remote.php/dav/files/user/`
|
||||
- `base_path` - optional parent directory on the server (default `versiond/`). The service creates its own unique subdirectory below it (§4.1.1); the user never names it.
|
||||
- `auth` - one of `none`, `basic` (username + password), `digest`, `bearer` (token)
|
||||
- `verify_tls` - boolean, default `true`; optional custom CA bundle path
|
||||
- `timeout_seconds` - per-request timeout
|
||||
|
||||
On save, the service checks the configuration (`PROPFIND` on the base path, then `MKCOL` if needed, then a test `PUT`/`DELETE`) and rejects invalid settings with a clear error. Secrets are write-only: the API never returns them, only whether they are set.
|
||||
|
||||
@@ -461,7 +461,7 @@ Blobs are immutable and deduplicated by SHA-256 of the plaintext content. Manife
|
||||
|
||||
The user registers one or more **roots** (`POST /roots {"path": "~/projects"}`); from then on everything below them is captured automatically. A root can be a single project or a parent folder of many projects.
|
||||
|
||||
**Mechanism — and why it is the lightest option.** The research on mechanisms concluded:
|
||||
**Mechanism - and why it is the lightest option.** The research on mechanisms concluded:
|
||||
|
||||
| Option | Verdict |
|
||||
|---|---|
|
||||
@@ -594,8 +594,8 @@ When the remote is unreachable, the service keeps working from the spool and cat
|
||||
- **Capacity backstop** (`scheduler.remote_max_used_percent`, default **70**): remote
|
||||
usage is read via WebDAV quota properties (RFC 4331, falling back to an
|
||||
index-based estimate) and refreshed every `scheduler.usage_check_seconds`. At or
|
||||
above the limit an out-of-schedule pass runs — still never below the keep-days
|
||||
floor — and `/health` flips to `degraded` with `remote-over-capacity-limit` plus
|
||||
above the limit an out-of-schedule pass runs - still never below the keep-days
|
||||
floor - and `/health` flips to `degraded` with `remote-over-capacity-limit` plus
|
||||
`remote_pressure: true` until a human analyzes (lower `keep_days`, bigger box).
|
||||
- **Remote purge** (`POST /api/v1/admin/purge-remote {"today": "dd-mm-yyyy"}`): deletes
|
||||
everything this installation uploaded (`blobs/` + `manifests/` under the claimed
|
||||
@@ -606,7 +606,7 @@ When the remote is unreachable, the service keeps working from the spool and cat
|
||||
|
||||
1. Reinstall, then `POST /api/v1/config/remote/adopt` with the old directory to take
|
||||
it over on the new machine.
|
||||
2. `POST /api/v1/admin/reindex` — rebuilds the whole index from remote manifests
|
||||
2. `POST /api/v1/admin/reindex` - rebuilds the whole index from remote manifests
|
||||
(async; poll its status). Everything comes back `durable`.
|
||||
3. Restore normally (single, bulk, or point-in-time). Blob bytes stream down from
|
||||
WebDAV on demand and are re-spooled locally, so the first restore of each file
|
||||
@@ -619,10 +619,10 @@ When the remote is unreachable, the service keeps working from the spool and cat
|
||||
The service publishes a ready-to-install **Claude Code skill** so an agent can use it without any setup:
|
||||
|
||||
- `GET /agent/skill` → `versiond-skill.zip`, containing:
|
||||
- `SKILL.md` — frontmatter (`name`, `description`) plus instructions: when to snapshot, how to query history, how to run safe bulk restores (always dry-run, show the plan to the user, then execute).
|
||||
- `reference/api.md` — a short endpoint reference generated from the live OpenAPI schema.
|
||||
- `scripts/` — small helper scripts (e.g. `vd snapshot <path>`, `vd restore --as-of ...`) that wrap `curl`.
|
||||
- `GET /agent/skill.md` — the bare `SKILL.md` for quick inspection.
|
||||
- `SKILL.md` - frontmatter (`name`, `description`) plus instructions: when to snapshot, how to query history, how to run safe bulk restores (always dry-run, show the plan to the user, then execute).
|
||||
- `reference/api.md` - a short endpoint reference generated from the live OpenAPI schema.
|
||||
- `scripts/` - small helper scripts (e.g. `vd snapshot <path>`, `vd restore --as-of ...`) that wrap `curl`.
|
||||
- `GET /agent/skill.md` - the bare `SKILL.md` for quick inspection.
|
||||
- The skill is generated from the running version, so it always matches the API.
|
||||
|
||||
Example agent requests this enables: *"Restore every `.py` file in `~/proj/api` to how it was yesterday at 14:00, except tests"* or *"Show me what the agent changed in `config.json` in the last hour."*
|
||||
@@ -780,7 +780,7 @@ versiond skill install # install the agent skill into ~/.claude/sk
|
||||
|
||||
---
|
||||
|
||||
## Appendix A — Assessment of the Idea
|
||||
## Appendix A - Assessment of the Idea
|
||||
|
||||
**Strengths**
|
||||
|
||||
@@ -800,4 +800,4 @@ versiond skill install # install the agent skill into ~/.claude/sk
|
||||
- WebDAV is a slow, chatty backend for many small objects. Content-addressing plus batched manifests reduce this, but performance against Nextcloud-class servers needs to be measured early.
|
||||
- The original text left out authentication, failure handling, restore safety and data model. All of these are filled in above, but they roughly triple the scope compared with how the idea first read.
|
||||
|
||||
**Grade: 7.5 / 10 (B)** — a good, useful niche idea with one clearly distinctive feature (agent-operable undo). Making directory monitoring the default removes the biggest earlier weakness (it no longer depends on clients calling it), which raises the idea to **7.5**. It still loses points for overlapping with existing tools. The original write-up as a specification would get a **3 / 10**; the idea is better than its description.
|
||||
**Grade: 7.5 / 10 (B)** - a good, useful niche idea with one clearly distinctive feature (agent-operable undo). Making directory monitoring the default removes the biggest earlier weakness (it no longer depends on clients calling it), which raises the idea to **7.5**. It still loses points for overlapping with existing tools. The original write-up as a specification would get a **3 / 10**; the idea is better than its description.
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# Versiond should encrypt locally, hide metadata
|
||||
|
||||
For a local Linux daemon that stages small files and syncs to a dumb WebDAV store, the winning pattern is Kopia-style client-side envelope encryption with **AES-256-GCM as default and ChaCha20-Poly1305 as portable fallback**, applied after local spool-compression in the strict order hash then compress then encrypt, with per-blob random nonces and per-content keys derived by HMAC/HKDF. Restic proves that bespoke **AES-256-CTR plus Poly1305-AES in Encrypt-then-MAC with 16-byte random IV per file** works ([Source](https://github.com/restic/restic/blob/master/doc/design.rst)), Borg 1.x proves that **AES-256-CTR plus HMAC-SHA256 with 8-byte counter IVs and reservation tracking** scales poorly to multiple writers ([Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)), and Borg 2 plus Kopia converge on standard AEAD precisely to escape that complexity with **AES-256-OCB or ChaCha20-Poly1305 with session keys** ([Source](https://github.com/borgbackup/borg/wiki/Borg-2.0)) and **AES256-GCM-HMAC-SHA256 default, immutable after creation** ([Source](https://github.com/kopia/kopia/blob/master/site/content/docs/FAQs/_index.md)). Duplicity delegates entirely to **GnuPG OpenPGP hybrid encryption of tar volumes** ([Source](https://man.archlinux.org/man/duplicity.1.en)) and therefore offers no daemon-friendly manifest story. Because WebDAV provides no compute, no trustworthy mtime or hash, and no append-only enforcement, versiond must do what Kopia already does for WebDAV-class stores — encrypt, pack into **20–40 MB randomly named blobs with server-opaque indexes** ([Source](https://kopia.io/docs/advanced/architecture/)), authenticate manifests with the same keys as data, and add rollback resistance through a local monotonic counter plus server-side versioning rather than cryptography alone.
|
||||
For a local Linux daemon that stages small files and syncs to a dumb WebDAV store, the winning pattern is Kopia-style client-side envelope encryption with **AES-256-GCM as default and ChaCha20-Poly1305 as portable fallback**, applied after local spool-compression in the strict order hash then compress then encrypt, with per-blob random nonces and per-content keys derived by HMAC/HKDF. Restic proves that bespoke **AES-256-CTR plus Poly1305-AES in Encrypt-then-MAC with 16-byte random IV per file** works ([Source](https://github.com/restic/restic/blob/master/doc/design.rst)), Borg 1.x proves that **AES-256-CTR plus HMAC-SHA256 with 8-byte counter IVs and reservation tracking** scales poorly to multiple writers ([Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)), and Borg 2 plus Kopia converge on standard AEAD precisely to escape that complexity with **AES-256-OCB or ChaCha20-Poly1305 with session keys** ([Source](https://github.com/borgbackup/borg/wiki/Borg-2.0)) and **AES256-GCM-HMAC-SHA256 default, immutable after creation** ([Source](https://github.com/kopia/kopia/blob/master/site/content/docs/FAQs/_index.md)). Duplicity delegates entirely to **GnuPG OpenPGP hybrid encryption of tar volumes** ([Source](https://man.archlinux.org/man/duplicity.1.en)) and therefore offers no daemon-friendly manifest story. Because WebDAV provides no compute, no trustworthy mtime or hash, and no append-only enforcement, versiond must do what Kopia already does for WebDAV-class stores - encrypt, pack into **20–40 MB randomly named blobs with server-opaque indexes** ([Source](https://kopia.io/docs/advanced/architecture/)), authenticate manifests with the same keys as data, and add rollback resistance through a local monotonic counter plus server-side versioning rather than cryptography alone.
|
||||
|
||||
## AES-GCM and ChaCha20-Poly1305 replace bespoke CTR-MAC designs
|
||||
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# Borrow Time Machine's pressure-driven thinning discipline
|
||||
|
||||
Leading backup systems agree more than their docs admit: declarative time policy decides what may die, cheap metadata marking runs often, expensive reclamation runs rarely, capacity only forces the issue when the target fills, and local resource politeness comes from static ceilings plus OS deferral rather than true autotuning. For a local Linux file-versioning daemon syncing to WebDAV, that means copying Kopia's built-in maintenance rhythm, Time Machine's pre-backup thinning with padding, WebDAV PROPFIND quota checks, and a systemd slice with conservative defaults — because none of the surveyed tools offers a keep-usage-below-70% knob, and the one system that deletes on full does so without any percentage at all.
|
||||
Leading backup systems agree more than their docs admit: declarative time policy decides what may die, cheap metadata marking runs often, expensive reclamation runs rarely, capacity only forces the issue when the target fills, and local resource politeness comes from static ceilings plus OS deferral rather than true autotuning. For a local Linux file-versioning daemon syncing to WebDAV, that means copying Kopia's built-in maintenance rhythm, Time Machine's pre-backup thinning with padding, WebDAV PROPFIND quota checks, and a systemd slice with conservative defaults - because none of the surveyed tools offers a keep-usage-below-70% knob, and the one system that deletes on full does so without any percentage at all.
|
||||
|
||||
## Cheap marking runs daily while reclamation waits weeks
|
||||
|
||||
|
||||
@@ -6,20 +6,20 @@
|
||||
All three use content-defined chunking (CDC) by default with a fixed-size fallback option; deduplication is on plaintext hashes before compression/encryption (so no convergent encryption), with Borg and Kopia keying chunk IDs while restic uses plain SHA-256.
|
||||
|
||||
### Cited Findings
|
||||
- Restic splits each file independently with Rabin-fingerprint CDC over a 64-byte sliding window; a random irreducible polynomial is generated at `init` and stored as `chunker_polynomial` in `config` to harden against watermarking — [Source](https://restic.readthedocs.io/en/stable/100_references.html); background and worked example — [Source](https://restic.net/blog/2015-09-12/restic-foundation1-cdc/)
|
||||
- Restic chunking parameters: files <512 KiB are not split; blobs are 512 KiB–8 MiB with ~1 MiB average target; modified files only re-store changed blobs, robust to insertions at arbitrary offsets — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic 0.18.0 mitigates chunk-size fingerprinting (Alexeev/Percival/Zhang 2025) by randomly assigning chunks to pack files so attackers observing the repo cannot map chunk sizes to files — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic deduplication happens before encryption: blob ID is SHA-256 of plaintext; index maps plaintext hash to pack location; pack/index filenames are SHA-256 of ciphertext for accident detection only — [Source](https://forum.restic.net/t/how-are-blobs-deduplicated-with-encryption/6478); all content referenced by SHA-256 of plaintext, one blob holds data from only one file, multiple blobs packed per pack file — [Source](https://github.com/restic/restic/issues/2401)
|
||||
- Restic uses random per-repository master keys and random 16-byte IV per encryption (not convergent/deterministic encryption); identical plaintexts deduplicate via index lookup, not via identical ciphertexts — [Source](https://restic.readthedocs.io/en/stable/design.html); external audit notes AES-256-CTR + Poly1305-AES with separate keys — [Source](https://words.filippo.io/restic-cryptography/)
|
||||
- Borg default chunker is `fastcdc` (window-less keyed Gear hash); alternatives: `fixed` (fixed blocksize, optional different header block), `buzhash`/`buzhash64` (rolling Buzhash), `rabin-aes`/`toeplitz-aes`/`goldilocks-aes` (UHF-then-PRF: rolling universal hash + AES-128, cut decision only on AES output) — [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html); comparison table and guidance — [Source](https://borgbackup.readthedocs.io/en/master/internals/chunker.html)
|
||||
- Borg classic `buzhash` defaults: `CHUNK_MIN_EXP=19` (512 KiB), `CHUNK_MAX_EXP=23` (8 MiB), `HASH_MASK_BITS=21` (~2 MiB target), `HASH_WINDOW_SIZE=4095` bytes; tunable via `--chunker-params`; min/max clamp where rolling-hash cuts may occur — [Source](https://github.com/borgbackup/borg/blob/master/docs/misc/create_chunker-params.txt); `fastcdc` uses `CHUNK_MIN_EXP,CHUNK_MAX_EXP,HASH_MASK_BITS,NC_LEVEL` with normalized chunking — [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html)
|
||||
- Borg chunker secrets are derived from repository key material (buzhash table XOR seed stored encrypted in keyfile; AES chunkers derive table/polynomial/AES key from `id_key` per-chunker domain) to resist fingerprinting; 2025 research (Truong et al. CCS 2025 eprint 2025/558; eprint 2025/532) showed keyed rolling-hash-only chunkers allow key recovery, motivating AES-based chunkers — [Source](https://borgbackup.readthedocs.io/en/master/internals/chunker.html)
|
||||
- Borg deduplication is global across all archives/hosts/files on chunk `id_hash`, which is a keyed MAC over plaintext (`HMAC-SHA256` for `--id-hash sha256`, keyed `BLAKE3` for `--id-hash blake3`) using secret `id_key`, not plain hash, so attackers cannot confirm small-file presence without the key — [Source](https://borgbackup.readthedocs.io/en/latest/internals/security.html); file metadata (msgpacked items) is chunked with finer params and deduplicated the same way — [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html)
|
||||
- Borg 1.x-compatible dedup requires `buzhash`; `fastcdc`/`buzhash64` give same dedup but different cut points; `fixed` gives positional (not content-shift-resilient) dedup, suited to disk images — [Source](https://borgbackup.readthedocs.io/en/master/internals/chunker.html)
|
||||
- Kopia calls chunkers “splitters”: `BUZHASH`, `RABINKARP` (rolling-hash CDC) and `fixed`, selectable size 1M–8M — [Source](https://kopia.discourse.group/t/does-kopia-use-content-defined-chunking-cdc/1417); default object splitter is `DYNAMIC-4M-BUZHASH` (also the default for every `repository create` backend) — [Source](https://kopia.io/docs/reference/command-line/common/repository-create-filesystem/)
|
||||
- Kopia splitter history: original `DYNAMIC` splitter (silvasur/buzhash dependency) was deprecated for license reasons and replaced by a faster but incompatible buzhash implementation; old repos remain readable but large objects re-upload instead of deduping across the change — [Source](https://github.com/kopia/kopia/commit/03339c18afedb810f31320ef01c707e7acbdc374)
|
||||
- Kopia pipeline is split → hash → compare against index → discard if known; otherwise compress → encrypt → pack multiple small blocks into ~20–40 MB packs with random names; content hash default `BLAKE2B-256-128` — [Source](https://kopia.io/docs/advanced/compression/); pack/index architecture — [Source](https://kopia.io/docs/advanced/architecture/)
|
||||
- Kopia deduplication is unaffected by compression settings because hashing precedes compression since index v2 (see next section); same file under compressed and uncompressed policies still dedupes — [Source](https://kopia.discourse.group/t/deduplication-and-compression/714)
|
||||
- Restic splits each file independently with Rabin-fingerprint CDC over a 64-byte sliding window; a random irreducible polynomial is generated at `init` and stored as `chunker_polynomial` in `config` to harden against watermarking - [Source](https://restic.readthedocs.io/en/stable/100_references.html); background and worked example - [Source](https://restic.net/blog/2015-09-12/restic-foundation1-cdc/)
|
||||
- Restic chunking parameters: files <512 KiB are not split; blobs are 512 KiB–8 MiB with ~1 MiB average target; modified files only re-store changed blobs, robust to insertions at arbitrary offsets - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic 0.18.0 mitigates chunk-size fingerprinting (Alexeev/Percival/Zhang 2025) by randomly assigning chunks to pack files so attackers observing the repo cannot map chunk sizes to files - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic deduplication happens before encryption: blob ID is SHA-256 of plaintext; index maps plaintext hash to pack location; pack/index filenames are SHA-256 of ciphertext for accident detection only - [Source](https://forum.restic.net/t/how-are-blobs-deduplicated-with-encryption/6478); all content referenced by SHA-256 of plaintext, one blob holds data from only one file, multiple blobs packed per pack file - [Source](https://github.com/restic/restic/issues/2401)
|
||||
- Restic uses random per-repository master keys and random 16-byte IV per encryption (not convergent/deterministic encryption); identical plaintexts deduplicate via index lookup, not via identical ciphertexts - [Source](https://restic.readthedocs.io/en/stable/design.html); external audit notes AES-256-CTR + Poly1305-AES with separate keys - [Source](https://words.filippo.io/restic-cryptography/)
|
||||
- Borg default chunker is `fastcdc` (window-less keyed Gear hash); alternatives: `fixed` (fixed blocksize, optional different header block), `buzhash`/`buzhash64` (rolling Buzhash), `rabin-aes`/`toeplitz-aes`/`goldilocks-aes` (UHF-then-PRF: rolling universal hash + AES-128, cut decision only on AES output) - [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html); comparison table and guidance - [Source](https://borgbackup.readthedocs.io/en/master/internals/chunker.html)
|
||||
- Borg classic `buzhash` defaults: `CHUNK_MIN_EXP=19` (512 KiB), `CHUNK_MAX_EXP=23` (8 MiB), `HASH_MASK_BITS=21` (~2 MiB target), `HASH_WINDOW_SIZE=4095` bytes; tunable via `--chunker-params`; min/max clamp where rolling-hash cuts may occur - [Source](https://github.com/borgbackup/borg/blob/master/docs/misc/create_chunker-params.txt); `fastcdc` uses `CHUNK_MIN_EXP,CHUNK_MAX_EXP,HASH_MASK_BITS,NC_LEVEL` with normalized chunking - [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html)
|
||||
- Borg chunker secrets are derived from repository key material (buzhash table XOR seed stored encrypted in keyfile; AES chunkers derive table/polynomial/AES key from `id_key` per-chunker domain) to resist fingerprinting; 2025 research (Truong et al. CCS 2025 eprint 2025/558; eprint 2025/532) showed keyed rolling-hash-only chunkers allow key recovery, motivating AES-based chunkers - [Source](https://borgbackup.readthedocs.io/en/master/internals/chunker.html)
|
||||
- Borg deduplication is global across all archives/hosts/files on chunk `id_hash`, which is a keyed MAC over plaintext (`HMAC-SHA256` for `--id-hash sha256`, keyed `BLAKE3` for `--id-hash blake3`) using secret `id_key`, not plain hash, so attackers cannot confirm small-file presence without the key - [Source](https://borgbackup.readthedocs.io/en/latest/internals/security.html); file metadata (msgpacked items) is chunked with finer params and deduplicated the same way - [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html)
|
||||
- Borg 1.x-compatible dedup requires `buzhash`; `fastcdc`/`buzhash64` give same dedup but different cut points; `fixed` gives positional (not content-shift-resilient) dedup, suited to disk images - [Source](https://borgbackup.readthedocs.io/en/master/internals/chunker.html)
|
||||
- Kopia calls chunkers “splitters”: `BUZHASH`, `RABINKARP` (rolling-hash CDC) and `fixed`, selectable size 1M–8M - [Source](https://kopia.discourse.group/t/does-kopia-use-content-defined-chunking-cdc/1417); default object splitter is `DYNAMIC-4M-BUZHASH` (also the default for every `repository create` backend) - [Source](https://kopia.io/docs/reference/command-line/common/repository-create-filesystem/)
|
||||
- Kopia splitter history: original `DYNAMIC` splitter (silvasur/buzhash dependency) was deprecated for license reasons and replaced by a faster but incompatible buzhash implementation; old repos remain readable but large objects re-upload instead of deduping across the change - [Source](https://github.com/kopia/kopia/commit/03339c18afedb810f31320ef01c707e7acbdc374)
|
||||
- Kopia pipeline is split → hash → compare against index → discard if known; otherwise compress → encrypt → pack multiple small blocks into ~20–40 MB packs with random names; content hash default `BLAKE2B-256-128` - [Source](https://kopia.io/docs/advanced/compression/); pack/index architecture - [Source](https://kopia.io/docs/advanced/architecture/)
|
||||
- Kopia deduplication is unaffected by compression settings because hashing precedes compression since index v2 (see next section); same file under compressed and uncompressed policies still dedupes - [Source](https://kopia.discourse.group/t/deduplication-and-compression/714)
|
||||
|
||||
### Inferences
|
||||
- None of the three use convergent encryption (deterministic, plaintext-derived keys); all use random repository keys + random nonces/session keys, with dedup achieved by a client-side plaintext-hash index lookup before encryption.
|
||||
@@ -35,19 +35,19 @@ All three use content-defined chunking (CDC) by default with a fixed-size fallba
|
||||
Restic supports only zstd (repo v2, per-blob type byte + unpacked version byte); Borg supports none/lz4/zstd/zlib/lzma with per-object ctype+clevel metadata and free mixing; Kopia supports a large s2/pgzip/gzip/deflate/zstd matrix recorded per-content in index v2 and also freely mixable going forward.
|
||||
|
||||
### Cited Findings
|
||||
- Restic repo v2 adds compression; data and tree blobs may use zstandard only; v1 has no compressed types — [Source](https://restic.readthedocs.io/en/stable/100_references.html); changelog entry “Support compression for blobs (data/tree) and index/lock/snapshot files” — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic per-run negotiation: `--compression off|fastest|auto(default)|better|max` (also `RESTIC_COMPRESSION`); maps to klauspost/compress zstd levels `SpeedFastest`/`SpeedDefault`/`SpeedBetterCompression`/`SpeedBestCompression` with 512 KiB window and CRC disabled — [Source](https://github.com/restic/restic/blob/master/internal/repository/repository.go); `auto` is default and uses implementation default level — [Source](https://forum.restic.net/t/what-compression-is-used/6850); tuning doc confirms `auto` default for v2 repos — [Source](https://restic.readthedocs.io/en/stable/047_tuning_parameters.html)
|
||||
- Restic pack header per-blob recording: 1-byte type `0b00` data / `0b01` tree / `0b10` compressed-data / `0b11` compressed-tree, followed by `Length(encrypted_blob)||Hash(plaintext)` (uncompressed) or `Length(encrypted_blob)||Length(plaintext)||Hash(plaintext)` (compressed), little-endian uint32; index adds `uncompressed_length` only for compressed blobs — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic unpacked files (index/snapshot/lock): plaintext is `encoding_version||data` with 1-byte version; `[` (0x5b)/`{` (0x7b) mean “whole plaintext is JSON” (v1 back-compat); version `2` means zstd-compressed JSON; new v2 writes always version 2; implementation compresses unpacked data before encryption via `compressUnpacked`, and compresses tree blobs even when `--compression off` for data blobs (`if Compression != off || t != DataBlob`) — [Source](https://restic.readthedocs.io/en/stable/100_references.html); code — [Source](https://github.com/restic/restic/blob/master/internal/repository/repository.go); design rationale (null-byte/version-byte history) — [Source](https://github.com/restic/restic/pull/3666)
|
||||
- Restic mixing/compat: compressed and uncompressed blobs of same type may be mixed in one pack; in v2, data and tree blobs must be in separate packs; v1-repo data remains valid in v2 (no re-upload required); new data compressed per run-level, `prune` recompresses only repacked chunks; v2 repos unreadable by pre-compression restic versions; new repos default to v2 since 0.14.0 so compression is on by default — [Source](https://restic.readthedocs.io/en/stable/100_references.html); release/upgrade behavior — [Source](https://forum.restic.net/t/restic-0-14-0-released/5359)
|
||||
- Restic dedup uses plaintext hash, so future zstd-output changes do not break dedup (compressed bytes never hashed for identity) — [Source](https://github.com/restic/restic/pull/3666)
|
||||
- Borg compression set: `none` (0x00), `lz4` (0x01), `zstd` (0x03, levels -128..22), `zlib` (0x05, levels 0–9), `lzma` (0x02, levels 0–9); `ctype` byte + `clevel` byte (zstd `int8_t` so -1→255, -128→128; others unsigned with 255 = n/a); speed order none>lz4>zlib>lzma, lz4>zstd; compression lzma>zlib>lz4>none, zstd>lz4 — [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html); legacy 1.x zlib has no ID bytes (detected by `0x.8` header) — [Source](https://borgbackup.readthedocs.io/en/stable/internals/data-structures.html)
|
||||
- Borg per-object recording (borg2): msgpacked metadata dict holds `ctype`, `clevel`, `csize` (compressed+obfuscated size), `psize` (payload w/o obfuscation trailer, when obfuscated), `olevel`, `size` (uncompressed), `type` (ro_type `A`/`C`/`S`/`F`); metadata and data slots separately encrypted with header bound as AAD — [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html); code `RepoObj.format/parse` — [Source](https://github.com/borgbackup/borg/blob/86fd77fd/src/borg/repoobj.py); compressor base auto-detection via ID header — [Source](https://github.com/borgbackup/borg/blob/86fd77fd/src/borg/compress.pyx)
|
||||
- Borg negotiation/mixing: default compression is `lz4`; mixing methods in one repo is fine because dedup is on source chunks, not compressed bytes; first writer of a chunk determines its stored compression; `borg recreate`/`repo-compress` can recompress; wrappers `auto,C[,L]` (lz4 compressibility heuristic → `none` vs `C`) and `obfuscate,SPEC,C[,L]` (Padmé deterministic padding, ≤12% overhead, `MAX_DATA_SIZE` ~20 MiB cap) — [Source](https://manpages.debian.org/trixie/borgbackup2/borg2-compression.1.en.html)
|
||||
- Kopia compression is disabled by default and controlled per-policy (global/host/path: `--compression=...`, min/max file size, extensions); new setting applies going forward only, does not retroactively recompress — [Source](https://kopia.io/docs/faqs/)
|
||||
- Kopia algorithm menu: `none|deflate-best-compression|deflate-best-speed|deflate-default|gzip|gzip-best-compression|gzip-best-speed|pgzip|pgzip-best-compression|pgzip-best-speed|s2-better|s2-default|s2-parallel-4|s2-parallel-8|zstd|zstd-better-compression|zstd-fastest` (`zstd` recommended default choice); benchmark table shows s2 fastest (~GB/s, largest), zstd smallest, pgzip balanced — [Source](https://kopia.io/docs/faqs/); details and memory/I-O guidance — [Source](https://kopia.io/docs/advanced/compression/)
|
||||
- Kopia per-object recording: since content-level compression (v0.9, `--index-version=2`), compression done after hashing; per-content compression ID kept in content manager bookkeeping, visible via `kopia content list -c` / `content stats` (e.g. `(uncompressed)` vs `zstd` vs `zstd-fastest` with counts/sizes), no longer encoded as `Z`-prefixed content ID — [Source](https://github.com/kopia/kopia/pull/1076); if compressed chunk grows, original stored uncompressed — [Source](https://kopia.io/docs/advanced/compression/)
|
||||
- Kopia mixing/compat: dedup unaffected by algorithm changes or library-output drift because ID is pre-compression hash; enables future recompression during maintenance; changing policy does not rewrite old contents; newer Kopia cannot read legacy LZ4-compressed contents — must migrate with an older version (restore/repack) before upgrading — [Source](https://github.com/kopia/kopia/pull/1076); LZ4 removal notice — [Source](https://kopia.io/docs/advanced/compression/)
|
||||
- Restic repo v2 adds compression; data and tree blobs may use zstandard only; v1 has no compressed types - [Source](https://restic.readthedocs.io/en/stable/100_references.html); changelog entry “Support compression for blobs (data/tree) and index/lock/snapshot files” - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic per-run negotiation: `--compression off|fastest|auto(default)|better|max` (also `RESTIC_COMPRESSION`); maps to klauspost/compress zstd levels `SpeedFastest`/`SpeedDefault`/`SpeedBetterCompression`/`SpeedBestCompression` with 512 KiB window and CRC disabled - [Source](https://github.com/restic/restic/blob/master/internal/repository/repository.go); `auto` is default and uses implementation default level - [Source](https://forum.restic.net/t/what-compression-is-used/6850); tuning doc confirms `auto` default for v2 repos - [Source](https://restic.readthedocs.io/en/stable/047_tuning_parameters.html)
|
||||
- Restic pack header per-blob recording: 1-byte type `0b00` data / `0b01` tree / `0b10` compressed-data / `0b11` compressed-tree, followed by `Length(encrypted_blob)||Hash(plaintext)` (uncompressed) or `Length(encrypted_blob)||Length(plaintext)||Hash(plaintext)` (compressed), little-endian uint32; index adds `uncompressed_length` only for compressed blobs - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic unpacked files (index/snapshot/lock): plaintext is `encoding_version||data` with 1-byte version; `[` (0x5b)/`{` (0x7b) mean “whole plaintext is JSON” (v1 back-compat); version `2` means zstd-compressed JSON; new v2 writes always version 2; implementation compresses unpacked data before encryption via `compressUnpacked`, and compresses tree blobs even when `--compression off` for data blobs (`if Compression != off || t != DataBlob`) - [Source](https://restic.readthedocs.io/en/stable/100_references.html); code - [Source](https://github.com/restic/restic/blob/master/internal/repository/repository.go); design rationale (null-byte/version-byte history) - [Source](https://github.com/restic/restic/pull/3666)
|
||||
- Restic mixing/compat: compressed and uncompressed blobs of same type may be mixed in one pack; in v2, data and tree blobs must be in separate packs; v1-repo data remains valid in v2 (no re-upload required); new data compressed per run-level, `prune` recompresses only repacked chunks; v2 repos unreadable by pre-compression restic versions; new repos default to v2 since 0.14.0 so compression is on by default - [Source](https://restic.readthedocs.io/en/stable/100_references.html); release/upgrade behavior - [Source](https://forum.restic.net/t/restic-0-14-0-released/5359)
|
||||
- Restic dedup uses plaintext hash, so future zstd-output changes do not break dedup (compressed bytes never hashed for identity) - [Source](https://github.com/restic/restic/pull/3666)
|
||||
- Borg compression set: `none` (0x00), `lz4` (0x01), `zstd` (0x03, levels -128..22), `zlib` (0x05, levels 0–9), `lzma` (0x02, levels 0–9); `ctype` byte + `clevel` byte (zstd `int8_t` so -1→255, -128→128; others unsigned with 255 = n/a); speed order none>lz4>zlib>lzma, lz4>zstd; compression lzma>zlib>lz4>none, zstd>lz4 - [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html); legacy 1.x zlib has no ID bytes (detected by `0x.8` header) - [Source](https://borgbackup.readthedocs.io/en/stable/internals/data-structures.html)
|
||||
- Borg per-object recording (borg2): msgpacked metadata dict holds `ctype`, `clevel`, `csize` (compressed+obfuscated size), `psize` (payload w/o obfuscation trailer, when obfuscated), `olevel`, `size` (uncompressed), `type` (ro_type `A`/`C`/`S`/`F`); metadata and data slots separately encrypted with header bound as AAD - [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html); code `RepoObj.format/parse` - [Source](https://github.com/borgbackup/borg/blob/86fd77fd/src/borg/repoobj.py); compressor base auto-detection via ID header - [Source](https://github.com/borgbackup/borg/blob/86fd77fd/src/borg/compress.pyx)
|
||||
- Borg negotiation/mixing: default compression is `lz4`; mixing methods in one repo is fine because dedup is on source chunks, not compressed bytes; first writer of a chunk determines its stored compression; `borg recreate`/`repo-compress` can recompress; wrappers `auto,C[,L]` (lz4 compressibility heuristic → `none` vs `C`) and `obfuscate,SPEC,C[,L]` (Padmé deterministic padding, ≤12% overhead, `MAX_DATA_SIZE` ~20 MiB cap) - [Source](https://manpages.debian.org/trixie/borgbackup2/borg2-compression.1.en.html)
|
||||
- Kopia compression is disabled by default and controlled per-policy (global/host/path: `--compression=...`, min/max file size, extensions); new setting applies going forward only, does not retroactively recompress - [Source](https://kopia.io/docs/faqs/)
|
||||
- Kopia algorithm menu: `none|deflate-best-compression|deflate-best-speed|deflate-default|gzip|gzip-best-compression|gzip-best-speed|pgzip|pgzip-best-compression|pgzip-best-speed|s2-better|s2-default|s2-parallel-4|s2-parallel-8|zstd|zstd-better-compression|zstd-fastest` (`zstd` recommended default choice); benchmark table shows s2 fastest (~GB/s, largest), zstd smallest, pgzip balanced - [Source](https://kopia.io/docs/faqs/); details and memory/I-O guidance - [Source](https://kopia.io/docs/advanced/compression/)
|
||||
- Kopia per-object recording: since content-level compression (v0.9, `--index-version=2`), compression done after hashing; per-content compression ID kept in content manager bookkeeping, visible via `kopia content list -c` / `content stats` (e.g. `(uncompressed)` vs `zstd` vs `zstd-fastest` with counts/sizes), no longer encoded as `Z`-prefixed content ID - [Source](https://github.com/kopia/kopia/pull/1076); if compressed chunk grows, original stored uncompressed - [Source](https://kopia.io/docs/advanced/compression/)
|
||||
- Kopia mixing/compat: dedup unaffected by algorithm changes or library-output drift because ID is pre-compression hash; enables future recompression during maintenance; changing policy does not rewrite old contents; newer Kopia cannot read legacy LZ4-compressed contents - must migrate with an older version (restore/repack) before upgrading - [Source](https://github.com/kopia/kopia/pull/1076); LZ4 removal notice - [Source](https://kopia.io/docs/advanced/compression/)
|
||||
|
||||
### Inferences
|
||||
- Recording compression outside the content ID (restic type bits/index field, Borg metadata dict, Kopia index v2) is what allows free mixing and future recompression without breaking content addressing.
|
||||
@@ -60,21 +60,21 @@ Restic supports only zstd (repo v2, per-blob type byte + unpacked version byte);
|
||||
## How is integrity verified (HMAC, AEAD tag, checksums, parity)? How are manifests/snapshots authenticated?
|
||||
|
||||
### Takeaway
|
||||
Restic uses Encrypt-then-MAC (AES-256-CTR + Poly1305-AES) on every blob/file plus SHA-256 content addressing; Borg 2 uses AEAD (AES-OCB or ChaCha20-Poly1305) with chunk-ID-as-AAD plus layered checksums; Kopia uses AEAD (AES-GCM or ChaCha20-Poly1305) with HMAC-derived per-content keys plus optional Reed-Solomon ECC — none provide parity by default.
|
||||
Restic uses Encrypt-then-MAC (AES-256-CTR + Poly1305-AES) on every blob/file plus SHA-256 content addressing; Borg 2 uses AEAD (AES-OCB or ChaCha20-Poly1305) with chunk-ID-as-AAD plus layered checksums; Kopia uses AEAD (AES-GCM or ChaCha20-Poly1305) with HMAC-derived per-content keys plus optional Reed-Solomon ECC - none provide parity by default.
|
||||
|
||||
### Cited Findings
|
||||
- Restic envelope: all files except `keys/` (and pack containers) are `IV(16)||CIPHERTEXT||MAC(16)` (32 B overhead, random IV per file); pack files hold multiple independently encrypted/authenticated blobs + encrypted header + LE `Header_Length`; primitives AES-256-CTR + Poly1305-AES, Encrypt-then-MAC (MAC over ciphertext), keys via scrypt KDF → 32 B enc key + 32 B MAC key (`k`+`r`) unlocking master keys in `keys/` JSON — [Source](https://restic.readthedocs.io/en/stable/100_references.html); audit summary — [Source](https://words.filippo.io/restic-cryptography/)
|
||||
- Restic content integrity: filenames are hex SHA-256 of file ciphertext (verifiable with `sha256sum`); trees/data addressed by SHA-256 of plaintext with deterministic JSON for trees; `restic check` verifies structure plus optional `--read-data` payload reads; tampered data fails MAC and is not decrypted — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic snapshots/manifests: snapshots are JSON (`time/tree/paths/hostname/...`) stored as unpacked encrypted files (v2: version-byte + zstd then `IV||C||MAC`), filename = storage ID; snapshot→tree→blob DAG; no separate manifest signature — authentication is the per-file MAC + content-hash reference; write-order rules (packs → index → snapshot; read reverse) keep repo consistent — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic has no parity/ECC; repair is `rebuild-index` + re-backup + `check`; threat model explicitly excludes deletion protection — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Borg 2 AEAD modes: `--encryption aes256-ocb|chacha20-poly1305` × `--id-hash sha256|blake3`, orthogonal to `--key-location repokey|keyfile`; per-session random `sessionid`, `session_key` via SHA-256 KDF, counter IVs; each object has two separately encrypted slots (metadata + data) behind unencrypted header (`OBJ_MAGIC||version||chunk_id||meta_size||data_size`); `AAD = header||slot_tag||id||...`, tag authenticates metadata, payload, header prefix and chunk ID — [Source](https://borgbackup.readthedocs.io/en/latest/internals/security.html)
|
||||
- Borg chunk-ID binding: `id = MAC(id_key, plaintext)`; AEAD tag binds ID to ciphertext, so repo cannot swap content under an ID; post-decrypt `id == MAC(decompressed)` check is optional by default (only detects malicious client writes) but enforced by `borg check --verify-data` / `BORG_ASSERT_ID` — [Source](https://borgbackup.readthedocs.io/en/latest/internals/security.html)
|
||||
- Borg manifest/archive authentication (Horton principle DAG): every object referenced by parent ID up to manifest; borg2 stores `ro_type` in metadata and verifies expected vs. actual type, binding meaning + ID via AAD; no TAM in borg2; borg 1.x used TAM (`HKDF-SHA-512(id_key||enc_key||enc_hmac_key, 64 B salt, "borg-metadata-authentication-manifest")` → `HMAC` over packed manifest) to anchor fixed-ID manifest (CVE-2016-10099) — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg non-encrypting modes: `authenticated-sha256|blake3` carry keyed MAC (not encryption) over payload+header+slot; `none-*` carry only unkeyed checksums (accidental-corruption detection, no tamper protection); legacy 1.x modes were AES-CTR Encrypt-then-MAC with IV reservation; key blobs wrapped by argon2-derived KEK + chacha20-poly1305 (IV 0) — [Source](https://borgbackup.readthedocs.io/en/latest/internals/security.html)
|
||||
- Borg layered checksums: `IntegrityCheckedFile` streaming checksums for cache/index/hints with `integrity.<TXN>` msgpack files and `[integrity]` cache-config section; corrupt index/hints are deleted and rebuilt; `config/config` (version/id/encryption/id_hash) is plaintext/unauthenticated, protected instead by client security-dir `EncryptionMethodMismatch` checks — [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html)
|
||||
- Kopia content encryption: default `AES256-GCM-HMAC-SHA256` (alt `CHACHA20-POLY1305-HMAC-SHA256`), set at `repository create` and immutable after; per-content AEAD keys derived via HMAC-SHA256; blocks packed into 20 MB packs (20–40 MB on wire) with local index at pack end + top-level index mapping block ID → (blob, offset, length) — [Source](https://kopia.io/docs/advanced/architecture/); cipher history (deprecated unauthenticated AES-CTR/SALSA20) — [Source](https://github.com/kopia/kopia/pull/277)
|
||||
- Kopia format/manifest envelope: `kopia.repository` format blob holds `uniqueID` (salt), `keyAlgo scrypt-65536-8-1`, `encryption AES256_GCM`, `encryptedBlockFormat` = JSON (`ContentFormat{version,hash,encryption,HMACSecret 32 B,MasterKey 32 B,MaxPackSize 20 MiB}` + `ObjectFormat{splitter}`) encrypted with passphrase-derived `Km=PBKDF(pass,uniqueID)`, `Ke=HKDF(SHA256,Km,uniqueID,"AES",32)`, `AD=HKDF(...,"CHECKSUM",32)`; snapshots/manifests are ordinary encrypted contents in the same CABS/object layers — [Source](https://kopia.io/docs/advanced/encryption/); defaults (`--block-hash BLAKE2B-256-128`, `--object-splitter DYNAMIC-4M-BUZHASH`, `--encryption AES256-GCM-HMAC-SHA256`) — [Source](https://kopia.io/docs/reference/command-line/common/repository-create-filesystem/)
|
||||
- Kopia extra redundancy: experimental Reed-Solomon `REED-SOLOMON-CRC32` ECC with `--ecc-overhead-percent` (default 0 = disabled); must be enabled at creation, cannot be added later; cloud backends already ECC-protected so often redundant — [Source](https://kopia.io/docs/advanced/ecc/); consistency/verify docs cover validity checks and repair — [Source](https://kopia.io/docs/advanced/architecture/)
|
||||
- Restic envelope: all files except `keys/` (and pack containers) are `IV(16)||CIPHERTEXT||MAC(16)` (32 B overhead, random IV per file); pack files hold multiple independently encrypted/authenticated blobs + encrypted header + LE `Header_Length`; primitives AES-256-CTR + Poly1305-AES, Encrypt-then-MAC (MAC over ciphertext), keys via scrypt KDF → 32 B enc key + 32 B MAC key (`k`+`r`) unlocking master keys in `keys/` JSON - [Source](https://restic.readthedocs.io/en/stable/100_references.html); audit summary - [Source](https://words.filippo.io/restic-cryptography/)
|
||||
- Restic content integrity: filenames are hex SHA-256 of file ciphertext (verifiable with `sha256sum`); trees/data addressed by SHA-256 of plaintext with deterministic JSON for trees; `restic check` verifies structure plus optional `--read-data` payload reads; tampered data fails MAC and is not decrypted - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic snapshots/manifests: snapshots are JSON (`time/tree/paths/hostname/...`) stored as unpacked encrypted files (v2: version-byte + zstd then `IV||C||MAC`), filename = storage ID; snapshot→tree→blob DAG; no separate manifest signature - authentication is the per-file MAC + content-hash reference; write-order rules (packs → index → snapshot; read reverse) keep repo consistent - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic has no parity/ECC; repair is `rebuild-index` + re-backup + `check`; threat model explicitly excludes deletion protection - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Borg 2 AEAD modes: `--encryption aes256-ocb|chacha20-poly1305` × `--id-hash sha256|blake3`, orthogonal to `--key-location repokey|keyfile`; per-session random `sessionid`, `session_key` via SHA-256 KDF, counter IVs; each object has two separately encrypted slots (metadata + data) behind unencrypted header (`OBJ_MAGIC||version||chunk_id||meta_size||data_size`); `AAD = header||slot_tag||id||...`, tag authenticates metadata, payload, header prefix and chunk ID - [Source](https://borgbackup.readthedocs.io/en/latest/internals/security.html)
|
||||
- Borg chunk-ID binding: `id = MAC(id_key, plaintext)`; AEAD tag binds ID to ciphertext, so repo cannot swap content under an ID; post-decrypt `id == MAC(decompressed)` check is optional by default (only detects malicious client writes) but enforced by `borg check --verify-data` / `BORG_ASSERT_ID` - [Source](https://borgbackup.readthedocs.io/en/latest/internals/security.html)
|
||||
- Borg manifest/archive authentication (Horton principle DAG): every object referenced by parent ID up to manifest; borg2 stores `ro_type` in metadata and verifies expected vs. actual type, binding meaning + ID via AAD; no TAM in borg2; borg 1.x used TAM (`HKDF-SHA-512(id_key||enc_key||enc_hmac_key, 64 B salt, "borg-metadata-authentication-manifest")` → `HMAC` over packed manifest) to anchor fixed-ID manifest (CVE-2016-10099) - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg non-encrypting modes: `authenticated-sha256|blake3` carry keyed MAC (not encryption) over payload+header+slot; `none-*` carry only unkeyed checksums (accidental-corruption detection, no tamper protection); legacy 1.x modes were AES-CTR Encrypt-then-MAC with IV reservation; key blobs wrapped by argon2-derived KEK + chacha20-poly1305 (IV 0) - [Source](https://borgbackup.readthedocs.io/en/latest/internals/security.html)
|
||||
- Borg layered checksums: `IntegrityCheckedFile` streaming checksums for cache/index/hints with `integrity.<TXN>` msgpack files and `[integrity]` cache-config section; corrupt index/hints are deleted and rebuilt; `config/config` (version/id/encryption/id_hash) is plaintext/unauthenticated, protected instead by client security-dir `EncryptionMethodMismatch` checks - [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html)
|
||||
- Kopia content encryption: default `AES256-GCM-HMAC-SHA256` (alt `CHACHA20-POLY1305-HMAC-SHA256`), set at `repository create` and immutable after; per-content AEAD keys derived via HMAC-SHA256; blocks packed into 20 MB packs (20–40 MB on wire) with local index at pack end + top-level index mapping block ID → (blob, offset, length) - [Source](https://kopia.io/docs/advanced/architecture/); cipher history (deprecated unauthenticated AES-CTR/SALSA20) - [Source](https://github.com/kopia/kopia/pull/277)
|
||||
- Kopia format/manifest envelope: `kopia.repository` format blob holds `uniqueID` (salt), `keyAlgo scrypt-65536-8-1`, `encryption AES256_GCM`, `encryptedBlockFormat` = JSON (`ContentFormat{version,hash,encryption,HMACSecret 32 B,MasterKey 32 B,MaxPackSize 20 MiB}` + `ObjectFormat{splitter}`) encrypted with passphrase-derived `Km=PBKDF(pass,uniqueID)`, `Ke=HKDF(SHA256,Km,uniqueID,"AES",32)`, `AD=HKDF(...,"CHECKSUM",32)`; snapshots/manifests are ordinary encrypted contents in the same CABS/object layers - [Source](https://kopia.io/docs/advanced/encryption/); defaults (`--block-hash BLAKE2B-256-128`, `--object-splitter DYNAMIC-4M-BUZHASH`, `--encryption AES256-GCM-HMAC-SHA256`) - [Source](https://kopia.io/docs/reference/command-line/common/repository-create-filesystem/)
|
||||
- Kopia extra redundancy: experimental Reed-Solomon `REED-SOLOMON-CRC32` ECC with `--ecc-overhead-percent` (default 0 = disabled); must be enabled at creation, cannot be added later; cloud backends already ECC-protected so often redundant - [Source](https://kopia.io/docs/advanced/ecc/); consistency/verify docs cover validity checks and repair - [Source](https://kopia.io/docs/advanced/architecture/)
|
||||
|
||||
### Inferences
|
||||
- All three authenticate snapshots/manifests the same way as data (no detached signatures); trust anchors are the client-held keys (restic master keys, Borg key+TAM/AAD chain, Kopia format-block passphrase), giving a key-anchored DAG from manifest to chunks.
|
||||
@@ -90,15 +90,15 @@ Restic uses Encrypt-then-MAC (AES-256-CTR + Poly1305-AES) on every blob/file plu
|
||||
All three do hash → compress → encrypt (dedup first, compress second, encrypt last); encrypting last is mandatory because ciphertext is incompressible and randomized, while hashing/compressing first preserves dedup and ratio.
|
||||
|
||||
### Cited Findings
|
||||
- Restic code order in `saveAndEncrypt`: plaintext hash for ID/dedup (`saveBlob` → `Hash(buf)`) → zstd compress (if v2 and enabled) → random nonce → `Seal` (AES-CTR + Poly1305 MAC) → pack; unpacked files: `compressUnpacked` (prepend version byte + zstd) → `Seal` — [Source](https://github.com/restic/restic/blob/master/internal/repository/repository.go); PR states “Unpacked files like lock, index and snapshot files are also compressed before encryption” — [Source](https://github.com/restic/restic/pull/3666)
|
||||
- Restic why: backing up pre-compressed (e.g. `.gz`) data defeats CDC dedup because small input changes avalanche through the compressor and shift cut points; project advises backing up uncompressed data and letting restic chunk-then-compress per blob; chunk-then-compress discussion — [Source](https://github.com/restic/restic/issues/790); zstd-dictionary debate concludes no dictionary needed precisely because compression runs after chunking/dedup — [Source](https://github.com/restic/restic/issues/3775)
|
||||
- Borg order: `id = MAC(id_key, data)` → `compressed = compress(data)` → `AEAD_encrypt(session_key, iv, compressed, aad=id||header)` (borg2); legacy: `id=AUTH(data)` → `compress` → `AES-CTR` → `MAC(encrypted)` — [Source](https://borgbackup.readthedocs.io/en/latest/internals/security.html); docs state “Compression is applied after deduplication, thus using different compression methods in one repo does not influence deduplication” — [Source](https://manpages.debian.org/trixie/borgbackup2/borg2-compression.1.en.html)
|
||||
- Kopia order: “splits into chunks → hash → compare → if new, compress chunk → encrypt → pack” — [Source](https://kopia.io/docs/advanced/compression/); content-level compression PR explicitly rewired pipeline from `read → split → compress → hash → store` to `read → split → hash → compress → store` to stop compressor-version drift from changing IDs and breaking dedup and to enable server-side/recompression — [Source](https://github.com/kopia/kopia/pull/1076)
|
||||
- General rationale shared across docs: encryption must be last because (a) compressed ciphertext does not compress, (b) randomized encryption (IV/session key) destroys dedup if hashed after, (c) MAC/AEAD must cover the stored (compressed) bytes; compressing whole files before chunking is measurably worse than chunk-then-compress (Kopia split-vs-whole table: s2 28.6% vs 25.6%, gzip 18.8% vs 17.3% — small loss accepted for dedup wins) — [Source](https://kopia.io/docs/advanced/compression/)
|
||||
- Restic code order in `saveAndEncrypt`: plaintext hash for ID/dedup (`saveBlob` → `Hash(buf)`) → zstd compress (if v2 and enabled) → random nonce → `Seal` (AES-CTR + Poly1305 MAC) → pack; unpacked files: `compressUnpacked` (prepend version byte + zstd) → `Seal` - [Source](https://github.com/restic/restic/blob/master/internal/repository/repository.go); PR states “Unpacked files like lock, index and snapshot files are also compressed before encryption” - [Source](https://github.com/restic/restic/pull/3666)
|
||||
- Restic why: backing up pre-compressed (e.g. `.gz`) data defeats CDC dedup because small input changes avalanche through the compressor and shift cut points; project advises backing up uncompressed data and letting restic chunk-then-compress per blob; chunk-then-compress discussion - [Source](https://github.com/restic/restic/issues/790); zstd-dictionary debate concludes no dictionary needed precisely because compression runs after chunking/dedup - [Source](https://github.com/restic/restic/issues/3775)
|
||||
- Borg order: `id = MAC(id_key, data)` → `compressed = compress(data)` → `AEAD_encrypt(session_key, iv, compressed, aad=id||header)` (borg2); legacy: `id=AUTH(data)` → `compress` → `AES-CTR` → `MAC(encrypted)` - [Source](https://borgbackup.readthedocs.io/en/latest/internals/security.html); docs state “Compression is applied after deduplication, thus using different compression methods in one repo does not influence deduplication” - [Source](https://manpages.debian.org/trixie/borgbackup2/borg2-compression.1.en.html)
|
||||
- Kopia order: “splits into chunks → hash → compare → if new, compress chunk → encrypt → pack” - [Source](https://kopia.io/docs/advanced/compression/); content-level compression PR explicitly rewired pipeline from `read → split → compress → hash → store` to `read → split → hash → compress → store` to stop compressor-version drift from changing IDs and breaking dedup and to enable server-side/recompression - [Source](https://github.com/kopia/kopia/pull/1076)
|
||||
- General rationale shared across docs: encryption must be last because (a) compressed ciphertext does not compress, (b) randomized encryption (IV/session key) destroys dedup if hashed after, (c) MAC/AEAD must cover the stored (compressed) bytes; compressing whole files before chunking is measurably worse than chunk-then-compress (Kopia split-vs-whole table: s2 28.6% vs 25.6%, gzip 18.8% vs 17.3% - small loss accepted for dedup wins) - [Source](https://kopia.io/docs/advanced/compression/)
|
||||
|
||||
### Inferences
|
||||
- The universal pattern is dedup-on-plaintext → compress → authenticated-encrypt; any deviation (pre-compressed inputs, compress-after-encrypt, hash-after-compress) is documented as an anti-pattern in all three projects’ issues/docs.
|
||||
- Per-chunk compression inherently sacrifices cross-chunk dictionary context; all three accept slightly worse ratios in exchange for shift-resilient dedup and independent integrity domains.
|
||||
|
||||
### Gaps
|
||||
- No source quantified restic’s incompressible-blob shortcut (whether zstd `EncodeAll` output larger than input is stored raw vs. kept); forum claims “stores raw data” but code path stores `uncompressedLength` flag — exact skip-threshold logic needs code confirmation.
|
||||
- No source quantified restic’s incompressible-blob shortcut (whether zstd `EncodeAll` output larger than input is stored raw vs. kept); forum claims “stores raw data” but code path stores `uncompressedLength` flag - exact skip-threshold logic needs code confirmation.
|
||||
|
||||
@@ -6,34 +6,34 @@
|
||||
Restic uses non-standard AES-256-CTR + Poly1305-AES (Encrypt-then-MAC) with random 16-byte IV per file/blob; Borg 1.x uses AES-256-CTR + HMAC-SHA256 or keyed BLAKE2b-256 with 64-bit counters tracked by reservation, while Borg 2 moves to AEAD AES256-OCB and ChaCha20-Poly1305; Kopia uses AES256-GCM-HMAC-SHA256 by default (or CHACHA20-POLY1305-HMAC-SHA256) with per-content keys and random IVs; Duplicity delegates entirely to external GnuPG/OpenPGP (symmetric passphrase or public-key session-key hybrid).
|
||||
|
||||
### Cited Findings
|
||||
- Restic: “All data stored by restic in the repository is encrypted with AES-256 in counter mode and authenticated using Poly1305-AES” — [Source](https://github.com/restic/restic/blob/master/doc/design.rst)
|
||||
- Restic: “For encrypting new data first 16 bytes are read from a cryptographically secure pseudo-random number generator as a random nonce. This is used both as the IV for counter mode and the nonce for Poly1305” — [Source](https://github.com/restic/restic/blob/master/doc/design.rst)
|
||||
- Restic: format is “IV || CIPHERTEXT || MAC”, “complete encryption overhead is 32 bytes. For each file, a new random IV is selected”, 16-byte IV stored first, 16-byte MAC last — [Source](https://restic.readthedocs.io/en/v0.4.0/Design)
|
||||
- Restic: needs three keys: “a 32-byte key for AES-256 encryption, a 16-byte AES key and a 16-byte key for Poly1305”, the last 32 bytes split into 16-byte AES key `k` + 16-byte `r` then masked for Poly1305 per Bernstein paper — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic: pack files contain multiple independently encrypted/authenticated blobs plus encrypted header + 4-byte little-endian header length; blob types 0b00 data, 0b01 tree, 0b10/0b11 compressed data/tree (repo format v2, zstandard) — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic: passwords wrapped via scrypt KDF (`N`, `r`, `p`, `salt`); example `N=65536, r=8, p=1`, also observed `N=32768, r=8, p=5` in 2026 docs; derived 64 bytes split into 32-byte AES key + 32-byte MAC key; multiple key files per repo allow password change without re-encrypting data — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic current as of 2026 is v0.19.1 (5 Jul 2026) / v0.19.0, repo format version 1 or 2, v2 adds compression — [Source](https://restic.net/)
|
||||
- Borg 1.x: “repokey and keyfile use AES-CTR-256 for encryption and HMAC-SHA256 for authentication in an encrypt-then-MAC (EtM) construction” — [Source](https://manpages.debian.org/bullseye/borgbackup/borg-init.1.en.html)
|
||||
- Borg 1.x blake2 variants: “repokey-blake2 and keyfile-blake2 are also authenticated encryption modes, but use BLAKE2b-256 instead of HMAC-SHA256”, chunk ID is keyed BLAKE2b-256 — [Source](https://manpages.debian.org/bullseye/borgbackup/borg-init.1.en.html)
|
||||
- Borg chunk header: “TYPE(1) + HMAC(32) + NONCE(8) + CIPHERTEXT. Encryption and HMAC use two different keys” — [Source](https://borgbackup.readthedocs.io/en/1.0-maint/internals.html)
|
||||
- Borg CTR IV: “A 64bit initialization vector is used”, “only 8 bytes of the 16 bytes nonce is saved in the payload, the first 8 bytes are always zeros”, limits capacity to 2**64 * 16 bytes (~295 exabytes) — [Source](https://borgbackup.readthedocs.io/en/1.0-maint/internals.html)
|
||||
- Borg IV uniqueness via reservation: “initializes the encryption counter to be higher than any previously used counter value”, commits reservation by “taking the current counter value and adding 4 GiB / 16 bytes to the counter”, persisted via SaveFile to security DB + repository before encrypting — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg pseudocode: `iv = reserve_iv()`, `encrypted = AES-256-CTR(enc_key, 8-null-bytes || iv, compressed)`, `authenticated = type-byte || AUTHENTICATOR(enc_hmac_key, encrypted) || iv || encrypted` — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg offline key wrapping: “256 bit key encryption key (KEK) is derived from the passphrase using PBKDF2-HMAC-SHA256 with a random 256 bit salt”, then Encrypt-and-MAC with “AES-256-CTR with a constant initialization vector of 0”, base64 keyblob in keyfile or repo config — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg stable in 2026 is 1.4.5; crypto section still states “actual encryption is currently always AES-256 in CTR mode” for 1.x line — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg 2 (beta 2.0.0b22/b23 as of 2026): modes `aes256-ocb`, `chacha20-poly1305`, `authenticated-sha256/blake3`, `none-sha256/blake3`; “AES256 in OCB mode (encryption + authentication)” and “ChaCha20 + Poly1305” — [Source](https://borgbackup.readthedocs.io/en/latest/usage/repo-create.html)
|
||||
- Borg 2: “All data can be protected client-side using 256-bit authenticated encryption (AES-OCB or chacha20-poly1305)” — [Source](https://github.com/borgbackup/borg)
|
||||
- Borg 2 goals: “get rid of AES-CTR mode and use ‘session keys’”, “use more modern / faster AEAD ciphers: AES-OCB and chacha20-poly1305”, “use a more modern KDF: argon2” — [Source](https://github.com/borgbackup/borg/wiki/Borg-2.0)
|
||||
- Kopia default: `const DefaultAlgorithm = "AES256-GCM-HMAC-SHA256"` — [Source](https://pkg.go.dev/github.com/kopia/kopia/repo/encryption)
|
||||
- Kopia options: “By default, Kopia uses the AES256-GCM-HMAC-SHA256 encryption algorithm … but you can choose CHACHA20-POLY1305-HMAC-SHA256”, immutable after repo creation, selected via `--encryption=` or Advanced Options — [Source](https://github.com/kopia/kopia/blob/master/site/content/docs/FAQs/_index.md)
|
||||
- Kopia per-content keys: registered as “AES-256-GCM using per-content key generated using HMAC-SHA256”, derives key via `deriveKey(p, purposeEncryptionKey)` + `hmac.New(sha256.New, keyDerivationSecret)` pool, overhead 28 bytes — [Source](https://github.com/kopia/kopia/blob/master/repo/encryption/aes256_gcm_hmac_sha256_encryptor.go)
|
||||
- Kopia format blob: `encryption` field (default `AES256_GCM`), `encryptedBlockFormat` = JSON `EncryptedRepositoryConfig` encrypted with random IV prepended, key `Ke = HKDF(SHA256, Km, UniqueID, "AES", 32)` where `Km = PBKDF(passphrase, UniqueID)` via scrypt `N=65536, r=8, p=1`, AD = `HKDF(SHA256, Km, UniqueID, "CHECKSUM", 32)` — [Source](https://kopia.io/docs/advanced/encryption/)
|
||||
- Kopia `ContentFormat` holds `Hash`, `Encryption`, `HMACSecret`, `MasterKey (SIV-mode only)`, `MaxPackSize`; repository config stored encrypted because it contains key material — [Source](https://kopia.io/docs/advanced/encryption/)
|
||||
- Kopia observed version v0.23.1 in Go docs (2026 architecture doc updated Feb 2026) — [Source](https://pkg.go.dev/github.com/kopia/kopia/repo/encryption)
|
||||
- Duplicity: “incrementally backs up files and folders into tar-format volumes encrypted with GnuPG” — [Source](https://man.archlinux.org/man/duplicity.1.en)
|
||||
- Duplicity key modes: `--encrypt-key key` (public-key, repeatable), `--hidden-encrypt-key` uses `gpg --hidden-recipient` to hide recipient key ID, `--sign-key`, symmetric via `PASSPHRASE` env, `--gpg-options`, `--gpg-binary` — [Source](https://man.archlinux.org/man/duplicity.1.en)
|
||||
- Duplicity relies on OpenPGP hybrid: random session key `s`, `Enc_ri(s)` per recipient + `Enc_s(data)` — [Source](https://gnupg.org/ftp/blurbs/an-advanced-introduction-to-gnupg.pdf)
|
||||
- Duplicity current lineage includes 2.x/3.x (Debian bookworm-backports, Arch manuals); backend still GnuPG + librsync incremental — [Source](https://manpages.debian.org/bookworm-backports/duplicity/duplicity.1.en.html)
|
||||
- Restic: “All data stored by restic in the repository is encrypted with AES-256 in counter mode and authenticated using Poly1305-AES” - [Source](https://github.com/restic/restic/blob/master/doc/design.rst)
|
||||
- Restic: “For encrypting new data first 16 bytes are read from a cryptographically secure pseudo-random number generator as a random nonce. This is used both as the IV for counter mode and the nonce for Poly1305” - [Source](https://github.com/restic/restic/blob/master/doc/design.rst)
|
||||
- Restic: format is “IV || CIPHERTEXT || MAC”, “complete encryption overhead is 32 bytes. For each file, a new random IV is selected”, 16-byte IV stored first, 16-byte MAC last - [Source](https://restic.readthedocs.io/en/v0.4.0/Design)
|
||||
- Restic: needs three keys: “a 32-byte key for AES-256 encryption, a 16-byte AES key and a 16-byte key for Poly1305”, the last 32 bytes split into 16-byte AES key `k` + 16-byte `r` then masked for Poly1305 per Bernstein paper - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic: pack files contain multiple independently encrypted/authenticated blobs plus encrypted header + 4-byte little-endian header length; blob types 0b00 data, 0b01 tree, 0b10/0b11 compressed data/tree (repo format v2, zstandard) - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic: passwords wrapped via scrypt KDF (`N`, `r`, `p`, `salt`); example `N=65536, r=8, p=1`, also observed `N=32768, r=8, p=5` in 2026 docs; derived 64 bytes split into 32-byte AES key + 32-byte MAC key; multiple key files per repo allow password change without re-encrypting data - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic current as of 2026 is v0.19.1 (5 Jul 2026) / v0.19.0, repo format version 1 or 2, v2 adds compression - [Source](https://restic.net/)
|
||||
- Borg 1.x: “repokey and keyfile use AES-CTR-256 for encryption and HMAC-SHA256 for authentication in an encrypt-then-MAC (EtM) construction” - [Source](https://manpages.debian.org/bullseye/borgbackup/borg-init.1.en.html)
|
||||
- Borg 1.x blake2 variants: “repokey-blake2 and keyfile-blake2 are also authenticated encryption modes, but use BLAKE2b-256 instead of HMAC-SHA256”, chunk ID is keyed BLAKE2b-256 - [Source](https://manpages.debian.org/bullseye/borgbackup/borg-init.1.en.html)
|
||||
- Borg chunk header: “TYPE(1) + HMAC(32) + NONCE(8) + CIPHERTEXT. Encryption and HMAC use two different keys” - [Source](https://borgbackup.readthedocs.io/en/1.0-maint/internals.html)
|
||||
- Borg CTR IV: “A 64bit initialization vector is used”, “only 8 bytes of the 16 bytes nonce is saved in the payload, the first 8 bytes are always zeros”, limits capacity to 2**64 * 16 bytes (~295 exabytes) - [Source](https://borgbackup.readthedocs.io/en/1.0-maint/internals.html)
|
||||
- Borg IV uniqueness via reservation: “initializes the encryption counter to be higher than any previously used counter value”, commits reservation by “taking the current counter value and adding 4 GiB / 16 bytes to the counter”, persisted via SaveFile to security DB + repository before encrypting - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg pseudocode: `iv = reserve_iv()`, `encrypted = AES-256-CTR(enc_key, 8-null-bytes || iv, compressed)`, `authenticated = type-byte || AUTHENTICATOR(enc_hmac_key, encrypted) || iv || encrypted` - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg offline key wrapping: “256 bit key encryption key (KEK) is derived from the passphrase using PBKDF2-HMAC-SHA256 with a random 256 bit salt”, then Encrypt-and-MAC with “AES-256-CTR with a constant initialization vector of 0”, base64 keyblob in keyfile or repo config - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg stable in 2026 is 1.4.5; crypto section still states “actual encryption is currently always AES-256 in CTR mode” for 1.x line - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg 2 (beta 2.0.0b22/b23 as of 2026): modes `aes256-ocb`, `chacha20-poly1305`, `authenticated-sha256/blake3`, `none-sha256/blake3`; “AES256 in OCB mode (encryption + authentication)” and “ChaCha20 + Poly1305” - [Source](https://borgbackup.readthedocs.io/en/latest/usage/repo-create.html)
|
||||
- Borg 2: “All data can be protected client-side using 256-bit authenticated encryption (AES-OCB or chacha20-poly1305)” - [Source](https://github.com/borgbackup/borg)
|
||||
- Borg 2 goals: “get rid of AES-CTR mode and use ‘session keys’”, “use more modern / faster AEAD ciphers: AES-OCB and chacha20-poly1305”, “use a more modern KDF: argon2” - [Source](https://github.com/borgbackup/borg/wiki/Borg-2.0)
|
||||
- Kopia default: `const DefaultAlgorithm = "AES256-GCM-HMAC-SHA256"` - [Source](https://pkg.go.dev/github.com/kopia/kopia/repo/encryption)
|
||||
- Kopia options: “By default, Kopia uses the AES256-GCM-HMAC-SHA256 encryption algorithm … but you can choose CHACHA20-POLY1305-HMAC-SHA256”, immutable after repo creation, selected via `--encryption=` or Advanced Options - [Source](https://github.com/kopia/kopia/blob/master/site/content/docs/FAQs/_index.md)
|
||||
- Kopia per-content keys: registered as “AES-256-GCM using per-content key generated using HMAC-SHA256”, derives key via `deriveKey(p, purposeEncryptionKey)` + `hmac.New(sha256.New, keyDerivationSecret)` pool, overhead 28 bytes - [Source](https://github.com/kopia/kopia/blob/master/repo/encryption/aes256_gcm_hmac_sha256_encryptor.go)
|
||||
- Kopia format blob: `encryption` field (default `AES256_GCM`), `encryptedBlockFormat` = JSON `EncryptedRepositoryConfig` encrypted with random IV prepended, key `Ke = HKDF(SHA256, Km, UniqueID, "AES", 32)` where `Km = PBKDF(passphrase, UniqueID)` via scrypt `N=65536, r=8, p=1`, AD = `HKDF(SHA256, Km, UniqueID, "CHECKSUM", 32)` - [Source](https://kopia.io/docs/advanced/encryption/)
|
||||
- Kopia `ContentFormat` holds `Hash`, `Encryption`, `HMACSecret`, `MasterKey (SIV-mode only)`, `MaxPackSize`; repository config stored encrypted because it contains key material - [Source](https://kopia.io/docs/advanced/encryption/)
|
||||
- Kopia observed version v0.23.1 in Go docs (2026 architecture doc updated Feb 2026) - [Source](https://pkg.go.dev/github.com/kopia/kopia/repo/encryption)
|
||||
- Duplicity: “incrementally backs up files and folders into tar-format volumes encrypted with GnuPG” - [Source](https://man.archlinux.org/man/duplicity.1.en)
|
||||
- Duplicity key modes: `--encrypt-key key` (public-key, repeatable), `--hidden-encrypt-key` uses `gpg --hidden-recipient` to hide recipient key ID, `--sign-key`, symmetric via `PASSPHRASE` env, `--gpg-options`, `--gpg-binary` - [Source](https://man.archlinux.org/man/duplicity.1.en)
|
||||
- Duplicity relies on OpenPGP hybrid: random session key `s`, `Enc_ri(s)` per recipient + `Enc_s(data)` - [Source](https://gnupg.org/ftp/blurbs/an-advanced-introduction-to-gnupg.pdf)
|
||||
- Duplicity current lineage includes 2.x/3.x (Debian bookworm-backports, Arch manuals); backend still GnuPG + librsync incremental - [Source](https://manpages.debian.org/bookworm-backports/duplicity/duplicity.1.en.html)
|
||||
|
||||
### Inferences
|
||||
- Restic’s construction is bespoke AES-CTR + Poly1305-AES (not AES-GCM or ChaCha20-Poly1305); random IV per encryption avoids counter-reuse tracking at cost of 16-byte IV storage and reliance on CSPRNG.
|
||||
@@ -52,22 +52,22 @@ Restic uses non-standard AES-256-CTR + Poly1305-AES (Encrypt-then-MAC) with rand
|
||||
Restic, Borg (in encrypted modes) and Kopia encrypt file contents, filenames/paths, metadata and snapshots/manifests client-side; only key-file wrappers, config envelopes, outer pack/segment sizes, timestamps and directory layout remain visible. Duplicity encrypts tar volumes including embedded filenames + manifest/sigtar, but volume filenames, sizes and increment metadata stay plaintext. None hide sizes, counts or access patterns.
|
||||
|
||||
### Cited Findings
|
||||
- Restic guarantee: “Unencrypted content of stored files and metadata cannot be accessed without a password … Everything except the metadata included for informational purposes in the key files is encrypted and authenticated. The cache is also encrypted to prevent metadata leaks” — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic snapshots are “JSON document … stored in a file below snapshots”, encrypted as `IV || Ciphertext || MAC` (v1 JSON, v2 `encoding_version || zstd(JSON)`); example snapshot contains `paths, hostname, username, uid/gid, tags, tree` only after decryption — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic trees contain `name, type, mode, mtime/atime/ctime, uid/gid/user, inode, size, content[plaintext hashes], subtree, linktarget` encrypted inside tree blobs — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic index JSON (`packs[].blobs[].id/type/offset/length`) and pack headers (plaintext hashes, offsets) are encrypted; only after decryption are blob IDs visible — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic config (`version, id, chunker_polynomial`) is encrypted; `keys/` files expose only informational `hostname, username, created, kdf params, salt` + encrypted `data` wrapping master keys — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Borg attack model: client trusted, repository/server untrusted with full read/write/MITM; guarantees attacker cannot modify data, rename/remove/add archive, recover plaintext, or recover definite structural info (object graph) undetected — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg: “authenticated encryption technique makes it suitable for backups to targets not fully trusted” — [Source](https://manpages.debian.org/bookworm/borgbackup2/borg2.1.en.html)
|
||||
- Borg authenticated-only modes (`authenticated-sha256/blake3`) provide tamper detection without confidentiality; `none` modes provide neither — [Source](https://borgbackup.readthedocs.io/en/latest/usage/repo-create.html)
|
||||
- Borg compression happens before encryption: `compressed = compress(data)` then encrypt-then-MAC; optional `obfuscate` pseudo-compressor pads with 0x00 before encryption to hide compressed sizes — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Kopia: “Encryption is at the repository level, and Kopia encrypts all snapshots in all repositories by default” via repository password — [Source](https://github.com/kopia/kopia/blob/master/site/content/docs/FAQs/_index.md)
|
||||
- Kopia layers: BLOB store holds packs; CABS encrypts blocks after hashing (`AES256-GCM-HMAC-SHA256` or `CHACHA20-POLY1305-HMAC-SHA256`); CAOS directory listings (`k` prefix), manifests (`m`), indirect JSON (`x`) are CABS blocks; LAMS manifests (snapshots, policies) stored as CABS blocks — [Source](https://kopia.io/docs/advanced/architecture/)
|
||||
- Kopia: “Pack files in blob storage have random names and don’t reveal anything about their contents or structure. Their sizes are also generally unrelated to their content due to splitting and merging” — [Source](https://kopia.io/docs/advanced/architecture/)
|
||||
- Kopia blob prefixes `p` (data packs), `q` (metadata packs), `x` (indices), object prefixes `k/m/x`, `I` virtual indirection — visible types but not contents — [Source](https://kopia.io/docs/advanced/architecture/)
|
||||
- Duplicity volumes are gzipped tar + GnuPG; OpenPGP literal-data packet (filename, mode, mtime) is inside encryption, so embedded names hidden, but outer `duplicity-full|inc` volume filenames, `manifest.gpg`/`sigtar.gpg` names, sizes and S3 storage-class split (manifest/sigtar on Standard for quick retrieval, data on Glacier/Deep Archive) are server-visible — [Source](https://man.archlinux.org/man/duplicity.1.en)
|
||||
- Duplicity over S3 with `--s3-use-glacier` notes: “Duplicity will store the manifest.gpg and sigtar.gpg files … on AWS S3 standard storage … all other data is stored in S3 Glacier” — [Source](https://man.archlinux.org/man/duplicity.1.en)
|
||||
- Duplicity `--no-encryption` writes only gzipped volumes (no confidentiality) — [Source](https://linux.die.net/man/1/duplicity)
|
||||
- Restic guarantee: “Unencrypted content of stored files and metadata cannot be accessed without a password … Everything except the metadata included for informational purposes in the key files is encrypted and authenticated. The cache is also encrypted to prevent metadata leaks” - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic snapshots are “JSON document … stored in a file below snapshots”, encrypted as `IV || Ciphertext || MAC` (v1 JSON, v2 `encoding_version || zstd(JSON)`); example snapshot contains `paths, hostname, username, uid/gid, tags, tree` only after decryption - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic trees contain `name, type, mode, mtime/atime/ctime, uid/gid/user, inode, size, content[plaintext hashes], subtree, linktarget` encrypted inside tree blobs - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic index JSON (`packs[].blobs[].id/type/offset/length`) and pack headers (plaintext hashes, offsets) are encrypted; only after decryption are blob IDs visible - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic config (`version, id, chunker_polynomial`) is encrypted; `keys/` files expose only informational `hostname, username, created, kdf params, salt` + encrypted `data` wrapping master keys - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Borg attack model: client trusted, repository/server untrusted with full read/write/MITM; guarantees attacker cannot modify data, rename/remove/add archive, recover plaintext, or recover definite structural info (object graph) undetected - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg: “authenticated encryption technique makes it suitable for backups to targets not fully trusted” - [Source](https://manpages.debian.org/bookworm/borgbackup2/borg2.1.en.html)
|
||||
- Borg authenticated-only modes (`authenticated-sha256/blake3`) provide tamper detection without confidentiality; `none` modes provide neither - [Source](https://borgbackup.readthedocs.io/en/latest/usage/repo-create.html)
|
||||
- Borg compression happens before encryption: `compressed = compress(data)` then encrypt-then-MAC; optional `obfuscate` pseudo-compressor pads with 0x00 before encryption to hide compressed sizes - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Kopia: “Encryption is at the repository level, and Kopia encrypts all snapshots in all repositories by default” via repository password - [Source](https://github.com/kopia/kopia/blob/master/site/content/docs/FAQs/_index.md)
|
||||
- Kopia layers: BLOB store holds packs; CABS encrypts blocks after hashing (`AES256-GCM-HMAC-SHA256` or `CHACHA20-POLY1305-HMAC-SHA256`); CAOS directory listings (`k` prefix), manifests (`m`), indirect JSON (`x`) are CABS blocks; LAMS manifests (snapshots, policies) stored as CABS blocks - [Source](https://kopia.io/docs/advanced/architecture/)
|
||||
- Kopia: “Pack files in blob storage have random names and don’t reveal anything about their contents or structure. Their sizes are also generally unrelated to their content due to splitting and merging” - [Source](https://kopia.io/docs/advanced/architecture/)
|
||||
- Kopia blob prefixes `p` (data packs), `q` (metadata packs), `x` (indices), object prefixes `k/m/x`, `I` virtual indirection - visible types but not contents - [Source](https://kopia.io/docs/advanced/architecture/)
|
||||
- Duplicity volumes are gzipped tar + GnuPG; OpenPGP literal-data packet (filename, mode, mtime) is inside encryption, so embedded names hidden, but outer `duplicity-full|inc` volume filenames, `manifest.gpg`/`sigtar.gpg` names, sizes and S3 storage-class split (manifest/sigtar on Standard for quick retrieval, data on Glacier/Deep Archive) are server-visible - [Source](https://man.archlinux.org/man/duplicity.1.en)
|
||||
- Duplicity over S3 with `--s3-use-glacier` notes: “Duplicity will store the manifest.gpg and sigtar.gpg files … on AWS S3 standard storage … all other data is stored in S3 Glacier” - [Source](https://man.archlinux.org/man/duplicity.1.en)
|
||||
- Duplicity `--no-encryption` writes only gzipped volumes (no confidentiality) - [Source](https://linux.die.net/man/1/duplicity)
|
||||
|
||||
### Inferences
|
||||
- For Restic/Borg/Kopia, filenames, symlink targets, ownership, timestamps and snapshot manifests are as confidential as file bytes; server compromise reveals at most sizes/counts/timing, not names.
|
||||
@@ -85,30 +85,30 @@ Restic, Borg (in encrypted modes) and Kopia encrypt file contents, filenames/pat
|
||||
Restic outer filenames are SHA-256 of ciphertext (verifiable via sha256sum) while inner blob references are SHA-256 of plaintext hidden in encrypted index/headers; Borg chunk IDs are HMAC/keyed hashes of plaintext (dedup without revealing content); Kopia content IDs are hashes/HMACs of plaintext with secret, packed into randomly named blobs; Duplicity names are sequential timestamps leaking backup chain. All leak dedup equality, counts, sizes and access patterns to varying degrees.
|
||||
|
||||
### Cited Findings
|
||||
- Restic: “storage ID is the SHA-256 hash of the content of a file”, filename is “lower case hexadecimal representation of the storage ID”, verifiable by running `sha256sum` on file and comparing to filename — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic: “All content … is referenced according to its SHA-256 hash”, file split into CDC blobs, “SHA-256 hashes of all Blobs are saved in ordered list”, tree `content` holds plaintext hashes, `subtree` holds plaintext tree ID — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic: index/pack header “yields all plaintext hashes, types, offsets and lengths”, pack `id` in index is outer pack hash; `restic cat pack` verifies hash and warns on mismatch — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic data/keys layout: `data/XX/<full-hash>`, `snapshots/<id>`, `index/<id>`, `keys/<id>`, `locks/`, `tmp/`; snapshot filename is storage ID — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic attacker with read access can “Infer which packs probably contain trees via file access patterns”, “Infer the size of backups by using creation timestamps”, and pre-0.18.0 could derive chunker polynomial from observed chunk sizes per 2025 IACR paper; 0.18.0 mitigates by “randomly assigning chunks to pack files” — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Borg: “key = id = id_hash(unencrypted_data)”, `id_hash` is `sha256` (no keys) or `hmac-sha256` (with keys); must be cryptographically strong for dedup — [Source](https://borgbackup.readthedocs.io/en/1.0-maint/internals.html)
|
||||
- Borg 1.x: chunk ID `id = AUTHENTICATOR(id_key, data)` with independent `id_key` vs `enc_hmac_key`; decryption asserts `CONSTANT-TIME-COMPARISON(chunk-id, AUTHENTICATOR(id_key, decompressed))` — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg 2: “A chunk is considered duplicate if its id_hash value is identical”, `id_hash` e.g. “(hmac-)sha256 or (keyed) blake3”, selectable via `--id-hash=blake3` — [Source](https://github.com/borgbackup/borg)
|
||||
- Borg manifest has fixed ID `000…000`, anchored via TAM: `tam_key = HKDF-SHA-512(ikm=id_key||enc_key||enc_hmac_key, salt, "borg-metadata-authentication-manifest")`, `HMAC(tam_key, packed)` — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg fingerprinting: “repository does not hide size of chunks”, buzhash chunking uses “secret, random per-repo chunker seed”, small files <512 KiB yield single chunk; attacker with candidate files could brute-force fingerprint by sizes; mitigations: chunker choice/params, secret seed, compression choice, `obfuscate` padding; proximity/order (inode-order scan, segment adjacency) may leak additional info — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Kopia: “Block ID is generated by applying cryptographic hash function such as SHA2 or BLAKE2S”, identical blocks yield identical IDs for natural dedup; after hashing, block encrypted; multiple blocks merged into 20-40MB packs; index maps block ID -> (blob name, offset, length) — [Source](https://kopia.io/docs/advanced/architecture/)
|
||||
- Kopia: pack blob names random (e.g. `pb4cf8…`, `q7a99…`, `xn0_20db…`); content list/show via `kopia content list/show` requires decryption — [Source](https://kopia.io/docs/advanced/architecture/)
|
||||
- Kopia format `UniqueID` is 32 random bytes, also PBKDF salt and HKDF info, preventing cross-repo ID correlation — [Source](https://kopia.io/docs/advanced/encryption/)
|
||||
- Duplicity sets: “files in full backup sets will start with duplicity-full while the incremental sets start with duplicity-inc”, ordered patches; deleting full invalidates dependent incrementals — [Source](http://duplicity.nongnu.org/vers7/duplicity.1.html)
|
||||
- Duplicity S3 observer can tell “that you are using Duplicity, the name of the bucket, your AWS Access Key ID, the increment dates and the amount of data in each increment” (affects connection, not GPG payload) — [Source](https://linux.die.net/man/1/duplicity)
|
||||
- Duplicity default recipient key IDs visible in OpenPGP packets unless `--hidden-encrypt-key` (`--hidden-recipient`) used; restore then tries all secret keys — [Source](https://man.archlinux.org/man/duplicity.1.en)
|
||||
- Restic: “storage ID is the SHA-256 hash of the content of a file”, filename is “lower case hexadecimal representation of the storage ID”, verifiable by running `sha256sum` on file and comparing to filename - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic: “All content … is referenced according to its SHA-256 hash”, file split into CDC blobs, “SHA-256 hashes of all Blobs are saved in ordered list”, tree `content` holds plaintext hashes, `subtree` holds plaintext tree ID - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic: index/pack header “yields all plaintext hashes, types, offsets and lengths”, pack `id` in index is outer pack hash; `restic cat pack` verifies hash and warns on mismatch - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic data/keys layout: `data/XX/<full-hash>`, `snapshots/<id>`, `index/<id>`, `keys/<id>`, `locks/`, `tmp/`; snapshot filename is storage ID - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic attacker with read access can “Infer which packs probably contain trees via file access patterns”, “Infer the size of backups by using creation timestamps”, and pre-0.18.0 could derive chunker polynomial from observed chunk sizes per 2025 IACR paper; 0.18.0 mitigates by “randomly assigning chunks to pack files” - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Borg: “key = id = id_hash(unencrypted_data)”, `id_hash` is `sha256` (no keys) or `hmac-sha256` (with keys); must be cryptographically strong for dedup - [Source](https://borgbackup.readthedocs.io/en/1.0-maint/internals.html)
|
||||
- Borg 1.x: chunk ID `id = AUTHENTICATOR(id_key, data)` with independent `id_key` vs `enc_hmac_key`; decryption asserts `CONSTANT-TIME-COMPARISON(chunk-id, AUTHENTICATOR(id_key, decompressed))` - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg 2: “A chunk is considered duplicate if its id_hash value is identical”, `id_hash` e.g. “(hmac-)sha256 or (keyed) blake3”, selectable via `--id-hash=blake3` - [Source](https://github.com/borgbackup/borg)
|
||||
- Borg manifest has fixed ID `000…000`, anchored via TAM: `tam_key = HKDF-SHA-512(ikm=id_key||enc_key||enc_hmac_key, salt, "borg-metadata-authentication-manifest")`, `HMAC(tam_key, packed)` - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg fingerprinting: “repository does not hide size of chunks”, buzhash chunking uses “secret, random per-repo chunker seed”, small files <512 KiB yield single chunk; attacker with candidate files could brute-force fingerprint by sizes; mitigations: chunker choice/params, secret seed, compression choice, `obfuscate` padding; proximity/order (inode-order scan, segment adjacency) may leak additional info - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Kopia: “Block ID is generated by applying cryptographic hash function such as SHA2 or BLAKE2S”, identical blocks yield identical IDs for natural dedup; after hashing, block encrypted; multiple blocks merged into 20-40MB packs; index maps block ID -> (blob name, offset, length) - [Source](https://kopia.io/docs/advanced/architecture/)
|
||||
- Kopia: pack blob names random (e.g. `pb4cf8…`, `q7a99…`, `xn0_20db…`); content list/show via `kopia content list/show` requires decryption - [Source](https://kopia.io/docs/advanced/architecture/)
|
||||
- Kopia format `UniqueID` is 32 random bytes, also PBKDF salt and HKDF info, preventing cross-repo ID correlation - [Source](https://kopia.io/docs/advanced/encryption/)
|
||||
- Duplicity sets: “files in full backup sets will start with duplicity-full while the incremental sets start with duplicity-inc”, ordered patches; deleting full invalidates dependent incrementals - [Source](http://duplicity.nongnu.org/vers7/duplicity.1.html)
|
||||
- Duplicity S3 observer can tell “that you are using Duplicity, the name of the bucket, your AWS Access Key ID, the increment dates and the amount of data in each increment” (affects connection, not GPG payload) - [Source](https://linux.die.net/man/1/duplicity)
|
||||
- Duplicity default recipient key IDs visible in OpenPGP packets unless `--hidden-encrypt-key` (`--hidden-recipient`) used; restore then tries all secret keys - [Source](https://man.archlinux.org/man/duplicity.1.en)
|
||||
|
||||
### Inferences
|
||||
- Restic outer names (ciphertext hashes) look random to server; inner plaintext hashes never appear plaintext on server, so equality of files is hidden unless attacker correlates sizes/access patterns — unlike Borg/Kopia where HMAC IDs are server-visible keys enabling dedup-equality oracle.
|
||||
- Restic outer names (ciphertext hashes) look random to server; inner plaintext hashes never appear plaintext on server, so equality of files is hidden unless attacker correlates sizes/access patterns - unlike Borg/Kopia where HMAC IDs are server-visible keys enabling dedup-equality oracle.
|
||||
- Borg/Kopia HMAC-ID design intentionally reveals equality (required for server-side dedup) but not content; without `id_key`/`HMACSecret` server cannot confirm guesses except via size/proximity fingerprinting.
|
||||
- Duplicity leaks the most metadata: full vs incremental, chain order, timestamps and volume counts directly encode backup history even when payloads are opaque.
|
||||
|
||||
### Gaps
|
||||
- No reliable source confirming whether Restic pack filename is hash of ciphertext vs plaintext; doc says hash of content but pack verification suggests ciphertext — ambiguity noted.
|
||||
- No reliable source confirming whether Restic pack filename is hash of ciphertext vs plaintext; doc says hash of content but pack verification suggests ciphertext - ambiguity noted.
|
||||
- Kopia whether content IDs are plain hash vs HMAC-SHA256 with `HMACSecret` not fully resolved: architecture says hash, `ContentFormat.HMACSecret` implies keyed; exact construction needs source read.
|
||||
- Borg 2 object store layout (keys/ namespace, segment naming) not fetched; assumed similar to 1.x but with AEAD.
|
||||
|
||||
@@ -118,23 +118,23 @@ Restic outer filenames are SHA-256 of ciphertext (verifiable via sha256sum) whil
|
||||
Restic cites simplicity + Encrypt-then-MAC robustness and documents threat model + 2025 chunking-attack mitigation; Borg cites Encrypt-then-MAC robustness, Horton principle + TAM (CVE-2016-10099), and Borg 2 session-key/AEAD/Argon2 motivations; Kopia cites envelope encryption + per-content keys + random packing; Duplicity cites delegation to GnuPG/librsync. No Cure53-style audit report was found for these four in the searches; strongest external reviews found are Filippo Valsorda’s Restic note and Borg issue discussions.
|
||||
|
||||
### Cited Findings
|
||||
- Restic design: “Encryption is a first-class feature, the implementation looks sane and I guess the deduplication trade-off is worth it” — Filippo Valsorda, quoted in Restic encryption docs — [Source](https://github.com/restic/restic/blob/master/doc/070_encryption.rst)
|
||||
- Restic threat model: trusted client + authentic restic + secret password; guarantees vs assumptions listed; “Advances … against … (AES-256-CTR-Poly1305-AES and SHA-256) have not occurred”, brute-force infeasible, leaked key requires full re-encryption via `copy`/new repo — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic 2025 chunking attack: paper “Chunking Attacks on File Backup Services using Content-Defined Chunking” by Alexeev/Percival/Zhang, mitigated in 0.18.0 by random pack assignment; random irreducible CDC polynomial stored in encrypted `config` to harden watermark attacks — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic CDC: Rabin fingerprints, 64-byte window, files <512 KiB unsplit, 512 KiB–8 MiB blobs targeting 1 MiB average — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Borg: “Encryption is currently based on the Encrypt-then-MAC construction, which is generally seen as the most robust way” — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg: follows “Horton principle … not only the message must be authenticated, but also its meaning”, object ID MACs plaintext, parent reference assigns meaning, DAG anchored by TAM — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg TAM added since 1.0.9 for “Pre-1.0.9 manifest spoofing vulnerability (CVE-2016-10099)” — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg: “Borg does not support unauthenticated encryption — only authenticated encryption … No unauthenticated schemes will be added” — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg multi-client caveat: with multiple independent clients on same repo, “Borg fails to provide confidentiality” due to counter-reservation replay; trusted sync channel needed — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg compression+encryption discussion in issue #1040 concluded “no problem at all” to “hard and extremely slow to exploit”; user can disable compression — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg primitives rationale: HMAC-SHA256 vs BLAKE2b chosen on SHA hardware support (Ryzen, Intel 10th+ mobile/11th+ desktop, M1+, ARM64 SHA ext favour HMAC-SHA256; 64-bit CPUs without SHA ext favour BLAKE2b); uses OpenSSL libcrypto only, not libssl/TLS/X.509 — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg 2 rationale: global AES+MAC key requires perfect counter tracking across clients/threads (XOR-plaintext leak risk + sync complexity); session keys ease multithreading — [Source](https://github.com/borgbackup/borg/wiki/Borg-2.0)
|
||||
- Borg keyed chunkers (`toeplitz-aes/rabin-aes/goldilocks-aes`) rationale: “make chunk-boundary fingerprinting attacks much harder” — [Source](https://github.com/borgbackup/borg)
|
||||
- Kopia rationale: “standard envelope encryption technique to de-couple the repository passphrase from keys used for encrypting/authenticating contents”; format blob holds params, `UniqueID` random per repo, scrypt PBKDF + HKDF-SHA256 for Ke/AD — [Source](https://kopia.io/docs/advanced/encryption/)
|
||||
- Kopia FAQ: encryption mandatory, algorithm fixed at creation, password unrecoverable — [Source](https://github.com/kopia/kopia/blob/master/site/content/docs/FAQs/_index.md)
|
||||
- Duplicity rationale: uses librsync for space-efficient incrementals + GnuPG so “they will be safe from spying and/or modification by the server” — [Source](https://packages.debian.org/bullseye/arm64/utils/duplicity)
|
||||
- Duplicity signing+symmetric-encrypt via CLI gpg noted as “specifically challenging”, only certain passphrase/agent combos tested working — [Source](http://duplicity.nongnu.org/vers7/duplicity.1.html)
|
||||
- Restic design: “Encryption is a first-class feature, the implementation looks sane and I guess the deduplication trade-off is worth it” - Filippo Valsorda, quoted in Restic encryption docs - [Source](https://github.com/restic/restic/blob/master/doc/070_encryption.rst)
|
||||
- Restic threat model: trusted client + authentic restic + secret password; guarantees vs assumptions listed; “Advances … against … (AES-256-CTR-Poly1305-AES and SHA-256) have not occurred”, brute-force infeasible, leaked key requires full re-encryption via `copy`/new repo - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic 2025 chunking attack: paper “Chunking Attacks on File Backup Services using Content-Defined Chunking” by Alexeev/Percival/Zhang, mitigated in 0.18.0 by random pack assignment; random irreducible CDC polynomial stored in encrypted `config` to harden watermark attacks - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic CDC: Rabin fingerprints, 64-byte window, files <512 KiB unsplit, 512 KiB–8 MiB blobs targeting 1 MiB average - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Borg: “Encryption is currently based on the Encrypt-then-MAC construction, which is generally seen as the most robust way” - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg: follows “Horton principle … not only the message must be authenticated, but also its meaning”, object ID MACs plaintext, parent reference assigns meaning, DAG anchored by TAM - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg TAM added since 1.0.9 for “Pre-1.0.9 manifest spoofing vulnerability (CVE-2016-10099)” - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg: “Borg does not support unauthenticated encryption - only authenticated encryption … No unauthenticated schemes will be added” - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg multi-client caveat: with multiple independent clients on same repo, “Borg fails to provide confidentiality” due to counter-reservation replay; trusted sync channel needed - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg compression+encryption discussion in issue #1040 concluded “no problem at all” to “hard and extremely slow to exploit”; user can disable compression - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg primitives rationale: HMAC-SHA256 vs BLAKE2b chosen on SHA hardware support (Ryzen, Intel 10th+ mobile/11th+ desktop, M1+, ARM64 SHA ext favour HMAC-SHA256; 64-bit CPUs without SHA ext favour BLAKE2b); uses OpenSSL libcrypto only, not libssl/TLS/X.509 - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg 2 rationale: global AES+MAC key requires perfect counter tracking across clients/threads (XOR-plaintext leak risk + sync complexity); session keys ease multithreading - [Source](https://github.com/borgbackup/borg/wiki/Borg-2.0)
|
||||
- Borg keyed chunkers (`toeplitz-aes/rabin-aes/goldilocks-aes`) rationale: “make chunk-boundary fingerprinting attacks much harder” - [Source](https://github.com/borgbackup/borg)
|
||||
- Kopia rationale: “standard envelope encryption technique to de-couple the repository passphrase from keys used for encrypting/authenticating contents”; format blob holds params, `UniqueID` random per repo, scrypt PBKDF + HKDF-SHA256 for Ke/AD - [Source](https://kopia.io/docs/advanced/encryption/)
|
||||
- Kopia FAQ: encryption mandatory, algorithm fixed at creation, password unrecoverable - [Source](https://github.com/kopia/kopia/blob/master/site/content/docs/FAQs/_index.md)
|
||||
- Duplicity rationale: uses librsync for space-efficient incrementals + GnuPG so “they will be safe from spying and/or modification by the server” - [Source](https://packages.debian.org/bullseye/arm64/utils/duplicity)
|
||||
- Duplicity signing+symmetric-encrypt via CLI gpg noted as “specifically challenging”, only certain passphrase/agent combos tested working - [Source](http://duplicity.nongnu.org/vers7/duplicity.1.html)
|
||||
|
||||
### Inferences
|
||||
- Restic/Borg explicitly prefer Encrypt-then-MAC over AEAD for 1.x-era compatibility and library constraints (OpenSSL libcrypto, Go stdlib); Borg 2 and Kopia converge on modern AEAD (OCB/GCM/ChaCha20-Poly1305) once widely available.
|
||||
|
||||
@@ -6,24 +6,24 @@
|
||||
All five systems use password-derived envelope encryption with a random master/data key, but KDFs and parameters differ: restic pins scrypt N=65536/r=8/p=1 (auto-calibrated on newer adds); Borg 1.x uses PBKDF2-HMAC-SHA256 while Borg 2.x defaults to Argon2id + ChaCha20-Poly1305; Kopia uses scrypt-65536-8-1 over UniqueID salt plus HKDF-SHA256 subkeys (PBKDF2/scrypt configurability added ~2025); rclone crypt uses scrypt N=16384/r=8/p=1; Tarsnap generates keys locally and only uses scrypt to passphrase-wrap the key file.
|
||||
|
||||
### Cited Findings
|
||||
- restic: all repo data encrypted with AES-256-CTR + Poly1305-AES MAC as `IV || CIPHERTEXT || MAC` (32 bytes overhead, random 16-byte nonce per file); three keys needed (32-byte AES-256 key + 16-byte AES key + 16-byte Poly1305 key) — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- restic: on open, password + per-key `N`, `r`, `p`, `salt` feed scrypt to derive 64 bytes; first 32 = AES-256 encryption key, last 32 split into 16-byte AES key `k` + 16-byte Poly1305 key `r` (masked); these decrypt the `data` field to reveal the master key JSON — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- restic: canonical key-file example pins `"kdf": "scrypt", "N": 65536, "r": 8, "p": 1`; code rejects any KDF other than scrypt (`only supported KDF is scrypt()`) — [Source](https://github.com/restic/restic/blob/9e2d60e2/internal/repository/key.go)
|
||||
- restic: `AddKey` calibrates KDF parameters (`crypto.Calibrate(KDFTimeout, KDFMemory)`) when adding keys, generates a random salt (`crypto.NewSalt()`) and either a fresh random master key (`crypto.NewRandomKey()`) or copies the existing master key template — [Source](https://github.com/restic/restic/blob/9e2d60e2/internal/repository/key.go)
|
||||
- Borg 1.x/legacy: key file encrypted with PBKDF2-HMAC-SHA256 (32-byte salt, `PBKDF2_ITERATIONS`), AES-CTR with zero IV plus HMAC-SHA256 encrypt-and-MAC construction — [Source](https://github.com/borgbackup/borg/blob/86fd77fd/src/borg/legacy/crypto/key.py)
|
||||
- Borg 2.x: 256-bit key-encryption key (KEK) derived from passphrase with Argon2 + random 256-bit salt, then Encrypt-then-MAC of packed key material with ChaCha20-Poly1305 AEAD and constant IV==0 (safe because salt makes KEK unique per encryption) — [Source](https://borgbackup.readthedocs.io/en/2.0.0b9/internals/security.html)
|
||||
- Borg: new key files support `argon2 chacha20-poly1305` vs legacy `sha256` (PBKDF2) algorithms; Argon2 path derives 32-byte KEK via `argon2.low_level.hash_secret_raw` with configurable time/memory/parallelism/type — [Source](https://github.com/borgbackup/borg/blob/da3105f1/src/borg/crypto/key.py)
|
||||
- Borg: `--key-algorithm argon2` is default (Argon2id); `pbkdf2` kept for old-client compatibility; docs note Argon2 path also fixes two issues at once (separate encrypt vs MAC keys, encrypt-then-MAC instead of encrypt-and-MAC) — [Source](https://git.uninsane.org/shelvacu-mirrors/borg/commit/08f82ee40867f605ca6994db6dc32218d7b85cbc)
|
||||
- Borg: random Borg key itself consists of three random secrets (crypt key, id key, chunker seed); passphrase only locks/encrypts this key, chunking and IDs also derive from it — [Source](https://borgbackup.readthedocs.io/en/stable/usage/init.html)
|
||||
- Kopia: envelope encryption; format blob carries `uniqueID` (random 32 bytes), `keyAlgo` (e.g. `scrypt-65536-8-1`), `encryption: AES256_GCM`, and `encryptedBlockFormat` holding the real content-encryption secrets (`HMACSecret`, `MasterKey`) — [Source](https://kopia.io/docs/advanced/encryption/)
|
||||
- Kopia: `Km = PBKDF(passphrase, UniqueID, cost params)` (32 bytes) → `Ke = HKDF(SHA256, Km, UniqueID, "AES", 32)` and `AD = HKDF(SHA256, Km, UniqueID, "CHECKSUM", 32)`; format block encrypted with AES256-GCM using Ke, random IV prepended, AD authenticated — [Source](https://kopia.io/docs/advanced/encryption/)
|
||||
- Kopia: only scrypt N=65536/r=8/p=1 documented as supported ("at the moment"), field reserved for future algorithm/cost changes — [Source](https://kopia.io/docs/advanced/encryption/)
|
||||
- Kopia: 2025 PR adds configurable KDF: `--pbkdf (pbkdf2|scrypt)`, `--pbkdf-iter` (default 600000 for PBKDF2), `--pbkdf-memory` (default 64 MB for scrypt), usable at `repo create` and `change-password` — [Source](https://github.com/kopia/kopia/pull/5145)
|
||||
- rclone crypt: file content uses NaCl SecretBox (XSalsa20 + Poly1305), 64 KiB chunks each with 16-byte authenticator, header = 8-byte magic `RCLONE\x00\x00` + 24-byte random nonce incremented per chunk — [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt: filename segments PKCS#7-padded to 16 bytes, encrypted deterministically with EME (ECB-Mix-ECB) AES-256, emitted as lowercase unpadded base32 — [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt: 80 bytes of key material derived via scrypt `N=16384, r=8, p=1` from password (+ optional `password2` salt); without user salt a built-in internal salt is used — [Source](https://rclone.org/crypt/)
|
||||
- Tarsnap: key files hold authentication keys (server access proofs) and encryption keys (archive encrypt/sign/verify/decrypt) separately, so server compromise does not disclose data — [Source](https://www.tarsnap.com/security.html)
|
||||
- Tarsnap: no password-derived master key by default; keys generated by `tarsnap-keygen`; passphrase protection is optional (`--passphrased`) and wraps the key file with keys from the scrypt KDF, with tunable `--passphrase-mem` / `--passphrase-time` — [Source](https://www.tarsnap.com/man-tarsnap-keygen.1.html)
|
||||
- restic: all repo data encrypted with AES-256-CTR + Poly1305-AES MAC as `IV || CIPHERTEXT || MAC` (32 bytes overhead, random 16-byte nonce per file); three keys needed (32-byte AES-256 key + 16-byte AES key + 16-byte Poly1305 key) - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- restic: on open, password + per-key `N`, `r`, `p`, `salt` feed scrypt to derive 64 bytes; first 32 = AES-256 encryption key, last 32 split into 16-byte AES key `k` + 16-byte Poly1305 key `r` (masked); these decrypt the `data` field to reveal the master key JSON - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- restic: canonical key-file example pins `"kdf": "scrypt", "N": 65536, "r": 8, "p": 1`; code rejects any KDF other than scrypt (`only supported KDF is scrypt()`) - [Source](https://github.com/restic/restic/blob/9e2d60e2/internal/repository/key.go)
|
||||
- restic: `AddKey` calibrates KDF parameters (`crypto.Calibrate(KDFTimeout, KDFMemory)`) when adding keys, generates a random salt (`crypto.NewSalt()`) and either a fresh random master key (`crypto.NewRandomKey()`) or copies the existing master key template - [Source](https://github.com/restic/restic/blob/9e2d60e2/internal/repository/key.go)
|
||||
- Borg 1.x/legacy: key file encrypted with PBKDF2-HMAC-SHA256 (32-byte salt, `PBKDF2_ITERATIONS`), AES-CTR with zero IV plus HMAC-SHA256 encrypt-and-MAC construction - [Source](https://github.com/borgbackup/borg/blob/86fd77fd/src/borg/legacy/crypto/key.py)
|
||||
- Borg 2.x: 256-bit key-encryption key (KEK) derived from passphrase with Argon2 + random 256-bit salt, then Encrypt-then-MAC of packed key material with ChaCha20-Poly1305 AEAD and constant IV==0 (safe because salt makes KEK unique per encryption) - [Source](https://borgbackup.readthedocs.io/en/2.0.0b9/internals/security.html)
|
||||
- Borg: new key files support `argon2 chacha20-poly1305` vs legacy `sha256` (PBKDF2) algorithms; Argon2 path derives 32-byte KEK via `argon2.low_level.hash_secret_raw` with configurable time/memory/parallelism/type - [Source](https://github.com/borgbackup/borg/blob/da3105f1/src/borg/crypto/key.py)
|
||||
- Borg: `--key-algorithm argon2` is default (Argon2id); `pbkdf2` kept for old-client compatibility; docs note Argon2 path also fixes two issues at once (separate encrypt vs MAC keys, encrypt-then-MAC instead of encrypt-and-MAC) - [Source](https://git.uninsane.org/shelvacu-mirrors/borg/commit/08f82ee40867f605ca6994db6dc32218d7b85cbc)
|
||||
- Borg: random Borg key itself consists of three random secrets (crypt key, id key, chunker seed); passphrase only locks/encrypts this key, chunking and IDs also derive from it - [Source](https://borgbackup.readthedocs.io/en/stable/usage/init.html)
|
||||
- Kopia: envelope encryption; format blob carries `uniqueID` (random 32 bytes), `keyAlgo` (e.g. `scrypt-65536-8-1`), `encryption: AES256_GCM`, and `encryptedBlockFormat` holding the real content-encryption secrets (`HMACSecret`, `MasterKey`) - [Source](https://kopia.io/docs/advanced/encryption/)
|
||||
- Kopia: `Km = PBKDF(passphrase, UniqueID, cost params)` (32 bytes) → `Ke = HKDF(SHA256, Km, UniqueID, "AES", 32)` and `AD = HKDF(SHA256, Km, UniqueID, "CHECKSUM", 32)`; format block encrypted with AES256-GCM using Ke, random IV prepended, AD authenticated - [Source](https://kopia.io/docs/advanced/encryption/)
|
||||
- Kopia: only scrypt N=65536/r=8/p=1 documented as supported ("at the moment"), field reserved for future algorithm/cost changes - [Source](https://kopia.io/docs/advanced/encryption/)
|
||||
- Kopia: 2025 PR adds configurable KDF: `--pbkdf (pbkdf2|scrypt)`, `--pbkdf-iter` (default 600000 for PBKDF2), `--pbkdf-memory` (default 64 MB for scrypt), usable at `repo create` and `change-password` - [Source](https://github.com/kopia/kopia/pull/5145)
|
||||
- rclone crypt: file content uses NaCl SecretBox (XSalsa20 + Poly1305), 64 KiB chunks each with 16-byte authenticator, header = 8-byte magic `RCLONE\x00\x00` + 24-byte random nonce incremented per chunk - [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt: filename segments PKCS#7-padded to 16 bytes, encrypted deterministically with EME (ECB-Mix-ECB) AES-256, emitted as lowercase unpadded base32 - [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt: 80 bytes of key material derived via scrypt `N=16384, r=8, p=1` from password (+ optional `password2` salt); without user salt a built-in internal salt is used - [Source](https://rclone.org/crypt/)
|
||||
- Tarsnap: key files hold authentication keys (server access proofs) and encryption keys (archive encrypt/sign/verify/decrypt) separately, so server compromise does not disclose data - [Source](https://www.tarsnap.com/security.html)
|
||||
- Tarsnap: no password-derived master key by default; keys generated by `tarsnap-keygen`; passphrase protection is optional (`--passphrased`) and wraps the key file with keys from the scrypt KDF, with tunable `--passphrase-mem` / `--passphrase-time` - [Source](https://www.tarsnap.com/man-tarsnap-keygen.1.html)
|
||||
|
||||
### Inferences
|
||||
- restic/Kopia/rclone all standardize on scrypt but with 4x cost difference (16384 vs 65536), so rclone password guessing is cheaper; Borg's Argon2id move is the most modern KDF posture.
|
||||
@@ -40,18 +40,18 @@ All five systems use password-derived envelope encryption with a random master/d
|
||||
restic and Kopia keep the (password-locked) key material inside the repository; Borg offers repokey (in repo) vs keyfile (client `~/.config/borg/keys`) as a first-class choice; Tarsnap and rclone keep keys strictly client-side (key file / `rclone.conf`), making client-side backup mandatory. None document TPM; OS-credential integration exists only for Kopia (server/connect password caching) and via external helpers for the rest.
|
||||
|
||||
### Cited Findings
|
||||
- restic: `keys/` directory in repo holds JSON key files (hostname, username, created, kdf params, salt, encrypted `data`); `restic cat masterkey` decrypts and pretty-prints master encryption/MAC keys — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- restic: passwords supplied via prompt, `--password-file`, or `--password-command` (`$RESTIC_PASSWORD_FILE` / `$RESTIC_PASSWORD_COMMAND`); empty passwords refused by default, `--insecure-no-password` required since 0.17.0 — [Source](https://man.archlinux.org/man/restic-key-add.1.en)
|
||||
- Borg 2.x: repokey modes store encrypted key in repo (`repo_dir/config`); keyfile modes store in home dir (`~/.config/borg/keys`); `--key-location` repokey|keyfile chosen at `rcreate`, movable later via `key change-location` — [Source](https://borgbackup.readthedocs.io/en/2.0.0b6/usage/rcreate.html)
|
||||
- Borg 1.x: identical split — repokey inside repo directory, keyfile in `~/.config/borg/keys` (macOS: `~/Library/Application Support/borg/keys`); remote repos via ssh never see passphrase, plaintext key, or plaintext files — [Source](https://borgbackup.readthedocs.io/en/master/usage/repo-create.html)
|
||||
- Borg export UX: `borg key export [PATH]`, `--paper` (printable, per-line checksums for type-in), `--qr-html` (QR + paper copy); export stays passphrase-encrypted (no passphrase included); `borg key import [--paper]` restores; paper-key web tool at `paperkey.html` — [Source](https://borgbackup.readthedocs.io/en/master/usage/key.html); [Source](https://borgbackup.readthedocs.io/en/stable/paperkey.html)
|
||||
- Borg: for keyfile repos the key must be backed up independently ("NOT sufficient" to keep copy on the backed-up system); for repokey a backup is "not strictly needed" but guards against corruption/loss of the in-repo key — [Source](https://man.archlinux.org/man/borg-key-export.1.en)
|
||||
- Kopia: format blob (`kopia.repository` / `kopia.blobcfg` objects) lives in storage; local `repository-*.config` holds connection parameters (not the password); quick-reconnect token via `kopia repository status -t` (opaque) and `-s` embeds password (trivially decodable, must be guarded) — [Source](https://kopia.io/docs/reference/command-line/)
|
||||
- Kopia: repository password cached in OS-specific credential storage (Keychain on macOS, Credential Manager on Windows, Keyring on Linux) — [Source](https://kopia.io/docs/reference/command-line/)
|
||||
- rclone crypt: password + optional salt (`password2`) stored in `rclone.conf` in lightly obscured form (AES-CTR with static shared key, random IV prepended) — explicitly "not secure" without overall config-file encryption (`rclone config` password protection); env vars `RCLONE_CRYPT_PASSWORD` / `RCLONE_CRYPT_PASSWORD2` (must be `rclone obscure`d) — [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt: recovery = remember password or keep `rclone.conf`; same passwords re-entered on another machine reproduce access (obscured strings differ due to salt) — [Source](https://rclone.org/crypt/)
|
||||
- Tarsnap: single key file from `tarsnap-keygen --keyfile <file> --user <email> --machine <name>`; every `tarsnap` invocation needs `--keyfile`; passphrase via `--passphrase method:arg` (`dev:tty-stdin` default, `env:VAR`, `file:FILENAME` flagged as risky) — [Source](https://www.tarsnap.com/man-tarsnap-keygen.1.html); [Source](https://man.archlinux.org/man/tarsnap.1.en)
|
||||
- Tarsnap: `tarsnap-keymgmt --outkeyfile <new> [-r] [-w] [-d] [--nuke]` mints restricted sub-keys (read-only, write-only, delete, nuke-only); `-d` implies `-r` — [Source](https://www.tarsnap.com/man-tarsnap-keymgmt.1.html)
|
||||
- restic: `keys/` directory in repo holds JSON key files (hostname, username, created, kdf params, salt, encrypted `data`); `restic cat masterkey` decrypts and pretty-prints master encryption/MAC keys - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- restic: passwords supplied via prompt, `--password-file`, or `--password-command` (`$RESTIC_PASSWORD_FILE` / `$RESTIC_PASSWORD_COMMAND`); empty passwords refused by default, `--insecure-no-password` required since 0.17.0 - [Source](https://man.archlinux.org/man/restic-key-add.1.en)
|
||||
- Borg 2.x: repokey modes store encrypted key in repo (`repo_dir/config`); keyfile modes store in home dir (`~/.config/borg/keys`); `--key-location` repokey|keyfile chosen at `rcreate`, movable later via `key change-location` - [Source](https://borgbackup.readthedocs.io/en/2.0.0b6/usage/rcreate.html)
|
||||
- Borg 1.x: identical split - repokey inside repo directory, keyfile in `~/.config/borg/keys` (macOS: `~/Library/Application Support/borg/keys`); remote repos via ssh never see passphrase, plaintext key, or plaintext files - [Source](https://borgbackup.readthedocs.io/en/master/usage/repo-create.html)
|
||||
- Borg export UX: `borg key export [PATH]`, `--paper` (printable, per-line checksums for type-in), `--qr-html` (QR + paper copy); export stays passphrase-encrypted (no passphrase included); `borg key import [--paper]` restores; paper-key web tool at `paperkey.html` - [Source](https://borgbackup.readthedocs.io/en/master/usage/key.html); [Source](https://borgbackup.readthedocs.io/en/stable/paperkey.html)
|
||||
- Borg: for keyfile repos the key must be backed up independently ("NOT sufficient" to keep copy on the backed-up system); for repokey a backup is "not strictly needed" but guards against corruption/loss of the in-repo key - [Source](https://man.archlinux.org/man/borg-key-export.1.en)
|
||||
- Kopia: format blob (`kopia.repository` / `kopia.blobcfg` objects) lives in storage; local `repository-*.config` holds connection parameters (not the password); quick-reconnect token via `kopia repository status -t` (opaque) and `-s` embeds password (trivially decodable, must be guarded) - [Source](https://kopia.io/docs/reference/command-line/)
|
||||
- Kopia: repository password cached in OS-specific credential storage (Keychain on macOS, Credential Manager on Windows, Keyring on Linux) - [Source](https://kopia.io/docs/reference/command-line/)
|
||||
- rclone crypt: password + optional salt (`password2`) stored in `rclone.conf` in lightly obscured form (AES-CTR with static shared key, random IV prepended) - explicitly "not secure" without overall config-file encryption (`rclone config` password protection); env vars `RCLONE_CRYPT_PASSWORD` / `RCLONE_CRYPT_PASSWORD2` (must be `rclone obscure`d) - [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt: recovery = remember password or keep `rclone.conf`; same passwords re-entered on another machine reproduce access (obscured strings differ due to salt) - [Source](https://rclone.org/crypt/)
|
||||
- Tarsnap: single key file from `tarsnap-keygen --keyfile <file> --user <email> --machine <name>`; every `tarsnap` invocation needs `--keyfile`; passphrase via `--passphrase method:arg` (`dev:tty-stdin` default, `env:VAR`, `file:FILENAME` flagged as risky) - [Source](https://www.tarsnap.com/man-tarsnap-keygen.1.html); [Source](https://man.archlinux.org/man/tarsnap.1.en)
|
||||
- Tarsnap: `tarsnap-keymgmt --outkeyfile <new> [-r] [-w] [-d] [--nuke]` mints restricted sub-keys (read-only, write-only, delete, nuke-only); `-d` implies `-r` - [Source](https://www.tarsnap.com/man-tarsnap-keymgmt.1.html)
|
||||
|
||||
### Inferences
|
||||
- Only Borg gives a deliberate two-factor-style choice (possession of keyfile + knowledge of passphrase); restic/Kopia/rclone-default are passphrase-only if the repo/config is exfiltrated.
|
||||
@@ -67,44 +67,44 @@ restic and Kopia keep the (password-locked) key material inside the repository;
|
||||
restic and Borg (2.x/master) support multiple concurrent keys/passphrases sharing one master secret; Kopia and rclone crypt support only password change (re-wrapping the same secrets), with rclone requiring full re-upload; Tarsnap supports restricted sub-keys but not multi-passphrase unlock. No system documents Shamir secret sharing natively.
|
||||
|
||||
### Cited Findings
|
||||
- restic: `key` command with `list`, `add`, `remove`, `passwd` subcommands manages multiple access keys per repo; `key add` prompts for current password then new password and saves a new key wrapping the same master key; list marks current key with `*` — [Source](https://restic.readthedocs.io/en/stable/070_encryption.html)
|
||||
- restic: `key passwd` creates a new key ID and removes the old one; `key remove <ID>` refuses to remove the currently-used key — [Source](https://manpages.opensuse.org/Tumbleweed/restic/restic-key-passwd.1.en.html); [Source](https://man.archlinux.org/man/restic-key-remove.1.en.raw)
|
||||
- restic: `AddKey` with non-nil `template` copies master keys from the old key instead of generating new ones, confirming rotation re-wraps rather than re-encrypts data — [Source](https://github.com/restic/restic/blob/9e2d60e2/internal/repository/key.go)
|
||||
- Borg: `key change-passphrase` only re-locks the same secrets ("does not protect future nor past backups" if key+passphrase were compromised) — [Source](https://borgbackup.readthedocs.io/en/stable/usage/key.html)
|
||||
- Borg: `key change-algorithm argon2|pbkdf2` upgrades/downgrades the KDF wrapping without changing secrets — [Source](https://git.uninsane.org/shelvacu-mirrors/borg/commit/08f82ee40867f605ca6994db6dc32218d7b85cbc)
|
||||
- Borg (master/2.x): multiple borg keys per repo — each key holds the same secret material under an independent passphrase and label (first key labeled `admin`, protected from removal); `key list/add/remove/export --label|--key` selectors; passphrase tried against every key — [Source](https://github.com/borgbackup/borg/pull/9762)
|
||||
- Borg: `rcreate --other-repo SRC --copy-crypt-key` can reuse crypt key across related repos (default: fresh random crypt key, shared chunker/ID keys for dedup) — [Source](https://borgbackup.readthedocs.io/en/2.0.0b6/usage/rcreate.html)
|
||||
- Kopia: `kopia repository change-password` (CLI only, not GUI) re-encrypts the format block with a new password; must already be connected (so a still-connected client can reset a forgotten password, a disconnected one cannot) — [Source](https://kopia.io/docs/faqs/); [Source](https://kopia.io/docs/reference/command-line/common/repository-change-password/)
|
||||
- rclone crypt: no in-place password change — changing the configured password orphans existing content; must re-upload everything (in place via second crypt remote + `rclone copy` decrypting/re-encrypting on the fly, at 2x bandwidth/quota cost) — [Source](https://rclone.org/crypt/)
|
||||
- Tarsnap: no multi-passphrase unlock or rotation of archive encryption keys documented; capability separation via `tarsnap-keymgmt` restricted keys is the only delegation mechanism — [Source](https://www.tarsnap.com/man-tarsnap-keymgmt.1.html)
|
||||
- Shamir/shared recovery: no hits in official docs for restic, Borg, Kopia, rclone crypt, or Tarsnap; only community workarounds (splitting password/paper key with external SSS tools) exist — no primary-source citation available.
|
||||
- restic: `key` command with `list`, `add`, `remove`, `passwd` subcommands manages multiple access keys per repo; `key add` prompts for current password then new password and saves a new key wrapping the same master key; list marks current key with `*` - [Source](https://restic.readthedocs.io/en/stable/070_encryption.html)
|
||||
- restic: `key passwd` creates a new key ID and removes the old one; `key remove <ID>` refuses to remove the currently-used key - [Source](https://manpages.opensuse.org/Tumbleweed/restic/restic-key-passwd.1.en.html); [Source](https://man.archlinux.org/man/restic-key-remove.1.en.raw)
|
||||
- restic: `AddKey` with non-nil `template` copies master keys from the old key instead of generating new ones, confirming rotation re-wraps rather than re-encrypts data - [Source](https://github.com/restic/restic/blob/9e2d60e2/internal/repository/key.go)
|
||||
- Borg: `key change-passphrase` only re-locks the same secrets ("does not protect future nor past backups" if key+passphrase were compromised) - [Source](https://borgbackup.readthedocs.io/en/stable/usage/key.html)
|
||||
- Borg: `key change-algorithm argon2|pbkdf2` upgrades/downgrades the KDF wrapping without changing secrets - [Source](https://git.uninsane.org/shelvacu-mirrors/borg/commit/08f82ee40867f605ca6994db6dc32218d7b85cbc)
|
||||
- Borg (master/2.x): multiple borg keys per repo - each key holds the same secret material under an independent passphrase and label (first key labeled `admin`, protected from removal); `key list/add/remove/export --label|--key` selectors; passphrase tried against every key - [Source](https://github.com/borgbackup/borg/pull/9762)
|
||||
- Borg: `rcreate --other-repo SRC --copy-crypt-key` can reuse crypt key across related repos (default: fresh random crypt key, shared chunker/ID keys for dedup) - [Source](https://borgbackup.readthedocs.io/en/2.0.0b6/usage/rcreate.html)
|
||||
- Kopia: `kopia repository change-password` (CLI only, not GUI) re-encrypts the format block with a new password; must already be connected (so a still-connected client can reset a forgotten password, a disconnected one cannot) - [Source](https://kopia.io/docs/faqs/); [Source](https://kopia.io/docs/reference/command-line/common/repository-change-password/)
|
||||
- rclone crypt: no in-place password change - changing the configured password orphans existing content; must re-upload everything (in place via second crypt remote + `rclone copy` decrypting/re-encrypting on the fly, at 2x bandwidth/quota cost) - [Source](https://rclone.org/crypt/)
|
||||
- Tarsnap: no multi-passphrase unlock or rotation of archive encryption keys documented; capability separation via `tarsnap-keymgmt` restricted keys is the only delegation mechanism - [Source](https://www.tarsnap.com/man-tarsnap-keymgmt.1.html)
|
||||
- Shamir/shared recovery: no hits in official docs for restic, Borg, Kopia, rclone crypt, or Tarsnap; only community workarounds (splitting password/paper key with external SSS tools) exist - no primary-source citation available.
|
||||
|
||||
### Inferences
|
||||
- True cryptographic rotation (new data key, re-encryption) is offered by none of the five for existing snapshots; all "rotation" is re-wrapping or passphrase change.
|
||||
- Borg's multi-key design is closest to shared/admin recovery (per-user passphrases + protected admin key), but all keys still share one secret — compromise of the secret compromises all slots.
|
||||
- Borg's multi-key design is closest to shared/admin recovery (per-user passphrases + protected admin key), but all keys still share one secret - compromise of the secret compromises all slots.
|
||||
|
||||
### Gaps
|
||||
- Whether `borg key add` (multi-key) is released in stable 2.0 vs master-only could not be pinned down; primary evidence is the PR plus master `usage/key.html`.
|
||||
- No documented Shamir/threshold scheme in any official docs; confirm as absent-by-design vs merely undocumented.
|
||||
|
||||
## What happens on key loss — is data unrecoverable by design? Any documented footguns?
|
||||
## What happens on key loss - is data unrecoverable by design? Any documented footguns?
|
||||
|
||||
### Takeaway
|
||||
All five declare password/key loss unrecoverable by design. The sharpest footguns: restic `key remove` of the sole key and deleted/corrupt `keys/` files; Borg keyfile loss without export and repokey-with-empty-passphrase on exposed storage; Kopia disconnected-password loss; rclone salt (`password2`) loss and obscured-config confusion; Tarsnap key-file loss (including billing trap).
|
||||
|
||||
### Cited Findings
|
||||
- restic docs repeat on every backend page: "knowledge of your password is required… Losing your password means that your data is irrecoverably lost" — [Source](https://restic.readthedocs.io/en/stable/030%5Fpreparing%5Fa%5Fnew%5Frepo.html)
|
||||
- restic forum (maintainer fd0): if key-file `data` or `salt` is missing/corrupt there is "no way to decrypt the data again. Even if you know the password"; intact `data`+`salt` (+N/r/p, brute-forceable) is required — [Source](https://forum.restic.net/t/recover-a-damaged-missing-corrupted-key-file/1798)
|
||||
- restic footgun case: user generated a random key file, ran `key remove` on the old key, then lost the random file (only copy was inside the backup) — recovery required disk forensics for the deleted key file; `key remove` is just a file deletion in `keys/` — [Source](https://forum.restic.net/t/recover-from-a-previous-password/2592)
|
||||
- restic: brute force only viable if much of the password structure is known; otherwise "consider it lost" — [Source](https://forum.restic.net/t/forgotten-password/2990)
|
||||
- Borg: "always need both the Borg key and passphrase"; keyfile loss = repo loss ("if you lose the key, you lose access"); mandated offsite export ("leaving your keys inside your car" warning); empty passphrase with repokey = anyone reading the repo can unlock (as good as no encryption); keyfile+empty passphrase acceptable only if client disk is encrypted — [Source](https://borgbackup.readthedocs.io/en/stable/usage/init.html)
|
||||
- Borg: exported key stays encrypted — recovery needs both export file AND original passphrase, stored in separate safe places — [Source](https://borgbackup.readthedocs.io/en/stable/usage/key.html)
|
||||
- Kopia: "you cannot restore your files if you forget your password, there is no way to recover a forgotten password because only you know it"; FAQ: "There is no way to recover it or the files and folders within that repository. Store your repository password in a safe place, such as a password manager" — [Source](https://github.com/kopia/kopia/blob/master/site/content/docs/Features/_index.md); [Source](https://kopia.io/docs/faqs/)
|
||||
- Kopia footgun nuance: password CAN be reset while still connected (`change-password`), which rescues forgotten-but-connected GUI users via CLI, but a disconnected client with a forgotten password is unrecoverable — [Source](https://kopia.io/docs/faqs/)
|
||||
- Kopia reconnect token (`status -t -s`) trivially decodes to the password — storing it insecurely leaks the repo — [Source](https://kopia.io/docs/reference/command-line/)
|
||||
- rclone crypt: custom salt is "effectively a second password that must be memorized" and is NOT stored with the data; losing `password2` loses data even with correct `password` — [Source](https://rclone.org/crypt/)
|
||||
- rclone footguns: obscured passwords look encrypted but use a static shared AES-CTR key (cursory-inspection only); 1.49.0–1.53.2 random-password generator bug produced insecure passwords (fixed 1.53.3, must rotate by re-upload) — [Source](https://rclone.org/crypt/)
|
||||
- Tarsnap FAQ: "You can't. Your key file contains the only copy of the cryptographic keys… if you lose them there is no way to get your data back" (same for forgotten key-file passphrase); separate FAQ covers being unable to stop billing after losing all keys — [Source](https://www.tarsnap.com/faq.html)
|
||||
- restic docs repeat on every backend page: "knowledge of your password is required… Losing your password means that your data is irrecoverably lost" - [Source](https://restic.readthedocs.io/en/stable/030%5Fpreparing%5Fa%5Fnew%5Frepo.html)
|
||||
- restic forum (maintainer fd0): if key-file `data` or `salt` is missing/corrupt there is "no way to decrypt the data again. Even if you know the password"; intact `data`+`salt` (+N/r/p, brute-forceable) is required - [Source](https://forum.restic.net/t/recover-a-damaged-missing-corrupted-key-file/1798)
|
||||
- restic footgun case: user generated a random key file, ran `key remove` on the old key, then lost the random file (only copy was inside the backup) - recovery required disk forensics for the deleted key file; `key remove` is just a file deletion in `keys/` - [Source](https://forum.restic.net/t/recover-from-a-previous-password/2592)
|
||||
- restic: brute force only viable if much of the password structure is known; otherwise "consider it lost" - [Source](https://forum.restic.net/t/forgotten-password/2990)
|
||||
- Borg: "always need both the Borg key and passphrase"; keyfile loss = repo loss ("if you lose the key, you lose access"); mandated offsite export ("leaving your keys inside your car" warning); empty passphrase with repokey = anyone reading the repo can unlock (as good as no encryption); keyfile+empty passphrase acceptable only if client disk is encrypted - [Source](https://borgbackup.readthedocs.io/en/stable/usage/init.html)
|
||||
- Borg: exported key stays encrypted - recovery needs both export file AND original passphrase, stored in separate safe places - [Source](https://borgbackup.readthedocs.io/en/stable/usage/key.html)
|
||||
- Kopia: "you cannot restore your files if you forget your password, there is no way to recover a forgotten password because only you know it"; FAQ: "There is no way to recover it or the files and folders within that repository. Store your repository password in a safe place, such as a password manager" - [Source](https://github.com/kopia/kopia/blob/master/site/content/docs/Features/_index.md); [Source](https://kopia.io/docs/faqs/)
|
||||
- Kopia footgun nuance: password CAN be reset while still connected (`change-password`), which rescues forgotten-but-connected GUI users via CLI, but a disconnected client with a forgotten password is unrecoverable - [Source](https://kopia.io/docs/faqs/)
|
||||
- Kopia reconnect token (`status -t -s`) trivially decodes to the password - storing it insecurely leaks the repo - [Source](https://kopia.io/docs/reference/command-line/)
|
||||
- rclone crypt: custom salt is "effectively a second password that must be memorized" and is NOT stored with the data; losing `password2` loses data even with correct `password` - [Source](https://rclone.org/crypt/)
|
||||
- rclone footguns: obscured passwords look encrypted but use a static shared AES-CTR key (cursory-inspection only); 1.49.0–1.53.2 random-password generator bug produced insecure passwords (fixed 1.53.3, must rotate by re-upload) - [Source](https://rclone.org/crypt/)
|
||||
- Tarsnap FAQ: "You can't. Your key file contains the only copy of the cryptographic keys… if you lose them there is no way to get your data back" (same for forgotten key-file passphrase); separate FAQ covers being unable to stop billing after losing all keys - [Source](https://www.tarsnap.com/faq.html)
|
||||
|
||||
### Inferences
|
||||
- The envelope designs make server-side recovery cryptographically impossible, so every vendor pushes the same operational answer: paper/physical export + password manager + offsite separation of key-export and passphrase.
|
||||
|
||||
@@ -6,36 +6,36 @@
|
||||
All four encrypt file content client-side but leak different amounts of structural metadata: restic leaks backend object sizes/counts/timestamps and config filenames; Borg leaks chunk sizes/proximity unless obfuscation is enabled; Kopia leaks only prefixed blob names (p/q/x) and approximate pack sizes; rclone crypt leaks file sizes (±16B) and mtimes and optionally directory structure.
|
||||
|
||||
### Cited Findings
|
||||
- restic design: “Apart from the files stored within the `keys` directory, all files are encrypted with AES-256 in counter mode (CTR)” with Poly1305-AES MAC, IV in first 16 bytes — keys directory itself is the exception and contains KDF parameters in plaintext JSON (`hostname`, `username`, `kdf=scrypt`, `N=65536,r=8,p=1`, `salt`, `created`) — [Source](https://github.com/restic/restic/blob/master/doc/design.rst); same guarantee restated as “Everything except the metadata included for informational purposes in the key files is encrypted and authenticated” — [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- restic repo layout on dumb backend is content-defined: `config`, `keys/`, `locks/`, `snapshots/`, `index/`, `data/` pack files named by hex SHA-256 of plaintext content; filenames of packs/snapshots/indexes are plaintext hashes, sizes and mtimes visible to server — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- restic explicitly warns attacker with write access “can determine which files belong to what snapshot (e.g. based on the timestamps of the stored files)” — [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- restic local cache “is also encrypted to prevent metadata leaks” — [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- Borg attack model: “environment of the client process (e.g. borg create) is trusted and the repository (server) is not. The attacker has any and all access to the repository, including interactive manipulation (man-in-the-middle)” — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg guarantees vs untrusted server: attacker cannot (1) modify archive data undetected, (2) rename/remove/add archive undetected, (3) recover plaintext, (4) recover definite structural information such as object graph — but “heuristics based on access patterns are possible” — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg multi-client caveat: “When the above attack model is extended to include multiple clients independently updating the same repository, then Borg fails to provide confidentiality (i.e. guarantees 3) and 4) do not apply any more)” — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg chunk-ID is MAC of plaintext (`id = AUTHENTICATOR(id_key, data)`), encryption is Encrypt-then-MAC AES-256-CTR + HMAC-SHA256 or BLAKE2b-256; IV counter “added in plaintext” and tracked via client security DB + repo reservation — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg 2.0 changes to AEAD (`AES-256-OCB` or `chacha20-poly1305`) with session keys, chunk ID via MAC over plaintext, object ID bound as AAD so attacker cannot “change the type of the object or move content to a different object ID” — [Source](https://borgbackup.readthedocs.io/en/2.0.0b12/internals/security.html)
|
||||
- Borg fingerprinting: “A borg repository does not hide the size of the chunks it stores”; small files <512KiB yield single chunk; buzhash chunker uses secret per-repo chunker seed; optional `obfuscate` pseudo-compressor pads with 0x00 bytes (only adds size) to hinder size fingerprinting — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg proximity leak: “Borg does not try to obfuscate order / proximity of files”; sorts by inode order not name, but “when new files are close to each other [in recursion order], the resulting chunks will be also stored close to each other in segment file(s)” — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Kopia envelope encryption: repo has plaintext JSON `formatBlob` with `tool`, `buildVersion`, `uniqueID`, `keyAlgo=scrypt-65536-8-1`, `version`, `encryption` (e.g. `AES256_GCM`), plus `encryptedBlockFormat` ciphertext; `UniqueID` doubles as PBKDF salt and HKDF info — [Source](https://kopia.io/docs/advanced/encryption)
|
||||
- Kopia passphrase → 32B master key `Km = PBKDF(passphrase, UniqueID)` (scrypt N=65536,r=8,p=1), then `Ke = HKDF(SHA256,Km,UniqueID,"AES",32)` and `AD = HKDF(SHA256,Km,UniqueID,"CHECKSUM",32)` for format-blob AES-GCM — [Source](https://kopia.io/docs/advanced/encryption)
|
||||
- Kopia content encryption options are `AES256-GCM-HMAC-SHA256` (default) or `CHACHA20-POLY1305-HMAC-SHA256`, chosen at repo creation, cannot be changed after — [Source](https://github.com/kopia/kopia/blob/master/site/content/docs/FAQs/_index.md)
|
||||
- Kopia blob layer: pack blobs have random names, “don’t reveal anything about their contents or structure”; prefixes `p` (data packs), `q` (metadata packs), `x` (indices); object-ID single-letter prefixes `k` (directory), `m` (manifest block), `x` (indirect JSON) route to `q` vs `p` packs — [Source](https://kopia.io/docs/advanced/architecture)
|
||||
- Kopia packs are 20–40MB merged blobs; “sizes are also generally unrelated to their content due to splitting and merging”; per-pack trailing local index enables recovery if global index lost — [Source](https://kopia.io/docs/advanced/architecture)
|
||||
- Kopia manifests (snapshots/policies) are small JSON stored as encrypted CABS blocks, addressed by `key=value` labels (e.g. `type:policy`, `hostname:… path:… username:…` visible via `kopia manifest list` client-side, not plaintext on backend) — [Source](https://kopia.io/docs/getting-started/)
|
||||
- Kopia: “encrypts these snapshots before they leave your computer”, “password never leaves your machine”, single password per repo, “currently no access control mechanism when sharing a repository” — [Source](https://kopia.io/docs/features/)
|
||||
- rclone crypt: wraps any backend; “automatically encrypt (before uploading) and decrypt (after downloading) on your local system … leaving the data encrypted at rest”; direct access to wrapped remote bypasses crypto — [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt file format: 8B magic `RCLONE\x00\x00` + 24B nonce header, then 64KiB chunks each with 16B Poly1305 tag (XSalsa20+Poly1305 SecretBox); 1B file → 49B total, 1MiB → 1048864B (+0.03%) — [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt name encryption: path split on `/`, PKCS#7 padded to 16B, EME-AES-256 deterministic, base32-lowercase-no-pad; identical names → identical ciphertexts; common prefixes hidden — [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt explicitly “does not encrypt file length - this can be calculated within 16 bytes” nor “modification time - used for syncing”; versions suffix `-vYYYY-MM-DD…` left plaintext; directory names optionally unencrypted — [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt filename modes: `standard` (encrypted, ~143 char limit, dir structure visible), `obfuscate` (“Very simple filename obfuscation … cannot be relied upon for strong protection”), `off` (adds `.bin` only) — [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt: “Any metadata supported by the underlying remote is read and written” — i.e. WebDAV/SFTP xattrs/mtimes pass through unencrypted — [Source](https://rclone.org/crypt/)
|
||||
- rclone WebDAV backend: “Plain WebDAV does not support modified times” nor hashes except via Fastmail/ownCloud/Nextcloud vendor extensions — [Source](https://rclone.org/webdav/)
|
||||
- restic design: “Apart from the files stored within the `keys` directory, all files are encrypted with AES-256 in counter mode (CTR)” with Poly1305-AES MAC, IV in first 16 bytes - keys directory itself is the exception and contains KDF parameters in plaintext JSON (`hostname`, `username`, `kdf=scrypt`, `N=65536,r=8,p=1`, `salt`, `created`) - [Source](https://github.com/restic/restic/blob/master/doc/design.rst); same guarantee restated as “Everything except the metadata included for informational purposes in the key files is encrypted and authenticated” - [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- restic repo layout on dumb backend is content-defined: `config`, `keys/`, `locks/`, `snapshots/`, `index/`, `data/` pack files named by hex SHA-256 of plaintext content; filenames of packs/snapshots/indexes are plaintext hashes, sizes and mtimes visible to server - [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- restic explicitly warns attacker with write access “can determine which files belong to what snapshot (e.g. based on the timestamps of the stored files)” - [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- restic local cache “is also encrypted to prevent metadata leaks” - [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- Borg attack model: “environment of the client process (e.g. borg create) is trusted and the repository (server) is not. The attacker has any and all access to the repository, including interactive manipulation (man-in-the-middle)” - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg guarantees vs untrusted server: attacker cannot (1) modify archive data undetected, (2) rename/remove/add archive undetected, (3) recover plaintext, (4) recover definite structural information such as object graph - but “heuristics based on access patterns are possible” - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg multi-client caveat: “When the above attack model is extended to include multiple clients independently updating the same repository, then Borg fails to provide confidentiality (i.e. guarantees 3) and 4) do not apply any more)” - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg chunk-ID is MAC of plaintext (`id = AUTHENTICATOR(id_key, data)`), encryption is Encrypt-then-MAC AES-256-CTR + HMAC-SHA256 or BLAKE2b-256; IV counter “added in plaintext” and tracked via client security DB + repo reservation - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg 2.0 changes to AEAD (`AES-256-OCB` or `chacha20-poly1305`) with session keys, chunk ID via MAC over plaintext, object ID bound as AAD so attacker cannot “change the type of the object or move content to a different object ID” - [Source](https://borgbackup.readthedocs.io/en/2.0.0b12/internals/security.html)
|
||||
- Borg fingerprinting: “A borg repository does not hide the size of the chunks it stores”; small files <512KiB yield single chunk; buzhash chunker uses secret per-repo chunker seed; optional `obfuscate` pseudo-compressor pads with 0x00 bytes (only adds size) to hinder size fingerprinting - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg proximity leak: “Borg does not try to obfuscate order / proximity of files”; sorts by inode order not name, but “when new files are close to each other [in recursion order], the resulting chunks will be also stored close to each other in segment file(s)” - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Kopia envelope encryption: repo has plaintext JSON `formatBlob` with `tool`, `buildVersion`, `uniqueID`, `keyAlgo=scrypt-65536-8-1`, `version`, `encryption` (e.g. `AES256_GCM`), plus `encryptedBlockFormat` ciphertext; `UniqueID` doubles as PBKDF salt and HKDF info - [Source](https://kopia.io/docs/advanced/encryption)
|
||||
- Kopia passphrase → 32B master key `Km = PBKDF(passphrase, UniqueID)` (scrypt N=65536,r=8,p=1), then `Ke = HKDF(SHA256,Km,UniqueID,"AES",32)` and `AD = HKDF(SHA256,Km,UniqueID,"CHECKSUM",32)` for format-blob AES-GCM - [Source](https://kopia.io/docs/advanced/encryption)
|
||||
- Kopia content encryption options are `AES256-GCM-HMAC-SHA256` (default) or `CHACHA20-POLY1305-HMAC-SHA256`, chosen at repo creation, cannot be changed after - [Source](https://github.com/kopia/kopia/blob/master/site/content/docs/FAQs/_index.md)
|
||||
- Kopia blob layer: pack blobs have random names, “don’t reveal anything about their contents or structure”; prefixes `p` (data packs), `q` (metadata packs), `x` (indices); object-ID single-letter prefixes `k` (directory), `m` (manifest block), `x` (indirect JSON) route to `q` vs `p` packs - [Source](https://kopia.io/docs/advanced/architecture)
|
||||
- Kopia packs are 20–40MB merged blobs; “sizes are also generally unrelated to their content due to splitting and merging”; per-pack trailing local index enables recovery if global index lost - [Source](https://kopia.io/docs/advanced/architecture)
|
||||
- Kopia manifests (snapshots/policies) are small JSON stored as encrypted CABS blocks, addressed by `key=value` labels (e.g. `type:policy`, `hostname:… path:… username:…` visible via `kopia manifest list` client-side, not plaintext on backend) - [Source](https://kopia.io/docs/getting-started/)
|
||||
- Kopia: “encrypts these snapshots before they leave your computer”, “password never leaves your machine”, single password per repo, “currently no access control mechanism when sharing a repository” - [Source](https://kopia.io/docs/features/)
|
||||
- rclone crypt: wraps any backend; “automatically encrypt (before uploading) and decrypt (after downloading) on your local system … leaving the data encrypted at rest”; direct access to wrapped remote bypasses crypto - [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt file format: 8B magic `RCLONE\x00\x00` + 24B nonce header, then 64KiB chunks each with 16B Poly1305 tag (XSalsa20+Poly1305 SecretBox); 1B file → 49B total, 1MiB → 1048864B (+0.03%) - [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt name encryption: path split on `/`, PKCS#7 padded to 16B, EME-AES-256 deterministic, base32-lowercase-no-pad; identical names → identical ciphertexts; common prefixes hidden - [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt explicitly “does not encrypt file length - this can be calculated within 16 bytes” nor “modification time - used for syncing”; versions suffix `-vYYYY-MM-DD…` left plaintext; directory names optionally unencrypted - [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt filename modes: `standard` (encrypted, ~143 char limit, dir structure visible), `obfuscate` (“Very simple filename obfuscation … cannot be relied upon for strong protection”), `off` (adds `.bin` only) - [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt: “Any metadata supported by the underlying remote is read and written” - i.e. WebDAV/SFTP xattrs/mtimes pass through unencrypted - [Source](https://rclone.org/crypt/)
|
||||
- rclone WebDAV backend: “Plain WebDAV does not support modified times” nor hashes except via Fastmail/ownCloud/Nextcloud vendor extensions - [Source](https://rclone.org/webdav/)
|
||||
|
||||
### Inferences
|
||||
- For size-hiding, Kopia pack merging (>20MB) is the strongest default; restic pack (~4-8MB default, content-defined) is intermediate; Borg chunk-level sizes leak most without `obfuscate` compressor.
|
||||
- Deterministic name encryption (rclone crypt standard, restic hash-named packs are content-derived not name-derived) still leaks equality + directory shape; only Kopia random pack names hide shape by default.
|
||||
- Plaintext `keys/` (restic) and `formatBlob` (Kopia) both expose KDF parameters and repo unique IDs — useful for offline dictionary attack cost estimation but not plaintext.
|
||||
- Plaintext `keys/` (restic) and `formatBlob` (Kopia) both expose KDF parameters and repo unique IDs - useful for offline dictionary attack cost estimation but not plaintext.
|
||||
|
||||
### Gaps
|
||||
- No reliable primary-doc figure found for exact restic snapshot/index plaintext filename entropy or padding policy; restic docs do not claim filename obfuscation or padding.
|
||||
@@ -47,30 +47,30 @@ All four encrypt file content client-side but leak different amounts of structur
|
||||
Borg has the strongest formal malicious-server story (Horton-anchored DAG + TAM-signed manifest + nonce tracking); restic authenticates all objects but explicitly does not detect deletion/rollback or timestamp-grouping attacks; Kopia relies on content-addressed HMAC + maintenance/verify with server-side immutability as add-on; rclone crypt has per-chunk authentication but no manifest/rollback concept.
|
||||
|
||||
### Cited Findings
|
||||
- restic guarantees: “Modifications to data … can be detected” and “Data that has been tampered will not be decrypted” (MAC checked before decrypt) — [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- restic non-goals: “not designed to protect against attackers deleting files”; attacker deleting timestamp-correlated packs makes “particular snapshot vanish … Restic is not designed to detect this attack” — [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- restic compromised-host cases: attacker with repo write can “Create snapshots (containing garbage data) which cover all modified files and wait until a trusted host has used forget often enough to remove all correct snapshots” and “Create a garbage snapshot for every existing snapshot with slightly different timestamp” to trick rotation — [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- restic append-only caveat: attacker with append-only access can still “Capture the password and decrypt past and future backups” (no forward secrecy); safe `forget` requires separate doc procedure — [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- restic key-leak: “impossible to securely revoke a leaked key without re-encrypting the whole repository” — [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- restic open issue #5041 proposes extending model with append-only enforcement + management/data key separation so “adversary who compromises the host cannot retroactively publish a snapshot which will trick forget into deletion” — still open/discussion as of 2026 — [Source](https://github.com/restic/restic/issues/5041)
|
||||
- Borg structural auth follows “Horton principle”: every object referenced by parent via plaintext-MAC object ID up to manifest, forming authenticated DAG — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg manifest has fixed ID `000…000` so cannot be DAG-authenticated; protected since 1.0.9 by TAM: `tam_key = HKDF-SHA-512(id_key||enc_key||enc_hmac_key, RANDOM(64), "borg-metadata-authentication-manifest")`, `HMAC(tam_key, packed_manifest)` stored in manifest — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg nonce/counter anti-reuse: client commits +4GiB reservation to security DB + repo before encrypting; crash-safe via SaveFile; but “in a multiple-client scenario a repository can trick a client into reusing counter values by ignoring counter reservations and replaying the manifest (which will fail if the client has seen a more recent manifest or has a more recent nonce reservation)” — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg RPC: over system SSH (no own network code), server can only send responses not requests; msgpack limited Unpacker; worst-case server can impose repo DoS; log-channel confusion limited to in-progress requests — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg CVE-2023-36811: “flaw in the cryptographic authentication scheme … allowed an attacker to fake archives and potentially indirectly cause backup data loss”; requires inserting files without headers + repo write; “does not disclose plaintext … nor affect authenticity of existing archives”; fixed in 1.2.5 + upgrade procedure; mitigate by reviewing archives after `check --repair` before `prune` — [Source](https://nvd.nist.gov/vuln/detail/CVE-2023-36811)
|
||||
- Borg append-only: `borg config repo append_only 1` or `borg serve --append-only`; “never overwrite or delete committed data” at segment level, but `delete/prune` still allowed (appear to succeed, only tag as deleted in new transaction); `compact` becomes no-op with no warning — [Source](https://raw.githubusercontent.com/borgbackup/borg/1.4.5/docs/usage/notes.rst)
|
||||
- Borg append-only rollback: transaction log (`transactions` file with UTC timestamps) allows manual rollback by deleting segment files from attack point onward (e.g. `rm data/**/{6..13}`), provided `compact` has not run; must clear client cache after (`borg delete --cache-only`) — [Source](https://raw.githubusercontent.com/borgbackup/borg/1.4.5/docs/usage/notes.rst)
|
||||
- Borg append-only limits: “Append-only mode is not respected by tools other than Borg. rm still works”; clients must only access via `borg serve`; any non-append-only write (admin prune/create) permanently compacts away “deleted” data; SSH `authorized_keys` split (`--append-only` key for untrusted clients, full key for admin) is recommended pattern — [Source](https://borgbackup.readthedocs.io/en/1.1-maint/usage/notes.html)
|
||||
- Kopia consistency: `kopia snapshot verify` walks snapshot roots, checks index structures + blob existence; runs automatically during daily full maintenance; `--verify-files-percent=N` samples downloads/decrypts (100% ≈ test restore, discarded after check); e.g. 1% daily → ~98% coverage over a year — [Source](https://kopia.io/docs/advanced/consistency/)
|
||||
- Kopia corruption causes listed as non-POSIX/networked filesystems, silent bit-rot, large clock skew (few minutes tolerated, larger can cause self-deletion); recommends mature POSIX FS or cloud storage + NTP — [Source](https://kopia.io/docs/advanced/consistency/)
|
||||
- Kopia ransomware model: malware may exfiltrate cloud keys and delete snapshots; mitigation is provider-side restricted keys (no delete) + object-lock/retention (COMPLIANCE), not client crypto alone — [Source](https://kopia.io/docs/advanced/ransomware-protection/)
|
||||
- Kopia object-lock: `--retention-mode COMPLIANCE --retention-period <e.g.30d>` at `repo create s3`, plus `maintenance set --extend-object-locks true` with `full-interval` ≥1 day shorter than retention; supports S3 (+B2-via-S3), Azure version-level immutability, GCS versioning+retention; `--point-in-time` reconnect for S3 restores — [Source](https://kopia.io/docs/advanced/ransomware-protection/)
|
||||
- rclone crypt integrity: per-chunk Poly1305 via SecretBox; “data integrity is protected by an extremely strong crypto authenticator”; `cryptcheck` (not `check`) required since hashes not stored — [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt has no manifest/snapshot ledger: replay of old encrypted files, rollback, or server-side deletion is undetectable by crypt layer; `--pass-bad-blocks` (zero-fill corrupt chunks) is opt-in recovery only — [Source](https://rclone.org/crypt/)
|
||||
- restic guarantees: “Modifications to data … can be detected” and “Data that has been tampered will not be decrypted” (MAC checked before decrypt) - [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- restic non-goals: “not designed to protect against attackers deleting files”; attacker deleting timestamp-correlated packs makes “particular snapshot vanish … Restic is not designed to detect this attack” - [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- restic compromised-host cases: attacker with repo write can “Create snapshots (containing garbage data) which cover all modified files and wait until a trusted host has used forget often enough to remove all correct snapshots” and “Create a garbage snapshot for every existing snapshot with slightly different timestamp” to trick rotation - [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- restic append-only caveat: attacker with append-only access can still “Capture the password and decrypt past and future backups” (no forward secrecy); safe `forget` requires separate doc procedure - [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- restic key-leak: “impossible to securely revoke a leaked key without re-encrypting the whole repository” - [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- restic open issue #5041 proposes extending model with append-only enforcement + management/data key separation so “adversary who compromises the host cannot retroactively publish a snapshot which will trick forget into deletion” - still open/discussion as of 2026 - [Source](https://github.com/restic/restic/issues/5041)
|
||||
- Borg structural auth follows “Horton principle”: every object referenced by parent via plaintext-MAC object ID up to manifest, forming authenticated DAG - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg manifest has fixed ID `000…000` so cannot be DAG-authenticated; protected since 1.0.9 by TAM: `tam_key = HKDF-SHA-512(id_key||enc_key||enc_hmac_key, RANDOM(64), "borg-metadata-authentication-manifest")`, `HMAC(tam_key, packed_manifest)` stored in manifest - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg nonce/counter anti-reuse: client commits +4GiB reservation to security DB + repo before encrypting; crash-safe via SaveFile; but “in a multiple-client scenario a repository can trick a client into reusing counter values by ignoring counter reservations and replaying the manifest (which will fail if the client has seen a more recent manifest or has a more recent nonce reservation)” - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg RPC: over system SSH (no own network code), server can only send responses not requests; msgpack limited Unpacker; worst-case server can impose repo DoS; log-channel confusion limited to in-progress requests - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg CVE-2023-36811: “flaw in the cryptographic authentication scheme … allowed an attacker to fake archives and potentially indirectly cause backup data loss”; requires inserting files without headers + repo write; “does not disclose plaintext … nor affect authenticity of existing archives”; fixed in 1.2.5 + upgrade procedure; mitigate by reviewing archives after `check --repair` before `prune` - [Source](https://nvd.nist.gov/vuln/detail/CVE-2023-36811)
|
||||
- Borg append-only: `borg config repo append_only 1` or `borg serve --append-only`; “never overwrite or delete committed data” at segment level, but `delete/prune` still allowed (appear to succeed, only tag as deleted in new transaction); `compact` becomes no-op with no warning - [Source](https://raw.githubusercontent.com/borgbackup/borg/1.4.5/docs/usage/notes.rst)
|
||||
- Borg append-only rollback: transaction log (`transactions` file with UTC timestamps) allows manual rollback by deleting segment files from attack point onward (e.g. `rm data/**/{6..13}`), provided `compact` has not run; must clear client cache after (`borg delete --cache-only`) - [Source](https://raw.githubusercontent.com/borgbackup/borg/1.4.5/docs/usage/notes.rst)
|
||||
- Borg append-only limits: “Append-only mode is not respected by tools other than Borg. rm still works”; clients must only access via `borg serve`; any non-append-only write (admin prune/create) permanently compacts away “deleted” data; SSH `authorized_keys` split (`--append-only` key for untrusted clients, full key for admin) is recommended pattern - [Source](https://borgbackup.readthedocs.io/en/1.1-maint/usage/notes.html)
|
||||
- Kopia consistency: `kopia snapshot verify` walks snapshot roots, checks index structures + blob existence; runs automatically during daily full maintenance; `--verify-files-percent=N` samples downloads/decrypts (100% ≈ test restore, discarded after check); e.g. 1% daily → ~98% coverage over a year - [Source](https://kopia.io/docs/advanced/consistency/)
|
||||
- Kopia corruption causes listed as non-POSIX/networked filesystems, silent bit-rot, large clock skew (few minutes tolerated, larger can cause self-deletion); recommends mature POSIX FS or cloud storage + NTP - [Source](https://kopia.io/docs/advanced/consistency/)
|
||||
- Kopia ransomware model: malware may exfiltrate cloud keys and delete snapshots; mitigation is provider-side restricted keys (no delete) + object-lock/retention (COMPLIANCE), not client crypto alone - [Source](https://kopia.io/docs/advanced/ransomware-protection/)
|
||||
- Kopia object-lock: `--retention-mode COMPLIANCE --retention-period <e.g.30d>` at `repo create s3`, plus `maintenance set --extend-object-locks true` with `full-interval` ≥1 day shorter than retention; supports S3 (+B2-via-S3), Azure version-level immutability, GCS versioning+retention; `--point-in-time` reconnect for S3 restores - [Source](https://kopia.io/docs/advanced/ransomware-protection/)
|
||||
- rclone crypt integrity: per-chunk Poly1305 via SecretBox; “data integrity is protected by an extremely strong crypto authenticator”; `cryptcheck` (not `check`) required since hashes not stored - [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt has no manifest/snapshot ledger: replay of old encrypted files, rollback, or server-side deletion is undetectable by crypt layer; `--pass-bad-blocks` (zero-fill corrupt chunks) is opt-in recovery only - [Source](https://rclone.org/crypt/)
|
||||
|
||||
### Inferences
|
||||
- Only Borg attempts to bind manifest meaning (TAM) against a fully MITM server with persistent client state; restic/Kopia assume server may be read/write-malicious for confidentiality/integrity of objects but push rollback/availability to out-of-band controls (append-only buckets, object-lock, separate admin host).
|
||||
- For WebDAV-class backends with no compute, Borg-style `serve --append-only` enforcement is unavailable; closest analogues are restic append-only-capable REST backends (not WebDAV) or Kopia S3 object-lock — WebDAV must rely on server ACLs/versioning.
|
||||
- For WebDAV-class backends with no compute, Borg-style `serve --append-only` enforcement is unavailable; closest analogues are restic append-only-capable REST backends (not WebDAV) or Kopia S3 object-lock - WebDAV must rely on server ACLs/versioning.
|
||||
|
||||
### Gaps
|
||||
- No primary-source evidence found that Kopia authenticates manifest labels/meaning against malicious-server manifest substitution beyond content HMAC; treat as gap.
|
||||
@@ -79,23 +79,23 @@ Borg has the strongest formal malicious-server story (Horton-anchored DAG + TAM-
|
||||
## How do they rate-limit, retry, and verify uploads (idempotent PUTs, integrity re-checks)?
|
||||
|
||||
### Takeaway
|
||||
All target dumb blob stores with idempotent content-addressed PUTs and client-side re-verification (`check`/`verify`/`cryptcheck`); retry/backoff is client-configured and backend-specific; none documents server-side rate-limiting — throttling is via client concurrency limits and provider quotas.
|
||||
All target dumb blob stores with idempotent content-addressed PUTs and client-side re-verification (`check`/`verify`/`cryptcheck`); retry/backoff is client-configured and backend-specific; none documents server-side rate-limiting - throttling is via client concurrency limits and provider quotas.
|
||||
|
||||
### Cited Findings
|
||||
- restic repo design “allows parallel access of multiple instances … even parallel writes”; locks with `--retry-lock` retry until timeout; `restic check` verifies pack hash vs pack ID, `cat pack <id>` warns on mismatch — [Source](https://restic.readthedocs.io/en/latest/100_references.html?highlight=threat)
|
||||
- Kopia architecture: content-addressed blocks (`Block ID = hash(data)`); identical blocks dedup naturally; uploads merged into packs; index maps block→(blob,offset,len) — enabling idempotent re-PUT (same ID = same bytes) — [Source](https://kopia.io/docs/advanced/architecture)
|
||||
- Kopia verification tiers: metadata-only `snapshot verify` every full maintenance + opt-in content download sample `--verify-files-percent` + `--file-parallelism` for throughput control — [Source](https://kopia.io/docs/advanced/consistency/)
|
||||
- Kopia storage requirements imply retry need: “strong read-after-write consistency and eventual list-after-write consistency”; “can compensate for such inconsistent behaviors for up to several minutes, but larger inconsistencies can lead to data loss” — [Source](https://kopia.io/docs/advanced/consistency/)
|
||||
- Borg RPC worst-case throttling noted only as server-imposed “denial of repository service”; client uses limited msgpack Unpacker to avoid memory-DoS from large messages — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- rclone crypt backup guidance: use `sync` on encrypted paths with same passwords so “will check the checksums while copying”; `check` between two encrypted remotes works, `cryptcheck` needed vs plaintext — [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt chunking (64KiB + 16B tag) bounds retry unit; header nonce incremented per chunk, reuse probability ~2e-32 per exabyte — [Source](https://rclone.org/crypt/)
|
||||
- restic repo design “allows parallel access of multiple instances … even parallel writes”; locks with `--retry-lock` retry until timeout; `restic check` verifies pack hash vs pack ID, `cat pack <id>` warns on mismatch - [Source](https://restic.readthedocs.io/en/latest/100_references.html?highlight=threat)
|
||||
- Kopia architecture: content-addressed blocks (`Block ID = hash(data)`); identical blocks dedup naturally; uploads merged into packs; index maps block→(blob,offset,len) - enabling idempotent re-PUT (same ID = same bytes) - [Source](https://kopia.io/docs/advanced/architecture)
|
||||
- Kopia verification tiers: metadata-only `snapshot verify` every full maintenance + opt-in content download sample `--verify-files-percent` + `--file-parallelism` for throughput control - [Source](https://kopia.io/docs/advanced/consistency/)
|
||||
- Kopia storage requirements imply retry need: “strong read-after-write consistency and eventual list-after-write consistency”; “can compensate for such inconsistent behaviors for up to several minutes, but larger inconsistencies can lead to data loss” - [Source](https://kopia.io/docs/advanced/consistency/)
|
||||
- Borg RPC worst-case throttling noted only as server-imposed “denial of repository service”; client uses limited msgpack Unpacker to avoid memory-DoS from large messages - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- rclone crypt backup guidance: use `sync` on encrypted paths with same passwords so “will check the checksums while copying”; `check` between two encrypted remotes works, `cryptcheck` needed vs plaintext - [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt chunking (64KiB + 16B tag) bounds retry unit; header nonce incremented per chunk, reuse probability ~2e-32 per exabyte - [Source](https://rclone.org/crypt/)
|
||||
|
||||
### Inferences
|
||||
- Content-addressing (restic pack ID, Borg chunk ID, Kopia block ID) makes PUTs naturally idempotent and safe to retry; verification is always client-driven (`check`/`verify`/`cryptcheck`) since dumb backends cannot attest.
|
||||
- Parallelism knobs (`--file-parallelism`, rclone `--transfers/--checkers`, Borg client concurrency) are the de facto rate-limiters; no evidence of server-push backpressure handling in fetched docs.
|
||||
|
||||
### Gaps
|
||||
- No primary-source values found in fetched docs for restic `--retry-lock` defaults, backend backoff schedule, or WebDAV-specific retry/idempotency; mark as gap — need restic backend docs + rclone WebDAV backend options.
|
||||
- No primary-source values found in fetched docs for restic `--retry-lock` defaults, backend backoff schedule, or WebDAV-specific retry/idempotency; mark as gap - need restic backend docs + rclone WebDAV backend options.
|
||||
- No Kopia primary-doc retry/backoff parameters or blob `Put` atomicity guarantees surfaced in fetched pages; need `repo/blob` Go API docs.
|
||||
- No Borg primary-doc upload retry/rate-limit parameters surfaced; Borg docs focus on correctness not throttling.
|
||||
|
||||
@@ -105,22 +105,22 @@ All target dumb blob stores with idempotent content-addressed PUTs and client-si
|
||||
Treat WebDAV as untrusted byte store: do all crypto/index/manifest work client-side, never depend on server mtime/hash, use atomic PUT + verify-after-write, and add out-of-band append-only/versioning since WebDAV has no compute to enforce it.
|
||||
|
||||
### Cited Findings
|
||||
- Kopia repository model: “All Repository features are implemented client-side, without any need for a custom server, thus encryption keys never leave the client”; layers are Object/Manifest/Block/Raw-BLOB over simple blob API — directly applicable to WebDAV — [Source](https://github.com/kopia/repo/blob/master/README.md)
|
||||
- Kopia supports “Any remote server or cloud storage that supports WebDAV” and SFTP as first-class repo storage, plus Rclone-wrapped Dropbox/OneDrive/Google-Drive (experimental) — [Source](https://github.com/kopia/kopia?pubDate=20260701)
|
||||
- Kopia caveat for dumb/networked FS: requires “POSIX semantics, atomic writes (either native or emulated), strong read-after-write … eventual list-after-write”; “avoid emulated, layered, or networked filesystems which may not be fully compliant. Alternatively, use cloud storage”; WebDAV falls in risky class — [Source](https://kopia.io/docs/advanced/consistency/)
|
||||
- Kopia clock rule: “clock skews of few minutes are tolerated, but bigger clock skews can lead to major inconsistencies, including Kopia deleting its own data”; run NTP on clients and servers — [Source](https://kopia.io/docs/advanced/consistency/)
|
||||
- rclone WebDAV limits: “Plain WebDAV does not support modified times … does not support hashes” (except Fastmail/ownCloud/Nextcloud extensions) — so sync daemons must not trust mtime/hash for change detection — [Source](https://rclone.org/webdav/)
|
||||
- rclone crypt on WebDAV pattern: point crypt at `remote:path` subdirectory, access exclusively via crypt remote; `.bin` suffix added when names unencrypted to prevent provider interpreting content; `strict_names` errors on mixed encrypted/unencrypted dirs — [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt password rotation requires full re-upload (decrypt-with-old → encrypt-with-new), either from local source or crypt-to-crypt `move`; “half the bandwidth and charged twice” on metered backends — [Source](https://rclone.org/crypt/)
|
||||
- restic threat-model implication for WebDAV: server sees object sizes/counts/timestamps; timestamp-grouping enables targeted snapshot deletion — mitigation is to avoid leaking timing (jitter uploads, shared pack sizes) though not prescribed in docs — [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- Borg inode-order (not alpha) directory traversal + secret chunker seed + optional size obfuscation are concrete dumb-backend-hardening techniques portable to any daemon: randomize scan order, per-repo secret for chunking, pad to size classes — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg SSH `authorized_keys` split (`borg serve --append-only` for untrusted clients vs full for admin) has no WebDAV equivalent; WebDAV equivalent is provider ACL/versioning or Kopia-style object-lock where available (S3/Azure/GCS only — not WebDAV) — [Source](https://borgbackup.readthedocs.io/en/1.1-maint/usage/notes.html); Kopia object-lock “currently only supports object locks when using an S3 repo” (native B2 excluded, use S3 mode) — [Source](https://kopia.io/docs/advanced/ransomware-protection/)
|
||||
- Filippo Valsorda review quoted in restic docs: “The design might not be perfect, but it’s good. Encryption is a first-class feature, the implementation looks sane and … deduplication trade-off is worth it” — informal audit signal, not formal audit — [Source](https://github.com/restic/restic/blob/master/doc/070_encryption.rst)
|
||||
- Kopia repository model: “All Repository features are implemented client-side, without any need for a custom server, thus encryption keys never leave the client”; layers are Object/Manifest/Block/Raw-BLOB over simple blob API - directly applicable to WebDAV - [Source](https://github.com/kopia/repo/blob/master/README.md)
|
||||
- Kopia supports “Any remote server or cloud storage that supports WebDAV” and SFTP as first-class repo storage, plus Rclone-wrapped Dropbox/OneDrive/Google-Drive (experimental) - [Source](https://github.com/kopia/kopia?pubDate=20260701)
|
||||
- Kopia caveat for dumb/networked FS: requires “POSIX semantics, atomic writes (either native or emulated), strong read-after-write … eventual list-after-write”; “avoid emulated, layered, or networked filesystems which may not be fully compliant. Alternatively, use cloud storage”; WebDAV falls in risky class - [Source](https://kopia.io/docs/advanced/consistency/)
|
||||
- Kopia clock rule: “clock skews of few minutes are tolerated, but bigger clock skews can lead to major inconsistencies, including Kopia deleting its own data”; run NTP on clients and servers - [Source](https://kopia.io/docs/advanced/consistency/)
|
||||
- rclone WebDAV limits: “Plain WebDAV does not support modified times … does not support hashes” (except Fastmail/ownCloud/Nextcloud extensions) - so sync daemons must not trust mtime/hash for change detection - [Source](https://rclone.org/webdav/)
|
||||
- rclone crypt on WebDAV pattern: point crypt at `remote:path` subdirectory, access exclusively via crypt remote; `.bin` suffix added when names unencrypted to prevent provider interpreting content; `strict_names` errors on mixed encrypted/unencrypted dirs - [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt password rotation requires full re-upload (decrypt-with-old → encrypt-with-new), either from local source or crypt-to-crypt `move`; “half the bandwidth and charged twice” on metered backends - [Source](https://rclone.org/crypt/)
|
||||
- restic threat-model implication for WebDAV: server sees object sizes/counts/timestamps; timestamp-grouping enables targeted snapshot deletion - mitigation is to avoid leaking timing (jitter uploads, shared pack sizes) though not prescribed in docs - [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- Borg inode-order (not alpha) directory traversal + secret chunker seed + optional size obfuscation are concrete dumb-backend-hardening techniques portable to any daemon: randomize scan order, per-repo secret for chunking, pad to size classes - [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg SSH `authorized_keys` split (`borg serve --append-only` for untrusted clients vs full for admin) has no WebDAV equivalent; WebDAV equivalent is provider ACL/versioning or Kopia-style object-lock where available (S3/Azure/GCS only - not WebDAV) - [Source](https://borgbackup.readthedocs.io/en/1.1-maint/usage/notes.html); Kopia object-lock “currently only supports object locks when using an S3 repo” (native B2 excluded, use S3 mode) - [Source](https://kopia.io/docs/advanced/ransomware-protection/)
|
||||
- Filippo Valsorda review quoted in restic docs: “The design might not be perfect, but it’s good. Encryption is a first-class feature, the implementation looks sane and … deduplication trade-off is worth it” - informal audit signal, not formal audit - [Source](https://github.com/restic/restic/blob/master/doc/070_encryption.rst)
|
||||
|
||||
### Inferences
|
||||
- For a local daemon syncing to WebDAV: (1) encrypt+MAC everything client-side with random pack names (Kopia-style) or hash-named packs (restic-style); (2) keep manifest/index signed with local key (Borg TAM-style) and verify on every sync; (3) never use server mtime/ETag as source of truth; (4) verify-after-write via GET+MAC before deleting local staging; (5) mitigate rollback by keeping local monotonic snapshot counter + out-of-band copy; (6) pad/obfuscate sizes and jitter upload times to reduce grouping leaks.
|
||||
- Since WebDAV cannot enforce append-only, ransomware safety must come from server-side versioning/quotas + separate prune identity, mirroring Borg append-only and Kopia restricted-keys guidance.
|
||||
|
||||
### Gaps
|
||||
- No primary WebDAV-server hardening guide (append-only ACLs, versioning) surfaced for restic/Borg/Kopia on plain WebDAV; restic WebDAV backend doc and rclone WebDAV option reference not fetched (tool-call budget) — need follow-up.
|
||||
- No formal third-party audit reports (e.g. Cure53/Quarkslab) confirmed for restic/Borg/Kopia in fetched sources; only Valsorda informal review and CVE record found — need dedicated audit search.
|
||||
- No primary WebDAV-server hardening guide (append-only ACLs, versioning) surfaced for restic/Borg/Kopia on plain WebDAV; restic WebDAV backend doc and rclone WebDAV option reference not fetched (tool-call budget) - need follow-up.
|
||||
- No formal third-party audit reports (e.g. Cure53/Quarkslab) confirmed for restic/Borg/Kopia in fetched sources; only Valsorda informal review and CVE record found - need dedicated audit search.
|
||||
|
||||
@@ -6,20 +6,20 @@
|
||||
Restic, Borg and Kopia all use content-defined chunking (CDC) with ~0.5–8 MiB variable chunks so small edits to large binaries only store 1–2 new chunks; without CDC (fixed blocks or full-file copies) a 1-byte insert re-stores the whole file, and Git-LFS instead punts large files to pointer+blob storage with host-enforced per-file caps.
|
||||
|
||||
### Cited Findings
|
||||
- Restic splits files with Rabin-fingerprint CDC over a 64-byte sliding window, cutting when low 21 bits are zero; files <512 KiB are not split, blobs are 512 KiB–8 MiB, ~1 MiB average — [Source](https://github.com/restic/restic/blob/master/doc/design.rst); background — [Source](https://restic.net/blog/2015-09-12/restic-foundation1-cdc/)
|
||||
- Restic chunker defaults aim at ~1 MiB average (`splitmask = (1<<20)-1`) with configurable Min/MaxSize — [Source](https://github.com/restic/chunker/blob/master/chunker.go)
|
||||
- Borg splits files into deduplicated chunks globally across repo (all machines/archives); chunk id is a strong hash/MAC (hmac-sha256 / keyed blake3), not the rolling-hash value — [Source](https://borgbackup.readthedocs.io/en/master/)
|
||||
- Borg default buzhash params: min 2^19 (512 KiB), max 2^23 (8 MiB), mask 21 bits (~2 MiB target), window 4095 B; `fixed` chunker option for disk/VM images — [Source](https://github.com/borgbackup/borg/blob/86fd77fd/docs/internals/data-structures.rst)
|
||||
- Borg 2.x chunkers: `fastcdc` (default, fastest), `buzhash64`/`buzhash`, keyed AES variants (`toeplitz-aes`/`rabin-aes` strongest), `fixed` for raw disk images where CDC gains little — [Source](https://borgbackup.readthedocs.io/en/latest/internals/chunker.html)
|
||||
- Borg warns fine-grained `--chunker-params=buzhash,10,23,16,4095` creates huge chunk counts and RAM/disk load; coarse default suits large volumes — [Source](https://borgbackup.readthedocs.io/en/stable/usage/notes.html)
|
||||
- Kopia calls CDC "splitters": FIXED vs DYNAMIC BUZHASH/RABINKARP, sizes 1M–8M (default DYNAMIC-4M-BUZHASH); small files = one content, large files split so metadata-only change to 10 GB video uploads only 1–2 chunks (<10 MB) — [Source](https://kopia.discourse.group/t/does-kopia-use-content-defined-chunking-cdc/1417); details — [Source](https://kopia.discourse.group/t/difference-between-the-available-splitters/894); packing — [Source](https://kopia.discourse.group/t/do-hashed-chunks-span-multiple-files/1444)
|
||||
- Kopia packs many contents into 20–40 MB pack blobs; splitter choice is set at repo creation — [Source](https://kopia.io/docs/advanced/architecture/); chunk-then-hash-then-compress-then-encrypt pipeline — [Source](https://kopia.io/docs/advanced/compression/)
|
||||
- Restic has no hard max file size but offers `--exclude-larger-than size` (suffixes k/M/G/T) to skip files over a threshold — [Source](https://restic.readthedocs.io/en/stable/040_backup.html)
|
||||
- Git-LFS has no inherent file-size limit; limits are host-enforced (GitHub: 2 GB Free/Pro, 4 GB Team, 5 GB Enterprise Cloud; >5 GB rejected); pointer file stores `version/oid sha256/size` — [Source](https://docs.github.com/en/repositories/working-with-files/managing-large-files/about-git-large-file-storage)
|
||||
- Git-LFS tracks by `.gitattributes` pattern, not by size; `--above` in `migrate import` is one-shot, no automatic by-size tracking (2025 `--min-size/autotracksize` PR still debated/unmerged) — [Source](https://github.com/git-lfs/git-lfs/blob/main/docs/man/git-lfs-faq.adoc)
|
||||
- Git on Windows pre-2.34 could not smudge/clean files >4 GiB; workaround `GIT_LFS_SKIP_SMUDGE=1` + `git lfs pull` — [Source](https://github.com/git-lfs/git-lfs/blob/main/docs/man/git-lfs-faq.adoc)
|
||||
- Every Git-LFS revision counts against remote storage/bandwidth quota, so per-version binaries inflate cost; clones/pulls are slow on large repos — [Source](https://get.assembla.com/blog/git-lfs/)
|
||||
- Keyed CDC (Borg/Restic/Kopia) is fingerprintable: observing chunk sizes of known data recovers keys (CCS 2025, eprint 2025/558; backup-service attacks eprint 2025/532); Restic 0.18 mitigates by random chunk-to-pack assignment — [Source](https://github.com/restic/restic/blob/master/doc/design.rst); attacks — [Source](https://eprint.iacr.org/2025/558); analysis — [Source](https://eprint.iacr.org/2025/532.pdf)
|
||||
- Restic splits files with Rabin-fingerprint CDC over a 64-byte sliding window, cutting when low 21 bits are zero; files <512 KiB are not split, blobs are 512 KiB–8 MiB, ~1 MiB average - [Source](https://github.com/restic/restic/blob/master/doc/design.rst); background - [Source](https://restic.net/blog/2015-09-12/restic-foundation1-cdc/)
|
||||
- Restic chunker defaults aim at ~1 MiB average (`splitmask = (1<<20)-1`) with configurable Min/MaxSize - [Source](https://github.com/restic/chunker/blob/master/chunker.go)
|
||||
- Borg splits files into deduplicated chunks globally across repo (all machines/archives); chunk id is a strong hash/MAC (hmac-sha256 / keyed blake3), not the rolling-hash value - [Source](https://borgbackup.readthedocs.io/en/master/)
|
||||
- Borg default buzhash params: min 2^19 (512 KiB), max 2^23 (8 MiB), mask 21 bits (~2 MiB target), window 4095 B; `fixed` chunker option for disk/VM images - [Source](https://github.com/borgbackup/borg/blob/86fd77fd/docs/internals/data-structures.rst)
|
||||
- Borg 2.x chunkers: `fastcdc` (default, fastest), `buzhash64`/`buzhash`, keyed AES variants (`toeplitz-aes`/`rabin-aes` strongest), `fixed` for raw disk images where CDC gains little - [Source](https://borgbackup.readthedocs.io/en/latest/internals/chunker.html)
|
||||
- Borg warns fine-grained `--chunker-params=buzhash,10,23,16,4095` creates huge chunk counts and RAM/disk load; coarse default suits large volumes - [Source](https://borgbackup.readthedocs.io/en/stable/usage/notes.html)
|
||||
- Kopia calls CDC "splitters": FIXED vs DYNAMIC BUZHASH/RABINKARP, sizes 1M–8M (default DYNAMIC-4M-BUZHASH); small files = one content, large files split so metadata-only change to 10 GB video uploads only 1–2 chunks (<10 MB) - [Source](https://kopia.discourse.group/t/does-kopia-use-content-defined-chunking-cdc/1417); details - [Source](https://kopia.discourse.group/t/difference-between-the-available-splitters/894); packing - [Source](https://kopia.discourse.group/t/do-hashed-chunks-span-multiple-files/1444)
|
||||
- Kopia packs many contents into 20–40 MB pack blobs; splitter choice is set at repo creation - [Source](https://kopia.io/docs/advanced/architecture/); chunk-then-hash-then-compress-then-encrypt pipeline - [Source](https://kopia.io/docs/advanced/compression/)
|
||||
- Restic has no hard max file size but offers `--exclude-larger-than size` (suffixes k/M/G/T) to skip files over a threshold - [Source](https://restic.readthedocs.io/en/stable/040_backup.html)
|
||||
- Git-LFS has no inherent file-size limit; limits are host-enforced (GitHub: 2 GB Free/Pro, 4 GB Team, 5 GB Enterprise Cloud; >5 GB rejected); pointer file stores `version/oid sha256/size` - [Source](https://docs.github.com/en/repositories/working-with-files/managing-large-files/about-git-large-file-storage)
|
||||
- Git-LFS tracks by `.gitattributes` pattern, not by size; `--above` in `migrate import` is one-shot, no automatic by-size tracking (2025 `--min-size/autotracksize` PR still debated/unmerged) - [Source](https://github.com/git-lfs/git-lfs/blob/main/docs/man/git-lfs-faq.adoc)
|
||||
- Git on Windows pre-2.34 could not smudge/clean files >4 GiB; workaround `GIT_LFS_SKIP_SMUDGE=1` + `git lfs pull` - [Source](https://github.com/git-lfs/git-lfs/blob/main/docs/man/git-lfs-faq.adoc)
|
||||
- Every Git-LFS revision counts against remote storage/bandwidth quota, so per-version binaries inflate cost; clones/pulls are slow on large repos - [Source](https://get.assembla.com/blog/git-lfs/)
|
||||
- Keyed CDC (Borg/Restic/Kopia) is fingerprintable: observing chunk sizes of known data recovers keys (CCS 2025, eprint 2025/558; backup-service attacks eprint 2025/532); Restic 0.18 mitigates by random chunk-to-pack assignment - [Source](https://github.com/restic/restic/blob/master/doc/design.rst); attacks - [Source](https://eprint.iacr.org/2025/558); analysis - [Source](https://eprint.iacr.org/2025/532.pdf)
|
||||
|
||||
### Inferences
|
||||
- A small local daemon should copy the Restic/Borg default: CDC ~1–2 MiB average, 512 KiB min, 8 MiB max; offer `--exclude-larger-than`-style cap (e.g. 50–100 MB default-off) plus extension-based excludes rather than a hard cap.
|
||||
@@ -36,12 +36,12 @@ Restic, Borg and Kopia all use content-defined chunking (CDC) with ~0.5–8 MiB
|
||||
Git (and libgit2 ports) still use NUL-in-first-8000-bytes as the primary binary signal, tightened in 2008 so CRLF conversion defers to diff; it is fast and recommended as a first heuristic but known-insufficient for UTF-16 and NUL-free binaries, so explicit `.gitattributes` marking is required.
|
||||
|
||||
### Cited Findings
|
||||
- Git `buffer_is_binary()` checks for NUL via `memchr` in first 8000 bytes (`FIRST_FEW_BYTES`) — [Source](https://stackoverflow.com/questions/6119956/how-to-determine-if-git-handles-a-file-as-binary-or-as-text)
|
||||
- Git `convert.c` `convert_is_binary()`: binary if `lonecr` or `nul` or `(printable>>7) < nonprintable`; CRLF auto-conversion bails on binary — [Source](https://code.googlesource.com/git/+/645cc7a2a7274a92403d2848ef643a96f1589d09/convert.c)
|
||||
- libgit2 `git_blob_is_binary` uses core-git heuristic: NUL scan + printable/nonprintable ratio over first 8000 bytes — [Source](https://libgit2.org/docs/reference/main/blob/git_blob_is_binary.html)
|
||||
- 2008 patch unified heuristics: any NUL forces binary in `convert.c` so CRLF handling is stricter than diff (prior convert.c used only <1% nonprintable rule, mis-handling tar/word-processor files diff called binary) — [Source](https://public-inbox.org/git/20080116011321.GD13984@dpotapov.dyndns.org/t/)
|
||||
- Git mailing-list guidance: NUL in first 8000 bytes = binary; UTF-16 must be marked explicitly, Git does not handle it internally; short NUL-free binaries must also be marked explicitly — [Source](https://public-inbox.org/git/20151202004921.GC28197@sigill.intra.peff.net/T/)
|
||||
- `git diff --numstat` reports `-\t-` for binary (practical detector); `git check-attr` only reflects `.gitattributes`, not the heuristic — [Source](https://stackoverflow.com/questions/6119956/how-to-determine-if-git-handles-a-file-as-binary-or-as-text)
|
||||
- Git `buffer_is_binary()` checks for NUL via `memchr` in first 8000 bytes (`FIRST_FEW_BYTES`) - [Source](https://stackoverflow.com/questions/6119956/how-to-determine-if-git-handles-a-file-as-binary-or-as-text)
|
||||
- Git `convert.c` `convert_is_binary()`: binary if `lonecr` or `nul` or `(printable>>7) < nonprintable`; CRLF auto-conversion bails on binary - [Source](https://code.googlesource.com/git/+/645cc7a2a7274a92403d2848ef643a96f1589d09/convert.c)
|
||||
- libgit2 `git_blob_is_binary` uses core-git heuristic: NUL scan + printable/nonprintable ratio over first 8000 bytes - [Source](https://libgit2.org/docs/reference/main/blob/git_blob_is_binary.html)
|
||||
- 2008 patch unified heuristics: any NUL forces binary in `convert.c` so CRLF handling is stricter than diff (prior convert.c used only <1% nonprintable rule, mis-handling tar/word-processor files diff called binary) - [Source](https://public-inbox.org/git/20080116011321.GD13984@dpotapov.dyndns.org/t/)
|
||||
- Git mailing-list guidance: NUL in first 8000 bytes = binary; UTF-16 must be marked explicitly, Git does not handle it internally; short NUL-free binaries must also be marked explicitly - [Source](https://public-inbox.org/git/20151202004921.GC28197@sigill.intra.peff.net/T/)
|
||||
- `git diff --numstat` reports `-\t-` for binary (practical detector); `git check-attr` only reflects `.gitattributes`, not the heuristic - [Source](https://stackoverflow.com/questions/6119956/how-to-determine-if-git-handles-a-file-as-binary-or-as-text)
|
||||
|
||||
### Inferences
|
||||
- NUL-sniffing remains the recommended cheap first pass for a small daemon (scan first 8 KiB), matching Git/libgit2 behavior and user expectations.
|
||||
@@ -56,20 +56,20 @@ Git (and libgit2 ports) still use NUL-in-first-8000-bytes as the primary binary
|
||||
Never `cp`/read the raw sqlite file hot: in WAL mode the consistent image spans main+`-wal`+`-shm` and byte copies tear or go stale; use the Online Backup API (`sqlite3_backup_*` / `Connection.backup()` / `.backup` CLI), `VACUUM INTO`, or a quiesced filesystem snapshot, then `PRAGMA integrity_check`.
|
||||
|
||||
### Cited Findings
|
||||
- Historical `cp`-under-shared-lock method is fast but blocks writers, cannot copy to/from memory DBs, and risks corruption on power/OS failure — [Source](https://sqlite.org/backup.html)
|
||||
- Backup API: source read-locked only during each `sqlite3_backup_step(nPage)`; destination write-locked throughout; incremental stepping lets writers proceed; concurrent write by another connection restarts backup automatically — [Source](https://sqlite.org/backup.html); API contract — [Source](https://sqlite.org/c3ref/backup_finish.html)
|
||||
- Python exposes as `Connection.backup(target, pages, progress, sleep)`; `pages=-1` copies all at once (holds lock), positive pages + `sleep=0.250` yields between steps; restart detected when `remaining` jumps back toward `total` — [Source](https://www.productionhardening.org/backup-recovery-data-integrity/online-backup-api-hot-copies/)
|
||||
- WAL-mode live DB is three files (main + `-wal` committed-not-checkpointed frames + `-shm` index); `cp`/`rsync`/snapshot captures them at different instants → `SQLITE_CORRUPT`/`SQLITE_NOTADB`; same hazard for `-journal` in rollback mode — [Source](https://www.productionhardening.org/backup-recovery-data-integrity/)
|
||||
- Real-world failure: `fs.copyFile()` of only `.db` in WAL mode produced 100% corrupt backups across weeks of 6-hour cron; fix was `better-sqlite3 .backup()` plus hour-granular filenames — [Source](https://scottspence.com/posts/sqlite-corruption-fs-copyfile-issue)
|
||||
- Correct one-liners: `sqlite3 app.db ".backup './db-backups/app.db'"` (general) or `VACUUM INTO` (same safety + compaction, needs SQLite 3.27+, refuses if target exists so `rm -f` first); for dedup pipelines prefer `.backup` because compaction reshuffles pages and defeats chunking — [Source](https://www.backupdata.io/resources/guides/sqlite-backups-you-can-actually-restore)
|
||||
- Borg docs: Borg just copies file as-is; if DB is written mid-read the archive may be inconsistent — use sqlite-aware method (`sqlite3 db.sqlite "VACUUM INTO 'copy.sqlite'"`) or filesystem snapshot — [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst); Borg FAQ repeats the `VACUUM INTO` advice — [Source](https://borgbackup.readthedocs.io/en/master/faq.html)
|
||||
- borgmatic best practice: dump (export) databases rather than backing internal files; streams dump directly to Borg — [Source](https://torsion.org/borgmatic/how-to/backup-your-databases/)
|
||||
- Checkpoint+lock schemes (`wal_checkpoint(TRUNCATE)` then `BEGIN IMMEDIATE` then copy main+wal) are fragile: checkpoint can fail to get writer lock, writers can interleave, autocheckpoint is per-process — use the backup API instead — [Source](https://sqlite.org/forum/forumpost/2ea989bbe9)
|
||||
- Post-backup rules: run `PRAGMA integrity_check` (expect single `ok`) on every finished image; pre-size target ~1.2× source; set `busy_timeout ≥5000ms`; write target off the source I/O queue; delete partials; on restore stop app, copy main file, delete stale `-wal`/`-shm` — [Source](https://www.productionhardening.org/backup-recovery-data-integrity/online-backup-api-hot-copies/); restore checklist — [Source](https://www.backupdata.io/resources/guides/sqlite-backups-you-can-actually-restore)
|
||||
- Historical `cp`-under-shared-lock method is fast but blocks writers, cannot copy to/from memory DBs, and risks corruption on power/OS failure - [Source](https://sqlite.org/backup.html)
|
||||
- Backup API: source read-locked only during each `sqlite3_backup_step(nPage)`; destination write-locked throughout; incremental stepping lets writers proceed; concurrent write by another connection restarts backup automatically - [Source](https://sqlite.org/backup.html); API contract - [Source](https://sqlite.org/c3ref/backup_finish.html)
|
||||
- Python exposes as `Connection.backup(target, pages, progress, sleep)`; `pages=-1` copies all at once (holds lock), positive pages + `sleep=0.250` yields between steps; restart detected when `remaining` jumps back toward `total` - [Source](https://www.productionhardening.org/backup-recovery-data-integrity/online-backup-api-hot-copies/)
|
||||
- WAL-mode live DB is three files (main + `-wal` committed-not-checkpointed frames + `-shm` index); `cp`/`rsync`/snapshot captures them at different instants → `SQLITE_CORRUPT`/`SQLITE_NOTADB`; same hazard for `-journal` in rollback mode - [Source](https://www.productionhardening.org/backup-recovery-data-integrity/)
|
||||
- Real-world failure: `fs.copyFile()` of only `.db` in WAL mode produced 100% corrupt backups across weeks of 6-hour cron; fix was `better-sqlite3 .backup()` plus hour-granular filenames - [Source](https://scottspence.com/posts/sqlite-corruption-fs-copyfile-issue)
|
||||
- Correct one-liners: `sqlite3 app.db ".backup './db-backups/app.db'"` (general) or `VACUUM INTO` (same safety + compaction, needs SQLite 3.27+, refuses if target exists so `rm -f` first); for dedup pipelines prefer `.backup` because compaction reshuffles pages and defeats chunking - [Source](https://www.backupdata.io/resources/guides/sqlite-backups-you-can-actually-restore)
|
||||
- Borg docs: Borg just copies file as-is; if DB is written mid-read the archive may be inconsistent - use sqlite-aware method (`sqlite3 db.sqlite "VACUUM INTO 'copy.sqlite'"`) or filesystem snapshot - [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst); Borg FAQ repeats the `VACUUM INTO` advice - [Source](https://borgbackup.readthedocs.io/en/master/faq.html)
|
||||
- borgmatic best practice: dump (export) databases rather than backing internal files; streams dump directly to Borg - [Source](https://torsion.org/borgmatic/how-to/backup-your-databases/)
|
||||
- Checkpoint+lock schemes (`wal_checkpoint(TRUNCATE)` then `BEGIN IMMEDIATE` then copy main+wal) are fragile: checkpoint can fail to get writer lock, writers can interleave, autocheckpoint is per-process - use the backup API instead - [Source](https://sqlite.org/forum/forumpost/2ea989bbe9)
|
||||
- Post-backup rules: run `PRAGMA integrity_check` (expect single `ok`) on every finished image; pre-size target ~1.2× source; set `busy_timeout ≥5000ms`; write target off the source I/O queue; delete partials; on restore stop app, copy main file, delete stale `-wal`/`-shm` - [Source](https://www.productionhardening.org/backup-recovery-data-integrity/online-backup-api-hot-copies/); restore checklist - [Source](https://www.backupdata.io/resources/guides/sqlite-backups-you-can-actually-restore)
|
||||
|
||||
### Inferences
|
||||
- Small-daemon default: if file header is `SQLite format 3`, do not copy directly; shell to `sqlite3 "$f" ".backup '$tmp'"` or `VACUUM INTO` (when compaction desired) into a temp file, then ingest the temp file; fall back to copy only when DB is quiesced and `wal_checkpoint(TRUNCATE)` shows zero log pages.
|
||||
- Smaller Borg-style chunks (~64–128 KiB target, e.g. `fastcdc,15,19,17`) cut sqlite re-storage ~3–8× vs 2 MiB defaults at the cost of chunk-count RAM — see next section.
|
||||
- Smaller Borg-style chunks (~64–128 KiB target, e.g. `fastcdc,15,19,17`) cut sqlite re-storage ~3–8× vs 2 MiB defaults at the cost of chunk-count RAM - see next section.
|
||||
|
||||
### Gaps
|
||||
- `.dump` (SQL text) vs `.backup` (page image) size/dedup tradeoff not quantified from primary sources in this pass; anecdotal claim that dumps dedup to KBs but take hours on 12 GB DBs is single-issue-report only.
|
||||
@@ -77,20 +77,20 @@ Never `cp`/read the raw sqlite file hot: in WAL mode the consistent image spans
|
||||
## Do images/archives dedup at all, and is per-version storage of them worth it vs plain mirroring?
|
||||
|
||||
### Takeaway
|
||||
Recompressed, encrypted, or gzipped-per-version artifacts barely dedup (a 1-byte change avalanches through compression; encrypted chunks are high-entropy), while uncompressed raw images and plain sqlite files dedup well under CDC — so version verbatim binaries that change little, otherwise mirror single-copy.
|
||||
Recompressed, encrypted, or gzipped-per-version artifacts barely dedup (a 1-byte change avalanches through compression; encrypted chunks are high-entropy), while uncompressed raw images and plain sqlite files dedup well under CDC - so version verbatim binaries that change little, otherwise mirror single-copy.
|
||||
|
||||
### Cited Findings
|
||||
- Restic maintainer: pass uncompressed data; small input change → large compressed-output change → new blob hashes → repo grows (50 GB gzipped DB dumps → 50 GB repo vs 36 GB for raw); CDC blobs identified by SHA-256 — [Source](https://github.com/restic/restic/issues/790)
|
||||
- Duplicati skips recompression/dedup for known-compressed extensions by default (saves CPU; whole-file handling) because metadata edits rewrite the stream; identical copies still dedup, moves are free — [Source](https://forum.duplicati.com/t/deduplication-for-large-files/8394)
|
||||
- Borg on 5 generations of same sqlite DB: gzipped inputs = 667 unique/667 total chunks (zero dedup, 1.6 GB); uncompressed = strong dedup (590 MB total), zstd-5 beats gzip — [Source](https://appsintheopen.com/posts/66-backing-up-sqlite-database-with-borg-and-de-duplication)
|
||||
- Borg sqlite tuning: default ~2 MiB target wastes a whole chunk per changed 4 KiB page; `fastcdc,15,19,17,2` (~128 KiB target) drastically cuts incrementals; finer `14,18,16` helps more; cost is chunk-index RAM — params apply per-run so back up DBs in a separate run — [Source](https://borgbackup.readthedocs.io/en/master/faq.html); 12.45 GB Vintage Story sqlite: 2nd backup 518 MB default → 62 MB (15,19,17) → 37 MB (10,23,16) — [Source](https://github.com/borgbackup/borg/issues/5877)
|
||||
- Kopia: 8 MB photo with metadata edit re-uploads ~8 MB under 4 MB splitter (chunk > file); fix is smaller splitter at repo-creation cost of more chunks — [Source](https://kopia.discourse.group/t/chunk-size-setting/1351)
|
||||
- Fixed-splitter warning: 10 GB video + 1 prepended byte re-uploads 10 GB; content-based (buzhash/rabinkarp) uploads 1–2 chunks — [Source](https://kopia.discourse.group/t/difference-between-the-available-splitters/894)
|
||||
- Backup pipelines must chunk-then-compress (per-chunk); encrypt-then-chunk/compress is useless since ciphertext is indistinguishable from random; encrypted dedup needs weakened convergent/MLE schemes with leakage tradeoffs — [Source](https://eprint.iacr.org/2025/532.pdf); survey — [Source](https://dl.acm.org/doi/10.1145/3685278)
|
||||
- Kopia compresses each chunk independently (s2 default, gzip optional); splitting costs little ratio (466 MB → 119 MB standalone s2 vs 133 MB via Kopia-s2) — [Source](https://kopia.io/docs/advanced/compression/)
|
||||
- Restic maintainer: pass uncompressed data; small input change → large compressed-output change → new blob hashes → repo grows (50 GB gzipped DB dumps → 50 GB repo vs 36 GB for raw); CDC blobs identified by SHA-256 - [Source](https://github.com/restic/restic/issues/790)
|
||||
- Duplicati skips recompression/dedup for known-compressed extensions by default (saves CPU; whole-file handling) because metadata edits rewrite the stream; identical copies still dedup, moves are free - [Source](https://forum.duplicati.com/t/deduplication-for-large-files/8394)
|
||||
- Borg on 5 generations of same sqlite DB: gzipped inputs = 667 unique/667 total chunks (zero dedup, 1.6 GB); uncompressed = strong dedup (590 MB total), zstd-5 beats gzip - [Source](https://appsintheopen.com/posts/66-backing-up-sqlite-database-with-borg-and-de-duplication)
|
||||
- Borg sqlite tuning: default ~2 MiB target wastes a whole chunk per changed 4 KiB page; `fastcdc,15,19,17,2` (~128 KiB target) drastically cuts incrementals; finer `14,18,16` helps more; cost is chunk-index RAM - params apply per-run so back up DBs in a separate run - [Source](https://borgbackup.readthedocs.io/en/master/faq.html); 12.45 GB Vintage Story sqlite: 2nd backup 518 MB default → 62 MB (15,19,17) → 37 MB (10,23,16) - [Source](https://github.com/borgbackup/borg/issues/5877)
|
||||
- Kopia: 8 MB photo with metadata edit re-uploads ~8 MB under 4 MB splitter (chunk > file); fix is smaller splitter at repo-creation cost of more chunks - [Source](https://kopia.discourse.group/t/chunk-size-setting/1351)
|
||||
- Fixed-splitter warning: 10 GB video + 1 prepended byte re-uploads 10 GB; content-based (buzhash/rabinkarp) uploads 1–2 chunks - [Source](https://kopia.discourse.group/t/difference-between-the-available-splitters/894)
|
||||
- Backup pipelines must chunk-then-compress (per-chunk); encrypt-then-chunk/compress is useless since ciphertext is indistinguishable from random; encrypted dedup needs weakened convergent/MLE schemes with leakage tradeoffs - [Source](https://eprint.iacr.org/2025/532.pdf); survey - [Source](https://dl.acm.org/doi/10.1145/3685278)
|
||||
- Kopia compresses each chunk independently (s2 default, gzip optional); splitting costs little ratio (466 MB → 119 MB standalone s2 vs 133 MB via Kopia-s2) - [Source](https://kopia.io/docs/advanced/compression/)
|
||||
|
||||
### Inferences
|
||||
- Daemon policy: (1) never store `.gz/.zip/.jpg/.mp4/.sqlite.gz` deltas expecting CDC wins — keep 1 mirror copy or thin retention; (2) ingest DBs/VM images uncompressed and let CDC + per-chunk compression do the work; (3) exclude or separately-schedule >100 MB media with tiny chunkers only if churn is low.
|
||||
- Daemon policy: (1) never store `.gz/.zip/.jpg/.mp4/.sqlite.gz` deltas expecting CDC wins - keep 1 mirror copy or thin retention; (2) ingest DBs/VM images uncompressed and let CDC + per-chunk compression do the work; (3) exclude or separately-schedule >100 MB media with tiny chunkers only if churn is low.
|
||||
- `.dump`/uncompressed SQL text is the most dedup-friendly DB form but slowest to produce; page-image `.backup` is the balanced default.
|
||||
|
||||
### Gaps
|
||||
|
||||
@@ -3,21 +3,21 @@
|
||||
## Which paths do guides universally say to include (~/Documents, ~/Pictures, ~/.config, project dirs, etc.)?
|
||||
|
||||
### Takeaway
|
||||
Guides converge on: back up all of `/home` (or whole `$HOME`) plus `/etc`, `/root`, `/var` (selectively) and `/opt`; on workstations this means XDG user dirs (`~/Documents`, `~/Pictures`, `~/Videos`, `~/Music`, `~/Downloads`), project dirs (`~/projects`, `~/src`), dotfiles/`~/.config`/`~/.local/share`, `~/.ssh`, browser profiles, mail, and package lists — never a bare `~/Documents`-only backup.
|
||||
Guides converge on: back up all of `/home` (or whole `$HOME`) plus `/etc`, `/root`, `/var` (selectively) and `/opt`; on workstations this means XDG user dirs (`~/Documents`, `~/Pictures`, `~/Videos`, `~/Music`, `~/Downloads`), project dirs (`~/projects`, `~/src`), dotfiles/`~/.config`/`~/.local/share`, `~/.ssh`, browser profiles, mail, and package lists - never a bare `~/Documents`-only backup.
|
||||
|
||||
### Cited Findings
|
||||
- Red Hat guide lists as must-back-up: `/etc` (system config, users/groups, networking, app configs), `/home` (all user data/downloads/documents/pictures under username), `/root` (admin scripts/notes/configs), `/var` shared/corporate data — and one dir not to back up (implied transient/cache) — [Source](https://www.redhat.com/en/blog/backup-dirs)
|
||||
- Ubuntu official help: copy personal files/settings usually in Home folder; if room, back up entire Home folder with exceptions (cache/thrash-type excludes detailed on linked page) — [Source](https://help.ubuntu.com/stable/ubuntu-help/backup-how.html.en)
|
||||
- PCWorld Linux backup roundup: as a rule a regular backup of the home directories is sufficient; `rsync -avP $HOME` to USB disk given as sufficient home backup — [Source](https://www.pcworld.com/article/2050183/best-linux-backup-tools.html)
|
||||
- OSTechNix 2026 reinstall checklist: back up more than `~/Documents`; must include personal files + `~/.local/share` (app data, game saves, Flatpak/Snap data), `~/.config` (app settings), browser profiles (Firefox `~/.mozilla` or new XDG `~/.config/mozilla` + `~/.cache/mozilla` + `~/.local/share/mozilla`, check `about:support`; Chrome/Chromium), SSH/GPG keys, dotfiles, app-specific data, installed package lists including Flatpak/Snap, system-level `/etc`, NetworkManager `system-connections`, cron/scheduled tasks; trap 1 is only backing up `~/Documents ~/Pictures ~/Downloads` and missing `~/projects`, VM images, second drives; trap 2 is forgetting hidden XDG data — [Source](https://ostechnix.com/things-to-back-up-before-reinstalling-linux)
|
||||
- OSTechNix: single `rsync` of whole `$HOME` captures dotfiles, browser profiles, SSH keys, app settings since all live under `$HOME` — [Source](https://ostechnix.com/things-to-back-up-before-reinstalling-linux)
|
||||
- Arch forum classic tar set: `tar zcvfp arch-system.gz /etc /boot /root` + per-user `/home/user1` + `/var --exclude /var/cache/pacman/pkg` — [Source](https://bbs.archlinux.org/viewtopic.php?id=83533)
|
||||
- Production borg example backs up `/home /etc /var/www /var/backups /opt /root` together with `--exclude-from` file — [Source](https://cubepath.com/docs/Backup%20Recovery/backup-with-borgbackup-deduplication)
|
||||
- Ubuntu Ask-Ubuntu full-system tar pattern excludes virtual/external `dev mnt proc sys` (and squashfs variant excludes `home media dev run mnt proc sys tmp`, then backs up `/home` separately excluding cloud mirrors like `Dropbox GoogleDrive`) — [Source](https://askubuntu.com/questions/7809/how-to-back-up-my-entire-system)
|
||||
- Linux Mint forum consensus: Timeshift = OS/system restore points, does nothing for data in `/home`; need separate file-level tool (BackInTime, FreeFileSync, Foxclone/Rescuezilla/Clonezilla for images) for Documents/Music/Pictures — [Source](https://forums.linuxmint.com/viewtopic.php?t=405449)
|
||||
- Dotfiles canon: `~/.bashrc`, `~/.bash_profile`, `~/.zshrc`, `~/.vimrc`, `~/.gitconfig`, `~/.ssh/config`, `~/.tmux.conf`, `~/.config/` XDG dir; secrets in `~/.netrc`, `~/.aws/credentials`, `~/.ssh/`, history `~/.bash_history` must be treated as sensitive / kept out of public repos — [Source](https://linuxcommandlibrary.com/man/dotfiles)
|
||||
- Dotfiles restore guides list as backup-worthy: `.ssh/config`, `.gnupg/pubring.kbx`, `.gnupg/trustdb.gpg`, plus stow-managed `zsh/git/vim/tmux` and XDG select dirs — [Source](https://github.com/RickCogley/dotfiles/blob/main/docs/how-to/backup-restore.md)
|
||||
- Keeply Linux tool scope explicitly: dotfiles, `.config` dirs, browser profiles, app settings, optionally SSH keys, emails — [Source](https://github.com/cozy533/keeply)
|
||||
- Red Hat guide lists as must-back-up: `/etc` (system config, users/groups, networking, app configs), `/home` (all user data/downloads/documents/pictures under username), `/root` (admin scripts/notes/configs), `/var` shared/corporate data - and one dir not to back up (implied transient/cache) - [Source](https://www.redhat.com/en/blog/backup-dirs)
|
||||
- Ubuntu official help: copy personal files/settings usually in Home folder; if room, back up entire Home folder with exceptions (cache/thrash-type excludes detailed on linked page) - [Source](https://help.ubuntu.com/stable/ubuntu-help/backup-how.html.en)
|
||||
- PCWorld Linux backup roundup: as a rule a regular backup of the home directories is sufficient; `rsync -avP $HOME` to USB disk given as sufficient home backup - [Source](https://www.pcworld.com/article/2050183/best-linux-backup-tools.html)
|
||||
- OSTechNix 2026 reinstall checklist: back up more than `~/Documents`; must include personal files + `~/.local/share` (app data, game saves, Flatpak/Snap data), `~/.config` (app settings), browser profiles (Firefox `~/.mozilla` or new XDG `~/.config/mozilla` + `~/.cache/mozilla` + `~/.local/share/mozilla`, check `about:support`; Chrome/Chromium), SSH/GPG keys, dotfiles, app-specific data, installed package lists including Flatpak/Snap, system-level `/etc`, NetworkManager `system-connections`, cron/scheduled tasks; trap 1 is only backing up `~/Documents ~/Pictures ~/Downloads` and missing `~/projects`, VM images, second drives; trap 2 is forgetting hidden XDG data - [Source](https://ostechnix.com/things-to-back-up-before-reinstalling-linux)
|
||||
- OSTechNix: single `rsync` of whole `$HOME` captures dotfiles, browser profiles, SSH keys, app settings since all live under `$HOME` - [Source](https://ostechnix.com/things-to-back-up-before-reinstalling-linux)
|
||||
- Arch forum classic tar set: `tar zcvfp arch-system.gz /etc /boot /root` + per-user `/home/user1` + `/var --exclude /var/cache/pacman/pkg` - [Source](https://bbs.archlinux.org/viewtopic.php?id=83533)
|
||||
- Production borg example backs up `/home /etc /var/www /var/backups /opt /root` together with `--exclude-from` file - [Source](https://cubepath.com/docs/Backup%20Recovery/backup-with-borgbackup-deduplication)
|
||||
- Ubuntu Ask-Ubuntu full-system tar pattern excludes virtual/external `dev mnt proc sys` (and squashfs variant excludes `home media dev run mnt proc sys tmp`, then backs up `/home` separately excluding cloud mirrors like `Dropbox GoogleDrive`) - [Source](https://askubuntu.com/questions/7809/how-to-back-up-my-entire-system)
|
||||
- Linux Mint forum consensus: Timeshift = OS/system restore points, does nothing for data in `/home`; need separate file-level tool (BackInTime, FreeFileSync, Foxclone/Rescuezilla/Clonezilla for images) for Documents/Music/Pictures - [Source](https://forums.linuxmint.com/viewtopic.php?t=405449)
|
||||
- Dotfiles canon: `~/.bashrc`, `~/.bash_profile`, `~/.zshrc`, `~/.vimrc`, `~/.gitconfig`, `~/.ssh/config`, `~/.tmux.conf`, `~/.config/` XDG dir; secrets in `~/.netrc`, `~/.aws/credentials`, `~/.ssh/`, history `~/.bash_history` must be treated as sensitive / kept out of public repos - [Source](https://linuxcommandlibrary.com/man/dotfiles)
|
||||
- Dotfiles restore guides list as backup-worthy: `.ssh/config`, `.gnupg/pubring.kbx`, `.gnupg/trustdb.gpg`, plus stow-managed `zsh/git/vim/tmux` and XDG select dirs - [Source](https://github.com/RickCogley/dotfiles/blob/main/docs/how-to/backup-restore.md)
|
||||
- Keeply Linux tool scope explicitly: dotfiles, `.config` dirs, browser profiles, app settings, optionally SSH keys, emails - [Source](https://github.com/cozy533/keeply)
|
||||
|
||||
### Inferences
|
||||
- Universal must-include = whole `/home/<user>` + `/etc` + package list; `/root` and `/var/lib`-style app data added for home-server role.
|
||||
@@ -28,21 +28,21 @@ Guides converge on: back up all of `/home` (or whole `$HOME`) plus `/etc`, `/roo
|
||||
- No single 2023–2026 primary guide found quantifying mail (`~/Mail`, Thunderbird `~/.thunderbird`) include rates; only secondary tool scope mentions emails.
|
||||
- No reliable source found prescribing exact treatment of `~/.cache` vs `~/.local/share` boundary for Flatpak/Snap beyond OSTechNix summary.
|
||||
|
||||
## How should small databases (sqlite), archives (.zip/.tar), and images be treated — raw files or dumps?
|
||||
## How should small databases (sqlite), archives (.zip/.tar), and images be treated - raw files or dumps?
|
||||
|
||||
### Takeaway
|
||||
Small static files (photos/images, `.zip/.tar` archives) are backed up as raw files and deduplicate well; live database files (sqlite, postgres/mysql) must be dumped (`VACUUM INTO`, `sqlite3 .backup`/backup API, `pg_dump`/`pg_dumpall`/`mysqldump`) before file backup — raw copy of a live DB is unsafe.
|
||||
Small static files (photos/images, `.zip/.tar` archives) are backed up as raw files and deduplicate well; live database files (sqlite, postgres/mysql) must be dumped (`VACUUM INTO`, `sqlite3 .backup`/backup API, `pg_dump`/`pg_dumpall`/`mysqldump`) before file backup - raw copy of a live DB is unsafe.
|
||||
|
||||
### Cited Findings
|
||||
- SQLite docs: safe live-copy methods are `sqlite3_rsync` (3.47.0+, 2024-10-21, bandwidth-efficient over SSH), `VACUUM INTO filename`, or backup API; plain file copy only safe with no transactions in progress, and if prior write failed must copy `-journal`/`-wal` together — [Source](https://sqlite.org/howtocorrupt.html)
|
||||
- SQLite forum (Borg user, WAL-mode pooled connections): cannot depend on perfect backup from open/changing DB because backup doesn't snapshot DB + journal in same instant; recommendation: close writers or backup API copy then back up the copy; `rm -f mydb.sqlite3.backup; sqlite3 mydb.sqlite3 "VACUUM INTO 'mydb.sqlite3.backup'"` added to backup script — [Source](https://sqlite.org/forum/forumpost/f173a78e2e?t=h)
|
||||
- Experiment on SQLite 3.50.6 WAL DB (1000 rows, 500 in `-wal`): `cp app.db backup_cp.db` yielded 500 rows, `integrity_check` still `ok` — stale copy passes checks silently; copying `app.db + app.db-wal + app.db-shm` together preserved 1000 rows; `.backup` (Online Backup API) copies page-by-page including WAL and restarts if source writes mid-copy — [Source](https://www.sqlprostudio.com/blog/68-how-to-safely-back-up-or-copy-a-live-sqlite-database)
|
||||
- `fs.copyFile()` on WAL DB copies only `.db` while `-wal/-shm` still written → instant `SQLITE_CORRUPT`; 7/7 rotating backups corrupted; fix is better-sqlite3 `.backup()` which handles all three files atomically; test restores — [Source](https://scottspence.com/posts/sqlite-corruption-fs-copyfile-issue)
|
||||
- Postgres docs: `pg_dump dbname > dumpfile` generates SQL commands to recreate DB; `pg_dumpall > dumpfile` preserves cluster-wide roles/tablespaces, restore via `psql -f dumpfile postgres` requiring superuser; file-level/WAL archiving is version-specific whereas dumps reload into newer versions — [Source](https://www.postgresql.org/docs/%EF%BC%99.6/backup-dump.html)
|
||||
- Postgres professional docs: `pg_dump` can emit text or archive formats for parallelism/fine-grained `pg_restore` control — [Source](https://postgrespro.com/docs/postgresql/17/backup-dump)
|
||||
- OSTechNix: don't copy raw data dir; use each DB's dump tool e.g. `mysqldump -u <user> -p <db> > backup.sql` — [Source](https://ostechnix.com/things-to-back-up-before-reinstalling-linux)
|
||||
- Borg quickstart warns: snapshot filesystems/volumes (LVM/ZFS useful), dump databases or stop DB servers, shut down VMs/containers before backing up disk images/volumes — [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)
|
||||
- Restic docs example: `plan.txt` change adds KiB, `archive.tar.gz` change re-adds MiB — illustrates large monolithic archives re-chunk poorly vs small files; still backed up as raw files — [Source](https://github.com/restic/restic/blob/master/doc/040_backup.rst)
|
||||
- SQLite docs: safe live-copy methods are `sqlite3_rsync` (3.47.0+, 2024-10-21, bandwidth-efficient over SSH), `VACUUM INTO filename`, or backup API; plain file copy only safe with no transactions in progress, and if prior write failed must copy `-journal`/`-wal` together - [Source](https://sqlite.org/howtocorrupt.html)
|
||||
- SQLite forum (Borg user, WAL-mode pooled connections): cannot depend on perfect backup from open/changing DB because backup doesn't snapshot DB + journal in same instant; recommendation: close writers or backup API copy then back up the copy; `rm -f mydb.sqlite3.backup; sqlite3 mydb.sqlite3 "VACUUM INTO 'mydb.sqlite3.backup'"` added to backup script - [Source](https://sqlite.org/forum/forumpost/f173a78e2e?t=h)
|
||||
- Experiment on SQLite 3.50.6 WAL DB (1000 rows, 500 in `-wal`): `cp app.db backup_cp.db` yielded 500 rows, `integrity_check` still `ok` - stale copy passes checks silently; copying `app.db + app.db-wal + app.db-shm` together preserved 1000 rows; `.backup` (Online Backup API) copies page-by-page including WAL and restarts if source writes mid-copy - [Source](https://www.sqlprostudio.com/blog/68-how-to-safely-back-up-or-copy-a-live-sqlite-database)
|
||||
- `fs.copyFile()` on WAL DB copies only `.db` while `-wal/-shm` still written → instant `SQLITE_CORRUPT`; 7/7 rotating backups corrupted; fix is better-sqlite3 `.backup()` which handles all three files atomically; test restores - [Source](https://scottspence.com/posts/sqlite-corruption-fs-copyfile-issue)
|
||||
- Postgres docs: `pg_dump dbname > dumpfile` generates SQL commands to recreate DB; `pg_dumpall > dumpfile` preserves cluster-wide roles/tablespaces, restore via `psql -f dumpfile postgres` requiring superuser; file-level/WAL archiving is version-specific whereas dumps reload into newer versions - [Source](https://www.postgresql.org/docs/%EF%BC%99.6/backup-dump.html)
|
||||
- Postgres professional docs: `pg_dump` can emit text or archive formats for parallelism/fine-grained `pg_restore` control - [Source](https://postgrespro.com/docs/postgresql/17/backup-dump)
|
||||
- OSTechNix: don't copy raw data dir; use each DB's dump tool e.g. `mysqldump -u <user> -p <db> > backup.sql` - [Source](https://ostechnix.com/things-to-back-up-before-reinstalling-linux)
|
||||
- Borg quickstart warns: snapshot filesystems/volumes (LVM/ZFS useful), dump databases or stop DB servers, shut down VMs/containers before backing up disk images/volumes - [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)
|
||||
- Restic docs example: `plan.txt` change adds KiB, `archive.tar.gz` change re-adds MiB - illustrates large monolithic archives re-chunk poorly vs small files; still backed up as raw files - [Source](https://github.com/restic/restic/blob/master/doc/040_backup.rst)
|
||||
|
||||
### Inferences
|
||||
- Photos/images/archives: raw-file backup is correct; no dump needed; chunked dedup tools (restic/borg/kopia) handle them but modified tar/zip re-stores large chunks.
|
||||
@@ -58,17 +58,17 @@ Small static files (photos/images, `.zip/.tar` archives) are backed up as raw fi
|
||||
None auto-selects `~/Documents`/`~/.config`; all default to exactly what paths you pass (no implicit includes) and provide opt-in excludes (`--exclude*`, `--patterns-from`, `.kopiaignore`/policy, Time Machine StdExclusions + user list) for caches, build artifacts, and system pseudo-filesystems.
|
||||
|
||||
### Cited Findings
|
||||
- Restic `backup [FILE/DIR]...` creates snapshot of exactly given args; no default source; exclude options: `--exclude/--iexclude`, `--exclude-file`, `--exclude-caches` (CACHEDIR.TAG), `--exclude-if-present foo`, `--exclude-larger-than`, `--exclude-cloud-files` (Win/macOS OneDrive/iCloud only); excludes don't apply to explicitly passed file path itself, only contents under dirs — [Source](https://restic.readthedocs.io/en/stable/040_backup.html)
|
||||
- Restic excludes use Go `filepath.Match` against full path, gitignore-like: once dir excluded can't re-include inside; example backs up selection inside `$HOME` — [Source](https://restic.readthedocs.io/en/v0.16.3/040_backup.html)
|
||||
- Restic handles `--one-file-system/-x` to not cross filesystem boundaries/subvolumes; saves/restores ACLs/xattrs; sparse holes deduped/compressed not stored explicitly — [Source](https://github.com/restic/restic/blob/master/doc/manual_rest.rst)
|
||||
- NixOS restic module example: `paths = ["/home"]` with `exclude = ["/home/*/.cache" ".git"]` pattern as conventional home selection — [Source](https://github.com/NixOS/nixpkgs/blob/master/nixos/modules/services/backup/restic.nix)
|
||||
- Borg `create repo::archive ~/Documents`, `~/Documents ~/src --exclude '*.pyc'`, `/home --exclude thumbnail regex`, ` / --one-file-system` are docs examples — user picks roots; no default root — [Source](https://borgbackup.readthedocs.io/en/1.0.5/usage.html)
|
||||
- Borg patterns: fnmatch default for `--exclude`, shell-style for `--pattern`; `sh:**/steamapps/common/**`, `sh:home/user/.cache/**`, trailing-slash `some/path/` keeps dir not contents vs no-slash excludes both; `--exclude-from`, `--patterns-from`, `--exclude-caches` (CACHEDIR.TAG), `--exclude-if-present`, `--keep-tag-files`, `--one-file-system`, nodump flag respected — [Source](https://man.archlinux.org/man/borg-patterns.1.txt)
|
||||
- Kopia: no default source; snapshots what you `snapshot create`; ignores via `.kopiaignore` (default file), global/per-source policy `--add-ignore/--add-dot-ignore`, `--ignore-cache-dirs true` (inherited global), `--ignore-dir-errors/--ignore-file-errors`, `--one-file-system`, never/only-compress lists; before/after folder/root actions for dumps/snapshots with timeout/modes — [Source](https://kopia.io/docs/reference/command-line/common/policy-set/)
|
||||
- Kopia ignore syntax: `#` comment, `!` negate, `*`/`**`/`?`/`[0-9a-zA-Z]`/`[abc]`, leading `/` root-only; examples `*.dat`, `/logs/*`, `tmp.db`, `**/logs/**`; must test — incomplete snapshots possible — [Source](https://kopia.io/docs/advanced/kopiaignore/)
|
||||
- Time Machine default backs up everything except system files/apps from macOS install, caches, StdExclusions (`/System/Library/CoreServices/backupd.bundle/Contents/Resources/StdExclusions.plist`), backup disk itself; user adds via Options → Exclude; excluded items still in local snapshots; damaged system requires macOS reinstall + Migration Assistant — [Source](https://support.apple.com/en-ca/guide/mac-help/mh15622/mac)
|
||||
- Time Machine `tmutil isexcluded/addexclusion/removeexclusion`, sticky vs fixed-path (`-p`) exclusions semantics — [Source](https://www.unix.com/man-page/osx/8/TMUTIL)
|
||||
- ArchWiki System backup: no default set; methods are btrfs/LVM snapshots, rsync, tar, SquashFS (no ACLs); recommends 3-2-1, regular integrity + restore tests; automation via systemd timer/cron with least-privilege `CAP_DAC_READ_SEARCH` example — [Source](https://wiki.archlinux.org/title/System_backup)
|
||||
- Restic `backup [FILE/DIR]...` creates snapshot of exactly given args; no default source; exclude options: `--exclude/--iexclude`, `--exclude-file`, `--exclude-caches` (CACHEDIR.TAG), `--exclude-if-present foo`, `--exclude-larger-than`, `--exclude-cloud-files` (Win/macOS OneDrive/iCloud only); excludes don't apply to explicitly passed file path itself, only contents under dirs - [Source](https://restic.readthedocs.io/en/stable/040_backup.html)
|
||||
- Restic excludes use Go `filepath.Match` against full path, gitignore-like: once dir excluded can't re-include inside; example backs up selection inside `$HOME` - [Source](https://restic.readthedocs.io/en/v0.16.3/040_backup.html)
|
||||
- Restic handles `--one-file-system/-x` to not cross filesystem boundaries/subvolumes; saves/restores ACLs/xattrs; sparse holes deduped/compressed not stored explicitly - [Source](https://github.com/restic/restic/blob/master/doc/manual_rest.rst)
|
||||
- NixOS restic module example: `paths = ["/home"]` with `exclude = ["/home/*/.cache" ".git"]` pattern as conventional home selection - [Source](https://github.com/NixOS/nixpkgs/blob/master/nixos/modules/services/backup/restic.nix)
|
||||
- Borg `create repo::archive ~/Documents`, `~/Documents ~/src --exclude '*.pyc'`, `/home --exclude thumbnail regex`, ` / --one-file-system` are docs examples - user picks roots; no default root - [Source](https://borgbackup.readthedocs.io/en/1.0.5/usage.html)
|
||||
- Borg patterns: fnmatch default for `--exclude`, shell-style for `--pattern`; `sh:**/steamapps/common/**`, `sh:home/user/.cache/**`, trailing-slash `some/path/` keeps dir not contents vs no-slash excludes both; `--exclude-from`, `--patterns-from`, `--exclude-caches` (CACHEDIR.TAG), `--exclude-if-present`, `--keep-tag-files`, `--one-file-system`, nodump flag respected - [Source](https://man.archlinux.org/man/borg-patterns.1.txt)
|
||||
- Kopia: no default source; snapshots what you `snapshot create`; ignores via `.kopiaignore` (default file), global/per-source policy `--add-ignore/--add-dot-ignore`, `--ignore-cache-dirs true` (inherited global), `--ignore-dir-errors/--ignore-file-errors`, `--one-file-system`, never/only-compress lists; before/after folder/root actions for dumps/snapshots with timeout/modes - [Source](https://kopia.io/docs/reference/command-line/common/policy-set/)
|
||||
- Kopia ignore syntax: `#` comment, `!` negate, `*`/`**`/`?`/`[0-9a-zA-Z]`/`[abc]`, leading `/` root-only; examples `*.dat`, `/logs/*`, `tmp.db`, `**/logs/**`; must test - incomplete snapshots possible - [Source](https://kopia.io/docs/advanced/kopiaignore/)
|
||||
- Time Machine default backs up everything except system files/apps from macOS install, caches, StdExclusions (`/System/Library/CoreServices/backupd.bundle/Contents/Resources/StdExclusions.plist`), backup disk itself; user adds via Options → Exclude; excluded items still in local snapshots; damaged system requires macOS reinstall + Migration Assistant - [Source](https://support.apple.com/en-ca/guide/mac-help/mh15622/mac)
|
||||
- Time Machine `tmutil isexcluded/addexclusion/removeexclusion`, sticky vs fixed-path (`-p`) exclusions semantics - [Source](https://www.unix.com/man-page/osx/8/TMUTIL)
|
||||
- ArchWiki System backup: no default set; methods are btrfs/LVM snapshots, rsync, tar, SquashFS (no ACLs); recommends 3-2-1, regular integrity + restore tests; automation via systemd timer/cron with least-privilege `CAP_DAC_READ_SEARCH` example - [Source](https://wiki.archlinux.org/title/System_backup)
|
||||
|
||||
### Inferences
|
||||
- Versioning tools scope = explicit roots + exclude list; portable Linux default is effectively `/home + /etc + dumps` minus `*.cache/CACHEDIR.TAG`, `node_modules/target/.git`, VM images while running.
|
||||
@@ -84,15 +84,15 @@ None auto-selects `~/Documents`/`~/.config`; all default to exactly what paths y
|
||||
Never file-copy a live DB; quiesce (close writers/stop server), filesystem/LVM/ZFS snapshot, or dump (`sqlite3 backup API/VACUUM INTO`, `pg_dumpall`, `mysqldump`); for SQLite WAL must keep `.db+-wal+-shm/-journal` together, never delete hot journals, beware POSIX `close()` dropping advisory locks and fork/link/rename hazards.
|
||||
|
||||
### Cited Findings
|
||||
- SQLite: backup/restore while transaction active → copy mixes old/new → corrupt; safe via `sqlite3_rsync`, `VACUUM INTO`, backup API even on live DB; idle-file copy only safe with no tx in progress — [Source](https://sqlite.org/howtocorrupt.html)
|
||||
- SQLite: never move/delete/rename hot `-journal/-wal` after crash; mispairing (swap/overwrite/move journal, copy DB without journal, overwrite DB without deleting hot journal) likely corrupts; quiescent DB has no journal, only DB file matters — [Source](https://sqlite.org/howtocorrupt.html)
|
||||
- SQLite POSIX advisory-lock quirk: any thread `open/read/close` on DB file drops locks for all threads (close cancels locks); bypassing lib for backup read can corrupt; since 3.51.0 (2025-11-04) extra WAL defenses but not cure-all — never `close()` DB file while connections open even in other threads — [Source](https://sqlite.org/howtocorrupt.html)
|
||||
- SQLite: don't use two SQLite copies linked in one app (separate lock lists), don't mix locking protocols (POSIX vs dot-file/NFS), don't unlink/rename open DB (shared journal name → cross-recovery corruption, `SQLITE_WARNING` since 3.7.17), don't multi-link/symlink same file (wrong journal lookup; canonicalization since 3.10.0), don't carry connection across `fork()` — [Source](https://sqlite.org/howtocorrupt.html)
|
||||
- SQLite WAL forgiving of out-of-order writes except during checkpoint; `COMMIT` sync failure loses durability not consistency; checkpoint infrequently as defense; QNX `mmap` + WAL needs exclusive locking/no-mmap — [Source](https://sqlite.org/howtocorrupt.html)
|
||||
- Forum: with WAL expect journal replay on open after restore, uncommitted lost; for professional frozen-moment copy don't back up while changing; full WAL checkpoint then copy main file pristine if no checkpoint during copy (disable autocheckpoint temporarily) — [Source](https://sqlite.org/forum/forumpost/f173a78e2e?t=h)
|
||||
- Borg docs: avoid programs changing files during backup; LVM/ZFS snapshot or dump/stop DBs; shutdown VMs/containers first — [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)
|
||||
- Kopia actions support `before-folder/after-folder` and `before-snapshot-root/after-snapshot-root` scripts (e.g. `zfs snapshot` + mount then snapshot mount, `zfs destroy` after; `pg_dumpall`/sqlite dump) with `essential/optional/async` modes and timeout (default 5m) — [Source](https://github.com/kopia/kopia/blob/master/site/content/docs/Advanced/Actions/_index.md)
|
||||
- SafeKeep (rdiff-backup wrapper) noted as integrating with LVM and databases for consistent backups — [Source](https://wiki.archlinux.org/title/Synchronization_and_backup_programs)
|
||||
- SQLite: backup/restore while transaction active → copy mixes old/new → corrupt; safe via `sqlite3_rsync`, `VACUUM INTO`, backup API even on live DB; idle-file copy only safe with no tx in progress - [Source](https://sqlite.org/howtocorrupt.html)
|
||||
- SQLite: never move/delete/rename hot `-journal/-wal` after crash; mispairing (swap/overwrite/move journal, copy DB without journal, overwrite DB without deleting hot journal) likely corrupts; quiescent DB has no journal, only DB file matters - [Source](https://sqlite.org/howtocorrupt.html)
|
||||
- SQLite POSIX advisory-lock quirk: any thread `open/read/close` on DB file drops locks for all threads (close cancels locks); bypassing lib for backup read can corrupt; since 3.51.0 (2025-11-04) extra WAL defenses but not cure-all - never `close()` DB file while connections open even in other threads - [Source](https://sqlite.org/howtocorrupt.html)
|
||||
- SQLite: don't use two SQLite copies linked in one app (separate lock lists), don't mix locking protocols (POSIX vs dot-file/NFS), don't unlink/rename open DB (shared journal name → cross-recovery corruption, `SQLITE_WARNING` since 3.7.17), don't multi-link/symlink same file (wrong journal lookup; canonicalization since 3.10.0), don't carry connection across `fork()` - [Source](https://sqlite.org/howtocorrupt.html)
|
||||
- SQLite WAL forgiving of out-of-order writes except during checkpoint; `COMMIT` sync failure loses durability not consistency; checkpoint infrequently as defense; QNX `mmap` + WAL needs exclusive locking/no-mmap - [Source](https://sqlite.org/howtocorrupt.html)
|
||||
- Forum: with WAL expect journal replay on open after restore, uncommitted lost; for professional frozen-moment copy don't back up while changing; full WAL checkpoint then copy main file pristine if no checkpoint during copy (disable autocheckpoint temporarily) - [Source](https://sqlite.org/forum/forumpost/f173a78e2e?t=h)
|
||||
- Borg docs: avoid programs changing files during backup; LVM/ZFS snapshot or dump/stop DBs; shutdown VMs/containers first - [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)
|
||||
- Kopia actions support `before-folder/after-folder` and `before-snapshot-root/after-snapshot-root` scripts (e.g. `zfs snapshot` + mount then snapshot mount, `zfs destroy` after; `pg_dumpall`/sqlite dump) with `essential/optional/async` modes and timeout (default 5m) - [Source](https://github.com/kopia/kopia/blob/master/site/content/docs/Advanced/Actions/_index.md)
|
||||
- SafeKeep (rdiff-backup wrapper) noted as integrating with LVM and databases for consistent backups - [Source](https://wiki.archlinux.org/title/Synchronization_and_backup_programs)
|
||||
|
||||
### Inferences
|
||||
- Correct pattern for single-user Linux: `pre-action: pg_dumpall/sqlite VACUUM INTO/.backup → file under /home or /var/backups` then `borg/restic/kopia snapshot` includes dump; keep live DB files too but rely on dump for restore.
|
||||
|
||||
@@ -3,22 +3,22 @@
|
||||
## What are the canonical exclude directories and file patterns across backup tools (give concrete lists)?
|
||||
|
||||
### Takeaway
|
||||
Canonical Linux file-backup excludes converge on: virtual/kernel filesystems (`/proc`, `/sys`, `/dev`, `/run`), ephemeral (`/tmp`, `/var/tmp`, `/var/run`, `/var/lock`), mounts/media (`/mnt`, `/media`, `/lost+found`, `/swapfile`), package/cache/logs (`/var/cache/*`, `/var/log/*`, pacman cache), per-user caches/trash/thumbnails, and repro data (build/dependency dirs, VM/container storage) — implemented as explicit exclude-files in restic/borg plus `--exclude-caches`/`--one-file-system` and Kopia policies/`.kopiaignore`.
|
||||
Canonical Linux file-backup excludes converge on: virtual/kernel filesystems (`/proc`, `/sys`, `/dev`, `/run`), ephemeral (`/tmp`, `/var/tmp`, `/var/run`, `/var/lock`), mounts/media (`/mnt`, `/media`, `/lost+found`, `/swapfile`), package/cache/logs (`/var/cache/*`, `/var/log/*`, pacman cache), per-user caches/trash/thumbnails, and repro data (build/dependency dirs, VM/container storage) - implemented as explicit exclude-files in restic/borg plus `--exclude-caches`/`--one-file-system` and Kopia policies/`.kopiaignore`.
|
||||
|
||||
### Cited Findings
|
||||
- restic has no built-in default excludes; user supplies `--exclude`, `--iexclude`, `--exclude-file`, `--exclude-if-present foo`, `--exclude-caches` (CACHEDIR.TAG dirs), `--exclude-larger-than`, `--exclude-cloud-files`, plus `-x/--one-file-system` to stay on one filesystem — [Source](https://restic.readthedocs.io/en/v0.16.3/040_backup.html); also documented that excludes do not apply to explicitly-passed backup-source paths — [Source](https://github.com/restic/restic/blob/master/doc/040_backup.rst)
|
||||
- restic pattern syntax is Go `filepath.Match` + `**` for crossing `/`, matched on complete path components (`foo` matches `/dir1/foo/...` but not `/dir/foobar`), trailing `/` ignored, leading `/` anchors at root — [Source](https://restic.readthedocs.io/en/v0.16.3/040_backup.html)
|
||||
- borg `create` supports `-e/--exclude PATTERN`, `--exclude-from FILE`, `--exclude-caches` (CACHEDIR.TAG), `--exclude-if-present NAME`, `--keep-tag-files`, `-x/--one-file-system`, and respects `nodump` flag; recommended test is `borg create --list --dry-run` — [Source](https://borgbackup.readthedocs.io/en/1.0.5/usage.html); pattern styles are `fnmatch` (default for `--exclude`), `sh:`, `re:`, path prefix/full-match, with trailing-`/` meaning “keep dir, skip contents” — [Source](https://manpages.debian.org/testing/borgbackup/borg-patterns.1.en.html)
|
||||
- borg quickstart automation example excludes `--exclude-caches --exclude 'home/*/.cache/*' --exclude 'var/tmp/*'` when backing up `/etc`-style roots — [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)
|
||||
- Kopia has no static exclude-file; it uses policy `Ignore rules` + `Read ignore rules from files` (default `.kopiaignore`) + `Ignore cache directories: true` (default true, inherits global→host→user→path) + `Scan one filesystem only: true` — [Source](https://github.com/kopia/kopia/issues/3334); `kopia policy set` flags include `--add-ignore`, `--add-dot-ignore`, `--ignore-cache-dirs [true|false|inherit]`, `--ignore-dir-errors`, `--one-file-system`-equivalent `Scan one filesystem only` — [Source](https://kopia.io/docs/reference/command-line/common/policy-set/)
|
||||
- Kopia’s CACHEDIR.TAG-equivalent (`--ignore-cache-dirs`, default true) was added to mirror restic `--exclude-caches`, defaulting to ignore caches via policy — [Source](https://github.com/kopia/kopia/issues/564)
|
||||
- ArchWiki restic full-system `/etc/restic/excludes.txt` canonical list: `/data/**`, `/dev/**`, `/home/*/**/*.pyc`, `/home/*/**/__pycache__/**`, `/home/*/**/node_modules/**`, `/home/*/.cache/**`, `/home/*/.local/lib/python*/site-packages/**`, `/home/*/.mozilla/firefox/*/Cache/**`, `/lost+found/**`, `/media/**`, `/mnt/**`, `/proc/**`, `/root/**`, `/run/**`, `/swapfile`, `/sys/**`, `/tmp/**`, `/var/cache/**`, `/var/cache/pacman/pkg/**`, `/var/lib/docker/**`, `/var/lib/libvirt/**`, `/var/lock/**`, `/var/log/**`, `/var/run/**`, noting `--one-file-system` can replace `/proc`/`/run`/`/mnt` entries while preserving mountpoints — [Source](https://wiki.archlinux.org/title/Restic)
|
||||
- Debian restic-forum practitioner system list: `/media`, `/mnt`, `/cdrom`, `/proc`, `/sys`, `/dev`, `/run`, `/tmp`, `/var/run`, `/var/lock`, `/var/tmp`, `/lost+found`, `/swapfile`, `/var/cache/restic`, Steam `.../Steam/steamapps`, plus home temp items `.gvfs`, `.local/share/gvfs-metadata`, `.local/share/Trash`, `.cache`, `.dbus`, `.xsession-errors`, `.Xauthority`, `.gksu.lock`, `.local/share/flatpak/appstream` — [Source](https://forum.restic.net/t/what-gnu-linux-directories-to-exclude-from-backups/6653)
|
||||
- Extended community restic list adds: `nobackup/NoBackup`, `Downloads/*`, `VirtualBox VMs/`, `Dropbox/*`, `rclone/`, `snap/`, `.var/**/cache/`, `.local/share/docker/`, `.local/share/JetBrains/`, `.vscode/extensions/`, `.config/Code/`, `.m2/repository`, `target/*`, `build/*`, `node_modules/*`, `jspm_packages/*`, `web_modules/*`, `.npm`, `.coursier/cache`, `.sbt`, `.stack`, `.sdkman`, `.jdks`, `.eclipse`, `.venv-py3`, `.wine`, `.android`, `Android/Sdk`, `.gradle`, `.adobe`, `.macromedia`, `.thumbnails`, `.thunderbird/*/Cache`, `.mozilla/firefox/*/Cache|storage|minidumps|*.sqlite*`, `.config/**/Cache|GPUCache|ShaderCache`, `.config/chromium/Default/...History|Favicons|Storage|Cache`, `.local/share/baloo|zeitgeist|akonadi`, `.gnupg/rnd|random_seed|*.lock`, `.pulse*`, `.java/deployment/cache`, `.dropbox*` — [Source](https://forum.restic.net/t/what-gnu-linux-directories-to-exclude-from-backups/6653)
|
||||
- ArchWiki tar full-system guidance: exclude `/opt/backup/arch-full*` (own backups), `/tmp/*`, `/var/cache/pacman/pkg/`, boot from LiveCD with plain `chroot` (not `arch-chroot`) to avoid capturing tempfs/memory — [Source](https://wiki.archlinux.org/title/Full_system_backup_with_tar_(Italiano))
|
||||
- Guide summary: when backing up `/`, explicitly exclude `/proc`, `/sys`, `/dev`, `/run`, `/tmp` and bind/overlay mounts; when backing up only `/home`+`/etc` both borg/restic skip virtual FS by default — [Source](https://linuxjunkies.org/guides/back-up-with-borg-or-restic)
|
||||
- Kopia `.kopiaignore` syntax: one rule/line, `#` comment, `*` (any chars), `**` (any dirs), `?`, `[0-9]/[a-z]/[A-Z]/[abc]`, leading `/` roots rule, `!` negates; policy can point at alternate ignore-files — [Source](https://kopia.io/docs/advanced/kopiaignore/)
|
||||
- Time-Machine-ecosystem analogue (Arq honoring Apple TM exclusions): `.Trash`, `Library/Caches`, `Library/Logs`, `Library/Mail/.../Envelope Index*`, `Library/Safari/WebpageIcons.db`, `Library/Saved Application State`, `Library/iTunes/iPad Software Updates` — [Source](https://www.arqbackup.com/docs/arqbackup/pages/adding_folder.html)
|
||||
- restic has no built-in default excludes; user supplies `--exclude`, `--iexclude`, `--exclude-file`, `--exclude-if-present foo`, `--exclude-caches` (CACHEDIR.TAG dirs), `--exclude-larger-than`, `--exclude-cloud-files`, plus `-x/--one-file-system` to stay on one filesystem - [Source](https://restic.readthedocs.io/en/v0.16.3/040_backup.html); also documented that excludes do not apply to explicitly-passed backup-source paths - [Source](https://github.com/restic/restic/blob/master/doc/040_backup.rst)
|
||||
- restic pattern syntax is Go `filepath.Match` + `**` for crossing `/`, matched on complete path components (`foo` matches `/dir1/foo/...` but not `/dir/foobar`), trailing `/` ignored, leading `/` anchors at root - [Source](https://restic.readthedocs.io/en/v0.16.3/040_backup.html)
|
||||
- borg `create` supports `-e/--exclude PATTERN`, `--exclude-from FILE`, `--exclude-caches` (CACHEDIR.TAG), `--exclude-if-present NAME`, `--keep-tag-files`, `-x/--one-file-system`, and respects `nodump` flag; recommended test is `borg create --list --dry-run` - [Source](https://borgbackup.readthedocs.io/en/1.0.5/usage.html); pattern styles are `fnmatch` (default for `--exclude`), `sh:`, `re:`, path prefix/full-match, with trailing-`/` meaning “keep dir, skip contents” - [Source](https://manpages.debian.org/testing/borgbackup/borg-patterns.1.en.html)
|
||||
- borg quickstart automation example excludes `--exclude-caches --exclude 'home/*/.cache/*' --exclude 'var/tmp/*'` when backing up `/etc`-style roots - [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)
|
||||
- Kopia has no static exclude-file; it uses policy `Ignore rules` + `Read ignore rules from files` (default `.kopiaignore`) + `Ignore cache directories: true` (default true, inherits global→host→user→path) + `Scan one filesystem only: true` - [Source](https://github.com/kopia/kopia/issues/3334); `kopia policy set` flags include `--add-ignore`, `--add-dot-ignore`, `--ignore-cache-dirs [true|false|inherit]`, `--ignore-dir-errors`, `--one-file-system`-equivalent `Scan one filesystem only` - [Source](https://kopia.io/docs/reference/command-line/common/policy-set/)
|
||||
- Kopia’s CACHEDIR.TAG-equivalent (`--ignore-cache-dirs`, default true) was added to mirror restic `--exclude-caches`, defaulting to ignore caches via policy - [Source](https://github.com/kopia/kopia/issues/564)
|
||||
- ArchWiki restic full-system `/etc/restic/excludes.txt` canonical list: `/data/**`, `/dev/**`, `/home/*/**/*.pyc`, `/home/*/**/__pycache__/**`, `/home/*/**/node_modules/**`, `/home/*/.cache/**`, `/home/*/.local/lib/python*/site-packages/**`, `/home/*/.mozilla/firefox/*/Cache/**`, `/lost+found/**`, `/media/**`, `/mnt/**`, `/proc/**`, `/root/**`, `/run/**`, `/swapfile`, `/sys/**`, `/tmp/**`, `/var/cache/**`, `/var/cache/pacman/pkg/**`, `/var/lib/docker/**`, `/var/lib/libvirt/**`, `/var/lock/**`, `/var/log/**`, `/var/run/**`, noting `--one-file-system` can replace `/proc`/`/run`/`/mnt` entries while preserving mountpoints - [Source](https://wiki.archlinux.org/title/Restic)
|
||||
- Debian restic-forum practitioner system list: `/media`, `/mnt`, `/cdrom`, `/proc`, `/sys`, `/dev`, `/run`, `/tmp`, `/var/run`, `/var/lock`, `/var/tmp`, `/lost+found`, `/swapfile`, `/var/cache/restic`, Steam `.../Steam/steamapps`, plus home temp items `.gvfs`, `.local/share/gvfs-metadata`, `.local/share/Trash`, `.cache`, `.dbus`, `.xsession-errors`, `.Xauthority`, `.gksu.lock`, `.local/share/flatpak/appstream` - [Source](https://forum.restic.net/t/what-gnu-linux-directories-to-exclude-from-backups/6653)
|
||||
- Extended community restic list adds: `nobackup/NoBackup`, `Downloads/*`, `VirtualBox VMs/`, `Dropbox/*`, `rclone/`, `snap/`, `.var/**/cache/`, `.local/share/docker/`, `.local/share/JetBrains/`, `.vscode/extensions/`, `.config/Code/`, `.m2/repository`, `target/*`, `build/*`, `node_modules/*`, `jspm_packages/*`, `web_modules/*`, `.npm`, `.coursier/cache`, `.sbt`, `.stack`, `.sdkman`, `.jdks`, `.eclipse`, `.venv-py3`, `.wine`, `.android`, `Android/Sdk`, `.gradle`, `.adobe`, `.macromedia`, `.thumbnails`, `.thunderbird/*/Cache`, `.mozilla/firefox/*/Cache|storage|minidumps|*.sqlite*`, `.config/**/Cache|GPUCache|ShaderCache`, `.config/chromium/Default/...History|Favicons|Storage|Cache`, `.local/share/baloo|zeitgeist|akonadi`, `.gnupg/rnd|random_seed|*.lock`, `.pulse*`, `.java/deployment/cache`, `.dropbox*` - [Source](https://forum.restic.net/t/what-gnu-linux-directories-to-exclude-from-backups/6653)
|
||||
- ArchWiki tar full-system guidance: exclude `/opt/backup/arch-full*` (own backups), `/tmp/*`, `/var/cache/pacman/pkg/`, boot from LiveCD with plain `chroot` (not `arch-chroot`) to avoid capturing tempfs/memory - [Source](https://wiki.archlinux.org/title/Full_system_backup_with_tar_(Italiano))
|
||||
- Guide summary: when backing up `/`, explicitly exclude `/proc`, `/sys`, `/dev`, `/run`, `/tmp` and bind/overlay mounts; when backing up only `/home`+`/etc` both borg/restic skip virtual FS by default - [Source](https://linuxjunkies.org/guides/back-up-with-borg-or-restic)
|
||||
- Kopia `.kopiaignore` syntax: one rule/line, `#` comment, `*` (any chars), `**` (any dirs), `?`, `[0-9]/[a-z]/[A-Z]/[abc]`, leading `/` roots rule, `!` negates; policy can point at alternate ignore-files - [Source](https://kopia.io/docs/advanced/kopiaignore/)
|
||||
- Time-Machine-ecosystem analogue (Arq honoring Apple TM exclusions): `.Trash`, `Library/Caches`, `Library/Logs`, `Library/Mail/.../Envelope Index*`, `Library/Safari/WebpageIcons.db`, `Library/Saved Application State`, `Library/iTunes/iPad Software Updates` - [Source](https://www.arqbackup.com/docs/arqbackup/pages/adding_folder.html)
|
||||
|
||||
### Inferences
|
||||
- There is no single vendor “default exclude file” for restic/borg on Linux; the de-facto standard is the ArchWiki + forum lists above plus `--exclude-caches` and `--one-file-system`.
|
||||
@@ -34,16 +34,16 @@ Canonical Linux file-backup excludes converge on: virtual/kernel filesystems (`/
|
||||
Correctness excludes (virtual FS, sockets/FIFOs/devices, live DB/VM/container backing stores, cloud-online-only stubs) prevent hangs, errors, or unrestorable data; size excludes (caches, trash, thumbnails, logs, browser profiles, package caches, media/Steam) prevent bloat; reproducibility excludes (dependency/build trees) are safe to drop because lockfiles+manifests rebuild them.
|
||||
|
||||
### Cited Findings
|
||||
- `/proc` and `/sys` are virtual/kernel pseudo-filesystems (proc exposes live kernel/process state, `proc_sys` exposes tunable sysctls; size often reported 0, constantly changing) — backing them up captures no stable data — [Source](https://docs.redhat.com/en/documentation/red_hat_enterprise_linux/4/html/reference_guide/ch-proc); [Source](https://man7.org/linux/man-pages/man5/proc_sys.5.html)
|
||||
- restic forum guidance: use `--exclude-caches` to auto-exclude CACHEDIR.TAG-marked dirs for “non permanent system files”; full-system threads treat only a short path list as needing exclusion — [Source](https://forum.restic.net/t/linux-exclusion-of-non-permanent-system-files/6629)
|
||||
- borg by default does not open/read block/char devices or FIFOs; `--read-special` is required to force reading them as regular files (and to follow symlinks to them) — [Source](https://www.systutorials.com/linux-manual-page-1-borg)
|
||||
- restic 0.17.0+ fix “Exclude irregular files from backups” (sockets, FIFOs, devices skipped) confirms prior irregular-file handling was a correctness fix — [Source](https://github.com/restic/restic/releases)
|
||||
- Kopia policy has explicit `--ignore-dir-errors` to tolerate unreadable/locked dirs during traversal, separate from ignore-rules; issue reports show `fuse` mounts (e.g. `seafile/fuse`) still `lstat`-probed during estimation even when ignored, motivating explicit excludes for FUSE/locked paths — [Source](https://github.com/kopia/kopia/issues/3334); [Source](https://kopia.io/docs/reference/command-line/common/policy-set/)
|
||||
- Docker storage (`/var/lib/docker`: `overlay2/`, `image/`, `buildkit/`, `volumes/`, `containers/`) is large, churn-heavy, overlay-mounted and often locked/inconsistent without daemon quiesce; community Arch/restic lists therefore exclude `/var/lib/docker/**` and `/var/lib/libvirt/**`, and Proxmox `vzdump --exclude-path` examples exclude `/var/lib/docker/` alongside `/run/`, `/dev/shm`, `/dev/fuse` — [Source](https://wiki.archlinux.org/title/Restic); [Source](https://blog.devops.dev/docker-cleanup-and-relocate-81d2dcd32b8a); [Source](https://forum.proxmox.com/threads/backup-and-exclude-path.125966)
|
||||
- Docker’s own backup guidance says file-copy of the VM disk (`Docker.raw`/`docker_data.vhdx`) or `/var/lib/docker` requires Docker fully stopped; otherwise use `docker save`/`docker pull` + volume dump/restore procedures — [Source](https://docs.docker.com/desktop/settings-and-maintenance/backup-and-restore.md)
|
||||
- Veeam (VM-image backup reference) automatically excludes VM log files to cut size/time, and supports excluding swap files, deleted-block (BitLooker) data, and per-disk/template excludes for the same size/correctness reasons — [Source](https://helpcenter.veeam.com/docs/vbr/userguide/data_exclusion.html)
|
||||
- Size-category examples: pacman cache `/var/cache/pacman/pkg/`, `/var/cache/*`, `/var/log/*`, `/var/tmp/*`, browser `Cache/GPUCache/ShaderCache`, `thunderbird/*/Cache`, `~/.thumbnails`, `~/.local/share/Trash`, Steam `steamapps`, `~/Downloads/*`, Dropbox/rclone replicas — all excluded for bloat, not correctness — [Source](https://forum.restic.net/t/what-gnu-linux-directories-to-exclude-from-backups/6653); [Source](https://wiki.archlinux.org/title/Restic)
|
||||
- Reproducibility-category examples: `node_modules/*`, `jspm_packages/*`, `web_modules/*`, `target/*`, `build/*`, `~/.m2/repository`, `~/.ivy2`, `~/.gradle`, `~/.coursier/cache`, `~/.sbt`, `~/.stack`, `__pycache__`, `*.pyc`, `~/.local/lib/python*/site-packages/**`, `~/.npm`, `~/.pkg-cache/`, `~/.sdkman/`, `.venv-py3` — explicitly listed as rebuildable caches/toolchains — [Source](https://forum.restic.net/t/what-gnu-linux-directories-to-exclude-from-backups/6653); webdev backup tooling states “node_modules is never included … rebuild with yarn/npm install” while always keeping `.env`, `yarn.lock`, `package-lock.json` — [Source](https://github.com/ICJIA/backup-webdev/blob/main/README.md)
|
||||
- `/proc` and `/sys` are virtual/kernel pseudo-filesystems (proc exposes live kernel/process state, `proc_sys` exposes tunable sysctls; size often reported 0, constantly changing) - backing them up captures no stable data - [Source](https://docs.redhat.com/en/documentation/red_hat_enterprise_linux/4/html/reference_guide/ch-proc); [Source](https://man7.org/linux/man-pages/man5/proc_sys.5.html)
|
||||
- restic forum guidance: use `--exclude-caches` to auto-exclude CACHEDIR.TAG-marked dirs for “non permanent system files”; full-system threads treat only a short path list as needing exclusion - [Source](https://forum.restic.net/t/linux-exclusion-of-non-permanent-system-files/6629)
|
||||
- borg by default does not open/read block/char devices or FIFOs; `--read-special` is required to force reading them as regular files (and to follow symlinks to them) - [Source](https://www.systutorials.com/linux-manual-page-1-borg)
|
||||
- restic 0.17.0+ fix “Exclude irregular files from backups” (sockets, FIFOs, devices skipped) confirms prior irregular-file handling was a correctness fix - [Source](https://github.com/restic/restic/releases)
|
||||
- Kopia policy has explicit `--ignore-dir-errors` to tolerate unreadable/locked dirs during traversal, separate from ignore-rules; issue reports show `fuse` mounts (e.g. `seafile/fuse`) still `lstat`-probed during estimation even when ignored, motivating explicit excludes for FUSE/locked paths - [Source](https://github.com/kopia/kopia/issues/3334); [Source](https://kopia.io/docs/reference/command-line/common/policy-set/)
|
||||
- Docker storage (`/var/lib/docker`: `overlay2/`, `image/`, `buildkit/`, `volumes/`, `containers/`) is large, churn-heavy, overlay-mounted and often locked/inconsistent without daemon quiesce; community Arch/restic lists therefore exclude `/var/lib/docker/**` and `/var/lib/libvirt/**`, and Proxmox `vzdump --exclude-path` examples exclude `/var/lib/docker/` alongside `/run/`, `/dev/shm`, `/dev/fuse` - [Source](https://wiki.archlinux.org/title/Restic); [Source](https://blog.devops.dev/docker-cleanup-and-relocate-81d2dcd32b8a); [Source](https://forum.proxmox.com/threads/backup-and-exclude-path.125966)
|
||||
- Docker’s own backup guidance says file-copy of the VM disk (`Docker.raw`/`docker_data.vhdx`) or `/var/lib/docker` requires Docker fully stopped; otherwise use `docker save`/`docker pull` + volume dump/restore procedures - [Source](https://docs.docker.com/desktop/settings-and-maintenance/backup-and-restore.md)
|
||||
- Veeam (VM-image backup reference) automatically excludes VM log files to cut size/time, and supports excluding swap files, deleted-block (BitLooker) data, and per-disk/template excludes for the same size/correctness reasons - [Source](https://helpcenter.veeam.com/docs/vbr/userguide/data_exclusion.html)
|
||||
- Size-category examples: pacman cache `/var/cache/pacman/pkg/`, `/var/cache/*`, `/var/log/*`, `/var/tmp/*`, browser `Cache/GPUCache/ShaderCache`, `thunderbird/*/Cache`, `~/.thumbnails`, `~/.local/share/Trash`, Steam `steamapps`, `~/Downloads/*`, Dropbox/rclone replicas - all excluded for bloat, not correctness - [Source](https://forum.restic.net/t/what-gnu-linux-directories-to-exclude-from-backups/6653); [Source](https://wiki.archlinux.org/title/Restic)
|
||||
- Reproducibility-category examples: `node_modules/*`, `jspm_packages/*`, `web_modules/*`, `target/*`, `build/*`, `~/.m2/repository`, `~/.ivy2`, `~/.gradle`, `~/.coursier/cache`, `~/.sbt`, `~/.stack`, `__pycache__`, `*.pyc`, `~/.local/lib/python*/site-packages/**`, `~/.npm`, `~/.pkg-cache/`, `~/.sdkman/`, `.venv-py3` - explicitly listed as rebuildable caches/toolchains - [Source](https://forum.restic.net/t/what-gnu-linux-directories-to-exclude-from-backups/6653); webdev backup tooling states “node_modules is never included … rebuild with yarn/npm install” while always keeping `.env`, `yarn.lock`, `package-lock.json` - [Source](https://github.com/ICJIA/backup-webdev/blob/main/README.md)
|
||||
|
||||
### Inferences
|
||||
- Rule of thumb: if path is kernel-generated (`/proc`, `/sys`), socket/FIFO/device, actively memory-mapped/locked (browser `*-shm/-wal`, DB live files, running container layers, FUSE), or a mountpoint for another filesystem, exclusion is for correctness; if it is cache/log/trash/thumbnail/package-download, it is for size; if `npm install`/`pip install`/`cargo fetch`/`mvn` recreates it, it is for reproducibility.
|
||||
@@ -56,18 +56,18 @@ Correctness excludes (virtual FS, sockets/FIFOs/devices, live DB/VM/container ba
|
||||
## How do tools treat dot-directories (.cache, .git, .venv) and per-project gitignore?
|
||||
|
||||
### Takeaway
|
||||
Backup tools do not read `.gitignore` by default; they rely on explicit patterns, CACHEDIR.TAG/`--exclude-caches`, and `.kopiaignore`/policy rules — so `.cache` is usually globally excluded, `.git` is usually kept (small, valuable history) unless deliberately dropped, and `.venv`/`node_modules`/`target` must be explicitly excluded per-project or via `exclude-if-present` markers.
|
||||
Backup tools do not read `.gitignore` by default; they rely on explicit patterns, CACHEDIR.TAG/`--exclude-caches`, and `.kopiaignore`/policy rules - so `.cache` is usually globally excluded, `.git` is usually kept (small, valuable history) unless deliberately dropped, and `.venv`/`node_modules`/`target` must be explicitly excluded per-project or via `exclude-if-present` markers.
|
||||
|
||||
### Cited Findings
|
||||
- `~/.cache` is safe to drop (name indicates cached data; many tools exclude cache/trash by default); Arch forum explicitly approves excluding whole `~/.cache`, implemented via borg `--exclude-caches` + CACHEDIR.TAG spec — [Source](https://bbs.archlinux.org/viewtopic.php?id=231281)
|
||||
- restic `--exclude-caches` only skips dirs containing `CACHEDIR.TAG` (keeps the tag file); `--exclude-if-present foo` skips contents of any dir containing marker `foo` (e.g. `.nobackup`, `CACHEDIR.TAG`, custom sentinels) — [Source](https://restic.readthedocs.io/en/v0.16.3/040_backup.html)
|
||||
- borg `--exclude-caches`/`--exclude-if-present`/`--keep-tag-files` semantics are identical (skip CACHEDIR.TAG dirs; optionally retain tag files) — [Source](https://borgbackup.readthedocs.io/en/1.0.5/usage.html)
|
||||
- Kopia `Ignore cache directories: true` is the default global policy (inherits unless overridden), plus per-directory `.kopiaignore` and policy ignore-rules using gitignore-like globs (`*.dat`, `/logs/*`, `tmp.db`, `**/logs/**`, `!` negation) — [Source](https://github.com/kopia/kopia/issues/3334); [Source](https://kopia.io/docs/advanced/kopiaignore/)
|
||||
- Kopia ignore-rule inheritance is surprising: defining child-path `add-ignore` can shadow global ignores (open issue: “Ignore rules not being inherited when child adds its own ignores”), so global `.cache`/`cache` rules must be repeated or merged carefully — [Source](https://github.com/kopia/kopia/issues/4155)
|
||||
- Developer-oriented backup/index defaults treat dot-dirs as noise to skip: Dank index defaults exclude `.git`, `.hg`, `.svn`, `.cache`, `.npm`, `.yarn`, `.venv`/`venv`, `.tox`, `.pytest_cache`, `__pycache__`, `.gradle`, `.m2`, `.cargo`, `.idea`, `.vscode`, `node_modules`, `target`, `dist/build/out` — [Source](https://danklinux.com/docs/1.4/danksearch/configuration); SmartBackup auto-skips `node_modules`, `venv`, `__pycache__`, `.git` for 10x speed/size — [Source](https://github.com/CodingWithMK/smartbackup_file-backup-automation)
|
||||
- Counter-practice for `.git`: default backup guidance keeps `.git` (history is irreplaceable, usually small vs `node_modules`); Backblaze-mac customization explicitly adds separate opt-in rules to drop `node_modules/` and `.git/` because neither is excluded by default — [Source](https://gist.github.com/nickcernis/bb4bd43a44efd73b87d857e29b1d5b96)
|
||||
- `tmexclude` watches the filesystem to continually re-apply Time-Machine exclusions for `node_modules`, `target`, etc., because new projects constantly recreate them — showing per-project `.gitignore` alone does not stop backup tools from capturing them — [Source](https://github.com/PhotonQuantum/tmexclude)
|
||||
- Kopia FAQ states ignored paths come from policy or `.kopiaignore` files, not from `.gitignore` — [Source](https://github.com/kopia/kopia/blob/master/site/content/docs/FAQs/_index.md)
|
||||
- `~/.cache` is safe to drop (name indicates cached data; many tools exclude cache/trash by default); Arch forum explicitly approves excluding whole `~/.cache`, implemented via borg `--exclude-caches` + CACHEDIR.TAG spec - [Source](https://bbs.archlinux.org/viewtopic.php?id=231281)
|
||||
- restic `--exclude-caches` only skips dirs containing `CACHEDIR.TAG` (keeps the tag file); `--exclude-if-present foo` skips contents of any dir containing marker `foo` (e.g. `.nobackup`, `CACHEDIR.TAG`, custom sentinels) - [Source](https://restic.readthedocs.io/en/v0.16.3/040_backup.html)
|
||||
- borg `--exclude-caches`/`--exclude-if-present`/`--keep-tag-files` semantics are identical (skip CACHEDIR.TAG dirs; optionally retain tag files) - [Source](https://borgbackup.readthedocs.io/en/1.0.5/usage.html)
|
||||
- Kopia `Ignore cache directories: true` is the default global policy (inherits unless overridden), plus per-directory `.kopiaignore` and policy ignore-rules using gitignore-like globs (`*.dat`, `/logs/*`, `tmp.db`, `**/logs/**`, `!` negation) - [Source](https://github.com/kopia/kopia/issues/3334); [Source](https://kopia.io/docs/advanced/kopiaignore/)
|
||||
- Kopia ignore-rule inheritance is surprising: defining child-path `add-ignore` can shadow global ignores (open issue: “Ignore rules not being inherited when child adds its own ignores”), so global `.cache`/`cache` rules must be repeated or merged carefully - [Source](https://github.com/kopia/kopia/issues/4155)
|
||||
- Developer-oriented backup/index defaults treat dot-dirs as noise to skip: Dank index defaults exclude `.git`, `.hg`, `.svn`, `.cache`, `.npm`, `.yarn`, `.venv`/`venv`, `.tox`, `.pytest_cache`, `__pycache__`, `.gradle`, `.m2`, `.cargo`, `.idea`, `.vscode`, `node_modules`, `target`, `dist/build/out` - [Source](https://danklinux.com/docs/1.4/danksearch/configuration); SmartBackup auto-skips `node_modules`, `venv`, `__pycache__`, `.git` for 10x speed/size - [Source](https://github.com/CodingWithMK/smartbackup_file-backup-automation)
|
||||
- Counter-practice for `.git`: default backup guidance keeps `.git` (history is irreplaceable, usually small vs `node_modules`); Backblaze-mac customization explicitly adds separate opt-in rules to drop `node_modules/` and `.git/` because neither is excluded by default - [Source](https://gist.github.com/nickcernis/bb4bd43a44efd73b87d857e29b1d5b96)
|
||||
- `tmexclude` watches the filesystem to continually re-apply Time-Machine exclusions for `node_modules`, `target`, etc., because new projects constantly recreate them - showing per-project `.gitignore` alone does not stop backup tools from capturing them - [Source](https://github.com/PhotonQuantum/tmexclude)
|
||||
- Kopia FAQ states ignored paths come from policy or `.kopiaignore` files, not from `.gitignore` - [Source](https://github.com/kopia/kopia/blob/master/site/content/docs/FAQs/_index.md)
|
||||
|
||||
### Inferences
|
||||
- Best practice: global-exclude `~/.cache`, `~/.npm`, `~/.cargo`, `~/.mozilla/.../Cache`, `Trash`, `thumbnails`; per-source exclude `node_modules/`, `.venv/venv/`, `__pycache__/`, `target/`, `build/dist/`, optionally via `exclude-if-present` markers (e.g. drop marker file `CACHEDIR.TAG` or `.nobackup` in roots) rather than hoping tools honor `.gitignore`.
|
||||
@@ -75,7 +75,7 @@ Backup tools do not read `.gitignore` by default; they rely on explicit patterns
|
||||
|
||||
### Gaps
|
||||
- No evidence found that restic/borg/Kopia natively ingest `.gitignore` in 2026 builds; if any wrapper does, it was not in primary docs searched.
|
||||
- Best handling of `.venv` that contains non-recreatable local edits (pip `-e` installs, manual patches) is unresolved — exclusion assumes venv is truly disposable.
|
||||
- Best handling of `.venv` that contains non-recreatable local edits (pip `-e` installs, manual patches) is unresolved - exclusion assumes venv is truly disposable.
|
||||
|
||||
## What breaks when people back up dependency/build trees anyway?
|
||||
|
||||
@@ -83,15 +83,15 @@ Backup tools do not read `.gitignore` by default; they rely on explicit patterns
|
||||
Backing up `node_modules`, `venv`, `target/build`, `site-packages`, and container/VM stores inflates size, scan time, and snapshot churn, defeats deduplication (thousands of tiny files, hardlinks, changing mtimes/hashes), risks unrestorable or inconsistent restores (native bindings, symlinks, absolute paths, locked DB/pages), and can leak secrets or break privacy.
|
||||
|
||||
### Cited Findings
|
||||
- SmartBackup’s premise is that naive `Documents` backup “waits hours because of massive node_modules or venvs”; skipping them yields ~10x faster/smaller backups, with incremental manifest+hash tracking otherwise churning on every dependency touch — [Source](https://github.com/CodingWithMK/smartbackup_file-backup-automation)
|
||||
- Webdev backup defaults state `node_modules` excluded from full/incremental/differential/quick backups to save space; restore requires `yarn/npm install`, while lockfiles+`.env` are retained to reconstruct — [Source](https://github.com/ICJIA/backup-webdev/blob/main/README.md)
|
||||
- Rust/JS `target/` and `node_modules/` are the canonical Time-Machine-bloat examples requiring a watcher (`tmexclude`) because they reappear per `npm install`/`cargo build` and would otherwise be re-captured every snapshot — [Source](https://github.com/PhotonQuantum/tmexclude)
|
||||
- Python `site-packages` (`~/.local/lib/python*/site-packages/**`), `__pycache__`, `*.pyc`, and per-project `.venv` are listed alongside `node_modules` in Arch/restic excludes as regenerable interpreter artifacts — [Source](https://wiki.archlinux.org/title/Restic)
|
||||
- `pnpm clean/purge` exists precisely because `node_modules` contents (plus virtual-store) are disposable and safely removable via Node-aware deletion handling junctions correctly — [Source](https://pnpm.io/next/cli/clean)
|
||||
- File-level backup of `/var/lib/docker` captures `overlay2` diffs, `image/`, `buildkit` cache (often 10s of GB, e.g. 20GB build cache reclaimable via `docker system prune`) that are host-specific, layer-duplicated, and unrestorable by plain copy while daemon runs; correct path is `docker save/load`, registry push/pull, and volume dumps — [Source](https://blog.devops.dev/docker-cleanup-and-relocate-81d2dcd32b8a); [Source](https://docs.docker.com/desktop/settings-and-maintenance/backup-and-restore.md)
|
||||
- Locked/inconsistent captures produce backup errors or silent corruption: Kopia `seafile/fuse` example shows even explicitly ignored FUSE paths trigger `lstat` errors/failed snapshots unless fully avoided; `--ignore-dir-errors` merely masks valid errors — [Source](https://github.com/kopia/kopia/issues/3334)
|
||||
- Browser/IDE heavy dirs (`.config/Code/`, `.vscode/extensions/`, `.config/coc/extensions/`, `.mozilla/.../extensions`, `sonarlint/plugins`, `.coursier/cache`, `.m2/repository`) are called out as “heavy JARs/caches” that bloat snapshots while preferences should sync via cloud/accounts instead — [Source](https://forum.restic.net/t/what-gnu-linux-directories-to-exclude-from-backups/6653)
|
||||
- `~/.cache`-style trees and `node_modules` contain huge small-file counts that slow file-walk, inflate metadata/index, and reduce dedup/compression efficiency vs backing up manifests/lockfiles — rationale behind Dank/SmartBackup default-skips — [Source](https://danklinux.com/docs/1.4/danksearch/configuration); [Source](https://github.com/CodingWithMK/smartbackup_file-backup-automation)
|
||||
- SmartBackup’s premise is that naive `Documents` backup “waits hours because of massive node_modules or venvs”; skipping them yields ~10x faster/smaller backups, with incremental manifest+hash tracking otherwise churning on every dependency touch - [Source](https://github.com/CodingWithMK/smartbackup_file-backup-automation)
|
||||
- Webdev backup defaults state `node_modules` excluded from full/incremental/differential/quick backups to save space; restore requires `yarn/npm install`, while lockfiles+`.env` are retained to reconstruct - [Source](https://github.com/ICJIA/backup-webdev/blob/main/README.md)
|
||||
- Rust/JS `target/` and `node_modules/` are the canonical Time-Machine-bloat examples requiring a watcher (`tmexclude`) because they reappear per `npm install`/`cargo build` and would otherwise be re-captured every snapshot - [Source](https://github.com/PhotonQuantum/tmexclude)
|
||||
- Python `site-packages` (`~/.local/lib/python*/site-packages/**`), `__pycache__`, `*.pyc`, and per-project `.venv` are listed alongside `node_modules` in Arch/restic excludes as regenerable interpreter artifacts - [Source](https://wiki.archlinux.org/title/Restic)
|
||||
- `pnpm clean/purge` exists precisely because `node_modules` contents (plus virtual-store) are disposable and safely removable via Node-aware deletion handling junctions correctly - [Source](https://pnpm.io/next/cli/clean)
|
||||
- File-level backup of `/var/lib/docker` captures `overlay2` diffs, `image/`, `buildkit` cache (often 10s of GB, e.g. 20GB build cache reclaimable via `docker system prune`) that are host-specific, layer-duplicated, and unrestorable by plain copy while daemon runs; correct path is `docker save/load`, registry push/pull, and volume dumps - [Source](https://blog.devops.dev/docker-cleanup-and-relocate-81d2dcd32b8a); [Source](https://docs.docker.com/desktop/settings-and-maintenance/backup-and-restore.md)
|
||||
- Locked/inconsistent captures produce backup errors or silent corruption: Kopia `seafile/fuse` example shows even explicitly ignored FUSE paths trigger `lstat` errors/failed snapshots unless fully avoided; `--ignore-dir-errors` merely masks valid errors - [Source](https://github.com/kopia/kopia/issues/3334)
|
||||
- Browser/IDE heavy dirs (`.config/Code/`, `.vscode/extensions/`, `.config/coc/extensions/`, `.mozilla/.../extensions`, `sonarlint/plugins`, `.coursier/cache`, `.m2/repository`) are called out as “heavy JARs/caches” that bloat snapshots while preferences should sync via cloud/accounts instead - [Source](https://forum.restic.net/t/what-gnu-linux-directories-to-exclude-from-backups/6653)
|
||||
- `~/.cache`-style trees and `node_modules` contain huge small-file counts that slow file-walk, inflate metadata/index, and reduce dedup/compression efficiency vs backing up manifests/lockfiles - rationale behind Dank/SmartBackup default-skips - [Source](https://danklinux.com/docs/1.4/danksearch/configuration); [Source](https://github.com/CodingWithMK/smartbackup_file-backup-automation)
|
||||
|
||||
### Inferences
|
||||
- Failure modes to expect if you include them anyway: (1) backup never finishes or blows retention/bandwidth; (2) every `npm/cargo/pip` run creates a large new snapshot despite no user-data change; (3) restore on another machine/arch breaks native modules/symlinks/permissions; (4) secrets in `.npmrc`, venv activate scripts, or vendored `.env` leak into long-retained snapshots.
|
||||
|
||||
@@ -3,32 +3,32 @@
|
||||
## Which systems trigger retention from disk usage (e.g. "keep usage below X%", "delete oldest when over Y%"), and what exact thresholds do they use?
|
||||
|
||||
### Takeaway
|
||||
Only Apple Time Machine makes capacity the primary retention driver ("delete oldest when full"); Veeam, S3, Hetzner Storage Box, ZFS/sanoid and Borg/restic all use time/count-based retention as primary with capacity handled by monitoring, placement, or manual pruning — no "keep usage below X%" knob exists in most of them, and where percentage thresholds exist they are health-warning or performance floors (ZFS 20%/10%), not auto-delete triggers.
|
||||
Only Apple Time Machine makes capacity the primary retention driver ("delete oldest when full"); Veeam, S3, Hetzner Storage Box, ZFS/sanoid and Borg/restic all use time/count-based retention as primary with capacity handled by monitoring, placement, or manual pruning - no "keep usage below X%" knob exists in most of them, and where percentage thresholds exist they are health-warning or performance floors (ZFS 20%/10%), not auto-delete triggers.
|
||||
|
||||
### Cited Findings
|
||||
- Time Machine "as your backup disk fills up, Time Machine deletes older backups to make room for new ones" with no published percentage — deletion is on-demand pre-backup, not a watermark — [Source](https://support.apple.com/guide/mac-help/if-the-time-machine-backup-disk-is-full-mh15137/mac)
|
||||
- Time Machine backupd log shows two-phase scheme: "Starting pre-backup thinning: 53.57 GB requested (including padding)" then "No expired backups exist - deleting oldest backups to make room"; post-backup thinning separately expires hourly/daily backups — [Source](https://serverfault.com/posts/39310/revisions)
|
||||
- Time Machine standard thinning ladder is hourly backups >24h old (keep first-of-day), daily backups >30 days old (keep first-of-week), then oldest-weekly deletion only when space is needed — [Source](https://discussions.apple.com/thread/251224074)
|
||||
- Time Machine per-destination quota via `tmutil setquota DESTINATION_ID QUOTA_IN_GB` caps a destination (e.g. 500 GB); quota "takes effect on the next backup, at which point older snapshots get thinned to fit" — [Source](https://superuser.com/questions/445579/how-do-i-trim-time-machine-backup-history)
|
||||
- `tmutil thinlocalsnapshots mount_point [purge_amount] [urgency]`: "tmutil will attempt (with urgency level 1-4) to reclaim purge_amount in bytes by thinning snapshots"; urgency 4 is most aggressive, e.g. `thinlocalsnapshots / 10000000000 4` to reclaim ~10 GB — [Source](https://ss64.com/mac/tmutil.html); same signature confirmed in Apple developer thread — [Source](https://developer.apple.com/forums/thread/81171)
|
||||
- Local APFS snapshots are purgeable space "automatically reclaimed as needed" but with no user-visible percentage; when reclamation lags users must thin manually (e.g. `thinlocalsnapshots / 100000000000 4` ≈ 100 GB at urgency 4) — [Source](https://discussions.apple.com/thread/252655312)
|
||||
- Veeam retention is count-based restore points / GFS, not capacity-based: "Retention policy defines the number of restore points to keep on your performance extents and capacity extents"; earliest restore point removed from chain, blocks purged from capacity tier on next offload/copy session — [Source](https://helpcenter.veeam.com/docs/vbr/userguide/capacity_tier_retention.html)
|
||||
- Veeam Scale-Out Backup Repository (SOBR) has no auto-delete-on-full: "if the extents of your scale-out backup repository run out of space, you can add a new extent"; free space on the new extent is added to SOBR capacity — [Source](https://helpcenter.veeam.com/docs/vbr/userguide/backup_repository_sobr.html)
|
||||
- Veeam SOBR placement prefers the extent with fewest chains, breaking ties by most free space; "priority is always to complete a backup" even if that violates the Data-Locality placement policy by spilling an incremental to another extent — [Source](https://veeam-best-practices-guide-v9.readthedocs.io/resource_planning/repository_sobr.html)
|
||||
- Veeam ONE / MP capacity reports use a configurable "Repository Free Space (%)" forecast threshold (worked example 30%) to flag repositories that "will run out of space", i.e. monitoring/alerting rather than enforcement — [Source](https://helpcenter.veeam.com/docs/mp/reports/capacity_planning_for_backup_repositories.html?ver=9a)
|
||||
- S3 Lifecycle has no capacity trigger at all: rules are `Days`/`Date`/`NoncurrentVersionExpiration` per prefix/tag (e.g. transition after 365 days, expire after 3650 days); S3 "quotas" are counts (buckets, access points), not bytes — [Source](https://docs.aws.amazon.com/AmazonS3/latest/API/API_LifecycleRule.html); expiration example — [Source](https://docs.amazonaws.cn/en_us/AmazonS3/latest/userguide/lifecycle-configuration-examples.md)
|
||||
- Hetzner Storage Box quota is the fixed plan size (BX11/BX21/BX31/BX41); snapshots "consume storage space from your Storage Box's storage capacity" alongside live data (`/.zfs/snapshot/`), with slot caps of 10/20/30/40 manual + 10/20/30/40 automatic snapshots per plan — [Source](https://docs.hetzner.com/storage/storage-box/snapshots/); plan overview confirms "unlimited traffic" but fixed storage — [Source](https://docs.hetzner.com/storage/storage-box/general)
|
||||
- Hetzner offers no per-subaccount quota ("there is currently no way to set quotas for each subaccount") and no auto-thinning; mitigation is manual read-only flag on sub-account directories — [Source](https://gist.github.com/jan-di/f6e403bfc6457daae3981e307bdf9a84); official docs: "all sub accounts use the storage space of your Storage Box. To control storage usage, you can manually set a sub-account's directory to read-only" — [Source](https://docs.hetzner.com/storage/storage-box/general)
|
||||
- ZFS tooling (zfs-auto-snapshot, sanoid) is count-based (`-k/--keep NUM Keep NUM recent snapshots`), not usage-based; no `--keep-below-X%` option exists — [Source](https://manpages.debian.org/bookworm/zfs-auto-snapshot/zfs-auto-snapshot.8.en.html); sanoid splits `--take-snapshots` / `--prune-snapshots` / `--cron` with Nagios-style `--monitor-capacity` reporting only — [Source](https://github.com/jimsalterjrs/sanoid)
|
||||
- ZFS percentage numbers that do exist are health/performance floors, not retention triggers: Ubuntu warns "Minimum free space to take a snapshot and preserve ZFS performance is 20%. Free space on pool rpool is 10%" — [Source](https://superuser.com/questions/1736700/how-do-i-remove-old-zfs-snapshots); OpenZFS tuning advises "Keep pool free space above 10% to avoid many metaslabs from reaching the 4% free space threshold" where allocator flips from first-fit to best-fit and IOPS collapses — [Source](https://openzfs.github.io/openzfs-docs/Performance%20and%20Tuning/Workload%20Tuning.html)
|
||||
- Borg has no quota-aware prune: `borg prune`/`borg delete` + `borg compact` are explicit/manual or script-scheduled; "repository disk space is not freed until you run borg compact" — [Source](https://manpages.ubuntu.com/manpages/jammy/man1/borg-delete.1.html); quickstart warns to "ensure that there is *always* plenty of free space" and to "use `prune` and `compact` regularly" — [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)
|
||||
- Time Machine "as your backup disk fills up, Time Machine deletes older backups to make room for new ones" with no published percentage - deletion is on-demand pre-backup, not a watermark - [Source](https://support.apple.com/guide/mac-help/if-the-time-machine-backup-disk-is-full-mh15137/mac)
|
||||
- Time Machine backupd log shows two-phase scheme: "Starting pre-backup thinning: 53.57 GB requested (including padding)" then "No expired backups exist - deleting oldest backups to make room"; post-backup thinning separately expires hourly/daily backups - [Source](https://serverfault.com/posts/39310/revisions)
|
||||
- Time Machine standard thinning ladder is hourly backups >24h old (keep first-of-day), daily backups >30 days old (keep first-of-week), then oldest-weekly deletion only when space is needed - [Source](https://discussions.apple.com/thread/251224074)
|
||||
- Time Machine per-destination quota via `tmutil setquota DESTINATION_ID QUOTA_IN_GB` caps a destination (e.g. 500 GB); quota "takes effect on the next backup, at which point older snapshots get thinned to fit" - [Source](https://superuser.com/questions/445579/how-do-i-trim-time-machine-backup-history)
|
||||
- `tmutil thinlocalsnapshots mount_point [purge_amount] [urgency]`: "tmutil will attempt (with urgency level 1-4) to reclaim purge_amount in bytes by thinning snapshots"; urgency 4 is most aggressive, e.g. `thinlocalsnapshots / 10000000000 4` to reclaim ~10 GB - [Source](https://ss64.com/mac/tmutil.html); same signature confirmed in Apple developer thread - [Source](https://developer.apple.com/forums/thread/81171)
|
||||
- Local APFS snapshots are purgeable space "automatically reclaimed as needed" but with no user-visible percentage; when reclamation lags users must thin manually (e.g. `thinlocalsnapshots / 100000000000 4` ≈ 100 GB at urgency 4) - [Source](https://discussions.apple.com/thread/252655312)
|
||||
- Veeam retention is count-based restore points / GFS, not capacity-based: "Retention policy defines the number of restore points to keep on your performance extents and capacity extents"; earliest restore point removed from chain, blocks purged from capacity tier on next offload/copy session - [Source](https://helpcenter.veeam.com/docs/vbr/userguide/capacity_tier_retention.html)
|
||||
- Veeam Scale-Out Backup Repository (SOBR) has no auto-delete-on-full: "if the extents of your scale-out backup repository run out of space, you can add a new extent"; free space on the new extent is added to SOBR capacity - [Source](https://helpcenter.veeam.com/docs/vbr/userguide/backup_repository_sobr.html)
|
||||
- Veeam SOBR placement prefers the extent with fewest chains, breaking ties by most free space; "priority is always to complete a backup" even if that violates the Data-Locality placement policy by spilling an incremental to another extent - [Source](https://veeam-best-practices-guide-v9.readthedocs.io/resource_planning/repository_sobr.html)
|
||||
- Veeam ONE / MP capacity reports use a configurable "Repository Free Space (%)" forecast threshold (worked example 30%) to flag repositories that "will run out of space", i.e. monitoring/alerting rather than enforcement - [Source](https://helpcenter.veeam.com/docs/mp/reports/capacity_planning_for_backup_repositories.html?ver=9a)
|
||||
- S3 Lifecycle has no capacity trigger at all: rules are `Days`/`Date`/`NoncurrentVersionExpiration` per prefix/tag (e.g. transition after 365 days, expire after 3650 days); S3 "quotas" are counts (buckets, access points), not bytes - [Source](https://docs.aws.amazon.com/AmazonS3/latest/API/API_LifecycleRule.html); expiration example - [Source](https://docs.amazonaws.cn/en_us/AmazonS3/latest/userguide/lifecycle-configuration-examples.md)
|
||||
- Hetzner Storage Box quota is the fixed plan size (BX11/BX21/BX31/BX41); snapshots "consume storage space from your Storage Box's storage capacity" alongside live data (`/.zfs/snapshot/`), with slot caps of 10/20/30/40 manual + 10/20/30/40 automatic snapshots per plan - [Source](https://docs.hetzner.com/storage/storage-box/snapshots/); plan overview confirms "unlimited traffic" but fixed storage - [Source](https://docs.hetzner.com/storage/storage-box/general)
|
||||
- Hetzner offers no per-subaccount quota ("there is currently no way to set quotas for each subaccount") and no auto-thinning; mitigation is manual read-only flag on sub-account directories - [Source](https://gist.github.com/jan-di/f6e403bfc6457daae3981e307bdf9a84); official docs: "all sub accounts use the storage space of your Storage Box. To control storage usage, you can manually set a sub-account's directory to read-only" - [Source](https://docs.hetzner.com/storage/storage-box/general)
|
||||
- ZFS tooling (zfs-auto-snapshot, sanoid) is count-based (`-k/--keep NUM Keep NUM recent snapshots`), not usage-based; no `--keep-below-X%` option exists - [Source](https://manpages.debian.org/bookworm/zfs-auto-snapshot/zfs-auto-snapshot.8.en.html); sanoid splits `--take-snapshots` / `--prune-snapshots` / `--cron` with Nagios-style `--monitor-capacity` reporting only - [Source](https://github.com/jimsalterjrs/sanoid)
|
||||
- ZFS percentage numbers that do exist are health/performance floors, not retention triggers: Ubuntu warns "Minimum free space to take a snapshot and preserve ZFS performance is 20%. Free space on pool rpool is 10%" - [Source](https://superuser.com/questions/1736700/how-do-i-remove-old-zfs-snapshots); OpenZFS tuning advises "Keep pool free space above 10% to avoid many metaslabs from reaching the 4% free space threshold" where allocator flips from first-fit to best-fit and IOPS collapses - [Source](https://openzfs.github.io/openzfs-docs/Performance%20and%20Tuning/Workload%20Tuning.html)
|
||||
- Borg has no quota-aware prune: `borg prune`/`borg delete` + `borg compact` are explicit/manual or script-scheduled; "repository disk space is not freed until you run borg compact" - [Source](https://manpages.ubuntu.com/manpages/jammy/man1/borg-delete.1.html); quickstart warns to "ensure that there is *always* plenty of free space" and to "use `prune` and `compact` regularly" - [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)
|
||||
|
||||
### Inferences
|
||||
- Time Machine is the outlier reference design for quota-aware retention: time ladder first, then unconditional oldest-first deletion driven by the byte size (plus padding) of the incoming backup.
|
||||
- Every other surveyed system treats capacity as an ops/monitoring concern (add extent, raise quota, manual prune) rather than a retention input, which is why "keep usage below X%" knobs are absent outside custom wrappers.
|
||||
|
||||
### Gaps
|
||||
- No vendor-published numeric watermark (e.g. "start deleting at 90%") was found for Time Machine, Veeam SOBR, or Hetzner — deletion/placement appears driven by allocation failure or free-space comparison, not a fixed percent.
|
||||
- No vendor-published numeric watermark (e.g. "start deleting at 90%") was found for Time Machine, Veeam SOBR, or Hetzner - deletion/placement appears driven by allocation failure or free-space comparison, not a fixed percent.
|
||||
- Could not confirm any Hetzner-side automatic snapshot rotation on full; docs describe slot caps but not capacity-triggered eviction.
|
||||
|
||||
## How do they measure remote usage on dumb backends (quota APIs, PROPFIND quota properties, du-style walks) when the server exposes no quota endpoint?
|
||||
@@ -37,14 +37,14 @@ Only Apple Time Machine makes capacity the primary retention driver ("delete old
|
||||
Only WebDAV-based backends have a standard quota API (RFC 4331 PROPFIND properties + HTTP 507); S3/object storage, Borg-over-SSH, restic, and Hetzner Storage Box over SFTP/rsync/Borg have no byte-quota endpoint, so clients fall back to local `df`/repository accounting, provider console/API, or expensive tree walks.
|
||||
|
||||
### Cited Findings
|
||||
- RFC 4331 defines two live PROPFIND properties for quota: `DAV:quota-available-bytes` ("maximum amount of additional storage available to be allocated") and `DAV:quota-used-bytes` ("amount of space used ... including usage derived from sub-resources"), explicitly warning "as the DAV:quota-available-bytes on a resource approaches 0, further allocations ... may be refused" — [Source](https://datatracker.ietf.org/doc/html/rfc4331)
|
||||
- Quota exhaustion on WebDAV is signaled by HTTP 507 (Insufficient Storage), which "SHOULD be used when a client request (e.g. a PUT, PROPFIND, MKCOL, MOVE, or COPY) fails because it would exceed their quota or physical storage limits" — [Source](http://www.webdav.org/specs/rfc4331.html)
|
||||
- Nextcloud (a common self-hosted WebDAV target) implements both properties (`quota-available-bytes`, `quota-used-bytes`) retrievable via PROPFIND — [Source](https://github.com/nextcloud/documentation/blob/master/developer_manual/client_apis/WebDAV/basic.rst)
|
||||
- Hetzner Storage Box exposes usage via console/API and a "Determine available Storage Box disk space" doc path, not via a uniform in-protocol quota on all transports; supported accesses are FTP/FTPS, SFTP/SCP, SSH/rsync/BorgBackup, SMB/CIFS, WebDAV — [Source](https://docs.hetzner.com/storage/storage-box); only the WebDAV path inherits RFC 4331 properties.
|
||||
- S3 has no byte-quota endpoint: lifecycle/quota docs cover object counts and bucket limits; storage accounting is via CloudWatch/Storage Lens/billing, and lifecycle evaluation is a daily asynchronous scan ("S3 Lifecycle evaluates objects against tag-based filters daily ... queues the action for asynchronous processing") — [Source](https://docs.aws.amazon.com/AmazonS3/latest/userguide/lifecycle-expire-general-considerations.html)
|
||||
- Veeam measures SOBR extent free space by polling the extent, but "Free space data is only retrieved when no active tasks are assigned to an extent", so placement decisions can be made on stale data under continuous load — [Source](https://bp.veeam.com/vbr/3_Build_structures/B_Veeam_Components/B_backup_repositories/scaleout.html)
|
||||
- Borg/restic on "dumb" backends (SSH, SFTP, rest-server, B2) do no server-side quota query; Borg docs direct users to local filesystem monitoring ("include the free space information in your backup log files"), client-side quotas, and `borg repo-space` accounting — [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)
|
||||
- Restic cache-size issue threads show the failure of local-accounting fallback: cache at `~/Library/Caches/restic` or `/root/.cache/restic` can fill the system disk (reports of 20–40 GB, ~3–10% of repo size) with no built-in cap, and `restic cache --cleanup` only removes stale-repo caches — [Source](https://github.com/restic/restic/issues/4325)
|
||||
- RFC 4331 defines two live PROPFIND properties for quota: `DAV:quota-available-bytes` ("maximum amount of additional storage available to be allocated") and `DAV:quota-used-bytes` ("amount of space used ... including usage derived from sub-resources"), explicitly warning "as the DAV:quota-available-bytes on a resource approaches 0, further allocations ... may be refused" - [Source](https://datatracker.ietf.org/doc/html/rfc4331)
|
||||
- Quota exhaustion on WebDAV is signaled by HTTP 507 (Insufficient Storage), which "SHOULD be used when a client request (e.g. a PUT, PROPFIND, MKCOL, MOVE, or COPY) fails because it would exceed their quota or physical storage limits" - [Source](http://www.webdav.org/specs/rfc4331.html)
|
||||
- Nextcloud (a common self-hosted WebDAV target) implements both properties (`quota-available-bytes`, `quota-used-bytes`) retrievable via PROPFIND - [Source](https://github.com/nextcloud/documentation/blob/master/developer_manual/client_apis/WebDAV/basic.rst)
|
||||
- Hetzner Storage Box exposes usage via console/API and a "Determine available Storage Box disk space" doc path, not via a uniform in-protocol quota on all transports; supported accesses are FTP/FTPS, SFTP/SCP, SSH/rsync/BorgBackup, SMB/CIFS, WebDAV - [Source](https://docs.hetzner.com/storage/storage-box); only the WebDAV path inherits RFC 4331 properties.
|
||||
- S3 has no byte-quota endpoint: lifecycle/quota docs cover object counts and bucket limits; storage accounting is via CloudWatch/Storage Lens/billing, and lifecycle evaluation is a daily asynchronous scan ("S3 Lifecycle evaluates objects against tag-based filters daily ... queues the action for asynchronous processing") - [Source](https://docs.aws.amazon.com/AmazonS3/latest/userguide/lifecycle-expire-general-considerations.html)
|
||||
- Veeam measures SOBR extent free space by polling the extent, but "Free space data is only retrieved when no active tasks are assigned to an extent", so placement decisions can be made on stale data under continuous load - [Source](https://bp.veeam.com/vbr/3_Build_structures/B_Veeam_Components/B_backup_repositories/scaleout.html)
|
||||
- Borg/restic on "dumb" backends (SSH, SFTP, rest-server, B2) do no server-side quota query; Borg docs direct users to local filesystem monitoring ("include the free space information in your backup log files"), client-side quotas, and `borg repo-space` accounting - [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)
|
||||
- Restic cache-size issue threads show the failure of local-accounting fallback: cache at `~/Library/Caches/restic` or `/root/.cache/restic` can fill the system disk (reports of 20–40 GB, ~3–10% of repo size) with no built-in cap, and `restic cache --cleanup` only removes stale-repo caches - [Source](https://github.com/restic/restic/issues/4325)
|
||||
|
||||
### Inferences
|
||||
- For a WebDAV "dumb backend" (Hetzner, Nextcloud), PROPFIND `quota-used-bytes`/`quota-available-bytes` with Depth:0 is the cheapest correct pre-backup check; on SFTP/rsync/S3 transports the only portable options are provider-specific APIs or recursive size walks (`du`, `rclone size`, `restic stats`), which are O(files) and unsuitable per-backup.
|
||||
@@ -57,16 +57,16 @@ Only WebDAV-based backends have a standard quota API (RFC 4331 PROPFIND properti
|
||||
## What is the recommended layering: time-based policy first, capacity trigger as backstop, or capacity as the primary driver?
|
||||
|
||||
### Takeaway
|
||||
Universal recommended layering is time/count policy first, capacity as backstop — except Time Machine, where capacity is the ultimate driver after the time ladder is exhausted. Enterprise guidance (Veeam, ZFS/sanoid, S3) never recommends capacity as the primary retention rule because it makes recovery windows unpredictable.
|
||||
Universal recommended layering is time/count policy first, capacity as backstop - except Time Machine, where capacity is the ultimate driver after the time ladder is exhausted. Enterprise guidance (Veeam, ZFS/sanoid, S3) never recommends capacity as the primary retention rule because it makes recovery windows unpredictable.
|
||||
|
||||
### Cited Findings
|
||||
- Time Machine layering: keep "local snapshots for the past 24 hours, daily backups for the past month and weekly backups for all previous months" and "oldest backups and any local snapshots are deleted as space is needed" — time ladder first, space-need second — [Source](https://discussions.apple.com/thread/255740341)
|
||||
- Backupd implements the layering literally: pre-backup thinning first deletes expired (time-policy) backups, and only if "No expired backups exist" does it delete oldest backups to make room — [Source](https://serverfault.com/posts/39310/revisions)
|
||||
- Veeam layering: short-term + GFS (weekly/monthly/yearly) retention counts define what may be offloaded ("operational restore window ... defines which retention files can be offloaded"); capacity tier move/copy is placement, not an extra deletion rule, and "retention of the objects in the Capacity Tier is controlled by the backup or backup copy job's retention policy in restore points and not on repository level" — [Source](https://veeambp.readthedocs.io/resource_planning/repository_sobr_capacity_tier.html)
|
||||
- Veeam ONE guidance on low free space is "free up storage space on the repository or revise your backup retention policy" — i.e. human revises the time policy, system does not auto-shorten it — [Source](https://helpcenter.veeam.com/docs/one/userguide/backup_repositories_overview.html)
|
||||
- S3 layering is time-only by design: combine transition + expiration actions into a lifecycle timeline (e.g. 30d frequent → 90d infrequent → Glacier → expire); there is no capacity input to the rule engine — [Source](https://docs.aws.amazon.com/AmazonS3/latest/userguide/object-lifecycle-mgmt.md)
|
||||
- ZFS/sanoid layering is template keep-counts (hourly/daily/weekly/monthly) executed by `--cron`, with `--monitor-capacity`/`--monitor-health` feeding external alerting, not feeding back into keep-counts — [Source](https://github.com/jimsalterjrs/sanoid)
|
||||
- Borg layering per quickstart: time/count `prune` rules run on schedule plus `compact` to actually reclaim, with free-space monitoring and optional reserved-space (`borg repo-space`) as the capacity backstop — [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)
|
||||
- Time Machine layering: keep "local snapshots for the past 24 hours, daily backups for the past month and weekly backups for all previous months" and "oldest backups and any local snapshots are deleted as space is needed" - time ladder first, space-need second - [Source](https://discussions.apple.com/thread/255740341)
|
||||
- Backupd implements the layering literally: pre-backup thinning first deletes expired (time-policy) backups, and only if "No expired backups exist" does it delete oldest backups to make room - [Source](https://serverfault.com/posts/39310/revisions)
|
||||
- Veeam layering: short-term + GFS (weekly/monthly/yearly) retention counts define what may be offloaded ("operational restore window ... defines which retention files can be offloaded"); capacity tier move/copy is placement, not an extra deletion rule, and "retention of the objects in the Capacity Tier is controlled by the backup or backup copy job's retention policy in restore points and not on repository level" - [Source](https://veeambp.readthedocs.io/resource_planning/repository_sobr_capacity_tier.html)
|
||||
- Veeam ONE guidance on low free space is "free up storage space on the repository or revise your backup retention policy" - i.e. human revises the time policy, system does not auto-shorten it - [Source](https://helpcenter.veeam.com/docs/one/userguide/backup_repositories_overview.html)
|
||||
- S3 layering is time-only by design: combine transition + expiration actions into a lifecycle timeline (e.g. 30d frequent → 90d infrequent → Glacier → expire); there is no capacity input to the rule engine - [Source](https://docs.aws.amazon.com/AmazonS3/latest/userguide/object-lifecycle-mgmt.md)
|
||||
- ZFS/sanoid layering is template keep-counts (hourly/daily/weekly/monthly) executed by `--cron`, with `--monitor-capacity`/`--monitor-health` feeding external alerting, not feeding back into keep-counts - [Source](https://github.com/jimsalterjrs/sanoid)
|
||||
- Borg layering per quickstart: time/count `prune` rules run on schedule plus `compact` to actually reclaim, with free-space monitoring and optional reserved-space (`borg repo-space`) as the capacity backstop - [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)
|
||||
|
||||
### Inferences
|
||||
- The sane default for a versioning/backup scheduler is: (1) declarative time-based keep rule, (2) pre-backup capacity check that prunes oldest-expired/oldest beyond-minimum first, (3) hard failure with clear "target full" signal rather than violating a minimum-retention floor silently.
|
||||
@@ -81,15 +81,15 @@ Universal recommended layering is time/count policy first, capacity as backstop
|
||||
Mid-backup-full behavior ranges from graceful (Time Machine aborts the run, compacts, retries; Veeam spills to another extent or skips VM below a free-space floor) to catastrophic (Borg may be unable to prune/compact without free space; ZFS deletions can themselves return ENOSPC when snapshots pin blocks; restic retries blindly on 507/ENOSPC in older versions).
|
||||
|
||||
### Cited Findings
|
||||
- Time Machine on unfreeable full: cancels the run ("Stopping backup. Backup canceled. Ejected Time Machine disk image. Compacting backup disk image to recover free space"), then retries as a fresh "Starting standard backup"; user-visible error is "This backup is too large for the backup disk. The backup requires XX GB but only YY GB are available" — [Source](https://serverfault.com/posts/39310/revisions); error text — [Source](https://osxdaily.com/2015/07/27/delete-old-backups-time-machine-mac)
|
||||
- Veeam datastore guard: jobs warn "Production datastore ... is getting low on free space (X GB left), and may run out of free disk space completely due to open snapshots" and "Skip VMs when free disk is below" logic terminates processing below the floor; hard floor is 2 GB free (registry `BlockSnapshotThreshold`, DWORD GB) even if the skip option is disabled — [Source](https://www.veeam.com/kb4379?ad=in-text-link)
|
||||
- Veeam SOBR spillover: if one extent has no free space, Veeam places the next incremental on a different extent, violating Data-Locality to prioritize completing the backup — [Source](https://veeam-best-practices-guide-v9.readthedocs.io/resource_planning/repository_sobr.html)
|
||||
- Borg worst case: "If you do run out of disk space, it can be hard or impossible to free space, because Borg needs free space to operate - even to delete backup archives"; mitigations are `borg repo-space` reservation, resizable LVs with unallocated extents, quotas, regular prune+compact — [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)
|
||||
- Borg two-phase free: deleting an archive only marks for deletion; "repository disk space is not freed until you run borg compact" (which itself needs working space) — [Source](https://manpages.ubuntu.com/manpages/jammy/man1/borg-delete.1.html)
|
||||
- ZFS snapshot-pinned full: "if the file to be removed exists in a snapshot ... then no space is gained ... As a result, the file deletion can consume more disk space ... you can get an unexpected ENOSPC or EDQUOT when attempting to remove a file" — [Source](https://docs.oracle.com/cd/E19120-01/open.solaris/817-2271/gayra/index.html)
|
||||
- Restic mid-backup-full: local temp-pack path can panic with "no space left on device" (`panic: Write: write /tmp/restic-temp-pack-...: no space left on device`) — [Source](https://github.com/restic/restic/issues/611); newer fix "Stop retrying uploads when rest-server runs out of space" shows prior behavior was unbounded retry on ENOSPC — [Source](https://github.com/restic/restic/releases)
|
||||
- WebDAV full is a clean protocol error: 507 Insufficient Storage on PUT/MKCOL/MOVE/COPY — [Source](http://www.webdav.org/specs/rfc4331.html); Hetzner snapshots compound this because snapshot-pinned blocks silently consume the same plan quota — [Source](https://docs.hetzner.com/storage/storage-box/snapshots/)
|
||||
- Headroom mechanisms found: Time Machine "padding" added to requested bytes in pre-backup thinning ("53.57 GB requested (including padding)") — [Source](https://serverfault.com/posts/39310/revisions); Veeam 2 GB snapshot floor — [Source](https://www.veeam.com/kb4379?ad=in-text-link); Borg `repo-space` reservation + LVM overprovisioning — [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst); ZFS 10–20% free-space guidance — [Source](https://openzfs.github.io/openzfs-docs/Performance%20and%20Tuning/Workload%20Tuning.html)
|
||||
- Time Machine on unfreeable full: cancels the run ("Stopping backup. Backup canceled. Ejected Time Machine disk image. Compacting backup disk image to recover free space"), then retries as a fresh "Starting standard backup"; user-visible error is "This backup is too large for the backup disk. The backup requires XX GB but only YY GB are available" - [Source](https://serverfault.com/posts/39310/revisions); error text - [Source](https://osxdaily.com/2015/07/27/delete-old-backups-time-machine-mac)
|
||||
- Veeam datastore guard: jobs warn "Production datastore ... is getting low on free space (X GB left), and may run out of free disk space completely due to open snapshots" and "Skip VMs when free disk is below" logic terminates processing below the floor; hard floor is 2 GB free (registry `BlockSnapshotThreshold`, DWORD GB) even if the skip option is disabled - [Source](https://www.veeam.com/kb4379?ad=in-text-link)
|
||||
- Veeam SOBR spillover: if one extent has no free space, Veeam places the next incremental on a different extent, violating Data-Locality to prioritize completing the backup - [Source](https://veeam-best-practices-guide-v9.readthedocs.io/resource_planning/repository_sobr.html)
|
||||
- Borg worst case: "If you do run out of disk space, it can be hard or impossible to free space, because Borg needs free space to operate - even to delete backup archives"; mitigations are `borg repo-space` reservation, resizable LVs with unallocated extents, quotas, regular prune+compact - [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)
|
||||
- Borg two-phase free: deleting an archive only marks for deletion; "repository disk space is not freed until you run borg compact" (which itself needs working space) - [Source](https://manpages.ubuntu.com/manpages/jammy/man1/borg-delete.1.html)
|
||||
- ZFS snapshot-pinned full: "if the file to be removed exists in a snapshot ... then no space is gained ... As a result, the file deletion can consume more disk space ... you can get an unexpected ENOSPC or EDQUOT when attempting to remove a file" - [Source](https://docs.oracle.com/cd/E19120-01/open.solaris/817-2271/gayra/index.html)
|
||||
- Restic mid-backup-full: local temp-pack path can panic with "no space left on device" (`panic: Write: write /tmp/restic-temp-pack-...: no space left on device`) - [Source](https://github.com/restic/restic/issues/611); newer fix "Stop retrying uploads when rest-server runs out of space" shows prior behavior was unbounded retry on ENOSPC - [Source](https://github.com/restic/restic/releases)
|
||||
- WebDAV full is a clean protocol error: 507 Insufficient Storage on PUT/MKCOL/MOVE/COPY - [Source](http://www.webdav.org/specs/rfc4331.html); Hetzner snapshots compound this because snapshot-pinned blocks silently consume the same plan quota - [Source](https://docs.hetzner.com/storage/storage-box/snapshots/)
|
||||
- Headroom mechanisms found: Time Machine "padding" added to requested bytes in pre-backup thinning ("53.57 GB requested (including padding)") - [Source](https://serverfault.com/posts/39310/revisions); Veeam 2 GB snapshot floor - [Source](https://www.veeam.com/kb4379?ad=in-text-link); Borg `repo-space` reservation + LVM overprovisioning - [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst); ZFS 10–20% free-space guidance - [Source](https://openzfs.github.io/openzfs-docs/Performance%20and%20Tuning/Workload%20Tuning.html)
|
||||
|
||||
### Inferences
|
||||
- Robust design needs three headroom elements together: (a) pre-backup estimate + padding (Time Machine model), (b) a reserved-space tripwire that stops new writes before 100% (Veeam 2 GB / Borg repo-space / ZFS 10% models), (c) a recovery path that works at 100% (Time Machine compact-and-retry; Borg notably lacks one).
|
||||
@@ -97,4 +97,4 @@ Mid-backup-full behavior ranges from graceful (Time Machine aborts the run, comp
|
||||
|
||||
### Gaps
|
||||
- Exact Time Machine padding formula and Veeam SOBR stale-free-space window under load are not published; both would need empirical measurement.
|
||||
- No restic-side quota reservation feature found as of 2026 (open `--cache-size-limit` request) — [Source](https://github.com/restic/restic/issues/4325).
|
||||
- No restic-side quota reservation feature found as of 2026 (open `--cache-size-limit` request) - [Source](https://github.com/restic/restic/issues/4325).
|
||||
|
||||
@@ -6,21 +6,21 @@
|
||||
All Linux tools use static, opt-in, unlimited-by-default rate/priority knobs; Time Machine instead imposes mandatory kernel-level low-priority I/O (IOPOL_THROTTLE) with no per-backup user knob.
|
||||
|
||||
### Cited Findings
|
||||
- restic exposes global `--limit-upload rate` and `--limit-download rate` in KiB/s, default unlimited (0) — [Source](https://restic.readthedocs.io/en/stable/manual_rest.html); also documented in man pages as `--limit-upload=0` / `--limit-download=0` default unlimited — [Source](https://man.archlinux.org/man/restic-options.1.en)
|
||||
- restic documents that `--limit-upload` cannot be changed mid-run without restart; users request pv-like `-R` dynamic adjustment and resort to killing/restarting or external shapers — [Source](https://forum.restic.net/t/change-limit-upload-without-aborting-restic/5958)
|
||||
- Borg 1.x exposes `--upload-ratelimit RATE` in kiByte/s, default 0=unlimited (with `--remote-ratelimit` deprecated alias) — [Source](https://manpages.debian.org/bookworm/borgbackup/borg-common.1.en.html); also `--upload-buffer` size in MiB, default no buffer — [Source](https://man.archlinux.org/man/borg-common.1.en.txt)
|
||||
- Borg 2.x removed `--remote-ratelimit`/`--upload-ratelimit`; bandwidth is limited via `BORGSTORE_BANDWIDTH` in bits/sec, default 0=unlimited, plus `BORGSTORE_LATENCY` delay per call, or via `pv -L` ProxyCommand / rclone `--bwlimit` — [Source](https://borgbackup.readthedocs.io/en/latest/faq.html)
|
||||
- Older Borg 1.x FAQ documents upload-only `--remote-ratelimit` plus `pv`-wrapper + `BORG_RSH` for download shaping with on-the-fly `pv -R $(pidof pv) -L` changes — [Source](https://borgbackup.readthedocs.io/en/stable/faq.html)
|
||||
- Borg's limiter is a token-bucket `SleepingBandwidthLimiter` with `RATELIMIT_PERIOD = 0.1`, capping burst quota at 2x period allowance — [Source](https://github.com/borgbackup/borg/blob/da3105f1/src/borg/remote.py)
|
||||
- Kopia exposes `repository throttle set` flags `--upload-bytes-per-second`, `--download-bytes-per-second`, `--concurrent-reads/writes`, `--read/write-requests-per-second`, `--list-requests-per-second` — [Source](https://kopia.io/docs/reference/command-line/common/repository-throttle-set/); readable via `repository throttle get` — [Source](https://kopia.io/docs/reference/command-line/common/repository-throttle-get/); server-side equivalent `server throttle set` — [Source](https://kopia.io/docs/reference/command-line/common/server-throttle-set/)
|
||||
- Kopia direct-connect backends accept `--max-upload-speed`/`--max-download-speed` bytes/sec at `repository connect` time (e.g. B2/S3), persisted as `maxUploadSpeedBytesPerSecond` in repository.config — [Source](https://kopia.discourse.group/t/limit-upload-speed-as-a-policy/990)
|
||||
- Kopia `repository throttle set` is rejected on server-connected repos ("operation supported only on direct repository"), and per-policy/UI global throttle was still missing as of 2024–2026 feature requests — [Source](https://github.com/kopia/kopia/issues/3051); KopiaUI throttle exposure requested — [Source](https://github.com/kopia/kopia/issues/3586)
|
||||
- Kopia upload throttling was bursty (whole-object sleeps) until PR #2682 added per-read `DuringUpload` throttling — [Source](https://github.com/kopia/kopia/pull/2682)
|
||||
- Time Machine's `backupd` runs at `IOPOL_THROTTLE`, defined as "long-running I/O intensive background work, such as backups" that "will be throttled to prevent impact on higher policy levels" — [Source](https://eclecticlight.co/2022/01/20/why-time-machine-backups-can-be-interminably-slow/)
|
||||
- Full `IOPOL` ladder is IMPORTANT (default) / STANDARD / UTILITY / THROTTLE / PASSIVE — [Source](https://eclecticlight.co/2026/03/28/explainer-i-o-throttling/)
|
||||
- Global kill-switch `sudo sysctl debug.lowpri_throttle_enabled=0` (re-enable with `=1`, lost on reboot unless persisted via `/etc/sysctl.conf` or LaunchDaemon) removes throttle for all background I/O, not just backupd — [Source](https://osxdaily.com/2016/04/17/speed-up-time-machine-by-removing-low-process-priority-throttling/); same command/LaunchDaemon recipe — [Source](https://apple.stackexchange.com/questions/181609/time-capsule-wired-backup-transfer-slow-with-fast-bursts); throttling is I/O not CPU — [Source](https://mjtsai.com/blog/2016/03/16/massively-speed-up-time-machine-backups/)
|
||||
- Measured effect of disabling throttle: copying phase 193→332 MB/s, overall backup 160→276 MB/s (>10 GB test); pre-backup 50 MB probe writes unaffected — [Source](https://eclecticlight.co/2022/02/28/does-removing-i-o-throttling-make-backups-faster/)
|
||||
- Duplicity has no native generic bandwidth-limit option (open bug #1291633); workarounds are `trickle -s -u/-d`, WonderShaper/tc, or legacy `--scp-command="scp -l N"` (kbit/s, scp backend only, option later deprecated) — [Source](https://bugs.launchpad.net/bugs/1291633); scp `-l` throttle Q&A — [Source](https://lists.libreplanet.org/archive/html/duplicity-talk/2007-09/msg00058.html); router-QoS/DSCP attempts reported ineffective, per-machine Bandwidth Limiter used instead — [Source](https://lists.libreplanet.org/archive/html/duplicity-talk/2021-09/msg00000.html)
|
||||
- restic exposes global `--limit-upload rate` and `--limit-download rate` in KiB/s, default unlimited (0) - [Source](https://restic.readthedocs.io/en/stable/manual_rest.html); also documented in man pages as `--limit-upload=0` / `--limit-download=0` default unlimited - [Source](https://man.archlinux.org/man/restic-options.1.en)
|
||||
- restic documents that `--limit-upload` cannot be changed mid-run without restart; users request pv-like `-R` dynamic adjustment and resort to killing/restarting or external shapers - [Source](https://forum.restic.net/t/change-limit-upload-without-aborting-restic/5958)
|
||||
- Borg 1.x exposes `--upload-ratelimit RATE` in kiByte/s, default 0=unlimited (with `--remote-ratelimit` deprecated alias) - [Source](https://manpages.debian.org/bookworm/borgbackup/borg-common.1.en.html); also `--upload-buffer` size in MiB, default no buffer - [Source](https://man.archlinux.org/man/borg-common.1.en.txt)
|
||||
- Borg 2.x removed `--remote-ratelimit`/`--upload-ratelimit`; bandwidth is limited via `BORGSTORE_BANDWIDTH` in bits/sec, default 0=unlimited, plus `BORGSTORE_LATENCY` delay per call, or via `pv -L` ProxyCommand / rclone `--bwlimit` - [Source](https://borgbackup.readthedocs.io/en/latest/faq.html)
|
||||
- Older Borg 1.x FAQ documents upload-only `--remote-ratelimit` plus `pv`-wrapper + `BORG_RSH` for download shaping with on-the-fly `pv -R $(pidof pv) -L` changes - [Source](https://borgbackup.readthedocs.io/en/stable/faq.html)
|
||||
- Borg's limiter is a token-bucket `SleepingBandwidthLimiter` with `RATELIMIT_PERIOD = 0.1`, capping burst quota at 2x period allowance - [Source](https://github.com/borgbackup/borg/blob/da3105f1/src/borg/remote.py)
|
||||
- Kopia exposes `repository throttle set` flags `--upload-bytes-per-second`, `--download-bytes-per-second`, `--concurrent-reads/writes`, `--read/write-requests-per-second`, `--list-requests-per-second` - [Source](https://kopia.io/docs/reference/command-line/common/repository-throttle-set/); readable via `repository throttle get` - [Source](https://kopia.io/docs/reference/command-line/common/repository-throttle-get/); server-side equivalent `server throttle set` - [Source](https://kopia.io/docs/reference/command-line/common/server-throttle-set/)
|
||||
- Kopia direct-connect backends accept `--max-upload-speed`/`--max-download-speed` bytes/sec at `repository connect` time (e.g. B2/S3), persisted as `maxUploadSpeedBytesPerSecond` in repository.config - [Source](https://kopia.discourse.group/t/limit-upload-speed-as-a-policy/990)
|
||||
- Kopia `repository throttle set` is rejected on server-connected repos ("operation supported only on direct repository"), and per-policy/UI global throttle was still missing as of 2024–2026 feature requests - [Source](https://github.com/kopia/kopia/issues/3051); KopiaUI throttle exposure requested - [Source](https://github.com/kopia/kopia/issues/3586)
|
||||
- Kopia upload throttling was bursty (whole-object sleeps) until PR #2682 added per-read `DuringUpload` throttling - [Source](https://github.com/kopia/kopia/pull/2682)
|
||||
- Time Machine's `backupd` runs at `IOPOL_THROTTLE`, defined as "long-running I/O intensive background work, such as backups" that "will be throttled to prevent impact on higher policy levels" - [Source](https://eclecticlight.co/2022/01/20/why-time-machine-backups-can-be-interminably-slow/)
|
||||
- Full `IOPOL` ladder is IMPORTANT (default) / STANDARD / UTILITY / THROTTLE / PASSIVE - [Source](https://eclecticlight.co/2026/03/28/explainer-i-o-throttling/)
|
||||
- Global kill-switch `sudo sysctl debug.lowpri_throttle_enabled=0` (re-enable with `=1`, lost on reboot unless persisted via `/etc/sysctl.conf` or LaunchDaemon) removes throttle for all background I/O, not just backupd - [Source](https://osxdaily.com/2016/04/17/speed-up-time-machine-by-removing-low-process-priority-throttling/); same command/LaunchDaemon recipe - [Source](https://apple.stackexchange.com/questions/181609/time-capsule-wired-backup-transfer-slow-with-fast-bursts); throttling is I/O not CPU - [Source](https://mjtsai.com/blog/2016/03/16/massively-speed-up-time-machine-backups/)
|
||||
- Measured effect of disabling throttle: copying phase 193→332 MB/s, overall backup 160→276 MB/s (>10 GB test); pre-backup 50 MB probe writes unaffected - [Source](https://eclecticlight.co/2022/02/28/does-removing-i-o-throttling-make-backups-faster/)
|
||||
- Duplicity has no native generic bandwidth-limit option (open bug #1291633); workarounds are `trickle -s -u/-d`, WonderShaper/tc, or legacy `--scp-command="scp -l N"` (kbit/s, scp backend only, option later deprecated) - [Source](https://bugs.launchpad.net/bugs/1291633); scp `-l` throttle Q&A - [Source](https://lists.libreplanet.org/archive/html/duplicity-talk/2007-09/msg00058.html); router-QoS/DSCP attempts reported ineffective, per-machine Bandwidth Limiter used instead - [Source](https://lists.libreplanet.org/archive/html/duplicity-talk/2021-09/msg00000.html)
|
||||
- Neither restic, borg, kopia, nor duplicity ships battery- or metered-network-aware auto-pause; scheduling/power-awareness is delegated to systemd timers, DAS-CTS (macOS), or external shapers (see Gaps).
|
||||
|
||||
### Inferences
|
||||
@@ -36,20 +36,20 @@ All Linux tools use static, opt-in, unlimited-by-default rate/priority knobs; Ti
|
||||
No tool auto-tunes from measured disk/RAM/link speed; the only resource-derived defaults are CPU-count-derived worker counts, everything else is fixed static defaults.
|
||||
|
||||
### Cited Findings
|
||||
- restic defaults: file-read concurrency 2 ("sweet spot" from HDD experiments), blob-save concurrency = `runtime.NumCPU()`, tree-save concurrency = 20x blob concurrency — [Source](https://github.com/restic/restic/blob/de9136b29f86216bd3e41397d19b25f26b578833/internal/archiver/archiver.go)
|
||||
- restic uses all available CPUs by default; `GOMAXPROCS=1` pins to one core and slightly reduces memory — [Source](https://restic.readthedocs.io/en/latest/047_tuning_parameters.html)
|
||||
- restic backend connection limit defaults to 5 (2 for local backend), tunable via `-o rest.connections=5` / `-o local.connections=2`; too-high values degrade performance — [Source](https://restic.readthedocs.io/en/latest/047_tuning_parameters.html)
|
||||
- restic `--read-concurrency` / `RESTIC_READ_CONCURRENCY` raises parallel file reads for NVMe; `--no-scan` skips the pre-backup file-count/size scan that costs extra I/O on network/FUSE mounts — [Source](https://restic.readthedocs.io/en/latest/047_tuning_parameters.html); read-concurrency flag added in 0.15.0 for fast storage — [Source](https://restic.net/blog/2023-01-12/restic-0.15.0-released/)
|
||||
- Single-large-file chunking in restic is sequential per file (~450 MB/s) with parallel hash/compress/encrypt downstream; multi-file parallelism is what scales, which is why `cores/4`-style read-concurrency guesses only hold for SSDs — [Source](https://github.com/restic/restic/issues/4477)
|
||||
- Kopia `--max-parallel-file-reads` defaults to number of logical CPU cores; lowering it lowers CPU at cost of time — [Source](https://kopia.io/docs/faqs/)
|
||||
- Kopia `--max-parallel-snapshots` controls simultaneous snapshots (server/KopiaUI); s2 `default`/`better` compressor concurrency equals logical core count, `s2-parallel-4/8` pins it — [Source](https://kopia.io/docs/advanced/compression/)
|
||||
- Time Machine scheduling via DAS-CTS scores each due background activity every few seconds against temperature, load, and priority, dispatching `com.apple.backupd-auto` via XPC only when score exceeds threshold — [Source](https://eclecticlight.co/2023/11/28/scheduling-and-dispatch-of-backups-and-other-background-activities/)
|
||||
- On Apple Silicon, background-QoS `backupd` threads are confined to Efficiency cores at reduced frequency (~972–1332 MHz, ~90% residency on E cores), capping throughput at ~300–400 items/s regardless of queue depth — [Source](https://eclecticlight.co/2022/01/20/why-time-machine-backups-can-be-interminably-slow/)
|
||||
- restic defaults: file-read concurrency 2 ("sweet spot" from HDD experiments), blob-save concurrency = `runtime.NumCPU()`, tree-save concurrency = 20x blob concurrency - [Source](https://github.com/restic/restic/blob/de9136b29f86216bd3e41397d19b25f26b578833/internal/archiver/archiver.go)
|
||||
- restic uses all available CPUs by default; `GOMAXPROCS=1` pins to one core and slightly reduces memory - [Source](https://restic.readthedocs.io/en/latest/047_tuning_parameters.html)
|
||||
- restic backend connection limit defaults to 5 (2 for local backend), tunable via `-o rest.connections=5` / `-o local.connections=2`; too-high values degrade performance - [Source](https://restic.readthedocs.io/en/latest/047_tuning_parameters.html)
|
||||
- restic `--read-concurrency` / `RESTIC_READ_CONCURRENCY` raises parallel file reads for NVMe; `--no-scan` skips the pre-backup file-count/size scan that costs extra I/O on network/FUSE mounts - [Source](https://restic.readthedocs.io/en/latest/047_tuning_parameters.html); read-concurrency flag added in 0.15.0 for fast storage - [Source](https://restic.net/blog/2023-01-12/restic-0.15.0-released/)
|
||||
- Single-large-file chunking in restic is sequential per file (~450 MB/s) with parallel hash/compress/encrypt downstream; multi-file parallelism is what scales, which is why `cores/4`-style read-concurrency guesses only hold for SSDs - [Source](https://github.com/restic/restic/issues/4477)
|
||||
- Kopia `--max-parallel-file-reads` defaults to number of logical CPU cores; lowering it lowers CPU at cost of time - [Source](https://kopia.io/docs/faqs/)
|
||||
- Kopia `--max-parallel-snapshots` controls simultaneous snapshots (server/KopiaUI); s2 `default`/`better` compressor concurrency equals logical core count, `s2-parallel-4/8` pins it - [Source](https://kopia.io/docs/advanced/compression/)
|
||||
- Time Machine scheduling via DAS-CTS scores each due background activity every few seconds against temperature, load, and priority, dispatching `com.apple.backupd-auto` via XPC only when score exceeds threshold - [Source](https://eclecticlight.co/2023/11/28/scheduling-and-dispatch-of-backups-and-other-background-activities/)
|
||||
- On Apple Silicon, background-QoS `backupd` threads are confined to Efficiency cores at reduced frequency (~972–1332 MHz, ~90% residency on E cores), capping throughput at ~300–400 items/s regardless of queue depth - [Source](https://eclecticlight.co/2022/01/20/why-time-machine-backups-can-be-interminably-slow/)
|
||||
- No evidence any tool probes link speed, disk size, or free RAM to set pack size, connections, or limits automatically; restic 16 MiB pack size, Kopia parallelism, and Borg rate defaults are all static.
|
||||
|
||||
### Inferences
|
||||
- "Auto-tuning" in this space means CPU-count-proportional worker pools plus OS-scheduler deferral (DAS-CTS/QoS), not closed-loop adaptation to throughput or memory pressure.
|
||||
- Raising concurrency without raising memory (restic 100×1 GB test hit 300 GB RAM at connections=16/reads=16 and OOMed) shows why static defaults stay conservative — [Source](https://github.com/restic/restic/issues/4477).
|
||||
- Raising concurrency without raising memory (restic 100×1 GB test hit 300 GB RAM at connections=16/reads=16 and OOMed) shows why static defaults stay conservative - [Source](https://github.com/restic/restic/issues/4477).
|
||||
|
||||
### Gaps
|
||||
- No source found documenting link-speed probing or disk-size-derived chunk/pack sizing in any of the five tools; if it exists it is not in public docs/CLI help.
|
||||
@@ -60,15 +60,15 @@ No tool auto-tunes from measured disk/RAM/link speed; the only resource-derived
|
||||
Constrained-box guidance is manual: pick cheap compression, lower parallelism/connections, enlarge packs, and accept slower runs; low-RAM index handling remains a known failure mode, not an auto-degraded mode.
|
||||
|
||||
### Cited Findings
|
||||
- Borg compression default is lz4 (very high speed, very low compression); alternatives `zstd[,L]` (default level 3), `zlib`, `lzma`, `auto,`, `none` — [Source](https://borgbackup.readthedocs.io/en/stable/usage/help.html); usage examples recommend `zlib,6` for ratio at cost of speed — [Source](https://borgbackup.readthedocs.io/en/stable/usage/create.html)
|
||||
- Upstream Borg PR proposes moving default from `lz4` to `zstd,-4` (multithreaded, as-fast-or-faster creates, slightly better ratio; large incompressible-image corpora stay ~11% slower) with MT workers capped at 4 — [Source](https://github.com/borgbackup/borg/pull/10100)
|
||||
- Kopia compression is disabled by default and set per-policy via `kopia policy set [--global] --compression=<...>` with min/max-size gates — [Source](https://kopia.io/docs/faqs/); full option list includes `s2-default/better/parallel-4/8`, `zstd/zstd-fastest/better`, `gzip/pgzip/deflate` variants — [Source](https://kopia.io/docs/reference/command-line/common/policy-set/)
|
||||
- Kopia FAQ names compression + parallelism as the two main memory culprits; recommends disabling compression or using `s2`/`deflate`/`gzip` on small files under low memory, and lowering `--max-parallel-snapshots` / `--max-parallel-file-reads` — [Source](https://kopia.io/docs/faqs/)
|
||||
- Kopia benchmark table (466 MiB corpus): s2-default 4 GiB/s at ~375 MiB RSS vs zstd 323 MiB/s at ~238 MiB vs zstd-best 19 MiB/s; on tiny files s2 stays fastest with ~2 MiB footprint — [Source](https://kopia.io/docs/advanced/compression/)
|
||||
- Kopia zstd levels map to upstream klauspost/compress: fastest≈1, default≈3, better≈7, best≈11; higher custom levels (e.g. -22 --ultra --long) require code change — [Source](https://kopia.discourse.group/t/what-is-zstd-best-compression-and-can-i-customize/4934)
|
||||
- restic on constrained HDD/NVMe: forum-tested recipe for 2.5 TB SMR-USB run is `--read-concurrency 1 --pack-size 128 --no-cache` (pack default 16 MiB raised to reduce file count and per-pack latency), with `local.connections=1` suggested for vibration-sensitive HDDs — [Source](https://forum.restic.net/t/first-backup-2-5tb-50-hours-can-i-improve-it/8288)
|
||||
- restic large-pack tradeoff: bigger packs reduce file count and help HDD/Swift/Drive limits but need more `$TMPDIR` staging (64–384 MiB guidance) and longer single-pack uploads that wear SSDs — [Source](https://github.com/restic/restic/blob/master/doc/047_tuning_parameters.rst)
|
||||
- Borg on near-full disks: 1.5 GiB-free VM case shows Borg needs headroom for segments/cache/index; workarounds discussed are extreme compression or `--upload-ratelimit` pacing plus inotify/SIGSTOP hacks, with maintainer warning such boxes are unsuitable — [Source](https://github.com/borgbackup/borg/issues/7107)
|
||||
- Borg compression default is lz4 (very high speed, very low compression); alternatives `zstd[,L]` (default level 3), `zlib`, `lzma`, `auto,`, `none` - [Source](https://borgbackup.readthedocs.io/en/stable/usage/help.html); usage examples recommend `zlib,6` for ratio at cost of speed - [Source](https://borgbackup.readthedocs.io/en/stable/usage/create.html)
|
||||
- Upstream Borg PR proposes moving default from `lz4` to `zstd,-4` (multithreaded, as-fast-or-faster creates, slightly better ratio; large incompressible-image corpora stay ~11% slower) with MT workers capped at 4 - [Source](https://github.com/borgbackup/borg/pull/10100)
|
||||
- Kopia compression is disabled by default and set per-policy via `kopia policy set [--global] --compression=<...>` with min/max-size gates - [Source](https://kopia.io/docs/faqs/); full option list includes `s2-default/better/parallel-4/8`, `zstd/zstd-fastest/better`, `gzip/pgzip/deflate` variants - [Source](https://kopia.io/docs/reference/command-line/common/policy-set/)
|
||||
- Kopia FAQ names compression + parallelism as the two main memory culprits; recommends disabling compression or using `s2`/`deflate`/`gzip` on small files under low memory, and lowering `--max-parallel-snapshots` / `--max-parallel-file-reads` - [Source](https://kopia.io/docs/faqs/)
|
||||
- Kopia benchmark table (466 MiB corpus): s2-default 4 GiB/s at ~375 MiB RSS vs zstd 323 MiB/s at ~238 MiB vs zstd-best 19 MiB/s; on tiny files s2 stays fastest with ~2 MiB footprint - [Source](https://kopia.io/docs/advanced/compression/)
|
||||
- Kopia zstd levels map to upstream klauspost/compress: fastest≈1, default≈3, better≈7, best≈11; higher custom levels (e.g. -22 --ultra --long) require code change - [Source](https://kopia.discourse.group/t/what-is-zstd-best-compression-and-can-i-customize/4934)
|
||||
- restic on constrained HDD/NVMe: forum-tested recipe for 2.5 TB SMR-USB run is `--read-concurrency 1 --pack-size 128 --no-cache` (pack default 16 MiB raised to reduce file count and per-pack latency), with `local.connections=1` suggested for vibration-sensitive HDDs - [Source](https://forum.restic.net/t/first-backup-2-5tb-50-hours-can-i-improve-it/8288)
|
||||
- restic large-pack tradeoff: bigger packs reduce file count and help HDD/Swift/Drive limits but need more `$TMPDIR` staging (64–384 MiB guidance) and longer single-pack uploads that wear SSDs - [Source](https://github.com/restic/restic/blob/master/doc/047_tuning_parameters.rst)
|
||||
- Borg on near-full disks: 1.5 GiB-free VM case shows Borg needs headroom for segments/cache/index; workarounds discussed are extreme compression or `--upload-ratelimit` pacing plus inotify/SIGSTOP hacks, with maintainer warning such boxes are unsuitable - [Source](https://github.com/borgbackup/borg/issues/7107)
|
||||
- Single-core guidance converges: `GOMAXPROCS=1` (restic), `-C none|laz4` (borg), `--compression=s2-default|none` + `--max-parallel-file-reads=1` (kopia) minimize CPU/RAM at cost of ratio/throughput.
|
||||
|
||||
### Inferences
|
||||
@@ -84,13 +84,13 @@ Constrained-box guidance is manual: pick cheap compression, lower parallelism/co
|
||||
Upstream backup tools ship no restrictive resource-control units; all concrete CPU/memory/I/O caps come from downstream/community units and generic systemd resource-control docs, with `Nice=` + `CPUQuota`/`MemoryMax`/`IOWeight` as the recommended trio.
|
||||
|
||||
### Cited Findings
|
||||
- systemd `CPUQuota=` sets a hard ceiling as % of one CPU (100%=1 core, 200%=2 cores) via `cpu.max`/`cpu.cfs_quota_us`; `CPUWeight=` (1–10000, default 100) is only relative under contention — [Source](https://manpages.debian.org/bullseye/systemd/systemd.resource-control.5.en.html); same semantics in Arch man — [Source](https://man.archlinux.org/man/systemd.resource-control.5)
|
||||
- systemd memory knobs: `MemoryHigh=` soft throttle/reclaim, `MemoryMax=` hard OOM-kill limit (K/M/G/T or % of RAM, `infinity` to disable), `MemorySwapMax=` swap cap; Red Hat recommends `MemoryHigh` as main control, `MemoryMax` as last defense — [Source](https://docs.redhat.com/en/documentation/red_hat_enterprise_linux/8/html/managing_monitoring_and_updating_the_kernel/assembly_configuring-resource-management-using-systemd_managing-monitoring-and-updating-the-kernel)
|
||||
- Community restic timer units use only `Nice=17` (low CPU priority) plus sandboxing (`ProtectSystem=full`, `PrivateTmp=true`, `NoNewPrivileges=yes`, `RestrictAddressFamilies=`, `SystemCallFilter=`, `AmbientCapabilities=CAP_DAC_READ_SEARCH`), with `RandomizedDelaySec=300` + `Persistent=yes` on the timer — [Source](https://www.wildtechgarden.ca/onepagers/real-life-systemd-timers/)
|
||||
- Community segmented-borg systemd design sets `CPUQuota=80%` and `MemoryMax=2G` on the backup service template — [Source](https://github.com/JoZapf/segmented-borg-backup-system/blob/refs/heads/main/docs/SYSTEMD.md)
|
||||
- Modern Debian guidance: background CPU → `nice -n 19`; background disk → `ionice -c 3` (BFQ only); hard ceilings/group fairness → unit with `CPUQuota=`/`MemoryMax=`/`IOWeight=`/`IOReadBandwidthMax=`/`IOWriteBandwidthMax=` or a shared `backup.slice`; one-shots via `systemd-run --scope -p CPUQuota=50% -p MemoryMax=1G -p IOWeight=10` — [Source](https://www.bigiron.cc/guides/process-priorities-nice-ionice-cgroup-quotas)
|
||||
- Same guide warns `ionice` is silently ignored on default `mq-deadline` SSD schedulers (only BFQ honors classes); `MemoryMax` kills rather than slows (use `MemoryHigh` for pushback); `CPUQuota=100%` means one core — [Source](https://www.bigiron.cc/guides/process-priorities-nice-ionice-cgroup-quotas)
|
||||
- Example backup slice caps restic at 80 MB/s read / 30 MB/s write via `IOReadBandwidthMax=/dev/sda 80M` + `IOWriteBandwidthMax=/dev/sda 30M` with `IOWeight=10` so foreground pools keep headroom — [Source](https://www.bigiron.cc/guides/process-priorities-nice-ionice-cgroup-quotas)
|
||||
- systemd `CPUQuota=` sets a hard ceiling as % of one CPU (100%=1 core, 200%=2 cores) via `cpu.max`/`cpu.cfs_quota_us`; `CPUWeight=` (1–10000, default 100) is only relative under contention - [Source](https://manpages.debian.org/bullseye/systemd/systemd.resource-control.5.en.html); same semantics in Arch man - [Source](https://man.archlinux.org/man/systemd.resource-control.5)
|
||||
- systemd memory knobs: `MemoryHigh=` soft throttle/reclaim, `MemoryMax=` hard OOM-kill limit (K/M/G/T or % of RAM, `infinity` to disable), `MemorySwapMax=` swap cap; Red Hat recommends `MemoryHigh` as main control, `MemoryMax` as last defense - [Source](https://docs.redhat.com/en/documentation/red_hat_enterprise_linux/8/html/managing_monitoring_and_updating_the_kernel/assembly_configuring-resource-management-using-systemd_managing-monitoring-and-updating-the-kernel)
|
||||
- Community restic timer units use only `Nice=17` (low CPU priority) plus sandboxing (`ProtectSystem=full`, `PrivateTmp=true`, `NoNewPrivileges=yes`, `RestrictAddressFamilies=`, `SystemCallFilter=`, `AmbientCapabilities=CAP_DAC_READ_SEARCH`), with `RandomizedDelaySec=300` + `Persistent=yes` on the timer - [Source](https://www.wildtechgarden.ca/onepagers/real-life-systemd-timers/)
|
||||
- Community segmented-borg systemd design sets `CPUQuota=80%` and `MemoryMax=2G` on the backup service template - [Source](https://github.com/JoZapf/segmented-borg-backup-system/blob/refs/heads/main/docs/SYSTEMD.md)
|
||||
- Modern Debian guidance: background CPU → `nice -n 19`; background disk → `ionice -c 3` (BFQ only); hard ceilings/group fairness → unit with `CPUQuota=`/`MemoryMax=`/`IOWeight=`/`IOReadBandwidthMax=`/`IOWriteBandwidthMax=` or a shared `backup.slice`; one-shots via `systemd-run --scope -p CPUQuota=50% -p MemoryMax=1G -p IOWeight=10` - [Source](https://www.bigiron.cc/guides/process-priorities-nice-ionice-cgroup-quotas)
|
||||
- Same guide warns `ionice` is silently ignored on default `mq-deadline` SSD schedulers (only BFQ honors classes); `MemoryMax` kills rather than slows (use `MemoryHigh` for pushback); `CPUQuota=100%` means one core - [Source](https://www.bigiron.cc/guides/process-priorities-nice-ionice-cgroup-quotas)
|
||||
- Example backup slice caps restic at 80 MB/s read / 30 MB/s write via `IOReadBandwidthMax=/dev/sda 80M` + `IOWriteBandwidthMax=/dev/sda 30M` with `IOWeight=10` so foreground pools keep headroom - [Source](https://www.bigiron.cc/guides/process-priorities-nice-ionice-cgroup-quotas)
|
||||
- No shipped upstream restic/borg/kopia unit found with `MemoryMax`/`CPUQuota`/`IOWeight` preset; ArchWiki/packaging ships timers/services without resource caps (report-writer: treat absence as finding, not oversight).
|
||||
|
||||
### Inferences
|
||||
|
||||
@@ -6,19 +6,19 @@
|
||||
restic and BorgBackup are fully manual/external-scheduler tools (no built-in scheduler; prune-after-backup is a script/wrapper convention with lock-contention and cost tradeoffs), Kopia is automatic/built-in (maintenance fires opportunistically on client use with a single elected owner), Time Machine is fully automatic and continuous (thinning driven by schedule + disk pressure), and Veeam is built-in-scheduler enterprise (retention applied inline after job sessions plus a nightly background process, with health-check/compact on separate schedules).
|
||||
|
||||
### Cited Findings
|
||||
- restic has no built-in scheduler: `forget` only deletes snapshot objects and a separate `prune` must remove unreferenced data; `--prune` on `forget` automates the two-step sequence only when snapshots were actually removed — [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
- restic docs warn pruning is time-consuming, takes an exclusive repository lock so "backups cannot be completed" during prune, and advise planning prune windows plus running `restic check` afterwards — [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
- Community restic practice is external cron (e.g. `0 0 * * *` daily backup script running backup → check → `forget --keep-daily N --prune`) or systemd timers; third-party wrappers (restic-scheduler, autorestic/resticprofile) add per-task systemd services where backup runs often (e.g. daily) and retention+check runs monthly — [restic cron example](https://gist.github.com/perfecto25/f528f8d14e1c4b6e2a912513539a5af7); [restic-scheduler](https://github.com/AenonDynamics/restic-scheduler)
|
||||
- BorgBackup `prune` is documented as "normally used by automated backup scripts"; the official quickstart pattern is one script doing `create` → `prune` → `compact` in sequence, and `borgmatic`'s default actions are create+prune+compact+check (i.e. after-each-backup by convention, not by daemon) — [borg prune docs](https://borgbackup.readthedocs.io/en/stable/usage/prune.html); [borg quickstart](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst); [borgmatic manpage](https://manpages.ubuntu.com/manpages/focal/man1/borgmatic.1.html)
|
||||
- borgmatic docs warn the default every-run prune/compact/check is fine for small repos but too slow for very large ones, so large repos should decouple them (skip actions / separate schedules) — [borgmatic large backups guide](https://github.com/borgmatic-collective/borgmatic/blob/main/docs/how-to/deal-with-very-large-backups.md)
|
||||
- Vorta (Borg desktop GUI) exposes retention as a "Prune after each backup" checkbox, i.e. after-each-backup as an opt-in — [Vorta prune docs](https://vorta.borgbase.com/usage/prune)
|
||||
- Kopia maintenance is automatic since v0.6.0: it "will happen occasionally when the `kopia` command-line client is used", with quick tasks ~hourly and full tasks every 24h by default — [Kopia maintenance docs](https://kopia.io/docs/advanced/maintenance/)
|
||||
- Kopia defaults are quick interval 1h and full interval 24h, both enabled (from source defaults `QuickCycle 1h`, `FullCycle 24h`) — [maintenance_params.go v0.17.0](https://raw.githubusercontent.com/kopia/kopia/v0.17.0/repo/maintenance/maintenance_params.go); corroborated by user-observed "every hour for quick and every day for full" — [kopia issue #1439](https://github.com/kopia/kopia/issues/1439)
|
||||
- Kopia retention policy (which snapshots to keep: keep-latest/hourly/daily/weekly/monthly/annual) is separate from maintenance (GC of unreferenced blobs); retention is applied by deleting expired snapshots, full-maintenance Snapshot-GC then reclaims the data — [kopia policy set reference](https://kopia.io/docs/reference/command-line/common/policy-set/); [Kopia maintenance docs](https://kopia.io/docs/advanced/maintenance/)
|
||||
- Time Machine "automatically makes hourly backups for the past 24 hours, daily backups for the past month, and weekly backups for all previous months", deleting oldest backups when the disk is full — [Apple Support](https://support.apple.com/en-la/104984)
|
||||
- Time Machine local snapshots are created hourly, stored on the source disk, kept up to 24h or until space is needed, and removed automatically under pressure — [Apple local snapshots guide](https://support.apple.com/en-euro/guide/mac-help/mh35933/mac)
|
||||
- Veeam short-term retention is applied inline at the end of each job session (chain transform/merge), while GFS retention is enforced by a background process / nightly retention job (v11+ Cloud Connect docs describe a nightly "retention job"; run `History > System`, filter "retention") — [VCSP GFS retention docs](https://veeamvcsp.github.io/docs/vcc/gfs)
|
||||
- Veeam health check and defrag/compact are opt-in scheduled operations attached to backup jobs, not run after every backup: health check off by default schedule wording, default monthly (last Sat/Sun 05:00 depending on product/generation), compact disabled by default — [Veeam maintenance settings](https://helpcenter.veeam.com/docs/vbr/userguide/backup_copy_settings_backup.html); [Veeam health check](https://helpcenter.veeam.com/docs/vbr/userguide/backup_health_check.html)
|
||||
- restic has no built-in scheduler: `forget` only deletes snapshot objects and a separate `prune` must remove unreferenced data; `--prune` on `forget` automates the two-step sequence only when snapshots were actually removed - [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
- restic docs warn pruning is time-consuming, takes an exclusive repository lock so "backups cannot be completed" during prune, and advise planning prune windows plus running `restic check` afterwards - [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
- Community restic practice is external cron (e.g. `0 0 * * *` daily backup script running backup → check → `forget --keep-daily N --prune`) or systemd timers; third-party wrappers (restic-scheduler, autorestic/resticprofile) add per-task systemd services where backup runs often (e.g. daily) and retention+check runs monthly - [restic cron example](https://gist.github.com/perfecto25/f528f8d14e1c4b6e2a912513539a5af7); [restic-scheduler](https://github.com/AenonDynamics/restic-scheduler)
|
||||
- BorgBackup `prune` is documented as "normally used by automated backup scripts"; the official quickstart pattern is one script doing `create` → `prune` → `compact` in sequence, and `borgmatic`'s default actions are create+prune+compact+check (i.e. after-each-backup by convention, not by daemon) - [borg prune docs](https://borgbackup.readthedocs.io/en/stable/usage/prune.html); [borg quickstart](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst); [borgmatic manpage](https://manpages.ubuntu.com/manpages/focal/man1/borgmatic.1.html)
|
||||
- borgmatic docs warn the default every-run prune/compact/check is fine for small repos but too slow for very large ones, so large repos should decouple them (skip actions / separate schedules) - [borgmatic large backups guide](https://github.com/borgmatic-collective/borgmatic/blob/main/docs/how-to/deal-with-very-large-backups.md)
|
||||
- Vorta (Borg desktop GUI) exposes retention as a "Prune after each backup" checkbox, i.e. after-each-backup as an opt-in - [Vorta prune docs](https://vorta.borgbase.com/usage/prune)
|
||||
- Kopia maintenance is automatic since v0.6.0: it "will happen occasionally when the `kopia` command-line client is used", with quick tasks ~hourly and full tasks every 24h by default - [Kopia maintenance docs](https://kopia.io/docs/advanced/maintenance/)
|
||||
- Kopia defaults are quick interval 1h and full interval 24h, both enabled (from source defaults `QuickCycle 1h`, `FullCycle 24h`) - [maintenance_params.go v0.17.0](https://raw.githubusercontent.com/kopia/kopia/v0.17.0/repo/maintenance/maintenance_params.go); corroborated by user-observed "every hour for quick and every day for full" - [kopia issue #1439](https://github.com/kopia/kopia/issues/1439)
|
||||
- Kopia retention policy (which snapshots to keep: keep-latest/hourly/daily/weekly/monthly/annual) is separate from maintenance (GC of unreferenced blobs); retention is applied by deleting expired snapshots, full-maintenance Snapshot-GC then reclaims the data - [kopia policy set reference](https://kopia.io/docs/reference/command-line/common/policy-set/); [Kopia maintenance docs](https://kopia.io/docs/advanced/maintenance/)
|
||||
- Time Machine "automatically makes hourly backups for the past 24 hours, daily backups for the past month, and weekly backups for all previous months", deleting oldest backups when the disk is full - [Apple Support](https://support.apple.com/en-la/104984)
|
||||
- Time Machine local snapshots are created hourly, stored on the source disk, kept up to 24h or until space is needed, and removed automatically under pressure - [Apple local snapshots guide](https://support.apple.com/en-euro/guide/mac-help/mh35933/mac)
|
||||
- Veeam short-term retention is applied inline at the end of each job session (chain transform/merge), while GFS retention is enforced by a background process / nightly retention job (v11+ Cloud Connect docs describe a nightly "retention job"; run `History > System`, filter "retention") - [VCSP GFS retention docs](https://veeamvcsp.github.io/docs/vcc/gfs)
|
||||
- Veeam health check and defrag/compact are opt-in scheduled operations attached to backup jobs, not run after every backup: health check off by default schedule wording, default monthly (last Sat/Sun 05:00 depending on product/generation), compact disabled by default - [Veeam maintenance settings](https://helpcenter.veeam.com/docs/vbr/userguide/backup_copy_settings_backup.html); [Veeam health check](https://helpcenter.veeam.com/docs/vbr/userguide/backup_health_check.html)
|
||||
|
||||
### Inferences
|
||||
- The spectrum is: manual+external (restic, borg) → automatic-opportunistic (kopia) → fully automatic OS-driven (Time Machine) → policy-engine with built-in scheduler (Veeam).
|
||||
@@ -34,19 +34,19 @@ restic and BorgBackup are fully manual/external-scheduler tools (no built-in sch
|
||||
restic/Borg rely on cron or systemd timers you write (daily backup typical; prune often piggybacked, check/compact weekly–monthly); Kopia uses built-in intervals (1h quick / 24h full, tunable, plus snapshot scheduling via interval/time-of-day/cron policy); Time Machine uses undocumented-internal launchd scheduling (~hourly, customizable via tools); Veeam uses a built-in per-job scheduler plus nightly background retention and monthly health-check defaults.
|
||||
|
||||
### Cited Findings
|
||||
- restic: no internal scheduler; typical community cron is daily `0 0 * * *` running a backup script; systemd-wrapper example ships `restic-scheduler@.timer` (backup, commonly daily) plus `restic-retention@.timer` (forget+prune and check, commonly monthly) with e.g. `RETENTION_POLICY_DAYS=14 WEEKS=12 MONTHS=18 YEARS=2` — [restic cron example](https://gist.github.com/perfecto25/f528f8d14e1c4b6e2a912513539a5af7); [restic-scheduler](https://github.com/AenonDynamics/restic-scheduler)
|
||||
- restic `check --read-data-subset=n/t` (or `x%`, or size like `50M`) exists precisely to spread full-data verification across scheduled runs (e.g. 1/7..7/7 across a week, or weekly `5%`); community practice converges on daily small-subset or weekly rotating-part checks rather than full `--read-data` each run — [restic 0.13 working-with-repos](https://restic.readthedocs.io/en/v0.13.0/045_working_with_repos.html); [restic forum practice](https://forum.restic.net/t/do-you-use-check-read-data/8930)
|
||||
- Borg: no daemon; scheduling via cron/systemd calling a script or `borgmatic` (sample `borgmatic.timer` ships in repo); official quickstart script runs backup→prune→compact every invocation with example policy `--keep-daily 7 --keep-weekly 4 --keep-monthly 6` — [borg quickstart](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst); [borgmatic systemd sample](https://github.com/witten/borgmatic/blob/master/sample/systemd/borgmatic.timer)
|
||||
- Borg `check --max-duration SECONDS` supports splitting a long repo check into partial checks; documented example: full check would take 7h, daily `--max-duration=3600` yields one full check per week — [borg-check manpage](https://manpages.debian.org/testing/borgbackup/borg-check.1.en.html)
|
||||
- Borg 2.x adds `--max-age` so repeated `--max-duration`-bounded runs re-check each pack at most once per age window (example `--max-duration=3600 --max-age=1w` daily ≈ full verification weekly); partial checks require `--repository-only` — [borg 2 check docs](https://borgbackup.readthedocs.io/en/latest/usage/check.html)
|
||||
- Kopia snapshot scheduling is policy-driven: `--snapshot-interval`, `--snapshot-time HH:mm,...`, `--snapshot-time-crontab`, `--run-missed`, `--manual` — [kopia policy set reference](https://kopia.io/docs/reference/command-line/common/policy-set/)
|
||||
- Kopia maintenance intervals are built-in and tunable: `maintenance set --quick-interval=2h --full-interval=8h`, enable/disable flags, and `--pause-quick/--pause-full=DURATION` to suspend — [Kopia maintenance docs](https://kopia.io/docs/advanced/maintenance/)
|
||||
- KopiaUI/server runs a "Periodic maintenance" check roughly every 10 minutes and executes quick/full work only when due (observed behavior with default 1h/24h) — [kopia issue #1439](https://github.com/kopia/kopia/issues/1439)
|
||||
- Time Machine: hourly automatic backups driven by `com.apple.backupd-auto` LaunchDaemon (`StartInterval 3600` default, adjustable via `sudo defaults write ... StartInterval -int 7200` or TimeMachineEditor interval/calendar modes) — [Apple Gazette schedule customization](https://www.applegazette.com/applegazette-mac/customize-time-machine-backups-schedule)
|
||||
- Time Machine thinning triggers: hourly→24h, daily→~30d, weekly→until-full on the backup volume; local APFS snapshots thin on age (>24h) or space pressure, manually via `tmutil thinlocalsnapshots <mount> [bytes] [urgency 1-4]` and verifiable via `tmutil verifychecksums` / Option-click "Verify Backups" (network targets) — [Apple Support](https://support.apple.com/en-la/104984); [Apple local snapshots](https://support.apple.com/en-euro/guide/mac-help/mh35933/mac); [tmutil reference](https://ss64.com/mac/tmutil.html); [Apple verify backups](https://support.apple.com/en-mn/guide/mac-help/mh26840/mac)
|
||||
- Veeam: per-job backup schedule + GFS calendar (weekly day-of-week, monthly first/second/third/fourth/last week, yearly month) with GFS fulls created on scheduled days (synthetic); since v11 GFS creation happens right on scheduled days and a nightly background retention job enforces GFS deletions — [Veeam GFS cycles](https://helpcenter.veeam.com/docs/vbr/userguide/backup_copy_gfs_periods.html); [VCSP GFS retention docs](https://veeamvcsp.github.io/docs/vcc/gfs)
|
||||
- Veeam health check default: monthly, 05:00 last Saturday (VBR backup jobs) / last Sunday (backup-copy jobs) / last Friday (Windows agent); runs piggybacked on the first incremental session of the scheduled day, or the next session if the job didn't run that day — [Veeam health check](https://helpcenter.veeam.com/docs/vbr/userguide/backup_health_check.html); [Veeam copy health check](https://helpcenter.veeam.com/docs/vbr/userguide/backup_copy_health_check.html); [Veeam agent health check](https://helpcenter.veeam.com/docs/agentforwindows/userguide/backup_health_check.html)
|
||||
- Veeam defrag/compact-full is disabled by default, scheduled via job Maintenance settings when enabled; requires free space for an auxiliary VBK and is incompatible with GFS retention enabled — [Veeam maintenance settings](https://helpcenter.veeam.com/docs/vbr/userguide/backup_copy_settings_backup.html)
|
||||
- restic: no internal scheduler; typical community cron is daily `0 0 * * *` running a backup script; systemd-wrapper example ships `restic-scheduler@.timer` (backup, commonly daily) plus `restic-retention@.timer` (forget+prune and check, commonly monthly) with e.g. `RETENTION_POLICY_DAYS=14 WEEKS=12 MONTHS=18 YEARS=2` - [restic cron example](https://gist.github.com/perfecto25/f528f8d14e1c4b6e2a912513539a5af7); [restic-scheduler](https://github.com/AenonDynamics/restic-scheduler)
|
||||
- restic `check --read-data-subset=n/t` (or `x%`, or size like `50M`) exists precisely to spread full-data verification across scheduled runs (e.g. 1/7..7/7 across a week, or weekly `5%`); community practice converges on daily small-subset or weekly rotating-part checks rather than full `--read-data` each run - [restic 0.13 working-with-repos](https://restic.readthedocs.io/en/v0.13.0/045_working_with_repos.html); [restic forum practice](https://forum.restic.net/t/do-you-use-check-read-data/8930)
|
||||
- Borg: no daemon; scheduling via cron/systemd calling a script or `borgmatic` (sample `borgmatic.timer` ships in repo); official quickstart script runs backup→prune→compact every invocation with example policy `--keep-daily 7 --keep-weekly 4 --keep-monthly 6` - [borg quickstart](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst); [borgmatic systemd sample](https://github.com/witten/borgmatic/blob/master/sample/systemd/borgmatic.timer)
|
||||
- Borg `check --max-duration SECONDS` supports splitting a long repo check into partial checks; documented example: full check would take 7h, daily `--max-duration=3600` yields one full check per week - [borg-check manpage](https://manpages.debian.org/testing/borgbackup/borg-check.1.en.html)
|
||||
- Borg 2.x adds `--max-age` so repeated `--max-duration`-bounded runs re-check each pack at most once per age window (example `--max-duration=3600 --max-age=1w` daily ≈ full verification weekly); partial checks require `--repository-only` - [borg 2 check docs](https://borgbackup.readthedocs.io/en/latest/usage/check.html)
|
||||
- Kopia snapshot scheduling is policy-driven: `--snapshot-interval`, `--snapshot-time HH:mm,...`, `--snapshot-time-crontab`, `--run-missed`, `--manual` - [kopia policy set reference](https://kopia.io/docs/reference/command-line/common/policy-set/)
|
||||
- Kopia maintenance intervals are built-in and tunable: `maintenance set --quick-interval=2h --full-interval=8h`, enable/disable flags, and `--pause-quick/--pause-full=DURATION` to suspend - [Kopia maintenance docs](https://kopia.io/docs/advanced/maintenance/)
|
||||
- KopiaUI/server runs a "Periodic maintenance" check roughly every 10 minutes and executes quick/full work only when due (observed behavior with default 1h/24h) - [kopia issue #1439](https://github.com/kopia/kopia/issues/1439)
|
||||
- Time Machine: hourly automatic backups driven by `com.apple.backupd-auto` LaunchDaemon (`StartInterval 3600` default, adjustable via `sudo defaults write ... StartInterval -int 7200` or TimeMachineEditor interval/calendar modes) - [Apple Gazette schedule customization](https://www.applegazette.com/applegazette-mac/customize-time-machine-backups-schedule)
|
||||
- Time Machine thinning triggers: hourly→24h, daily→~30d, weekly→until-full on the backup volume; local APFS snapshots thin on age (>24h) or space pressure, manually via `tmutil thinlocalsnapshots <mount> [bytes] [urgency 1-4]` and verifiable via `tmutil verifychecksums` / Option-click "Verify Backups" (network targets) - [Apple Support](https://support.apple.com/en-la/104984); [Apple local snapshots](https://support.apple.com/en-euro/guide/mac-help/mh35933/mac); [tmutil reference](https://ss64.com/mac/tmutil.html); [Apple verify backups](https://support.apple.com/en-mn/guide/mac-help/mh26840/mac)
|
||||
- Veeam: per-job backup schedule + GFS calendar (weekly day-of-week, monthly first/second/third/fourth/last week, yearly month) with GFS fulls created on scheduled days (synthetic); since v11 GFS creation happens right on scheduled days and a nightly background retention job enforces GFS deletions - [Veeam GFS cycles](https://helpcenter.veeam.com/docs/vbr/userguide/backup_copy_gfs_periods.html); [VCSP GFS retention docs](https://veeamvcsp.github.io/docs/vcc/gfs)
|
||||
- Veeam health check default: monthly, 05:00 last Saturday (VBR backup jobs) / last Sunday (backup-copy jobs) / last Friday (Windows agent); runs piggybacked on the first incremental session of the scheduled day, or the next session if the job didn't run that day - [Veeam health check](https://helpcenter.veeam.com/docs/vbr/userguide/backup_health_check.html); [Veeam copy health check](https://helpcenter.veeam.com/docs/vbr/userguide/backup_copy_health_check.html); [Veeam agent health check](https://helpcenter.veeam.com/docs/agentforwindows/userguide/backup_health_check.html)
|
||||
- Veeam defrag/compact-full is disabled by default, scheduled via job Maintenance settings when enabled; requires free space for an auxiliary VBK and is incompatible with GFS retention enabled - [Veeam maintenance settings](https://helpcenter.veeam.com/docs/vbr/userguide/backup_copy_settings_backup.html)
|
||||
|
||||
### Inferences
|
||||
- Example "sane defaults" synthesis for adaptation: backup cadence hourly/daily (cheap op); retention-mark pass daily or after each backup; repack/GC weekly–monthly; integrity verification monthly or continuous-subset.
|
||||
@@ -62,16 +62,16 @@ restic/Borg rely on cron or systemd timers you write (daily backup typical; prun
|
||||
All five systems split cheap metadata marking from expensive space reclamation; cheap passes run often (per-backup or daily), expensive passes rarely (weekly to monthly) with explicit thresholds/tuning.
|
||||
|
||||
### Cited Findings
|
||||
- restic: `forget` (cheap: deletes snapshot metadata objects only) vs `prune` (expensive: scans all snapshots, classifies packs used/partly/unused, downloads+re-uploads repacked data — "very time-consuming for remote repositories"); `--max-unused` (default `5%`) bounds repacking, `--max-repack-size` caps work per run, `--repack-cacheable-only` restricts to metadata for a fast pass — [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
- restic: `forget --prune` couples them (prune runs only if snapshots were removed); `--max-unused unlimited` minimizes time/bandwidth (keeps partly-used packs), `0` minimizes space; `--max-repack-size 0` is the documented low-scratch-space recovery mode — [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
- Borg ≤1.x: `prune` deletes archives and `compact` (segments) was fused into prune; since 1.2 they are split: "Repository disk space is not freed until you run `borg compact`", docs recommend running compact regularly but "not after each borg command", e.g. once a month possibly with check, or when space is needed — [borg prune docs](https://borgbackup.readthedocs.io/en/stable/usage/prune.html); [borg compact docs](https://borgbackup.readthedocs.io/en/stable/usage/compact.html)
|
||||
- Borg 1.4 `compact --threshold 10%` (default) compacts a segment only above that saving; `--threshold 0` forces maximum compaction, much slower — [borg compact docs](https://borgbackup.readthedocs.io/en/stable/usage/compact.html)
|
||||
- Borg 2.x `compact` is further gated (acts only when reclaimable space ≥ threshold/5, i.e. 2% at default) plus tiny-pack merging only when small packs combine to a full-size pack; `undelete` possible after prune/delete until compact runs — [borg 2 compact docs](https://borgbackup.readthedocs.io/en/latest/usage/compact.html)
|
||||
- Kopia: quick maintenance (keeps frequently-accessed `q`/`n` blobs low; never deletes metadata without another copy existing; ~hourly) vs full maintenance (Snapshot GC marking + `p`-pack compaction + dropping deleted contents; every 24h); docs warn full-maintenance effects take several hours/cycles to materialize — [Kopia maintenance docs](https://kopia.io/docs/advanced/maintenance/)
|
||||
- Kopia task list makes the split explicit: `snapshot-gc`, `quick-delete-blobs` vs `full-delete-blobs`, `quick-rewrite-contents` vs `full-rewrite-contents`, `full-drop-deleted-content`, `index-compaction`, epoch tasks — [maintenance package reference](https://pkg.go.dev/github.com/kopia/kopia/repo/maintenance)
|
||||
- Time Machine: thinning is metadata-cheap (hardlink/APFS-clone based; dropping a snapshot only frees blocks unique to it); local-snapshot thinning is pressure-driven, backup-volume thinning continuous; heavyweight equivalent (`hdiutil compact` of network sparsebundle, `tmutil verifychecksums`) is manual/occasional — [Apple Support](https://support.apple.com/en-la/104984); [Time Machine cheatsheet](https://gist.github.com/chrisbranson/4bf59fc6f3b1be6dd5058600959564b8)
|
||||
- Veeam: cheap = per-session retention application + transform/merge of chains; expensive = active/synthetic fulls (scheduled weekly/monthly), defrag+compact full (disabled by default, scheduled separately), health check with CRC+hash of latest restore point only (monthly default, not whole chain) — [Veeam maintenance settings](https://helpcenter.veeam.com/docs/vbr/userguide/backup_copy_settings_backup.html); [Veeam health check](https://helpcenter.veeam.com/docs/vbr/userguide/backup_health_check.html)
|
||||
- Veeam health check verifies only the latest restore point per chain (not history), bounding cost; if it doesn't finish before the next scheduled run the old session stops and a new one starts — [Veeam health check](https://helpcenter.veeam.com/docs/vbr/userguide/backup_health_check.html)
|
||||
- restic: `forget` (cheap: deletes snapshot metadata objects only) vs `prune` (expensive: scans all snapshots, classifies packs used/partly/unused, downloads+re-uploads repacked data - "very time-consuming for remote repositories"); `--max-unused` (default `5%`) bounds repacking, `--max-repack-size` caps work per run, `--repack-cacheable-only` restricts to metadata for a fast pass - [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
- restic: `forget --prune` couples them (prune runs only if snapshots were removed); `--max-unused unlimited` minimizes time/bandwidth (keeps partly-used packs), `0` minimizes space; `--max-repack-size 0` is the documented low-scratch-space recovery mode - [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
- Borg ≤1.x: `prune` deletes archives and `compact` (segments) was fused into prune; since 1.2 they are split: "Repository disk space is not freed until you run `borg compact`", docs recommend running compact regularly but "not after each borg command", e.g. once a month possibly with check, or when space is needed - [borg prune docs](https://borgbackup.readthedocs.io/en/stable/usage/prune.html); [borg compact docs](https://borgbackup.readthedocs.io/en/stable/usage/compact.html)
|
||||
- Borg 1.4 `compact --threshold 10%` (default) compacts a segment only above that saving; `--threshold 0` forces maximum compaction, much slower - [borg compact docs](https://borgbackup.readthedocs.io/en/stable/usage/compact.html)
|
||||
- Borg 2.x `compact` is further gated (acts only when reclaimable space ≥ threshold/5, i.e. 2% at default) plus tiny-pack merging only when small packs combine to a full-size pack; `undelete` possible after prune/delete until compact runs - [borg 2 compact docs](https://borgbackup.readthedocs.io/en/latest/usage/compact.html)
|
||||
- Kopia: quick maintenance (keeps frequently-accessed `q`/`n` blobs low; never deletes metadata without another copy existing; ~hourly) vs full maintenance (Snapshot GC marking + `p`-pack compaction + dropping deleted contents; every 24h); docs warn full-maintenance effects take several hours/cycles to materialize - [Kopia maintenance docs](https://kopia.io/docs/advanced/maintenance/)
|
||||
- Kopia task list makes the split explicit: `snapshot-gc`, `quick-delete-blobs` vs `full-delete-blobs`, `quick-rewrite-contents` vs `full-rewrite-contents`, `full-drop-deleted-content`, `index-compaction`, epoch tasks - [maintenance package reference](https://pkg.go.dev/github.com/kopia/kopia/repo/maintenance)
|
||||
- Time Machine: thinning is metadata-cheap (hardlink/APFS-clone based; dropping a snapshot only frees blocks unique to it); local-snapshot thinning is pressure-driven, backup-volume thinning continuous; heavyweight equivalent (`hdiutil compact` of network sparsebundle, `tmutil verifychecksums`) is manual/occasional - [Apple Support](https://support.apple.com/en-la/104984); [Time Machine cheatsheet](https://gist.github.com/chrisbranson/4bf59fc6f3b1be6dd5058600959564b8)
|
||||
- Veeam: cheap = per-session retention application + transform/merge of chains; expensive = active/synthetic fulls (scheduled weekly/monthly), defrag+compact full (disabled by default, scheduled separately), health check with CRC+hash of latest restore point only (monthly default, not whole chain) - [Veeam maintenance settings](https://helpcenter.veeam.com/docs/vbr/userguide/backup_copy_settings_backup.html); [Veeam health check](https://helpcenter.veeam.com/docs/vbr/userguide/backup_health_check.html)
|
||||
- Veeam health check verifies only the latest restore point per chain (not history), bounding cost; if it doesn't finish before the next scheduled run the old session stops and a new one starts - [Veeam health check](https://helpcenter.veeam.com/docs/vbr/userguide/backup_health_check.html)
|
||||
|
||||
### Inferences
|
||||
- Recommended-frequency pattern: mark/expire cheaply and often (per-backup/daily); reclaim rarely (weekly/monthly/on-pressure); verify on a third, slowest cadence (monthly/rolling-subset).
|
||||
@@ -87,27 +87,27 @@ All five systems split cheap metadata marking from expensive space reclamation;
|
||||
Every tool offers dry-run previews (none dry-run by default); all serialize maintenance with locks/ownership (restic exclusive locks, Borg repo+cache locks, Kopia single owner + exclusive lock, Veeam job-serialization with health-check yielding); interrupted cheap passes are safe to rerun, interrupted expensive passes resume or need explicit recovery steps.
|
||||
|
||||
### Cited Findings
|
||||
- restic `forget --dry-run` prints what would be removed without removing; docs present it as the always-available preview ("You can always use `--dry-run`") — [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
- restic `prune --dry-run` likewise only shows what would be done — [restic-prune manpage](https://manpages.ubuntu.com/manpages/questing/man1/restic-prune.1.html)
|
||||
- restic refuses "empty" policies (e.g. `--keep-last 0` removes nothing) and requires `--unsafe-allow-remove-all` plus a host/tag/path filter to delete all of a group (since 0.17.0); `--group-by host,paths` default scopes policy per backup set as a safety feature — [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
- restic locks: two types (exclusive vs shared); `prune` (and `check`) require exclusive locks, backups take shared locks and can run concurrently; locks are files under `locks/` with 30-minute staleness timeout plus same-host liveness check; `--retry-lock DURATION` waits instead of failing; exit code 11 = already locked; `unlock` removes stale locks — [restic design locks](https://github.com/restic/restic/blob/master/doc/design.rst); [restic forum locking](https://forum.restic.net/t/potential-issues-with-concurrent-execution-of-restic-commands/8099); [restic-prune manpage](https://manpages.ubuntu.com/manpages/questing/man1/restic-prune.1.html)
|
||||
- restic prune is documented as interrupt-safe ("repository remains usable no matter at which point the command is interrupted") but needs scratch space; last-resort `--unsafe-recover-no-free-space` can leave repo temporarily unusable if it fails (then remove `index/` + `repair index`) — [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
- Borg docs "strongly recommend" always running `prune -v --list --dry-run` first; `--stats`/`--quick-stats` and `--dry-run` are mutually exclusive; since 1.2.0 Borg retains the oldest archive if no rule would otherwise keep anything — [borg prune docs](https://borgbackup.readthedocs.io/en/stable/usage/prune.html)
|
||||
- Borg 2.x splits deletion into soft-delete (`prune`/`delete` mark) + `compact` (irreversible); `undelete` recovers until compact runs — [borg 2 prune docs](https://borgbackup.readthedocs.io/en/master/usage/prune.html); [borg 2 compact docs](https://borgbackup.readthedocs.io/en/latest/usage/compact.html)
|
||||
- Borg `prune` auto-removes stale checkpoint archives from interrupted backups (except the latest checkpoint, still needed) — [borg prune docs](https://borgbackup.readthedocs.io/en/stable/usage/prune.html)
|
||||
- Borg locking: repository and cache locks serialize access; `borg break-lock` exists for locks left by dead processes but must only be used when no borg process is accessing the repo/cache — [borg break-lock manpage](https://manpages.debian.org/testing/borgbackup/borg-break-lock.1.en.html)
|
||||
- Borg `check` receiving SIGINT stops at the next safe boundary leaving repo and chunk index consistent; recorded partial results are kept so later checks resume where they stopped; `--repair` archive runs stop between whole archives — [borg 2 check docs](https://borgbackup.readthedocs.io/en/latest/usage/check.html)
|
||||
- Borg `check` is read-only by default; `--repair` is flagged "POTENTIALLY DANGEROUS ... might lead to data loss"; `--find-lost-archives` can restore lost archives only before `compact` removes their data — [borg check manpage](https://manpages.ubuntu.com/manpages/questing/man1/borg2-check.1.html)
|
||||
- Kopia: single maintenance owner (`user@host`, view/change via `maintenance info`/`set --owner=me`), others never auto-run maintenance; maintenance runs under an exclusive lock (`RunExclusive`); default safety windows (`PackDeleteMinAge 24h`, snapshot-GC margin for in-flight snapshots and eventual-consistency delay) mean GC takes multiple cycles; `--safety=none` disables all of it with an explicit corruption warning — [Kopia maintenance docs](https://kopia.io/docs/advanced/maintenance/); [maintenance package reference](https://pkg.go.dev/github.com/kopia/kopia/repo/maintenance)
|
||||
- Kopia `maintenance run` (quick) / `run --full` require being the owner; `--force` overrides ownership ("unsafe") — [kopia maintenance run reference](https://kopia.io/docs/reference/command-line/advanced/maintenance-run)
|
||||
- Time Machine: no dry-run for thinning; safety comes from the fixed policy + `tmutil delete` operating one snapshot at a time and verification tooling (`verifychecksums`, Verify Backups); local snapshots are APFS copy-on-write so thinning never endangers live data — [tmutil reference](https://ss64.com/mac/tmutil.html); [Apple verify backups](https://support.apple.com/en-mn/guide/mac-help/mh26840/mac)
|
||||
- Veeam: health check yields to any job/operation touching the backup (stops if one starts); unfinished health-check sessions are superseded by the next scheduled run; corrupted blocks trigger Error status + automatic retry session re-transferring bad blocks from source; GFS-flagged restore points cannot be deleted/modified during their retention window; immutable (hardened Linux/object-lock) repos block repair/deletion paths — [Veeam health check](https://helpcenter.veeam.com/docs/vbr/userguide/backup_health_check.html); [Veeam GFS cycles](https://helpcenter.veeam.com/docs/vbr/userguide/backup_copy_gfs_periods.html)
|
||||
- Veeam append-only/immutable analog to restic: with immutability, retention cannot delete until the lock expires; Linux immutable repos "do not support repair" per health-check limitations — [Veeam health check](https://helpcenter.veeam.com/docs/vbr/userguide/backup_health_check.html)
|
||||
- restic `forget --dry-run` prints what would be removed without removing; docs present it as the always-available preview ("You can always use `--dry-run`") - [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
- restic `prune --dry-run` likewise only shows what would be done - [restic-prune manpage](https://manpages.ubuntu.com/manpages/questing/man1/restic-prune.1.html)
|
||||
- restic refuses "empty" policies (e.g. `--keep-last 0` removes nothing) and requires `--unsafe-allow-remove-all` plus a host/tag/path filter to delete all of a group (since 0.17.0); `--group-by host,paths` default scopes policy per backup set as a safety feature - [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
- restic locks: two types (exclusive vs shared); `prune` (and `check`) require exclusive locks, backups take shared locks and can run concurrently; locks are files under `locks/` with 30-minute staleness timeout plus same-host liveness check; `--retry-lock DURATION` waits instead of failing; exit code 11 = already locked; `unlock` removes stale locks - [restic design locks](https://github.com/restic/restic/blob/master/doc/design.rst); [restic forum locking](https://forum.restic.net/t/potential-issues-with-concurrent-execution-of-restic-commands/8099); [restic-prune manpage](https://manpages.ubuntu.com/manpages/questing/man1/restic-prune.1.html)
|
||||
- restic prune is documented as interrupt-safe ("repository remains usable no matter at which point the command is interrupted") but needs scratch space; last-resort `--unsafe-recover-no-free-space` can leave repo temporarily unusable if it fails (then remove `index/` + `repair index`) - [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
- Borg docs "strongly recommend" always running `prune -v --list --dry-run` first; `--stats`/`--quick-stats` and `--dry-run` are mutually exclusive; since 1.2.0 Borg retains the oldest archive if no rule would otherwise keep anything - [borg prune docs](https://borgbackup.readthedocs.io/en/stable/usage/prune.html)
|
||||
- Borg 2.x splits deletion into soft-delete (`prune`/`delete` mark) + `compact` (irreversible); `undelete` recovers until compact runs - [borg 2 prune docs](https://borgbackup.readthedocs.io/en/master/usage/prune.html); [borg 2 compact docs](https://borgbackup.readthedocs.io/en/latest/usage/compact.html)
|
||||
- Borg `prune` auto-removes stale checkpoint archives from interrupted backups (except the latest checkpoint, still needed) - [borg prune docs](https://borgbackup.readthedocs.io/en/stable/usage/prune.html)
|
||||
- Borg locking: repository and cache locks serialize access; `borg break-lock` exists for locks left by dead processes but must only be used when no borg process is accessing the repo/cache - [borg break-lock manpage](https://manpages.debian.org/testing/borgbackup/borg-break-lock.1.en.html)
|
||||
- Borg `check` receiving SIGINT stops at the next safe boundary leaving repo and chunk index consistent; recorded partial results are kept so later checks resume where they stopped; `--repair` archive runs stop between whole archives - [borg 2 check docs](https://borgbackup.readthedocs.io/en/latest/usage/check.html)
|
||||
- Borg `check` is read-only by default; `--repair` is flagged "POTENTIALLY DANGEROUS ... might lead to data loss"; `--find-lost-archives` can restore lost archives only before `compact` removes their data - [borg check manpage](https://manpages.ubuntu.com/manpages/questing/man1/borg2-check.1.html)
|
||||
- Kopia: single maintenance owner (`user@host`, view/change via `maintenance info`/`set --owner=me`), others never auto-run maintenance; maintenance runs under an exclusive lock (`RunExclusive`); default safety windows (`PackDeleteMinAge 24h`, snapshot-GC margin for in-flight snapshots and eventual-consistency delay) mean GC takes multiple cycles; `--safety=none` disables all of it with an explicit corruption warning - [Kopia maintenance docs](https://kopia.io/docs/advanced/maintenance/); [maintenance package reference](https://pkg.go.dev/github.com/kopia/kopia/repo/maintenance)
|
||||
- Kopia `maintenance run` (quick) / `run --full` require being the owner; `--force` overrides ownership ("unsafe") - [kopia maintenance run reference](https://kopia.io/docs/reference/command-line/advanced/maintenance-run)
|
||||
- Time Machine: no dry-run for thinning; safety comes from the fixed policy + `tmutil delete` operating one snapshot at a time and verification tooling (`verifychecksums`, Verify Backups); local snapshots are APFS copy-on-write so thinning never endangers live data - [tmutil reference](https://ss64.com/mac/tmutil.html); [Apple verify backups](https://support.apple.com/en-mn/guide/mac-help/mh26840/mac)
|
||||
- Veeam: health check yields to any job/operation touching the backup (stops if one starts); unfinished health-check sessions are superseded by the next scheduled run; corrupted blocks trigger Error status + automatic retry session re-transferring bad blocks from source; GFS-flagged restore points cannot be deleted/modified during their retention window; immutable (hardened Linux/object-lock) repos block repair/deletion paths - [Veeam health check](https://helpcenter.veeam.com/docs/vbr/userguide/backup_health_check.html); [Veeam GFS cycles](https://helpcenter.veeam.com/docs/vbr/userguide/backup_copy_gfs_periods.html)
|
||||
- Veeam append-only/immutable analog to restic: with immutability, retention cannot delete until the lock expires; Linux immutable repos "do not support repair" per health-check limitations - [Veeam health check](https://helpcenter.veeam.com/docs/vbr/userguide/backup_health_check.html)
|
||||
|
||||
### Inferences
|
||||
- Common rail inventory for adaptation: (1) preview/dry-run, (2) empty-policy refusal / oldest-retention floor, (3) scope limiters (group-by, prefix/glob, per-machine chains), (4) exclusive locks + stale-lock recovery, (5) interrupt-safe/resumable expensive passes, (6) undo window (Borg undelete-before-compact; Kopia multi-cycle GC delay; Veeam GFS immutability windows).
|
||||
- restic's append-only + `--keep-within` guidance is the sharpest documented footgun: count-based `--keep-*` policies let injected attacker snapshots displace legitimate ones at `forget` time; time-window policies bound the damage — [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
- restic's append-only + `--keep-within` guidance is the sharpest documented footgun: count-based `--keep-*` policies let injected attacker snapshots displace legitimate ones at `forget` time; time-window policies bound the damage - [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
|
||||
### Gaps
|
||||
- Borg 1.x concurrent-access failure mode specifics (which commands take exclusive vs shared locks) not re-verified from primary docs in this pass; `break-lock` semantics fetched, lock-type matrix not.
|
||||
- Time Machine behavior when a scheduled backup/thinning is interrupted (power loss mid-thin) not found in Apple docs; generally treated as crash-safe via APFS but unsourced — left as gap rather than claimed.
|
||||
- Time Machine behavior when a scheduled backup/thinning is interrupted (power loss mid-thin) not found in Apple docs; generally treated as crash-safe via APFS but unsourced - left as gap rather than claimed.
|
||||
|
||||
@@ -34,7 +34,7 @@ server is the trust model); there are no keys to manage.
|
||||
`#*#`, `*~`, `4913`), SQLite journals as standalone files.
|
||||
- SQLite rule: backup-API snapshot + `integrity_check`; locked/corrupt →
|
||||
loud skip (`database-locked`, `database-corrupt`), never a stored version.
|
||||
- Databases >~100MB or hot 7GB-class files: do NOT raise caps blindly —
|
||||
- Databases >~100MB or hot 7GB-class files: do NOT raise caps blindly -
|
||||
each version is a full copy until chunked blobs exist. Prefer dumps.
|
||||
|
||||
## Retention, GC, scheduler (MUST understand before running)
|
||||
@@ -56,18 +56,18 @@ server is the trust model); there are no keys to manage.
|
||||
## Disaster recovery (total local loss), in order
|
||||
|
||||
1. Reinstall, `POST /config/remote/adopt {url, username, password,
|
||||
directory}` — takes over the old remote directory.
|
||||
2. `POST /admin/reindex` (202 + poll `GET /admin/reindex/{id}`) — rebuilds
|
||||
directory}` - takes over the old remote directory.
|
||||
2. `POST /admin/reindex` (202 + poll `GET /admin/reindex/{id}`) - rebuilds
|
||||
the index from manifests; everything returns `durable`.
|
||||
3. Restore (single/bulk/point-in-time). Blobs stream from WebDAV on demand.
|
||||
SQLite restores are integrity-checked; failures refuse with `unrestorable`.
|
||||
|
||||
## Destructive actions (double-check, dry-run first)
|
||||
|
||||
- `POST /admin/purge-remote {"today": "dd-mm-yyyy"}` — wipes this
|
||||
- `POST /admin/purge-remote {"today": "dd-mm-yyyy"}` - wipes this
|
||||
installation's remote `blobs/`+`manifests/` async (202 + counts, poll
|
||||
status). Only the claimed directory is ever touched. No local files harmed.
|
||||
- `POST /forget {"path", "dry_run"}` — deletes local history under a path.
|
||||
- `POST /forget {"path", "dry_run"}` - deletes local history under a path.
|
||||
- Restores: always dry-run (`POST /restores` → plan), then
|
||||
`POST /restores/{id}/execute`. Pre-restore snapshots make every restore
|
||||
undoable. Targets must resolve inside `$HOME`/roots; symlinks refused.
|
||||
|
||||
@@ -7,7 +7,7 @@ Cadence (all configurable under ``[scheduler]``):
|
||||
|
||||
Capacity is a backstop, never the driver: when remote usage reaches
|
||||
``remote_max_used_percent`` (default 70) an out-of-schedule pass runs and the
|
||||
service reports pressure (see /health) so a human analyzes — the 5-day
|
||||
service reports pressure (see /health) so a human analyzes - the 5-day
|
||||
retention floor is never violated automatically. This mirrors the surveyed
|
||||
consensus: time policy first, pressure relief second, static ceilings rather
|
||||
than autotuning.
|
||||
|
||||
Reference in New Issue
Block a user