Remove encryption; user-data filters; SQLite safety; recovery, retention, scheduler
- No client-side encryption: plaintext content-addressed remote (SHA-256 names), no backup key, no cryptography dependency (server disk encryption is the trust model). - Filters back user data: images, archives, databases, PDFs accepted; 10 MiB cap; binary sniffing removed; temp names hardened (~$, #..#, .temp). - SQLite zero-error policy: backup-API snapshots + integrity_check, journal folding, locked/corrupt loud skips, verified restores. - Recovery: reindex from manifests, remote adopt, on-demand blob fetch, 5-day retention + thinning, GC, date-guarded remote purge, metrics. - Scheduler with WebDAV quota signal and 70% pressure backstop (floor kept). - 37 tests incl. live-monitor capture safety and DB safety.
This commit is contained in:
@@ -12,7 +12,6 @@ dependencies = [
|
||||
"fastapi>=0.110",
|
||||
"uvicorn>=0.29",
|
||||
"httpx>=0.27",
|
||||
"cryptography>=42",
|
||||
]
|
||||
|
||||
[project.optional-dependencies]
|
||||
|
||||
@@ -11,10 +11,10 @@
|
||||
| Milestone | State |
|
||||
|---|---|
|
||||
| M1 – Core (monitor, filters, index, spool, history, diff, restore, systemd install) | **Implemented** |
|
||||
| M2 – WebDAV remote, unique remote directory, encryption, manifests | **Implemented** |
|
||||
| M2 – WebDAV remote, unique remote directory, manifests (unencrypted) | **Implemented** |
|
||||
| Statistics, progress (incl. live stream) and web dashboard | **Implemented** |
|
||||
| M3 – Retention, thinning, remote GC | Not started (pinning and `forget` exist) |
|
||||
| M4 – Agent skill, Prometheus metrics, reindex from remote | Not started |
|
||||
| M3 – Retention (5-day + thinning), remote GC, remote purge, scheduler + capacity backstop | **Implemented** (built-in loop; no external cron needed) |
|
||||
| M4 – Agent skill, reindex from remote, adopt, metrics | **Implemented** except agent skill |
|
||||
|
||||
```bash
|
||||
pipx install -e . # installs the `versiond` command
|
||||
@@ -26,7 +26,6 @@ versiond diff ~/projects/app/main.py # last change
|
||||
versiond restore '~/projects/app/src/*' --as-of 2026-10-08T14:00 # dry run
|
||||
versiond restore '~/projects/app/src/*' --as-of 2026-10-08T14:00 --execute
|
||||
versiond remote set --url https://u123456.your-storagebox.de --user u123456 # prompts for the password
|
||||
versiond key export # store the backup key somewhere safe, off this machine
|
||||
versiond progress --follow # scan + upload progress, rate and ETA
|
||||
versiond stats # files, versions, storage, activity
|
||||
versiond dashboard # opens the web dashboard, already signed in
|
||||
@@ -56,7 +55,7 @@ The primary use case is a safety net for workflows where files are rewritten fre
|
||||
|
||||
### 1.2 Non-goals
|
||||
|
||||
- Backing up large or binary assets (hard limit: 200 KB per file).
|
||||
- Backing up VM/container disk images of running machines (use dumps/snapshots, not live file copies).
|
||||
- Replacing Git. `versiond` versions working-tree states, not commits, branches or merges.
|
||||
- Multi-user or network-exposed operation. The service binds to loopback only.
|
||||
- Full-disk or system backup.
|
||||
@@ -188,12 +187,15 @@ Monitored roots do not need their own remote directories: blobs are shared (cont
|
||||
|
||||
```
|
||||
<base_path>/<hostname-slug>-<id8>/
|
||||
blobs/<name[:2]>/<name> # name = HMAC-SHA256(key, sha256); zstd, then AES-256-GCM
|
||||
manifests/<yyyy>/<mm>/<dd>/<batch-id>.jsonl.zst.enc # append-only version + rename records, encrypted
|
||||
meta/owner.json # installation id, hostname, key id (plaintext)
|
||||
blobs/<sha256[:2]>/<sha256> # name = SHA-256 of content; compressed, NOT encrypted
|
||||
manifests/<yyyy>/<mm>/<dd>/<batch-id>.jsonl.zst # append-only version + rename records, compressed, NOT encrypted
|
||||
meta/owner.json # installation id, hostname (plaintext)
|
||||
meta/format.json # storage format version (plaintext)
|
||||
```
|
||||
|
||||
Remote storage is deliberately unencrypted: it relies on the storage machine's
|
||||
disk encryption. There is no backup key and nothing to lose.
|
||||
|
||||
One directory level under `blobs/` (256 prefixes) keeps the number of `MKCOL` requests small.
|
||||
|
||||
Blobs are immutable and deduplicated by SHA-256 of the plaintext content. Manifests are append-only journals. Together they make it possible to fully rebuild the index (`POST /admin/reindex`) after local data loss.
|
||||
@@ -263,12 +265,12 @@ Each file is linked to a *project root*, found by walking up (within its monitor
|
||||
|
||||
A file is rejected (HTTP `422` with a reason code for push requests; silently skipped by the monitor) if any of these apply. The same rules decide which **directories** get an inotify watch at all:
|
||||
|
||||
- **Size** > 200 KB (204 800 bytes) → `413 Payload Too Large`.
|
||||
- **Size** > 10 MiB (10 485 760 bytes, configurable via `limits.max_file_bytes`) → `413 Payload Too Large`.
|
||||
- **Any path component starts with `.`** (e.g. `.git/`, `.venv/`, `.idea/`, `.cache/`), **except** dotfiles that are themselves project files at the leaf (e.g. `.env`, `.gitignore`, `.editorconfig`, `.dockerignore`). Hidden *directories* are always ignored; hidden *files* follow an allowlist.
|
||||
- **Dependency, build and cache directories**, default list:
|
||||
`node_modules`, `bower_components`, `jspm_packages`, `vendor`, `__pycache__`, `venv`, `env`, `site-packages`, `.tox`, `build`, `dist`, `target`, `out`, `bin`, `obj`, `.gradle`, `Pods`, `Carthage`, `DerivedData`, `.next`, `.nuxt`, `.svelte-kit`, `coverage`, `.terraform`, `_build`, `deps`, `elm-stuff`, `zig-cache`, `zig-out`.
|
||||
- **Compiled/binary artifacts:** `*.pyc`, `*.pyo`, `*.class`, `*.o`, `*.obj`, `*.so`, `*.dylib`, `*.dll`, `*.exe`, `*.a`, `*.lib`, `*.wasm`, `*.jar`, `*.war`, `*.whl`, `*.egg`, archives, images, media, `*.lock` files above the size limit, `*.min.js`, `*.map`.
|
||||
- **Binary content:** a file with a NUL byte in its first 8 KB is rejected.
|
||||
- **Compiled artifacts and generated bundles:** `*.pyc`, `*.pyo`, `*.class`, `*.o`, `*.obj`, `*.so`, `*.dylib`, `*.dll`, `*.exe`, `*.a`, `*.lib`, `*.wasm`, `*.jar`, `*.war`, `*.whl`, `*.egg`, `*.min.js`, `*.map`. Everything else is user data and is kept: images, audio/video, archives (`.zip`, `.tar.*`, ...), databases including SQLite journals (`-wal`/`-shm`), documents (`.pdf`), fonts. There is no binary sniffing; any bytes up to the size cap are versioned, served back as `application/octet-stream` when they are not UTF-8.
|
||||
- **Live SQLite databases (zero-error policy):** any file with the SQLite magic is never raw-copied. It is snapshotted read-only through the SQLite Online Backup API (no write lock, no WAL recovery, live writers unaffected), then `PRAGMA integrity_check` must return `ok`, otherwise nothing is stored (`database-corrupt`, `database-locked` past a 5 s deadline, `database-unreadable` surface as skips, never as versions). Journal files (`*-wal`, `*-shm`, `*-journal`) are not versioned standalone; they are folded into the main file's verified snapshot. For non-SQLite engines only a crash-consistent raw copy is possible, so dump-then-backup remains required: back up the `.dump`/`VACUUM INTO`/`pg_dump` output alongside the live file.
|
||||
- **User rules:** `.gitignore`-style patterns in `config.toml` (`[ignore] patterns = [...]`), plus optional respect of the project's own `.gitignore` (`respect_gitignore = true`, but `.env` is still captured unless explicitly excluded).
|
||||
|
||||
All filter decisions can be checked with `POST /filters/test` without storing anything.
|
||||
@@ -321,17 +323,41 @@ When the remote is unreachable, the service keeps working from the spool and cat
|
||||
- **Path safety:** targets are resolved and must lie inside an allowed root; symlink escapes and `..` traversal are refused.
|
||||
- Restoring a file that no longer exists recreates it, including parent directories.
|
||||
|
||||
### 4.8 Retention and purge
|
||||
### 4.8 Retention and purge (implemented)
|
||||
|
||||
Retention policies run on a schedule (default: daily) and on demand. Every purge supports `dry_run=true` and returns what would be removed and the bytes reclaimed.
|
||||
- **Retention** (`POST /api/v1/admin/retention/run`, dry run by default): keeps every
|
||||
version younger than `retention.keep_days` (default **5**), thins older ones to
|
||||
one per file per UTC day, and always keeps each file's latest version and all
|
||||
pins. Follow with GC to reclaim the freed blobs.
|
||||
- **Garbage collection** (`POST /api/v1/admin/gc`, dry run by default): deletes blobs
|
||||
no version references anymore, from the local spool and (unless `remote: false`)
|
||||
from WebDAV. Only unreferenced data is ever touched, so a crash can leak garbage
|
||||
but never lose a referenced blob.
|
||||
- **Built-in scheduler** (no cron needed): every `scheduler.retention_interval_seconds`
|
||||
(default daily) a retention pass runs; GC is folded in every
|
||||
`scheduler.gc_interval_seconds` (default weekly). Runs are serialized and audited.
|
||||
- **Capacity backstop** (`scheduler.remote_max_used_percent`, default **70**): remote
|
||||
usage is read via WebDAV quota properties (RFC 4331, falling back to an
|
||||
index-based estimate) and refreshed every `scheduler.usage_check_seconds`. At or
|
||||
above the limit an out-of-schedule pass runs — still never below the keep-days
|
||||
floor — and `/health` flips to `degraded` with `remote-over-capacity-limit` plus
|
||||
`remote_pressure: true` until a human analyzes (lower `keep_days`, bigger box).
|
||||
- **Remote purge** (`POST /api/v1/admin/purge-remote {"today": "dd-mm-yyyy"}`): deletes
|
||||
everything this installation uploaded (`blobs/` + `manifests/` under the claimed
|
||||
directory only). Replies 202 immediately with file/byte counts; deletes async.
|
||||
- **Pinning:** versions can be pinned (`POST /versions/{id}/pin`) and are never removed.
|
||||
|
||||
- **Thinning (grandfather-father-son):** keep all versions from the last 24 h, hourly for 7 days, daily for 30 days, weekly for 6 months, monthly after that. Fully configurable.
|
||||
- **Stale file purge:** remove files whose path no longer exists on disk *and* that have had no new version for N days (default 90).
|
||||
- **Stale project purge:** remove a whole project that has had no activity for N days.
|
||||
- **Intermediate purge:** for a file or glob, remove all versions between two timestamps, or all except the first and last of each day.
|
||||
- **Explicit delete:** a specific version, file, or project.
|
||||
- **Garbage collection:** blobs not referenced by any remaining version are deleted from WebDAV in a separate mark-and-sweep pass with a grace period (default 7 days), so a crash during purge can never lose a referenced blob.
|
||||
- **Pinning:** versions can be pinned (`POST /versions/{id}/pin`) and are never removed by automatic policies.
|
||||
### 4.8.1 Disaster recovery (total local loss)
|
||||
|
||||
1. Reinstall, then `POST /api/v1/config/remote/adopt` with the old directory to take
|
||||
it over on the new machine.
|
||||
2. `POST /api/v1/admin/reindex` — rebuilds the whole index from remote manifests
|
||||
(async; poll its status). Everything comes back `durable`.
|
||||
3. Restore normally (single, bulk, or point-in-time). Blob bytes stream down from
|
||||
WebDAV on demand and are re-spooled locally, so the first restore of each file
|
||||
needs the remote reachable; after that it is local again.
|
||||
4. Restored SQLite files are integrity-checked before being written; a stored
|
||||
version that fails the check is refused (`unrestorable`) instead of written.
|
||||
|
||||
### 4.9 Agent skill download
|
||||
|
||||
@@ -364,9 +390,9 @@ All endpoints are JSON, under `/api/v1` (the docs, schema and skill endpoints ar
|
||||
| GET / DELETE | `/roots/{id}` | Root status (watch count, mode `inotify`/`polling`/`missing`, baseline progress) / stop monitoring |
|
||||
| POST | `/roots/{id}/rescan` | Force a reconciliation scan |
|
||||
| POST | `/config/remote/test` | Test the remote without saving |
|
||||
| GET / PATCH | `/config/upload` | Concurrency, rate and retry settings |
|
||||
| GET / PATCH | `/config/filters` | Ignore rules |
|
||||
| GET / PATCH | `/config/retention` | Retention policies |
|
||||
| GET / PATCH | `/config/upload` | Not implemented |
|
||||
| GET / PATCH | `/config/filters` | Not implemented |
|
||||
| GET / PATCH | `/config/retention` | Not implemented (see `retention.keep_days` in config.toml) |
|
||||
| POST | `/filters/test` | Check whether paths would be accepted |
|
||||
| POST | `/snapshots` | Submit a snapshot (single) |
|
||||
| POST | `/snapshots/batch` | Submit up to 100 snapshots |
|
||||
@@ -377,21 +403,26 @@ All endpoints are JSON, under `/api/v1` (the docs, schema and skill endpoints ar
|
||||
| GET | `/versions/{id}/content` | Raw content |
|
||||
| GET | `/diff?from=&to=` | Diff two versions (`to=disk` for the working copy) |
|
||||
| GET | `/projects/{id}/tree?as_of=` | Project state at a point in time |
|
||||
| GET | `/search?q=` | Full-text search |
|
||||
| GET | `/search?q=` | Not implemented |
|
||||
| POST | `/restores` | Create a restore plan (dry run) |
|
||||
| POST | `/restores/{plan_id}/execute` | Run a restore plan |
|
||||
| POST | `/versions/{id}/pin` / `DELETE` | Pin / unpin |
|
||||
| POST | `/purge` | Run a purge (criteria + `dry_run`) |
|
||||
| DELETE | `/versions/{id}`, `/files?path=`, `/projects/{id}` | Explicit deletion |
|
||||
| POST | `/purge` | Not implemented (use `/admin/retention/run` + `/admin/gc`) |
|
||||
| DELETE | `/versions/{id}`, `/files?path=`, `/projects/{id}` | Not implemented (use `/forget`) |
|
||||
| GET | `/stats` | Files, versions (per source, per hour for 24 h), storage, dedup/compression, top files and projects |
|
||||
| GET | `/progress` | Active and last scans (percent, files/s, ETA), upload queue (percent, rate, ETA, errors), monitor counters |
|
||||
| GET | `/progress/stream` | Server-sent events with `/progress` every `interval` seconds |
|
||||
| GET | `/dashboard` | Web dashboard (token via `#token=` fragment or prompt) |
|
||||
| POST | `/forget` | Delete all history below a path (`dry_run` by default) |
|
||||
| POST | `/admin/gc` | Garbage-collect remote blobs |
|
||||
| POST | `/admin/reindex` | Rebuild the index from remote manifests |
|
||||
| GET | `/admin/queue` | Upload queue status |
|
||||
| GET | `/agent/skill`, `/agent/skill.md` | Agent skill package |
|
||||
| POST | `/admin/gc` | Garbage-collect orphan blobs, local + remote (`dry_run` by default) |
|
||||
| POST | `/admin/reindex` | Rebuild the index from remote manifests (async, 202 + status) |
|
||||
| GET | `/admin/reindex/{task_id}` | Reindex task status |
|
||||
| POST | `/admin/retention/run` | Thin history per keep-days (default 5, `dry_run` by default) |
|
||||
| POST | `/admin/purge-remote` | Delete this installation's remote data (date-guarded, async, 202) |
|
||||
| GET | `/admin/purge-remote/{task_id}` | Purge task status |
|
||||
| GET | `/metrics` | Prometheus text metrics (behind bearer auth) |
|
||||
| GET | `/admin/queue` | Not implemented (use `/progress` upload section) |
|
||||
| GET | `/agent/skill`, `/agent/skill.md` | Not implemented |
|
||||
| GET | `/docs`, `/redoc`, `/openapi.json` | API documentation |
|
||||
|
||||
Errors use RFC 9457 problem details (`application/problem+json`) with stable `type` codes (e.g. `file-too-large`, `path-ignored`, `remote-unreachable`, `rate-limited`).
|
||||
@@ -420,7 +451,12 @@ Indexes on `(file_id, captured_at)`, `files(path)`, `versions(blob_sha256)`. Sch
|
||||
|
||||
- **Loopback only.** The server refuses to start if configured to bind to a non-loopback address.
|
||||
- **Local authentication.** Other local users and processes (including browsers, via DNS rebinding) can reach `127.0.0.1`. Every request therefore needs a bearer token stored in `~/.config/versiond/credentials` (`0600`). The `Host` header is checked against `127.0.0.1:9922`/`localhost:9922`, and CORS is disabled.
|
||||
- **Secrets in backed-up files.** `.env` and similar files contain credentials and are sent off-machine. Client-side encryption is therefore **always on**: blobs and manifests are encrypted with AES-256-GCM (random 96-bit nonce) using a 64-byte key in `~/.config/versiond/backup.key` (`0600`). Remote blob names are HMAC-SHA-256 of the content hash, so the server cannot confirm guesses of known content. **The key is the only way to read the remote copy**: `versiond key export` prints it, and it must be stored off the machine.
|
||||
- **Remote storage is unencrypted by design.** `.env` and similar files are sent
|
||||
off-machine as compressed (not encrypted) blobs, and blob names are plain
|
||||
SHA-256 content hashes. The threat model is a disk-encrypted storage box on a
|
||||
trusted network (HTTPS/Basic-auth WebDAV), not an untrusted server. There is
|
||||
no backup key and nothing extra to lose: any WebDAV client can read the
|
||||
remote, and the local spool + SQLite index is sufficient to restore.
|
||||
- **Credential storage.** WebDAV secrets are kept in the `0600` credentials file or, when available, the Secret Service/keyring. They never appear in logs or API responses.
|
||||
- **Restore path safety.** See §4.7.
|
||||
- **Audit log.** All configuration changes, restores and purges are recorded with timestamp and client identity.
|
||||
@@ -476,7 +512,7 @@ versiond skill install # install the agent skill into ~/.claude/sk
|
||||
## 11. Milestones
|
||||
|
||||
1. **M1 – Core:** API skeleton, filters, SQLite index, local spool, root registration, inotify monitor with pruned watches, baseline and reconciliation scan, push snapshots, history, diff, single restore.
|
||||
2. **M2 – Remote:** WebDAV configuration, automatic unique remote directory and claim, upload workers, manifests, durability tracking, encryption.
|
||||
2. **M2 – Remote:** WebDAV configuration, automatic unique remote directory and claim, upload workers, manifests, durability tracking (no encryption).
|
||||
3. **M3 – Control:** coalescing, rate limits, bulk restore plans, retention, purge, GC.
|
||||
4. **M4 – Agent & ops:** skill package, CLI, `install` command, metrics, polling fallback, reindex/adopt.
|
||||
|
||||
@@ -485,7 +521,6 @@ versiond skill install # install the agent skill into ~/.claude/sk
|
||||
## 12. Open Questions
|
||||
|
||||
- Should `~/` be offered as a one-click root, or should the user be pushed toward choosing specific folders (lower watch count, less noise)?
|
||||
- Is 200 KB a hard product limit, or should it be configurable with 200 KB as the default?
|
||||
- Should other backends (S3, SFTP, local directory) be supported behind the same storage interface? WebDAV-only is the v1 scope.
|
||||
|
||||
---
|
||||
@@ -496,7 +531,7 @@ versiond skill install # install the agent skill into ~/.claude/sk
|
||||
|
||||
- It solves a real, current problem: AI agents rewrite files quickly and destructively, often outside Git commits. A "before every write" safety net with bulk point-in-time restore is useful, and editor local-history features (JetBrains, VS Code Timeline) do not cover agent edits across tools.
|
||||
- Shipping a skill so the agent can operate its own undo system is a smart idea and the most distinctive part of the design.
|
||||
- The 200 KB cap and the ignore rules keep the problem small and cheap, so it can stay running all the time.
|
||||
- The 10 MiB cap and the ignore rules (deps/build/caches out, user data in) keep the problem small and cheap, so it can stay running all the time.
|
||||
- Coalescing burst writes (keep first and last) is the right approach to versioning noise.
|
||||
|
||||
**Weaknesses and risks**
|
||||
@@ -504,7 +539,9 @@ versiond skill install # install the agent skill into ~/.claude/sk
|
||||
- Overlaps heavily with existing tools: Git (plus `git stash`/autocommit tools), restic/borg/kopia, Nextcloud's own file versioning on the very WebDAV server it targets, and editor local history. The differentiation has to be the agent integration and pre-write capture, and the README should say so.
|
||||
- inotify only reports changes after they happen, so capturing the "before" state relies on a complete baseline and a reconciliation scan. Changes made while the service is down are captured only as their final state.
|
||||
- The inotify watch limit is shared with IDEs and dev servers. Very large roots fall back to polling, which costs more.
|
||||
- Sending `.env` files to a remote server without encryption would be a security liability. That is why encryption is on by default in this spec.
|
||||
- `.env` files are stored unencrypted on the remote by design. The storage box is
|
||||
assumed disk-encrypted on a trusted network; do not point versiond at an
|
||||
untrusted server.
|
||||
- WebDAV is a slow, chatty backend for many small objects. Content-addressing plus batched manifests reduce this, but performance against Nextcloud-class servers needs to be measured early.
|
||||
- The original text left out authentication, failure handling, restore safety and data model. All of these are filled in above, but they roughly triple the scope compared with how the idea first read.
|
||||
|
||||
|
||||
@@ -0,0 +1,39 @@
|
||||
# Versiond should encrypt locally, hide metadata
|
||||
|
||||
For a local Linux daemon that stages small files and syncs to a dumb WebDAV store, the winning pattern is Kopia-style client-side envelope encryption with **AES-256-GCM as default and ChaCha20-Poly1305 as portable fallback**, applied after local spool-compression in the strict order hash then compress then encrypt, with per-blob random nonces and per-content keys derived by HMAC/HKDF. Restic proves that bespoke **AES-256-CTR plus Poly1305-AES in Encrypt-then-MAC with 16-byte random IV per file** works ([Source](https://github.com/restic/restic/blob/master/doc/design.rst)), Borg 1.x proves that **AES-256-CTR plus HMAC-SHA256 with 8-byte counter IVs and reservation tracking** scales poorly to multiple writers ([Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)), and Borg 2 plus Kopia converge on standard AEAD precisely to escape that complexity with **AES-256-OCB or ChaCha20-Poly1305 with session keys** ([Source](https://github.com/borgbackup/borg/wiki/Borg-2.0)) and **AES256-GCM-HMAC-SHA256 default, immutable after creation** ([Source](https://github.com/kopia/kopia/blob/master/site/content/docs/FAQs/_index.md)). Duplicity delegates entirely to **GnuPG OpenPGP hybrid encryption of tar volumes** ([Source](https://man.archlinux.org/man/duplicity.1.en)) and therefore offers no daemon-friendly manifest story. Because WebDAV provides no compute, no trustworthy mtime or hash, and no append-only enforcement, versiond must do what Kopia already does for WebDAV-class stores — encrypt, pack into **20–40 MB randomly named blobs with server-opaque indexes** ([Source](https://kopia.io/docs/advanced/architecture/)), authenticate manifests with the same keys as data, and add rollback resistance through a local monotonic counter plus server-side versioning rather than cryptography alone.
|
||||
|
||||
## AES-GCM and ChaCha20-Poly1305 replace bespoke CTR-MAC designs
|
||||
|
||||
Restic encrypts every repository object except key wrappers as **IV(16) || CIPHERTEXT || MAC(16) for 32 bytes overhead** with a fresh CSPRNG nonce per encryption ([Source](https://restic.readthedocs.io/en/v0.4.0/Design), and derives its three subkeys by splitting 64 scrypt-derived bytes into a **32-byte AES-256 key plus 16-byte AES key k plus 16-byte Poly1305 key r** ([Source](https://restic.readthedocs.io/en/stable/100_references.html)). Pack files then carry multiple independently encrypted blobs plus an encrypted header and little-endian header length, with blob type bits distinguishing plain versus compressed data and tree objects ([Source](https://restic.readthedocs.io/en/stable/100_references.html)). External review found the construction sane, quoting that **encryption is a first-class feature and the deduplication trade-off is worth it** ([Source](https://github.com/restic/restic/blob/master/doc/070_encryption.rst)), but the design is deliberately non-standard and carries Poly1305-AES masking complexity that modern AEAD avoids. Borg 1.x chose a different bespoke path with **AES-CTR-256 plus HMAC-SHA256 in Encrypt-then-MAC** ([Source](https://manpages.debian.org/bullseye/borgbackup/borg-init.1.en.html)) and a **TYPE(1) + HMAC(32) + NONCE(8) + CIPHERTEXT envelope using two different keys** ([Source](https://borgbackup.readthedocs.io/en/1.0-maint/internals.html)), where only 8 stored nonce bytes cap capacity at **2**64 times 16 bytes** ([Source](https://borgbackup.readthedocs.io/en/1.0-maint/internals.html)) and uniqueness depends on **reserving counter ranges by adding 4 GiB divided by 16 bytes and persisting through SaveFile** ([Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)). That reservation breaks under independent writers, with Borg documenting that **with multiple independent clients on the same repo Borg fails to provide confidentiality** ([Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)), which is exactly why Borg 2 abandons counters for session keys and standard AEAD. Borg 2 offers **aes256-ocb, chacha20-poly1305, authenticated-only and none modes** ([Source](https://borgbackup.readthedocs.io/en/latest/usage/repo-create.html)) to get **256-bit authenticated encryption client-side** ([Source](https://github.com/borgbackup/borg)) with goals stated as getting rid of CTR, using session keys, and moving to **Argon2 plus AES-OCB and ChaCha20-Poly1305** ([Source](https://github.com/borgbackup/borg/wiki/Borg-2.0)). Kopia made the same modern choice earlier with **DefaultAlgorithm AES256-GCM-HMAC-SHA256** ([Source](https://pkg.go.dev/github.com/kopia/kopia/repo/encryption)) and an explicit option for **CHACHA20-POLY1305-HMAC-SHA256 fixed at creation** ([Source](https://github.com/kopia/kopia/blob/master/site/content/docs/FAQs/_index.md)), where per-content keys come from **HMAC-SHA256 derivation with 28 bytes overhead** ([Source](https://github.com/kopia/kopia/blob/master/repo/encryption/aes256_gcm_hmac_sha256_encryptor.go)) and the repository format block itself is sealed with **Ke equals HKDF-SHA256 of Km over UniqueID for AES plus AD equals HKDF for CHECKSUM** ([Source](https://kopia.io/docs/advanced/encryption/)). For versiond the implication is direct: use a Go-available AEAD with random 12-byte nonces per blob and HKDF-separated data versus manifest keys, pick AES-GCM where **hardware AES favours HMAC-SHA256 paths and portable CPUs favour BLAKE2/ChaCha thinking** already documented by Borg ([Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)), and never invent a CTR-plus-MAC variant or shell out to GnuPG per sync like Duplicity does.
|
||||
|
||||
| System | Cipher and integrity | Nonce and key separation |
|
||||
|---|---|---|
|
||||
| restic | AES-256-CTR + Poly1305-AES Encrypt-then-MAC ([Source](https://github.com/restic/restic/blob/master/doc/design.rst)) | 16-byte random IV per file, 32 B overhead ([Source](https://restic.readthedocs.io/en/v0.4.0/Design)) |
|
||||
| Borg 1.x | AES-CTR-256 + HMAC-SHA256 or keyed BLAKE2b-256 ([Source](https://manpages.debian.org/bullseye/borgbackup/borg-init.1.en.html)) | 64-bit counter with reservation commit ([Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)) |
|
||||
| Borg 2 | AES-256-OCB or ChaCha20-Poly1305 AEAD ([Source](https://borgbackup.readthedocs.io/en/latest/usage/repo-create.html)) | Session keys to avoid global counter tracking ([Source](https://github.com/borgbackup/borg/wiki/Borg-2.0)) |
|
||||
| Kopia | AES256-GCM-HMAC-SHA256 default, ChaCha20 variant optional ([Source](https://github.com/kopia/kopia/blob/master/site/content/docs/FAQs/_index.md)) | Per-content HMAC-derived keys, format block with Ke and AD separation ([Source](https://kopia.io/docs/advanced/encryption/)) |
|
||||
| Duplicity | External GnuPG OpenPGP hybrid with per-recipient session keys ([Source](https://gnupg.org/ftp/blurbs/an-advanced-introduction-to-gnupg.pdf)) | Random session key s plus Enc_ri(s) per recipient, cipher negotiated by GnuPG ([Source](https://gnupg.org/ftp/blurbs/an-advanced-introduction-to-gnupg.pdf)) |
|
||||
|
||||
In transit the same client-side property carries over, which matters for dumb WebDAV over plain HTTP. Borg assumes **client trusted and repository untrusted with full read-write man-in-the-middle** ([Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)) and runs RPC **over system SSH with no own network code** ([Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)), while Kopia states that **all repository features are implemented client-side without any need for a custom server, thus encryption keys never leave the client** ([Source](https://github.com/kopia/repo/blob/master/README.md)). Restic promises that **everything except informational key-file metadata is encrypted and authenticated** ([Source](https://restic.readthedocs.io/en/stable/100_references.html)) and Duplicity claims volumes **will be safe from spying and modification by the server** via librsync plus GnuPG ([Source](https://packages.debian.org/bullseye/arm64/utils/duplicity)). Versiond therefore gets confidentiality even over untrusted HTTP, but still needs TLS or SSH tunnelling to blunt traffic analysis and to stop a network attacker from deleting or replaying blobs undetected, since object MACs detect tampering but not deletion.
|
||||
|
||||
## Hashing before compressing before encrypting preserves deduplication
|
||||
|
||||
All three content-defined systems converge on **split then hash then compress then encrypt then pack**, and every deviation is documented as an anti-pattern. Restic hashes plaintext for identity in saveBlob before zstd compression and Seal, and compresses unpacked index and snapshot JSON before encryption ([Source](https://github.com/restic/restic/blob/master/internal/repository/repository.go)), explicitly noting that **unpacked files like lock, index and snapshot files are also compressed before encryption** ([Source](https://github.com/restic/restic/pull/3666)). Borg follows **id equals MAC of id_key over data, then compress, then AEAD encrypt with id and header as associated data** ([Source](https://borgbackup.readthedocs.io/en/latest/internals/security.html)) and documents that **compression is applied after deduplication, thus different methods do not influence deduplication** ([Source](https://manpages.debian.org/trixie/borgbackup2/borg2-compression.1.en.html)). Kopia pipelines **splits into chunks, hash, compare, if new compress chunk, encrypt, pack** ([Source](https://kopia.io/docs/advanced/compression/)) after deliberately rewiring from read-split-compress-hash-store to **read-split-hash-compress-store to stop compressor drift from changing IDs** ([Source](https://github.com/kopia/kopia/pull/1076)). The rationale is physical: ciphertext from randomized encryption does not compress and destroys equality, so encryption must be last while hashing must be first, and compressing whole files before chunking avalanches small edits through the compressor and shifts cut points ([Source](https://github.com/restic/restic/issues/790)). Kopia even quantifies the accepted cost of chunk-then-compress versus whole-file compression with s2 and gzip ratios differing by a few points in exchange for shift-resilient dedup ([Source](https://kopia.io/docs/advanced/compression/)). Versiond spool-compressing many small files locally should copy this exactly: chunk the spool with a content-defined splitter, hash plaintext for the dedup index, compress each new chunk independently, then AEAD-encrypt.
|
||||
|
||||
Codec handling differs in ways versiond can borrow selectively. Restic repository v2 supports **zstandard only for data and tree blobs** ([Source](https://restic.readthedocs.io/en/stable/100_references.html)) negotiated per run as **off, fastest, auto default, better, max** mapped to klauspost compress levels ([Source](https://github.com/restic/restic/blob/master/internal/repository/repository.go)), recording compressed blobs with an extra **plaintext length field plus uncompressed_length in the index** ([Source](https://restic.readthedocs.io/en/stable/100_references.html)) so that **compressed and uncompressed blobs may be mixed in one pack** without re-uploading old data ([Source](https://restic.readthedocs.io/en/stable/100_references.html)). Borg records **ctype plus clevel plus csize plus size per object** for **none, lz4, zstd, zlib and lzma** ([Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html)), defaults to **lz4** and explicitly allows the first writer to fix a chunk's stored compression while later runs mix freely ([Source](https://manpages.debian.org/trixie/borgbackup2/borg2-compression.1.en.html)), plus wrappers for **auto selection and obfuscate padding with about 12 percent overhead** ([Source](https://manpages.debian.org/trixie/borgbackup2/borg2-compression.1.en.html)). Kopia disables compression by default, controls it per policy path, offers a wide menu from **s2-default through pgzip to zstd-better-compression with zstd recommended** ([Source](https://kopia.io/docs/faqs/)), and since index v2 keeps **per-content compression IDs in manager bookkeeping visible via content list and stats** rather than in the ID itself ([Source](https://github.com/kopia/kopia/pull/1076)), so **dedup is unaffected because hashing precedes compression** ([Source](https://kopia.discourse.group/t/deduplication-and-compression/714)). For versiond the practical synthesis is to store one zstd default for text-like spool files plus s2 for speed-sensitive paths, record codec and level outside the content ID exactly as Borg and Kopia do, freely mix codecs across blobs, and skip compression when output grows, mirroring Kopia storing the original when compressed output is larger ([Source](https://kopia.io/docs/advanced/compression/)). Integrity then comes free from the AEAD tag plus content addressing: restic verifies **filenames as hex SHA-256 of ciphertext with sha256sum** and checks structure with optional read-data payload reads ([Source](https://restic.readthedocs.io/en/stable/100_references.html)), Borg binds **chunk ID as associated data so the server cannot swap content under an ID** ([Source](https://borgbackup.readthedocs.io/en/latest/internals/security.html)), and Kopia packs blocks with a trailing local index plus top-level index for rebuilds ([Source](https://kopia.io/docs/advanced/architecture/)).
|
||||
|
||||
## Envelope encryption with scrypt or Argon2id decides recoverability
|
||||
|
||||
Every system separates a random master secret from a password-derived wrapping key, but parameters and ergonomics decide whether versiond operators survive key loss. Restic pins its key-file example to **scrypt N equals 65536, r equals 8, p equals 1** and rejects any other KDF ([Source](https://github.com/restic/restic/blob/9e2d60e2/internal/repository/key.go)), calibrating stronger parameters only when adding keys ([Source](https://github.com/restic/restic/blob/9e2d60e2/internal/repository/key.go)). Borg 1.x wraps keys with **PBKDF2-HMAC-SHA256 plus AES-CTR with zero IV and HMAC** ([Source](https://github.com/borgbackup/borg/blob/86fd77fd/src/borg/legacy/crypto/key.py)), while Borg 2 derives a **256-bit key-encryption key with Argon2 plus random 256-bit salt then ChaCha20-Poly1305 with constant IV zero** ([Source](https://borgbackup.readthedocs.io/en/2.0.0b9/internals/security.html)) and makes **Argon2id the default with PBKDF2 kept for compatibility** ([Source](https://git.uninsane.org/shelvacu-mirrors/borg/commit/08f82ee40867f605ca6994db6dc32218d7b85cbc)). Kopia seals its format block with **Km equals PBKDF over UniqueID then Ke and AD via HKDF-SHA256 for AES and CHECKSUM domains** ([Source](https://kopia.io/docs/advanced/encryption/)), documents only **scrypt-65536-8-1 as supported at the moment** ([Source](https://kopia.io/docs/advanced/encryption/)), and in 2025 added **configurable PBKDF with 600000 iterations default for PBKDF2 and 64 MB default for scrypt** ([Source](https://github.com/kopia/kopia/pull/5145)). Ancillary systems sharpen the contrast: rclone crypt derives **80 bytes via scrypt N equals 16384, r equals 8, p equals 1** ([Source](https://rclone.org/crypt/)) and Tarsnap generates keys locally, using scrypt only to optionally wrap the key file ([Source](https://www.tarsnap.com/man-tarsnap-keygen.1.html)). Versiond should therefore follow Borg 2 on new installs with Argon2id where libargon2 is available and otherwise scrypt-65536-8-1, always with a random 32-byte UniqueID salt that doubles as HKDF info to prevent cross-repo correlation ([Source](https://kopia.io/docs/advanced/encryption/)).
|
||||
|
||||
Storage location and rotation complete the choice. Restic keeps **keys directory JSON in the repo with hostname, KDF params, salt and encrypted data** ([Source](https://restic.readthedocs.io/en/stable/100_references.html)), Kopia keeps a **format blob with UniqueID, keyAlgo and encryptedBlockFormat in storage plus a local connection config without the password** ([Source](https://kopia.io/docs/reference/command-line/)), and Borg alone offers **repokey in the repo versus keyfile under home with move via key change-location** ([Source](https://borgbackup.readthedocs.io/en/2.0.0b6/usage/rcreate.html)). True rotation by re-encrypting data exists nowhere; all rotation re-wraps the same secret, whether through **restic key add and passwd managing multiple access keys sharing one master** ([Source](https://restic.readthedocs.io/en/stable/070_encryption.html)), **Borg change-passphrase only re-locking the same secrets** ([Source](https://borgbackup.readthedocs.io/en/stable/usage/key.html)), **Kopia change-password re-encrypting the format block while still connected** ([Source](https://kopia.io/docs/faqs/)), or rclone requiring full re-upload to change passwords ([Source](https://rclone.org/crypt/)). No system documents native Shamir sharing. Loss is therefore uniformly fatal by design: restic repeats that **losing your password means data is irrecoverably lost** ([Source](https://restic.readthedocs.io/en/stable/030%5Fpreparing%5Fa%5Fnew%5Frepo.html)), Kopia states **there is no way to recover a forgotten password because only you know it** ([Source](https://kopia.io/docs/faqs/)), Borg demands **both key and passphrase with offsite export** ([Source](https://borgbackup.readthedocs.io/en/stable/usage/init.html)), and Tarsnap warns that **the key file contains the only copy** ([Source](https://www.tarsnap.com/faq.html)). For versiond this means shipping multi-key wrapping from day one so a daemon key and an offline admin recovery key share the same master, exporting with **paper printouts with per-line checksums or QR plus separate passphrase storage** as Borg documents ([Source](https://borgbackup.readthedocs.io/en/master/usage/key.html)), caching the daemon password in the Linux keyring only as Kopia does transiently ([Source](https://kopia.io/docs/reference/command-line/)), and documenting that a leaked master forces new-repo migration because **revocation without re-encrypting the whole repository is impossible** ([Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)).
|
||||
|
||||
## Random packing and padding hide far more than encryption alone
|
||||
|
||||
Encryption hides bytes but every system leaks sizes, counts, timing and directory shape to a curious WebDAV host, and versiond must treat that as its main residual risk. Restic filenames are **lowercase hex storage IDs verifiable by sha256sum** ([Source](https://restic.readthedocs.io/en/stable/100_references.html)) under predictable **config, keys, locks, snapshots, index and data prefixes** ([Source](https://restic.readthedocs.io/en/stable/100_references.html)), with the server able to **infer which packs probably contain trees via access patterns and infer backup sizes via timestamps** ([Source](https://restic.readthedocs.io/en/stable/100_references.html)), including pre-0.18 chunk-size fingerprinting only partly fixed by **randomly assigning chunks to pack files** ([Source](https://restic.readthedocs.io/en/stable/100_references.html)). Borg is explicit that **the repository does not hide the size of chunks** and that single-chunk small files plus inode-order proximity let an observer correlate neighbours unless the operator enables the **obfuscate pseudo-compressor that pads with zero bytes** ([Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)). Rclone crypt is weakest here by design, stating it **does not encrypt file length within 16 bytes nor modification time used for syncing** ([Source](https://rclone.org/crypt/)) and encrypting names deterministically with **EME-AES-256 plus base32 so identical names yield identical ciphertexts** ([Source](https://rclone.org/crypt/)), while **plain WebDAV does not support modified times nor hashes** except vendor extensions ([Source](https://rclone.org/webdav/)). Kopia leaks least by default because **pack files have random names and sizes generally unrelated to content due to splitting and merging** ([Source](https://kopia.io/docs/advanced/architecture/)), exposing only **p versus q versus x prefixes for data, metadata and indices** ([Source](https://kopia.io/docs/advanced/architecture/)).
|
||||
|
||||
Active-server protection is where versiond needs append-only manifests plus local verification. Borg builds a Horton DAG where **every object is referenced by its parent via plaintext-MAC ID up to a TAM-signed manifest** protecting the fixed **000…000 manifest ID with HKDF-derived TAM keys** ([Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)), tracks nonces to catch replay, and ships **append-only mode that never overwrites committed segments** with transaction-log rollback ([Source](https://raw.githubusercontent.com/borgbackup/borg/1.4.5/docs/usage/notes.rst)), yet even Borg suffered **CVE-2023-36811 allowing faked archives with repo write** ([Source](https://nvd.nist.gov/vuln/detail/CVE-2023-36811)). Restic authenticates objects but concedes it is **not designed to protect against attackers deleting files and not designed to detect timestamp-grouped deletion** ([Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)), warns that write access lets an attacker **create garbage snapshots covering modified files and wait for forget to prune correct ones** ([Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)), and leaves management-key separation as open issue discussion ([Source](https://github.com/restic/restic/issues/5041)). Kopia relies on **snapshot verify walks plus sampled content downloads during maintenance** ([Source](https://kopia.io/docs/advanced/consistency/)) and pushes ransomware safety to **provider restricted keys plus COMPLIANCE object-lock with retention periods** ([Source](https://kopia.io/docs/advanced/ransomware-protection/)), which on WebDAV has no equivalent since **object-lock currently only supports S3 repos** ([Source](https://kopia.io/docs/advanced/ransomware-protection/)). For versiond on dumb WebDAV the synthesis is concrete: pack spool files into fixed-size-class encrypted blobs with random names, pad small manifests, jitter upload times, never trust server mtime, verify after write with GET plus AEAD open before deleting staging, keep an **encrypted local cache to prevent metadata leaks** as restic does ([Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)), sign each append-only manifest with the master key and chain it to the prior manifest hash, enforce a local monotonic epoch to detect rollback, and compensate for missing server append-only with WebDAV versioning or a second prune identity that alone may delete, mirroring the split Borg recommends for serve keys.
|
||||
|
||||
## Conclusion
|
||||
|
||||
The durable lesson is that cipher negotiation matters less than pipeline order, identifier keying, and operational separation of sync versus prune credentials. A daemon that hashes with a keyed MAC, compresses per chunk with an explicitly recorded codec, seals with a standard AEAD under HKDF-separated keys, and appends chained manifests into randomly sized packs inherits the best of Borg's authentication thinking and Kopia's WebDAV-ready packing without repeating restic's bespoke crypto or Duplicity's GnuPG fragility. What remains honestly unsolved is metadata-minimal versioning on a store that willingly reveals sizes and timestamps; the next hardening step is not a stronger cipher but deterministic padding classes, falsified access ordering, and provider-enforced immutability that make deletion the only move left to an untrusted host.
|
||||
@@ -0,0 +1,34 @@
|
||||
# Borrow Time Machine's pressure-driven thinning discipline
|
||||
|
||||
Leading backup systems agree more than their docs admit: declarative time policy decides what may die, cheap metadata marking runs often, expensive reclamation runs rarely, capacity only forces the issue when the target fills, and local resource politeness comes from static ceilings plus OS deferral rather than true autotuning. For a local Linux file-versioning daemon syncing to WebDAV, that means copying Kopia's built-in maintenance rhythm, Time Machine's pre-backup thinning with padding, WebDAV PROPFIND quota checks, and a systemd slice with conservative defaults — because none of the surveyed tools offers a keep-usage-below-70% knob, and the one system that deletes on full does so without any percentage at all.
|
||||
|
||||
## Cheap marking runs daily while reclamation waits weeks
|
||||
|
||||
The universal split is cheap expiration versus expensive reclamation, with very different cadences. In restic **forget only deletes snapshot metadata while prune repacks data and is very time-consuming for remote repositories**, bounded by **--max-unused default 5%** ([Source](https://restic.readthedocs.io/en/stable/060_forget.html)). Borg made the same split explicit after 1.2 where **repository disk space is not freed until you run borg compact** and docs tell operators to run it regularly but not after every command, roughly monthly or when space is needed ([Source](https://borgbackup.readthedocs.io/en/stable/usage/compact.html)). Borg's default **compact --threshold 10%** compacts only segments above that saving, with 2.x gating further to reclaimable space above threshold divided by five ([Source](https://borgbackup.readthedocs.io/en/stable/usage/compact.html)). Kopia mirrors the design with **quick maintenance hourly and full maintenance every 24h by default**, where quick work never deletes metadata without another copy existing and full work does Snapshot GC plus pack compaction over several cycles ([Source](https://kopia.io/docs/advanced/maintenance/)). Veeam applies short-term retention inline at the end of each job session while GFS deletions fall to a nightly background job, with health check verifying **only the latest restore point per chain on a monthly default** rather than whole history ([Source](https://helpcenter.veeam.com/docs/vbr/userguide/backup_health_check.html)).
|
||||
|
||||
Scheduling ownership falls on a spectrum from manual to fully automatic that the daemon should learn from. Restic and Borg ship **no built-in scheduler at all**, so retention becomes a wrapper convention of backup then forget plus prune or create then prune then compact, typically driven by daily cron such as **0 0 * * * with retention plus check monthly** ([Source](https://gist.github.com/perfecto25/f528f8d14e1c4b6e2a912513539a5af7)). Borgmatic warns its every-run prune plus compact plus check default suits small repos but large repos must decouple them to separate schedules ([Source](https://github.com/borgmatic-collective/borgmatic/blob/main/docs/how-to/deal-with-very-large-backups.md). Kopia inverted this by making maintenance automatic since v0.6.0, firing opportunistically whenever the client runs with **QuickCycle 1h and FullCycle 24h** tunable via maintenance set and pausable via pause flags ([Source](https://kopia.io/docs/advanced/maintenance/)). KopiaUI or server checks roughly every 10 minutes and does work only when due ([Source](https://github.com/kopia/kopia/issues/1439). Time Machine goes furthest with zero user-visible retention scheduling, keeping **hourly backups for 24 hours, daily for a month, weekly for previous months** and thinning continuously plus under pressure ([Source](https://support.apple.com/en-la/104984)). Veeam sits in the enterprise middle with a per-job scheduler plus GFS calendar and monthly health-check piggybacked on the first session of the scheduled day ([Source](https://helpcenter.veeam.com/docs/vbr/userguide/backup_health_check.html)). Safety rails are strikingly convergent and worth copying wholesale: every tool offers dry-run previews without making them default, restic refuses empty policies and requires **--unsafe-allow-remove-all plus filters** to wipe a group, Borg retains the oldest archive when no rule keeps anything and since 2.x offers **undelete until compact runs**, and Kopia elects a **single maintenance owner under exclusive lock with PackDeleteMinAge 24h** so GC takes multiple cycles unless defeated with safety none ([Source](https://kopia.io/docs/advanced/maintenance/)). Restic prune holds an exclusive lock that blocks backups while Borg and Kopia serialize through repo plus cache locks with break-lock and force overrides, and interrupted cheap passes rerun safely while expensive passes resume from checkpoints or recorded partial results ([Source](https://github.com/restic/restic/blob/master/doc/design.rst)).
|
||||
|
||||
## Time policy leads, capacity follows except when disk fills
|
||||
|
||||
Only Apple Time Machine makes capacity the primary retention driver, and even it leads with time. Backupd logs show the literal two-phase layering of **Starting pre-backup thinning: 53.57 GB requested (including padding)** then only if **No expired backups exist - deleting oldest backups to make room** ([Source](https://serverfault.com/posts/39310/revisions)). The published ladder keeps first-of-day beyond 24 hours and first-of-week beyond 30 days, deleting oldest weeklies only when space is needed ([Source](https://discussions.apple.com/thread/251224074)). There is **no published percentage watermark**, deletion is on-demand pre-backup rather than a high-water mark ([Source](https://support.apple.com/guide/mac-help/if-the-time-machine-backup-disk-is-full-mh15137/mac)). Per-destination caps come only from **tmutil setquota DESTINATION_ID QUOTA_IN_GB**, which thins to fit on the next backup ([Source](https://superuser.com/questions/445579/how-do-i-trim-time-machine-backup-history)). Local snapshots are purgeable space automatically reclaimed as needed with urgency-controlled manual relief via **thinlocalsnapshots mount_point purge_amount urgency 1-4** ([Source](https://ss64.com/mac/tmutil.html)). Every other surveyed system treats capacity as monitoring, placement, or manual pruning rather than a retention input. Veeam retention defines **the number of restore points to keep on performance and capacity extents**, purging capacity-tier blocks on the next offload session, and when extents fill guidance is simply to add a new extent ([Source](https://helpcenter.veeam.com/docs/vbr/userguide/capacity_tier_retention.html)). Veeam placement prefers the extent with fewest chains breaking ties by most free space, with **priority always to complete a backup even by violating Data-Locality** ([Source](https://veeam-best-practices-guide-v9.readthedocs.io/resource_planning/repository_sobr.html)). Forecast thresholds such as a worked **Repository Free Space 30%** example flag repositories that will run out rather than enforcing deletion ([Source](https://helpcenter.veeam.com/docs/mp/reports/capacity_planning_for_backup_repositories.html?ver=9a)). S3 Lifecycle rules use **Days, Date, and NoncurrentVersionExpiration per prefix or tag** with no capacity trigger at all ([Source](https://docs.aws.amazon.com/AmazonS3/latest/API/API_LifecycleRule.html)). Hetzner Storage Box quotas are fixed plan sizes with **10, 20, 30, or 40 manual plus automatic snapshot slots** consuming the same plan capacity alongside live data, offering no per-subaccount quota and no auto-thinning beyond setting directories read-only ([Source](https://docs.hetzner.com/storage/storage-box/snapshots/)). ZFS tooling is count-based with **-k keep NUM recent snapshots** and splits take, prune, and cron while capacity monitoring only feeds Nagios-style alerts ([Source](https://manpages.debian.org/bookworm/zfs-auto-snapshot/zfs-auto-snapshot.8.en.html)). The percentage numbers that do exist are health or performance floors, not auto-delete triggers, with Ubuntu warning that **minimum free space to preserve ZFS performance is 20%** while the pool sits at 10% ([Source](https://superuser.com/questions/1736700/how-do-i-remove-old-zfs-snapshots)) and OpenZFS advising to **keep pool free space above 10% to avoid metaslabs hitting 4%** where the allocator collapses from first-fit to best-fit ([Source](https://openzfs.github.io/openzfs-docs/Performance%20and%20Tuning/Workload%20Tuning.html)). Borg likewise has no quota-aware prune and warns that **repository disk space is not freed until you run borg compact**, urging operators to ensure always plenty of free space and to prune plus compact regularly ([Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)).
|
||||
|
||||
| System | Primary retention driver | Capacity role | Published percentage trigger |
|
||||
|---|---|---|---|
|
||||
| Time Machine | Time ladder then oldest-first on full | Pre-backup thinning with padding | None, allocation-driven |
|
||||
| Veeam SOBR | Restore-point counts plus GFS | Placement plus alerting, add extent | 30% forecast example only |
|
||||
| S3 Lifecycle | Days and versions | None | None |
|
||||
| ZFS, sanoid, Borg, restic | Counts and time windows | Monitoring plus manual prune | 20%, 10%, 4% health floors only |
|
||||
|
||||
Failure at 100% full explains why a headroom design matters more than any target percentage. Time Machine cancels the run with **Stopping backup. Backup canceled. Compacting backup disk image to recover free space** then retries as a fresh standard backup, surfacing **This backup is too large for the backup disk. The backup requires XX GB but only YY GB are available** when nothing is freeable ([Source](https://serverfault.com/posts/39310/revisions)). Veeam guards production datastores with **Skip VMs when free disk is below plus a hard 2 GB floor** even when skipping is disabled ([Source](https://www.veeam.com/kb4379?ad=in-text-link)). Borg is the cautionary tale that **if you do run out of disk space, it can be hard or impossible to free space, because Borg needs free space to operate - even to delete backup archives** ([Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)). ZFS adds the snapshot-pinned trap where deleting a file present in a snapshot gains no space and can itself return unexpected ENOSPC ([Source](https://docs.oracle.com/cd/E19120-01/open.solaris/817-2271/gayra/index.html)). Restic historically panicked on temp-pack writes with **no space left on device** and retried blindly on ENOSPC before rest-server fixes ([Source](https://github.com/restic/restic/issues/611)). WebDAV at least fails cleanly with **507 Insufficient Storage** on PUT, MKCOL, MOVE, or COPY ([Source](http://www.webdav.org/specs/rfc4331.html)). The robust pattern combining all three headroom elements is pre-backup estimate plus padding as Time Machine does, a reserved-space tripwire as Veeam 2 GB, Borg repo-space, and ZFS 10% do, and a recovery path that works at 100% as Time Machine compact-and-retry does and Borg notably lacks. Crucially for snapshot-capable targets, a capacity pruner must delete snapshots themselves oldest-first rather than thinning live data, because Hetzner snapshots silently consume the same plan quota ([Source](https://docs.hetzner.com/storage/storage-box/snapshots/)).
|
||||
|
||||
## WebDAV PROPFIND gives daemon a cheap capacity signal
|
||||
|
||||
WebDAV is unusually lucky among dumb backends because RFC 4331 standardizes quota discovery. Servers expose **DAV:quota-available-bytes as maximum additional storage allocatable and DAV:quota-used-bytes including sub-resources**, warning that as available bytes approaches zero further allocations may be refused ([Source](https://datatracker.ietf.org/doc/html/rfc4331)). Nextcloud implements both properties retrievable via PROPFIND ([Source](https://github.com/nextcloud/documentation/blob/master/developer_manual/client_apis/WebDAV/basic.rst)). That makes a **Depth:0 PROPFIND for quota-used-bytes plus quota-available-bytes the cheapest correct pre-backup check** for Hetzner or Nextcloud targets, avoiding any tree walk. Alternatives on other transports show why this matters. S3 has no byte-quota endpoint with accounting left to CloudWatch and Storage Lens and lifecycle evaluated by a daily asynchronous scan ([Source](https://docs.aws.amazon.com/AmazonS3/latest/userguide/lifecycle-expire-general-considerations.html)). Hetzner exposes usage via console or API rather than uniformly in-protocol across its FTP, SFTP, rsync, Borg, SMB, and WebDAV paths ([Source](https://docs.hetzner.com/storage/storage-box)). Veeam polls extent free space but **free space data is only retrieved when no active tasks are assigned to an extent**, so placement under continuous load uses stale data ([Source](https://bp.veeam.com/vbr/3_Build_structures/B_Veeam_Components/B_backup_repositories/scaleout.html)). Borg and restic on SSH, SFTP, or B2 do no server quota query at all, directing users to local filesystem monitoring and **borg repo-space accounting** ([Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)). Restic cache mismanagement shows the cost of local-only accounting when cache at system paths can grow to **20 to 40 GB or roughly 3 to 10% of repo size** with no built-in cap ([Source](https://github.com/restic/restic/issues/4325)). Time Machine sidesteps the whole problem by asking the local filesystem how many bytes plus padding the next backup needs and thinning until that fits rather than tracking a remote percentage, a model the daemon should emulate by combining local size estimation with a PROPFIND confirmation before each sync.
|
||||
|
||||
## Static ceilings plus OS deferral beat true autotuning
|
||||
|
||||
No surveyed tool probes link speed, disk size, or free RAM to size packs, connections, or limits. Rate knobs are consistently static, opt-in, and unlimited by default. Restic exposes **--limit-upload and --limit-download in KiB per second defaulting to 0 meaning unlimited** and cannot change them mid-run without restart ([Source](https://restic.readthedocs.io/en/stable/manual_rest.html)). Borg 1.x exposes **--upload-ratelimit in kiByte per second default 0** over a token-bucket SleepingBandwidthLimiter with 0.1 second period capping bursts at twice the allowance ([Source](https://manpages.debian.org/bookworm/borgbackup/borg-common.1.en.html)). Borg 2.x moved shaping to **BORGSTORE_BANDWIDTH in bits per second default 0 plus BORGSTORE_LATENCY** ([Source](https://borgbackup.readthedocs.io/en/latest/faq.html)). Kopia exposes **repository throttle set with upload and download bytes per second plus concurrent reads, writes, and request rates** ([Source](https://kopia.io/docs/reference/command-line/common/repository-throttle-set/)). Duplicity has no native generic limit with workarounds via trickle or tc ([Source](https://bugs.launchpad.net/bugs/1291633)). Time Machine instead imposes mandatory kernel-level low-priority I/O where **backupd runs at IOPOL_THROTTLE for long-running background work throttled to protect higher policies** ([Source](https://eclecticlight.co/2022/01/20/why-time-machine-backups-can-be-interminably-slow/)). Disabling that global throttle lifted a measured copy phase from **193 to 332 MB per second and overall backup from 160 to 276 MB per second** ([Source](https://eclecticlight.co/2022/02/28/does-removing-i-o-throttling-make-backups-faster/)). The only resource-derived defaults anywhere are CPU-count-derived worker pools. Restic defaults to **file-read concurrency 2, blob-save concurrency equal to runtime.NumCPU, tree-save at 20 times blob concurrency, and backend connections 5 or 2 for local** ([Source](https://github.com/restic/restic/blob/de9136b29f86216bd3e41397d19b25f26b578833/internal/archiver/archiver.go)). Kopia defaults **--max-parallel-file-reads to logical core count** and scales s2 compressor concurrency similarly ([Source](https://kopia.io/docs/faqs/)). Time Machine scheduling scores each due activity every few seconds against temperature and load via DAS-CTS, dispatching only above threshold, with Apple Silicon background threads confined to Efficiency cores at **972 to 1332 MHz and roughly 90% E-core residency capping throughput near 300 to 400 items per second** ([Source](https://eclecticlight.co/2023/11/28/scheduling-and-dispatch-of-backups-and-other-background-activities/)). Raising concurrency without memory awareness is dangerous, with one restic test at connections 16 and reads 16 hitting **300 GB RAM and OOM** ([Source](https://github.com/restic/restic/issues/4477)). Constrained-box guidance therefore stays manual with cheap compression, lower parallelism, and larger packs. Borg defaults to **lz4 with a proposal to move to zstd level -4 capped at 4 threads** ([Source](https://github.com/borgbackup/borg/pull/10100). Kopia disables compression by default and benchmarks **s2-default at 4 GiB per second and 375 MiB RSS versus zstd at 323 MiB per second** on a 466 MiB corpus ([Source](https://kopia.io/docs/advanced/compression/)). Restic on SMR USB disks converges on **--read-concurrency 1 plus --pack-size 128 plus --no-cache with local connections 1**, trading longer single-pack uploads and 64 to 384 MiB staging for fewer files ([Source](https://forum.restic.net/t/first-backup-2-5tb-50-hours-can-i-improve-it/8288)). Upstream tools ship no restrictive systemd units at all, leaving the enforceable layer to downstream slices. Systemd semantics give **CPUQuota as percent of one CPU, CPUWeight 1 to 10000 default 100 for contention only, MemoryHigh as soft throttle versus MemoryMax as hard OOM kill, and IOWeight plus per-device bandwidth caps** ([Source](https://manpages.debian.org/bullseye/systemd/systemd.resource-control.5.en.html)). Community practice fills the gap with **Nice 17 plus sandboxing and RandomizedDelaySec 300** on restic timers and **CPUQuota 80% plus MemoryMax 2G** on segmented Borg services ([Source](https://github.com/JoZapf/segmented-borg-backup-system/blob/refs/heads/main/docs/SYSTEMD.md)). Modern guidance recommends **Nice 19 plus ionice class idle for soft priority with a backup.slice carrying CPUQuota 50 to 80%, MemoryHigh and Max 1 to 4G, IOWeight 10, and IOWriteBandwidthMax such as 30 MB per second**, noting ionice is ignored on default mq-deadline SSD schedulers and MemoryMax kills rather than slows ([Source](https://www.bigiron.cc/guides/process-priorities-nice-ionice-cgroup-quotas)).
|
||||
|
||||
## Conclusion
|
||||
|
||||
The evidence reframes the daemon as a Kopia-style insider with Time Machine instincts rather than a restic-style script kit: own its scheduler with hourly cheap and daily full passes, enforce time policy first with capacity-triggered oldest-first eviction guarded by a minimum-retention floor, and treat PROPFIND plus local size-plus-padding as the single pre-sync gate instead of inventing a 70% watermark no vendor validates. Resource modesty then becomes configuration rather than cleverness, with static unlimited-by-default throttles, core-count workers, s2 or lz4 compression, and a systemd slice doing the throttling the backup code cannot do for itself.
|
||||
@@ -0,0 +1,104 @@
|
||||
# Compression, Deduplication, and Integrity in Restic, BorgBackup, and Kopia
|
||||
|
||||
## Fixed-file vs content-defined chunking: algorithm, average chunk sizes, how dedup works with encryption
|
||||
|
||||
### Takeaway
|
||||
All three use content-defined chunking (CDC) by default with a fixed-size fallback option; deduplication is on plaintext hashes before compression/encryption (so no convergent encryption), with Borg and Kopia keying chunk IDs while restic uses plain SHA-256.
|
||||
|
||||
### Cited Findings
|
||||
- Restic splits each file independently with Rabin-fingerprint CDC over a 64-byte sliding window; a random irreducible polynomial is generated at `init` and stored as `chunker_polynomial` in `config` to harden against watermarking — [Source](https://restic.readthedocs.io/en/stable/100_references.html); background and worked example — [Source](https://restic.net/blog/2015-09-12/restic-foundation1-cdc/)
|
||||
- Restic chunking parameters: files <512 KiB are not split; blobs are 512 KiB–8 MiB with ~1 MiB average target; modified files only re-store changed blobs, robust to insertions at arbitrary offsets — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic 0.18.0 mitigates chunk-size fingerprinting (Alexeev/Percival/Zhang 2025) by randomly assigning chunks to pack files so attackers observing the repo cannot map chunk sizes to files — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic deduplication happens before encryption: blob ID is SHA-256 of plaintext; index maps plaintext hash to pack location; pack/index filenames are SHA-256 of ciphertext for accident detection only — [Source](https://forum.restic.net/t/how-are-blobs-deduplicated-with-encryption/6478); all content referenced by SHA-256 of plaintext, one blob holds data from only one file, multiple blobs packed per pack file — [Source](https://github.com/restic/restic/issues/2401)
|
||||
- Restic uses random per-repository master keys and random 16-byte IV per encryption (not convergent/deterministic encryption); identical plaintexts deduplicate via index lookup, not via identical ciphertexts — [Source](https://restic.readthedocs.io/en/stable/design.html); external audit notes AES-256-CTR + Poly1305-AES with separate keys — [Source](https://words.filippo.io/restic-cryptography/)
|
||||
- Borg default chunker is `fastcdc` (window-less keyed Gear hash); alternatives: `fixed` (fixed blocksize, optional different header block), `buzhash`/`buzhash64` (rolling Buzhash), `rabin-aes`/`toeplitz-aes`/`goldilocks-aes` (UHF-then-PRF: rolling universal hash + AES-128, cut decision only on AES output) — [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html); comparison table and guidance — [Source](https://borgbackup.readthedocs.io/en/master/internals/chunker.html)
|
||||
- Borg classic `buzhash` defaults: `CHUNK_MIN_EXP=19` (512 KiB), `CHUNK_MAX_EXP=23` (8 MiB), `HASH_MASK_BITS=21` (~2 MiB target), `HASH_WINDOW_SIZE=4095` bytes; tunable via `--chunker-params`; min/max clamp where rolling-hash cuts may occur — [Source](https://github.com/borgbackup/borg/blob/master/docs/misc/create_chunker-params.txt); `fastcdc` uses `CHUNK_MIN_EXP,CHUNK_MAX_EXP,HASH_MASK_BITS,NC_LEVEL` with normalized chunking — [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html)
|
||||
- Borg chunker secrets are derived from repository key material (buzhash table XOR seed stored encrypted in keyfile; AES chunkers derive table/polynomial/AES key from `id_key` per-chunker domain) to resist fingerprinting; 2025 research (Truong et al. CCS 2025 eprint 2025/558; eprint 2025/532) showed keyed rolling-hash-only chunkers allow key recovery, motivating AES-based chunkers — [Source](https://borgbackup.readthedocs.io/en/master/internals/chunker.html)
|
||||
- Borg deduplication is global across all archives/hosts/files on chunk `id_hash`, which is a keyed MAC over plaintext (`HMAC-SHA256` for `--id-hash sha256`, keyed `BLAKE3` for `--id-hash blake3`) using secret `id_key`, not plain hash, so attackers cannot confirm small-file presence without the key — [Source](https://borgbackup.readthedocs.io/en/latest/internals/security.html); file metadata (msgpacked items) is chunked with finer params and deduplicated the same way — [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html)
|
||||
- Borg 1.x-compatible dedup requires `buzhash`; `fastcdc`/`buzhash64` give same dedup but different cut points; `fixed` gives positional (not content-shift-resilient) dedup, suited to disk images — [Source](https://borgbackup.readthedocs.io/en/master/internals/chunker.html)
|
||||
- Kopia calls chunkers “splitters”: `BUZHASH`, `RABINKARP` (rolling-hash CDC) and `fixed`, selectable size 1M–8M — [Source](https://kopia.discourse.group/t/does-kopia-use-content-defined-chunking-cdc/1417); default object splitter is `DYNAMIC-4M-BUZHASH` (also the default for every `repository create` backend) — [Source](https://kopia.io/docs/reference/command-line/common/repository-create-filesystem/)
|
||||
- Kopia splitter history: original `DYNAMIC` splitter (silvasur/buzhash dependency) was deprecated for license reasons and replaced by a faster but incompatible buzhash implementation; old repos remain readable but large objects re-upload instead of deduping across the change — [Source](https://github.com/kopia/kopia/commit/03339c18afedb810f31320ef01c707e7acbdc374)
|
||||
- Kopia pipeline is split → hash → compare against index → discard if known; otherwise compress → encrypt → pack multiple small blocks into ~20–40 MB packs with random names; content hash default `BLAKE2B-256-128` — [Source](https://kopia.io/docs/advanced/compression/); pack/index architecture — [Source](https://kopia.io/docs/advanced/architecture/)
|
||||
- Kopia deduplication is unaffected by compression settings because hashing precedes compression since index v2 (see next section); same file under compressed and uncompressed policies still dedupes — [Source](https://kopia.discourse.group/t/deduplication-and-compression/714)
|
||||
|
||||
### Inferences
|
||||
- None of the three use convergent encryption (deterministic, plaintext-derived keys); all use random repository keys + random nonces/session keys, with dedup achieved by a client-side plaintext-hash index lookup before encryption.
|
||||
- Keyed chunk IDs (Borg MAC, Kopia HMAC-derived keys/format secret) hide confirmation-of-file attacks from repo-only attackers; restic's plain-SHA-256 blob IDs do not, which is part of why restic added pack-mixing and per-repo random polynomials.
|
||||
|
||||
### Gaps
|
||||
- Exact numeric defaults for Kopia `DYNAMIC-4M-BUZHASH` (min/max/expected size, window bytes) were not found in fetched docs; only the 4M-average naming and 1M–8M range were confirmed.
|
||||
- Whether Kopia content IDs themselves are HMAC-keyed (vs. plain hash + separately HMAC-derived encryption keys) could not be confirmed from fetched pages; per-content key derivation via HMAC-SHA256 is confirmed but ID-keying needs source-code confirmation.
|
||||
|
||||
## Which compression algorithms are supported, how is the algorithm recorded per object, and what happens mixing versions
|
||||
|
||||
### Takeaway
|
||||
Restic supports only zstd (repo v2, per-blob type byte + unpacked version byte); Borg supports none/lz4/zstd/zlib/lzma with per-object ctype+clevel metadata and free mixing; Kopia supports a large s2/pgzip/gzip/deflate/zstd matrix recorded per-content in index v2 and also freely mixable going forward.
|
||||
|
||||
### Cited Findings
|
||||
- Restic repo v2 adds compression; data and tree blobs may use zstandard only; v1 has no compressed types — [Source](https://restic.readthedocs.io/en/stable/100_references.html); changelog entry “Support compression for blobs (data/tree) and index/lock/snapshot files” — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic per-run negotiation: `--compression off|fastest|auto(default)|better|max` (also `RESTIC_COMPRESSION`); maps to klauspost/compress zstd levels `SpeedFastest`/`SpeedDefault`/`SpeedBetterCompression`/`SpeedBestCompression` with 512 KiB window and CRC disabled — [Source](https://github.com/restic/restic/blob/master/internal/repository/repository.go); `auto` is default and uses implementation default level — [Source](https://forum.restic.net/t/what-compression-is-used/6850); tuning doc confirms `auto` default for v2 repos — [Source](https://restic.readthedocs.io/en/stable/047_tuning_parameters.html)
|
||||
- Restic pack header per-blob recording: 1-byte type `0b00` data / `0b01` tree / `0b10` compressed-data / `0b11` compressed-tree, followed by `Length(encrypted_blob)||Hash(plaintext)` (uncompressed) or `Length(encrypted_blob)||Length(plaintext)||Hash(plaintext)` (compressed), little-endian uint32; index adds `uncompressed_length` only for compressed blobs — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic unpacked files (index/snapshot/lock): plaintext is `encoding_version||data` with 1-byte version; `[` (0x5b)/`{` (0x7b) mean “whole plaintext is JSON” (v1 back-compat); version `2` means zstd-compressed JSON; new v2 writes always version 2; implementation compresses unpacked data before encryption via `compressUnpacked`, and compresses tree blobs even when `--compression off` for data blobs (`if Compression != off || t != DataBlob`) — [Source](https://restic.readthedocs.io/en/stable/100_references.html); code — [Source](https://github.com/restic/restic/blob/master/internal/repository/repository.go); design rationale (null-byte/version-byte history) — [Source](https://github.com/restic/restic/pull/3666)
|
||||
- Restic mixing/compat: compressed and uncompressed blobs of same type may be mixed in one pack; in v2, data and tree blobs must be in separate packs; v1-repo data remains valid in v2 (no re-upload required); new data compressed per run-level, `prune` recompresses only repacked chunks; v2 repos unreadable by pre-compression restic versions; new repos default to v2 since 0.14.0 so compression is on by default — [Source](https://restic.readthedocs.io/en/stable/100_references.html); release/upgrade behavior — [Source](https://forum.restic.net/t/restic-0-14-0-released/5359)
|
||||
- Restic dedup uses plaintext hash, so future zstd-output changes do not break dedup (compressed bytes never hashed for identity) — [Source](https://github.com/restic/restic/pull/3666)
|
||||
- Borg compression set: `none` (0x00), `lz4` (0x01), `zstd` (0x03, levels -128..22), `zlib` (0x05, levels 0–9), `lzma` (0x02, levels 0–9); `ctype` byte + `clevel` byte (zstd `int8_t` so -1→255, -128→128; others unsigned with 255 = n/a); speed order none>lz4>zlib>lzma, lz4>zstd; compression lzma>zlib>lz4>none, zstd>lz4 — [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html); legacy 1.x zlib has no ID bytes (detected by `0x.8` header) — [Source](https://borgbackup.readthedocs.io/en/stable/internals/data-structures.html)
|
||||
- Borg per-object recording (borg2): msgpacked metadata dict holds `ctype`, `clevel`, `csize` (compressed+obfuscated size), `psize` (payload w/o obfuscation trailer, when obfuscated), `olevel`, `size` (uncompressed), `type` (ro_type `A`/`C`/`S`/`F`); metadata and data slots separately encrypted with header bound as AAD — [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html); code `RepoObj.format/parse` — [Source](https://github.com/borgbackup/borg/blob/86fd77fd/src/borg/repoobj.py); compressor base auto-detection via ID header — [Source](https://github.com/borgbackup/borg/blob/86fd77fd/src/borg/compress.pyx)
|
||||
- Borg negotiation/mixing: default compression is `lz4`; mixing methods in one repo is fine because dedup is on source chunks, not compressed bytes; first writer of a chunk determines its stored compression; `borg recreate`/`repo-compress` can recompress; wrappers `auto,C[,L]` (lz4 compressibility heuristic → `none` vs `C`) and `obfuscate,SPEC,C[,L]` (Padmé deterministic padding, ≤12% overhead, `MAX_DATA_SIZE` ~20 MiB cap) — [Source](https://manpages.debian.org/trixie/borgbackup2/borg2-compression.1.en.html)
|
||||
- Kopia compression is disabled by default and controlled per-policy (global/host/path: `--compression=...`, min/max file size, extensions); new setting applies going forward only, does not retroactively recompress — [Source](https://kopia.io/docs/faqs/)
|
||||
- Kopia algorithm menu: `none|deflate-best-compression|deflate-best-speed|deflate-default|gzip|gzip-best-compression|gzip-best-speed|pgzip|pgzip-best-compression|pgzip-best-speed|s2-better|s2-default|s2-parallel-4|s2-parallel-8|zstd|zstd-better-compression|zstd-fastest` (`zstd` recommended default choice); benchmark table shows s2 fastest (~GB/s, largest), zstd smallest, pgzip balanced — [Source](https://kopia.io/docs/faqs/); details and memory/I-O guidance — [Source](https://kopia.io/docs/advanced/compression/)
|
||||
- Kopia per-object recording: since content-level compression (v0.9, `--index-version=2`), compression done after hashing; per-content compression ID kept in content manager bookkeeping, visible via `kopia content list -c` / `content stats` (e.g. `(uncompressed)` vs `zstd` vs `zstd-fastest` with counts/sizes), no longer encoded as `Z`-prefixed content ID — [Source](https://github.com/kopia/kopia/pull/1076); if compressed chunk grows, original stored uncompressed — [Source](https://kopia.io/docs/advanced/compression/)
|
||||
- Kopia mixing/compat: dedup unaffected by algorithm changes or library-output drift because ID is pre-compression hash; enables future recompression during maintenance; changing policy does not rewrite old contents; newer Kopia cannot read legacy LZ4-compressed contents — must migrate with an older version (restore/repack) before upgrading — [Source](https://github.com/kopia/kopia/pull/1076); LZ4 removal notice — [Source](https://kopia.io/docs/advanced/compression/)
|
||||
|
||||
### Inferences
|
||||
- Recording compression outside the content ID (restic type bits/index field, Borg metadata dict, Kopia index v2) is what allows free mixing and future recompression without breaking content addressing.
|
||||
- Restic's single-codec choice simplifies negotiation to levels only; Borg/Kopia pay codec-matrix complexity for finer speed/ratio control.
|
||||
|
||||
### Gaps
|
||||
- Exact zstd numeric levels behind restic `auto/better/max` (klauspost `SpeedDefault` etc. numeric equivalents) were not pinned to zstd CLI levels in fetched sources.
|
||||
- Kopia on-disk bytes for per-content compression IDs (enum values/header layout) were not found; only CLI-visible behavior is documented.
|
||||
|
||||
## How is integrity verified (HMAC, AEAD tag, checksums, parity)? How are manifests/snapshots authenticated?
|
||||
|
||||
### Takeaway
|
||||
Restic uses Encrypt-then-MAC (AES-256-CTR + Poly1305-AES) on every blob/file plus SHA-256 content addressing; Borg 2 uses AEAD (AES-OCB or ChaCha20-Poly1305) with chunk-ID-as-AAD plus layered checksums; Kopia uses AEAD (AES-GCM or ChaCha20-Poly1305) with HMAC-derived per-content keys plus optional Reed-Solomon ECC — none provide parity by default.
|
||||
|
||||
### Cited Findings
|
||||
- Restic envelope: all files except `keys/` (and pack containers) are `IV(16)||CIPHERTEXT||MAC(16)` (32 B overhead, random IV per file); pack files hold multiple independently encrypted/authenticated blobs + encrypted header + LE `Header_Length`; primitives AES-256-CTR + Poly1305-AES, Encrypt-then-MAC (MAC over ciphertext), keys via scrypt KDF → 32 B enc key + 32 B MAC key (`k`+`r`) unlocking master keys in `keys/` JSON — [Source](https://restic.readthedocs.io/en/stable/100_references.html); audit summary — [Source](https://words.filippo.io/restic-cryptography/)
|
||||
- Restic content integrity: filenames are hex SHA-256 of file ciphertext (verifiable with `sha256sum`); trees/data addressed by SHA-256 of plaintext with deterministic JSON for trees; `restic check` verifies structure plus optional `--read-data` payload reads; tampered data fails MAC and is not decrypted — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic snapshots/manifests: snapshots are JSON (`time/tree/paths/hostname/...`) stored as unpacked encrypted files (v2: version-byte + zstd then `IV||C||MAC`), filename = storage ID; snapshot→tree→blob DAG; no separate manifest signature — authentication is the per-file MAC + content-hash reference; write-order rules (packs → index → snapshot; read reverse) keep repo consistent — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic has no parity/ECC; repair is `rebuild-index` + re-backup + `check`; threat model explicitly excludes deletion protection — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Borg 2 AEAD modes: `--encryption aes256-ocb|chacha20-poly1305` × `--id-hash sha256|blake3`, orthogonal to `--key-location repokey|keyfile`; per-session random `sessionid`, `session_key` via SHA-256 KDF, counter IVs; each object has two separately encrypted slots (metadata + data) behind unencrypted header (`OBJ_MAGIC||version||chunk_id||meta_size||data_size`); `AAD = header||slot_tag||id||...`, tag authenticates metadata, payload, header prefix and chunk ID — [Source](https://borgbackup.readthedocs.io/en/latest/internals/security.html)
|
||||
- Borg chunk-ID binding: `id = MAC(id_key, plaintext)`; AEAD tag binds ID to ciphertext, so repo cannot swap content under an ID; post-decrypt `id == MAC(decompressed)` check is optional by default (only detects malicious client writes) but enforced by `borg check --verify-data` / `BORG_ASSERT_ID` — [Source](https://borgbackup.readthedocs.io/en/latest/internals/security.html)
|
||||
- Borg manifest/archive authentication (Horton principle DAG): every object referenced by parent ID up to manifest; borg2 stores `ro_type` in metadata and verifies expected vs. actual type, binding meaning + ID via AAD; no TAM in borg2; borg 1.x used TAM (`HKDF-SHA-512(id_key||enc_key||enc_hmac_key, 64 B salt, "borg-metadata-authentication-manifest")` → `HMAC` over packed manifest) to anchor fixed-ID manifest (CVE-2016-10099) — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg non-encrypting modes: `authenticated-sha256|blake3` carry keyed MAC (not encryption) over payload+header+slot; `none-*` carry only unkeyed checksums (accidental-corruption detection, no tamper protection); legacy 1.x modes were AES-CTR Encrypt-then-MAC with IV reservation; key blobs wrapped by argon2-derived KEK + chacha20-poly1305 (IV 0) — [Source](https://borgbackup.readthedocs.io/en/latest/internals/security.html)
|
||||
- Borg layered checksums: `IntegrityCheckedFile` streaming checksums for cache/index/hints with `integrity.<TXN>` msgpack files and `[integrity]` cache-config section; corrupt index/hints are deleted and rebuilt; `config/config` (version/id/encryption/id_hash) is plaintext/unauthenticated, protected instead by client security-dir `EncryptionMethodMismatch` checks — [Source](https://borgbackup.readthedocs.io/en/latest/internals/data-structures.html)
|
||||
- Kopia content encryption: default `AES256-GCM-HMAC-SHA256` (alt `CHACHA20-POLY1305-HMAC-SHA256`), set at `repository create` and immutable after; per-content AEAD keys derived via HMAC-SHA256; blocks packed into 20 MB packs (20–40 MB on wire) with local index at pack end + top-level index mapping block ID → (blob, offset, length) — [Source](https://kopia.io/docs/advanced/architecture/); cipher history (deprecated unauthenticated AES-CTR/SALSA20) — [Source](https://github.com/kopia/kopia/pull/277)
|
||||
- Kopia format/manifest envelope: `kopia.repository` format blob holds `uniqueID` (salt), `keyAlgo scrypt-65536-8-1`, `encryption AES256_GCM`, `encryptedBlockFormat` = JSON (`ContentFormat{version,hash,encryption,HMACSecret 32 B,MasterKey 32 B,MaxPackSize 20 MiB}` + `ObjectFormat{splitter}`) encrypted with passphrase-derived `Km=PBKDF(pass,uniqueID)`, `Ke=HKDF(SHA256,Km,uniqueID,"AES",32)`, `AD=HKDF(...,"CHECKSUM",32)`; snapshots/manifests are ordinary encrypted contents in the same CABS/object layers — [Source](https://kopia.io/docs/advanced/encryption/); defaults (`--block-hash BLAKE2B-256-128`, `--object-splitter DYNAMIC-4M-BUZHASH`, `--encryption AES256-GCM-HMAC-SHA256`) — [Source](https://kopia.io/docs/reference/command-line/common/repository-create-filesystem/)
|
||||
- Kopia extra redundancy: experimental Reed-Solomon `REED-SOLOMON-CRC32` ECC with `--ecc-overhead-percent` (default 0 = disabled); must be enabled at creation, cannot be added later; cloud backends already ECC-protected so often redundant — [Source](https://kopia.io/docs/advanced/ecc/); consistency/verify docs cover validity checks and repair — [Source](https://kopia.io/docs/advanced/architecture/)
|
||||
|
||||
### Inferences
|
||||
- All three authenticate snapshots/manifests the same way as data (no detached signatures); trust anchors are the client-held keys (restic master keys, Borg key+TAM/AAD chain, Kopia format-block passphrase), giving a key-anchored DAG from manifest to chunks.
|
||||
- Only Kopia offers built-in parity (ECC); restic/Borg rely on backend durability + authenticated detection and rebuild/re-upload repair.
|
||||
|
||||
### Gaps
|
||||
- Kopia manifest-object specifics (manifest content type, index-epoch authentication, `kopia snapshot verify` coverage) were not fetched; snapshot authentication is inferred from “all contents encrypted/authenticated” architecture.
|
||||
- Exact Borg `MAX_DATA_SIZE`/obfuscation interaction with AEAD tag and `BORG_ASSERT_ID` default scope need code-level confirmation beyond docs snippets.
|
||||
|
||||
## Order of operations: compress-then-encrypt vs alternatives, and why
|
||||
|
||||
### Takeaway
|
||||
All three do hash → compress → encrypt (dedup first, compress second, encrypt last); encrypting last is mandatory because ciphertext is incompressible and randomized, while hashing/compressing first preserves dedup and ratio.
|
||||
|
||||
### Cited Findings
|
||||
- Restic code order in `saveAndEncrypt`: plaintext hash for ID/dedup (`saveBlob` → `Hash(buf)`) → zstd compress (if v2 and enabled) → random nonce → `Seal` (AES-CTR + Poly1305 MAC) → pack; unpacked files: `compressUnpacked` (prepend version byte + zstd) → `Seal` — [Source](https://github.com/restic/restic/blob/master/internal/repository/repository.go); PR states “Unpacked files like lock, index and snapshot files are also compressed before encryption” — [Source](https://github.com/restic/restic/pull/3666)
|
||||
- Restic why: backing up pre-compressed (e.g. `.gz`) data defeats CDC dedup because small input changes avalanche through the compressor and shift cut points; project advises backing up uncompressed data and letting restic chunk-then-compress per blob; chunk-then-compress discussion — [Source](https://github.com/restic/restic/issues/790); zstd-dictionary debate concludes no dictionary needed precisely because compression runs after chunking/dedup — [Source](https://github.com/restic/restic/issues/3775)
|
||||
- Borg order: `id = MAC(id_key, data)` → `compressed = compress(data)` → `AEAD_encrypt(session_key, iv, compressed, aad=id||header)` (borg2); legacy: `id=AUTH(data)` → `compress` → `AES-CTR` → `MAC(encrypted)` — [Source](https://borgbackup.readthedocs.io/en/latest/internals/security.html); docs state “Compression is applied after deduplication, thus using different compression methods in one repo does not influence deduplication” — [Source](https://manpages.debian.org/trixie/borgbackup2/borg2-compression.1.en.html)
|
||||
- Kopia order: “splits into chunks → hash → compare → if new, compress chunk → encrypt → pack” — [Source](https://kopia.io/docs/advanced/compression/); content-level compression PR explicitly rewired pipeline from `read → split → compress → hash → store` to `read → split → hash → compress → store` to stop compressor-version drift from changing IDs and breaking dedup and to enable server-side/recompression — [Source](https://github.com/kopia/kopia/pull/1076)
|
||||
- General rationale shared across docs: encryption must be last because (a) compressed ciphertext does not compress, (b) randomized encryption (IV/session key) destroys dedup if hashed after, (c) MAC/AEAD must cover the stored (compressed) bytes; compressing whole files before chunking is measurably worse than chunk-then-compress (Kopia split-vs-whole table: s2 28.6% vs 25.6%, gzip 18.8% vs 17.3% — small loss accepted for dedup wins) — [Source](https://kopia.io/docs/advanced/compression/)
|
||||
|
||||
### Inferences
|
||||
- The universal pattern is dedup-on-plaintext → compress → authenticated-encrypt; any deviation (pre-compressed inputs, compress-after-encrypt, hash-after-compress) is documented as an anti-pattern in all three projects’ issues/docs.
|
||||
- Per-chunk compression inherently sacrifices cross-chunk dictionary context; all three accept slightly worse ratios in exchange for shift-resilient dedup and independent integrity domains.
|
||||
|
||||
### Gaps
|
||||
- No source quantified restic’s incompressible-blob shortcut (whether zstd `EncodeAll` output larger than input is stored raw vs. kept); forum claims “stores raw data” but code path stores `uncompressedLength` flag — exact skip-threshold logic needs code confirmation.
|
||||
@@ -0,0 +1,147 @@
|
||||
# Backup encryption security choices
|
||||
|
||||
## Which cipher and mode does each use? How are nonces/IVs generated and stored?
|
||||
|
||||
### Takeaway
|
||||
Restic uses non-standard AES-256-CTR + Poly1305-AES (Encrypt-then-MAC) with random 16-byte IV per file/blob; Borg 1.x uses AES-256-CTR + HMAC-SHA256 or keyed BLAKE2b-256 with 64-bit counters tracked by reservation, while Borg 2 moves to AEAD AES256-OCB and ChaCha20-Poly1305; Kopia uses AES256-GCM-HMAC-SHA256 by default (or CHACHA20-POLY1305-HMAC-SHA256) with per-content keys and random IVs; Duplicity delegates entirely to external GnuPG/OpenPGP (symmetric passphrase or public-key session-key hybrid).
|
||||
|
||||
### Cited Findings
|
||||
- Restic: “All data stored by restic in the repository is encrypted with AES-256 in counter mode and authenticated using Poly1305-AES” — [Source](https://github.com/restic/restic/blob/master/doc/design.rst)
|
||||
- Restic: “For encrypting new data first 16 bytes are read from a cryptographically secure pseudo-random number generator as a random nonce. This is used both as the IV for counter mode and the nonce for Poly1305” — [Source](https://github.com/restic/restic/blob/master/doc/design.rst)
|
||||
- Restic: format is “IV || CIPHERTEXT || MAC”, “complete encryption overhead is 32 bytes. For each file, a new random IV is selected”, 16-byte IV stored first, 16-byte MAC last — [Source](https://restic.readthedocs.io/en/v0.4.0/Design)
|
||||
- Restic: needs three keys: “a 32-byte key for AES-256 encryption, a 16-byte AES key and a 16-byte key for Poly1305”, the last 32 bytes split into 16-byte AES key `k` + 16-byte `r` then masked for Poly1305 per Bernstein paper — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic: pack files contain multiple independently encrypted/authenticated blobs plus encrypted header + 4-byte little-endian header length; blob types 0b00 data, 0b01 tree, 0b10/0b11 compressed data/tree (repo format v2, zstandard) — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic: passwords wrapped via scrypt KDF (`N`, `r`, `p`, `salt`); example `N=65536, r=8, p=1`, also observed `N=32768, r=8, p=5` in 2026 docs; derived 64 bytes split into 32-byte AES key + 32-byte MAC key; multiple key files per repo allow password change without re-encrypting data — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic current as of 2026 is v0.19.1 (5 Jul 2026) / v0.19.0, repo format version 1 or 2, v2 adds compression — [Source](https://restic.net/)
|
||||
- Borg 1.x: “repokey and keyfile use AES-CTR-256 for encryption and HMAC-SHA256 for authentication in an encrypt-then-MAC (EtM) construction” — [Source](https://manpages.debian.org/bullseye/borgbackup/borg-init.1.en.html)
|
||||
- Borg 1.x blake2 variants: “repokey-blake2 and keyfile-blake2 are also authenticated encryption modes, but use BLAKE2b-256 instead of HMAC-SHA256”, chunk ID is keyed BLAKE2b-256 — [Source](https://manpages.debian.org/bullseye/borgbackup/borg-init.1.en.html)
|
||||
- Borg chunk header: “TYPE(1) + HMAC(32) + NONCE(8) + CIPHERTEXT. Encryption and HMAC use two different keys” — [Source](https://borgbackup.readthedocs.io/en/1.0-maint/internals.html)
|
||||
- Borg CTR IV: “A 64bit initialization vector is used”, “only 8 bytes of the 16 bytes nonce is saved in the payload, the first 8 bytes are always zeros”, limits capacity to 2**64 * 16 bytes (~295 exabytes) — [Source](https://borgbackup.readthedocs.io/en/1.0-maint/internals.html)
|
||||
- Borg IV uniqueness via reservation: “initializes the encryption counter to be higher than any previously used counter value”, commits reservation by “taking the current counter value and adding 4 GiB / 16 bytes to the counter”, persisted via SaveFile to security DB + repository before encrypting — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg pseudocode: `iv = reserve_iv()`, `encrypted = AES-256-CTR(enc_key, 8-null-bytes || iv, compressed)`, `authenticated = type-byte || AUTHENTICATOR(enc_hmac_key, encrypted) || iv || encrypted` — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg offline key wrapping: “256 bit key encryption key (KEK) is derived from the passphrase using PBKDF2-HMAC-SHA256 with a random 256 bit salt”, then Encrypt-and-MAC with “AES-256-CTR with a constant initialization vector of 0”, base64 keyblob in keyfile or repo config — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg stable in 2026 is 1.4.5; crypto section still states “actual encryption is currently always AES-256 in CTR mode” for 1.x line — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg 2 (beta 2.0.0b22/b23 as of 2026): modes `aes256-ocb`, `chacha20-poly1305`, `authenticated-sha256/blake3`, `none-sha256/blake3`; “AES256 in OCB mode (encryption + authentication)” and “ChaCha20 + Poly1305” — [Source](https://borgbackup.readthedocs.io/en/latest/usage/repo-create.html)
|
||||
- Borg 2: “All data can be protected client-side using 256-bit authenticated encryption (AES-OCB or chacha20-poly1305)” — [Source](https://github.com/borgbackup/borg)
|
||||
- Borg 2 goals: “get rid of AES-CTR mode and use ‘session keys’”, “use more modern / faster AEAD ciphers: AES-OCB and chacha20-poly1305”, “use a more modern KDF: argon2” — [Source](https://github.com/borgbackup/borg/wiki/Borg-2.0)
|
||||
- Kopia default: `const DefaultAlgorithm = "AES256-GCM-HMAC-SHA256"` — [Source](https://pkg.go.dev/github.com/kopia/kopia/repo/encryption)
|
||||
- Kopia options: “By default, Kopia uses the AES256-GCM-HMAC-SHA256 encryption algorithm … but you can choose CHACHA20-POLY1305-HMAC-SHA256”, immutable after repo creation, selected via `--encryption=` or Advanced Options — [Source](https://github.com/kopia/kopia/blob/master/site/content/docs/FAQs/_index.md)
|
||||
- Kopia per-content keys: registered as “AES-256-GCM using per-content key generated using HMAC-SHA256”, derives key via `deriveKey(p, purposeEncryptionKey)` + `hmac.New(sha256.New, keyDerivationSecret)` pool, overhead 28 bytes — [Source](https://github.com/kopia/kopia/blob/master/repo/encryption/aes256_gcm_hmac_sha256_encryptor.go)
|
||||
- Kopia format blob: `encryption` field (default `AES256_GCM`), `encryptedBlockFormat` = JSON `EncryptedRepositoryConfig` encrypted with random IV prepended, key `Ke = HKDF(SHA256, Km, UniqueID, "AES", 32)` where `Km = PBKDF(passphrase, UniqueID)` via scrypt `N=65536, r=8, p=1`, AD = `HKDF(SHA256, Km, UniqueID, "CHECKSUM", 32)` — [Source](https://kopia.io/docs/advanced/encryption/)
|
||||
- Kopia `ContentFormat` holds `Hash`, `Encryption`, `HMACSecret`, `MasterKey (SIV-mode only)`, `MaxPackSize`; repository config stored encrypted because it contains key material — [Source](https://kopia.io/docs/advanced/encryption/)
|
||||
- Kopia observed version v0.23.1 in Go docs (2026 architecture doc updated Feb 2026) — [Source](https://pkg.go.dev/github.com/kopia/kopia/repo/encryption)
|
||||
- Duplicity: “incrementally backs up files and folders into tar-format volumes encrypted with GnuPG” — [Source](https://man.archlinux.org/man/duplicity.1.en)
|
||||
- Duplicity key modes: `--encrypt-key key` (public-key, repeatable), `--hidden-encrypt-key` uses `gpg --hidden-recipient` to hide recipient key ID, `--sign-key`, symmetric via `PASSPHRASE` env, `--gpg-options`, `--gpg-binary` — [Source](https://man.archlinux.org/man/duplicity.1.en)
|
||||
- Duplicity relies on OpenPGP hybrid: random session key `s`, `Enc_ri(s)` per recipient + `Enc_s(data)` — [Source](https://gnupg.org/ftp/blurbs/an-advanced-introduction-to-gnupg.pdf)
|
||||
- Duplicity current lineage includes 2.x/3.x (Debian bookworm-backports, Arch manuals); backend still GnuPG + librsync incremental — [Source](https://manpages.debian.org/bookworm-backports/duplicity/duplicity.1.en.html)
|
||||
|
||||
### Inferences
|
||||
- Restic’s construction is bespoke AES-CTR + Poly1305-AES (not AES-GCM or ChaCha20-Poly1305); random IV per encryption avoids counter-reuse tracking at cost of 16-byte IV storage and reliance on CSPRNG.
|
||||
- Borg 1.x’s 8-byte stored counter is the scaling bottleneck that motivated Borg 2 session keys + AEAD nonces; Borg 2 OCB vs ChaCha choice is performance-driven (hardware AES vs portable).
|
||||
- Kopia’s “-HMAC-SHA256” suffix denotes per-content key derivation + outer HMAC layer, not just plain GCM; format-blob AD binding to passphrase-derived Km prevents config swapping without password.
|
||||
- Duplicity inherits whatever GnuPG/OpenPGP negotiates (modern GnuPG 2.x defaults to AES-256 + MDC/AEAD), so its cipher agility is a GnuPG property, not a Duplicity design choice.
|
||||
|
||||
### Gaps
|
||||
- Exact GnuPG symmetric cipher/MDC vs AEAD (OCB/EAX) negotiated by Duplicity 3.x in 2026 not confirmed from Duplicity docs; depends on installed GnuPG version and `--gpg-options`.
|
||||
- Kopia nonce length/placement for content blobs (vs format blob) not fully detailed in fetched docs; source states overhead 28 bytes but byte layout not quoted.
|
||||
- Borg 2 session-key derivation and nonce construction details not yet in stable internals doc (beta docs only list mode names).
|
||||
|
||||
## What is encrypted vs plaintext (content, filenames, metadata, manifests/snapshots)?
|
||||
|
||||
### Takeaway
|
||||
Restic, Borg (in encrypted modes) and Kopia encrypt file contents, filenames/paths, metadata and snapshots/manifests client-side; only key-file wrappers, config envelopes, outer pack/segment sizes, timestamps and directory layout remain visible. Duplicity encrypts tar volumes including embedded filenames + manifest/sigtar, but volume filenames, sizes and increment metadata stay plaintext. None hide sizes, counts or access patterns.
|
||||
|
||||
### Cited Findings
|
||||
- Restic guarantee: “Unencrypted content of stored files and metadata cannot be accessed without a password … Everything except the metadata included for informational purposes in the key files is encrypted and authenticated. The cache is also encrypted to prevent metadata leaks” — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic snapshots are “JSON document … stored in a file below snapshots”, encrypted as `IV || Ciphertext || MAC` (v1 JSON, v2 `encoding_version || zstd(JSON)`); example snapshot contains `paths, hostname, username, uid/gid, tags, tree` only after decryption — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic trees contain `name, type, mode, mtime/atime/ctime, uid/gid/user, inode, size, content[plaintext hashes], subtree, linktarget` encrypted inside tree blobs — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic index JSON (`packs[].blobs[].id/type/offset/length`) and pack headers (plaintext hashes, offsets) are encrypted; only after decryption are blob IDs visible — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic config (`version, id, chunker_polynomial`) is encrypted; `keys/` files expose only informational `hostname, username, created, kdf params, salt` + encrypted `data` wrapping master keys — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Borg attack model: client trusted, repository/server untrusted with full read/write/MITM; guarantees attacker cannot modify data, rename/remove/add archive, recover plaintext, or recover definite structural info (object graph) undetected — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg: “authenticated encryption technique makes it suitable for backups to targets not fully trusted” — [Source](https://manpages.debian.org/bookworm/borgbackup2/borg2.1.en.html)
|
||||
- Borg authenticated-only modes (`authenticated-sha256/blake3`) provide tamper detection without confidentiality; `none` modes provide neither — [Source](https://borgbackup.readthedocs.io/en/latest/usage/repo-create.html)
|
||||
- Borg compression happens before encryption: `compressed = compress(data)` then encrypt-then-MAC; optional `obfuscate` pseudo-compressor pads with 0x00 before encryption to hide compressed sizes — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Kopia: “Encryption is at the repository level, and Kopia encrypts all snapshots in all repositories by default” via repository password — [Source](https://github.com/kopia/kopia/blob/master/site/content/docs/FAQs/_index.md)
|
||||
- Kopia layers: BLOB store holds packs; CABS encrypts blocks after hashing (`AES256-GCM-HMAC-SHA256` or `CHACHA20-POLY1305-HMAC-SHA256`); CAOS directory listings (`k` prefix), manifests (`m`), indirect JSON (`x`) are CABS blocks; LAMS manifests (snapshots, policies) stored as CABS blocks — [Source](https://kopia.io/docs/advanced/architecture/)
|
||||
- Kopia: “Pack files in blob storage have random names and don’t reveal anything about their contents or structure. Their sizes are also generally unrelated to their content due to splitting and merging” — [Source](https://kopia.io/docs/advanced/architecture/)
|
||||
- Kopia blob prefixes `p` (data packs), `q` (metadata packs), `x` (indices), object prefixes `k/m/x`, `I` virtual indirection — visible types but not contents — [Source](https://kopia.io/docs/advanced/architecture/)
|
||||
- Duplicity volumes are gzipped tar + GnuPG; OpenPGP literal-data packet (filename, mode, mtime) is inside encryption, so embedded names hidden, but outer `duplicity-full|inc` volume filenames, `manifest.gpg`/`sigtar.gpg` names, sizes and S3 storage-class split (manifest/sigtar on Standard for quick retrieval, data on Glacier/Deep Archive) are server-visible — [Source](https://man.archlinux.org/man/duplicity.1.en)
|
||||
- Duplicity over S3 with `--s3-use-glacier` notes: “Duplicity will store the manifest.gpg and sigtar.gpg files … on AWS S3 standard storage … all other data is stored in S3 Glacier” — [Source](https://man.archlinux.org/man/duplicity.1.en)
|
||||
- Duplicity `--no-encryption` writes only gzipped volumes (no confidentiality) — [Source](https://linux.die.net/man/1/duplicity)
|
||||
|
||||
### Inferences
|
||||
- For Restic/Borg/Kopia, filenames, symlink targets, ownership, timestamps and snapshot manifests are as confidential as file bytes; server compromise reveals at most sizes/counts/timing, not names.
|
||||
- Kopia’s p/q/x prefix separation intentionally leaks data-vs-metadata partitioning to aid storage management while hiding actual object graphs.
|
||||
- Duplicity’s S3 Glacier optimization trades metadata confidentiality (manifests pinned to Standard, easily listable) for cheaper restores.
|
||||
|
||||
### Gaps
|
||||
- Whether Borg archive names themselves (manifest entries) are fully encrypted vs length-leaking was not explicitly quoted; inferred encrypted from manifest DAG but needs source confirmation.
|
||||
- Kopia policy/manifest label (key=value) visibility to server (for server-managed repos) not clarified in fetched docs.
|
||||
- Duplicity filename encryption inside OpenPGP (literal packet) assumed from GnuPG behavior, not Duplicity doc; no Duplicity doc explicitly states embedded names are hidden.
|
||||
|
||||
## How are object names derived (plaintext hash, HMAC, random IDs) and what does that leak to an untrusted server?
|
||||
|
||||
### Takeaway
|
||||
Restic outer filenames are SHA-256 of ciphertext (verifiable via sha256sum) while inner blob references are SHA-256 of plaintext hidden in encrypted index/headers; Borg chunk IDs are HMAC/keyed hashes of plaintext (dedup without revealing content); Kopia content IDs are hashes/HMACs of plaintext with secret, packed into randomly named blobs; Duplicity names are sequential timestamps leaking backup chain. All leak dedup equality, counts, sizes and access patterns to varying degrees.
|
||||
|
||||
### Cited Findings
|
||||
- Restic: “storage ID is the SHA-256 hash of the content of a file”, filename is “lower case hexadecimal representation of the storage ID”, verifiable by running `sha256sum` on file and comparing to filename — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic: “All content … is referenced according to its SHA-256 hash”, file split into CDC blobs, “SHA-256 hashes of all Blobs are saved in ordered list”, tree `content` holds plaintext hashes, `subtree` holds plaintext tree ID — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic: index/pack header “yields all plaintext hashes, types, offsets and lengths”, pack `id` in index is outer pack hash; `restic cat pack` verifies hash and warns on mismatch — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic data/keys layout: `data/XX/<full-hash>`, `snapshots/<id>`, `index/<id>`, `keys/<id>`, `locks/`, `tmp/`; snapshot filename is storage ID — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic attacker with read access can “Infer which packs probably contain trees via file access patterns”, “Infer the size of backups by using creation timestamps”, and pre-0.18.0 could derive chunker polynomial from observed chunk sizes per 2025 IACR paper; 0.18.0 mitigates by “randomly assigning chunks to pack files” — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Borg: “key = id = id_hash(unencrypted_data)”, `id_hash` is `sha256` (no keys) or `hmac-sha256` (with keys); must be cryptographically strong for dedup — [Source](https://borgbackup.readthedocs.io/en/1.0-maint/internals.html)
|
||||
- Borg 1.x: chunk ID `id = AUTHENTICATOR(id_key, data)` with independent `id_key` vs `enc_hmac_key`; decryption asserts `CONSTANT-TIME-COMPARISON(chunk-id, AUTHENTICATOR(id_key, decompressed))` — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg 2: “A chunk is considered duplicate if its id_hash value is identical”, `id_hash` e.g. “(hmac-)sha256 or (keyed) blake3”, selectable via `--id-hash=blake3` — [Source](https://github.com/borgbackup/borg)
|
||||
- Borg manifest has fixed ID `000…000`, anchored via TAM: `tam_key = HKDF-SHA-512(ikm=id_key||enc_key||enc_hmac_key, salt, "borg-metadata-authentication-manifest")`, `HMAC(tam_key, packed)` — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg fingerprinting: “repository does not hide size of chunks”, buzhash chunking uses “secret, random per-repo chunker seed”, small files <512 KiB yield single chunk; attacker with candidate files could brute-force fingerprint by sizes; mitigations: chunker choice/params, secret seed, compression choice, `obfuscate` padding; proximity/order (inode-order scan, segment adjacency) may leak additional info — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Kopia: “Block ID is generated by applying cryptographic hash function such as SHA2 or BLAKE2S”, identical blocks yield identical IDs for natural dedup; after hashing, block encrypted; multiple blocks merged into 20-40MB packs; index maps block ID -> (blob name, offset, length) — [Source](https://kopia.io/docs/advanced/architecture/)
|
||||
- Kopia: pack blob names random (e.g. `pb4cf8…`, `q7a99…`, `xn0_20db…`); content list/show via `kopia content list/show` requires decryption — [Source](https://kopia.io/docs/advanced/architecture/)
|
||||
- Kopia format `UniqueID` is 32 random bytes, also PBKDF salt and HKDF info, preventing cross-repo ID correlation — [Source](https://kopia.io/docs/advanced/encryption/)
|
||||
- Duplicity sets: “files in full backup sets will start with duplicity-full while the incremental sets start with duplicity-inc”, ordered patches; deleting full invalidates dependent incrementals — [Source](http://duplicity.nongnu.org/vers7/duplicity.1.html)
|
||||
- Duplicity S3 observer can tell “that you are using Duplicity, the name of the bucket, your AWS Access Key ID, the increment dates and the amount of data in each increment” (affects connection, not GPG payload) — [Source](https://linux.die.net/man/1/duplicity)
|
||||
- Duplicity default recipient key IDs visible in OpenPGP packets unless `--hidden-encrypt-key` (`--hidden-recipient`) used; restore then tries all secret keys — [Source](https://man.archlinux.org/man/duplicity.1.en)
|
||||
|
||||
### Inferences
|
||||
- Restic outer names (ciphertext hashes) look random to server; inner plaintext hashes never appear plaintext on server, so equality of files is hidden unless attacker correlates sizes/access patterns — unlike Borg/Kopia where HMAC IDs are server-visible keys enabling dedup-equality oracle.
|
||||
- Borg/Kopia HMAC-ID design intentionally reveals equality (required for server-side dedup) but not content; without `id_key`/`HMACSecret` server cannot confirm guesses except via size/proximity fingerprinting.
|
||||
- Duplicity leaks the most metadata: full vs incremental, chain order, timestamps and volume counts directly encode backup history even when payloads are opaque.
|
||||
|
||||
### Gaps
|
||||
- No reliable source confirming whether Restic pack filename is hash of ciphertext vs plaintext; doc says hash of content but pack verification suggests ciphertext — ambiguity noted.
|
||||
- Kopia whether content IDs are plain hash vs HMAC-SHA256 with `HMACSecret` not fully resolved: architecture says hash, `ContentFormat.HMACSecret` implies keyed; exact construction needs source read.
|
||||
- Borg 2 object store layout (keys/ namespace, segment naming) not fetched; assumed similar to 1.x but with AEAD.
|
||||
|
||||
## Any documented rationale for their choices (audit reports, design docs)?
|
||||
|
||||
### Takeaway
|
||||
Restic cites simplicity + Encrypt-then-MAC robustness and documents threat model + 2025 chunking-attack mitigation; Borg cites Encrypt-then-MAC robustness, Horton principle + TAM (CVE-2016-10099), and Borg 2 session-key/AEAD/Argon2 motivations; Kopia cites envelope encryption + per-content keys + random packing; Duplicity cites delegation to GnuPG/librsync. No Cure53-style audit report was found for these four in the searches; strongest external reviews found are Filippo Valsorda’s Restic note and Borg issue discussions.
|
||||
|
||||
### Cited Findings
|
||||
- Restic design: “Encryption is a first-class feature, the implementation looks sane and I guess the deduplication trade-off is worth it” — Filippo Valsorda, quoted in Restic encryption docs — [Source](https://github.com/restic/restic/blob/master/doc/070_encryption.rst)
|
||||
- Restic threat model: trusted client + authentic restic + secret password; guarantees vs assumptions listed; “Advances … against … (AES-256-CTR-Poly1305-AES and SHA-256) have not occurred”, brute-force infeasible, leaked key requires full re-encryption via `copy`/new repo — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic 2025 chunking attack: paper “Chunking Attacks on File Backup Services using Content-Defined Chunking” by Alexeev/Percival/Zhang, mitigated in 0.18.0 by random pack assignment; random irreducible CDC polynomial stored in encrypted `config` to harden watermark attacks — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Restic CDC: Rabin fingerprints, 64-byte window, files <512 KiB unsplit, 512 KiB–8 MiB blobs targeting 1 MiB average — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- Borg: “Encryption is currently based on the Encrypt-then-MAC construction, which is generally seen as the most robust way” — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg: follows “Horton principle … not only the message must be authenticated, but also its meaning”, object ID MACs plaintext, parent reference assigns meaning, DAG anchored by TAM — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg TAM added since 1.0.9 for “Pre-1.0.9 manifest spoofing vulnerability (CVE-2016-10099)” — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg: “Borg does not support unauthenticated encryption — only authenticated encryption … No unauthenticated schemes will be added” — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg multi-client caveat: with multiple independent clients on same repo, “Borg fails to provide confidentiality” due to counter-reservation replay; trusted sync channel needed — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg compression+encryption discussion in issue #1040 concluded “no problem at all” to “hard and extremely slow to exploit”; user can disable compression — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg primitives rationale: HMAC-SHA256 vs BLAKE2b chosen on SHA hardware support (Ryzen, Intel 10th+ mobile/11th+ desktop, M1+, ARM64 SHA ext favour HMAC-SHA256; 64-bit CPUs without SHA ext favour BLAKE2b); uses OpenSSL libcrypto only, not libssl/TLS/X.509 — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg 2 rationale: global AES+MAC key requires perfect counter tracking across clients/threads (XOR-plaintext leak risk + sync complexity); session keys ease multithreading — [Source](https://github.com/borgbackup/borg/wiki/Borg-2.0)
|
||||
- Borg keyed chunkers (`toeplitz-aes/rabin-aes/goldilocks-aes`) rationale: “make chunk-boundary fingerprinting attacks much harder” — [Source](https://github.com/borgbackup/borg)
|
||||
- Kopia rationale: “standard envelope encryption technique to de-couple the repository passphrase from keys used for encrypting/authenticating contents”; format blob holds params, `UniqueID` random per repo, scrypt PBKDF + HKDF-SHA256 for Ke/AD — [Source](https://kopia.io/docs/advanced/encryption/)
|
||||
- Kopia FAQ: encryption mandatory, algorithm fixed at creation, password unrecoverable — [Source](https://github.com/kopia/kopia/blob/master/site/content/docs/FAQs/_index.md)
|
||||
- Duplicity rationale: uses librsync for space-efficient incrementals + GnuPG so “they will be safe from spying and/or modification by the server” — [Source](https://packages.debian.org/bullseye/arm64/utils/duplicity)
|
||||
- Duplicity signing+symmetric-encrypt via CLI gpg noted as “specifically challenging”, only certain passphrase/agent combos tested working — [Source](http://duplicity.nongnu.org/vers7/duplicity.1.html)
|
||||
|
||||
### Inferences
|
||||
- Restic/Borg explicitly prefer Encrypt-then-MAC over AEAD for 1.x-era compatibility and library constraints (OpenSSL libcrypto, Go stdlib); Borg 2 and Kopia converge on modern AEAD (OCB/GCM/ChaCha20-Poly1305) once widely available.
|
||||
- Borg’s extensive fingerprinting/TAM documentation is the most explicit untrusted-server rationale among the four; Restic’s 2025 mitigation shows similar concerns emerging later.
|
||||
- Duplicity offloads rationale to GnuPG/OpenPGP; its docs contain no cipher-selection rationale beyond backend/storage-class notes.
|
||||
|
||||
### Gaps
|
||||
- No Cure53 or similar third-party audit report for restic/Borg/Kopia/Duplicity found in searches (results returned unrelated Ente/Project11/Monocypher audits); if audits exist, they were not retrievable with queries used.
|
||||
- Kopia independent security review or design-rationale doc beyond envelope-encryption page not found.
|
||||
- Duplicity successor scope (e.g. duplicacy, duplicity 3.x forks) not clarified; findings cover classic Duplicity only.
|
||||
@@ -0,0 +1,114 @@
|
||||
# Backup encryption security choices: key management
|
||||
|
||||
## How is the master key generated and derived from a password (scrypt, Argon2, PBKDF2 parameters)?
|
||||
|
||||
### Takeaway
|
||||
All five systems use password-derived envelope encryption with a random master/data key, but KDFs and parameters differ: restic pins scrypt N=65536/r=8/p=1 (auto-calibrated on newer adds); Borg 1.x uses PBKDF2-HMAC-SHA256 while Borg 2.x defaults to Argon2id + ChaCha20-Poly1305; Kopia uses scrypt-65536-8-1 over UniqueID salt plus HKDF-SHA256 subkeys (PBKDF2/scrypt configurability added ~2025); rclone crypt uses scrypt N=16384/r=8/p=1; Tarsnap generates keys locally and only uses scrypt to passphrase-wrap the key file.
|
||||
|
||||
### Cited Findings
|
||||
- restic: all repo data encrypted with AES-256-CTR + Poly1305-AES MAC as `IV || CIPHERTEXT || MAC` (32 bytes overhead, random 16-byte nonce per file); three keys needed (32-byte AES-256 key + 16-byte AES key + 16-byte Poly1305 key) — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- restic: on open, password + per-key `N`, `r`, `p`, `salt` feed scrypt to derive 64 bytes; first 32 = AES-256 encryption key, last 32 split into 16-byte AES key `k` + 16-byte Poly1305 key `r` (masked); these decrypt the `data` field to reveal the master key JSON — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- restic: canonical key-file example pins `"kdf": "scrypt", "N": 65536, "r": 8, "p": 1`; code rejects any KDF other than scrypt (`only supported KDF is scrypt()`) — [Source](https://github.com/restic/restic/blob/9e2d60e2/internal/repository/key.go)
|
||||
- restic: `AddKey` calibrates KDF parameters (`crypto.Calibrate(KDFTimeout, KDFMemory)`) when adding keys, generates a random salt (`crypto.NewSalt()`) and either a fresh random master key (`crypto.NewRandomKey()`) or copies the existing master key template — [Source](https://github.com/restic/restic/blob/9e2d60e2/internal/repository/key.go)
|
||||
- Borg 1.x/legacy: key file encrypted with PBKDF2-HMAC-SHA256 (32-byte salt, `PBKDF2_ITERATIONS`), AES-CTR with zero IV plus HMAC-SHA256 encrypt-and-MAC construction — [Source](https://github.com/borgbackup/borg/blob/86fd77fd/src/borg/legacy/crypto/key.py)
|
||||
- Borg 2.x: 256-bit key-encryption key (KEK) derived from passphrase with Argon2 + random 256-bit salt, then Encrypt-then-MAC of packed key material with ChaCha20-Poly1305 AEAD and constant IV==0 (safe because salt makes KEK unique per encryption) — [Source](https://borgbackup.readthedocs.io/en/2.0.0b9/internals/security.html)
|
||||
- Borg: new key files support `argon2 chacha20-poly1305` vs legacy `sha256` (PBKDF2) algorithms; Argon2 path derives 32-byte KEK via `argon2.low_level.hash_secret_raw` with configurable time/memory/parallelism/type — [Source](https://github.com/borgbackup/borg/blob/da3105f1/src/borg/crypto/key.py)
|
||||
- Borg: `--key-algorithm argon2` is default (Argon2id); `pbkdf2` kept for old-client compatibility; docs note Argon2 path also fixes two issues at once (separate encrypt vs MAC keys, encrypt-then-MAC instead of encrypt-and-MAC) — [Source](https://git.uninsane.org/shelvacu-mirrors/borg/commit/08f82ee40867f605ca6994db6dc32218d7b85cbc)
|
||||
- Borg: random Borg key itself consists of three random secrets (crypt key, id key, chunker seed); passphrase only locks/encrypts this key, chunking and IDs also derive from it — [Source](https://borgbackup.readthedocs.io/en/stable/usage/init.html)
|
||||
- Kopia: envelope encryption; format blob carries `uniqueID` (random 32 bytes), `keyAlgo` (e.g. `scrypt-65536-8-1`), `encryption: AES256_GCM`, and `encryptedBlockFormat` holding the real content-encryption secrets (`HMACSecret`, `MasterKey`) — [Source](https://kopia.io/docs/advanced/encryption/)
|
||||
- Kopia: `Km = PBKDF(passphrase, UniqueID, cost params)` (32 bytes) → `Ke = HKDF(SHA256, Km, UniqueID, "AES", 32)` and `AD = HKDF(SHA256, Km, UniqueID, "CHECKSUM", 32)`; format block encrypted with AES256-GCM using Ke, random IV prepended, AD authenticated — [Source](https://kopia.io/docs/advanced/encryption/)
|
||||
- Kopia: only scrypt N=65536/r=8/p=1 documented as supported ("at the moment"), field reserved for future algorithm/cost changes — [Source](https://kopia.io/docs/advanced/encryption/)
|
||||
- Kopia: 2025 PR adds configurable KDF: `--pbkdf (pbkdf2|scrypt)`, `--pbkdf-iter` (default 600000 for PBKDF2), `--pbkdf-memory` (default 64 MB for scrypt), usable at `repo create` and `change-password` — [Source](https://github.com/kopia/kopia/pull/5145)
|
||||
- rclone crypt: file content uses NaCl SecretBox (XSalsa20 + Poly1305), 64 KiB chunks each with 16-byte authenticator, header = 8-byte magic `RCLONE\x00\x00` + 24-byte random nonce incremented per chunk — [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt: filename segments PKCS#7-padded to 16 bytes, encrypted deterministically with EME (ECB-Mix-ECB) AES-256, emitted as lowercase unpadded base32 — [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt: 80 bytes of key material derived via scrypt `N=16384, r=8, p=1` from password (+ optional `password2` salt); without user salt a built-in internal salt is used — [Source](https://rclone.org/crypt/)
|
||||
- Tarsnap: key files hold authentication keys (server access proofs) and encryption keys (archive encrypt/sign/verify/decrypt) separately, so server compromise does not disclose data — [Source](https://www.tarsnap.com/security.html)
|
||||
- Tarsnap: no password-derived master key by default; keys generated by `tarsnap-keygen`; passphrase protection is optional (`--passphrased`) and wraps the key file with keys from the scrypt KDF, with tunable `--passphrase-mem` / `--passphrase-time` — [Source](https://www.tarsnap.com/man-tarsnap-keygen.1.html)
|
||||
|
||||
### Inferences
|
||||
- restic/Kopia/rclone all standardize on scrypt but with 4x cost difference (16384 vs 65536), so rclone password guessing is cheaper; Borg's Argon2id move is the most modern KDF posture.
|
||||
- Kopia's double-HKDF (separate Ke and AD from one Km) is the cleanest domain separation; restic's split-and-mask of 64 scrypt bytes is older Poly1305-AES style.
|
||||
|
||||
### Gaps
|
||||
- restic's current auto-calibration targets (`KDFTimeout`, `KDFMemory`) were not retrieved; documented example still shows 65536/8/1 but newly added keys may use higher N.
|
||||
- Borg's exact default Argon2 time/memory/parallelism constants (`ARGON2_ARGS`, `ARGON2_SALT_BYTES`) were not captured.
|
||||
- Tarsnap's default scrypt mem/time when `--passphrase-mem/--passphrase-time` are omitted was not found.
|
||||
|
||||
## Where is key material stored (repo config, local key file, OS keyring, TPM)? What is the exact UX for export/backup of keys?
|
||||
|
||||
### Takeaway
|
||||
restic and Kopia keep the (password-locked) key material inside the repository; Borg offers repokey (in repo) vs keyfile (client `~/.config/borg/keys`) as a first-class choice; Tarsnap and rclone keep keys strictly client-side (key file / `rclone.conf`), making client-side backup mandatory. None document TPM; OS-credential integration exists only for Kopia (server/connect password caching) and via external helpers for the rest.
|
||||
|
||||
### Cited Findings
|
||||
- restic: `keys/` directory in repo holds JSON key files (hostname, username, created, kdf params, salt, encrypted `data`); `restic cat masterkey` decrypts and pretty-prints master encryption/MAC keys — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- restic: passwords supplied via prompt, `--password-file`, or `--password-command` (`$RESTIC_PASSWORD_FILE` / `$RESTIC_PASSWORD_COMMAND`); empty passwords refused by default, `--insecure-no-password` required since 0.17.0 — [Source](https://man.archlinux.org/man/restic-key-add.1.en)
|
||||
- Borg 2.x: repokey modes store encrypted key in repo (`repo_dir/config`); keyfile modes store in home dir (`~/.config/borg/keys`); `--key-location` repokey|keyfile chosen at `rcreate`, movable later via `key change-location` — [Source](https://borgbackup.readthedocs.io/en/2.0.0b6/usage/rcreate.html)
|
||||
- Borg 1.x: identical split — repokey inside repo directory, keyfile in `~/.config/borg/keys` (macOS: `~/Library/Application Support/borg/keys`); remote repos via ssh never see passphrase, plaintext key, or plaintext files — [Source](https://borgbackup.readthedocs.io/en/master/usage/repo-create.html)
|
||||
- Borg export UX: `borg key export [PATH]`, `--paper` (printable, per-line checksums for type-in), `--qr-html` (QR + paper copy); export stays passphrase-encrypted (no passphrase included); `borg key import [--paper]` restores; paper-key web tool at `paperkey.html` — [Source](https://borgbackup.readthedocs.io/en/master/usage/key.html); [Source](https://borgbackup.readthedocs.io/en/stable/paperkey.html)
|
||||
- Borg: for keyfile repos the key must be backed up independently ("NOT sufficient" to keep copy on the backed-up system); for repokey a backup is "not strictly needed" but guards against corruption/loss of the in-repo key — [Source](https://man.archlinux.org/man/borg-key-export.1.en)
|
||||
- Kopia: format blob (`kopia.repository` / `kopia.blobcfg` objects) lives in storage; local `repository-*.config` holds connection parameters (not the password); quick-reconnect token via `kopia repository status -t` (opaque) and `-s` embeds password (trivially decodable, must be guarded) — [Source](https://kopia.io/docs/reference/command-line/)
|
||||
- Kopia: repository password cached in OS-specific credential storage (Keychain on macOS, Credential Manager on Windows, Keyring on Linux) — [Source](https://kopia.io/docs/reference/command-line/)
|
||||
- rclone crypt: password + optional salt (`password2`) stored in `rclone.conf` in lightly obscured form (AES-CTR with static shared key, random IV prepended) — explicitly "not secure" without overall config-file encryption (`rclone config` password protection); env vars `RCLONE_CRYPT_PASSWORD` / `RCLONE_CRYPT_PASSWORD2` (must be `rclone obscure`d) — [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt: recovery = remember password or keep `rclone.conf`; same passwords re-entered on another machine reproduce access (obscured strings differ due to salt) — [Source](https://rclone.org/crypt/)
|
||||
- Tarsnap: single key file from `tarsnap-keygen --keyfile <file> --user <email> --machine <name>`; every `tarsnap` invocation needs `--keyfile`; passphrase via `--passphrase method:arg` (`dev:tty-stdin` default, `env:VAR`, `file:FILENAME` flagged as risky) — [Source](https://www.tarsnap.com/man-tarsnap-keygen.1.html); [Source](https://man.archlinux.org/man/tarsnap.1.en)
|
||||
- Tarsnap: `tarsnap-keymgmt --outkeyfile <new> [-r] [-w] [-d] [--nuke]` mints restricted sub-keys (read-only, write-only, delete, nuke-only); `-d` implies `-r` — [Source](https://www.tarsnap.com/man-tarsnap-keymgmt.1.html)
|
||||
|
||||
### Inferences
|
||||
- Only Borg gives a deliberate two-factor-style choice (possession of keyfile + knowledge of passphrase); restic/Kopia/rclone-default are passphrase-only if the repo/config is exfiltrated.
|
||||
- No system in scope documents TPM / hardware-backed key storage as of 2026; OS keyring appears only as a convenience cache (Kopia), not as a KDF or sealing mechanism.
|
||||
|
||||
### Gaps
|
||||
- No reliable source found on native TPM2 / Secure Enclave / Windows Hello integration for any of the four; community wrappers (pass, keyring scripts, systemd-creds) exist but were not verified as documented behavior.
|
||||
- restic's recommended master-key backup workflow beyond `cat masterkey` (e.g. official paper-key guidance) was not found.
|
||||
|
||||
## Do any support key rotation, multiple keys/passphrases, or Shamir/shared recovery? How?
|
||||
|
||||
### Takeaway
|
||||
restic and Borg (2.x/master) support multiple concurrent keys/passphrases sharing one master secret; Kopia and rclone crypt support only password change (re-wrapping the same secrets), with rclone requiring full re-upload; Tarsnap supports restricted sub-keys but not multi-passphrase unlock. No system documents Shamir secret sharing natively.
|
||||
|
||||
### Cited Findings
|
||||
- restic: `key` command with `list`, `add`, `remove`, `passwd` subcommands manages multiple access keys per repo; `key add` prompts for current password then new password and saves a new key wrapping the same master key; list marks current key with `*` — [Source](https://restic.readthedocs.io/en/stable/070_encryption.html)
|
||||
- restic: `key passwd` creates a new key ID and removes the old one; `key remove <ID>` refuses to remove the currently-used key — [Source](https://manpages.opensuse.org/Tumbleweed/restic/restic-key-passwd.1.en.html); [Source](https://man.archlinux.org/man/restic-key-remove.1.en.raw)
|
||||
- restic: `AddKey` with non-nil `template` copies master keys from the old key instead of generating new ones, confirming rotation re-wraps rather than re-encrypts data — [Source](https://github.com/restic/restic/blob/9e2d60e2/internal/repository/key.go)
|
||||
- Borg: `key change-passphrase` only re-locks the same secrets ("does not protect future nor past backups" if key+passphrase were compromised) — [Source](https://borgbackup.readthedocs.io/en/stable/usage/key.html)
|
||||
- Borg: `key change-algorithm argon2|pbkdf2` upgrades/downgrades the KDF wrapping without changing secrets — [Source](https://git.uninsane.org/shelvacu-mirrors/borg/commit/08f82ee40867f605ca6994db6dc32218d7b85cbc)
|
||||
- Borg (master/2.x): multiple borg keys per repo — each key holds the same secret material under an independent passphrase and label (first key labeled `admin`, protected from removal); `key list/add/remove/export --label|--key` selectors; passphrase tried against every key — [Source](https://github.com/borgbackup/borg/pull/9762)
|
||||
- Borg: `rcreate --other-repo SRC --copy-crypt-key` can reuse crypt key across related repos (default: fresh random crypt key, shared chunker/ID keys for dedup) — [Source](https://borgbackup.readthedocs.io/en/2.0.0b6/usage/rcreate.html)
|
||||
- Kopia: `kopia repository change-password` (CLI only, not GUI) re-encrypts the format block with a new password; must already be connected (so a still-connected client can reset a forgotten password, a disconnected one cannot) — [Source](https://kopia.io/docs/faqs/); [Source](https://kopia.io/docs/reference/command-line/common/repository-change-password/)
|
||||
- rclone crypt: no in-place password change — changing the configured password orphans existing content; must re-upload everything (in place via second crypt remote + `rclone copy` decrypting/re-encrypting on the fly, at 2x bandwidth/quota cost) — [Source](https://rclone.org/crypt/)
|
||||
- Tarsnap: no multi-passphrase unlock or rotation of archive encryption keys documented; capability separation via `tarsnap-keymgmt` restricted keys is the only delegation mechanism — [Source](https://www.tarsnap.com/man-tarsnap-keymgmt.1.html)
|
||||
- Shamir/shared recovery: no hits in official docs for restic, Borg, Kopia, rclone crypt, or Tarsnap; only community workarounds (splitting password/paper key with external SSS tools) exist — no primary-source citation available.
|
||||
|
||||
### Inferences
|
||||
- True cryptographic rotation (new data key, re-encryption) is offered by none of the five for existing snapshots; all "rotation" is re-wrapping or passphrase change.
|
||||
- Borg's multi-key design is closest to shared/admin recovery (per-user passphrases + protected admin key), but all keys still share one secret — compromise of the secret compromises all slots.
|
||||
|
||||
### Gaps
|
||||
- Whether `borg key add` (multi-key) is released in stable 2.0 vs master-only could not be pinned down; primary evidence is the PR plus master `usage/key.html`.
|
||||
- No documented Shamir/threshold scheme in any official docs; confirm as absent-by-design vs merely undocumented.
|
||||
|
||||
## What happens on key loss — is data unrecoverable by design? Any documented footguns?
|
||||
|
||||
### Takeaway
|
||||
All five declare password/key loss unrecoverable by design. The sharpest footguns: restic `key remove` of the sole key and deleted/corrupt `keys/` files; Borg keyfile loss without export and repokey-with-empty-passphrase on exposed storage; Kopia disconnected-password loss; rclone salt (`password2`) loss and obscured-config confusion; Tarsnap key-file loss (including billing trap).
|
||||
|
||||
### Cited Findings
|
||||
- restic docs repeat on every backend page: "knowledge of your password is required… Losing your password means that your data is irrecoverably lost" — [Source](https://restic.readthedocs.io/en/stable/030%5Fpreparing%5Fa%5Fnew%5Frepo.html)
|
||||
- restic forum (maintainer fd0): if key-file `data` or `salt` is missing/corrupt there is "no way to decrypt the data again. Even if you know the password"; intact `data`+`salt` (+N/r/p, brute-forceable) is required — [Source](https://forum.restic.net/t/recover-a-damaged-missing-corrupted-key-file/1798)
|
||||
- restic footgun case: user generated a random key file, ran `key remove` on the old key, then lost the random file (only copy was inside the backup) — recovery required disk forensics for the deleted key file; `key remove` is just a file deletion in `keys/` — [Source](https://forum.restic.net/t/recover-from-a-previous-password/2592)
|
||||
- restic: brute force only viable if much of the password structure is known; otherwise "consider it lost" — [Source](https://forum.restic.net/t/forgotten-password/2990)
|
||||
- Borg: "always need both the Borg key and passphrase"; keyfile loss = repo loss ("if you lose the key, you lose access"); mandated offsite export ("leaving your keys inside your car" warning); empty passphrase with repokey = anyone reading the repo can unlock (as good as no encryption); keyfile+empty passphrase acceptable only if client disk is encrypted — [Source](https://borgbackup.readthedocs.io/en/stable/usage/init.html)
|
||||
- Borg: exported key stays encrypted — recovery needs both export file AND original passphrase, stored in separate safe places — [Source](https://borgbackup.readthedocs.io/en/stable/usage/key.html)
|
||||
- Kopia: "you cannot restore your files if you forget your password, there is no way to recover a forgotten password because only you know it"; FAQ: "There is no way to recover it or the files and folders within that repository. Store your repository password in a safe place, such as a password manager" — [Source](https://github.com/kopia/kopia/blob/master/site/content/docs/Features/_index.md); [Source](https://kopia.io/docs/faqs/)
|
||||
- Kopia footgun nuance: password CAN be reset while still connected (`change-password`), which rescues forgotten-but-connected GUI users via CLI, but a disconnected client with a forgotten password is unrecoverable — [Source](https://kopia.io/docs/faqs/)
|
||||
- Kopia reconnect token (`status -t -s`) trivially decodes to the password — storing it insecurely leaks the repo — [Source](https://kopia.io/docs/reference/command-line/)
|
||||
- rclone crypt: custom salt is "effectively a second password that must be memorized" and is NOT stored with the data; losing `password2` loses data even with correct `password` — [Source](https://rclone.org/crypt/)
|
||||
- rclone footguns: obscured passwords look encrypted but use a static shared AES-CTR key (cursory-inspection only); 1.49.0–1.53.2 random-password generator bug produced insecure passwords (fixed 1.53.3, must rotate by re-upload) — [Source](https://rclone.org/crypt/)
|
||||
- Tarsnap FAQ: "You can't. Your key file contains the only copy of the cryptographic keys… if you lose them there is no way to get your data back" (same for forgotten key-file passphrase); separate FAQ covers being unable to stop billing after losing all keys — [Source](https://www.tarsnap.com/faq.html)
|
||||
|
||||
### Inferences
|
||||
- The envelope designs make server-side recovery cryptographically impossible, so every vendor pushes the same operational answer: paper/physical export + password manager + offsite separation of key-export and passphrase.
|
||||
- rclone's model is the most fragile operationally (two secrets + weak-by-default config obscuring + no re-wrap), Tarsnap second (single local file, no passphrase fallback).
|
||||
|
||||
### Gaps
|
||||
- Corrupt-repo recovery (bitrot, backend errors) vs key loss is covered by Kopia consistency docs and restic `rebuild-index`/`recover`, but systematic per-tool data on partial recovery without keys was out of scope and not collected.
|
||||
@@ -0,0 +1,126 @@
|
||||
# Untrusted-Server Threat Model: restic, BorgBackup, Kopia, rclone crypt on dumb backends
|
||||
|
||||
## What plaintext metadata does each system leave on the backend (repo layout, config files, sizes, timestamps)?
|
||||
|
||||
### Takeaway
|
||||
All four encrypt file content client-side but leak different amounts of structural metadata: restic leaks backend object sizes/counts/timestamps and config filenames; Borg leaks chunk sizes/proximity unless obfuscation is enabled; Kopia leaks only prefixed blob names (p/q/x) and approximate pack sizes; rclone crypt leaks file sizes (±16B) and mtimes and optionally directory structure.
|
||||
|
||||
### Cited Findings
|
||||
- restic design: “Apart from the files stored within the `keys` directory, all files are encrypted with AES-256 in counter mode (CTR)” with Poly1305-AES MAC, IV in first 16 bytes — keys directory itself is the exception and contains KDF parameters in plaintext JSON (`hostname`, `username`, `kdf=scrypt`, `N=65536,r=8,p=1`, `salt`, `created`) — [Source](https://github.com/restic/restic/blob/master/doc/design.rst); same guarantee restated as “Everything except the metadata included for informational purposes in the key files is encrypted and authenticated” — [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- restic repo layout on dumb backend is content-defined: `config`, `keys/`, `locks/`, `snapshots/`, `index/`, `data/` pack files named by hex SHA-256 of plaintext content; filenames of packs/snapshots/indexes are plaintext hashes, sizes and mtimes visible to server — [Source](https://restic.readthedocs.io/en/stable/100_references.html)
|
||||
- restic explicitly warns attacker with write access “can determine which files belong to what snapshot (e.g. based on the timestamps of the stored files)” — [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- restic local cache “is also encrypted to prevent metadata leaks” — [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- Borg attack model: “environment of the client process (e.g. borg create) is trusted and the repository (server) is not. The attacker has any and all access to the repository, including interactive manipulation (man-in-the-middle)” — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg guarantees vs untrusted server: attacker cannot (1) modify archive data undetected, (2) rename/remove/add archive undetected, (3) recover plaintext, (4) recover definite structural information such as object graph — but “heuristics based on access patterns are possible” — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg multi-client caveat: “When the above attack model is extended to include multiple clients independently updating the same repository, then Borg fails to provide confidentiality (i.e. guarantees 3) and 4) do not apply any more)” — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg chunk-ID is MAC of plaintext (`id = AUTHENTICATOR(id_key, data)`), encryption is Encrypt-then-MAC AES-256-CTR + HMAC-SHA256 or BLAKE2b-256; IV counter “added in plaintext” and tracked via client security DB + repo reservation — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg 2.0 changes to AEAD (`AES-256-OCB` or `chacha20-poly1305`) with session keys, chunk ID via MAC over plaintext, object ID bound as AAD so attacker cannot “change the type of the object or move content to a different object ID” — [Source](https://borgbackup.readthedocs.io/en/2.0.0b12/internals/security.html)
|
||||
- Borg fingerprinting: “A borg repository does not hide the size of the chunks it stores”; small files <512KiB yield single chunk; buzhash chunker uses secret per-repo chunker seed; optional `obfuscate` pseudo-compressor pads with 0x00 bytes (only adds size) to hinder size fingerprinting — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg proximity leak: “Borg does not try to obfuscate order / proximity of files”; sorts by inode order not name, but “when new files are close to each other [in recursion order], the resulting chunks will be also stored close to each other in segment file(s)” — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Kopia envelope encryption: repo has plaintext JSON `formatBlob` with `tool`, `buildVersion`, `uniqueID`, `keyAlgo=scrypt-65536-8-1`, `version`, `encryption` (e.g. `AES256_GCM`), plus `encryptedBlockFormat` ciphertext; `UniqueID` doubles as PBKDF salt and HKDF info — [Source](https://kopia.io/docs/advanced/encryption)
|
||||
- Kopia passphrase → 32B master key `Km = PBKDF(passphrase, UniqueID)` (scrypt N=65536,r=8,p=1), then `Ke = HKDF(SHA256,Km,UniqueID,"AES",32)` and `AD = HKDF(SHA256,Km,UniqueID,"CHECKSUM",32)` for format-blob AES-GCM — [Source](https://kopia.io/docs/advanced/encryption)
|
||||
- Kopia content encryption options are `AES256-GCM-HMAC-SHA256` (default) or `CHACHA20-POLY1305-HMAC-SHA256`, chosen at repo creation, cannot be changed after — [Source](https://github.com/kopia/kopia/blob/master/site/content/docs/FAQs/_index.md)
|
||||
- Kopia blob layer: pack blobs have random names, “don’t reveal anything about their contents or structure”; prefixes `p` (data packs), `q` (metadata packs), `x` (indices); object-ID single-letter prefixes `k` (directory), `m` (manifest block), `x` (indirect JSON) route to `q` vs `p` packs — [Source](https://kopia.io/docs/advanced/architecture)
|
||||
- Kopia packs are 20–40MB merged blobs; “sizes are also generally unrelated to their content due to splitting and merging”; per-pack trailing local index enables recovery if global index lost — [Source](https://kopia.io/docs/advanced/architecture)
|
||||
- Kopia manifests (snapshots/policies) are small JSON stored as encrypted CABS blocks, addressed by `key=value` labels (e.g. `type:policy`, `hostname:… path:… username:…` visible via `kopia manifest list` client-side, not plaintext on backend) — [Source](https://kopia.io/docs/getting-started/)
|
||||
- Kopia: “encrypts these snapshots before they leave your computer”, “password never leaves your machine”, single password per repo, “currently no access control mechanism when sharing a repository” — [Source](https://kopia.io/docs/features/)
|
||||
- rclone crypt: wraps any backend; “automatically encrypt (before uploading) and decrypt (after downloading) on your local system … leaving the data encrypted at rest”; direct access to wrapped remote bypasses crypto — [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt file format: 8B magic `RCLONE\x00\x00` + 24B nonce header, then 64KiB chunks each with 16B Poly1305 tag (XSalsa20+Poly1305 SecretBox); 1B file → 49B total, 1MiB → 1048864B (+0.03%) — [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt name encryption: path split on `/`, PKCS#7 padded to 16B, EME-AES-256 deterministic, base32-lowercase-no-pad; identical names → identical ciphertexts; common prefixes hidden — [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt explicitly “does not encrypt file length - this can be calculated within 16 bytes” nor “modification time - used for syncing”; versions suffix `-vYYYY-MM-DD…` left plaintext; directory names optionally unencrypted — [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt filename modes: `standard` (encrypted, ~143 char limit, dir structure visible), `obfuscate` (“Very simple filename obfuscation … cannot be relied upon for strong protection”), `off` (adds `.bin` only) — [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt: “Any metadata supported by the underlying remote is read and written” — i.e. WebDAV/SFTP xattrs/mtimes pass through unencrypted — [Source](https://rclone.org/crypt/)
|
||||
- rclone WebDAV backend: “Plain WebDAV does not support modified times” nor hashes except via Fastmail/ownCloud/Nextcloud vendor extensions — [Source](https://rclone.org/webdav/)
|
||||
|
||||
### Inferences
|
||||
- For size-hiding, Kopia pack merging (>20MB) is the strongest default; restic pack (~4-8MB default, content-defined) is intermediate; Borg chunk-level sizes leak most without `obfuscate` compressor.
|
||||
- Deterministic name encryption (rclone crypt standard, restic hash-named packs are content-derived not name-derived) still leaks equality + directory shape; only Kopia random pack names hide shape by default.
|
||||
- Plaintext `keys/` (restic) and `formatBlob` (Kopia) both expose KDF parameters and repo unique IDs — useful for offline dictionary attack cost estimation but not plaintext.
|
||||
|
||||
### Gaps
|
||||
- No reliable primary-doc figure found for exact restic snapshot/index plaintext filename entropy or padding policy; restic docs do not claim filename obfuscation or padding.
|
||||
- Borg `obfuscate` compressor size-overhead curve (“few percent [cheap] to ridiculously larger [expensive]”) has no numeric table in fetched docs.
|
||||
|
||||
## How do they handle an actively malicious or compromised server (authentication of manifests, rollback protection, replay of old snapshots)?
|
||||
|
||||
### Takeaway
|
||||
Borg has the strongest formal malicious-server story (Horton-anchored DAG + TAM-signed manifest + nonce tracking); restic authenticates all objects but explicitly does not detect deletion/rollback or timestamp-grouping attacks; Kopia relies on content-addressed HMAC + maintenance/verify with server-side immutability as add-on; rclone crypt has per-chunk authentication but no manifest/rollback concept.
|
||||
|
||||
### Cited Findings
|
||||
- restic guarantees: “Modifications to data … can be detected” and “Data that has been tampered will not be decrypted” (MAC checked before decrypt) — [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- restic non-goals: “not designed to protect against attackers deleting files”; attacker deleting timestamp-correlated packs makes “particular snapshot vanish … Restic is not designed to detect this attack” — [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- restic compromised-host cases: attacker with repo write can “Create snapshots (containing garbage data) which cover all modified files and wait until a trusted host has used forget often enough to remove all correct snapshots” and “Create a garbage snapshot for every existing snapshot with slightly different timestamp” to trick rotation — [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- restic append-only caveat: attacker with append-only access can still “Capture the password and decrypt past and future backups” (no forward secrecy); safe `forget` requires separate doc procedure — [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- restic key-leak: “impossible to securely revoke a leaked key without re-encrypting the whole repository” — [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- restic open issue #5041 proposes extending model with append-only enforcement + management/data key separation so “adversary who compromises the host cannot retroactively publish a snapshot which will trick forget into deletion” — still open/discussion as of 2026 — [Source](https://github.com/restic/restic/issues/5041)
|
||||
- Borg structural auth follows “Horton principle”: every object referenced by parent via plaintext-MAC object ID up to manifest, forming authenticated DAG — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg manifest has fixed ID `000…000` so cannot be DAG-authenticated; protected since 1.0.9 by TAM: `tam_key = HKDF-SHA-512(id_key||enc_key||enc_hmac_key, RANDOM(64), "borg-metadata-authentication-manifest")`, `HMAC(tam_key, packed_manifest)` stored in manifest — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg nonce/counter anti-reuse: client commits +4GiB reservation to security DB + repo before encrypting; crash-safe via SaveFile; but “in a multiple-client scenario a repository can trick a client into reusing counter values by ignoring counter reservations and replaying the manifest (which will fail if the client has seen a more recent manifest or has a more recent nonce reservation)” — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg RPC: over system SSH (no own network code), server can only send responses not requests; msgpack limited Unpacker; worst-case server can impose repo DoS; log-channel confusion limited to in-progress requests — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg CVE-2023-36811: “flaw in the cryptographic authentication scheme … allowed an attacker to fake archives and potentially indirectly cause backup data loss”; requires inserting files without headers + repo write; “does not disclose plaintext … nor affect authenticity of existing archives”; fixed in 1.2.5 + upgrade procedure; mitigate by reviewing archives after `check --repair` before `prune` — [Source](https://nvd.nist.gov/vuln/detail/CVE-2023-36811)
|
||||
- Borg append-only: `borg config repo append_only 1` or `borg serve --append-only`; “never overwrite or delete committed data” at segment level, but `delete/prune` still allowed (appear to succeed, only tag as deleted in new transaction); `compact` becomes no-op with no warning — [Source](https://raw.githubusercontent.com/borgbackup/borg/1.4.5/docs/usage/notes.rst)
|
||||
- Borg append-only rollback: transaction log (`transactions` file with UTC timestamps) allows manual rollback by deleting segment files from attack point onward (e.g. `rm data/**/{6..13}`), provided `compact` has not run; must clear client cache after (`borg delete --cache-only`) — [Source](https://raw.githubusercontent.com/borgbackup/borg/1.4.5/docs/usage/notes.rst)
|
||||
- Borg append-only limits: “Append-only mode is not respected by tools other than Borg. rm still works”; clients must only access via `borg serve`; any non-append-only write (admin prune/create) permanently compacts away “deleted” data; SSH `authorized_keys` split (`--append-only` key for untrusted clients, full key for admin) is recommended pattern — [Source](https://borgbackup.readthedocs.io/en/1.1-maint/usage/notes.html)
|
||||
- Kopia consistency: `kopia snapshot verify` walks snapshot roots, checks index structures + blob existence; runs automatically during daily full maintenance; `--verify-files-percent=N` samples downloads/decrypts (100% ≈ test restore, discarded after check); e.g. 1% daily → ~98% coverage over a year — [Source](https://kopia.io/docs/advanced/consistency/)
|
||||
- Kopia corruption causes listed as non-POSIX/networked filesystems, silent bit-rot, large clock skew (few minutes tolerated, larger can cause self-deletion); recommends mature POSIX FS or cloud storage + NTP — [Source](https://kopia.io/docs/advanced/consistency/)
|
||||
- Kopia ransomware model: malware may exfiltrate cloud keys and delete snapshots; mitigation is provider-side restricted keys (no delete) + object-lock/retention (COMPLIANCE), not client crypto alone — [Source](https://kopia.io/docs/advanced/ransomware-protection/)
|
||||
- Kopia object-lock: `--retention-mode COMPLIANCE --retention-period <e.g.30d>` at `repo create s3`, plus `maintenance set --extend-object-locks true` with `full-interval` ≥1 day shorter than retention; supports S3 (+B2-via-S3), Azure version-level immutability, GCS versioning+retention; `--point-in-time` reconnect for S3 restores — [Source](https://kopia.io/docs/advanced/ransomware-protection/)
|
||||
- rclone crypt integrity: per-chunk Poly1305 via SecretBox; “data integrity is protected by an extremely strong crypto authenticator”; `cryptcheck` (not `check`) required since hashes not stored — [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt has no manifest/snapshot ledger: replay of old encrypted files, rollback, or server-side deletion is undetectable by crypt layer; `--pass-bad-blocks` (zero-fill corrupt chunks) is opt-in recovery only — [Source](https://rclone.org/crypt/)
|
||||
|
||||
### Inferences
|
||||
- Only Borg attempts to bind manifest meaning (TAM) against a fully MITM server with persistent client state; restic/Kopia assume server may be read/write-malicious for confidentiality/integrity of objects but push rollback/availability to out-of-band controls (append-only buckets, object-lock, separate admin host).
|
||||
- For WebDAV-class backends with no compute, Borg-style `serve --append-only` enforcement is unavailable; closest analogues are restic append-only-capable REST backends (not WebDAV) or Kopia S3 object-lock — WebDAV must rely on server ACLs/versioning.
|
||||
|
||||
### Gaps
|
||||
- No primary-source evidence found that Kopia authenticates manifest labels/meaning against malicious-server manifest substitution beyond content HMAC; treat as gap.
|
||||
- No restic primary doc found describing rollback detection (e.g. monotonic snapshot counter); restic docs explicitly disclaim it.
|
||||
|
||||
## How do they rate-limit, retry, and verify uploads (idempotent PUTs, integrity re-checks)?
|
||||
|
||||
### Takeaway
|
||||
All target dumb blob stores with idempotent content-addressed PUTs and client-side re-verification (`check`/`verify`/`cryptcheck`); retry/backoff is client-configured and backend-specific; none documents server-side rate-limiting — throttling is via client concurrency limits and provider quotas.
|
||||
|
||||
### Cited Findings
|
||||
- restic repo design “allows parallel access of multiple instances … even parallel writes”; locks with `--retry-lock` retry until timeout; `restic check` verifies pack hash vs pack ID, `cat pack <id>` warns on mismatch — [Source](https://restic.readthedocs.io/en/latest/100_references.html?highlight=threat)
|
||||
- Kopia architecture: content-addressed blocks (`Block ID = hash(data)`); identical blocks dedup naturally; uploads merged into packs; index maps block→(blob,offset,len) — enabling idempotent re-PUT (same ID = same bytes) — [Source](https://kopia.io/docs/advanced/architecture)
|
||||
- Kopia verification tiers: metadata-only `snapshot verify` every full maintenance + opt-in content download sample `--verify-files-percent` + `--file-parallelism` for throughput control — [Source](https://kopia.io/docs/advanced/consistency/)
|
||||
- Kopia storage requirements imply retry need: “strong read-after-write consistency and eventual list-after-write consistency”; “can compensate for such inconsistent behaviors for up to several minutes, but larger inconsistencies can lead to data loss” — [Source](https://kopia.io/docs/advanced/consistency/)
|
||||
- Borg RPC worst-case throttling noted only as server-imposed “denial of repository service”; client uses limited msgpack Unpacker to avoid memory-DoS from large messages — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- rclone crypt backup guidance: use `sync` on encrypted paths with same passwords so “will check the checksums while copying”; `check` between two encrypted remotes works, `cryptcheck` needed vs plaintext — [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt chunking (64KiB + 16B tag) bounds retry unit; header nonce incremented per chunk, reuse probability ~2e-32 per exabyte — [Source](https://rclone.org/crypt/)
|
||||
|
||||
### Inferences
|
||||
- Content-addressing (restic pack ID, Borg chunk ID, Kopia block ID) makes PUTs naturally idempotent and safe to retry; verification is always client-driven (`check`/`verify`/`cryptcheck`) since dumb backends cannot attest.
|
||||
- Parallelism knobs (`--file-parallelism`, rclone `--transfers/--checkers`, Borg client concurrency) are the de facto rate-limiters; no evidence of server-push backpressure handling in fetched docs.
|
||||
|
||||
### Gaps
|
||||
- No primary-source values found in fetched docs for restic `--retry-lock` defaults, backend backoff schedule, or WebDAV-specific retry/idempotency; mark as gap — need restic backend docs + rclone WebDAV backend options.
|
||||
- No Kopia primary-doc retry/backoff parameters or blob `Put` atomicity guarantees surfaced in fetched pages; need `repo/blob` Go API docs.
|
||||
- No Borg primary-doc upload retry/rate-limit parameters surfaced; Borg docs focus on correctness not throttling.
|
||||
|
||||
## Any guidance specific to WebDAV-class dumb backends (no server compute) that applies to a local daemon syncing to WebDAV?
|
||||
|
||||
### Takeaway
|
||||
Treat WebDAV as untrusted byte store: do all crypto/index/manifest work client-side, never depend on server mtime/hash, use atomic PUT + verify-after-write, and add out-of-band append-only/versioning since WebDAV has no compute to enforce it.
|
||||
|
||||
### Cited Findings
|
||||
- Kopia repository model: “All Repository features are implemented client-side, without any need for a custom server, thus encryption keys never leave the client”; layers are Object/Manifest/Block/Raw-BLOB over simple blob API — directly applicable to WebDAV — [Source](https://github.com/kopia/repo/blob/master/README.md)
|
||||
- Kopia supports “Any remote server or cloud storage that supports WebDAV” and SFTP as first-class repo storage, plus Rclone-wrapped Dropbox/OneDrive/Google-Drive (experimental) — [Source](https://github.com/kopia/kopia?pubDate=20260701)
|
||||
- Kopia caveat for dumb/networked FS: requires “POSIX semantics, atomic writes (either native or emulated), strong read-after-write … eventual list-after-write”; “avoid emulated, layered, or networked filesystems which may not be fully compliant. Alternatively, use cloud storage”; WebDAV falls in risky class — [Source](https://kopia.io/docs/advanced/consistency/)
|
||||
- Kopia clock rule: “clock skews of few minutes are tolerated, but bigger clock skews can lead to major inconsistencies, including Kopia deleting its own data”; run NTP on clients and servers — [Source](https://kopia.io/docs/advanced/consistency/)
|
||||
- rclone WebDAV limits: “Plain WebDAV does not support modified times … does not support hashes” (except Fastmail/ownCloud/Nextcloud extensions) — so sync daemons must not trust mtime/hash for change detection — [Source](https://rclone.org/webdav/)
|
||||
- rclone crypt on WebDAV pattern: point crypt at `remote:path` subdirectory, access exclusively via crypt remote; `.bin` suffix added when names unencrypted to prevent provider interpreting content; `strict_names` errors on mixed encrypted/unencrypted dirs — [Source](https://rclone.org/crypt/)
|
||||
- rclone crypt password rotation requires full re-upload (decrypt-with-old → encrypt-with-new), either from local source or crypt-to-crypt `move`; “half the bandwidth and charged twice” on metered backends — [Source](https://rclone.org/crypt/)
|
||||
- restic threat-model implication for WebDAV: server sees object sizes/counts/timestamps; timestamp-grouping enables targeted snapshot deletion — mitigation is to avoid leaking timing (jitter uploads, shared pack sizes) though not prescribed in docs — [Source](https://rustic.cli.rs/dev-docs/design/threat_model.html)
|
||||
- Borg inode-order (not alpha) directory traversal + secret chunker seed + optional size obfuscation are concrete dumb-backend-hardening techniques portable to any daemon: randomize scan order, per-repo secret for chunking, pad to size classes — [Source](https://borgbackup.readthedocs.io/en/stable/internals/security.html)
|
||||
- Borg SSH `authorized_keys` split (`borg serve --append-only` for untrusted clients vs full for admin) has no WebDAV equivalent; WebDAV equivalent is provider ACL/versioning or Kopia-style object-lock where available (S3/Azure/GCS only — not WebDAV) — [Source](https://borgbackup.readthedocs.io/en/1.1-maint/usage/notes.html); Kopia object-lock “currently only supports object locks when using an S3 repo” (native B2 excluded, use S3 mode) — [Source](https://kopia.io/docs/advanced/ransomware-protection/)
|
||||
- Filippo Valsorda review quoted in restic docs: “The design might not be perfect, but it’s good. Encryption is a first-class feature, the implementation looks sane and … deduplication trade-off is worth it” — informal audit signal, not formal audit — [Source](https://github.com/restic/restic/blob/master/doc/070_encryption.rst)
|
||||
|
||||
### Inferences
|
||||
- For a local daemon syncing to WebDAV: (1) encrypt+MAC everything client-side with random pack names (Kopia-style) or hash-named packs (restic-style); (2) keep manifest/index signed with local key (Borg TAM-style) and verify on every sync; (3) never use server mtime/ETag as source of truth; (4) verify-after-write via GET+MAC before deleting local staging; (5) mitigate rollback by keeping local monotonic snapshot counter + out-of-band copy; (6) pad/obfuscate sizes and jitter upload times to reduce grouping leaks.
|
||||
- Since WebDAV cannot enforce append-only, ransomware safety must come from server-side versioning/quotas + separate prune identity, mirroring Borg append-only and Kopia restricted-keys guidance.
|
||||
|
||||
### Gaps
|
||||
- No primary WebDAV-server hardening guide (append-only ACLs, versioning) surfaced for restic/Borg/Kopia on plain WebDAV; restic WebDAV backend doc and rclone WebDAV option reference not fetched (tool-call budget) — need follow-up.
|
||||
- No formal third-party audit reports (e.g. Cure53/Quarkslab) confirmed for restic/Borg/Kopia in fetched sources; only Valsorda informal review and CVE record found — need dedicated audit search.
|
||||
@@ -0,0 +1,98 @@
|
||||
# Binary, Large, Image, Archive and Database Files in Versioning/Backup Tools
|
||||
|
||||
## What size caps or chunking strategies do these tools use, and what breaks with multi-MB binaries in a per-version system (storage blowup, bandwidth)?
|
||||
|
||||
### Takeaway
|
||||
Restic, Borg and Kopia all use content-defined chunking (CDC) with ~0.5–8 MiB variable chunks so small edits to large binaries only store 1–2 new chunks; without CDC (fixed blocks or full-file copies) a 1-byte insert re-stores the whole file, and Git-LFS instead punts large files to pointer+blob storage with host-enforced per-file caps.
|
||||
|
||||
### Cited Findings
|
||||
- Restic splits files with Rabin-fingerprint CDC over a 64-byte sliding window, cutting when low 21 bits are zero; files <512 KiB are not split, blobs are 512 KiB–8 MiB, ~1 MiB average — [Source](https://github.com/restic/restic/blob/master/doc/design.rst); background — [Source](https://restic.net/blog/2015-09-12/restic-foundation1-cdc/)
|
||||
- Restic chunker defaults aim at ~1 MiB average (`splitmask = (1<<20)-1`) with configurable Min/MaxSize — [Source](https://github.com/restic/chunker/blob/master/chunker.go)
|
||||
- Borg splits files into deduplicated chunks globally across repo (all machines/archives); chunk id is a strong hash/MAC (hmac-sha256 / keyed blake3), not the rolling-hash value — [Source](https://borgbackup.readthedocs.io/en/master/)
|
||||
- Borg default buzhash params: min 2^19 (512 KiB), max 2^23 (8 MiB), mask 21 bits (~2 MiB target), window 4095 B; `fixed` chunker option for disk/VM images — [Source](https://github.com/borgbackup/borg/blob/86fd77fd/docs/internals/data-structures.rst)
|
||||
- Borg 2.x chunkers: `fastcdc` (default, fastest), `buzhash64`/`buzhash`, keyed AES variants (`toeplitz-aes`/`rabin-aes` strongest), `fixed` for raw disk images where CDC gains little — [Source](https://borgbackup.readthedocs.io/en/latest/internals/chunker.html)
|
||||
- Borg warns fine-grained `--chunker-params=buzhash,10,23,16,4095` creates huge chunk counts and RAM/disk load; coarse default suits large volumes — [Source](https://borgbackup.readthedocs.io/en/stable/usage/notes.html)
|
||||
- Kopia calls CDC "splitters": FIXED vs DYNAMIC BUZHASH/RABINKARP, sizes 1M–8M (default DYNAMIC-4M-BUZHASH); small files = one content, large files split so metadata-only change to 10 GB video uploads only 1–2 chunks (<10 MB) — [Source](https://kopia.discourse.group/t/does-kopia-use-content-defined-chunking-cdc/1417); details — [Source](https://kopia.discourse.group/t/difference-between-the-available-splitters/894); packing — [Source](https://kopia.discourse.group/t/do-hashed-chunks-span-multiple-files/1444)
|
||||
- Kopia packs many contents into 20–40 MB pack blobs; splitter choice is set at repo creation — [Source](https://kopia.io/docs/advanced/architecture/); chunk-then-hash-then-compress-then-encrypt pipeline — [Source](https://kopia.io/docs/advanced/compression/)
|
||||
- Restic has no hard max file size but offers `--exclude-larger-than size` (suffixes k/M/G/T) to skip files over a threshold — [Source](https://restic.readthedocs.io/en/stable/040_backup.html)
|
||||
- Git-LFS has no inherent file-size limit; limits are host-enforced (GitHub: 2 GB Free/Pro, 4 GB Team, 5 GB Enterprise Cloud; >5 GB rejected); pointer file stores `version/oid sha256/size` — [Source](https://docs.github.com/en/repositories/working-with-files/managing-large-files/about-git-large-file-storage)
|
||||
- Git-LFS tracks by `.gitattributes` pattern, not by size; `--above` in `migrate import` is one-shot, no automatic by-size tracking (2025 `--min-size/autotracksize` PR still debated/unmerged) — [Source](https://github.com/git-lfs/git-lfs/blob/main/docs/man/git-lfs-faq.adoc)
|
||||
- Git on Windows pre-2.34 could not smudge/clean files >4 GiB; workaround `GIT_LFS_SKIP_SMUDGE=1` + `git lfs pull` — [Source](https://github.com/git-lfs/git-lfs/blob/main/docs/man/git-lfs-faq.adoc)
|
||||
- Every Git-LFS revision counts against remote storage/bandwidth quota, so per-version binaries inflate cost; clones/pulls are slow on large repos — [Source](https://get.assembla.com/blog/git-lfs/)
|
||||
- Keyed CDC (Borg/Restic/Kopia) is fingerprintable: observing chunk sizes of known data recovers keys (CCS 2025, eprint 2025/558; backup-service attacks eprint 2025/532); Restic 0.18 mitigates by random chunk-to-pack assignment — [Source](https://github.com/restic/restic/blob/master/doc/design.rst); attacks — [Source](https://eprint.iacr.org/2025/558); analysis — [Source](https://eprint.iacr.org/2025/532.pdf)
|
||||
|
||||
### Inferences
|
||||
- A small local daemon should copy the Restic/Borg default: CDC ~1–2 MiB average, 512 KiB min, 8 MiB max; offer `--exclude-larger-than`-style cap (e.g. 50–100 MB default-off) plus extension-based excludes rather than a hard cap.
|
||||
- Per-version full copies of multi-MB binaries blow up storage and bandwidth linearly with versions; CDC reduces this to delta-of-chunks but still re-reads/re-hashes the whole file each run.
|
||||
- For VM/disk images with fixed internal layout, a `fixed`-block mode is faster and equally effective.
|
||||
|
||||
### Gaps
|
||||
- No reliable 2026 source found on VSCode/JetBrains Local History size caps or binary handling (searches rate-limited); editor-local-history defaults remain unverified.
|
||||
- Exact Kopia default content size bounds (min/max around 4 MB mean) not confirmed from primary docs; forum states 1–8 MB range only.
|
||||
|
||||
## NUL-byte sniffing for binary detection: who does it and is it still recommended?
|
||||
|
||||
### Takeaway
|
||||
Git (and libgit2 ports) still use NUL-in-first-8000-bytes as the primary binary signal, tightened in 2008 so CRLF conversion defers to diff; it is fast and recommended as a first heuristic but known-insufficient for UTF-16 and NUL-free binaries, so explicit `.gitattributes` marking is required.
|
||||
|
||||
### Cited Findings
|
||||
- Git `buffer_is_binary()` checks for NUL via `memchr` in first 8000 bytes (`FIRST_FEW_BYTES`) — [Source](https://stackoverflow.com/questions/6119956/how-to-determine-if-git-handles-a-file-as-binary-or-as-text)
|
||||
- Git `convert.c` `convert_is_binary()`: binary if `lonecr` or `nul` or `(printable>>7) < nonprintable`; CRLF auto-conversion bails on binary — [Source](https://code.googlesource.com/git/+/645cc7a2a7274a92403d2848ef643a96f1589d09/convert.c)
|
||||
- libgit2 `git_blob_is_binary` uses core-git heuristic: NUL scan + printable/nonprintable ratio over first 8000 bytes — [Source](https://libgit2.org/docs/reference/main/blob/git_blob_is_binary.html)
|
||||
- 2008 patch unified heuristics: any NUL forces binary in `convert.c` so CRLF handling is stricter than diff (prior convert.c used only <1% nonprintable rule, mis-handling tar/word-processor files diff called binary) — [Source](https://public-inbox.org/git/20080116011321.GD13984@dpotapov.dyndns.org/t/)
|
||||
- Git mailing-list guidance: NUL in first 8000 bytes = binary; UTF-16 must be marked explicitly, Git does not handle it internally; short NUL-free binaries must also be marked explicitly — [Source](https://public-inbox.org/git/20151202004921.GC28197@sigill.intra.peff.net/T/)
|
||||
- `git diff --numstat` reports `-\t-` for binary (practical detector); `git check-attr` only reflects `.gitattributes`, not the heuristic — [Source](https://stackoverflow.com/questions/6119956/how-to-determine-if-git-handles-a-file-as-binary-or-as-text)
|
||||
|
||||
### Inferences
|
||||
- NUL-sniffing remains the recommended cheap first pass for a small daemon (scan first 8 KiB), matching Git/libgit2 behavior and user expectations.
|
||||
- It must be paired with an explicit override list (extensions + `.gitattributes`-style `binary`/`-text`) because UTF-16 text false-positives as binary and small high-entropy binaries without NUL false-negative as text.
|
||||
|
||||
### Gaps
|
||||
- No 2023–2026 primary source found revising or deprecating NUL-sniffing; whether modern editors still use exactly 8000 bytes vs 4000/8000 variants (libgit2 history) is unresolved.
|
||||
|
||||
## How to safely snapshot live sqlite files (WAL checkpoint, .dump, filesystem snapshot) vs copying the raw file?
|
||||
|
||||
### Takeaway
|
||||
Never `cp`/read the raw sqlite file hot: in WAL mode the consistent image spans main+`-wal`+`-shm` and byte copies tear or go stale; use the Online Backup API (`sqlite3_backup_*` / `Connection.backup()` / `.backup` CLI), `VACUUM INTO`, or a quiesced filesystem snapshot, then `PRAGMA integrity_check`.
|
||||
|
||||
### Cited Findings
|
||||
- Historical `cp`-under-shared-lock method is fast but blocks writers, cannot copy to/from memory DBs, and risks corruption on power/OS failure — [Source](https://sqlite.org/backup.html)
|
||||
- Backup API: source read-locked only during each `sqlite3_backup_step(nPage)`; destination write-locked throughout; incremental stepping lets writers proceed; concurrent write by another connection restarts backup automatically — [Source](https://sqlite.org/backup.html); API contract — [Source](https://sqlite.org/c3ref/backup_finish.html)
|
||||
- Python exposes as `Connection.backup(target, pages, progress, sleep)`; `pages=-1` copies all at once (holds lock), positive pages + `sleep=0.250` yields between steps; restart detected when `remaining` jumps back toward `total` — [Source](https://www.productionhardening.org/backup-recovery-data-integrity/online-backup-api-hot-copies/)
|
||||
- WAL-mode live DB is three files (main + `-wal` committed-not-checkpointed frames + `-shm` index); `cp`/`rsync`/snapshot captures them at different instants → `SQLITE_CORRUPT`/`SQLITE_NOTADB`; same hazard for `-journal` in rollback mode — [Source](https://www.productionhardening.org/backup-recovery-data-integrity/)
|
||||
- Real-world failure: `fs.copyFile()` of only `.db` in WAL mode produced 100% corrupt backups across weeks of 6-hour cron; fix was `better-sqlite3 .backup()` plus hour-granular filenames — [Source](https://scottspence.com/posts/sqlite-corruption-fs-copyfile-issue)
|
||||
- Correct one-liners: `sqlite3 app.db ".backup './db-backups/app.db'"` (general) or `VACUUM INTO` (same safety + compaction, needs SQLite 3.27+, refuses if target exists so `rm -f` first); for dedup pipelines prefer `.backup` because compaction reshuffles pages and defeats chunking — [Source](https://www.backupdata.io/resources/guides/sqlite-backups-you-can-actually-restore)
|
||||
- Borg docs: Borg just copies file as-is; if DB is written mid-read the archive may be inconsistent — use sqlite-aware method (`sqlite3 db.sqlite "VACUUM INTO 'copy.sqlite'"`) or filesystem snapshot — [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst); Borg FAQ repeats the `VACUUM INTO` advice — [Source](https://borgbackup.readthedocs.io/en/master/faq.html)
|
||||
- borgmatic best practice: dump (export) databases rather than backing internal files; streams dump directly to Borg — [Source](https://torsion.org/borgmatic/how-to/backup-your-databases/)
|
||||
- Checkpoint+lock schemes (`wal_checkpoint(TRUNCATE)` then `BEGIN IMMEDIATE` then copy main+wal) are fragile: checkpoint can fail to get writer lock, writers can interleave, autocheckpoint is per-process — use the backup API instead — [Source](https://sqlite.org/forum/forumpost/2ea989bbe9)
|
||||
- Post-backup rules: run `PRAGMA integrity_check` (expect single `ok`) on every finished image; pre-size target ~1.2× source; set `busy_timeout ≥5000ms`; write target off the source I/O queue; delete partials; on restore stop app, copy main file, delete stale `-wal`/`-shm` — [Source](https://www.productionhardening.org/backup-recovery-data-integrity/online-backup-api-hot-copies/); restore checklist — [Source](https://www.backupdata.io/resources/guides/sqlite-backups-you-can-actually-restore)
|
||||
|
||||
### Inferences
|
||||
- Small-daemon default: if file header is `SQLite format 3`, do not copy directly; shell to `sqlite3 "$f" ".backup '$tmp'"` or `VACUUM INTO` (when compaction desired) into a temp file, then ingest the temp file; fall back to copy only when DB is quiesced and `wal_checkpoint(TRUNCATE)` shows zero log pages.
|
||||
- Smaller Borg-style chunks (~64–128 KiB target, e.g. `fastcdc,15,19,17`) cut sqlite re-storage ~3–8× vs 2 MiB defaults at the cost of chunk-count RAM — see next section.
|
||||
|
||||
### Gaps
|
||||
- `.dump` (SQL text) vs `.backup` (page image) size/dedup tradeoff not quantified from primary sources in this pass; anecdotal claim that dumps dedup to KBs but take hours on 12 GB DBs is single-issue-report only.
|
||||
|
||||
## Do images/archives dedup at all, and is per-version storage of them worth it vs plain mirroring?
|
||||
|
||||
### Takeaway
|
||||
Recompressed, encrypted, or gzipped-per-version artifacts barely dedup (a 1-byte change avalanches through compression; encrypted chunks are high-entropy), while uncompressed raw images and plain sqlite files dedup well under CDC — so version verbatim binaries that change little, otherwise mirror single-copy.
|
||||
|
||||
### Cited Findings
|
||||
- Restic maintainer: pass uncompressed data; small input change → large compressed-output change → new blob hashes → repo grows (50 GB gzipped DB dumps → 50 GB repo vs 36 GB for raw); CDC blobs identified by SHA-256 — [Source](https://github.com/restic/restic/issues/790)
|
||||
- Duplicati skips recompression/dedup for known-compressed extensions by default (saves CPU; whole-file handling) because metadata edits rewrite the stream; identical copies still dedup, moves are free — [Source](https://forum.duplicati.com/t/deduplication-for-large-files/8394)
|
||||
- Borg on 5 generations of same sqlite DB: gzipped inputs = 667 unique/667 total chunks (zero dedup, 1.6 GB); uncompressed = strong dedup (590 MB total), zstd-5 beats gzip — [Source](https://appsintheopen.com/posts/66-backing-up-sqlite-database-with-borg-and-de-duplication)
|
||||
- Borg sqlite tuning: default ~2 MiB target wastes a whole chunk per changed 4 KiB page; `fastcdc,15,19,17,2` (~128 KiB target) drastically cuts incrementals; finer `14,18,16` helps more; cost is chunk-index RAM — params apply per-run so back up DBs in a separate run — [Source](https://borgbackup.readthedocs.io/en/master/faq.html); 12.45 GB Vintage Story sqlite: 2nd backup 518 MB default → 62 MB (15,19,17) → 37 MB (10,23,16) — [Source](https://github.com/borgbackup/borg/issues/5877)
|
||||
- Kopia: 8 MB photo with metadata edit re-uploads ~8 MB under 4 MB splitter (chunk > file); fix is smaller splitter at repo-creation cost of more chunks — [Source](https://kopia.discourse.group/t/chunk-size-setting/1351)
|
||||
- Fixed-splitter warning: 10 GB video + 1 prepended byte re-uploads 10 GB; content-based (buzhash/rabinkarp) uploads 1–2 chunks — [Source](https://kopia.discourse.group/t/difference-between-the-available-splitters/894)
|
||||
- Backup pipelines must chunk-then-compress (per-chunk); encrypt-then-chunk/compress is useless since ciphertext is indistinguishable from random; encrypted dedup needs weakened convergent/MLE schemes with leakage tradeoffs — [Source](https://eprint.iacr.org/2025/532.pdf); survey — [Source](https://dl.acm.org/doi/10.1145/3685278)
|
||||
- Kopia compresses each chunk independently (s2 default, gzip optional); splitting costs little ratio (466 MB → 119 MB standalone s2 vs 133 MB via Kopia-s2) — [Source](https://kopia.io/docs/advanced/compression/)
|
||||
|
||||
### Inferences
|
||||
- Daemon policy: (1) never store `.gz/.zip/.jpg/.mp4/.sqlite.gz` deltas expecting CDC wins — keep 1 mirror copy or thin retention; (2) ingest DBs/VM images uncompressed and let CDC + per-chunk compression do the work; (3) exclude or separately-schedule >100 MB media with tiny chunkers only if churn is low.
|
||||
- `.dump`/uncompressed SQL text is the most dedup-friendly DB form but slowest to produce; page-image `.backup` is the balanced default.
|
||||
|
||||
### Gaps
|
||||
- Quantitative JPEG/MP4/ZIP re-storage ratios under Restic/Borg/Kopia CDC not found in primary docs; image-specific guidance (EXIF-only edits) relies on forum anecdotes.
|
||||
- Whether per-chunk compression (Kopia/Restic 0.16+) rescues any dedup on already-compressed inputs remains unmeasured in sources reviewed.
|
||||
@@ -0,0 +1,103 @@
|
||||
# Must-Include Data for Linux Workstation / Home Server Backups
|
||||
|
||||
## Which paths do guides universally say to include (~/Documents, ~/Pictures, ~/.config, project dirs, etc.)?
|
||||
|
||||
### Takeaway
|
||||
Guides converge on: back up all of `/home` (or whole `$HOME`) plus `/etc`, `/root`, `/var` (selectively) and `/opt`; on workstations this means XDG user dirs (`~/Documents`, `~/Pictures`, `~/Videos`, `~/Music`, `~/Downloads`), project dirs (`~/projects`, `~/src`), dotfiles/`~/.config`/`~/.local/share`, `~/.ssh`, browser profiles, mail, and package lists — never a bare `~/Documents`-only backup.
|
||||
|
||||
### Cited Findings
|
||||
- Red Hat guide lists as must-back-up: `/etc` (system config, users/groups, networking, app configs), `/home` (all user data/downloads/documents/pictures under username), `/root` (admin scripts/notes/configs), `/var` shared/corporate data — and one dir not to back up (implied transient/cache) — [Source](https://www.redhat.com/en/blog/backup-dirs)
|
||||
- Ubuntu official help: copy personal files/settings usually in Home folder; if room, back up entire Home folder with exceptions (cache/thrash-type excludes detailed on linked page) — [Source](https://help.ubuntu.com/stable/ubuntu-help/backup-how.html.en)
|
||||
- PCWorld Linux backup roundup: as a rule a regular backup of the home directories is sufficient; `rsync -avP $HOME` to USB disk given as sufficient home backup — [Source](https://www.pcworld.com/article/2050183/best-linux-backup-tools.html)
|
||||
- OSTechNix 2026 reinstall checklist: back up more than `~/Documents`; must include personal files + `~/.local/share` (app data, game saves, Flatpak/Snap data), `~/.config` (app settings), browser profiles (Firefox `~/.mozilla` or new XDG `~/.config/mozilla` + `~/.cache/mozilla` + `~/.local/share/mozilla`, check `about:support`; Chrome/Chromium), SSH/GPG keys, dotfiles, app-specific data, installed package lists including Flatpak/Snap, system-level `/etc`, NetworkManager `system-connections`, cron/scheduled tasks; trap 1 is only backing up `~/Documents ~/Pictures ~/Downloads` and missing `~/projects`, VM images, second drives; trap 2 is forgetting hidden XDG data — [Source](https://ostechnix.com/things-to-back-up-before-reinstalling-linux)
|
||||
- OSTechNix: single `rsync` of whole `$HOME` captures dotfiles, browser profiles, SSH keys, app settings since all live under `$HOME` — [Source](https://ostechnix.com/things-to-back-up-before-reinstalling-linux)
|
||||
- Arch forum classic tar set: `tar zcvfp arch-system.gz /etc /boot /root` + per-user `/home/user1` + `/var --exclude /var/cache/pacman/pkg` — [Source](https://bbs.archlinux.org/viewtopic.php?id=83533)
|
||||
- Production borg example backs up `/home /etc /var/www /var/backups /opt /root` together with `--exclude-from` file — [Source](https://cubepath.com/docs/Backup%20Recovery/backup-with-borgbackup-deduplication)
|
||||
- Ubuntu Ask-Ubuntu full-system tar pattern excludes virtual/external `dev mnt proc sys` (and squashfs variant excludes `home media dev run mnt proc sys tmp`, then backs up `/home` separately excluding cloud mirrors like `Dropbox GoogleDrive`) — [Source](https://askubuntu.com/questions/7809/how-to-back-up-my-entire-system)
|
||||
- Linux Mint forum consensus: Timeshift = OS/system restore points, does nothing for data in `/home`; need separate file-level tool (BackInTime, FreeFileSync, Foxclone/Rescuezilla/Clonezilla for images) for Documents/Music/Pictures — [Source](https://forums.linuxmint.com/viewtopic.php?t=405449)
|
||||
- Dotfiles canon: `~/.bashrc`, `~/.bash_profile`, `~/.zshrc`, `~/.vimrc`, `~/.gitconfig`, `~/.ssh/config`, `~/.tmux.conf`, `~/.config/` XDG dir; secrets in `~/.netrc`, `~/.aws/credentials`, `~/.ssh/`, history `~/.bash_history` must be treated as sensitive / kept out of public repos — [Source](https://linuxcommandlibrary.com/man/dotfiles)
|
||||
- Dotfiles restore guides list as backup-worthy: `.ssh/config`, `.gnupg/pubring.kbx`, `.gnupg/trustdb.gpg`, plus stow-managed `zsh/git/vim/tmux` and XDG select dirs — [Source](https://github.com/RickCogley/dotfiles/blob/main/docs/how-to/backup-restore.md)
|
||||
- Keeply Linux tool scope explicitly: dotfiles, `.config` dirs, browser profiles, app settings, optionally SSH keys, emails — [Source](https://github.com/cozy533/keeply)
|
||||
|
||||
### Inferences
|
||||
- Universal must-include = whole `/home/<user>` + `/etc` + package list; `/root` and `/var/lib`-style app data added for home-server role.
|
||||
- `~/Documents`-only backup is consistently called out as insufficient; project dirs and XDG hidden dirs hold irreplaceable state.
|
||||
- SSH private keys and Wi-Fi `psk=` files are must-include but must be encrypted, not pushed to public dotfiles repos.
|
||||
|
||||
### Gaps
|
||||
- No single 2023–2026 primary guide found quantifying mail (`~/Mail`, Thunderbird `~/.thunderbird`) include rates; only secondary tool scope mentions emails.
|
||||
- No reliable source found prescribing exact treatment of `~/.cache` vs `~/.local/share` boundary for Flatpak/Snap beyond OSTechNix summary.
|
||||
|
||||
## How should small databases (sqlite), archives (.zip/.tar), and images be treated — raw files or dumps?
|
||||
|
||||
### Takeaway
|
||||
Small static files (photos/images, `.zip/.tar` archives) are backed up as raw files and deduplicate well; live database files (sqlite, postgres/mysql) must be dumped (`VACUUM INTO`, `sqlite3 .backup`/backup API, `pg_dump`/`pg_dumpall`/`mysqldump`) before file backup — raw copy of a live DB is unsafe.
|
||||
|
||||
### Cited Findings
|
||||
- SQLite docs: safe live-copy methods are `sqlite3_rsync` (3.47.0+, 2024-10-21, bandwidth-efficient over SSH), `VACUUM INTO filename`, or backup API; plain file copy only safe with no transactions in progress, and if prior write failed must copy `-journal`/`-wal` together — [Source](https://sqlite.org/howtocorrupt.html)
|
||||
- SQLite forum (Borg user, WAL-mode pooled connections): cannot depend on perfect backup from open/changing DB because backup doesn't snapshot DB + journal in same instant; recommendation: close writers or backup API copy then back up the copy; `rm -f mydb.sqlite3.backup; sqlite3 mydb.sqlite3 "VACUUM INTO 'mydb.sqlite3.backup'"` added to backup script — [Source](https://sqlite.org/forum/forumpost/f173a78e2e?t=h)
|
||||
- Experiment on SQLite 3.50.6 WAL DB (1000 rows, 500 in `-wal`): `cp app.db backup_cp.db` yielded 500 rows, `integrity_check` still `ok` — stale copy passes checks silently; copying `app.db + app.db-wal + app.db-shm` together preserved 1000 rows; `.backup` (Online Backup API) copies page-by-page including WAL and restarts if source writes mid-copy — [Source](https://www.sqlprostudio.com/blog/68-how-to-safely-back-up-or-copy-a-live-sqlite-database)
|
||||
- `fs.copyFile()` on WAL DB copies only `.db` while `-wal/-shm` still written → instant `SQLITE_CORRUPT`; 7/7 rotating backups corrupted; fix is better-sqlite3 `.backup()` which handles all three files atomically; test restores — [Source](https://scottspence.com/posts/sqlite-corruption-fs-copyfile-issue)
|
||||
- Postgres docs: `pg_dump dbname > dumpfile` generates SQL commands to recreate DB; `pg_dumpall > dumpfile` preserves cluster-wide roles/tablespaces, restore via `psql -f dumpfile postgres` requiring superuser; file-level/WAL archiving is version-specific whereas dumps reload into newer versions — [Source](https://www.postgresql.org/docs/%EF%BC%99.6/backup-dump.html)
|
||||
- Postgres professional docs: `pg_dump` can emit text or archive formats for parallelism/fine-grained `pg_restore` control — [Source](https://postgrespro.com/docs/postgresql/17/backup-dump)
|
||||
- OSTechNix: don't copy raw data dir; use each DB's dump tool e.g. `mysqldump -u <user> -p <db> > backup.sql` — [Source](https://ostechnix.com/things-to-back-up-before-reinstalling-linux)
|
||||
- Borg quickstart warns: snapshot filesystems/volumes (LVM/ZFS useful), dump databases or stop DB servers, shut down VMs/containers before backing up disk images/volumes — [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)
|
||||
- Restic docs example: `plan.txt` change adds KiB, `archive.tar.gz` change re-adds MiB — illustrates large monolithic archives re-chunk poorly vs small files; still backed up as raw files — [Source](https://github.com/restic/restic/blob/master/doc/040_backup.rst)
|
||||
|
||||
### Inferences
|
||||
- Photos/images/archives: raw-file backup is correct; no dump needed; chunked dedup tools (restic/borg/kopia) handle them but modified tar/zip re-stores large chunks.
|
||||
- SQLite/postgres/mysql live files: dump-then-backup is mandatory; raw file backup alone is at best crash-consistent, at worst silently stale/corrupt.
|
||||
|
||||
### Gaps
|
||||
- No primary 2023–2026 benchmark found for dedup efficiency on `.zip/.tar` vs extracted trees under restic/borg/kopia; only illustrative restic example.
|
||||
- No reliable source prescribing whether to keep both live sqlite file + dump in same backup set or dump-only.
|
||||
|
||||
## What do tools like restic, borg, kopia, Time Machine include by default?
|
||||
|
||||
### Takeaway
|
||||
None auto-selects `~/Documents`/`~/.config`; all default to exactly what paths you pass (no implicit includes) and provide opt-in excludes (`--exclude*`, `--patterns-from`, `.kopiaignore`/policy, Time Machine StdExclusions + user list) for caches, build artifacts, and system pseudo-filesystems.
|
||||
|
||||
### Cited Findings
|
||||
- Restic `backup [FILE/DIR]...` creates snapshot of exactly given args; no default source; exclude options: `--exclude/--iexclude`, `--exclude-file`, `--exclude-caches` (CACHEDIR.TAG), `--exclude-if-present foo`, `--exclude-larger-than`, `--exclude-cloud-files` (Win/macOS OneDrive/iCloud only); excludes don't apply to explicitly passed file path itself, only contents under dirs — [Source](https://restic.readthedocs.io/en/stable/040_backup.html)
|
||||
- Restic excludes use Go `filepath.Match` against full path, gitignore-like: once dir excluded can't re-include inside; example backs up selection inside `$HOME` — [Source](https://restic.readthedocs.io/en/v0.16.3/040_backup.html)
|
||||
- Restic handles `--one-file-system/-x` to not cross filesystem boundaries/subvolumes; saves/restores ACLs/xattrs; sparse holes deduped/compressed not stored explicitly — [Source](https://github.com/restic/restic/blob/master/doc/manual_rest.rst)
|
||||
- NixOS restic module example: `paths = ["/home"]` with `exclude = ["/home/*/.cache" ".git"]` pattern as conventional home selection — [Source](https://github.com/NixOS/nixpkgs/blob/master/nixos/modules/services/backup/restic.nix)
|
||||
- Borg `create repo::archive ~/Documents`, `~/Documents ~/src --exclude '*.pyc'`, `/home --exclude thumbnail regex`, ` / --one-file-system` are docs examples — user picks roots; no default root — [Source](https://borgbackup.readthedocs.io/en/1.0.5/usage.html)
|
||||
- Borg patterns: fnmatch default for `--exclude`, shell-style for `--pattern`; `sh:**/steamapps/common/**`, `sh:home/user/.cache/**`, trailing-slash `some/path/` keeps dir not contents vs no-slash excludes both; `--exclude-from`, `--patterns-from`, `--exclude-caches` (CACHEDIR.TAG), `--exclude-if-present`, `--keep-tag-files`, `--one-file-system`, nodump flag respected — [Source](https://man.archlinux.org/man/borg-patterns.1.txt)
|
||||
- Kopia: no default source; snapshots what you `snapshot create`; ignores via `.kopiaignore` (default file), global/per-source policy `--add-ignore/--add-dot-ignore`, `--ignore-cache-dirs true` (inherited global), `--ignore-dir-errors/--ignore-file-errors`, `--one-file-system`, never/only-compress lists; before/after folder/root actions for dumps/snapshots with timeout/modes — [Source](https://kopia.io/docs/reference/command-line/common/policy-set/)
|
||||
- Kopia ignore syntax: `#` comment, `!` negate, `*`/`**`/`?`/`[0-9a-zA-Z]`/`[abc]`, leading `/` root-only; examples `*.dat`, `/logs/*`, `tmp.db`, `**/logs/**`; must test — incomplete snapshots possible — [Source](https://kopia.io/docs/advanced/kopiaignore/)
|
||||
- Time Machine default backs up everything except system files/apps from macOS install, caches, StdExclusions (`/System/Library/CoreServices/backupd.bundle/Contents/Resources/StdExclusions.plist`), backup disk itself; user adds via Options → Exclude; excluded items still in local snapshots; damaged system requires macOS reinstall + Migration Assistant — [Source](https://support.apple.com/en-ca/guide/mac-help/mh15622/mac)
|
||||
- Time Machine `tmutil isexcluded/addexclusion/removeexclusion`, sticky vs fixed-path (`-p`) exclusions semantics — [Source](https://www.unix.com/man-page/osx/8/TMUTIL)
|
||||
- ArchWiki System backup: no default set; methods are btrfs/LVM snapshots, rsync, tar, SquashFS (no ACLs); recommends 3-2-1, regular integrity + restore tests; automation via systemd timer/cron with least-privilege `CAP_DAC_READ_SEARCH` example — [Source](https://wiki.archlinux.org/title/System_backup)
|
||||
|
||||
### Inferences
|
||||
- Versioning tools scope = explicit roots + exclude list; portable Linux default is effectively `/home + /etc + dumps` minus `*.cache/CACHEDIR.TAG`, `node_modules/target/.git`, VM images while running.
|
||||
- Kopia `ignore-cache-dirs=true` and restic/borg `--exclude-caches` are the closest to a built-in default exclude.
|
||||
|
||||
### Gaps
|
||||
- No primary source found stating Kopia ships global ignore rules beyond `ignore-cache-dirs`; default policy contents not fully enumerated in fetched docs.
|
||||
- Time Machine StdExclusions full list not fetched (plist path only); Linux analogue must be inferred.
|
||||
|
||||
## Any special handling for live database files (WAL mode, locking, dump-then-backup)?
|
||||
|
||||
### Takeaway
|
||||
Never file-copy a live DB; quiesce (close writers/stop server), filesystem/LVM/ZFS snapshot, or dump (`sqlite3 backup API/VACUUM INTO`, `pg_dumpall`, `mysqldump`); for SQLite WAL must keep `.db+-wal+-shm/-journal` together, never delete hot journals, beware POSIX `close()` dropping advisory locks and fork/link/rename hazards.
|
||||
|
||||
### Cited Findings
|
||||
- SQLite: backup/restore while transaction active → copy mixes old/new → corrupt; safe via `sqlite3_rsync`, `VACUUM INTO`, backup API even on live DB; idle-file copy only safe with no tx in progress — [Source](https://sqlite.org/howtocorrupt.html)
|
||||
- SQLite: never move/delete/rename hot `-journal/-wal` after crash; mispairing (swap/overwrite/move journal, copy DB without journal, overwrite DB without deleting hot journal) likely corrupts; quiescent DB has no journal, only DB file matters — [Source](https://sqlite.org/howtocorrupt.html)
|
||||
- SQLite POSIX advisory-lock quirk: any thread `open/read/close` on DB file drops locks for all threads (close cancels locks); bypassing lib for backup read can corrupt; since 3.51.0 (2025-11-04) extra WAL defenses but not cure-all — never `close()` DB file while connections open even in other threads — [Source](https://sqlite.org/howtocorrupt.html)
|
||||
- SQLite: don't use two SQLite copies linked in one app (separate lock lists), don't mix locking protocols (POSIX vs dot-file/NFS), don't unlink/rename open DB (shared journal name → cross-recovery corruption, `SQLITE_WARNING` since 3.7.17), don't multi-link/symlink same file (wrong journal lookup; canonicalization since 3.10.0), don't carry connection across `fork()` — [Source](https://sqlite.org/howtocorrupt.html)
|
||||
- SQLite WAL forgiving of out-of-order writes except during checkpoint; `COMMIT` sync failure loses durability not consistency; checkpoint infrequently as defense; QNX `mmap` + WAL needs exclusive locking/no-mmap — [Source](https://sqlite.org/howtocorrupt.html)
|
||||
- Forum: with WAL expect journal replay on open after restore, uncommitted lost; for professional frozen-moment copy don't back up while changing; full WAL checkpoint then copy main file pristine if no checkpoint during copy (disable autocheckpoint temporarily) — [Source](https://sqlite.org/forum/forumpost/f173a78e2e?t=h)
|
||||
- Borg docs: avoid programs changing files during backup; LVM/ZFS snapshot or dump/stop DBs; shutdown VMs/containers first — [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)
|
||||
- Kopia actions support `before-folder/after-folder` and `before-snapshot-root/after-snapshot-root` scripts (e.g. `zfs snapshot` + mount then snapshot mount, `zfs destroy` after; `pg_dumpall`/sqlite dump) with `essential/optional/async` modes and timeout (default 5m) — [Source](https://github.com/kopia/kopia/blob/master/site/content/docs/Advanced/Actions/_index.md)
|
||||
- SafeKeep (rdiff-backup wrapper) noted as integrating with LVM and databases for consistent backups — [Source](https://wiki.archlinux.org/title/Synchronization_and_backup_programs)
|
||||
|
||||
### Inferences
|
||||
- Correct pattern for single-user Linux: `pre-action: pg_dumpall/sqlite VACUUM INTO/.backup → file under /home or /var/backups` then `borg/restic/kopia snapshot` includes dump; keep live DB files too but rely on dump for restore.
|
||||
- Filesystem snapshot (btrfs/LVM/ZFS) gives crash-consistent base but DB dump still needed for logical restore across versions.
|
||||
|
||||
### Gaps
|
||||
- No 2026 primary guide found for `sqlite3_rsync` adoption in borg/restic/kopia hooks; only SQLite 3.47.0+ release note reference.
|
||||
- No reliable latency/lock Hold-time guidance for `VACUUM INTO` vs backup API on large WAL DBs under active writers.
|
||||
@@ -0,0 +1,102 @@
|
||||
# Never Include in Linux Workstation/Home-Server Backups
|
||||
|
||||
## What are the canonical exclude directories and file patterns across backup tools (give concrete lists)?
|
||||
|
||||
### Takeaway
|
||||
Canonical Linux file-backup excludes converge on: virtual/kernel filesystems (`/proc`, `/sys`, `/dev`, `/run`), ephemeral (`/tmp`, `/var/tmp`, `/var/run`, `/var/lock`), mounts/media (`/mnt`, `/media`, `/lost+found`, `/swapfile`), package/cache/logs (`/var/cache/*`, `/var/log/*`, pacman cache), per-user caches/trash/thumbnails, and repro data (build/dependency dirs, VM/container storage) — implemented as explicit exclude-files in restic/borg plus `--exclude-caches`/`--one-file-system` and Kopia policies/`.kopiaignore`.
|
||||
|
||||
### Cited Findings
|
||||
- restic has no built-in default excludes; user supplies `--exclude`, `--iexclude`, `--exclude-file`, `--exclude-if-present foo`, `--exclude-caches` (CACHEDIR.TAG dirs), `--exclude-larger-than`, `--exclude-cloud-files`, plus `-x/--one-file-system` to stay on one filesystem — [Source](https://restic.readthedocs.io/en/v0.16.3/040_backup.html); also documented that excludes do not apply to explicitly-passed backup-source paths — [Source](https://github.com/restic/restic/blob/master/doc/040_backup.rst)
|
||||
- restic pattern syntax is Go `filepath.Match` + `**` for crossing `/`, matched on complete path components (`foo` matches `/dir1/foo/...` but not `/dir/foobar`), trailing `/` ignored, leading `/` anchors at root — [Source](https://restic.readthedocs.io/en/v0.16.3/040_backup.html)
|
||||
- borg `create` supports `-e/--exclude PATTERN`, `--exclude-from FILE`, `--exclude-caches` (CACHEDIR.TAG), `--exclude-if-present NAME`, `--keep-tag-files`, `-x/--one-file-system`, and respects `nodump` flag; recommended test is `borg create --list --dry-run` — [Source](https://borgbackup.readthedocs.io/en/1.0.5/usage.html); pattern styles are `fnmatch` (default for `--exclude`), `sh:`, `re:`, path prefix/full-match, with trailing-`/` meaning “keep dir, skip contents” — [Source](https://manpages.debian.org/testing/borgbackup/borg-patterns.1.en.html)
|
||||
- borg quickstart automation example excludes `--exclude-caches --exclude 'home/*/.cache/*' --exclude 'var/tmp/*'` when backing up `/etc`-style roots — [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)
|
||||
- Kopia has no static exclude-file; it uses policy `Ignore rules` + `Read ignore rules from files` (default `.kopiaignore`) + `Ignore cache directories: true` (default true, inherits global→host→user→path) + `Scan one filesystem only: true` — [Source](https://github.com/kopia/kopia/issues/3334); `kopia policy set` flags include `--add-ignore`, `--add-dot-ignore`, `--ignore-cache-dirs [true|false|inherit]`, `--ignore-dir-errors`, `--one-file-system`-equivalent `Scan one filesystem only` — [Source](https://kopia.io/docs/reference/command-line/common/policy-set/)
|
||||
- Kopia’s CACHEDIR.TAG-equivalent (`--ignore-cache-dirs`, default true) was added to mirror restic `--exclude-caches`, defaulting to ignore caches via policy — [Source](https://github.com/kopia/kopia/issues/564)
|
||||
- ArchWiki restic full-system `/etc/restic/excludes.txt` canonical list: `/data/**`, `/dev/**`, `/home/*/**/*.pyc`, `/home/*/**/__pycache__/**`, `/home/*/**/node_modules/**`, `/home/*/.cache/**`, `/home/*/.local/lib/python*/site-packages/**`, `/home/*/.mozilla/firefox/*/Cache/**`, `/lost+found/**`, `/media/**`, `/mnt/**`, `/proc/**`, `/root/**`, `/run/**`, `/swapfile`, `/sys/**`, `/tmp/**`, `/var/cache/**`, `/var/cache/pacman/pkg/**`, `/var/lib/docker/**`, `/var/lib/libvirt/**`, `/var/lock/**`, `/var/log/**`, `/var/run/**`, noting `--one-file-system` can replace `/proc`/`/run`/`/mnt` entries while preserving mountpoints — [Source](https://wiki.archlinux.org/title/Restic)
|
||||
- Debian restic-forum practitioner system list: `/media`, `/mnt`, `/cdrom`, `/proc`, `/sys`, `/dev`, `/run`, `/tmp`, `/var/run`, `/var/lock`, `/var/tmp`, `/lost+found`, `/swapfile`, `/var/cache/restic`, Steam `.../Steam/steamapps`, plus home temp items `.gvfs`, `.local/share/gvfs-metadata`, `.local/share/Trash`, `.cache`, `.dbus`, `.xsession-errors`, `.Xauthority`, `.gksu.lock`, `.local/share/flatpak/appstream` — [Source](https://forum.restic.net/t/what-gnu-linux-directories-to-exclude-from-backups/6653)
|
||||
- Extended community restic list adds: `nobackup/NoBackup`, `Downloads/*`, `VirtualBox VMs/`, `Dropbox/*`, `rclone/`, `snap/`, `.var/**/cache/`, `.local/share/docker/`, `.local/share/JetBrains/`, `.vscode/extensions/`, `.config/Code/`, `.m2/repository`, `target/*`, `build/*`, `node_modules/*`, `jspm_packages/*`, `web_modules/*`, `.npm`, `.coursier/cache`, `.sbt`, `.stack`, `.sdkman`, `.jdks`, `.eclipse`, `.venv-py3`, `.wine`, `.android`, `Android/Sdk`, `.gradle`, `.adobe`, `.macromedia`, `.thumbnails`, `.thunderbird/*/Cache`, `.mozilla/firefox/*/Cache|storage|minidumps|*.sqlite*`, `.config/**/Cache|GPUCache|ShaderCache`, `.config/chromium/Default/...History|Favicons|Storage|Cache`, `.local/share/baloo|zeitgeist|akonadi`, `.gnupg/rnd|random_seed|*.lock`, `.pulse*`, `.java/deployment/cache`, `.dropbox*` — [Source](https://forum.restic.net/t/what-gnu-linux-directories-to-exclude-from-backups/6653)
|
||||
- ArchWiki tar full-system guidance: exclude `/opt/backup/arch-full*` (own backups), `/tmp/*`, `/var/cache/pacman/pkg/`, boot from LiveCD with plain `chroot` (not `arch-chroot`) to avoid capturing tempfs/memory — [Source](https://wiki.archlinux.org/title/Full_system_backup_with_tar_(Italiano))
|
||||
- Guide summary: when backing up `/`, explicitly exclude `/proc`, `/sys`, `/dev`, `/run`, `/tmp` and bind/overlay mounts; when backing up only `/home`+`/etc` both borg/restic skip virtual FS by default — [Source](https://linuxjunkies.org/guides/back-up-with-borg-or-restic)
|
||||
- Kopia `.kopiaignore` syntax: one rule/line, `#` comment, `*` (any chars), `**` (any dirs), `?`, `[0-9]/[a-z]/[A-Z]/[abc]`, leading `/` roots rule, `!` negates; policy can point at alternate ignore-files — [Source](https://kopia.io/docs/advanced/kopiaignore/)
|
||||
- Time-Machine-ecosystem analogue (Arq honoring Apple TM exclusions): `.Trash`, `Library/Caches`, `Library/Logs`, `Library/Mail/.../Envelope Index*`, `Library/Safari/WebpageIcons.db`, `Library/Saved Application State`, `Library/iTunes/iPad Software Updates` — [Source](https://www.arqbackup.com/docs/arqbackup/pages/adding_folder.html)
|
||||
|
||||
### Inferences
|
||||
- There is no single vendor “default exclude file” for restic/borg on Linux; the de-facto standard is the ArchWiki + forum lists above plus `--exclude-caches` and `--one-file-system`.
|
||||
- Kopia inverts the model (opt-out caches by default + per-dir `.kopiaignore`) vs restic/borg (opt-in excludes), so migrated exclude-lists must be translated, not copied verbatim.
|
||||
|
||||
### Gaps
|
||||
- No authoritative 2023–2026 duplicity built-in exclusion list found for Linux workstation context; duplicity relies on `--exclude` user patterns rather than published defaults.
|
||||
- Exact current Apple `StdExclusions.plist` contents not verified for Linux relevance; only Arq’s documented subset was citable.
|
||||
|
||||
## Which exclusions exist for correctness (locked files, sockets) vs size vs reproducibility?
|
||||
|
||||
### Takeaway
|
||||
Correctness excludes (virtual FS, sockets/FIFOs/devices, live DB/VM/container backing stores, cloud-online-only stubs) prevent hangs, errors, or unrestorable data; size excludes (caches, trash, thumbnails, logs, browser profiles, package caches, media/Steam) prevent bloat; reproducibility excludes (dependency/build trees) are safe to drop because lockfiles+manifests rebuild them.
|
||||
|
||||
### Cited Findings
|
||||
- `/proc` and `/sys` are virtual/kernel pseudo-filesystems (proc exposes live kernel/process state, `proc_sys` exposes tunable sysctls; size often reported 0, constantly changing) — backing them up captures no stable data — [Source](https://docs.redhat.com/en/documentation/red_hat_enterprise_linux/4/html/reference_guide/ch-proc); [Source](https://man7.org/linux/man-pages/man5/proc_sys.5.html)
|
||||
- restic forum guidance: use `--exclude-caches` to auto-exclude CACHEDIR.TAG-marked dirs for “non permanent system files”; full-system threads treat only a short path list as needing exclusion — [Source](https://forum.restic.net/t/linux-exclusion-of-non-permanent-system-files/6629)
|
||||
- borg by default does not open/read block/char devices or FIFOs; `--read-special` is required to force reading them as regular files (and to follow symlinks to them) — [Source](https://www.systutorials.com/linux-manual-page-1-borg)
|
||||
- restic 0.17.0+ fix “Exclude irregular files from backups” (sockets, FIFOs, devices skipped) confirms prior irregular-file handling was a correctness fix — [Source](https://github.com/restic/restic/releases)
|
||||
- Kopia policy has explicit `--ignore-dir-errors` to tolerate unreadable/locked dirs during traversal, separate from ignore-rules; issue reports show `fuse` mounts (e.g. `seafile/fuse`) still `lstat`-probed during estimation even when ignored, motivating explicit excludes for FUSE/locked paths — [Source](https://github.com/kopia/kopia/issues/3334); [Source](https://kopia.io/docs/reference/command-line/common/policy-set/)
|
||||
- Docker storage (`/var/lib/docker`: `overlay2/`, `image/`, `buildkit/`, `volumes/`, `containers/`) is large, churn-heavy, overlay-mounted and often locked/inconsistent without daemon quiesce; community Arch/restic lists therefore exclude `/var/lib/docker/**` and `/var/lib/libvirt/**`, and Proxmox `vzdump --exclude-path` examples exclude `/var/lib/docker/` alongside `/run/`, `/dev/shm`, `/dev/fuse` — [Source](https://wiki.archlinux.org/title/Restic); [Source](https://blog.devops.dev/docker-cleanup-and-relocate-81d2dcd32b8a); [Source](https://forum.proxmox.com/threads/backup-and-exclude-path.125966)
|
||||
- Docker’s own backup guidance says file-copy of the VM disk (`Docker.raw`/`docker_data.vhdx`) or `/var/lib/docker` requires Docker fully stopped; otherwise use `docker save`/`docker pull` + volume dump/restore procedures — [Source](https://docs.docker.com/desktop/settings-and-maintenance/backup-and-restore.md)
|
||||
- Veeam (VM-image backup reference) automatically excludes VM log files to cut size/time, and supports excluding swap files, deleted-block (BitLooker) data, and per-disk/template excludes for the same size/correctness reasons — [Source](https://helpcenter.veeam.com/docs/vbr/userguide/data_exclusion.html)
|
||||
- Size-category examples: pacman cache `/var/cache/pacman/pkg/`, `/var/cache/*`, `/var/log/*`, `/var/tmp/*`, browser `Cache/GPUCache/ShaderCache`, `thunderbird/*/Cache`, `~/.thumbnails`, `~/.local/share/Trash`, Steam `steamapps`, `~/Downloads/*`, Dropbox/rclone replicas — all excluded for bloat, not correctness — [Source](https://forum.restic.net/t/what-gnu-linux-directories-to-exclude-from-backups/6653); [Source](https://wiki.archlinux.org/title/Restic)
|
||||
- Reproducibility-category examples: `node_modules/*`, `jspm_packages/*`, `web_modules/*`, `target/*`, `build/*`, `~/.m2/repository`, `~/.ivy2`, `~/.gradle`, `~/.coursier/cache`, `~/.sbt`, `~/.stack`, `__pycache__`, `*.pyc`, `~/.local/lib/python*/site-packages/**`, `~/.npm`, `~/.pkg-cache/`, `~/.sdkman/`, `.venv-py3` — explicitly listed as rebuildable caches/toolchains — [Source](https://forum.restic.net/t/what-gnu-linux-directories-to-exclude-from-backups/6653); webdev backup tooling states “node_modules is never included … rebuild with yarn/npm install” while always keeping `.env`, `yarn.lock`, `package-lock.json` — [Source](https://github.com/ICJIA/backup-webdev/blob/main/README.md)
|
||||
|
||||
### Inferences
|
||||
- Rule of thumb: if path is kernel-generated (`/proc`, `/sys`), socket/FIFO/device, actively memory-mapped/locked (browser `*-shm/-wal`, DB live files, running container layers, FUSE), or a mountpoint for another filesystem, exclusion is for correctness; if it is cache/log/trash/thumbnail/package-download, it is for size; if `npm install`/`pip install`/`cargo fetch`/`mvn` recreates it, it is for reproducibility.
|
||||
- Live DB/VM/container volumes need application-consistent dump/snapshot (e.g. `docker save`, DB dump, hypervisor snapshot), not file-level copy of the running store.
|
||||
|
||||
### Gaps
|
||||
- No citable 2024–2026 benchmark found quantifying restore-corruption rate when file-copying live SQLite/WAL or overlay2 without quiesce; guidance remains consensus-based.
|
||||
- restic/borg behavior on Windows/macOS cloud-online-only stubs (`--exclude-cloud-files`) verified in flags but Linux relevance is none; no Linux equivalent stub mechanism found.
|
||||
|
||||
## How do tools treat dot-directories (.cache, .git, .venv) and per-project gitignore?
|
||||
|
||||
### Takeaway
|
||||
Backup tools do not read `.gitignore` by default; they rely on explicit patterns, CACHEDIR.TAG/`--exclude-caches`, and `.kopiaignore`/policy rules — so `.cache` is usually globally excluded, `.git` is usually kept (small, valuable history) unless deliberately dropped, and `.venv`/`node_modules`/`target` must be explicitly excluded per-project or via `exclude-if-present` markers.
|
||||
|
||||
### Cited Findings
|
||||
- `~/.cache` is safe to drop (name indicates cached data; many tools exclude cache/trash by default); Arch forum explicitly approves excluding whole `~/.cache`, implemented via borg `--exclude-caches` + CACHEDIR.TAG spec — [Source](https://bbs.archlinux.org/viewtopic.php?id=231281)
|
||||
- restic `--exclude-caches` only skips dirs containing `CACHEDIR.TAG` (keeps the tag file); `--exclude-if-present foo` skips contents of any dir containing marker `foo` (e.g. `.nobackup`, `CACHEDIR.TAG`, custom sentinels) — [Source](https://restic.readthedocs.io/en/v0.16.3/040_backup.html)
|
||||
- borg `--exclude-caches`/`--exclude-if-present`/`--keep-tag-files` semantics are identical (skip CACHEDIR.TAG dirs; optionally retain tag files) — [Source](https://borgbackup.readthedocs.io/en/1.0.5/usage.html)
|
||||
- Kopia `Ignore cache directories: true` is the default global policy (inherits unless overridden), plus per-directory `.kopiaignore` and policy ignore-rules using gitignore-like globs (`*.dat`, `/logs/*`, `tmp.db`, `**/logs/**`, `!` negation) — [Source](https://github.com/kopia/kopia/issues/3334); [Source](https://kopia.io/docs/advanced/kopiaignore/)
|
||||
- Kopia ignore-rule inheritance is surprising: defining child-path `add-ignore` can shadow global ignores (open issue: “Ignore rules not being inherited when child adds its own ignores”), so global `.cache`/`cache` rules must be repeated or merged carefully — [Source](https://github.com/kopia/kopia/issues/4155)
|
||||
- Developer-oriented backup/index defaults treat dot-dirs as noise to skip: Dank index defaults exclude `.git`, `.hg`, `.svn`, `.cache`, `.npm`, `.yarn`, `.venv`/`venv`, `.tox`, `.pytest_cache`, `__pycache__`, `.gradle`, `.m2`, `.cargo`, `.idea`, `.vscode`, `node_modules`, `target`, `dist/build/out` — [Source](https://danklinux.com/docs/1.4/danksearch/configuration); SmartBackup auto-skips `node_modules`, `venv`, `__pycache__`, `.git` for 10x speed/size — [Source](https://github.com/CodingWithMK/smartbackup_file-backup-automation)
|
||||
- Counter-practice for `.git`: default backup guidance keeps `.git` (history is irreplaceable, usually small vs `node_modules`); Backblaze-mac customization explicitly adds separate opt-in rules to drop `node_modules/` and `.git/` because neither is excluded by default — [Source](https://gist.github.com/nickcernis/bb4bd43a44efd73b87d857e29b1d5b96)
|
||||
- `tmexclude` watches the filesystem to continually re-apply Time-Machine exclusions for `node_modules`, `target`, etc., because new projects constantly recreate them — showing per-project `.gitignore` alone does not stop backup tools from capturing them — [Source](https://github.com/PhotonQuantum/tmexclude)
|
||||
- Kopia FAQ states ignored paths come from policy or `.kopiaignore` files, not from `.gitignore` — [Source](https://github.com/kopia/kopia/blob/master/site/content/docs/FAQs/_index.md)
|
||||
|
||||
### Inferences
|
||||
- Best practice: global-exclude `~/.cache`, `~/.npm`, `~/.cargo`, `~/.mozilla/.../Cache`, `Trash`, `thumbnails`; per-source exclude `node_modules/`, `.venv/venv/`, `__pycache__/`, `target/`, `build/dist/`, optionally via `exclude-if-present` markers (e.g. drop marker file `CACHEDIR.TAG` or `.nobackup` in roots) rather than hoping tools honor `.gitignore`.
|
||||
- Keep `.git/` and lockfiles (`package-lock.json`, `yarn.lock`, `Cargo.lock`, `requirements*.txt`) while dropping the fetched trees; keep `.env` only deliberately (secret-sprawl risk vs rebuild need).
|
||||
|
||||
### Gaps
|
||||
- No evidence found that restic/borg/Kopia natively ingest `.gitignore` in 2026 builds; if any wrapper does, it was not in primary docs searched.
|
||||
- Best handling of `.venv` that contains non-recreatable local edits (pip `-e` installs, manual patches) is unresolved — exclusion assumes venv is truly disposable.
|
||||
|
||||
## What breaks when people back up dependency/build trees anyway?
|
||||
|
||||
### Takeaway
|
||||
Backing up `node_modules`, `venv`, `target/build`, `site-packages`, and container/VM stores inflates size, scan time, and snapshot churn, defeats deduplication (thousands of tiny files, hardlinks, changing mtimes/hashes), risks unrestorable or inconsistent restores (native bindings, symlinks, absolute paths, locked DB/pages), and can leak secrets or break privacy.
|
||||
|
||||
### Cited Findings
|
||||
- SmartBackup’s premise is that naive `Documents` backup “waits hours because of massive node_modules or venvs”; skipping them yields ~10x faster/smaller backups, with incremental manifest+hash tracking otherwise churning on every dependency touch — [Source](https://github.com/CodingWithMK/smartbackup_file-backup-automation)
|
||||
- Webdev backup defaults state `node_modules` excluded from full/incremental/differential/quick backups to save space; restore requires `yarn/npm install`, while lockfiles+`.env` are retained to reconstruct — [Source](https://github.com/ICJIA/backup-webdev/blob/main/README.md)
|
||||
- Rust/JS `target/` and `node_modules/` are the canonical Time-Machine-bloat examples requiring a watcher (`tmexclude`) because they reappear per `npm install`/`cargo build` and would otherwise be re-captured every snapshot — [Source](https://github.com/PhotonQuantum/tmexclude)
|
||||
- Python `site-packages` (`~/.local/lib/python*/site-packages/**`), `__pycache__`, `*.pyc`, and per-project `.venv` are listed alongside `node_modules` in Arch/restic excludes as regenerable interpreter artifacts — [Source](https://wiki.archlinux.org/title/Restic)
|
||||
- `pnpm clean/purge` exists precisely because `node_modules` contents (plus virtual-store) are disposable and safely removable via Node-aware deletion handling junctions correctly — [Source](https://pnpm.io/next/cli/clean)
|
||||
- File-level backup of `/var/lib/docker` captures `overlay2` diffs, `image/`, `buildkit` cache (often 10s of GB, e.g. 20GB build cache reclaimable via `docker system prune`) that are host-specific, layer-duplicated, and unrestorable by plain copy while daemon runs; correct path is `docker save/load`, registry push/pull, and volume dumps — [Source](https://blog.devops.dev/docker-cleanup-and-relocate-81d2dcd32b8a); [Source](https://docs.docker.com/desktop/settings-and-maintenance/backup-and-restore.md)
|
||||
- Locked/inconsistent captures produce backup errors or silent corruption: Kopia `seafile/fuse` example shows even explicitly ignored FUSE paths trigger `lstat` errors/failed snapshots unless fully avoided; `--ignore-dir-errors` merely masks valid errors — [Source](https://github.com/kopia/kopia/issues/3334)
|
||||
- Browser/IDE heavy dirs (`.config/Code/`, `.vscode/extensions/`, `.config/coc/extensions/`, `.mozilla/.../extensions`, `sonarlint/plugins`, `.coursier/cache`, `.m2/repository`) are called out as “heavy JARs/caches” that bloat snapshots while preferences should sync via cloud/accounts instead — [Source](https://forum.restic.net/t/what-gnu-linux-directories-to-exclude-from-backups/6653)
|
||||
- `~/.cache`-style trees and `node_modules` contain huge small-file counts that slow file-walk, inflate metadata/index, and reduce dedup/compression efficiency vs backing up manifests/lockfiles — rationale behind Dank/SmartBackup default-skips — [Source](https://danklinux.com/docs/1.4/danksearch/configuration); [Source](https://github.com/CodingWithMK/smartbackup_file-backup-automation)
|
||||
|
||||
### Inferences
|
||||
- Failure modes to expect if you include them anyway: (1) backup never finishes or blows retention/bandwidth; (2) every `npm/cargo/pip` run creates a large new snapshot despite no user-data change; (3) restore on another machine/arch breaks native modules/symlinks/permissions; (4) secrets in `.npmrc`, venv activate scripts, or vendored `.env` leak into long-retained snapshots.
|
||||
- Mitigation if you must include one tree (e.g. offline/air-gapped rebuild): snapshot it to a separate low-frequency archive/tag, not the daily home backup, and document the toolchain version needed to use it.
|
||||
|
||||
### Gaps
|
||||
- No controlled measurement found of dedup-ratio collapse or file-count slowdown specifically for `node_modules` vs `target/` vs `.venv` under borg/restic/Kopia 2026 versions; claims remain qualitative (“10x”, “hours”).
|
||||
- Privacy impact of vendored credentials inside dependency trees (e.g. `.pypirc`, npm tokens in `~/.npm/_cacache`) is asserted in practice guides but no incident stat was found.
|
||||
@@ -0,0 +1,100 @@
|
||||
# Quota-Driven Retention
|
||||
|
||||
## Which systems trigger retention from disk usage (e.g. "keep usage below X%", "delete oldest when over Y%"), and what exact thresholds do they use?
|
||||
|
||||
### Takeaway
|
||||
Only Apple Time Machine makes capacity the primary retention driver ("delete oldest when full"); Veeam, S3, Hetzner Storage Box, ZFS/sanoid and Borg/restic all use time/count-based retention as primary with capacity handled by monitoring, placement, or manual pruning — no "keep usage below X%" knob exists in most of them, and where percentage thresholds exist they are health-warning or performance floors (ZFS 20%/10%), not auto-delete triggers.
|
||||
|
||||
### Cited Findings
|
||||
- Time Machine "as your backup disk fills up, Time Machine deletes older backups to make room for new ones" with no published percentage — deletion is on-demand pre-backup, not a watermark — [Source](https://support.apple.com/guide/mac-help/if-the-time-machine-backup-disk-is-full-mh15137/mac)
|
||||
- Time Machine backupd log shows two-phase scheme: "Starting pre-backup thinning: 53.57 GB requested (including padding)" then "No expired backups exist - deleting oldest backups to make room"; post-backup thinning separately expires hourly/daily backups — [Source](https://serverfault.com/posts/39310/revisions)
|
||||
- Time Machine standard thinning ladder is hourly backups >24h old (keep first-of-day), daily backups >30 days old (keep first-of-week), then oldest-weekly deletion only when space is needed — [Source](https://discussions.apple.com/thread/251224074)
|
||||
- Time Machine per-destination quota via `tmutil setquota DESTINATION_ID QUOTA_IN_GB` caps a destination (e.g. 500 GB); quota "takes effect on the next backup, at which point older snapshots get thinned to fit" — [Source](https://superuser.com/questions/445579/how-do-i-trim-time-machine-backup-history)
|
||||
- `tmutil thinlocalsnapshots mount_point [purge_amount] [urgency]`: "tmutil will attempt (with urgency level 1-4) to reclaim purge_amount in bytes by thinning snapshots"; urgency 4 is most aggressive, e.g. `thinlocalsnapshots / 10000000000 4` to reclaim ~10 GB — [Source](https://ss64.com/mac/tmutil.html); same signature confirmed in Apple developer thread — [Source](https://developer.apple.com/forums/thread/81171)
|
||||
- Local APFS snapshots are purgeable space "automatically reclaimed as needed" but with no user-visible percentage; when reclamation lags users must thin manually (e.g. `thinlocalsnapshots / 100000000000 4` ≈ 100 GB at urgency 4) — [Source](https://discussions.apple.com/thread/252655312)
|
||||
- Veeam retention is count-based restore points / GFS, not capacity-based: "Retention policy defines the number of restore points to keep on your performance extents and capacity extents"; earliest restore point removed from chain, blocks purged from capacity tier on next offload/copy session — [Source](https://helpcenter.veeam.com/docs/vbr/userguide/capacity_tier_retention.html)
|
||||
- Veeam Scale-Out Backup Repository (SOBR) has no auto-delete-on-full: "if the extents of your scale-out backup repository run out of space, you can add a new extent"; free space on the new extent is added to SOBR capacity — [Source](https://helpcenter.veeam.com/docs/vbr/userguide/backup_repository_sobr.html)
|
||||
- Veeam SOBR placement prefers the extent with fewest chains, breaking ties by most free space; "priority is always to complete a backup" even if that violates the Data-Locality placement policy by spilling an incremental to another extent — [Source](https://veeam-best-practices-guide-v9.readthedocs.io/resource_planning/repository_sobr.html)
|
||||
- Veeam ONE / MP capacity reports use a configurable "Repository Free Space (%)" forecast threshold (worked example 30%) to flag repositories that "will run out of space", i.e. monitoring/alerting rather than enforcement — [Source](https://helpcenter.veeam.com/docs/mp/reports/capacity_planning_for_backup_repositories.html?ver=9a)
|
||||
- S3 Lifecycle has no capacity trigger at all: rules are `Days`/`Date`/`NoncurrentVersionExpiration` per prefix/tag (e.g. transition after 365 days, expire after 3650 days); S3 "quotas" are counts (buckets, access points), not bytes — [Source](https://docs.aws.amazon.com/AmazonS3/latest/API/API_LifecycleRule.html); expiration example — [Source](https://docs.amazonaws.cn/en_us/AmazonS3/latest/userguide/lifecycle-configuration-examples.md)
|
||||
- Hetzner Storage Box quota is the fixed plan size (BX11/BX21/BX31/BX41); snapshots "consume storage space from your Storage Box's storage capacity" alongside live data (`/.zfs/snapshot/`), with slot caps of 10/20/30/40 manual + 10/20/30/40 automatic snapshots per plan — [Source](https://docs.hetzner.com/storage/storage-box/snapshots/); plan overview confirms "unlimited traffic" but fixed storage — [Source](https://docs.hetzner.com/storage/storage-box/general)
|
||||
- Hetzner offers no per-subaccount quota ("there is currently no way to set quotas for each subaccount") and no auto-thinning; mitigation is manual read-only flag on sub-account directories — [Source](https://gist.github.com/jan-di/f6e403bfc6457daae3981e307bdf9a84); official docs: "all sub accounts use the storage space of your Storage Box. To control storage usage, you can manually set a sub-account's directory to read-only" — [Source](https://docs.hetzner.com/storage/storage-box/general)
|
||||
- ZFS tooling (zfs-auto-snapshot, sanoid) is count-based (`-k/--keep NUM Keep NUM recent snapshots`), not usage-based; no `--keep-below-X%` option exists — [Source](https://manpages.debian.org/bookworm/zfs-auto-snapshot/zfs-auto-snapshot.8.en.html); sanoid splits `--take-snapshots` / `--prune-snapshots` / `--cron` with Nagios-style `--monitor-capacity` reporting only — [Source](https://github.com/jimsalterjrs/sanoid)
|
||||
- ZFS percentage numbers that do exist are health/performance floors, not retention triggers: Ubuntu warns "Minimum free space to take a snapshot and preserve ZFS performance is 20%. Free space on pool rpool is 10%" — [Source](https://superuser.com/questions/1736700/how-do-i-remove-old-zfs-snapshots); OpenZFS tuning advises "Keep pool free space above 10% to avoid many metaslabs from reaching the 4% free space threshold" where allocator flips from first-fit to best-fit and IOPS collapses — [Source](https://openzfs.github.io/openzfs-docs/Performance%20and%20Tuning/Workload%20Tuning.html)
|
||||
- Borg has no quota-aware prune: `borg prune`/`borg delete` + `borg compact` are explicit/manual or script-scheduled; "repository disk space is not freed until you run borg compact" — [Source](https://manpages.ubuntu.com/manpages/jammy/man1/borg-delete.1.html); quickstart warns to "ensure that there is *always* plenty of free space" and to "use `prune` and `compact` regularly" — [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)
|
||||
|
||||
### Inferences
|
||||
- Time Machine is the outlier reference design for quota-aware retention: time ladder first, then unconditional oldest-first deletion driven by the byte size (plus padding) of the incoming backup.
|
||||
- Every other surveyed system treats capacity as an ops/monitoring concern (add extent, raise quota, manual prune) rather than a retention input, which is why "keep usage below X%" knobs are absent outside custom wrappers.
|
||||
|
||||
### Gaps
|
||||
- No vendor-published numeric watermark (e.g. "start deleting at 90%") was found for Time Machine, Veeam SOBR, or Hetzner — deletion/placement appears driven by allocation failure or free-space comparison, not a fixed percent.
|
||||
- Could not confirm any Hetzner-side automatic snapshot rotation on full; docs describe slot caps but not capacity-triggered eviction.
|
||||
|
||||
## How do they measure remote usage on dumb backends (quota APIs, PROPFIND quota properties, du-style walks) when the server exposes no quota endpoint?
|
||||
|
||||
### Takeaway
|
||||
Only WebDAV-based backends have a standard quota API (RFC 4331 PROPFIND properties + HTTP 507); S3/object storage, Borg-over-SSH, restic, and Hetzner Storage Box over SFTP/rsync/Borg have no byte-quota endpoint, so clients fall back to local `df`/repository accounting, provider console/API, or expensive tree walks.
|
||||
|
||||
### Cited Findings
|
||||
- RFC 4331 defines two live PROPFIND properties for quota: `DAV:quota-available-bytes` ("maximum amount of additional storage available to be allocated") and `DAV:quota-used-bytes` ("amount of space used ... including usage derived from sub-resources"), explicitly warning "as the DAV:quota-available-bytes on a resource approaches 0, further allocations ... may be refused" — [Source](https://datatracker.ietf.org/doc/html/rfc4331)
|
||||
- Quota exhaustion on WebDAV is signaled by HTTP 507 (Insufficient Storage), which "SHOULD be used when a client request (e.g. a PUT, PROPFIND, MKCOL, MOVE, or COPY) fails because it would exceed their quota or physical storage limits" — [Source](http://www.webdav.org/specs/rfc4331.html)
|
||||
- Nextcloud (a common self-hosted WebDAV target) implements both properties (`quota-available-bytes`, `quota-used-bytes`) retrievable via PROPFIND — [Source](https://github.com/nextcloud/documentation/blob/master/developer_manual/client_apis/WebDAV/basic.rst)
|
||||
- Hetzner Storage Box exposes usage via console/API and a "Determine available Storage Box disk space" doc path, not via a uniform in-protocol quota on all transports; supported accesses are FTP/FTPS, SFTP/SCP, SSH/rsync/BorgBackup, SMB/CIFS, WebDAV — [Source](https://docs.hetzner.com/storage/storage-box); only the WebDAV path inherits RFC 4331 properties.
|
||||
- S3 has no byte-quota endpoint: lifecycle/quota docs cover object counts and bucket limits; storage accounting is via CloudWatch/Storage Lens/billing, and lifecycle evaluation is a daily asynchronous scan ("S3 Lifecycle evaluates objects against tag-based filters daily ... queues the action for asynchronous processing") — [Source](https://docs.aws.amazon.com/AmazonS3/latest/userguide/lifecycle-expire-general-considerations.html)
|
||||
- Veeam measures SOBR extent free space by polling the extent, but "Free space data is only retrieved when no active tasks are assigned to an extent", so placement decisions can be made on stale data under continuous load — [Source](https://bp.veeam.com/vbr/3_Build_structures/B_Veeam_Components/B_backup_repositories/scaleout.html)
|
||||
- Borg/restic on "dumb" backends (SSH, SFTP, rest-server, B2) do no server-side quota query; Borg docs direct users to local filesystem monitoring ("include the free space information in your backup log files"), client-side quotas, and `borg repo-space` accounting — [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)
|
||||
- Restic cache-size issue threads show the failure of local-accounting fallback: cache at `~/Library/Caches/restic` or `/root/.cache/restic` can fill the system disk (reports of 20–40 GB, ~3–10% of repo size) with no built-in cap, and `restic cache --cleanup` only removes stale-repo caches — [Source](https://github.com/restic/restic/issues/4325)
|
||||
|
||||
### Inferences
|
||||
- For a WebDAV "dumb backend" (Hetzner, Nextcloud), PROPFIND `quota-used-bytes`/`quota-available-bytes` with Depth:0 is the cheapest correct pre-backup check; on SFTP/rsync/S3 transports the only portable options are provider-specific APIs or recursive size walks (`du`, `rclone size`, `restic stats`), which are O(files) and unsuitable per-backup.
|
||||
- Time Machine's approach sidesteps remote accounting entirely: it asks the local filesystem how many bytes (plus padding) the next backup needs and thins until that fits, rather than tracking a remote percentage.
|
||||
|
||||
### Gaps
|
||||
- No evidence found that Borg, restic, or Veeam natively issue WebDAV PROPFIND quota checks; quota-awareness would have to be added in wrapper/scheduler code.
|
||||
- Hetzner Storage Box per-protocol quota visibility (whether SFTP/SSH exposes the same numbers as WebDAV PROPFIND) is undocumented in fetched sources.
|
||||
|
||||
## What is the recommended layering: time-based policy first, capacity trigger as backstop, or capacity as the primary driver?
|
||||
|
||||
### Takeaway
|
||||
Universal recommended layering is time/count policy first, capacity as backstop — except Time Machine, where capacity is the ultimate driver after the time ladder is exhausted. Enterprise guidance (Veeam, ZFS/sanoid, S3) never recommends capacity as the primary retention rule because it makes recovery windows unpredictable.
|
||||
|
||||
### Cited Findings
|
||||
- Time Machine layering: keep "local snapshots for the past 24 hours, daily backups for the past month and weekly backups for all previous months" and "oldest backups and any local snapshots are deleted as space is needed" — time ladder first, space-need second — [Source](https://discussions.apple.com/thread/255740341)
|
||||
- Backupd implements the layering literally: pre-backup thinning first deletes expired (time-policy) backups, and only if "No expired backups exist" does it delete oldest backups to make room — [Source](https://serverfault.com/posts/39310/revisions)
|
||||
- Veeam layering: short-term + GFS (weekly/monthly/yearly) retention counts define what may be offloaded ("operational restore window ... defines which retention files can be offloaded"); capacity tier move/copy is placement, not an extra deletion rule, and "retention of the objects in the Capacity Tier is controlled by the backup or backup copy job's retention policy in restore points and not on repository level" — [Source](https://veeambp.readthedocs.io/resource_planning/repository_sobr_capacity_tier.html)
|
||||
- Veeam ONE guidance on low free space is "free up storage space on the repository or revise your backup retention policy" — i.e. human revises the time policy, system does not auto-shorten it — [Source](https://helpcenter.veeam.com/docs/one/userguide/backup_repositories_overview.html)
|
||||
- S3 layering is time-only by design: combine transition + expiration actions into a lifecycle timeline (e.g. 30d frequent → 90d infrequent → Glacier → expire); there is no capacity input to the rule engine — [Source](https://docs.aws.amazon.com/AmazonS3/latest/userguide/object-lifecycle-mgmt.md)
|
||||
- ZFS/sanoid layering is template keep-counts (hourly/daily/weekly/monthly) executed by `--cron`, with `--monitor-capacity`/`--monitor-health` feeding external alerting, not feeding back into keep-counts — [Source](https://github.com/jimsalterjrs/sanoid)
|
||||
- Borg layering per quickstart: time/count `prune` rules run on schedule plus `compact` to actually reclaim, with free-space monitoring and optional reserved-space (`borg repo-space`) as the capacity backstop — [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)
|
||||
|
||||
### Inferences
|
||||
- The sane default for a versioning/backup scheduler is: (1) declarative time-based keep rule, (2) pre-backup capacity check that prunes oldest-expired/oldest beyond-minimum first, (3) hard failure with clear "target full" signal rather than violating a minimum-retention floor silently.
|
||||
- Time Machine's "delete oldest even if within time window" is acceptable for a single-user mirror but violates enterprise expectations (Veeam/S3) where immutability and GFS floors must survive capacity pressure.
|
||||
|
||||
### Gaps
|
||||
- No authoritative source found prescribing a specific headroom percentage (e.g. "always keep 10% free") as part of a retention policy; headroom numbers come from filesystem-tuning docs, not backup-policy docs.
|
||||
|
||||
## Failure modes: what happens when the target is 100% full mid-backup, and how do systems reserve headroom?
|
||||
|
||||
### Takeaway
|
||||
Mid-backup-full behavior ranges from graceful (Time Machine aborts the run, compacts, retries; Veeam spills to another extent or skips VM below a free-space floor) to catastrophic (Borg may be unable to prune/compact without free space; ZFS deletions can themselves return ENOSPC when snapshots pin blocks; restic retries blindly on 507/ENOSPC in older versions).
|
||||
|
||||
### Cited Findings
|
||||
- Time Machine on unfreeable full: cancels the run ("Stopping backup. Backup canceled. Ejected Time Machine disk image. Compacting backup disk image to recover free space"), then retries as a fresh "Starting standard backup"; user-visible error is "This backup is too large for the backup disk. The backup requires XX GB but only YY GB are available" — [Source](https://serverfault.com/posts/39310/revisions); error text — [Source](https://osxdaily.com/2015/07/27/delete-old-backups-time-machine-mac)
|
||||
- Veeam datastore guard: jobs warn "Production datastore ... is getting low on free space (X GB left), and may run out of free disk space completely due to open snapshots" and "Skip VMs when free disk is below" logic terminates processing below the floor; hard floor is 2 GB free (registry `BlockSnapshotThreshold`, DWORD GB) even if the skip option is disabled — [Source](https://www.veeam.com/kb4379?ad=in-text-link)
|
||||
- Veeam SOBR spillover: if one extent has no free space, Veeam places the next incremental on a different extent, violating Data-Locality to prioritize completing the backup — [Source](https://veeam-best-practices-guide-v9.readthedocs.io/resource_planning/repository_sobr.html)
|
||||
- Borg worst case: "If you do run out of disk space, it can be hard or impossible to free space, because Borg needs free space to operate - even to delete backup archives"; mitigations are `borg repo-space` reservation, resizable LVs with unallocated extents, quotas, regular prune+compact — [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst)
|
||||
- Borg two-phase free: deleting an archive only marks for deletion; "repository disk space is not freed until you run borg compact" (which itself needs working space) — [Source](https://manpages.ubuntu.com/manpages/jammy/man1/borg-delete.1.html)
|
||||
- ZFS snapshot-pinned full: "if the file to be removed exists in a snapshot ... then no space is gained ... As a result, the file deletion can consume more disk space ... you can get an unexpected ENOSPC or EDQUOT when attempting to remove a file" — [Source](https://docs.oracle.com/cd/E19120-01/open.solaris/817-2271/gayra/index.html)
|
||||
- Restic mid-backup-full: local temp-pack path can panic with "no space left on device" (`panic: Write: write /tmp/restic-temp-pack-...: no space left on device`) — [Source](https://github.com/restic/restic/issues/611); newer fix "Stop retrying uploads when rest-server runs out of space" shows prior behavior was unbounded retry on ENOSPC — [Source](https://github.com/restic/restic/releases)
|
||||
- WebDAV full is a clean protocol error: 507 Insufficient Storage on PUT/MKCOL/MOVE/COPY — [Source](http://www.webdav.org/specs/rfc4331.html); Hetzner snapshots compound this because snapshot-pinned blocks silently consume the same plan quota — [Source](https://docs.hetzner.com/storage/storage-box/snapshots/)
|
||||
- Headroom mechanisms found: Time Machine "padding" added to requested bytes in pre-backup thinning ("53.57 GB requested (including padding)") — [Source](https://serverfault.com/posts/39310/revisions); Veeam 2 GB snapshot floor — [Source](https://www.veeam.com/kb4379?ad=in-text-link); Borg `repo-space` reservation + LVM overprovisioning — [Source](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst); ZFS 10–20% free-space guidance — [Source](https://openzfs.github.io/openzfs-docs/Performance%20and%20Tuning/Workload%20Tuning.html)
|
||||
|
||||
### Inferences
|
||||
- Robust design needs three headroom elements together: (a) pre-backup estimate + padding (Time Machine model), (b) a reserved-space tripwire that stops new writes before 100% (Veeam 2 GB / Borg repo-space / ZFS 10% models), (c) a recovery path that works at 100% (Time Machine compact-and-retry; Borg notably lacks one).
|
||||
- Snapshot-pinned-full (ZFS, Hetzner Storage Box, APFS) is the nastiest failure mode: deleting live files does not free bytes, so a capacity-triggered pruner must delete snapshots themselves, oldest-first, not just thin live data.
|
||||
|
||||
### Gaps
|
||||
- Exact Time Machine padding formula and Veeam SOBR stale-free-space window under load are not published; both would need empirical measurement.
|
||||
- No restic-side quota reservation feature found as of 2026 (open `--cache-size-limit` request) — [Source](https://github.com/restic/restic/issues/4325).
|
||||
@@ -0,0 +1,102 @@
|
||||
# Resource Adaptation in Backup Tools
|
||||
|
||||
## Which knobs exist (restic --limit-upload, borg --upload-ratelimit, kopia throttling, Time Machine low-priority I/O) and what are their defaults?
|
||||
|
||||
### Takeaway
|
||||
All Linux tools use static, opt-in, unlimited-by-default rate/priority knobs; Time Machine instead imposes mandatory kernel-level low-priority I/O (IOPOL_THROTTLE) with no per-backup user knob.
|
||||
|
||||
### Cited Findings
|
||||
- restic exposes global `--limit-upload rate` and `--limit-download rate` in KiB/s, default unlimited (0) — [Source](https://restic.readthedocs.io/en/stable/manual_rest.html); also documented in man pages as `--limit-upload=0` / `--limit-download=0` default unlimited — [Source](https://man.archlinux.org/man/restic-options.1.en)
|
||||
- restic documents that `--limit-upload` cannot be changed mid-run without restart; users request pv-like `-R` dynamic adjustment and resort to killing/restarting or external shapers — [Source](https://forum.restic.net/t/change-limit-upload-without-aborting-restic/5958)
|
||||
- Borg 1.x exposes `--upload-ratelimit RATE` in kiByte/s, default 0=unlimited (with `--remote-ratelimit` deprecated alias) — [Source](https://manpages.debian.org/bookworm/borgbackup/borg-common.1.en.html); also `--upload-buffer` size in MiB, default no buffer — [Source](https://man.archlinux.org/man/borg-common.1.en.txt)
|
||||
- Borg 2.x removed `--remote-ratelimit`/`--upload-ratelimit`; bandwidth is limited via `BORGSTORE_BANDWIDTH` in bits/sec, default 0=unlimited, plus `BORGSTORE_LATENCY` delay per call, or via `pv -L` ProxyCommand / rclone `--bwlimit` — [Source](https://borgbackup.readthedocs.io/en/latest/faq.html)
|
||||
- Older Borg 1.x FAQ documents upload-only `--remote-ratelimit` plus `pv`-wrapper + `BORG_RSH` for download shaping with on-the-fly `pv -R $(pidof pv) -L` changes — [Source](https://borgbackup.readthedocs.io/en/stable/faq.html)
|
||||
- Borg's limiter is a token-bucket `SleepingBandwidthLimiter` with `RATELIMIT_PERIOD = 0.1`, capping burst quota at 2x period allowance — [Source](https://github.com/borgbackup/borg/blob/da3105f1/src/borg/remote.py)
|
||||
- Kopia exposes `repository throttle set` flags `--upload-bytes-per-second`, `--download-bytes-per-second`, `--concurrent-reads/writes`, `--read/write-requests-per-second`, `--list-requests-per-second` — [Source](https://kopia.io/docs/reference/command-line/common/repository-throttle-set/); readable via `repository throttle get` — [Source](https://kopia.io/docs/reference/command-line/common/repository-throttle-get/); server-side equivalent `server throttle set` — [Source](https://kopia.io/docs/reference/command-line/common/server-throttle-set/)
|
||||
- Kopia direct-connect backends accept `--max-upload-speed`/`--max-download-speed` bytes/sec at `repository connect` time (e.g. B2/S3), persisted as `maxUploadSpeedBytesPerSecond` in repository.config — [Source](https://kopia.discourse.group/t/limit-upload-speed-as-a-policy/990)
|
||||
- Kopia `repository throttle set` is rejected on server-connected repos ("operation supported only on direct repository"), and per-policy/UI global throttle was still missing as of 2024–2026 feature requests — [Source](https://github.com/kopia/kopia/issues/3051); KopiaUI throttle exposure requested — [Source](https://github.com/kopia/kopia/issues/3586)
|
||||
- Kopia upload throttling was bursty (whole-object sleeps) until PR #2682 added per-read `DuringUpload` throttling — [Source](https://github.com/kopia/kopia/pull/2682)
|
||||
- Time Machine's `backupd` runs at `IOPOL_THROTTLE`, defined as "long-running I/O intensive background work, such as backups" that "will be throttled to prevent impact on higher policy levels" — [Source](https://eclecticlight.co/2022/01/20/why-time-machine-backups-can-be-interminably-slow/)
|
||||
- Full `IOPOL` ladder is IMPORTANT (default) / STANDARD / UTILITY / THROTTLE / PASSIVE — [Source](https://eclecticlight.co/2026/03/28/explainer-i-o-throttling/)
|
||||
- Global kill-switch `sudo sysctl debug.lowpri_throttle_enabled=0` (re-enable with `=1`, lost on reboot unless persisted via `/etc/sysctl.conf` or LaunchDaemon) removes throttle for all background I/O, not just backupd — [Source](https://osxdaily.com/2016/04/17/speed-up-time-machine-by-removing-low-process-priority-throttling/); same command/LaunchDaemon recipe — [Source](https://apple.stackexchange.com/questions/181609/time-capsule-wired-backup-transfer-slow-with-fast-bursts); throttling is I/O not CPU — [Source](https://mjtsai.com/blog/2016/03/16/massively-speed-up-time-machine-backups/)
|
||||
- Measured effect of disabling throttle: copying phase 193→332 MB/s, overall backup 160→276 MB/s (>10 GB test); pre-backup 50 MB probe writes unaffected — [Source](https://eclecticlight.co/2022/02/28/does-removing-i-o-throttling-make-backups-faster/)
|
||||
- Duplicity has no native generic bandwidth-limit option (open bug #1291633); workarounds are `trickle -s -u/-d`, WonderShaper/tc, or legacy `--scp-command="scp -l N"` (kbit/s, scp backend only, option later deprecated) — [Source](https://bugs.launchpad.net/bugs/1291633); scp `-l` throttle Q&A — [Source](https://lists.libreplanet.org/archive/html/duplicity-talk/2007-09/msg00058.html); router-QoS/DSCP attempts reported ineffective, per-machine Bandwidth Limiter used instead — [Source](https://lists.libreplanet.org/archive/html/duplicity-talk/2021-09/msg00000.html)
|
||||
- Neither restic, borg, kopia, nor duplicity ships battery- or metered-network-aware auto-pause; scheduling/power-awareness is delegated to systemd timers, DAS-CTS (macOS), or external shapers (see Gaps).
|
||||
|
||||
### Inferences
|
||||
- Rate limiting is consistently an operator-set static ceiling, not adaptive congestion control; only pv/WonderShaper/BORGSTORE allow out-of-band changes without restarting the backup.
|
||||
- Borg 2.x shifts shaping out of borg CLI into the borgstore layer, unifying ssh/sftp/rclone paths but breaking old `--upload-ratelimit` scripts.
|
||||
|
||||
### Gaps
|
||||
- No reliable source found for a built-in metered-WiFi or on-battery auto-suspend in restic/borg/kopia/duplicity as of 2026; Time Machine power/thermal inputs are internal to DAS scoring, not user flags.
|
||||
|
||||
## Does any tool auto-tune from measured resources (disk size, RAM, link speed) rather than static config, and how?
|
||||
|
||||
### Takeaway
|
||||
No tool auto-tunes from measured disk/RAM/link speed; the only resource-derived defaults are CPU-count-derived worker counts, everything else is fixed static defaults.
|
||||
|
||||
### Cited Findings
|
||||
- restic defaults: file-read concurrency 2 ("sweet spot" from HDD experiments), blob-save concurrency = `runtime.NumCPU()`, tree-save concurrency = 20x blob concurrency — [Source](https://github.com/restic/restic/blob/de9136b29f86216bd3e41397d19b25f26b578833/internal/archiver/archiver.go)
|
||||
- restic uses all available CPUs by default; `GOMAXPROCS=1` pins to one core and slightly reduces memory — [Source](https://restic.readthedocs.io/en/latest/047_tuning_parameters.html)
|
||||
- restic backend connection limit defaults to 5 (2 for local backend), tunable via `-o rest.connections=5` / `-o local.connections=2`; too-high values degrade performance — [Source](https://restic.readthedocs.io/en/latest/047_tuning_parameters.html)
|
||||
- restic `--read-concurrency` / `RESTIC_READ_CONCURRENCY` raises parallel file reads for NVMe; `--no-scan` skips the pre-backup file-count/size scan that costs extra I/O on network/FUSE mounts — [Source](https://restic.readthedocs.io/en/latest/047_tuning_parameters.html); read-concurrency flag added in 0.15.0 for fast storage — [Source](https://restic.net/blog/2023-01-12/restic-0.15.0-released/)
|
||||
- Single-large-file chunking in restic is sequential per file (~450 MB/s) with parallel hash/compress/encrypt downstream; multi-file parallelism is what scales, which is why `cores/4`-style read-concurrency guesses only hold for SSDs — [Source](https://github.com/restic/restic/issues/4477)
|
||||
- Kopia `--max-parallel-file-reads` defaults to number of logical CPU cores; lowering it lowers CPU at cost of time — [Source](https://kopia.io/docs/faqs/)
|
||||
- Kopia `--max-parallel-snapshots` controls simultaneous snapshots (server/KopiaUI); s2 `default`/`better` compressor concurrency equals logical core count, `s2-parallel-4/8` pins it — [Source](https://kopia.io/docs/advanced/compression/)
|
||||
- Time Machine scheduling via DAS-CTS scores each due background activity every few seconds against temperature, load, and priority, dispatching `com.apple.backupd-auto` via XPC only when score exceeds threshold — [Source](https://eclecticlight.co/2023/11/28/scheduling-and-dispatch-of-backups-and-other-background-activities/)
|
||||
- On Apple Silicon, background-QoS `backupd` threads are confined to Efficiency cores at reduced frequency (~972–1332 MHz, ~90% residency on E cores), capping throughput at ~300–400 items/s regardless of queue depth — [Source](https://eclecticlight.co/2022/01/20/why-time-machine-backups-can-be-interminably-slow/)
|
||||
- No evidence any tool probes link speed, disk size, or free RAM to set pack size, connections, or limits automatically; restic 16 MiB pack size, Kopia parallelism, and Borg rate defaults are all static.
|
||||
|
||||
### Inferences
|
||||
- "Auto-tuning" in this space means CPU-count-proportional worker pools plus OS-scheduler deferral (DAS-CTS/QoS), not closed-loop adaptation to throughput or memory pressure.
|
||||
- Raising concurrency without raising memory (restic 100×1 GB test hit 300 GB RAM at connections=16/reads=16 and OOMed) shows why static defaults stay conservative — [Source](https://github.com/restic/restic/issues/4477).
|
||||
|
||||
### Gaps
|
||||
- No source found documenting link-speed probing or disk-size-derived chunk/pack sizing in any of the five tools; if it exists it is not in public docs/CLI help.
|
||||
|
||||
## How do they behave on constrained boxes (low RAM SQLite/index handling, single-core compression choices)?
|
||||
|
||||
### Takeaway
|
||||
Constrained-box guidance is manual: pick cheap compression, lower parallelism/connections, enlarge packs, and accept slower runs; low-RAM index handling remains a known failure mode, not an auto-degraded mode.
|
||||
|
||||
### Cited Findings
|
||||
- Borg compression default is lz4 (very high speed, very low compression); alternatives `zstd[,L]` (default level 3), `zlib`, `lzma`, `auto,`, `none` — [Source](https://borgbackup.readthedocs.io/en/stable/usage/help.html); usage examples recommend `zlib,6` for ratio at cost of speed — [Source](https://borgbackup.readthedocs.io/en/stable/usage/create.html)
|
||||
- Upstream Borg PR proposes moving default from `lz4` to `zstd,-4` (multithreaded, as-fast-or-faster creates, slightly better ratio; large incompressible-image corpora stay ~11% slower) with MT workers capped at 4 — [Source](https://github.com/borgbackup/borg/pull/10100)
|
||||
- Kopia compression is disabled by default and set per-policy via `kopia policy set [--global] --compression=<...>` with min/max-size gates — [Source](https://kopia.io/docs/faqs/); full option list includes `s2-default/better/parallel-4/8`, `zstd/zstd-fastest/better`, `gzip/pgzip/deflate` variants — [Source](https://kopia.io/docs/reference/command-line/common/policy-set/)
|
||||
- Kopia FAQ names compression + parallelism as the two main memory culprits; recommends disabling compression or using `s2`/`deflate`/`gzip` on small files under low memory, and lowering `--max-parallel-snapshots` / `--max-parallel-file-reads` — [Source](https://kopia.io/docs/faqs/)
|
||||
- Kopia benchmark table (466 MiB corpus): s2-default 4 GiB/s at ~375 MiB RSS vs zstd 323 MiB/s at ~238 MiB vs zstd-best 19 MiB/s; on tiny files s2 stays fastest with ~2 MiB footprint — [Source](https://kopia.io/docs/advanced/compression/)
|
||||
- Kopia zstd levels map to upstream klauspost/compress: fastest≈1, default≈3, better≈7, best≈11; higher custom levels (e.g. -22 --ultra --long) require code change — [Source](https://kopia.discourse.group/t/what-is-zstd-best-compression-and-can-i-customize/4934)
|
||||
- restic on constrained HDD/NVMe: forum-tested recipe for 2.5 TB SMR-USB run is `--read-concurrency 1 --pack-size 128 --no-cache` (pack default 16 MiB raised to reduce file count and per-pack latency), with `local.connections=1` suggested for vibration-sensitive HDDs — [Source](https://forum.restic.net/t/first-backup-2-5tb-50-hours-can-i-improve-it/8288)
|
||||
- restic large-pack tradeoff: bigger packs reduce file count and help HDD/Swift/Drive limits but need more `$TMPDIR` staging (64–384 MiB guidance) and longer single-pack uploads that wear SSDs — [Source](https://github.com/restic/restic/blob/master/doc/047_tuning_parameters.rst)
|
||||
- Borg on near-full disks: 1.5 GiB-free VM case shows Borg needs headroom for segments/cache/index; workarounds discussed are extreme compression or `--upload-ratelimit` pacing plus inotify/SIGSTOP hacks, with maintainer warning such boxes are unsuitable — [Source](https://github.com/borgbackup/borg/issues/7107)
|
||||
- Single-core guidance converges: `GOMAXPROCS=1` (restic), `-C none|laz4` (borg), `--compression=s2-default|none` + `--max-parallel-file-reads=1` (kopia) minimize CPU/RAM at cost of ratio/throughput.
|
||||
|
||||
### Inferences
|
||||
- Low-RAM behavior is fail-slow/fail-OOM rather than graceful: operators must pre-lower parallelism and compression before the run.
|
||||
- s2-family (kopia) and lz4/zstd,-4 (borg) are the single-core-friendly choices; zlib/lzma/zstd-best are ratio-first and can dominate a weak CPU.
|
||||
|
||||
### Gaps
|
||||
- No citable SQLite/index RAM formula (bytes-per-file or per-GB-repo) found for current restic/borg/kopia versions; vendor docs describe symptoms and knobs, not a sizing equation.
|
||||
|
||||
## What systemd-level controls (MemoryMax, CPUQuota, IOWeight) do shipped units use, and what do docs recommend for background backup daemons?
|
||||
|
||||
### Takeaway
|
||||
Upstream backup tools ship no restrictive resource-control units; all concrete CPU/memory/I/O caps come from downstream/community units and generic systemd resource-control docs, with `Nice=` + `CPUQuota`/`MemoryMax`/`IOWeight` as the recommended trio.
|
||||
|
||||
### Cited Findings
|
||||
- systemd `CPUQuota=` sets a hard ceiling as % of one CPU (100%=1 core, 200%=2 cores) via `cpu.max`/`cpu.cfs_quota_us`; `CPUWeight=` (1–10000, default 100) is only relative under contention — [Source](https://manpages.debian.org/bullseye/systemd/systemd.resource-control.5.en.html); same semantics in Arch man — [Source](https://man.archlinux.org/man/systemd.resource-control.5)
|
||||
- systemd memory knobs: `MemoryHigh=` soft throttle/reclaim, `MemoryMax=` hard OOM-kill limit (K/M/G/T or % of RAM, `infinity` to disable), `MemorySwapMax=` swap cap; Red Hat recommends `MemoryHigh` as main control, `MemoryMax` as last defense — [Source](https://docs.redhat.com/en/documentation/red_hat_enterprise_linux/8/html/managing_monitoring_and_updating_the_kernel/assembly_configuring-resource-management-using-systemd_managing-monitoring-and-updating-the-kernel)
|
||||
- Community restic timer units use only `Nice=17` (low CPU priority) plus sandboxing (`ProtectSystem=full`, `PrivateTmp=true`, `NoNewPrivileges=yes`, `RestrictAddressFamilies=`, `SystemCallFilter=`, `AmbientCapabilities=CAP_DAC_READ_SEARCH`), with `RandomizedDelaySec=300` + `Persistent=yes` on the timer — [Source](https://www.wildtechgarden.ca/onepagers/real-life-systemd-timers/)
|
||||
- Community segmented-borg systemd design sets `CPUQuota=80%` and `MemoryMax=2G` on the backup service template — [Source](https://github.com/JoZapf/segmented-borg-backup-system/blob/refs/heads/main/docs/SYSTEMD.md)
|
||||
- Modern Debian guidance: background CPU → `nice -n 19`; background disk → `ionice -c 3` (BFQ only); hard ceilings/group fairness → unit with `CPUQuota=`/`MemoryMax=`/`IOWeight=`/`IOReadBandwidthMax=`/`IOWriteBandwidthMax=` or a shared `backup.slice`; one-shots via `systemd-run --scope -p CPUQuota=50% -p MemoryMax=1G -p IOWeight=10` — [Source](https://www.bigiron.cc/guides/process-priorities-nice-ionice-cgroup-quotas)
|
||||
- Same guide warns `ionice` is silently ignored on default `mq-deadline` SSD schedulers (only BFQ honors classes); `MemoryMax` kills rather than slows (use `MemoryHigh` for pushback); `CPUQuota=100%` means one core — [Source](https://www.bigiron.cc/guides/process-priorities-nice-ionice-cgroup-quotas)
|
||||
- Example backup slice caps restic at 80 MB/s read / 30 MB/s write via `IOReadBandwidthMax=/dev/sda 80M` + `IOWriteBandwidthMax=/dev/sda 30M` with `IOWeight=10` so foreground pools keep headroom — [Source](https://www.bigiron.cc/guides/process-priorities-nice-ionice-cgroup-quotas)
|
||||
- No shipped upstream restic/borg/kopia unit found with `MemoryMax`/`CPUQuota`/`IOWeight` preset; ArchWiki/packaging ships timers/services without resource caps (report-writer: treat absence as finding, not oversight).
|
||||
|
||||
### Inferences
|
||||
- Recommended Linux server pattern: `Nice=17–19` + optional `IOSchedulingClass=idle` for soft priority, plus a `backup.slice` with `CPUQuota=50–80%`, `MemoryHigh/Max=1–4G`, `IOWeight=10` and per-device `IOWriteBandwidthMax` for hard isolation.
|
||||
- `nice`/`ionice` alone are insufficient for noisy-neighbor backups on cgroup-v2/BFQ-mixed fleets; systemd slice quotas are the enforceable layer.
|
||||
|
||||
### Gaps
|
||||
- Could not locate a distribution-shipped (Debian/Fedora/Arch package) backup unit that presets `MemoryMax`/`CPUQuota`/`IOWeight`; all cited values are community/blog recommendations, not vendor defaults.
|
||||
- No citable official restic/borg/kopia doc page prescribing a canonical systemd resource-control stanza as of 2026.
|
||||
@@ -0,0 +1,113 @@
|
||||
# Retention / pruning / maintenance scheduling models
|
||||
|
||||
## On-demand vs scheduled vs automatic-after-each-backup: which model does each tool use, and what are the documented tradeoffs?
|
||||
|
||||
### Takeaway
|
||||
restic and BorgBackup are fully manual/external-scheduler tools (no built-in scheduler; prune-after-backup is a script/wrapper convention with lock-contention and cost tradeoffs), Kopia is automatic/built-in (maintenance fires opportunistically on client use with a single elected owner), Time Machine is fully automatic and continuous (thinning driven by schedule + disk pressure), and Veeam is built-in-scheduler enterprise (retention applied inline after job sessions plus a nightly background process, with health-check/compact on separate schedules).
|
||||
|
||||
### Cited Findings
|
||||
- restic has no built-in scheduler: `forget` only deletes snapshot objects and a separate `prune` must remove unreferenced data; `--prune` on `forget` automates the two-step sequence only when snapshots were actually removed — [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
- restic docs warn pruning is time-consuming, takes an exclusive repository lock so "backups cannot be completed" during prune, and advise planning prune windows plus running `restic check` afterwards — [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
- Community restic practice is external cron (e.g. `0 0 * * *` daily backup script running backup → check → `forget --keep-daily N --prune`) or systemd timers; third-party wrappers (restic-scheduler, autorestic/resticprofile) add per-task systemd services where backup runs often (e.g. daily) and retention+check runs monthly — [restic cron example](https://gist.github.com/perfecto25/f528f8d14e1c4b6e2a912513539a5af7); [restic-scheduler](https://github.com/AenonDynamics/restic-scheduler)
|
||||
- BorgBackup `prune` is documented as "normally used by automated backup scripts"; the official quickstart pattern is one script doing `create` → `prune` → `compact` in sequence, and `borgmatic`'s default actions are create+prune+compact+check (i.e. after-each-backup by convention, not by daemon) — [borg prune docs](https://borgbackup.readthedocs.io/en/stable/usage/prune.html); [borg quickstart](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst); [borgmatic manpage](https://manpages.ubuntu.com/manpages/focal/man1/borgmatic.1.html)
|
||||
- borgmatic docs warn the default every-run prune/compact/check is fine for small repos but too slow for very large ones, so large repos should decouple them (skip actions / separate schedules) — [borgmatic large backups guide](https://github.com/borgmatic-collective/borgmatic/blob/main/docs/how-to/deal-with-very-large-backups.md)
|
||||
- Vorta (Borg desktop GUI) exposes retention as a "Prune after each backup" checkbox, i.e. after-each-backup as an opt-in — [Vorta prune docs](https://vorta.borgbase.com/usage/prune)
|
||||
- Kopia maintenance is automatic since v0.6.0: it "will happen occasionally when the `kopia` command-line client is used", with quick tasks ~hourly and full tasks every 24h by default — [Kopia maintenance docs](https://kopia.io/docs/advanced/maintenance/)
|
||||
- Kopia defaults are quick interval 1h and full interval 24h, both enabled (from source defaults `QuickCycle 1h`, `FullCycle 24h`) — [maintenance_params.go v0.17.0](https://raw.githubusercontent.com/kopia/kopia/v0.17.0/repo/maintenance/maintenance_params.go); corroborated by user-observed "every hour for quick and every day for full" — [kopia issue #1439](https://github.com/kopia/kopia/issues/1439)
|
||||
- Kopia retention policy (which snapshots to keep: keep-latest/hourly/daily/weekly/monthly/annual) is separate from maintenance (GC of unreferenced blobs); retention is applied by deleting expired snapshots, full-maintenance Snapshot-GC then reclaims the data — [kopia policy set reference](https://kopia.io/docs/reference/command-line/common/policy-set/); [Kopia maintenance docs](https://kopia.io/docs/advanced/maintenance/)
|
||||
- Time Machine "automatically makes hourly backups for the past 24 hours, daily backups for the past month, and weekly backups for all previous months", deleting oldest backups when the disk is full — [Apple Support](https://support.apple.com/en-la/104984)
|
||||
- Time Machine local snapshots are created hourly, stored on the source disk, kept up to 24h or until space is needed, and removed automatically under pressure — [Apple local snapshots guide](https://support.apple.com/en-euro/guide/mac-help/mh35933/mac)
|
||||
- Veeam short-term retention is applied inline at the end of each job session (chain transform/merge), while GFS retention is enforced by a background process / nightly retention job (v11+ Cloud Connect docs describe a nightly "retention job"; run `History > System`, filter "retention") — [VCSP GFS retention docs](https://veeamvcsp.github.io/docs/vcc/gfs)
|
||||
- Veeam health check and defrag/compact are opt-in scheduled operations attached to backup jobs, not run after every backup: health check off by default schedule wording, default monthly (last Sat/Sun 05:00 depending on product/generation), compact disabled by default — [Veeam maintenance settings](https://helpcenter.veeam.com/docs/vbr/userguide/backup_copy_settings_backup.html); [Veeam health check](https://helpcenter.veeam.com/docs/vbr/userguide/backup_health_check.html)
|
||||
|
||||
### Inferences
|
||||
- The spectrum is: manual+external (restic, borg) → automatic-opportunistic (kopia) → fully automatic OS-driven (Time Machine) → policy-engine with built-in scheduler (Veeam).
|
||||
- Documented tradeoff axis is consistent across tools: running retention/GC after every backup minimizes staleness but costs time, I/O, lock contention and (cloud) egress; decoupling expensive steps to daily/weekly/monthly windows is the universally recommended mitigation.
|
||||
|
||||
### Gaps
|
||||
- No authoritative 2026 statement found on restic ever adding a built-in scheduler; docs and ecosystem still assume cron/systemd as of current stable docs.
|
||||
- Exact Veeam short-term-retention merge timing (end-of-session inline vs deferred background) varies by chain type (forever-forward vs forward-incremental vs reverse); per-type timing not fully pinned down from fetched pages.
|
||||
|
||||
## What are the concrete scheduling mechanisms (systemd timers, cron, built-in scheduler, maintenance windows) with example intervals?
|
||||
|
||||
### Takeaway
|
||||
restic/Borg rely on cron or systemd timers you write (daily backup typical; prune often piggybacked, check/compact weekly–monthly); Kopia uses built-in intervals (1h quick / 24h full, tunable, plus snapshot scheduling via interval/time-of-day/cron policy); Time Machine uses undocumented-internal launchd scheduling (~hourly, customizable via tools); Veeam uses a built-in per-job scheduler plus nightly background retention and monthly health-check defaults.
|
||||
|
||||
### Cited Findings
|
||||
- restic: no internal scheduler; typical community cron is daily `0 0 * * *` running a backup script; systemd-wrapper example ships `restic-scheduler@.timer` (backup, commonly daily) plus `restic-retention@.timer` (forget+prune and check, commonly monthly) with e.g. `RETENTION_POLICY_DAYS=14 WEEKS=12 MONTHS=18 YEARS=2` — [restic cron example](https://gist.github.com/perfecto25/f528f8d14e1c4b6e2a912513539a5af7); [restic-scheduler](https://github.com/AenonDynamics/restic-scheduler)
|
||||
- restic `check --read-data-subset=n/t` (or `x%`, or size like `50M`) exists precisely to spread full-data verification across scheduled runs (e.g. 1/7..7/7 across a week, or weekly `5%`); community practice converges on daily small-subset or weekly rotating-part checks rather than full `--read-data` each run — [restic 0.13 working-with-repos](https://restic.readthedocs.io/en/v0.13.0/045_working_with_repos.html); [restic forum practice](https://forum.restic.net/t/do-you-use-check-read-data/8930)
|
||||
- Borg: no daemon; scheduling via cron/systemd calling a script or `borgmatic` (sample `borgmatic.timer` ships in repo); official quickstart script runs backup→prune→compact every invocation with example policy `--keep-daily 7 --keep-weekly 4 --keep-monthly 6` — [borg quickstart](https://github.com/borgbackup/borg/blob/master/docs/quickstart.rst); [borgmatic systemd sample](https://github.com/witten/borgmatic/blob/master/sample/systemd/borgmatic.timer)
|
||||
- Borg `check --max-duration SECONDS` supports splitting a long repo check into partial checks; documented example: full check would take 7h, daily `--max-duration=3600` yields one full check per week — [borg-check manpage](https://manpages.debian.org/testing/borgbackup/borg-check.1.en.html)
|
||||
- Borg 2.x adds `--max-age` so repeated `--max-duration`-bounded runs re-check each pack at most once per age window (example `--max-duration=3600 --max-age=1w` daily ≈ full verification weekly); partial checks require `--repository-only` — [borg 2 check docs](https://borgbackup.readthedocs.io/en/latest/usage/check.html)
|
||||
- Kopia snapshot scheduling is policy-driven: `--snapshot-interval`, `--snapshot-time HH:mm,...`, `--snapshot-time-crontab`, `--run-missed`, `--manual` — [kopia policy set reference](https://kopia.io/docs/reference/command-line/common/policy-set/)
|
||||
- Kopia maintenance intervals are built-in and tunable: `maintenance set --quick-interval=2h --full-interval=8h`, enable/disable flags, and `--pause-quick/--pause-full=DURATION` to suspend — [Kopia maintenance docs](https://kopia.io/docs/advanced/maintenance/)
|
||||
- KopiaUI/server runs a "Periodic maintenance" check roughly every 10 minutes and executes quick/full work only when due (observed behavior with default 1h/24h) — [kopia issue #1439](https://github.com/kopia/kopia/issues/1439)
|
||||
- Time Machine: hourly automatic backups driven by `com.apple.backupd-auto` LaunchDaemon (`StartInterval 3600` default, adjustable via `sudo defaults write ... StartInterval -int 7200` or TimeMachineEditor interval/calendar modes) — [Apple Gazette schedule customization](https://www.applegazette.com/applegazette-mac/customize-time-machine-backups-schedule)
|
||||
- Time Machine thinning triggers: hourly→24h, daily→~30d, weekly→until-full on the backup volume; local APFS snapshots thin on age (>24h) or space pressure, manually via `tmutil thinlocalsnapshots <mount> [bytes] [urgency 1-4]` and verifiable via `tmutil verifychecksums` / Option-click "Verify Backups" (network targets) — [Apple Support](https://support.apple.com/en-la/104984); [Apple local snapshots](https://support.apple.com/en-euro/guide/mac-help/mh35933/mac); [tmutil reference](https://ss64.com/mac/tmutil.html); [Apple verify backups](https://support.apple.com/en-mn/guide/mac-help/mh26840/mac)
|
||||
- Veeam: per-job backup schedule + GFS calendar (weekly day-of-week, monthly first/second/third/fourth/last week, yearly month) with GFS fulls created on scheduled days (synthetic); since v11 GFS creation happens right on scheduled days and a nightly background retention job enforces GFS deletions — [Veeam GFS cycles](https://helpcenter.veeam.com/docs/vbr/userguide/backup_copy_gfs_periods.html); [VCSP GFS retention docs](https://veeamvcsp.github.io/docs/vcc/gfs)
|
||||
- Veeam health check default: monthly, 05:00 last Saturday (VBR backup jobs) / last Sunday (backup-copy jobs) / last Friday (Windows agent); runs piggybacked on the first incremental session of the scheduled day, or the next session if the job didn't run that day — [Veeam health check](https://helpcenter.veeam.com/docs/vbr/userguide/backup_health_check.html); [Veeam copy health check](https://helpcenter.veeam.com/docs/vbr/userguide/backup_copy_health_check.html); [Veeam agent health check](https://helpcenter.veeam.com/docs/agentforwindows/userguide/backup_health_check.html)
|
||||
- Veeam defrag/compact-full is disabled by default, scheduled via job Maintenance settings when enabled; requires free space for an auxiliary VBK and is incompatible with GFS retention enabled — [Veeam maintenance settings](https://helpcenter.veeam.com/docs/vbr/userguide/backup_copy_settings_backup.html)
|
||||
|
||||
### Inferences
|
||||
- Example "sane defaults" synthesis for adaptation: backup cadence hourly/daily (cheap op); retention-mark pass daily or after each backup; repack/GC weekly–monthly; integrity verification monthly or continuous-subset.
|
||||
- Time Machine is the only system with zero user-visible retention scheduling: policy is fixed, enforcement is event-driven (post-backup + on-full).
|
||||
|
||||
### Gaps
|
||||
- Precise modern macOS scheduler internals (whether current macOS still uses `StartInterval`-based `backupd-auto` vs newer XPC/Activity scheduler) not verified from a primary Apple source; third-party descriptions may lag macOS 14–15 changes.
|
||||
- borgmatic's current recommended check frequency knobs (e.g. `checks:` list with `frequency:` per check) were not fetched in detail; schema confirms `checks` and `compact_threshold` exist but exact frequency syntax needs a docs fetch.
|
||||
|
||||
## How do they separate cheap policy runs (mark/delete metadata) from expensive ones (repack/GC), and how often is each recommended?
|
||||
|
||||
### Takeaway
|
||||
All five systems split cheap metadata marking from expensive space reclamation; cheap passes run often (per-backup or daily), expensive passes rarely (weekly to monthly) with explicit thresholds/tuning.
|
||||
|
||||
### Cited Findings
|
||||
- restic: `forget` (cheap: deletes snapshot metadata objects only) vs `prune` (expensive: scans all snapshots, classifies packs used/partly/unused, downloads+re-uploads repacked data — "very time-consuming for remote repositories"); `--max-unused` (default `5%`) bounds repacking, `--max-repack-size` caps work per run, `--repack-cacheable-only` restricts to metadata for a fast pass — [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
- restic: `forget --prune` couples them (prune runs only if snapshots were removed); `--max-unused unlimited` minimizes time/bandwidth (keeps partly-used packs), `0` minimizes space; `--max-repack-size 0` is the documented low-scratch-space recovery mode — [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
- Borg ≤1.x: `prune` deletes archives and `compact` (segments) was fused into prune; since 1.2 they are split: "Repository disk space is not freed until you run `borg compact`", docs recommend running compact regularly but "not after each borg command", e.g. once a month possibly with check, or when space is needed — [borg prune docs](https://borgbackup.readthedocs.io/en/stable/usage/prune.html); [borg compact docs](https://borgbackup.readthedocs.io/en/stable/usage/compact.html)
|
||||
- Borg 1.4 `compact --threshold 10%` (default) compacts a segment only above that saving; `--threshold 0` forces maximum compaction, much slower — [borg compact docs](https://borgbackup.readthedocs.io/en/stable/usage/compact.html)
|
||||
- Borg 2.x `compact` is further gated (acts only when reclaimable space ≥ threshold/5, i.e. 2% at default) plus tiny-pack merging only when small packs combine to a full-size pack; `undelete` possible after prune/delete until compact runs — [borg 2 compact docs](https://borgbackup.readthedocs.io/en/latest/usage/compact.html)
|
||||
- Kopia: quick maintenance (keeps frequently-accessed `q`/`n` blobs low; never deletes metadata without another copy existing; ~hourly) vs full maintenance (Snapshot GC marking + `p`-pack compaction + dropping deleted contents; every 24h); docs warn full-maintenance effects take several hours/cycles to materialize — [Kopia maintenance docs](https://kopia.io/docs/advanced/maintenance/)
|
||||
- Kopia task list makes the split explicit: `snapshot-gc`, `quick-delete-blobs` vs `full-delete-blobs`, `quick-rewrite-contents` vs `full-rewrite-contents`, `full-drop-deleted-content`, `index-compaction`, epoch tasks — [maintenance package reference](https://pkg.go.dev/github.com/kopia/kopia/repo/maintenance)
|
||||
- Time Machine: thinning is metadata-cheap (hardlink/APFS-clone based; dropping a snapshot only frees blocks unique to it); local-snapshot thinning is pressure-driven, backup-volume thinning continuous; heavyweight equivalent (`hdiutil compact` of network sparsebundle, `tmutil verifychecksums`) is manual/occasional — [Apple Support](https://support.apple.com/en-la/104984); [Time Machine cheatsheet](https://gist.github.com/chrisbranson/4bf59fc6f3b1be6dd5058600959564b8)
|
||||
- Veeam: cheap = per-session retention application + transform/merge of chains; expensive = active/synthetic fulls (scheduled weekly/monthly), defrag+compact full (disabled by default, scheduled separately), health check with CRC+hash of latest restore point only (monthly default, not whole chain) — [Veeam maintenance settings](https://helpcenter.veeam.com/docs/vbr/userguide/backup_copy_settings_backup.html); [Veeam health check](https://helpcenter.veeam.com/docs/vbr/userguide/backup_health_check.html)
|
||||
- Veeam health check verifies only the latest restore point per chain (not history), bounding cost; if it doesn't finish before the next scheduled run the old session stops and a new one starts — [Veeam health check](https://helpcenter.veeam.com/docs/vbr/userguide/backup_health_check.html)
|
||||
|
||||
### Inferences
|
||||
- Recommended-frequency pattern: mark/expire cheaply and often (per-backup/daily); reclaim rarely (weekly/monthly/on-pressure); verify on a third, slowest cadence (monthly/rolling-subset).
|
||||
- Kopia's quick/full and Borg's prune/compact and restic's forget/prune are structurally the same two-tier design with different default couplings.
|
||||
|
||||
### Gaps
|
||||
- No hard vendor numbers for "how often to prune" in restic docs (only "plan so it completes without interfering"); frequencies above are community practice, not vendor prescription.
|
||||
- Veeam synthetic-full vs compact interaction details (when compact is worthwhile vs transform) rely on forum lore more than docs; not fully sourced here.
|
||||
|
||||
## What safety rails exist (dry-run defaults, locking against concurrent runs, what happens if a scheduled run is interrupted)?
|
||||
|
||||
### Takeaway
|
||||
Every tool offers dry-run previews (none dry-run by default); all serialize maintenance with locks/ownership (restic exclusive locks, Borg repo+cache locks, Kopia single owner + exclusive lock, Veeam job-serialization with health-check yielding); interrupted cheap passes are safe to rerun, interrupted expensive passes resume or need explicit recovery steps.
|
||||
|
||||
### Cited Findings
|
||||
- restic `forget --dry-run` prints what would be removed without removing; docs present it as the always-available preview ("You can always use `--dry-run`") — [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
- restic `prune --dry-run` likewise only shows what would be done — [restic-prune manpage](https://manpages.ubuntu.com/manpages/questing/man1/restic-prune.1.html)
|
||||
- restic refuses "empty" policies (e.g. `--keep-last 0` removes nothing) and requires `--unsafe-allow-remove-all` plus a host/tag/path filter to delete all of a group (since 0.17.0); `--group-by host,paths` default scopes policy per backup set as a safety feature — [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
- restic locks: two types (exclusive vs shared); `prune` (and `check`) require exclusive locks, backups take shared locks and can run concurrently; locks are files under `locks/` with 30-minute staleness timeout plus same-host liveness check; `--retry-lock DURATION` waits instead of failing; exit code 11 = already locked; `unlock` removes stale locks — [restic design locks](https://github.com/restic/restic/blob/master/doc/design.rst); [restic forum locking](https://forum.restic.net/t/potential-issues-with-concurrent-execution-of-restic-commands/8099); [restic-prune manpage](https://manpages.ubuntu.com/manpages/questing/man1/restic-prune.1.html)
|
||||
- restic prune is documented as interrupt-safe ("repository remains usable no matter at which point the command is interrupted") but needs scratch space; last-resort `--unsafe-recover-no-free-space` can leave repo temporarily unusable if it fails (then remove `index/` + `repair index`) — [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
- Borg docs "strongly recommend" always running `prune -v --list --dry-run` first; `--stats`/`--quick-stats` and `--dry-run` are mutually exclusive; since 1.2.0 Borg retains the oldest archive if no rule would otherwise keep anything — [borg prune docs](https://borgbackup.readthedocs.io/en/stable/usage/prune.html)
|
||||
- Borg 2.x splits deletion into soft-delete (`prune`/`delete` mark) + `compact` (irreversible); `undelete` recovers until compact runs — [borg 2 prune docs](https://borgbackup.readthedocs.io/en/master/usage/prune.html); [borg 2 compact docs](https://borgbackup.readthedocs.io/en/latest/usage/compact.html)
|
||||
- Borg `prune` auto-removes stale checkpoint archives from interrupted backups (except the latest checkpoint, still needed) — [borg prune docs](https://borgbackup.readthedocs.io/en/stable/usage/prune.html)
|
||||
- Borg locking: repository and cache locks serialize access; `borg break-lock` exists for locks left by dead processes but must only be used when no borg process is accessing the repo/cache — [borg break-lock manpage](https://manpages.debian.org/testing/borgbackup/borg-break-lock.1.en.html)
|
||||
- Borg `check` receiving SIGINT stops at the next safe boundary leaving repo and chunk index consistent; recorded partial results are kept so later checks resume where they stopped; `--repair` archive runs stop between whole archives — [borg 2 check docs](https://borgbackup.readthedocs.io/en/latest/usage/check.html)
|
||||
- Borg `check` is read-only by default; `--repair` is flagged "POTENTIALLY DANGEROUS ... might lead to data loss"; `--find-lost-archives` can restore lost archives only before `compact` removes their data — [borg check manpage](https://manpages.ubuntu.com/manpages/questing/man1/borg2-check.1.html)
|
||||
- Kopia: single maintenance owner (`user@host`, view/change via `maintenance info`/`set --owner=me`), others never auto-run maintenance; maintenance runs under an exclusive lock (`RunExclusive`); default safety windows (`PackDeleteMinAge 24h`, snapshot-GC margin for in-flight snapshots and eventual-consistency delay) mean GC takes multiple cycles; `--safety=none` disables all of it with an explicit corruption warning — [Kopia maintenance docs](https://kopia.io/docs/advanced/maintenance/); [maintenance package reference](https://pkg.go.dev/github.com/kopia/kopia/repo/maintenance)
|
||||
- Kopia `maintenance run` (quick) / `run --full` require being the owner; `--force` overrides ownership ("unsafe") — [kopia maintenance run reference](https://kopia.io/docs/reference/command-line/advanced/maintenance-run)
|
||||
- Time Machine: no dry-run for thinning; safety comes from the fixed policy + `tmutil delete` operating one snapshot at a time and verification tooling (`verifychecksums`, Verify Backups); local snapshots are APFS copy-on-write so thinning never endangers live data — [tmutil reference](https://ss64.com/mac/tmutil.html); [Apple verify backups](https://support.apple.com/en-mn/guide/mac-help/mh26840/mac)
|
||||
- Veeam: health check yields to any job/operation touching the backup (stops if one starts); unfinished health-check sessions are superseded by the next scheduled run; corrupted blocks trigger Error status + automatic retry session re-transferring bad blocks from source; GFS-flagged restore points cannot be deleted/modified during their retention window; immutable (hardened Linux/object-lock) repos block repair/deletion paths — [Veeam health check](https://helpcenter.veeam.com/docs/vbr/userguide/backup_health_check.html); [Veeam GFS cycles](https://helpcenter.veeam.com/docs/vbr/userguide/backup_copy_gfs_periods.html)
|
||||
- Veeam append-only/immutable analog to restic: with immutability, retention cannot delete until the lock expires; Linux immutable repos "do not support repair" per health-check limitations — [Veeam health check](https://helpcenter.veeam.com/docs/vbr/userguide/backup_health_check.html)
|
||||
|
||||
### Inferences
|
||||
- Common rail inventory for adaptation: (1) preview/dry-run, (2) empty-policy refusal / oldest-retention floor, (3) scope limiters (group-by, prefix/glob, per-machine chains), (4) exclusive locks + stale-lock recovery, (5) interrupt-safe/resumable expensive passes, (6) undo window (Borg undelete-before-compact; Kopia multi-cycle GC delay; Veeam GFS immutability windows).
|
||||
- restic's append-only + `--keep-within` guidance is the sharpest documented footgun: count-based `--keep-*` policies let injected attacker snapshots displace legitimate ones at `forget` time; time-window policies bound the damage — [restic forget docs](https://restic.readthedocs.io/en/stable/060_forget.html)
|
||||
|
||||
### Gaps
|
||||
- Borg 1.x concurrent-access failure mode specifics (which commands take exclusive vs shared locks) not re-verified from primary docs in this pass; `break-lock` semantics fetched, lock-type matrix not.
|
||||
- Time Machine behavior when a scheduled backup/thinning is interrupted (power loss mid-thin) not found in Apple docs; generally treated as crash-safe via APFS but unsourced — left as gap rather than claimed.
|
||||
+358
-12
@@ -23,17 +23,19 @@ from pydantic import BaseModel, Field
|
||||
from importlib import resources
|
||||
|
||||
from . import __version__
|
||||
from . import dbsafe
|
||||
from . import diff as difflib_
|
||||
from . import maintenance
|
||||
from . import recovery
|
||||
from . import stats as stats_
|
||||
from .config import Config
|
||||
from .crypto import KeyRing
|
||||
from .db import Database
|
||||
from .filters import Filters
|
||||
from .ingest import Coalescer, Rejected, Repository, TokenBucket, tree_range
|
||||
from .monitor import Monitor, RootError
|
||||
from .remote import RemoteError, RemoteSettings
|
||||
from .remote import RemoteError, RemoteSettings, WebDAV, join as remote_join, list_tree
|
||||
from .restore import Criteria, Restorer, RestoreError
|
||||
from .store import BlobStore
|
||||
from .store import BlobStore, decompress
|
||||
from .uploader import Uploader
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
@@ -49,9 +51,9 @@ class Problem(Exception):
|
||||
REJECT_STATUS = {"file-too-large": 413}
|
||||
ROOT_STATUS = {"not-found": 404, "not-a-directory": 400, "root-overlap": 409, "watch-limit": 409}
|
||||
REMOTE_STATUS = {"remote-auth": 400, "remote-unreachable": 502, "remote-http": 502,
|
||||
"remote-full": 507, "key-mismatch": 409, "claim-failed": 409}
|
||||
"remote-full": 507, "claim-failed": 409}
|
||||
RESTORE_STATUS = {"not-found": 404, "unsafe-path": 400, "criteria-too-broad": 400,
|
||||
"plan-expired": 410, "plan-used": 409, "conflict": 409}
|
||||
"plan-expired": 410, "plan-used": 409, "conflict": 409, "unrestorable": 422}
|
||||
|
||||
|
||||
def _problem(status: int, code: str, detail: str) -> JSONResponse:
|
||||
@@ -92,6 +94,63 @@ def sd_notify(message: str) -> None:
|
||||
log.debug("sd_notify failed", exc_info=True)
|
||||
|
||||
|
||||
async def _run_purge(s: "Services", task_id: str, paths: list[str]) -> None:
|
||||
"""Background half of purge-remote: WebDAV DELETEs only, no local writes."""
|
||||
task = s.purge_tasks.get(task_id)
|
||||
if task is None:
|
||||
return
|
||||
dav = WebDAV(s.uploader.settings(), transport=s.uploader.transport)
|
||||
try:
|
||||
for path in paths:
|
||||
try:
|
||||
await dav.delete(path)
|
||||
task["files_deleted"] += 1
|
||||
except Exception as exc: # noqa: BLE001 - recorded, never raised
|
||||
task["errors"].append(f"{path}: {exc}")
|
||||
if len(task["errors"]) > 50:
|
||||
task["errors"].append("... truncated")
|
||||
break
|
||||
await asyncio.sleep(0)
|
||||
# Remove now-empty subdirectories, deepest first. Best-effort: ignore failures.
|
||||
parents = sorted({p.rsplit("/", 1)[0] for p in paths if "/" in p[1:]},
|
||||
key=lambda p: p.count("/"), reverse=True)
|
||||
top = task["directory"]
|
||||
for parent in parents:
|
||||
if parent == top or not parent.startswith(top + "/"):
|
||||
continue
|
||||
try:
|
||||
await dav.delete(parent)
|
||||
except Exception: # noqa: BLE001 - empty-dir cleanup is best-effort
|
||||
pass
|
||||
task["status"] = "failed" if task["errors"] else "done"
|
||||
except Exception as exc: # noqa: BLE001 - task record is the error channel
|
||||
task["status"] = "failed"
|
||||
task["errors"].append(str(exc))
|
||||
finally:
|
||||
await dav.close()
|
||||
|
||||
|
||||
async def _run_reindex(s: "Services", task_id: str) -> None:
|
||||
"""Background half of reindex: downloads manifests, rebuilds the index."""
|
||||
task = s.reindex_tasks.get(task_id)
|
||||
if task is None:
|
||||
return
|
||||
dav = WebDAV(s.uploader.settings(), transport=s.uploader.transport)
|
||||
try:
|
||||
versions, renames, batches = await recovery.collect_manifest_records(dav, s.uploader.directory)
|
||||
task["manifests"] = len(batches)
|
||||
task["versions"] = len(versions)
|
||||
counts = recovery.rebuild_index(s.db, versions, renames)
|
||||
s.repo.reload_roots()
|
||||
task.update(counts)
|
||||
task["status"] = "done"
|
||||
except Exception as exc: # noqa: BLE001 - task record is the error channel
|
||||
task["status"] = "failed"
|
||||
task["errors"].append(str(exc))
|
||||
finally:
|
||||
await dav.close()
|
||||
|
||||
|
||||
# request models
|
||||
|
||||
class SnapshotIn(BaseModel):
|
||||
@@ -134,6 +193,28 @@ class ForgetIn(BaseModel):
|
||||
dry_run: bool = True
|
||||
|
||||
|
||||
class PurgeRemoteIn(BaseModel):
|
||||
today: str = Field(..., description="today's date as dd-mm-yyyy (safety confirmation)")
|
||||
|
||||
|
||||
class AdoptIn(BaseModel):
|
||||
url: str
|
||||
username: str
|
||||
password: str | None = None
|
||||
base_path: str = "/versioned/"
|
||||
directory: str = Field(..., description="existing remote directory to take over")
|
||||
|
||||
|
||||
class RetentionIn(BaseModel):
|
||||
dry_run: bool = True
|
||||
keep_days: float | None = Field(None, description="override config retention.keep_days")
|
||||
|
||||
|
||||
class GcIn(BaseModel):
|
||||
dry_run: bool = True
|
||||
remote: bool = Field(True, description="also delete orphan blobs from WebDAV")
|
||||
|
||||
|
||||
class RestoreIn(BaseModel):
|
||||
as_of: str | float | None = None
|
||||
paths: list[str] = []
|
||||
@@ -171,13 +252,50 @@ class Services:
|
||||
max_watch_fraction=float(cfg.get("monitor", "max_watch_fraction")),
|
||||
)
|
||||
self.restorer = Restorer(self.db, self.repo, self.blobs)
|
||||
self.keys = KeyRing.load_or_create(cfg.key_file)
|
||||
self.uploader = Uploader(cfg, self.db, self.blobs, self.keys, transport=remote_transport)
|
||||
self.uploader = Uploader(cfg, self.db, self.blobs, transport=remote_transport)
|
||||
self.repo.on_commit = self.uploader.wake
|
||||
rate = float(cfg.get("limits", "ingest_requests_per_second"))
|
||||
self.ingest_limit = TokenBucket(rate, 1.0, time.monotonic)
|
||||
self.token = cfg.api_token()
|
||||
self.started_at = time.time()
|
||||
self.purge_tasks: dict[str, dict[str, Any]] = {}
|
||||
self.reindex_tasks: dict[str, dict[str, Any]] = {}
|
||||
self.maintenance_lock = asyncio.Lock()
|
||||
self.last_maintenance: dict[str, float] = {}
|
||||
self.restorer.fetch_blob = self.blob_bytes
|
||||
|
||||
async def blob_bytes(self, sha256: str) -> bytes:
|
||||
"""Blob content, fetching from WebDAV on demand after index loss.
|
||||
|
||||
Fast path is the local spool; the remote fetch only triggers when the
|
||||
index says `uploaded` but no local copy exists (fresh reindex).
|
||||
"""
|
||||
try:
|
||||
return self.blobs.get(sha256)
|
||||
except FileNotFoundError:
|
||||
pass
|
||||
row = self.db.one("SELECT remote_state FROM blobs WHERE sha256 = ?", (sha256,))
|
||||
if row is None or row["remote_state"] != "uploaded":
|
||||
raise FileNotFoundError(sha256)
|
||||
settings = self.uploader.settings()
|
||||
if not settings.configured or not settings.directory:
|
||||
raise FileNotFoundError(sha256)
|
||||
dav = WebDAV(settings, transport=self.uploader.transport)
|
||||
try:
|
||||
data = await dav.get(remote_join(settings.directory, "blobs", sha256[:2], sha256))
|
||||
finally:
|
||||
await dav.close()
|
||||
if data is None:
|
||||
raise FileNotFoundError(sha256)
|
||||
target = self.blobs.path_for(sha256)
|
||||
target.parent.mkdir(parents=True, exist_ok=True)
|
||||
tmp = target.with_name(f".{target.name}.{os.getpid()}.tmp")
|
||||
with open(tmp, "wb") as fh:
|
||||
fh.write(data)
|
||||
fh.flush()
|
||||
os.fsync(fh.fileno())
|
||||
os.replace(tmp, target)
|
||||
return decompress(data)
|
||||
|
||||
|
||||
def create_app(cfg: Config | None = None, allowed_hosts: set[str] | None = None,
|
||||
@@ -191,12 +309,15 @@ def create_app(cfg: Config | None = None, allowed_hosts: set[str] | None = None,
|
||||
app.state.services = services
|
||||
await services.monitor.start()
|
||||
await services.uploader.start()
|
||||
maint = asyncio.get_running_loop().create_task(maintenance.maintenance_loop(services))
|
||||
sd_notify("READY=1")
|
||||
log.info("versiond %s ready on %s:%s", __version__, cfg.host, cfg.port)
|
||||
try:
|
||||
yield
|
||||
finally:
|
||||
sd_notify("STOPPING=1")
|
||||
maint.cancel()
|
||||
await asyncio.gather(maint, return_exceptions=True)
|
||||
await services.monitor.stop()
|
||||
await services.uploader.stop()
|
||||
flushed = services.coalescer.flush()
|
||||
@@ -252,6 +373,10 @@ def create_app(cfg: Config | None = None, allowed_hosts: set[str] | None = None,
|
||||
async def health(s: Services = Depends(services)) -> dict[str, Any]:
|
||||
roots = s.db.roots()
|
||||
degraded = [r["path"] for r in roots if r["mode"] in ("degraded", "missing")]
|
||||
usage = dict(getattr(s.uploader, "remote_usage", {}) or {})
|
||||
pressure = maintenance.over_limit(s)
|
||||
if pressure:
|
||||
degraded = degraded + ["remote-over-capacity-limit"]
|
||||
return {
|
||||
"status": "degraded" if degraded else "ok",
|
||||
"version": __version__,
|
||||
@@ -261,6 +386,8 @@ def create_app(cfg: Config | None = None, allowed_hosts: set[str] | None = None,
|
||||
"pending_paths": sum(1 for p in s.coalescer.pending.values() if p.content is not None),
|
||||
"watches": s.monitor.watch_total,
|
||||
"remote": s.uploader.state,
|
||||
"remote_usage": usage,
|
||||
"remote_pressure": pressure,
|
||||
}
|
||||
|
||||
@app.get("/dashboard", include_in_schema=False)
|
||||
@@ -538,7 +665,10 @@ def create_app(cfg: Config | None = None, allowed_hosts: set[str] | None = None,
|
||||
@api.get("/versions/{version_id}/content", tags=["versions"])
|
||||
async def get_content(version_id: int, s: Services = Depends(services)) -> Response:
|
||||
version = _version(s, version_id)
|
||||
content = s.blobs.get(version["blob_sha256"])
|
||||
try:
|
||||
content = await s.blob_bytes(version["blob_sha256"])
|
||||
except (FileNotFoundError, ValueError):
|
||||
raise Problem(404, "not-found", "blob is gone locally and remotely")
|
||||
try:
|
||||
return PlainTextResponse(content.decode("utf-8"))
|
||||
except UnicodeDecodeError:
|
||||
@@ -568,7 +698,16 @@ def create_app(cfg: Config | None = None, allowed_hosts: set[str] | None = None,
|
||||
if body.conflict == "rename":
|
||||
stem, ext = os.path.splitext(destination)
|
||||
destination = f"{stem}.restored-{time.strftime('%Y%m%d-%H%M%S')}{ext}"
|
||||
content = s.blobs.get(version["blob_sha256"])
|
||||
content: bytes
|
||||
try:
|
||||
content = await s.blob_bytes(version["blob_sha256"])
|
||||
except (FileNotFoundError, ValueError):
|
||||
raise RestoreError("unrestorable", f"blob {version['blob_sha256']} is gone or corrupt")
|
||||
if dbsafe.is_sqlite_image(content):
|
||||
try:
|
||||
dbsafe.verify_sqlite_bytes(content)
|
||||
except dbsafe.DatabaseUnsafe as exc:
|
||||
raise RestoreError("unrestorable", f"stored version fails integrity check: {exc.reason}")
|
||||
result = s.restorer.write(destination, content, version_id)
|
||||
s.db.audit("restore.version", version_id=version_id, destination=destination)
|
||||
return {"result": "restored", "destination": destination, **result}
|
||||
@@ -586,7 +725,10 @@ def create_app(cfg: Config | None = None, allowed_hosts: set[str] | None = None,
|
||||
s: Services = Depends(services),
|
||||
):
|
||||
a_version = _version(s, from_version)
|
||||
a = s.blobs.get(a_version["blob_sha256"])
|
||||
try:
|
||||
a = await s.blob_bytes(a_version["blob_sha256"])
|
||||
except (FileNotFoundError, ValueError):
|
||||
raise Problem(404, "not-found", "blob is gone locally and remotely")
|
||||
from_label = f"{a_version['path']}@{from_version}"
|
||||
if to == "disk":
|
||||
try:
|
||||
@@ -599,7 +741,10 @@ def create_app(cfg: Config | None = None, allowed_hosts: set[str] | None = None,
|
||||
if not to.isdigit():
|
||||
raise Problem(400, "invalid-target", "to must be a version id or 'disk'")
|
||||
b_version = _version(s, int(to))
|
||||
b = s.blobs.get(b_version["blob_sha256"])
|
||||
try:
|
||||
b = await s.blob_bytes(b_version["blob_sha256"])
|
||||
except (FileNotFoundError, ValueError):
|
||||
raise Problem(404, "not-found", "blob is gone locally and remotely")
|
||||
to_label = f"{b_version['path']}@{to}"
|
||||
if format == "html":
|
||||
return HTMLResponse(difflib_.side_by_side_html(a, b, from_label, to_label, context))
|
||||
@@ -633,6 +778,207 @@ def create_app(cfg: Config | None = None, allowed_hosts: set[str] | None = None,
|
||||
return {"path": path, "dry_run": False, "files_removed": files,
|
||||
"versions_removed": counts["versions"], "blobs_removed": blobs, "projects_removed": projects}
|
||||
|
||||
# remote purge (remote data only; never touches the local index or spool)
|
||||
|
||||
@api.post("/admin/purge-remote", tags=["admin"], status_code=202)
|
||||
async def purge_remote(body: PurgeRemoteIn, s: Services = Depends(services)) -> dict[str, Any]:
|
||||
"""Delete everything this installation uploaded (blobs/ + manifests/).
|
||||
|
||||
Safety: `today` must be today's date as dd-mm-yyyy. Only the claimed
|
||||
remote directory is ever deleted, never the base path or anything else.
|
||||
The listing happens synchronously so the response can state how many
|
||||
files/bytes will go; the actual deletes run in the background.
|
||||
Nothing local (index, spool, config) is read or modified here.
|
||||
"""
|
||||
try:
|
||||
given = datetime.strptime(body.today.strip(), "%d-%m-%Y").date()
|
||||
except ValueError:
|
||||
raise Problem(400, "invalid-date", "today must be today's date as dd-mm-yyyy (e.g. "
|
||||
f"{datetime.now().strftime('%d-%m-%Y')})")
|
||||
if given != datetime.now().date():
|
||||
raise Problem(400, "invalid-date", "today must be today's date as dd-mm-yyyy; "
|
||||
f"today is {datetime.now().strftime('%d-%m-%Y')}")
|
||||
settings = s.uploader.settings()
|
||||
directory = settings.directory.rstrip("/") or ""
|
||||
base = remote_join(settings.base_path)
|
||||
if not settings.configured or not directory:
|
||||
raise Problem(409, "remote-unconfigured", "no WebDAV remote is configured")
|
||||
if directory == base or directory == "/" or not directory.startswith(base + "/"):
|
||||
raise Problem(409, "refusing-purge", f"refusing to purge {directory!r}: outside the claimed directory")
|
||||
dav = WebDAV(settings, transport=s.uploader.transport)
|
||||
try:
|
||||
files: list = []
|
||||
for subtree in ("blobs", "manifests"):
|
||||
top = remote_join(directory, subtree)
|
||||
if await dav.exists(top):
|
||||
sub_files, _ = await list_tree(dav, top)
|
||||
files.extend(sub_files)
|
||||
finally:
|
||||
await dav.close()
|
||||
total_bytes: int | None = sum(f.size for f in files if f.size is not None)
|
||||
if any(f.size is None for f in files):
|
||||
total_bytes = None # server omitted some sizes; file count is exact
|
||||
task_id = secrets.token_hex(8)
|
||||
task: dict[str, Any] = {
|
||||
"task_id": task_id, "directory": directory, "status": "running",
|
||||
"files": len(files), "bytes": total_bytes,
|
||||
"files_deleted": 0, "errors": [],
|
||||
}
|
||||
s.purge_tasks[task_id] = task
|
||||
asyncio.get_running_loop().create_task(_run_purge(s, task_id, [f.path for f in files]))
|
||||
return {"task_id": task_id, "directory": directory, "status": "running",
|
||||
"files": len(files), "bytes": total_bytes,
|
||||
"note": "deletion runs in the background; poll GET /api/v1/admin/purge-remote/{task_id}"}
|
||||
|
||||
@api.get("/admin/purge-remote/{task_id}", tags=["admin"])
|
||||
async def purge_remote_status(task_id: str, s: Services = Depends(services)) -> dict[str, Any]:
|
||||
task = s.purge_tasks.get(task_id)
|
||||
if task is None:
|
||||
raise Problem(404, "not-found", f"purge task {task_id} does not exist")
|
||||
return task
|
||||
|
||||
# disaster recovery: rebuild the index from remote manifests
|
||||
|
||||
@api.post("/admin/reindex", tags=["admin"], status_code=202)
|
||||
async def start_reindex(s: Services = Depends(services)) -> dict[str, Any]:
|
||||
"""Rebuild the local index from remote manifests (fresh machine / lost index).
|
||||
|
||||
Lists manifests, replays version+rename records, marks everything
|
||||
durable. Responds immediately; poll the status endpoint. Local files
|
||||
are only re-registered as metadata; nothing is downloaded except
|
||||
manifests (blobs stay remote until restored).
|
||||
"""
|
||||
settings = s.uploader.settings()
|
||||
if not settings.configured or not settings.directory:
|
||||
raise Problem(409, "remote-unconfigured", "no WebDAV remote is configured")
|
||||
task_id = secrets.token_hex(8)
|
||||
task: dict[str, Any] = {"task_id": task_id, "status": "running",
|
||||
"manifests": 0, "versions": 0, "files": 0, "errors": []}
|
||||
s.reindex_tasks[task_id] = task
|
||||
asyncio.get_running_loop().create_task(_run_reindex(s, task_id))
|
||||
return {**task, "note": "poll GET /api/v1/admin/reindex/{task_id}"}
|
||||
|
||||
@api.get("/admin/reindex/{task_id}", tags=["admin"])
|
||||
async def reindex_status(task_id: str, s: Services = Depends(services)) -> dict[str, Any]:
|
||||
task = s.reindex_tasks.get(task_id)
|
||||
if task is None:
|
||||
raise Problem(404, "not-found", f"reindex task {task_id} does not exist")
|
||||
return task
|
||||
|
||||
@api.post("/config/remote/adopt", tags=["remote"])
|
||||
async def adopt_remote(body: AdoptIn, s: Services = Depends(services)) -> dict[str, Any]:
|
||||
"""Take over an existing remote directory (new machine after total loss)."""
|
||||
password = body.password if body.password is not None else s.cfg.credential("remote_password")
|
||||
if not password:
|
||||
raise Problem(400, "missing-password", "password is required")
|
||||
settings = RemoteSettings(url=body.url, username=body.username, password=password,
|
||||
base_path=body.base_path, verify_tls=True, timeout_seconds=30.0,
|
||||
directory=body.directory.rstrip("/") or "")
|
||||
if not settings.directory or settings.directory == remote_join(settings.base_path):
|
||||
raise Problem(400, "invalid-directory", "directory must be an existing claimed directory")
|
||||
dav = WebDAV(settings, transport=s.uploader.transport)
|
||||
try:
|
||||
owner_raw = await dav.get(remote_join(settings.directory, "meta", "owner.json"))
|
||||
if owner_raw is None:
|
||||
raise Problem(404, "not-found",
|
||||
f"{settings.directory} has no owner.json; nothing to adopt")
|
||||
try:
|
||||
owner = json.loads(owner_raw)
|
||||
except ValueError:
|
||||
owner = {}
|
||||
finally:
|
||||
await dav.close()
|
||||
await s.uploader.stop()
|
||||
s.cfg.update("remote", {
|
||||
"url": settings.url, "username": settings.username, "base_path": settings.base_path,
|
||||
"verify_tls": True, "timeout_seconds": 30.0, "directory": settings.directory,
|
||||
})
|
||||
s.cfg.set_credential("remote_password", password)
|
||||
s.db.audit("remote.adopt", url=settings.url, directory=settings.directory, owner=owner)
|
||||
await s.uploader.start()
|
||||
return {**s.uploader.public_settings(), "adopted": True, "previous_owner": owner}
|
||||
|
||||
# retention + garbage collection
|
||||
|
||||
@api.post("/admin/retention/run", tags=["admin"])
|
||||
async def run_retention(body: RetentionIn, s: Services = Depends(services)) -> dict[str, Any]:
|
||||
"""Thin old history: keep everything younger than keep_days (default 5),
|
||||
one version per file per older day, plus every per-file latest and all
|
||||
pins. Dry run by default; follow with POST /admin/gc to reclaim blobs."""
|
||||
keep_days = body.keep_days if body.keep_days is not None else float(s.cfg.get("retention", "keep_days"))
|
||||
if keep_days < 0:
|
||||
raise Problem(400, "invalid-keep-days", "keep_days must be >= 0")
|
||||
plan = recovery.retention_plan(s.db, keep_days)
|
||||
if body.dry_run:
|
||||
return {"dry_run": True, **{k: v for k, v in plan.items() if k != "delete_ids"},
|
||||
"versions_to_delete": len(plan["delete_ids"])}
|
||||
deleted = await recovery.delete_versions(s.db, plan["delete_ids"])
|
||||
s.db.audit("retention.run", keep_days=keep_days, deleted=deleted)
|
||||
return {"dry_run": False, "keep_days": keep_days, "candidates": plan["candidates"],
|
||||
"kept_daily": plan["kept_daily"], "versions_deleted": deleted}
|
||||
|
||||
@api.post("/admin/gc", tags=["admin"])
|
||||
async def run_gc(body: GcIn, s: Services = Depends(services)) -> dict[str, Any]:
|
||||
"""Delete blobs no version references (local spool + optionally remote)."""
|
||||
orphans = recovery.orphan_blobs(s.db)
|
||||
total_bytes = sum(o["stored_size"] or 0 for o in orphans)
|
||||
if body.dry_run:
|
||||
return {"dry_run": True, "blobs": len(orphans), "bytes": total_bytes}
|
||||
dav = None
|
||||
directory = ""
|
||||
if body.remote:
|
||||
settings = s.uploader.settings()
|
||||
if settings.configured and settings.directory:
|
||||
dav = WebDAV(settings, transport=s.uploader.transport)
|
||||
directory = settings.directory
|
||||
try:
|
||||
result = await recovery.collect_orphans(dav, s.db, s.blobs, directory, orphans)
|
||||
finally:
|
||||
if dav is not None:
|
||||
await dav.close()
|
||||
s.db.audit("gc", **{k: v for k, v in result.items() if k != "errors"})
|
||||
return {"dry_run": False, "blobs": len(orphans), "bytes": total_bytes, **result}
|
||||
|
||||
# metrics
|
||||
|
||||
@api.get("/metrics", tags=["status"])
|
||||
async def get_metrics(s: Services = Depends(services)):
|
||||
"""Prometheus text exposition (behind the same bearer auth)."""
|
||||
files = s.db.one("SELECT count(*) AS n FROM files")["n"]
|
||||
versions = s.db.one("SELECT count(*) AS n FROM versions")["n"]
|
||||
local = s.db.one("SELECT count(*) AS n FROM versions WHERE durability = 'local'")["n"]
|
||||
pending = s.db.one("SELECT count(*) AS n FROM blobs WHERE remote_state IN ('pending','failed')")["n"]
|
||||
failed = s.db.one("SELECT count(*) AS n FROM blobs WHERE remote_state = 'failed'")["n"]
|
||||
events = s.monitor.counters.get("events", 0)
|
||||
uptime = round(time.time() - s.started_at)
|
||||
body = (
|
||||
"# HELP versiond_files_tracked Files in the index\n"
|
||||
"# TYPE versiond_files_tracked gauge\n"
|
||||
f"versiond_files_tracked {files}\n"
|
||||
"# HELP versiond_versions_total Stored versions\n"
|
||||
"# TYPE versiond_versions_total gauge\n"
|
||||
f"versiond_versions_total {versions}\n"
|
||||
"# HELP versiond_versions_local Versions not yet durable remotely\n"
|
||||
"# TYPE versiond_versions_local gauge\n"
|
||||
f"versiond_versions_local {local}\n"
|
||||
"# HELP versiond_blobs_pending Blobs awaiting (re)upload\n"
|
||||
"# TYPE versiond_blobs_pending gauge\n"
|
||||
f"versiond_blobs_pending {pending}\n"
|
||||
"# HELP versiond_blobs_failed Blobs in failed/backoff state\n"
|
||||
"# TYPE versiond_blobs_failed gauge\n"
|
||||
f"versiond_blobs_failed {failed}\n"
|
||||
"# HELP versiond_monitor_events_total Inotify events seen\n"
|
||||
"# TYPE versiond_monitor_events_total counter\n"
|
||||
f"versiond_monitor_events_total {events}\n"
|
||||
"# HELP versiond_remote_state Remote state (1 when online)\n"
|
||||
"# TYPE versiond_remote_state gauge\n"
|
||||
f"versiond_remote_state {1 if s.uploader.state == 'online' else 0}\n"
|
||||
"# HELP versiond_uptime_seconds Process uptime\n"
|
||||
"# TYPE versiond_uptime_seconds counter\n"
|
||||
f"versiond_uptime_seconds {uptime}\n"
|
||||
)
|
||||
return Response(content=body, media_type="text/plain; version=0.0.4")
|
||||
|
||||
# bulk restore
|
||||
|
||||
@api.post("/restores", tags=["restore"], status_code=201)
|
||||
@@ -650,7 +996,7 @@ def create_app(cfg: Config | None = None, allowed_hosts: set[str] | None = None,
|
||||
|
||||
@api.post("/restores/{plan_id}/execute", tags=["restore"])
|
||||
async def execute_restore(plan_id: str, s: Services = Depends(services)) -> dict[str, Any]:
|
||||
return s.restorer.execute(plan_id)
|
||||
return await s.restorer.execute(plan_id)
|
||||
|
||||
app.include_router(api)
|
||||
return app
|
||||
|
||||
+2
-16
@@ -93,7 +93,7 @@ def cmd_install(args: argparse.Namespace) -> int:
|
||||
print(f"wrote {unit_path}")
|
||||
Config.load().api_token()
|
||||
|
||||
user = os.environ.get("USER") or os.getlogin()
|
||||
user = os.environ.get("USER") or os.environ.get("LOGNAME") or getpass.getuser()
|
||||
linger = _run("loginctl", "show-user", user, "-p", "Linger").stdout.strip()
|
||||
if linger != "Linger=yes":
|
||||
result = _run("loginctl", "enable-linger", user)
|
||||
@@ -306,8 +306,7 @@ def cmd_remote(args: argparse.Namespace) -> int:
|
||||
result = c.request("PUT", "/api/v1/config/remote", body)
|
||||
verb = "adopted existing" if result.get("adopted") else "claimed new"
|
||||
print(f"remote configured; {verb} directory {result['directory']} on {result['url']}")
|
||||
print(f"encryption: {result['encryption']}, key id {result['key_id']}")
|
||||
print("IMPORTANT: back up your key with `versiond key export`; without it the remote copy cannot be read.")
|
||||
print("remote storage is unencrypted; it relies on the server's disk encryption.")
|
||||
return 0
|
||||
|
||||
|
||||
@@ -323,15 +322,6 @@ def cmd_forget(args: argparse.Namespace) -> int:
|
||||
return 0
|
||||
|
||||
|
||||
def cmd_key(args: argparse.Namespace) -> int:
|
||||
from .crypto import KeyRing
|
||||
|
||||
keys = KeyRing.load_or_create(Config.load().key_file)
|
||||
print(keys.export())
|
||||
print(f"# key id {keys.key_id} - store this somewhere safe and off this machine", file=sys.stderr)
|
||||
return 0
|
||||
|
||||
|
||||
def cmd_stats(args: argparse.Namespace) -> int:
|
||||
s = Client().request("GET", "/api/v1/stats")
|
||||
if args.json:
|
||||
@@ -468,10 +458,6 @@ def build_parser() -> argparse.ArgumentParser:
|
||||
p.add_argument("--execute", action="store_true")
|
||||
p.set_defaults(func=cmd_forget)
|
||||
|
||||
p = sub.add_parser("key", help="backup encryption key")
|
||||
p.add_argument("action", choices=["export"])
|
||||
p.set_defaults(func=cmd_key)
|
||||
|
||||
p = sub.add_parser("stats", help="statistics: files, versions, storage, activity")
|
||||
p.add_argument("--json", action="store_true")
|
||||
p.set_defaults(func=cmd_stats)
|
||||
|
||||
+11
-5
@@ -17,7 +17,7 @@ DEFAULT_PORT = 9922
|
||||
DEFAULTS: dict[str, Any] = {
|
||||
"server": {"host": DEFAULT_HOST, "port": DEFAULT_PORT},
|
||||
"limits": {
|
||||
"max_file_bytes": 204_800,
|
||||
"max_file_bytes": 10_485_760,
|
||||
"ingest_requests_per_second": 50,
|
||||
},
|
||||
"coalesce": {
|
||||
@@ -50,6 +50,16 @@ DEFAULTS: dict[str, Any] = {
|
||||
"reconcile_interval_seconds": 86_400.0,
|
||||
"max_watch_fraction": 0.5,
|
||||
},
|
||||
"retention": {
|
||||
"keep_days": 5,
|
||||
},
|
||||
"scheduler": {
|
||||
"enabled": True,
|
||||
"retention_interval_seconds": 86_400.0,
|
||||
"gc_interval_seconds": 604_800.0,
|
||||
"usage_check_seconds": 600.0,
|
||||
"remote_max_used_percent": 70.0,
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
@@ -196,7 +206,3 @@ class Config:
|
||||
token = secrets.token_urlsafe(32)
|
||||
self.set_credential("api_token", token)
|
||||
return token
|
||||
|
||||
@property
|
||||
def key_file(self):
|
||||
return self.paths.config_dir / "backup.key"
|
||||
|
||||
@@ -1,57 +0,0 @@
|
||||
"""Client-side encryption for everything sent to the remote."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import base64
|
||||
import hashlib
|
||||
import hmac
|
||||
import os
|
||||
from pathlib import Path
|
||||
|
||||
from cryptography.hazmat.primitives.ciphers.aead import AESGCM
|
||||
|
||||
MAGIC = b"VD1"
|
||||
NONCE_BYTES = 12
|
||||
AAD = b"versiond"
|
||||
|
||||
|
||||
class KeyRing:
|
||||
"""A 64-byte master key: 32 bytes for AES-256-GCM, 32 bytes for keyed naming."""
|
||||
|
||||
def __init__(self, key: bytes):
|
||||
if len(key) != 64:
|
||||
raise ValueError("backup key must be 64 bytes")
|
||||
self._aead = AESGCM(key[:32])
|
||||
self._mac_key = key[32:]
|
||||
self.key = key
|
||||
|
||||
@classmethod
|
||||
def load_or_create(cls, path: Path) -> "KeyRing":
|
||||
if path.exists():
|
||||
return cls(base64.b64decode(path.read_text().strip()))
|
||||
key = os.urandom(64)
|
||||
fd = os.open(path, os.O_WRONLY | os.O_CREAT | os.O_EXCL, 0o600)
|
||||
with os.fdopen(fd, "w") as fh:
|
||||
fh.write(base64.b64encode(key).decode() + "\n")
|
||||
return cls(key)
|
||||
|
||||
def export(self) -> str:
|
||||
return base64.b64encode(self.key).decode()
|
||||
|
||||
@property
|
||||
def key_id(self) -> str:
|
||||
return hmac.new(self._mac_key, b"versiond-key-id", hashlib.sha256).hexdigest()[:16]
|
||||
|
||||
def blob_name(self, digest: str) -> str:
|
||||
"""Remote name for a blob; keyed so the server cannot confirm guessed content."""
|
||||
return hmac.new(self._mac_key, digest.encode(), hashlib.sha256).hexdigest()
|
||||
|
||||
def encrypt(self, data: bytes) -> bytes:
|
||||
nonce = os.urandom(NONCE_BYTES)
|
||||
return MAGIC + nonce + self._aead.encrypt(nonce, data, AAD)
|
||||
|
||||
def decrypt(self, data: bytes) -> bytes:
|
||||
if not data.startswith(MAGIC):
|
||||
raise ValueError("not a versiond encrypted object")
|
||||
nonce = data[len(MAGIC): len(MAGIC) + NONCE_BYTES]
|
||||
return self._aead.decrypt(nonce, data[len(MAGIC) + NONCE_BYTES:], AAD)
|
||||
@@ -0,0 +1,142 @@
|
||||
"""Enterprise-safe snapshots of live SQLite databases.
|
||||
|
||||
A raw byte copy of a live (especially WAL-mode) SQLite file can be torn or
|
||||
silently stale: the main file, -wal and -shm must be captured in one instant,
|
||||
which a plain read cannot do. So any file with the SQLite magic is snapshotted
|
||||
through the SQLite Online Backup API instead of read directly:
|
||||
|
||||
1. Open the source read-only (never triggers WAL recovery or checkpointing,
|
||||
never takes a write lock, never mutates the live DB).
|
||||
2. Copy page-by-page with ``Connection.backup()`` (restarts safely if the
|
||||
source is being written to; retries while the source is locked).
|
||||
3. Run ``PRAGMA integrity_check`` on the snapshot; only ``ok`` is stored.
|
||||
4. Anything else (locked past the deadline, unreadable, corrupt) raises
|
||||
``DatabaseUnsafe`` and nothing is versioned -- a loud skip beats a silent
|
||||
corrupt version. Monitor/API surfaces it as ``database-locked`` /
|
||||
``database-corrupt`` instead of storing garbage.
|
||||
|
||||
Journal files (``*-wal``, ``*-shm``, ``*-journal``) are never versioned on
|
||||
their own (see filters: ``database-journal``); they are folded into the main
|
||||
file's snapshot. For non-SQLite engines (postgres/mysql data files) only a
|
||||
crash-consistent raw copy is possible -- dump-then-backup remains required.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import os
|
||||
import queue
|
||||
import sqlite3
|
||||
import tempfile
|
||||
import threading
|
||||
from pathlib import Path
|
||||
|
||||
SQLITE_MAGIC = b"SQLite format 3\x00"
|
||||
|
||||
SNAPSHOT_TIMEOUT_SECONDS = 5.0
|
||||
|
||||
|
||||
class DatabaseUnsafe(Exception):
|
||||
"""Raised when no consistent snapshot can be produced. Never store partial data."""
|
||||
|
||||
def __init__(self, reason: str, detail: str = ""):
|
||||
super().__init__(detail or reason)
|
||||
self.reason = reason
|
||||
self.detail = detail
|
||||
|
||||
|
||||
def is_sqlite_image(data: bytes) -> bool:
|
||||
return data[: len(SQLITE_MAGIC)] == SQLITE_MAGIC
|
||||
|
||||
|
||||
def _integrity_ok(conn: sqlite3.Connection) -> bool:
|
||||
row = conn.execute("PRAGMA integrity_check").fetchone()
|
||||
return row is not None and row[0] == "ok"
|
||||
|
||||
|
||||
def verify_sqlite_bytes(data: bytes) -> None:
|
||||
"""Raise DatabaseUnsafe unless data is a fully consistent SQLite image."""
|
||||
if not is_sqlite_image(data):
|
||||
raise DatabaseUnsafe("not-a-database", "missing SQLite magic")
|
||||
fd, tmp = tempfile.mkstemp(prefix="versiond-verify-", suffix=".sqlite")
|
||||
try:
|
||||
with os.fdopen(fd, "wb") as fh:
|
||||
fh.write(data)
|
||||
conn = sqlite3.connect(f"file:{tmp}?mode=ro", uri=True)
|
||||
try:
|
||||
if not _integrity_ok(conn):
|
||||
raise DatabaseUnsafe("database-corrupt", "integrity_check failed on pushed content")
|
||||
finally:
|
||||
conn.close()
|
||||
except DatabaseUnsafe:
|
||||
raise
|
||||
except Exception as exc:
|
||||
raise DatabaseUnsafe("database-unreadable", str(exc)) from exc
|
||||
finally:
|
||||
try:
|
||||
os.unlink(tmp)
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
|
||||
def snapshot_sqlite(path: str, max_bytes: int, timeout: float = SNAPSHOT_TIMEOUT_SECONDS) -> bytes:
|
||||
"""Consistent, verified snapshot of a live SQLite DB file. Read-only on source.
|
||||
|
||||
The backup runs on a daemon worker thread bounded by ``timeout``: a source
|
||||
locked past the deadline fails loudly (``database-locked``) instead of
|
||||
blocking the monitor loop forever. SQLite connections are thread-local to
|
||||
the worker.
|
||||
"""
|
||||
out: queue.Queue = queue.Queue(maxsize=1)
|
||||
|
||||
def _work() -> None:
|
||||
try:
|
||||
out.put((True, _snapshot_once(path, max_bytes)))
|
||||
except Exception as exc: # noqa: BLE001 - ferried back to the caller
|
||||
out.put((False, exc))
|
||||
|
||||
worker = threading.Thread(target=_work, daemon=True)
|
||||
worker.start()
|
||||
try:
|
||||
ok, payload = out.get(timeout=timeout)
|
||||
except queue.Empty as exc:
|
||||
raise DatabaseUnsafe("database-locked", f"{path}: still locked after {timeout}s") from exc
|
||||
if ok:
|
||||
return payload
|
||||
if isinstance(payload, DatabaseUnsafe):
|
||||
raise payload
|
||||
raise DatabaseUnsafe("database-unreadable", f"{path}: {payload}") from payload
|
||||
|
||||
|
||||
def _snapshot_once(path: str, max_bytes: int) -> bytes:
|
||||
fd, tmp = tempfile.mkstemp(prefix="versiond-db-", suffix=".sqlite")
|
||||
os.close(fd)
|
||||
try:
|
||||
try:
|
||||
src = sqlite3.connect(f"file:{path}?mode=ro", uri=True, timeout=1.0)
|
||||
except Exception as exc:
|
||||
raise DatabaseUnsafe("database-unreadable", f"{path}: {exc}") from exc
|
||||
try:
|
||||
dst = sqlite3.connect(tmp)
|
||||
try:
|
||||
src.backup(dst)
|
||||
except sqlite3.OperationalError as exc:
|
||||
raise DatabaseUnsafe("database-locked", f"{path}: {exc}") from exc
|
||||
finally:
|
||||
dst.close()
|
||||
finally:
|
||||
src.close()
|
||||
conn = sqlite3.connect(f"file:{tmp}?mode=ro", uri=True)
|
||||
try:
|
||||
if not _integrity_ok(conn):
|
||||
raise DatabaseUnsafe("database-corrupt", f"integrity_check failed for {path}")
|
||||
finally:
|
||||
conn.close()
|
||||
data = Path(tmp).read_bytes()
|
||||
if len(data) > max_bytes:
|
||||
raise DatabaseUnsafe("file-too-large", f"{len(data)} bytes")
|
||||
return data
|
||||
finally:
|
||||
try:
|
||||
os.unlink(tmp)
|
||||
except OSError:
|
||||
pass
|
||||
+21
-20
@@ -8,31 +8,30 @@ from pathlib import PurePath
|
||||
|
||||
IGNORED_DIRS = frozenset({
|
||||
"node_modules", "bower_components", "jspm_packages", "vendor", "__pycache__",
|
||||
"venv", "env", "site-packages", "build", "dist", "target", "out", "bin", "obj",
|
||||
"Pods", "Carthage", "DerivedData", "coverage", "_build", "deps", "elm-stuff",
|
||||
"venv", "env", "site-packages", ".tox", "build", "dist", "target", "out", "bin", "obj",
|
||||
".gradle", "Pods", "Carthage", "DerivedData", ".next", ".nuxt", ".svelte-kit",
|
||||
"coverage", ".terraform", "_build", "deps", "elm-stuff",
|
||||
"zig-cache", "zig-out", "__pypackages__", "htmlcov",
|
||||
})
|
||||
|
||||
IGNORED_DIR_SUFFIXES = (".egg-info", ".dist-info", ".xcodeproj", ".xcworkspace")
|
||||
|
||||
IGNORED_SUFFIXES = (
|
||||
# compiled / object code
|
||||
# compiled / object code (reproducible from source; lockfiles/manifests are kept)
|
||||
".pyc", ".pyo", ".pyd", ".class", ".o", ".obj", ".so", ".dylib", ".dll", ".exe",
|
||||
".a", ".lib", ".wasm", ".jar", ".war", ".ear", ".whl", ".egg", ".beam", ".elc",
|
||||
# archives
|
||||
".zip", ".tar", ".gz", ".tgz", ".bz2", ".xz", ".zst", ".7z", ".rar",
|
||||
# media
|
||||
".png", ".jpg", ".jpeg", ".gif", ".webp", ".ico", ".bmp", ".tiff", ".psd",
|
||||
".mp3", ".mp4", ".wav", ".ogg", ".flac", ".mov", ".avi", ".mkv", ".webm",
|
||||
".pdf", ".ttf", ".otf", ".woff", ".woff2", ".eot",
|
||||
# databases
|
||||
".sqlite", ".sqlite3", ".db", ".db-wal", ".db-shm", ".sqlite-wal", ".sqlite-shm",
|
||||
# generated
|
||||
# generated bundles (rebuilt by the toolchain)
|
||||
".min.js", ".min.css", ".map",
|
||||
# editor temp files
|
||||
".swp", ".swo", ".swx", ".tmp", ".part", ".crdownload", ".orig", ".rej",
|
||||
# editor / downloader temp files (unfinished work: never versioned)
|
||||
".swp", ".swo", ".swx", ".kate-swp", ".tmp", ".temp", ".part", ".crdownload",
|
||||
".orig", ".rej",
|
||||
)
|
||||
|
||||
# NOTE: images, audio/video, archives, databases (incl. -wal/-shm journals),
|
||||
# documents (.pdf) and fonts are deliberately ACCEPTED: they are user data that
|
||||
# backup guides list as must-include. Dependency/build/cache directories above
|
||||
# stay excluded; only the per-file content types changed.
|
||||
|
||||
IGNORED_NAMES = frozenset({"4913"}) # vim's write-permission probe file
|
||||
|
||||
ALLOWED_DOTFILES = frozenset({
|
||||
@@ -47,7 +46,7 @@ ALLOWED_DOTFILES = frozenset({
|
||||
|
||||
@dataclass
|
||||
class Filters:
|
||||
max_file_bytes: int = 204_800
|
||||
max_file_bytes: int = 10_485_760 # 10 MiB: covers images, archives, small DBs
|
||||
extra_dirs: frozenset[str] = field(default_factory=frozenset)
|
||||
patterns: tuple[str, ...] = ()
|
||||
extra_dotfiles: frozenset[str] = field(default_factory=frozenset)
|
||||
@@ -97,10 +96,16 @@ class Filters:
|
||||
if self.dir_path_ignored("/".join(parts[:depth])):
|
||||
return "ignored-pattern"
|
||||
name = parts[-1]
|
||||
if name.endswith(("-wal", "-shm", "-journal")):
|
||||
# SQLite journals are folded into the main file's verified snapshot
|
||||
# (see dbsafe); they are never independently restorable versions.
|
||||
return "database-journal"
|
||||
if name.startswith("."):
|
||||
if not self._dotfile_allowed(name):
|
||||
return "hidden-file"
|
||||
elif name.endswith("~") or name in IGNORED_NAMES:
|
||||
elif (name.endswith("~") or name in IGNORED_NAMES or name.startswith("~$")
|
||||
or (name.startswith("#") and name.endswith("#"))):
|
||||
# trailing-~ backups, vim probe, Office locks, Emacs autosaves
|
||||
return "temporary-file"
|
||||
lower = name.lower()
|
||||
if lower.endswith(IGNORED_SUFFIXES):
|
||||
@@ -113,7 +118,3 @@ class Filters:
|
||||
|
||||
def size_reason(self, size: int) -> str | None:
|
||||
return "file-too-large" if size > self.max_file_bytes else None
|
||||
|
||||
@staticmethod
|
||||
def content_reason(content: bytes) -> str | None:
|
||||
return "binary-content" if b"\x00" in content[:8192] else None
|
||||
|
||||
+35
-3
@@ -12,6 +12,7 @@ from pathlib import Path, PurePath
|
||||
from typing import Any, Callable
|
||||
|
||||
from .db import Database
|
||||
from . import dbsafe
|
||||
from .filters import Filters
|
||||
from .store import BlobStore, sha256
|
||||
|
||||
@@ -83,12 +84,22 @@ class Repository:
|
||||
raise Rejected(reason, path)
|
||||
|
||||
def check_content(self, content: bytes) -> None:
|
||||
reason = self.filters.size_reason(len(content)) or self.filters.content_reason(content)
|
||||
reason = self.filters.size_reason(len(content))
|
||||
if reason:
|
||||
raise Rejected(reason, f"{len(content)} bytes")
|
||||
if dbsafe.is_sqlite_image(content):
|
||||
# Pushed bytes claiming to be a database: verify before storing.
|
||||
try:
|
||||
dbsafe.verify_sqlite_bytes(content)
|
||||
except dbsafe.DatabaseUnsafe as exc:
|
||||
raise Rejected(exc.reason, exc.detail) from exc
|
||||
|
||||
def read_file(self, path: str) -> tuple[bytes, os.stat_result]:
|
||||
"""Read a regular file without following a final symlink."""
|
||||
"""Read a regular file without following a final symlink.
|
||||
|
||||
SQLite images never go through the raw path: they are snapshotted via
|
||||
the Online Backup API + integrity_check (see dbsafe).
|
||||
"""
|
||||
fd = os.open(path, os.O_RDONLY | os.O_NOFOLLOW | os.O_CLOEXEC)
|
||||
try:
|
||||
st = os.fstat(fd)
|
||||
@@ -97,12 +108,29 @@ class Repository:
|
||||
reason = self.filters.size_reason(st.st_size)
|
||||
if reason:
|
||||
raise Rejected(reason, f"{st.st_size} bytes")
|
||||
magic = os.pread(fd, len(dbsafe.SQLITE_MAGIC), 0)
|
||||
if dbsafe.is_sqlite_image(magic):
|
||||
os.close(fd)
|
||||
return self._read_live_db(path)
|
||||
with os.fdopen(fd, "rb", closefd=False) as fh:
|
||||
content = fh.read(self.filters.max_file_bytes + 1)
|
||||
finally:
|
||||
os.close(fd)
|
||||
try:
|
||||
os.close(fd)
|
||||
except OSError:
|
||||
pass
|
||||
return content, st
|
||||
|
||||
def _read_live_db(self, path: str) -> tuple[bytes, os.stat_result]:
|
||||
try:
|
||||
data = dbsafe.snapshot_sqlite(path, self.filters.max_file_bytes)
|
||||
except dbsafe.DatabaseUnsafe as exc:
|
||||
raise Rejected(exc.reason, exc.detail or path) from exc
|
||||
st = os.stat(path)
|
||||
fields = list(st)
|
||||
fields[6] = len(data) # st_size tracks the verified snapshot, not the live file
|
||||
return data, os.stat_result(fields)
|
||||
|
||||
# projects
|
||||
|
||||
def _detect_project(self, path: str, root: dict[str, Any] | None) -> tuple[str, str]:
|
||||
@@ -365,6 +393,10 @@ class Coalescer:
|
||||
def _bucket(self, path: str) -> TokenBucket:
|
||||
bucket = self.buckets.get(path)
|
||||
if bucket is None:
|
||||
# Bound memory: the bucket table would otherwise grow with every
|
||||
# distinct path ever seen. Evict the oldest entry first.
|
||||
if len(self.buckets) >= 10_000:
|
||||
self.buckets.pop(next(iter(self.buckets)))
|
||||
bucket = self.buckets[path] = TokenBucket(self.per_hour, 3600, self.clock)
|
||||
return bucket
|
||||
|
||||
|
||||
@@ -0,0 +1,124 @@
|
||||
"""Built-in maintenance scheduler.
|
||||
|
||||
Cadence (all configurable under ``[scheduler]``):
|
||||
- retention pass every ``retention_interval_seconds`` (default daily),
|
||||
- GC folded into the pass every ``gc_interval_seconds`` (default weekly),
|
||||
- remote usage refreshed every ``usage_check_seconds`` (default 10 min).
|
||||
|
||||
Capacity is a backstop, never the driver: when remote usage reaches
|
||||
``remote_max_used_percent`` (default 70) an out-of-schedule pass runs and the
|
||||
service reports pressure (see /health) so a human analyzes — the 5-day
|
||||
retention floor is never violated automatically. This mirrors the surveyed
|
||||
consensus: time policy first, pressure relief second, static ceilings rather
|
||||
than autotuning.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import logging
|
||||
import time
|
||||
from typing import Any
|
||||
|
||||
from . import recovery
|
||||
from .remote import WebDAV, join as remote_join, quota as remote_quota
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
|
||||
def _scheduler_cfg(s: Any, key: str) -> float:
|
||||
return float(s.cfg.get("scheduler", key))
|
||||
|
||||
|
||||
async def refresh_usage(s: Any) -> dict[str, Any]:
|
||||
"""Snapshot remote usage; cached on the uploader for /health and /progress."""
|
||||
usage: dict[str, Any] = {"percent": None, "used": None, "total": None,
|
||||
"source": "unconfigured", "at": time.time()}
|
||||
settings = s.uploader.settings()
|
||||
if settings.configured and settings.directory:
|
||||
dav = WebDAV(settings, transport=s.uploader.transport)
|
||||
try:
|
||||
try:
|
||||
used, available = await remote_quota(dav, settings.directory)
|
||||
except Exception as exc:
|
||||
usage = {**usage, "source": "error", "error": str(exc)}
|
||||
else:
|
||||
if used is not None:
|
||||
total = used + available if available is not None else None
|
||||
usage = {
|
||||
"percent": round(100 * used / total, 1) if total else None,
|
||||
"used": used, "total": total, "source": "webdav-quota",
|
||||
"at": time.time(),
|
||||
}
|
||||
else:
|
||||
row = s.db.one(
|
||||
"SELECT coalesce(sum(b.stored_size), 0) AS n FROM versions v "
|
||||
"JOIN blobs b ON b.sha256 = v.blob_sha256 WHERE v.durability = 'durable'")
|
||||
usage = {"percent": None, "used": row["n"], "total": None,
|
||||
"source": "index-estimate", "at": time.time()}
|
||||
finally:
|
||||
await dav.close()
|
||||
s.uploader.remote_usage = usage
|
||||
return usage
|
||||
|
||||
|
||||
def over_limit(s: Any) -> bool:
|
||||
usage = getattr(s.uploader, "remote_usage", None) or {}
|
||||
percent = usage.get("percent")
|
||||
if percent is None:
|
||||
return False
|
||||
return percent >= _scheduler_cfg(s, "remote_max_used_percent")
|
||||
|
||||
|
||||
async def run_scheduled_pass(s: Any, reason: str) -> dict[str, Any]:
|
||||
"""One maintenance pass: refresh usage, retention, GC if due. Serialized."""
|
||||
async with s.maintenance_lock:
|
||||
await refresh_usage(s)
|
||||
keep_days = float(s.cfg.get("retention", "keep_days"))
|
||||
plan = recovery.retention_plan(s.db, keep_days)
|
||||
deleted = await recovery.delete_versions(s.db, plan["delete_ids"])
|
||||
now = time.time()
|
||||
gc_due = now - s.last_maintenance.get("gc", 0) >= _scheduler_cfg(s, "gc_interval_seconds")
|
||||
gc_result: dict[str, Any] = {"blobs_removed": 0, "remote_removed": 0, "errors": []}
|
||||
if gc_due:
|
||||
orphans = recovery.orphan_blobs(s.db)
|
||||
dav = None
|
||||
settings = s.uploader.settings()
|
||||
if settings.configured and settings.directory:
|
||||
dav = WebDAV(settings, transport=s.uploader.transport)
|
||||
try:
|
||||
gc_result = await recovery.collect_orphans(
|
||||
dav, s.db, s.blobs, settings.directory if dav else "", orphans)
|
||||
finally:
|
||||
if dav is not None:
|
||||
await dav.close()
|
||||
s.last_maintenance["gc"] = now
|
||||
s.last_maintenance["retention"] = now
|
||||
s.db.audit("maintenance.pass", reason=reason, deleted=deleted,
|
||||
gc=gc_result["blobs_removed"])
|
||||
await refresh_usage(s)
|
||||
return {"reason": reason, "keep_days": keep_days,
|
||||
"versions_deleted": deleted, "kept_daily": plan["kept_daily"],
|
||||
"gc": gc_result, "pressure": over_limit(s)}
|
||||
|
||||
|
||||
async def maintenance_loop(s: Any) -> None:
|
||||
"""Daemon loop; cancelled on shutdown. Never raises out."""
|
||||
if not s.cfg.data.get("scheduler", {}).get("enabled", True):
|
||||
return
|
||||
await refresh_usage(s)
|
||||
while True:
|
||||
try:
|
||||
await asyncio.sleep(_scheduler_cfg(s, "usage_check_seconds"))
|
||||
await refresh_usage(s)
|
||||
now = time.time()
|
||||
if over_limit(s):
|
||||
log.warning("remote usage over limit; running pressure maintenance")
|
||||
await run_scheduled_pass(s, "capacity-pressure")
|
||||
elif now - s.last_maintenance.get("retention", 0) >= _scheduler_cfg(
|
||||
s, "retention_interval_seconds"):
|
||||
await run_scheduled_pass(s, "scheduled")
|
||||
except asyncio.CancelledError:
|
||||
raise
|
||||
except Exception:
|
||||
log.exception("maintenance pass failed; will retry next interval")
|
||||
@@ -0,0 +1,200 @@
|
||||
"""Disaster recovery (reindex from remote), retention and garbage collection.
|
||||
|
||||
Reindex rebuilds the local SQLite index purely from remote manifests, so a
|
||||
total local loss (disk dead, fresh machine after adopt) is recoverable.
|
||||
Retention keeps every version younger than ``keep_days`` (default 5), thins
|
||||
older ones to one-per-day plus the per-file latest, and never touches pinned
|
||||
versions. GC deletes blobs no version references anymore, locally and (on
|
||||
request) remotely. All three are additive-safe: they can only leak remote
|
||||
garbage, never delete referenced data.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import json
|
||||
import time
|
||||
from datetime import datetime, timezone
|
||||
from typing import Any
|
||||
|
||||
from .db import Database
|
||||
from .remote import WebDAV, join as remote_join, list_tree
|
||||
from .store import BlobStore, decompress
|
||||
|
||||
MANIFEST_SUFFIX = ".jsonl.zst"
|
||||
|
||||
|
||||
async def collect_manifest_records(
|
||||
dav: WebDAV, directory: str
|
||||
) -> tuple[list[dict[str, Any]], list[dict[str, Any]], list[str]]:
|
||||
"""Download and parse every manifest. Returns (versions, renames, batches)."""
|
||||
top = remote_join(directory, "manifests")
|
||||
if not await dav.exists(top):
|
||||
return [], [], []
|
||||
files, _ = await list_tree(dav, top)
|
||||
versions: list[dict[str, Any]] = []
|
||||
renames: list[dict[str, Any]] = []
|
||||
batches: list[str] = []
|
||||
for entry in sorted(files, key=lambda e: e.path):
|
||||
if not entry.path.endswith(MANIFEST_SUFFIX):
|
||||
continue
|
||||
data = await dav.get(entry.path)
|
||||
if data is None:
|
||||
continue
|
||||
batch = entry.path.rsplit("/", 1)[-1][: -len(MANIFEST_SUFFIX)]
|
||||
batches.append(batch)
|
||||
for line in decompress(data).decode().splitlines():
|
||||
if not line.strip():
|
||||
continue
|
||||
try:
|
||||
rec = json.loads(line)
|
||||
except ValueError:
|
||||
continue
|
||||
rec["manifest_batch"] = batch
|
||||
(versions if rec.get("type") == "version" else renames).append(rec)
|
||||
return versions, renames, batches
|
||||
|
||||
|
||||
def rebuild_index(
|
||||
db: Database, versions: list[dict[str, Any]], renames: list[dict[str, Any]]
|
||||
) -> dict[str, int]:
|
||||
"""Replace index contents from manifest records. Returns counts."""
|
||||
with db.transaction():
|
||||
db.execute("DELETE FROM versions")
|
||||
db.execute("DELETE FROM renames")
|
||||
db.execute("DELETE FROM files")
|
||||
db.execute("DELETE FROM projects")
|
||||
db.execute("DELETE FROM blobs")
|
||||
now = time.time()
|
||||
n_blobs = 0
|
||||
for rec in versions:
|
||||
sha = rec.get("sha256", "")
|
||||
if not sha:
|
||||
continue
|
||||
cur = db.execute(
|
||||
"INSERT OR IGNORE INTO blobs(sha256, size, stored_size, created_at, remote_state, uploaded_at)"
|
||||
" VALUES (?, ?, ?, ?, 'uploaded', ?)",
|
||||
(sha, rec.get("size", 0), rec.get("stored_size", 0) or 0,
|
||||
rec.get("captured_at", now), now),
|
||||
)
|
||||
n_blobs += cur.rowcount
|
||||
file_ids: dict[str, int] = {}
|
||||
n_versions = 0
|
||||
for rec in sorted(versions, key=lambda r: (r.get("captured_at", 0), r.get("id", 0))):
|
||||
path, sha = rec.get("path", ""), rec.get("sha256", "")
|
||||
if not path or not sha:
|
||||
continue
|
||||
file_id = file_ids.get(path)
|
||||
if file_id is None:
|
||||
db.execute(
|
||||
"INSERT OR IGNORE INTO files(path, exists_on_disk, size, first_seen_at) "
|
||||
"VALUES (?, 0, ?, ?)",
|
||||
(path, rec.get("size", 0), rec.get("captured_at", now)),
|
||||
)
|
||||
row = db.one("SELECT id FROM files WHERE path = ?", (path,))
|
||||
file_id = file_ids[path] = row["id"]
|
||||
db.execute(
|
||||
"INSERT INTO versions(file_id, blob_sha256, captured_at, source, reason, pinned,"
|
||||
" durability, manifest_batch) VALUES (?, ?, ?, ?, ?, ?, 'durable', ?)",
|
||||
(file_id, sha, rec.get("captured_at", now), rec.get("source", "?"),
|
||||
rec.get("reason", "?"), 1 if rec.get("pinned") else 0, rec.get("manifest_batch")),
|
||||
)
|
||||
n_versions += 1
|
||||
for file_id, path in ((v, k) for k, v in file_ids.items()):
|
||||
row = db.one(
|
||||
"SELECT count(*) AS n, min(captured_at) AS first, max(captured_at) AS last "
|
||||
"FROM versions WHERE file_id = ?",
|
||||
(file_id,),
|
||||
)
|
||||
db.execute(
|
||||
"UPDATE files SET first_seen_at = ?, last_version_at = ? WHERE id = ?",
|
||||
(row["first"], row["last"], file_id),
|
||||
)
|
||||
n_renames = 0
|
||||
for rec in renames:
|
||||
row = db.one("SELECT id FROM files WHERE path = ?", (rec.get("new_path", ""),))
|
||||
if row is None:
|
||||
continue
|
||||
db.execute(
|
||||
"INSERT INTO renames(file_id, old_path, new_path, at, manifest_batch) "
|
||||
"VALUES (?, ?, ?, ?, ?)",
|
||||
(row["id"], rec.get("old_path", ""), rec.get("new_path", ""),
|
||||
rec.get("at", now), rec.get("manifest_batch")),
|
||||
)
|
||||
n_renames += 1
|
||||
db.audit("reindex", files=len(file_ids), versions=n_versions, renames=n_renames, blobs=n_blobs)
|
||||
return {"files": len(file_ids), "versions": n_versions, "renames": n_renames, "blobs": n_blobs}
|
||||
|
||||
|
||||
def retention_plan(db: Database, keep_days: float) -> dict[str, Any]:
|
||||
"""Versions safe to delete: older than keep_days, not pinned, not per-file latest.
|
||||
|
||||
Of those, one per file per UTC day is kept (thinning); the rest go.
|
||||
"""
|
||||
cutoff = time.time() - keep_days * 86_400
|
||||
latest = {r["file_id"]: r["latest"] for r in db.all(
|
||||
"SELECT file_id, max(id) AS latest FROM versions GROUP BY file_id")}
|
||||
cands = db.all(
|
||||
"SELECT id, file_id, captured_at FROM versions "
|
||||
"WHERE captured_at < ? AND pinned = 0 ORDER BY file_id, captured_at",
|
||||
(cutoff,),
|
||||
)
|
||||
cands = [c for c in cands if c["id"] != latest.get(c["file_id"])]
|
||||
keep: set[int] = set()
|
||||
by_day: dict[tuple[int, str], dict[str, Any]] = {}
|
||||
for cand in cands:
|
||||
day = datetime.fromtimestamp(cand["captured_at"], timezone.utc).strftime("%Y-%m-%d")
|
||||
key = (cand["file_id"], day)
|
||||
if key not in by_day or cand["id"] > by_day[key]["id"]:
|
||||
by_day[key] = cand
|
||||
keep = {c["id"] for c in by_day.values()}
|
||||
delete = [c["id"] for c in cands if c["id"] not in keep]
|
||||
return {"cutoff": cutoff, "keep_days": keep_days, "candidates": len(cands),
|
||||
"kept_daily": len(keep), "delete_ids": delete}
|
||||
|
||||
|
||||
async def delete_versions(db: Database, ids: list[int], chunk: int = 500) -> int:
|
||||
deleted = 0
|
||||
for i in range(0, len(ids), chunk):
|
||||
part = ids[i: i + chunk]
|
||||
with db.transaction():
|
||||
db.execute(f"DELETE FROM versions WHERE id IN ({','.join('?' * len(part))})", part)
|
||||
deleted += len(part)
|
||||
await asyncio.sleep(0)
|
||||
return deleted
|
||||
|
||||
|
||||
def orphan_blobs(db: Database) -> list[dict[str, Any]]:
|
||||
return db.all(
|
||||
"SELECT sha256, stored_size FROM blobs b WHERE NOT EXISTS "
|
||||
"(SELECT 1 FROM versions v WHERE v.blob_sha256 = b.sha256)")
|
||||
|
||||
|
||||
async def collect_orphans(
|
||||
dav: WebDAV | None, db: Database, blobs: BlobStore, directory: str, orphans: list[dict[str, Any]],
|
||||
task: dict[str, Any] | None = None,
|
||||
) -> dict[str, Any]:
|
||||
"""Delete orphan blobs locally and remotely. Only unreferenced data is touched."""
|
||||
removed = remote_removed = 0
|
||||
errors: list[str] = []
|
||||
for orphan in orphans:
|
||||
sha = orphan["sha256"]
|
||||
for cand in blobs._candidates(sha):
|
||||
try:
|
||||
cand.unlink()
|
||||
except FileNotFoundError:
|
||||
pass
|
||||
except OSError as exc:
|
||||
errors.append(f"{sha}: {exc}")
|
||||
db.execute("DELETE FROM blobs WHERE sha256 = ?", (sha,))
|
||||
removed += 1
|
||||
if dav is not None:
|
||||
try:
|
||||
await dav.delete(remote_join(directory, "blobs", sha[:2], sha))
|
||||
remote_removed += 1
|
||||
except Exception as exc: # noqa: BLE001 - recorded per blob
|
||||
errors.append(f"remote {sha}: {exc}")
|
||||
if task is not None:
|
||||
task["blobs_removed"] = removed
|
||||
await asyncio.sleep(0)
|
||||
return {"blobs_removed": removed, "remote_removed": remote_removed, "errors": errors[:50]}
|
||||
+135
-13
@@ -10,13 +10,15 @@ import re
|
||||
import secrets
|
||||
import socket
|
||||
import time
|
||||
import xml.etree.ElementTree as ET
|
||||
from dataclasses import dataclass
|
||||
from typing import Any
|
||||
from urllib.parse import unquote, urlparse
|
||||
|
||||
import httpx
|
||||
|
||||
INSTALL_APP_ID = b"versiond:4c1b9e0f7a2d4e8b"
|
||||
FORMAT_VERSION = 1
|
||||
FORMAT_VERSION = 2
|
||||
|
||||
|
||||
class RemoteError(Exception):
|
||||
@@ -53,8 +55,12 @@ def installation_id() -> str:
|
||||
machine_id = ""
|
||||
if not machine_id:
|
||||
return secrets.token_hex(8)
|
||||
try:
|
||||
key = bytes.fromhex(machine_id)
|
||||
except ValueError:
|
||||
return secrets.token_hex(8)
|
||||
msg = INSTALL_APP_ID + b":" + str(os.getuid()).encode()
|
||||
return hmac.new(bytes.fromhex(machine_id), msg, hashlib.sha256).hexdigest()[:16]
|
||||
return hmac.new(key, msg, hashlib.sha256).hexdigest()[:16]
|
||||
|
||||
|
||||
def hostname_slug() -> str:
|
||||
@@ -155,11 +161,13 @@ async def check_connection(dav: WebDAV) -> None:
|
||||
await dav.delete(probe)
|
||||
|
||||
|
||||
async def claim_directory(dav: WebDAV, key_id: str) -> dict[str, Any]:
|
||||
async def claim_directory(dav: WebDAV) -> dict[str, Any]:
|
||||
"""Pick and claim this installation's unique directory below base_path.
|
||||
|
||||
Returns {"directory": ..., "adopted": bool}. Never takes over a directory that
|
||||
belongs to a different installation id.
|
||||
belongs to a different installation id. Remote data is stored unencrypted
|
||||
(relying on server disk encryption); blobs are plain content-addressed
|
||||
objects named by their SHA-256.
|
||||
"""
|
||||
base = join(dav.settings.base_path)
|
||||
await dav.makedirs(base)
|
||||
@@ -177,12 +185,11 @@ async def claim_directory(dav: WebDAV, key_id: str) -> dict[str, Any]:
|
||||
"hostname": socket.gethostname(),
|
||||
"user": os.environ.get("USER", ""),
|
||||
"created_at": time.time(),
|
||||
"key_id": key_id,
|
||||
}
|
||||
if not await dav.put(owner_path, json.dumps(owner, indent=2).encode(), if_none_match=True):
|
||||
continue # someone created it between our GET and PUT; re-read
|
||||
fmt = {"format": FORMAT_VERSION, "encryption": "aes-256-gcm", "compression": "zstd",
|
||||
"blob_layout": "blobs/<hmac[:2]>/<hmac>", "key_id": key_id}
|
||||
fmt = {"format": FORMAT_VERSION, "encryption": "none", "compression": "zstd-or-zlib",
|
||||
"blob_layout": "blobs/<sha256[:2]>/<sha256>"}
|
||||
await dav.put(join(directory, "meta", "format.json"), json.dumps(fmt, indent=2).encode())
|
||||
return {"directory": directory, "adopted": False}
|
||||
try:
|
||||
@@ -190,12 +197,127 @@ async def claim_directory(dav: WebDAV, key_id: str) -> dict[str, Any]:
|
||||
except ValueError:
|
||||
owner = {}
|
||||
if owner.get("installation_id") == install_id:
|
||||
if owner.get("key_id") not in (None, key_id):
|
||||
raise RemoteError(
|
||||
"key-mismatch",
|
||||
f"{directory} was written with a different backup key; restore that key "
|
||||
"(~/.config/versiond/backup.key) or remove the directory",
|
||||
)
|
||||
return {"directory": directory, "adopted": True}
|
||||
name = f"{hostname_slug()}-{install_id[:8]}-{secrets.token_hex(2)}"
|
||||
raise RemoteError("claim-failed", "could not claim a unique remote directory")
|
||||
|
||||
|
||||
# remote listing (read-only PROPFIND) for the purge endpoint
|
||||
|
||||
DAV_NS = "DAV:"
|
||||
_PROPFIND_BODY = (
|
||||
'<?xml version="1.0" encoding="utf-8"?>'
|
||||
'<d:propfind xmlns:d="DAV:"><d:prop>'
|
||||
"<d:resourcetype/><d:getcontentlength/>"
|
||||
"</d:prop></d:propfind>"
|
||||
).encode()
|
||||
|
||||
_QUOTA_BODY = (
|
||||
'<?xml version="1.0" encoding="utf-8"?>'
|
||||
'<d:propfind xmlns:d="DAV:"><d:prop>'
|
||||
"<d:quota-used-bytes/><d:quota-available-bytes/>"
|
||||
"</d:prop></d:propfind>"
|
||||
).encode()
|
||||
|
||||
|
||||
@dataclass
|
||||
class RemoteEntry:
|
||||
path: str
|
||||
is_dir: bool
|
||||
size: int | None = None
|
||||
|
||||
|
||||
def _href_to_path(href: str) -> str:
|
||||
path = unquote(urlparse(href).path or href)
|
||||
return path.rstrip("/") or "/"
|
||||
|
||||
|
||||
async def list_dir(dav: "WebDAV", path: str) -> list[RemoteEntry]:
|
||||
"""Immediate children of a collection via PROPFIND Depth:1 (read-only)."""
|
||||
resp = await dav.request(
|
||||
"PROPFIND", path,
|
||||
headers={"Depth": "1", "Content-Type": "application/xml"},
|
||||
content=_PROPFIND_BODY,
|
||||
)
|
||||
if resp.status_code == 404:
|
||||
return []
|
||||
WebDAV._check(resp, 207)
|
||||
try:
|
||||
root = ET.fromstring(resp.content)
|
||||
except ET.ParseError as exc:
|
||||
raise RemoteError("remote-http", f"PROPFIND {path} returned unparseable XML") from exc
|
||||
entries: list[RemoteEntry] = []
|
||||
for response in root.findall(f"{{{DAV_NS}}}response"):
|
||||
href = response.findtext(f"{{{DAV_NS}}}href")
|
||||
if not href:
|
||||
continue
|
||||
child = _href_to_path(href)
|
||||
if child == path.rstrip("/") or child == path:
|
||||
continue # the collection itself, not a child
|
||||
is_dir = False
|
||||
size: int | None = None
|
||||
for propstat in response.findall(f"{{{DAV_NS}}}propstat"):
|
||||
prop = propstat.find(f"{{{DAV_NS}}}prop")
|
||||
if prop is None:
|
||||
continue
|
||||
restype = prop.find(f"{{{DAV_NS}}}resourcetype")
|
||||
if restype is not None and restype.find(f"{{{DAV_NS}}}collection") is not None:
|
||||
is_dir = True
|
||||
length = prop.findtext(f"{{{DAV_NS}}}getcontentlength")
|
||||
if length is not None:
|
||||
try:
|
||||
size = int(length)
|
||||
except ValueError:
|
||||
size = None
|
||||
entries.append(RemoteEntry(child, is_dir, size))
|
||||
return entries
|
||||
|
||||
|
||||
async def list_tree(dav: "WebDAV", top: str) -> tuple[list[RemoteEntry], list[RemoteEntry]]:
|
||||
"""All files and directories strictly below top. Returns (files, dirs)."""
|
||||
files: list[RemoteEntry] = []
|
||||
dirs: list[RemoteEntry] = []
|
||||
stack = [top]
|
||||
while stack:
|
||||
current = stack.pop()
|
||||
for entry in await list_dir(dav, current):
|
||||
if entry.is_dir:
|
||||
dirs.append(entry)
|
||||
stack.append(entry.path)
|
||||
else:
|
||||
files.append(entry)
|
||||
return files, dirs
|
||||
|
||||
|
||||
async def quota(dav: "WebDAV", path: str) -> tuple[int | None, int | None]:
|
||||
"""RFC 4331 quota properties. Returns (used_bytes, available_bytes or None).
|
||||
|
||||
(None, None) when the server exposes no quota information.
|
||||
"""
|
||||
resp = await dav.request(
|
||||
"PROPFIND", path,
|
||||
headers={"Depth": "0", "Content-Type": "application/xml"},
|
||||
content=_QUOTA_BODY,
|
||||
)
|
||||
if resp.status_code == 404:
|
||||
return None, None
|
||||
WebDAV._check(resp, 207)
|
||||
try:
|
||||
root = ET.fromstring(resp.content)
|
||||
except ET.ParseError:
|
||||
return None, None
|
||||
used = available = None
|
||||
for prop in root.findall(f".//{{{DAV_NS}}}prop"):
|
||||
used_text = prop.findtext(f"{{{DAV_NS}}}quota-used-bytes")
|
||||
avail_text = prop.findtext(f"{{{DAV_NS}}}quota-available-bytes")
|
||||
if used_text is not None:
|
||||
try:
|
||||
used = int(used_text)
|
||||
except ValueError:
|
||||
used = None
|
||||
if avail_text is not None:
|
||||
try:
|
||||
available = int(avail_text)
|
||||
except ValueError:
|
||||
available = None
|
||||
return used, available
|
||||
|
||||
+35
-5
@@ -12,6 +12,7 @@ from pathlib import Path
|
||||
from typing import Any, Literal
|
||||
|
||||
from .db import Database
|
||||
from . import dbsafe
|
||||
from .ingest import Repository, is_within
|
||||
from .store import BlobStore, sha256
|
||||
|
||||
@@ -48,12 +49,24 @@ class Restorer:
|
||||
self.db = db
|
||||
self.repo = repo
|
||||
self.blobs = blobs
|
||||
# Async (sha256) -> bytes; set by the service layer for on-demand
|
||||
# remote fetch after reindex. Defaults to local spool only.
|
||||
self.fetch_blob = None
|
||||
|
||||
# path safety
|
||||
|
||||
def allowed_bases(self) -> list[str]:
|
||||
return [str(Path.home())] + [r["path"] for r in self.db.roots()]
|
||||
|
||||
@staticmethod
|
||||
def _protected_dirs() -> list[str]:
|
||||
try:
|
||||
from .config import Paths
|
||||
p = Paths.resolve()
|
||||
return [str(p.config_dir), str(p.data_dir), str(p.cache_dir)]
|
||||
except Exception:
|
||||
return []
|
||||
|
||||
def safe_target(self, target: str) -> str:
|
||||
"""Resolve a write target; refuse traversal, symlink escapes and symlink targets."""
|
||||
if not os.path.isabs(target):
|
||||
@@ -61,6 +74,9 @@ class Restorer:
|
||||
normalized = os.path.normpath(target)
|
||||
if normalized != target.rstrip("/") or ".." in Path(target).parts:
|
||||
raise RestoreError("unsafe-path", f"{target} is not a normalized path")
|
||||
for protected in self._protected_dirs():
|
||||
if is_within(normalized, protected):
|
||||
raise RestoreError("unsafe-path", f"{target} is inside versiond's own data directory {protected}")
|
||||
parent = Path(normalized).parent
|
||||
existing = parent
|
||||
while not existing.exists() and existing != existing.parent:
|
||||
@@ -172,7 +188,7 @@ class Restorer:
|
||||
|
||||
# execution
|
||||
|
||||
def execute(self, plan_id: str) -> dict[str, Any]:
|
||||
async def execute(self, plan_id: str) -> dict[str, Any]:
|
||||
row = self.db.one("SELECT * FROM restore_plans WHERE id = ?", (plan_id,))
|
||||
if row is None:
|
||||
raise RestoreError("not-found", f"restore plan {plan_id} does not exist")
|
||||
@@ -187,14 +203,19 @@ class Restorer:
|
||||
results = []
|
||||
for action in plan["actions"]:
|
||||
try:
|
||||
results.append(self._apply(action))
|
||||
results.append(await self._apply(action))
|
||||
except (OSError, RestoreError) as exc:
|
||||
results.append({**action, "result": "error", "error": str(exc)})
|
||||
self.db.execute("UPDATE restore_plans SET status = 'executed' WHERE id = ?", (plan_id,))
|
||||
self.db.audit("restore.execute", plan_id=plan_id, count=len(results))
|
||||
return {"plan_id": plan_id, "results": results}
|
||||
|
||||
def _apply(self, action: dict[str, Any]) -> dict[str, Any]:
|
||||
async def _blob(self, sha256: str) -> bytes:
|
||||
if self.fetch_blob is not None:
|
||||
return await self.fetch_blob(sha256)
|
||||
return self.blobs.get(sha256)
|
||||
|
||||
async def _apply(self, action: dict[str, Any]) -> dict[str, Any]:
|
||||
kind = action["action"]
|
||||
if kind in ("unchanged", "skip-conflict", "skip-not-existing-at-as-of"):
|
||||
return {**action, "result": "skipped"}
|
||||
@@ -205,7 +226,16 @@ class Restorer:
|
||||
self.repo.mark_deleted(destination)
|
||||
return {**action, "result": "removed"}
|
||||
version = self.db.version(action["version_id"])
|
||||
content = self.blobs.get(version["blob_sha256"])
|
||||
try:
|
||||
content = await self._blob(version["blob_sha256"])
|
||||
except (FileNotFoundError, ValueError) as exc:
|
||||
raise RestoreError("unrestorable", f"blob {version['blob_sha256']} is gone or corrupt") from exc
|
||||
if dbsafe.is_sqlite_image(content):
|
||||
# Never write an unverified database back to disk.
|
||||
try:
|
||||
dbsafe.verify_sqlite_bytes(content)
|
||||
except dbsafe.DatabaseUnsafe as exc:
|
||||
raise RestoreError("unrestorable", f"stored version fails integrity check: {exc.reason}")
|
||||
if kind == "write-renamed":
|
||||
stem, ext = os.path.splitext(destination)
|
||||
destination = f"{stem}.restored-{time.strftime('%Y%m%d-%H%M%S')}{ext}"
|
||||
@@ -217,7 +247,7 @@ class Restorer:
|
||||
content, st = self.repo.read_file(path)
|
||||
except (OSError, Exception):
|
||||
return None
|
||||
if self.repo.filters.size_reason(len(content)) or self.repo.filters.content_reason(content):
|
||||
if self.repo.filters.size_reason(len(content)):
|
||||
return None
|
||||
return self.repo.commit(path, content, "restore", "pre-restore", st)["version_id"]
|
||||
|
||||
|
||||
+53
-7
@@ -7,12 +7,37 @@ import os
|
||||
from pathlib import Path
|
||||
|
||||
try:
|
||||
from compression import zstd as _codec # Python 3.14+
|
||||
from compression import zstd as _preferred # Python 3.14+
|
||||
_EXT = ".zst"
|
||||
except ImportError: # pragma: no cover - older Pythons
|
||||
import zlib as _codec
|
||||
_preferred = None
|
||||
_EXT = ".z"
|
||||
|
||||
import zlib as _fallback
|
||||
|
||||
# Back-compat: external code (and old tests) import _codec.
|
||||
_codec = _preferred if _preferred is not None else _fallback
|
||||
|
||||
_EXTS = (".zst", ".z")
|
||||
|
||||
|
||||
def compress(content: bytes) -> bytes:
|
||||
return _codec.compress(content)
|
||||
|
||||
|
||||
def decompress(data: bytes) -> bytes:
|
||||
"""Decompress spool/manifest bytes; raises ValueError uniformly on garbage."""
|
||||
codecs = [ _codec ]
|
||||
if _preferred is not None and _fallback is not _codec:
|
||||
codecs.append(_fallback)
|
||||
last: Exception | None = None
|
||||
for codec in codecs:
|
||||
try:
|
||||
return codec.decompress(data)
|
||||
except Exception as exc: # noqa: BLE001 - normalized below
|
||||
last = exc
|
||||
raise ValueError(f"cannot decompress blob: {last}")
|
||||
|
||||
|
||||
def sha256(content: bytes) -> str:
|
||||
return hashlib.sha256(content).hexdigest()
|
||||
@@ -26,17 +51,28 @@ class BlobStore:
|
||||
def path_for(self, digest: str) -> Path:
|
||||
return self.root / digest[:2] / digest[2:4] / f"{digest}{_EXT}"
|
||||
|
||||
def _candidates(self, digest: str) -> list[Path]:
|
||||
base = self.root / digest[:2] / digest[2:4]
|
||||
return [base / f"{digest}{ext}" for ext in _EXTS]
|
||||
|
||||
def _existing(self, digest: str) -> Path | None:
|
||||
for cand in self._candidates(digest):
|
||||
if cand.exists():
|
||||
return cand
|
||||
return None
|
||||
|
||||
def has(self, digest: str) -> bool:
|
||||
return self.path_for(digest).exists()
|
||||
return self._existing(digest) is not None
|
||||
|
||||
def put(self, content: bytes, digest: str | None = None) -> tuple[str, int]:
|
||||
"""Store content; returns (digest, stored_size). Idempotent."""
|
||||
digest = digest or sha256(content)
|
||||
found = self._existing(digest)
|
||||
if found is not None:
|
||||
return digest, found.stat().st_size
|
||||
target = self.path_for(digest)
|
||||
if target.exists():
|
||||
return digest, target.stat().st_size
|
||||
target.parent.mkdir(parents=True, exist_ok=True)
|
||||
data = _codec.compress(content)
|
||||
data = compress(content)
|
||||
tmp = target.with_name(f".{target.name}.{os.getpid()}.tmp")
|
||||
with open(tmp, "wb") as fh:
|
||||
fh.write(data)
|
||||
@@ -46,4 +82,14 @@ class BlobStore:
|
||||
return digest, len(data)
|
||||
|
||||
def get(self, digest: str) -> bytes:
|
||||
return _codec.decompress(self.path_for(digest).read_bytes())
|
||||
found = self._existing(digest)
|
||||
if found is None:
|
||||
raise FileNotFoundError(self.path_for(digest))
|
||||
return decompress(found.read_bytes())
|
||||
|
||||
def read_stored(self, digest: str) -> bytes:
|
||||
"""Raw (compressed) bytes as stored, regardless of codec extension."""
|
||||
found = self._existing(digest)
|
||||
if found is None:
|
||||
raise FileNotFoundError(self.path_for(digest))
|
||||
return found.read_bytes()
|
||||
|
||||
+22
-17
@@ -14,10 +14,9 @@ from typing import Any
|
||||
import httpx
|
||||
|
||||
from .config import Config
|
||||
from .crypto import KeyRing
|
||||
from .db import Database
|
||||
from .remote import RemoteError, RemoteSettings, WebDAV, check_connection, claim_directory, join
|
||||
from .store import BlobStore, _codec
|
||||
from .store import BlobStore, compress as _compress
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
@@ -48,12 +47,11 @@ class Throttle:
|
||||
|
||||
|
||||
class Uploader:
|
||||
def __init__(self, cfg: Config, db: Database, blobs: BlobStore, keys: KeyRing,
|
||||
def __init__(self, cfg: Config, db: Database, blobs: BlobStore,
|
||||
transport: httpx.AsyncBaseTransport | None = None):
|
||||
self.cfg = cfg
|
||||
self.db = db
|
||||
self.blobs = blobs
|
||||
self.keys = keys
|
||||
self.transport = transport
|
||||
self.dav: WebDAV | None = None
|
||||
self.task: asyncio.Task | None = None
|
||||
@@ -71,6 +69,8 @@ class Uploader:
|
||||
self._last_manifest = 0.0
|
||||
self._wake = asyncio.Event()
|
||||
self.offline_sleep = OFFLINE_SLEEP
|
||||
self.remote_usage: dict[str, Any] = {"percent": None, "used": None, "total": None,
|
||||
"source": "unknown", "at": None}
|
||||
|
||||
# settings
|
||||
|
||||
@@ -91,7 +91,7 @@ class Uploader:
|
||||
"configured": s.configured, "url": s.url, "username": s.username,
|
||||
"password_set": bool(s.password), "base_path": s.base_path, "directory": s.directory,
|
||||
"verify_tls": s.verify_tls, "timeout_seconds": s.timeout_seconds,
|
||||
"encryption": "aes-256-gcm", "key_id": self.keys.key_id, "state": self.state,
|
||||
"encryption": "none", "state": self.state,
|
||||
}
|
||||
|
||||
async def test(self, settings: RemoteSettings) -> None:
|
||||
@@ -106,7 +106,7 @@ class Uploader:
|
||||
dav = WebDAV(settings, transport=self.transport)
|
||||
try:
|
||||
await check_connection(dav)
|
||||
claim = await claim_directory(dav, self.keys.key_id)
|
||||
claim = await claim_directory(dav)
|
||||
finally:
|
||||
await dav.close()
|
||||
await self.stop()
|
||||
@@ -156,7 +156,7 @@ class Uploader:
|
||||
async def _ensure_directory(self) -> None:
|
||||
assert self.dav is not None
|
||||
if not self.directory:
|
||||
claim = await claim_directory(self.dav, self.keys.key_id)
|
||||
claim = await claim_directory(self.dav)
|
||||
self.directory = claim["directory"]
|
||||
self.cfg.update("remote", {"directory": self.directory})
|
||||
await self.dav.makedirs(join(self.directory, "blobs"))
|
||||
@@ -182,7 +182,7 @@ class Uploader:
|
||||
raise
|
||||
except RemoteError as exc:
|
||||
self._failed(exc)
|
||||
self.state = "error" if exc.code in ("remote-auth", "key-mismatch", "remote-full") else "offline"
|
||||
self.state = "error" if exc.code in ("remote-auth", "remote-full") else "offline"
|
||||
log.warning("remote %s: %s; retrying in %ss", self.state, exc.detail, self.offline_sleep)
|
||||
await asyncio.sleep(self.offline_sleep)
|
||||
except Exception as exc:
|
||||
@@ -208,8 +208,11 @@ class Uploader:
|
||||
fatal: list[RemoteError] = []
|
||||
|
||||
async def worker() -> None:
|
||||
while not queue.empty() and not fatal:
|
||||
row = queue.get_nowait()
|
||||
while not fatal:
|
||||
try:
|
||||
row = queue.get_nowait()
|
||||
except asyncio.QueueEmpty:
|
||||
return
|
||||
await requests.take()
|
||||
try:
|
||||
await self._upload_blob(row, bandwidth)
|
||||
@@ -226,14 +229,14 @@ class Uploader:
|
||||
assert self.dav is not None
|
||||
digest = row["sha256"]
|
||||
try:
|
||||
stored = self.blobs.path_for(digest).read_bytes()
|
||||
stored = self.blobs.read_stored(digest)
|
||||
except FileNotFoundError:
|
||||
self.db.execute("UPDATE blobs SET remote_state = 'missing-local', last_error = ? WHERE sha256 = ?",
|
||||
("local blob file is missing", digest))
|
||||
return
|
||||
name = self.keys.blob_name(digest)
|
||||
name = digest
|
||||
prefix = join(self.directory, "blobs", name[:2])
|
||||
payload = self.keys.encrypt(stored)
|
||||
payload = stored
|
||||
await bandwidth.take(len(payload))
|
||||
try:
|
||||
await self.dav.mkcol(prefix)
|
||||
@@ -266,7 +269,8 @@ class Uploader:
|
||||
assert self.dav is not None
|
||||
self._last_manifest = time.monotonic()
|
||||
versions = self.db.all(
|
||||
"SELECT v.id, f.path, v.blob_sha256 AS sha256, b.size, v.captured_at, v.source, v.reason, v.pinned "
|
||||
"SELECT v.id, f.path, v.blob_sha256 AS sha256, b.size, b.stored_size, v.captured_at, "
|
||||
"v.source, v.reason, v.pinned "
|
||||
"FROM versions v JOIN files f ON f.id = v.file_id JOIN blobs b ON b.sha256 = v.blob_sha256 "
|
||||
"WHERE v.durability = 'local' AND b.remote_state = 'uploaded' ORDER BY v.id LIMIT ?",
|
||||
(MANIFEST_MAX_RECORDS,),
|
||||
@@ -280,12 +284,12 @@ class Uploader:
|
||||
return False
|
||||
stamp = time.gmtime()
|
||||
batch_id = time.strftime("%Y%m%dT%H%M%SZ", stamp) + "-" + secrets.token_hex(4)
|
||||
lines = [json.dumps({"type": "version", **v, "blob": self.keys.blob_name(v["sha256"])}) for v in versions]
|
||||
lines = [json.dumps({"type": "version", **v, "blob": v["sha256"]}) for v in versions]
|
||||
lines += [json.dumps({"type": "rename", **r}) for r in renames]
|
||||
payload = self.keys.encrypt(_codec.compress(("\n".join(lines) + "\n").encode()))
|
||||
payload = _compress(("\n".join(lines) + "\n").encode())
|
||||
folder = join(self.directory, "manifests", time.strftime("%Y/%m/%d", stamp))
|
||||
await self.dav.makedirs(folder)
|
||||
await self.dav.put(join(folder, f"{batch_id}.jsonl.zst.enc"), payload)
|
||||
await self.dav.put(join(folder, f"{batch_id}.jsonl.zst"), payload)
|
||||
with self.db.transaction():
|
||||
for chunk in _chunks([v["id"] for v in versions], 500):
|
||||
self.db.execute(
|
||||
@@ -340,6 +344,7 @@ class Uploader:
|
||||
"last_error": self.last_error,
|
||||
"last_error_at": self.last_error_at,
|
||||
"consecutive_failures": self.consecutive_failures,
|
||||
"usage": dict(self.remote_usage),
|
||||
"session": dict(self.counters),
|
||||
}
|
||||
|
||||
|
||||
@@ -0,0 +1,109 @@
|
||||
"""Bone-level proof: temp files and unfinished writes are never versioned.
|
||||
|
||||
Runs against the real inotify monitor through the HTTP API (same fixture
|
||||
shape as test_service). Each test fails if a single temp/unfinished state
|
||||
is ever stored.
|
||||
"""
|
||||
|
||||
import base64
|
||||
import os
|
||||
import time
|
||||
|
||||
import pytest
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from versiond.api import create_app
|
||||
from versiond.config import Config, Paths
|
||||
|
||||
from test_service import CONFIG, history, wait_for
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def env(tmp_path, monkeypatch):
|
||||
monkeypatch.setenv("VERSIOND_HOME", str(tmp_path / "home"))
|
||||
paths = Paths.resolve()
|
||||
paths.ensure()
|
||||
paths.config_file.write_text(CONFIG)
|
||||
cfg = Config.load(paths)
|
||||
work = tmp_path / "work"
|
||||
(work / "src").mkdir(parents=True)
|
||||
(work / "src" / "app.py").write_text("v0\n")
|
||||
with TestClient(create_app(cfg), base_url="http://127.0.0.1:9922") as client:
|
||||
client.headers["Authorization"] = f"Bearer {cfg.api_token()}"
|
||||
yield client, work
|
||||
|
||||
|
||||
def _add_root(client, work):
|
||||
r = client.post("/api/v1/roots", json={"path": str(work)})
|
||||
assert r.status_code == 201, r.text
|
||||
root = r.json()
|
||||
wait_for(lambda: client.get(f"/api/v1/roots/{root['id']}").json()["baseline_state"] == "done")
|
||||
return root
|
||||
|
||||
|
||||
def test_temp_names_never_stored(env):
|
||||
client, work = env
|
||||
_add_root(client, work)
|
||||
temps = ["a.tmp", "b.temp", "c.swp", "d.kate-swp", "e~", "f.part",
|
||||
"g.crdownload", "~$lock.docx", "#autosave#", "4913", "h.orig"]
|
||||
for name in temps:
|
||||
(work / "src" / name).write_text("unfinished\n")
|
||||
time.sleep(0.8) # past the 0.2s coalescing window: anything stored would show
|
||||
for name in temps:
|
||||
assert history(client, work / "src" / name) == [], name
|
||||
# Same via the push API: every one must be refused, none stored.
|
||||
for name in temps:
|
||||
body = {"path": str(work / "src" / name),
|
||||
"content_b64": base64.b64encode(b"unfinished\n").decode()}
|
||||
assert client.post("/api/v1/snapshots", json=body).status_code == 422, name
|
||||
|
||||
|
||||
def test_open_file_not_captured_until_close(env):
|
||||
client, work = env
|
||||
_add_root(client, work)
|
||||
target = work / "src" / "slow.py"
|
||||
fh = open(target, "w")
|
||||
try:
|
||||
fh.write("partial-1\n")
|
||||
fh.flush()
|
||||
os.fsync(fh.fileno()) # bytes on disk, but writer still holds it open
|
||||
time.sleep(0.6)
|
||||
assert history(client, target) == [] # no CLOSE_WRITE yet: nothing stored
|
||||
fh.write("partial-2\n")
|
||||
fh.flush()
|
||||
finally:
|
||||
fh.close() # only now is the state finished and capturable
|
||||
wait_for(lambda: history(client, target))
|
||||
versions = history(client, target)
|
||||
assert len(versions) == 1
|
||||
assert client.get(f"/api/v1/versions/{versions[0]['id']}/content").text == "partial-1\npartial-2\n"
|
||||
|
||||
|
||||
def test_burst_collapses_to_first_and_last(env):
|
||||
client, work = env
|
||||
_add_root(client, work)
|
||||
target = work / "src" / "busy.py"
|
||||
target.write_text("v0\n")
|
||||
wait_for(lambda: len(history(client, target)) == 1) # first observed state committed
|
||||
for i in range(1, 6):
|
||||
target.write_text(f"v{i}\n") # burst inside the coalescing window
|
||||
wait_for(lambda: len(history(client, target)) == 2, timeout=8)
|
||||
time.sleep(1.5) # past max_hold: no third version may appear
|
||||
versions = history(client, target)
|
||||
assert len(versions) == 2
|
||||
bodies = [client.get(f"/api/v1/versions/{v['id']}/content").text for v in reversed(versions)]
|
||||
assert bodies == ["v0\n", "v5\n"]
|
||||
|
||||
|
||||
def test_atomic_save_is_one_version(env):
|
||||
client, work = env
|
||||
_add_root(client, work)
|
||||
target = work / "src" / "app.py"
|
||||
wait_for(lambda: history(client, target))
|
||||
tmp = work / "src" / "app.py.tmp"
|
||||
tmp.write_text("atomic\n")
|
||||
os.replace(tmp, target)
|
||||
wait_for(lambda: len(history(client, target)) == 2)
|
||||
time.sleep(0.6)
|
||||
assert len(history(client, target)) == 2 # temp state never stored
|
||||
assert history(client, work / "src" / "app.py.tmp") == []
|
||||
@@ -0,0 +1,126 @@
|
||||
"""Zero-error policy for live databases: only verified snapshots are stored."""
|
||||
|
||||
import sqlite3
|
||||
|
||||
import pytest
|
||||
|
||||
from versiond import dbsafe
|
||||
from versiond.db import Database
|
||||
from versiond.filters import Filters
|
||||
from versiond.ingest import Rejected, Repository
|
||||
from versiond.store import BlobStore
|
||||
|
||||
|
||||
def _repo(tmp_path):
|
||||
db = Database(tmp_path / "index.sqlite")
|
||||
return Repository(db, BlobStore(tmp_path / "spool"), Filters())
|
||||
|
||||
|
||||
def _open_db(path, **kw):
|
||||
conn = sqlite3.connect(path, **kw)
|
||||
conn.execute("CREATE TABLE IF NOT EXISTS t(id INTEGER PRIMARY KEY, v TEXT)")
|
||||
return conn
|
||||
|
||||
|
||||
def _read_snapshot_bytes(repo, path):
|
||||
data, _ = repo.read_file(path)
|
||||
return data
|
||||
|
||||
|
||||
def _rows_in_snapshot(data):
|
||||
import tempfile, os
|
||||
fd, tmp = tempfile.mkstemp(suffix=".sqlite")
|
||||
try:
|
||||
with os.fdopen(fd, "wb") as fh:
|
||||
fh.write(data)
|
||||
conn = sqlite3.connect(f"file:{tmp}?mode=ro", uri=True)
|
||||
try:
|
||||
assert conn.execute("PRAGMA integrity_check").fetchone()[0] == "ok"
|
||||
return conn.execute("SELECT v FROM t ORDER BY id").fetchall()
|
||||
finally:
|
||||
conn.close()
|
||||
finally:
|
||||
os.unlink(tmp)
|
||||
|
||||
|
||||
def test_sqlite_snapshot_roundtrip(tmp_path):
|
||||
repo = _repo(tmp_path)
|
||||
db_path = str(tmp_path / "app.db")
|
||||
conn = _open_db(db_path)
|
||||
conn.execute("INSERT INTO t(v) VALUES ('a'), ('b')")
|
||||
conn.commit()
|
||||
data = _read_snapshot_bytes(repo, db_path)
|
||||
assert _rows_in_snapshot(data) == [("a",), ("b",)]
|
||||
result = repo.commit(db_path, data, "test", "edit")
|
||||
assert result["status"] == "committed"
|
||||
conn.close()
|
||||
|
||||
|
||||
def test_snapshot_during_uncommitted_write_is_consistent(tmp_path):
|
||||
"""A writer with an open transaction must not produce a torn version."""
|
||||
repo = _repo(tmp_path)
|
||||
db_path = str(tmp_path / "live.db")
|
||||
writer = _open_db(db_path, isolation_level=None)
|
||||
writer.execute("INSERT INTO t(v) VALUES ('committed')")
|
||||
writer.execute("BEGIN IMMEDIATE")
|
||||
writer.execute("INSERT INTO t(v) VALUES ('uncommitted')")
|
||||
try:
|
||||
data = _read_snapshot_bytes(repo, db_path)
|
||||
finally:
|
||||
writer.execute("ROLLBACK")
|
||||
writer.close()
|
||||
rows = _rows_in_snapshot(data)
|
||||
assert rows == [("committed",)] # committed state only, verified clean
|
||||
|
||||
|
||||
def test_wal_mode_snapshot(tmp_path):
|
||||
repo = _repo(tmp_path)
|
||||
db_path = str(tmp_path / "wal.db")
|
||||
conn = _open_db(db_path)
|
||||
conn.execute("PRAGMA journal_mode=WAL")
|
||||
conn.execute("INSERT INTO t(v) VALUES ('w1')")
|
||||
conn.commit()
|
||||
conn.execute("INSERT INTO t(v) VALUES ('w2')") # may sit in -wal
|
||||
data = _read_snapshot_bytes(repo, db_path)
|
||||
conn.close()
|
||||
rows = _rows_in_snapshot(data)
|
||||
assert ("w1",) in rows # committed data present, snapshot verified
|
||||
|
||||
|
||||
def test_corrupt_sqlite_magic_rejected(tmp_path):
|
||||
repo = _repo(tmp_path)
|
||||
bad = tmp_path / "bad.db"
|
||||
bad.write_bytes(dbsafe.SQLITE_MAGIC + b"\x00" * 100 + b"garbage")
|
||||
with pytest.raises(Rejected) as exc:
|
||||
repo.read_file(str(bad))
|
||||
assert exc.value.reason in ("database-corrupt", "database-unreadable")
|
||||
|
||||
|
||||
def test_locked_database_rejected_not_stored(tmp_path):
|
||||
repo = _repo(tmp_path)
|
||||
db_path = str(tmp_path / "locked.db")
|
||||
holder = _open_db(db_path, isolation_level=None)
|
||||
holder.execute("INSERT INTO t(v) VALUES ('x')")
|
||||
holder.execute("BEGIN EXCLUSIVE")
|
||||
try:
|
||||
with pytest.raises(dbsafe.DatabaseUnsafe) as exc:
|
||||
dbsafe.snapshot_sqlite(db_path, 10_485_760, timeout=0.3)
|
||||
assert exc.value.reason == "database-locked"
|
||||
finally:
|
||||
holder.execute("ROLLBACK")
|
||||
holder.close()
|
||||
# After the lock clears, the same file snapshots fine.
|
||||
assert _rows_in_snapshot(_read_snapshot_bytes(repo, db_path)) == [("x",)]
|
||||
|
||||
|
||||
def test_pushed_sqlite_bytes_verified(tmp_path):
|
||||
repo = _repo(tmp_path)
|
||||
db_path = str(tmp_path / "push.db")
|
||||
with pytest.raises(Rejected):
|
||||
repo.check_content(dbsafe.SQLITE_MAGIC + b"\x00" * 100 + b"garbage")
|
||||
# Genuine DB bytes pass.
|
||||
conn = _open_db(db_path)
|
||||
conn.execute("INSERT INTO t(v) VALUES ('ok')")
|
||||
conn.commit()
|
||||
conn.close()
|
||||
repo.check_content((tmp_path / "push.db").read_bytes())
|
||||
@@ -0,0 +1,103 @@
|
||||
"""Scheduler, quota signal and capacity backstop."""
|
||||
|
||||
import asyncio
|
||||
import time
|
||||
|
||||
import httpx
|
||||
import pytest
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from versiond import maintenance
|
||||
from versiond.api import create_app
|
||||
from versiond.config import Config, Paths
|
||||
from versiond.remote import quota as remote_quota
|
||||
|
||||
from test_service import CONFIG, history, wait_for
|
||||
from test_purge import FakeDAV, REMOTE
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def maint_env(tmp_path, monkeypatch):
|
||||
monkeypatch.setenv("VERSIOND_HOME", str(tmp_path / "home"))
|
||||
paths = Paths.resolve()
|
||||
paths.ensure()
|
||||
paths.config_file.write_text(CONFIG + "\n[upload]\nmanifest_interval_seconds = 0.2\n")
|
||||
cfg = Config.load(paths)
|
||||
dav = FakeDAV()
|
||||
project = tmp_path / "proj"
|
||||
project.mkdir()
|
||||
(project / "main.py").write_text("print('hello')\n")
|
||||
app = create_app(cfg, remote_transport=httpx.MockTransport(dav.handler))
|
||||
with TestClient(app, base_url="http://127.0.0.1:9922") as client:
|
||||
client.headers["Authorization"] = f"Bearer {cfg.api_token()}"
|
||||
app.state.services.uploader.offline_sleep = 0.3
|
||||
yield client, dav, project, cfg, app.state.services
|
||||
|
||||
|
||||
def test_quota_signal_and_fallback(maint_env):
|
||||
client, dav, project, cfg, svc = maint_env
|
||||
directory = client.put("/api/v1/config/remote", json=REMOTE).json()["directory"]
|
||||
|
||||
async def get_quota():
|
||||
from versiond.remote import WebDAV
|
||||
settings = svc.uploader.settings()
|
||||
dav_client = WebDAV(settings, transport=svc.uploader.transport)
|
||||
try:
|
||||
return await remote_quota(dav_client, directory)
|
||||
finally:
|
||||
await dav_client.close()
|
||||
|
||||
used, available = asyncio.run(get_quota())
|
||||
assert used is not None and used >= 0 and available == dav.capacity_bytes - used
|
||||
|
||||
async def refresh():
|
||||
return await maintenance.refresh_usage(svc)
|
||||
|
||||
usage = asyncio.run(refresh())
|
||||
assert usage["source"] == "webdav-quota"
|
||||
assert usage["percent"] == round(100 * used / dav.capacity_bytes, 1)
|
||||
assert maintenance.over_limit(svc) is False
|
||||
assert client.get("/health").json()["remote_pressure"] is False
|
||||
|
||||
|
||||
def test_pressure_triggers_pass_but_floor_holds(maint_env):
|
||||
client, dav, project, cfg, svc = maint_env
|
||||
directory = client.put("/api/v1/config/remote", json=REMOTE).json()["directory"]
|
||||
client.post("/api/v1/roots", json={"path": str(project)})
|
||||
wait_for(lambda: history(client, project / "main.py"), timeout=10)
|
||||
# Tiny remote: everything ever uploaded already exceeds 70%.
|
||||
dav.capacity_bytes = 1
|
||||
|
||||
result = asyncio.run(maintenance.run_scheduled_pass(svc, "test-pressure"))
|
||||
assert result["pressure"] is True
|
||||
assert result["versions_deleted"] == 0 # all young: the 5-day floor holds
|
||||
assert len(history(client, project / "main.py")) == 1
|
||||
assert client.get("/health").json()["remote_pressure"] is True
|
||||
assert "remote-over-capacity-limit" in client.get("/health").json()["degraded_roots"]
|
||||
upload = client.get("/api/v1/progress").json()["upload"]
|
||||
assert upload["usage"]["source"] == "webdav-quota"
|
||||
assert upload["usage"]["percent"] >= 70
|
||||
|
||||
|
||||
def test_scheduled_pass_cleans_old_and_reports(maint_env):
|
||||
import sqlite3
|
||||
client, dav, project, cfg, svc = maint_env
|
||||
client.put("/api/v1/config/remote", json=REMOTE)
|
||||
client.post("/api/v1/roots", json={"path": str(project)})
|
||||
wait_for(lambda: history(client, project / "main.py"), timeout=10)
|
||||
(project / "main.py").write_text("v2\n")
|
||||
wait_for(lambda: len(history(client, project / "main.py")) == 2, timeout=10)
|
||||
import time as _t
|
||||
_t.sleep(0.5)
|
||||
(project / "main.py").write_text("v3\n")
|
||||
wait_for(lambda: len(history(client, project / "main.py")) == 3, timeout=10)
|
||||
old = time.time() - 10 * 86400
|
||||
conn = sqlite3.connect(cfg.paths.index_file, timeout=5.0)
|
||||
vids = [v["id"] for v in reversed(history(client, project / "main.py"))]
|
||||
for vid in vids[:2]: # two oldest share one old day: keep daily + latest
|
||||
conn.execute("UPDATE versions SET captured_at = ? WHERE id = ?", (old, vid))
|
||||
conn.commit()
|
||||
conn.close()
|
||||
result = asyncio.run(maintenance.run_scheduled_pass(svc, "test-scheduled"))
|
||||
assert result["versions_deleted"] == 1
|
||||
assert len(history(client, project / "main.py")) == 2
|
||||
@@ -0,0 +1,198 @@
|
||||
"""Purge-remote: date-guarded, async, remote-only deletion of this system's backup."""
|
||||
|
||||
import base64
|
||||
import json
|
||||
import time
|
||||
from datetime import datetime
|
||||
|
||||
import httpx
|
||||
import pytest
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from versiond.api import create_app
|
||||
from versiond.config import Config, Paths
|
||||
|
||||
from test_service import CONFIG, history, wait_for
|
||||
|
||||
|
||||
class FakeDAV:
|
||||
"""In-memory WebDAV with PROPFIND Depth:0/1 (XML) and recursive DELETE."""
|
||||
|
||||
def __init__(self, user="u1", password="secret"):
|
||||
self.auth = "Basic " + base64.b64encode(f"{user}:{password}".encode()).decode()
|
||||
self.dirs = {"/"}
|
||||
self.files: dict[str, bytes] = {}
|
||||
self.down = False
|
||||
self.capacity_bytes = 10 ** 12 # tiny servers set this low to simulate pressure
|
||||
|
||||
def _usage(self, path):
|
||||
used = sum(len(v) for k, v in self.files.items() if k == path or k.startswith(path + "/"))
|
||||
return used, max(0, self.capacity_bytes - used)
|
||||
|
||||
@staticmethod
|
||||
def parent(path):
|
||||
return path.rsplit("/", 1)[0] or "/"
|
||||
|
||||
def _children(self, path):
|
||||
kids = []
|
||||
for d in sorted(self.dirs):
|
||||
if d != path and self.parent(d) == path:
|
||||
kids.append((d, True))
|
||||
for f in sorted(self.files):
|
||||
if self.parent(f) == path:
|
||||
kids.append((f, False))
|
||||
return kids
|
||||
|
||||
def _multistatus(self, path):
|
||||
parts = ['<?xml version="1.0"?><d:multistatus xmlns:d="DAV:">']
|
||||
parts.append(f"<d:response><d:href>{path}</d:href><d:propstat><d:prop>"
|
||||
"<d:resourcetype><d:collection/></d:resourcetype>"
|
||||
"</d:prop><d:status>HTTP/1.1 200 OK</d:status></d:propstat></d:response>")
|
||||
for child, is_dir in self._children(path):
|
||||
rtype = "<d:resourcetype><d:collection/></d:resourcetype>" if is_dir else "<d:resourcetype/>"
|
||||
size = "" if is_dir else f"<d:getcontentlength>{len(self.files[child])}</d:getcontentlength>"
|
||||
parts.append(f"<d:response><d:href>{child}</d:href><d:propstat><d:prop>"
|
||||
f"{rtype}{size}</d:prop><d:status>HTTP/1.1 200 OK</d:status>"
|
||||
"</d:propstat></d:response>")
|
||||
parts.append("</d:multistatus>")
|
||||
return "".join(parts).encode()
|
||||
|
||||
def handler(self, request: httpx.Request) -> httpx.Response:
|
||||
if self.down:
|
||||
raise httpx.ConnectError("connection refused", request=request)
|
||||
if request.headers.get("authorization") != self.auth:
|
||||
return httpx.Response(401)
|
||||
path = request.url.path.rstrip("/") or "/"
|
||||
method = request.method
|
||||
if method == "PROPFIND":
|
||||
if path in self.files:
|
||||
body = (f'<?xml version="1.0"?><d:multistatus xmlns:d="DAV:">'
|
||||
f"<d:response><d:href>{path}</d:href><d:propstat><d:prop>"
|
||||
f"<d:resourcetype/><d:getcontentlength>{len(self.files[path])}</d:getcontentlength>"
|
||||
"</d:prop><d:status>HTTP/1.1 200 OK</d:status></d:propstat></d:response>"
|
||||
"</d:multistatus>").encode()
|
||||
return httpx.Response(207, content=body)
|
||||
if path not in self.dirs:
|
||||
return httpx.Response(404)
|
||||
if request.headers.get("depth", "0") == "1":
|
||||
return httpx.Response(207, content=self._multistatus(path),
|
||||
headers={"Content-Type": "application/xml"})
|
||||
used, available = self._usage(path)
|
||||
body = (f'<?xml version="1.0"?><d:multistatus xmlns:d="DAV:">'
|
||||
f"<d:response><d:href>{path}</d:href><d:propstat><d:prop>"
|
||||
f"<d:quota-used-bytes>{used}</d:quota-used-bytes>"
|
||||
f"<d:quota-available-bytes>{available}</d:quota-available-bytes>"
|
||||
"</d:prop><d:status>HTTP/1.1 200 OK</d:status></d:propstat></d:response>"
|
||||
"</d:multistatus>").encode()
|
||||
return httpx.Response(207, content=body)
|
||||
if method == "MKCOL":
|
||||
if path in self.dirs:
|
||||
return httpx.Response(405)
|
||||
if self.parent(path) not in self.dirs:
|
||||
return httpx.Response(409)
|
||||
self.dirs.add(path)
|
||||
return httpx.Response(201)
|
||||
if method == "PUT":
|
||||
if self.parent(path) not in self.dirs:
|
||||
return httpx.Response(409)
|
||||
if request.headers.get("if-none-match") == "*" and path in self.files:
|
||||
return httpx.Response(412)
|
||||
self.files[path] = request.read()
|
||||
return httpx.Response(201)
|
||||
if method == "GET":
|
||||
return httpx.Response(200, content=self.files[path]) if path in self.files else httpx.Response(404)
|
||||
if method == "DELETE":
|
||||
for f in [f for f in self.files if f == path or f.startswith(path + "/")]:
|
||||
del self.files[f]
|
||||
for d in [d for d in self.dirs if d != "/" and (d == path or d.startswith(path + "/"))]:
|
||||
self.dirs.discard(d)
|
||||
return httpx.Response(204)
|
||||
return httpx.Response(405)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def purge_env(tmp_path, monkeypatch):
|
||||
monkeypatch.setenv("VERSIOND_HOME", str(tmp_path / "home"))
|
||||
paths = Paths.resolve()
|
||||
paths.ensure()
|
||||
paths.config_file.write_text(CONFIG + "\n[upload]\nmanifest_interval_seconds = 0.2\n")
|
||||
cfg = Config.load(paths)
|
||||
dav = FakeDAV()
|
||||
project = tmp_path / "proj"
|
||||
project.mkdir()
|
||||
(project / "main.py").write_text("print('hello')\n")
|
||||
app = create_app(cfg, remote_transport=httpx.MockTransport(dav.handler))
|
||||
with TestClient(app, base_url="http://127.0.0.1:9922") as client:
|
||||
client.headers["Authorization"] = f"Bearer {cfg.api_token()}"
|
||||
app.state.services.uploader.offline_sleep = 0.3
|
||||
yield client, dav, project
|
||||
|
||||
|
||||
REMOTE = {"url": "https://dav.example", "username": "u1", "password": "secret", "base_path": "/versioned/"}
|
||||
|
||||
|
||||
def _configure_and_fill(client, project):
|
||||
r = client.put("/api/v1/config/remote", json=REMOTE)
|
||||
assert r.status_code == 200, r.text
|
||||
directory = r.json()["directory"]
|
||||
client.post("/api/v1/roots", json={"path": str(project)})
|
||||
wait_for(lambda: (lambda p: p["upload"]["versions_local"] == 0
|
||||
and p["upload"]["versions_durable"] >= 1)(
|
||||
client.get("/api/v1/progress").json()), timeout=10)
|
||||
return directory
|
||||
|
||||
|
||||
def test_purge_rejects_bad_date(purge_env):
|
||||
client, _, project = purge_env
|
||||
_configure_and_fill(client, project)
|
||||
r = client.post("/api/v1/admin/purge-remote", json={"today": "01-01-2000"})
|
||||
assert r.status_code == 400, r.text
|
||||
r = client.post("/api/v1/admin/purge-remote", json={"today": "today"})
|
||||
assert r.status_code == 400, r.text
|
||||
|
||||
|
||||
def test_purge_unconfigured_is_409(tmp_path, monkeypatch):
|
||||
monkeypatch.setenv("VERSIOND_HOME", str(tmp_path / "home2"))
|
||||
paths = Paths.resolve()
|
||||
paths.ensure()
|
||||
paths.config_file.write_text(CONFIG)
|
||||
cfg = Config.load(paths)
|
||||
app = create_app(cfg)
|
||||
with TestClient(app, base_url="http://127.0.0.1:9922") as client:
|
||||
client.headers["Authorization"] = f"Bearer {cfg.api_token()}"
|
||||
today = datetime.now().strftime("%d-%m-%Y")
|
||||
r = client.post("/api/v1/admin/purge-remote", json={"today": today})
|
||||
assert r.status_code == 409, r.text
|
||||
|
||||
|
||||
def test_purge_deletes_only_this_system(purge_env):
|
||||
client, dav, project = purge_env
|
||||
directory = _configure_and_fill(client, project)
|
||||
before = [p for p in dav.files
|
||||
if p.startswith(directory + "/blobs/") or p.startswith(directory + "/manifests/")]
|
||||
assert before, "expected uploaded blobs/manifests before purge"
|
||||
|
||||
today = datetime.now().strftime("%d-%m-%Y")
|
||||
expected_sizes = {p: len(dav.files[p]) for p in before}
|
||||
r = client.post("/api/v1/admin/purge-remote", json={"today": today})
|
||||
assert r.status_code == 202, r.text
|
||||
body = r.json()
|
||||
assert body["directory"] == directory
|
||||
assert body["files"] == len(before) and body["files"] > 0
|
||||
assert body["bytes"] == sum(expected_sizes.values())
|
||||
|
||||
task_id = body["task_id"]
|
||||
|
||||
def done():
|
||||
s = client.get(f"/api/v1/admin/purge-remote/{task_id}").json()
|
||||
return s if s["status"] == "done" else None
|
||||
|
||||
status = wait_for(done, timeout=10)
|
||||
assert status["files_deleted"] == body["files"]
|
||||
assert status["errors"] == []
|
||||
assert [p for p in dav.files if p.startswith(directory + "/blobs/")] == []
|
||||
assert [p for p in dav.files if p.startswith(directory + "/manifests/")] == []
|
||||
# meta (claim) is kept; local history is untouched
|
||||
assert directory + "/meta/owner.json" in dav.files
|
||||
assert history(client, project / "main.py"), "local index must be untouched"
|
||||
assert client.get(f"/api/v1/admin/purge-remote/nope").status_code == 404
|
||||
@@ -0,0 +1,197 @@
|
||||
"""Disaster recovery, retention, GC, adopt, metrics, verified restores."""
|
||||
|
||||
import json
|
||||
import sqlite3
|
||||
import time
|
||||
from datetime import datetime
|
||||
|
||||
import httpx
|
||||
import pytest
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from versiond import dbsafe
|
||||
from versiond.api import create_app
|
||||
from versiond.config import Config, Paths
|
||||
from versiond.store import BlobStore
|
||||
|
||||
from test_service import CONFIG, history, wait_for
|
||||
from test_purge import FakeDAV, REMOTE
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def rec_env(tmp_path, monkeypatch):
|
||||
monkeypatch.setenv("VERSIOND_HOME", str(tmp_path / "home"))
|
||||
paths = Paths.resolve()
|
||||
paths.ensure()
|
||||
paths.config_file.write_text(CONFIG + "\n[upload]\nmanifest_interval_seconds = 0.2\n")
|
||||
cfg = Config.load(paths)
|
||||
dav = FakeDAV()
|
||||
project = tmp_path / "proj"
|
||||
project.mkdir()
|
||||
(project / "main.py").write_text("print('hello')\n")
|
||||
conn = sqlite3.connect(project / "app.db")
|
||||
conn.execute("CREATE TABLE t(id INTEGER PRIMARY KEY, v TEXT)")
|
||||
conn.execute("INSERT INTO t(v) VALUES ('one')")
|
||||
conn.commit()
|
||||
conn.close()
|
||||
app = create_app(cfg, remote_transport=httpx.MockTransport(dav.handler))
|
||||
with TestClient(app, base_url="http://127.0.0.1:9922") as client:
|
||||
client.headers["Authorization"] = f"Bearer {cfg.api_token()}"
|
||||
app.state.services.uploader.offline_sleep = 0.3
|
||||
yield client, dav, project, cfg
|
||||
|
||||
|
||||
def _configure_and_fill(client, project, files=2):
|
||||
r = client.put("/api/v1/config/remote", json=REMOTE)
|
||||
assert r.status_code == 200, r.text
|
||||
directory = r.json()["directory"]
|
||||
client.post("/api/v1/roots", json={"path": str(project)})
|
||||
wait_for(lambda: (lambda p: p["upload"]["versions_local"] == 0
|
||||
and p["upload"]["versions_durable"] == files)(
|
||||
client.get("/api/v1/progress").json()), timeout=15)
|
||||
return directory
|
||||
|
||||
|
||||
def _db_index(cfg):
|
||||
return sqlite3.connect(cfg.paths.index_file, timeout=5.0)
|
||||
|
||||
|
||||
def test_reindex_rebuilds_lost_index(rec_env):
|
||||
client, dav, project, cfg = rec_env
|
||||
directory = _configure_and_fill(client, project)
|
||||
before = {v["id"]: v for v in history(client, project / "main.py")}
|
||||
assert before
|
||||
|
||||
# Simulate total local loss of metadata (spool stays, like a surviving disk).
|
||||
conn = _db_index(cfg)
|
||||
conn.execute("DELETE FROM versions")
|
||||
conn.execute("DELETE FROM files")
|
||||
conn.execute("DELETE FROM blobs")
|
||||
conn.commit()
|
||||
conn.close()
|
||||
assert history(client, project / "main.py") == []
|
||||
|
||||
r = client.post("/api/v1/admin/reindex")
|
||||
assert r.status_code == 202, r.text
|
||||
task_id = r.json()["task_id"]
|
||||
status = wait_for(lambda: (lambda s: s if s["status"] == "done" else None)(
|
||||
client.get(f"/api/v1/admin/reindex/{task_id}").json()), timeout=15)
|
||||
assert status["versions"] >= 2 and status["files"] == 2
|
||||
|
||||
after = history(client, project / "main.py")
|
||||
assert len(after) == len(before)
|
||||
assert after[0]["durability"] == "durable"
|
||||
assert client.get(f"/api/v1/versions/{after[0]['id']}/content").text == "print('hello')\n"
|
||||
# Restore works end to end from the rebuilt index.
|
||||
(project / "main.py").write_text("garbage\n")
|
||||
vid = after[0]["id"]
|
||||
out = str(project / "restored.py")
|
||||
assert client.post(f"/api/v1/versions/{vid}/restore", json={"target_path": out}).status_code == 200
|
||||
assert open(out).read() == "print('hello')\n"
|
||||
|
||||
|
||||
def test_reindex_fetches_blobs_from_remote(rec_env):
|
||||
"""True disaster: index AND spool gone; content streams back on demand."""
|
||||
client, dav, project, cfg = rec_env
|
||||
_configure_and_fill(client, project)
|
||||
conn = _db_index(cfg)
|
||||
conn.execute("DELETE FROM versions")
|
||||
conn.execute("DELETE FROM files")
|
||||
conn.execute("DELETE FROM blobs")
|
||||
conn.commit()
|
||||
conn.close()
|
||||
for p in (cfg.paths.spool_dir).rglob("*"):
|
||||
if p.is_file():
|
||||
p.unlink()
|
||||
task_id = client.post("/api/v1/admin/reindex").json()["task_id"]
|
||||
wait_for(lambda: (lambda s: s if s["status"] == "done" else None)(
|
||||
client.get(f"/api/v1/admin/reindex/{task_id}").json()), timeout=15)
|
||||
after = history(client, project / "main.py")
|
||||
assert after
|
||||
assert client.get(f"/api/v1/versions/{after[0]['id']}/content").text == "print('hello')\n"
|
||||
|
||||
|
||||
def test_adopt_takes_over_directory(rec_env):
|
||||
client, dav, project, cfg = rec_env
|
||||
directory = _configure_and_fill(client, project)
|
||||
# A fresh machine against the same server adopts the existing directory.
|
||||
import tempfile
|
||||
from pathlib import Path
|
||||
home2 = Path(tempfile.mkdtemp())
|
||||
(home2 / "config").mkdir()
|
||||
(home2 / "config" / "config.toml").write_text(CONFIG)
|
||||
cfg2 = Config.load(Paths(config_dir=home2 / "config", data_dir=home2 / "data", cache_dir=home2 / "cache"))
|
||||
app2 = create_app(cfg2, remote_transport=httpx.MockTransport(dav.handler))
|
||||
with TestClient(app2, base_url="http://127.0.0.1:9922") as c2:
|
||||
c2.headers["Authorization"] = f"Bearer {cfg2.api_token()}"
|
||||
r = c2.post("/api/v1/config/remote/adopt",
|
||||
json={**REMOTE, "password": "secret", "directory": directory})
|
||||
assert r.status_code == 200, r.text
|
||||
assert r.json()["directory"] == directory and r.json()["adopted"] is True
|
||||
|
||||
|
||||
def test_retention_thins_old_keeps_five_days(rec_env):
|
||||
client, dav, project, cfg = rec_env
|
||||
_configure_and_fill(client, project)
|
||||
for i in range(4):
|
||||
(project / "main.py").write_text(f"v{i}\n")
|
||||
time.sleep(0.5) # outside the 0.2s coalescing window: each is a version
|
||||
wait_for(lambda: len(history(client, project / "main.py")) == 5, timeout=10)
|
||||
now = time.time()
|
||||
conn = _db_index(cfg)
|
||||
# Age 3 versions 10 days; pin the oldest; keep the newest fresh.
|
||||
vids = [v["id"] for v in reversed(history(client, project / "main.py"))]
|
||||
for vid in vids[:3]:
|
||||
conn.execute("UPDATE versions SET captured_at = ? WHERE id = ?", (now - 10 * 86400, vid))
|
||||
conn.execute("UPDATE versions SET pinned = 1 WHERE id = ?", (vids[0],))
|
||||
conn.commit()
|
||||
conn.close()
|
||||
|
||||
dry = client.post("/api/v1/admin/retention/run", json={"dry_run": True}).json()
|
||||
assert dry["dry_run"] and dry["versions_to_delete"] == 1 # 3 old minus pinned minus daily-kept
|
||||
done = client.post("/api/v1/admin/retention/run", json={"dry_run": False}).json()
|
||||
assert done["versions_deleted"] == 1
|
||||
remaining = history(client, project / "main.py")
|
||||
assert len(remaining) == 4 # latest + pinned + one-per-day + fresh middle
|
||||
assert remaining[0]["pinned"] == 0 # newest still the newest
|
||||
assert any(v["pinned"] == 1 for v in remaining)
|
||||
|
||||
|
||||
def test_gc_reclaims_orphans_local_and_remote(rec_env):
|
||||
client, dav, project, cfg = rec_env
|
||||
directory = _configure_and_fill(client, project)
|
||||
(project / "main.py").write_text("v2\n")
|
||||
wait_for(lambda: len(history(client, project / "main.py")) == 2, timeout=10)
|
||||
conn = _db_index(cfg)
|
||||
conn.execute("DELETE FROM versions WHERE id = (SELECT max(id) FROM versions)")
|
||||
conn.commit()
|
||||
conn.close()
|
||||
dry = client.post("/api/v1/admin/gc", json={"dry_run": True}).json()
|
||||
assert dry["blobs"] >= 1
|
||||
done = client.post("/api/v1/admin/gc", json={"dry_run": False}).json()
|
||||
assert done["blobs_removed"] >= 1 and done["errors"] == []
|
||||
assert [p for p in dav.files if p.startswith(directory + "/blobs/")] != [] # referenced stay
|
||||
# Remaining history still restorable.
|
||||
assert client.get(f"/api/v1/versions/{history(client, project / 'main.py')[0]['id']}/content").status_code == 200
|
||||
|
||||
|
||||
def test_metrics_and_db_restore_guard(rec_env):
|
||||
client, dav, project, cfg = rec_env
|
||||
_configure_and_fill(client, project)
|
||||
m = client.get("/api/v1/metrics")
|
||||
assert m.status_code == 200 and "versiond_files_tracked" in m.text
|
||||
|
||||
db_hist = history(client, project / "app.db")
|
||||
assert db_hist
|
||||
vid = db_hist[0]["id"]
|
||||
row = _db_index(cfg).execute(
|
||||
"SELECT blob_sha256 FROM versions WHERE id = ?", (vid,)).fetchone()
|
||||
blob_path = None
|
||||
for cand in BlobStore(cfg.paths.spool_dir)._candidates(row[0]):
|
||||
if cand.exists():
|
||||
blob_path = cand
|
||||
assert blob_path is not None
|
||||
blob_path.write_bytes(dbsafe.SQLITE_MAGIC + b"\x00" * 200) # corrupt the stored bytes
|
||||
r = client.post(f"/api/v1/versions/{vid}/restore",
|
||||
json={"target_path": str(project / "evil.db")})
|
||||
assert r.status_code == 422 and "unrestorable" in r.text
|
||||
@@ -10,8 +10,7 @@ from fastapi.testclient import TestClient
|
||||
|
||||
from versiond.api import create_app
|
||||
from versiond.config import Config, Paths
|
||||
from versiond.crypto import KeyRing
|
||||
from versiond.store import _codec
|
||||
from versiond.store import decompress
|
||||
|
||||
from test_service import CONFIG, history, wait_for
|
||||
|
||||
@@ -100,23 +99,26 @@ def test_backup_to_webdav(remote_env):
|
||||
directory = remote["directory"]
|
||||
assert directory.startswith("/versioned/") and remote["adopted"] is False
|
||||
assert "password" not in remote and remote["password_set"] is True
|
||||
assert remote["encryption"] == "none"
|
||||
owner = json.loads(dav.files[directory + "/meta/owner.json"])
|
||||
assert owner["key_id"] == remote["key_id"]
|
||||
assert "installation_id" in owner
|
||||
assert "secret" not in cfg.paths.config_file.read_text() # password only in credentials
|
||||
|
||||
client.post("/api/v1/roots", json={"path": str(project)})
|
||||
wait_for(lambda: progress(client)["upload"]["versions_local"] == 0
|
||||
and progress(client)["upload"]["versions_durable"] == 2, timeout=10)
|
||||
|
||||
keys = KeyRing.load_or_create(cfg.key_file)
|
||||
blob_paths = [p for p in dav.files if p.startswith(directory + "/blobs/")]
|
||||
contents = {_codec.decompress(keys.decrypt(dav.files[p])) for p in blob_paths}
|
||||
# Unencrypted: remote blobs are plain compressed file contents, named by SHA-256.
|
||||
import re
|
||||
assert blob_paths and all(re.fullmatch(r"[0-9a-f]{64}", p.rsplit("/", 1)[-1]) for p in blob_paths)
|
||||
contents = {decompress(dav.files[p]) for p in blob_paths}
|
||||
assert contents == {b"print('hello')\n", b"SECRET=1\n"}
|
||||
assert all(b"SECRET" not in dav.files[p] for p in blob_paths) # encrypted at rest
|
||||
|
||||
manifests = [p for p in dav.files if p.startswith(directory + "/manifests/")]
|
||||
assert manifests and all(p.endswith(".jsonl.zst") for p in manifests)
|
||||
records = [json.loads(line) for p in manifests
|
||||
for line in _codec.decompress(keys.decrypt(dav.files[p])).decode().splitlines()]
|
||||
for line in decompress(dav.files[p]).decode().splitlines()]
|
||||
assert {r["path"] for r in records if r["type"] == "version"} == {str(project / "main.py"), str(project / ".env")}
|
||||
|
||||
# reconfiguring the same machine adopts its own directory
|
||||
|
||||
@@ -155,7 +155,7 @@ def test_push_snapshots_and_limits(env):
|
||||
assert client.post("/api/v1/snapshots", json=body).json()["status"] == "committed"
|
||||
assert client.post("/api/v1/snapshots", json=body).json()["status"] == "unchanged"
|
||||
|
||||
big = {**body, "content_b64": base64.b64encode(b"x" * 204_801).decode()}
|
||||
big = {**body, "content_b64": base64.b64encode(b"x" * 10_485_761).decode()}
|
||||
r = client.post("/api/v1/snapshots", json=big)
|
||||
assert r.status_code == 413 and r.headers["content-type"] == "application/problem+json"
|
||||
|
||||
|
||||
+13
-4
@@ -10,9 +10,15 @@ from versiond.store import BlobStore
|
||||
|
||||
def test_filters_paths():
|
||||
f = Filters()
|
||||
ok = ["app.py", "src/main.go", ".env", ".env.production", "Makefile", "pkg/package.json", ".gitignore"]
|
||||
ok = ["app.py", "src/main.go", ".env", ".env.production", "Makefile", "pkg/package.json", ".gitignore",
|
||||
# user data: images, archives, databases, documents (backed up as raw files)
|
||||
"photos/vacation.jpg", "img/logo.png", "backup/site.tar.gz", "data/archive.zip",
|
||||
"app.db", "app.sqlite3", "report.pdf", "song.mp3", "clip.mp4"]
|
||||
for p in ok:
|
||||
assert f.path_reason(PurePath(p)) is None, p
|
||||
journals = ["app.db-wal", "app.db-shm", "app.sqlite-journal", "data/x-wal"]
|
||||
for p in journals:
|
||||
assert f.path_reason(PurePath(p)) == "database-journal", p
|
||||
rejected = {
|
||||
"node_modules/x/index.js": "ignored-directory",
|
||||
".git/config": "ignored-directory",
|
||||
@@ -21,6 +27,10 @@ def test_filters_paths():
|
||||
".bashrc_local": "hidden-file",
|
||||
"notes.txt~": "temporary-file",
|
||||
"main.py.swp": "ignored-extension",
|
||||
"draft.temp": "ignored-extension",
|
||||
"x.kate-swp": "ignored-extension",
|
||||
"~$lock.docx": "temporary-file",
|
||||
"#autosave#": "temporary-file",
|
||||
"venv/lib/x.py": "ignored-directory",
|
||||
"target/debug/build.rs": "ignored-directory",
|
||||
"4913": "temporary-file",
|
||||
@@ -28,9 +38,8 @@ def test_filters_paths():
|
||||
}
|
||||
for p, reason in rejected.items():
|
||||
assert f.path_reason(PurePath(p)) == reason, p
|
||||
assert f.size_reason(204_800) is None
|
||||
assert f.size_reason(204_801) == "file-too-large"
|
||||
assert f.content_reason(b"abc\x00def") == "binary-content"
|
||||
assert f.size_reason(10_485_760) is None
|
||||
assert f.size_reason(10_485_761) == "file-too-large"
|
||||
|
||||
|
||||
def test_filter_patterns():
|
||||
|
||||
Reference in New Issue
Block a user