Remove encryption; user-data filters; SQLite safety; recovery, retention, scheduler
- No client-side encryption: plaintext content-addressed remote (SHA-256 names), no backup key, no cryptography dependency (server disk encryption is the trust model). - Filters back user data: images, archives, databases, PDFs accepted; 10 MiB cap; binary sniffing removed; temp names hardened (~$, #..#, .temp). - SQLite zero-error policy: backup-API snapshots + integrity_check, journal folding, locked/corrupt loud skips, verified restores. - Recovery: reindex from manifests, remote adopt, on-demand blob fetch, 5-day retention + thinning, GC, date-guarded remote purge, metrics. - Scheduler with WebDAV quota signal and 70% pressure backstop (floor kept). - 37 tests incl. live-monitor capture safety and DB safety.
This commit is contained in:
@@ -11,10 +11,10 @@
|
||||
| Milestone | State |
|
||||
|---|---|
|
||||
| M1 – Core (monitor, filters, index, spool, history, diff, restore, systemd install) | **Implemented** |
|
||||
| M2 – WebDAV remote, unique remote directory, encryption, manifests | **Implemented** |
|
||||
| M2 – WebDAV remote, unique remote directory, manifests (unencrypted) | **Implemented** |
|
||||
| Statistics, progress (incl. live stream) and web dashboard | **Implemented** |
|
||||
| M3 – Retention, thinning, remote GC | Not started (pinning and `forget` exist) |
|
||||
| M4 – Agent skill, Prometheus metrics, reindex from remote | Not started |
|
||||
| M3 – Retention (5-day + thinning), remote GC, remote purge, scheduler + capacity backstop | **Implemented** (built-in loop; no external cron needed) |
|
||||
| M4 – Agent skill, reindex from remote, adopt, metrics | **Implemented** except agent skill |
|
||||
|
||||
```bash
|
||||
pipx install -e . # installs the `versiond` command
|
||||
@@ -26,7 +26,6 @@ versiond diff ~/projects/app/main.py # last change
|
||||
versiond restore '~/projects/app/src/*' --as-of 2026-10-08T14:00 # dry run
|
||||
versiond restore '~/projects/app/src/*' --as-of 2026-10-08T14:00 --execute
|
||||
versiond remote set --url https://u123456.your-storagebox.de --user u123456 # prompts for the password
|
||||
versiond key export # store the backup key somewhere safe, off this machine
|
||||
versiond progress --follow # scan + upload progress, rate and ETA
|
||||
versiond stats # files, versions, storage, activity
|
||||
versiond dashboard # opens the web dashboard, already signed in
|
||||
@@ -56,7 +55,7 @@ The primary use case is a safety net for workflows where files are rewritten fre
|
||||
|
||||
### 1.2 Non-goals
|
||||
|
||||
- Backing up large or binary assets (hard limit: 200 KB per file).
|
||||
- Backing up VM/container disk images of running machines (use dumps/snapshots, not live file copies).
|
||||
- Replacing Git. `versiond` versions working-tree states, not commits, branches or merges.
|
||||
- Multi-user or network-exposed operation. The service binds to loopback only.
|
||||
- Full-disk or system backup.
|
||||
@@ -188,12 +187,15 @@ Monitored roots do not need their own remote directories: blobs are shared (cont
|
||||
|
||||
```
|
||||
<base_path>/<hostname-slug>-<id8>/
|
||||
blobs/<name[:2]>/<name> # name = HMAC-SHA256(key, sha256); zstd, then AES-256-GCM
|
||||
manifests/<yyyy>/<mm>/<dd>/<batch-id>.jsonl.zst.enc # append-only version + rename records, encrypted
|
||||
meta/owner.json # installation id, hostname, key id (plaintext)
|
||||
blobs/<sha256[:2]>/<sha256> # name = SHA-256 of content; compressed, NOT encrypted
|
||||
manifests/<yyyy>/<mm>/<dd>/<batch-id>.jsonl.zst # append-only version + rename records, compressed, NOT encrypted
|
||||
meta/owner.json # installation id, hostname (plaintext)
|
||||
meta/format.json # storage format version (plaintext)
|
||||
```
|
||||
|
||||
Remote storage is deliberately unencrypted: it relies on the storage machine's
|
||||
disk encryption. There is no backup key and nothing to lose.
|
||||
|
||||
One directory level under `blobs/` (256 prefixes) keeps the number of `MKCOL` requests small.
|
||||
|
||||
Blobs are immutable and deduplicated by SHA-256 of the plaintext content. Manifests are append-only journals. Together they make it possible to fully rebuild the index (`POST /admin/reindex`) after local data loss.
|
||||
@@ -263,12 +265,12 @@ Each file is linked to a *project root*, found by walking up (within its monitor
|
||||
|
||||
A file is rejected (HTTP `422` with a reason code for push requests; silently skipped by the monitor) if any of these apply. The same rules decide which **directories** get an inotify watch at all:
|
||||
|
||||
- **Size** > 200 KB (204 800 bytes) → `413 Payload Too Large`.
|
||||
- **Size** > 10 MiB (10 485 760 bytes, configurable via `limits.max_file_bytes`) → `413 Payload Too Large`.
|
||||
- **Any path component starts with `.`** (e.g. `.git/`, `.venv/`, `.idea/`, `.cache/`), **except** dotfiles that are themselves project files at the leaf (e.g. `.env`, `.gitignore`, `.editorconfig`, `.dockerignore`). Hidden *directories* are always ignored; hidden *files* follow an allowlist.
|
||||
- **Dependency, build and cache directories**, default list:
|
||||
`node_modules`, `bower_components`, `jspm_packages`, `vendor`, `__pycache__`, `venv`, `env`, `site-packages`, `.tox`, `build`, `dist`, `target`, `out`, `bin`, `obj`, `.gradle`, `Pods`, `Carthage`, `DerivedData`, `.next`, `.nuxt`, `.svelte-kit`, `coverage`, `.terraform`, `_build`, `deps`, `elm-stuff`, `zig-cache`, `zig-out`.
|
||||
- **Compiled/binary artifacts:** `*.pyc`, `*.pyo`, `*.class`, `*.o`, `*.obj`, `*.so`, `*.dylib`, `*.dll`, `*.exe`, `*.a`, `*.lib`, `*.wasm`, `*.jar`, `*.war`, `*.whl`, `*.egg`, archives, images, media, `*.lock` files above the size limit, `*.min.js`, `*.map`.
|
||||
- **Binary content:** a file with a NUL byte in its first 8 KB is rejected.
|
||||
- **Compiled artifacts and generated bundles:** `*.pyc`, `*.pyo`, `*.class`, `*.o`, `*.obj`, `*.so`, `*.dylib`, `*.dll`, `*.exe`, `*.a`, `*.lib`, `*.wasm`, `*.jar`, `*.war`, `*.whl`, `*.egg`, `*.min.js`, `*.map`. Everything else is user data and is kept: images, audio/video, archives (`.zip`, `.tar.*`, ...), databases including SQLite journals (`-wal`/`-shm`), documents (`.pdf`), fonts. There is no binary sniffing; any bytes up to the size cap are versioned, served back as `application/octet-stream` when they are not UTF-8.
|
||||
- **Live SQLite databases (zero-error policy):** any file with the SQLite magic is never raw-copied. It is snapshotted read-only through the SQLite Online Backup API (no write lock, no WAL recovery, live writers unaffected), then `PRAGMA integrity_check` must return `ok`, otherwise nothing is stored (`database-corrupt`, `database-locked` past a 5 s deadline, `database-unreadable` surface as skips, never as versions). Journal files (`*-wal`, `*-shm`, `*-journal`) are not versioned standalone; they are folded into the main file's verified snapshot. For non-SQLite engines only a crash-consistent raw copy is possible, so dump-then-backup remains required: back up the `.dump`/`VACUUM INTO`/`pg_dump` output alongside the live file.
|
||||
- **User rules:** `.gitignore`-style patterns in `config.toml` (`[ignore] patterns = [...]`), plus optional respect of the project's own `.gitignore` (`respect_gitignore = true`, but `.env` is still captured unless explicitly excluded).
|
||||
|
||||
All filter decisions can be checked with `POST /filters/test` without storing anything.
|
||||
@@ -321,17 +323,41 @@ When the remote is unreachable, the service keeps working from the spool and cat
|
||||
- **Path safety:** targets are resolved and must lie inside an allowed root; symlink escapes and `..` traversal are refused.
|
||||
- Restoring a file that no longer exists recreates it, including parent directories.
|
||||
|
||||
### 4.8 Retention and purge
|
||||
### 4.8 Retention and purge (implemented)
|
||||
|
||||
Retention policies run on a schedule (default: daily) and on demand. Every purge supports `dry_run=true` and returns what would be removed and the bytes reclaimed.
|
||||
- **Retention** (`POST /api/v1/admin/retention/run`, dry run by default): keeps every
|
||||
version younger than `retention.keep_days` (default **5**), thins older ones to
|
||||
one per file per UTC day, and always keeps each file's latest version and all
|
||||
pins. Follow with GC to reclaim the freed blobs.
|
||||
- **Garbage collection** (`POST /api/v1/admin/gc`, dry run by default): deletes blobs
|
||||
no version references anymore, from the local spool and (unless `remote: false`)
|
||||
from WebDAV. Only unreferenced data is ever touched, so a crash can leak garbage
|
||||
but never lose a referenced blob.
|
||||
- **Built-in scheduler** (no cron needed): every `scheduler.retention_interval_seconds`
|
||||
(default daily) a retention pass runs; GC is folded in every
|
||||
`scheduler.gc_interval_seconds` (default weekly). Runs are serialized and audited.
|
||||
- **Capacity backstop** (`scheduler.remote_max_used_percent`, default **70**): remote
|
||||
usage is read via WebDAV quota properties (RFC 4331, falling back to an
|
||||
index-based estimate) and refreshed every `scheduler.usage_check_seconds`. At or
|
||||
above the limit an out-of-schedule pass runs — still never below the keep-days
|
||||
floor — and `/health` flips to `degraded` with `remote-over-capacity-limit` plus
|
||||
`remote_pressure: true` until a human analyzes (lower `keep_days`, bigger box).
|
||||
- **Remote purge** (`POST /api/v1/admin/purge-remote {"today": "dd-mm-yyyy"}`): deletes
|
||||
everything this installation uploaded (`blobs/` + `manifests/` under the claimed
|
||||
directory only). Replies 202 immediately with file/byte counts; deletes async.
|
||||
- **Pinning:** versions can be pinned (`POST /versions/{id}/pin`) and are never removed.
|
||||
|
||||
- **Thinning (grandfather-father-son):** keep all versions from the last 24 h, hourly for 7 days, daily for 30 days, weekly for 6 months, monthly after that. Fully configurable.
|
||||
- **Stale file purge:** remove files whose path no longer exists on disk *and* that have had no new version for N days (default 90).
|
||||
- **Stale project purge:** remove a whole project that has had no activity for N days.
|
||||
- **Intermediate purge:** for a file or glob, remove all versions between two timestamps, or all except the first and last of each day.
|
||||
- **Explicit delete:** a specific version, file, or project.
|
||||
- **Garbage collection:** blobs not referenced by any remaining version are deleted from WebDAV in a separate mark-and-sweep pass with a grace period (default 7 days), so a crash during purge can never lose a referenced blob.
|
||||
- **Pinning:** versions can be pinned (`POST /versions/{id}/pin`) and are never removed by automatic policies.
|
||||
### 4.8.1 Disaster recovery (total local loss)
|
||||
|
||||
1. Reinstall, then `POST /api/v1/config/remote/adopt` with the old directory to take
|
||||
it over on the new machine.
|
||||
2. `POST /api/v1/admin/reindex` — rebuilds the whole index from remote manifests
|
||||
(async; poll its status). Everything comes back `durable`.
|
||||
3. Restore normally (single, bulk, or point-in-time). Blob bytes stream down from
|
||||
WebDAV on demand and are re-spooled locally, so the first restore of each file
|
||||
needs the remote reachable; after that it is local again.
|
||||
4. Restored SQLite files are integrity-checked before being written; a stored
|
||||
version that fails the check is refused (`unrestorable`) instead of written.
|
||||
|
||||
### 4.9 Agent skill download
|
||||
|
||||
@@ -364,9 +390,9 @@ All endpoints are JSON, under `/api/v1` (the docs, schema and skill endpoints ar
|
||||
| GET / DELETE | `/roots/{id}` | Root status (watch count, mode `inotify`/`polling`/`missing`, baseline progress) / stop monitoring |
|
||||
| POST | `/roots/{id}/rescan` | Force a reconciliation scan |
|
||||
| POST | `/config/remote/test` | Test the remote without saving |
|
||||
| GET / PATCH | `/config/upload` | Concurrency, rate and retry settings |
|
||||
| GET / PATCH | `/config/filters` | Ignore rules |
|
||||
| GET / PATCH | `/config/retention` | Retention policies |
|
||||
| GET / PATCH | `/config/upload` | Not implemented |
|
||||
| GET / PATCH | `/config/filters` | Not implemented |
|
||||
| GET / PATCH | `/config/retention` | Not implemented (see `retention.keep_days` in config.toml) |
|
||||
| POST | `/filters/test` | Check whether paths would be accepted |
|
||||
| POST | `/snapshots` | Submit a snapshot (single) |
|
||||
| POST | `/snapshots/batch` | Submit up to 100 snapshots |
|
||||
@@ -377,21 +403,26 @@ All endpoints are JSON, under `/api/v1` (the docs, schema and skill endpoints ar
|
||||
| GET | `/versions/{id}/content` | Raw content |
|
||||
| GET | `/diff?from=&to=` | Diff two versions (`to=disk` for the working copy) |
|
||||
| GET | `/projects/{id}/tree?as_of=` | Project state at a point in time |
|
||||
| GET | `/search?q=` | Full-text search |
|
||||
| GET | `/search?q=` | Not implemented |
|
||||
| POST | `/restores` | Create a restore plan (dry run) |
|
||||
| POST | `/restores/{plan_id}/execute` | Run a restore plan |
|
||||
| POST | `/versions/{id}/pin` / `DELETE` | Pin / unpin |
|
||||
| POST | `/purge` | Run a purge (criteria + `dry_run`) |
|
||||
| DELETE | `/versions/{id}`, `/files?path=`, `/projects/{id}` | Explicit deletion |
|
||||
| POST | `/purge` | Not implemented (use `/admin/retention/run` + `/admin/gc`) |
|
||||
| DELETE | `/versions/{id}`, `/files?path=`, `/projects/{id}` | Not implemented (use `/forget`) |
|
||||
| GET | `/stats` | Files, versions (per source, per hour for 24 h), storage, dedup/compression, top files and projects |
|
||||
| GET | `/progress` | Active and last scans (percent, files/s, ETA), upload queue (percent, rate, ETA, errors), monitor counters |
|
||||
| GET | `/progress/stream` | Server-sent events with `/progress` every `interval` seconds |
|
||||
| GET | `/dashboard` | Web dashboard (token via `#token=` fragment or prompt) |
|
||||
| POST | `/forget` | Delete all history below a path (`dry_run` by default) |
|
||||
| POST | `/admin/gc` | Garbage-collect remote blobs |
|
||||
| POST | `/admin/reindex` | Rebuild the index from remote manifests |
|
||||
| GET | `/admin/queue` | Upload queue status |
|
||||
| GET | `/agent/skill`, `/agent/skill.md` | Agent skill package |
|
||||
| POST | `/admin/gc` | Garbage-collect orphan blobs, local + remote (`dry_run` by default) |
|
||||
| POST | `/admin/reindex` | Rebuild the index from remote manifests (async, 202 + status) |
|
||||
| GET | `/admin/reindex/{task_id}` | Reindex task status |
|
||||
| POST | `/admin/retention/run` | Thin history per keep-days (default 5, `dry_run` by default) |
|
||||
| POST | `/admin/purge-remote` | Delete this installation's remote data (date-guarded, async, 202) |
|
||||
| GET | `/admin/purge-remote/{task_id}` | Purge task status |
|
||||
| GET | `/metrics` | Prometheus text metrics (behind bearer auth) |
|
||||
| GET | `/admin/queue` | Not implemented (use `/progress` upload section) |
|
||||
| GET | `/agent/skill`, `/agent/skill.md` | Not implemented |
|
||||
| GET | `/docs`, `/redoc`, `/openapi.json` | API documentation |
|
||||
|
||||
Errors use RFC 9457 problem details (`application/problem+json`) with stable `type` codes (e.g. `file-too-large`, `path-ignored`, `remote-unreachable`, `rate-limited`).
|
||||
@@ -420,7 +451,12 @@ Indexes on `(file_id, captured_at)`, `files(path)`, `versions(blob_sha256)`. Sch
|
||||
|
||||
- **Loopback only.** The server refuses to start if configured to bind to a non-loopback address.
|
||||
- **Local authentication.** Other local users and processes (including browsers, via DNS rebinding) can reach `127.0.0.1`. Every request therefore needs a bearer token stored in `~/.config/versiond/credentials` (`0600`). The `Host` header is checked against `127.0.0.1:9922`/`localhost:9922`, and CORS is disabled.
|
||||
- **Secrets in backed-up files.** `.env` and similar files contain credentials and are sent off-machine. Client-side encryption is therefore **always on**: blobs and manifests are encrypted with AES-256-GCM (random 96-bit nonce) using a 64-byte key in `~/.config/versiond/backup.key` (`0600`). Remote blob names are HMAC-SHA-256 of the content hash, so the server cannot confirm guesses of known content. **The key is the only way to read the remote copy**: `versiond key export` prints it, and it must be stored off the machine.
|
||||
- **Remote storage is unencrypted by design.** `.env` and similar files are sent
|
||||
off-machine as compressed (not encrypted) blobs, and blob names are plain
|
||||
SHA-256 content hashes. The threat model is a disk-encrypted storage box on a
|
||||
trusted network (HTTPS/Basic-auth WebDAV), not an untrusted server. There is
|
||||
no backup key and nothing extra to lose: any WebDAV client can read the
|
||||
remote, and the local spool + SQLite index is sufficient to restore.
|
||||
- **Credential storage.** WebDAV secrets are kept in the `0600` credentials file or, when available, the Secret Service/keyring. They never appear in logs or API responses.
|
||||
- **Restore path safety.** See §4.7.
|
||||
- **Audit log.** All configuration changes, restores and purges are recorded with timestamp and client identity.
|
||||
@@ -476,7 +512,7 @@ versiond skill install # install the agent skill into ~/.claude/sk
|
||||
## 11. Milestones
|
||||
|
||||
1. **M1 – Core:** API skeleton, filters, SQLite index, local spool, root registration, inotify monitor with pruned watches, baseline and reconciliation scan, push snapshots, history, diff, single restore.
|
||||
2. **M2 – Remote:** WebDAV configuration, automatic unique remote directory and claim, upload workers, manifests, durability tracking, encryption.
|
||||
2. **M2 – Remote:** WebDAV configuration, automatic unique remote directory and claim, upload workers, manifests, durability tracking (no encryption).
|
||||
3. **M3 – Control:** coalescing, rate limits, bulk restore plans, retention, purge, GC.
|
||||
4. **M4 – Agent & ops:** skill package, CLI, `install` command, metrics, polling fallback, reindex/adopt.
|
||||
|
||||
@@ -485,7 +521,6 @@ versiond skill install # install the agent skill into ~/.claude/sk
|
||||
## 12. Open Questions
|
||||
|
||||
- Should `~/` be offered as a one-click root, or should the user be pushed toward choosing specific folders (lower watch count, less noise)?
|
||||
- Is 200 KB a hard product limit, or should it be configurable with 200 KB as the default?
|
||||
- Should other backends (S3, SFTP, local directory) be supported behind the same storage interface? WebDAV-only is the v1 scope.
|
||||
|
||||
---
|
||||
@@ -496,7 +531,7 @@ versiond skill install # install the agent skill into ~/.claude/sk
|
||||
|
||||
- It solves a real, current problem: AI agents rewrite files quickly and destructively, often outside Git commits. A "before every write" safety net with bulk point-in-time restore is useful, and editor local-history features (JetBrains, VS Code Timeline) do not cover agent edits across tools.
|
||||
- Shipping a skill so the agent can operate its own undo system is a smart idea and the most distinctive part of the design.
|
||||
- The 200 KB cap and the ignore rules keep the problem small and cheap, so it can stay running all the time.
|
||||
- The 10 MiB cap and the ignore rules (deps/build/caches out, user data in) keep the problem small and cheap, so it can stay running all the time.
|
||||
- Coalescing burst writes (keep first and last) is the right approach to versioning noise.
|
||||
|
||||
**Weaknesses and risks**
|
||||
@@ -504,7 +539,9 @@ versiond skill install # install the agent skill into ~/.claude/sk
|
||||
- Overlaps heavily with existing tools: Git (plus `git stash`/autocommit tools), restic/borg/kopia, Nextcloud's own file versioning on the very WebDAV server it targets, and editor local history. The differentiation has to be the agent integration and pre-write capture, and the README should say so.
|
||||
- inotify only reports changes after they happen, so capturing the "before" state relies on a complete baseline and a reconciliation scan. Changes made while the service is down are captured only as their final state.
|
||||
- The inotify watch limit is shared with IDEs and dev servers. Very large roots fall back to polling, which costs more.
|
||||
- Sending `.env` files to a remote server without encryption would be a security liability. That is why encryption is on by default in this spec.
|
||||
- `.env` files are stored unencrypted on the remote by design. The storage box is
|
||||
assumed disk-encrypted on a trusted network; do not point versiond at an
|
||||
untrusted server.
|
||||
- WebDAV is a slow, chatty backend for many small objects. Content-addressing plus batched manifests reduce this, but performance against Nextcloud-class servers needs to be measured early.
|
||||
- The original text left out authentication, failure handling, restore safety and data model. All of these are filled in above, but they roughly triple the scope compared with how the idea first read.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user