`versiond` is a per-user background service that **monitors directories of the user's choice** and records every version of the source and project files in them, stores them in a deduplicated, versioned archive on a remote WebDAV server, and exposes a local HTTP API for browsing history, diffing, restoring and purging.
The primary use case is a safety net for workflows where files are rewritten frequently and automatically — most notably AI coding agents — so that any previous state of any project file can be inspected and restored, individually or in bulk.
### 1.1 Goals
- Automatically monitor user-selected directories and capture every meaningful version of small text/project files (source code, `.env`, `.json`, `Makefile`, configs, etc.) with no client integration needed.
- Choose and claim a unique remote directory on its own; the user supplies only the WebDAV server and credentials.
- Store versions durably on a user-configured WebDAV server.
- Provide full history, diffs and point-in-time restore through a documented REST API.
- Be cheap to run continuously: bounded CPU, memory, disk and network usage.
- Be operable entirely by an AI agent via a downloadable skill package.
| API docs | Swagger UI at `/docs`, ReDoc at `/redoc`, schema at `/openapi.json` |
### 2.2 systemd user service
The service runs as a **systemd user unit** so it runs with the user's own permissions and needs no root. **Lingering** is enabled so the service starts at boot and keeps running when the user is not logged in.
```bash
loginctl enable-linger "$USER"
systemctl --user daemon-reload
systemctl --user enable --now versiond.service
```
`~/.config/systemd/user/versiond.service`:
```ini
[Unit]
Description=versiond - local file versioning and backup service
Note: filesystem sandboxing (`ProtectSystem=`, `PrivateTmp=`, `ReadWritePaths=`) is deliberately left out. In a *user* unit these options need unprivileged user namespaces, which many distributions restrict (e.g. Ubuntu's AppArmor userns policy), so the unit would fail to start. Restores must also be able to write anywhere under the monitored roots. The service already runs unprivileged as the user, binds to loopback only, and refuses restore targets outside `$HOME` and the roots (§4.7).
1.**API server** — FastAPI application; all operations go through it.
2.**Ingest pipeline** — validates, filters, hashes, coalesces and records snapshots.
3.**Directory monitor** — watches every user-selected root with inotify and submits a snapshot whenever a file changes. This is the primary capture path; see §4.2.
4.**Upload scheduler** — a bounded async worker pool that moves blobs from the spool to WebDAV.
5.**Retention engine** — runs purge and thinning policies on a schedule or on request.
6.**Index** — SQLite database that is the source of truth for metadata. It can be rebuilt from remote manifests.
---
## 4. Functional Specification
### 4.1 Remote storage configuration
The WebDAV target is configured at runtime through the API and saved in the config:
-`base_url` — e.g. `https://dav.example.com/remote.php/dav/files/user/`
-`base_path` — optional parent directory on the server (default `versiond/`). The service creates its own unique subdirectory below it (§4.1.1); the user never names it.
-`auth` — one of `none`, `basic` (username + password), `digest`, `bearer` (token)
-`verify_tls` — boolean, default `true`; optional custom CA bundle path
-`timeout_seconds` — per-request timeout
On save, the service checks the configuration (`PROPFIND` on the base path, then `MKCOL` if needed, then a test `PUT`/`DELETE`) and rejects invalid settings with a clear error. Secrets are write-only: the API never returns them, only whether they are set.
#### 4.1.1 Automatic unique remote directory
Each installation chooses and claims its own remote directory, so several machines (or several users on one machine) can share one WebDAV account without colliding:
1.**Installation id.** At first start, the service derives a stable id with the same method as systemd's `sd_id128_get_machine_app_specific()`: `HMAC-SHA256(key=/etc/machine-id, msg=<versiond app id + $UID>)`, truncated to 16 hex characters. The raw machine id never leaves the host, and the derived id cannot be traced back to it. If `/etc/machine-id` is missing, a random UUIDv4 is used.
2.**Directory name.**`<base_path>/<hostname-slug>-<id[:8]>/`, e.g. `versiond/laptop-3f9a1c07/`. The hostname is only there to be readable; the id makes the name unique.
3.**Claim.** The service writes `meta/owner.json` (installation id, hostname, user, created_at) with an `If-None-Match: *` conditional `PUT`. If the directory already exists with a *different* id (e.g. a cloned VM image with the same machine-id), the service adds a random suffix and tries again. If it exists with the *same* id (reinstall), the service adopts it and runs `reindex`.
4.**Persistence.** The chosen path is saved in `config.toml`, so it stays the same if the hostname changes. `GET /config/remote` shows it; `POST /config/remote/adopt` lets a new machine take over an existing directory explicitly (restore after reinstall on new hardware).
Monitored roots do not need their own remote directories: blobs are shared (content-addressed) and each manifest record carries its root id.
Blobs are immutable and deduplicated by SHA-256 of the plaintext content. Manifests are append-only journals. Together they make it possible to fully rebuild the index (`POST /admin/reindex`) after local data loss.
### 4.2 Ingestion
#### 4.2.1 Directory monitoring (primary)
The user registers one or more **roots** (`POST /roots {"path": "~/projects"}`); from then on everything below them is captured automatically. A root can be a single project or a parent folder of many projects.
**Mechanism — and why it is the lightest option.** The research on mechanisms concluded:
| **Raw inotify (own ctypes binding)** | **Chosen.** One kernel watch per *directory* (not per file), costing about 1 KB of kernel memory each. Events are delivered on one file descriptor read directly by the asyncio loop (`add_reader`): no threads, no polling, ~0 % CPU when idle. Watched directories are kept as a tree of `(wd, parent, name)` nodes, so a directory rename is one pointer change and memory stays at ~100 bytes per watch. `asyncinotify` was tried first; its per-watch `Path` objects cost ~260 MB RSS at 136k watches versus ~90 MB with the tree. |
| `watchfiles` (Rust `notify`) | Good library, but its recursive mode adds a watch to *every* directory, including `node_modules`/`.venv`; filters only drop events afterwards. It can use up the watch limit on large trees, and it pulls in a compiled dependency. |
| `watchdog` | Heavier (emitter threads, snapshot objects), same one-watch-per-directory cost, cross-platform layer we don't need. |
| fanotify | Unprivileged use (kernel ≥ 5.13) does not allow filesystem- or mount-wide marks, so it still needs one mark per directory, with a more complex API. No benefit without root. |
| Polling (mtime scans) | Used only as a fallback (see below). |
The key saving comes from **controlling where watches are placed**: the monitor walks each root with `os.scandir` and **prunes ignored directories (§4.3) before adding watches**. Dot-directories, `node_modules`, `venv`, `target`, etc. never get a watch. For a typical developer tree this cuts the watch count by 90 % or more compared with a naive recursive watch.
-`IN_MODIFY` is deliberately **not** watched: it fires on every `write()` call. `IN_CLOSE_WRITE` fires once per save, and `IN_MOVED_TO` covers editors and agents that save atomically (write temp file, then `rename`).
**Event handling:**
-`IN_CLOSE_WRITE` / `IN_MOVED_TO` on a file → apply filters, then send it to the ingest pipeline (§4.4 coalescing absorbs bursts). Editor temp/swap files (`*.swp`, `*~`, `4913`, `.#*`, `*.tmp`) are filtered.
-`IN_CREATE` on a directory → if it isn't ignored, add a watch **first, then scan it**, so files created inside before the watch existed are not missed (the race window).
-`IN_MOVED_FROM`/`IN_MOVED_TO` pairs (matched by cookie) → record a rename in the index, keeping history linked across the move. An unpaired `MOVED_FROM` counts as a delete.
-`IN_DELETE` → mark the file `exists_on_disk = false` (history is kept; the stale purge in §4.8 handles it later).
-`IN_DELETE_SELF` / `IN_MOVE_SELF` on a root → mark the root as `missing` and poll for it to come back.
-`IN_Q_OVERFLOW` → events were lost; run a reconciliation scan of the affected roots.
**Before/after semantics.** inotify reports changes *after* they happen. Previous states are still preserved because every root gets a **baseline snapshot** of all qualifying files when it is registered. From then on each change stores the new state, so "the version before the change" is always the previous stored version. The baseline is uploaded in the background at low priority.
**Reconciliation scan.** inotify sees nothing while the service is stopped. At startup, after a queue overflow, and once a day, a scan compares each file's `(mtime_ns, size, inode)` with the index and hashes only the files that differ. Cost: one `stat` per file, no reads.
**Watch limit.** The kernel limit `fs.inotify.max_user_watches` is shared by all of the user's processes (IDEs, `tsc --watch`, etc.). Since Linux 5.11 the default scales with RAM (8 192 – 1 048 576). The monitor:
- counts the directories it needs before registering a root and refuses with a clear error if the root would use more than 50 % of the remaining limit (`GET /roots/{id}` shows the counts);
- if `ENOSPC` still occurs, switches only the subtrees it could not watch to **polling** (one mtime scan every 60 s) and reports `degraded` in `/health`, together with the `sysctl` command needed to raise the limit.
**Restrictions:** roots on network filesystems (NFS, SMB, sshfs) and FUSE do not deliver reliable inotify events; they automatically use polling mode. Symlinks are not followed.
#### 4.2.2 Push API (secondary)
Clients can still submit a snapshot explicitly. This is useful for paths outside any root, or for a **true pre-write** capture by an agent hook:
POST /snapshots { "path": "/home/u/proj/app.py", "read_from_disk": true }
```
Push snapshots go through the same filters and coalescing. A `source` tag (e.g. `claude-agent`) is stored on the version so restores can target "everything the agent changed".
#### 4.2.3 Project detection
Each file is linked to a *project root*, found by walking up (within its monitored root) to the nearest `.git`, `pyproject.toml`, `package.json`, `go.mod`, `Cargo.toml`, `Makefile`, or the monitored root itself. This allows history and restore per project, even when a monitored root contains many projects.
### 4.3 Filtering
A file is rejected (HTTP `422` with a reason code for push requests; silently skipped by the monitor) if any of these apply. The same rules decide which **directories** get an inotify watch at all:
- **Compiled artifacts and generated bundles:** `*.pyc`, `*.pyo`, `*.class`, `*.o`, `*.obj`, `*.so`, `*.dylib`, `*.dll`, `*.exe`, `*.a`, `*.lib`, `*.wasm`, `*.jar`, `*.war`, `*.whl`, `*.egg`, `*.min.js`, `*.map`. Everything else is user data and is kept: images, audio/video, archives (`.zip`, `.tar.*`, ...), databases including SQLite journals (`-wal`/`-shm`), documents (`.pdf`), fonts. There is no binary sniffing; any bytes up to the size cap are versioned, served back as `application/octet-stream` when they are not UTF-8.
- **Live SQLite databases (zero-error policy):** any file with the SQLite magic is never raw-copied. It is snapshotted read-only through the SQLite Online Backup API (no write lock, no WAL recovery, live writers unaffected), then `PRAGMA integrity_check` must return `ok`, otherwise nothing is stored (`database-corrupt`, `database-locked` past a 5 s deadline, `database-unreadable` surface as skips, never as versions). Journal files (`*-wal`, `*-shm`, `*-journal`) are not versioned standalone; they are folded into the main file's verified snapshot. For non-SQLite engines only a crash-consistent raw copy is possible, so dump-then-backup remains required: back up the `.dump`/`VACUUM INTO`/`pg_dump` output alongside the live file.
- **User rules:** `.gitignore`-style patterns in `config.toml` (`[ignore] patterns = [...]`), plus optional respect of the project's own `.gitignore` (`respect_gitignore = true`, but `.env` is still captured unless explicitly excluded).
All filter decisions can be checked with `POST /filters/test` without storing anything.
### 4.4 Coalescing and rate limiting
Purpose: when a file is snapshotted many times in a short burst (e.g. an agent saving 40 times in a minute), keep the meaningful states and drop the noise.
- **Deduplication:** a snapshot whose hash equals the file's latest stored version is a no-op (`200`, `"status": "unchanged"`).
- **Coalescing window** (default 10 s, per path): while a snapshot for a path is pending in the window, new snapshots replace it. When the window closes, **only the first and the last state are committed**; intermediate states are dropped. The window is a debounce with a maximum hold (default 60 s), so a file that never stops changing is still captured.
- **Per-path rate limit:** token bucket (default 30 committed versions/hour/path). Above the limit, snapshots are coalesced into the next free slot rather than refused, so the most recent state is never lost.
- **Global rate limit:** token bucket on ingest requests (default 50 req/s) → `429 Too Many Requests` with a `Retry-After` header.
### 4.5 Upload and concurrency
All upload behaviour is configurable (`[upload]` section and `PATCH /config/upload`):
| Setting | Default | Meaning |
|---|---|---|
| `concurrency` | 4 | Max parallel WebDAV requests |
| `max_requests_per_second` | 10 | Upstream request rate cap |
| `max_bandwidth_kbps` | 0 (unlimited) | Upload bandwidth cap |
| `retry_max_attempts` | 8 | Per blob |
| `retry_backoff` | exponential, 1 s → 5 min, with jitter | |
Uploads are idempotent (content-addressed names, so a repeated `PUT` writes identical data). No `HEAD` is sent before `PUT`; that would double the requests for the common case where the blob is new. A version counts as **durable** only after its blob and manifest entry are both stored remotely. The API shows this state per version (`pending` / `durable`).
- Point-in-time view: the state of an entire project (or glob) **as of** a timestamp.
- Full-text search across stored versions (SQLite FTS5 over the latest N versions per file; configurable).
### 4.7 Restore
- **Single restore:** write version X of a file back to its original path or to another target path.
- **Bulk restore** by criteria: project, path glob(s), extension, time point (`as_of`), source/reason, "files changed by source `claude-agent` after T", etc.
- **Always dry-run first:** a bulk restore returns a plan (file list, actions, conflicts) and a `plan_id`; running the plan requires `POST /restores/{plan_id}/execute`. Plans expire after 15 minutes.
- **Safety:** before overwriting, the current on-disk content is itself snapshotted (`reason: "pre-restore"`), so every restore can be undone.
- **Conflict policy:** `overwrite`, `skip`, `rename` (`file.restored-<ts>.ext`), or `fail`.
- **Path safety:** targets are resolved and must lie inside an allowed root; symlink escapes and `..` traversal are refused.
- Restoring a file that no longer exists recreates it, including parent directories.
-`SKILL.md` — frontmatter (`name`, `description`) plus instructions: when to snapshot, how to query history, how to run safe bulk restores (always dry-run, show the plan to the user, then execute).
-`reference/api.md` — a short endpoint reference generated from the live OpenAPI schema.
-`scripts/` — small helper scripts (e.g. `vd snapshot <path>`, `vd restore --as-of ...`) that wrap `curl`.
-`GET /agent/skill.md` — the bare `SKILL.md` for quick inspection.
- The skill is generated from the running version, so it always matches the API.
Example agent requests this enables: *"Restore every `.py` file in `~/proj/api` to how it was yesterday at 14:00, except tests"* or *"Show me what the agent changed in `config.json` in the last hour."*
Optional extra: a Claude Code `PreToolUse` hook on `Write`/`Edit` that calls `POST /snapshots` with `read_from_disk: true` and `source: "claude-agent"`. The monitor already captures these edits; the hook adds a guaranteed pre-write state (even if the service was down) and tags the change as the agent's.
---
## 5. API Reference (summary)
All endpoints are JSON, under `/api/v1` (the docs, schema and skill endpoints are at the root).
Indexes on `(file_id, captured_at)`, `files(path)`, `versions(blob_sha256)`. Schema migrations are versioned and applied at startup.
---
## 7. Security
- **Loopback only.** The server refuses to start if configured to bind to a non-loopback address.
- **Local authentication.** Other local users and processes (including browsers, via DNS rebinding) can reach `127.0.0.1`. Every request therefore needs a bearer token stored in `~/.config/versiond/credentials` (`0600`). The `Host` header is checked against `127.0.0.1:9922`/`localhost:9922`, and CORS is disabled.
- **Credential storage.** WebDAV secrets are kept in the `0600` credentials file or, when available, the Secret Service/keyring. They never appear in logs or API responses.
- **Restore path safety.** See §4.7.
- **Audit log.** All configuration changes, restores and purges are recorded with timestamp and client identity.
versiond skill install # install the agent skill into ~/.claude/skills/
```
---
## 10. Testing Strategy
- Unit tests: filter rules (table-driven over a large path corpus), coalescing state machine, retention planner, path-safety resolver.
- Integration tests against a containerized WebDAV server (e.g. `rclone serve webdav` or Apache `mod_dav`), including fault injection (timeouts, 5xx, partial writes).
- Crash-consistency tests: kill the process at random points during ingest, upload and GC; verify no referenced blob is lost and the index stays consistent.
- Disaster-recovery test: delete the local index, run `reindex`, and compare.
- Contract test: the generated skill's examples run against the live API.
---
## 11. Milestones
1.**M1 – Core:** API skeleton, filters, SQLite index, local spool, root registration, inotify monitor with pruned watches, baseline and reconciliation scan, push snapshots, history, diff, single restore.
- Should `~/` be offered as a one-click root, or should the user be pushed toward choosing specific folders (lower watch count, less noise)?
- Should other backends (S3, SFTP, local directory) be supported behind the same storage interface? WebDAV-only is the v1 scope.
---
## Appendix A — Assessment of the Idea
**Strengths**
- It solves a real, current problem: AI agents rewrite files quickly and destructively, often outside Git commits. A "before every write" safety net with bulk point-in-time restore is useful, and editor local-history features (JetBrains, VS Code Timeline) do not cover agent edits across tools.
- Shipping a skill so the agent can operate its own undo system is a smart idea and the most distinctive part of the design.
- Coalescing burst writes (keep first and last) is the right approach to versioning noise.
**Weaknesses and risks**
- Overlaps heavily with existing tools: Git (plus `git stash`/autocommit tools), restic/borg/kopia, Nextcloud's own file versioning on the very WebDAV server it targets, and editor local history. The differentiation has to be the agent integration and pre-write capture, and the README should say so.
- inotify only reports changes after they happen, so capturing the "before" state relies on a complete baseline and a reconciliation scan. Changes made while the service is down are captured only as their final state.
- The inotify watch limit is shared with IDEs and dev servers. Very large roots fall back to polling, which costs more.
- WebDAV is a slow, chatty backend for many small objects. Content-addressing plus batched manifests reduce this, but performance against Nextcloud-class servers needs to be measured early.
- The original text left out authentication, failure handling, restore safety and data model. All of these are filled in above, but they roughly triple the scope compared with how the idea first read.
**Grade: 7.5 / 10 (B)** — a good, useful niche idea with one clearly distinctive feature (agent-operable undo). Making directory monitoring the default removes the biggest earlier weakness (it no longer depends on clients calling it), which raises the idea to **7.5**. It still loses points for overlapping with existing tools. The original write-up as a specification would get a **3 / 10**; the idea is better than its description.