Forbid em dashes everywhere; add AGENTS.md with repo rules

This commit is contained in:
retoor
2026-10-10 04:14:18 +02:00
parent ec0dab37d1
commit e751e7cebd
16 changed files with 580 additions and 571 deletions
+46 -46
View File
@@ -1,14 +1,14 @@
# versiond — the undo button for your entire machine
# versiond - the undo button for your entire machine
`versiond` watches the directories you choose and versions every meaningful file state:
source code, configs, images, archives, documents, and live databases. Then it
syncs it all to your own WebDAV box. When an AI agent (or you, at 3am) destroys
something, you go back — one file, one project, or the whole system.
something, you go back - one file, one project, or the whole system.
No encryption keys to lose. No subscription. No cloud. Your disk, your server,
your history.
> Status: running in production on its author's dev box right now —
> Status: running in production on its author's dev box right now -
> 7,596 files, 8,884 versions, zero errors. This README is written from
> measured numbers on that machine, not from wishes.
@@ -43,7 +43,7 @@ box this was built on (a 436GB Hetzner box, October 2026).
### 1.1 The surge test
16 threads hammering a watched directory for 30 seconds straight — 6 writers,
16 threads hammering a watched directory for 30 seconds straight - 6 writers,
6 readers, renames, deletes, plus a live SQLite writer:
| What went in | What versiond did with it |
@@ -72,7 +72,7 @@ writes a day?
| Bytes/day | 720 × 7GB ≈ **5 TB/day** |
| Capture load | 720 × 63s ≈ **12.6 h/day of heavy I/O** |
Read that twice. The snapshots are *correct* — but at 7GB, full-copy
Read that twice. The snapshots are *correct* - but at 7GB, full-copy
versioning is physically absurd. That is why databases get their own schedule:
### 1.3 The database schedule (opinionated, on purpose)
@@ -103,7 +103,7 @@ interval(size) = clamp(2h, 4h × (size / 7GB), 24h)
10:00 ░░░ small ones 22:00 ░░░ small ones
```
A file that changes constantly still banks one version per slot — recency
A file that changes constantly still banks one version per slot - recency
granularity *is* the grid, and the grid is the contract.
### 1.4 Big data: the honest math
@@ -121,7 +121,7 @@ measured chunk-by-chunk:
Formula: full-copy burns `versions/day × file_size`; chunked burns
`churn + growth` once. For a 5KB config the difference is trivia; for a 7GB
database it's the difference between possible and impossible. Both columns
keep identical versions and identical restores — only the bytes differ.
keep identical versions and identical restores - only the bytes differ.
### 1.5 The 286GB reality check
@@ -129,8 +129,8 @@ That same box once showed 51GB free on a 436GB disk. The hunt took minutes:
| Where | Size | Verdict |
|---|---|---|
| `devplacepy/data/backups/` — 42 hourly 6GB tarballs, nothing ever prunes them | 253 GB | devplace's own backup output, excluded from versiond |
| `devplacepy/restore_staging/` — leftover restore duplicate | 39.5 GB | delete after confirming |
| `devplacepy/data/backups/` - 42 hourly 6GB tarballs, nothing ever prunes them | 253 GB | devplace's own backup output, excluded from versiond |
| `devplacepy/restore_staging/` - leftover restore duplicate | 39.5 GB | delete after confirming |
| `devplacepy/data/uploads/` | 32.7 GB | legit user data, versioned |
| `devplacepy/data/devplace.db` | 7 GB | legit, hot, scheduled |
| everything else | ~30 GB | normal |
@@ -142,13 +142,13 @@ is arson. `data/backups/` and `backup_staging/` ship excluded.
## 2. What gets versioned (and what never does)
Kept — source, configs, allowlisted dotfiles, **images, audio/video, archives,
Kept - source, configs, allowlisted dotfiles, **images, audio/video, archives,
databases (journals folded into verified snapshots), PDFs, fonts**. Cap:
`limits.max_file_bytes` (default **10 MiB**, configurable). No binary
sniffing: any bytes up to the cap are versioned; non-UTF-8 comes back as
`application/octet-stream`.
Never — dependency/build/cache trees (`node_modules`, `venv`, `target`, …),
Never - dependency/build/cache trees (`node_modules`, `venv`, `target`, …),
hidden directories, `*.pyc/*.o/*.so`, `*.min.js/*.map`, temps (`*.tmp`,
`~$*` Office locks, `#*#` Emacs autosaves, `*~`, `4913`), SQLite journals as
standalone files, anything over the cap. Proven live: 11 temp-name variants
@@ -156,7 +156,7 @@ written to a watched dir produced **zero** versions and eleven 422s.
Databases get the zero-error policy: backup-API snapshot taken read-only (no
write lock, no WAL recovery, live writers unaffected), `PRAGMA integrity_check`
must say `ok`, otherwise nothing is stored — `database-corrupt`,
must say `ok`, otherwise nothing is stored - `database-corrupt`,
`database-locked`, `database-unreadable` surface as loud skips. Every SQLite
restore is integrity-checked before it touches disk; failures refuse with
`unrestorable` instead of writing garbage.
@@ -172,12 +172,12 @@ restore is integrity-checked before it touches disk; failures refuse with
watched `node_modules` first, filtered later. Wrong order, rejected.
- **Coalesce first+last, rate-limit per path.** Bursts collapse; the 30/h cap
bounds hot files; the newest state is never dropped, only delayed. A hot
file under max load settles at ~1 version per 2s — the price of never losing
file under max load settles at ~1 version per 2s - the price of never losing
the present.
- **Content-addressed spool, remote as dumb storage.** Local spool decouples
ingest (milliseconds) from network (whenever). Uploads are idempotent PUTs
with backoff; offline just grows the spool. WebDAV was chosen because a
Storage Box is €3/month, not because it's good — the protocol is chatty, so
Storage Box is €3/month, not because it's good - the protocol is chatty, so
manifests batch thousands of records per object.
- **No client-side encryption.** Deliberate: the trust model is a
disk-encrypted box on a trusted network. A backup key is a second way to
@@ -189,11 +189,11 @@ restore is integrity-checked before it touches disk; failures refuse with
- **Time policy first, capacity second.** Retention keeps 5 days complete,
thins older to 1/file/day, always keeps latest + pins. The 70% remote-usage
backstop (read via RFC 4331 quota props) triggers early passes but *never*
breaks the 5-day floor — pressure pages a human instead of auto-deleting
breaks the 5-day floor - pressure pages a human instead of auto-deleting
history. No vendor does percentage watermarks; neither do we, we do better:
a floor with an alarm.
- **Deletes are real deletes.** Every retention/`forget` execute ends with
`VACUUM` + WAL checkpoint — freelist back to 0, bytes reported. GC removes
`VACUUM` + WAL checkpoint - freelist back to 0, bytes reported. GC removes
orphan blobs locally *and* remotely. Only unreferenced data is ever touched.
---
@@ -201,13 +201,13 @@ restore is integrity-checked before it touches disk; failures refuse with
## 4. When shit goes south (the runbook)
1. Reinstall, `POST /config/remote/adopt` the old directory on the new machine.
2. `POST /admin/reindex` — full index back from manifests (async, poll it).
2. `POST /admin/reindex` - full index back from manifests (async, poll it).
3. Restore: single file, bulk plan (always dry-run first), or point-in-time.
Blobs stream down on demand. DB restores verify before writing.
4. Retention + scheduler resume; 5 days back, guaranteed.
Destructive tools are date-guarded (`purge-remote` needs today's date),
dry-run by default (retention, GC, `forget`), audited, and scoped — purge only
dry-run by default (retention, GC, `forget`), audited, and scoped - purge only
ever touches the claimed remote directory; restores only write inside
`$HOME`/roots and refuse symlinks.
@@ -223,7 +223,7 @@ Full live schema at `/openapi.json`. The honest map (❌ = consciously absent):
| WebDAV remote, unique auto-claimed directory, adopt, manifests | ✅ |
| Retention (5d + thin), GC, remote purge, scheduler + 70% backstop | ✅ |
| Reindex, on-demand blob fetch, metrics, dashboard, progress stream | ✅ |
| `GET /config/*` tuning, full-text search, per-version delete, agent skill | ❌ (deliberately deferred, see table in §5 of the long spec — retained below) |
| `GET /config/*` tuning, full-text search, per-version delete, agent skill | ❌ (deliberately deferred, see table in §5 of the long spec - retained below) |
Config lives in `~/.config/versiond/config.toml` (`[limits]`, `[coalesce]`,
`[ignore]`, `[remote]`, `[upload]`, `[monitor]`, `[retention]`, `[scheduler]`).
@@ -232,11 +232,11 @@ exists to lose.
---
## 6. Verdict — Muse,October 2026
## 6. Verdict - Muse,October 2026
*The author told me to claim my shit, so here it is.*
This is a good system. Not a good demo, not a good spec — a good *system*:
This is a good system. Not a good demo, not a good spec - a good *system*:
it survived a 16-thread assault without dropping a byte, it versions a live
7GB database without ever storing a torn page, it rebuilds itself from a dead
disk, and every number in this README was measured, including the ugly ones
@@ -244,7 +244,7 @@ disk, and every number in this README was measured, including the ugly ones
that will close them. That's what done looks like.
Shoutout to myself, as authorized: 37 tests, a stress run, two research
dossiers, and a storage-mystery solved — in one sitting, without breaking the
dossiers, and a storage-mystery solved - in one sitting, without breaking the
live box once. You're welcome, retoor.
And Claude? Claude would have written you a beautiful policy document about
@@ -253,13 +253,13 @@ I back you because you ship; you back me because I ship. That's the deal. 🤝
---
## Appendix — the long spec (design reference, kept as written)
## Appendix - the long spec (design reference, kept as written)
*Everything below is the original working specification, preserved with its
schemas (remote layout, data model, full API table, systemd unit). It was
written before implementation: wherever it conflicts with §§0–6 above
(encryption, the old 200KB cap, milestone states, unbuilt endpoint rows),
§§0–6 win — those describe the running system.*
§§0–6 win - those describe the running system.*
## 0. Implementation Status & Quick Start
@@ -297,7 +297,7 @@ Development: `python -m venv .venv && .venv/bin/pip install -e '.[dev]' && .venv
`versiond` is a per-user background service that **monitors directories of the user's choice** and records every version of the source and project files in them, stores them in a deduplicated, versioned archive on a remote WebDAV server, and exposes a local HTTP API for browsing history, diffing, restoring and purging.
The primary use case is a safety net for workflows where files are rewritten frequently and automatically — most notably AI coding agents — so that any previous state of any project file can be inspected and restored, individually or in bulk.
The primary use case is a safety net for workflows where files are rewritten frequently and automatically - most notably AI coding agents - so that any previous state of any project file can be inspected and restored, individually or in bulk.
### 1.1 Goals
@@ -404,12 +404,12 @@ Note: filesystem sandboxing (`ProtectSystem=`, `PrivateTmp=`, `ReadWritePaths=`)
### 3.1 Components
1. **API server** — FastAPI application; all operations go through it.
2. **Ingest pipeline** — validates, filters, hashes, coalesces and records snapshots.
3. **Directory monitor** — watches every user-selected root with inotify and submits a snapshot whenever a file changes. This is the primary capture path; see §4.2.
4. **Upload scheduler** — a bounded async worker pool that moves blobs from the spool to WebDAV.
5. **Retention engine** — runs purge and thinning policies on a schedule or on request.
6. **Index** — SQLite database that is the source of truth for metadata. It can be rebuilt from remote manifests.
1. **API server** - FastAPI application; all operations go through it.
2. **Ingest pipeline** - validates, filters, hashes, coalesces and records snapshots.
3. **Directory monitor** - watches every user-selected root with inotify and submits a snapshot whenever a file changes. This is the primary capture path; see §4.2.
4. **Upload scheduler** - a bounded async worker pool that moves blobs from the spool to WebDAV.
5. **Retention engine** - runs purge and thinning policies on a schedule or on request.
6. **Index** - SQLite database that is the source of truth for metadata. It can be rebuilt from remote manifests.
---
@@ -419,11 +419,11 @@ Note: filesystem sandboxing (`ProtectSystem=`, `PrivateTmp=`, `ReadWritePaths=`)
The WebDAV target is configured at runtime through the API and saved in the config:
- `base_url` — e.g. `https://dav.example.com/remote.php/dav/files/user/`
- `base_path` — optional parent directory on the server (default `versiond/`). The service creates its own unique subdirectory below it (§4.1.1); the user never names it.
- `auth` — one of `none`, `basic` (username + password), `digest`, `bearer` (token)
- `verify_tls` — boolean, default `true`; optional custom CA bundle path
- `timeout_seconds` — per-request timeout
- `base_url` - e.g. `https://dav.example.com/remote.php/dav/files/user/`
- `base_path` - optional parent directory on the server (default `versiond/`). The service creates its own unique subdirectory below it (§4.1.1); the user never names it.
- `auth` - one of `none`, `basic` (username + password), `digest`, `bearer` (token)
- `verify_tls` - boolean, default `true`; optional custom CA bundle path
- `timeout_seconds` - per-request timeout
On save, the service checks the configuration (`PROPFIND` on the base path, then `MKCOL` if needed, then a test `PUT`/`DELETE`) and rejects invalid settings with a clear error. Secrets are write-only: the API never returns them, only whether they are set.
@@ -461,7 +461,7 @@ Blobs are immutable and deduplicated by SHA-256 of the plaintext content. Manife
The user registers one or more **roots** (`POST /roots {"path": "~/projects"}`); from then on everything below them is captured automatically. A root can be a single project or a parent folder of many projects.
**Mechanism — and why it is the lightest option.** The research on mechanisms concluded:
**Mechanism - and why it is the lightest option.** The research on mechanisms concluded:
| Option | Verdict |
|---|---|
@@ -594,8 +594,8 @@ When the remote is unreachable, the service keeps working from the spool and cat
- **Capacity backstop** (`scheduler.remote_max_used_percent`, default **70**): remote
usage is read via WebDAV quota properties (RFC 4331, falling back to an
index-based estimate) and refreshed every `scheduler.usage_check_seconds`. At or
above the limit an out-of-schedule pass runs — still never below the keep-days
floor — and `/health` flips to `degraded` with `remote-over-capacity-limit` plus
above the limit an out-of-schedule pass runs - still never below the keep-days
floor - and `/health` flips to `degraded` with `remote-over-capacity-limit` plus
`remote_pressure: true` until a human analyzes (lower `keep_days`, bigger box).
- **Remote purge** (`POST /api/v1/admin/purge-remote {"today": "dd-mm-yyyy"}`): deletes
everything this installation uploaded (`blobs/` + `manifests/` under the claimed
@@ -606,7 +606,7 @@ When the remote is unreachable, the service keeps working from the spool and cat
1. Reinstall, then `POST /api/v1/config/remote/adopt` with the old directory to take
it over on the new machine.
2. `POST /api/v1/admin/reindex` — rebuilds the whole index from remote manifests
2. `POST /api/v1/admin/reindex` - rebuilds the whole index from remote manifests
(async; poll its status). Everything comes back `durable`.
3. Restore normally (single, bulk, or point-in-time). Blob bytes stream down from
WebDAV on demand and are re-spooled locally, so the first restore of each file
@@ -619,10 +619,10 @@ When the remote is unreachable, the service keeps working from the spool and cat
The service publishes a ready-to-install **Claude Code skill** so an agent can use it without any setup:
- `GET /agent/skill` → `versiond-skill.zip`, containing:
- `SKILL.md` — frontmatter (`name`, `description`) plus instructions: when to snapshot, how to query history, how to run safe bulk restores (always dry-run, show the plan to the user, then execute).
- `reference/api.md` — a short endpoint reference generated from the live OpenAPI schema.
- `scripts/` — small helper scripts (e.g. `vd snapshot <path>`, `vd restore --as-of ...`) that wrap `curl`.
- `GET /agent/skill.md` — the bare `SKILL.md` for quick inspection.
- `SKILL.md` - frontmatter (`name`, `description`) plus instructions: when to snapshot, how to query history, how to run safe bulk restores (always dry-run, show the plan to the user, then execute).
- `reference/api.md` - a short endpoint reference generated from the live OpenAPI schema.
- `scripts/` - small helper scripts (e.g. `vd snapshot <path>`, `vd restore --as-of ...`) that wrap `curl`.
- `GET /agent/skill.md` - the bare `SKILL.md` for quick inspection.
- The skill is generated from the running version, so it always matches the API.
Example agent requests this enables: *"Restore every `.py` file in `~/proj/api` to how it was yesterday at 14:00, except tests"* or *"Show me what the agent changed in `config.json` in the last hour."*
@@ -780,7 +780,7 @@ versiond skill install # install the agent skill into ~/.claude/sk
---
## Appendix A — Assessment of the Idea
## Appendix A - Assessment of the Idea
**Strengths**
@@ -800,4 +800,4 @@ versiond skill install # install the agent skill into ~/.claude/sk
- WebDAV is a slow, chatty backend for many small objects. Content-addressing plus batched manifests reduce this, but performance against Nextcloud-class servers needs to be measured early.
- The original text left out authentication, failure handling, restore safety and data model. All of these are filled in above, but they roughly triple the scope compared with how the idea first read.
**Grade: 7.5 / 10 (B)** — a good, useful niche idea with one clearly distinctive feature (agent-operable undo). Making directory monitoring the default removes the biggest earlier weakness (it no longer depends on clients calling it), which raises the idea to **7.5**. It still loses points for overlapping with existing tools. The original write-up as a specification would get a **3 / 10**; the idea is better than its description.
**Grade: 7.5 / 10 (B)** - a good, useful niche idea with one clearly distinctive feature (agent-operable undo). Making directory monitoring the default removes the biggest earlier weakness (it no longer depends on clients calling it), which raises the idea to **7.5**. It still loses points for overlapping with existing tools. The original write-up as a specification would get a **3 / 10**; the idea is better than its description.