Add versiond-admin skill: deploy, capture policy, retention, DR runbook
This commit is contained in:
@@ -0,0 +1,80 @@
|
||||
# versiond admin skill
|
||||
|
||||
Operate the local `versiond` file-versioning service. API base is
|
||||
`http://127.0.0.1:9922`, every `/api/v1/*` call needs
|
||||
`Authorization: Bearer $(versiond token)`. State: `ok` is healthy,
|
||||
`degraded` means act. Remote storage is **unencrypted** (disk-encrypted
|
||||
server is the trust model); there are no keys to manage.
|
||||
|
||||
## Health first (read these before touching anything)
|
||||
|
||||
- `GET /health` → `status`, `degraded_roots`, `pending_paths`, `remote`,
|
||||
`remote_usage {percent,used,total,source}`, `remote_pressure`.
|
||||
- `GET /api/v1/progress` → scans, upload queue, monitor counters.
|
||||
- `GET /api/v1/stats` → files, versions, storage, top files.
|
||||
- `GET /api/v1/metrics` → Prometheus text.
|
||||
- `journalctl --user -u versiond -p err` → service errors.
|
||||
|
||||
## Deploy / service
|
||||
|
||||
- Install/reinstall: `pipx reinstall versiond` (from the repo), then
|
||||
`systemctl --user restart versiond`. Verify: new routes appear in
|
||||
`/openapi.json` (`/admin/*`, `/metrics`).
|
||||
- Never bind non-loopback: the server refuses to start otherwise.
|
||||
- `versiond status`, `versiond roots`, `versiond add <dir> [--force]`,
|
||||
`versiond progress`, `versiond stats`, `versiond dashboard`.
|
||||
|
||||
## Capture policy (what gets versioned)
|
||||
|
||||
- Kept: source, configs, dotfiles (allowlist), images, audio/video,
|
||||
archives, databases incl. SQLite journals (folded into verified snapshots),
|
||||
PDFs, fonts. Cap: `limits.max_file_bytes` (default 10 MiB).
|
||||
- Never: dependency/build/cache dirs (`node_modules`, `venv`, `target`, …),
|
||||
hidden dirs, `*.pyc/*.o/*.so`, `*.min.js/*.map`, temps (`*.tmp`, `~$*`,
|
||||
`#*#`, `*~`, `4913`), SQLite journals as standalone files.
|
||||
- SQLite rule: backup-API snapshot + `integrity_check`; locked/corrupt →
|
||||
loud skip (`database-locked`, `database-corrupt`), never a stored version.
|
||||
- Databases >~100MB or hot 7GB-class files: do NOT raise caps blindly —
|
||||
each version is a full copy until chunked blobs exist. Prefer dumps.
|
||||
|
||||
## Retention, GC, scheduler (MUST understand before running)
|
||||
|
||||
- `retention.keep_days` (default **5**): everything younger kept; older
|
||||
thinned to 1/file/day; per-file latest + pins always kept.
|
||||
- `POST /admin/retention/run` (dry_run default true) then
|
||||
`POST /admin/gc` (dry_run default true, `remote: true` also wipes WebDAV).
|
||||
- Built-in loop: retention daily, GC weekly (`[scheduler]`). No cron needed.
|
||||
- Capacity backstop `scheduler.remote_max_used_percent` (default 70):
|
||||
over-limit runs a pass but NEVER breaks the keep-days floor; health shows
|
||||
`remote_pressure: true` + `remote-over-capacity-limit` → human analyzes
|
||||
(lower keep_days, bigger box). Never auto-delete below the floor.
|
||||
|
||||
## Disaster recovery (total local loss), in order
|
||||
|
||||
1. Reinstall, `POST /config/remote/adopt {url, username, password,
|
||||
directory}` — takes over the old remote directory.
|
||||
2. `POST /admin/reindex` (202 + poll `GET /admin/reindex/{id}`) — rebuilds
|
||||
the index from manifests; everything returns `durable`.
|
||||
3. Restore (single/bulk/point-in-time). Blobs stream from WebDAV on demand.
|
||||
SQLite restores are integrity-checked; failures refuse with `unrestorable`.
|
||||
|
||||
## Destructive actions (double-check, dry-run first)
|
||||
|
||||
- `POST /admin/purge-remote {"today": "dd-mm-yyyy"}` — wipes this
|
||||
installation's remote `blobs/`+`manifests/` async (202 + counts, poll
|
||||
status). Only the claimed directory is ever touched. No local files harmed.
|
||||
- `POST /forget {"path", "dry_run"}` — deletes local history under a path.
|
||||
- Restores: always dry-run (`POST /restores` → plan), then
|
||||
`POST /restores/{id}/execute`. Pre-restore snapshots make every restore
|
||||
undoable. Targets must resolve inside `$HOME`/roots; symlinks refused.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
- `pending_paths` stuck high → coalescer backpressure or rate limits; check
|
||||
`capture:coalesced` vs `capture:committed` counters.
|
||||
- `remote: offline` → WebDAV down; spool grows, catches up alone. Check
|
||||
`last_error`, credentials file, server quota.
|
||||
- `remote_pressure: true` → over 70%: analyze, don't auto-purge.
|
||||
- Root `missing` → path vanished; service polls for return.
|
||||
- `watch-limit` on add → smaller root, raise `fs.inotify.max_user_watches`,
|
||||
or `--force` (degraded polling for the overflow).
|
||||
Reference in New Issue
Block a user