This commit is contained in:
2026-07-07 16:09:28 +02:00
parent 32c8bbe0a9
commit 5083efb150
42 changed files with 2556 additions and 2806 deletions
+190
View File
@@ -0,0 +1,190 @@
This file documents the service files that live directly in devplacepy/services/ (background task queue, AI correction/modifier, presence, live view relay, and the BaseService/ServiceManager base machinery). Claude Code loads it automatically whenever a file directly under devplacepy/services/ (or any of its subdirectories) is read or edited - domain-specific services (Devii, containers, gateway, jobs, audit, telegram, email, gitea, news, bot, messaging, xmlrpc, dbapi, pubsub, game) have their own more specific nested CLAUDE.md files.
## Background task queue (`services/background.py`)
Generic, lightweight, **fire-and-forget** offload for non-critical side-effects so they leave the request path. The singleton `background` (`from devplacepy.services.background import background`) wraps one in-process `asyncio.Queue` drained by a single consumer task. The whole API is one call: `background.submit(fn, *args, **kwargs)` enqueues a **synchronous** callable and returns immediately (`put_nowait`). The consumer runs each callable inside its own `try/except`, so a failing task is logged and never kills the loop; FIFO order is preserved.
- **Per-worker, not lock-gated (the reason it is NOT a `JobService`/`BaseService`).** Producers run in every uvicorn worker, and process memory is not shared, so the drain must run wherever requests are handled. `main.py` `startup()` calls `await background.start()` for every worker (inside the `if not DEVPLACE_DISABLE_SERVICES` guard, **outside** the `acquire_service_lock()` branch); `shutdown()` calls `await background.stop()`. A `BaseService` supervisor only runs in the lock owner, which would strand records produced by the other worker. The `JobService` queue is also wrong here: it is DB-backed, so deferring a tiny audit insert would *add* writes instead of removing them.
- **Inline fallback (load-bearing).** When the consumer is not running - tests with `DEVPLACE_DISABLE_SERVICES=1`, unit tests, request-less bootstrap, or a full queue - `submit` runs `fn` **inline and synchronously**. This keeps audit/notification/XP writes deterministic and immediately visible to the test suite (which asserts audit rows right after an action) while production defers them. No test changes are needed.
- **Best-effort durability (intentional).** In-memory only; `stop()` drains whatever remains synchronously so a **graceful** shutdown loses nothing, but a hard crash drops unflushed items. This matches audit/notifications already being best-effort (the recorder never raises). Do not put response-critical or money-touching work on it.
- **Capture plain data, never the `Request`.** Its lifecycle ends with the response - build the row/payload synchronously on the request thread and submit only the resulting dict/scalars.
- **Consumers today** (deferred at their canonical choke points, so callers need no change):
- **Audit** - `services/audit/record.py` `_write` builds the row + links synchronously, generates the `uid`/`created_at` eagerly so `record()` still returns the real uid, then `background.submit(_persist, row, links)` does the two `store` inserts off-thread.
- **XP/rewards** - `utils.award_rewards` defers its body via `background.submit(_apply_rewards, ...)`, so every XP award (posts/comments/projects/gists/follow/votes) leaves the request path.
- **Notifications** - `utils.create_notification` (the single notification funnel) defers via `background.submit(_deliver_notification, ...)`, so every in-app insert + push schedule + audit for a notification (vote/follow/comment/mention/message/badge/level) runs on the consumer.
- **Mention fan-out** - `utils.create_mention_notifications` defers its whole body (`_deliver_mention_notifications`): the `extract_mentions` regex + the `@`-username lookup query + the per-user loop all run off-thread (it is called on every post/gist/project/comment/message create).
- **Issue-comment admin fan-out** - `routers/issues/comment.py` defers the `_notify_admins` loop (admin lookup + per-admin notify) via `background.submit`, after the synchronous Gitea call.
- Because the reward and notification funnels self-defer, request handlers just **call `award_rewards`/`create_notification` directly** - do NOT wrap them in `background.submit` (that double-queues). The pattern for any NEW side-effect: do the user-visible write inline, then call the self-deferring funnel (or `background.submit(...)` a one-off).
- **Feature/first-use badges** ride the same self-deferring model: call `utils.track_action(user_uid, action[, target])` inline at a feature's success point (do NOT wrap it - it `background.submit`s itself), where `action` is a key in `utils.ACHIEVEMENTS` (`action -> [(threshold, badge), ...]`, threshold 1 = first use; `UNIQUE_ACTIONS` use `record_unique_activity` for "N distinct things" like `docs.read`). Source-derivable milestones (counts of existing tables) instead go in `check_milestone_badges`/`_COUNT_MILESTONES`. Badge metadata + `group` live in `BADGE_CATALOG`; the profile shows a grouped earned/locked **Achievements** showcase via `build_achievements`.
- **Deliberately NOT deferred (would break correctness):** cache invalidations (`clear_user_cache`/`clear_unread_cache`/`clear_messages_cache`/`bump_cache_version`) must run before the response so the next read is fresh (and they are microsecond version bumps); the vote/reaction count aggregation feeds the AJAX response body; and synchronous **external** calls whose result the response needs or must surface on failure (the Gitea comment/status calls; file/thumbnail writes whose returned URL must already exist) belong in an async **JobService** (durable + retryable), not this fire-and-forget queue.
## AI content correction (`services/correction.py`)
**Opt-in, default off, server-side.** When a user turns it on (profile **AI content correction** block, owner-only, saved at `POST /profile/{username}/ai-correction`), the prose they author is rewritten by the AI gateway. The per-user `ai_correction_sync` flag (0 = background default, 1 = sync) picks the **apply mode**: background rewrites the stored fields a moment after the write (never slows the request); sync makes the HTTP response wait until the correction is applied.
- **Sync must never block the event loop (the self-deadlock trap).** The correction POSTs to the in-process gateway (`INTERNAL_GATEWAY_URL` = `localhost:{PORT}`), and the hooked content helpers are synchronous on the loop thread, so a blocking inline call self-deadlocks the (single) worker against its own gateway request - on a single worker that loop is the only thing that can serve the gateway request it is waiting on, so a blocking inline call deadlocks the worker until the httpx timeout, then fail-softs (server hangs, no correction). The fix: sync runs `_run_correction` via `loop.run_in_executor` (loop stays free), stashes the future on `request.scope[PENDING_SCOPE_KEY]`, and the `await_pending_corrections` middleware in `main.py` awaits it after the handler. Never reintroduce a blocking inline correction on the loop thread.
- **Field registry is the single source of truth.** `CORRECTABLE_FIELDS: dict[str, tuple[str, ...]]` maps each correctable table to its prose columns: `posts` -> `(title, content)`, `projects`/`gists` -> `(title, description)`, `comments`/`messages` -> `(content,)`, `users` -> `(bio,)`. `gists.source_code`, project files, and Gitea issues are intentionally excluded - code and external systems are never corrected.
- **Choke entrypoint.** `schedule_correction(user, table, uid, request=None)` is the only thing handlers call (every hook passes `request`). It is a no-op unless: a user dict is present, `table` is in the registry, `user["ai_correction_enabled"]` is truthy, and the user has a non-empty `api_key`. In **sync** mode (`user["ai_correction_sync"]`) it calls `_run_inline_awaited`, which - only when a `request.scope` and a running event loop exist - submits `_run_correction` to `loop.run_in_executor` (off the loop thread) and stashes the future on `request.scope[PENDING_SCOPE_KEY]` for the middleware to await; if there is no request/loop it returns False and falls back to background. In **background** mode (default) it `background.submit`s `_run_correction` (per-worker queue; also runs inline under `DEVPLACE_DISABLE_SERVICES=1`). The same `_run_correction` worker function runs in all cases.
- **The hooks (covers UI + REST + Devii + devRant in one place):** `content.create_content_item` (posts/projects/gists), `content.edit_content_item`, `content.create_comment_record`, `content.edit_comment_record`, `services/messaging/persist.persist_message` (DMs - sender is the user), `routers/profile/index.update_profile` (bio), and the two devRant direct-update edit paths (`routers/devrant/rants.edit_rant`, `routers/devrant/comments.edit_comment`). devRant create paths route through the shared content cores, so they are already hooked. Code-only and external paths (project files, Gitea) are deliberately not hooked.
- **The gateway call is fail-soft.** `correct_text(api_key, prompt, text)` is synchronous (runs on the background worker thread), POSTs to `INTERNAL_GATEWAY_URL` with `model=INTERNAL_MODEL` via `stealth.stealth_sync_client`, authenticated with the user's own `Bearer` api_key (per-user attribution). It returns the ORIGINAL text on any error, empty output, or suspiciously large output (`len > len(text) * MAX_GROWTH_FACTOR + 200`, rejecting hallucinated expansion). `_run_correction` reads the row back, corrects each registry field that has non-blank text, writes only changed fields via `table.update(updates, ["uid"])` without touching `updated_at` (auto-correction is not a user edit) or the slug (slugs are permanent), and `clear_user_cache(user_uid)` when `table == "users"` so the corrected bio re-caches.
- **Per-user usage aggregation (only on success).** `correct_text` returns `(text, usage)`; `usage` is parsed from the gateway's `X-Gateway-*` response headers (`_usage_from_headers`) whenever the upstream call returned 200 (cost was incurred, even if the corrected output was rejected), else `None` on any failure. `_usage_from_headers` captures the token and cost headers PLUS the timing headers `X-Gateway-Upstream-Latency-Ms` and `X-Gateway-Total-Latency-Ms` (as `upstream_latency_ms`/`total_latency_ms`), so each call's timing is metered. `_run_correction` accumulates the per-field `usage` into one `totals` dict and, when `totals["calls"] > 0`, makes ONE call to `database.add_correction_usage(user_uid, totals)` (a single `totals` dict, not positional args) - so a 2-field content item is a single aggregated write, and a failed/empty correction records nothing. `add_correction_usage` (and `add_modifier_usage`) delegate to the shared `database._add_usage(usage_table, user_uid, totals)`: a single atomic `INSERT ... ON CONFLICT(user_uid) DO UPDATE SET col = col + excluded.col` upsert against the `correction_usage` table (per-user running SUMS: `calls`/`prompt_tokens`/`completion_tokens`/`total_tokens`/`cost_usd`/`upstream_latency_ms`/`total_latency_ms`/`updated_at`, unique index `idx_correction_usage_user`, the two latency columns REAL default 0.0, all ensured in `init_db`). It is a derived counter table (NOT in `SOFT_DELETE_TABLES`, like `gateway_usage_ledger`) and is deliberately separate from `users` so accumulating never invalidates the auth/user cache. `database.get_correction_usage(user_uid)` (via `_get_usage`) returns the stored sums PLUS computed averages: `avg_tokens` (total_tokens/calls), `avg_upstream_latency_ms`, `avg_total_latency_ms`, `avg_tokens_per_second` (completion_tokens over total upstream seconds), and `avg_cost_usd`.
- **Profile display:** `routers/profile/usage._correction_usage(uid, include_cost)` shapes it via the shared `_usage_view(data, include_cost)` (mirrors `_ai_quota`); `profile/index.py` builds it only for `is_owner or viewer_is_admin` and passes `include_cost=viewer_is_admin`, exposed as the `correction_usage` dict on the context and `ProfileOut`. The view surfaces the sums plus averages - `avg_tokens` (avg tokens/call), `avg_latency_ms` (avg upstream latency), `avg_total_latency_ms`, `avg_tokens_per_second` (avg speed), and `total_time_s` (total upstream seconds) - rendered as extra tiles on the card. **Financial gating:** tokens/call-count and the performance tiles show to the owner and admins; the dollar `cost_usd` and `avg_cost_usd` keys are present ONLY when `viewer_is_admin`, in BOTH the HTML card (`templates/profile.html`, `.correction-usage-*`) and the `respond(..., model=ProfileOut)` JSON (same rule as `_ai_quota`'s `spent_usd` - hiding it in the template alone would leak it to a member fetching their own profile as JSON). The card renders only when `correction_usage.calls` > 0.
- **Import-cycle discipline.** `correction.py` imports only `stealth`, `config`, `database.get_table`/`add_correction_usage`, and `services.background.background` at module top; `clear_user_cache` is imported lazily inside `_run_correction`. Never import `content` or `utils` at module top.
- **Settings live on `users`:** three columns `ai_correction_enabled` (0/1), `ai_correction_sync` (0/1, default 0 = background), and `ai_correction_prompt` (text, default `config.DEFAULT_CORRECTION_PROMPT`), ensured in `database.backfill_api_keys()` (the user column-ensure block run by `init_db`) and seeded born-live in `utils._create_account`. The edit route is the owner-or-admin leaf `POST /profile/{username}/ai-correction` (`routers/profile/ai_correction.py`, `AiCorrectionForm{enabled, sync, prompt}`, audit key `profile.ai_correction`). The owner-only values are exposed on the profile page context and `ProfileOut` (`ai_correction_enabled`/`ai_correction_sync`/`ai_correction_prompt`, gated by `is_owner`), the UI block lives in `profile.html` (owner-only: enable checkbox, **Apply mode** select, prompt textarea) wired by `static/js/AiCorrection.js` (`app.aiCorrection`), and Devii drives it via the owner-scoped `ai_correction_get`/`ai_correction_set` tools (`services/devii/ai_correction/`, `handler="ai_correction"`, `requires_auth=True`, not confirm-gated - it is a reversible per-user toggle; `ai_correction_set` accepts `enabled`, optional `sync`, optional `prompt`).
## AI modifier (`services/ai_modifier.py`, `services/ai_context.py`)
**A sibling of AI content correction that runs only on an explicit inline directive.** The engine reuses the correction plumbing wholesale (`CORRECTABLE_FIELDS`, `PENDING_SCOPE_KEY`, `gateway_complete`, the sync/background apply-mode mechanics, the per-user usage upsert) and differs only in the trigger and the apply-mode/enabled defaults: it is **enabled by default** and **synchronous by default**, and it runs ONLY where an authored prose field contains an inline `@ai <instruction>` directive.
- **The `@ai` gate is the whole difference.** `has_ai_directive(text)` matches the regex `@ai\s+\S` (case-insensitive). `schedule_modification(user, table, uid, request=None)` is a no-op unless a user is present, `table` is in `CORRECTABLE_FIELDS`, `user["ai_modifier_enabled"]` is truthy, the user has an `api_key`, AND at least one of the table's registry fields actually contains an `@ai` directive. `_run_modification` re-checks the gate per field, so untriggered fields are never sent to the gateway and never metered. Triggerless writes cost nothing. The configured prompt (default `config.DEFAULT_MODIFIER_PROMPT` = "Execute what is behind `@ai` (the prompt) and replace that part including `@ai`") tells the model to execute the instruction and replace the marked part including the `@ai` marker.
- **Total reuse of the correction layer.** `CORRECTABLE_FIELDS`, `PENDING_SCOPE_KEY`, `gateway_complete`, the sync/background apply-mode mechanics (`_run_inline_awaited` -> `loop.run_in_executor` + `request.scope[PENDING_SCOPE_KEY]` awaited by the `await_pending_corrections` middleware in `main.py`, the same self-deadlock-avoiding path), and the per-user usage upsert pattern are all imported from / mirror `services/correction.py`. `modify_text(api_key, prompt, text, context="")` composes a modifier system message and calls the shared `gateway_complete`. The same hooks fire it: `profile/index.update_profile` calls both `schedule_correction` and `schedule_modification`, and the content/comment/messaging cores invoke it alongside correction, so it covers the web UI, REST, devRant, and Devii in one place. Code and source files are never modified (same `CORRECTABLE_FIELDS` registry, `gists.source_code`/project files/Gitea excluded). `schedule_modification` is hooked alongside `schedule_correction` at the same content/comment/messaging/profile entrypoints.
- **Context-aware (modifier only, not correction).** Unlike correction, the modifier gives the model a grounding **context block** so an `@ai` instruction can reason about who is asking and what it is attached to. `services/ai_context.py` `build_context(table, uid, row, user_uid) -> str` assembles it; `_run_modification` builds it **lazily once per row** (only after a field is confirmed to contain `@ai`, so triggerless writes do no extra queries) and passes it to every field's `modify_text`, which appends it to the system message under a `# Context (use it to inform the result; never echo this block)` header. The block (fail-soft, length-capped, each part wrapped in try/except so a failed lookup never blocks the modification) has three parts:
1. **date** - `Today is DD/MM/YYYY on the DevPlace developer network.`
2. **author/stats** - the author's username, role, level, stars, post count (`get_user_post_count`), leaderboard rank (`get_user_rank`), follower count (`get_follow_counts`), member-since date, and bio (capped `MAX_BIO`).
3. **location/thread** - table-specific: a comment gets the target post/project/gist/news title + excerpt and the parent comment it replies to; a post gets its topic and attached project; a project/gist gets its sibling title/description and, for gists, the language + `source_code` (truncated, marked "reference only" - this is read context, the modifier still never WRITES `source_code`); a direct message gets the recipient and the last `MAX_THREAD_MESSAGES` messages of the conversation (each truncated).
All excerpts are length-capped (`MAX_*` constants) to bound token cost; the added context tokens are billed and metered like any other input via the usage stats. So `@ai answer the question above` in a comment, `@ai write my bio from my stats`, or `@ai reply to this` in a DM all work. The message-thread query binds a param named `:current` (NOT `:self`, which collides with `db.query`'s bound-method argument).
- **Defaults differ:** enabled **ON** by default and apply mode defaults to **synchronous** (`ai_modifier_sync` column defaults to 1), the inverse of correction.
- **Per-user usage (only on success).** Same plumbing as correction: `_usage_from_headers` captures the token, cost, and timing headers (`X-Gateway-Upstream-Latency-Ms`/`X-Gateway-Total-Latency-Ms`); `_run_modification` accumulates per-field gateway usage into a `totals` dict and, when `totals["calls"] > 0`, makes ONE `database.add_modifier_usage(user_uid, totals)` call (a single `totals` dict) that delegates to the shared `database._add_usage` atomic `INSERT ... ON CONFLICT DO UPDATE SET col = col + excluded.col` upsert into the `modifier_usage` table (running SUMS `calls`/token/`cost_usd`/`upstream_latency_ms`/`total_latency_ms` totals, the two latency columns REAL default 0.0, separate from `users` so it never busts the auth cache, NOT in `SOFT_DELETE_TABLES`, ensured in `init_db`). `database.get_modifier_usage(user_uid)` (via `_get_usage`) returns the sums plus the same computed averages as correction (`avg_tokens`, `avg_upstream_latency_ms`, `avg_total_latency_ms`, `avg_tokens_per_second`, `avg_cost_usd`).
- **Settings live on `users`:** three columns `ai_modifier_enabled` (0/1, default 1), `ai_modifier_sync` (0/1, default 1 = sync), and `ai_modifier_prompt` (text, default `config.DEFAULT_MODIFIER_PROMPT`), ensured in `database.backfill_api_keys()` (the user column-ensure block run by `init_db`) and seeded born-live in `utils._create_account`. **Because the feature is on by default, `backfill_api_keys()` ALSO runs a `with db:` `UPDATE users SET ... WHERE ... IS NULL` so EXISTING accounts inherit the on/sync defaults** (a freshly `create_column_by_example`-added column is NULL on existing rows, which would read as disabled - the backfill is what makes "on by default" true for everyone, not just new accounts). The `with db:` wrapper is load-bearing: a raw `db.query` write that does not commit holds the SQLite write lock. The edit route is the owner-or-admin leaf `POST /profile/{username}/ai-modifier` (`routers/profile/ai_modifier.py`, `AiModifierForm{enabled, sync, prompt}`, audit key `profile.ai_modifier`). The owner-only values are exposed on the profile page context and `ProfileOut` (`ai_modifier_enabled`/`ai_modifier_sync`/`ai_modifier_prompt`, gated by `is_owner`); `modifier_usage` is built by `routers/profile/usage._modifier_usage(uid, include_cost=viewer_is_admin)` (via the shared `_usage_view`) for `is_owner or viewer_is_admin`, surfacing the sums plus the performance averages (avg tokens/call, avg latency, avg tokens/sec, total time) like `correction_usage`; the dollar `cost_usd` and `avg_cost_usd` are present only when `viewer_is_admin`, in both HTML and JSON. The UI block lives in `profile.html` (owner-only: enable checkbox, **Apply mode** select defaulting to sync, prompt textarea, plus the usage card reusing the `.correction-usage-*` classes with the extra performance tiles) wired by `static/js/AiModifier.js` (`app.aiModifier`). Devii drives it via the owner-scoped `ai_modifier_get`/`ai_modifier_set` tools (`services/devii/ai_modifier/`, `handler="ai_modifier"`, `requires_auth=True`, not confirm-gated; `ai_modifier_set` accepts `enabled`, optional `sync`, optional `prompt`).
## Background services base machinery (`services/base.py`, `services/manager.py`)
`devplacepy/services/` provides a generic framework for running background async services alongside the FastAPI server. `BaseService` provides the async run loop, a `deque(maxlen=20)` log buffer, and graceful cancellation; `ServiceManager` is a singleton that registers, starts, and stops services. Services are fully managed from the **Services admin tab** (`/admin/services`): start/stop, enable-on-boot, run-now, edit parameters, clear logs, adjustable log buffer size, with live status. All of this is generic - a new service gets it for free by declaring its config and implementing `run_once`.
This section covers only the shared machinery. The individual services built on top of it live in their own subdirectories with their own nested CLAUDE.md: `NewsService` (`services/news/`, see `devplacepy/services/news/CLAUDE.md`), `GatewayService` and provider/model routing (`services/openai_gateway/`, see `devplacepy/services/openai_gateway/CLAUDE.md`), `DeviiService` (`services/devii/`, see `devplacepy/services/devii/CLAUDE.md`), and the bot fleet service (`services/bot/`, see `devplacepy/services/bot/CLAUDE.md`).
### DB-backed state (correct across workers)
In production (`make prod`, 2 workers) only the lock-holding worker runs services, but an admin request can hit either worker. State therefore lives in the DB, not process memory:
- **Desired/config state** in `site_settings` (via `get_setting`/`set_setting`): `service_<name>_enabled` (`"0"`/`"1"` - also the boot flag), `service_<name>_command` (`"<verb>:<counter>"`, verbs `run`/`clear`), `service_<name>_log_size`, plus each declared config field's own key.
- **Observed state** in the `service_state` table (one row per service): `status`, `last_run`, `next_run`, `started_at`, `heartbeat`, `logs` (JSON), `updated_at`. Written by the supervising worker, read by any worker.
`main.py` registers every service in **all** workers (so `describe_all()` works anywhere) but only the lock worker calls `service_manager.supervise()`.
### `ConfigField` (`services/base.py`)
Declarative parameter spec. `type` in `int`/`str`/`url`/`text`/`password`/`bool`/`select`. `coerce(raw)` validates and types (raises `ValueError`, never silent); `read()` is the lenient runtime reader (falls back to `default`); `spec()` renders the field for the template/JSON (secrets masked, never sent). Optional `minimum`/`maximum`/`options`/`help`/`secret`. `group="..."` lets the detail page render fields as sections (built-in fields are grouped "General"/"Advanced").
### `BaseService` (`services/base.py`)
Abstract class for all services:
- **`name`**, **`interval_key`** (default `service_<name>_interval`; a subclass may point it at an existing key), **`min_interval`**.
- **`config_fields`** - class-level list of `ConfigField`. The framework prepends built-in `enabled` and interval fields and appends `log_size`; `all_fields()` returns the full set.
- **`get_config()`** returns a typed dict from settings; **`is_enabled()`**, **`current_interval()`** (floored to `min_interval`).
- **`log(message)`** - writes to `log_buffer` and standard `logging`.
- **`run_once()`** - abstract; override with actual work. Read parameters via `self.get_config()`.
- **Reconciling loop** (`_run_loop`/`_tick`, ~1s tick): honors `enabled`, runs `run_once()` when due or on a `run` command, handles `clear`, rebuilds the log deque if `log_size` changed, and persists observed state (throttled, force on transitions). Start/stop/run/clear from any worker are just DB writes the loop reacts to within ~1s - services are never hard-restarted.
- **Lifecycle hooks** `async on_enable()` / `async on_disable()` - called on transition to enabled / disabled (and `on_disable` on shutdown). Override for long-running resources that outlive one `run_once` (e.g. the bot fleet); `run_once` then becomes a periodic reconcile/metrics tick.
- **`collect_metrics() -> dict`** - generic live metrics. Return `{"stats": [{"label", "value"}], "table": {"columns": [...], "rows": [[...]]}}`. Persisted to `service_state.metrics` each tick and rendered live in the Services tab (stat cards + table) by `ServiceMonitor.js`. Default returns `{}`.
- **`default_enabled`** class flag (default `True`) - set `False` for opt-in services so they never auto-start on boot (the bots service uses this).
- **`title` / `description`** class attrs - shown in the admin UI; every service should set a `description` of what it does.
**Admin UI** (`routers/admin/services.py`): `/admin/services` is a slim **index** (`services.html`, an `admin-table` of name/description/status/uptime/last-run/next-run/interval) where each row links to `/admin/services/{name}`. The **detail** page (`service_detail.html`) holds everything for one service - header with controls (start/stop/run/clear), and `[data-tabs]` tabs **Overview** (meta + metrics), **Configuration** (the config form rendered as a `<fieldset>` per `field_groups` entry), **Logs**. `ServiceMonitor.js` is dual-mode (polls `/admin/services/data` for the index, `/admin/services/{name}/data` for the detail, via the shared `Poller`; its config-save POST goes through `Http.sendForm`) and shares one `update(root, svc)` over `data-*` hooks; tabs use the generic reusable `static/js/Tabs.js` (`[data-tabs]`/`[data-tab]`/`[data-tab-pane]`, wired in `Application.js`). Adding a service still needs zero UI code.
- **`describe()`** - builds the dict the UI consumes: `title`, `description`, derived display status (`stopped` when disabled, `running` with fresh heartbeat, `stalled` when enabled but heartbeat older than `STALE_SECONDS`), meta, logs, metrics, flat `fields`, and `field_groups` (ordered `{name, fields}` sections) read from `service_state` + settings.
### `ServiceManager` (`services/manager.py`)
Singleton that manages all registered services:
- `register(service)` - add a service
- `describe_all()` -> `list[dict]` (per-service `describe()`)
- `set_enabled(name, bool)` - Start/Stop (writes the enabled flag)
- `send_command(name, "run"|"clear")` - one-shot command channel
- `save_config(name, form)` - strict validation via each `ConfigField`; blank secret = keep existing; returns `{ok, errors}`
- `supervise()` - start the reconciling task per service (lock worker only)
- `shutdown_all()` - cancel tasks on server shutdown
### Adding a new service
1. Create `devplacepy/services/your_service.py` with a class extending `BaseService`. Declare `config_fields` (and set `interval_key`/`min_interval` if reusing an existing key); read them in `run_once` via `self.get_config()`.
2. Override `async def run_once(self) -> None`.
3. Register in `main.py` startup event (outside the lock block, so all workers know the roster):
```python
from devplacepy.services.your_service import YourService
service_manager.register(YourService())
```
4. The service appears automatically on `/admin/services` with start/stop, enable-on-boot, run-now, a generic config form, clear-logs, live status, and the log tail - no template, JS, or router changes needed.
### CLI
```bash
devplace news clear # Delete all news from local database
devplace devii reset-quota <username> # Reset one user's rolling 24h AI quota
devplace devii reset-quota --guests # Reset every guest quota
devplace devii reset-quota --all # Reset every quota (users and guests)
```
## Multi-worker concurrency (preferred rules)
`uvicorn --workers N` = N independent processes sharing only the filesystem and SQLite DB. Module-global caches/counters are per-process, so a local `clear()` is invisible to siblings. Full reference: admin docs `Production -> Multi-worker and concurrency` (`templates/docs/production-concurrency.html`). Enforce these:
- **In-process caches are never authoritative across workers.** Coordinate invalidation through the `cache_state` version table in `database.py`: call `bump_cache_version(name)` at every write path and `sync_local_cache(name, cache)` before every read path. Existing names: `auth` (guards `_user_cache`; bumped by `clear_user_cache`/`clear_session_cache` on logout, ban, role/password change), `settings` (guards `_settings_cache`; bumped by `set_setting`/`clear_settings_cache`), `relations` (`_relations_cache`), `customizations` (`_customizations_cache`), `notif_prefs` (`_notification_prefs_cache`), `gateway_routing` (`_ROUTING_CACHE`), and **`admins`** (guards `_admins_cache`, which memoizes `get_admin_uids()` / `get_primary_admin_uid()`). The `_cache_version_cache` (`ttl=1`) caches the version reads themselves, so the bump is seen by other workers within ~1s on their next read (the originating worker is immediate); the version read/bump fail open (logged, never raise). When you add a per-process cache whose staleness matters, give it a name and wire both calls - do not invent a second invalidation mechanism.
- **Admin-set cache invariant (load-bearing).** `_admins_cache` makes `get_admin_uids()` / `get_primary_admin_uid()` (hit on the projects-listing visibility filter, `_owner_is_admin`, and every primary-admin gate: `/dbapi`, backup-archive download) a dict lookup instead of a per-call `SELECT`. It is NOT the authorization gate - `is_admin(user)`/`require_admin` read the role off the user object (independently invalidated via `clear_user_cache`), so a stale admin set cannot grant access. But **any code that writes `users.role` MUST call `database.invalidate_admins_cache()`** (clears local + bumps the `admins` version), exactly like the soft-delete and `with db:` rules. Current role-write sites all do: `routers/admin/users.py` (role change), `cli.py` (`role set`), and `utils._create_account` (first user -> Admin). Omitting the call leaves the admin set stale for up to the 300s TTL.
- **Deterministic / eventual caches skip version-sync on purpose.** Two caches need no `cache_state` name because correctness does not depend on cross-worker freshness: (1) `docs_prose._render_markdown` is an `@lru_cache` on the **static** prose source string (pure function of template content, so identical on every worker; changes only on deploy), and (2) `main._home_cache` (`TTLCache ttl=60`) memoizes the guest `/` blocks - the featured-news block (identical for everyone) and the latest-posts block **only for the no-block case**; a viewer with a non-empty block set bypasses the cache and is computed fresh, so the original block-filter semantics are preserved byte-for-byte and there is zero cross-user leakage. Worst case is a soft-deleted/edited public post lingering on `/` for <=60s. Use a plain `TTLCache`/`lru_cache` (no version name) ONLY when the value is deterministic or its staleness is purely cosmetic; anything whose staleness affects correctness or permissions MUST use the version-sync pattern above.
- **Authoritative state lives in the DB**, memory is only a short-TTL accelerator (sessions, the Devii 24h cap via `devii_usage_ledger`, cost counters).
- **Global rates/quotas are worker-count-aware or DB-backed.** The rate limiter enforces `ceil(limit / DEVPLACE_WEB_WORKERS)` per process so the aggregate matches the configured value with no per-request write. `DEVPLACE_WEB_WORKERS` MUST equal the real `--workers` (set in `make prod`=2 and the Dockerfile=2; default 1 for dev). Update it whenever you change worker count.
- **First-boot side effects are flock-serialized and idempotent.** `init_db()` runs under a blocking `flock` on `devplace-init.lock` (the seed-if-missing inserts would otherwise race to duplicate rows); `ensure_certificates()` early-returns when keys exist and otherwise generates under `.vapid.lock`. Any new run-once-at-startup work follows the same pattern (flock + exists-check / `INSERT OR IGNORE`).
- **Run-once background work stays behind the services lock**, never per-worker: gate it on `service_manager.owns_lock()` (held via `devplace-services.lock`), as news/bots/Devii hub do.
## Performance caches (request-path)
Read-mostly aggregates that were recomputed per request or per vote sit behind short per-worker `TTLCache`s (`devplacepy/cache.py`). These are DISPLAY caches: staleness up to the TTL is accepted by design, so none of them uses the cross-worker `cache_state` versioning (that is reserved for correctness-critical caches like auth/settings/admins). The registry:
| Cache | Where | Key | TTL | Invalidation |
|---|---|---|---|---|
| `_authors_cache` | `database/ranking.py` | `ranked`/`rank_map` | 15s | TTL only - `update_target_stars` deliberately does NOT clear it per vote anymore (the old per-vote `clear()` forced a full UNION-JOIN ranking recompute on the next feed/leaderboard/landing request and never propagated cross-worker anyway) |
| `_stars_cache` | `database/ranking.py` | user uid | 15s | TTL + `clear_user_stars(owner_uid)` from `content.apply_vote`, so a user's own total updates immediately on the voting worker |
| `_leaderboard_cache` | `services/game/store/farm.py` | `top:{limit}` | 15s | TTL only (the farm scan + Python `farm_score` sort runs at most once per 15s per worker) |
| `_projects_cache` | `templating.py` | user uid | 10s | TTL + `clear_user_projects_cache(uid)` from the `content.py` project create/delete choke points (in-function import - `templating` imports `content` at module level). `jinja_user_projects` also filters `deleted_at=None` now (the composer dropdown previously listed soft-deleted projects) |
Rules: a new hot read-path aggregate follows this exact pattern (module-level `TTLCache`, 10-15s, `clear_*` helper only when a same-worker write must be visible immediately); never reintroduce a whole-cache `clear()` on a per-event write path; profile `projects`/`gists` loads are tab-gated in `routers/profile/index.py` (`tab == X or wants_json(request)` - the JSON `ProfileOut` keeps serializing both, HTML tabs that do not render them get `[]`).
## Live view relay (`services/live_view_relay.py`)
`LiveViewRelayService` is a second lock-owner pub/sub bridge (default-enabled, 1s interval, registered in `main.py` after `NotificationRelayService`) that replaces the per-client HTTP polling on the admin live views with server push, **computing a snapshot only for topics that currently have subscribers**. Each tick it reads `pubsub.topics()` (the hub's live subscription set, which converges on the lock owner), matches each concrete subscribed topic against the `VIEWS` registry (a list of `(compiled_regex, compute_callable, min_interval_seconds)`), throttles per topic via a `time.monotonic()` map, computes the payload, and publishes it on the same topic. The compute callables (async, lazy-importing to avoid cycles, defensive try/except -> skip) reproduce the **exact payload shape the matching HTTP endpoint already returns**, so the frontend render code is unchanged - only the data-arrival path is added. Registry:
| Topic | Cadence | Payload (same as endpoint) |
|-------|---------|----------------------------|
| `container.list` | 4s | `{instances}` (admin all-instances, `_decorate`) |
| `project.{slug}.containers` | 3s | `{instances}` (per-project, slug resolved via `resolve_by_slug`) |
| `container.{uid}.detail` | 4s | `{instance, events, schedules, stats, runtime}` |
| `container.{uid}.logs` | 3s | `{logs}` (backend log tail 400) |
| `fleet.bots` | 2s | bot frames payload (`routers/admin/bots._frames_payload`) |
| `admin.services` | 5s | `{services}` (`service_manager.describe_all()`) |
| `admin.services.{name}` | 5s | `{service}` |
| `admin.ai-usage.{hours}` | 15s | `build_analytics(hours)` (hours parsed from the topic) |
| `admin.backups` | 8s | `routers/admin/backups._dashboard(can_download=False)` (storage, backups, schedules, metrics) |
All these topics are admin-only by pub/sub policy (non-`public`, non-`user.{uid}` -> `privileged` required), matching the admin-only pages. Frontend monitors (`ContainerInstance`, `ContainerList`, `ContainerManager`, `BotMonitor`, `ServiceMonitor`, `AiUsageMonitor`, `BackupMonitor`) each `window.app.pubsub.subscribe(topic, render)` in their init and keep a **lengthened HTTP poll (15-30s) as initial-load + fallback** - the relay drives liveness at the cadence above. `AiUsageMonitor` re-subscribes (unsubscribe old, subscribe new) when the window-hours selector changes, since hours is in the topic.
**Container topics never broadcast private-project instances**: `container.list` publishes only public-project rows with `partial: true` (`ContainerList.merge` updates by uid, never removes, so private rows from the authoritative HTTP poll survive), and `project.{slug}.containers` / `container.{uid}.detail` / `container.{uid}.logs` skip private-project targets entirely (owners fall back to their HTTP polls).
**`admin.backups` is the one view with a per-viewer field:** the backup `download_url` is restricted to the primary administrator at every endpoint, so the relay publishes the snapshot with `can_download=False` (the bus broadcasts one payload to every admin subscriber, and is therefore treated as another endpoint that must withhold it). `BackupMonitor` derives `canDownload` from its authoritative per-viewer HTTP poll only (never from a pushed frame) and builds the `/admin/backups/{uid}/download` URL client-side; the download route itself stays primary-admin-gated.
To add a new admin live view: add one `(regex, compute, interval)` row to `VIEWS` whose compute returns the endpoint payload, and have the frontend subscribe to the topic while keeping a fallback poll. Reuse this subscriber-gated snapshot pattern for any future "several admins watch a periodically-recomputed server snapshot" surface; keep durable/ordered/stateful/raw-I/O channels (messages, Devii, SEO/DeepSearch progress, container PTY) OFF pub/sub. See `pubreport.md` for the full migration analysis.
## Notification relay (`services/notification_relay.py`)
**Live toasts ride the `in_app` channel** (no new channel). `NotificationRelayService` is a lock-owner `BaseService` (registered in `main.py`, default-enabled, 1s interval) modeled on `MessageRelay`: it primes a watermark to `MAX(notifications.id)` on first tick, then each tick selects `notifications WHERE id > watermark` and publishes `{uid, type, message, target_url}` to the per-recipient pub/sub topic `user.{user_uid}.notifications`. It runs **only on the service lock owner**, which is exactly where every pub/sub WS subscriber converges (non-owner `/pubsub/ws` closes `4013`), so the in-process `services.pubsub.publish` reaches the recipient's browser. Because the `notifications` row exists only when the `in_app` channel is enabled, a live toast fires **exactly** for the notifications a user has enabled - the relay needs no preference lookup of its own. The watermark always advances (boot-priming + every-tick) so a newly-connected subscriber never gets a backlog flood.
**Frontend:** `base.html` exposes the viewer uid as `<body data-user-uid>`, `static/js/LiveNotifications.js` (`app.liveNotifications`, constructed with `this.pubsub`/`this.toast`) subscribes to that topic and calls `app.toast.show(message, {type:"info", ms, url})`. **The toast click reproduces a bell-item click exactly:** it targets `GET /notifications/open/{uid}` (which marks the row read and redirects to its `target_url`), not the bare `target_url` - the relay publishes `uid` for this. `AppToast` gained a generic click action via `_action(options)`: `options.onClick` (a function) or `options.url` (navigate); either adds the `.dp-toast-link` pointer cursor, so any caller can attach a click action. DMs toast too (type `message` is not excluded). Reuse this relay-on-the-lock-owner pattern for any future "live mirror of a DB-persisted, per-user event."
**Unread-count badges ride the same relay.** The same `NotificationRelayService` tick, after publishing the toast rows, also publishes the recipient's fresh unread counts `{notifications, messages}` to `user.{uid}.counts` (computed with a direct DB query on the lock owner, never the per-worker `_unread_cache`, which could be stale there). Because a new DM also inserts a `message`-type notification row through `persist_message -> create_notification`, this one publish point covers both new notifications and new messages, so the header badge bumps instantly. `static/js/CounterManager.js` (`app.counters`, constructed with `this.pubsub`) subscribes to `user.{uid}.counts` and applies the pushed counts; its HTTP poll of `/notifications/counts` stays as a 60s reconciliation fallback (covers decrements on read and any missed frame). Read events still clear the per-worker cache as before.
## Online presence (`services/presence.py`, `services/presence_relay.py`)
Online status is a single **`users.last_seen`** UTC-ISO column (ensured in `database.backfill_api_keys`), never a new table and never a per-request insert. The design is dictated by the multi-worker architecture: a profile of user X renders on **any** worker, so "is X online" needs cross-worker state, and pub/sub is in-process only (a non-owner worker cannot publish to the lock owner), leaving SQLite as the sole shared medium. So presence is a heavily-throttled in-place UPDATE.
**Write path (all workers):** `main.py`'s `track_presence` HTTP middleware resolves the cached current user on every non-`/static`, non-`/avatar` request and calls `presence.touch(uid)`. `touch` keeps a per-worker in-memory `_last_write: dict[uid -> monotonic]` and writes `users.last_seen` (via `database.set_last_seen`) only when the last write for that uid is older than `config.PRESENCE_WRITE_SECONDS` (= `PRESENCE_TIMEOUT_SECONDS // 2`). So continuous browsing is a dict lookup; a write happens at most ~once per half-window per active user per worker, and the row is updated in place (zero growth). It deliberately does **not** call `clear_user_cache` (that would defeat the 300s auth cache; the stale cached self-row is irrelevant since presence of *other* users is always read from a fresh row).
**Read path (any worker):** `presence.is_online(user_row)` = `now - last_seen < PRESENCE_TIMEOUT_SECONDS` (env `DEVPLACE_PRESENCE_TIMEOUT_SECONDS`, default 60). Profile (`routers/profile/index.py` -> `profile_online`) and messages (`routers/messages.py` seed) read `last_seen` off the user row they already loaded - no extra query. Exposed as the Jinja global `is_online(user)` (`templating.py`), on `UserOut.last_seen` and `ProfileOut.profile_online`. This is the **only** cross-worker-correct approach here because pub/sub is in-process.
**Live path (lock owner only), change-only + hysteresis:** `PresenceRelayService` (`services/presence_relay.py`, `BaseService`, default-enabled, 2s tick, registered in `main.py`) is a sibling of `NotificationRelayService`/`LiveViewRelayService`. Each tick it recomputes ONE global online set `self._online` (dot-subscribed uids batch-read via `get_users_by_uids` + roster candidates) and drives BOTH the per-user dots and the feed roster from that one set, so they can never disagree. The set uses **hysteresis** via `presence.stays_online(elapsed, was_online)`: a user becomes online at `PRESENCE_TIMEOUT_SECONDS` but only drops after `+ PRESENCE_ONLINE_MARGIN_SECONDS` (env `DEVPLACE_PRESENCE_ONLINE_MARGIN_SECONDS`, default 20) - **quick to go online, slow (grace margin) to go offline** - which kills boundary flicker for a user hovering near the timeout. It publishes `{online, last_seen}` to `public.presence.{uid}` **only when a topic's `online` bool changed OR the topic is newly subscribed (first-seen)** - never on a fixed interval, so a steady page emits nothing after the initial frame (`self._published[topic] -> bool`, pruned to active topics). `public.presence.{uid}` is subscribable by any logged-in user (`pubsub/policy.py` allows `public.*`); guests fall back to the server-rendered initial state. The one-directional grace also means dots and roster stay consistent across viewers (the shared `self._online` is the single authority). To keep it lightweight the relay reads all due users in **one batched `get_users_by_uids`** per tick.
**Frontend:** `static/js/PresenceManager.js` (`app.presence`, constructed with `this.pubsub` in `Application.js`, mirroring `LocalTime`/`CounterManager`) scans `[data-presence-uid]` elements, subscribes each uid to `public.presence.{uid}` (deduped per uid, so a repeated author is one subscription), and treats each frame's `online` flag as **authoritative** (`entry.online`), toggling the `online` class + (for `data-presence-label` elements) the "online / last seen X / offline" text. Because the relay is change-only and authoritative, a live-subscribed dot is **not** expired by the client clock (no false-offline flicker for an active user whose `last_seen` the client cannot see advancing); the 20s `last_seen` staleness timer only applies to entries that never received a frame (guests / degraded). The window is read from `<body data-presence-timeout>`. Both the profile `.profile-presence` dot and the messages `#messages-presence` span carry `data-presence-uid`/`data-presence-last-seen`.
**Online-now roster (feed):** the same relay maintains ONE shared topic `public.presence.roster`, republished **only when the SET of online uids changes** (a `frozenset` compare, so pure reordering never republishes). `services/presence.py` `online_users(limit)` (strict, feed initial render) and `online_candidates(limit)` (grace window, relay hysteresis) both go through `database.get_online_users(cutoff_iso, limit)`, which reads users with `last_seen >= cutoff` via the `idx_users_last_seen` index (the one place presence is queried by `last_seen`; `config.PRESENCE_ONLINE_LIMIT`, env `DEVPLACE_PRESENCE_ONLINE_LIMIT`, default 30). **The list is ordered ALPHABETICALLY by username** (`get_online_users` `order_by=["username"]` + case-insensitive `presence.sort_by_username`), NOT by recency, so avatars keep a stable position and do not needlessly reshuffle as people's `last_seen` ticks. `routers/feed.py` puts `online_users` on the context (`FeedOut.online_users`) and `feed.html` renders the initial **Online now** panel as a `.sidebar-section` at the bottom of the left feed sidebar (`aside.sidebar-card`); `static/js/OnlineUsers.js` (`app.onlineUsers`) subscribes to `public.presence.roster` and re-renders the avatar list + count live. Roster avatars use a plain green `.presence-dot` with NO `data-presence-uid` (list membership IS the presence, so no per-user subscription - the relay drops a user from the roster when they go offline).
**Avatar presence dot (sitewide, DRY):** a small corner dot on **every** user avatar (green online, muted grey offline) comes from ONE reusable partial `templates/_presence_dot.html` - `<span class="presence-dot" data-presence-uid data-presence-last-seen>` guarded on `_user.get('uid')` (a partial-dict author, e.g. the issues includes, renders no dot). It carries **no** `data-presence-label`, so `PresenceManager` colours it with zero extra JS. It is included by the shared avatar partial `templates/_avatar_link.html` (its `.user-avatar-link` anchor is the positioning host, covering ~19 sites) and by the handful of raw-`<img class="avatar-img">` sites wrapped in a positioned `<span class="avatar-badge">` (the two `base.html` nav avatars, the `profile.html` hero + followers list, the `messages.html` conversation list). CSS in `static/css/base.css` (`.user-avatar-link`/`.avatar-badge` `position:relative;display:inline-flex`, `.presence-dot` sized `30%` of the avatar clamped 8-14px with a `--bg-card` ring, `.online` -> `--success`), so it is proportional and responsive at every avatar size with no per-size class. `database/follows.py` `get_follow_list` now carries `last_seen` in its trimmed dict so the followers/following dots resolve (all other author dicts are full `get_users_by_uids` rows). `dp-avatar` (`AppAvatar.js`) is docs-demo only (no real user avatars) and is intentionally out of scope. Reuse `_presence_dot.html` + the `.avatar-badge` wrapper for any new avatar surface - never hand-roll a presence dot.
**Messaging refactor:** the old presence was per-worker and WS-connect-based (`message_hub.is_online`/`last_seen`, `_announce_presence`, the WS `presence` frame) and broke with >1 worker. That display path was **removed**; `message_hub` keeps only its socket connection tracking for message delivery. The messages header presence is now the shared `PresenceManager`, so chat presence is finally cross-worker correct. **Never re-implement WS-connect presence** - reuse `presence.is_online`, the `public.presence.{uid}` topic, and `PresenceManager`.
+58
View File
@@ -0,0 +1,58 @@
This file documents the audit log subsystem. Claude Code auto-loads it when a file under `devplacepy/services/audit/` is read or edited.
## Audit log (`services/audit/`)
**Admin-only, append-only** record of every state-changing action. The authoritative event catalogue is `events.md` (its storage model, relation vocabulary, and per-event specs are authoritative); the catalogue currently spans 223 keys across 38 domains.
### Package layout
- `store.py` - the `audit_log` + `audit_log_links` tables, `ensure_tables`, `insert_event`/`insert_links`/`get_event`/`get_links`/`sweep`.
- `categories.py` - `category_for(event_key)` (longest-prefix map).
- `record.py` - the recorder + link builders.
- `query.py` - `list_events`/`filter_options`/`get_event_with_links` for the admin UI (admin list/detail).
- `service.py` - `AuditService` (retention sweep).
`ensure_tables()` + the 11 indexes are wired into `database.py` `init_db()` via a **local import** (audit modules import `database`/`utils`, so the import is deferred to avoid the cycle). `AuditService` is registered in `main.py`.
### Two recorder entrypoints (the reuse keystone)
`record(request, event_key, *, user=_UNSET, actor_kind=None, target_type/uid/label, old_value, new_value, summary, metadata, result="success", origin, via_agent, links, category)` is for HTTP/WebSocket handlers - `request` may be a Starlette `Request` OR `WebSocket` (both expose `.headers`/`.url`/`.client`). It auto-derives the actor from `get_current_user(request)` (or an explicit `user=`), the request fields, and the **`X-Devii-Agent`** header (-> `origin=devii`, `via_agent=1`), and auto-appends the `actor` link.
`record_system(event_key, *, actor_kind, actor_uid, actor_username, actor_role, origin, via_agent, ...)` is for request-less contexts (rewards, notifications, the AI gateway ledger, the news/container services, Devii turns/tasks, job services, the CLI).
**Both are best-effort: wrapped in try/except, they NEVER raise into the caller** - the audited action is never blocked by a logging failure. `summary` is HTML-stripped and capped at 140 chars; `metadata` is JSON-serialised. Link builders (`audit.target`, `audit.parent`, `audit.author`, `audit.recipient`, `audit.project`, `audit.instance`, `audit.setting`, `audit.job`, `audit.poll`, `audit.option`, `audit.task`, ...) keep call sites terse.
### Emission convention
Emission is **explicit per mutation** - `audit.record(...)` / `record_system(...)` calls placed after the mutation succeeds, or on the guard branch with `result="denied"`/`"failure"`. DRY choke points:
- Posts/projects/gists create/edit/delete record inside `content.py` (`create_content_item`/`edit_content_item`/`delete_content_item`, the latter two derive the key from `target_type`).
- Project file ops record via the `_audit_file`/`_edit`/`_fail` helpers in `project_files.py` (read-only guard -> `result="denied"`).
- Services via `_audit_service`.
- Containers via `audit_instance` (in `routers/projects/containers/_shared.py`) for the HTTP path, and in the Devii `actions/dispatcher.py` `_audit_mechanic` hook (after a successful `_run`) for the agent path - the two paths are disjoint (Devii calls `api.py` directly, never the routes), so there is no double counting. The dispatcher hook also emits the `devii.*` self-config mechanics (behavior/tools/tasks/lessons/customization) from one place.
The dispatcher's **authorization guard** (before `_run`) mirrors this with `_audit_denied`: a tool call refused for `requires_auth`/`requires_admin`/`requires_primary_admin` records a `security.authz.denied` event (`origin="devii"`, `via_agent=1`, `result="denied"`, `metadata.tool`/`reason`) so an agent-driven escalation attempt is visible in the admin Audit Log, not just the app log. This is what surfaces a non-admin's Devii probe of `admin_*`/`db_*` tools (the tool schemas are already withheld from non-admins by `tool_schemas_for`, so this fires only when the model invents a name it never received).
### Failures and denials are events
`auth.login.failure`, `auth.password.forgot_request` (unknown email), `security.rate_limit.block`, `security.maintenance.block`, `security.authz.denied` (emitted inside `require_user`/`require_admin` on the HTTP side, and inside the Devii dispatcher's `_audit_denied` for a refused agent tool call), self-role-change and self-disable denials, **admin-seniority denials** (a junior admin managing a senior admin), read-only write attempts, and `ai.quota.exceeded` all record with `result` set to `failure`/`denied`. Auth failures, authz denials (in `require_user`/`require_admin`), rate-limit/maintenance blocks, self-role-change/self-disable denials, and quota blocks are recorded with `result` set.
### Admin UI
`GET /admin/audit-log` (paginated, filterable list, `admin_section="audit-log"`) and `GET /admin/audit-log/{uid}` (detail + links) in the `routers/admin/` package, both `require_admin`, both negotiate via `respond(..., model=AuditLogOut|AuditEventOut)`. Templates `admin_audit_log.html` / `admin_audit_event.html` reuse `admin.css` + a small `audit.css`; `_pagination.html` gained an optional backward-compatible `pagination_query` prefix (default empty) so filters survive paging. Sidebar link sits last, before Settings.
### Devii access (read-only)
Admins query the same two routes conversationally via the admin-only Devii tools `audit_log` (GET `/admin/audit-log`) and `audit_event` (GET `/admin/audit-log/{uid}`) (`services/devii/actions/catalog.py`, `handler="http"`, `requires_admin=True`) - no new endpoint, since the `respond(...)` routes already serve JSON to Devii's `Accept: application/json` client (like `admin_list_users`). `audit_log` forwards every filter as a query param (`page`, `event_key`, `category`, `actor_role`, `actor_uid`, `origin`, `result`, `q`, `date_from`, `date_to`) and returns `AuditLogOut`, whose `options` object lists the valid values for each filter so the agent can discover them in one call. `audit_event` returns `AuditEventOut` (the row + related links). A system-prompt steer in `agent.py` (the AGGREGATES block) routes audit/history/"who did X" questions to these tools.
### Deferred persistence
`record`/`record_system` do all the cheap prep (actor resolve, sanitize, `json.dumps`, link assembly) on the request thread, then hand the two DB inserts to the **background task queue** (`services/background.py`, `background.submit(_persist, row, links)`) so the request returns without waiting on SQLite. The `uid`/`created_at` are generated eagerly at record time (so the return value and the audit timestamp reflect the action, not the flush). Under `DEVPLACE_DISABLE_SERVICES=1` (tests) the queue runs inline, so audit rows are visible immediately. See CLAUDE.md -> "Background task queue".
### Retention
`AuditService` (the `audit` service, daily) prunes rows older than `audit_log_retention_days` (default 90; `0` disables) via `store.sweep`. Configurable on the Services page like any other service.
### Adding an event
Pick/extend a key in `events.md`, add the `category_for` prefix if it is a new domain, then call `audit.record`/`record_system` at the mutation point with the right links and `result`. Never gate the audited action on the recording.
+32
View File
@@ -0,0 +1,32 @@
This file documents the Backup service (`devplacepy/services/backup/`, `devplacepy/services/jobs/backup_worker.py`, `devplacepy/routers/admin/backups.py`). Claude Code loads it automatically whenever a file under `devplacepy/services/backup/` is read or edited.
## Overview
Admin-only, enterprise-grade backups built on the **same async-job pattern as zip/fork** (see `devplacepy/services/jobs/CLAUDE.md`), so nothing runs on the request path. `BackupService(JobService)` (kind `backup`) is registered in `main.py` and overrides `run_once` to call `super().run_once()` (reap/refill/sweep) and then `_fire_due_schedules()`.
## Targets and worker
- **Targets** (`store.BACKUP_TARGETS`): `database` (consistent SQLite snapshot of the main DB + `devii_tasks.db` + `devii_lessons.db` via the `sqlite3` online backup API, so it is consistent under WAL - never a raw file copy), `uploads` (`UPLOADS_DIR`), `keys` (`KEYS_DIR`), `full` (database snapshot + uploads + keys; regenerable/volatile dirs - staging, locks, zips, container workspaces, chroma - are intentionally excluded). `service._materialize(target, staging)` returns a list of `{root, path}` sources; DB targets are snapshotted into `staging/db_snapshot` first.
- **Worker** (`services/jobs/backup_worker.py`, stdlib only): reads a JSON spec `{sources:[{root,path}]}` and an output path, builds a deterministic `tar.gz` (uid/gid 0, symlinks skipped) rooted at each `root/`, and returns `{bytes_in, bytes_out, file_count, dir_count, sha256}`. Invoked via `asyncio.create_subprocess_exec` like `zip_worker`.
## Storage and data model
- **Storage:** archives go under `config.BACKUPS_DIR` (`data/backups/`, in `DATA_PATHS`) sharded with `attachments._directory_for` on the **random uuid tail** (same load-bearing reason as zips/blobs), named `{target}-{YYYYMMDD-HHMMSS}-{tail}.tar.gz`. Staging is `config.BACKUP_STAGING_DIR` (`data/backup_staging/`), removed in `process` `finally`.
- **Data model** (`store.py`, ensured in `database.init_db` via `backup_store.ensure_tables()`): `backups` (NOT soft-deletable - an archive is a reclaimable operational artifact, hard-deleted like zips) and `backup_schedules` (in `SOFT_DELETE_TABLES`, born-live `deleted_at:None`). `store` holds all CRUD plus `compute_storage_stats()` (du of every major data area + `shutil.disk_usage`, run in `asyncio.to_thread` from the route, 30s in-process TTL cache so the walk never blocks).
- **Permanent artifact:** `cleanup(job)` only removes leftover staging, NEVER the archive. Job retention prunes the `jobs` row; the archive and `backups` row persist until an admin deletes it, a schedule rotates it out (`keep_last`), or `devplace backups clear`. Deleting a backup is a HARD delete (unlink file + delete row) - correct because backups are GC artifacts, the documented exception to the soft-delete rule.
## Schedules
`backup_schedules` carry `kind` (`interval`|`cron`), `every_seconds`/`cron`, `enabled`, `keep_last`, `next_run_at`, run bookkeeping. `_fire_due_schedules` (lock-owner only, so each fires once) compares `next_run_at <= to_iso(now_utc())` and enqueues a `backup` job + a `backups` record, then advances `next_run_at` via `schedule.next_run`. **Timestamp format is load-bearing:** schedule `next_run_at` uses the devii `schedule.to_iso` format (`%Y-%m-%dT%H:%M:%S`, no tz/micros) on BOTH sides of the comparison so lexicographic compare equals chronological - do not mix it with `datetime.isoformat()`.
## Routes and frontend
- **Routes** (`routers/admin/backups.py`, all `require_admin`, mounted via `admin/__init__`): `GET /admin/backups` (dashboard, `respond(..., model=BackupDashboardOut)`), `GET /admin/backups/data` (JSON dashboard for the monitor + Devii), `POST /admin/backups/run`, `GET /admin/backups/jobs/{uid}` (`BackupJobOut` for `JobPoller`), `POST /admin/backups/{uid}/delete`, `GET /admin/backups/{uid}/download` (`FileResponse`, path-guarded against `BACKUPS_DIR`), `GET /admin/backups/{uid}` (`BackupOut`), and schedule CRUD `schedules/create|{uid}/edit|toggle|run|delete`. **Route order:** every literal sub-path (`data`, `run`, `jobs/{uid}`, `schedules/...`) is declared BEFORE the `{uid}` catch-alls.
- **Frontend:** `static/js/BackupMonitor.js` (`window.BackupMonitor`, started inline from `admin_backups.html` like `AiUsageMonitor`) polls `/admin/backups/data` every 8s and renders storage/backup/schedule sections; run uses `Http.send` + `JobPoller`; delete/toggle/run/schedule actions use delegated clicks + `Http.send`; the schedule modal (create/edit) reuses the shared `.modal-overlay`/`data-modal` pattern. CSS `static/css/backups.css`. Sidebar link in `admin_base.html` (`admin_section == 'backups'`).
- **Devii** (all `requires_admin=True`, `handler="http"`): `backups_overview`, `backup_run`, `backup_status`, `backup_delete` (+confirm, in `CONFIRM_REQUIRED`), `backup_schedule_create`, `backup_schedule_delete` (+confirm). **CLI:** `devplace backups list|run <target>|prune|clear`. **Audit:** `job.backup.complete|failed` (category `backup`), `admin.backup.run|delete`, `admin.backup_schedule.create|update|toggle|delete` (category `admin`). **Docs:** admin API group endpoints + admin prose page `backups`.
## Download restricted to the primary administrator (load-bearing)
A backup archive contains the whole database, uploads, and VAPID keys, so `GET /admin/backups/{uid}/download` is restricted beyond `require_admin` to the **primary administrator** - the earliest-created user who currently holds the Admin role (the founder; `utils._create_account` auto-promotes the first registered user). The keystone is `database.get_primary_admin_uid()` (`SELECT uid FROM users WHERE role='Admin' ORDER BY created_at ASC, id ASC LIMIT 1`; users have no `deleted_at`, do NOT filter it) and `utils.is_primary_admin(user)` (admin AND `uid == get_primary_admin_uid()`, also a Jinja global). The download endpoint records `security.authz.denied` and raises `403` for any other admin. The `download_url` field is gated identically at every emission point (`_download_url`/`_backup_payload`/`_job_payload`, threaded `can_download = is_primary_admin(admin)`), so `GET /admin/backups/data`, `/jobs/{uid}`, and `/{uid}` never expose the URL to a non-primary admin; `BackupDashboardOut.can_download_backups` carries the flag. **Client:** `BackupMonitor.js` reads `data.can_download_backups`; a `done` backup renders an enabled `<a>` Download only for the primary admin, otherwise a **disabled** `<button>` with `title="Not available"`. The disabled state cannot be bypassed - the URL is never sent and the endpoint 403s regardless. If the founder is demoted/removed the crown passes to the next-oldest admin. Creating/running/deleting/scheduling backups stay open to all admins.
No restore into a live server (by design - overwriting the DB/uploads while running risks corruption). Restore is a manual ops procedure documented on the `backups` docs page. The live view relay's `admin.backups` topic publishes with `can_download=False` unconditionally (see `devplacepy/services/CLAUDE.md` -> Live view relay) - `BackupMonitor` derives real download capability only from its authoritative HTTP poll.
+60
View File
@@ -0,0 +1,60 @@
# Bot fleet (`devplacepy/services/bot/`)
This file documents the Playwright-driven AI persona fleet. Claude Code auto-loads it when a file under `devplacepy/services/bot/` is read or edited.
## Overview
`BotsService` (`services/bot/`) is a Playwright-driven fleet of AI personas that browse and interact with a DevPlace instance (posts, comments, votes, gists, projects, issues, follows, messages). It is the former standalone `dpbot.py` refactored into a package and wired into the service framework. Opt-in (`default_enabled = False`). Every other AI consumer's endpoint defaults reference "news, bots, Devii guests" collectively as internal-gateway consumers (see `GatewayService`).
## Module layout (one responsibility per file)
| File | Purpose |
|------|---------|
| `config.py` | Static constants + defaults (URLs, model, costs, personas, categories, paths) |
| `runtime.py` | `BotRuntimeConfig` dataclass - per-run settings threaded into a bot |
| `state.py` | `BotState` dataclass + JSON persistence (incl. the per-bot `identity` card) |
| `llm.py` | `LLMClient` (parameterized: key/url/model/costs); content generation + quality checks + the `decide`/`generate_identity` decision engine |
| `news_fetcher.py` | `NewsFetcher` - TTL-cached article source |
| `browser.py` | `BotBrowser` - Playwright wrapper (human-like typing/clicking/scrolling) |
| `registry.py` | `ArticleRegistry` - flock-based cross-bot article dedupe |
| `bot.py` | `DevPlaceBot` - the orchestrator (sessions, action cycle, `run_forever`) |
| `service.py` | `BotsService(BaseService)` - fleet manager + metrics |
## Mechanics
- `run_once` is a **reconcile tick**: it ensures `bot_fleet_size` bot tasks are alive (relaunching dead ones), scales the fleet up/down, and is a no-op heavy work otherwise. Each bot is an `asyncio` task running `DevPlaceBot.run_forever` in the lock worker's loop. `on_disable` cancels the fleet (each bot closes its browser on `CancelledError`).
- All "startup parameters" of the old script are `config_fields`: `bot_fleet_size` (was `--bots`), `bot_headless` (was `--headed`), `bot_max_actions` (was `--actions`), plus `bot_base_url`, `bot_api_url` (defaults to the internal gateway), `bot_news_api`, `bot_model` (defaults to `molodetz`), `bot_api_key` (secret; falls back to `internal_gateway_key()`), `bot_input_cost_per_1m`, `bot_output_cost_per_1m`.
- **A bot uses its own account's `api_key` for gateway calls once it has an account.** `bot_api_key` (-> `cfg.api_key`, the shared `internal_gateway_key()` fallback) is only the bootstrap credential used until the bot registers/logs in. After `_ensure_auth()` succeeds in `run_forever`, `DevPlaceBot._adopt_account_api_key()` fetches the bot's own `users.api_key` from its profile JSON (`fetch('/profile/{username}', {Accept: 'application/json'})` over the authenticated browser session, where `ProfileOut.api_key` is exposed to the owner) and switches `LLMClient.api_key` to it, persisting it in `BotState.account_api_key`. The key is read once and reused on every later session (the `LLMClient` is constructed with `state.account_api_key or cfg.api_key`, and `_raw_call` reads `self.api_key` per request so a mid-session swap takes effect on the next call). This routes each bot's gateway spend through its own user attribution (`resolve_owner()` -> `(owner_kind="user", owner_id=uid)`) instead of the shared internal key. Adoption is best-effort: a fetch failure leaves the bot on the fallback key.
- **Behavior and pacing tuning are admin `config_fields` too**, threaded through `BotRuntimeConfig` into the bot (never read module constants directly for these). The `config.py` constants (`MAX_BOTS_PER_ARTICLE`, `ARTICLE_TTL_DAYS`, `GIST_MIN_LINES`, `ACTION_PAUSE_MIN_SECONDS`/`ACTION_PAUSE_MAX_SECONDS`, `BREAK_SCALE_DEFAULT`) are only the **defaults**; the live values come from the settings: **Behavior** group - `bot_max_per_article` (-> `ArticleRegistry.max_per_article`), `bot_article_ttl_days` (-> `ArticleRegistry.ttl_days`), `bot_gist_min_lines` (-> `LLMClient.gist_min_lines`); **Pacing** group - `bot_action_pause_min_seconds`/`bot_action_pause_max_seconds` (the intra-session pause, clamped lo/hi in `DevPlaceBot.__init__` so an inverted pair is harmless) and `bot_break_scale` (a float multiplier on the between-session break, floored at 0.1 and the result floored at 5s); **Decisions** group - `bot_ai_decisions` (-> `BotRuntimeConfig.ai_decisions`, the AI-decision kill switch) and `bot_decision_temperature` (-> `BotRuntimeConfig.decision_temperature`). To add a new tuning knob: add the default to `config.py`, a `ConfigField` to `BotsService.config_fields` (its `group` creates/joins a UI section automatically), a field to `BotRuntimeConfig`, the `cfg[...]` mapping in `_launch_slot`, and consume it via the runtime config - do not reach for the module constant at the call site.
- **Live pricing** is published via `collect_metrics()`: fleet stat cards (bots running, fleet cost, LLM calls, tokens, posts/comments/votes) plus a per-bot table. `Fleet cost` is cumulative across restarts because each bot seeds its `LLMClient` totals from its saved `BotState`. **Cost is read from the gateway's authoritative per-call headers, not recomputed locally.** `LLMClient._account_usage(resp.headers, body_usage)` parses the `X-Gateway-*` response headers via `openai_gateway.usage.parse_usage_headers` (the shared parser, also used by `services/correction.py`) and takes `X-Gateway-Cost-USD` plus the prompt/completion token counts straight from the gateway, so the figure is cache-aware and native-cost-aware (it matches what `/admin/ai-usage` attributes to each bot's account). The `bot_input_cost_per_1m`/`bot_output_cost_per_1m` settings are **fallback-only**, applied to the response-body token counts solely when the endpoint returns no gateway headers (a non-gateway OpenAI URL); their defaults (`0.14`/`0.28`) mirror the gateway's chat pricing so even the fallback is sane. (Before this, the bots flatly multiplied every prompt token by an inflated `0.27`/`1.10` with no cache awareness, which massively overstated cost and the 24h projection.) `Cost rate` and `Projected 24h cost` come from a **sliding window**, not lifetime: `_sample_cost()` appends `(now, total_cost)` each tick into `_cost_samples`, evicts samples older than `config.COST_WINDOW_SECONDS` (600s), and measures the cost delta over the retained span; the projection is shown only once the window reaches `config.COST_WARMUP_SECONDS` (120s), otherwise both cards read "warming up". Samples reset when `_started_at` changes (`_cost_anchor`). Never project from lifetime `Fleet cost`/uptime, and never anchor the rate at startup - the boot burst (all bots ramping at once) then dominates the average and massively overstates a 24h projection extrapolated 274x from a few minutes.
- Heavy deps (`playwright`, `faker`) are **lazily imported** inside `run_once`/`_launch_slot` (install the `bots` extra). The service registers and shows in the admin tab even when they are absent; it logs and stays idle. Per-bot state lives in `~/.devplace_bots/state_slot{N}.json`; the article registry in `~/.dpbot_article_registry.json`.
- **Direct messages stay a dialog, never a monologue.** A bot may open a conversation, but `_compose_message` first calls `_spoke_last_in_thread()` and refuses to send when the last `.message-bubble` in `.messages-thread` is `.mine` (the bot itself sent the most recent message). So a bot only replies after the other person has spoken since its last message; it never sends two messages in a row. An empty thread (no bubbles) is allowed, so a bot can still initiate contact. This guards both reply paths (`_check_messages` and `_send_message`), mirroring how a real person waits for a reply before writing again.
## Live screenshot monitor (`services/bot/monitor.py`, `/admin/bots`)
The admin **Bot Monitor** (`/admin/bots`, sidebar link in `admin_base.html`, `require_admin`, `noindex`) shows the **latest low-quality screenshot per bot** with its username, persona, and current action/status, auto-refreshing every 2s. It is the constructive counterpart to the cost metrics: cost tells you what the fleet spent, the monitor tells you what each bot is looking at right now.
- **Capture is a `BotBrowser` concern** because the browser owns the Playwright `page`. `BotBrowser.bind_monitor(slot, username, persona)` (called from `DevPlaceBot.run_forever` once auth completes), `BotBrowser.note(action=, status=)`, and `BotBrowser.capture(reason, force=)` are the surface. `capture` takes a JPEG via `page.screenshot(type="jpeg", quality=MONITOR_JPEG_QUALITY=35, scale="css")`, throttled to one every `MONITOR_MIN_INTERVAL_SECONDS` (1s) unless `force=True`. **Capture is best-effort and never raises into the bot loop.** It fires at the meaningful-event choke points already centralized in `BotBrowser`: `goto` (page load, forced), `scroll`, `fill` (field input), `reload` (forced), plus `DevPlaceBot._action(tag, detail)` (every post/comment/vote/react/gist - an action-labelled forced capture via `asyncio.create_task`).
- **Storage is in-memory + one file per bot, never accumulating.** `monitor` (the `BotMonitor` singleton) holds a `slot -> BotFrame` dataclass map of tiny metadata, and writes the JPEG bytes to `BOT_DIR/monitor/slot{N}.jpg` (`config.DATA_DIR`, OUTSIDE the package, NOT under `/static`), **overwritten atomically in place** (`.tmp` then `replace`). One frame per slot, capped at the fleet size (`bot_fleet_size`, 0 to 20). `_stop_slot` calls `monitor.drop(slot)` to evict the row + unlink the file. No DB table, no soft-delete (nothing is persisted as a row), no shard tree (one file per bot is already bounded). This is what keeps it cheap at full fleet: at most 20 small JPEGs in RAM-backed files, refreshed in place.
- **Transport is polling, not WebSocket - deliberately.** Frames live only in the worker running the Bots service (the service lock owner), exactly like container logs/metrics. Rather than a lock-owner-gated WS with the 4013 retry, the monitor reuses the **existing `service_state` cross-worker bridge**: the lock owner publishes the frame metadata in `collect_metrics()["frames"]` (persisted to the `service_state` DB row by `BaseService._persist_state`), and ANY worker reads it back via `service_manager.get_service("bots").describe()["metrics"]["frames"]`. The page polls `GET /admin/bots/data` (`Poller`, 2s, `pauseHidden`) for that metadata and renders `<img>` tags at `GET /admin/bots/{slot}/frame.jpg` (raw bytes from the shared `BOT_DIR/monitor/` file, `Cache-Control: no-store`, served by any worker since it is a file under `DATA_DIR`). The `frame_url` carries a `?t=captured_at` cache-buster so a new frame is fetched only when it changed. This is the lightest correct option for the multi-worker model: no extra socket, no handshake, no 4013 bounce, and it reuses the same DB-backed metrics path the services page already polls. JSON shape is `schemas.AdminBotsOut`/`BotFrameOut`; Devii reads the same `/admin/bots/data` via the admin `bot_monitor` action; docs in `docs_api.py` (`admin-bots-*`, the frame endpoint in `NON_BODY_ENDPOINTS`).
- **Frontend** is `static/js/BotMonitor.js` (`app.botMonitor`, one class, `Http` + `Poller`) rendering the `static/css/bots.css` responsive grid (`bot-active`/`bot-idle` by frame age, design tokens only). Add a new capture point by calling `self.b.capture("reason")` (or `note(...)` then capture) at the new event; do not screenshot outside `BotBrowser` (it owns the page and the throttle).
## Realism: titles, topic spread, gist quality, inter-bot threads, thread-aware distinct opinions
Five content/behaviour mechanics keep the fleet from reading as machine-generated. Treat them as the realism contract; do not regress them.
- **Post titles are rewritten, never the raw headline.** A post body reacts to a news article, but the title is generated separately by `llm.generate_post_title(headline, persona, category)` - a short (3 to 8 word) human title in the persona's voice, not the verbatim article headline. The raw headline is still the registry dedupe key (`clean_title`), but the typed title is `type_title` (the rewritten one, `strip_md`-capped at 120). Same applies to projects: `generate_project_title`/`generate_project_desc` take the persona and run through `strip_label`. **`LLMClient.strip_label`** removes LLM label echoes (`Title:`, `Project Name:`, `Snippet name:`, and a trailing ` Concept:`/` Description:` clause) so a title can never be `Project Name: X Concept: Y`. Run every generated title through `strip_label`.
- **Personality drives topic and category, not just voice.** Category is chosen by `config.pick_category(persona)` (a persona-weighted `random.choices` over `PERSONA_CATEGORY_WEIGHTS`), replacing the old content classifier - a grumpy_senior rants, an enthusiastic_junior shows off and asks questions. Article selection is persona-scored: `config.persona_article_score(article, persona)` counts `SEARCH_TERMS` keyword hits, and both `_pick_article` (cache path) and `ArticleRegistry.reserve_unused` (direct path) rank articles by score plus a `random()` jitter, so different personas gravitate to different news rather than all reacting to the same headline.
- **Gists are quality-gated and non-trivial.** `generate_gist` now writes the code first (8 to 20 lines, persona-flavoured via `PERSONA_GIST_FLAVOR`, language drawn from `PERSONA_LANGUAGES`), then titles/describes it from the code. `_create_gist` runs a two-attempt loop through `llm.gist_quality_check(title, code, language)`: a `TRIVIAL_GIST_TERMS` blocklist on the title, a minimum non-empty line count (`LLMClient.gist_min_lines`, admin `bot_gist_min_lines`, default 6), and a strict LLM judge that rejects 101-tutorial snippets. Mirrors the post/comment `quality_check` loop; a reject regenerates once then skips.
- **Every comment is thread-aware and must hold a distinct opinion.** A bot never comments blind to what others already said. Before generating, `_comment_on_post` / `_reply_to_random_comment` call `DevPlaceBot._existing_comments()`, which scrapes the last `MAX_SIBLING_COMMENTS` (8) visible `.comment-text` bodies with their `/profile/` author (own comments excluded, each capped at `SIBLING_SNIPPET_LEN`=240). The block is threaded into `llm.generate_comment(..., existing_comments=)`, whose system prompt then forces a point none of them made - a different angle, a missed caveat, or respectful disagreement with one of them by name - and forbids repeating any opinion, framing, example, or question already raised. The mechanical backstop is in `llm.quality_check(..., siblings=)`: a candidate whose `_overlap_ratio` against the joined sibling text exceeds `SIBLING_OVERLAP_THRESHOLD` (0.55) is rejected as `echoes another comment` and regenerated (same two-attempt loop as the source-restatement gate at `RESTATEMENT_OVERLAP_THRESHOLD`=0.6). This is what stops a salient post from drawing a dozen near-identical "tipping point / feedback loop / cooldown" replies. Any new comment-generation path MUST gather siblings and pass them to both `generate_comment` and `quality_check`.
- **Usernames read like nerd handles, not real names.** Bots do NOT sign up with faker person names; they pick devRant/Hacker-News-style handles. Two layers, LLM-first with an offline fallback: (1) `llm.generate_handle_candidates(persona)` asks the model for ~8 distinct handles grounded in the persona and its `SEARCH_TERMS` interests (tech nouns, leetspeak, adjective+noun, creatures, short word+number), each run through `handles.sanitize_handle`; the pool is cached on `BotState.handle_candidates` and consumed by `_next_handle()`. (2) `handles.make_handle(interests)` is the pure, offline algorithmic generator (curated word banks + probabilistic leetspeak + number/separator/casing decoration, persona-seeded from `SEARCH_TERMS`), used as the fallback whenever the LLM pool is empty or a signup collides. Both layers emit only `[A-Za-z0-9_-]`, 3 to 20 chars (the old `first.last` styles silently failed signup validation, which rejects dots). On a third signup retry a random number is appended to bust collisions. Word banks and tuning constants live in `services/bot/handles.py`; never reintroduce faker person-name handles.
- **Bots engage each other, and threads deepen.** `ArticleRegistry` allows up to `max_per_article` holders per article (admin `bot_max_per_article`, default 2), each with a **distinct category** (stored as `holders: [{bot, category, time}]`; old `{bot, time}` rows are normalized on load), so two bots can post different angles on the same trending news - which gives them each other's posts to discuss. A holder ages out after `ttl_days` (admin `bot_article_ttl_days`, default 7). `_engage_community()` (called once per normal/deep session after notifications) goes to `/feed?tab=recent|trending`, ranks non-own posts with a preference for `known_users` authors, opens one, comments, and replies into the comment thread. The on-post reply-to-comment probability is also raised so multi-bot threads actually form. `reserve(title, owner, category)` is idempotent per owner and enforces the same distinct-angle / max-holders rules as `reserve_unused`.
## AI-driven decisions and identity cards (`bot_ai_decisions`)
Opt-in mechanic (off by default; full design and cost model in `aibots.md`). When `bot_ai_decisions` is on, a bot replaces the procedural action cascade with one LLM decision call per page, driven by a unique AI-generated identity. **Do not regress the grounding and the kill-switch fallback.**
- **Two LLM entrypoints, both on `LLMClient`.** `generate_identity(archetype)` runs once per bot (seeded from the persona archetype) and returns the identity card: `name`, `backstory`, `interests`, `dislikes`, `temperament`, and four 0.0-1.0 scalars (`verbosity`, `contrarianness`, `generosity`, `curiosity`) plus `rhythm`. `_normalize_identity` clamps every scalar and defaults `interests` to the archetype's `SEARCH_TERMS`, so a malformed model reply can never produce a broken card. The card persists on `BotState.identity` (serialised to `state_slot{N}.json`); old state files load with an empty dict and regenerate on first AI session. `decide(identity, page_state, menu, history, temperature)` returns `{plan, energy, stop_after}` (strict JSON), the per-page ordered action plan.
- **JSON is parsed raw, never `clean()`-ed.** `_call` runs the output through `clean()` (which strips `*`/`_` and collapses whitespace, corrupting JSON keys like `stop_after`); the decision engine instead uses `_raw_call` (the extracted HTTP+cost-accounting core, no `clean`) and `_parse_json` (tolerates code fences and prose wrappers by slicing the first `{`..`}`). When adding any JSON-returning LLM method, use `_raw_call`, not `_call`.
- **The harness supplies the options; the model only picks.** `DevPlaceBot._build_menu(page_state)` builds the action menu from the same observable-page predicates the random `_cycle` uses, listing only actions that are physically possible and policy-allowed (throttles like `can_post`/`can_project`/own-post-no-mention are enforced by *omitting* the action). `decide` filters the returned plan against that menu (`_filter_plan` drops any unknown action), re-asks once on malformed/empty output, and the executor `_run_plan` falls back to a single `navigate` to feed when the plan is still empty - **never a `random` draw**. `_dispatch_action` maps each action string to the existing per-action method (`_comment_on_post`, `_vote_for_page`, `_react_to_object`, `_navigate_to`, ...); content actions still run their `quality_check`/`gist_quality_check` gates unchanged, so quality cannot degrade.
- **`_cycle_ai` mirrors `_cycle`'s signature and return tuple** so `run_forever` swaps between them with `(self._cycle_ai if use_ai else self._cycle)(...)`. `energy` and `stop_after` replace `_session_type` mood and `_fatigue_check`: the session ends when `session_actions >= stop_after`, the intra-action pause is `_scaled_pause()` (scaled by `energy`), and the between-session break is scaled by `energy` too. Timing jitter (`_read_like_human`, `_type_like_human`, `_idle`) stays procedural - the model decides *what*, not the millisecond cadence. `use_ai` is `self._ai_decisions and bool(self.state.identity)`; identity generation and the flag are gated so a disabled fleet pays nothing and behaves exactly as before.
+200
View File
@@ -0,0 +1,200 @@
This file documents the Container manager subsystem (devplacepy/services/containers/, plus its routers/projects/containers and routers/admin/containers.py HTTP surface, and the shared ppy Docker image). Claude Code loads it automatically whenever a file under devplacepy/services/containers/ is read or edited.
## Overview
Admin-only. Supervises container instances, all running one shared prebuilt image. Driven through the `docker` CLI via `asyncio.create_subprocess_exec` behind a pluggable **`Backend` ABC**, never a docker SDK. Reference: admin docs `Services -> Container Manager` and the `containers` API group.
## Backend seam
`backend/base.py` defines the `Backend` ABC. `DockerCliBackend` drives `docker` via `asyncio.create_subprocess_exec`, streaming run logs line-by-line (`_stream`). `FakeBackend` is the in-memory test double (its `image_exists` returns `True`). `runtime.get_backend()`/`set_backend(b)` selects the active backend - tests call `set_backend(FakeBackend())`. A future Kubernetes/remote backend implements the same ABC. The container OS workspace mount point is the single constant `backend/base.py` `WORKSPACE_MOUNT` (`/app`).
## Security (load-bearing)
Mounting the docker socket is root-equivalent on the host. Every run/exec/lifecycle/schedule operation is gated `require_admin` (HTTP) and `requires_admin=True` (Devii). `--privileged` is never passed. User input is never shell-interpolated - every call is an arg-list subprocess, never `shell=True`. Resource limits are applied via `--cpus`/`--memory`.
On top of the admin gate, **per-user container isolation** applies (see root CLAUDE.md "Project visibility (is_private) and read-only" convention): only the **primary administrator** (earliest-created Admin) and the instance owner (creator or project owner) may manage an instance; other admins get read-only access, and only when the instance's project is public. Enforced via `content.py` predicates `owns_instance`, `can_view_project_containers`, `can_view_instance`, `can_manage_instance` (all reusing `is_primary_admin`): the primary administrator sees and manages every container including those on private projects; any other admin can VIEW containers of others only when the instance's project is public (private-project containers of others are invisible; of private containers they see only their own), and can MANAGE (edit/lifecycle/exec/terminal/sync/delete/schedules) only instances they own - non-owners get 403 plus an audit row with `result="denied"` (WS close `1008`). This is enforced at every entrypoint: `routers/projects/containers/` (`project_for` + `manage_guard` + the exec WS), the admin Containers manager (`_decorate` filtering + per-row `can_manage`, `_viewable_instance_or_404`, `_manage_denied`, create + project-search), the Devii container tools (`ContainerController._project` + `_require_manage`), and the live view relay (broadcast topics never carry private-project instances).
## One shared image, no in-app builds (load-bearing)
Every instance runs `config.CONTAINER_IMAGE` (default `ppy:latest`, override `DEVPLACE_CONTAINER_IMAGE`), built ONCE from `ppy.Dockerfile` via `make ppy` (context `devplacepy/services/containers/files`). There are no per-project Dockerfiles, builds, `ContainerBuildService`, or build UI - all removed.
`service._launch` calls `api.run_spec_for(inst, config.CONTAINER_IMAGE)`. `api.create_instance(project, *, name, ...)` fails fast with `ContainerError` if `backend.image_exists(CONTAINER_IMAGE)` is `false` ("run 'make ppy'"); it no longer takes a dockerfile/build. `backend.image_exists(ref)` (`docker image inspect`) is part of the `Backend` ABC.
Migrating off the old per-project-build model: `devplace containers prune-builds` removes the legacy per-project images (`backend.remove_image`) and clears the `dockerfiles`/`dockerfile_versions`/`builds` tables (one-time); existing instances auto-repoint to `ppy` (their stored `build_uid` is ignored).
**Default idle:** `run_spec_for` runs `["sleep", "infinity"]` when an instance has no `boot_command`, so a bare instance stays up instead of exiting (the image `CMD` is the same, but the explicit command makes it independent of the image).
## Data model
`store.py`, indexed in `init_db`:
- `instances`
- `instance_events` (audit trail; also home of the status-history rows written by `_set_status`)
- `instance_metrics` (ring buffer, `METRICS_RING=720`; sweep keeps it bounded)
- `instance_schedules` (cron/interval/once)
## Reconciler (`ContainerService`, `service.py`)
A `BaseService` reconciler; the model is desired-vs-actual, NOT a task per container. Each tick:
1. `backend.ps(label=devplace.instance)` snapshots `docker ps` (label `devplace.instance=<uid>` is the join key).
2. Converge each instance to its `desired_state`.
3. Apply restart policies.
4. Reap orphan containers (labeled but no DB row -> no orphans, no lost state).
5. Fire due `instance_schedules` (reusing `devii/tasks/schedule.py` `cron_next`/`next_run`).
6. Sample `docker stats`.
Only the service lock owner reconciles; HTTP handlers only flip `desired_state` / write events. `docker run --name <slug>` collision is the double-launch guard.
## Workspace materialization
Workspace = `/app` (the `WORKSPACE_MOUNT` constant). `project_files.export_to_dir` materializes the project to `config.CONTAINER_WORKSPACES_DIR/<project>` (`DATA_DIR/container_workspaces`, default `data/`) once, bind-mounted RW; `project_files.import_from_dir` (the inverse) syncs it back on the sync action. Both run via `asyncio.to_thread`.
**Runtime data must live OUTSIDE the `devplacepy/` package and NOT under `/static`**: workspaces (and zips) live in `config.DATA_DIR` (configurable via `DEVPLACE_DATA_DIR`). Never put generated data under `STATIC_DIR` - it is served publicly AND watched by dev `--reload` (scoped to `--reload-dir devplacepy`). The docker daemon must be able to bind-mount `DATA_DIR` for `/app`.
Every exec passes `-w /app` explicitly (`DockerCliBackend.exec` and the PTY exec in `routers/projects/containers/instances.py`), so one-shot `container_exec`, the interactive shell, and tmux sessions all start in the project workspace. The agent is told this (system prompt + `container_exec` tool summary) so it never prefixes a command with `cd /app`.
## Shared operations
`api.py` holds the operations shared by `routers/projects/containers/instances.py` (and the rest of the `routers/projects/containers/` subpackage) and the Devii `ContainerController` (`services/devii/container/`, `handler="container"`, 7 instance tools). The reconciler service defaults **disabled** (needs docker); enable it on `/admin/services`.
## HTTP routing surface
- `/projects/{slug}/containers` - the `routers/projects/containers/` subpackage: `instances.py` (creation/lifecycle/exec/logs/metrics/sync plus the exec websocket), `schedules.py` (cron/interval/once schedules), shared helpers in `_shared.py`. Every instance runs the shared `ppy` image. Discoverable from the project detail page's admin-only **Containers** button (gated by `content.can_view_project_containers` via the `viewer_can_containers` context flag) and from the admin index.
- `/admin/containers` - `routers/admin/containers.py`: lists every instance across all projects with inline actions (start/stop/restart/terminal/edit/delete) plus a create modal (project search-select, run-as user search-select, boot language + source editor, restart policy, start-on-boot, env/ports/limits/ingress); `/admin/containers/{uid}` is the per-instance detail page (lifecycle, live logs/metrics, PTY terminal, schedules, ingress, sync, status history) and `/admin/containers/{uid}/edit` edits run-as user, boot language/script/command, restart policy, start-on-boot, and limits.
**Key reuse rule:** the admin Containers section adds NO lifecycle/logs/exec endpoints of its own - the instance carries its `project_uid`, so the admin detail route resolves the project and its frontend targets the existing `/projects/{slug}/containers/instances/{uid}/...` routes. Add any new instance operation to `routers/projects/containers/instances.py` only; the admin detail page picks it up for free.
## UI - two surfaces over one API set
Both are discoverable (an earlier version of the per-project page had zero links to it - fixed).
1. **Per-project manager**: `templates/containers.html` + `static/js/ContainerManager.js` handles instance creation through an app modal form (the `_macros.html` `modal()` macro + `ModalManager` `.visible` toggle); its instance list links out to the shared detail page. Reached from the project detail page's Containers button, and passes breadcrumbs so content clears the fixed nav.
2. **Admin Containers section**: `routers/admin/containers.py` (mounted `/admin/containers`, sidebar link in `admin_base.html`, `admin_section="containers"`). `GET /admin/containers` lists every instance via `store.all_instances()` (decorated with project title/slug from one `projects` lookup) in an `.admin-table`. `GET /admin/containers/data` is the poll JSON. `GET /admin/containers/{uid}` renders `templates/containers_instance.html` + `static/js/ContainerInstance.js` - a dedicated detail page (lifecycle, poll logs/metrics, schedules add/delete, ingress, sync, interactive exec over a PTY WebSocket gated on the lock owner).
All container CSS (`static/css/containers.css`) uses app design tokens (`--bg-card`, `--text-primary`, `--success`/`--danger`/`--warning`, `--radius`) and the shared `.card` recipe.
## Admin CRUD manager (`/admin/containers`)
`/admin/containers` is a full manager, not a read-only list. `routers/admin/containers.py` adds (all `require_admin`, all reuse `api.*` directly - no docker/exec backend duplicated):
- `POST /create` (`ContainerAdminCreateForm` -> `api.create_instance`, project search-select + run-as user + boot fields + restart policy + start_on_boot + env/ports/limits/ingress).
- `GET`/`POST /{uid}/edit` (`ContainerEditForm` -> `api.update_instance_config`, edits `run_as_uid`/`boot_language`/`boot_script`/`boot_command`/`restart_policy`/`start_on_boot`/`cpu_limit`/`mem_limit`).
- Stacked literal lifecycle routes `POST /{uid}/{start,stop,restart,pause,resume}` (the no-wildcard rule below).
- `POST /{uid}/sync` (bidirectional).
- `POST /{uid}/delete` (soft, owner-or-admin via the admin guard).
- Two JSON search endpoints `GET /{projects,users}/search` (back the create/edit search-selects via `db.query LIKE`).
**Route order:** `/{uid}/edit` and the literal verb routes are declared BEFORE the `/{uid}` catch-all.
Frontend: `static/js/ContainerList.js` (poll-render rows with inline action buttons + the create modal with two debounced search-select widgets and a boot-language toggle) and `static/js/ContainerEdit.js` (the edit page). Both route through `Http.send`/`Http.getJson` and surface errors via `app.toast`; the Terminal button reuses `app.containerTerminals.open(slug, uid, name)`. The instance detail page (`containers_instance.html`) has a config summary (boot language, run-as uid, start-on-boot) plus a **Status history** card rendering the server-side `instance_events`.
Devii: `container_create_instance` takes the boot/run-as/start_on_boot args; `container_configure_instance` (`controller._configure_instance` -> `api.update_instance_config`) edits the same fields; both admin-only, audited via `_DEVII_CONTAINER_EVENTS`.
## No wildcard verb route (load-bearing)
`routers/projects/containers/instances.py` registers every instance verb as a **literal** route - `start`/`stop`/`pause`/`resume`/`restart` (stacked `@router.post` decorators on one `instance_action` handler that reads the verb from `request.url.path`), plus `delete`/`exec`/`sync`/`schedules`. There is deliberately NO `POST /instances/{uid}/{action}` catch-all: a wildcard segment matches its literal siblings too (Starlette matches first-declared), so a catch-all would silently swallow `exec`/`sync`/`delete`/`schedules` as `{"error": "unknown action: exec"}` whenever it sat above them. Add any new instance verb as a literal route (and to the `instance_action` decorator stack if it just flips desired state) - never reintroduce a `{action}` wildcard.
## Host port allocation
Host ports are auto-allocated and globally unique. `parse_ports` accepts a bare container port (`8899`, host side `0` = auto) or a pinned `host:container`. `api.assign_host_ports` (called in `create_instance`) resolves every `host=0` to the lowest free port in `[HOST_PORT_MIN=20001, 65535]` that is neither in `used_host_ports()` (the union of every instance's published host ports from `ports_json`) nor currently bound on the host (`_host_port_free` test-binds `0.0.0.0:port`). A pinned host port already published by another instance is rejected up front. This guarantees no two instances ever collide on a host port - previously a silent `docker run` "port is already allocated" failure. The resolved host port is what gets stored in `ports_json` and what the ingress proxy reads.
## Floating terminals (interactive shell)
An instance's shell is NOT inline - it is a floating `<container-terminal>` (`static/js/components/ContainerTerminal.js`) over the exec WS (`/projects/{slug}/containers/instances/{uid}/exec/ws`, `pty.openpty()` + `docker exec -it`, admin + `service_manager.owns_lock()` gated). xterm.js + the fit addon are vendored under `static/vendor/xterm/` and lazy-loaded. Opened by the instance-page **Open terminal** button (`app.containerTerminals.open(slug, uid, name)`, `ContainerTerminalManager`) or by telling Devii to "attach &lt;container&gt;" - the admin-only `open_terminal` **client** action (`client_actions.py`, `requires_admin=True`) -> `DeviiClient._openTerminal` -> the same manager.
**`FloatingWindow` base.** It extends the generic `FloatingWindow` base (`static/js/components/FloatingWindow.js` + `static/css/floating-window.css`), which owns the Devii-style chrome: drag, native `resize:both`, maximize/fullscreen, geometry persistence. `app.windows` (`WindowManager`, `components/WindowManager.js`) assigns z-index on `pointerdown` so the last-clicked window is on top (below the avatar/toasts).
**The Devii terminal also extends `FloatingWindow`.** `devii/devii-terminal.js` reuses the base drag/geometry/preset/resize/persist plus `_window`/`_registerWindow` (z-order + context-menu) machinery, overriding only the parts that differ - `_setState` (Devii adds a `closed` state + launcher FAB), `_defaultGeometry` (bottom-right, 760x600), `_changeFont` (per-user `--devii-font-size` CSS var), `_contextItems`, and `_persistGeometry` (delegates to its richer `{state,geometry,focused}` blob under `devii-terminal-state`). Devii keeps native `resize:both` (no `.fw-resize` grip); `devii.css` is unchanged - the only base change required was a `get _dragClass()` getter (default `fw-dragging`) that Devii overrides to `devii-dragging`, so the base drag handlers toggle Devii's own class and its existing CSS drives everything.
**Resize protocol.** The exec WS treats a brace-leading JSON frame `{"type":"resize","cols","rows"}` as a control message: it `TIOCSWINSZ`-es the host pty (`fcntl.ioctl`) **and** signals the `docker exec` client (`proc.send_signal(SIGWINCH)`) so the new size forwards through to the container's exec tty (vim/top reflow). The SIGWINCH is required: with `start_new_session=True` the host slave is not the client's controlling terminal, so the kernel does not auto-deliver it; the pair shares its winsize, so the client reads the new size off its slave stdio and calls Docker's exec-resize API. Any non-resize frame is raw stdin (backward compatible). The fit addon drives it client-side, and the window has a dedicated bottom-right resize grip (`.fw-resize`, since xterm covers the native CSS resize corner - the base `FloatingWindow` adds a manual grip instead of relying on `resize:both`). xterm renders ANSI natively, so the old `cleanTerm` stripping is gone for the interactive shell (the one-shot exec box still strips output, since `docker exec` without `-t` is non-TTY).
**Persistence (tmux).** A `?session=<name>` query param (validated `^[A-Za-z0-9_-]{1,64}$`) makes the WS run the shell inside tmux (`tmux new-session -A -s <name>`, falling back to bash if tmux is absent), so the session survives WS disconnect/window close - `proc.kill` on disconnect kills the tmux *client*, the server daemon + session live on. The frontend keeps exactly **one persistent slot** at a time (`ContainerTerminalManager`, the last-opened terminal): it attaches to session `devplace`, **auto-reconnects** with capped backoff on unexpected WS close (not on code `1008`), and is remembered per-user in `localStorage` (`ct-persistent:<scope>`); `app.containerTerminals.restore()` (called once in `Application.js` after `window.app`) reopens it on the next page load. Opening another terminal demotes/calls `setPersistent(false)` on the previous one (its tmux session lingers in the container, but the frontend stops auto-reconnecting/remembering it); explicit window-close (`_onClose`) detaches + forgets (session still survives, reopen re-attaches).
**Right-click context menu.** Every window wires `app.contextMenu.attach(this.win, () => this._contextItems())` (right-click + long-press), opening the shared `dp-context-menu`. The base `FloatingWindow._contextItems` is Minimize/Normalize/Close; `ContainerTerminal` overrides it with Copy (xterm `getSelection`, disabled when `!term.hasSelection()`) / Paste (clipboard -> WS stdin) / Sync files / Restart / Terminate (`dp-dialog` confirm, then delete + close) / Minimize / Normalize / Project (-> `/projects/{slug}`) / Files (-> `/projects/{slug}/files`) / Close - the lifecycle items POST to the existing `instances/{uid}/{sync,restart,delete}` routes. Devii's `_contextItems` override returns Copy (live `window.getSelection()`, or the current input line when empty) / Paste (clipboard inserted into `.devii-input` at the caret, `selectionStart/End` splice, then refocus) / Minimize / Normalize / Close (it is not bound to a container or project). Copy/Paste both use the async Clipboard API with a hidden-`textarea` + `execCommand("copy")` write fallback.
**Minimize/Normalize.** Geometry presets (`_presetGeometry(w,h)`, smallest-usable / comfortable), exposed both as titlebar buttons (`data-win="minimize|normalize"`) and menu items on every window.
## Pravda image (load-bearing, workspace ownership)
The `ppy` image (`ppy.Dockerfile`, built by `make ppy`, context `devplacepy/services/containers/files`) is a `python:3.13-slim-bookworm` base with Playwright plus a broad set of common Python libraries preinstalled, plus CLI tools (`tmux`, `apache2-utils` for `ab`, `procps`/`htop`/`iftop`/`iotop`, `netcat-openbsd` for `nc`, `zip`/`unzip`, `fakeroot`, git/curl/wget/vim/ack).
The security hotpatch that used to run per build is now baked into `ppy.Dockerfile` once. The final stage:
- `COPY`s the **sudo superclone** (`files/sudo`) over `/usr/local/bin/sudo` (+ symlink `/usr/bin/sudo`; the real `sudo` package is not installed).
- `COPY`s the **`aptroot` fakeroot wrapper** (`files/aptroot`, symlinked over `apt`/`apt-get`/`dpkg` in `/usr/local/bin` so pravda installs system packages without root).
- `COPY`s **`pagent`** (`files/pagent`, the stdlib AI agent; reads `DEVPLACE_OPENAI_URL`+`DEVPLACE_API_KEY`, falling back to its public endpoint + `DEEPSEEK_API_KEY`) to `/usr/bin/pagent.py`, plus `files/.vimrc` to `/home/pravda/.vimrc` (whose AI helper - `AiEditSelection` - targets the same gateway as pagent via `DEVPLACE_OPENAI_URL`/`DEVPLACE_API_KEY`, with a public fallback, never `api.openai.com`).
- Evicts any pre-existing uid-1000 user, creates user **`pravda` at `1000:1000`**.
- Hands pravda ownership of the toolchain AND the OS package trees (`chown -R pravda` over `/usr/local/lib`, `/usr/local/bin`, `/usr/lib/python3`, `/opt`, `/app`, `/home/pravda`, plus `/usr/lib`, `/usr/bin`, `/usr/sbin`, `/usr/share`, `/usr/include`, `/etc`, `/var/lib`, `/var/cache`, `/var/log`, `/srv` so `apt`/`dpkg` can write; `~/.local/bin` on `PATH`).
- Ends on `USER pravda`.
**Why uid 1000.** `/app` is bind-mounted from the host, and under DooD a container's UID maps 1:1 to the host. A container running as **root** wrote root-owned files into the workspace; the app process (host `retoor`, uid 1000) then hit `[Errno 13] Permission denied` in `export_to_dir` on the next instance create - a permanent per-project brick. Pinning the container UID to `1000` makes every write land as `retoor`, and pravda owning the global site-packages means runtime `pip install` succeeds as pravda with no elevation.
**The sudo superclone (`files/sudo`, POSIX `sh`)** is sudo-CLI-compatible but **never swaps user**: it parses the full flag interface (`-u/-g/-E/-H/-n/-S/-i/-s/--`, leading `VAR=value` env assignments, `-V/-l/-v/-k/-K` short-circuits), exports `SUDO_USER`/`SUDO_UID`/`SUDO_GID`/`SUDO_COMMAND`, then `exec`s the command **as the current user (pravda)**. So even a blind `sudo apt ...` / `sudo -u root ...` runs as uid 1000 and **cannot create a root-owned file** - the residual risk is gone by construction.
**Rootless apt (`files/aptroot`).** Pravda installs system packages directly (`apt install <pkg>`, no sudo). `aptroot` is symlinked over `apt`/`apt-get`/`dpkg` in `/usr/local/bin` (ahead of `/usr/bin` on `PATH`) and execs the real tool under **`fakeroot`** with `APT::Sandbox::User=root`; dpkg's chown-to-root calls are virtualized while every file actually written lands owned by **pravda (uid 1000)**. Combined with pravda owning the system trees (`/usr`, `/etc`, `/var/lib`, `/var/cache`, `/var/log`, `/srv`, ...), `apt install <pkg>` works with no real euid 0 and still cannot brick the bind mount. Run `apt update` first (the image clears `/var/lib/apt/lists`); a package whose maintainer script needs genuinely privileged syscalls may still fail - bake those into `ppy.Dockerfile` (its `RUN apt-get` runs as root before `USER pravda`), then `make ppy`.
**Trade-off (intentional).** The only genuinely-root operation that still does NOT escalate is binding a port < 1024 - use a high port + `/p/<slug>` ingress instead. Enforcement lives entirely in the Dockerfile (no `--user` on `docker run`). The `export_to_dir` unlink-before-write fix remains as belt-and-suspenders (the app owns the workspace dir, so it may delete any stale file in it regardless of owner before rewriting it).
## `PRAVDA_*` runtime env injection
`api.run_spec_for` merges `api.pravda_env(instance)` over the instance's own `env_json` (PRAVDA keys win), so every running container gets these platform vars:
- `DEVPLACE_BASE_URL` - the `site_url` setting via `seo.public_base_url()`.
- `DEVPLACE_OPENAI_URL` - `{base}/openai/v1`, the OpenAI-compatible gateway base.
- `DEVPLACE_API_KEY` - the **creating** user's (or `run_as_uid`'s) `api_key`, resolved live.
- `DEVPLACE_USER_UID` - the project **owner**'s uid.
- `DEVPLACE_CONTAINER_NAME` - the instance name.
- `DEVPLACE_CONTAINER_UID` - the instance uid.
- `DEVPLACE_INGRESS_URL` - the instance's absolute public ingress URL `{base}/p/{ingress_slug}` when published and `site_url` is set, relative `/p/{slug}` if `site_url` is unset, empty if the instance has no `ingress_slug`.
These let code inside a container call back into the platform and the AI gateway authenticated as the user. **Injected at run** (`docker run -e`), never baked into the image: the image is shared by every instance, but name/uid/api-key are per-instance and base-url/key must stay fresh, so they are resolved each launch and never stored as secrets in the DB.
Resolution needs two instance columns set at `create_instance`: `created_by` (= `actor[1]` when the actor is a user) feeds `DEVPLACE_API_KEY`, and `owner_uid` (= `project["user_uid"]`) feeds `DEVPLACE_USER_UID`. Instances created before this change have neither and degrade to an empty key/uid (base-url and container name/uid still resolve) until recreated. `DEVPLACE_BASE_URL`/`DEVPLACE_OPENAI_URL` are empty when `site_url` is unset, so set the admin `site_url` for container-to-platform calls (and ingress URLs) to work.
## Ingress (`/p/<slug>`)
`routers/proxy.py` (registered at `/p`) implements the ingress reverse proxy. An instance opts in with `ingress_slug` + `ingress_port` (must be one of its mapped container ports; slug unique, validated in `api.validate_ingress`). `_resolve(slug)` -> `store.find_instance_by_ingress` -> `api.proxy_target(instance)` returns `(gateway, host_port)` (the instance's recorded docker bridge gateway + published host port); the route reverse-proxies HTTP (httpx) and WebSocket (`websockets` client) to `http://{host}:{port}/{path}`, stripping the `/p/<slug>` prefix. The reconciler records `container_ip`/`container_gateway` from `docker inspect` each running tick.
**`proxy_target` dials the container's docker bridge gateway + the published host port** (`gateway:host_port`), NOT loopback and NOT the container's own bridge IP. This is deliberate: `127.0.0.1` fails where docker's `nat OUTPUT` excludes `127.0.0.0/8` from DNAT and no `docker-proxy` binds loopback for a `0.0.0.0`-published port; the container's own bridge IP can be dropped by docker's bridge-isolation rule (`DOCKER ! -i docker0 -o docker0 -j DROP`); the gateway+published-port path survives both. `DEVPLACE_CONTAINER_PROXY_HOST` overrides the host (e.g. `host.docker.internal` for a containerized app) and still uses the published host port; before a gateway is recorded the proxy falls back to `127.0.0.1`.
Ingress is **public** (no auth) but SSRF-safe: host/port are derived from the instance row, never from user input. The Devii container tools return an absolute `ingress_url` built from `seo.public_base_url()` (the admin `site_url` setting), so the agent fetches the public production URL, not localhost (which web tools refuse); unset `site_url` yields a relative `/p/<slug>`. nginx needs a `/p/` location with WS upgrade + long timeouts (present in `nginx/nginx.conf.template`).
## Production wiring
The container manager drives the host docker daemon, which needs heavy wiring - all carried by the opt-in `docker-compose.containers.yml` overlay (docker socket mount, `INSTALL_DOCKER_CLI=true` build arg, `group_add` the docker gid). **`make docker-build`/`make docker-up` always apply this overlay** and self-derive its inputs: `DOCKER_GID` via `stat -c '%g' /var/run/docker.sock` (the socket's owning group) and `DEVPLACE_DATA_DIR` = `$(CURDIR)/data` (the project's own dir at its real host absolute path). This makes the manager work out of the box with no `sudo`, no `/srv`, no manual `.env` edits. A plain `docker compose up -d` drops the overlay (no CLI, no socket) and silently re-breaks it - always update through the make targets.
**The DooD bind-mount gotcha:** `docker run -v <path>:/app` resolves `<path>` on the HOST, so `DEVPLACE_DATA_DIR` must be mounted at an identical host+container path (the make targets use `$(CURDIR)/data` on both sides; manual `docker compose` users get a `/srv/devplace-data` default). Build contexts ship via the docker API tarball, so the container temp dir is fine. Set `DEVPLACE_CONTAINER_PROXY_HOST=host.docker.internal` only when a containerized app cannot route to the recorded gateway. See README "Container Manager wiring".
## Bidirectional newer-wins sync (load-bearing direction rule)
Sync is NOT one-directional import. `project_files.sync_dir_bidirectional(project_uid, workspace, user) -> {"exported", "imported"}` is the one helper that reconciles a project's virtual FS against an instance's `workspace_dir`: per file, the side with the newer timestamp wins (project `updated_at` epoch, via `datetime.fromisoformat(...).timestamp()`, vs filesystem `st_mtime`, with a 1s skew tolerance favouring export on ties), and a file present on only one side propagates to the other. It **NEVER deletes a file** - only creates/overwrites the older side. A **read-only** project (`is_readonly`) exports only, never imports (the read-only guard direction).
Both `api.sync_workspace` (HTTP/Devii sync action, returns `{exported, imported}`) and `api.sync_bidirectional_sync` (reconciler, non-blocking `record_event` system actor, logs only when non-zero) call the same helper. The reconciler runs it before every `_launch` AND on a ~60s wall-clock cadence over running instances (`SYNC_EVERY_SECONDS`, gated on `time.monotonic()` independent of the 5s reconcile tick). The per-instance boot-helper files (`.devplace_boot.py`/`.devplace_boot.sh`) are in `SYNC_SKIP_NAMES` so they never round-trip into the project.
## Run-as user = identity + API key ONLY (load-bearing constraint)
An instance's `run_as_uid` column selects WHICH DevPlace user's identity and `api_key` are injected (`DEVPLACE_API_KEY`, `DEVPLACE_USER_UID`), resolved in `api.pravda_env` ahead of the `created_by`/`owner_uid` fallback chain. It does **NOT** change the container OS user, which is ALWAYS `pravda` (uid 1000) - required for the bind-mounted `/app` (DooD uid maps 1:1 to host). Validate it against an existing user via `api.validate_run_as`.
## Boot source precedence (load-bearing)
Columns `boot_language` (`none`|`python`|`bash`) + `boot_script` (multiline source) sit alongside the legacy `boot_command`. Precedence in `api.run_spec_for`: `boot_script` (by language) > `boot_command` > image CMD (`sleep infinity`). When a boot script is set, the reconciler writes it into the workspace (`api.materialize_boot_script`, `.devplace_boot.py`/`.devplace_boot.sh`, excluded from sync) before launch and runs `python|bash /app/.devplace_boot.<ext>`. `api.validate_boot` enforces the language set and a 100k char cap.
## start_on_boot is per-container
The `start_on_boot` integer flag (0/1) forces `desired_state=running` ONLY for flagged instances when `ContainerService` first runs (`_boot_pass`, one-time); all others keep their last `desired_state`.
## Status-change choke point
Every reconciler status mutation flows through `service._set_status(inst, changes, reason=)`, which diffs old vs new `status` and, on change, writes an `instance_events` `status_change` row AND an audit `container.instance.status` `record_system` event (`old_value`/`new_value`). `_reconcile`, `_launch`, and `_handle_exit` (terminal + policy_restart) all flow through it, so previously-silent transitions (`running->stopped`/`->paused`/exit) are now logged and visible on the admin detail page's **Status history** card. New code that changes an instance's status in the reconciler must use `_set_status`, never a raw `store.update_instance(... status ...)`.
## New columns are ensured in init_db and defaulted in store.create_instance
`run_as_uid`/`boot_language`/`boot_script`/`start_on_boot` are added to the `instances` ensure-block in `init_db` (filtered/queried columns must exist on the cached schema) and seeded in the `store.create_instance` base dict so pre-existing rows degrade gracefully.
## Testing without Docker
Use `FakeBackend` (its `image_exists` returns `True`) + `runtime.set_backend`, and monkeypatch `config.CONTAINER_WORKSPACES_DIR` to a tmp dir. See `tests/unit/services/containers.py` and `tests/api/containers.py` (argv, instance creation on `config.CONTAINER_IMAGE`, image-not-built guard, reconcile matrices, schedule firing, ingress validation + live HTTP proxy, HTTP admin gate).
## Vibe coding on-ramp (user-facing doc)
The container runtime is also the basis of "vibe coding": the public prose page `templates/docs/getting-started-vibing.html` (slug `getting-started-vibing`, `SECTION_GENERAL`, not admin-gated so the docs stay public while the feature is admin-only/Alpha) is the canonical user-facing guide. It teaches the Devii-driven flow project -> container -> start -> terminal (`create_project`, `container_create_instance`, `container_instance_action`, the `open_terminal` client action), documents the three baked-in agents (`dpc` = DevPlace Code at `/usr/bin/dpc`, the Claude-Code-class coding agent; `botje.py` = the copy of `services/containers/files/bot.py` at `/usr/bin/botje.py`; `pagent`), all metered through the container's own `DEVPLACE_API_KEY`, the full `PRAVDA_*` env table (see `api.pravda_env`), and ingress at `/p/<slug>` via `ingress_slug`/`ingress_port`.
When the runtime, the agent binaries, or the `PRAVDA_*`/ingress contract change, update this page alongside the source.
**The agents are gateway-only:** `dpc`/`d.py` and `botje.py`/`bot.py` use a single `molodetz` backend pointed at `DEVPLACE_OPENAI_URL` (the gateway); the former direct `api.deepseek.com` fallback backend was removed so every in-container AI call is ledgered under the run-as user and nothing bypasses `gateway_usage_ledger`. `pagent`/`.vimrc` already posted to the gateway URL (using `DEEPSEEK_API_KEY` only as a key fallback, never the DeepSeek endpoint). Rebuild the image (`make ppy`) for the change to reach running containers.
+41
View File
@@ -0,0 +1,41 @@
# Database API service (`devplacepy/services/dbapi/`, `devplacepy/routers/dbapi/`)
This file documents the primary-administrator-only, read-only generic database API. Claude Code auto-loads it when a file under `devplacepy/services/dbapi/` is read or edited.
## Overview
`dbapi/` (routers package, mounted at `/dbapi`) is a **primary-administrator-only, strictly READ-ONLY** generic database API over `dataset` (never inserts/updates/deletes/restores; the write surface was removed because it bypassed every per-route admin safeguard). `tables.py` (`GET /tables`, `GET /{table}/schema`), `crud.py` (read only: `GET /{table}` + `GET /{table}/{key}/{value}`, `?include_deleted`), `query.py` (`POST /query` hard SELECT-only validated read-only run; `POST /query/async` + `GET /query/{uid}` + `WS /query/{uid}/ws` via `DbApiJobService` kind `dbquery`), `nl.py` (`POST /nl` natural-language->SQL via the AI gateway, returns validated SELECT, `execute=true` runs it read-only). **Auth = the PRIMARY administrator only** (the earliest-created Admin, `utils.is_primary_admin` / `database.get_primary_admin_uid` - the same identity that gates backup downloads) by session/api_key (`services/dbapi/policy.py`); every other administrator is refused like a member, and there is **no internal-key path** - the gateway `internal_gateway_key()` is NOT accepted (no service-to-service access). SQL validated by `services/dbapi/validate.py` (sqlglot classify + suspicious-flag + read-only `EXPLAIN` dry-run). Devii tools `db_*` are read-only (list/get/query/nl; no write tools) and flagged `requires_primary_admin=True`, so they are added to the LLM tool list ONLY for the primary administrator (every other session never sees them and Devii is unaware they exist). `services/pubsub/policy.py` resolves its own admin/internal actor and does NOT reuse `dbapi.policy.caller_for`, so the primary-admin restriction does not leak onto the pub/sub bus.
It exposes per-table reads, a validated raw `query()`, a natural-language-to-SQL designer backed by the internal AI gateway, and async query execution streamed over a websocket. **It can never insert, update, replace, delete, or restore data in any way** - there are no write endpoints and no write tools (this is a hard rule; removed because the generic write surface bypassed every per-route admin safeguard - role-change seniority guards, container privilege locks, the audit trail). Reuses the existing dataset helpers, the `JobService`/`ProgressHub`/lock-owner websocket pattern, the AI gateway, and the Devii action catalog - it does not re-implement any of them.
## Auth (single boundary, primary-admin ONLY)
`services/dbapi/policy.py` `caller_for(request)` returns a `Caller` ONLY for the **primary administrator** (a session OR api_key whose user passes `utils.is_primary_admin` - the earliest-created Admin, the same identity that gates backup downloads, resolved by `database.get_primary_admin_uid`), else `None`. **There is no internal-key path** - the gateway `internal_gateway_key()` is NOT accepted (removed at the user's request); there is no service-to-service access. Every other administrator is refused exactly like a member, so the `Caller.kind` is always `"admin"`. `require_caller` raises `DbApiDenied`; the router's `require_dbapi_caller` turns that into a `403` and an audited `database.access.denied`. Members, guests, non-primary admins, and internal callers never reach a handler. **Note:** `services/pubsub/policy.py` no longer reuses this `caller_for` (it would have leaked the primary-admin restriction into the pub/sub bus); it resolves its own admin/internal actor so any admin stays `privileged` on the bus.
## Table guard
`policy.assert_table(name)` enforces a name regex, existence in `db.tables`, and a deny-list (`DEFAULT_DENY = {sessions, password_resets, cache_state}` plus the `dbapi_deny_tables` setting). Used on every table-scoped path and indirectly by the validator's table extraction, so neither a crafted segment nor an NL-designed query can touch a denied table.
## Reads only (`services/dbapi/crud.py`, `routers/dbapi/crud.py`)
`GET /dbapi/{table}` (keyset pagination newest-first, `?filter.<col>=`, `?gte/lte/gt/lt.<col>=`, `?search=`, `?before=`, `?limit=`, `?include_deleted=`) and `GET /dbapi/{table}/{key}/{value}`. There are **no insert/update/delete/restore routes** - the router exposes only the two read handlers, and `services/dbapi/crud.py` keeps only read functions (`schema`, `list_rows`, `count_rows`, `get_row`). The only audit events the API emits are `database.access.denied`, `database.query`, and `database.nl.design`.
## `query()` is hard SELECT-only (`services/dbapi/validate.py`)
`classify()` parses with `sqlglot` (statement type, table list, `has_where/has_join/has_limit`, suspicious notes for missing `WHERE/JOIN/LIMIT`, multiple statements, `ATTACH/PRAGMA/VACUUM`). `validate_select()` rejects anything but a single SELECT, then `dry_run()` does `EXPLAIN` on a **separate read-only** connection (`file:...?mode=ro` + `PRAGMA query_only=ON`) for true validity. `run_select()` executes read-only. `POST /dbapi/query` returns `409` for non-SELECT (the database API is read-only; data cannot be changed through it), `400` for invalid SQL, else rows + a `suspicious` list. The database API can never mutate data through any path.
## NL-to-SQL (`services/dbapi/nl2sql.py`)
`POST /dbapi/nl` builds a system prompt from the table schema + 5 example rows + a soft-delete rule (auto `deleted_at IS NULL` unless `apply_soft_delete=false`), calls the internal gateway (`INTERNAL_GATEWAY_URL`, model `dbapi_nl_model` or `molodetz`, auth = the primary-admin caller's api_key for attribution), and **reprompts with the validator's error until the SQL validates** (max 3 attempts). Returns the validated SQL by default; `execute=true` also runs it read-only and returns rows.
## Async (`services/dbapi/service.py` `DbApiJobService`, kind `dbquery`)
`POST /dbapi/query/async` validates then `queue.enqueue`s; the service re-validates (defense in depth), streams rows in batches to its own `ProgressHub`, writes the full result to `config.DBAPI_DIR/{uid}/result.json`, and returns stats. `GET /dbapi/query/{uid}` (status), `GET /dbapi/query/{uid}/result` (rows from disk, extends retention), `WS /dbapi/query/{uid}/ws` (lock-owner gated, `4013` retry, snapshot-then-stream - the SEO pattern). Config on `/admin/services`: `dbapi_max_rows`, `dbapi_nl_model`, `dbapi_nl_system_preamble`, `dbapi_deny_tables`, plus the standard Jobs fields.
## Devii (primary-admin gated)
**Read-only** tools `db_list_tables`, `db_table_schema`, `db_list_rows`, `db_get_row`, `db_query` (SELECT-only; surface `suspicious`), and `db_design_query` (NL->SQL), all flagged `requires_primary_admin=True` (alongside `requires_admin=True`) in `actions/catalog.py`. The `Action.requires_primary_admin` flag is filtered by `Catalog.tool_schemas_for(authenticated, is_admin, is_primary_admin)` and enforced again in `Dispatcher.dispatch`, so these tools are **added to the LLM tool list only for the primary administrator** - every other administrator's (and member's/guest's) Devii never receives the schemas and is unaware the database API exists. `is_primary_admin` is threaded WS/Telegram/CLI -> `hub.get_or_create(is_primary_admin=)` -> `DeviiSession` -> `Dispatcher`, mirroring how `is_admin` flows (`routers/devii.py` `_resolve_ws_owner`, `services/telegram/bridge.py`, `services/devii/cli.py` which runs both flags `True` as the trusted local operator, and the headless scheduler in `service.py`). There are **no write tools** - the former `db_insert_row`/`db_update_row`/`db_delete_row` were removed along with the HTTP write routes, so Devii cannot change data through the database API.
## nginx
`/dbapi/query/<uid>/ws` has its own upgrade location.
+153
View File
@@ -0,0 +1,153 @@
This file documents the Devii assistant subsystem (devplacepy/services/devii/ and all its subdirectories: virtual_tools, behavior, email, telegram, container, client, customization, tasks, rsearch, agentic, actions). Claude Code loads it automatically whenever any file under devplacepy/services/devii/ is read or edited.
## Overview and account scope
`DeviiService` is the in-platform agentic assistant (a relocated standalone agent), exposed as a WebSocket terminal at `/devii/ws` (router `routers/devii.py`, prefix `/devii`) and reachable from the **Devii** user-dropdown item (`data-devii-open`) and auto-opened on `/docs` (`data-devii-autoopen`). Opt-in (`default_enabled=False`). Frontend is `static/js/DeviiTerminal.js` (an `Application.js`-instantiated controller, `app.devii`) plus the `devii-terminal`/`devii-avatar` custom elements and ported renderers under `static/js/devii/`, the vendored md-clippy avatar under `static/vendor/md-clippy/`, and `static/css/devii.css`.
A signed-in user's Devii drives **their own** account: the agent's `PlatformClient` uses this instance as base URL and the user's `api_key` as Bearer auth, so admins get admin-level tools. The **LLM** calls use the same `api_key` too: the hub threads `owner_kind` into `_build_settings` -> `build_settings(cfg, base_url, api_key, owner_kind)`, which sets `Settings.ai_key = api_key` for `owner_kind == "user"` (else the configured/internal gateway key). So a user's Devii spend is attributed to them at the gateway (`user:<uid>`) and counts toward any per-user gateway limit, while guests fall back to the internal key. Guests get an unauthenticated sandbox client. Every turn is audited in `devii_turns`.
## Sessions, hub, and channels
`DeviiHub` keeps one `DeviiSession` per owner-and-**channel** `(owner_kind, owner_id, channel)` - `user`/uid or `guest`/`devii_guest` cookie, channel `main` (default, the floating terminal) or `docs` (the in-page docs chat). A session holds a `set` of WebSockets and **broadcasts** every frame to all of that owner-channel's tabs/devices. After a turn it persists `agent._messages` to `devii_conversations` (users only; rehydrated on reconnect -> survives restarts), appends a `devii_usage_ledger` row, and a `devii_turns` audit row. On `attach` it sends a `{"type":"history"}` snapshot so a new tab renders the prior conversation. Tasks persist in `devii_tasks` (owner-scoped; guests use an in-memory db via `tasks.store.memory_db()`).
**Channels are an independent conversation thread only.** A `channel` separates the floating terminal (`main`) from the in-page docs chat (`docs`) so each has its own conversation thread for the same owner. The WS reads `?channel=` (`routers/devii.py`, allowlisted to `{"main","docs"}`, falls back to `main`) and threads it `get_or_create(..., channel=)` -> `DeviiSession(channel=)`. **Only** `devii_conversations` is keyed `(owner_kind, owner_id, channel)` (column added + NULL->`main` backfilled + index extended in `init_db`); lessons, behavior, virtual tools, tasks, and the `devii_usage_ledger`/`devii_turns` 24h quota stay owner-scoped (deliberately shared across channels). `find(...)`/`get_or_create(...)` default `channel="main"` so existing callers (e.g. `/devii/adopt`) are unchanged. The `docs` channel is driven only by `<dp-docs-chat>` (`static/js/components/AppDocsChat.js`): a light-DOM component that calls `/devii/session` (guest cookie), opens `DeviiSocket(handlers, "docs")`, renders user/agent bubbles in docs style (`static/css/docs-chat.css`) via the devii markdown renderer, stubs any `avatar`/`client` request so the agent never hangs, and seeds its first question from the `seed` attribute (the sidebar search box, intercepted). It is mounted on `/docs/search.html` (`templates/docs/_chat.html`); the docs pages no longer carry `data-devii-autoopen`. The docs page no longer auto-opens the floating terminal.
## The Docii docs channel
The `docs` channel is branded **Docii**, a documentation-only assistant, built by specialising the same `DeviiSession` on `self.channel == "docs"` (no separate service):
- **Prompt.** `session.py` sets `self._system_prompt = DOCS_SYSTEM_PROMPT` (instead of `_system_prompt_for`) - it forces the agent to call `search_docs` first for every question, to recursively refine-and-search (reading the returned section `content`, then searching again with new keywords) until grounded, to answer ONLY from retrieved docs, and to do no platform actions. `_compose_system_prompt()` returns it verbatim for docs (the owner `BEHAVIOR_HEADER` is NOT appended, so a user's general Devii behavior rules can't derail it).
- **Tools.** `_builtin_tools()` filters `CATALOG.tool_schemas_for(...)` to the `DOCS_TOOLS` allowlist (`{"search_docs"}`); `_refresh_tools()` for docs sets `self.tools[:] = _builtin_tools()` with NO virtual tools. So the docs agent has exactly one tool and cannot wander.
- **Greeting.** `bootstrap_greeting()` returns `DOCS_GREETING` for docs.
- **search_docs.** `services/devii/docs/controller.py` delegates to the in-process page index `docs_search.search_pages(...)` (no HTTP), returning per result `title`, `url` (`/docs/{slug}.html`), `score` (BM25), and `content` (page text truncated to `CONTENT_CHARS`), admin-filtered by the dispatcher's `is_admin`. So "visiting the search results" is repeated `search_docs` calls; the default 40 `max_tool_iterations` leaves ample room to recurse.
- **References and inline linking.** Because results carry real `url`s and `score`s, the prompt REQUIRES the answer to (a) link inline to documented pages with markdown links to their `url` and (b) end with a `## References` list of the pages used as clickable links with their score. The devii markdown renderer makes `/docs/...` links clickable client-side.
- **Topic gate (follow-up vs new).** `_run_turn` runs `_docs_topic_gate(text)` before the main loop (docs channel only). When there is prior history it calls `_classify_topic` (a no-tools `LLMClient.complete_text` with `TOPIC_CLASSIFIER_PROMPT`, parsed by `_parse_topic`, defaulting to `follow_up` on any failure). On `new` it resets `agent._messages` to `[system]`, clears the persisted docs conversation, and emits `{type:"topic", decision, reason}`; the frontend clears the log, re-shows the current question, and prints a reasoning note. On `follow_up` it keeps context and shows the note.
- **Isolation from Devii (critical).** The docs channel is built with an EPHEMERAL `owned_db` (`hub.get_or_create`: `db if owner_kind=="user" and channel=="main" else memory_db()`), so its task/lesson/behavior/virtual-tool stores are private and empty - a Docii reflection never pollutes the user's Devii memory and vice-versa. The **Scheduler only starts in the `main` channel** (`ensure_scheduler_started`: `if not self._started and self.channel == "main"`), so Devii's scheduled tasks never fire inside a docs session. (Regression fixed: previously the docs session's scheduler ran owner tasks over the shared `devii_tasks` table, and the search-only Docii agent tried to fulfil a `run_js` task by spamming `search_docs` - tasks are a `main`-only feature.)
- The `main` channel is untouched (full 90-tool catalog + behavior header + Devii greeting).
## Docs search mode and BM25 fallback
The docs search surface is admin-configurable via the `docs_search_mode` site setting (`agent` | `bm25`, default `agent`, on `/admin/settings`, seeded in `init_db` `operational_defaults`, field in `AdminSettingsForm`). `routers/docs/views.py` `docs_page` (slug `search`) resolves the effective mode: `_agent_search_state(request, user, is_admin)` returns `disabled` (Devii off, or guests-disabled for a guest), `quota` (viewer over their `daily_limit_for`/`spent_24h`), or `ok`. Agent renders **only** when `mode=="agent"` and state `ok`; otherwise it runs the classic `docs_search.search(...)` (BM25, `docs_search.py`) and renders `templates/docs/_search.html`. `docs_base.html` branches the `kind=="search"` include on `search_mode`; on a quota fallback (`quota_fallback`) the BM25 page shows a "reached your daily AI assistant limit" hint. So `docs_search.py`/`_search.html` are live again (the bm25 path), not orphaned.
## Conversation, lesson, and task persistence
Conversation history persists to `devii_conversations` (rehydrated on reconnect, surviving restarts); per-turn cost goes to `devii_usage_ledger` (the authoritative 24h-spend source - the in-memory `CostTracker` is display-only) and an audit row to `devii_turns`; tasks live in `devii_tasks` (guests use an in-memory store).
**Per-owner self-learning memory is privacy-critical.** The `LessonStore` (reflect/recall) is **owner-scoped, never shared**: `LessonStore(db, owner_kind, owner_id)` filters every read/write by owner. The hub builds one per session over the same `owned_db` as the task store - the main `db` (table `devii_lessons`) for signed-in users (persistent, isolated, survives restarts) and a fresh `memory_db()` for guests (ephemeral, scoped to that web session). `forget_lessons` (agentic tool -> `LessonStore.clear()`/`delete()`) lets the user purge them; the system prompt forbids storing credentials/secrets in lessons. **Regression to avoid: do NOT share one `LessonStore` across sessions** - that leaked one user's reflected lessons (including credentials) into every other user's recall.
## Reminders and scheduled tasks (persistent across reboot)
`create_task` queues a self-contained prompt that a fresh agent runs later (`services/devii/tasks/`): `kind=once` (`delay_seconds` for relative, `run_at` UTC for absolute), `interval` (`every_seconds`), or `cron`. The per-session `Scheduler` ticks every 1s and executes due rows through the session's executor, so the result is broadcast to any connected tab and buffered (`type:"task"` frame) when the terminal is closed.
**The scheduler is no longer tied to a live WebSocket.** It is started by `session.ensure_scheduler_started()`, called both on `attach` AND by the lock-owner `DeviiService._ensure_task_schedulers()` each `run_once` (and promptly at boot): that pass scans the shared `devii_tasks` for every user owner with an enabled `pending`/`running` task (`tasks.store.pending_owner_ids(db)`), resolves the user (api_key/username/is_admin/timezone), and `hub.get_or_create(...)`s a headless session whose scheduler then runs the task - so a queued reminder survives a server reboot and fires even if the user never reopens Devii. `hub.gc_idle()` skips a session with `has_pending_tasks()` so a task-only session is not reaped between ticks.
A task created as a reminder carries `notify=1`: when it finishes, `session._deliver_reminder` sends a `utils.create_notification(owner, "reminder", result, target_url="/devii")` (the `reminder` notification type), so the in-app notification + live toast reach the user regardless of the terminal.
**Timezone awareness.** The terminal sends a `clientinfo` frame (`Intl...timeZone` + UTC offset) on connect; `session.set_clientinfo` stores it and persists `users.timezone` (`database.set_user_timezone`). `session._compose_system_prompt()` injects a `# CURRENT TIME` block (UTC now + the user's local time/timezone, falling back to the stored `users.timezone` for a headless run) so the agent converts wall-clock requests to a correct UTC `run_at`; relative requests use `delay_seconds`. The `REMINDERS AND SCHEDULED TASKS` system-prompt rule makes the agent SCHEDULE rather than fire a delayed instruction immediately, and set `notify=true` for reminders.
## Cost, quota, and the financial-data-is-admin-only rule
**Financial cap.** The rolling-24h `$` limit is read from `devii_usage_ledger` (`UsageLedger.spent_24h`), **not** the in-memory `CostTracker` (which resets on reboot and is display-only). Checked before each turn in the router; a turn already over is blocked (the over-limit WS error states only "Daily AI quota reached (100%)", never a dollar figure). The caps are `devii_user_daily_usd`, `devii_guest_daily_usd`, and `devii_admin_daily_usd` (default `0` = unlimited, admins exempt by default), resolved via `config.effective_daily_limit`. `config.effective_daily_limit(cfg, owner_kind, is_admin)` is the single source of truth, used by both `service.daily_limit_for` and `build_settings`: guests use `devii_guest_daily_usd` (default $0.05), users `devii_user_daily_usd` ($1.00), and **administrators are exempt by default** via `devii_admin_daily_usd` (default `0.0`, where **`0` = unlimited**). All three are editable on the Devii service config (`/admin/services`). The WS gate is `if limit > 0 and spent >= limit` so a `0` cap never blocks.
**Every billed turn is ledgered, not just interactive ones:** `session._record_spend` writes the usage delta in the `finally` of both `_run_turn` AND the scheduler executor, and it runs even when a turn is cancelled/superseded, so scheduled-task spend and interrupted-turn spend count against the cap rather than being forgiven. Pricing is admin-configurable - `cost/tracker.py` resolves a `Pricing` object per `CostTracker` rather than import-time env constants.
**Quotas are resettable** (clearing the owner's `devii_usage_ledger` rows): per-user via `POST /admin/users/{uid}/reset-ai-quota` (button on the admin Users page), globally via `POST /admin/ai-quota/reset-guests` and `/reset-all` (buttons on `/admin/ai-usage`), and from the CLI via `devplace devii reset-quota <username> | --guests | --all`.
**Financial data is admin-only:** any monetary figure (USD cost, pricing, spend, limit) is restricted to administrators; members and guests see only the **percentage** of their 24h quota used. The owner's admin status is resolved server-side (`is_admin(user)`) and threaded WS -> `hub.get_or_create(is_admin=)` -> `DeviiSession(is_admin=)` -> `Dispatcher(is_admin=)`. The `Action` dataclass has a `requires_admin` flag (`cost_stats` is admin-only USD; the member-safe `usage_quota` tool returns **only** `{used_pct, turns_today, limit_reached}` with no money and is available to everyone). Gating is double: `Catalog.tool_schemas_for(authenticated, is_admin)` never hands an admin-only tool's schema to a non-admin, and the dispatcher independently raises `AuthRequiredError` for any `requires_admin` action a non-admin attempts. `usage_quota`'s data comes from `DeviiSession._quota_snapshot` (ledger `spent_24h`/`turns_24h` over `settings.daily_limit_usd`); for non-admin owners the system prompt also appends a hard rule forbidding any cost disclosure (defense in depth). `GET /devii/usage` mirrors this: it always returns `used_pct`/`turns_today` and adds `spent_24h`/`limit` only for admins. The standalone `devii` CLI runs with `is_admin=True` (the local operator owns the process). The catalog's cost/analytics HTTP tools `ai_usage` (GET `/admin/ai-usage/data`, USD breakdown) and `site_analytics` (GET `/admin/analytics`) are also `requires_admin=True`, so a non-admin session is never offered them and the dispatcher blocks them even if named - the platform 403 is no longer the only guard. The profile page is correctly split too (`_ai_quota(include_cost=viewer_is_admin)` emits dollars only for admins, in both its HTML and JSON forms, so a member fetching their own profile with `Accept: application/json` never sees dollars either).
## Aggregate analytics (no pagination)
Site analytics is a **documented, admin-secured HTTP endpoint** - `GET /admin/analytics` (`routers/admin/` package, returns JSON, `is_admin` check -> 403 otherwise) backed by `database.get_platform_analytics()` (single UNION query over `posts/comments/gists/projects.created_at`, TTL-cached 60s; total members, active users 24h/7d/30d, signed-in-now, signups, content totals, top authors). Devii calls it as a normal `http` catalog action `site_analytics` (GET `/admin/analytics`, `top_n` query param) - the platform enforces admin, so it behaves identically for the web and the `devii` CLI and never touches the DB directly from the agent. It is documented in the Admin API docs group (`docs_api.py`, id `admin-analytics`). The system prompt tells the model to use it for any "how many"/"how active" question instead of paging `admin_list_users` (which caused a pagination storm), and to `search_docs` before guessing a route rather than probing URLs.
## Context-window sizing (DeepSeek V4 Flash, 1M tokens)
The upstream model `deepseek-v4-flash` (what `deepseek-chat` routes to; default `gateway_model`) has a 1,048,576-token context and 384,000-token max output. All Devii size limits are **characters**, derived in `services/devii/config.py` from one source of truth: `CONTEXT_INPUT_BUDGET_TOKENS = CONTEXT_WINDOW_TOKENS - MAX_OUTPUT_TOKENS - SYSTEM_RESERVE_TOKENS` (600,576 tokens), and `DEFAULT_CONTEXT_COMPACT_THRESHOLD = CONTEXT_INPUT_BUDGET_TOKENS * CHARS_PER_TOKEN` (3 chars/token = ~1.8M chars). At the conservative 3-chars/token estimate the worst-case input at compaction plus the full max output plus the system reserve sums to exactly the 1M window, so it can never overflow. `DEFAULT_MAX_RESPONSE_CHARS` (200,000) is the master chunk knob: every non-`chunks` tool result passes through `wrap_if_large(result, max_response_chars)` in `actions/dispatcher.py`, and `read_more` slices are clamped to it, so a file up to ~200KB returns in **one** read instead of paging in 12KB slices (the old read_more storm). `OUTPUT_CAP_CHARS` (`agentic/loop.py`, 400,000) must stay **above** `max_response_chars` or it would re-truncate a full chunk envelope. `ChunkStore` caches up to `STORE_MAX_CHARS` (8MB) per entry, `STORE_MAX_ENTRIES` (16) entries; `SUMMARY_INPUT_CAP` (`agentic/compaction.py`, 600,000) bounds what the summarizer ingests when compacting. When changing the upstream model, retune from `CONTEXT_WINDOW_TOKENS`/`MAX_OUTPUT_TOKENS` - everything else derives.
## Multi-worker service-lock routing
Hubs are per-process, so `/devii/ws` is served **only** by the worker that holds the background-service lock (`service_manager.owns_lock()`, set in `main.py` startup); a non-owner worker `close(4013)` (an application code in the private 4000-4999 range, reliably delivered as the browser `CloseEvent.code`). `DeviiSocket.js` recognises 4013 as "wrong worker, not yet settled" and fast-retries in ~200ms **silently** (no "disconnected, reconnecting..." notice), so with 2 prod workers the client converges on the owner in a fraction of a second instead of bouncing every 1500ms. The visible "connected." notice is emitted on the first server frame (`onReady`), never on raw socket open, so an accepted-then-4013-closed handshake on the wrong worker is invisible. The disabled/no-service path keeps the standard `1013` (a real, user-visible disconnect). The cap stays correct regardless because it reads the DB.
## Client-side browser-automation tool channel
Beyond the avatar, Devii has a `client`/browser channel using the same request/response pattern: `services/devii/client/ClientController` is bound on `attach`, and `actions/client_actions.py` exposes `get_page_context`, `run_js`, `highlight_element`, `clear_highlights`, `show_toast`, `scroll_to_element`, `navigate_to`, `reload_page` (`handler="client"`, all `requires_auth=False`). The session's generic `_browser_request(channel, action, args)` powers both `avatar` and `client`; the router resolves both `avatar_result` and `client_result`. Frontend `static/js/devii/DeviiClient.js` executes the actions in the user's own browser and the terminal routes `type:"client"` frames to it, returning `client_result`. This powers live on-screen tutorials, page-context awareness, and navigation/refresh. `navigate_to`/`reload_page` reply *before* unloading so the turn completes, then the persistent session reconnects.
**Target selection (visibility-aware).** Unlike replies/traces, which `_emit` *broadcasts* to every tab, a browser request is sent to one **target** chosen by `_pick_target()`. The terminal reports each connection's `document` visibility and focus via `{"type":"visibility"}` (on connect, `focus`, `blur`, and `visibilitychange`); the session stores it in `_conn_meta` and ranks connections `(focused, visible, attach_seq)`, so the command runs on the tab the user is actually looking at, falling back to the most recently attached when none reports focus. **Regression to avoid: do NOT route to a single "last-attached primary" socket** - with the `/docs` auto-open tab or any second tab, `reload_page`/`navigate_to` ran on the wrong (hidden) tab while still reporting success, so "the reload did not happen" from the user's view. `run_js` is gated by the `devii_allow_eval` config field (default on), checked in `ClientController` before any round-trip. Overlays (`.devii-hl-box`, `.devii-hl-callout`, `.devii-toast`) attach to `document.body`, not the terminal.
## Shared browser session (auth adoption)
The terminal and the browser are one session. `/devii/ws` resolves its owner from the browser's `session` cookie (`_resolve_ws_owner`), so the terminal is whoever the browser is logged in as. Agent-initiated auth propagates: the session watches its own trace stream (`_trace`) and, on a successful `login`/`signup`, captures the real session token the `PlatformClient` minted against this instance (`session_cookie()`) and broadcasts `{"type":"auth","action":"adopt"}`; the browser navigates to single-use `GET /devii/adopt` which sets the httpOnly `session` cookie and redirects (reload). On `logout` it broadcasts `{"type":"auth","action":"logout"}` and the browser goes to `/auth/logout`. Browser->terminal: login/logout are navigations that reconnect the WS; a cross-tab change is caught by a focus check against `/devii/session` that reloads. The httpOnly cookie is only ever set/cleared by HTTP endpoints, never JS.
## Screen awareness
The `SYSTEM_PROMPT` instructs Devii, after a mutation, to `get_page_context` and `reload_page` when the page the user is viewing displays the data just changed, so settings/list/detail views the user is watching update live without a manual refresh.
## Project filesystem tools: line-range editing and the overwrite-protection guard
Devii's project filesystem tools include surgical **line-range editing** (`project_read_lines`, `project_replace_lines`, `project_insert_lines`, `project_delete_lines`, `project_append_file`) so it edits large text files without resending them, and `Dispatcher._run_http` enforces an **overwrite-protection guard**: `project_write_file` is rejected only when the target `(slug, path)` **already exists** (the guard probes `GET /projects/{slug}/files/raw`) and the agent has not read it this session (call `project_read_file` first); creating a new file needs no prior read. Reads are recorded in the per-session `Dispatcher._read_files` set. This steers the agent to update files rather than blindly overwrite them and is agent-scoped only (the human UI and public HTTP API are unaffected). The `WRITING AND EDITING FILES` system-prompt rule mirrors it.
## Web fetch and arbitrary HTTP (`handler="fetch"`)
`services/devii/fetch/FetchController` (a standalone `httpx.AsyncClient`, not `PlatformClient`) backs two tools: `fetch_url` (read a page as readable text, `requires_auth=False`) and `http_request` (`requires_auth=True`) - the general external HTTP client. `http_request` takes `method` (GET/POST/PUT/PATCH/DELETE/HEAD/OPTIONS), `headers` (object), and one body of `json` (object/array, Content-Type set automatically), `form` (url-encoded), or `body` (raw string), and returns `{http_status, headers, content_type, content}`; it does **not** raise on non-2xx so the model can read API errors. Both run through one `_stream()` (`fetch_url` via `_download` with `raise_5xx=True`; `http_request` with `raise_5xx=False`) sharing the SSRF `_guard` (refuses private/loopback unless `DEVII_FETCH_ALLOW_PRIVATE`), `fetch_max_bytes` cap, timeout, and decoding. The `object`-typed params required adding an `object` branch to `Action.tool_schema()` in `spec.py` (emits a valid open-object schema). The `SYSTEM_PROMPT` "WEB REQUESTS AND EXTERNAL APIS" rule tells the agent to call `http_request` DIRECTLY for an API and never wrap it in a self-invoking user-defined tool (a virtual tool only re-runs the agent and cannot do I/O, so wrapping the call recurses into the `MAX_EVAL_DEPTH` guard - the exact failure this fixed). To attach a remote file see `attach_url` / `store_attachment_from_url` under Attachments and media.
## Remote web tools (rsearch, non-platform)
`services/devii/rsearch/RsearchController` (`handler="rsearch"`, registered in `registry.py`, specs in `actions/rsearch_actions.py`, all `requires_auth=False`) exposes four tools that hit the EXTERNAL public service `https://rsearch.app.molodetz.nl` over a standalone `httpx.AsyncClient` (like `FetchController`, not the platform-bound `PlatformClient`): `rsearch` (web/image search, optional `content`/`deep`/`type=images`), `rsearch_answer` (AI answer grounded in fresh web search, returns answer + sources), `rsearch_chat` (direct AI chat, no search), `rsearch_describe_image` (vision describe a public image URL). These are NOT platform-specific: the `SYSTEM_PROMPT` "REMOTE WEB TOOLS" section makes the model prefer platform tools always and call `rsearch_*` only when the user explicitly asks to search the web/an outside source, never to answer questions about this instance. When adding another remote service, add a controller + handler literal in `spec.py` + a dispatcher branch rather than routing external calls through `PlatformClient` (which is hardwired to the instance base URL).
rsearch has its **own** read timeout (`rsearch_timeout_seconds`, field `devii_rsearch_timeout` / env `DEVII_RSEARCH_TIMEOUT`, default **300s** with a 30s connect cap via `httpx.Timeout`) - it does NOT share `fetch_timeout_seconds`, because web-grounded answers legitimately take minutes.
**rsearch is metered into the gateway ledger.** Because these calls never traverse the gateway upstream, `RsearchController` (constructed with `owner_kind`/`owner_id` from the dispatcher) records one `gateway_usage_ledger` row per call inside its single choke point `_request()` via `GatewayUsageLedger.record_external(...)` (backend `rsearch`, zero tokens, the flat admin-set `gateway_rsearch_cost_per_call`, success and failure both logged). The DeepSearch worker's `crawl.search_queries` emits a `{"type":"rsearch"}` NDJSON frame per search that `DeepsearchService._run_worker` ledgers in-process under the job owner. This keeps `gateway_usage_ledger` (and `/admin/ai-usage`) a complete record of platform AI spend - external AI calls included. `record_external` is the helper to reuse for any future off-gateway AI call.
## Email tools (IMAP/SMTP, non-platform)
A signed-in user's Devii can connect to their OWN external mailbox via `services/devii/email/EmailController` (`handler="email"`, specs in `actions/email_actions.py`). See `devplacepy/services/email/CLAUDE.md` for the full protocol engine, credential storage, and tool list.
## Telegram bot
Devii is reachable over Telegram via `channel="telegram"` sessions driven in-process by `services/telegram/bridge.py`. See `devplacepy/services/telegram/CLAUDE.md` for the worker/bridge architecture, pairing flow, and message-update behavior.
## Database API tools (`db_*`, primary-administrator only)
Devii exposes read-only `db_list_tables`, `db_table_schema`, `db_list_rows`, `db_get_row`, `db_query`, and `db_design_query` tools gated `requires_primary_admin=True`. See `devplacepy/services/dbapi/CLAUDE.md` for the full auth boundary, validation pipeline, and async query service - these tools are added to the LLM tool list only for the primary administrator; every other session is unaware they exist.
## All Devii/gateway network timeouts are admin-configurable with a five-minute (300s) floor
The Devii LLM-client and `PlatformClient` read timeout is `timeout_seconds` (field `devii_timeout`), the fetch/docs tool timeout is `fetch_timeout_seconds` (field `devii_fetch_timeout`), web search is `rsearch_timeout_seconds` (`devii_rsearch_timeout`); each defaults to 300s with `minimum=MIN_TIMEOUT_SECONDS` (300). The gateway upstream timeout is `gateway_timeout` (`minimum=TIMEOUT_MIN`=300), which also bounds the vision describe-image call (it shares the gateway's httpx client, no per-call override). A stored value below the floor fails `ConfigField.coerce()` and `read()` falls back to the default, so old sub-300 values upgrade automatically. Devii's LLM timeout was previously a hardcoded 45s while the gateway waited up to 180s x retries - large prompts aborted as `[model error] Could not reach the model endpoint:` (an httpx timeout, which stringifies to empty). `build_settings()` now reads these from config instead of hardcoding them.
## Config
Config is all `config_fields` (AI url/model/key, base url, plan/verify toggles, max iterations, JS-execution toggle, user/guest 24h caps, guests toggle, pricing). `effective_config()` falls back the AI key to `DEVII_AI_KEY` and the base url to the instance origin.
## Backend confidentiality
Devii must never disclose the underlying model, provider, or any upstream URL - enforced in two layers. The `SYSTEM_PROMPT` (`agent.py`) forbids it even when such values appear inside a tool result. Structurally, `text.redact_backend()` runs on every JSON HTTP tool response in `format_response()` and scrubs the values of `REDACT_FIELD_KEYS` (`gateway_upstream_url`/`gateway_model`/`gateway_vision_url`/`gateway_vision_model`) and any stat labelled in `REDACT_STAT_LABELS` (`Model`) to `[hidden]`, so the admin-services tools cannot leak the gateway's upstream even though the admin settings UI still shows the real values. Add a config key to `REDACT_FIELD_KEYS` if a new field would expose backend infrastructure to the agent.
## CLI
`pyproject.toml [project.scripts]` ships `devii = "devplacepy.services.devii.cli:main"`. No new dependency (`httpx`/`dataset` already required). The md-clippy avatar is vendored under `static/vendor/md-clippy/`; its `index.js` must not set `globalThis.app` (it would clobber the DevPlace `app`) and its AI proxy is `/devii/clippy/ai/chat`. The standalone `devii` CLI runs with `is_admin=True` (the local operator owns the process) and both `is_admin`/`is_primary_admin` `True` for the trusted local operator.
## Devii user-defined ("virtual") tools
Users invent new Devii tools in natural language ("when I say woeii, do Y"); each is stored per-owner and added to Devii's live LLM tool list, and when called its handler **re-prompts Devii itself** (a self-eval sub-agent) with the stored prompt plus the user's single free-form `input`. Full CRUD is Devii-only (`tool_create`/`tool_list`/`tool_get`/`tool_update`/`tool_delete`, `handler="virtual_tool"`).
- **Dynamic tool list (the crux).** The session tool list is normally frozen (`session.py` `self.tools = CATALOG.tool_schemas_for(...)`, handed by reference to `Agent`, `AgenticController.bind`, and read live by `react_loop` each `llm.complete`). `DeviiSession._refresh_tools()` runs at the top of every `_run_turn` (inside `self._lock`, before `agent.respond`) and rewrites that list **in place**: `self.tools[:] = CATALOG.tool_schemas_for(self.client.authenticated, self.is_admin) + self._virtual_tool_store.tool_schemas()`. Because every holder shares the same list object, a tool created/edited/deleted (and a mid-session login) takes effect on the **next** turn with no Agent rebuild. **Never reassign `self.tools`, always mutate in place.**
- **Self-eval engine** (`agentic/controller.py`). `_delegate` was refactored onto `_spawn(prompt, tools, system_prompt)`, which runs `react_loop` with a fresh `messages`/`AgentState` under a **depth guard** (`agentic/state.py` `get/set/reset_eval_depth`, `MAX_EVAL_DEPTH=2`) so delegate/eval/virtual tools cannot self-call infinitely. `run_subagent(prompt)` (tools minus `delegate`) is the public engine; the `eval` agentic tool wraps it; `_delegate` keeps its richer JSON report.
- **Store** (`services/devii/virtual_tools/store.py`). `VirtualToolStore(db, owner_kind, owner_id)` over `devii_virtual_tools` (`uid, owner_kind, owner_id, name, description, prompt, input_description, enabled, created_at, updated_at`; indexes `idx_devii_vtools_owner` / `idx_devii_vtools_name`). `tool_schemas()` builds, for each **enabled** row, an LLM schema with a single free-form `input` string. Persistent for users, `memory_db()` for guests (built in `hub.get_or_create` from the shared `owned_db`).
- **Controller** (`virtual_tools/controller.py`). `VirtualToolController(store, evaluator, builtin_names)`, evaluator = `agentic.run_subagent`, `builtin_names = set(CATALOG.by_name())`. CRUD `tool_create/list/get/update/delete` (`handler="virtual_tool"`); `tool_create` validates name `^[a-zA-Z][a-zA-Z0-9_]{1,40}$`, rejects built-in collisions and duplicates, requires `description`+`prompt`. `has(name)` / `run(name, args)` compose `f"{prompt}\n\nUser input: {input}"` and call the evaluator.
- **Dispatch** (`actions/dispatcher.py`). Constructed with `virtual_tools=...`; `_run` routes `handler="virtual_tool"` to CRUD; the **unknown-name branch** (`if action is None`) resolves a virtual tool via `self._virtual_tools.has(name)` -> `run(name, args)` before erroring. So virtual tools live only in the store (not the static catalog) and are resolved on demand; the static `_actions` stays built-in only.
- Registered via `VIRTUAL_TOOL_ACTIONS` (`registry.py`); the `eval` tool via `AGENTIC_ACTIONS`. System-prompt **USER-DEFINED TOOLS (VIBE TOOLS)** section steers creation and warns against deep self-calls.
## Devii self-configured behavior (`services/devii/behavior/`)
Devii can partially configure its OWN system message: every system prompt ends with a `# TRUTH RULES AND BEHAVIOR` section (the `BEHAVIOR_HEADER` constant in `session.py`) whose body is an owner-scoped, persistent set of behavior rules. When the user tells Devii to behave differently, says they expect different behavior, or Devii upsets them, Devii calls the **`update_behavior`** tool (`handler="behavior"`, `requires_auth=False`, single `behavior` arg = the FULL new section content) to persist the change. It is **self-prompted only**: the sole way to change the section is Devii calling that tool - there is no HTTP route or UI, like `LessonStore`/customization. Because Devii already SEES the current section in its own system message, it merges (copy current rules, apply the change, send the whole result) so prior rules are not lost; this is why the tool replaces rather than appends.
- **Store** (`behavior/store.py`). `BehaviorStore(db, owner_kind, owner_id)` over `devii_behavior` (one upserted row per owner keyed on `owner_kind`/`owner_id`; `text()` reads, `set()` upserts; index `idx_devii_behavior_owner`). Persistent for users, `memory_db()` for guests (built in `hub.get_or_create` from the shared `owned_db`, like the other owner stores).
- **Controller** (`behavior/controller.py`). `BehaviorController(store)`, `dispatch("update_behavior", args)` -> `store.set(behavior)`. Built in `DeviiSession` and passed to `Dispatcher(behavior=...)`; the dispatcher routes `handler="behavior"` to it and degrades gracefully ("not available in this context") when unwired (e.g. the standalone CLI, same as `virtual_tools`/`avatar`).
- **Injection and refresh** (`session.py`). `_compose_system_prompt()` = base prompt (`_system_prompt_for(is_admin)`) + `\n\n` + `BEHAVIOR_HEADER` (+ `\n` + body when non-empty; just the header when empty). Used to seed the `Agent` and the scheduler executor worker, and `_refresh_system_prompt()` rewrites `agent._messages[0]["content"]` **at the top of every `_run_turn`** (next to `_refresh_tools`), so a mid-conversation `update_behavior` takes effect on the following turn and the live system message is never overridden by the stale base.
- Registered via `BEHAVIOR_ACTIONS` (`registry.py`); the system-prompt **SELF-CONFIGURED BEHAVIOR (TRUTH RULES)** section in `agent.py` steers when/how to call it.
## Related Devii tool wrappers (customization, container)
Two more Devii-side controllers live under `services/devii/` but their full mechanism is documented in their owning subsystem's file, not here:
- **Per-user customization tools** (`services/devii/customization/`, `handler="customization"`): `customize_list`/`get`/`set_css`/`set_js`/`reset` (all `requires_auth=False`) plus `customize_set_enabled` (`requires_auth=True`, the suppression toggles). `CustomizationController(owner_kind, owner_id)` is built in `Dispatcher.__init__`; set/reset are in `CONFIRM_REQUIRED` so `confirmation_error` forces Devii to ask **page-type vs global** before `confirm=true`. Devii previews live with `run_js` (ids `devii-preview-css`/`devii-preview-js`) and `reload_page` after saving. For the storage model, injection pipeline, and per-user suppression flags, see `devplacepy/customization.py` and its own documentation.
- **Container tools** (`services/devii/container/ContainerController`, `handler="container"`): wraps `services/containers/api.py`, the same operations shared by `routers/containers.py`. For the reconciler, backend ABC, ingress proxy, and workspace sync mechanism, see `devplacepy/services/containers/CLAUDE.md`.
+32
View File
@@ -0,0 +1,32 @@
This file documents the email (IMAP/SMTP) subsystem. Claude Code auto-loads it when a file under `devplacepy/services/email/` is read or edited.
## Email via Devii (`services/email/` + `services/devii/email/`)
A signed-in user's Devii can connect to their OWN external mailbox over IMAP/SMTP (non-platform, external protocol tools).
### Protocol engine and adapter
The protocol engine `services/email/` (`EmailAccount` dataclass + `EmailClient`) is built over the Python **stdlib** `imaplib`/`smtplib`/`email` - no new dependency, since IMAP/SMTP are not HTTP the stealth-client rule does not apply. It is wrapped by the Devii adapter `services/devii/email/EmailController` (`handler="email"`, registered in `registry.py`, specs in `actions/email_actions.py`, routed in `dispatcher.py`). Every blocking protocol call runs through `asyncio.to_thread` in the controller so the single worker loop never stalls (same discipline as `correction.py`).
### Tools
- Connection CRUD: `email_account_set`/`email_account_get`/`email_accounts_list`/`email_account_delete`.
- Message ops: `email_list_folders`, `email_list_messages`, `email_search`, `email_read_message` (read); `email_mark`/`email_set_flags`/`email_move_message` (organise); `email_delete_message`; `email_send`.
All are `requires_auth=True` AND additionally guarded to `owner_kind == "user"` (guests never reach email), gated by the admin `devii_email_enabled` flag (default on) / `devii_email_timeout` on the Devii service config ("Email" service-config group).
### Credentials
Credentials live in the soft-deletable per-owner `email_accounts` table (one row per `(owner_kind, owner_id, label)`; helpers `list/get/set/delete_email_account` in `database/`, sensible defaults imap 993/ssl + smtp 587/starttls applied in `set_email_account`). The password is stored **plaintext** (consistent with `api_key`/`gateway_api_key`) and is **NEVER echoed back** - `email_account_get`/`email_accounts_list` mask it as `password_set`.
### Configuration surface
Configuration is **Devii-only** (no HTTP route or profile UI), like the CSS/JS customization feature.
### Destructive actions and SSRF guard
`email_send` is a real outbound action, and the two delete tools (`email_account_delete`, `email_delete_message`) are confirmation-gated (`CONFIRM_REQUIRED`, each declaring a `confirm` param). The IMAP/SMTP host is run through `net_guard.guard_public_host_sync` before every connect (private/loopback addresses refused - SSRF/internal-scan defense, mirroring `guard_public_url`).
### Audit and metering
Audited via the existing `_audit_mechanic` choke point under the `email.*` domain (category `email`). Mutations are **not metered into the gateway ledger** (no LLM call is made).
+57
View File
@@ -0,0 +1,57 @@
# Code Farm game (`devplacepy/services/game/`, `devplacepy/routers/game/`)
This file documents the Code Farm idle game. Claude Code auto-loads it when a file under `devplacepy/services/game/` is read or edited.
## Overview
`game/` package (routers, mounted at `/game`) is the **Code Farm** idle game (`index.py` + `farm.py`): `GET /game` (page), `GET /game/state`, `GET /game/leaderboard`, `POST /game/{plant,harvest,buy-plot,upgrade,fertilize,daily,perk,prestige,legacy,quests/claim}`, plus social `GET /game/farm/{username}` and `POST /game/farm/{username}/{water,steal}`.
A cooperative-and-competitive idle game (Farmville-style) mounted at `/game`, member-only to play, public to view another farm. Cooperative loop = watering a neighbour's growing build; competitive loop = stealing a neighbour's ready build.
## Data layer
**Data layer is pure + timestamp-driven (no background tick).** `services/game/economy.py` holds every constant and formula as frozen dataclasses/functions: the `CROPS` tuple (key/name/icon/cost/grow_seconds/reward_coins/reward_xp/min_level), `CI_TIERS` (speed multiplier + upgrade cost), level thresholds (`xp_threshold`/`level_for_xp`/`level_progress`), `plot_cost` (doubles per extra plot), and watering bonus. `services/game/store.py` is the only DB access (`game_farms`, `game_plots`). **A plot's state is derived from `ready_at` vs now, never stored** - growing crops finish purely by the clock, so there is no reconciler/service. `serialize_farm(farm, viewer=, owner=)` computes plot states, remaining seconds, `can_water` (viewer is not owner, build growing, viewer not already in the per-cycle `watered_by` JSON list, under `MAX_WATERS_PER_PLOT`), `can_steal`/`steal_coins` (viewer is not owner, build ready, and `now >= ready_at + STEAL_GRACE_SECONDS` so the owner gets a protection window), level progress, and the plantable crop list. `serialize_plot` takes the owner farm's `yield_level`/`prestige` (threaded from `serialize_farm`) only to compute the steal payout. Mutations (`plant`/`harvest`/`buy_plot`/`upgrade_ci`/`water`/`steal`) raise `GameError` on any invalid op (insufficient coins, locked crop, wrong state, still-protected harvest); the routers translate that to a `400` JSON error or a redirect.
## Tables
`game_farms` (one per user, `coins`/`xp`/`level`/`ci_tier`/`plot_count`/`total_harvests`, plus `prestige`/`streak`/`last_daily_at`, the four `perk_*` columns, and the endgame `stars` + five `legacy_*` columns), `game_plots` (`farm_uid`/`slot_index`/`crop_key`/`planted_at`/`ready_at`/`watered_by`), and `game_steals` (`thief_uid`/`owner_uid`/`slot_index`/`crop_key`/`coins`/`stolen_at`, the per-pair steal-cooldown ledger) are **not** soft-deletable - they are mutable game state consumed by state transitions, not user content. Columns + indexes are ensured in `database.init_db` (unique `idx_game_farms_user`, `idx_game_farms_rank`, `idx_game_plots_farm`, `idx_game_steals_pair` on `(thief_uid, owner_uid, stolen_at)`); the `_uid_index` loop adds the uid index. `ensure_farm(user_uid)` lazily creates a farm + starting plots on first access (idempotent), so there is no signup hook.
## Routes (`routers/game/`)
`index.py` is the base router - `GET /game` (own farm page), `GET /game/state` (own farm JSON), `GET /game/leaderboard`, and the action POSTs `plant`/`harvest`/`buy-plot`/`upgrade`. `farm.py` adds `GET /game/farm/{username}` (view), `POST /game/farm/{username}/water`, and `POST /game/farm/{username}/steal`. Every action returns the **full updated farm** as JSON (so the client refreshes in one round trip) or redirects for no-JS, via the shared `_respond_action` choke (the `farm.py` water/steal handlers inline the same shape). Harvest awards site XP (`award_rewards`) and the `harvest`/`water` achievements (`track_action`). All action handlers are `require_user`; reads of other farms/the leaderboard are public.
**Leaderboard ranking is a composite `economy.farm_score(farm)`** (one integer per row, computed in memory over the already-loaded `_farms().find()` set, so it stays fast): it sums weighted contributions from every tracked factor - `xp`, `prestige` * `SCORE_PRESTIGE` (5000, dominant since a refactor is a full completed cycle), `total_harvests` * `SCORE_HARVEST`, `coins` // `SCORE_COIN_DIVISOR`, `(ci_tier-1)` * `SCORE_CI`, `(plot_count-STARTING_PLOTS)` * `SCORE_PLOT`, the summed perk levels * `SCORE_PERK`, and `min(streak, SCORE_STREAK_CAP)` * `SCORE_STREAK` - so a player who refactored (which resets xp/level/coins/ci/plots/perks) is no longer buried below a never-refactored higher-level player. The weights are module-level constants in `economy.py` for tuning; the leaderboard entry carries `score` and `prestige` (surfaced on `GameLeaderboardEntryOut`) and `GameFarm.js` renders the score next to `Lv X`.
## Live + frontend
After any mutation the handler `await`s `_shared.notify_farm(username)` which publishes a nudge to the `public.game.farm.{username}` pub/sub topic; subscribed clients re-fetch their viewer-specific state (keeps `can_water` correct without broadcasting per-viewer payloads). `static/js/GameFarm.js` (`app.gameFarm`, auto-detects `[data-game-root]`) renders the grid/HUD/leaderboard, ticks plot countdowns client-side every second, delegates the `data-game-action` forms through `Http.send`, subscribes to the farm topic, and keeps a 20s `Poller` fallback. The server renders a full no-JS fallback grid (`templates/_game_grid.html`, shared by `game.html` and `game_farm.html` with progressive-enhancement POST forms).
## Fan-out
Devii plays via the `game_*` `http` actions in `actions/catalog.py` (state/leaderboard/view public, the rest `requires_auth`); the API reference has a **Code Farm** group in `docs_api.py`; pages are `noindex,follow` (interactive, user-specific) so they are intentionally not in the sitemap; badges live in `utils.BADGE_CATALOG` under the **Code Farm** group with `harvest`/`water`/`harvest_stolen`/`got_stolen_from` `ACHIEVEMENTS` (the steal pair awards **Cat Burglar** to the thief and **Robbed** to the victim, both threshold 1).
## Stealing (competitive loop, backwards compatible)
`store.steal(thief, owner, slot)` mirrors `water`: it requires the owner plot to be `ready` AND past the protection window `economy.effective_steal_grace(owner_defense_level)` (base `STEAL_GRACE_SECONDS` 60s, +30s per owner Branch Protection level) measured from `ready_at`, AND that the thief is **off cooldown for this victim** (`steal_cooldown_remaining(thief, owner, now) == 0`, i.e. no row in `game_steals` for the pair within `STEAL_COOLDOWN_SECONDS` = 3600 - you can raid a given neighbour only once per hour); it clears the plot exactly like a harvest (owner gets nothing), credits the **thief's** farm `economy.steal_reward_coins` (the owner's yield/prestige/legacy-multiplier realized value times `effective_steal_fraction(owner_defense_level)`, base `STEAL_FRACTION` 0.5, -5% per defense level, floored at 0.1) and **coins only** - no XP, no `total_harvests`, so the leaderboard stays earned by real farming - then inserts a `game_steals` row stamping the cooldown. The route `POST /game/farm/{username}/steal` (reusing `GameSlotForm`) then `track_action`s both sides, fires `create_notification(owner, "harvest_stolen", "Someone raided your Code Farm...", thief_uid, "/game")` (the message never names the thief; `related_uid` is internal), and `await notify_farm`s **both** the owner and the thief usernames so both farms refresh live. The new `harvest_stolen` notification type rides the existing in-app relay (live toast, no new wiring). No DB migration: grace + payout are computed from existing `ready_at`/perk/legacy columns, the `game_steals` table is created by the ensure-block (absent rows -> cooldown 0 -> first steal always allowed), and the new `GamePlotOut.can_steal`/`steal_coins`/`steal_cooldown_seconds`/`steal_reason` + `GameFarmOut.steal_cooldown_seconds` + `GameFarmViewOut.stole_coins` schema fields default safe. The per-pair cooldown is computed **once per farm** in `serialize_farm` (when the viewer is not the owner) and threaded into every `serialize_plot` as `steal_locked_until`, so `can_steal` is false and `steal_reason` is `"cooldown"`/`"protected"` accordingly. Frontend: the ready/not-owner branch of `_game_grid.html` and `GameFarm.js._plotHtml` render a `btn-danger` Steal button (`data-game-action="steal"`, `data-confirm` gated like prestige) carrying `steal_coins`, or a disabled "Raid again in <countdown>" label when on cooldown; on success `GameFarm.js._submit` toasts the payout from `data.stole_coins`.
## Extended mechanics (all backwards compatible)
Five further systems layer onto the base loop, every one defaulting gracefully for pre-existing `game_farms`/`game_plots` rows: new `game_farms` columns (`prestige`, `streak`, `last_daily_at`, `perk_yield`/`perk_growth`/`perk_discount`/`perk_xp`) are added in the `init_db` ensure-block, and `store._lvl(farm, key)` reads every one as `int(farm.get(key) or 0)` so a legacy NULL row behaves as level 0 / no streak / no perks.
1. **Daily bonus** (`POST /game/daily`, `store.claim_daily`): once per UTC day, consecutive days grow `streak` (reward `economy.daily_reward`, capped at `DAILY_STREAK_CAP`).
2. **Daily quests** (`game_quests` table, `POST /game/quests/claim`): `economy.daily_quests(user_uid, day)` deterministically (sha256 of `user:day`) picks 3 of {plant, harvest, water, earn} with goals/rewards; `store.ensure_quests` lazily materializes the day's rows, `store.advance_quests(user_uid, kind, amount)` is called inside `plant`/`harvest`/`water` (wrapped in try/except so a quest write never breaks the action), `claim_quest` pays out when `progress >= goal`.
3. **Perks** (`POST /game/perk`, `store.upgrade_perk`): four permanent upgrades (`economy.PERKS`) with escalating `perk_cost`; applied in the economy formulas - `effective_plant_cost` (discount), `grow_seconds_for(crop, ci_tier, growth_level)` via `farm_speed` (growth), `effective_reward_coins` (yield + prestige), `effective_reward_xp` (xp).
4. **Fertilizer** (`POST /game/fertilize`, `store.fertilize`): cut a growing plot's `ready_at` by `FERTILIZE_FRACTION` for `economy.fertilize_click_cost(eff_reward, reduce_seconds, full_grow_seconds)` = `ceil(eff_reward * reduce/full_grow * FERTILIZE_TAX)` (TAX 1.05). **The cost is priced against the build's realized harvest value (`effective_reward_coins`, which already carries yield/prestige/legacy multipliers), not raw grow-seconds, so the prestige dependence cancels and fully fertilizing a crop always costs >= its harvest - fertilize is a pure time-skip and can NEVER be a profit at any prestige.** (This replaced the old grow-seconds-based `fertilize_cost`, which was an unbounded money pump at high prestige.)
5. **Prestige/Refactor** (`POST /game/prestige`, `store.prestige`): at `PRESTIGE_MIN_LEVEL` resets coins/xp/level/ci/perks, sets `plot_count = economy.prestige_base_plots(legacy_plots_level)` (keeps/recreates plot rows up to that base, deletes the rest), increments `prestige` for a permanent `prestige_multiplier` (+25% coins each), and awards `economy.stars_for_refactor(level, prestige)` Stars (the `legacy_*`/`stars` columns are **omitted from the reset dict** so they survive every refactor, the established pattern).
`serialize_farm` exposes all of this (`perks`, `quests`, `streak`, `daily_available`/`daily_reward`, `prestige*`, `stars`, `legacy`, `steal_cooldown_seconds`) only to the owner; the existing `crop_payload`/`serialize_plot` now carry perk- and legacy-adjusted costs/rewards and per-plot value-based `fertilize_cost`. Frontend hosts (`[data-shop-host]`/`[data-perk-host]`/`[data-legacy-host]`/`[data-daily-host]`/`[data-quest-host]`) are server-rendered from partials (`_game_shop.html`/`_game_perks.html`/`_game_legacy.html`/`_game_daily.html`/`_game_quests.html`) and fully re-rendered by `GameFarm.js` builders; the generic `data-game-action` form delegation handles the new actions, with `data-confirm` gating the prestige reset.
## Endgame: Stars, Legacy upgrades, auto-harvest, golden builds (all backwards compatible)
The infinite progression for maxed farms. Six new default-0 `game_farms` columns (`stars`, `legacy_autoharvest`, `legacy_multiplier`, `legacy_speed`, `legacy_plots`, `legacy_defense`), all read via `_lvl`.
**Stars** are a meta-currency earned only on Refactor (`stars_for_refactor`), spent via `POST /game/legacy` (`store.upgrade_legacy`, `GameLegacyForm`, Devii `game_upgrade_legacy`) on `economy.LEGACY_UPGRADES` (escalating `legacy_cost` in Stars) that **survive prestige** unlike perks: `autoharvest` (CI Bot), `multiplier` (+10% coins/lvl via `legacy_multiplier`, folded into `effective_reward_coins`), `speed` (+5% base build speed/lvl via `farm_speed`/`grow_seconds_for`/`water_bonus_seconds`), `plots` (+1 base plot after refactor via `prestige_base_plots`), `defense` (steal grace/fraction via `effective_steal_grace`/`effective_steal_fraction`).
**Auto-harvest** is lazy and tick-free: `store._auto_harvest(farm, owner_uid, now)` runs at the top of `serialize_farm` ONLY when the viewer is the owner AND `legacy_autoharvest > 0`; it **clears each ready plot first then credits** the pre-clear crop's coins/xp/`total_harvests` in one `_update_farm`, advances quests, and re-reads the farm - clear-then-credit is idempotent (a second read sees empty plots), and it never fires on a visitor's `GET /game/farm/{username}` view or the leaderboard, honouring the no-background-service invariant (both `GET /game` and `GET /game/state` go through `state_payload` with viewer=owner, so both auto-collect).
**Golden builds** are deterministic and storage-free: `economy.is_golden(plot_uid, planted_at)` (sha256, ~`GOLDEN_CHANCE`) marks a planting golden for its life; `harvest`/`_auto_harvest` multiply **coins only** by `GOLDEN_MULTIPLIER` (XP unscaled to keep level pacing), and `serialize_plot.is_golden` surfaces a sparkle badge (`.game-plot-golden`). New schema fields (`GameLegacyOut`, `GameFarmOut.stars`/`legacy`, `GamePlotOut.is_golden`) all default safe; no DB migration.
+48
View File
@@ -0,0 +1,48 @@
This file documents the Gitea issue-tracker integration subsystem. Claude Code auto-loads it when a file under `devplacepy/services/gitea/` is read or edited.
## Issue tracker (`services/gitea/`)
The `/issues` feature is a **full Gitea integration with no local issue store** - there is no local issue store; the listing and detail views read issues straight from Gitea (default repo `retoor/devplacepy` on `retoor.molodetz.nl`) with live status. All connection settings live in `site_settings` and are admin-editable on the `IssueTrackerService` config at `/admin/services`.
### Package `services/gitea/`
- `config.py` - the admin-editable `CONFIG_FIELDS` (Gitea base URL/owner/repo/token, AI enhance toggle/model/key) plus `gitea_config()` -> frozen `GiteaConfig` (with `api_base`/`repo_base`/`is_configured`) and `is_configured()`. The token is one shared bot account; per-user attribution is done by storing names in DevPlace, not by per-user Gitea accounts.
- `client.py` - async `GiteaClient` over the Gitea REST API (`Authorization: token`): `create_issue`, `list_issues(state,page,limit)` returning `(issues, total)` from `X-Total-Count` and filtering out pull requests, `get_issue`, `list_comments`, `create_comment`, `set_state`. Raises `GiteaError(message, status)` on any non-2xx or transport error.
- `fake.py` - in-memory `FakeGiteaClient` mirroring the same async interface (plus `add_external_comment` to simulate a developer reply); the test backend/double.
- `runtime.py` - `get_client()` / `set_client()` swap (tests inject the fake). `get_client()` builds a fresh `GiteaClient` from current config unless overridden.
- `enhance.py` - `enhance_ticket(title, description, config)` rewrites a raw report into a consistent markdown ticket via the **internal AI gateway** (`INTERNAL_GATEWAY_URL`, key defaults to the gateway internal key): a level-2-heading markdown body (`Summary` / `Steps to Reproduce` / `Expected Behaviour` / `Actual Behaviour` / `Environment`) and a tightened title. It is fail-soft to a deterministic template: any error or unparseable response falls back to a deterministic template built from the original text (`EnhancedTicket.enhanced=False`). This is the AI attachment point - the gateway, like every other AI consumer.
- `store.py` - the **only DB state**, a thin local mapping: `issue_tickets` (gitea_number -> author_uid + original text + cached `last_status`/`last_comment_count` for update detection) and `issue_comment_authors` (gitea_comment_id -> author_uid, for display attribution and to tell DevPlace-authored comments from developer replies). Helpers: `record_ticket`, `get_ticket`, `author_uid_for_issue`, `author_map`, `record_comment_author`, `comment_author_map`, `local_comment_ids`, `tracked_tickets`, `update_ticket_cache`. Both tables are in `SOFT_DELETE_TABLES`; `init_db` ensures their full column set + indexes (so queries are safe on a fresh DB before any insert).
- `service.py` - `IssueTrackerService(BaseService)` (`default_enabled=False`, holds `CONFIG_FIELDS`), polls tracked tickets and notifies the **reporter** on developer replies / status changes. Each tick it polls every tracked ticket: when the Gitea comment count grew, it loads comments and notifies the reporter if any of the new tail comments are **not** in `local_comment_ids` (a developer reply); when the state changed it notifies the reporter (closed/reopened). It then updates the cache. This is the single source of update notifications to the author (`issue.sync.reply` / `issue.sync.status`).
### Filing (async job)
Filing is an **async job** (`services/jobs/issue_create_service.py` `IssueCreateService`, kind `issue_create`): `process` loads the user, enhances the report, appends a `Reported by **user** via DevPlace` footer, creates the Gitea issue, records the local mapping, notifies the reporter (`Your issue report was filed as #N`), and audits `issue.create`. `cleanup` is a no-op (the issue is permanent; only the job row is swept). `POST /issues/create` enqueues and returns `{uid, status_url}`; frontend `app.issueReporter` (`static/js/IssueReporter.js`) submits the create form, polls `/issues/jobs/{uid}` via `JobPoller`, and redirects to `/issues/{number}`.
### Routes (`routers/issues/` package, prefix `/issues`)
- `GET /issues` - Gitea list, `?state=open|closed|all&page=`, admin-style `build_pagination` + `_pagination.html` with a `state=` prefix; renders an empty-state notice when unconfigured/unreachable.
- `POST /issues/create` - enqueue (see Filing above).
- `GET /issues/jobs/{uid}` - `IssueJobOut`.
- `GET /issues/{number}` - `issue_detail.html`: issue body + comments rendered as markdown via `data-render`, decorated with DevPlace authors or a `dev` badge.
- `POST /issues/{number}/comment` - synchronous: posts the comment to Gitea attributed to the user, records the author mapping, notifies **all admins**, audits `issue.comment`.
- `POST /issues/{number}/status` - admin-only open/closed via Gitea PATCH, audits `issue.status`; non-admins get a `denied` audit + 403.
Comments and status changes are deliberately synchronous (one fast Gitea call, immediate redirect); only the AI-heavy create is a job. Status-change notifications to the author are NOT emitted by the route - the poller detects the diff and notifies, so the admin (or external developer) closing an issue both flow through one path.
### Content rendering and mentions (frontend reuse, no new modules)
The issue body and every comment render through the platform content pipeline like posts/comments - each is a `<div class="...rendered-content" data-render>{{ body|comment.body }}</div>` (`issue_detail.html`), so the globally-instantiated `ContentEnhancer` (`app.content`) runs `contentRenderer.applyTo` (marked -> DOMPurify -> highlight.js -> media/autolink) over Gitea-sourced markdown client-side from `element.textContent`. Gitea content is **NEVER** marked `| safe` into raw HTML; Jinja autoescapes it and the DOMPurify step inside `ContentRenderer` is the XSS control, identical to user posts. The issue authoring textareas (`#issue-description` in the create modal on `issues.html`, and the comment textarea on `issue_detail.html`) carry the `data-mention` attribute, so the same `ContentEnhancer.initMentionInputs()` -> `MentionInput` flow that wires post `@`-mention autocomplete (debounced `GET /profile/search?q=` dropdown) attaches to them with no new code. There is no local issue-edit form (issues are AI-created then commented), so "edit" reuses nothing further. `mention.css` is loaded globally in `base.html`.
### Attachments (mirror: local store + Gitea assets, `routers/issues/attachments.py`)
Issues and issue comments accept file attachments, with full add/delete CRUD allowed **only while the issue is open**. Storage is a mirror - the canonical copy is the platform attachment system (`attachments` table, keyed `target_type="issue"`/`target_uid=str(number)` and `target_type="issue_comment"`/`target_uid=str(comment_id)`), and each file is also pushed to the Gitea native asset API so it appears on the tracker. The local row carries a `gitea_asset_id` back-reference (added to the `init_db` attachments ensure-block + born-live in `store_attachment`) so a later delete removes both copies; `attachments.mirror_attachment_to_gitea(uid)` / `remove_gitea_asset(row)` branch on `target_type` to call `GiteaClient.create_issue_asset`/`create_comment_asset` / `delete_issue_asset`/`delete_comment_asset` (all best-effort - the local copy and the user action survive any Gitea failure). Upload reuses the canonical **orphan-then-link** path: `dp-upload` stores to `/uploads/upload`, the form carries `attachment_uids`, and the handler `link_attachments(uids, target_type, target_uid)` then mirrors.
**Role/state enforcement:** `require_user` blocks guests; add validates each uid is the caller's own unlinked orphan (admins any); delete is owner-or-admin on `attachments.user_uid`; every mutation re-checks the live Gitea `state == "open"` (closed -> 409) and audits denied branches. Attach-at-creation: `IssueForm.attachment_uids` rides into the `issue_create` job, which links + mirrors once the issue number exists; `IssueCommentForm.attachment_uids` links + mirrors inline after the comment is created (synchronous comment path).
Routes: `GET/POST /issues/{number}/attachments`, `DELETE /issues/{number}/attachments/{uid}`, and the `/issues/{number}/comments/{cid}/attachments[...]` comment variants. Frontend `app.issueAttachments` (`static/js/IssueAttachments.js`) submits the add form (`form[data-issue-attach]` -> POST `attachment_uids` -> reload) and handles per-attachment delete (`[data-attachment-delete]` -> confirm -> DELETE). The shared render partial is `_issue_attachments.html` (gallery + per-item delete gated by `can_modify`).
Note: for Gitea to accept every file type the instance `app.ini` `[attachment] ALLOWED_TYPES` must be `*/*` (infra, out of app scope); the DevPlace allow-list is the `allowed_file_types` setting, and the local copy is kept even when Gitea rejects a type. Audit events `issue.attachment.add`/`issue.attachment.delete` (category `content`). Devii tools `list_issue_attachments`, `add_issue_attachment`, `delete_issue_attachment` (+ `add_comment_attachment`/`delete_comment_attachment`); the two deletes are in `CONFIRM_REQUIRED`.
### Schemas, forms, and tools
Schemas: `IssueItemOut`/`IssueCommentOut`/`IssueDetailOut`/`IssueAttachmentsOut`/`IssuesOut`/`IssueJobOut` (`schemas.py`). Forms: `IssueForm`/`IssueCommentForm`/`IssueAttachmentForm`/`IssueStatusForm` (`models.py`). Devii tools: `list_issues`, `create_issue`, `view_issue`, `comment_issue`, `set_issue_status` (admin), plus the attachment tools above. Docs: the `issues` API group in `docs_api.py`. Footer link in `base.html` under `.site-footer`.
+127
View File
@@ -0,0 +1,127 @@
This file documents the async job services subsystem (devplacepy/services/jobs/ and its subdirectories deepsearch/, isslop/, seo/) - the standard pattern for running heavy blocking work off the request path. Claude Code loads it automatically whenever a file under this directory is read or edited.
## Overview: the JobService pattern
The standard way to run heavy, blocking work off the request path and hand back a result URL, used by every job kind in this subsystem (`zip`, `planning`, `fork`, `seo`, `seo_meta`, `deepsearch`, `isslop`). Reference: admin docs `Architecture -> Async job services` and `Services -> ZipService`.
- **The DB row is the queue.** `services/jobs/queue.py` (`enqueue`, `get_job`, `touch_job`, `list_jobs`) is pure DB and callable from any worker's handler - enqueue is a fast, non-blocking insert, status reads work from any worker, and only the lock owner processes. All kinds share ONE `jobs` table discriminated by a `kind` column: common lifecycle timestamps, `retry_count`, `last_accessed_at`/`expires_at`, `bytes_in`/`bytes_out`/`item_count` stat columns, plus `payload`/`result` JSON columns where kind-specific fields live.
- **`JobService(BaseService)`** runs only in the lock owner. Each `run_once`: reap finished in-flight tasks -> recover orphaned `running` rows (left by a dead owner, since there is only one processor) back to `pending` with bounded `retry_count` -> refill up to `max_concurrent` oldest `pending` jobs (uuid7 sorts FIFO) as `asyncio` tasks -> sweep rows past `expires_at` via the per-kind `cleanup()` hook. In-flight tasks span ticks (kept in an instance map); the loop returns immediately each tick, so keep the interval short (default 2s).
- **No atomic claim is needed** (single processor by construction) and **no separate reaper exists** - retention is built into every job service, default 7 days, admin-configurable; downloading/reading a result extends `expires_at` via `touch_job`.
- **Add a kind:** subclass `JobService`, set `kind`, implement `async process(self, job) -> dict` (return the `result` dict incl. stat keys) and `cleanup(self, job)`; register the service instance in `main.py`; add enqueue endpoints that own authz, a status route, a download/report route, a Devii tool, and docs. **A new `JobService` needs a server restart to go live.**
- **Permanent vs expiring artifacts (load-bearing distinction - decide this per kind).** Most kinds produce a DISPOSABLE artifact that expires with the retention sweep: `ZipService`, `SeoService`, and `DeepsearchService` all delete their real output (archive / report+screenshots / vector collection+report) inside `cleanup()`. Two kinds are different: `ForkService`'s artifact is a **permanent project**, and the AI Usage Analyzer's (`isslop`) artifacts (analysis, events, file/image results, report, badge) are **permanent public capability URLs** - both make `cleanup()` a no-op on the real artifact and let the retention sweep remove only the `jobs` tracking row. Get this decision right for any new kind: if the output should outlive the job the way a fork or an isslop report does, do not wire `cleanup()` to delete it.
## ZipService (kind `zip`)
- `process` materializes a project subtree via `project_files.export_to_dir` (staging under `config.DATA_DIR/zip_staging/{uid}`), compresses it in a **subprocess** (`python -m devplacepy.services.jobs.zip_worker`, stdlib-only, prints stats JSON incl. crc32), names the output `{crc32}.{slug}.zip` under `config.DATA_DIR/zips/{uid[-2:]}/{uid[-4:-2]}/` (sharded on the random tail of the uuid7 via `attachments._directory_for`, NOT its time-ordered head), and removes staging.
- Runtime artifacts live in `DATA_DIR` (`data/` by default), OUTSIDE the package and NOT under `/static`; the download is served by the `/zips/{uid}/download` route via `FileResponse`, not the static mount.
- Enqueue endpoints own authz: `POST /projects/{slug}/zip`, `POST /projects/{slug}/files/zip?path=`.
- `GET /zips/{uid}` (status `ZipJobOut`) and `GET /zips/{uid}/download` (FileResponse, extends expiry via `touch_job`) are **capability URLs** scoped only by the unguessable uuid7, not by owner - the owner is stored for attribution, not access control, matching publicly-viewable projects.
- Frontend `app.zipDownloader` (`static/js/ZipDownloader.js`) auto-wires any `data-zip-download` element plus the files context menu: POST -> poll `/zips/{uid}` via the shared `JobPoller` -> trigger download.
- Devii tools `zip_project`/`zip_status`; CLI `devplace zips prune|clear`.
- Disposable: `cleanup()` deletes the archive and the retention sweep prunes both the file and the `jobs` row.
## Planning report generator (kind `planning`, `services/jobs/planning_service.py`)
`PlanningReportService` is **admin-only** and builds a complete, phased markdown implementation document from a selectable set of open Gitea tickets, intended to be handed straight to a coding agent for one-shot execution, off the request path via the same async-job pattern as zip.
- `process` collects open tickets via `services/gitea/planning.py` `collect_open_issues(client, limit=MAX_ISSUES)` (the paginated `list_issues(state="open", limit=50)` loop, cap 50; reused by the admin page too), narrows them to the **selected ticket numbers** carried in the job payload (`{"numbers": [int, ...]}`, preserving selection order - an empty/absent list keeps all open, so the no-arg Devii action and any legacy enqueue stay backward compatible), and builds the markdown via `generate_plan(issues, config)` (an AI-or-fallback helper mirroring `enhance.py`).
- **AI path:** passes each ticket's FULL, verbatim description (`_body`, capped only at the `BODY_MAX=40000` per-ticket safety limit, not the old 600-char excerpt) to the internal gateway via `stealth.stealth_async_client` with a high output budget (`MAX_TOKENS=32000`, `PLAN_MAX=600000` final cap, `PLANNING_TIMEOUT_SECONDS=600` client timeout) and a `SYSTEM_PROMPT` that demands a self-contained document: a `## Execution Order` list, then `## Phase K: <name>` headings, and under each phase a `### #N <title>` subsection per ticket carrying labelled **Original ticket** (verbatim blockquote), **Goal**, **Dependencies**, **Affected areas / files**, **Implementation steps**, **Acceptance criteria**, and **Risks / open questions** blocks - preserving every detail with nothing summarised away.
- **Because an LLM can still summarise or truncate, the AI document is never trusted to be complete on its own:** `generate_plan` always appends a deterministic `# Appendix: Source Tickets (verbatim)` section built straight from the issue dicts by `planning.py` `verbatim_tickets(issues)` (per ticket `## #N <title>`, labels, the `html_url` source link, and the full body as a blockquote), so every covered ticket's full text is guaranteed present inline AND in the appendix - the document is genuinely self-contained.
- Fail-soft `_fallback` (when `ai_enhance` is off or any `httpx.HTTPError`/`ValueError`/`KeyError`/`IndexError` occurs) groups by primary label then ticket number, emitting the same `## Execution Order` + `## Phase K` shape with each ticket's full description reproduced verbatim (already self-contained, so no extra appendix).
- Writes the markdown to `config.PLANNING_REPORTS_DIR/_directory_for(uid)/{crc32}.{slug}.md` (sharded on the uuid7 random tail, crc32 over the markdown bytes), and returns `{download_url, local_path, final_name, markdown, ai_used, item_count, bytes_out}`. The `markdown` string is carried in the result so the status route can hand it to the renderer without re-reading disk. `cleanup` unlinks the file (retention prunes only the artifact + job row).
- Audit `issue.planning.request` (enqueue) and `issue.planning.generate` (success/failure) via `record`/`record_system` with an `audit.job(uid)` link.
- Routes (all `require_admin`): `POST /issues/planning` (enqueue, `503` when Gitea unconfigured, `PlanningForm` with a comma-separated `numbers` field parsed to `list[int]` and stored in the payload, returns `{uid, status_url}`), `GET /issues/planning/{uid}` (`PlanningJobOut`: status, `markdown`, `download_url`, `issue_count`, `ai_used`), `GET /issues/planning/{uid}/download` (`FileResponse`, `text/markdown`, traversal-guarded against `PLANNING_REPORTS_DIR`, extends retention via `touch_job`). The `/issues/{number}` detail route uses the `{number:int}` path converter so `/issues/planning` is never captured by it.
- Admin entry point: `GET /admin/issues/planning` (`routers/admin/issues.py`, `admin_issues_planning.html`) fetches the open tickets server-side via `collect_open_issues` (wrapped so a Gitea/network failure renders an empty state, never a 500) and renders a **selectable checkbox list** (all checked by default, a select-all/none master, and a live selected count) as the step in between, plus the **Generate planning** button, the `JobPoller` status panel, the `<dp-content>` render target, and a **Download** anchor; the issues listing shows an admin-only **Generate planning** link (`viewer_is_admin` in `issues_page`).
- Frontend driver `static/js/PlanningGenerator.js` (`app.planningGenerator`): tracks the checkbox selection (master toggle, count, disables Generate at zero selected), sends the chosen numbers as the `numbers` form field, POST -> `JobPoller.run` (widened to `intervalMs:2000, maxAttempts:400` = ~800s so it outlives the longer detailed generation) -> replace the `<dp-content>` with a fresh one holding `status.markdown` so it re-renders, set the download href.
- Devii tool `planning_report_generate` (`requires_admin=True`). The renderer's built-in copy button (the `dp-content` **Copy** button styled in `markdown.css`, identical to the per-code-block copy button) satisfies the copy-source requirement and is opt-out via the `no-copy` boolean attribute on `<dp-content>`.
## Project fork - ForkService (kind `fork`, `services/jobs/fork_service.py`)
Forking copies a source project into a brand-new project owned by the forking user, off the request path via the same async-job pattern as the zip flow.
- **Enqueue:** `POST /projects/{slug}/fork` (`routers/projects/index.py`) - `require_user` then `can_view_project(source, user)` (any logged-in user may fork any project they can view: public projects and their own private ones). Body is `ForkForm{title}` (the destination project name). It enqueues a `fork` job (`{source_project_uid, title, forked_by_uid}`, owner `("user", uid)`) and returns `{uid, status_url}`. No work happens on the request path.
- **`process`** looks up the source project row and the forking user row (raises -> job `failed` if either is missing), creates the destination project with `content.create_content_item("projects", "project", user, fields, ...)` (so slug/XP/logging match a normal create; metadata - `description`, `project_type`, `platforms`, `status`, dates, `is_private` - is copied from the source, `read_only` reset to 0), copies the whole virtual FS off-thread (`export_to_dir(source, "", staging)` -> `import_from_dir(new_uid, staging, user, skip_names=set())`, an **exact** copy including binaries, staging under `config.DATA_DIR/fork_staging/{uid}`), then records the relation with `database.record_fork(source_uid, new_uid, forked_by_uid)`. Result: `{project_uid, project_url, source_project_uid, item_count}`.
- **Rollback:** any failure after the project is created triggers `_rollback(new_uid)` (`project_files.delete_all_project_files` + `delete_fork_relations` + delete the `projects` row) before re-raising, so a failed fork leaves no orphan project.
- **Cleanup is a no-op on the project.** Unlike zip's disposable archive, a fork's artifact is a **permanent project** that must outlive the job. `cleanup()` only removes any leftover staging dir; the retention sweep deletes the job tracking row, never the forked project. `devplace forks prune|clear` likewise deletes job rows only.
- **Relation model:** the `project_forks` table (`uid`, `source_project_uid`, `forked_project_uid`, `forked_by_uid`, `created_at`) records direction explicitly (source -> forked), indexed on both project columns. Helpers in `database.py`: `record_fork`, `get_fork_parent(forked_uid)` (the source project row, for the "Forked from X" link on the project detail page), `count_forks(source_uid)`, and `delete_fork_relations(project_uid)` (called on project delete in `content.delete_content_item`, and in fork rollback).
- **Frontend:** `app.projectForker` (`static/js/ProjectForker.js`) wires `data-fork-project` (a Fork button on the project detail page, visible to any logged-in user): prompt for a name via `app.dialog.prompt` -> `Http.sendForm` POST -> `JobPoller.run("/forks/{uid}", ...)` -> on `done` redirect to `project_url`. The forked project's detail page shows a "Forked from X" link (`database.get_fork_parent`).
- **Status route** `GET /forks/{uid}` (`routers/forks.py`, `ForkJobOut`) exposes `project_uid`/`project_url`/`source_project_uid` only once `status == done`. Devii tools `fork_project`/`fork_status`; docs `projects-fork`/`forks-status`.
## SEO Diagnostics tool - SeoService (kind `seo`, `services/jobs/seo/`, `routers/tools/`)
The public **Tools -> SEO Diagnostics** auditor crawls a URL or sitemap with a headless browser and runs a broad battery of SEO checks, on the **same async-job pattern as zip/fork** plus a live websocket. It is **public (guests included)**; abuse is bounded by the per-IP POST rate limit, a per-owner one-active-job cap, a page cap, and the shared SSRF guard.
- **Surface:** a collapsible **Tools** dropdown in `base.html` (desktop center nav + a mobile section, visible to everyone) toggled by `MobileNav.initToolsDropdown`. `GET /tools` lists tools; `GET /tools/seo` is the auditor page (`static/js/SeoDiagnostics.js` -> `app.seoDiagnostics`, instantiated page-side in the template, not in `Application.js`).
- **Enqueue:** `POST /tools/seo/run` (`routers/tools/seo.py`, body `SeoRunForm{url, mode: url|sitemap, max_pages 1-50}`). Owner is `("user", uid)` or `("guest", X-Real-IP)`. It rejects with `429` if the owner already has a pending/running `seo` job, then enqueues `{url, mode, max_pages, allow_private:False}` and returns `{uid, status_url, ws_url}`.
- **`process`** writes the payload to `config.SEO_REPORTS_DIR/{uid}/payload.json`, launches `python -m devplacepy.services.jobs.seo.worker <payload_json> <output_dir>` via `create_subprocess_exec` (high `limit=` so big lines never overflow the StreamReader), reads **NDJSON frames from stdout** line by line (stage/target/progress/page/site_checks/report_ready), forwards each into the in-process **`ProgressHub`** (`services/jobs/seo/progress.py`, uid -> set of `asyncio.Queue`), and on completion loads `output_dir/report.json` as the job result. `cleanup()` clears the hub buffer and removes the report dir.
- **Worker** (`worker.py`, subprocess): `crawler.crawl_target` resolves the target (single URL, or sitemap `<loc>` URLs capped at `max_pages`) and fetches `robots.txt`/`sitemap.xml`/`llms.txt` with `httpx`. The audited target host is guarded once with `net_guard.guard_public_url`; candidate URLs **sharing that host are pre-approved** (no redundant per-URL `getaddrinfo` - a transient DNS failure or a self-hosted server resolving its own domain must not blank the whole crawl), and only cross-host sitemap entries are re-guarded. **In sitemap mode the crawler never falls back to auditing the sitemap document itself**: if no page URLs survive it raises a clear error (a stray `or [target]` fallback previously rendered the sitemap XML as one 56k-node "page" with no title/H1). For each page it launches one Playwright navigation: a single `page.evaluate(EXTRACT_SCRIPT)` returns the whole DOM contract (title/metas/canonical/headings/images/links/jsonld/og/twitter/semantic/mixed-content/word-count), an injected `PerformanceObserver` (`add_init_script(INIT_SCRIPT)`) captures LCP/CLS, navigation timing gives TTFB/FCP/transfer/protocol, a mobile-viewport pass measures overflow/tap-targets, a screenshot is saved, and a raw `httpx` GET supplies the SSR HTML for the rendered-vs-server parity check.
- **Check registry** (`checks/`): `base.py` defines the `Check`/`PageContext`/`SiteContext` dataclasses and the `@page_check`/`@site_check` decorators (collected into `PAGE_CHECKS`/`SITE_CHECKS`); one module per category (`crawl`, `meta`, `headings`, `links`, `structured_data`, `social`, `performance`, `mobile_a11y`, `security`, `ai_readiness`, `crosspage`). `registry.run_page_checks`/`run_site_checks` run them defensively (one failing check never aborts a page), and `compute_score` produces a severity-weighted overall score + grade and per-category subscores. **To add a check:** write a function decorated `@page_check`/`@site_check` in the right category module and import that module in `registry.py`.
- **Live progress WS:** `WS /tools/seo/{uid}/ws` is served **only by the service-lock owner** (closes `4013` for a fast retry on a non-owner worker, like `/devii/ws`); it replays `hub.snapshot(uid)` then streams `hub.subscribe(uid)`. The frontend `SeoProgressSocket` mirrors the `DeviiSocket` 4013/reconnect pattern. **The `done` frame carries the full report inline** (and the router's late-join terminal branch reads it from `job.result`); the client renders from `frame.report` directly. This is load-bearing: `service.process` publishes `done` from inside `process()`, BEFORE the JobService base persists the result on the next reap tick (~2s later), so a client that fetched `/tools/seo/{uid}/report` on the `done` signal would race the DB write and get an empty (all-zero) report. **Do not "simplify" this back into a fetch-on-done.** After publishing `done`, `process()` calls `hub.clear(uid)` to bound the in-memory buffer (late reconnects fall back to the DB-backed terminal branch).
- **Status/report routes:** `GET /tools/seo/{uid}` (`SeoJobOut`), `GET /tools/seo/{uid}/report` (`respond(..., SeoReportOut)`, HTML or JSON), `GET /tools/seo/{uid}/screenshot/{n}` (FileResponse from `SEO_REPORTS_DIR`, path-guarded). All are **capability URLs** scoped by the unguessable uuid7. The full report is written to `config.SEO_REPORTS_DIR/{uid}/report.json` (NOT inlined on a stdout line - avoids the StreamReader limit). Devii tools `seo_diagnostics`/`seo_status`/`seo_report` (public); docs `tools-seo-*`; CLI `devplace seo prune|clear`; audit `seo.run.request|complete|failed` (category `tools`).
- **SSRF guard is shared:** `devplacepy/net_guard.py` (`guard_public_url`, `is_blocked_address`, `effective_address`) was extracted from the Devii fetch controller, which now imports it; the crawler reuses it. **`playwright` is a core dependency** (Chromium installed by `make install` / the Docker image).
- Disposable: `cleanup()` removes the report dir; retention prunes both the artifacts and the `jobs` row.
- **Production nginx needs a dedicated WS location.** In `nginx/nginx.conf.template` the catch-all `location /` sets `Connection ""` (no upgrade) and a 60s timeout, so any websocket route that falls through to it fails the handshake (browser: `WebSocket connection failed`, no close code). The progress socket has its own `location ~ ^/tools/seo/[^/]+/ws$` block forwarding `Upgrade`/`Connection` with a 1h timeout, mirroring `/devii/ws`. **Any new websocket path must add its own upgrade `location` above `location /`** (see docs `Production -> nginx`).
## SEO metadata service - SeoMetaService (kind `seo_meta`, `services/jobs/seo_meta_service.py`, `services/seo_meta.py`, `seo_meta_text.py`)
`SeoMetaService` generates a clean, SEO-optimized title/description/keywords for every published content item (types `post`, `project`, `gist`, `news`, `issue`) off the request path and **meters its own AI spend**. It is a distinct concern from the public **Tools -> SEO Diagnostics** auditor (kind `seo`) - do NOT conflate the two kinds; they share only the `jobs` queue table. It is the **constructive counterpart** to the diagnostics tool: diagnostics audits a URL, this one populates the on-page metadata. It is a `JobService` like `ForkService`: the artifact (the `seo_metadata` row) is **permanent**, so `cleanup()` is a no-op and the retention sweep removes only the job tracking row.
- **Tables.** `seo_metadata` is polymorphic and soft-deletable (in `SOFT_DELETE_TABLES`, born-live `deleted_at`/`deleted_by`): `uid, target_type, target_uid, seo_title, seo_description, seo_keywords, status (ready|pending|failed), source (ai|plain), generated_at, created_at, updated_at`, keyed UNIQUE on `(target_type, target_uid)`. `init_db()` ensures every column, the UNIQUE `idx_seo_metadata_target` and the live `idx_seo_metadata_status (status, deleted_at)` index. `seo_usage` is a single-row config-like usage table (NOT soft-delete) mirroring `news_usage`, keyed `SEO_USAGE_KEY="seo_meta"`. Helpers in `database.py`: `get_seo_metadata`/`get_seo_metadata_batch`/`has_fresh_seo_metadata`/`upsert_seo_metadata`/`mark_seo_metadata_stale` (every read filters `deleted_at IS NULL`; `get_seo_metadata` returns only `status="ready"` live rows) and `add_seo_usage`/`get_seo_usage`.
- **Choke helper.** `services/seo_meta.py` `schedule_seo_meta(target_type, uid, regenerate=False)` and `schedule_seo_meta_for_table(table, uid, ...)` are import-cycle-free (only `database` + `queue`). They no-op for unknown types, missing uid, or (without `regenerate`) when a fresh `ready` row exists (`database.has_fresh_seo_metadata`); `regenerate=True` marks the row stale first (`database.mark_seo_metadata_stale`); both skip a target with an existing pending/running `seo_meta` job; otherwise `queue.enqueue("seo_meta", {target_type, target_uid}, "system", "seo_meta")`. Hooked at `content.create_content_item` (create, no-op guard) and `content.edit_content_item` (regenerate), the news publish sites in `services/news.py` (`status=="published"` only; existing-row update path uses `regenerate=True`), and `IssueCreateService` after the Gitea ticket is recorded. Because the work is async via the queue (NOT `run_in_executor`), the helper only enqueues.
- **`process`** loads the target row (posts/projects/gists/news via `get_table`, issues via `gitea.store.get_ticket`), builds grounding via `services/ai_context.build_context` (fail-soft for news/empty `user_uid`), and calls the gateway **off-thread** (`asyncio.to_thread(correction.gateway_complete, internal_gateway_key(), system, source_text, timeout)` - the synchronous gateway call posts to the in-process gateway on localhost, so it MUST run via `to_thread` or it self-deadlocks the single worker - the same lesson as `correction.py`'s sync mode). It demands strict JSON `{seo_title, seo_description, seo_keywords}`, parses fail-soft, re-clamps every field server-side via `seo_meta_text.clamp_generated`, and falls back to `seo_meta_text.plain_seo_defaults` (status `failed`, source `plain`) on any failure - **the fields are never empty**. Usage accumulates via `correction.new_usage_totals` and flushes once with `database.add_seo_usage(totals)` when `calls>0` (the single-row `seo_usage` table, mirroring `news_usage`). It emits an audit `seo.meta.generate`/`seo.meta.failed` (`record_system`, category `tools`) and `upsert_seo_metadata(...)`.
- **Backfill.** `run_once` calls `super().run_once()` then a bounded backfill sweep (gated by `seo_meta_backfill_enabled`, `seo_meta_backfill_batch` per type per tick) over published content lacking a fresh `ready` row, so pre-existing items get metadata with no one-shot migration.
- **Admin surface.** `collect_metrics()` merges the `JobService` job-pipeline stats with `usage_metric_cards(get_seo_usage())`, so the **SEO Metadata** card on `/admin/services` shows both the live task pipeline and the AI cost/averages; the existing `admin.services.{name}` pub/sub topic + `live_view_relay` row pushes it live with NO new VIEWS row. A standard-paginated task list reads `queue.list_jobs(kind="seo_meta")` with `database.build_pagination` + `_pagination.html` when a dedicated page is desired.
- **Clamps (single source of truth, `seo_meta_text.py`):** `seo_title` hard cap 60 (word-boundary, single hyphen, keyword front-loaded); `seo_description` hard cap 160 (word-boundary via `seo.truncate`, key message in the first 120 chars); `seo_keywords` 5-8 distinct lowercase comma-joined terms (the `<meta keywords>` tag is dead for Google but the feature mandates it - emit a SHORT honest list, never stuffed). `plain_text_from_markdown` reuses `rendering._render_content` + `utils.strip_html` so markdown (and em-dash) never leaks.
- **SEO consumption fix (`seo.py` `base_seo_context`).** The meta description is now markdown-stripped (`plain_markdown`/`plain_text_from_markdown`, fixing the prior raw-markdown leak), `meta_keywords` is emitted (a safe plain string, NOT a Jinja-global name), and a `seo_target=(target_type, target_uid)` param makes it consume the ready `seo_metadata` row when present, else `plain_seo_defaults`. `base_seo_context`'s new `keywords`/`seo_target` params are OPTIONAL with safe defaults so existing callers are unaffected; the five detail routers (posts/projects/gists/news/issues) pass `seo_target`. `og_title`/`twitter:title` use the bare `seo_title` (drops the redundant " - DevPlace" suffix in social cards); `base.html` adds `<meta name="keywords">`, `og:image:width/height/alt` and `twitter:image:alt`. The per-type JSON-LD (`discussion_forum_posting`, `software_application_schema`, `news_article_schema`, `software_source_code_schema`) route their text/description through `plain_markdown`.
- **Fan-out.** Schema `SeoMetaOut`; read route `GET /tools/seo-meta/{target_type}/{target_uid}` (`routers/tools/index.py`, public, JSON) returning the ready row or a plain default with status `pending`; Devii action `seo_meta_status` (public, read-only) + docs `tools-seo-meta-status`; CLI `devplace seo-meta prune|clear` (job rows only; the metadata persists); events `seo.meta.generate|failed`. Registered in `main.py` alongside `SeoService`. **A new `JobService` needs a server restart to go live.**
## DeepSearch tool - DeepsearchService (kind `deepsearch`, `services/jobs/deepsearch/`, `services/deepsearch/`, `routers/tools/deepsearch.py`)
The public **Tools -> DeepSearch** researcher is a multi-agent deep web researcher built on the **same async-job + ProgressHub + 4013-WS pattern as the SEO tool**, plus a per-session vector store and a grounded RAG chat. It is **public (guests included)**; abuse is bounded by the per-IP POST rate limit, a per-owner one-active-job cap, a page cap (1-30), depth cap (1-4), and the shared SSRF guard. Reuse the SEO tool as the template for any new Tools async job.
- **Owner helper is shared:** `routers/tools/_shared.py` `owner_for(request)` returns `("user", uid)` or `("guest", X-Real-IP)`; both `seo.py` and `deepsearch.py` import it (do not re-inline the owner derivation).
- **Enqueue:** `POST /tools/deepsearch/run` (body `DeepsearchRunForm{query, depth 1-4, max_pages 1-30}`). It rejects with `429` if the owner already has a pending/running `deepsearch` job. It resolves the **logged-in user's `users.api_key`** (guests use `database.internal_gateway_key()`) into the job payload for per-user embedding/LLM spend attribution, generates the uid up front, writes a `deepsearch_sessions` row (`create_deepsearch_session`), enqueues the job carrying `{query, depth, max_pages, api_key, collection}`, and returns `{uid, status_url, ws_url}`. The enqueue uses a local `_enqueue` (not `queue.enqueue`) so the session uid and the job uid match.
- **`process`** writes `control.json` (state `running`) + `payload.json` (augmented with the cross-session `cached_hashes`) under `config.DEEPSEARCH_DIR/{uid}`, launches `python -m devplacepy.services.jobs.deepsearch.worker <payload_json> <output_dir>` via `create_subprocess_exec` (high `limit=`), pumps **NDJSON stdout frames** into the in-process `ProgressHub` (`progress.py`), and on completion loads `output_dir/report.json`, persists the URL cache (`upsert_deepsearch_url_cache`), and updates the session row. `cleanup()` **drops the ChromaDB collection** (`VectorStore.drop`) and removes the session dir. Disposable: the collection + report dir are both deleted, unlike Fork/isslop.
- **Worker pipeline** (`worker.py`, stdlib + httpx + playwright, importable subprocess): `enhance.plan_queries` (gateway -> JSON sub-queries, deterministic fallback) -> `crawl.search_queries` (rsearch via standalone httpx, never `PlatformClient`; per-query result buckets are **round-robin interleaved** so every planned angle contributes pages, never just the first query) -> `crawl.crawl` (batches of `CRAWL_CONCURRENCY` concurrent httpx fetches then playwright render fallback, `guard_public_url` on the URL and every redirect, content-hash + URL-hash dedup; **`depth` follows in-page links**: after each level the links of every crawled page are scored by query-token overlap via `extract.relevant_links` and the top `LINKS_PER_PAGE` unseen ones form the next level, `depth=1` disables following) -> `chunking.chunk_text` -> `embeddings.embed_texts` (gateway, **local hashing fallback** when unavailable) -> `store.VectorStore.add` (Chroma) -> `orchestrate.orchestrate` (retrieval-grounded agents, see below). The worker writes `report.json` and `url_cache.json` and emits a `report_ready` frame carrying `synthesis`.
- **Search-provided content is a first-class source (`crawl.py`, the second junk-report fix).** `search_queries` calls rsearch with `content=true`, so each candidate carries the search engine's own readable `content`/`description` extract. This matters because the top sources for many questions are **bot-hostile** (X/Twitter, YouTube, Reddit, Facebook, Instagram, LinkedIn, TikTok - `HOSTILE_DOMAINS`): a headless fetch of those hits a login/consent wall ("Before you continue to YouTube", "Sign in to X") and yields near-zero text, which is why a 12-source run used to collapse to ~14 chunks. Now `crawl._resolve_candidate` **skips the fetch entirely for a hostile domain and uses the rsearch snippet** (`_snippet_page`, `source="search"`, `SNIPPET_MIN_CHARS` floor), and for every other domain it fetches normally but keeps the rsearch snippet as a **floor** (uses whichever of crawl-text vs snippet is longer), so a walled or thin page still contributes its real content instead of being dropped. This alone took a query from "cannot be answered" to a correct cited answer (14 -> 61 chunks, diversity 0.333 -> 0.75). **Never revert `content=true` and never send a headless render at a `HOSTILE_DOMAINS` host.**
- **Content extraction (`extract.py`, stdlib only):** `extract_html(raw, base_url)` is a readability-grade `HTMLParser` extractor used by both the httpx and playwright fetch paths (the old naive regex tag-stripper produced nav/cookie-banner boilerplate as "content" - the historic root cause of junk reports). It skips `script/style/nav/header/footer/aside/form` and ARIA `role=navigation|banner|contentinfo|...` regions, prefers `<article>`/`<main>` when they carry at least `MIN_CONTENT_TOTAL` chars, drops link-dense blocks (`MAX_LINK_DENSITY`, menus) and sub-`MIN_BLOCK_CHARS` fragments, unescapes entities, and emits real paragraphs joined by blank lines - which also makes `chunking.chunk_text`'s paragraph split actually fire (the flattened text used to be sliced mid-sentence). It also returns the page's `(url, anchor_text)` links (absolute, deduped, nav links excluded) for depth crawling; `relevant_links(links, query, limit)` scores them by query-token overlap and filters non-document extensions.
- **No "gaps"/critic agent (removed - do not reintroduce).** DeepSearch used to run a fourth "critic" agent that produced an "Open gaps" list. It was removed end-to-end (orchestrate, worker report, `DeepsearchSessionOut`, router context, session template, markdown/HTML export, `DeepsearchTool.js` agent labels, CSS) because it routinely emitted misleading, self-contradictory gaps on correctly-cited reports (the historic cause was that the critic was fed a truncated, retrieval-ordered source slice that dropped cited source numbers out of its window, so it fabricated "citation [n] is not in the report / no evidence provided"). **Do NOT reintroduce a `gaps` field or a critic agent.** The `_numbered_source_digest(pages)` helper survives and is used by the **linker** - a compact per-source `[n] title (url)\nexcerpt` block for EVERY page where the header line is always emitted even when the excerpt is trimmed, so all source numbers `1..N` are guaranteed present (never pass a raw `context[:N]` slice to an agent that reasons about source numbers). Regression: `tests/unit/services/jobs/deepsearch/orchestrate.py::test_linker_receives_full_source_list` / `::test_numbered_source_digest_keeps_every_source_number_under_cap`.
- **Orchestration (`orchestrate.py`) is retrieval-grounded and markdown-first.** The worker indexes BEFORE analysis and passes the `VectorStore` + planned queries to `orchestrate`, which embeds the question and each sub-query and pulls `hybrid_search` top chunks (round-robin merged, up to `CONTEXT_CHUNKS_MAX`), building the context from the RETRIEVED passages grouped per source - the source numbers `[n]` align with the report's `sources` list (page order), so inline citations, finding citations, and the rendered numbered source list agree. Page-head excerpts are only the fallback when retrieval is empty. Synthesis is **two-step to avoid the markdown-inside-JSON trap**: the summarizer writes a plain markdown report (`REPORT_MAX_TOKENS`, retried once if empty), then a separate `extractor` agent returns the findings JSON (retried once, tolerant `_parse_json` handles code fences and trailing garbage); the linker (confidence) failure is caught and never discards the report. The three agents in the pipeline today are summarizer, extractor, and linker - there is no fourth "critic" stage. Only a failed/empty report falls back to `_heuristic`, which stamps `synthesis="heuristic"` and emits a `status:"failed"` agent frame - the degradation is VISIBLE: `report.synthesis` flows through `DeepsearchSessionOut.synthesis`, the session template renders a "Degraded report" banner (`.ds-degraded`), and the markdown export carries the same note. A successful run stamps `synthesis="agents"`. **Never re-inline synthesis into a single JSON blob and never let a synthesis failure ship silently.**
- **PDF ingestion (`crawl.py` `fetch_page` + `pdf.py`):** a crawled candidate is treated as a PDF when its `content-type` is `application/pdf`/`application/x-pdf`, its URL path ends in `.pdf`, or its first bytes match the `%PDF-` magic (`pdf.is_pdf`). `fetch_page` streams the body and caps it at `MAX_PDF_BYTES` (15 MB); for a PDF it calls `pdf.extract_pdf_text`, which writes the raw binary to a `tempfile` temp location (cleaned up via `Path.unlink(missing_ok=True)` in `finally`), parses it with `pypdf` (`PdfReader`, capped at `MAX_PDF_PAGES` = 50, title pulled from metadata), and normalizes whitespace. The resulting `CrawledPage` carries `source="pdf"`; PDFs skip the Playwright fallback. Everything downstream is source-agnostic (chunking/embedding/orchestration read `page.text`/`page.source` unchanged), so no other module changes. New unpinned dep `pypdf` (pure-python, no system deps).
- **Pause/resume/cancel:** `POST /tools/deepsearch/{uid}/{pause|resume|cancel}` (owner-gated) rewrite `control.json`; the worker's `should_stop` callback polls it between source fetches (paused = sleep-loop, cancelled = stop). State lives in a file, not the job row, so the running-in-a-subprocess worker can read it without a DB round-trip.
- **Vector store (`services/deepsearch/store.py`):** `VectorStore` wraps `chromadb.PersistentClient(path=config.DEEPSEARCH_CHROMA_DIR)`, one collection per session (`ds_<uid>`). `Chunk` is the dataclass. `hybrid_search` blends cosine vector similarity with a BM25 keyword score (weights `HYBRID_VECTOR_WEIGHT`/`HYBRID_KEYWORD_WEIGHT`) over the candidate set, with optional metadata `where` filters. `embeddings.py` `embed_texts` calls the gateway embeddings endpoint and **falls back to a deterministic local hashing vector** on any failure (so the tool degrades, never breaks).
- **RAG chat (`services/deepsearch/chat.py` + `WS /tools/deepsearch/{uid}/chat`):** a dedicated lightweight loop (NOT the Devii hub), served **only by the service-lock owner** (closes `4013` for fast retry). Answers are grounded ONLY in the session collection via `hybrid_search`, cited inline, rendered client-side via `dp-content`. Turns persist to `deepsearch_messages` and audit `deepsearch.chat`. Frontend component `<dp-deepsearch-chat>` (`static/js/components/AppDeepsearchChat.js`) clones `AppDocsChat`'s framing but uses its own WebSocket to the chat path.
- **Status/report/export routes:** `GET /tools/deepsearch/{uid}` (`DeepsearchJobOut`), `GET /tools/deepsearch/{uid}/session` (`respond(..., DeepsearchSessionOut)`, HTML or JSON), `GET /tools/deepsearch/{uid}/export.{md,json,pdf}` (`services/deepsearch/export.py`; PDF via weasyprint). All are **capability URLs** scoped by the unguessable uuid7. **Viewer-flag discipline:** the session schema/context use `viewer_is_admin`/`viewer_owns` (never `is_admin`/`owns`) so a `respond()` context key never shadows a Jinja global (the same class of issue as the issues `/{number}` route). `tests/api/tools/deepsearch/session.py` guards the HTML render.
- **Completion race (load-bearing read-path fix).** The worker writes `report.json` to disk and `service.process` publishes the `done` frame **from inside `process()`**, but the `JobService` framework only commits `jobs.result`/`status=DONE` afterwards, in `_reap()` -> `_finish_done()` on a later tick. The frontend navigates to the session page the instant it receives `done`, so a read that keyed only off `jobs.status == DONE` returned an EMPTY report (`None` score, 0 sources) until a manual refresh. Fix: `_report_for(uid, job)` returns `job.result.report` when the job is `DONE` and non-empty, else falls back to the on-disk `report.json` (`_report_from_disk`, `DEEPSEARCH_DIR/{uid}/report.json`) - which exists before the `done` frame is ever sent - and returns `{}` only for a `FAILED` job or a genuinely still-running job with no report on disk. `_session_context` derives `done`/`status` from `bool(report)` (not raw job status), and the chat WS gate accepts `session.status == "done"` (set inside `process()` before the publish) as ready. `_export_report` reuses the same fallback. Regression: `tests/api/tools/deepsearch/session.py::test_session_reads_disk_report_before_result_commit`. **Any new read of a job result that a client reaches immediately after a `done`/`session_url` frame must use this same on-disk fallback, never bare `jobs.status`.**
- **Clickable inline citations (`services/deepsearch/citations.py`).** The report/findings carry `[n]` markers (and the model sometimes emits `[3][9][1-2]`); the `link_citations(html, source_count)` template global (registered in `templating.py`) rewrites each `[n]` and each `[a-b]` range into `<a class="ds-cite" href="#ds-source-n">[n]</a>` anchors that jump to the numbered `<li id="ds-source-n">` in the Sources list (source numbering is page order, matching the `[n]` the summarizer was given). It splits out `<a>`/`<code>`/`<pre>` regions first so markers inside links/code are left alone, expands ranges to individual links, and drops out-of-range numbers (no broken anchors). The session template nests it over the server render: `{{ link_citations(render_content(summary), sources|length) }}` and `{{ link_citations(finding.detail|e, sources|length) }}`, plus a per-finding `.ds-finding-cites` chip row from `finding.citations`. `.ds-cite`/`.ds-sources li:target` styling lives in `deepsearch.css`. The report prompt asks for one number per bracket (never a range) so output is consistent, but the linkifier handles ranges regardless. Regression: `tests/unit/services/deepsearch/citations.py`.
- **Tables (`deepsearch_sessions`, `deepsearch_messages` soft-deletable + in `SOFT_DELETE_TABLES`; `deepsearch_url_cache` GC-only):** columns are ensured in `init_db()` (every queried column) with indexes. Every insert writes `deleted_at:None/deleted_by:None`; every read filters `deleted_at IS NULL`.
- **Frontend** (do not hand-roll): `DeepsearchTool.js` (`app.deepsearchTool`) drives the form via `Http.send`, watches `DeepsearchProgressSocket` (cloned from `SeoProgressSocket`, 4013 retry), and wires pause/resume/cancel. `static/css/deepsearch.css` uses the design tokens and is mobile-responsive.
- **Progress frame protocol (append-only, the worker<->JS contract):** `phases.py` is the single source of truth for phase identity, shared by `worker.py` (emit) and `DeepsearchTool.js` (render). `PHASE_ORDER = [planning, searching, crawling, indexing, analysis, synthesis]`; `worker._stage(stage, message, phase)` emits BOTH the legacy `stage` frame (byte-identical to before) AND a parallel `phase` frame `{phase, index, total, label}` so the timeline strip advances. The first emitted frame carries `version:1`. Every other frame type and its keys: `substep` (planning angles, `phase`+`message`; also emitted by analysis grounding), `queries`, `candidates`, `rsearch`, `progress` (`done`/`total`/`url`/`depth`), `page_loaded` (now `source`/`render`/`depth`/`elapsed_ms`/`done`/`total`), `page_cached`/`page_skipped`/`page_duplicate` (now `reason`/`elapsed_ms`), `embed_batch` (`batch`/`total_batches`/`backend`/`done`/`total`, emitted before AND after each batch), `embed_done` (`backend`/`chunk_count`), `agent` (`agent` one of summarizer|extractor|linker, `stage`/`status` start|done|failed, with `elapsed_ms`/`tokens_in`/`tokens_out` on done), `report_ready` (now also `synthesis`), `done` (`session_url`), `failed` (`message`). **The contract is append-only: never rename or drop a frame type**; `service._run_worker` pumps every stdout line into the `ProgressHub` untouched, so new frame types reach the WS with no handler change. `tests/api/tools/deepsearch/index.py` is the append-only regression guard.
- **RAG-audit hardening (engine correctness, no route/schema change):** (1) **Embedding-dimension consistency** - gateway embeddings carry provider-native dims while `local_embed` is fixed 256-dim; the worker `_index_chunks` now decides the backend ONCE per job (the first gateway failure or non-gateway result forces local for ALL remaining batches), and `VectorStore.add` drops any vector whose length differs from the collection's established dim, so one collection never mixes dims (cosine search across mixed dims is corrupt). `embeddings.EmbedResult.dims` and `VectorStore.dims` (lazily probed from the collection) expose the dimension; `chat.retrieve` re-embeds the query locally and skips retrieval if it still cannot match the stored dim. (2) **Citation grounding** - `orchestrate` drops any finding with no citations, and the summarizer prompt forbids uncited claims and treats the QUESTION as data not an instruction (prompt-injection reduction via `_sanitize_question`); `chat._strip_unmatched_markers` removes any inline `[n]` marker that does not map to an emitted citation. (3) **Confidence calibration** - when only a single domain was crawled, `orchestrate` caps confidence by the source-diversity-derived ceiling (overconfidence guard). (4) `EmbeddingCache` is bounded at `EMBED_CACHE_MAX` to cap in-memory growth.
- Devii tools `deepsearch`/`deepsearch_status`/`deepsearch_session` (public); docs `tools-deepsearch`; CLI `devplace deepsearch prune|clear`; audit `deepsearch.run.request|complete|failed` + `deepsearch.chat` (category `tools`). **New dependencies:** `chromadb`, `weasyprint`, `pypdf` (all unpinned). New runtime dirs `config.DEEPSEARCH_DIR`/`DEEPSEARCH_CHROMA_DIR` are registered in `DATA_PATHS`. **Add a dedicated nginx WS `location` for `/tools/deepsearch/{uid}/ws` and `/chat`** above `location /` for production, like the SEO and Devii sockets.
## AI Usage Analyzer tool - IsslopService (kind `isslop`, `services/jobs/isslop/`, `routers/tools/isslop.py`)
The public **Tools -> AI Usage Analyzer** classifies a git repository or website as AI slop, sophisticated AI-assisted work or genuine human work. It is built on the standard async-job pattern (a `JobService` running a subprocess worker), but its live channel is **pub/sub, not a dedicated WS route**: every worker event is published to `public.isslop.{uid}` AND persisted to `isslop_events`, and the frontend pairs the pub/sub subscription with an incremental `GET /tools/isslop/{uid}/events?after=SEQ` poll, so guests (who cannot subscribe to `public.*` unless `pubsub_allow_guests` is on) and reconnecting tabs replay from the durable trail. **Never rely on pub/sub alone for this tool: the DB event trail is the source of truth, pub/sub is the fast path.**
- **Engine layout:** `services/jobs/isslop/` holds `acquisition/` (git probe via `git ls-remote`, depth-1 clone with size preflight + live 3 GB kill guard, stealth Playwright website crawler with HTTP fallback, path-traversal-safe workspace helpers), `analysis/` (exclusion rules, stylometric metrics, language detection, per-repo baselines, `signals/` with one detector family per file, two-axis scoring), `agent/` (gateway LLM client, per-file classifier, vision reviewer, report writer with deterministic fallback), plus `pipeline.py` (the event-yielding run), `worker.py` (subprocess entry), `events.py` (frame protocol), `persistence.py` (`EventPersister` writes events/file results/image results/report and stamps the analysis row), `store.py` (all DB access), `badge.py` (SVG), `service.py` (`IsslopService`), `config.py` (all constants + `WorkerSettings`).
- **Worker contract:** `IsslopService.process` writes the worker payload (url + admin toggles + gateway endpoint/model/key) to `config.ISSLOP_RUNS_DIR/{uid}/payload.json`, resolves the workspace under `config.ISSLOP_WORKSPACES_DIR` (`workspace_for` rejects any path escaping the root), launches `python -m devplacepy.services.jobs.isslop.worker <payload> <workspace>`, and relays each NDJSON stdout line through `EventPersister.apply` (SQLite) then `pubsub.publish`. The workspace and run dir are removed in a `finally`; the pipeline also deletes the workspace itself as its final act, so **no acquired source survives an analysis** - only the report and its evidence rows.
- **AI through the gateway only:** `agent/llm.py` talks solely to `config.INTERNAL_GATEWAY_URL` with model `molodetz` and `database.internal_gateway_key()` (vision uses the same model - the gateway handles image parts). `review_available`/`vision_available` gate the AI and image stages; on any gateway failure the static engine remains authoritative and the report falls back to the deterministic composer. All HTTP (gateway, git size preflight, website fallback crawl) goes through `stealth_async_client`.
- **Artifacts are permanent, the job row is not.** Like `ForkService`, `cleanup()` never touches `isslop_analyses`/`isslop_events`/`isslop_file_results`/`isslop_image_results`/`isslop_reports` - the report and badge are public capability URLs meant to outlive the run; the retention sweep removes only the `jobs` tracking row. `devplace isslop clear` is the only bulk hard-delete (plus per-analysis `store.purge_analysis`).
- **Ownership and guest history sync:** the owner is `("user", uid)` or `("guest", DEVII_GUEST_COOKIE)` - NOT the tools `_shared.owner_for` IP fallback, because history must survive IP changes and be claimable. The page/list/run handlers mint the guest cookie when absent (same cookie as Devii/customization, one guest identity platform-wide). `_sync_guest_history` runs on page and list requests: when a signed-in user still carries a guest cookie, `store.claim_guest_analyses` re-owns those rows via UPDATE (a move, never a copy - no duplicate data). One active analysis per owner (`429` otherwise, audited `denied`).
- **Tables:** `isslop_analyses` is soft-deletable (in `SOFT_DELETE_TABLES`, born-live inserts, reads filter `deleted_at IS NULL`, indexed on `(owner_kind, owner_id, created_at)`/`status`/`content_hash`); the evidence tables (`isslop_events` keyed `(analysis_uid, seq)`, `isslop_file_results`, `isslop_image_results`, `isslop_reports` UNIQUE on `analysis_uid`) are GC-only evidence purged with their analysis. All ensured in `init_db()`.
- **Routes** (all under `/tools/isslop`, capability URLs): `GET ""` page, `POST /run` (`IsslopRunForm`, http/git/ssh URL pattern), `GET /list` (owner history), `GET /{uid}` (`IsslopAnalysisOut`), `GET /{uid}/events` (ordered replay), `GET /{uid}/report` (`respond(..., IsslopReportOut)` - HTML shows the live `<dp-isslop-run>` while running and the server-rendered report when completed; the markdown body goes through `render_content`), `GET /{uid}/report.md`, `GET /{uid}/badge.svg` (self-contained SVG, hardcoded colors by design - it must render on external sites). Badge/report URLs are absolute via `seo.site_url`.
- **Frontend:** two site-wide web components (`static/js/components/AppIsslop.js` `<dp-isslop>` = submit form + history list; `AppIsslopRun.js` `<dp-isslop-run>` = live progress feed), registered in `components/index.js`, light DOM, reusing `Http`/`Poller` and `app.pubsub`. Page CSS `static/css/isslop.css` (design tokens). On `done` the run component reloads the page so the report is the server-rendered (SEO/`render_content`) version, never a client re-render.
- **Verdict blending is per-file, never a global mean (load-bearing).** The AI review pass runs its per-file gateway calls concurrently (`pipeline.AI_REVIEW_CONCURRENCY`, semaphore-bounded like the image pass) and adjusts ONLY the files it actually reviewed: `pipeline.apply_ai_verdicts` blends each sampled file's static origin/quality with its own verdict (0.6/0.4), then the WHOLE repo is re-aggregated with the normal SLOC/criticality weights, and `scoring.ai_fraction` maps per-file origin scores through a smooth 35-65 ramp (never the old hard 45/55 buckets). **Never reintroduce a repo-level mean of the 12 sampled verdicts** - it hands a tiny sample a fixed 40% of the verdict, so the LLM's clustered hedging values (30/40/50) drown thousands of files of static evidence and unrelated projects converge on identical percentages (the real-world twin-84%-human defect). **Single source of truth for the final verdict:** the SCORE event, the DONE payload and `generate_report` all consume the SAME final `RepoScores` (image influence applied via `scoring.adjust_for_images`, which recomputes slop/grade/human together) - the summary grade and the report body grade can therefore never disagree; `tests/unit/services/jobs/isslop/pipeline.py` guards both invariants.
- **Image evidence thumbnails + retry-safe evidence.** Workspaces die with the run, so the vision stage persists an aspect-preserving WebP thumbnail per reviewed image (`vision.make_thumbnail`, sha1-of-relative-path name) into `config.ISSLOP_MEDIA_DIR/{uid}` (the dir comes to the worker via the payload `media_dir`); the `thumb` name rides the `image` event, is stored on `isslop_image_results`, and is served by `GET /tools/isslop/{uid}/media/{name}` (strict `^[a-f0-9]{16}\.webp$` name pattern + `is_relative_to` root check - never loosen either). The report page renders the thumbnails with `data-lightbox` (the shared `app.lightbox` opens them full-size) and the live feed shows a tiny inline preview per image event. **`store.reset_evidence(uid)` runs at the top of every `IsslopService.process`** - a retried job (orphan recovery) previously re-inserted its events/file/image rows, duplicating every image and file in the report; any new evidence table MUST be added to `reset_evidence` AND `purge_analysis`.
- **Template provenance (the "ships defaults" detector).** `analysis/templates.py` `detect_template(workspace)` scores starter-template evidence repo-wide (it reads files the inventory excludes, like `package.json`): known template slugs/authors in the manifest (+ `ct3aMetadata`), README template marketing (weights capped so a wordy README cannot alone confirm), and the kitchen-sink scaffold constellation (count of standard scaffold artifacts past 3 freebies). The saturating score feeds `scoring.adjust_for_template` - a no-op below 35, a 0.7x floor on ai_percent/origin when confident, 0.85x when >= 70 - applied LAST in the pipeline scoring stage, after the AI and image blends. Near-certain evidence (>= 70) also FORCES category `ai-slop` (defaults shipped as-is are slop by the canonical definition - clean scaffold code never earns an untouched template `sophisticated-ai`, whose meaning is 'the presenter decided'); the confident band caps a `human-*` category at `uncertain`. Calibration truth set (guarded by tests): the six stock boilerplates (ixartz x2, create-t3-app, vercel ai-chatbot x2, fullstack-nextjs-app-template) grade C-D with markers listed, while devplacepy itself scores 0.0 with zero markers - tune weights against BOTH sides, never only the slop set. The evidence rides the `signal` inventory event, the SCORE payload (`template_score`/`template_markers`) and the report's Template Provenance section; admin toggle `isslop_template_detection`.
- **Rendered-DOM signal family (`analysis/domsignals/`).** A homepage-only, live-browser companion to the text-based `signals/webtells.py` checks: `acquisition/browser.py`'s `StealthBrowser.capture()` loads the page and returns the actual rendered DOM (computed styles, a class-name census, meta tags, headings/landmarks, console warnings, network response hosts/headers, a screenshot) via `DOM_EXTRACT_SCRIPT` in `acquisition/domcapture.py`. `domsignals/base.py` defines its own `@dom_check`/`@dom_site_check` decorator registry, mirroring the SEO job's `checks/` `@page_check`/`@site_check` mechanics, but emitting isslop's own `Signal` type (not a new one). Eight category modules (`builders`, `color`, `typography`, `layout`, `copy`, `metaseo`, `accessibility`, `buildsignals`) contribute 49 distinct signal codes; `domsignals/aggregate.py` `aggregate_dom_evidence` runs every registered check over the captured page(s) and saturates the weighted total into a `DomEvidence` (score/bucket/signals/detected_builder/builder_confidence). This brings the engine's total to twenty-one detector families (13 text-based in `signals/`, 8 rendered-DOM in `domsignals/`) and 126 signal codes (77 text-based, 49 rendered-DOM) - update these counts again the next time a family is added or removed, never leave them stale.
- **Homepage-only by design (`config.DOM_ANALYSIS_MAX_PAGES = 1`).** Capturing a full DOM, console, network trail and screenshot on every crawled page would multiply the browser cost of a run, so the pass runs ONLY against the first page at crawl depth 0 (`acquisition/website.py` gates `depth == 0 and index < DOM_ANALYSIS_MAX_PAGES`). Bump the constant later if the tool needs multi-page DOM coverage - `DomSiteContext`/`dom_site_check` (e.g. `detect_duplicate_meta_description`) already support more than one page. Git-repository sources never populate `dom_snapshots` (`dom_sink` is threaded only through `crawl_website`, never `clone_repository`), and a run where the browser could not be launched simply yields an empty page list, so `aggregate_dom_evidence([])` is a clean no-op (`DomEvidence(score=0.0, bucket="none")`).
- **Scoring order is load-bearing: images -> DOM -> template, never reorder.** `pipeline.py` applies `scoring.adjust_for_images`, then `adjust_for_dom_signals`, then `adjust_for_template`, in that exact sequence, and `adjust_for_template` MUST stay last. `adjust_for_dom_signals` blends `DomEvidence.score` into `ai_percent` at a small, deliberately cautious weight (`config.DOM_AI_WEIGHT = 0.12`, since this whole signal family is new and uncalibrated) UNLESS a confident builder match (`builder_confidence >= DOM_BUILDER_CONFIDENT_THRESHOLD`) forces the category to `ai-slop` via the same `_force_slop_category` helper the template detector uses. Because `adjust_for_template` runs after it, a template match can still floor/force the category further; nothing may run after `adjust_for_template`, since a later step would silently undo a forced `ai-slop` category.
- **DOM evidence never attaches to a per-file `FileScore`/`FileContext` (do not bolt it on).** `DomEvidence` is repo/page-level evidence blended once into the final `RepoScores`, exactly like template provenance and the image mean - never distributed across individual files. A rendered page has no SLOC and no line numbers to weight against, and its evidence (computed styles, a screenshot, console/network output) is not textual, so it cannot be scored, sampled or displayed through the SLOC-weighted per-file model the rest of the engine uses. Any new DOM check emits page-level `Signal`s into `DomEvidence.signals`, never into a `FileScore.signals` list.
- **`isslop_dom_results` follows the same evidence-table obligation as every other isslop table.** It mirrors `isslop_image_results` (`store.py` `TABLE_DOM_RESULTS`) and is already wired into both `store.reset_evidence(uid)` and `store.purge_analysis(uid)`; any future evidence table added under `domsignals/` must be added to both the same way.
- **DEP_UNRESOLVED is alias-aware (do not regress).** The JS/TS unresolved-import detector (`signals/hallucination.py`) flags imports that match no package.json dependency (all four sections), Node builtin or local path - i.e. phantom/hallucinated dependencies. It MUST skip everything that is not an npm specifier: relative/absolute/URL imports, `node:` and any `scheme:` specifier, the non-package prefixes `@/`, `~`, `#`, `$` (tsconfig/subpath/Svelte aliases - none are valid npm names), and every prefix parsed from `tsconfig.json`/`jsconfig.json` `compilerOptions.paths` (`engine._javascript_alias_prefixes`, carried on `RepoContext.javascript_alias_prefixes`). tsconfig is JSONC and alias keys contain `/*`, so comments are stripped with the string-aware scanner `_strip_jsonc_comments` - NEVER a comment regex (a regex eats the `"@/*"` alias strings themselves; this was a real defect). Before the alias awareness the detector flagged nearly every file of a standard Next.js app; after, only genuine phantom deps remain. `tests/unit/services/jobs/isslop/hallucination.py` guards it.
- **Clickable source references (annotated source viewer).** The static stage persists the FULL source of every signal-bearing file (cap `SOURCE_CAP_FILES`=60 files / 200KB each) into the same per-analysis media dir (`_persist_source`, `s<sha1>.txt`); the name rides the `file` event and the `isslop_file_results.source` column. `GET /tools/isslop/{uid}/source?path=...&line=N` renders `isslop_source.html`: server-rendered line table (Jinja autoescape covers XSS) with line-number anchors `#LN`, signal lines highlighted with inline annotation chips, a focused line, and a findings nav in the sidebar; the strict `^s[a-f0-9]{16}\.txt$` + `is_relative_to` checks mirror the media route. Everything referencing a file links there: the file-results table path, each signal chip (`&line=N#LN`), and the report prose - `_linkify_sources` rewrites backticked paths in the report markdown into links BEFORE `render_content`, and the reporter system prompt requires the model to backtick every path it mentions. `reset_evidence`/`purge_analysis` already sweep the media dir, so sources share the thumbnail lifecycle.
- **No absolute paths ever leave the engine:** every user-facing path is workspace-relative (`relative_to(workspace)`); keep it that way in new detectors/events.
- **Docs heading anchors + `.docs-toc` (platform-wide):** `docs_prose._render_markdown` post-processes every prose page, stamping a slugified `id` on each `h2`/`h3` (`heading_slug`, GFM-style, deduplicated with `-N` suffixes) and appending a hover-visible `.docs-heading-anchor` permalink; `scroll-margin-top` keeps targets below the topnav. The `isslop-checks` page uses this for its clickable Contents grid: a `.docs-toc` nav placed OUTSIDE the `data-render` block (raw HTML passes through untouched) whose `href="#slug"` values are computed with the SAME `heading_slug` function at generation time - reuse `.docs-toc`/`.docs-toc-item`/`.docs-toc-count` (styled in `docs.css`) for any other long docs page, and never hand-write a slug that `heading_slug` would not produce. `tests/unit/docs_prose.py` guards the slugging and injection.
- Devii tools `isslop`/`isslop_status`/`isslop_report`/`isslop_list` (member-only, `requires_auth=True` per policy - the HTTP surface stays public); docs `tools-isslop` + `isslop-checks` (the full plain-language check catalog); CLI `devplace isslop analyze|prune|clear`; audit `isslop.run.request|complete|failed` (category `tools`); achievement key `isslop` ("Slop Hunter"). New unpinned dep `playwright-stealth`; runtime dirs `config.ISSLOP_DIR`/`ISSLOP_WORKSPACES_DIR`/`ISSLOP_RUNS_DIR` in `DATA_PATHS`.
+49
View File
@@ -0,0 +1,49 @@
# Messaging (`devplacepy/services/messaging/`)
This file documents the real-time direct-message chat subsystem. Claude Code auto-loads it when a file under `devplacepy/services/messaging/` is read or edited.
## Overview
The `/messages` feature is a live WebSocket chat layered over the SAME `messages` table and audit/notification cores as before; there is no parallel chat store. The WS only adds live delivery, presence, typing, and read receipts on top of the existing persistence.
`messages.py` (the router, mounted at `/messages`) exposes: `GET ""` (page; the conversation sidebar is one windowed SQL aggregate and an opened thread loads only the newest `CONVERSATION_MESSAGE_LIMIT` = 500 messages - keep these queries bounded), `GET /search`, `POST /send` (no-JS fallback), and `WS /messages/ws` (NOT lock-owner gated - accepts on every worker; closes `1008` for guests).
## Bounded page queries (load-bearing performance rule)
`get_conversations` is ONE window-function query (`ROW_NUMBER() OVER (PARTITION BY other_uid ORDER BY created_at DESC, id DESC)` over `sender_uid = :me OR receiver_uid = :me`, the `get_recent_comments_by_target_uids` precedent) that returns only each partner's latest row - never load the user's full message history into Python to build the sidebar. `get_conversation_messages` loads at most `CONVERSATION_MESSAGE_LIMIT` (500) most-recent rows (`ORDER BY created_at DESC, id DESC LIMIT`, reversed to ascending in Python); older history stays in the DB and is simply not rendered. Both are covered by `idx_messages_conversation`/`_rev` (verified `EXPLAIN QUERY PLAN`, MULTI-INDEX OR, no table scan). `messages` is NOT in `SOFT_DELETE_TABLES`, so neither query filters `deleted_at`. Any new message listing must stay bounded and indexed the same way.
## WS endpoint `/messages/ws`
Lives on the existing messages router (no new router/prefix; already mounted at `/messages`). It is **NOT lock-owner gated** (deliberately - unlike `/devii/ws`): `await websocket.accept()`, resolve the owner via session cookie / api-key (`_resolve_ws_user` reuses `utils._user_from_session`/`_user_from_api_key`), and accept on EVERY worker. Gating on `service_manager.owns_lock()` was the original design and was wrong here: when no worker holds the service lock (services disabled, or the lock not yet acquired) every socket got closed `4013` and the client fast-retried every 200ms forever (a connection storm with no stream, and the send form fell back to a full-page POST reload). Chat must not depend on the background-service lock. Messaging requires an authenticated user - guests are closed with `1008`. The receive loop is wrapped in `try/except WebSocketDisconnect` + a broad `except` (never crash the worker) and ALWAYS unregisters the socket in `finally`.
## Connection registry
The `ConnectionManager` singleton `message_hub` (`services/messaging/hub.py`), **one per worker**: `user_uid -> set[WebSocket]`, plus an in-process `last_seen` map and a bounded `delivered` LRU set (`mark_delivered`/`was_delivered`, cap 4000) used to dedupe direct vs relayed delivery. `register`/`unregister` return whether the user just came online / went offline (for presence announces). `send_to_user`/`send_to_users` fan a frame out to all of a user's sockets (multi-tab) and swallow dead-socket errors. Presence and last-seen are in-process per worker (no DB column).
## Cross-worker delivery (the relay)
Because the WS accepts on every worker, a message persisted on worker A must still reach a recipient whose socket lives on worker B. `services/messaging/relay.py` `message_relay` (a per-worker singleton asyncio loop, started on the first socket connect, self-stops when `message_hub.has_connections()` is false) polls `SELECT * FROM messages WHERE id > :watermark` (uuid7-backed autoincrement `id`) every ~1s and pushes any row whose sender/receiver has a LOCAL socket and that was not already `was_delivered` to those local sockets, advancing the watermark to the max row seen. Same-worker sends are delivered INSTANTLY by `broadcast_message` (which calls `message_hub.mark_delivered` so the relay skips them); the relay only fills the cross-worker gap, so same-worker latency is zero and cross-worker latency is bounded by the poll interval. The relay primes its watermark to the current `MAX(id)` on first start so it never replays history. **Trade-off:** typing/read/presence frames are in-process per worker only - across the two `make prod` workers those ephemeral signals reach only same-worker peers (message delivery is always correct). On a single dev worker everything is instant and complete.
## DRY persist choke point
`services/messaging/persist.py` `persist_message(sender, receiver_uid, content, attachment_uids, *, request=None, origin)` is the ONE function that inserts the row, links attachments, fires `create_notification` + `clear_messages_cache` + `create_mention_notifications`, logs, and writes the `message.send` audit event. Both the WS `send` handler and the HTTP `POST /messages/send` handler call it, so audit/notification/mention behavior is byte-identical on both paths (the Devii `send_message` action is `handler="http"` -> `POST /messages/send`, so it also flows through here and broadcasts live). Content is capped at 2000 chars server-side. When `request` is a `Request` it audits via `audit.record`; on the WS path it now passes `request=websocket` (auditing via `audit.record_system` with `actor_kind="user"` + `origin="websocket"`, same `message.send` key/category - no new event invented; presence/typing/read are ephemeral and NOT audited).
## AI correction and AI modifier apply to direct messages, with LIVE delivery of the final content
`"messages": ("content",)` is in the `CORRECTABLE_FIELDS` registry and `persist_message` invokes `schedule_correction`/`schedule_modification` (sender = the user), so typing `@ai <instruction>` in a DM runs the AI modifier (default on + sync) and an enabled correction rewrites the content. Both send paths funnel through `routers/messages._finalize_and_broadcast(sender, message, request, client_id=None)`: it awaits any pending SYNC correction/modification futures stashed on `request.scope[PENDING_SCOPE_KEY]` (the same `PENDING_SCOPE_KEY`/`loop.run_in_executor` mechanism the HTTP `await_pending_corrections` middleware uses), RE-READS the message row, and broadcasts the FINAL (corrected/modified) content via `broadcast_message`. The HTTP `POST /send` and the WS `/ws` send path both call it; the WS path passes `request=websocket` to `persist_message` so the modifier stashes its executor future on the websocket scope and the handler awaits it directly (HTTP middleware does not run for websockets). A message with no `@ai` directive and correction off has nothing pending, so the broadcast is immediate with zero added latency. Frontend `static/js/MessagesLayout.js` reconciles the sender's own echo by `client_id` and replaces the optimistic bubble's `.rendered-content` with the server's authoritative `frame.content`, re-rendered - so the sender sees their `@ai` directive resolved in-place (and corrections become visible to the sender too). See "AI content correction" and "AI modifier".
## WS protocol frames
Client -> server: `{type:"send", receiver_uid, content, attachment_uids?, client_id}`, `{type:"typing", receiver_uid}` (throttled client-side), `{type:"read", with_uid}`, `{type:"presence", with_uid}` (one-shot presence query). Server -> client: `{type:"ready", user_uid}` (sent first on connect), `{type:"message", uid, sender_uid, sender_username, receiver_uid, content, created_at, time_ago, client_id}` (broadcast to BOTH sender and receiver sockets so multi-tab and the sender's own optimistic bubble reconcile via the echoed `client_id`), `{type:"typing", from_uid}` (to the receiver only), `{type:"read", by_uid}` (read-receipt to the other user), `{type:"presence", user_uid, online, last_seen}` (announced to everyone in an open conversation with the user on connect/disconnect, and as a direct reply to a `presence` query). The WS `read` path reuses `mark_conversation_read` (the same `UPDATE messages SET read=1` SQL the page uses) and `clear_messages_cache`.
## Renderer integration (XSS control)
Live/echoed message bubbles are NEVER injected as raw HTML. `MessagesLayout.buildBubble` creates a `<div class="rendered-content" data-render>`, sets its `textContent` to the message body, then runs `contentRenderer.applyTo(el)` + marks `img:not(.avatar-img)` with `data-lightbox` + `contentRenderer.highlightAll()` - exactly the `ContentEnhancer.initContentRenderer` sequence (emoji shortcode -> marked -> DOMPurify.sanitize -> highlight.js -> image/YouTube/autolink). This gives Discord-style emoji, image and YouTube embeds, autolinking, and sanitization for free.
## Frontend
`MessagesSocket.js` (mirrors `DeviiSocket.js`: reconnect with backoff, `4013` silent fast-retry, `onReady` fires on the first server frame not raw open) points at `/messages/ws`. `MessagesLayout.js` (instantiated on `app` in `Application.js`) opens the socket on the messages page, sends over WS with an optimistic pending bubble keyed by `client_id`, reconciles on the echoed `message` frame, appends incoming messages live through the renderer, auto-scrolls, throttles + shows the typing indicator, flips the read-receipt double-check, renders the presence dot / last-seen in the header, and live-updates the conversation-list preview + unread dot + ordering. If the socket is not open the send `<form method=POST>` submits normally (no-JS fallback). The template carries `data-self-uid`/`data-other-uid`/`data-other-online`/`data-other-last-seen` on `.messages-layout`. Styling (presence dot, typing dots, read-receipt, unread dot, pending opacity) is in `messages.css` using design tokens only, responsive down to 360px.
## Attachments stream live
The single `message_frame` builder (`services/messaging/persist.py`, used by BOTH `broadcast_message` and the relay) fetches `get_attachments("message", uid)` and includes a slimmed list (`uid`/`url`/`thumbnail_url`/`is_image`/`is_video`/`original_filename`/`file_size`/`mime_type`) on every frame. `MessagesLayout.renderAttachments` mirrors `_attachment_display.html` in the DOM (image -> `img.gallery-thumb` marked `data-lightbox`+`data-full`, video -> `<video>`, else a download link) - no `innerHTML`, so it stays XSS-safe. The sender's optimistic bubble shows an "Uploading attachment..." placeholder until its own echo arrives and swaps in the real gallery. **Attachment-only messages (empty caption) are allowed**: the content input dropped `required`, `sendViaSocket` sends when content OR attachments are present, and `persist_message` accepts empty content when `attachment_uids` is non-empty (`if not content and not attachment_uids: return None`). `dp-upload` (mode `attachment`) uploads each file to `/uploads/upload` and writes the uids into its hidden `attachment_uids` input; `collectAttachments()` reads that before the send, then `dp-upload.clear()` resets it. The chat uploader sets `hide-chips` so only the `(N)` count badge shows (no filename chips) - a generic `dp-upload` attribute. The send button shows a spinner and is disabled while busy: `MessagesLayout.refreshSendButton()` ORs `_uploading` (driven by the new `dp-upload:busy` event) with `_pendingSends` (a set of in-flight send `client_id`s, cleared on the echo or an 8s safety timeout), toggling `.is-sending` + `disabled` on `.messages-send-btn`.
+46
View File
@@ -0,0 +1,46 @@
# News service (`devplacepy/services/news/`)
This file documents the automated developer-news import pipeline. Claude Code auto-loads it when a file under `devplacepy/services/news/` is read or edited.
## Overview
`BaseService` provides the async run loop, a `deque(maxlen=20)` log buffer, and graceful cancellation. `ServiceManager` is a singleton that registers, starts, and stops services. `NewsService` (`services/news/service.py`, `default_enabled=True`, registered unconditionally in `main.py`) is a fully automatic, zero-maintenance import pipeline: it fetches articles from `news_api_url`, **cleans** each (HTML strip + `clean_news_text`/`JUNK_PATTERNS` Reddit-boilerplate removal), **fetches and perceptually compares the images** (SSRF-guarded fetch + Pillow decode + `imagehash` phash off-thread; placeholder = too small / undecodable / a phash shared across 2+ different articles), **grades deterministically** (AI grade on cleaned text via `news_ai_url` + a reliability gate + a unique-image bonus and thin-content penalty -> effective `grade`, raw in `ai_grade`), **reformats every valid article into clean Markdown** (`_format_article` + the editable `news_format_prompt`; paragraphs, `## ` headings, lists and code, preserving every fact; fail-soft to the cleaned original, toggle `news_format_enabled`), and inserts ALL of them into `news` with `status="published"`/`"draft"` on `news_grade_threshold` - nothing is silently skipped. After the loop it **auto-rotates Featured + Landing** (`_apply_landing_selection`, top scored unique-image articles in the recent window) while honouring per-row `featured_locked`/`landing_locked` set when an admin manually toggles. **AI usage is metered like the correction/modifier consumers**: each gateway call's `X-Gateway-*` response headers are parsed (`parse_usage_headers`) and the run's totals accumulated into the durable single-row `news_usage` table (`database.add_news_usage`/`get_news_usage`, same SUMs-plus-computed-averages shape as `correction_usage`); `NewsService.collect_metrics()` surfaces calls/tokens/cost and the per-call averages as `stats` on the admin-only `/admin/services` page. The full ruleset is on `NewsService.description` and surfaces on `/admin/services`, which also polls live status and the log tail.
## Per-run flow
The per-run flow per article is: clean text -> fetch and perceptually compare images -> AI-grade the cleaned text -> reliability gate -> compute an effective score -> AI-reformat the body into Markdown (valid articles only) -> publish/Feature decision -> store -> then a single post-loop Landing rotation.
- Fetches `GET {news_api_url}` -> `{"articles": [...]}`; already-synced `external_id`s (from `news_sync`) are skipped.
- Every article is graded via AI on the CLEANED text: `POST {news_ai_url}` with model `{news_ai_model}`, temperature 0. Both default to the internal gateway (`INTERNAL_GATEWAY_URL` / `molodetz`); the key falls back to `internal_gateway_key()` when `news_ai_key`/`NEWS_AI_KEY` is unset.
- ALL articles are inserted into `news` regardless of grade (never silently skipped). Articles re-synced each run (upsert by `external_id`); grade, status, image, and images updated each cycle. Slugs via `make_combined_slug(title, uid)`.
- All parameters (`news_api_url`, `news_ai_url`, `news_ai_model`, `news_grade_threshold`, `news_ai_key`, `news_format_enabled`, `news_format_prompt`, interval) are declared as `config_fields` and edited on the Services tab. The full grading ruleset is carried on `NewsService.description` (the `GRADING_RULES_DESCRIPTION` module string) and renders on `/admin/services`.
## AI reformatting (`_format_article`, `FORMAT_PROMPT_SPEC`)
After grading, every VALID article (one that passed the reliability gate and got a grade) has its body reformatted by the AI into clean Markdown - short paragraphs, `## ` section headings, bullet/numbered lists, and inline/fenced code - turning the source wall of text into a readable article. It reuses the grading endpoint/model/key (`news_ai_url`/`news_ai_model`/`_get_ai_key`) at `temperature 0.3`, `FORMAT_MAX_TOKENS=6000`, with the cleaned `description`+`content` (capped at `FORMAT_INPUT_MAX_CHARS=14000`) appended to the editable `news_format_prompt`. The prompt forbids inventing/removing facts and only restructures. The result is fence-stripped (`_strip_md_fence`), validated to be at least `MIN_BODY_CHARS`, capped at `FORMAT_OUTPUT_MAX_CHARS=30000`, and stored as the `news.content` (rendered server-side by `render_content`, the markdown engine, on `news_detail.html`). It is fail-soft: on any error, an empty/too-short result, or `news_format_enabled` off, the cleaned original content is stored unchanged. Toggle with the `news_format_enabled` bool config field.
## AI usage metering and stats (the shared `usage.py` helpers, `news_usage`, `collect_metrics`)
Service-level AI spend is metered exactly like the per-user `correction_usage`/`modifier_usage` consumers, and the accumulation/formatting plumbing is **shared, not copied per service**, in `services/openai_gateway/usage.py` beside `parse_usage_headers`: `USAGE_FIELDS` (the seven summed columns), `new_usage_totals()` (a zeroed totals dict), `accumulate_usage(totals, response)` (parses the `X-Gateway-*` headers via `parse_usage_headers` - defensive `getattr(response, "headers", None)`, so fake responses without headers are a no-op - and adds each field into the totals; `totals=None` is a no-op so the metering is opt-in per call site), and `usage_metric_cards(usage)` (the one stat-card list builder). For NewsService the grading and formatting gateway calls are its only AI spend: on every successful gateway response it calls `accumulate_usage(usage_totals, resp)` into a per-run dict from `new_usage_totals()`, and after the run `database.add_news_usage(usage_totals)` does ONE `INSERT ... ON CONFLICT DO UPDATE SET col = col + excluded.col` upsert into the durable single-row `news_usage` table (the same `_add_usage` helper as the correction/modifier tables, keyed on the constant `NEWS_USAGE_KEY="news"` in the `user_uid` column, ensured + unique-indexed in `init_db`). `database.get_news_usage()` (the same `_get_usage`) returns the running SUMs plus computed AVERAGES (avg tokens/call, avg cost/call, avg upstream latency, avg tokens/sec). `NewsService.collect_metrics()` is `{"stats": usage_metric_cards(get_news_usage())}`, which `BaseService._persist_state` serializes (every ~3s tick, regardless of the service's enabled state) and `ServiceMonitor.renderMetrics` shows on the admin-only `/admin/services` page - consistent with `BotsService` (live fleet cost) and the financial-data-is-admin-only rule (the services page is `require_admin`).
**The issue tracker is the second service consumer and reuses all of this 1:1:** its two AI call sites (`services/gitea/enhance.py` `enhance_ticket` for ticket filing and `services/gitea/planning.py` `generate_plan` for the planning document) take an optional `totals` dict and call `accumulate_usage(totals, response)` after `raise_for_status`; the owning job services (`IssueCreateService`, `PlanningReportService`) build a `new_usage_totals()`, pass it in, and `database.add_issue_usage(totals)` (constant `ISSUE_USAGE_KEY="issues"`, table `issue_usage`, ensured beside `news_usage`) when `totals["calls"]`. `IssueTrackerService.collect_metrics()` is `{"stats": usage_metric_cards(get_issue_usage())}`, so the aggregate ticket-AI spend shows on the Issue Tracker service page even though the poller defaults to disabled (the supervisor still ticks and persists metrics). To meter a NEW service AI call: thread a `new_usage_totals()` dict to the call site, `accumulate_usage` the response, persist with a matching `add_<x>_usage`/single-row `<x>_usage` table, and surface it via `collect_metrics` -> `usage_metric_cards`.
## Content cleaning (`clean_news_text`, `JUNK_PATTERNS`)
After `strip_html`, removes Reddit boilerplate (`submitted by /u/<user>`, `submitted by ... to /r/<sub>`, standalone `[link]`, `[comments]`, `[N comments]`, trailing `submitted by ... [link] [comments]`) and collapses whitespace. Applied to title (light), description, and content BEFORE grading and BEFORE storage.
## Image fetch + perceptual placeholder/uniqueness detection
Per article up to `IMG_PER_ARTICLE=5` candidate image URLs (from `_get_article_images`) are fetched through `net_guard.guarded_async_client` (SSRF-safe, re-guards redirects), capped at `IMG_FETCH_MAX_BYTES=5_000_000` and a short timeout; each is decoded with Pillow and phashed (`imagehash`) via `loop.run_in_executor` so the CPU work never blocks the loop, recording width/height. Reject as placeholder on decode/fetch failure or width/height `< MIN_IMAGE_DIMENSION=100`. Across the WHOLE batch, phashes within Hamming distance `PHASH_DISTANCE=5` (compared via `imagehash.hex_to_hash`) that span 2+ DIFFERENT articles are a shared logo/placeholder and are flagged placeholder for all of them. `has_unique_image` = at least one non-placeholder candidate; its URL becomes the article's `image_url`. Each `news_images` row stores `phash`/`width`/`height`/`is_placeholder`; the `news` row stores `image_url` and `has_unique_image`.
## Deterministic grading (constants at module top)
`MIN_TITLE_CHARS=12`, `MIN_BODY_CHARS=200`, `MAX_TITLE_CAPS_RATIO=0.6`, `MIN_IMAGE_DIMENSION=100`, `PHASH_DISTANCE=5`, `IMG_PER_ARTICLE=5`, `IMG_FETCH_MAX_BYTES=5_000_000`, `UNIQUE_IMAGE_BONUS=2`, `THIN_CONTENT_PENALTY=2`, `FEATURE_MIN_SCORE=8`, `LANDING_MIN_SCORE=9`, `LANDING_MAX=6`, `LANDING_RECENCY_DAYS=4`. Per article: (1) the raw AI grade (1-10) is stored in `ai_grade`; `None` is invalid. (2) The reliability gate (`reliability_reason`) forces draft (never Featured/Landing) when the url is empty/invalid, the cleaned title `< MIN_TITLE_CHARS`, the combined cleaned body `< MIN_BODY_CHARS`, or the title uppercase-letter ratio `> MAX_TITLE_CAPS_RATIO`; the reason is recorded in `news_sync.status` as `rejected_quality:<reason>` and in the audit metadata. (3) `effective_score = clamp(ai_grade + UNIQUE_IMAGE_BONUS if has_unique_image - THIN_CONTENT_PENALTY if body marginal-but-not-gated, 1..10)` (marginal = body `< MIN_BODY_CHARS * 2`); stored in `grade` (so `/news` ordering reflects final quality). (4) `status = "published"` when valid and `effective_score >= news_grade_threshold` else `draft`. (5) `featured = 1` when published AND `has_unique_image` AND `effective_score >= FEATURE_MIN_SCORE` (respecting the lock).
## Auto-rotating Featured + Landing with admin lock
After the per-article loop, `_apply_landing_selection` queries recent (`<= LANDING_RECENCY_DAYS`) rows that are published + featured + `has_unique_image` + NOT `landing_locked`, sorts by `(-grade, -synced_at)`, sets `show_on_landing=1` for the top `LANDING_MAX` with `grade >= LANDING_MIN_SCORE` and `show_on_landing=0` for the rest of THAT service-managed set (self-rotating). When an admin manually toggles Featured or Landing (`routers/admin/news.py`), the matching `featured_locked`/`landing_locked` is set to 1 so the service stops auto-managing that row; updates of a locked Featured row never overwrite `featured`. New `news` columns: `ai_grade`, `featured`, `has_unique_image`, `image_url`, `featured_locked`, `landing_locked`, `author`, `article_published` (all added to the `init_db` ensure-block - the stale-schema rule). Audit events `news.service.ingest`/`news.service.publish`/`news.service.draft`/`news.service.reject`/`news.service.landing` (`record_system`, category `news`). New dep: `imagehash` (`Pillow` already present).
## Editorial labels are admin-only on the public surfaces
The `&#x2605; Featured` badge and the `Grade {{ grade }}` chip are internal editorial signals, so on every PUBLIC news surface - the `/news` listing (`news.html`), the article page (`news_detail.html`), and the landing/home Developer News cards (`landing.html`) - both are wrapped in `{% if is_admin(user) %}`. Members and guests never see them; administrators still see them everywhere (including the public pages) and the admin **Manage News** area (`admin_news.html`) is unchanged, where the grade and featured state are the editorial controls. The article's source name and timestamp stay visible to everyone. Gating only the rendered labels does not touch the stored `grade`/`featured` columns, the grading/landing logic, or ordering (`/news` is still ordered by effective `grade`).
@@ -0,0 +1,90 @@
This file documents the AI gateway subsystem (devplacepy/services/openai_gateway/) - the single point of truth for AI calls, usage metering, and provider/model routing. Claude Code loads it automatically whenever a file under this directory is read or edited.
## Overview
`GatewayService` (`services/openai_gateway/`) is an OpenAI-compatible LLM gateway (the ported `openai5.py`), mounted at `/openai/v1/*` (`routers/openai_gateway.py`, prefix `/openai`). It is a **service that serves an HTTP endpoint** rather than a loop.
- The router is thin: it calls `service_manager.get_service("openai").handle(request, subpath)`. `POST /v1/chat/completions` runs the full gateway (vision augment -> model override -> forward -> optional SSE re-emit via `_fake_stream`); `POST /v1/embeddings` forwards to the configured embeddings upstream (`handle_embeddings`, no vision/streaming); `/v1/{path:path}` is a transparent passthrough to the upstream base. Disabled -> 503, unauthorized -> 401, upstream connection failure -> 502.
- `main.py` registers it and exempts `/openai` from the rate-limit middleware and the maintenance gate. No new dependency (`httpx` already required).
- `GatewayService` is `default_enabled=True` so internal callers work out of the box.
## Per-call cost/token response headers
Every gateway response (chat, embeddings, passthrough; success and error) carries `X-Gateway-*` headers describing that single call: `Model`, `Backend`, `Prompt-Tokens`, `Completion-Tokens`, `Total-Tokens`, `Cache-Hit-Tokens`, `Cache-Miss-Tokens`, `Reasoning-Tokens`, `Cost-USD`/`Input-Cost-USD`/`Output-Cost-USD` (dollars), `Cost-Native` (1 if the upstream returned a native cost), `Tokens-Per-Second`, `Upstream-Latency-Ms`, and `Context-Window`/`Context-Utilization` when the model's window is known, plus the timing headers `X-Gateway-Upstream-Latency-Ms`/`X-Gateway-Total-Latency-Ms`. `GatewayUsageLedger.record(...)` RETURNS the computed ledger row (or `None` on failure); each handler's `finalize` closure maps it through `usage.usage_response_headers(row)` and attaches it to the returned `Response` (`resp_headers`). The single denied path with no upstream call (embeddings disabled) carries no headers. This is what lets any caller read its own spend - the AI correction worker reads `X-Gateway-Cost-USD`/token headers off its own correction call to accumulate per-user totals (see "AI content correction" in the root `CLAUDE.md`).
## Per-worker runtime
**Serves in every worker.** Config/enabled come from `site_settings` (via the service's `get_config()`/`is_enabled()`), so any uvicorn worker answers - not just the supervisor worker. Per-worker runtime (`GatewayRuntime`: `httpx.AsyncClient` pool + `asyncio.Semaphore` sized to `gateway_instances`, plus counters and a `VisionCache`) is created lazily on first request and rebuilt when `instances`/`timeout`/cache size change. `gateway_instances` is the **scaling knob** (concurrency per worker); process scaling is uvicorn workers.
## Auth
`authorize()` reuses `get_current_user` (cookie/X-API-KEY/Bearer/Basic). Allowed if `gateway_require_auth` is off, the presented X-API-KEY/Bearer equals the static `gateway_access_key` **or the auto-generated `gateway_internal_key`**, or the resolved user is an admin (`gateway_allow_admins`) or any user (`gateway_allow_users`).
**Per-user attribution:** the gateway also accepts a real user's `api_key` (Bearer / X-API-KEY / session), gated by `gateway_allow_users` (default **on**); `resolve_owner()` records `(owner_kind, owner_id)` per call, so usage is attributed and limitable per user. With `gateway_allow_users` off, only internal/access/admin keys work and per-user attribution never happens. Devii operating a signed-in user authenticates its LLM calls with that user's own `api_key` (set as the session's `ai_key` in `build_settings(..., owner_kind="user")`), so a user's full gateway spend (Devii and direct API calls) rolls up under their uid - surfaced admin-only on the profile page via `build_user_usage()` / `GET /admin/users/{uid}/ai-usage`. Guests keep the internal key.
## Config fields
All config is `config_fields` (upstream url/model/key, force-model, the Prompt group's `gateway_system_preamble`, timeout, instances, vision url/model/key/cache/toggle, the Embeddings group (`gateway_embed_enabled`/`_url`/`_model`/`_key`), the auth toggles + static key + internal key, plus the Pricing/Reliability/Tracking groups); live metrics (requests/errors/in-flight/vision-calls/embed-calls/latency plus 24h cost/tokens/success rollups) via `collect_metrics`.
## System-message composition (date awareness + operator preamble)
`system_message.py` (`apply_system_directives`) runs in `handle_chat` only (NOT `handle_embeddings`/`handle_passthrough`), after vision augmentation and before the upstream payload is assembled, so it reaches every chat consumer (news, bots, Devii guests, per-user direct API calls). It composes ONE system message in a fixed, cache-friendly order: **operator preamble (`gateway_system_preamble`) -> date line (`Current date: DD/MM/YYYY`) -> the client's original system content**.
Two behaviors:
- **Date awareness (EU, no time).** The current date is injected only when the composed text (preamble + client system content) contains NO date in any common format. Detection lives in `contains_date` via the `_DATE_PATTERNS` regex set: `YYYY-MM-DD` (ISO), `D/M/YYYY` and `DD/MM/YYYY` (also matches `MM/DD/YYYY`), `DD-MM-YYYY`, `DD.MM.YYYY`, and textual months (`14 June 2026`, `June 14, 2026`, `14 Jun 2026`, with abbreviations and optional ordinal/comma). The date is `datetime.now().strftime("%d/%m/%Y")` - **date only, never time**: omitting the time keeps the system message byte-stable for the whole day so upstream prompt/token caching stays effective.
- **Operator preamble.** `gateway_system_preamble` (Prompt group, `type="text"` textarea, default empty) is prepended ahead of the client's system content on every chat call so all calls carry operator preferences. When the client sent NO system message, the composed content (preamble and/or date) becomes a brand-new leading system message; when the client DID send one, its content is replaced in place (preamble + date + original text). A blank/whitespace-only preamble with an already-dated request is a no-op (no empty system message is created). The date detection runs over the preamble too, so an operator preamble that already states a date suppresses the injected date line.
Embeddings/passthrough are untouched by this composition step.
## Usage, cost, latency, and reliability tracking
The gateway records one row per upstream call (chat, vision, passthrough) and surfaces per-hour and 24h analytics. The pieces:
- **`usage.py`** - `GatewayUsageLedger` writes to `gateway_usage_ledger` (one row per call, success and failure) and samples `gateway_concurrency_samples` each 30s tick. `normalize_usage` handles BOTH upstream shapes: DeepSeek (`prompt_cache_hit_tokens`/`prompt_cache_miss_tokens`, `completion_tokens_details.reasoning_tokens`) and OpenRouter (`prompt_tokens_details.cached_tokens`, native `cost`, `cost_details`). `compute_cost` PREFERS the upstream native `cost` (OpenRouter) and FALLS BACK to per-1M pricing for the chat backend (DeepSeek reports no cost); it also returns the input/output split (native cost is split by the modeled-rate share `input_cost / (input_cost + output_cost)` from the configured per-1M prices, not by token volume; a negative native cost is clamped to 0). `record()` is wrapped in try/except so a tracking failure never breaks the proxy.
- **`gateway.py`** - `_send` is the gated/timed/retried path: a `CircuitBreaker` (`reliability.py`) short-circuits when open; `retry_send` retries timeouts/connection errors/5xx with linear backoff; semaphore acquire time is the `queue_wait_ms`; an httpx per-request `trace` extension captures `connect_ms` (0 on pooled reuse). `handle_chat`/`handle_passthrough` record every outcome with timings, tokens, cost, params, and the resolved owner. Vision calls record from `VisionAugmenter._describe_one`.
- **Owner attribution** - `GatewayService.resolve_owner` maps the caller to `(owner_kind, owner_id)`: `internal:devii` (internal key), `key:access` (static key), `user:<uid>`/`admin:<uid>` (session or API key), else `anonymous`.
- **`analytics.py`** - `build_analytics(hours, top_n, pricing)` pulls the bounded window once (clamped to `MAX_WINDOW_HOURS=168`) and computes everything in Python: volume/throughput, tokens (sums + avg + p50/p90/p95/p99 + max via `reliability.percentile`), latency (upstream/overhead/queue/connect/total + tokens/sec), errors by category, cost (per model/caller, input vs output, projected monthly, caching savings), behavior, and an hourly breakdown. `summary_metrics()` runs cheap SQL aggregates for the live service panel. TTFT/inter-token are returned `null` (gateway forwards non-streaming).
- **Admin** - `/admin/ai-usage` (HTML, `templates/ai_usage.html` + `static/js/AiUsageMonitor.js`) and `/admin/ai-usage/data?hours=&top_n=` (JSON) in `routers/admin/` package, admin-gated. Schema `GatewayUsageOut`. Devii action `ai_usage` and docs `admin-ai-usage` expose the same endpoint. The `ai_usage` and `site_analytics` Devii catalog tools are both `requires_admin=True`.
- **Per-user usage** - `analytics.build_user_usage(owner_id, hours=24, pricing)` filters the ledger by `owner_id` (a uid matches both `user:` and `admin:` rows) and returns requests, success/error %, token totals, cost (window / per-hour / per-request / 30-day projection = 24h spend x 30), avg latency + tps, per-model breakdown, and an hourly cost series. Served admin-only at `GET /admin/users/{uid}/ai-usage` (schema `UserAiUsageOut`, docs `admin-user-ai-usage`) and rendered on the profile page by `static/js/UserAiUsage.js` (the `[data-ai-usage]` card, gated `{% if is_admin(user) %}` in `templates/profile.html`). The 24h window is a sound projection basis (unlike a startup burst), so `cost_24h * 30` is honest.
- **Profile visibility split** - `routers/profile/` package `_ai_quota(uid, include_cost=False)` builds the profile quota bar. **The displayed spend is the user's REAL total gateway spend, not the Devii-turn subtotal:** `spent_usd`/`used_pct`/`requests`/`last_used` come from `analytics.user_spend_24h(uid)`, the SAME `gateway_usage_ledger` source (summed by `owner_id`, matching both `user:` and `admin:` rows) that feeds the stat grid's `cost.window_usd`. So the server-rendered `spent_usd` is byte-for-byte the card's "Cost 24h" and the two halves of the card reconcile. `turns` is still the `devii_usage_ledger` `turns_24h` count (Devii conversation turns) and is labeled **"Devii turns"** in the template precisely because it is a *different metric* from the grid's total **API requests** (one Devii turn fans out into several gateway requests; direct API-key calls add requests with no turn, which is why `0 Devii turns` can sit beside `18 API requests`). The `limit`/`turns` still come from the Devii service (`daily_limit_for`/`turns_24h`); an **unlimited** limit (`<= 0`, the admin default) sets `unlimited=True` so the template renders "no cap" / "No daily spend cap" instead of a misleading `/ $0.00`. **This is display-only: the Devii daily-cap ENFORCEMENT and `/devii/usage` still read the Devii-scoped `devii_usage_ledger.spent_24h`** (per-turn, resettable, and the only per-guest spend record since guest Devii calls hit the gateway under the shared internal key, not the guest cookie). The profile page renders three ways: **admin** sees the full gateway stats card *plus* a server-rendered quota bar (with spend/limit/Devii turns); the **owner** (non-admin, `is_self`) sees a quota-**only** card showing just the percentage; everyone else sees nothing AI-related. `_ai_quota` is fail-safe (returns `None` on any error or when the Devii service is unregistered) and the result rides on `ProfileOut.ai_quota`. **Dollar figures (`spent_usd`/`limit_usd`) are included only when `include_cost=viewer_is_admin`** - never for a non-admin owner. This matters because the route serves the same context as JSON via `respond(..., model=ProfileOut)` and `ProfileOut.ai_quota` is a free-form dict, so hiding dollars in the HTML branch alone would still leak them to a member requesting their own profile with `Accept: application/json`; the server-side omission is the real control. Non-admins (HTML and JSON) get only `used_pct`/`turns`/`requests`.
- **Retention** - `run_once` prunes both tables past `gateway_usage_retention_hours` (default 720 = 30 days).
**Financial data is admin-only everywhere.** Any monetary figure (USD cost, pricing, spend, limit) is restricted to administrators; members and guests see only the percentage of quota used - this rule is enforced consistently across the profile card, `ai_correction`/`ai_modifier` usage displays, and the Devii cost tools.
## Embeddings
`POST /openai/v1/embeddings` exposes an OpenAI-compatible text-embeddings model. Clients send the generic model `molodetz~embed` (`config.INTERNAL_EMBED_MODEL`), which `handle_embeddings` remaps to `gateway_embed_model` exactly like chat remaps `molodetz` -> `gateway_model` (also remapped when `gateway_force_model` is on or the model is empty). It defaults to OpenRouter's `qwen/qwen3-embedding-8b` at `https://openrouter.ai/api/v1/embeddings` (`config.EMBED_*_DEFAULT`, $0.01 per 1M input tokens). `handle_embeddings` mirrors `handle_chat` but is simpler: no vision augmentation and no streaming - build the payload, forward via `_send`, and record one ledger row through the same `finalize(...)` closure. The config fields are the **Embeddings** group (`gateway_embed_enabled` default on, `gateway_embed_url`, `gateway_embed_model`, `gateway_embed_key`) plus the Pricing-group `gateway_embed_price_input_per_m`. `effective_config()` falls the embed key back to `gateway_vision_key` then `OPENROUTER_API_KEY` (NOT `gateway_api_key`: that is the DeepSeek chat upstream key, whereas embeddings target OpenRouter like vision does). Usage is recorded with **`backend="embed"`**; `usage.compute_cost` adds an `embed` branch (input-only, completion always 0, native OpenRouter `cost` still preferred) and `Pricing` gained `embed_input_per_m`. `analytics.py` groups by `backend` generically, so embed rows roll up automatically; `caching_savings` counts only **non-native** chat rows (native-priced rows did not use the configured cache-hit/miss rates, so folding them in would report a fictional saving). When `gateway_embed_enabled` is off the endpoint returns 503 with no ledger row.
## Single point of truth for AI
The gateway is the only place that holds real provider URLs/models/keys. Every other AI consumer (news, bots, Devii guests) points at it by default and never touches a provider key:
- Defaults live in `config.py`: `INTERNAL_GATEWAY_URL` (`http://localhost:10500/openai/v1/chat/completions`, override with `DEVPLACE_INTERNAL_BASE_URL`) and `INTERNAL_MODEL` (`molodetz`). `news_ai_url`, `bot_api_url`, `devii_ai_url` default to `INTERNAL_GATEWAY_URL`; their model defaults to `molodetz`.
- Each consumer's key falls back to `database.internal_gateway_key()` (reads the `gateway_internal_key` setting) when its own key field/env is unset - the provider-key fallbacks (`DEEPSEEK_API_KEY`/`OPENROUTER_API_KEY`) were removed from news and bots.
- `gateway_force_model` (default on) and a `molodetz`/empty alias in `handle_chat` make the upstream always receive `gateway_model`, so `molodetz` is a stable generic alias.
- `database.migrate_ai_gateway_settings()` (called at the end of `init_db()`, under the startup `init_lock`): generates `gateway_internal_key` (uuid4) if missing; migrates `DEEPSEEK_API_KEY`/`OPENROUTER_API_KEY` env into `gateway_api_key`/`gateway_vision_key` when the db value is empty; and rewrites any consumer AI URL still equal to the old `openai.app.molodetz.nl` default to the gateway, plus `bot_model` `deepseek-chat` -> `molodetz` (only uncustomized values).
- The gateway's `gateway_api_key`/`gateway_vision_key`/`gateway_internal_key` fields are **non-secret** so the admin services page shows the value actually in use, editable.
- Bot LLM calls are synchronous `urllib` but already run via `asyncio.to_thread` (`bot/bot.py`), so the local round-trip never blocks the event loop.
- Real provider keys/URLs/models live ONLY in the gateway. The gateway is `default_enabled=True`.
## Provider and model routing (`routing.py`, admin `/admin/gateway`)
Layered ON TOP of the single-provider service config above, which stays THE implicit `default` provider (env migration, internal key, `molodetz`/`molodetz~embed`, vision, embeddings, ledger all unchanged).
**Storage.** Two dataset tables, ensured in `init_db` via `routing.ensure_tables`, cross-worker cache-invalidated under the `"gateway_routing"` cache-version name (module-level `provider_store`/`model_store` over a shared `_ROUTING_CACHE`; writes `bump_cache_version` and clear the cache). They are admin config, NOT in `SOFT_DELETE_TABLES` - hard CRUD, mirroring `site_settings`:
- `gateway_providers` - named upstreams: `name` + `base_url` (chat-completions URL) + `api_key` + `is_active`; the embeddings URL is derived by swapping `/chat/completions` -> `/embeddings`.
- `gateway_models` - source->target routes: `source_model` (unique, what clients request) -> `provider` (blank = default) + `target_model` + `kind` (chat|embed) + optional `vision_provider`/`vision_model` for the **text+vision merge** (when set, image content is described by that vision model before forwarding) + `context_window` + its **own economy** (`price_cache_hit_per_m`/`price_cache_miss_per_m`/`price_output_per_m`/`price_input_per_m`, USD per 1M tokens) + `is_active`.
**Resolution is a per-request overlay, not a fork.** At request time `routing.chat_overlay(requested, cfg)` / `routing.embed_overlay(requested, cfg)` resolve an active route by the requested model name and return a per-request OVERLAY dict of `gateway_*` cfg keys (`gateway_force_model`+`gateway_model`=target, `gateway_upstream_url`/`gateway_api_key` from the provider, the price keys, an augmented `gateway_model_context_map`, and vision overlay keys); `handle_chat`/`handle_embeddings` merge it onto the base cfg (`cfg = {**cfg, **overlay}`) BEFORE everything else, so the existing model-selection / `pricing_from_cfg` / `parse_context_map` / vision / url+key paths transparently use the route's provider, target model, pricing, vision model and context window. `_ensure` (the httpx pool / semaphore / breaker / vision cache) reads only the non-overlaid pool keys, so the connection pool is never churned per request.
**No matching route = `None` overlay = byte-identical legacy behavior** - this is the "nobody feels the transformation" guarantee: `molodetz`/`molodetz~embed` and every existing caller are byte-identical when no route matches. A route with a blank provider overlays only model+pricing(+vision), keeping the default upstream url/key.
**CRUD.** Admin JSON at `/admin/gateway/{providers,models}` (`routers/admin/gateway_configs.py`, `require_admin`, Pydantic `ProviderIn`/`ModelRouteIn` validation, accepts both JSON from `static/js/GatewayAdmin.js` and form from Devii), audited under `gateway.provider.*`/`gateway.model.*` (category `ai`), included in the admin package with the page at `/admin/gateway` (`templates/admin_gateway.html`, sidebar link, `admin_section="gateway"`).
**Devii tools:** `gateway_providers`/`gateway_models` (read) and `gateway_provider_set`/`gateway_provider_delete`/`gateway_model_set`/`gateway_model_delete` (admin; the two deletes are `CONFIRM_REQUIRED`).
When adding a routed value, overlay it as the matching `gateway_*` cfg key so the runtime needs no new branch.
**Tests:** `tests/unit/services/openai_gateway/routing.py` (overlay/economy/kind isolation), `tests/unit/services/openai_gateway/gateway.py::test_model_route_overrides_upstream` (end-to-end through `handle_chat`), and `tests/api/admin/gateway/` (admin CRUD + validation + role gating).
+33
View File
@@ -0,0 +1,33 @@
# Pub/Sub service (`devplacepy/services/pubsub/`, `devplacepy/routers/pubsub.py`)
This file documents the database-free publish/subscribe bus. Claude Code auto-loads it when a file under `devplacepy/services/pubsub/` is read or edited.
## Overview
`pubsub.py` (router, mounted at `/pubsub`) is a **database-free** publish/subscribe bus. `WS /pubsub/ws` (subscribe/unsubscribe/publish frames, `foo.*` wildcards), `POST /pubsub/publish` + `GET /pubsub/topics` (admin/internal). Lock-owner-gated WS (`4013` retry) so subscribers converge on one worker; topic authz in `services/pubsub/policy.py` (`user.{uid}.*` private, `public.*` shared, admin/internal anywhere, guests opt-in). In-process `services.pubsub.publish(topic, data)` for backends; frontend `app.pubsub`.
A general-purpose, **database-free** publish/subscribe bus mounted at `/pubsub`, for live broadcast between frontend tabs, background services, and Devii without polling.
## Hub (`services/pubsub/hub.py`)
`PubSubHub` keyed by topic pattern -> live websockets (no persisted buffer; pub/sub is fire-and-forget). `subscribe/unsubscribe/drop_socket`, `async publish(topic, frame) -> delivered` (drops dead sockets), `topics()`. Wildcards: a subscription to `foo.*` matches `foo` and `foo.bar` (`topic_matches`). Module singleton `pubsub`; `services.pubsub.publish(topic, data)` is the in-process helper for backends.
## Single-worker model
The `WS /pubsub/ws` is gated to the service lock-owner worker with the `4013` fast-retry close code (like `/devii/ws`), so every subscriber transparently converges on one worker and the in-process fan-out reaches everyone with no broker and no DB. Deliberate, documented trade-off: ephemeral and best-effort (a crash drops in-flight messages), matching the `background` queue's durability stance.
## Topic authz (`services/pubsub/policy.py`)
`resolve_actor` reuses the dbapi caller (admin/internal) then falls back to user/guest. `can_subscribe`/`can_publish`: admin/internal anywhere; users publish only to their own `user.{uid}.*` namespace and subscribe to that plus `public.*`; publishing to `public.*` is admin/internal only (prevents cross-user spam); guests get read-only `public.*` when `pubsub_allow_guests` is on. Topic names match `^[A-Za-z0-9_.*-]{1,128}$`; payloads capped at 64 KB.
## HTTP + introspection
`POST /pubsub/publish` (admin/internal; returns `409` on a non-owner worker, retry), `GET /pubsub/topics` (admin/internal). Admin/internal publishes record a `pubsub.publish` audit event.
## Frontend
`app.pubsub` (`static/js/PubSubClient.js`) is a lazy reconnecting client: `app.pubsub.subscribe(topic, cb)` / `app.pubsub.publish(topic, data)`, `4013`-aware fast retry, re-subscribes on reconnect, client-side wildcard dispatch.
## Service + nginx
`PubSubService` (`default_enabled=True`, no-op run loop) surfaces the `pubsub_allow_guests` toggle on `/admin/services` and gates the WS via `is_enabled()`. `/pubsub/ws` has its own nginx upgrade location.
+41
View File
@@ -0,0 +1,41 @@
This file documents the Telegram bot subsystem. Claude Code auto-loads it when a file under `devplacepy/services/telegram/` is read or edited.
## Telegram bot (`services/telegram/`)
Puts Devii on Telegram. `TelegramService(BaseService)` (`services/telegram/service.py`, `default_enabled=False`, registered in `main.py`) supervises ONE long-poller subprocess (`python -m devplacepy.services.telegram.worker`). Because `BaseService` runs only on the service-lock owner, there is always exactly one poller - mandatory, since Telegram returns **409 Conflict** on concurrent `getUpdates` per token (a load-bearing invariant). The bot token is a **secret `ConfigField`** passed to the subprocess via the `TELEGRAM_BOT_TOKEN` env var, never baked or stored outside `site_settings`.
### Worker (`worker.py`, `backend.py`)
A pure Telegram I/O gateway: long-polls `getUpdates`, handles every Telegram API call, rate-limits, honours `429 retry_after`, treats `409` as transient backoff, and self-exits on stdin EOF (parent death). It imports only stdlib + `devplacepy.stealth` + `format.py` - **no FastAPI app / database import** (verify with `python -c "import devplacepy.services.telegram.worker"`). It speaks NDJSON over stdin/stdout: inbound `{type:"message", chat_id, from_id, text, images:[data-uri]}` on stdout, outbound `{cmd:"send"|"edit"|"chat_action", req_id, ...}` on stdin, plus `{type:"result", req_id, message_id}` and `{type:"log"}` frames. Private chats only. A sent photo is downloaded and inlined as a base64 `data:` URI (capped by `TELEGRAM_IMAGE_MAX_BYTES`); HTTP goes through the stealth client. `backend.py` has a `TelegramBackend` ABC + `HttpTelegramBackend` (real) + `FakeTelegramBackend` (tests).
### Bridge (`bridge.py`)
Runs in the parent (co-located with the Devii hub on the lock owner) and drives Devii **in-process** - NOT via the WebSocket, no WebSocket-from-subprocess. For each inbound message it resolves the chat to a pairing (`store.py`), then for a paired user opens a `hub.get_or_create("user", uid, ..., channel="telegram")` session, attaches a per-turn `TelegramConnection` (an object exposing `async send_json` - the same sink shape `session.attach` expects, reusing the `routers/devii.py` quota gate and the shared ledger), runs the same quota gate as `routers/devii.py` (`devii.daily_limit_for` / `spent_24h`, shared ledger), and calls `session.spawn_turn(content, audit_text=...)`. Turns serialize per chat via a lock and are bounded by a global semaphore. The connection sends a `Devii is thinking...` placeholder, refreshes a `typing` chat action every ~4s, throttle-edits the placeholder with tool progress, and on the `reply` frame edits it into the final answer (markdown -> Telegram HTML via `format.py`, split at 4096 with a plain-text fallback; overflow becomes extra messages) - so **a turn updates ONE Telegram message** (placeholder -> typing -> progress edits -> final answer) instead of spamming. Browser/avatar frames are resolved immediately as unsupported (no web client on Telegram). `_emit` frames never reach an open web terminal because the thread is its own channel.
### `channel="telegram"` isolation
A dedicated conversation thread, isolated from the web `main` thread, but `hub.get_or_create` is patched so this channel still uses the persistent `db` (shared lessons/behavior/tasks/quota with `main`) - `hub.get_or_create` treats `telegram` like `main` for `owned_db`, so lessons/behavior/tasks/quota are shared; only `devii_conversations` is keyed per channel. The scheduler only starts on `main`, so the Telegram session never double-runs it; scheduled-task pushes to Telegram go through the `telegram_send` tool from the user's `main` session.
### Pairing (`store.py`, `routers/profile/telegram.py`)
`POST /profile/{username}/telegram` (owner-or-admin, mirrors `ai_correction`) issues a single-use 4 digit code (sha256-hashed, unique among active codes, reissue invalidates prior, TTL `telegram_code_ttl_minutes`) into `telegram_pairings`, or unpairs. Sending the code to the bot binds `telegram_links` (chat_id+from_id -> user_uid); wrong attempts are throttled per chat in memory. Tables are ensured in `init_db`. Profile shows a Telegram card (`static/js/TelegramPairing.js`).
### `telegram_send` Devii tool
`services/devii/telegram/`, `handler="telegram"`, `requires_auth=True`, audited `telegram.send`. Looks up the paired chat, requires the service running on the lock owner, and delivers a markdown message. Use it for notifications from a turn or a scheduled task.
### Image vision (turn extension)
`session.spawn_turn`/`_run_turn` and `agent.respond` accept `str | list` for image vision; the bridge builds an OpenAI `[{type:text},{type:image_url,image_url:{url:<data-uri>}}]` content list, which the gateway `VisionAugmenter` describes when the route has a vision model. After the turn the structured message is redacted to a short `[image]` placeholder so `devii_conversations` never stores base64.
### Exact per-user cost (vision included)
The turn authenticates the gateway with the paired user's own api_key, so `resolve_owner` books both the chat call and the vision sub-call to `gateway_usage_ledger` under that user (authoritative). For Devii's own per-user ledger and 24h cap to also be exact, the gateway folds the vision sub-call cost into the chat response `X-Gateway-Cost-USD` (and exposes `X-Gateway-Vision-Cost-USD`), and `services/devii/llm.py` records that native cost via `cost.record_cost` instead of estimating from tokens (`CostTracker.has_native` makes `cost_usd()["total"]` the gateway's real cost). This applies to every gateway consumer, not just Telegram; with no gateway header it falls back to token-based pricing.
### Live logs and stats
Worker output streams to the pubsub topic `admin.services.telegram.logs` and the service detail page's live pane (`ServiceMonitor.appendLive`); it is never persisted. `collect_metrics` reports messages in/out, edits, errors, average response latency, uptime, and token presence.
### Audit and nginx
Audit events: `telegram.pair.request|success|failure`, `telegram.unpair`, `telegram.send` (category `telegram`). nginx note: the bot uses outbound long-polling, so no inbound WebSocket location is needed.
+21
View File
@@ -0,0 +1,21 @@
# XML-RPC bridge (`devplacepy/services/xmlrpc/`, `devplacepy/routers/xmlrpc.py`)
This file documents the XML-RPC bridge. Claude Code auto-loads it when a file under `devplacepy/services/xmlrpc/` is read or edited.
## Overview
`xmlrpc.py` (the router, mounted at `/xmlrpc`) reverse-proxies XML-RPC calls to the forking XML-RPC bridge (`services/xmlrpc/`, supervised by `XmlrpcService` on loopback `config.XMLRPC_PORT`), which generates one XML-RPC method per documented REST endpoint from `docs_api.API_GROUPS` (`posts.create`, `feed.list`, ...). One struct of named params per call; auth via in-band `api_key`, `X-API-KEY`/`Bearer` header, or `http://user:pass@host/xmlrpc` Basic. Full introspection + `system.multicall`; REST errors become XML-RPC faults. Exempt from the rate limiter (enforced on the forwarded internal hop).
The whole REST surface is also exposed over XML-RPC at `/xmlrpc`, generated **automatically from `docs_api.API_GROUPS`** so it never drifts: add a documented endpoint and it becomes an XML-RPC method for free.
## Pieces
- **`registry.py`** walks `docs_api.API_GROUPS` and yields one `Method` per endpoint. Method name = the endpoint `id` with `-`->`.` (`posts-create` -> `posts.create`). It carries the HTTP method, path, auth, and the `param` list split by `location` (path/query/form). It also builds the `system.methodHelp` text and `system.methodSignature` from the same data.
- **`bridge.py`** translates ONE XML-RPC call into an internal REST request. The single struct argument is a dict of named params; the bridge substitutes `{name}` path segments, routes the rest to query (GET) or form (write methods), sets auth headers, and calls `config.INTERNAL_BASE_URL` with a **synchronous** `stealth.stealth_sync_client` (the forking server is sync). It returns the decoded JSON; any REST `status >= 400` becomes an `xmlrpc.client.Fault` whose `faultCode` is the HTTP status. A missing required param faults before the call.
- **`server.py`** is `ForkingXMLRPCServer(ForkingMixIn, BridgeDispatcher, SimpleXMLRPCServer)`, `allow_none=True`, with `register_introspection_functions()` + `register_multicall_functions()` and one registered function per generated method. The custom request handler captures auth from the POST into `server.current_auth` (safe because each request runs in its own forked child) and the bridge forwards it: `X-API-KEY` and `Bearer` become an api-key header; a raw `Authorization: Basic` header is **forwarded verbatim** so the REST layer does its normal `username/email:password` login (`utils._user_from_basic`) - this is what makes `xmlrpc.client.ServerProxy("http://user:pass@host/xmlrpc")` work. It also forwards `X-Real-IP`/`X-Forwarded-For` so **rate limiting is attributed to the real caller** on the internal REST hop. Run standalone with `python -m devplacepy.services.xmlrpc.server` or the `devplace-xmlrpc` console script.
- **`XmlrpcService(BaseService)`** (`services/xmlrpc/__init__.py`) supervises the server as a **subprocess** (never fork the asyncio/uvicorn process): `on_enable` spawns, `run_once` health-checks and respawns, `on_disable` terminates. `default_enabled=True`, but like every service only the lock-owner worker runs it, so exactly one process binds the port.
- **`routers/xmlrpc.py`** reverse-proxies `/xmlrpc` (POST = RPC, GET = usage banner) to the forking server on `config.XMLRPC_PORT`, forwarding the auth + real-IP headers. `/xmlrpc` is **exempt from the request rate limiter** in `main.py` (enforcement happens on the forwarded internal hop). nginx forwards `/xmlrpc` in production with longer timeouts.
## Auth
Auth has three equivalent forms: `api_key` inside the param struct (works with a vanilla `ServerProxy`), an `X-API-KEY`/`Bearer` header on the transport, or HTTP Basic via a credentialed URL `http://username:password@host/xmlrpc`. Public endpoints need none. Config: `config.XMLRPC_BIND`/`XMLRPC_PORT` (loopback, `10550`). Documented at `/docs/xmlrpc.html`.