Add admin-unlimited workspaces, AI gateway model fallback, and real streaming/thinking control

Admin-unlimited Dev Workspaces: an admin-owned workspace is now exempt from the
max-workspace-count limit, the max-tunnel-count limit, and the whole
idle-stop/idle-warn/retention-delete lifecycle. Resolved once in
quota.resolve() as Limits.unlimited (owner uid checked against
get_admin_uids()), consumed at the three enforcement points
(provision.ensure, provision.publish_tunnel,
WorkspaceService._advance_lifecycle). Also hardens
get_admin_uids()/get_primary_admin_uid() against a partially-schemaed users
table (uid/role column guard), which a fresh test/init_db() path could hit.

AI gateway per-model automatic fallback: any gateway_models route
(chat/embed/image) can now name a fallback_model, picked on /admin/gateway
from a select box of other configured public model names of the same kind
only (never an internal upstream model id). When a route fails after its own
retries are exhausted, the gateway retries once, automatically, against the
fallback's own provider/pricing/key, before any bytes reach the client
(including for a streaming response). One hop only, no chains or cycles;
self-reference and cross-kind fallbacks are rejected at write time.

AI gateway real upstream streaming and thinking-default control: stream:true
is now forwarded to the upstream and relayed to the client as real SSE
chunks (measured TTFT/inter-token latency) instead of a simulated split
response, and every chat/vision call explicitly disables model "thinking" by
default (admin-overridable via gateway_thinking), with per-dialect handling
for DeepSeek, OpenRouter, and Ollama.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TjdKTWgWpW2SMNW8SFqxz5
This commit is contained in:
2026-09-03 08:47:57 +02:00
co-authored by Claude Sonnet 5
parent 8ae3f628c7
commit a693a6f4d8
33 changed files with 1310 additions and 110 deletions
@@ -120,6 +120,18 @@ class GatewayService(BaseService):
"content, all in one system message). Leave blank to disable.",
group="Prompt",
),
ConfigField(
"gateway_thinking",
"Enable thinking by default",
type="bool",
default=config.THINKING_DEFAULT,
help="Off by default: the gateway disables model thinking on every chat "
"call (DeepSeek thinking.type=disabled, OpenRouter reasoning.effort=none, "
"Ollama think=false) unless the client explicitly enables it with think, "
"thinking, or reasoning. Fastest path. On: thinking is enabled unless the "
"client disables it.",
group="Prompt",
),
ConfigField(
"gateway_vision_enabled",
"Vision augmentation",