Add skill install (copy/symlink) and tmux-gated chunked RAG search

- install_skill: adopt an existing skill from a local dir, SKILL.md,
  .zip, or git URL; symlinks when the source already lives in a
  recognized external skills folder (.claude/skills, .agents/skills)
  instead of vendoring a duplicate, with install provenance recorded
  per skill. Skill discovery now also reads those external folders
  read-only, so skills placed by other tools are visible with no
  install step at all.
- export_terminal_log: full tmux scrollback export, visible only
  inside a live tmux session (TOOL_ENV_GATES, a new env-conditional
  layer on top of the existing keyword-based lazy tool loading).
- Chunked lexical RAG (SQLite FTS5 + bm25, no vectors/embeddings):
  new chunks table wired into record_save, shell/subagent output
  spill, and terminal log export, searchable via the existing search
  tool (kind chunk) with full chunk text returned, not just a
  snippet. delete_record cleans up a record's chunks with it.
- Fix a regression where create_skill's default scope silently
  became context-dependent instead of always "project".
- Rewrite README.md to document all of the above plus the existing
  context/tool-selection and search/memory internals in depth.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
2026-10-09 20:05:11 +02:00
co-authored by Claude Sonnet 5
parent 2022e1082e
commit f078ab4731
2 changed files with 591 additions and 433 deletions
+187 -401
View File
@@ -2,20 +2,18 @@ retoor <retoor@molodetz.nl>
# tai
tai is a single-file autonomous AI agent written in Python. The entire
implementation lives in `tai.py` (about 5200 lines) and uses only the Python
standard library: no dependencies, no install step, no build system.
**One file. Zero dependencies. A real agent.**
The agent runs as an interactive REPL or as a one-shot command. It reasons
through an OpenAI-compatible backend, acts through thirty-six tools, keeps
per-profile memory, seals its stored state at rest, and can isolate shell and
file operations inside a container sandbox.
`tai` is a single-file autonomous AI agent written in Python — the entire implementation lives in `tai.py` (~7,600 lines) and imports **only the Python standard library**. No `pip install`, no `requirements.txt`, no build step, no lockfile to rot. It runs as an interactive REPL or a one-shot command, reasons through an OpenAI-compatible backend, acts through 38 tools, keeps per-profile memory in a self-healing SQLite vault, seals that vault at rest, ranks every search with real BM25 (no vector database, no embeddings bill, no GPU), chunks and indexes anything worth remembering for retrieval, and can wall shell/file work off inside a disposable container sandbox.
Nothing here is decorative. Every mechanism below exists because something specific was expensive — context tokens, API calls, attacker surface, disk writes, or your attention — and got engineered out.
## Requirements
- Python 3.10 or newer, no third-party packages.
- Optional: `podman` or `docker` for the sandbox and Telegram voice notes.
- Optional: `tmux` for terminal content capture.
- Optional: `tmux` for terminal content capture and full-history export.
- Optional: `git` only if you install a skill straight from a git URL — every other skill/install path is pure stdlib.
## Quick start
@@ -27,42 +25,42 @@ file operations inside a container sandbox.
./tai.py --version
./tai.py what is 2+3, use the shell
Trailing arguments form a one-shot prompt: the agent answers once and exits
with code 0. Without arguments, tai starts an interactive session.
Trailing arguments form a one-shot prompt: the agent answers once and exits with code 0. Without arguments, tai starts an interactive session.
## REPL commands
| Command | Effect |
|-------------------|---------------------------------------------------|
| `/profile [name]` | Show the current profile or switch to it |
| `/profiles` | List all profiles, current marked with `*` |
| `/bots` | List profile bots, current marked with `*` |
| `/bot [name]` | Switch to another bot, history resumes |
| `/env [target]` | Show or switch execution environment |
| `/skills` | List loaded skill files |
| `/secret` | Manage sealed secrets (set|list|delete) |
| `/sysinfo` | Show host environment checks |
| `/schedule` | Schedule a prompt for later (at|every) |
| `/schedules` | List scheduled prompts and outcomes |
| `/unschedule` | Delete a scheduled prompt by id |
| `/records` | Search saved records, optional query |
| `/record <id>` | Read one record page by mem id |
| `/graph <node>` | Show one vault node neighborhood |
| `/tags [prefix]` | List tags with usage counts |
| `/tools [name]` | Show core and lazy tools |
| `/install ...` | Install status, install, upgrade, reinstall |
| `/search <query>` | Ranked search over records, events, audit |
| `/audit [path]` | Show file audit trail |
| `/restore <id>` | Restore a file from an audit row |
| `/release` | Bump version, back up, log message |
| `/fork <task>` | Spawn a background subagent, REPL stays free |
| `/agents` | List background subagents |
| `/agent <id>` | Show one subagent result |
| `/agent clear` | Purge finished subagents |
| `/compact` | Compress history into a summary |
| `/clear` | Drop history, keep the system message |
| `/help` | Show the command overview |
| `/quit` | Exit |
| Command | Effect |
|-----------------------|---------------------------------------------------------|
| `/profile [name]` | Show the current profile or switch to it |
| `/profiles` | List all profiles, current marked with `*` |
| `/bots` | List profile bots, current marked with `*` |
| `/bot [name]` | Switch to another bot, history resumes |
| `/env [target]` | Show or switch execution environment |
| `/skills` | List loaded skills with provenance (native/copied/symlinked/external) |
| `/skills install ...` | Install an existing skill by copy or symlink, no scaffolding |
| `/secret` | Manage sealed secrets (set\|list\|delete) |
| `/sysinfo` | Show host environment checks |
| `/schedule` | Schedule a prompt for later (`at`\|`every`) |
| `/schedules` | List scheduled prompts and outcomes |
| `/unschedule` | Delete a scheduled prompt by id |
| `/records` | Search saved records, optional query |
| `/record <id>` | Read one record page by mem id |
| `/graph <node>` | Show one vault node neighborhood |
| `/tags [prefix]` | List tags with usage counts |
| `/tools [name]` | Show core and lazy tools, respecting environment gates |
| `/install ...` | Install status, install, upgrade, reinstall |
| `/search <query>` | Ranked search over records, events, audit, **and chunks** |
| `/audit [path]` | Show file audit trail |
| `/restore <id>` | Restore a file from an audit row |
| `/release` | Bump version, back up, log message |
| `/fork <task>` | Spawn a background subagent, REPL stays free |
| `/agents` | List background subagents |
| `/agent <id>` | Show one subagent result |
| `/agent clear` | Purge finished subagents |
| `/compact` | Compress history into a summary |
| `/clear` | Drop history, keep the system message |
| `/help` | Show the command overview |
| `/quit` | Exit |
Any other input is sent to the agent.
@@ -72,282 +70,157 @@ Any other input is sent to the agent.
/install status
/install upgrade scheduler-service
Six install targets exist side by side: `binary` (`~/.local/bin/tai.py`),
`bash-hook` (a guarded `command_not_found_handle` block in `~/.bashrc`,
backed up once to `~/.bashrc.bak-tai`, so unknown shell commands are
answered by the agent), `venv` (`~/.tai/venv`, created whenever the
venv module exists), `scheduler-service` and `telegram-service`
(systemd user units that prefer the venv python), and `container`
(the sandbox, whose image carries its own `/box/venv`). `/install
status` (or the `install` tool) reports exactly which of these exist,
with versions and service states. `install` adds missing pieces,
`upgrade` refreshes in place, `reinstall` rebuilds artifacts from
scratch, `uninstall` removes them; service changes restart or start
units immediately. Data handling is explicit: no action ever touches
the vault (`memory.db`, secrets, schedules, records, backups).
Six install targets exist side by side: `binary` (`~/.local/bin/tai.py`), `bash-hook` (a guarded `command_not_found_handle` block in `~/.bashrc`, backed up once to `~/.bashrc.bak-tai`, so unknown shell commands are answered by the agent), `venv` (`~/.tai/venv`, created whenever the venv module exists), `scheduler-service` and `telegram-service` (systemd user units that prefer the venv python), and `container` (the sandbox, whose image carries its own `/box/venv`). `/install status` (or the `install` tool) reports exactly which of these exist, with versions and service states. `install` adds missing pieces, `upgrade` refreshes in place, `reinstall` rebuilds artifacts from scratch, `uninstall` removes them; service changes restart or start units immediately. Data handling is explicit: no action ever touches the vault (`memory.db`, secrets, schedules, records, backups).
## Backends
The primary backend is `model.cloud.pravda.education`, an OpenAI-compatible
gateway that needs no API key and selects a free model per request. If a
request fails, tai retries it on `devplace.net/openai/v1`, which requires
`DEVPLACE_API_KEY`. Both endpoints speak `/chat/completions`, including
native tool calls and streaming.
The primary backend is `model.cloud.pravda.education`, an OpenAI-compatible gateway that needs no API key and selects a free model per request. If a request fails, tai retries it on `devplace.net/openai/v1`, which requires `DEVPLACE_API_KEY`. Both endpoints speak `/chat/completions`, including native tool calls and streaming.
---
## The context problem, and how tai solves it
Every tool schema you hand an LLM costs real tokens on every single turn, whether or not it gets used, and a model drowning in fifty tool definitions picks worse than one offered the five that actually matter. tai treats this as a first-class engineering problem, not an afterthought, with three independent layers stacked on top of each other:
**1. Core vs. lazy.** Of the 38 tools, only `CORE_TOOLS` — `shell`, `read_file`, `write_file`, `remember`, `recall`, `load_skill`, `fork`, `poll` — are sent to the model on every turn. The other 30 are *lazy*: their full JSON schema only enters the request payload when they're judged relevant to what's actually being discussed.
**2. Relevance by three independent signals, not a fixed list.** `select_tools()` scans the last 6000 characters of conversation (user, assistant, and tool-result turns — `conversation_text()`) and pulls in a lazy tool's full schema the moment any one of three things matches:
- a **name token** — `record_search` offers itself up the moment "record" or "search" appears, filtered through `LAZY_NAME_STOP` (`get`, `set`, `list`, `add`) so generic verbs don't trigger everything;
- a **trigger tag** — every tool carries a hand-picked synonym list in `TOOL_TAGS` (`cron` loads `schedule`, `undo` loads `restore`, `password` or `apikey` loads `store_secret`, `symlink` loads `install_skill`), so the model never has to know a tool's literal function name to reach for it;
- a **description word** (6+ characters) — a last-resort fallback so nothing is ever truly unreachable just because its tags happened to miss the exact phrasing.
Tool *results* feed this loop too: a spill pointer that mentions `record_read` loads that tool for the very next step, so the model is never stuck knowing a capability exists but unable to reach it.
**3. Environment gates, orthogonal to relevance.** `TOOL_ENV_GATES` is a second filter that runs *before* relevance is even checked. `export_terminal_log` is gated on `bool(os.environ.get("TMUX"))` — outside a live tmux session it is invisible in the system prompt, invisible in `select_tools()`, invisible in `/tools`'s no-arg listing, full stop, regardless of how insistently you ask for it. The handler re-checks the same condition at call time anyway, so even a stale tool-call from earlier context can't slip through. This is the general mechanism for "this capability is nonsensical outside environment X" — a second axis entirely separate from keyword relevance.
A compact name-plus-one-line-summary catalog of every lazy tool (`tool_catalog()`) still rides in the system prompt at all times, so nothing is ever *hidden* — only its verbose, multi-field JSON schema is deferred. `/tools` shows the core/lazy split live (respecting the same environment gates a real turn would see), `/tools <name>` shows one tool with its full description and trigger tags.
Net effect: the model sees 8 schemas by default instead of 38, picks up exactly the ones a conversation actually needs, and never has a tool pushed on it that can't possibly apply in the current environment.
## Context budget and compaction
Raw history isn't free either. Token count is estimated every turn (`estimate_tokens`, 4 chars ≈ 1 token) against a `CONTEXT_CAP` of 32,000. The instant usage crosses 80% of that (`COMPACT_RATIO`), `compact_messages()` fires automatically: it keeps the system message, keeps the last 6 turns verbatim (`KEEP_TURNS`, snapped backward to the nearest user turn so a conversation never resumes mid-answer), and fires everything in between — including a readable digest of which tool calls happened, not just their text — at a dedicated, separate LLM call whose only job is lossy compression (`COMPACT_SYSTEM`). The summary replaces the gutted middle. `/compact` runs the exact same path on demand. The system message itself is independently capped at `SYSTEM_MAX_CHARS` (12,000) so `remember` can't let it grow without bound either.
## Spill-to-vault: the other half of context control
Tool output is the other place context explodes. Any shell command or subagent poll result over `SPILL_LIMIT` (6000 characters) never reaches the model whole: `spill_or_truncate()` writes it to the vault as a tagged record instead and hands back a head/tail preview plus a `mem:<id>` pointer (`spilled_view()`). The model pages the rest back deliberately with `record_read`, offset and all, instead of having 40KB of build log forced into every subsequent turn of the conversation. And — see the RAG section below — that spilled record doesn't just sit there as a cold blob: it gets chunked and indexed the instant it's written, so "what was that error again" can be answered by search instead of by re-reading the whole thing.
---
## Keyword storage, tagging, and the knowledge graph
Every record, secret, and schedule in the vault carries **normalized tags**, and tai is deliberately strict about vocabulary so it doesn't fragment into a thousand near-duplicate labels:
- tags are lowercased, hyphenated, capped at 32 characters (`normalize_tag`), and a single item never carries more than 20 raw / effectively 5 auto-detected tags;
- a small set of genuinely irregular plurals (`news`, `means`, `series`, `species`, `physics`) is hard-excluded from singularization (`SINGULAR_KEEP`); everything else runs through a real little inflection engine (`singular_noun`): `-ies → -y`, `-ses/-xes/-zes/-ches/-shes → drop "es"`, trailing `-s` dropped unless it's `-ss/-us/-is/-os`;
- `canonical_tag()` always prefers whatever form *already exists* in that profile's vocabulary over whatever form you typed, so "skills" and "skill" collapse to whichever one got there first, instead of silently forking the taxonomy;
- the `tags` tool lists every known tag with usage counts — the explicit, correct way to reuse vocabulary instead of inventing synonyms, and the one both humans and the model are told to consult first.
**Auto-tagging is not keyword extraction, it's vocabulary matching.** `auto_tags_for()` scans a new record's title-plus-content for words (and adjacent word-pairs, so multi-word tags like `api-client` still fire) that are *already* real tags in that profile — up to `TAG_AUTO_MAX` (5) — and attaches them automatically. A brand-new vocabulary word never gets invented this way; only previously-established meaning spreads to new content, which is what keeps the tag space from drifting into noise over time.
**The graph builds itself.** Every new record automatically links to the `TAG_LINK_MAX` (3) most recently updated records sharing a real tag (`tag_neighbors()`, explicitly excluding the structural `record`/`file` tags so it's never self-referential noise) — an edge with relation `shares-<tag>` lands in the `edges` table with zero manual curation. `graph_link` lets you add arbitrary typed edges between any two vault nodes (`mem:<hash>` records, `secret:<name>`, `sched:<id>`), and `graph_query` walks the neighborhood breadth-first, capped at depth 4 so a query can never explode into the whole vault. `search`'s `expand` flag appends each record hit's nearest neighbors inline, for free.
---
## Search and RAG — BM25 and chunking, deliberately no vectors
This is the part worth being loud about: **tai does retrieval-augmented generation with zero embeddings, zero vector database, and zero extra dependency, on purpose.**
**One unified lexical index.** `fts_docs` is a single SQLite FTS5 virtual table covering `records`, `events`, `audit`, and `chunks` at once, kept in sync not by any application-level reindex job but by SQL triggers fired directly on insert/update/delete (`setup_fts()`). Query terms are tokenized safely server-side — no raw FTS MATCH syntax ever reaches SQLite from user input — stemmed by the Porter tokenizer when available, and ranked with SQLite's own built-in `bm25()`. `search`, `recall`, `record_search`, and `audit` all rank through this *one* index instead of each reinventing matching, falling back to plain `LIKE` only when a query has no tokens FTS can use.
**Chunking, for retrieval-grade granularity, without inventing a second storage system.** A document-level FTS hit tells you a 40KB record contains your answer somewhere; it doesn't hand you the answer. `chunk_lines()` splits any text into bounded, overlapping windows — up to 60 lines or 4000 characters, whichever comes first, with a 3-line overlap carried into the next chunk — and critically **never splits mid-line**, so a chunk boundary never severs a command from the middle of its own output. Every chunk lands in the `chunks` table and gets its own `fts_docs` row (`kind='chunk'`), ranked by the exact same BM25 engine as everything else — one index, one ranking function, no separate "vector similarity vs. keyword score" fusion problem to solve, because there's only ever one kind of score. `collapse_repeated_lines()` runs first on anything log-shaped, folding 3+ identical consecutive lines (progress-bar spam, retry loops) into one line plus a count note, so pathological repetition doesn't burn chunk budget on noise.
A chunk hit from `search` returns **the actual chunk text**, not just FTS's highlighted snippet — the whole point of chunk-level retrieval is to hand back something directly usable as context, and the parent record id rides along so `record_read` can pull more surrounding material on demand (hierarchical retrieval: chunk for precision, parent for context, one hop away, never both loaded by default).
**Why no vectors.** Embeddings mean a model call (cost, latency, an API dependency) or a local model (a real dependency, a real download, real RAM) for *every single thing written* — directly incompatible with the single-file, stdlib-only, zero-install rule this whole project is built on. BM25 over SQLite FTS5 costs nothing to build, nothing to maintain, and handles exactly what this vault actually needs to search well: command output, notes, transcripts, terminal history — content where the literal words (an error string, a flag name, a path) *are* the signal. This was an explicit, considered tradeoff, not a missing feature.
**Chunking is wired into the system, not bolted onto one feature.** `record_save`, the shell/subagent output spill path (`spill_output`), and terminal log export all chunk through the identical `chunk_lines()` call — "RAG" here means a property of the vault, not a special case for logs. It is deliberately *not* wired into `add_record`'s lower-level, highest-frequency caller — `upsert_file_record`, which fires on every single file write and edit for audit bookkeeping — because exploding the chunk table on every keystroke-adjacent save would be pure waste for content `read_file`/`grep` already serves perfectly well. And deleting a record deletes its chunks and their index rows with it (`Store.delete_record`); there is no orphan path to worry about.
**Sealed search stays ranked without ever touching plaintext to disk.** Encrypted-at-rest events can't be FTS-indexed in place — SQLite can only full-text-search what it can read as text. Every boot decrypts events into a `:memory:`-only FTS5 mirror (`open_mem_events`), capped at the newest `MEM_EVENTS_CAP` (10,000) rows and synced incrementally as new ones arrive, so sealed vaults get exactly the same BM25 ranking and `[marked]` snippets as plaintext ones — and nothing decrypted ever becomes a temp file, a spill table, or a persistent index on disk. Alternatives were researched and rejected on purpose, not just not-gotten-to: page-level encryption (SQLCipher) is a non-standard C extension, incompatible with the stdlib-only rule; blind indexes (deterministic HMACs of tokens, à la CipherStash) would persist in the vault file and leak term frequency and access patterns to anyone who steals it; academic searchable-encryption schemes leak access patterns too and are far heavier than this threat model needs. Decrypting into RAM, once, where the key and plaintext already live during normal operation anyway, is the only option that keeps the at-rest file fully opaque *and* the search fully ranked.
/search deploy ranked hits across records, events, audit, and chunks
/records deploy search records for "deploy"
/record mem:9f2c41aa77c3e5d1
/graph secret:db show what links to the db secret
/tags all tags by usage count
/tags de tags starting with "de"
---
## Terminal log export
`export_terminal_log` is only ever offered inside a live tmux session (see the environment-gating section above — this is the tool `TOOL_ENV_GATES` exists for). When it's available, it runs `tmux capture-pane -p -S -` — the **entire** pane scrollback, not just what's currently visible — writes it verbatim to `~/.tai/termlogs/<profile>/<pane>-<timestamp>.log`, saves it as a `termlog`-kind vault record for provenance and full-text paging, and immediately chunks and indexes it exactly like everything else above. From that point on, "what was that error three hundred lines back" is a `search` call, not a scrollback hunt — the parent record stays for full context, the chunks stay for precision, and the whole thing cost one tmux subprocess call and zero network requests.
Its quieter cousin, `get_current_terminal_content`, grabs just the visible pane (configurable line count) for a quick glance without writing or indexing anything — the difference between a peek and a capture is deliberate.
---
## Skills
Standard agent skill files (`SKILL.md` with `name` plus `description` frontmatter, optional `scripts/`, `references/`, `assets/`, per the open Agent Skills format) are discovered from four places, in ascending priority so the most tai-specific location always wins a name collision: read-only interop with `~/.claude/skills` and `./.agents/skills` first (so a skill another tool already dropped next to your project is visible with *zero* install step), then tai's own `~/.tai/skills`, then `./.tai/skills`. Descriptions stay in the system prompt catalog at all times; full instructions load only through `load_skill`, on demand — the same lazy-loading philosophy as tools, applied to skills.
Two independent, complementary ways to get a skill:
- **`create_skill`** authors a brand-new one from a brief: a dedicated deep-research subagent runs `sysinfo` first so the skill matches the actual machine, then researches with web search and shell until the material is independently verified, then writes `SKILL.md` plus supporting files — project-scoped by default. Six blueprints ship built in (`bot-creator`, `api-client`, `web-researcher`, `pdf-forms`, `data-wrangler`, `home-sysadmin`); `load_skill` builds a missing one through this exact same path the first time it's actually needed, so a skill finalizes lazily at the moment of use, not at install time.
- **`install_skill`** adopts a skill that already exists *somewhere else* instead of building one from scratch — a local directory, a bare `SKILL.md` file, a `.zip`, or a git URL (the only path in this entire feature that touches a non-stdlib binary, and it's checked for and refused cleanly if absent). It copies by default, vendoring a private snapshot — except when the source already sits inside one of the recognized external skill folders (`.claude/skills`, `.agents/skills`), in which case it **symlinks** instead, so you get one canonical copy instead of a silently drifting duplicate. Scope (project vs. home) defaults the same way `create_skill` does, and a sidecar `.installs/<name>.json` records exactly where a skill came from and how, which `/skills` and `load_skill` both surface (`[installed copy, project scope, from ...]`, `[native]`, or `[external, not managed by tai: ...]`) — every skill in the catalog is honest about its own provenance.
Both share one scope-resolution helper and one path-construction helper under the hood, so "where does a project-scoped vs. home-scoped skill actually live" has exactly one definition in the entire codebase.
The `sysinfo` tool itself runs os, python, venv, root, container, binaries, cpu, and disk checks **in parallel**, each with its own timing; `/sysinfo` prints the identical report in the REPL.
---
## Orchestration
`/fork <task>` spawns a background subagent with its own context while the
REPL stays free (the prompt shows a `+N` counter). `/agents` lists workers,
`/agent <id>` shows a result, `/agent clear` purges finished ones.
`/fork <task>` spawns a background subagent with its own context while the REPL stays free (the prompt shows a `+N` counter). `/agents` lists workers, `/agent <id>` shows a result, `/agent clear` purges finished ones. The model orchestrates through the exact same primitives: the `fork` tool (task, timeout up to one hour, profile) and the `poll` tool (id, wait up to two minutes). Workers get `WORKER_STEPS` (12) steps, a cooperative deadline, no session writes, and no interactive approval prompts — nesting caps at `FORK_MAX_DEPTH` (2) levels. Timeouts and errors always surface as an explicit status, never silently.
The model itself orchestrates through the `fork` tool (task, timeout up to
one hour, profile) and the `poll` tool (id, wait up to two minutes). Workers
get 12 steps, a cooperative deadline, no session writes, and no interactive
approval prompts. Nesting is capped at two levels. Timeouts and errors
surface as statuses, never silently.
## Execution policy
The agent decides who runs each unit of work by one ordered rule:
**The execution policy is one ordered rule, not a vibe:**
1. Future work is scheduled, never forked and waited on.
2. Quick, interactive, or memory-changing work runs on the main agent.
3. Independent, long, or context-heavy work goes to a forked subagent.
Forking buys parallelism and context isolation, not security isolation:
workers share the same machine, tools, and approval setting. Sandbox
mode is the separate axis that decides where untrusted or destructive
commands may run. Workers cannot schedule, cannot prompt for approval,
and stop nesting after two levels; scheduled prompts fire as subagents
that nobody waits on, with outcomes kept in the vault.
Forking buys parallelism and context isolation, *not* security isolation — workers share the same machine, tools, and approval setting. Sandbox mode is the separate axis that decides where untrusted or destructive commands may actually execute. Workers cannot schedule, cannot prompt for approval, and stop nesting after two levels; scheduled prompts fire as subagents that nobody waits on, with outcomes kept in the vault.
Turns run on a progress budget instead of a fixed step count. A step
counts as productive when it tries an unseen action, returns an
unseen result, or answers a user prompt, and any productive step
resets the stall counter, so a turn doing varied work can run as
long as it keeps moving. Repeating one action with identical
results earns a loop warning at 3 repeats and a stop at 5; wider
stalls (alternating actions, no new information) stop after the
patience budget (25 main, 12 worker, 40 skill builder); a 500-step
total cap backstops everything. Every user approval or typed
guidance resets the stall counter, since a supervised agent is a
safe agent.
**Turns run on a progress budget, not a step counter.** A step counts as productive the moment it tries an unseen action, returns an unseen result, or answers the user — and any productive step resets the stall counter, so a turn doing genuinely varied work can run as long as it keeps moving forward. Repeating one action with identical results earns a loop warning at `LOOP_NUDGE_AT` (3) repeats and a hard stop at `LOOP_STOP_AT` (5); wider stalls — alternating actions that produce no new information — stop after the patience budget runs out (`MAX_STEPS` 25 for the main agent, `WORKER_STEPS` 12 for a subagent, `CREATE_SKILL_STEPS` 40 for the skill-building worker, which legitimately needs the room); a `TOTAL_STEP_CAP` of 500 backstops absolutely everything regardless of how the patience math plays out. Every user approval or typed guidance resets the stall counter outright, on the theory that a supervised agent is a safe agent — a human steering is never treated as "no progress."
## Tools
---
| Tool | Purpose |
|------------------------------|------------------------------------------------------|
| `shell` | Run a shell command, big output spills to a record |
| `read_file` | Read a text file, large files truncated |
| `write_file` | Write content to a file, creating parent directories |
| `edit_file` | Replace one unique exact text match in a file |
| `web_search` | Search the web, optionally images or page content |
| `web_fetch` | HTTP client: methods, headers, bodies, status |
| `speak` | Synthesize speech, save MP3, play when possible |
| `listen` | Record from the microphone and transcribe it |
| `remember` | Merge knowledge into the profile system message |
| `recall` | Search past session memory by keyword |
| `load_skill` | Load a skill file by name |
| `get_current_terminal_content` | Capture the current tmux pane with scrollback |
| `fork` | Spawn a background subagent |
| `poll` | Collect a subagent result, big ones spill |
| `sysinfo` | Inspect the host, parallel checks with timing |
| `create_skill` | Deep-research and write a new skill file |
| `store_secret` | Store a password or token in the sealed vault |
| `list_secrets` | List vault secret names, values never shown |
| `delete_secret` | Delete a vault secret by name |
| `schedule` | Run a prompt later or on an interval |
| `unschedule` | Delete a scheduled prompt by id |
| `schedules` | List scheduled prompts and outcomes |
| `record_save` | Save text as a tagged record, get a mem id |
| `record_read` | Read one page of a record by mem id |
| `record_search` | Search records by text, kind, and tags |
| `record_delete` | Delete a record by mem id, asks first |
| `graph_link` | Link two vault nodes with a relation |
| `graph_query` | Show one vault node neighborhood |
| `delete_file` | Delete a file, pre-image stays audited |
| `audit` | Show the file audit trail for time travel |
| `restore` | Restore a file to an audit row image |
| `release` | Bump version, back up, log the message |
| `tags` | List vault tags with usage counts |
| `install` | Manage installs, status to uninstall |
| `create_bot` | Create a bot: system plus history, shared vault |
| `search` | Ranked full-text search over everything stored |
## Execution environment / sandbox
Web search runs on `rsearch.app.molodetz.nl`. Destructive shell commands ask
for confirmation unless `--yes` is given; read-only commands run directly.
Every prompt offers `[y]once [Y]always [n]o`: `Y` enables yolo mode for
the session, `--yolo` starts there, and `--auto` adds autonomous research
instead of ever asking. Answering no asks what to do instead: typed
guidance continues the turn while empty input aborts it. Overwriting an
existing file requires reading it first in the same session; new files
are always writable. Every write, edit, delete, and risky shell target
is audited with before/after images for time travel (see below).
`/env` shows the execution environment, `/env sandbox` switches shell and file tools into an isolated `tai-box` container (podman or docker, no mounts, no shared filesystem), `/env home` switches back. The image builds automatically on first use from an embedded Containerfile; every pip requirement the *agent's own features* need (`faster-whisper`, `edge-tts`) lives inside that image, while `tai.py` itself stays dependency-free — the sandbox is where "needs a real package" problems go, never the host process. Sandbox commands need no approval because the container is disposable; this is deliberately a *different axis* from the `--yes`/`--yolo`/`--auto` approval settings, not a replacement for them.
## Lazy tools
## Profiles and bots
With 30-plus tools, sending every schema on every turn would burn
context and blur tool selection, so tai loads lazily. Eight everyday
tools (`shell`, `read_file`, `write_file`, `remember`, `recall`,
`load_skill`, `fork`, `poll`) are always present; everything else
enters the payload only when the recent conversation names it. Each
tool carries trigger tags with synonyms (`cron` loads `schedule`,
`undo` loads `restore`, `password` loads `store_secret`), so ordinary
wording just works, and tool results feed selection too: a spill
pointer naming `record_read` loads it for the next step. A compact
name-plus-summary catalog stays in the system prompt so no capability
is ever hidden, only its verbose schema. `/tools` shows the split,
`/tools <name>` shows one tool with its tags.
A **profile** is a complete identity: system message plus session history under `~/.tai/profiles` (mode 0600), and its own slice of every single vault table — records, secrets, tags, graph edges, audit rows, schedules, episodic events, all keyed by profile. Switching profiles is total amnesia; the only cross-profile knowledge that exists is the identity list itself (`/profiles`). Read permissions and secret grants reset on every switch, subagents can't be polled across profiles, and tools flatly refuse to fork, schedule, or restore for another identity. Vaults from before this rule existed migrate automatically — unscoped rows join `default`.
## Skills
Standard agent skill files (`SKILL.md` with `name` plus `description`
frontmatter, optional `scripts/`, `references/`, `assets/`, per the Agent
Skills open format) are discovered in `~/.tai/skills/*/` and
`./.tai/skills/*/` (project wins on name collisions). Descriptions stay in
context; the agent loads full instructions through `load_skill` only when
needed. `/skills` lists what is available.
The `create_skill` tool authors new skills on demand: it runs a dedicated
deep-research worker on the prompt (sysinfo first, then web and shell
research until the material is verified against independent sources) and
writes the skill directory, project scope by default. Six blueprints ship
with the agent (`bot-creator`, `api-client`, `web-researcher`, `pdf-forms`,
`data-wrangler`, `home-sysadmin`): `load_skill` builds a missing one on
first use through the same deep-research path, so skills finalize lazily
the moment they are needed. The `sysinfo` tool
reports os, python, venv, root, container, binaries, cpu, and disk in
parallel, each check with its own timing; `/sysinfo` prints the same
report in the REPL.
## Sandbox
`/env` shows the execution environment, `/env sandbox` switches shell and
file tools into an isolated `tai-box` container (podman or docker, no mounts,
no shared filesystem), `/env home` switches back. The image is built
automatically on first use from an embedded Containerfile; every pip
requirement (`faster-whisper`, `edge-tts`) lives inside the image while
`tai.py` itself stays dependency-free. Sandbox commands need no approval
because the container is disposable.
## Telegram
./tai.py --install-telegram
./tai.py --uninstall-telegram
Install asks for the bot token up front, verifies it against `getMe`, stores
it in `~/.tai/telegram.env` (0600), builds the sandbox container (used for
voice transcription), and registers a `tai-telegram.service` systemd user
unit with linger enabled (failures ignored). The bot long-polls, answers
text, transcribes voice notes, and understands `/new`. It runs without
`--yes`, so destructive shell commands are denied. Uninstall stops and
removes the service and purges the container and image; the token file and
data stay.
## Profiles and memory
Each profile is a complete identity: system message plus session history
under `~/.tai/profiles` (mode 0600), and its own slice of every vault
table. Records, secrets, tags, graph edges, audit rows, schedules, and
episodic events are all keyed by profile; switching profiles is total
amnesia. The only cross-profile knowledge is the identity list itself
(`/profiles`). Read permissions and secret grants reset on every switch,
subagents cannot be polled across profiles, and tools refuse to fork,
schedule, or restore for another identity. Databases from before this
rule migrate automatically: unscoped rows join `default`.
A **bot** is a lighter-weight thing inside one profile: just a system message plus its own history, with the whole vault shared. Every profile starts with `main`; `create_bot` adds named bots from a description, rules, behavior, and optional nicknames (lowercased, with a short prefix auto-registering). `/bot` switches roles and resumes exactly where that bot left off; `@name` routes a single turn to another bot and files the question and answer into *both* histories, so the bot you're actually talking to stays aware a detour even happened.
/profile show current profile
/profile [name] switch profile, creating it when missing
/profiles list all profiles
/bots list bots with nicknames
/bot coder switch to the coder bot
@coder fix this one turn via coder, logged in both
The `remember` tool merges an instruction into the current profile system
message through the model itself: it adds facts, updates behavior, or removes
forgotten items while preserving the rest. It fires by default on new
passwords and behavior changes. `recall` searches the per-profile episodic
log in `~/.tai/memory.db` (SQLite). Context is budgeted at roughly 32k
tokens with automatic compaction at 80 percent. Credentials never enter
memory: secrets go to the vault through `store_secret` (or `/secret
set`), and profiles written before this rule are migrated to it
automatically on load.
`remember` merges an instruction into the current profile's system message *through the model itself* — adding facts, updating behavior, or removing forgotten items while leaving the rest intact — and fires by default on new passwords or behavior changes. `recall` searches the per-profile episodic log in `~/.tai/memory.db`. Credentials never enter memory this way: secrets go through `store_secret` (or `/secret set`), and any profile written before that rule existed gets migrated to it automatically on load.
## Bots
Where a profile is an identity, a bot is a lightweight role inside it:
only a system message plus its own session history, with the whole
vault shared. Each profile starts with `main`; `create_bot` adds named
bots from a description, rules, behavior, and optional nicknames (all
lowercased, the short prefix auto-registers). `/bot` switches roles
and resumes exactly where that bot left off; `@name` routes a single
turn to another bot and files the question and answer in both
histories, so the current bot stays aware of the detour.
/bots list bots with nicknames
/bot coder switch to the coder bot
@coder fix this one turn via coder, logged in both
---
## Sealed storage
Storage is sealed by default with a built-in key, which stops casual reads
but not a determined attacker, since the key ships in the source. Set
`TAI_PASSPHRASE` for real protection: a home sealed with the default key is
re-sealed to your passphrase automatically on first boot, with a notice.
Storage is sealed by default with a built-in key — this stops casual reads, not a determined attacker, since that key ships in the source. Set `TAI_PASSPHRASE` for real protection: a home sealed with the default key is re-sealed to your passphrase automatically on first boot, with a notice.
The key comes from PBKDF2-SHA256 (200k rounds) over a random salt in
`~/.tai/.seal`; values use a per-value nonce with HMAC-SHA256 encrypt-then-
MAC. SQLite access goes through custom `tai_enc`/`tai_dec` functions
registered with `create_function`, so inserts encrypt inline and recall
decrypts before matching. Existing plaintext data is sealed automatically on
first sealed start. A wrong passphrase refuses to start with exit code 2.
Set the variable empty for plaintext storage.
The real key derives from PBKDF2-SHA256 (200,000 rounds) over a random salt in `~/.tai/.seal`; values use a per-value nonce with HMAC-SHA256 in encrypt-then-MAC order. SQLite access goes through custom `tai_enc`/`tai_dec` functions registered with `create_function`, so inserts encrypt inline and reads decrypt before matching ever happens — there is no separate encrypt/decrypt pass bolted around the database. Existing plaintext data is sealed automatically on first sealed start. A wrong passphrase refuses to boot at all, exit code 2. Set the variable empty for deliberate plaintext storage.
This construction uses only the standard library and is honest file-theft
protection, not audited cryptography; high-value secrets still belong in a
dedicated manager.
This construction uses only the standard library and is honest file-theft protection, not audited cryptography — high-value secrets still belong in a dedicated manager.
The vault heals its own schema: every boot creates missing tables and
adds missing columns automatically, reporting what changed
(`vault schema upgraded: added secrets.meta, ...`). Old vaults from
any previous version open without manual steps, and SQLite's dynamic
typing means historic type drift in a column never blocks a boot.
## Sealed search
The seal guards against file theft: an attacker who copies `~/.tai`
learns nothing without the passphrase. Full-text search over sealed
events works without weakening that promise. Each boot decrypts
events into a memory-only FTS5 index (a `:memory:` database holding
the newest 10000 rows, synced incrementally as new rows arrive),
so sealed stores get BM25 ranking and marked snippets exactly like
plaintext ones. Nothing decrypted ever touches disk: no temp files,
no spill tables, no persistent helper index. The key and plaintext
already live in process RAM during operation, so a RAM-only index
adds no new exposure against the file-theft threat model, and
passphrase rotation needs no rebuild since plaintext is unchanged.
Alternatives were researched and rejected deliberately. Page-level
encryption (SQLCipher) keeps FTS5 working transparently but is a
non-standard C extension, incompatible with the single-file stdlib
rule. Blind indexes (deterministic HMACs of tokens stored next to
the ciphertext, as in CipherStash or IronCore cloaked search) would
persist in the vault file and leak term frequency plus search and
access patterns to anyone stealing it, while losing stemming and
BM25 ranking. Academic searchable-encryption schemes leak access
patterns too and are far heavier than this threat model needs.
Decrypting into RAM is the only option that keeps the at-rest file
fully opaque and the search fully ranked.
**The vault heals its own schema.** Every boot adds any missing table or column automatically and reports exactly what changed (`vault schema upgraded: added secrets.meta, ...`). The entire schema — tables, columns, primary keys, indexes — is declared once as a plain data structure (`SCHEMA` / `SCHEMA_INDEXES`) and reconciled against whatever's actually on disk (`ensure_schema()`); adding a feature that needs a new column is a one-line diff to that structure, never a hand-written migration script, and SQLite's dynamic typing means historic type drift in a column never blocks a boot either. Vaults from any previous version open with zero manual steps.
## Secrets
Passwords, tokens, and secrets live in a sealed `secrets` table inside
the same vault: encrypted at rest, migrated on passphrase rotation, and
never revealed by any tool. The model only handles names. `shell`
exposes chosen secrets as `TAI_SECRET_<NAME>` variables for one
command, `web_fetch` sends one as an authentication header, and every
result is scrubbed of known values before it reaches context, memory,
or display. Using or deleting secrets asks the user once per session
and scope; workers and the Telegram bot are denied unless started with
`--yes`, `--yolo`, or `--auto`, so unattended secret use is always an
explicit choice.
Passwords, tokens, and secrets live in a sealed `secrets` table inside the same vault: encrypted at rest, migrated automatically on passphrase rotation, and never revealed by any tool once stored. The model only ever handles names. `shell` exposes chosen secrets as `TAI_SECRET_<NAME>` environment variables for exactly one command invocation, `web_fetch` can send one as an auth header, and every tool result is scrubbed of known secret values (`Store.redact`) before it reaches context, memory, or display — so a secret can't leak sideways through some unrelated tool echoing it back. Using or deleting a secret asks the user once per session and scope; workers and the Telegram bot are denied outright unless the whole session was started with `--yes`, `--yolo`, or `--auto`, so unattended secret use is always a deliberate, explicit choice, never an accident of automation.
/secret set wifi store a value typed invisibly, never entering context
/secret list show names only
@@ -355,15 +228,7 @@ explicit choice.
## Scheduler
Schedules persist in the same vault with sealed prompts: one-shot
appointments (`at` an ISO datetime, naive means local time) or
repeating work (`every` 60 seconds or more). A background thread ticks
every 30 seconds in the REPL, the Telegram service, and `--scheduler`
mode, claims due rows atomically so parallel processes never double
fire, and runs each prompt as a subagent nobody waits on. Repeats
advance past missed windows instead of backfilling; outcomes land in
the row and stay visible through `/schedules`, `/agents`, and
`/agent <id>`.
Schedules persist in the same vault with sealed prompts: one-shot appointments (`at` an ISO datetime, naive means local time) or repeating work (`every` 60 seconds or more). A background thread ticks every `SCHEDULER_INTERVAL` (30) seconds in the REPL, the Telegram service, and `--scheduler` mode alike, and **claims due rows atomically** so multiple processes running at once (your REPL plus the systemd service, say) can never double-fire the same prompt. Repeats advance past any missed windows instead of backfilling a queue of stale ones; outcomes land right in the row and stay visible through `/schedules`, `/agents`, and `/agent <id>`.
/schedule at 2026-10-08T09:00 water the plants
/schedule every 1h check the inbox
@@ -373,73 +238,11 @@ the row and stay visible through `/schedules`, `/agents`, and
## Records and graph
Large or durable text lives in the same vault as tagged `records`,
each addressed by a `mem:<16 hex>` id: research notes, command
output, transcripts. Anything over 6000 chars returned by a shell
command or subagent is stored automatically and replaced by a pointer
with its size; `record_read` pages slices back without loading the
whole, `record_search` finds by text, kind, or tags, and
`record_delete` removes after confirmation. Secrets, schedules, and
records all carry normalized tags and join one graph: `graph_link`
connects nodes such as `mem:<hash>`, `secret:<name>`, and
`sched:<id>` with a relation, and `graph_query` walks the
neighborhood breadth-first, capped at depth 4 so context never
explodes. Records are working memory in cleartext, like session
events; true credentials belong in `store_secret`.
/records deploy search records for "deploy"
/record mem:9f2c41aa77c3e5d1
/graph secret:db show what links to the db secret
## Tags
Tag rules are strict so the vocabulary stays small: lowercase singular,
shortest common word for the subject, one canonical tag per subject, at
most 5 per item, and always reuse from the `tags` tool (which lists
every tag with usage counts) instead of inventing synonyms. Writes
merge plurals into a known singular automatically, while irregular
words (`news`, `glass`, `status`, `physics`) are never rewritten;
searches expand singular/plural variants so old splits still match.
Any content word that already exists as a tag attaches itself to a new
record (up to 5, plurals included), and each new record links itself to
the 3 most recent records sharing a tag, so the knowledge graph stays
connected without any manual work.
/tags all tags by usage count
/tags de tags starting with "de"
## Search
One FTS5 index covers records, episodic events, and the audit trail,
kept in sync by triggers and backfilled once for older rows. Queries
are tokenized safely (no raw MATCH syntax reaches SQLite), stemmed by
the porter tokenizer when available, ranked by BM25, and returned
with `[marked]` snippets. `search` queries all three stores at once
with kind filters and optional graph expansion that appends linked
neighbors to each record hit; `recall`, `record_search`, and `audit`
all rank through the same index, falling back to LIKE matching when
a query has no full-text match. Sealed events are no exception:
each boot decrypts them into a memory-only FTS5 index (newest 10000,
synced incrementally as rows arrive, never written to disk), so
sealed stores rank event hits with BM25 and snippets exactly like
plaintext ones while the at-rest seal stays untouched. The recipe
is deliberate: one ranked query plus graph hops keeps context
small and accurate instead of paging through stores.
/search deploy ranked hits across records, events, audit
Large or durable text lives in the vault as tagged `records`, each addressed by a `mem:<16 hex>` id: research notes, command output, transcripts, terminal captures. `record_read` pages slices back without ever loading the whole thing, `record_search` finds by text, kind, or tags (ranked through the same unified FTS5 index everything else uses), and `record_delete` removes a record — tags, graph edges, *and* chunks — after confirmation. Records are working memory in cleartext, exactly like session events; actual credentials belong in `store_secret`, never here.
## Audit and time travel
Every file mutation lands in an append-only `audit` table in the vault:
tool writes, edits, and deletes with before/after images, plus
pre-execution snapshots of shell targets (`rm`, `mv`, `cp`, `tee`,
`dd`, `truncate`, `shred`, and `>` redirections, globs expanded,
best-effort heuristic). Each row carries actor, timestamp, message,
true byte sizes, and tags, so history queries time-travel by path or
tag. Images cap at 20000 chars with an explicit truncation marker, and
`restore` refuses truncated images rather than writing partial
content. Every audited path also keeps a `file` record with its
absolute path and latest contents (50000 chars, marked when cut).
Every file mutation lands in an append-only `audit` table: tool writes, edits, and deletes with full before/after images, plus pre-execution snapshots of risky shell targets (`rm`, `mv`, `cp`, `tee`, `dd`, `truncate`, `shred`, and `>` redirections, globs expanded, best-effort heuristic). Each row carries actor, timestamp, message, true byte sizes, and tags, so history queries time-travel cleanly by path or by tag. Images cap at `AUDIT_MAX_CHARS` (20,000) with an explicit truncation marker rather than silently growing the vault forever, and `restore` flatly refuses to write back a truncated image rather than restore partial content and call it done. Every audited path also keeps a `file`-kind record with its absolute path and latest contents (`FILE_RECORD_MAX`, 50,000 chars, marked when cut).
/audit /etc/hosts history of one path
/audit latest rows across all paths
@@ -447,66 +250,49 @@ absolute path and latest contents (50000 chars, marked when cut).
## Releases and self-backup
The first thing every boot does is back the running script up to
`~/.tai/backups/` as `tai-<version>-<utcstamp>-<sha8>.py`, skipping
when the content hash already has a backup and pruning to the newest
ten. `/release <major|minor|patch> <message>` (or the `release` tool)
cuts a release: it rewrites the `VERSION` line atomically
(temp-plus-rename), snapshots the new script, and logs the message in
the audit trail tagged `release` and `v<version>`. The rule stays
constant: `patch` for fixes with no interface change, `minor` for
backwards-compatible features, `major` for breaking changes.
The very first thing every boot does — before anything else — is back the running script up to `~/.tai/backups/` as `tai-<version>-<utcstamp>-<sha8>.py`, skipping the write entirely when the content hash already has a backup on disk, and pruning down to the newest `BACKUP_KEEP` (10). `/release <major|minor|patch> <message>` (or the `release` tool) cuts an actual release: it rewrites the `VERSION` line atomically (temp file plus rename, never a partial write), snapshots the new script, and logs the message into the audit trail tagged `release` and `v<version>`. The rule stays constant regardless of who's cutting the release: `patch` for fixes with no interface change, `minor` for backwards-compatible features, `major` for breaking changes.
## Voice
`speak` synthesizes free neural speech via the Microsoft Edge Read Aloud
protocol, implemented with `socket` and `ssl` from the standard library. No
key, no package. MP3 files land in `~/.tai/audio` and play when an OS player
exists. `listen` records and transcribes when a recorder (`arecord`, `sox`,
`ffmpeg`) and a transcriber (`whisper-cpp`, `whisper`) are installed, and
reports exactly what is missing otherwise.
`speak` synthesizes free neural speech via the Microsoft Edge Read Aloud protocol, implemented with nothing but `socket` and `ssl` from the standard library — no key, no package, no third-party TTS SDK. MP3 files land in `~/.tai/audio` and play automatically when an OS player exists. `listen` records and transcribes when a recorder (`arecord`, `sox`, `ffmpeg`) and a transcriber (`whisper-cpp`, `whisper`) are both installed, and reports exactly what's missing when they aren't, instead of failing opaquely.
## Terminal output
Assistant replies render as formatted markdown on color terminals: aligned
tables with left, center, and right columns, verbatim fenced code blocks,
nested bullet and numbered lists with checkboxes, headings, blockquotes,
rules, and inline bold, italic, code, strikethrough, and links. Long lines
wrap to the terminal width without breaking styles. Piped output and
`NO_COLOR` stay raw markdown.
Assistant replies render as formatted markdown on color terminals: aligned tables with left/center/right columns, verbatim fenced code blocks, nested bullet and numbered lists with checkboxes, headings, blockquotes, rules, and inline bold/italic/code/strikethrough/links. Long lines wrap to terminal width without ever breaking a style mid-sequence. Piped output and `NO_COLOR` fall back to raw markdown, deliberately.
Shell output streams live while the command runs: the active agent
shows a rolling 4-line window inside the call box, then an
exit line with the return code, elapsed time, line count, and byte
size. Long lines trim to
the terminal width without breaking colors, progress-style output
keeps its latest segment, and the full text still reaches the model
for feedback. Workers, pipes, and capture mode stay silent.
Shell output streams live while a command runs — a rolling 4-line window inside the call box, then an exit line with return code, elapsed time, line count, and byte size. Long lines trim to terminal width without breaking color codes, progress-style output keeps only its latest segment on screen, and the model still receives the full text underneath for real feedback. Workers, pipes, and capture mode stay silent, on purpose.
File changes render as a unified diff right inside the call box:
line numbers, green `+` and red `-` rows on tinted backgrounds,
Python syntax colors, and `···` separators between hunks, capped
at 120 rows with a hidden-line note. New files show all green,
deletes all red, restores diff against current content.
File changes render as a unified diff right inside the call box: line numbers, green `+`/red `-` rows on tinted backgrounds, Python syntax coloring, and `···` separators between hunks, capped at 120 rows with an explicit hidden-line note rather than a wall of scroll. New files show all green, deletes show all red, restores diff against whatever's currently on disk.
## Configuration
| Variable | Purpose | Default |
|---------------------|--------------------------------------|--------------------------------------------|
| `TAI_HOME` | State directory | `~/.tai` |
| `TAI_MODEL` | Model id, ignored by primary gateway | `openrouter/free` |
| `DEVPLACE_API_KEY` | Fallback backend credential | Empty, fallback disabled |
| `TAI_VOICE` | Edge voice name | `en-US-EmmaMultilingualNeural` |
| `TAI_PASSPHRASE` | Personal seal key | Unset: built-in key; empty: plaintext |
| `TELEGRAM_BOT_TOKEN`| Bot token, or set during install | Empty |
| Variable | Purpose | Default |
|----------------------|---------------------------------------|---------------------------------------------|
| `TAI_HOME` | State directory | `~/.tai` |
| `TAI_MODEL` | Model id, ignored by primary gateway | `openrouter/free` |
| `DEVPLACE_API_KEY` | Fallback backend credential | Empty, fallback disabled |
| `TAI_VOICE` | Edge voice name | `en-US-EmmaMultilingualNeural` |
| `TAI_PASSPHRASE` | Personal seal key | Unset: built-in key; empty: plaintext |
| `TELEGRAM_BOT_TOKEN` | Bot token, or set during install | Empty |
## Telegram
./tai.py --install-telegram
./tai.py --uninstall-telegram
Install asks for the bot token up front, verifies it against `getMe`, stores it in `~/.tai/telegram.env` (mode 0600), builds the sandbox container (used for voice transcription), and registers a `tai-telegram.service` systemd user unit with linger enabled (failures ignored). The bot long-polls, answers text, transcribes voice notes, and understands `/new`. It runs without `--yes`, so destructive shell commands are denied outright. Uninstall stops and removes the service and purges the container and image; the token file and data stay put.
## Testing
python3 test_seal.py
python3 test_tai.py
python3 test_tai.py # full agent/tools/vault/scheduler suite
python3 test_seal.py # seal/encryption regression suite
# a single test, either suite, standard unittest addressing:
python3 -m unittest test_tai.CreateSkillTests
python3 -m unittest test_tai.CreateSkillTests.test_creates_and_refreshes
## Layout
tai.py the entire agent
tai.py the entire agent — tools, vault, search, skills, orchestration, sandbox, all of it
test_seal.py seal regression tests
test_tai.py agent, tools, records, graph, scheduler tests
test_tai.py agent, tools, records, graph, search/RAG, skills, scheduler tests
+404 -32
View File
@@ -30,6 +30,7 @@ import urllib.error
import urllib.parse
import urllib.request
import uuid
import zipfile
from datetime import datetime, timedelta, timezone
try:
@@ -291,7 +292,8 @@ HELP_TEXT = (
" /bots list bots, current marked with *\n"
" /bot <name> switch bot, history resumes\n"
" /env [home|sandbox] show or switch execution environment\n"
" /skills list loaded skill files\n"
" /skills list loaded skills with provenance\n"
" /skills install ... install an existing skill (copy or symlink)\n"
" /secret manage sealed secrets (set|list|delete)\n"
" /sysinfo show host environment checks\n"
" /models show model roster with speed health\n"
@@ -357,7 +359,7 @@ def truncate(text, head=1500, tail=500):
SPILL_LIMIT = 6000
RECORD_KINDS = ("note", "output", "file", "research", "transcript")
RECORD_KINDS = ("note", "output", "file", "research", "transcript", "termlog")
AUDIT_MAX_CHARS = 20000
FILE_RECORD_MAX = 50000
SHELL_SNAPSHOT_MAX_FILES = 25
@@ -494,7 +496,10 @@ def segment_targets(tokens):
def spill_output(store, kind, title, text, tags=(), profile=None):
return store.add_record(kind, title, text, tags, profile if profile is not None else store.profile)
profile = profile if profile is not None else store.profile
record_id = store.add_record(kind, title, text, tags, profile)
store.add_chunks("record", record_id, chunk_lines(text), profile)
return record_id
def spilled_view(text, record_id, size, head=1500, tail=500):
@@ -508,6 +513,70 @@ def spill_or_truncate(store, kind, title, text, tags=(), profile=None):
return spilled_view(text, record_id, len(text))
CHUNK_MAX_LINES = 60
CHUNK_OVERLAP_LINES = 3
CHUNK_MAX_CHARS = 4000
def chunk_lines(text, max_lines=CHUNK_MAX_LINES, overlap_lines=CHUNK_OVERLAP_LINES, max_chars=CHUNK_MAX_CHARS):
"""Generic RAG chunker: bounded windows of whole lines with small overlap, never splitting mid-line.
Used for any content that gets indexed for retrieval (terminal captures, long records, ...),
so there is one chunking rule in the codebase instead of one per content type.
"""
lines = text.splitlines()
n = len(lines)
if n == 0:
return []
chunks = []
start = 0
while start < n:
end = start
size = 0
count = 0
while end < n and count < max_lines and size < max_chars:
size += len(lines[end]) + 1
count += 1
end += 1
piece = "\n".join(lines[start:end]).strip()
if piece:
chunks.append(piece)
if end >= n:
break
start = max(start + 1, end - overlap_lines)
return chunks
def collapse_repeated_lines(text, threshold=3):
"""Collapses runs of 3+ identical lines (progress bars, retry spam) to one line + a count note."""
lines = text.splitlines()
out = []
i = 0
n = len(lines)
while i < n:
j = i
while j < n and lines[j] == lines[i]:
j += 1
run = j - i
if run >= threshold:
out.append(lines[i])
out.append("... (repeated %d more times) ..." % (run - 1))
else:
out.extend(lines[i:j])
i = j
return "\n".join(out)
def tmux_pane_header():
try:
info = subprocess.run(["tmux", "display-message", "-p", "#{session_name}:#{window_index}.#{pane_index} #{pane_current_command}"], stdout=subprocess.PIPE, stderr=subprocess.PIPE, universal_newlines=True, timeout=10)
except (OSError, subprocess.SubprocessError):
return ""
if info.returncode == 0 and info.stdout.strip():
return info.stdout.strip()
return ""
VERSION_RE = re.compile(r"^VERSION = \"(\d+)\.(\d+)\.(\d+)\"$", re.MULTILINE)
@@ -1405,11 +1474,74 @@ def parse_skill_file(path):
return {"name": name, "description": description, "body": raw[end + 4:].lstrip("\n")}
EXTERNAL_SKILL_DIRNAMES = (os.path.join(".claude", "skills"), os.path.join(".agents", "skills"))
SKILL_INSTALL_META_DIRNAME = ".installs"
def skill_scope_dir(scope, home_dir, project_dir, name=""):
"""Shared by create_skill and install_skill so both land skills in the same place."""
base = os.path.join(project_dir, ".tai", "skills") if scope == "project" else os.path.join(home_dir, "skills")
return os.path.join(base, name) if name else base
def default_skill_scope(project_dir):
for marker in (".git", ".tai") + EXTERNAL_SKILL_DIRNAMES:
if os.path.exists(os.path.join(project_dir, marker)):
return "project"
return "home"
def skill_extra_files(root):
extras = []
for sub in ("scripts", "references", "assets"):
folder = os.path.join(root, sub)
if os.path.isdir(folder):
for entry in sorted(os.listdir(folder)):
extras.append(os.path.join(root, sub, entry))
return extras
def skill_install_meta_path(scope_dir, entry):
return os.path.join(scope_dir, SKILL_INSTALL_META_DIRNAME, entry + ".json")
def load_skill_install_meta(scope_dir, entry):
try:
with open(skill_install_meta_path(scope_dir, entry), "r", encoding="utf-8") as handle:
return json.load(handle)
except (OSError, ValueError):
return None
def save_skill_install_meta(scope_dir, entry, meta):
path = skill_install_meta_path(scope_dir, entry)
os.makedirs(os.path.dirname(path), exist_ok=True)
with open(path, "w", encoding="utf-8") as handle:
json.dump(meta, handle)
def is_external_skill_root(path):
return os.path.dirname(os.path.normpath(path)).endswith(EXTERNAL_SKILL_DIRNAMES)
def skill_provenance_note(skill):
install = skill.get("install")
if install:
return "[installed %s, %s scope, from %s]" % (install.get("mode", "copy"), install.get("scope", "?"), install.get("source", "?"))
if not skill.get("managed", True):
return "[external, not managed by tai: %s]" % skill["root"]
return "[native]"
def discover_skills(home_dir, project_dir):
found = {}
for base in (os.path.join(home_dir, "skills"), os.path.join(project_dir, ".tai", "skills")):
tai_roots = (os.path.join(home_dir, "skills"), os.path.join(project_dir, ".tai", "skills"))
external_roots = (os.path.join(os.path.expanduser("~"), ".claude", "skills"),) + tuple(os.path.join(project_dir, dirname) for dirname in EXTERNAL_SKILL_DIRNAMES)
# External roots scan first so tai's own dirs always win on a name collision (dict assignment below overwrites).
for base in external_roots + tai_roots:
if not os.path.isdir(base):
continue
managed = base in tai_roots
for entry in sorted(os.listdir(base)):
path = os.path.join(base, entry, "SKILL.md")
if not os.path.isfile(path):
@@ -1421,6 +1553,8 @@ def discover_skills(home_dir, project_dir):
if skill is None:
continue
skill["root"] = os.path.join(base, entry)
skill["managed"] = managed
skill["install"] = load_skill_install_meta(base, entry) if managed else None
found[skill["name"]] = skill
return found
@@ -1658,6 +1792,12 @@ def setup_fts(db):
db.execute("""CREATE TRIGGER IF NOT EXISTS trg_audit_ai AFTER INSERT ON audit BEGIN
INSERT INTO fts_docs(item, kind, sub, title, body, profile) VALUES('audit:' || new.id, 'audit', new.action, new.path, new.message, new.profile);
END""")
db.execute("""CREATE TRIGGER IF NOT EXISTS trg_chunks_ai AFTER INSERT ON chunks BEGIN
INSERT INTO fts_docs(item, kind, sub, title, body, profile) VALUES('chunk:' || new.id, 'chunk', new.parent_kind, new.parent_id, new.text, new.profile);
END""")
db.execute("""CREATE TRIGGER IF NOT EXISTS trg_chunks_ad AFTER DELETE ON chunks BEGIN
DELETE FROM fts_docs WHERE item = 'chunk:' || old.id AND kind = 'chunk';
END""")
if db.execute("SELECT COUNT(*) FROM fts_docs").fetchone()[0] == 0:
db.execute("INSERT INTO fts_docs(item, kind, sub, title, body, profile) SELECT id, 'record', kind, title, content, profile FROM records")
db.execute("INSERT INTO fts_docs(item, kind, sub, title, body, profile) SELECT id, 'event', role || '/' || kind, ts || ' ' || role || '/' || kind, text, profile FROM events")
@@ -1706,6 +1846,7 @@ SCHEMA = (
("tags", ("item TEXT", "tag TEXT", "profile TEXT"), ("item", "tag", "profile")),
("edges", ("src TEXT", "dst TEXT", "relation TEXT", "profile TEXT", "created TEXT"), ("src", "dst", "relation", "profile")),
("audit", ("id INTEGER PRIMARY KEY", "profile TEXT", "ts TEXT", "actor TEXT", "action TEXT", "path TEXT", "message TEXT", "old_size INTEGER", "new_size INTEGER", "old TEXT", "new TEXT", "tags TEXT"), ()),
("chunks", ("id INTEGER PRIMARY KEY", "parent_kind TEXT", "parent_id TEXT", "idx INTEGER", "text TEXT", "profile TEXT", "created TEXT"), ()),
)
SCHEMA_INDEXES = (
@@ -1718,6 +1859,8 @@ SCHEMA_INDEXES = (
("idx_edges_profile", "edges(profile)"),
("idx_records_profile", "records(profile)"),
("idx_audit_profile", "audit(profile)"),
("idx_chunks_parent", "chunks(parent_kind, parent_id)"),
("idx_chunks_profile", "chunks(profile)"),
)
@@ -2094,6 +2237,23 @@ class Store:
self.db.commit()
return record_id
def add_chunks(self, parent_kind, parent_id, texts, profile=None):
"""Replaces any prior chunks for this parent, so re-indexing is idempotent."""
profile = self._scope(profile)
self.db.execute("DELETE FROM chunks WHERE parent_kind = ? AND parent_id = ? AND profile = ?", (parent_kind, parent_id, profile))
stamp = datetime.now(timezone.utc).isoformat()
for idx, text in enumerate(texts):
self.db.execute("INSERT INTO chunks(parent_kind, parent_id, idx, text, profile, created) VALUES (?, ?, ?, ?, ?, ?)", (parent_kind, parent_id, idx, text, profile, stamp))
self.db.commit()
def get_chunks_by_ids(self, ids, profile=None):
if not ids:
return {}
profile = self._scope(profile)
marks = ", ".join("?" * len(ids))
rows = self.db.execute("SELECT id, parent_kind, parent_id, idx, text FROM chunks WHERE profile = ? AND id IN (%s)" % marks, tuple([profile] + list(ids))).fetchall()
return {row[0]: {"parent_kind": row[1], "parent_id": row[2], "idx": row[3], "text": row[4]} for row in rows}
def get_record(self, record_id, profile=None):
profile = self._scope(profile)
row = self.db.execute("SELECT id, kind, title, content, size, reads, created, updated FROM records WHERE profile = ? AND id = ?", (profile, record_id,)).fetchone()
@@ -2170,6 +2330,7 @@ class Store:
if done.rowcount > 0:
self.db.execute("DELETE FROM tags WHERE item = ? AND profile = ?", (record_id, profile))
self.db.execute("DELETE FROM edges WHERE profile = ? AND (src = ? OR dst = ?)", (profile, record_id, record_id))
self.db.execute("DELETE FROM chunks WHERE parent_kind = 'record' AND parent_id = ? AND profile = ?", (record_id, profile))
self.db.commit()
return True
self.db.commit()
@@ -4264,10 +4425,12 @@ TOOL_SCHEMAS = [
tool_schema("recall", "Search past session memory of the current profile by keyword, optionally filtered by tags.", {"query": {"type": "string"}, "tags": {"type": "array"}}, ["query"]),
tool_schema("load_skill", "Load a skill by name. Returns full instructions plus bundled file paths. Missing skills with a blueprint are researched and built on demand.", {"name": {"type": "string"}}, ["name"]),
tool_schema("get_current_terminal_content", "Capture visible text of the current tmux pane including scrollback. Works inside tmux or against a running tmux server.", {"lines": {"type": "integer"}}, []),
tool_schema("export_terminal_log", "Only available inside a live tmux session. Exports the full pane scrollback (tmux capture-pane -pS -, the entire history, not just what's visible) to a log file under tai's home dir, then chunks it and indexes the chunks for retrieval with SQLite FTS5 + bm25 (lexical only, no embeddings). From then on, use the search tool (kind chunk) to answer questions about anything that happened in this terminal; it returns the actual chunk text, not just a snippet.", {"tags": {"type": "array"}}, []),
tool_schema("fork", "Spawn a background subagent with its own context that works while you continue. Returns an agent id immediately. Collect its summarized result with poll. Subagents get a smaller step budget and a time limit.", {"task": {"type": "string"}, "timeout": {"type": "integer"}, "profile": {"type": "string"}}, ["task"]),
tool_schema("poll", "Collect a background subagent result by id. Waits up to wait seconds, then reports running or the result.", {"id": {"type": "integer"}, "wait": {"type": "integer"}}, ["id"]),
tool_schema("sysinfo", "Inspect the host machine where tai runs. Runs os, python, venv, root, container, binaries, cpu, and disk checks in parallel and reports each with timing. Pass checks to run a named subset.", {"checks": {"type": "array"}}, []),
tool_schema("create_skill", "Create a new agent skill by name from a brief. Runs a dedicated deep-research worker on this machine (it calls sysinfo first, then researches with web search and shell until it has verified information) and writes SKILL.md plus supporting files into the skill directory. Runs synchronously and can take many minutes. Scope project writes under ./.tai/skills, scope home under ~/.tai/skills.", {"name": {"type": "string"}, "brief": {"type": "string"}, "scope": {"type": "string"}, "timeout": {"type": "integer"}}, ["name", "brief"]),
tool_schema("create_skill", "Create a new agent skill by name from a brief. Runs a dedicated deep-research worker on this machine (it calls sysinfo first, then researches with web search and shell until it has verified information) and writes SKILL.md plus supporting files into the skill directory. Runs synchronously and can take many minutes. Scope project writes under ./.tai/skills (the default), scope home under ~/.tai/skills.", {"name": {"type": "string"}, "brief": {"type": "string"}, "scope": {"type": "string"}, "timeout": {"type": "integer"}}, ["name", "brief"]),
tool_schema("install_skill", "Install a skill that already exists somewhere else, instead of building one from scratch. source can be a local directory containing SKILL.md (or one subdirectory that does), a bare SKILL.md file, a local .zip, or a git URL (needs git installed; most sources need nothing beyond the standard library). Validates the SKILL.md frontmatter, then lands it under tai's own skill directories: copy vendors a private snapshot, symlink points at the source in place (the default when the source already sits in a recognized skills folder like .claude/skills or .agents/skills, so there is one canonical copy instead of a drifting duplicate). Scope defaults like create_skill: project when the cwd looks like a project, else home. Overwriting an existing skill of the same name asks the user first unless force is set.", {"source": {"type": "string"}, "name": {"type": "string"}, "scope": {"type": "string"}, "mode": {"type": "string"}, "force": {"type": "boolean"}}, ["source"]),
tool_schema("store_secret", "Store a password, token, or secret in the sealed vault under a name. Values are encrypted at rest, never shown back, and only usable by reference; prefer /secret set in the REPL so the value never enters the conversation. Optional username, host, port, notes, expires (ISO datetime), and tags describe it; name and value suffice, never pester the user for more.", {"name": {"type": "string"}, "value": {"type": "string"}, "username": {"type": "string"}, "host": {"type": "string"}, "port": {"type": "integer"}, "notes": {"type": "string"}, "expires": {"type": "string"}, "tags": {"type": "array"}}, ["name", "value"]),
tool_schema("list_secrets", "List vault secret names with metadata and tags. Values are never revealed.", {}, []),
tool_schema("delete_secret", "Delete a vault secret by name. Asks the user first.", {"name": {"type": "string"}}, ["name"]),
@@ -4287,13 +4450,25 @@ TOOL_SCHEMAS = [
tool_schema("tags", "List vault tags with usage counts, most used first. Consult before tagging so new items reuse established tags instead of inventing synonyms.", {"prefix": {"type": "string"}, "limit": {"type": "integer"}}, []),
tool_schema("install", "Manage tai installations: binary, bash command-not-found hook, venv, scheduler service, telegram service, container. Status reports what exists. Install adds missing pieces, upgrade refreshes in place and restarts services, reinstall rebuilds artifacts, uninstall removes them. Service changes take effect immediately. The vault is never touched by any action. Non-status actions ask the user first.", {"action": {"type": "string"}, "targets": {"type": "array"}}, ["action"]),
tool_schema("create_bot", "Create a bot in the current profile: a name plus its own system message and history. Give at least one of description, rules, or behavior; optional nicknames register as @mention aliases (a short form of the name registers automatically). Names are lowercase.", {"name": {"type": "string"}, "description": {"type": "string"}, "rules": {"type": "string"}, "behavior": {"type": "string"}, "nicknames": {"type": "array"}}, ["name"]),
tool_schema("search", "Search everything at once with ranked full-text matching: records, episodic events, and the audit trail. One call replaces paging through each store separately. Returns snippets ordered by relevance; expand adds one graph hop per record hit.", {"query": {"type": "string"}, "kinds": {"type": "array"}, "limit": {"type": "integer"}, "expand": {"type": "boolean"}}, ["query"]),
tool_schema("search", "Search everything at once with ranked full-text matching (SQLite FTS5 + bm25, no vectors): records, episodic events, the audit trail, and RAG chunks. Chunks are the retrieval-grade granular hits (e.g. one slice of an exported terminal log) and come back with their actual text, not just a snippet; use the chunk's parent id with record_read to see more surrounding context. One call replaces paging through each store separately. expand adds one graph hop per record hit.", {"query": {"type": "string"}, "kinds": {"type": "array"}, "limit": {"type": "integer"}, "expand": {"type": "boolean"}}, ["query"]),
]
CORE_TOOLS = ("shell", "read_file", "write_file", "remember", "recall", "load_skill", "fork", "poll")
LAZY_NAME_STOP = ("get", "set", "list", "add")
# Tools gated on top of the lazy keyword match: a gate must pass before a tool can even be
# offered to the model, regardless of relevance. Used for tools that only make sense in a
# specific environment (e.g. a tmux pane to capture).
TOOL_ENV_GATES = {
"export_terminal_log": lambda: bool(os.environ.get("TMUX")),
}
def tool_env_ok(name):
gate = TOOL_ENV_GATES.get(name)
return gate is None or gate()
TOOL_TAGS = {
"shell": ("run", "execute", "command", "bash", "terminal", "script"),
"read_file": ("read", "open", "view", "contents"),
@@ -4307,10 +4482,12 @@ TOOL_TAGS = {
"recall": ("recall", "memory", "history", "past"),
"load_skill": ("skill", "capability"),
"get_current_terminal_content": ("terminal", "tmux", "pane", "scrollback", "screen"),
"export_terminal_log": ("terminal", "tmux", "export", "log", "capture", "history", "rag", "index"),
"fork": ("fork", "subagent", "background", "parallel", "delegate", "spawn"),
"poll": ("poll", "collect"),
"sysinfo": ("sysinfo", "system", "host", "machine", "hardware", "specs", "installed", "python", "container"),
"create_skill": ("skill", "create", "author", "blueprint"),
"install_skill": ("skill", "install", "import", "adopt", "copy", "symlink"),
"store_secret": ("secret", "password", "token", "credential", "passwd", "vault", "apikey"),
"list_secrets": ("secrets", "vault", "credentials"),
"delete_secret": ("secret", "remove", "revoke"),
@@ -4338,6 +4515,8 @@ def tool_catalog():
lines = ["", "", "## Tool catalog (core loads always; name a lazy tool or topic to load it)"]
for schema in TOOL_SCHEMAS:
name = schema["function"]["name"]
if not tool_env_ok(name):
continue
if name in CORE_TOOLS:
continue
first = schema["function"]["description"].split(".")[0][:80]
@@ -4359,6 +4538,8 @@ def select_tools(text):
picked = []
for schema in TOOL_SCHEMAS:
name = schema["function"]["name"]
if not tool_env_ok(name):
continue
if name in CORE_TOOLS:
picked.append(schema)
continue
@@ -4467,10 +4648,12 @@ class Tools:
"recall": self.run_recall,
"load_skill": self.run_load_skill,
"get_current_terminal_content": self.run_terminal_content,
"export_terminal_log": self.run_export_terminal_log,
"fork": self.run_fork,
"poll": self.run_poll,
"sysinfo": self.run_sysinfo,
"create_skill": self.run_create_skill,
"install_skill": self.run_install_skill,
"store_secret": self.run_store_secret,
"list_secrets": self.run_list_secrets,
"delete_secret": self.run_delete_secret,
@@ -4974,14 +5157,8 @@ class Tools:
known = ", ".join(sorted(self.app.skills)) or "none"
return "error: unknown skill, known: " + known
return "building skill '%s' from blueprint, this runs deep research and takes a while:\n%s" % (name, self.run_create_skill({"name": name, "brief": blueprint["brief"], "scope": blueprint["scope"]}))
parts = [skill["body"].strip()]
parts = [skill["body"].strip()]
extras = []
for sub in ("scripts", "references", "assets"):
folder = os.path.join(skill["root"], sub)
if os.path.isdir(folder):
for entry in sorted(os.listdir(folder)):
extras.append(os.path.join(skill["root"], sub, entry))
parts = [skill["body"].strip(), skill_provenance_note(skill)]
extras = skill_extra_files(skill["root"])
if extras:
parts.append("bundled files:\n" + "\n".join(extras))
return "\n\n".join(parts)
@@ -4993,13 +5170,8 @@ class Tools:
lines = 100
if not shutil.which("tmux"):
return "tmux not available"
header = []
try:
info = subprocess.run(["tmux", "display-message", "-p", "#{session_name}:#{window_index}.#{pane_index} #{pane_current_command}"], stdout=subprocess.PIPE, stderr=subprocess.PIPE, universal_newlines=True, timeout=10)
if info.returncode == 0 and info.stdout.strip():
header.append(info.stdout.strip())
except (OSError, subprocess.SubprocessError):
pass
pane_header = tmux_pane_header()
header = [pane_header] if pane_header else []
try:
done = subprocess.run(["tmux", "capture-pane", "-p", "-S", "-%d" % lines], stdout=subprocess.PIPE, stderr=subprocess.PIPE, universal_newlines=True, timeout=10)
except (OSError, subprocess.SubprocessError) as exc:
@@ -5010,6 +5182,43 @@ class Tools:
return "terminal pane is empty"
return "\n".join(header + [truncate(done.stdout, 6000, 2000)])
def run_export_terminal_log(self, args):
if not os.environ.get("TMUX"):
return "error: tmux is not active (this tool only works inside a live tmux session)"
if not shutil.which("tmux"):
return "error: tmux not available"
extra_tags = args.get("tags") or []
if not isinstance(extra_tags, list) or any(not isinstance(item, str) for item in extra_tags):
return "error: tags must be a list of strings"
try:
done = subprocess.run(["tmux", "capture-pane", "-p", "-S", "-"], stdout=subprocess.PIPE, stderr=subprocess.PIPE, universal_newlines=True, timeout=30)
except (OSError, subprocess.SubprocessError) as exc:
return "error: " + short_error(exc)
if done.returncode != 0:
return "error: tmux capture failed: " + (done.stderr or "").strip()[:200]
raw = done.stdout
if not raw.strip():
return "terminal pane history is empty"
profile = self.active_profile
pane_header = tmux_pane_header() or "pane"
stamp = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ")
slug = re.sub(r"[^A-Za-z0-9_.-]+", "-", pane_header).strip("-") or "pane"
folder = os.path.join(self.app.config.home, "termlogs", profile)
try:
os.makedirs(folder, exist_ok=True)
path = os.path.join(folder, "%s-%s.log" % (slug, stamp))
with open(path, "w", encoding="utf-8") as handle:
handle.write(raw)
except OSError as exc:
return "error: " + short_error(exc)
if self.app.store is None:
return "captured %s to %s (%d bytes); vault unavailable, not indexed for search" % (pane_header, path, len(raw.encode("utf-8")))
title = "tmux %s captured %s" % (pane_header, stamp)
record_id = self.app.store.add_record("termlog", title, raw, ["termlog"] + extra_tags, profile)
chunks = chunk_lines(collapse_repeated_lines(raw))
self.app.store.add_chunks("record", record_id, chunks, profile)
return "captured %s to %s (%d bytes), indexed as %s in %d chunks -- use search (kind chunk) to answer questions about this terminal session" % (pane_header, path, len(raw.encode("utf-8")), record_id, len(chunks))
def run_fork(self, args):
task = str(args.get("task") or "").strip()
if not task:
@@ -5048,6 +5257,7 @@ class Tools:
if not spilled:
task = str(record.get("task") or "")[:80]
spilled = self.app.store.add_record("output", "agent %d: %s" % (agent_id, task), full, ["agent-output"], self.active_profile)
self.app.store.add_chunks("record", spilled, chunk_lines(full), self.active_profile)
with AGENTS_LOCK:
if agent_id in AGENTS:
AGENTS[agent_id]["spilled"] = spilled
@@ -5082,10 +5292,7 @@ class Tools:
timeout = max(60, min(3600, int(args.get("timeout") or 1200)))
except (TypeError, ValueError):
timeout = 1200
if scope == "project":
skill_dir = os.path.join(os.getcwd(), ".tai", "skills", name)
else:
skill_dir = os.path.join(self.app.config.home, "skills", name)
skill_dir = skill_scope_dir(scope, self.app.config.home, os.getcwd(), name)
skill_file = os.path.join(skill_dir, "SKILL.md")
existed = os.path.isfile(skill_file)
spinner = Spinner("creating skill %s" % name)
@@ -5105,6 +5312,133 @@ class Tools:
status = "timeout" if "[time limit reached]" in result else "done"
return "skill '%s' was not created (%s), worker output:\n%s" % (name, status, truncate(result, 3000, 1000))
@staticmethod
def find_skill_root(base):
"""A source dir qualifies if it's SKILL.md itself, or has exactly one child that is."""
if os.path.isfile(os.path.join(base, "SKILL.md")):
return base
try:
entries = sorted(os.listdir(base))
except OSError:
return None
candidates = [entry for entry in entries if os.path.isfile(os.path.join(base, entry, "SKILL.md"))]
return os.path.join(base, candidates[0]) if len(candidates) == 1 else None
def resolve_skill_source(self, source, workdir):
"""Returns ((path, ephemeral), '') on success, (None, 'error: ...') otherwise.
path is a directory containing SKILL.md, or a bare SKILL.md file. ephemeral
means the path lives under workdir (git clone / zip extraction) and symlinking
to it would dangle once workdir is cleaned up, so install must copy it.
"""
if source.endswith(".git") or source.startswith("git@") or source.startswith("git+"):
git_url = source[4:] if source.startswith("git+") else source
if not shutil.which("git"):
return None, "error: git is not installed, supply a local path or .zip instead"
clone_dir = os.path.join(workdir, "clone")
try:
done = subprocess.run(["git", "clone", "--depth", "1", git_url, clone_dir], stdout=subprocess.PIPE, stderr=subprocess.PIPE, universal_newlines=True, timeout=120)
except (OSError, subprocess.SubprocessError) as exc:
return None, "error: git clone failed: " + short_error(exc)
if done.returncode != 0:
return None, "error: git clone failed: " + (done.stderr or "").strip()[:300]
found = self.find_skill_root(clone_dir)
if found is None:
return None, "error: no SKILL.md found in " + git_url
return (found, True), ""
path = os.path.expanduser(source)
if not os.path.exists(path):
return None, "error: no such file or directory: " + source
if os.path.isfile(path) and path.endswith(".zip"):
extract_dir = os.path.join(workdir, "zip")
try:
with zipfile.ZipFile(path) as archive:
archive.extractall(extract_dir)
except (OSError, zipfile.BadZipFile) as exc:
return None, "error: bad zip file: " + short_error(exc)
found = self.find_skill_root(extract_dir)
if found is None:
return None, "error: no SKILL.md found in " + source
return (found, True), ""
if os.path.isfile(path) and path.endswith(".md"):
return (path, False), ""
if os.path.isdir(path):
found = self.find_skill_root(path)
if found is None:
return None, "error: no SKILL.md found under " + source
return (found, False), ""
return None, "error: source must be a directory, a SKILL.md file, a .zip, or a git URL"
def run_install_skill(self, args):
source = str(args.get("source") or "").strip()
if not source:
return "error: empty source"
scope = str(args.get("scope") or "").strip().lower() or default_skill_scope(os.getcwd())
if scope not in ("project", "home"):
return "error: scope must be project or home"
mode = str(args.get("mode") or "").strip().lower() or None
if mode is not None and mode not in ("copy", "symlink"):
return "error: mode must be copy or symlink"
force = bool(args.get("force"))
workdir = tempfile.mkdtemp(prefix="tai-skill-")
try:
spinner = Spinner("installing skill from %s" % source)
spinner.start()
try:
resolved, failure = self.resolve_skill_source(source, workdir)
finally:
spinner.stop()
if failure:
return failure
real_source, ephemeral = resolved
source_dir = real_source if os.path.isdir(real_source) else None
skill_file = os.path.join(real_source, "SKILL.md") if source_dir else real_source
try:
skill = parse_skill_file(skill_file)
except OSError as exc:
return "error: " + short_error(exc)
if skill is None:
return "error: %s has no valid SKILL.md (needs name, description, and a name of lowercase letters, digits, hyphens)" % source
name = str(args.get("name") or "").strip() or skill["name"]
if not SKILL_NAME_RE.match(name):
return "error: invalid skill name, use lowercase letters, digits, and hyphens"
if mode is None:
mode = "symlink" if (source_dir and not ephemeral and is_external_skill_root(source_dir)) else "copy"
elif mode == "symlink" and (ephemeral or source_dir is None):
return "error: symlink mode needs a persistent source directory, not a git/zip source or a bare SKILL.md file"
scope_dir = skill_scope_dir(scope, self.app.config.home, os.getcwd())
dest = os.path.join(scope_dir, name)
existed = os.path.islink(dest) or os.path.exists(dest)
if existed and not force and not self.app.ask_approval("overwrite existing skill '%s' at %s" % (name, dest)):
return "denied by user"
if existed:
if os.path.islink(dest) or os.path.isfile(dest):
os.remove(dest)
else:
shutil.rmtree(dest)
os.makedirs(scope_dir, exist_ok=True)
if mode == "symlink":
os.symlink(os.path.abspath(source_dir), dest)
elif source_dir is not None:
shutil.copytree(source_dir, dest)
else:
os.makedirs(dest)
shutil.copy2(real_source, os.path.join(dest, "SKILL.md"))
save_skill_install_meta(scope_dir, name, {"source": source, "mode": mode, "scope": scope, "installed_at": datetime.now(timezone.utc).isoformat()})
if self.app.store is not None:
try:
self.app.store.audit_event(self.audit_actor(), "skill-install", dest, "installed skill '%s' from %s (%s, %s scope)" % (name, source, mode, scope), tags=["skill", "install"])
except (sqlite3.Error, OSError):
pass
self.app.skills = discover_skills(self.app.config.home, os.getcwd())
self.app.apply_system()
extras = skill_extra_files(dest)
tail = "\nbundled files: %s" % ", ".join(extras) if extras else ""
replaced = " (replacing the previous install)" if existed else ""
return "skill '%s' installed (%s, %s scope) at %s from %s%s\n%s%s" % (name, mode, scope, dest, source, replaced, skill["description"], tail)
finally:
shutil.rmtree(workdir, ignore_errors=True)
def run_store_secret(self, args):
name = str(args.get("name") or "").strip()
value = str(args.get("value") or "")
@@ -5260,7 +5594,9 @@ class Tools:
if tags is None:
return "error: tags must be a list of names"
record_id = self.app.store.add_record(kind, title, content, tags, self.active_profile)
saved = "saved %s (%d chars, kind %s, tags: %s)" % (record_id, len(content), kind, ", ".join(self.app.store.item_tags(record_id, self.active_profile)))
chunks = chunk_lines(content)
self.app.store.add_chunks("record", record_id, chunks, self.active_profile)
saved = "saved %s (%d chars, kind %s, %d chunks, tags: %s)" % (record_id, len(content), kind, len(chunks), ", ".join(self.app.store.item_tags(record_id, self.active_profile)))
links = sorted(edge["other"] for edge in self.app.store.edges_for(record_id, self.active_profile) if edge["direction"] == "out")
if links:
saved += " [linked: %s]" % ", ".join(links)
@@ -5330,9 +5666,9 @@ class Tools:
query = str(args.get("query") or "").strip()
if not query:
return "error: empty query"
kinds = args.get("kinds") or ["record", "event", "audit"]
if not isinstance(kinds, list) or not kinds or any(item not in ("record", "event", "audit") for item in kinds):
return "error: kinds must be a list of record, event, audit"
kinds = args.get("kinds") or ["record", "event", "audit", "chunk"]
if not isinstance(kinds, list) or not kinds or any(item not in ("record", "event", "audit", "chunk") for item in kinds):
return "error: kinds must be a list of record, event, audit, chunk"
try:
limit = int(args.get("limit") or 10)
except (TypeError, ValueError):
@@ -5346,6 +5682,14 @@ class Tools:
hits.sort(key=lambda hit: hit["rank"])
if not hits:
return "no matches for '%s'" % query[:80]
chunk_ids = []
for hit in hits:
if hit["kind"] == "chunk":
try:
chunk_ids.append(int(str(hit["item"]).split(":", 1)[1]))
except (IndexError, ValueError):
pass
chunk_rows = self.app.store.get_chunks_by_ids(chunk_ids, profile) if chunk_ids else {}
lines = []
for hit in hits:
if hit["kind"] == "record":
@@ -5359,6 +5703,17 @@ class Tools:
lines.append(" linked: %s" % "; ".join(neighbors))
elif hit["kind"] == "event":
lines.append("[event/%s] #%s %s -- %s" % (hit["sub"], hit["item"], hit["title"][:60], hit["snippet"]))
elif hit["kind"] == "chunk":
try:
chunk_id = int(str(hit["item"]).split(":", 1)[1])
except (IndexError, ValueError):
chunk_id = None
chunk = chunk_rows.get(chunk_id)
if chunk is None:
lines.append("[chunk/%s] %s -- %s" % (hit["sub"], hit["title"][:80], hit["snippet"]))
else:
preview = chunk["text"][:500] + ("..." if len(chunk["text"]) > 500 else "")
lines.append("[chunk/%s] %s#%d (parent %s, read full with record_read) -- %s" % (chunk["parent_kind"], chunk["parent_id"][:60], chunk["idx"], chunk["parent_id"], preview))
else:
lines.append("[audit/%s] #%s %s -- %s" % (hit["sub"], str(hit["item"]).split(":", 1)[1], hit["title"][:80], hit["snippet"]))
return "\n".join(lines)
@@ -6293,12 +6648,28 @@ def handle_command(agent, text):
agent.env = "home"
print("environment: home")
elif name == "skills":
bits = arg.split()
if bits and bits[0] == "install":
if len(bits) < 2:
print(paint("use /skills install <source> [name=x] [scope=project|home] [mode=copy|symlink] [force]", Ansi.RED))
return True
call_args = {"source": bits[1]}
for token in bits[2:]:
if token == "force":
call_args["force"] = True
elif "=" in token:
key, _, value = token.partition("=")
if key in ("name", "scope", "mode"):
call_args[key] = value
print(agent.tools.dispatch("install_skill", json.dumps(call_args)))
return True
agent.skills = discover_skills(agent.config.home, os.getcwd())
agent.apply_system()
if not agent.skills:
print("(no skills)")
for skill_name in sorted(agent.skills):
print("- %s: %s" % (skill_name, agent.skills[skill_name]["description"][:200]))
skill = agent.skills[skill_name]
print("- %s: %s %s" % (skill_name, skill["description"][:200], skill_provenance_note(skill)))
wanted = sorted(item for item in SKILL_BLUEPRINTS if item not in agent.skills)
if wanted:
print("blueprints, built on first load:")
@@ -6399,8 +6770,9 @@ def handle_command(agent, text):
print(match[0]["description"])
print("trigger tags: %s" % ", ".join(TOOL_TAGS.get(match[0]["name"], ())))
return True
cores = sorted(tool for tool in (schema["function"]["name"] for schema in TOOL_SCHEMAS) if tool in CORE_TOOLS)
lazy = sorted(tool for tool in (schema["function"]["name"] for schema in TOOL_SCHEMAS) if tool not in CORE_TOOLS)
names = [schema["function"]["name"] for schema in TOOL_SCHEMAS if tool_env_ok(schema["function"]["name"])]
cores = sorted(tool for tool in names if tool in CORE_TOOLS)
lazy = sorted(tool for tool in names if tool not in CORE_TOOLS)
print("core, always loaded:\n %s" % "\n ".join(cores))
print("lazy, loads when named:\n %s" % "\n ".join(lazy))
elif name == "search":