- install_skill: adopt an existing skill from a local dir, SKILL.md, .zip, or git URL; symlinks when the source already lives in a recognized external skills folder (.claude/skills, .agents/skills) instead of vendoring a duplicate, with install provenance recorded per skill. Skill discovery now also reads those external folders read-only, so skills placed by other tools are visible with no install step at all. - export_terminal_log: full tmux scrollback export, visible only inside a live tmux session (TOOL_ENV_GATES, a new env-conditional layer on top of the existing keyword-based lazy tool loading). - Chunked lexical RAG (SQLite FTS5 + bm25, no vectors/embeddings): new chunks table wired into record_save, shell/subagent output spill, and terminal log export, searchable via the existing search tool (kind chunk) with full chunk text returned, not just a snippet. delete_record cleans up a record's chunks with it. - Fix a regression where create_skill's default scope silently became context-dependent instead of always "project". - Rewrite README.md to document all of the above plus the existing context/tool-selection and search/memory internals in depth. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
36 KiB
retoor retoor@molodetz.nl
tai
One file. Zero dependencies. A real agent.
tai is a single-file autonomous AI agent written in Python — the entire implementation lives in tai.py (~7,600 lines) and imports only the Python standard library. No pip install, no requirements.txt, no build step, no lockfile to rot. It runs as an interactive REPL or a one-shot command, reasons through an OpenAI-compatible backend, acts through 38 tools, keeps per-profile memory in a self-healing SQLite vault, seals that vault at rest, ranks every search with real BM25 (no vector database, no embeddings bill, no GPU), chunks and indexes anything worth remembering for retrieval, and can wall shell/file work off inside a disposable container sandbox.
Nothing here is decorative. Every mechanism below exists because something specific was expensive — context tokens, API calls, attacker surface, disk writes, or your attention — and got engineered out.
Requirements
- Python 3.10 or newer, no third-party packages.
- Optional:
podmanordockerfor the sandbox and Telegram voice notes. - Optional:
tmuxfor terminal content capture and full-history export. - Optional:
gitonly if you install a skill straight from a git URL — every other skill/install path is pure stdlib.
Quick start
./tai.py
./tai.py --profile work
./tai.py --yes
./tai.py --yolo
./tai.py --auto
./tai.py --version
./tai.py what is 2+3, use the shell
Trailing arguments form a one-shot prompt: the agent answers once and exits with code 0. Without arguments, tai starts an interactive session.
REPL commands
| Command | Effect |
|---|---|
/profile [name] |
Show the current profile or switch to it |
/profiles |
List all profiles, current marked with * |
/bots |
List profile bots, current marked with * |
/bot [name] |
Switch to another bot, history resumes |
/env [target] |
Show or switch execution environment |
/skills |
List loaded skills with provenance (native/copied/symlinked/external) |
/skills install ... |
Install an existing skill by copy or symlink, no scaffolding |
/secret |
Manage sealed secrets (set|list|delete) |
/sysinfo |
Show host environment checks |
/schedule |
Schedule a prompt for later (at|every) |
/schedules |
List scheduled prompts and outcomes |
/unschedule |
Delete a scheduled prompt by id |
/records |
Search saved records, optional query |
/record <id> |
Read one record page by mem id |
/graph <node> |
Show one vault node neighborhood |
/tags [prefix] |
List tags with usage counts |
/tools [name] |
Show core and lazy tools, respecting environment gates |
/install ... |
Install status, install, upgrade, reinstall |
/search <query> |
Ranked search over records, events, audit, and chunks |
/audit [path] |
Show file audit trail |
/restore <id> |
Restore a file from an audit row |
/release |
Bump version, back up, log message |
/fork <task> |
Spawn a background subagent, REPL stays free |
/agents |
List background subagents |
/agent <id> |
Show one subagent result |
/agent clear |
Purge finished subagents |
/compact |
Compress history into a summary |
/clear |
Drop history, keep the system message |
/help |
Show the command overview |
/quit |
Exit |
Any other input is sent to the agent.
Installation
./tai.py --install
/install status
/install upgrade scheduler-service
Six install targets exist side by side: binary (~/.local/bin/tai.py), bash-hook (a guarded command_not_found_handle block in ~/.bashrc, backed up once to ~/.bashrc.bak-tai, so unknown shell commands are answered by the agent), venv (~/.tai/venv, created whenever the venv module exists), scheduler-service and telegram-service (systemd user units that prefer the venv python), and container (the sandbox, whose image carries its own /box/venv). /install status (or the install tool) reports exactly which of these exist, with versions and service states. install adds missing pieces, upgrade refreshes in place, reinstall rebuilds artifacts from scratch, uninstall removes them; service changes restart or start units immediately. Data handling is explicit: no action ever touches the vault (memory.db, secrets, schedules, records, backups).
Backends
The primary backend is model.cloud.pravda.education, an OpenAI-compatible gateway that needs no API key and selects a free model per request. If a request fails, tai retries it on devplace.net/openai/v1, which requires DEVPLACE_API_KEY. Both endpoints speak /chat/completions, including native tool calls and streaming.
The context problem, and how tai solves it
Every tool schema you hand an LLM costs real tokens on every single turn, whether or not it gets used, and a model drowning in fifty tool definitions picks worse than one offered the five that actually matter. tai treats this as a first-class engineering problem, not an afterthought, with three independent layers stacked on top of each other:
1. Core vs. lazy. Of the 38 tools, only CORE_TOOLS — shell, read_file, write_file, remember, recall, load_skill, fork, poll — are sent to the model on every turn. The other 30 are lazy: their full JSON schema only enters the request payload when they're judged relevant to what's actually being discussed.
2. Relevance by three independent signals, not a fixed list. select_tools() scans the last 6000 characters of conversation (user, assistant, and tool-result turns — conversation_text()) and pulls in a lazy tool's full schema the moment any one of three things matches:
- a name token —
record_searchoffers itself up the moment "record" or "search" appears, filtered throughLAZY_NAME_STOP(get,set,list,add) so generic verbs don't trigger everything; - a trigger tag — every tool carries a hand-picked synonym list in
TOOL_TAGS(cronloadsschedule,undoloadsrestore,passwordorapikeyloadsstore_secret,symlinkloadsinstall_skill), so the model never has to know a tool's literal function name to reach for it; - a description word (6+ characters) — a last-resort fallback so nothing is ever truly unreachable just because its tags happened to miss the exact phrasing.
Tool results feed this loop too: a spill pointer that mentions record_read loads that tool for the very next step, so the model is never stuck knowing a capability exists but unable to reach it.
3. Environment gates, orthogonal to relevance. TOOL_ENV_GATES is a second filter that runs before relevance is even checked. export_terminal_log is gated on bool(os.environ.get("TMUX")) — outside a live tmux session it is invisible in the system prompt, invisible in select_tools(), invisible in /tools's no-arg listing, full stop, regardless of how insistently you ask for it. The handler re-checks the same condition at call time anyway, so even a stale tool-call from earlier context can't slip through. This is the general mechanism for "this capability is nonsensical outside environment X" — a second axis entirely separate from keyword relevance.
A compact name-plus-one-line-summary catalog of every lazy tool (tool_catalog()) still rides in the system prompt at all times, so nothing is ever hidden — only its verbose, multi-field JSON schema is deferred. /tools shows the core/lazy split live (respecting the same environment gates a real turn would see), /tools <name> shows one tool with its full description and trigger tags.
Net effect: the model sees 8 schemas by default instead of 38, picks up exactly the ones a conversation actually needs, and never has a tool pushed on it that can't possibly apply in the current environment.
Context budget and compaction
Raw history isn't free either. Token count is estimated every turn (estimate_tokens, 4 chars ≈ 1 token) against a CONTEXT_CAP of 32,000. The instant usage crosses 80% of that (COMPACT_RATIO), compact_messages() fires automatically: it keeps the system message, keeps the last 6 turns verbatim (KEEP_TURNS, snapped backward to the nearest user turn so a conversation never resumes mid-answer), and fires everything in between — including a readable digest of which tool calls happened, not just their text — at a dedicated, separate LLM call whose only job is lossy compression (COMPACT_SYSTEM). The summary replaces the gutted middle. /compact runs the exact same path on demand. The system message itself is independently capped at SYSTEM_MAX_CHARS (12,000) so remember can't let it grow without bound either.
Spill-to-vault: the other half of context control
Tool output is the other place context explodes. Any shell command or subagent poll result over SPILL_LIMIT (6000 characters) never reaches the model whole: spill_or_truncate() writes it to the vault as a tagged record instead and hands back a head/tail preview plus a mem:<id> pointer (spilled_view()). The model pages the rest back deliberately with record_read, offset and all, instead of having 40KB of build log forced into every subsequent turn of the conversation. And — see the RAG section below — that spilled record doesn't just sit there as a cold blob: it gets chunked and indexed the instant it's written, so "what was that error again" can be answered by search instead of by re-reading the whole thing.
Keyword storage, tagging, and the knowledge graph
Every record, secret, and schedule in the vault carries normalized tags, and tai is deliberately strict about vocabulary so it doesn't fragment into a thousand near-duplicate labels:
- tags are lowercased, hyphenated, capped at 32 characters (
normalize_tag), and a single item never carries more than 20 raw / effectively 5 auto-detected tags; - a small set of genuinely irregular plurals (
news,means,series,species,physics) is hard-excluded from singularization (SINGULAR_KEEP); everything else runs through a real little inflection engine (singular_noun):-ies → -y,-ses/-xes/-zes/-ches/-shes → drop "es", trailing-sdropped unless it's-ss/-us/-is/-os; canonical_tag()always prefers whatever form already exists in that profile's vocabulary over whatever form you typed, so "skills" and "skill" collapse to whichever one got there first, instead of silently forking the taxonomy;- the
tagstool lists every known tag with usage counts — the explicit, correct way to reuse vocabulary instead of inventing synonyms, and the one both humans and the model are told to consult first.
Auto-tagging is not keyword extraction, it's vocabulary matching. auto_tags_for() scans a new record's title-plus-content for words (and adjacent word-pairs, so multi-word tags like api-client still fire) that are already real tags in that profile — up to TAG_AUTO_MAX (5) — and attaches them automatically. A brand-new vocabulary word never gets invented this way; only previously-established meaning spreads to new content, which is what keeps the tag space from drifting into noise over time.
The graph builds itself. Every new record automatically links to the TAG_LINK_MAX (3) most recently updated records sharing a real tag (tag_neighbors(), explicitly excluding the structural record/file tags so it's never self-referential noise) — an edge with relation shares-<tag> lands in the edges table with zero manual curation. graph_link lets you add arbitrary typed edges between any two vault nodes (mem:<hash> records, secret:<name>, sched:<id>), and graph_query walks the neighborhood breadth-first, capped at depth 4 so a query can never explode into the whole vault. search's expand flag appends each record hit's nearest neighbors inline, for free.
Search and RAG — BM25 and chunking, deliberately no vectors
This is the part worth being loud about: tai does retrieval-augmented generation with zero embeddings, zero vector database, and zero extra dependency, on purpose.
One unified lexical index. fts_docs is a single SQLite FTS5 virtual table covering records, events, audit, and chunks at once, kept in sync not by any application-level reindex job but by SQL triggers fired directly on insert/update/delete (setup_fts()). Query terms are tokenized safely server-side — no raw FTS MATCH syntax ever reaches SQLite from user input — stemmed by the Porter tokenizer when available, and ranked with SQLite's own built-in bm25(). search, recall, record_search, and audit all rank through this one index instead of each reinventing matching, falling back to plain LIKE only when a query has no tokens FTS can use.
Chunking, for retrieval-grade granularity, without inventing a second storage system. A document-level FTS hit tells you a 40KB record contains your answer somewhere; it doesn't hand you the answer. chunk_lines() splits any text into bounded, overlapping windows — up to 60 lines or 4000 characters, whichever comes first, with a 3-line overlap carried into the next chunk — and critically never splits mid-line, so a chunk boundary never severs a command from the middle of its own output. Every chunk lands in the chunks table and gets its own fts_docs row (kind='chunk'), ranked by the exact same BM25 engine as everything else — one index, one ranking function, no separate "vector similarity vs. keyword score" fusion problem to solve, because there's only ever one kind of score. collapse_repeated_lines() runs first on anything log-shaped, folding 3+ identical consecutive lines (progress-bar spam, retry loops) into one line plus a count note, so pathological repetition doesn't burn chunk budget on noise.
A chunk hit from search returns the actual chunk text, not just FTS's highlighted snippet — the whole point of chunk-level retrieval is to hand back something directly usable as context, and the parent record id rides along so record_read can pull more surrounding material on demand (hierarchical retrieval: chunk for precision, parent for context, one hop away, never both loaded by default).
Why no vectors. Embeddings mean a model call (cost, latency, an API dependency) or a local model (a real dependency, a real download, real RAM) for every single thing written — directly incompatible with the single-file, stdlib-only, zero-install rule this whole project is built on. BM25 over SQLite FTS5 costs nothing to build, nothing to maintain, and handles exactly what this vault actually needs to search well: command output, notes, transcripts, terminal history — content where the literal words (an error string, a flag name, a path) are the signal. This was an explicit, considered tradeoff, not a missing feature.
Chunking is wired into the system, not bolted onto one feature. record_save, the shell/subagent output spill path (spill_output), and terminal log export all chunk through the identical chunk_lines() call — "RAG" here means a property of the vault, not a special case for logs. It is deliberately not wired into add_record's lower-level, highest-frequency caller — upsert_file_record, which fires on every single file write and edit for audit bookkeeping — because exploding the chunk table on every keystroke-adjacent save would be pure waste for content read_file/grep already serves perfectly well. And deleting a record deletes its chunks and their index rows with it (Store.delete_record); there is no orphan path to worry about.
Sealed search stays ranked without ever touching plaintext to disk. Encrypted-at-rest events can't be FTS-indexed in place — SQLite can only full-text-search what it can read as text. Every boot decrypts events into a :memory:-only FTS5 mirror (open_mem_events), capped at the newest MEM_EVENTS_CAP (10,000) rows and synced incrementally as new ones arrive, so sealed vaults get exactly the same BM25 ranking and [marked] snippets as plaintext ones — and nothing decrypted ever becomes a temp file, a spill table, or a persistent index on disk. Alternatives were researched and rejected on purpose, not just not-gotten-to: page-level encryption (SQLCipher) is a non-standard C extension, incompatible with the stdlib-only rule; blind indexes (deterministic HMACs of tokens, à la CipherStash) would persist in the vault file and leak term frequency and access patterns to anyone who steals it; academic searchable-encryption schemes leak access patterns too and are far heavier than this threat model needs. Decrypting into RAM, once, where the key and plaintext already live during normal operation anyway, is the only option that keeps the at-rest file fully opaque and the search fully ranked.
/search deploy ranked hits across records, events, audit, and chunks
/records deploy search records for "deploy"
/record mem:9f2c41aa77c3e5d1
/graph secret:db show what links to the db secret
/tags all tags by usage count
/tags de tags starting with "de"
Terminal log export
export_terminal_log is only ever offered inside a live tmux session (see the environment-gating section above — this is the tool TOOL_ENV_GATES exists for). When it's available, it runs tmux capture-pane -p -S - — the entire pane scrollback, not just what's currently visible — writes it verbatim to ~/.tai/termlogs/<profile>/<pane>-<timestamp>.log, saves it as a termlog-kind vault record for provenance and full-text paging, and immediately chunks and indexes it exactly like everything else above. From that point on, "what was that error three hundred lines back" is a search call, not a scrollback hunt — the parent record stays for full context, the chunks stay for precision, and the whole thing cost one tmux subprocess call and zero network requests.
Its quieter cousin, get_current_terminal_content, grabs just the visible pane (configurable line count) for a quick glance without writing or indexing anything — the difference between a peek and a capture is deliberate.
Skills
Standard agent skill files (SKILL.md with name plus description frontmatter, optional scripts/, references/, assets/, per the open Agent Skills format) are discovered from four places, in ascending priority so the most tai-specific location always wins a name collision: read-only interop with ~/.claude/skills and ./.agents/skills first (so a skill another tool already dropped next to your project is visible with zero install step), then tai's own ~/.tai/skills, then ./.tai/skills. Descriptions stay in the system prompt catalog at all times; full instructions load only through load_skill, on demand — the same lazy-loading philosophy as tools, applied to skills.
Two independent, complementary ways to get a skill:
create_skillauthors a brand-new one from a brief: a dedicated deep-research subagent runssysinfofirst so the skill matches the actual machine, then researches with web search and shell until the material is independently verified, then writesSKILL.mdplus supporting files — project-scoped by default. Six blueprints ship built in (bot-creator,api-client,web-researcher,pdf-forms,data-wrangler,home-sysadmin);load_skillbuilds a missing one through this exact same path the first time it's actually needed, so a skill finalizes lazily at the moment of use, not at install time.install_skilladopts a skill that already exists somewhere else instead of building one from scratch — a local directory, a bareSKILL.mdfile, a.zip, or a git URL (the only path in this entire feature that touches a non-stdlib binary, and it's checked for and refused cleanly if absent). It copies by default, vendoring a private snapshot — except when the source already sits inside one of the recognized external skill folders (.claude/skills,.agents/skills), in which case it symlinks instead, so you get one canonical copy instead of a silently drifting duplicate. Scope (project vs. home) defaults the same waycreate_skilldoes, and a sidecar.installs/<name>.jsonrecords exactly where a skill came from and how, which/skillsandload_skillboth surface ([installed copy, project scope, from ...],[native], or[external, not managed by tai: ...]) — every skill in the catalog is honest about its own provenance.
Both share one scope-resolution helper and one path-construction helper under the hood, so "where does a project-scoped vs. home-scoped skill actually live" has exactly one definition in the entire codebase.
The sysinfo tool itself runs os, python, venv, root, container, binaries, cpu, and disk checks in parallel, each with its own timing; /sysinfo prints the identical report in the REPL.
Orchestration
/fork <task> spawns a background subagent with its own context while the REPL stays free (the prompt shows a +N counter). /agents lists workers, /agent <id> shows a result, /agent clear purges finished ones. The model orchestrates through the exact same primitives: the fork tool (task, timeout up to one hour, profile) and the poll tool (id, wait up to two minutes). Workers get WORKER_STEPS (12) steps, a cooperative deadline, no session writes, and no interactive approval prompts — nesting caps at FORK_MAX_DEPTH (2) levels. Timeouts and errors always surface as an explicit status, never silently.
The execution policy is one ordered rule, not a vibe:
- Future work is scheduled, never forked and waited on.
- Quick, interactive, or memory-changing work runs on the main agent.
- Independent, long, or context-heavy work goes to a forked subagent.
Forking buys parallelism and context isolation, not security isolation — workers share the same machine, tools, and approval setting. Sandbox mode is the separate axis that decides where untrusted or destructive commands may actually execute. Workers cannot schedule, cannot prompt for approval, and stop nesting after two levels; scheduled prompts fire as subagents that nobody waits on, with outcomes kept in the vault.
Turns run on a progress budget, not a step counter. A step counts as productive the moment it tries an unseen action, returns an unseen result, or answers the user — and any productive step resets the stall counter, so a turn doing genuinely varied work can run as long as it keeps moving forward. Repeating one action with identical results earns a loop warning at LOOP_NUDGE_AT (3) repeats and a hard stop at LOOP_STOP_AT (5); wider stalls — alternating actions that produce no new information — stop after the patience budget runs out (MAX_STEPS 25 for the main agent, WORKER_STEPS 12 for a subagent, CREATE_SKILL_STEPS 40 for the skill-building worker, which legitimately needs the room); a TOTAL_STEP_CAP of 500 backstops absolutely everything regardless of how the patience math plays out. Every user approval or typed guidance resets the stall counter outright, on the theory that a supervised agent is a safe agent — a human steering is never treated as "no progress."
Execution environment / sandbox
/env shows the execution environment, /env sandbox switches shell and file tools into an isolated tai-box container (podman or docker, no mounts, no shared filesystem), /env home switches back. The image builds automatically on first use from an embedded Containerfile; every pip requirement the agent's own features need (faster-whisper, edge-tts) lives inside that image, while tai.py itself stays dependency-free — the sandbox is where "needs a real package" problems go, never the host process. Sandbox commands need no approval because the container is disposable; this is deliberately a different axis from the --yes/--yolo/--auto approval settings, not a replacement for them.
Profiles and bots
A profile is a complete identity: system message plus session history under ~/.tai/profiles (mode 0600), and its own slice of every single vault table — records, secrets, tags, graph edges, audit rows, schedules, episodic events, all keyed by profile. Switching profiles is total amnesia; the only cross-profile knowledge that exists is the identity list itself (/profiles). Read permissions and secret grants reset on every switch, subagents can't be polled across profiles, and tools flatly refuse to fork, schedule, or restore for another identity. Vaults from before this rule existed migrate automatically — unscoped rows join default.
A bot is a lighter-weight thing inside one profile: just a system message plus its own history, with the whole vault shared. Every profile starts with main; create_bot adds named bots from a description, rules, behavior, and optional nicknames (lowercased, with a short prefix auto-registering). /bot switches roles and resumes exactly where that bot left off; @name routes a single turn to another bot and files the question and answer into both histories, so the bot you're actually talking to stays aware a detour even happened.
/profile show current profile
/profile [name] switch profile, creating it when missing
/profiles list all profiles
/bots list bots with nicknames
/bot coder switch to the coder bot
@coder fix this one turn via coder, logged in both
remember merges an instruction into the current profile's system message through the model itself — adding facts, updating behavior, or removing forgotten items while leaving the rest intact — and fires by default on new passwords or behavior changes. recall searches the per-profile episodic log in ~/.tai/memory.db. Credentials never enter memory this way: secrets go through store_secret (or /secret set), and any profile written before that rule existed gets migrated to it automatically on load.
Sealed storage
Storage is sealed by default with a built-in key — this stops casual reads, not a determined attacker, since that key ships in the source. Set TAI_PASSPHRASE for real protection: a home sealed with the default key is re-sealed to your passphrase automatically on first boot, with a notice.
The real key derives from PBKDF2-SHA256 (200,000 rounds) over a random salt in ~/.tai/.seal; values use a per-value nonce with HMAC-SHA256 in encrypt-then-MAC order. SQLite access goes through custom tai_enc/tai_dec functions registered with create_function, so inserts encrypt inline and reads decrypt before matching ever happens — there is no separate encrypt/decrypt pass bolted around the database. Existing plaintext data is sealed automatically on first sealed start. A wrong passphrase refuses to boot at all, exit code 2. Set the variable empty for deliberate plaintext storage.
This construction uses only the standard library and is honest file-theft protection, not audited cryptography — high-value secrets still belong in a dedicated manager.
The vault heals its own schema. Every boot adds any missing table or column automatically and reports exactly what changed (vault schema upgraded: added secrets.meta, ...). The entire schema — tables, columns, primary keys, indexes — is declared once as a plain data structure (SCHEMA / SCHEMA_INDEXES) and reconciled against whatever's actually on disk (ensure_schema()); adding a feature that needs a new column is a one-line diff to that structure, never a hand-written migration script, and SQLite's dynamic typing means historic type drift in a column never blocks a boot either. Vaults from any previous version open with zero manual steps.
Secrets
Passwords, tokens, and secrets live in a sealed secrets table inside the same vault: encrypted at rest, migrated automatically on passphrase rotation, and never revealed by any tool once stored. The model only ever handles names. shell exposes chosen secrets as TAI_SECRET_<NAME> environment variables for exactly one command invocation, web_fetch can send one as an auth header, and every tool result is scrubbed of known secret values (Store.redact) before it reaches context, memory, or display — so a secret can't leak sideways through some unrelated tool echoing it back. Using or deleting a secret asks the user once per session and scope; workers and the Telegram bot are denied outright unless the whole session was started with --yes, --yolo, or --auto, so unattended secret use is always a deliberate, explicit choice, never an accident of automation.
/secret set wifi store a value typed invisibly, never entering context
/secret list show names only
/secret delete wifi remove a value after confirmation
Scheduler
Schedules persist in the same vault with sealed prompts: one-shot appointments (at an ISO datetime, naive means local time) or repeating work (every 60 seconds or more). A background thread ticks every SCHEDULER_INTERVAL (30) seconds in the REPL, the Telegram service, and --scheduler mode alike, and claims due rows atomically so multiple processes running at once (your REPL plus the systemd service, say) can never double-fire the same prompt. Repeats advance past any missed windows instead of backfilling a queue of stale ones; outcomes land right in the row and stay visible through /schedules, /agents, and /agent <id>.
/schedule at 2026-10-08T09:00 water the plants
/schedule every 1h check the inbox
/schedules
/unschedule 2
./tai.py --scheduler
Records and graph
Large or durable text lives in the vault as tagged records, each addressed by a mem:<16 hex> id: research notes, command output, transcripts, terminal captures. record_read pages slices back without ever loading the whole thing, record_search finds by text, kind, or tags (ranked through the same unified FTS5 index everything else uses), and record_delete removes a record — tags, graph edges, and chunks — after confirmation. Records are working memory in cleartext, exactly like session events; actual credentials belong in store_secret, never here.
Audit and time travel
Every file mutation lands in an append-only audit table: tool writes, edits, and deletes with full before/after images, plus pre-execution snapshots of risky shell targets (rm, mv, cp, tee, dd, truncate, shred, and > redirections, globs expanded, best-effort heuristic). Each row carries actor, timestamp, message, true byte sizes, and tags, so history queries time-travel cleanly by path or by tag. Images cap at AUDIT_MAX_CHARS (20,000) with an explicit truncation marker rather than silently growing the vault forever, and restore flatly refuses to write back a truncated image rather than restore partial content and call it done. Every audited path also keeps a file-kind record with its absolute path and latest contents (FILE_RECORD_MAX, 50,000 chars, marked when cut).
/audit /etc/hosts history of one path
/audit latest rows across all paths
/restore 41 restore that row's image after confirmation
Releases and self-backup
The very first thing every boot does — before anything else — is back the running script up to ~/.tai/backups/ as tai-<version>-<utcstamp>-<sha8>.py, skipping the write entirely when the content hash already has a backup on disk, and pruning down to the newest BACKUP_KEEP (10). /release <major|minor|patch> <message> (or the release tool) cuts an actual release: it rewrites the VERSION line atomically (temp file plus rename, never a partial write), snapshots the new script, and logs the message into the audit trail tagged release and v<version>. The rule stays constant regardless of who's cutting the release: patch for fixes with no interface change, minor for backwards-compatible features, major for breaking changes.
Voice
speak synthesizes free neural speech via the Microsoft Edge Read Aloud protocol, implemented with nothing but socket and ssl from the standard library — no key, no package, no third-party TTS SDK. MP3 files land in ~/.tai/audio and play automatically when an OS player exists. listen records and transcribes when a recorder (arecord, sox, ffmpeg) and a transcriber (whisper-cpp, whisper) are both installed, and reports exactly what's missing when they aren't, instead of failing opaquely.
Terminal output
Assistant replies render as formatted markdown on color terminals: aligned tables with left/center/right columns, verbatim fenced code blocks, nested bullet and numbered lists with checkboxes, headings, blockquotes, rules, and inline bold/italic/code/strikethrough/links. Long lines wrap to terminal width without ever breaking a style mid-sequence. Piped output and NO_COLOR fall back to raw markdown, deliberately.
Shell output streams live while a command runs — a rolling 4-line window inside the call box, then an exit line with return code, elapsed time, line count, and byte size. Long lines trim to terminal width without breaking color codes, progress-style output keeps only its latest segment on screen, and the model still receives the full text underneath for real feedback. Workers, pipes, and capture mode stay silent, on purpose.
File changes render as a unified diff right inside the call box: line numbers, green +/red - rows on tinted backgrounds, Python syntax coloring, and ··· separators between hunks, capped at 120 rows with an explicit hidden-line note rather than a wall of scroll. New files show all green, deletes show all red, restores diff against whatever's currently on disk.
Configuration
| Variable | Purpose | Default |
|---|---|---|
TAI_HOME |
State directory | ~/.tai |
TAI_MODEL |
Model id, ignored by primary gateway | openrouter/free |
DEVPLACE_API_KEY |
Fallback backend credential | Empty, fallback disabled |
TAI_VOICE |
Edge voice name | en-US-EmmaMultilingualNeural |
TAI_PASSPHRASE |
Personal seal key | Unset: built-in key; empty: plaintext |
TELEGRAM_BOT_TOKEN |
Bot token, or set during install | Empty |
Telegram
./tai.py --install-telegram
./tai.py --uninstall-telegram
Install asks for the bot token up front, verifies it against getMe, stores it in ~/.tai/telegram.env (mode 0600), builds the sandbox container (used for voice transcription), and registers a tai-telegram.service systemd user unit with linger enabled (failures ignored). The bot long-polls, answers text, transcribes voice notes, and understands /new. It runs without --yes, so destructive shell commands are denied outright. Uninstall stops and removes the service and purges the container and image; the token file and data stay put.
Testing
python3 test_tai.py # full agent/tools/vault/scheduler suite
python3 test_seal.py # seal/encryption regression suite
# a single test, either suite, standard unittest addressing:
python3 -m unittest test_tai.CreateSkillTests
python3 -m unittest test_tai.CreateSkillTests.test_creates_and_refreshes
Layout
tai.py the entire agent — tools, vault, search, skills, orchestration, sandbox, all of it
test_seal.py seal regression tests
test_tai.py agent, tools, records, graph, search/RAG, skills, scheduler tests