Files
tai/README.md
T

481 lines
25 KiB
Markdown
Raw Normal View History

retoor <retoor@molodetz.nl>
# tai
tai is a single-file autonomous AI agent written in Python. The entire
implementation lives in `tai.py` (about 5200 lines) and uses only the Python
standard library: no dependencies, no install step, no build system.
The agent runs as an interactive REPL or as a one-shot command. It reasons
through an OpenAI-compatible backend, acts through thirty-six tools, keeps
per-profile memory, seals its stored state at rest, and can isolate shell and
file operations inside a container sandbox.
## Requirements
- Python 3.10 or newer, no third-party packages.
- Optional: `podman` or `docker` for the sandbox and Telegram voice notes.
- Optional: `tmux` for terminal content capture.
## Quick start
./tai.py
./tai.py --profile work
./tai.py --yes
./tai.py --yolo
./tai.py --auto
./tai.py --version
./tai.py what is 2+3, use the shell
Trailing arguments form a one-shot prompt: the agent answers once and exits
with code 0. Without arguments, tai starts an interactive session.
## REPL commands
| Command | Effect |
|-------------------|---------------------------------------------------|
| `/profile [name]` | Show the current profile or switch to it |
| `/profiles` | List all profiles, current marked with `*` |
| `/bots` | List profile bots, current marked with `*` |
| `/bot [name]` | Switch to another bot, history resumes |
| `/env [target]` | Show or switch execution environment |
| `/skills` | List loaded skill files |
| `/secret` | Manage sealed secrets (set|list|delete) |
| `/sysinfo` | Show host environment checks |
| `/schedule` | Schedule a prompt for later (at|every) |
| `/schedules` | List scheduled prompts and outcomes |
| `/unschedule` | Delete a scheduled prompt by id |
| `/records` | Search saved records, optional query |
| `/record <id>` | Read one record page by mem id |
| `/graph <node>` | Show one vault node neighborhood |
| `/tags [prefix]` | List tags with usage counts |
| `/tools [name]` | Show core and lazy tools |
| `/install ...` | Install status, install, upgrade, reinstall |
| `/search <query>` | Ranked search over records, events, audit |
| `/audit [path]` | Show file audit trail |
| `/restore <id>` | Restore a file from an audit row |
| `/release` | Bump version, back up, log message |
| `/fork <task>` | Spawn a background subagent, REPL stays free |
| `/agents` | List background subagents |
| `/agent <id>` | Show one subagent result |
| `/agent clear` | Purge finished subagents |
| `/compact` | Compress history into a summary |
| `/clear` | Drop history, keep the system message |
| `/help` | Show the command overview |
| `/quit` | Exit |
Any other input is sent to the agent.
## Installation
./tai.py --install
/install status
/install upgrade scheduler-service
Six install targets exist side by side: `binary` (`~/.local/bin/tai.py`),
`bash-hook` (a guarded `command_not_found_handle` block in `~/.bashrc`,
backed up once to `~/.bashrc.bak-tai`, so unknown shell commands are
answered by the agent), `venv` (`~/.tai/venv`, created whenever the
venv module exists), `scheduler-service` and `telegram-service`
(systemd user units that prefer the venv python), and `container`
(the sandbox, whose image carries its own `/box/venv`). `/install
status` (or the `install` tool) reports exactly which of these exist,
with versions and service states. `install` adds missing pieces,
`upgrade` refreshes in place, `reinstall` rebuilds artifacts from
scratch, `uninstall` removes them; service changes restart or start
units immediately. Data handling is explicit: no action ever touches
the vault (`memory.db`, secrets, schedules, records, backups).
## Backends
The primary backend is `model.cloud.pravda.education`, an OpenAI-compatible
gateway that needs no API key and selects a free model per request. If a
request fails, tai retries it on `devplace.net/openai/v1`, which requires
`DEVPLACE_API_KEY`. Both endpoints speak `/chat/completions`, including
native tool calls and streaming.
## Orchestration
`/fork <task>` spawns a background subagent with its own context while the
REPL stays free (the prompt shows a `+N` counter). `/agents` lists workers,
`/agent <id>` shows a result, `/agent clear` purges finished ones.
The model itself orchestrates through the `fork` tool (task, timeout up to
one hour, profile) and the `poll` tool (id, wait up to two minutes). Workers
get 12 steps, a cooperative deadline, no session writes, and no interactive
approval prompts. Nesting is capped at two levels. Timeouts and errors
surface as statuses, never silently.
## Execution policy
The agent decides who runs each unit of work by one ordered rule:
1. Future work is scheduled, never forked and waited on.
2. Quick, interactive, or memory-changing work runs on the main agent.
3. Independent, long, or context-heavy work goes to a forked subagent.
Forking buys parallelism and context isolation, not security isolation:
workers share the same machine, tools, and approval setting. Sandbox
mode is the separate axis that decides where untrusted or destructive
commands may run. Workers cannot schedule, cannot prompt for approval,
and stop nesting after two levels; scheduled prompts fire as subagents
that nobody waits on, with outcomes kept in the vault.
## Tools
| Tool | Purpose |
|------------------------------|------------------------------------------------------|
| `shell` | Run a shell command, big output spills to a record |
| `read_file` | Read a text file, large files truncated |
| `write_file` | Write content to a file, creating parent directories |
| `edit_file` | Replace one unique exact text match in a file |
| `web_search` | Search the web, optionally images or page content |
| `web_fetch` | Fetch a URL and return its text content |
| `speak` | Synthesize speech, save MP3, play when possible |
| `listen` | Record from the microphone and transcribe it |
| `remember` | Merge knowledge into the profile system message |
| `recall` | Search past session memory by keyword |
| `load_skill` | Load a skill file by name |
| `get_current_terminal_content` | Capture the current tmux pane with scrollback |
| `fork` | Spawn a background subagent |
| `poll` | Collect a subagent result, big ones spill |
| `sysinfo` | Inspect the host, parallel checks with timing |
| `create_skill` | Deep-research and write a new skill file |
| `store_secret` | Store a password or token in the sealed vault |
| `list_secrets` | List vault secret names, values never shown |
| `delete_secret` | Delete a vault secret by name |
| `schedule` | Run a prompt later or on an interval |
| `unschedule` | Delete a scheduled prompt by id |
| `schedules` | List scheduled prompts and outcomes |
| `record_save` | Save text as a tagged record, get a mem id |
| `record_read` | Read one page of a record by mem id |
| `record_search` | Search records by text, kind, and tags |
| `record_delete` | Delete a record by mem id, asks first |
| `graph_link` | Link two vault nodes with a relation |
| `graph_query` | Show one vault node neighborhood |
| `delete_file` | Delete a file, pre-image stays audited |
| `audit` | Show the file audit trail for time travel |
| `restore` | Restore a file to an audit row image |
| `release` | Bump version, back up, log the message |
| `tags` | List vault tags with usage counts |
| `install` | Manage installs, status to uninstall |
| `create_bot` | Create a bot: system plus history, shared vault |
| `search` | Ranked full-text search over everything stored |
Web search runs on `rsearch.app.molodetz.nl`. Destructive shell commands ask
for confirmation unless `--yes` is given; read-only commands run directly.
Every prompt offers `[y]once [Y]always [n]o`: `Y` enables yolo mode for
the session, `--yolo` starts there, and `--auto` adds autonomous research
instead of ever asking. Answering no asks what to do instead: typed
guidance continues the turn while empty input aborts it. Overwriting an
existing file requires reading it first in the same session; new files
are always writable. Every write, edit, delete, and risky shell target
is audited with before/after images for time travel (see below).
## Lazy tools
With 30-plus tools, sending every schema on every turn would burn
context and blur tool selection, so tai loads lazily. Eight everyday
tools (`shell`, `read_file`, `write_file`, `remember`, `recall`,
`load_skill`, `fork`, `poll`) are always present; everything else
enters the payload only when the recent conversation names it. Each
tool carries trigger tags with synonyms (`cron` loads `schedule`,
`undo` loads `restore`, `password` loads `store_secret`), so ordinary
wording just works, and tool results feed selection too: a spill
pointer naming `record_read` loads it for the next step. A compact
name-plus-summary catalog stays in the system prompt so no capability
is ever hidden, only its verbose schema. `/tools` shows the split,
`/tools <name>` shows one tool with its tags.
## Skills
Standard agent skill files (`SKILL.md` with `name` plus `description`
frontmatter, optional `scripts/`, `references/`, `assets/`, per the Agent
Skills open format) are discovered in `~/.tai/skills/*/` and
`./.tai/skills/*/` (project wins on name collisions). Descriptions stay in
context; the agent loads full instructions through `load_skill` only when
needed. `/skills` lists what is available.
The `create_skill` tool authors new skills on demand: it runs a dedicated
deep-research worker on the prompt (sysinfo first, then web and shell
research until the material is verified against independent sources) and
writes the skill directory, project scope by default. Six blueprints ship
with the agent (`bot-creator`, `api-client`, `web-researcher`, `pdf-forms`,
`data-wrangler`, `home-sysadmin`): `load_skill` builds a missing one on
first use through the same deep-research path, so skills finalize lazily
the moment they are needed. The `sysinfo` tool
reports os, python, venv, root, container, binaries, cpu, and disk in
parallel, each check with its own timing; `/sysinfo` prints the same
report in the REPL.
## Sandbox
`/env` shows the execution environment, `/env sandbox` switches shell and
file tools into an isolated `tai-box` container (podman or docker, no mounts,
no shared filesystem), `/env home` switches back. The image is built
automatically on first use from an embedded Containerfile; every pip
requirement (`faster-whisper`, `edge-tts`) lives inside the image while
`tai.py` itself stays dependency-free. Sandbox commands need no approval
because the container is disposable.
## Telegram
./tai.py --install-telegram
./tai.py --uninstall-telegram
Install asks for the bot token up front, verifies it against `getMe`, stores
it in `~/.tai/telegram.env` (0600), builds the sandbox container (used for
voice transcription), and registers a `tai-telegram.service` systemd user
unit with linger enabled (failures ignored). The bot long-polls, answers
text, transcribes voice notes, and understands `/new`. It runs without
`--yes`, so destructive shell commands are denied. Uninstall stops and
removes the service and purges the container and image; the token file and
data stay.
## Profiles and memory
Each profile is a complete identity: system message plus session history
under `~/.tai/profiles` (mode 0600), and its own slice of every vault
table. Records, secrets, tags, graph edges, audit rows, schedules, and
episodic events are all keyed by profile; switching profiles is total
amnesia. The only cross-profile knowledge is the identity list itself
(`/profiles`). Read permissions and secret grants reset on every switch,
subagents cannot be polled across profiles, and tools refuse to fork,
schedule, or restore for another identity. Databases from before this
rule migrate automatically: unscoped rows join `default`.
/profile show current profile
/profile [name] switch profile, creating it when missing
/profiles list all profiles
The `remember` tool merges an instruction into the current profile system
message through the model itself: it adds facts, updates behavior, or removes
forgotten items while preserving the rest. It fires by default on new
passwords and behavior changes. `recall` searches the per-profile episodic
log in `~/.tai/memory.db` (SQLite). Context is budgeted at roughly 32k
tokens with automatic compaction at 80 percent. Credentials never enter
memory: secrets go to the vault through `store_secret` (or `/secret
set`), and profiles written before this rule are migrated to it
automatically on load.
## Bots
Where a profile is an identity, a bot is a lightweight role inside it:
only a system message plus its own session history, with the whole
vault shared. Each profile starts with `main`; `create_bot` adds named
bots from a description, rules, behavior, and optional nicknames (all
lowercased, the short prefix auto-registers). `/bot` switches roles
and resumes exactly where that bot left off; `@name` routes a single
turn to another bot and files the question and answer in both
histories, so the current bot stays aware of the detour.
/bots list bots with nicknames
/bot coder switch to the coder bot
@coder fix this one turn via coder, logged in both
## Sealed storage
Storage is sealed by default with a built-in key, which stops casual reads
but not a determined attacker, since the key ships in the source. Set
`TAI_PASSPHRASE` for real protection: a home sealed with the default key is
re-sealed to your passphrase automatically on first boot, with a notice.
The key comes from PBKDF2-SHA256 (200k rounds) over a random salt in
`~/.tai/.seal`; values use a per-value nonce with HMAC-SHA256 encrypt-then-
MAC. SQLite access goes through custom `tai_enc`/`tai_dec` functions
registered with `create_function`, so inserts encrypt inline and recall
decrypts before matching. Existing plaintext data is sealed automatically on
first sealed start. A wrong passphrase refuses to start with exit code 2.
Set the variable empty for plaintext storage.
This construction uses only the standard library and is honest file-theft
protection, not audited cryptography; high-value secrets still belong in a
dedicated manager.
## Sealed search
The seal guards against file theft: an attacker who copies `~/.tai`
learns nothing without the passphrase. Full-text search over sealed
events works without weakening that promise. Each boot decrypts
events into a memory-only FTS5 index (a `:memory:` database holding
the newest 10000 rows, synced incrementally as new rows arrive),
so sealed stores get BM25 ranking and marked snippets exactly like
plaintext ones. Nothing decrypted ever touches disk: no temp files,
no spill tables, no persistent helper index. The key and plaintext
already live in process RAM during operation, so a RAM-only index
adds no new exposure against the file-theft threat model, and
passphrase rotation needs no rebuild since plaintext is unchanged.
Alternatives were researched and rejected deliberately. Page-level
encryption (SQLCipher) keeps FTS5 working transparently but is a
non-standard C extension, incompatible with the single-file stdlib
rule. Blind indexes (deterministic HMACs of tokens stored next to
the ciphertext, as in CipherStash or IronCore cloaked search) would
persist in the vault file and leak term frequency plus search and
access patterns to anyone stealing it, while losing stemming and
BM25 ranking. Academic searchable-encryption schemes leak access
patterns too and are far heavier than this threat model needs.
Decrypting into RAM is the only option that keeps the at-rest file
fully opaque and the search fully ranked.
## Secrets
Passwords, tokens, and secrets live in a sealed `secrets` table inside
the same vault: encrypted at rest, migrated on passphrase rotation, and
never revealed by any tool. The model only handles names. `shell`
exposes chosen secrets as `TAI_SECRET_<NAME>` variables for one
command, `web_fetch` sends one as an authentication header, and every
result is scrubbed of known values before it reaches context, memory,
or display. Using or deleting secrets asks the user once per session
and scope; workers and the Telegram bot are denied unless started with
`--yes`, `--yolo`, or `--auto`, so unattended secret use is always an
explicit choice.
/secret set wifi store a value typed invisibly, never entering context
/secret list show names only
/secret delete wifi remove a value after confirmation
## Scheduler
Schedules persist in the same vault with sealed prompts: one-shot
appointments (`at` an ISO datetime, naive means local time) or
repeating work (`every` 60 seconds or more). A background thread ticks
every 30 seconds in the REPL, the Telegram service, and `--scheduler`
mode, claims due rows atomically so parallel processes never double
fire, and runs each prompt as a subagent nobody waits on. Repeats
advance past missed windows instead of backfilling; outcomes land in
the row and stay visible through `/schedules`, `/agents`, and
`/agent <id>`.
/schedule at 2026-10-08T09:00 water the plants
/schedule every 1h check the inbox
/schedules
/unschedule 2
./tai.py --scheduler
## Records and graph
Large or durable text lives in the same vault as tagged `records`,
each addressed by a `mem:<16 hex>` id: research notes, command
output, transcripts. Anything over 6000 chars returned by a shell
command or subagent is stored automatically and replaced by a pointer
with its size; `record_read` pages slices back without loading the
whole, `record_search` finds by text, kind, or tags, and
`record_delete` removes after confirmation. Secrets, schedules, and
records all carry normalized tags and join one graph: `graph_link`
connects nodes such as `mem:<hash>`, `secret:<name>`, and
`sched:<id>` with a relation, and `graph_query` walks the
neighborhood breadth-first, capped at depth 4 so context never
explodes. Records are working memory in cleartext, like session
events; true credentials belong in `store_secret`.
/records deploy search records for "deploy"
/record mem:9f2c41aa77c3e5d1
/graph secret:db show what links to the db secret
## Tags
Tag rules are strict so the vocabulary stays small: lowercase singular,
shortest common word for the subject, one canonical tag per subject, at
most 5 per item, and always reuse from the `tags` tool (which lists
every tag with usage counts) instead of inventing synonyms. Writes
merge plurals into a known singular automatically, while irregular
words (`news`, `glass`, `status`, `physics`) are never rewritten;
searches expand singular/plural variants so old splits still match.
Any content word that already exists as a tag attaches itself to a new
record (up to 5, plurals included), and each new record links itself to
the 3 most recent records sharing a tag, so the knowledge graph stays
connected without any manual work.
/tags all tags by usage count
/tags de tags starting with "de"
## Search
One FTS5 index covers records, episodic events, and the audit trail,
kept in sync by triggers and backfilled once for older rows. Queries
are tokenized safely (no raw MATCH syntax reaches SQLite), stemmed by
the porter tokenizer when available, ranked by BM25, and returned
with `[marked]` snippets. `search` queries all three stores at once
with kind filters and optional graph expansion that appends linked
neighbors to each record hit; `recall`, `record_search`, and `audit`
all rank through the same index, falling back to LIKE matching when
a query has no full-text match. Sealed events are no exception:
each boot decrypts them into a memory-only FTS5 index (newest 10000,
synced incrementally as rows arrive, never written to disk), so
sealed stores rank event hits with BM25 and snippets exactly like
plaintext ones while the at-rest seal stays untouched. The recipe
is deliberate: one ranked query plus graph hops keeps context
small and accurate instead of paging through stores.
/search deploy ranked hits across records, events, audit
## Audit and time travel
Every file mutation lands in an append-only `audit` table in the vault:
tool writes, edits, and deletes with before/after images, plus
pre-execution snapshots of shell targets (`rm`, `mv`, `cp`, `tee`,
`dd`, `truncate`, `shred`, and `>` redirections, globs expanded,
best-effort heuristic). Each row carries actor, timestamp, message,
true byte sizes, and tags, so history queries time-travel by path or
tag. Images cap at 20000 chars with an explicit truncation marker, and
`restore` refuses truncated images rather than writing partial
content. Every audited path also keeps a `file` record with its
absolute path and latest contents (50000 chars, marked when cut).
/audit /etc/hosts history of one path
/audit latest rows across all paths
/restore 41 restore that row's image after confirmation
## Releases and self-backup
The first thing every boot does is back the running script up to
`~/.tai/backups/` as `tai-<version>-<utcstamp>-<sha8>.py`, skipping
when the content hash already has a backup and pruning to the newest
ten. `/release <major|minor|patch> <message>` (or the `release` tool)
cuts a release: it rewrites the `VERSION` line atomically
(temp-plus-rename), snapshots the new script, and logs the message in
the audit trail tagged `release` and `v<version>`. The rule stays
constant: `patch` for fixes with no interface change, `minor` for
backwards-compatible features, `major` for breaking changes.
## Voice
`speak` synthesizes free neural speech via the Microsoft Edge Read Aloud
protocol, implemented with `socket` and `ssl` from the standard library. No
key, no package. MP3 files land in `~/.tai/audio` and play when an OS player
exists. `listen` records and transcribes when a recorder (`arecord`, `sox`,
`ffmpeg`) and a transcriber (`whisper-cpp`, `whisper`) are installed, and
reports exactly what is missing otherwise.
## Terminal output
Assistant replies render as formatted markdown on color terminals: aligned
tables with left, center, and right columns, verbatim fenced code blocks,
nested bullet and numbered lists with checkboxes, headings, blockquotes,
rules, and inline bold, italic, code, strikethrough, and links. Long lines
wrap to the terminal width without breaking styles. Piped output and
`NO_COLOR` stay raw markdown.
## Configuration
| Variable | Purpose | Default |
|---------------------|--------------------------------------|--------------------------------------------|
| `TAI_HOME` | State directory | `~/.tai` |
| `TAI_MODEL` | Model id, ignored by primary gateway | `openrouter/free` |
| `DEVPLACE_API_KEY` | Fallback backend credential | Empty, fallback disabled |
| `TAI_VOICE` | Edge voice name | `en-US-EmmaMultilingualNeural` |
| `TAI_PASSPHRASE` | Personal seal key | Unset: built-in key; empty: plaintext |
| `TELEGRAM_BOT_TOKEN`| Bot token, or set during install | Empty |
## Testing
python3 test_seal.py
python3 test_tai.py
## Layout
tai.py the entire agent
test_seal.py seal regression tests
test_tai.py agent, tools, records, graph, scheduler tests