Files

194 lines
7.3 KiB
Markdown

# tmux-shot
Drive a TUI application inside tmux from Python, and capture pixel-accurate
PNG screenshots of it — no X server, no headless browser, no external
binaries. Pure `tmux` + `pyte` (terminal emulator) + `Pillow`.
Built for AI-driven workflows: an agent can create a session, type into it,
capture a PNG, look at it, decide what to do next, and repeat.
## How it works
1. `tmux_shot.TmuxApp` creates/attaches a tmux session running your TUI and
sends keystrokes to it (`libtmux` under the hood).
2. `tmux capture-pane -e -p` dumps the pane's exact character grid,
including ANSI color/style escape codes.
3. `pyte.Screen` replays that byte stream into an in-memory terminal buffer
(one cell per character, with fg/bg/bold/etc.).
4. `render.py` draws that buffer cell-by-cell with Pillow into a PNG using a
monospace font.
If the `freeze` CLI (https://github.com/charmbracelet/freeze) is installed,
you can use `render_with_freeze()` instead for nicer-looking output (window
chrome, better font shaping) — it's optional, not a hard dependency.
## Install
```bash
uv sync # or: pip install -e .
```
## Quick start
```python
from tmux_shot import TmuxApp
app = TmuxApp("demo", command="python3 examples/demo_tui.py", width=100, height=30)
app.screenshot("out/1.png")
app.send_keys("j") # drive the app
app.screenshot("out/2.png")
app.send_keys("q")
app.kill()
```
Or from the CLI:
```bash
tmux-shot new demo --command "top" --width 120 --height 40
tmux-shot shot demo out.png
tmux-shot send demo "q"
tmux-shot kill demo
```
## Demo TUI
`examples/demo_tui.py` is a self-contained curses app (arrow keys to move a
cursor, `q` to quit) used to exercise the pipeline without depending on any
other installed program.
```bash
uv run python examples/run_demo.py
```
produces `out/demo_*.png` frames showing the cursor moving and colors
rendering correctly.
## AI-driven "photo session" mode
`tmux-agent` is its own binary (also reachable as `tmux-shot agent`): give it
a plan in plain English and it drives the app itself (arrow keys, waits,
screenshots) via OpenAI-style function calling, looking at each screenshot as
it goes.
It talks to [devplace.net](https://devplace.net)'s OpenAI-compatible gateway
(`https://devplace.net/openai/v1`, model `molodetz`) by default. Set your key
first:
```bash
export DEVPLACE_API_KEY=<your key from your devplace.net profile page>
```
The prompt is the only required argument — everything else has a sensible
default:
```bash
uv run tmux-agent "open the app, find the settings panel, screenshot it"
```
Omit the prompt and it asks interactively instead. Every knob from
`tmux-shot`'s other commands is still available when you need it:
```bash
uv run tmux-agent \
--session myapp --command "myapp" --width 120 --height 40 \
--delay 1.5 --max-steps 40 --out-dir out/session1 --gif out/session1.gif \
"Open the settings menu, navigate to Network, screenshot it as \
'network-panel', then go back to the main menu and screenshot that \
as 'home', then finish."
```
Tools exposed to the model: `send_keys` (tmux key names like `Up`/`Enter`/
`C-c`, or literal text), `wait`, `read_text` (cheap text-only state check),
`screenshot` (saves + shows the model the PNG), and `finish`. `--delay`
controls the automatic pause after every `send_keys` so the TUI has time to
redraw before the next look; the model can also request extra waits itself.
Swap `--model`/`--base-url`/`--app-reference` to point at a different
OpenAI-compatible backend.
## "N minutes of X" — timed sessions with a guaranteed GIF/video
Asking for "a gif of 3 minutes of you doing X" doesn't map onto a normal
step-count loop, and a model can misread "video" as "produce a video file"
and just refuse. Two things fix this:
1. The system prompt is explicit that "video"/"gif"/"movie"/"N minutes of X"
requests are fully in scope — the model's job is to keep taking real
actions (`send_keys`) and periodic `screenshot`s, not to produce a file
itself.
2. `--duration SECONDS` makes the guarantee mechanical instead of relying on
the model cooperating: a background timer (stdlib `threading`, no extra
dependency) takes a screenshot every `--interval` seconds for the entire
requested duration, regardless of what the model does. If the model
stalls or refuses partway through, the run keeps waiting out the clock —
letting the ticker keep capturing — instead of ending early, so a 3-minute
request always produces roughly 3 minutes of real footage.
```bash
uv run tmux-agent \
--duration 180 --interval 5 --max-steps 80 \
--command "grok --cwd /tmp/scratch --always-approve" \
--gif out/vibecode.gif \
"actively vibecode something small and fun with grok the whole session"
```
`--always-approve` (a `grok` flag, not ours) is worth knowing about here: if
the target app itself asks for permission before editing files or running
commands, an unattended agent has no one to click "yes" for it and will
stall. Only use it against a disposable scratch directory (`--cwd
/tmp/scratch` above) — it removes the safety net that normally stops a
misbehaving app from doing something destructive.
## Testing
```bash
uv run pytest
```
- `tests/test_render.py`, `tests/test_tmux_control.py`, `tests/test_cli.py` — core
pipeline (colors, reverse video, `send_keys` correctness, screenshots, CLI
round-trips). No network, no API key needed.
- `tests/test_agent.py` — exercises the full LLM tool-calling loop
(`send_keys`/`wait`/`read_text`/`screenshot`/`finish`, image follow-up
messages, the max-steps safety cap) against a scripted fake OpenAI client —
validates the agent harness itself without spending real API calls.
- `tests/test_grok_target.py` — a real-world capability probe against the
`grok` CLI (a third-party TUI), auto-skipped if `grok` isn't installed.
Read-only by design: only ever sends arrow keys, never Enter, so no prompt
is ever submitted to grok's own model. Produces an actual animated GIF of
its welcome screen as a working example of the "movie" pipeline below.
`TmuxApp.wait_for(text, timeout=10)` polls the pane until a marker string
appears — use it instead of a fixed `sleep()` before screenshotting a TUI
whose startup/redraw time varies (this was discovered as real flakiness
while probing `grok`: its splash screen can take anywhere from ~2s to ~7s
depending on an update check).
## Making a "movie" (animated GIF) from a sequence of screenshots
```python
from tmux_shot import frames_to_gif
frames_to_gif(["01_a.png", "02_b.png", "03_c.png"], "session.gif", duration_ms=500)
```
or from the CLI:
```bash
tmux-shot gif out/01_a.png out/02_b.png out/03_c.png out/session.gif --duration-ms 500
```
`tmux-agent --gif out/session.gif "..."` stitches every screenshot the agent
took during that run into one GIF automatically.
## Notes for automating this with an AI agent
- `TmuxApp.screenshot()` returns the path it wrote — feed that straight to a
vision-capable model.
- `TmuxApp.capture_text()` gives you the plain-text grid (no image) when the
agent just needs to read state cheaply instead of "looking" at a picture.
- Sessions are named, so multiple agents/tasks can run against independent
tmux sessions concurrently without colliding.
- Everything is idempotent: `TmuxApp(name, command=...)` attaches to an
existing session with that name instead of erroring if it's already
running.