Tau — DevOps Manual¶
Cognition × Acquisition
How Tau is deployed and operated: where sessions live, how model credentials are supplied, and where the boundary between harness and agent actually falls.
Draft
Written from the source repo's own README, docs/PI-RPC-REPLACEMENT.md,
docs/NATS-BUS-EXTENSION.md, and docs/REMOTE-CONTROL.md. Not yet
checked against a deployment outside the maintainer's own.
Install¶
pip install 'ffwf-tau-coding-agent[tui]' # interactive
pip install ffwf-tau-coding-agent # headless only
The base distribution pulls ffwf-tau-agent-core and ffwf-tau-llm behind it.
[tui] adds Textual and Rich, and that extra is the only thing an interactive
tau needs that a headless one does not — without it, tau -p and
tau --mode rpc still run a full turn, extensions and all. Measured: 15
packages and 13 MB against 27 packages and 31 MB. That gap is the argument for
the split, because a container that only ever runs headless turns has no reason
to ship a terminal interface.
The other capabilities are extras on the same principle — [jmfts] for
--store jmfts, ffwf-tau-agent-core[bus] for the nats_bus extension. Each
reports its own absence with the install command rather than a traceback. The
Reference page carries the full table.
The ffwf- prefix is load-bearing rather than branding: tau-ai and tau-llm
on PyPI are unrelated third-party projects, so an install command missing the
prefix installs someone else's code.
Which command name to write¶
The install puts both tau and ffwf-tau on PATH — one entry point, two
wrappers pip owns and removes on uninstall. Type tau at a terminal; write
ffwf-tau in a systemd unit, a Dockerfile, or a cron line. PyPI reserves
distribution names but not command names, and an unrelated project ships its
own tau: in an environment holding both, whichever installed last owns the
name, and nothing tells you which one that was. ffwf-tau cannot be taken.
Five ways to run τ¶
| Mode | Entry point | Shape |
|---|---|---|
| Interactive | tau |
Textual TUI in a terminal. |
| Headless print | tau -p "..." |
One turn, prints a transcript, exits. |
| Headless JSON | tau -p --mode json "..." |
One turn, JSONL lifecycle events instead of text — the machine-readable equivalent of the TUI's stream. |
| RPC subprocess | tau --mode rpc |
A persistent JSON-RPC 2.0 server over stdio — τ as a process a host drives, not a library it imports. See the Reference page for the verb table. |
| Embedded (SDK) | create_agent_session(...) in Python |
In-process, no subprocess boundary. tau_agent_core never imports tau_coding_agent, so this path has no Textual dependency at all. |
tau -p and tau --mode rpc both write and resume real sessions — a
headless run shows up in the TUI's sidebar and can be picked up
interactively later. There is no separate "headless-only" session format.
Where sessions live¶
Default: ~/.tau/sessions for the TUI and tau -p; a private
<tmp>/.tau-<uid>/sessions for --mode rpc (a subprocess a host spawns and
tears down shouldn't litter the shared directory by default).
--session-dir DIR overrides this for the file store. --store {file,jmfts}
picks the backend for the run — ffwf-tau-jmfts is an optional package
(pip install 'ffwf-tau-coding-agent[jmfts]'), loaded lazily only when
selected, never a hard dependency of the TUI or CLI.
--no-session skips persistence entirely (ephemeral).
Sessions are append-only JSONL, walked by parent_id to build model input —
not a flat chat log. That is what makes fork, branch, and rollback safe
operations rather than special cases: AgentSession.submit()'s
multitask_strategy="rollback" navigates back to the pre-turn leaf without
deleting anything (the abandoned turn becomes a sibling branch), and
multitask_strategy="fork" branches instead of extending the active leaf.
See the Reference page
for the full submission-strategy table.
Model credentials¶
~/.tau/config.json selects the default model (out of the box, a
local-llm entry pointing at a local OpenAI-compatible server — vLLM,
Ollama, or llama.cpp's server). A missing API key raises (No API key
for provider: …) rather than running with a fabricated placeholder key —
there is no silent fallback to try to guess a credential.
Per-extension credentials follow the same shape: extensions.<name> in
config.json, overridable per-run with --ext-config NAME.KEY=VALUE
(CLI wins over config.json). An extension that needs a bearer token for a
downstream service (JMFTS, a bus) reads it from api.config, not from an
environment variable the model's own shell could printenv.
Running as a subprocess (RPC)¶
--mode rpc gives a host process the properties a plain Python import
cannot: a real hard kill. terminate()/kill() against a τ child works
the same way it works against any subprocess — a runaway tool-call loop is
stoppable from outside, not just cooperatively. The reader loop is strictly
serial, so abort stays answerable while a turn is in flight (measured:
get_state answered at +0.44s and abort at +0.46s against a 20-second
provider call in flight). Embedding τ in-process trades this away — abort
becomes cooperative only (agent_loop.py polls the abort signal per SSE
line), and a wedged tool or a CPU-bound stretch shares the host's own event
loop. Choose the RPC path over embedding whenever a runaway agent must not
be able to degrade its host.
A host driving tau --mode rpc should call get_capabilities first (it
publishes limits.max_request_line_bytes and the verb list) and give its
own subprocess reader an 8 MiB+ line limit before it does — get_capabilities's
own response is tens of kilobytes, well over the stdlib StreamReader
default of 64 KiB. See docs/PI-RPC-REPLACEMENT.md in the source repo (a
real integration writeup, not a design doc) for the full porting notes.
Extensions in a deployment¶
Extensions that touch a message bus declare TOUCHES_BUS = True and are
refused at load time unless the run explicitly opts in with --bus (CLI) or
bus_available=True (SDK) — a capability grant, not a default. The
declaration has a second half: such an extension must also declare a non-empty
SUBJECTS, naming the subjects it touches. "Leave it unset" is refused even
with --bus given, because the grant is per-subject rather than blanket. The shipped
example is nats_bus.py (tau-006/tau-007), which needs
pip install 'ffwf-tau-agent-core[bus]' for its NATS client: τ speaking NATS directly as a
bus-native agent node, bridging to Tectum's effector nodes and to a
simulation engine's world verbs. See docs/NATS-BUS-EXTENSION.md in the
source repo for the full config surface and verb table.
Discovery is ~/.tau/extensions/ plus any explicit -e PATH. There is no
project-local <cwd>/.tau/extensions/ discovery in a deployment today —
deliberately deferred pending a trust gate, not an oversight.
Testing without a live backend¶
τ's own test suite sandboxes config with monkeypatch.setattr against
tau_coding_agent.config.CONFIG_PATH/TAU_DIR rather than relying on
environment variables alone — a test that reads the real
~/.tau/config.json will happily talk to whatever real backend that config
points at (a real JMFTS server, a real model endpoint), which is a slow and
non-hermetic default worth guarding against explicitly in any Tau-embedding
project's own test setup, not just Tau's.
Dev loop¶
From a checkout of the source repo, rather than from PyPI:
python -m venv venv && source venv/bin/activate
pip install -e ./tau-llm -e './tau-agent-core[dev]' -e './tau-coding-agent[dev]' -e ./tau-jmfts
pytest # whole suite (config lives in the repo root pyproject.toml)
mypy tau-llm/src tau-agent-core/src tau-coding-agent/src tau-jmfts/src
The [dev] extras pull the TUI, JMFTS and bus extras transitively, so a plain
editable install is not enough to run the suite. mypy takes all four source
trees in one call — running it against a single package in isolation reports
errors that are artefacts of the missing siblings.
A pre-commit hook (ruff check, ruff format --check, mypy) hard-gates
commits on the source repo — git config core.hooksPath .githooks to
enable it locally.
Known gaps¶
create_agent_session(the documented SDK entry point) is not on the live TUI/headless path today —tau_coding_agent's backend constructsAgentSessiondirectly and never calls it. Real code, orphaned from the path that actually runs; worth knowing before assuming the SDK's own system-prompt-building helper is exercised in production.- The per-dispatch "thinking level" toggle some deployments want (switching
a single running model between fast/slow per request) has no first-class
field on
AgentLoopConfigyet; the workaround is two model config entries plusset_modelover RPC, not a true per-request toggle.
See the Reference page for exact signatures and the Cookbook for worked examples built from real code in the source tree.