Skip to content

tau_agent_core — the runtime

Cognition × Application

The loop that drives a conversation, the door every input arrives through, and the tree a session actually is. Headless: no Textual, no assumptions about stdin or stdout. It runs under the TUI, in a host process, or as an RPC subprocess, and cannot tell which.

Distribution ffwf-tau-agent-core. tau_agent_core never imports tau_coding_agent.

The agent loop

AgentLoop(
    config: AgentLoopConfig, emit=None, tools=None, model=None,
    abort_signal=None, hook_dispatcher=None, steer_queue=None,
)

config, tools, model and abort_signal are constructor arguments, not fields on one shared object. AgentLoopConfig is deliberately small: model, system_prompt, tool_execution_mode, max_retries, max_turns, temperature, api_key, reasoning.

Everything an earlier design put on the config as callbacks — before_tool_call, get_steering_messages, transform_context — is injected instead. Extension hooks go through hook_dispatcher, gated on has_hook_handlers() so a zero-extension session pays nothing for it; mid-turn steering goes through a shared steer_queue.

max_turns defaults to 50. It is a τ-original safeguard — pi has no turn bound at all — but it is a blunt one, and it does not notice that turns 2 through 50 are the same failure repeating.

Two event vocabularies

This is the thing most worth understanding, because it spans four files and two sets of names that describe the same turn.

Where the provider's events become the loop's eventsThree columns. On the left, under the heading tau underscore l l m streaming events, four boxes: TextDeltaEvent, ThinkingDeltaEvent and ToolCallDeltaEvent drawn in a light outline, and DoneEvent dot final drawn in a heavier one. Each of the three light boxes sends a thin arrow right into a tall box in the middle named AgentLoop dot run. DoneEvent dot final instead sends a red connector, squared off in red where it leaves the box, which runs right then up and enters an inner box named get underscore tool underscore calls, glossed as reading off the final message. Below that inner box sits a second one named underscore execute underscore tool underscore calls, glossed sequential or parallel. On the right, under the heading tau underscore agent underscore core AgentEvents, four more boxes receive arrows from the loop: message start with update and end, tool execution start with update and end, turn start and turn end, and agent start and agent end. Beneath the left column the red mark reads what actually runs, glossed that get underscore tool underscore calls reads DoneEvent dot final and never the accumulated deltas.tau_llm streaming eventstau_agent_core AgentEventsTextDeltaEventThinkingDeltaEventToolCallDeltaEventDoneEvent.finalAgentLoop.runone turn, then loopget_tool_calls()off the final message_execute_tool_callssequential or parallelmessage_startupdate, endtool_execution_startupdate, endturn_start / turn_endagent_start / agent_endwhat actually runsget_tool_calls() reads DoneEvent.final,never the accumulated deltas
The loop consumes one vocabulary and emits another, which is why a single turn has two sets of names for the same moment. Red marks the crossing that carries execution rather than display: the three delta events exist so a reader can watch a turn happen, but the tool calls that actually run are pulled off DoneEvent.final, whose arguments the provider has already assembled and parsed. Reconstructing a call from the deltas would work until the day a server fragments differently.

A tool call is transformed four times between HTTP bytes and a rendered widget:

  1. the provider's ToolCall block, off DoneEvent.final;
  2. a message dict block {"type": "toolCall", ...} — model_dump() at the loop boundary;
  3. a backend tool_calls_info dict;
  4. a TUI widget.

When tool calling misbehaves, trace the arguments value through all four hops rather than reading any one of them.

Beyond the base event set, τ stamps provenance on every event — submission_id, source, submitter, correlation — plus blocked and blocked_by on tool_execution_end for extension vetoes, and error on agent_end when the loop raised.

The one door

AgentSession.submit() is the real entry point every input source funnels through: TUI keystrokes, tau -p, the SDK, an extension, an RPC client. prompt() builds a Submission and calls submit() — it is a wrapper, not a second door.

multitask_strategy decides what happens when a submission arrives while a turn is already running. It is a policy rather than an answer improvised per caller:

Strategy Effect
reject Refuse the new submission outright; rejection_reason explains why.
enqueue Queue it; it runs after the current turn ends.
steer Inject it into the running turn without waiting.
rollback Navigate back to the pre-turn leaf and run as if from there. The running turn's output becomes an abandoned sibling branch, not a deletion.
fork Branch the conversation instead of extending the active leaf.

Sessions are a tree

Append-only entries, walked by parent_id. Model input for a turn is an ephemeral frame — system prompt, tool schemas — plus the exact linear path from the root to the active leaf. There is no hidden channel on either side of that.

That is what makes fork, branch and rollback ordinary operations rather than special cases, and it is why a rollback leaves the abandoned branch on disk instead of deleting it.

Piece Role
SessionLog (Protocol) The minimal contract: append_message, append_custom_message, append_custom_entry, append_compaction, append_elide, append_navigate, append_branch_summary, append_at, entries, cursor.
BranchView / open_branch() A lightweight view of one branch that does not disturb the parent log's cursor.
ConversationTree Walks parent_id chains to build model input.
SessionCatalog Lists and resolves sessions. File-backed or JMFTS-backed, chosen by --store.

There is no method named clone, and none named navigate — the latter is append_navigate, which appends a marker entry rather than mutating a cursor in place.

tau_agent_core.testing ships contract suites for both SessionCatalog and SessionLog. A new store costs about twenty lines of knobs to run them.

Release note for SessionLog implementors

The branchOf lane tag is gone. It recorded who wrote an entry, while three of its four consumers wanted does this entry belong to the conversation being looked at — which is ancestry from the cursor. The two agree for a sub-agent and disagree for a fork, so a three-way fork returned three mutually exclusive alternatives as one conversation, and a two-message session counted four messages.

Session.messages and session listing now walk the cursor's ancestry. resolve_cursor is the last-entry rule again; the guarantee dropped on purpose is crash-exact resume under a second concurrent writer, which τ does not buy. subtree_text is bounded by descendants of the node the caller named rather than by write provenance.

In the contract suite, one test is inverted rather than deleted: a store must not reintroduce a cursor filter.

SDK entry point

def create_agent_session(
    model: str | Model = "gpt-4o", provider: str = "openai",
    base_url: str | None = None, api_key: str | None = None,
    tools: list[str] | None = None, session_log: SessionLog | None = None,
    extensions: list[Callable] | None = None, system_prompt: str | None = None,
    no_context_files: bool = False,
    thinking_level: str = "off", cwd: str | None = None,
    tool_execution_mode: Literal["sequential", "parallel"] = "parallel",
    compaction_policy: CompactionPolicy | None = None,
    bus_available: bool = False,
    no_tools: Literal["all", "builtin"] | None = None,
) -> AgentSession

tools takes built-in name strings only — read, write, edit, bash, grep, find, ls. A custom AgentTool instance needs the AgentSession constructor directly.

There is deliberately no settings= parameter. An earlier version had one that was silently ignored, so it was removed rather than kept as a no-op; passing it raises TypeError.

no_tools is the SDK half of the CLI's two flags. "all" offers the model nothing at all; "builtin" drops the built-in set and keeps whatever extensions registered. Passing tools= and no_tools= together raises — they ask for opposite things, and neither outranks the other at a call site. tools=None and tools=[] stay legal alongside "all", which also withholds extension-registered tools.

This factory is not on the live TUI or headless path: tau_coding_agent's backend constructs AgentSession directly. What used to follow from that no longer does — the backend now calls the same prompt builder, so τ's base prompt and its context files reach the model on every path.

Project context files

The system prompt is τ's base prompt, then any discovered context files, then the tool schemas. system_prompt replaces the base text and nothing else; setting it does not switch context files off.

Discovery takes the agent directory's file first, then walks every ancestor of the working directory, root-most first so the nearest file is read last. At most one file per directory, deduplicated by resolved path, first match winning among:

AGENTS.override.md   AGENTS.md   AGENTS.MD   CLAUDE.md   CLAUDE.MD

τ's own .tau/SYSTEM.md is read from the working directory only and appended last. It is deliberately not in that tuple: inside it, it would compete with a sibling AGENTS.md under the one-file-per-directory rule, so a project carrying both — which τ has always read both of — would silently lose one.

A worktree nested inside its own main repository suppresses the main repository's same-named file, so the walk does not load both.

Three departures from pi, each because Fail Early asks for it:

  • A file that is found and cannot be read raises, naming the path and the escape hatch. pi warns to stderr and continues. A prompt silently missing its project instructions looks exactly like a model ignoring them.
  • Decoding is strict. Replacing undecodable bytes turns a mis-encoded instruction file into replacement characters the model still reads as instructions.
  • Every block is wrapped in <project_instructions path="…">, so a prompt cannot carry instructions whose origin it does not state.

The walk reaches /, so a CLAUDE.md in $HOME is read on every run. --no-context-files / -nc turns discovery off, and it is run-level: a mid-session /model switch cannot hand the files back.

Compaction

LLM-backed, with no fabricated-summary fallback. A compaction error raises rather than silently truncating the conversation.

should_compact(context_tokens, context_window, settings) -> bool
prepare_compaction(path_entries, settings) -> CompactionPreparation | None
async def compact(preparation, model, api_key, *,
                  custom_instructions=None, thinking_level=None) -> CompactionResult

Applying one is tree surgery rather than truncation: the first kept entry is re-parented onto the compaction entry, so the summary is on the path and the compacted span is still in the tree.

Built-in tools

read, write, edit, bash, grep, find, ls. --tools allowlists, --exclude-tools denylists, --no-builtin-tools drops the set while keeping extension-registered ones, and --no-tools offers the model nothing at all.

Extensions still load under --no-tools: hooks, commands, injections and subscriptions are untouched, and only callable tools are withheld.

See Extensions for the registration surface and RPC for driving all of this from another process.