Providers and the client¶
AbortSignal¶
tau_llm.abort.AbortSignal
Signal that can be checked during async operations.
When abort() is called, all subsequent is_aborted() checks return True. Async operations should check this periodically (e.g., every 100ms) and raise asyncio.CancelledError if aborted.
Reference: SUBPHASE-0.0.md, "3. AbortSignal" section.
abort¶
abort() -> None
tau_llm.abort.AbortSignal.abort
Signal that the current operation should be aborted.
Idempotent: calling multiple times has no additional effect. Thread-safe: uses threading.Lock internally.
is_aborted¶
is_aborted() -> bool
tau_llm.abort.AbortSignal.is_aborted
Check if the signal has been aborted.
Thread-safe: uses threading.Lock internally. Idempotent: multiple calls return the same value.
Returns
True if abort() has been called, False otherwise.
ApiFactory¶
tau_llm.providers.base.ApiFactory
Builds a :class:Provider for one wire protocol, bound to one endpoint.
Keyword-only, so a factory can ignore what it does not need and so adding a field later does not silently shift a positional argument.
call¶
__call__(*, provider_id: str, name: str, base_url: str, api_key: str | None) -> Provider
tau_llm.providers.base.ApiFactory.__call__
Build the provider.
Parameters
provider_id: str— Vendor id to stamp ontoProvider.id.name: str— Vendor display name to stamp ontoProvider.name.base_url: str— Fully resolved endpoint — the caller has already applied the model's ownbase_urland any vendor default, so a factory must NOT substitute one of its own.api_key: str | None— Resolved credential, or None when neither the call, the model nor the vendor's environment variables supplied one. A factory that requires a key raises at request time (Fail-Early: no fabricated credential).
Returns
A provider bound to that endpoint.
AssistantMessageEventStream¶
class AssistantMessageEventStream(provider_stream: AsyncIterable[Any], model: Any)
tau_llm.streaming.AssistantMessageEventStream
Async iterator over streaming events from the LLM.
Yields: TextDeltaEvent, ToolCallDeltaEvent, DoneEvent, ErrorEvent
This class wraps an underlying provider stream (an async iterator of typed
StreamEvents) and re-publishes those events through an internal queue. A
background _collect coroutine drains the provider stream into the queue
so result() and async for can be awaited independently; the main
coroutine yields events from that queue.
Constructor parameters
provider_stream: AsyncIterable[Any]— Async-iterable yielding typed provider StreamEvents.model: Any— The Model configuration.
Usage
stream = await stream_simple(model, context, options) async for event in stream: if event.type == "text_delta": print(event.delta, end="") elif event.type == "toolcall_delta": print(f"\nTool: {event.delta.get('id', '')}") final = await stream.result()
aiter¶
__aiter__() -> 'AssistantMessageEventStream'
tau_llm.streaming.AssistantMessageEventStream.__aiter__
Return self as the async iterator.
anext¶
__anext__() -> Any
tau_llm.streaming.AssistantMessageEventStream.__anext__
Yield the next event from the internal queue.
If the stream is done and the queue is empty, raises StopAsyncIteration.
Returns
The next StreamEvent (TextDeltaEvent, ToolCallDeltaEvent, DoneEvent, or ErrorEvent).
Raises
StopAsyncIteration— When the stream is complete.
abort¶
abort() -> None
tau_llm.streaming.AssistantMessageEventStream.abort
Stop consuming the stream locally. Does NOT cancel the HTTP request.
Against the real stack this cancels the collector task and nothing else.
OpenAICompletionsProvider.stream_chat() returns a bare async generator,
which has no abort attribute, so the propagation branch below never
fires — it exists for a provider stream that chooses to expose one, and is
exercised only by a fake in the tests.
Real cancellation of an in-flight request is a separate mechanism: an
:class:~tau_llm.abort.AbortSignal passed as options={"abort_signal": ...},
which the provider polls between chunks. Reach for that, not this.
The partial state is preserved either way.
result¶
result() -> AssistantMessage
tau_llm.streaming.AssistantMessageEventStream.result
Wait for the stream to complete and return the final message.
Returns
The fully accumulated AssistantMessage.
Raises
Exception— If the stream produced an ErrorEvent. The ORIGINAL exception is re-raised with its type intact (e.g.ConstraintViolation, which carries the offending output), not flattened to a bareException.
Compat¶
tau_llm.compat.Compat
Wire quirks an operator STATES for one endpoint.
Every field is optional, and None means "not stated" rather than "false"
— :func:resolve_compat falls through to the detected value for anything
left unset, so an operator can correct one field without restating the rest.
Set per-model in ~/.tau/config.json (models.<name>.compat).
max_tokens_field¶
tau_llm.compat.Compat.max_tokens_field: MaxTokensField | None
No description. This object is marked but undocumented.
supports_usage_in_streaming¶
tau_llm.compat.Compat.supports_usage_in_streaming: bool | None
No description. This object is marked but undocumented.
tool_call_schema¶
tau_llm.compat.Compat.tool_call_schema: ToolCallSchema | None
No description. This object is marked but undocumented.
DoneEvent¶
class DoneEvent(final: AssistantMessage, usage: Usage, type: Literal['done'] = 'done')
tau_llm.streaming.DoneEvent
A done event signaling the stream is complete.
Carries the fully accumulated AssistantMessage and token usage information.
Reference: SUBPHASE-0.0.md, "4. Streaming Events" section.
Constructor parameters
final: AssistantMessage— The fully accumulated AssistantMessage.usage: Usage— Token usage information for the response.type: Literal['done'] = 'done'— Always "done".
ErrorEvent¶
class ErrorEvent(message: str, is_error: Literal[True] = True, type: Literal['error'] = 'error')
tau_llm.streaming.ErrorEvent
An error event from the LLM stream.
Carries an error message. When the stream produces an ErrorEvent, no further events will be produced.
Reference: SUBPHASE-0.0.md, "4. Streaming Events" section.
Constructor parameters
message: str— Description of the error.is_error: Literal[True] = True— Always True.type: Literal['error'] = 'error'— Always "error".
Provider¶
tau_llm.providers.base.Provider
Abstract base class for LLM chat providers.
A Provider instance is bound to ONE endpoint: it bakes in a base URL and a
credential at construction (that is why tau_llm.client pools instances
per endpoint rather than per class). id and name say which vendor
that endpoint belongs to — pi's Provider.id/.name
(models.ts:98-99).
Reference: SUBPHASE-0.0.md, Phase 1 Subphase 0 — Provider interface.
aclose¶
aclose() -> None
tau_llm.providers.base.Provider.aclose
Release anything this instance holds open (HTTP connections, …).
tau_llm.client.aclose_providers calls this on EVERY pooled provider,
so the contract has to live on the base class — otherwise a provider
registered through :func:register_api would crash the shutdown path of
an application that never heard of it. The default is a no-op because a
provider that opens nothing has nothing to close; one that does (see
OpenAICompletionsProvider.aclose) overrides it, and must stay
idempotent — the pool may close an instance that never issued a request.
id¶
tau_llm.providers.base.Provider.id: str
No description. This object is marked but undocumented.
name¶
tau_llm.providers.base.Provider.name: str
No description. This object is marked but undocumented.
stream_chat¶
stream_chat(model: Model, messages: list[Any], tools: list[ToolSpec] | None = None, options: dict[str, Any] | None = None) -> StreamEventStream
tau_llm.providers.base.Provider.stream_chat
Stream chat completions from the LLM.
Parameters
model: Model— The Model configuration to use for the request.messages: list[Any]— List of τ message objects (user/assistant/toolResult).tools: list[ToolSpec] | None = None— Optional list of objects satisfying ToolSpec. The agent loop passesAgentToolwrappers, notToolDefinition; see ToolSpec for why this is a Protocol.options: dict[str, Any] | None = None— Optional provider-specific options (temperature, max_tokens, etc.).
Returns
A StreamEventStream — an async iterator of typed streaming events (TextDeltaEvent, ThinkingDeltaEvent, ToolCallDeltaEvent, DoneEvent, ErrorEvent). stream_simple wraps it in AssistantMessageEventStream, which exposes the terminal AssistantMessage via result().
Raises
NotImplementedError— If the provider hasn't implemented this method.
ProviderSpec¶
class ProviderSpec(id: str, api: str, name: str = '', base_url: str | None = None, api_key_env: tuple[str, ...] = tuple())
tau_llm.providers.base.ProviderSpec
One vendor, as data.
Everything τ needs to reach an OpenAI-compatible vendor that is not OpenAI.
pi's CreateProviderOptions (models.ts:738) minus the parts τ has no
consumer for — see this module's docstring and Provider above.
Constructor parameters
id: str— Vendor id. MatchesModel.provider; that is the lookup key.api: str— The ONE wire protocol this vendor speaks. A model claiming a differentapifor this vendor is a configuration error, not a request to be attempted — seeclient._resolve_request. A gateway that genuinely speaks two protocols registers two ids.name: str = ''— Display name. Defaults toid, as pi'screateProviderdoes.base_url: str | None = None— Default endpoint, used only when theModelcarries nobase_urlof its own. None means the vendor has no fixed endpoint (a self-hosted server), and a model must then supply one.api_key_env: tuple[str, ...] = tuple()— Environment variables searched, in order, for this vendor's credential. Empty means "this vendor has no known credential source"; a NON-empty tuple that resolves to nothing is a hard error rather than a silent fall-through to another vendor's key.
resolve_api_key¶
resolve_api_key() -> str | None
tau_llm.providers.base.ProviderSpec.resolve_api_key
First non-empty value among api_key_env, or None if none is set.
ResolvedCompat¶
tau_llm.compat.ResolvedCompat
:class:Compat with every field decided. What the provider reads.
max_tokens_field¶
tau_llm.compat.ResolvedCompat.max_tokens_field: MaxTokensField
No description. This object is marked but undocumented.
supports_usage_in_streaming¶
tau_llm.compat.ResolvedCompat.supports_usage_in_streaming: bool
No description. This object is marked but undocumented.
tool_call_schema¶
tau_llm.compat.ResolvedCompat.tool_call_schema: ToolCallSchema
No description. This object is marked but undocumented.
StreamEventStream¶
tau_llm.providers.base.StreamEventStream
Structural return type for Provider.stream_chat.
A provider stream is async-iterable over typed streaming events
(TextDeltaEvent / ThinkingDeltaEvent / ToolCallDeltaEvent / DoneEvent /
ErrorEvent). The client (stream_simple) wraps it once in
AssistantMessageEventStream (streaming.py) — the single stream type
that adds queue buffering and the terminal result().
aiter¶
__aiter__() -> AsyncIterator[Any]
tau_llm.providers.base.StreamEventStream.__aiter__
No description. This object is marked but undocumented.
TextDeltaEvent¶
class TextDeltaEvent(delta: str, partial: AssistantMessage, type: Literal['text_delta'] = 'text_delta')
tau_llm.streaming.TextDeltaEvent
A text delta event from the LLM stream.
Carries a partial text chunk and the partially accumulated message. The consumer should append delta to the partial message's text content.
Reference: SUBPHASE-0.0.md, "4. Streaming Events" section.
Constructor parameters
delta: str— The text chunk from this event.partial: AssistantMessage— The partially accumulated AssistantMessage.type: Literal['text_delta'] = 'text_delta'— Always "text_delta".
ThinkingDeltaEvent¶
class ThinkingDeltaEvent(delta: str, partial: AssistantMessage, type: Literal['thinking_delta'] = 'thinking_delta')
tau_llm.streaming.ThinkingDeltaEvent
A thinking/reasoning delta event from the LLM stream.
Mirrors :class:TextDeltaEvent but carries reasoning content — the
OpenAI-compatible reasoning_content / reasoning / reasoning_text
delta fields (llama.cpp, vLLM, DeepSeek, OpenRouter). Kept as a distinct
event so consumers can render reasoning separately from the answer and
collapse it once the answer/tool content begins.
Reference: pi openai-completions.ts thinking_delta event.
Constructor parameters
delta: str— The reasoning chunk from this event.partial: AssistantMessage— The partially accumulated AssistantMessage.type: Literal['thinking_delta'] = 'thinking_delta'— Always "thinking_delta".
ToolCallDeltaEvent¶
class ToolCallDeltaEvent(delta: dict[str, Any], partial: AssistantMessage, type: Literal['toolcall_delta'] = 'toolcall_delta')
tau_llm.streaming.ToolCallDeltaEvent
A tool call delta event from the LLM stream.
Carries a partial tool call update and the partially accumulated message. Multiple deltas for the same tool call are accumulated until DoneEvent.
Reference: SUBPHASE-0.0.md, "4. Streaming Events" section.
Constructor parameters
delta: dict[str, Any]— The OpenAI-style tool call delta dict.partial: AssistantMessage— The partially accumulated AssistantMessage.type: Literal['toolcall_delta'] = 'toolcall_delta'— Always "toolcall_delta".
aclose_providers¶
aclose_providers() -> None
tau_llm.client.aclose_providers
Close and drop every pooled provider for the CURRENT event loop.
Explicit teardown (docs/PROVIDER-LIFETIME.md §6.3: "closed explicitly,
not by GC"). Call this from a shutdown path that still runs on the loop
the providers were built on — a TUI's unmount handler, or a headless
run's finally before the driving asyncio.run() returns. Calling
it again (or with nothing pooled) is a no-op; a subsequent
stream_simple call rebuilds providers on demand.
complete_simple¶
complete_simple(model: Any, context: dict[str, Any], options: dict[str, Any] | None = None) -> AssistantMessage
tau_llm.client.complete_simple
Whole-message completion — drive a stream to its terminal message.
Faithful port of pi's completeSimple (stream.ts:67), which is simply
stream(...).result(). Used where the caller wants the whole
AssistantMessage and not the intermediate deltas — e.g. compaction's
summary generation, which has no streaming UI to feed.
This is about the CALLER's shape, not the transport: it collapses the event
stream for a caller that has no use for deltas, and the request underneath is
still whatever Model.stream / options["stream"] selected (PLAN-0.9.3
§4.1). A non-streaming BACKEND is a separate axis — set stream=False and
every entry point here, this one included, keeps working unchanged.
Parameters
model: Any— The Model configuration (has provider, id, etc.).context: dict[str, Any]— Context dict (same shape asstream_simple):messagesand optionaltools. A leading{"role": "system", ...}message sets the system prompt (client.py does not readsystem_prompt).options: dict[str, Any] | None = None— Optional provider options (max_tokens,api_key,reasoning,temperature, …).
Returns
The fully accumulated AssistantMessage.
Raises
Exception— If the stream produced an ErrorEvent (propagated byAssistantMessageEventStream.result). Fail-Early: no fabricated fallback message.
detect_compat¶
detect_compat(provider: str, base_url: str) -> ResolvedCompat
tau_llm.compat.detect_compat
Infer wire quirks from the endpoint URL.
Only the URL decides anything today. pi also matches on the provider NAME,
and τ cannot: build_model_from_config defaults an entry with no
backend key to provider="openai", so in τ that string means "the
operator did not say" far more often than it means OpenAI. Matching it would
switch a local llama.cpp to a spelling llama.cpp rejects, which is the
regression this whole module is arranged to avoid.
A proxy in front of real OpenAI therefore goes undetected. That is the
correct trade: it is one explicit compat.max_tokens_field in the config,
where the alternative silently breaks endpoints that work today.
Parameters
provider: str—Model.provider. Accepted for signature stability and to keep the pi correspondence readable; not consulted, for the reason above.base_url: str—Model.base_url— the endpoint the request goes to.
Returns
A fully-decided :class:ResolvedCompat.
get_api_factory¶
get_api_factory(api: str) -> ApiFactory
tau_llm.providers.base.get_api_factory
The factory for api.
Parameters
api: str— (no description)
Raises
ValueError— Ifapiis unknown. The message names what was asked for and what exists — an unimplemented protocol must fail loudly here, never fall through to whichever client happens to be built in.
get_provider_spec¶
get_provider_spec(provider_id: str) -> ProviderSpec | None
tau_llm.providers.base.get_provider_spec
The vendor's spec, or None if it was never registered.
None is not an error: an unregistered vendor id is a free-form label on a
Model that already carries its own endpoint. Only a model whose api
is unknown cannot be served.
Parameters
provider_id: str— (no description)
register_api¶
register_api(api: str, factory: ApiFactory, *, replace: bool = False) -> None
tau_llm.providers.base.register_api
Register the factory that builds clients for one wire protocol.
Parameters
api: str— Wire-protocol id, matchingModel.api(e.g."openai-completions").factory: ApiFactory— See :class:ApiFactory.replace: bool = False— Required to overwrite an existing registration. Without it a second registration of the same id raises: two libraries silently fighting over which client serves a protocol is the kind of invisible mis-routing this whole module exists to prevent.
Raises
ValueError— On an empty id, or on a duplicate withoutreplace=True.
register_provider¶
register_provider(spec: ProviderSpec, *, replace: bool = False) -> None
tau_llm.providers.base.register_provider
Register a vendor's defaults.
Parameters
spec: ProviderSpec— The vendor, as data.replace: bool = False— Required to overwrite an existing registration, for the same reason as :func:register_api— a silently redirected vendor sends prompts and credentials somewhere the operator did not choose.
Raises
ValueError— On a duplicate id withoutreplace=True.
registered_apis¶
registered_apis() -> tuple[str, ...]
tau_llm.providers.base.registered_apis
Every registered wire-protocol id, sorted.
registered_providers¶
registered_providers() -> tuple[str, ...]
tau_llm.providers.base.registered_providers
Every registered vendor id, sorted.
resolve_compat¶
resolve_compat(model: Model) -> ResolvedCompat
tau_llm.compat.resolve_compat
The compat τ will actually use for model.
Detection first, then the operator's stated :class:Compat over the top,
field by field — an unset field falls through to the detected value rather
than to a type default, so stating one quirk never silently resets another.
Mirrors pi's getCompat (openai-completions.ts:1631).
Parameters
model: Model— (no description)
split_tool_result_content¶
split_tool_result_content(content: Any) -> tuple[list[str], list[tuple[str, str]]]
tau_llm.providers.base.split_tool_result_content
A tool result's content as (text parts, [(mime_type, base64 data), ...]).
Every client needs the same split, because every client has to put the text
somewhere the wire format calls a tool result and the images somewhere it
does not. It lives here rather than three times over: the copies in
openai.py and google.py had already drifted on the join separator and
on which non-block shapes they tolerated, and a shape rule that differs per
provider is a shape rule nobody can state.
Accepts the three shapes a tool result reaches a client in: a bare string,
a list of pydantic blocks (the live path, ToolResultMessage.content), and
a list of raw dicts (the persisted path — a reloaded session arrives as
model_dump()ed messages).
Text parts come back unjoined, because the separator is the caller's: the OpenAI client has always joined with a space and the Google client with an empty string, and quietly changing how a multi-block text result reads is not something an image change should do.
Parameters
content: Any— The tool result's content, in any of the three shapes above.
Returns
A (text_parts, images) pair. images holds (mime_type, data) with data still base64-encoded.
Raises
TypeError— Ifcontentis neither a string nor an iterable of blocks, or if it holds a block of a type this cannot read. Fabricating text from an unreadable value is how an image became a filename in the first place; a wrong tool result must not reach the model looking like a right one.
stream_simple¶
stream_simple(model: Any, context: dict[str, Any], options: dict[str, Any] | None = None) -> AssistantMessageEventStream
tau_llm.client.stream_simple
Simple streaming client for the agent loop.
This is the ONLY entry point that τ-agent-core uses to talk to τ-llm.
Parameters
model: Any— The Model configuration (has provider, id, etc.).context: dict[str, Any]— Context dict with keys: - messages: List of message dicts (user/assistant/toolResult). - tools: Optional list of tool definitions. - system_prompt: Optional system prompt string.options: dict[str, Any] | None = None— Optional provider-specific options (temperature, etc.). Two of them are TRANSPORT settings the provider strips from the request body rather than sending:request_timeoutandstream(False = talk to a backend that does not implement SSE; the events below are produced either way — PLAN-0.9.3 §4.1).
Returns
AssistantMessageEventStream yielding TextDeltaEvent, ToolCallDeltaEvent, DoneEvent, and ErrorEvent instances.
Raises
ValueError— If the model names a wire protocol τ has no implementation for, contradicts its own vendor's protocol, resolves to no endpoint, or to no credential for a vendor that declares where its credential lives. See_resolve_request.
unregister_api¶
unregister_api(api: str) -> None
tau_llm.providers.base.unregister_api
Remove a wire-protocol registration.
Parameters
api: str— (no description)
Raises
KeyError— If nothing is registered under that id — undoing a registration that never happened means the caller's model of the registry is wrong, and saying so is cheaper than not.
unregister_provider¶
unregister_provider(provider_id: str) -> None
tau_llm.providers.base.unregister_provider
Remove a vendor registration.
Parameters
provider_id: str— (no description)
Raises
KeyError— If that vendor was never registered.