StreamingTTSConsumer — any platform adapter that opts into streaming audio.
Voice replies start speaking after the first clause instead of after full
generation + synthesis.
Architecture
The streaming pipeline has four parts:- Producer — the LLM emits text deltas as it generates a response
- Sentence chunker —
tools.tts_streaming.SentenceChunkeraccumulates deltas, strips<think>blocks (even split across deltas), and flushes complete sentences - TTS provider — a registered
StreamingTTSProviderturns each sentence into raw PCM chunks (int16 mono at the provider’s declaredsample_rate) - Audio sink —
sounddevice.OutputStreamfor local playback (tools.tts_tool_speaker.stream_tts_to_speaker), or a gateway platform adapter’swrite_streaming_ttsseam (gateway/streaming_tts_consumer.py)
text_to_speech_tool path, so edge (the default) is conversational too.
All spoken text is cleaned by tools.tts_text_normalize.prepare_spoken_text
(one cleaner, all paths).
How to pick a provider
By default the dispatcher streams with the provider you already configured (tts.provider) when that provider has a chunked API — it never silently
swaps your voice for a different provider just to get streaming.
To override, set tts.streaming.provider in your config.yaml:
- a provider name (
elevenlabs,gemini,openai,xai) pins that streamer autowalks the priority listelevenlabs → gemini → openai → xaiand uses the first one whose credentials resolve — an explicit opt-in to “best chunked voice available”
Capability matrix
All credential lookups go through
resolve_provider_secret()
(config > env/.env > credential pool) — never bare env reads. Streamed bodies
are capped at 16 MiB per sentence, mirroring the sync providers’ bounded
upstream-body invariant.
Adding a new streaming provider
- Subclass
StreamingTTSProviderintools/tts_streaming.py - Set
sample_rate(andchannels/sample_widthif not int16 mono) - Implement
available()(a pure probe — never install anything) andstream(self, text) -> Iterator[bytes]yielding raw PCM chunks - Decorate with
@register("yourname") - Add tests in
tests/tools/test_tts_streaming.py
stream_tts_to_speaker) and the gateway consumer handle the
sentence buffer, stop events, and audio sink for free.
Gateway streaming (platform adapters)
gateway/streaming_tts_consumer.py bridges agent deltas to an adapter’s
streaming-audio seam. Adapters opt in by overriding, on
BasePlatformAdapter:
supports_streaming_tts(chat_id, audio_format) -> boolbegin_streaming_tts / write_streaming_tts / finish_streaming_tts / abort_streaming_tts

