Text-to-Speech
Convert text to speech with eleven providers:Platform Delivery
Configuration
onnxruntime or PyTorch wheels for those platforms. Selecting it there reports the provider unavailable.
MiniMax TTS selects its region, endpoint, and credential together:
region: "global"useshttps://api.minimax.io/v1/t2a_v2withMINIMAX_API_KEY.region: "cn"useshttps://api.minimaxi.com/v1/t2a_v2withMINIMAX_CN_API_KEY.- If
regionis omitted,MINIMAX_API_KEYkeeps precedence for backward compatibility. If onlyMINIMAX_CN_API_KEYis configured, Mibyan selectscn. - An explicitly selected region must have its matching credential. Mibyan never borrows the other region’s key. A
base_urloverride does not change the selected credential, and an override pointing at the other region’s official endpoint is rejected.
tts.speed value applies to all providers by default. Each provider can override it with its own speed setting (e.g., tts.openai.speed: 1.5). Provider-specific speed takes precedence over the global value. Default is 1.0 (normal speed).
Gemini Persona Prompts
Gemini TTS can follow natural-language performance direction. Settts.gemini.persona_prompt_file to a local Markdown or text file that describes the voice persona. The file can include Gemini-style sections such as AUDIO PROFILE, SCENE, DIRECTOR'S NOTES, SAMPLE CONTEXT, and TRANSCRIPT.
If the file contains {transcript} or {{ transcript }}, Mibyan replaces that placeholder with the live TTS text. Otherwise, Mibyan appends a labeled TRANSCRIPT section automatically. The persona prompt stays local and is not shown in the chat reply.
Audio Tags (Gemini, xAI)
Google’s Gemini 3.1 Flash TTS and xAI’s Grok TTS support freeform square-bracket audio tags such as[whispers], [excitedly], [very slow], [laughs], and other expressive delivery notes. Enable tts.gemini.audio_tags or tts.xai.auto_speech_tags to have Mibyan run a hidden rewrite pass before TTS. The rewrite inserts inline tags into the TTS script only; the visible chat reply stays unchanged.
auxiliary.tts_audio_tags and defaults to your main chat model. Override that auxiliary task if you want tag insertion handled by a cheaper or faster model.
Streaming sample rate (OpenAI-compatible endpoints): streaming playback receives headerless raw PCM, so Mibyan must know its sample rate. The official OpenAI API emits 24 kHz. A compatible server that reports its rate — the X-Audio-Sample-Rate response header, or rate= in the Content-Type (audio/pcm; rate=44100) — is honored automatically: the speaker, the temp-WAV player and the gateway audio stream all open at the reported rate once the response arrives. For servers that report nothing, set tts.openai.pcm_sample_rate to the endpoint’s output rate (e.g. 22050 for Piper-backed servers); otherwise speech plays at the wrong speed and pitch. Invalid values log a warning and fall back to 24000.
Language (OpenAI-compatible endpoints): tts.openai.language is forwarded to the endpoint as a lang_code request parameter. It is intended for OpenAI-compatible TTS servers that support lang_code — for example Kokoro-FastAPI, where language: "es" selects the Spanish phonemizer instead of the English default. Leave it unset when using the official OpenAI API, which does not accept this parameter. When unset, nothing extra is sent.
Cloned-voice consent (OpenAI-compatible endpoints): some self-hosted OpenAI-compatible TTS servers reject a cloned voice with 400 consent_required unless the request carries a consent_attestation field. Set tts.openai.consent_attestation to the attestation text your server expects; Mibyan forwards it verbatim in the request body on every OpenAI-compatible path (whole-file synthesis, streaming, and the desktop’s client-direct voice). Leave it unset for the official OpenAI API — when unset, the field is not sent.
Input length limits
Each provider has a documented per-request input-character cap. Mibyan splits longer replies into ordered, sentence-aware chunks before calling the provider, so the full normalized text is preserved instead of silently truncated:
ElevenLabs picks a cap from the configured
model_id:
Override per provider with
max_text_length: under the provider section of your TTS config:
Telegram Voice Bubbles & ffmpeg
Telegram voice bubbles require Opus/OGG audio format:- OpenAI, ElevenLabs, and Mistral produce Opus natively — no extra setup
- Edge TTS (default) outputs MP3 and needs ffmpeg to convert:
- MiniMax TTS outputs MP3 and needs ffmpeg to convert for Telegram voice bubbles
- Google Gemini TTS outputs raw PCM and uses ffmpeg to encode Opus directly for Telegram voice bubbles
- xAI TTS outputs MP3 and needs ffmpeg to convert for Telegram voice bubbles
- NeuTTS outputs WAV and also needs ffmpeg to convert for Telegram voice bubbles
- KittenTTS outputs WAV and also needs ffmpeg to convert for Telegram voice bubbles
- Piper outputs WAV and also needs ffmpeg to convert for Telegram voice bubbles
xAI Custom Voices (voice cloning)
xAI supports cloning your voice and using it with TTS. Create a custom voice in the xAI Console, then set the resultingvoice_id in your config:
Piper (local, 44 languages)
Piper is a fast, local neural TTS engine from the Open Home Foundation (the Home Assistant maintainers). It runs entirely on CPU, supports 44 languages with pre-trained voices, and needs no API key. Install viamibyan tools → Voice & TTS → Piper. Mibyan requests the
piper extra through PM. From a prepared source checkout, the explicit command
is python -c "import pm; pm.sync_venv(['piper'], explicit=True)". Platform markers still apply.
Switch to Piper:
python -m piper.download_voices <name> and downloads the model (~20-90MB depending on quality tier) into ~/.mibyan/cache/piper-voices/. Subsequent calls reuse the cached model.
Picking a voice. The full voice catalog covers English, Spanish, French, German, Italian, Dutch, Portuguese, Russian, Polish, Turkish, Chinese, Arabic, Hindi, and more — each with x_low / low / medium / high quality tiers. Sample voices at rhasspy.github.io/piper-samples.
Using a pre-downloaded voice. Set tts.piper.voice to an absolute path ending in .onnx:
tts.piper.length_scale / noise_scale / noise_w_scale / volume / normalize_audio, use_cuda) correspond 1:1 to Piper’s SynthesisConfig. They’re ignored on older piper-tts versions.
Warm-up and unload via speech toggles (local engines)
Local engines (Piper, KittenTTS) load their model lazily, so without help the first spoken reply after you turn speech on pays the whole model load — and on a fresh install the voice download — as silence before the first word. Mibyan treats the speech-output toggles as the signal that TTS is about to be needed:- Desktop — Read replies aloud is a desktop-local preference, independent of the gateway’s
voice.auto_ttssetting in Settings → Voice. It migrates the shared value once, then later gateway configuration changes do not override the desktop toggle. If local storage is full or unavailable, the choice still lasts for this window; persistence across a reload remains best-effort. Turning on Read replies aloud, or starting a voice conversation, pre-loads the configured engine in the background right away. Turning both off again unloads the resident model (a Piper voice is tens of MB; KittenTTS up to ~80MB) so it isn’t parked in RAM for nothing. - CLI / TUI —
/voice tts(and/voice onwhenvoice.auto_ttsis set) do the same;/voice offreleases.
tts.keep_warm_seconds (default 60) after the last release, and any toggle turning speech back on within that window keeps the loaded model, so a wake-word loop or a quickly restarted voice conversation doesn’t reload the voice each time. Set it to 0 to unload immediately. For cloud providers there is no model to hold — the toggle only makes sure a lazily-installed SDK (edge-tts, ElevenLabs, Mistral) is present. Warm-up is best-effort: if the engine can’t load, the toggle still succeeds and the first reply falls back to loading on demand as before.
The Desktop calls POST /api/audio/tts-lease with {"lease": "<name>", "active": true|false}; other frontends can use the same endpoint.
The same lease also reaches user-declared providers, so a self-hosted TTS server can preload and unload its model on the toggles: a command provider runs its optional warm_command / release_command, and a Python plugin provider gets warm() / release().
Custom command providers
If a TTS engine you want isn’t natively supported (VoxCPM, MLX-Kokoro, XTTS CLI, a voice-cloning script, anything else that exposes a CLI), you can wire it in as a command-type provider without writing any Python. Mibyan writes the input text to a temp UTF-8 file, runs your shell command, and reads the audio file the command produced. Declare one or more providers undertts.providers.<name> and switch between them with tts.provider: <name> — the same way you switch between built-ins like edge and openai.
output_format values: mp3 (default), wav, ogg, flac, m4a, aac, amr, opus. Your command must actually produce that format (e.g. via ffmpeg); Mibyan only validates the declared value and names the output file accordingly. An unknown value falls back to mp3. The chosen format is also exposed to the command as the {format} placeholder.
Subprocess environment: command providers (TTS and STT) run with Mibyan secrets scrubbed from the child environment — gateway bot tokens, LLM provider API keys, and internal relay credentials are removed; PATH, HOME, locale, and other normal variables are kept. If your command template needs its own API key from the environment (e.g. a curl one-liner), list the variable names under env_passthrough in the provider config:
Example: Doubao (Chinese seed-tts-2.0)
For high-quality Chinese TTS via ByteDance’s seed-tts-2.0 bidirectional-streaming API, install thedoubao-speech PyPI package and wire it in as a command provider:
Install this external command provider in its own tool environment, not in
Mibyan’s Python environment. Make its executable available on PATH.
VOLCENGINE_APP_ID / VOLCENGINE_ACCESS_TOKEN) or ~/.doubao-speech/config.yaml. Pick a voice by adding --voice zh-female-warm (or any other alias from doubao-speech list-voices) to the command. doubao-speech also bundles streaming ASR — see the STT section below for Mibyan integration. Source and full docs: github.com/Hypnus-Yuan/doubao-speech.
Placeholders
Your command template can reference these placeholders. Mibyan substitutes them at render time and shell-quotes each value for the surrounding context (bare / single-quoted / double-quoted), so paths with spaces and other shell-sensitive characters are safe.
Use
{{ and }} for literal braces.
Optional keys
Behavior notes
- Built-in names always win. A
tts.providers.openaientry never shadows the native OpenAI provider, so no user config can silently replace a built-in. - Default delivery is a document. Command providers deliver as regular audio attachments on every platform. Opt in to voice-bubble delivery per-provider with
voice_compatible: true. - Command failures surface to the agent. Non-zero exit, empty output, or timeout all return an error with the command’s stderr/stdout included so you can debug the provider from the conversation.
type: commandis the default whencommand:is set. Writingtype: commandexplicitly is good practice but not required; an entry with a non-emptycommandstring is treated as a command provider.{input_path}/{text_path}are interchangeable. Use whichever reads better in your command.
Security
Command-type providers run whatever shell command you configure, with your user’s permissions. Mibyan quotes placeholder values and enforces the configured timeout, but the command template itself is trusted local input — treat it the same way you would a shell script on your PATH.Python plugin providers
For TTS engines that can’t be expressed as a single shell command — Python SDKs without a CLI, streaming engines, voice-listing APIs, OAuth-refreshing auth — register a Python plugin viactx.register_tts_provider(). The plugin coexists with (does not replace) the Custom command providers registry; pick the surface that fits your engine.
When to pick which
Built-ins always win, and command providers win over a same-name plugin — so plugins are safe to register against any non-built-in name without worrying about shadowing your existing config.
Minimal plugin
Drop this in~/.mibyan/plugins/my-tts/:
plugin.yaml:
__init__.py:
mibyan plugins enable my-tts), point tts.provider at it (tts.provider: my-tts in config.yaml), and the text_to_speech tool will route through your plugin.
Optional hooks
Override these on your provider class for richer integration:list_voices()→ list of{id, display, language, gender, preview_url}dicts shown inmibyan tools.list_models()→ list of{id, display, languages, max_text_length}dicts.get_setup_schema()→ return{name, badge, tag, env_vars: [{key, prompt, url}]}to power the picker row inmibyan tools/mibyan setup. Without this, the plugin still works but its row in the picker is minimal.stream(text, *, voice, model, format, **extra)→ iterator yielding audio bytes for streaming delivery (default raisesNotImplementedError).voice_compatibleproperty → setTrueif your output is Opus-compatible and the gateway should deliver it as a voice bubble (defaultFalse= regular audio attachment).warm()/release()→ called when a surface toggles speech output on / when the last lease across surfaces is released, while your provider is the configuredtts.provider— preload or unload a local model server here. Both default to no-ops; exceptions are logged at debug and never fail the toggle.
agent/tts_provider.py for the full ABC including docstrings.
Voice Message Transcription (STT)
Voice messages sent on Telegram, Discord, WhatsApp, Slack, or Signal are automatically transcribed and injected as text into the conversation. The agent sees the transcript as normal text.Zero ConfigLocal transcription works out of the box when
faster-whisper is installed. If that’s unavailable, Mibyan can also use a local whisper CLI from common install locations (like /opt/homebrew/bin) or a custom command via mibyan_LOCAL_STT_COMMAND.Configuration
Provider Details
Local (faster-whisper) — Runs Whisper locally via faster-whisper. Uses CPU by default, GPU if available. Model sizes:
The first use downloads the selected model from
huggingface.co; later loads prefer the local cache and do not require an online revision check. On networks where the Hub is unavailable, export an accessible mirror in the shell or service that starts Mibyan:
HF_HUB_DISABLE_XET=1 keeps downloads on the mirror’s regular HTTP path instead of contacting Xet CAS hosts that do not honor HF_ENDPOINT.
Groq API — Requires GROQ_API_KEY. Good cloud fallback when you want a free hosted STT option. Set stt.groq.language (or the global mibyan_LOCAL_STT_LANGUAGE env var) to skip Whisper’s auto-detect and reduce latency on known-language audio.
OpenAI API — Accepts VOICE_TOOLS_OPENAI_KEY first and falls back to OPENAI_API_KEY. Supports whisper-1, gpt-4o-mini-transcribe, gpt-4o-transcribe, and gpt-transcribe.
Mistral API (Voxtral Transcribe) — Requires MISTRAL_API_KEY. Uses Mistral’s Voxtral Transcribe models. Supports 13 languages, speaker diarization, and word-level timestamps. Install with cd ~/.mibyan/mibyan-agent && python -c "import pm; pm.sync_venv(['mistral'], explicit=True)".
xAI Grok STT — Requires XAI_API_KEY. Posts to https://api.x.ai/v1/stt as multipart/form-data. Good choice if you’re already using xAI for chat or TTS and want one API key for everything. Auto-detection order puts it after Groq — explicitly set stt.provider: xai to force it.
Custom local CLI fallback — Set mibyan_LOCAL_STT_COMMAND if you want Mibyan to call a local transcription command directly. The command template supports {input_path}, {output_dir}, {language}, and {model} placeholders. Mibyan tokenizes the rendered template into an argument list and executes it without a shell, so operators such as |, >, &&, and ; are passed as literal arguments. Your command must write a .txt transcript somewhere under {output_dir}.
Example: Doubao / Volcengine ASR
If you usedoubao-speech for Doubao TTS (see above), the same package handles speech-to-text via the local-command STT surface:
Install this external command provider in its own tool environment, not in
Mibyan’s Python environment. Make its executable available on PATH.
cmd /c or PowerShell wrapper instead. An explicit wrapper makes shell interpretation an opt-in part of the configured argv rather than an implicit property of every local STT template.
{input_path}, runs the command, and reads the .txt file produced under {output_dir}. Language is auto-detected by the Volcengine bigmodel endpoint.
Fallback Behavior
An explicitstt.provider selection (written in config.yaml, e.g. via mibyan tools) is honored strictly — if that provider can’t run, transcription fails with a clear error (stt is configured to use <provider> (set via mibyan tools), but <failure>. Run 'mibyan tools' to change it.) instead of silently switching engines. Note that stt.provider: local written in your config counts as an explicit selection.
When no provider has ever been selected, Mibyan auto-detects from what’s available:
- Local faster-whisper unavailable → Tries a local
whisperCLI ormibyan_LOCAL_STT_COMMANDbefore cloud providers - Groq key not set → Skipped; next available provider
- OpenAI key not set → Skipped; next available provider
- Mistral key/SDK not set → Skipped in auto-detect; falls through to next available provider
- Nothing available → Voice messages pass through with an accurate note to the user
STT custom command providers
If the STT engine you want isn’t natively supported (Doubao ASR, NVIDIA Parakeet, a whisper.cpp build, an open-source SenseVoice CLI, anything else that exposes a shell command), wire it in as a command-type provider without writing any Python. Mibyan runs your shell command against the audio file and reads back the transcript. Declare one or more providers understt.providers.<name> and switch between them with stt.provider: <name> — same shape as the TTS command-provider registry, adapted for the input=audio → output=transcript direction.
mibyan_LOCAL_STT_COMMAND escape hatch via the built-in local_command path. Unlike the shell-driven command-provider registry, the legacy template is tokenized into argv and runs without implicit shell interpretation. Use stt.providers.<name> when you want multiple shell-driven STT engines, a name you can pick via stt.provider, or anything that needs per-provider language / model / timeout.
STT placeholders
Your command template can reference these placeholders. Mibyan substitutes them at render time and shell-quotes each value for the surrounding context (bare / single-quoted / double-quoted), so paths with spaces are safe.
Use
{{ and }} for literal braces (handy when embedding JSON snippets in the command).
How the transcript is read back
After your command exits successfully:- If
{output_path}exists and is non-empty → Mibyan reads it as UTF-8 text. - Otherwise, if the command wrote to stdout → Mibyan uses that.
- Otherwise → error: “Command STT provider wrote no output file and produced no stdout”.
whisper-cli, parakeet-asr) and curl-style one-liners that emit transcript to stdout (curl … | jq -r .text).
For format: json / srt / vtt, Mibyan returns the raw file content as the transcript field. Extracting .text from JSON is out of scope for the runner — either configure format: txt, or post-process JSON downstream.
STT command-provider optional keys
STT command-provider behavior notes
- Built-ins always win. Declaring
stt.providers.openai: type: commanddoes NOT override the real OpenAI Whisper handler. The built-in name is short-circuited before the command-provider resolver runs. - Process-tree cleanup. A command running over
timeouthas its entire process tree killed, not just the shell wrapper. Long-running ASR pipelines that fork model-loading subprocesses are reaped reliably. - Shell-quoting is automatic. Placeholders inside
'…'get single-quote-safe escaping; inside"…"get$/`/"escaping; outside quotes getshlex.quote. Don’t pre-quote placeholder values.
STT command-provider security
The shell command runs under the same user as Mibyan with full filesystem access — same trust model astts.providers.<name>: type: command and mibyan_LOCAL_STT_COMMAND. Only declare command providers from sources you trust.
Python plugin providers (STT)
For STT engines that aren’t built-in AND can’t be expressed as a shell command (need a Python SDK, OAuth-refreshing auth, streaming chunks, etc.), register a Python plugin viactx.register_transcription_provider(). The plugin coexists with the 8 built-in providers (local, local_command, groq, openai, mistral, xai, elevenlabs, deepinfra) and the stt.providers.<name>: type: command registry — built-ins keep their native implementations and always win on name collision; command providers win over plugins of the same name (config is more local than plugin install).
When to pick which (STT)
Resolution order
stt.provideris a built-in name → built-in dispatch. Always wins.stt.providermatchesstt.providers.<name>withcommand:set → command-provider runner (see STT custom command providers). Wins over a same-name plugin.stt.providermatches a plugin-registeredTranscriptionProvider→ plugin dispatch:- if the plugin’s
is_available()returnsFalse(missing creds or SDK), the call surfaces an unavailability error envelope identifying the plugin — not the generic “No STT provider available” message. - otherwise the plugin’s
transcribe()is called withmodel(from the publicmodel=arg, falling back tostt.<provider>.model) andlanguage(fromstt.<provider>.language).
- if the plugin’s
- No match → “No STT provider available” error.
Per-provider config namespace
Plugins read their per-provider configuration fromstt.<provider> in config.yaml, mirroring how built-ins read stt.openai.model / stt.mistral.model:
model and language from this section; everything else, the plugin can read itself.
Minimal plugin
Drop this in~/.mibyan/plugins/my-stt/:
plugin.yaml:
__init__.py:
mibyan plugins enable my-stt), set stt.provider: my-stt in config.yaml, and voice-message transcription will route through your plugin.
Optional hooks
Override these on your provider class for richer integration:list_models()→ list of{id, display, languages, max_audio_seconds}dicts.default_model()→ string returned when the user doesn’t override the model.get_setup_schema()→ return{name, badge, tag, env_vars: [{key, prompt, url}]}to power picker rows inmibyan tools/mibyan setup(the picker category for STT is not yet shipped — this metadata is available to plugins for forward compatibility).
agent/transcription_provider.py for the full ABC including docstrings.
