Prerequisites
Before using voice features, make sure you have:- Mibyan installed — via the install script (see Installation)
- An LLM provider configured — run
mibyan modelor set your preferred provider credentials in~/.mibyan/.env - A working base setup — run
mibyanto verify the agent responds to text before enabling voice
Overview
Requirements
Python Packages
Usemibyan tools to configure voice providers. Missing built-in feature
requirements go through PM, subject to security.allow_lazy_installs and the
target’s dependency support. Restart Mibyan if the selected dependency
environment changes.
A bundled app includes its supported engine dependencies. Docker includes a
curated subset and disables on-demand installs. Do not use pip to modify a
signed payload or the system Python. For a manual development environment,
select the required extras in the
development setup.
Local Faster-Whisper is excluded on native Windows ARM64 and Intel macOS.
Use a cloud or command-based STT provider on those targets.
audio-io contains
the microphone/playback dependencies without local STT. The all extra does
not mean every voice or wake engine.
NeuTTS is a separate optional runtime and downloads models on first use.
Do not install its dependencies into a signed app or system Python.
discord.py[voice] installs PyNaCl (for voice encryption) and opus bindings automatically. This is required for Discord voice channel support.System Dependencies
API Keys
Add to~/.mibyan/.env:
huggingface.co. If that host is unavailable on your network, export an accessible mirror in the shell or service that starts Mibyan:
CLI Voice Mode
Voice mode is available in both the classic CLI (mibyan chat) and the TUI (mibyan --tui). Behavior is identical across both — same slash commands, same VAD silence detection, same streaming TTS, same hallucination filter. The TUI additionally forwards crash-forensic logs to ~/.mibyan/logs/ so push-to-talk failures on exotic audio backends can be reported with a full stack trace rather than disappearing silently.
Quick Start
Start the CLI and enable voice mode:How It Works
- Start the CLI with
mibyanand enable voice mode with/voice on - Press Ctrl+B — a beep plays (880Hz), recording starts
- Speak — a live audio level bar shows your input:
● [▁▂▃▅▇▇▅▂] ❯ - Stop speaking — after 3 seconds of silence, recording auto-stops
- Two beeps play (660Hz) confirming the recording ended
- Audio is transcribed via Whisper and sent to the agent
- If TTS is enabled, the agent’s reply is spoken aloud
- Recording automatically restarts — speak again without pressing any key
Silence Detection
Two-stage algorithm detects when you’ve finished speaking:- Speech confirmation — waits for audio above the RMS threshold (200) for at least 0.3s, tolerating brief dips between syllables
- End detection — once speech is confirmed, triggers after 3.0 seconds of continuous silence
silence_threshold and silence_duration are configurable in config.yaml. You can also disable the record start/stop beeps with voice.beep_enabled: false.
Ending a voice chat by voice
Say “stop” — and nothing else — to end the voice conversation hands-free. The match is deliberately strict: the whole utterance (case-insensitive, surrounding punctuation ignored) must equal a configured phrase, so “stop doing that and try X instead” still reaches the agent normally. Customize the phrase list withvoice.stop_phrases in config.yaml (e.g. ["stop", "goodbye mibyan"]), or set it to [] to disable. Phrases can be in any language (e.g. ["отбой", "стоп"] with stt.language: ru). The desktop app honours the same list; while voice.stop_phrases is left at its default it also accepts a few English extras (“goodbye”, “never mind”, “cancel”, …), and a customised list replaces them. A voice chat also ends on its own after three consecutive silent cycles (no speech detected).
Typing a bare stop phrase while a voice chat is active works the same way on every surface (CLI, TUI, desktop): the message ends the voice chat instead of being sent to the agent. Outside a voice chat, typed “stop” is an ordinary message.
Streaming TTS
When TTS is enabled, the agent speaks its reply sentence-by-sentence as it generates text — you don’t wait for the full response. This works with every TTS provider:- Buffers text deltas into complete sentences (min 20 chars)
- Strips markdown formatting, emoji, and
<think>blocks - Plays audio per sentence in real-time — providers with a chunked PCM API (ElevenLabs, OpenAI) stream raw audio for the lowest time-to-first-word; every other provider (including the default Edge) synthesizes and plays each sentence as it completes
Desktop remote: client-direct voice (lowest-hop path)
When Mibyan Desktop is connected to a remote gateway, audio does not need to be relayed through the gateway at all. At voice-session start the desktop fetches the active profile’s resolved STT/TTS settings (provider, model, language/voice, and credential) from the gateway over the authenticated REST channel (GET /api/audio/voice-config) and then calls the providers directly:
- Dictation / voice input: the mic recording goes straight from your desktop to the profile’s STT provider; only the resulting text is sent to the gateway as the prompt.
- Spoken replies: the reply text is already streaming to the desktop over the chat socket, so the desktop synthesizes it locally with the profile’s TTS provider and plays it — the gateway link never carries audio.
edge TTS, command providers, plugins) automatically fall back to the relay path (/api/audio/transcribe and the speech WebSocket), as does any older backend without the endpoint. To force the relay for every provider, set:
Desktop: GPT-Live voice chat mode (full duplex, delegates to Mibyan)
The chained loop above is one of two voice chat modes in the desktop app. The other replaces the whole STT → turn → TTS chain with one full-duplex voice model, OpenAI’sgpt-live-1: it listens while it speaks, handles interruptions, backchannels and background noise itself, and has no tools of its own. Whenever you ask for real work it delegates to Mibyan, which answers as usual — with whatever model and provider the session has selected, the full toolset, memory and approvals — and the voice paraphrases the answer aloud.
OPENAI_API_KEY, VOICE_TOOLS_OPENAI_KEY, or voice.gpt_live.api_key). The voice layer is billed by OpenAI at $0.05 per minute of session time (idle time counts); the Mibyan turn is billed on its own provider as always. The mode is also in Settings → Voice → Voice Chat Mode.
How it works: pressing the voice button opens a WebRTC session from the desktop to GPT-Live; the desktop only ever receives a session id and an SDP answer — the key stays on the gateway host, which performs the session creation (POST /api/audio/voice-live/session). Each session.delegation.created becomes a normal turn on the open chat (the bubble shows what you said; the recent spoken exchange rides the model input as a per-turn note, never the system prompt, so the reply is speakable prose). Tool activity is fed to the voice as quiet context (“Mibyan is working: terminal”) so it can tell you what is happening if you ask; the final answer is streamed back sentence by sentence. Saying the stop phrase ends the conversation. If gpt-live is selected but no key resolves, the button falls back to the chained mode with a notice.
Not supported in this mode: the Nous-managed audio proxy (direct key only), the CLI/TUI (/voice keeps the chained loop), and the tts tool (it keeps using tts.provider).
Barge-in
You can interrupt the agent at ANY point in its turn — the microphone stays live from the moment you finish speaking until the reply has fully played (full duplex):- Interject while it’s thinking — in continuous voice mode, speaking during LLM generation (before any audio plays) interrupts the in-flight turn and your interjection becomes the next message, the same as typing over a running turn.
- Talk over it — speaking while the agent’s reply plays cuts playback the moment you start talking and submits what you said. The detector calibrates its noise floor against the quiet room at turn start (never against the playback itself), so speaker bleed can’t deafen it and normal speech reliably trips it.
- Type or press the record key — sending a new message or hitting the push-to-talk key stops playback instantly on every surface.
- Say “stop” — the stop phrase works in both phases: mid-generation it interrupts the turn AND ends the voice chat; mid-playback it cuts the speech and ends the chat.
voice.barge_in: false disables it; voice.barge_in_threshold_multiplier (default 3.0) scales the speech trigger over the quiet-room floor — lower is more sensitive; the desktop app also scales its playback-phase trigger by it, so a quiet Bluetooth headset that can’t interrupt a reply can use e.g. 1.5; voice.barge_in_grace_seconds (default 0.5) suppresses trips right after playback starts. Set mibyan_VOICE_DEBUG=1 to stream per-block VAD diagnostics (calibrated floor, RMS, trip decisions) to stderr for live tuning.
The agent knows it was interrupted: the next message carries a short note telling the model its spoken reply was cut off, so it can react naturally (“rude!”) or pick up where it left off instead of being oblivious.
Hallucination Filter
Whisper sometimes generates phantom text from silence or background noise (“Thank you for watching”, “Subscribe”, etc.). The agent filters these out using a set of 26 known hallucination phrases across multiple languages, plus a regex pattern that catches repetitive variations.Gateway Voice Reply (Telegram & Discord)
If you haven’t set up your messaging bots yet, see the platform-specific guides: Start the gateway to connect to your messaging platforms:Discord: Channels vs DMs
The bot supports two interaction modes on Discord:
DM (recommended for personal use): Just open a DM with the bot and type — no @mention needed. Voice replies and all commands work the same as in channels.
Server channels: The bot only responds when you @mention it (e.g.
@mibyanbyt4 hello). Make sure you select the bot user from the mention popup, not the role with the same name.
Commands
These work in both Telegram and Discord (DMs and text channels):Modes
Voice mode setting is persisted across gateway restarts.
Platform Delivery
Discord Voice Channels
The most immersive voice feature: the bot joins a Discord voice channel, listens to users speaking, transcribes their speech, processes through the agent, and speaks the reply back in the voice channel.Setup
1. Discord Bot Permissions
If you already have a Discord bot set up for text (see Discord Setup Guide), you need to add voice permissions. Go to the Discord Developer Portal → your application → Installation → Default Install Settings → Guild Install: Add these permissions to the existing text permissions:
Updated Permissions Integer:
Re-invite the bot with the updated permissions URL:
YOUR_APP_ID with your Application ID from the Developer Portal.
2. Privileged Gateway Intents
In the Developer Portal → your application → Bot → Privileged Gateway Intents, enable all three:
Message Content Intent is required. Server Members Intent is only needed if your
DISCORD_ALLOWED_USERS list uses usernames — if you use numeric user IDs, you can leave it OFF. Voice-channel SSRC → user_id mapping comes from Discord’s SPEAKING opcode on the voice websocket and does not require the Server Members Intent.
3. Opus Codec
The Opus codec library must be installed on the machine running the gateway:- macOS:
/opt/homebrew/lib/libopus.dylib - Linux:
libopus.so.0
4. Environment Variables
Start the Gateway
Commands
Use these in the Discord text channel where the bot is present:You must be in a voice channel before running
/voice join. The bot joins the same VC you’re in.How It Works
When the bot joins a voice channel, it:- Listens to each user’s audio stream independently
- Detects silence — 1.5s of silence after at least 0.5s of speech triggers processing
- Transcribes the audio via Whisper STT (local, Groq, or OpenAI)
- Processes through the full agent pipeline (session, tools, memory)
- Speaks the reply back in the voice channel via TTS
Text Channel Integration
When the bot is in a voice channel:- Transcripts appear in the text channel:
[Voice] @user: what you said - Agent responses are sent as text in the channel AND spoken in the VC
- The text channel is the one where
/voice joinwas issued
Echo Prevention
The bot automatically pauses its audio listener while playing TTS replies, preventing it from hearing and re-processing its own output.Access Control
Only users listed inDISCORD_ALLOWED_USERS can interact via voice. Other users’ audio is silently ignored.
Configuration Reference
config.yaml
Environment Variables
STT Provider Comparison
Provider priority (automatic fallback): local > groq > openai
TTS Provider Comparison
NeuTTS uses the
tts.neutts config block above.
For openai, the text_to_speech tool accepts an optional instructions
argument that unlocks gpt-4o-mini-tts’s voice-design capability (tone,
emotion, pacing, accent, whispering). The same field also routes to
OpenAI-compatible voice-design servers mounted via tts.openai.base_url
(e.g. Qwen3-TTS-VoiceDesign via oMLX).
Troubleshooting
”No audio device found” (CLI)
PortAudio is not installed:Bot doesn’t respond in Discord server channels
The bot requires an @mention by default in server channels. Make sure you:- Type
@and select the bot user (with the #discriminator), not the role with the same name - Or use DMs instead — no mention needed
- Or set
DISCORD_REQUIRE_MENTION=falsein~/.mibyan/.env
Bot joins VC but doesn’t hear me
- Check your Discord user ID is in
DISCORD_ALLOWED_USERS - Make sure you’re not muted in Discord
- The bot needs a SPEAKING event from Discord before it can map your audio — start speaking within a few seconds of joining
Bot hears me but doesn’t respond
- Verify STT is available: install
faster-whisper(no key needed) or setGROQ_API_KEY/VOICE_TOOLS_OPENAI_KEY - Check the LLM model is configured and accessible
- Review gateway logs:
tail -f ~/.mibyan/logs/gateway.log
Bot responds in text but not in voice channel
- TTS provider may be failing — check API key and quota
- Edge TTS (free, no key) is the default fallback
- Check logs for TTS errors
Whisper returns garbage text
The hallucination filter catches most cases automatically. If you’re still getting phantom transcripts:- Use a quieter environment
- Adjust
silence_thresholdin config (higher = less sensitive) - Try a different STT model

