Skip to main content
Mibyan supports full voice interaction across CLI and messaging platforms. Talk to the agent using your microphone, hear spoken replies, and have live voice conversations in Discord voice channels. If you want a practical setup walkthrough with recommended configurations and real usage patterns, see Use Voice Mode with Mibyan. For hands-free session start — saying “hey mibyan” (or any phrase) to open a fresh voice session on the CLI, TUI, or desktop app — see Wake Word.

Prerequisites

Before using voice features, make sure you have:
  1. Mibyan installed — via the install script (see Installation)
  2. An LLM provider configured — run mibyan model or set your preferred provider credentials in ~/.mibyan/.env
  3. A working base setup — run mibyan to verify the agent responds to text before enabling voice
The ~/.mibyan/ directory and default config.yaml are created automatically the first time you run mibyan. You only need to create ~/.mibyan/.env manually for API keys.
Nous Portal covers bothA paid Nous Portal subscription supplies the LLM (step 2) and OpenAI TTS via the Tool Gateway — no separate OpenAI key needed. On a fresh install, mibyan setup --portal wires both up at once.

Overview

Requirements

Python Packages

Use mibyan tools to configure voice providers. Missing built-in feature requirements go through PM, subject to security.allow_lazy_installs and the target’s dependency support. Restart Mibyan if the selected dependency environment changes. A bundled app includes its supported engine dependencies. Docker includes a curated subset and disables on-demand installs. Do not use pip to modify a signed payload or the system Python. For a manual development environment, select the required extras in the development setup. Local Faster-Whisper is excluded on native Windows ARM64 and Intel macOS. Use a cloud or command-based STT provider on those targets. audio-io contains the microphone/playback dependencies without local STT. The all extra does not mean every voice or wake engine. NeuTTS is a separate optional runtime and downloads models on first use. Do not install its dependencies into a signed app or system Python.
discord.py[voice] installs PyNaCl (for voice encryption) and opus bindings automatically. This is required for Discord voice channel support.

System Dependencies

API Keys

Add to ~/.mibyan/.env:
If faster-whisper is installed, voice mode works with zero API keys for STT. The model (~150 MB for base) downloads automatically on first use.
The first download normally comes from huggingface.co. If that host is unavailable on your network, export an accessible mirror in the shell or service that starts Mibyan:
Disabling Xet avoids authentication failures from Xet’s separate CAS hosts when a mirror is in use. After the model is cached, Mibyan loads that snapshot without an online revision check.

CLI Voice Mode

Voice mode is available in both the classic CLI (mibyan chat) and the TUI (mibyan --tui). Behavior is identical across both — same slash commands, same VAD silence detection, same streaming TTS, same hallucination filter. The TUI additionally forwards crash-forensic logs to ~/.mibyan/logs/ so push-to-talk failures on exotic audio backends can be reported with a full stack trace rather than disappearing silently.

Quick Start

Start the CLI and enable voice mode:
Then use these commands inside the CLI:

How It Works

  1. Start the CLI with mibyan and enable voice mode with /voice on
  2. Press Ctrl+B — a beep plays (880Hz), recording starts
  3. Speak — a live audio level bar shows your input: ● [▁▂▃▅▇▇▅▂] ❯
  4. Stop speaking — after 3 seconds of silence, recording auto-stops
  5. Two beeps play (660Hz) confirming the recording ended
  6. Audio is transcribed via Whisper and sent to the agent
  7. If TTS is enabled, the agent’s reply is spoken aloud
  8. Recording automatically restarts — speak again without pressing any key
This loop continues until you press Ctrl+B during recording (exits continuous mode) or 3 consecutive recordings detect no speech.
The record key is configurable via voice.record_key in ~/.mibyan/config.yaml (default: ctrl+b).

Silence Detection

Two-stage algorithm detects when you’ve finished speaking:
  1. Speech confirmation — waits for audio above the RMS threshold (200) for at least 0.3s, tolerating brief dips between syllables
  2. End detection — once speech is confirmed, triggers after 3.0 seconds of continuous silence
If no speech is detected at all for 15 seconds, recording stops automatically. Both silence_threshold and silence_duration are configurable in config.yaml. You can also disable the record start/stop beeps with voice.beep_enabled: false.

Ending a voice chat by voice

Say “stop” — and nothing else — to end the voice conversation hands-free. The match is deliberately strict: the whole utterance (case-insensitive, surrounding punctuation ignored) must equal a configured phrase, so “stop doing that and try X instead” still reaches the agent normally. Customize the phrase list with voice.stop_phrases in config.yaml (e.g. ["stop", "goodbye mibyan"]), or set it to [] to disable. Phrases can be in any language (e.g. ["отбой", "стоп"] with stt.language: ru). The desktop app honours the same list; while voice.stop_phrases is left at its default it also accepts a few English extras (“goodbye”, “never mind”, “cancel”, …), and a customised list replaces them. A voice chat also ends on its own after three consecutive silent cycles (no speech detected). Typing a bare stop phrase while a voice chat is active works the same way on every surface (CLI, TUI, desktop): the message ends the voice chat instead of being sent to the agent. Outside a voice chat, typed “stop” is an ordinary message.

Streaming TTS

When TTS is enabled, the agent speaks its reply sentence-by-sentence as it generates text — you don’t wait for the full response. This works with every TTS provider:
  1. Buffers text deltas into complete sentences (min 20 chars)
  2. Strips markdown formatting, emoji, and <think> blocks
  3. Plays audio per sentence in real-time — providers with a chunked PCM API (ElevenLabs, OpenAI) stream raw audio for the lowest time-to-first-word; every other provider (including the default Edge) synthesizes and plays each sentence as it completes
The same pipeline runs in the classic CLI, the TUI, and the desktop app. In a desktop voice conversation the reply text is fed live into a per-reply speech WebSocket as the model generates it, so speech overlaps generation — one socket and one audio clock per reply, no per-sentence connection gaps.

Desktop remote: client-direct voice (lowest-hop path)

When Mibyan Desktop is connected to a remote gateway, audio does not need to be relayed through the gateway at all. At voice-session start the desktop fetches the active profile’s resolved STT/TTS settings (provider, model, language/voice, and credential) from the gateway over the authenticated REST channel (GET /api/audio/voice-config) and then calls the providers directly:
  • Dictation / voice input: the mic recording goes straight from your desktop to the profile’s STT provider; only the resulting text is sent to the gateway as the prompt.
  • Spoken replies: the reply text is already streaming to the desktop over the chat socket, so the desktop synthesizes it locally with the profile’s TTS provider and plays it — the gateway link never carries audio.
There is nothing to configure on the client: the profile you’re talking to is the single source of truth for providers and keys, exactly as if the gateway had done the work itself. Keys are held in the desktop’s memory for the session only — never written to disk on the client. Providers that can only run on the gateway host (local whisper, edge TTS, command providers, plugins) automatically fall back to the relay path (/api/audio/transcribe and the speech WebSocket), as does any older backend without the endpoint. To force the relay for every provider, set:
Client-direct wire support: OpenAI (incl. Nous-managed audio), Groq, Mistral, and DeepInfra via the OpenAI-compatible shapes, xAI Grok STT, and ElevenLabs STT + TTS. xAI configured through OAuth stays on the relay (the OAuth bearer refreshes server-side).

Desktop: GPT-Live voice chat mode (full duplex, delegates to Mibyan)

The chained loop above is one of two voice chat modes in the desktop app. The other replaces the whole STT → turn → TTS chain with one full-duplex voice model, OpenAI’s gpt-live-1: it listens while it speaks, handles interruptions, backchannels and background noise itself, and has no tools of its own. Whenever you ask for real work it delegates to Mibyan, which answers as usual — with whatever model and provider the session has selected, the full toolset, memory and approvals — and the voice paraphrases the answer aloud.
Requirements: an OpenAI API key (OPENAI_API_KEY, VOICE_TOOLS_OPENAI_KEY, or voice.gpt_live.api_key). The voice layer is billed by OpenAI at $0.05 per minute of session time (idle time counts); the Mibyan turn is billed on its own provider as always. The mode is also in Settings → Voice → Voice Chat Mode. How it works: pressing the voice button opens a WebRTC session from the desktop to GPT-Live; the desktop only ever receives a session id and an SDP answer — the key stays on the gateway host, which performs the session creation (POST /api/audio/voice-live/session). Each session.delegation.created becomes a normal turn on the open chat (the bubble shows what you said; the recent spoken exchange rides the model input as a per-turn note, never the system prompt, so the reply is speakable prose). Tool activity is fed to the voice as quiet context (“Mibyan is working: terminal”) so it can tell you what is happening if you ask; the final answer is streamed back sentence by sentence. Saying the stop phrase ends the conversation. If gpt-live is selected but no key resolves, the button falls back to the chained mode with a notice. Not supported in this mode: the Nous-managed audio proxy (direct key only), the CLI/TUI (/voice keeps the chained loop), and the tts tool (it keeps using tts.provider).

Barge-in

You can interrupt the agent at ANY point in its turn — the microphone stays live from the moment you finish speaking until the reply has fully played (full duplex):
  • Interject while it’s thinking — in continuous voice mode, speaking during LLM generation (before any audio plays) interrupts the in-flight turn and your interjection becomes the next message, the same as typing over a running turn.
  • Talk over it — speaking while the agent’s reply plays cuts playback the moment you start talking and submits what you said. The detector calibrates its noise floor against the quiet room at turn start (never against the playback itself), so speaker bleed can’t deafen it and normal speech reliably trips it.
  • Type or press the record key — sending a new message or hitting the push-to-talk key stops playback instantly on every surface.
  • Say “stop” — the stop phrase works in both phases: mid-generation it interrupts the turn AND ends the voice chat; mid-playback it cuts the speech and ends the chat.
Tuning (config.yaml): voice.barge_in: false disables it; voice.barge_in_threshold_multiplier (default 3.0) scales the speech trigger over the quiet-room floor — lower is more sensitive; the desktop app also scales its playback-phase trigger by it, so a quiet Bluetooth headset that can’t interrupt a reply can use e.g. 1.5; voice.barge_in_grace_seconds (default 0.5) suppresses trips right after playback starts. Set mibyan_VOICE_DEBUG=1 to stream per-block VAD diagnostics (calibrated floor, RMS, trip decisions) to stderr for live tuning. The agent knows it was interrupted: the next message carries a short note telling the model its spoken reply was cut off, so it can react naturally (“rude!”) or pick up where it left off instead of being oblivious.

Hallucination Filter

Whisper sometimes generates phantom text from silence or background noise (“Thank you for watching”, “Subscribe”, etc.). The agent filters these out using a set of 26 known hallucination phrases across multiple languages, plus a regex pattern that catches repetitive variations.

Gateway Voice Reply (Telegram & Discord)

If you haven’t set up your messaging bots yet, see the platform-specific guides: Start the gateway to connect to your messaging platforms:

Discord: Channels vs DMs

The bot supports two interaction modes on Discord: DM (recommended for personal use): Just open a DM with the bot and type — no @mention needed. Voice replies and all commands work the same as in channels. Server channels: The bot only responds when you @mention it (e.g. @mibyanbyt4 hello). Make sure you select the bot user from the mention popup, not the role with the same name.
To disable the mention requirement in server channels, add to ~/.mibyan/.env:
Or set specific channels as free-response (no mention needed):

Commands

These work in both Telegram and Discord (DMs and text channels):

Modes

Voice mode setting is persisted across gateway restarts.

Platform Delivery


Discord Voice Channels

The most immersive voice feature: the bot joins a Discord voice channel, listens to users speaking, transcribes their speech, processes through the agent, and speaks the reply back in the voice channel.

Setup

1. Discord Bot Permissions

If you already have a Discord bot set up for text (see Discord Setup Guide), you need to add voice permissions. Go to the Discord Developer Portal → your application → Installation → Default Install Settings → Guild Install: Add these permissions to the existing text permissions: Updated Permissions Integer: Re-invite the bot with the updated permissions URL:
Replace YOUR_APP_ID with your Application ID from the Developer Portal.
Re-inviting the bot to a server it’s already in will update its permissions without removing it. You won’t lose any data or configuration.

2. Privileged Gateway Intents

In the Developer Portal → your application → Bot → Privileged Gateway Intents, enable all three: Message Content Intent is required. Server Members Intent is only needed if your DISCORD_ALLOWED_USERS list uses usernames — if you use numeric user IDs, you can leave it OFF. Voice-channel SSRC → user_id mapping comes from Discord’s SPEAKING opcode on the voice websocket and does not require the Server Members Intent.

3. Opus Codec

The Opus codec library must be installed on the machine running the gateway:
The bot auto-loads the codec from:
  • macOS: /opt/homebrew/lib/libopus.dylib
  • Linux: libopus.so.0

4. Environment Variables

Start the Gateway

The bot should come online in Discord within a few seconds.

Commands

Use these in the Discord text channel where the bot is present:
You must be in a voice channel before running /voice join. The bot joins the same VC you’re in.

How It Works

When the bot joins a voice channel, it:
  1. Listens to each user’s audio stream independently
  2. Detects silence — 1.5s of silence after at least 0.5s of speech triggers processing
  3. Transcribes the audio via Whisper STT (local, Groq, or OpenAI)
  4. Processes through the full agent pipeline (session, tools, memory)
  5. Speaks the reply back in the voice channel via TTS

Text Channel Integration

When the bot is in a voice channel:
  • Transcripts appear in the text channel: [Voice] @user: what you said
  • Agent responses are sent as text in the channel AND spoken in the VC
  • The text channel is the one where /voice join was issued

Echo Prevention

The bot automatically pauses its audio listener while playing TTS replies, preventing it from hearing and re-processing its own output.

Access Control

Only users listed in DISCORD_ALLOWED_USERS can interact via voice. Other users’ audio is silently ignored.

Configuration Reference

config.yaml

Environment Variables

STT Provider Comparison

Provider priority (automatic fallback): local > groq > openai

TTS Provider Comparison

NeuTTS uses the tts.neutts config block above. For openai, the text_to_speech tool accepts an optional instructions argument that unlocks gpt-4o-mini-tts’s voice-design capability (tone, emotion, pacing, accent, whispering). The same field also routes to OpenAI-compatible voice-design servers mounted via tts.openai.base_url (e.g. Qwen3-TTS-VoiceDesign via oMLX).

Troubleshooting

”No audio device found” (CLI)

PortAudio is not installed:
If you are running Mibyan inside Docker on a Linux desktop, the container also needs access to your host audio socket. See the Docker audio bridge notes for a PulseAudio/PipeWire-compatible setup.

Bot doesn’t respond in Discord server channels

The bot requires an @mention by default in server channels. Make sure you:
  1. Type @ and select the bot user (with the #discriminator), not the role with the same name
  2. Or use DMs instead — no mention needed
  3. Or set DISCORD_REQUIRE_MENTION=false in ~/.mibyan/.env

Bot joins VC but doesn’t hear me

  • Check your Discord user ID is in DISCORD_ALLOWED_USERS
  • Make sure you’re not muted in Discord
  • The bot needs a SPEAKING event from Discord before it can map your audio — start speaking within a few seconds of joining

Bot hears me but doesn’t respond

  • Verify STT is available: install faster-whisper (no key needed) or set GROQ_API_KEY / VOICE_TOOLS_OPENAI_KEY
  • Check the LLM model is configured and accessible
  • Review gateway logs: tail -f ~/.mibyan/logs/gateway.log

Bot responds in text but not in voice channel

  • TTS provider may be failing — check API key and quota
  • Edge TTS (free, no key) is the default fallback
  • Check logs for TTS errors

Whisper returns garbage text

The hallucination filter catches most cases automatically. If you’re still getting phantom transcripts:
  • Use a quieter environment
  • Adjust silence_threshold in config (higher = less sensitive)
  • Try a different STT model