What voice mode is good for
Voice mode is especially useful when:- you want a hands-free CLI workflow
- you want spoken responses in Telegram or Discord
- you want Mibyan sitting in a Discord voice channel for live conversation
- you want quick idea capture, debugging, or back-and-forth while walking around instead of typing
Choose your voice mode setup
There are really three different voice experiences in Mibyan.
A good path is:
- get text working first
- enable voice replies second
- move to Discord voice channels last if you want the full experience
Step 1: make sure normal Mibyan works first
Before touching voice mode, verify that:- Mibyan starts
- your provider is configured
- the agent can answer text prompts normally
Step 2: install the right extras
CLI microphone + playback
Messaging platforms
Premium ElevenLabs TTS
Local NeuTTS (optional)
The declaredneutts dependency requires Python below 3.14. The Mibyan runtime
requires Python 3.14, so requesting this extra does not install NeuTTS there.
Choose a compatible provider. A separately managed NeuTTS command provider
needs its own supported Python environment.
Combined voice and messaging setup
all extra is not every optional feature. It does not include these voice
and messaging extras.
Step 3: install system dependencies
macOS
Ubuntu / Debian
portaudio→ microphone input / playback for CLI voice modeffmpeg→ audio conversion for TTS and messaging deliveryopus→ Discord voice codec supportespeak-ng→ phonemizer backend for NeuTTS
Step 4: choose STT and TTS providers
Mibyan supports both local and cloud speech stacks.Easiest / cheapest setup
Use local STT and free Edge TTS:- STT provider:
local - TTS provider:
edge
Environment file example
Add to~/.mibyan/.env:
Provider recommendations
Speech-to-text
local→ best default for privacy and zero-cost usegroq→ very fast cloud transcriptionopenai→ good paid fallback
Text-to-speech
edge→ free and good enough for most usersneutts→ free local/on-device TTSelevenlabs→ best qualityopenai→ good middle groundmistral→ multilingual, native Opus
If you use mibyan setup
Setup requests declared Python extras through PM. It cannot override their
Python-version or platform markers:
The declared neutts dependency requires Python below 3.14. The Mibyan runtime
requires Python 3.14, so requesting this extra does not install NeuTTS there.
Choose a compatible provider. A separately managed NeuTTS command provider
needs its own supported Python environment.
Select another provider if a dependency cannot run on your platform.
Step 5: recommended config
voice.submit_mode controls what happens after transcription:
direct(default) submits the transcript immediately.draftputs the transcript in the composer so you can edit or cancel it before pressing Enter.
tts block to:
Use case 1: CLI voice mode
Turn it on
Start Mibyan:Recording flow
Default key:Ctrl+B
- press
Ctrl+B - speak
- wait for silence detection to stop recording automatically
- Mibyan transcribes and responds
- if TTS is on, it speaks the answer
- the loop can automatically restart for continuous use
Useful commands
Good CLI workflows
Walk-up debugging
Say:- “Read the last error again”
- “Explain the root cause in simpler terms”
- “Now give me the exact fix”
Research / brainstorming
Great for:- walking around while thinking
- dictating half-formed ideas
- asking Mibyan to structure your thoughts in real time
Accessibility / low-typing sessions
If typing is inconvenient, voice mode is one of the fastest ways to stay in the full Mibyan loop.Tuning CLI behavior
Silence threshold
If Mibyan starts/stops too aggressively, tune:Silence duration
If you pause a lot between sentences, increase:Record key
IfCtrl+B conflicts with your terminal or tmux habits:
Use case 2: voice replies in Telegram or Discord
This mode is simpler than full voice channels. Mibyan stays a normal chat bot, but can speak replies.Start the gateway
Turn on voice replies
Inside Telegram or Discord:Modes
When to use which mode
/voice onif you want spoken replies only for voice-originating messages/voice ttsif you want a full spoken assistant all the time
Good messaging workflows
Telegram assistant on your phone
Use when:- you are away from your machine
- you want to send voice notes and get quick spoken replies
- you want Mibyan to function like a portable research or ops assistant
Discord DMs with spoken output
Useful when you want private interaction without server-channel mention behavior.Use case 3: Discord voice channels
This is the most advanced mode. Mibyan joins a Discord VC, listens to user speech, transcribes it, runs the normal agent pipeline, and speaks replies back into the channel.Required Discord permissions
In addition to the normal text-bot setup, make sure the bot has:- Connect
- Speak
- preferably Use Voice Activity
- Presence Intent
- Server Members Intent
- Message Content Intent
Join and leave
In a Discord text channel where the bot is present:What happens when joined
- users speak in the VC
- Mibyan detects speech boundaries
- transcripts are posted in the associated text channel
- Mibyan responds in text and audio
- the text channel is the one where
/voice joinwas issued
Best practices for Discord VC use
- keep
DISCORD_ALLOWED_USERStight - use a dedicated bot/testing channel at first
- verify STT and TTS work in ordinary text-chat voice mode before trying VC mode
Voice quality recommendations
Best quality setup
- STT: local
large-v3or Groqwhisper-large-v3 - TTS: ElevenLabs
Best speed / convenience setup
- STT: local
baseor Groq - TTS: Edge
Best zero-cost setup
- STT: local
- TTS: Edge
Common failure modes
”No audio device found”
Installportaudio.
”Bot joins but hears nothing”
Check:- your Discord user ID is in
DISCORD_ALLOWED_USERS - you are not muted
- privileged intents are enabled
- the bot has Connect/Speak permissions
”It transcribes but does not speak”
Check:- TTS provider config
- API key / quota for ElevenLabs or OpenAI
ffmpeginstall for Edge conversion paths
”Whisper outputs garbage”
Try:- quieter environment
- higher
silence_threshold - different STT provider/model
- shorter, clearer utterances
”It works in DMs but not in server channels”
That is often mention policy. By default, the bot needs an@mention in Discord server text channels unless configured otherwise.
Suggested first-week setup
If you want the shortest path to success:- get text Mibyan working
- run
mibyan setup ttsto enable voice support - use CLI voice mode with local STT + Edge TTS
- then enable
/voice onin Telegram or Discord - only after that, try Discord VC mode

