Commands, package names, and image names on this page come from the open-source project that Mibyan Desktop is built on, and can differ from the Mibyan Desktop installer. For the supported Mibyan install and update path, see Install and update.
The Problem
Cloud LLM APIs charge per token. A heavy coding session can cost $5–20. For personal projects, learning, or privacy-sensitive work, that adds up — and you’re sending every conversation to a third party.What This Guide Solves
You’ll set up Mibyan running entirely on your own hardware, using Ollama as the model backend. No API keys, no subscriptions, no data leaving your machine. Once configured, Mibyan works exactly like it does with OpenRouter or Anthropic — terminal commands, file editing, web browsing, delegation — but the model runs locally. By the end, you’ll have:- Ollama serving one or more open-weight models
- Mibyan connected to Ollama as a custom endpoint
- A working local agent that can edit files, run commands, and browse the web
- Optional: a Telegram/Discord bot powered entirely by your own hardware
What You Need
Step 1: Install Ollama
Step 2: Pull a Model
Choose based on your hardware:
Pull your chosen model:
Multiple modelsYou can pull several models and switch between them inside Mibyan with
/model. Ollama loads the active model into memory on demand and unloads idle ones automatically.Step 3: Configure Mibyan
Run the Mibyan setup wizard:- Base URL:
http://localhost:11434/v1 - API Key: Leave empty or type
no-key(Ollama doesn’t need one) - Model:
gemma4:31b(or whichever model you pulled)
~/.mibyan/config.yaml directly:
Step 4: Start Using Mibyan
Step 5: Pick the Right Model for Your Task
Not every task needs the biggest model. Here’s a practical guide:For full agentic work (editing files, running commands, browsing),
gemma4:31b is currently the best local option with tool-call support. Check Ollama’s model library for newer models — tool-calling support is expanding rapidly.Step 6: Optimize for Speed
Increase Ollama’s Context Window
By default, Ollama uses a 2048-token context. Mibyan requires at least 64,000 tokens for agentic work with tools:gemma4-64k as the model name.
Keep the Model Loaded
By default, Ollama unloads models after 5 minutes of inactivity. For a persistent gateway bot, keep it loaded:Use GPU Offloading (If Available)
If you have an NVIDIA GPU, Ollama automatically offloads layers to it. Check with:Step 7: Run as a Gateway Bot (Optional)
Once Mibyan works locally in the CLI, you can expose it as a Telegram or Discord bot — still running entirely on your hardware.Telegram
- Create a bot via @BotFather and get the token
- Add to your
~/.mibyan/config.yaml:
- Start the gateway:
Discord
- Create a Discord application at discord.com/developers
- Add to config:
- Start:
mibyan gateway
Step 8: Set Up Fallbacks (Optional)
Local models can struggle with complex tasks. Set up a cloud fallback that only activates when the local model fails:Troubleshooting
”provider ‘ollama’ has no endpoint configured”
mibyan chat --provider ollama (or vllm) stops with this error when no endpoint is configured for that alias anywhere — no providers.ollama.base_url, no model.base_url. Mibyan refuses to send the request rather than fall back to OpenRouter with a cloud key (OPENROUTER_API_KEY / OPENAI_API_KEY) that happens to be set. Add the endpoint:
“Connection refused” on startup
Ollama isn’t running. Start it:Slow responses
- Check model size vs RAM: If your model needs more RAM than available, it swaps to disk. Use a smaller model or add RAM.
- Check
ollama ps: If no GPU layers are offloaded, responses are CPU-bound. This is normal for CPU-only servers. - Reduce context: Large conversations slow down inference. Use
/compressregularly, or set a lower compression threshold in config.
Slow first response (prefill)
Mibyan sends a fixed payload on every API call — the system prompt plus the tool schemas for all enabled tools — before any of your conversation content. On CPU-only or low-VRAM setups, processing that prompt (the prefill phase) dominates the first turn: the model can sit silent for minutes while it works through the prompt, then generate at its normal pace. This is expected behaviour, not a hang. The Mac local-LLM guide documents the same effect — during prefill on large contexts, local models may produce no output for minutes while processing the prompt — and Mibyan automatically raises its stream read timeout from 120s to 1800s for local endpoints (mibyan_STREAM_READ_TIMEOUT).
What helps:
- Keep the model loaded — Ollama unloads idle models after 5 minutes, adding a full reload before the next prefill. Set
OLLAMA_KEEP_ALIVE=24h(see Step 6). - Widen the API timeout — set
mibyan_API_TIMEOUT=1800in~/.mibyan/.env(see What You Need). - Measure and trim the fixed prompt — run
mibyan prompt-sizefor a byte breakdown of the system prompt and tool schemas, then disable unused toolsets withmibyan toolsand uninstall skills you don’t need withmibyan skills. - Use GPU offloading — even a partial offload gives a significant speedup (see Step 6).
Model doesn’t follow tool calls
Models without tool-call support produce plain text instead of structured function calls. Solutions:- Use a model with tool-call support — of the models listed above, only
gemma4:31bhas reliable tool calling. - Mibyan has auto-repair — it detects malformed tool calls and attempts to fix them automatically.
- Set up a fallback — if the local model fails 3 times, Mibyan falls back to a cloud provider.
{"name": "web_search", ...} in its reply instead of actually running the tool, that’s usually the server, not the model — tool calling isn’t enabled or the tool-call format isn’t parsed. See the per-server fix table in Tool calls appear as text instead of executing (llama.cpp needs --jinja, vLLM needs --enable-auto-tool-choice --tool-call-parser mibyan, and so on).
Context window errors
The default Ollama context (2048 tokens) is too small for agentic work. See Step 6 to increase it.Cost Comparison
Here’s what running locally saves compared to cloud APIs, based on a typical coding session (~100K tokens input, ~20K tokens output):
Your only cost is electricity — roughly $0.01–0.05 per session depending on hardware.
What Works Well Locally
- File editing and code generation — models 9B+ handle this well
- Terminal commands — Mibyan wraps the command, runs it, reads output regardless of model
- Web browsing — the browser tool does the fetching; the model just interprets results
- Cron jobs and scheduled tasks — work identically to cloud setups
- Multi-platform gateway — Telegram, Discord, Slack all work with local models
What’s Better with Cloud Models
- Very complex multi-step reasoning — 70B+ or cloud models like Claude Opus are noticeably better
- Long context windows — cloud models offer 100K–1M tokens; local runtimes often default below Mibyan’ 64K minimum unless you configure them
- Speed on large responses — cloud inference is faster than CPU-only local for long generations

