Skip to main content
Commands, package names, and image names on this page come from the open-source project that Mibyan Desktop is built on, and can differ from the Mibyan Desktop installer. For the supported Mibyan install and update path, see Install and update.
Desktop users: there’s a one-click pathOn the Mibyan desktop app, Settings → Providers → Local Models installs and manages a local llama.cpp server for you — model downloads, memory fitting, and context sizing included. See Local Models. This guide is for manual setup: Ollama specifically, CLI-first workflows, or servers you want to run yourself.

The Problem

Cloud LLM APIs charge per token. A heavy coding session can cost $5–20. For personal projects, learning, or privacy-sensitive work, that adds up — and you’re sending every conversation to a third party.

What This Guide Solves

You’ll set up Mibyan running entirely on your own hardware, using Ollama as the model backend. No API keys, no subscriptions, no data leaving your machine. Once configured, Mibyan works exactly like it does with OpenRouter or Anthropic — terminal commands, file editing, web browsing, delegation — but the model runs locally. By the end, you’ll have:
  • Ollama serving one or more open-weight models
  • Mibyan connected to Ollama as a custom endpoint
  • A working local agent that can edit files, run commands, and browse the web
  • Optional: a Telegram/Discord bot powered entirely by your own hardware

What You Need

CPU-only works, but expect slower responsesOllama runs on CPU-only servers. A 9B model on a modern 8-core CPU gives ~10 tokens/sec. A 31B model on CPU is slower (~2–5 tokens/sec) — each response takes 30–120 seconds, but it works. A GPU dramatically improves this. For CPU-only setups, widen the API timeout via the env var (it’s not a config.yaml key):

Step 1: Install Ollama

Verify it’s running:

Step 2: Pull a Model

Choose based on your hardware:
Tool calling mattersMibyan is an agentic assistant — it edits files, runs commands, and browses the web through tool calls. Models without tool-call support can only chat; they can’t take actions. For the full Mibyan experience, use a model that supports tools (like gemma4:31b).
Pull your chosen model:
Multiple modelsYou can pull several models and switch between them inside Mibyan with /model. Ollama loads the active model into memory on demand and unloads idle ones automatically.
Verify the model works:
You should see a JSON response with the model’s reply.

Step 3: Configure Mibyan

Run the Mibyan setup wizard:
When prompted for a provider, select Custom Endpoint and enter:
  • Base URL: http://localhost:11434/v1
  • API Key: Leave empty or type no-key (Ollama doesn’t need one)
  • Model: gemma4:31b (or whichever model you pulled)
Alternatively, edit ~/.mibyan/config.yaml directly:

Step 4: Start Using Mibyan

That’s it. You’re now running a fully local agent. Try it out:
Mibyan will use the terminal tool, file operations, and your local model — no cloud calls.

Step 5: Pick the Right Model for Your Task

Not every task needs the biggest model. Here’s a practical guide:
For full agentic work (editing files, running commands, browsing), gemma4:31b is currently the best local option with tool-call support. Check Ollama’s model library for newer models — tool-calling support is expanding rapidly.
Switch models on the fly inside a session:

Step 6: Optimize for Speed

Increase Ollama’s Context Window

By default, Ollama uses a 2048-token context. Mibyan requires at least 64,000 tokens for agentic work with tools:
Then update your Mibyan config to use gemma4-64k as the model name.

Keep the Model Loaded

By default, Ollama unloads models after 5 minutes of inactivity. For a persistent gateway bot, keep it loaded:
Or set it globally in Ollama’s environment:

Use GPU Offloading (If Available)

If you have an NVIDIA GPU, Ollama automatically offloads layers to it. Check with:
For a 31B model on a 12 GB GPU, you’ll get partial offload (~40 layers on GPU, rest on CPU), which still gives a significant speedup.

Step 7: Run as a Gateway Bot (Optional)

Once Mibyan works locally in the CLI, you can expose it as a Telegram or Discord bot — still running entirely on your hardware.

Telegram

  1. Create a bot via @BotFather and get the token
  2. Add to your ~/.mibyan/config.yaml:
  1. Start the gateway:
Now message your bot on Telegram — it responds using your local model.

Discord

  1. Create a Discord application at discord.com/developers
  2. Add to config:
  1. Start: mibyan gateway

Step 8: Set Up Fallbacks (Optional)

Local models can struggle with complex tasks. Set up a cloud fallback that only activates when the local model fails:
This way, 90% of your usage is free (local), and only the hard tasks hit the paid API.

Troubleshooting

”provider ‘ollama’ has no endpoint configured”

mibyan chat --provider ollama (or vllm) stops with this error when no endpoint is configured for that alias anywhere — no providers.ollama.base_url, no model.base_url. Mibyan refuses to send the request rather than fall back to OpenRouter with a cloud key (OPENROUTER_API_KEY / OPENAI_API_KEY) that happens to be set. Add the endpoint:

“Connection refused” on startup

Ollama isn’t running. Start it:

Slow responses

  • Check model size vs RAM: If your model needs more RAM than available, it swaps to disk. Use a smaller model or add RAM.
  • Check ollama ps: If no GPU layers are offloaded, responses are CPU-bound. This is normal for CPU-only servers.
  • Reduce context: Large conversations slow down inference. Use /compress regularly, or set a lower compression threshold in config.

Slow first response (prefill)

Mibyan sends a fixed payload on every API call — the system prompt plus the tool schemas for all enabled tools — before any of your conversation content. On CPU-only or low-VRAM setups, processing that prompt (the prefill phase) dominates the first turn: the model can sit silent for minutes while it works through the prompt, then generate at its normal pace. This is expected behaviour, not a hang. The Mac local-LLM guide documents the same effect — during prefill on large contexts, local models may produce no output for minutes while processing the prompt — and Mibyan automatically raises its stream read timeout from 120s to 1800s for local endpoints (mibyan_STREAM_READ_TIMEOUT). What helps:
  • Keep the model loaded — Ollama unloads idle models after 5 minutes, adding a full reload before the next prefill. Set OLLAMA_KEEP_ALIVE=24h (see Step 6).
  • Widen the API timeout — set mibyan_API_TIMEOUT=1800 in ~/.mibyan/.env (see What You Need).
  • Measure and trim the fixed prompt — run mibyan prompt-size for a byte breakdown of the system prompt and tool schemas, then disable unused toolsets with mibyan tools and uninstall skills you don’t need with mibyan skills.
  • Use GPU offloading — even a partial offload gives a significant speedup (see Step 6).

Model doesn’t follow tool calls

Models without tool-call support produce plain text instead of structured function calls. Solutions:
  • Use a model with tool-call support — of the models listed above, only gemma4:31b has reliable tool calling.
  • Mibyan has auto-repair — it detects malformed tool calls and attempts to fix them automatically.
  • Set up a fallback — if the local model fails 3 times, Mibyan falls back to a cloud provider.
If the model prints raw JSON like {"name": "web_search", ...} in its reply instead of actually running the tool, that’s usually the server, not the model — tool calling isn’t enabled or the tool-call format isn’t parsed. See the per-server fix table in Tool calls appear as text instead of executing (llama.cpp needs --jinja, vLLM needs --enable-auto-tool-choice --tool-call-parser mibyan, and so on).

Context window errors

The default Ollama context (2048 tokens) is too small for agentic work. See Step 6 to increase it.

Cost Comparison

Here’s what running locally saves compared to cloud APIs, based on a typical coding session (~100K tokens input, ~20K tokens output): Your only cost is electricity — roughly $0.01–0.05 per session depending on hardware.

What Works Well Locally

  • File editing and code generation — models 9B+ handle this well
  • Terminal commands — Mibyan wraps the command, runs it, reads output regardless of model
  • Web browsing — the browser tool does the fetching; the model just interprets results
  • Cron jobs and scheduled tasks — work identically to cloud setups
  • Multi-platform gateway — Telegram, Discord, Slack all work with local models

What’s Better with Cloud Models

  • Very complex multi-step reasoning — 70B+ or cloud models like Claude Opus are noticeably better
  • Long context windows — cloud models offer 100K–1M tokens; local runtimes often default below Mibyan’ 64K minimum unless you configure them
  • Speed on large responses — cloud inference is faster than CPU-only local for long generations
The sweet spot: use local for everyday tasks, set up a cloud fallback for the hard stuff.