> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mibyanai.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Run Mibyan Locally with Ollama — Zero API Cost

> Step-by-step guide to running Mibyan entirely on your own machine with Ollama and open-weight models like Gemma 4, no cloud API keys or paid subscriptions needed

<Info>
  Commands, package names, and image names on this page come from the open-source project that Mibyan Desktop is built on, and can differ from the Mibyan Desktop installer. For the supported Mibyan install and update path, see [Install and update](/products/desktop-guide/install-and-update).
</Info>

<Tip>
  **Desktop users: there's a one-click path**

  On the Mibyan desktop app, **Settings → Providers → Local Models** installs
  and manages a local llama.cpp server for you — model downloads, memory
  fitting, and context sizing included. See [Local Models](/desktop/user-guide/local-models).
  This guide is for manual setup: Ollama specifically, CLI-first workflows,
  or servers you want to run yourself.
</Tip>

## The Problem

Cloud LLM APIs charge per token. A heavy coding session can cost \$5–20. For personal projects, learning, or privacy-sensitive work, that adds up — and you're sending every conversation to a third party.

## What This Guide Solves

You'll set up Mibyan running entirely on your own hardware, using [Ollama](https://ollama.com) as the model backend. No API keys, no subscriptions, no data leaving your machine. Once configured, Mibyan works exactly like it does with OpenRouter or Anthropic — terminal commands, file editing, web browsing, delegation — but the model runs locally.

By the end, you'll have:

* Ollama serving one or more open-weight models
* Mibyan connected to Ollama as a custom endpoint
* A working local agent that can edit files, run commands, and browse the web
* Optional: a Telegram/Discord bot powered entirely by your own hardware

## What You Need

| Component | Minimum | Recommended |
| - | - | - |
| **RAM** | 8 GB (for 3B models) | 32+ GB (for 27B+ models) |
| **Storage** | 5 GB free | 30+ GB (for multiple models) |
| **CPU** | 4 cores | 8+ cores (AMD EPYC, Ryzen, Intel Xeon) |
| **GPU** | Not required | NVIDIA GPU with 8+ GB VRAM speeds things up significantly |

<Tip>
  **CPU-only works, but expect slower responses**

  Ollama runs on CPU-only servers. A 9B model on a modern 8-core CPU gives \~10 tokens/sec. A 31B model on CPU is slower (\~2–5 tokens/sec) — each response takes 30–120 seconds, but it works. A GPU dramatically improves this. For CPU-only setups, widen the API timeout via the env var (it's not a `config.yaml` key):

  ```bash theme={null}
  # ~/.mibyan/.env
  mibyan_API_TIMEOUT=1800   # 30 minutes — generous for slow local models
  ```
</Tip>

## Step 1: Install Ollama

```bash theme={null}
curl -fsSL https://ollama.com/install.sh | sh
```

Verify it's running:

```bash theme={null}
ollama --version
curl http://localhost:11434/api/tags   # Should return {"models":[]}
```

## Step 2: Pull a Model

Choose based on your hardware:

| Model | Size on Disk | RAM Needed | Tool Calling | Best For |
| - | - | - | :-: | - |
| `gemma4:31b` | \~20 GB | 24+ GB | Yes | Best quality — strong tool use and reasoning |
| `gemma2:27b` | \~16 GB | 20+ GB | No | Conversational tasks, no tool use |
| `gemma2:9b` | \~5 GB | 8+ GB | No | Fast chat, Q\&A — cannot call tools |
| `llama3.2:3b` | \~2 GB | 4+ GB | No | Lightweight quick answers only |

<Warning>
  **Tool calling matters**

  Mibyan is an **agentic** assistant — it edits files, runs commands, and browses the web through tool calls. Models without tool-call support can only chat; they can't take actions. For the full Mibyan experience, use a model that supports tools (like `gemma4:31b`).
</Warning>

Pull your chosen model:

```bash theme={null}
ollama pull gemma4:31b
```

<Info>
  **Multiple models**

  You can pull several models and switch between them inside Mibyan with `/model`. Ollama loads the active model into memory on demand and unloads idle ones automatically.
</Info>

Verify the model works:

```bash theme={null}
curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma4:31b",
    "messages": [{"role": "user", "content": "Say hello"}],
    "max_tokens": 50
  }'
```

You should see a JSON response with the model's reply.

## Step 3: Configure Mibyan

Run the Mibyan setup wizard:

```bash theme={null}
mibyan setup
```

When prompted for a provider, select **Custom Endpoint** and enter:

* **Base URL:** `http://localhost:11434/v1`
* **API Key:** Leave empty or type `no-key` (Ollama doesn't need one)
* **Model:** `gemma4:31b` (or whichever model you pulled)

Alternatively, edit `~/.mibyan/config.yaml` directly:

```yaml theme={null}
model:
  default: "gemma4:31b"
  provider: "custom"
  base_url: "http://localhost:11434/v1"
```

## Step 4: Start Using Mibyan

```bash theme={null}
mibyan
```

That's it. You're now running a fully local agent. Try it out:

```
You: List all Python files in this directory and count the lines of code in each

You: Read the README.md and summarize what this project does

You: Create a Python script that fetches the weather for Ho Chi Minh City
```

Mibyan will use the terminal tool, file operations, and your local model — no cloud calls.

## Step 5: Pick the Right Model for Your Task

Not every task needs the biggest model. Here's a practical guide:

| Task | Recommended Model | Why |
| - | - | - |
| File edits, code, terminal commands | `gemma4:31b` | Only model with reliable tool calling |
| Quick Q\&A (no tool use needed) | `gemma2:9b` | Fast responses for conversational tasks |
| Lightweight chat | `llama3.2:3b` | Fastest, but very limited capabilities |

<Note>
  For full agentic work (editing files, running commands, browsing), `gemma4:31b` is currently the best local option with tool-call support. Check [Ollama's model library](https://ollama.com/library) for newer models — tool-calling support is expanding rapidly.
</Note>

Switch models on the fly inside a session:

```
/model gemma2:9b
```

## Step 6: Optimize for Speed

### Increase Ollama's Context Window

By default, Ollama uses a 2048-token context. Mibyan requires at least 64,000 tokens for agentic work with tools:

```bash theme={null}
# Create a Modelfile that extends context
cat > ~/.mibyan/cache/scratch/Modelfile << 'EOF'
FROM gemma4:31b
PARAMETER num_ctx 64000
EOF

ollama create gemma4-64k -f ~/.mibyan/cache/scratch/Modelfile
```

Then update your Mibyan config to use `gemma4-64k` as the model name.

### Keep the Model Loaded

By default, Ollama unloads models after 5 minutes of inactivity. For a persistent gateway bot, keep it loaded:

```bash theme={null}
# Set keep-alive to 24 hours
curl http://localhost:11434/api/generate \
  -d '{"model": "gemma4:31b", "keep_alive": "24h"}'
```

Or set it globally in Ollama's environment:

```bash theme={null}
# /etc/systemd/system/ollama.service.d/override.conf
[Service]
Environment="OLLAMA_KEEP_ALIVE=24h"
```

### Use GPU Offloading (If Available)

If you have an NVIDIA GPU, Ollama automatically offloads layers to it. Check with:

```bash theme={null}
ollama ps   # Shows which model is loaded and how many GPU layers
```

For a 31B model on a 12 GB GPU, you'll get partial offload (\~40 layers on GPU, rest on CPU), which still gives a significant speedup.

## Step 7: Run as a Gateway Bot (Optional)

Once Mibyan works locally in the CLI, you can expose it as a Telegram or Discord bot — still running entirely on your hardware.

### Telegram

1. Create a bot via [@BotFather](https://t.me/BotFather) and get the token
2. Add to your `~/.mibyan/config.yaml`:

```yaml theme={null}
model:
  default: "gemma4:31b"
  provider: "custom"
  base_url: "http://localhost:11434/v1"

platforms:
  telegram:
    enabled: true
    token: "YOUR_TELEGRAM_BOT_TOKEN"
```

3. Start the gateway:

```bash theme={null}
mibyan gateway
```

Now message your bot on Telegram — it responds using your local model.

### Discord

1. Create a Discord application at [discord.com/developers](https://discord.com/developers/applications)
2. Add to config:

```yaml theme={null}
platforms:
  discord:
    enabled: true
    token: "YOUR_DISCORD_BOT_TOKEN"
```

3. Start: `mibyan gateway`

## Step 8: Set Up Fallbacks (Optional)

Local models can struggle with complex tasks. Set up a cloud fallback that only activates when the local model fails:

```yaml theme={null}
model:
  default: "gemma4:31b"
  provider: "custom"
  base_url: "http://localhost:11434/v1"

fallback_providers:
  - provider: openrouter
    model: anthropic/claude-sonnet-4
```

This way, 90% of your usage is free (local), and only the hard tasks hit the paid API.

## Troubleshooting

### "provider 'ollama' has no endpoint configured"

`mibyan chat --provider ollama` (or `vllm`) stops with this error when no endpoint is configured for that alias anywhere — no `providers.ollama.base_url`, no `model.base_url`. Mibyan refuses to send the request rather than fall back to OpenRouter with a cloud key (`OPENROUTER_API_KEY` / `OPENAI_API_KEY`) that happens to be set. Add the endpoint:

```yaml theme={null}
providers:
  ollama:
    base_url: "http://localhost:11434/v1"
```

### "Connection refused" on startup

Ollama isn't running. Start it:

```bash theme={null}
sudo systemctl start ollama
# or
ollama serve
```

### Slow responses

* **Check model size vs RAM:** If your model needs more RAM than available, it swaps to disk. Use a smaller model or add RAM.
* **Check `ollama ps`:** If no GPU layers are offloaded, responses are CPU-bound. This is normal for CPU-only servers.
* **Reduce context:** Large conversations slow down inference. Use `/compress` regularly, or set a lower compression threshold in config.

### Slow first response (prefill)

Mibyan sends a fixed payload on every API call — the system prompt plus the tool schemas for all enabled tools — before any of your conversation content. On CPU-only or low-VRAM setups, processing that prompt (the *prefill* phase) dominates the first turn: the model can sit silent for minutes while it works through the prompt, then generate at its normal pace. This is expected behaviour, not a hang. The [Mac local-LLM guide](/desktop/guides/local-llm-on-mac#timeouts) documents the same effect — during prefill on large contexts, local models may produce no output for minutes while processing the prompt — and Mibyan automatically raises its stream read timeout from 120s to 1800s for local endpoints (`mibyan_STREAM_READ_TIMEOUT`).

What helps:

* **Keep the model loaded** — Ollama unloads idle models after 5 minutes, adding a full reload before the next prefill. Set `OLLAMA_KEEP_ALIVE=24h` (see [Step 6](#keep-the-model-loaded)).
* **Widen the API timeout** — set `mibyan_API_TIMEOUT=1800` in `~/.mibyan/.env` (see [What You Need](#what-you-need)).
* **Measure and trim the fixed prompt** — run `mibyan prompt-size` for a byte breakdown of the system prompt and tool schemas, then disable unused toolsets with `mibyan tools` and uninstall skills you don't need with `mibyan skills`.
* **Use GPU offloading** — even a partial offload gives a significant speedup (see [Step 6](#use-gpu-offloading-if-available)).

### Model doesn't follow tool calls

Models without tool-call support produce plain text instead of structured function calls. Solutions:

* **Use a model with tool-call support** — of the models listed above, only `gemma4:31b` has reliable tool calling.
* **Mibyan has auto-repair** — it detects malformed tool calls and attempts to fix them automatically.
* **Set up a fallback** — if the local model fails 3 times, Mibyan falls back to a cloud provider.

If the model prints raw JSON like `{"name": "web_search", ...}` in its reply instead of actually running the tool, that's usually the *server*, not the model — tool calling isn't enabled or the tool-call format isn't parsed. See the per-server fix table in [Tool calls appear as text instead of executing](/desktop/integrations/providers#tool-calls-appear-as-text-instead-of-executing) (llama.cpp needs `--jinja`, vLLM needs `--enable-auto-tool-choice --tool-call-parser mibyan`, and so on).

### Context window errors

The default Ollama context (2048 tokens) is too small for agentic work. See [Step 6](#step-6-optimize-for-speed) to increase it.

## Cost Comparison

Here's what running locally saves compared to cloud APIs, based on a typical coding session (\~100K tokens input, \~20K tokens output):

| Provider | Cost per Session | Monthly (daily use) |
| - | - | - |
| Anthropic Claude Sonnet | \~\$0.80 | \~\$24 |
| OpenRouter (GPT-4o) | \~\$0.60 | \~\$18 |
| **Ollama (local)** | **\$0.00** | **\$0.00** |

Your only cost is electricity — roughly \$0.01–0.05 per session depending on hardware.

## What Works Well Locally

* **File editing and code generation** — models 9B+ handle this well
* **Terminal commands** — Mibyan wraps the command, runs it, reads output regardless of model
* **Web browsing** — the browser tool does the fetching; the model just interprets results
* **Cron jobs and scheduled tasks** — work identically to cloud setups
* **Multi-platform gateway** — Telegram, Discord, Slack all work with local models

## What's Better with Cloud Models

* **Very complex multi-step reasoning** — 70B+ or cloud models like Claude Opus are noticeably better
* **Long context windows** — cloud models offer 100K–1M tokens; local runtimes often default below Mibyan' 64K minimum unless you configure them
* **Speed on large responses** — cloud inference is faster than CPU-only local for long generations

The sweet spot: use local for everyday tasks, set up a cloud fallback for the hard stuff.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.