Skip to main content
Desktop users: there’s a one-click pathOn the Mibyan desktop app, Settings → Providers → Local Models installs and manages a local llama.cpp server for you — model downloads, memory fitting, and context sizing included. See Local Models. This guide is for manual setup: MLX, custom builds, or servers you want to run yourself.
This guide walks you through running a local LLM server on macOS with an OpenAI-compatible API. You get full privacy, zero API costs, and surprisingly good performance on Apple Silicon. We cover two backends: Both expose an OpenAI-compatible /v1/chat/completions endpoint. Mibyan works with either one — just point it at http://localhost:8080 or http://localhost:8000.
Apple Silicon onlyThis guide targets Macs with Apple Silicon (M1 and later). Intel Macs will work with llama.cpp but without GPU acceleration — expect significantly slower performance.

Choosing a model

For getting started, we recommend Qwen3.5-9B — it’s a strong reasoning model that fits comfortably in 8GB+ of unified memory with quantization. Memory rule of thumb: model size + KV cache. A 9B Q4 model is ~5 GB. The KV cache at 128K context with Q4 quantization adds ~4-5 GB. With default (f16) KV cache, that balloons to ~16 GB. The quantized KV cache flags in llama.cpp are the key trick for memory-constrained systems. For larger models (27B, 35B), you’ll need 32 GB+ of unified memory. The 9B is the sweet spot for 8-16 GB machines.

Option A: llama.cpp

llama.cpp is the most portable local LLM runtime. On macOS it uses Metal for GPU acceleration out of the box.

Install

This gives you the llama-server command globally.

Download the model

You need a GGUF-format model. The easiest source is Hugging Face via the huggingface-cli:
Then download:
Gated modelsSome models on Hugging Face require authentication. Run huggingface-cli login first if you get a 401 or 404 error.

Start the server

Here’s what each flag does: The server is ready when you see:

Memory optimization for constrained systems

The --cache-type-k q4_0 --cache-type-v q4_0 flags are the most important optimization for systems with limited memory. Here’s the impact at 128K context: On an 8 GB Mac, use q4_0 KV cache and choose a smaller model that can still fit Mibyan’ 64K minimum context. On 16 GB, you can comfortably do 128K context. On 32 GB+, you can run larger models or multiple parallel slots. If you’re still running out of memory, reduce context only while staying at or above Mibyan’ 64K minimum; otherwise switch to a smaller model or smaller quantization (Q3_K_M instead of Q4_K_M).

Test it

Get the model name

If you forget the model name, query the models endpoint:

Option B: MLX via omlx

omlx is a macOS-native app that manages and serves MLX models. MLX is Apple’s own machine learning framework, optimized specifically for Apple Silicon’s unified memory architecture.

Install

Download and install from omlx.ai. It provides a GUI for model management and a built-in server.

Download the model

Use the omlx app to browse and download models. Search for Qwen3.5-9B-mlx-lm-mxfp4 and download it. Models are stored locally (typically in ~/.omlx/models/).

Start the server

omlx serves models on http://127.0.0.1:8000 by default. Start serving from the app UI, or use the CLI if available.

Test it

List available models

omlx can serve multiple models simultaneously:

Benchmarks: llama.cpp vs MLX

Both backends tested on the same machine (Apple M5 Max, 128 GB unified memory) running the same model (Qwen3.5-9B) at comparable quantization levels (Q4_K_M for GGUF, mxfp4 for MLX). Five diverse prompts, three runs each, backends tested sequentially to avoid resource contention.

Results

What this means

  • llama.cpp excels at prompt processing — its flash attention + quantized KV cache pipeline gets you the first token in ~66ms. If you’re building interactive applications where perceived responsiveness matters (chatbots, autocomplete), this is a meaningful advantage.
  • MLX generates tokens ~37% faster once it gets going. For batch workloads, long-form generation, or any task where total completion time matters more than initial latency, MLX finishes sooner.
  • Both backends are extremely consistent — variance across runs was negligible. You can rely on these numbers.

Which one should you pick?


Connect to Mibyan

Once your local server is running:
Select Custom endpoint and follow the prompts. It will ask for the base URL and model name — use the values from whichever backend you set up above.

Timeouts

Mibyan automatically detects local endpoints (localhost, LAN IPs) and relaxes its streaming timeouts. No configuration needed for most setups. If you still hit timeout errors (e.g. very large contexts on slow hardware), you can override the streaming read timeout:
The stream read timeout is the one most likely to cause issues — it’s the socket-level deadline for receiving the next chunk of data. During prefill on large contexts, local models may produce no output for minutes while processing the prompt. The auto-detection handles this transparently.
A silent first turn is usually prefill, not a hangMibyan sends its system prompt and tool schemas on every call, so on slower hardware the first turn can involve minutes of silence while the model processes that prompt before generating anything. That’s prefill at work, not a stalled session. See Slow first response (prefill) in the Ollama guide for mitigations like keeping the model loaded and trimming the fixed prompt with mibyan prompt-size.