Both expose an OpenAI-compatible
/v1/chat/completions endpoint. Mibyan works with either one — just point it at http://localhost:8080 or http://localhost:8000.
Apple Silicon onlyThis guide targets Macs with Apple Silicon (M1 and later). Intel Macs will work with llama.cpp but without GPU acceleration — expect significantly slower performance.
Choosing a model
For getting started, we recommend Qwen3.5-9B — it’s a strong reasoning model that fits comfortably in 8GB+ of unified memory with quantization.
Memory rule of thumb: model size + KV cache. A 9B Q4 model is ~5 GB. The KV cache at 128K context with Q4 quantization adds ~4-5 GB. With default (f16) KV cache, that balloons to ~16 GB. The quantized KV cache flags in llama.cpp are the key trick for memory-constrained systems.
For larger models (27B, 35B), you’ll need 32 GB+ of unified memory. The 9B is the sweet spot for 8-16 GB machines.
Option A: llama.cpp
llama.cpp is the most portable local LLM runtime. On macOS it uses Metal for GPU acceleration out of the box.Install
llama-server command globally.
Download the model
You need a GGUF-format model. The easiest source is Hugging Face via thehuggingface-cli:
Start the server
The server is ready when you see:
Memory optimization for constrained systems
The--cache-type-k q4_0 --cache-type-v q4_0 flags are the most important optimization for systems with limited memory. Here’s the impact at 128K context:
On an 8 GB Mac, use
q4_0 KV cache and choose a smaller model that can still fit Mibyan’ 64K minimum context. On 16 GB, you can comfortably do 128K context. On 32 GB+, you can run larger models or multiple parallel slots.
If you’re still running out of memory, reduce context only while staying at or above Mibyan’ 64K minimum; otherwise switch to a smaller model or smaller quantization (Q3_K_M instead of Q4_K_M).
Test it
Get the model name
If you forget the model name, query the models endpoint:Option B: MLX via omlx
omlx is a macOS-native app that manages and serves MLX models. MLX is Apple’s own machine learning framework, optimized specifically for Apple Silicon’s unified memory architecture.Install
Download and install from omlx.ai. It provides a GUI for model management and a built-in server.Download the model
Use the omlx app to browse and download models. Search forQwen3.5-9B-mlx-lm-mxfp4 and download it. Models are stored locally (typically in ~/.omlx/models/).
Start the server
omlx serves models onhttp://127.0.0.1:8000 by default. Start serving from the app UI, or use the CLI if available.
Test it
List available models
omlx can serve multiple models simultaneously:Benchmarks: llama.cpp vs MLX
Both backends tested on the same machine (Apple M5 Max, 128 GB unified memory) running the same model (Qwen3.5-9B) at comparable quantization levels (Q4_K_M for GGUF, mxfp4 for MLX). Five diverse prompts, three runs each, backends tested sequentially to avoid resource contention.Results
What this means
- llama.cpp excels at prompt processing — its flash attention + quantized KV cache pipeline gets you the first token in ~66ms. If you’re building interactive applications where perceived responsiveness matters (chatbots, autocomplete), this is a meaningful advantage.
- MLX generates tokens ~37% faster once it gets going. For batch workloads, long-form generation, or any task where total completion time matters more than initial latency, MLX finishes sooner.
- Both backends are extremely consistent — variance across runs was negligible. You can rely on these numbers.
Which one should you pick?
Connect to Mibyan
Once your local server is running:Timeouts
Mibyan automatically detects local endpoints (localhost, LAN IPs) and relaxes its streaming timeouts. No configuration needed for most setups. If you still hit timeout errors (e.g. very large contexts on slow hardware), you can override the streaming read timeout:
The stream read timeout is the one most likely to cause issues — it’s the socket-level deadline for receiving the next chunk of data. During prefill on large contexts, local models may produce no output for minutes while processing the prompt. The auto-detection handles this transparently.

