Inference Providers
You need at least one way to connect to an LLM. Usemibyan model to switch providers and models interactively, or configure directly:
Both built-in OpenCode providers send an opaque, per-conversation
x-opencode-session header on every request (main turns on every transport plus auxiliary calls such as compression, titles, approval checks, skills-hub lookups and /btw side questions — including the ones that run in the background after the turn has ended; headless Kanban specify/decompose and dashboard estimate calls use a per-task key; one-shots with no live session at all, such as Desktop commit-message generation from the review panel, send a fresh ephemeral key). OpenCode uses it to pin a conversation to one backend so its prompt cache stays warm; the value is derived from the Mibyan session id (or the Kanban task id) and carries no personal data.
The two built-in OpenCode providers each pin their own relay on opencode.ai (opencode-zen → /zen/v1, opencode-go → /zen/go/v1). A model.base_url left behind by the other relay is healed to the selected provider’s relay, and the model you pick (-m, /model, a fallback entry or a channel override) decides which relay is used — so switching from a Zen model to a Go-only one never sends the request to Zen. The same per-model routing is re-applied when a session is resumed (mibyan --resume, /resume, the TUI and desktop resume paths): a wire format or relay URL recorded while the session ran a different OpenCode model never carries over to the model the session is reopened on. OpenCode models whose id carries a -vision marker (for example deepseek-v4-flash-vision-exp) are treated as vision-capable even before models.dev lists them, so agent.image_input_mode: auto attaches images natively without a supports_vision override. A custom provider you define under providers: whose name extends a family slug (for example opencode-go-bridge) still gets the family’s per-model API-mode routing and /v1 handling, but its base_url is taken as declared: name it after the relay it actually points at. Auxiliary tasks (auxiliary.compression, titles, vision, MoA) pointed at an OpenCode provider follow the same per-model table, so a Responses-only model such as gpt-5.6-luna or an Anthropic-wire one such as minimax-m2.5 works there exactly as it does for the main conversation.
For the official API-key path, see the dedicated Google Gemini guide.
Nous Portal
Nous Portal is Nous Research’s unified subscription gateway and the recommended way to run Mibyan. One OAuth login covers 300+ frontier agentic models (Claude, GPT, Gemini, DeepSeek, Qwen, Kimi, GLM, MiniMax, Grok, …) plus the Tool Gateway (web search, image generation, TTS, browser automation) — billed against your Nous subscription instead of separate per-provider accounts.client=mibyan-client-v<version> tag (e.g. client=mibyan-client-v0.13.0) auto-aligned to your installed release. This is sent on all Portal pathways — main chat loop, auxiliary calls, compression summarizer, web extraction — and lets Portal-side telemetry distinguish Mibyan traffic from other clients. No config required; the tag updates automatically when you mibyan update.
JWT auth (automatic). Mibyan prefers scoped inference:invoke JWTs for Portal requests with the legacy opaque session-key path as a fallback. No configuration is required — credentials are managed by the OAuth flow and rotate transparently. Revoked refresh tokens are quarantined to avoid replay loops.
Codex NoteThe OpenAI Codex provider authenticates via device code by default (open a URL, enter a code). Organizations that disable the device-code grant can opt in to the browser authorization-code + PKCE flow instead:
mibyan auth add openai-codex --browser (one login) or auth.codex_login_flow: browser in config.yaml (every Codex login, including mibyan model). That flow listens on http://localhost:1455/auth/callback — the redirect URI registered for the Codex client, so the port is fixed; if it is already taken (a Codex CLI sign-in in progress) Mibyan says so and falls back to device code. Over SSH the listener needs a tunnel (ssh -N -L 1455:127.0.0.1:1455 user@host, see OAuth over SSH). Mibyan stores the resulting credentials in its own auth store under ~/.mibyan/auth.json and can import existing Codex CLI credentials from ~/.codex/auth.json when present. No Codex CLI installation is required. Automatic adoption of the Codex CLI login (when Mibyan’ own refresh fails) is controlled by auth.adopt_external_logins — see Borrowed CLI logins.If a token refresh fails with a terminal error (HTTP 4xx, invalid_grant, revoked grant, etc.), Mibyan marks the refresh token as dead and stops replaying it so you don’t see a flood of identical auth failures. The next request surfaces a typed re-auth message instead. Run mibyan auth add openai-codex (or mibyan model → ChatGPT or Codex Subscription) to start a fresh login (device code, or --browser for the loopback PKCE flow); the quarantine clears on the next successful exchange.Device login can fail with [SSL: UNEXPECTED_EOF_WHILE_READING] or a TLS handshake timeout on Python/OpenSSL 3.5+ when a middlebox rejects post-quantum groups such as X25519MLKEM768 (curl may still work). A one-off dropped connection is not fatal: while waiting for your browser approval Mibyan keeps polling through up to six consecutive transport errors (and retries the device-code request and token exchange twice) before giving up, so only a persistently broken network surfaces this error. Mibyan does not change default TLS policy. Point OPENSSL_CONF at a config that restricts Groups to classic curves before running mibyan model, or diagnose with TLS 1.2:Two Commands for Model Management
Mibyan has two model commands that serve different purposes:
If you’re trying to switch to a provider you haven’t set up yet (e.g. you only have OpenRouter configured and want to use Anthropic), you need
mibyan model, not /model. Exit your session first (Ctrl+C or /quit), run mibyan model, complete the provider setup, then start a new session.
Subscription plans: what your plan pays for
Several providers let you sign in to Mibyan with a consumer subscription (Claude Max, ChatGPT, SuperGrok / X Premium+, …) instead of an API key. What that subscription actually pays for — and what it doesn’t — differs per provider, and it’s the single most common source of billing surprises. The table below is the short version; each provider’s own section has the details.Cells marked not currently documented mean exactly that: Mibyan docs do not yet specify the behavior. Don’t assume — check your provider’s billing dashboard, and treat these as open questions.
Anthropic. The OAuth path routes as Claude Code against your Anthropic account and only works on a Claude Max plan with purchased extra usage credits — the base Max allowance is never consumed by Mibyan, only the extra/overage credits on top. Claude Pro subscribers cannot use this path; the supported alternative is an
ANTHROPIC_API_KEY, billed pay-per-token against that key’s organization at standard API pricing. See Anthropic (Native) below.
OpenAI Codex. Mibyan authenticates via ChatGPT device-code OAuth, stores credentials in ~/.mibyan/auth.json, and can import existing Codex CLI credentials from ~/.codex/auth.json. Which ChatGPT plan tiers are eligible, and how Mibyan usage counts against your plan’s Codex limits, are not currently documented — the Codex note under Nous Portal covers authentication and token-refresh behavior only.
xAI (SuperGrok / X Premium+). Browser OAuth works with either an active SuperGrok subscription or an X Premium+ subscription on the linked X account, and the same bearer token is reused by direct-to-xAI tools (TTS, image gen, video gen, transcription, X Search). If inference returns HTTP 403 after a successful login, that’s a tier/entitlement restriction on xAI’s side, not a stale token — the workaround is switching to an XAI_API_KEY. See xAI (Grok) below and the xAI Grok OAuth guide.
Google Gemini. There is currently no way to sign in to Mibyan with a consumer Gemini subscription — the gemini provider takes an API key, and Google Vertex AI bills to your GCP project. A billing-enabled Google Cloud project is recommended for agent use; free-tier quotas are too small for long-running agent sessions. See the Google Gemini guide.
Anthropic (Native)
Use Claude models directly through the Anthropic API — no OpenRouter proxy needed. Supports three auth methods: When no explicit environment credential is selected, Mibyan-owned OAuth grants in the credential pool take precedence over a borrowed Claude Code login. The borrowed login remains the fallback when no owned OAuth grant is available — unlessauth.adopt_external_logins: false is set, in which case Mibyan never
reads or refreshes Claude Code’s credentials (see
Borrowed CLI logins).
Auxiliary authentication recovery refreshes the credential used by the failed
request, not an unrelated ambient login; rotating a borrowed login can otherwise
invalidate its owner’s refresh token.
mibyan model, Mibyan prefers Claude Code’s own credential store over copying the token into ~/.mibyan/.env. That keeps refreshable Claude credentials refreshable.
Or set it permanently:
GitHub Copilot
Mibyan supports GitHub Copilot as a first-class provider with two modes:copilot — Direct Copilot API (recommended). Uses your GitHub Copilot subscription to access GPT-5.x, Claude, Gemini, and other models through the Copilot API.
COPILOT_GITHUB_TOKENenvironment variableGH_TOKENenvironment variableGITHUB_TOKENenvironment variablegh auth tokenCLI fallback
mibyan model offers an OAuth device code login — the same flow used by the Copilot CLI and opencode.
Copilot auth behavior in MibyanMibyan sends a supported GitHub token (
gho_*, github_pat_*, or ghu_*) directly to api.githubcopilot.com and includes Copilot-specific headers (Editor-Version, Copilot-Integration-Id, Openai-Intent, x-initiator).On HTTP 401, Mibyan now performs a one-shot credential recovery before fallback:- Re-resolve token via the normal priority chain (
COPILOT_GITHUB_TOKEN→GH_TOKEN→GITHUB_TOKEN→gh auth token) - Rebuild the shared OpenAI client with refreshed headers
- Retry the request once
api.github.com/copilot_internal/v2/token exchange flows. That endpoint can be unavailable for some account types (returns 404). Mibyan therefore keeps direct-token auth as the primary path and relies on runtime credential refresh + retry for robustness.gpt-5-mini) automatically use the Responses API. All other models (GPT-4o, Claude, Gemini, etc.) use Chat Completions. Models are auto-detected from the live Copilot catalog.
copilot-acp — Copilot ACP agent backend. Spawns the local Copilot CLI as a subprocess:
First-Class API-Key Providers
These providers have built-in support with dedicated provider IDs. Set the API key and use--provider to select:
accounts/fireworks/models/kimi-k2p6. Run mibyan model, choose Fireworks AI, and select from the live catalog or enter another Fireworks model ID. The default endpoint is https://api.fireworks.ai/inference/v1; configure a different endpoint through model.base_url in config.yaml, not .env.
Or set the provider permanently in config.yaml:
NOVITA_BASE_URL, GLM_BASE_URL, KIMI_BASE_URL, MINIMAX_BASE_URL, MINIMAX_CN_BASE_URL, DASHSCOPE_BASE_URL, XIAOMI_BASE_URL, GMI_BASE_URL, META_BASE_URL, or TOKENHUB_BASE_URL environment variables.
Meta contributor tier
muse-spark-1.2-contributor and muse-spark-1.3-contributor are Meta’s contributor tiers — Meta may train on your prompts and completions, so interactive model selection asks for confirmation before using either. For current pricing and rate limits, see Meta Model API pricing and rate limits. Use the standard muse-spark-1.2 / muse-spark-1.3 (no training) for confidential work.Z.AI Endpoint Auto-DetectionWhen using the Z.AI / GLM provider, Mibyan automatically probes multiple endpoints (global, China, coding variants) to find one that accepts your API key. You don’t need to set
GLM_BASE_URL manually — the working endpoint is detected and cached automatically.xAI (Grok) — Responses API + Prompt Caching
xAI is wired through the Responses API (codex_responses transport) for automatic reasoning support on Grok 4 models — no reasoning_effort parameter needed, the server reasons by default. Set XAI_API_KEY in ~/.mibyan/.env and pick xAI in mibyan model, or drop grok as a shortcut into /model grok-4-fast-reasoning.
SuperGrok and X Premium+ subscribers can sign in with browser OAuth instead of using an API key — pick xAI Grok OAuth (SuperGrok / Premium+) in mibyan model, or run mibyan auth add xai-oauth. The same OAuth bearer token is automatically reused by direct-to-xAI tools (TTS, image gen, video gen, transcription). See the xAI Grok OAuth guide for the full flow — and if Mibyan runs on a remote host, also see OAuth over SSH / Remote Hosts for the required ssh -L tunnel.
When using xAI as a provider (any base URL containing x.ai), Mibyan automatically enables prompt caching by sending the x-grok-conv-id header with every API request. This routes requests to the same server within a conversation session, allowing xAI’s infrastructure to reuse cached system prompts and conversation history.
No configuration is needed — caching activates automatically when an xAI endpoint is detected and a session ID is available. This reduces latency and cost for multi-turn conversations.
xAI also ships a dedicated TTS endpoint (/v1/tts). Select xAI TTS in mibyan tools → Voice & TTS, or see the Voice & TTS page for config.
Retired xAI model migration (May 15, 2026): xAI is retiring grok-4*, grok-3, grok-code-fast-1, and grok-imagine-image-pro on 2026-05-15. mibyan doctor and mibyan chat startup both detect any config still pointing at a retired ref and print the recommended replacement. Use mibyan migrate xai for a one-shot config rewrite — dry-run by default, add --apply to write changes (a timestamped copy of the previous config lands in backups/config/ first).
web.backend: xai routes search through xAI’s hosted search endpoint using the same XAI_API_KEY / OAuth credentials. No additional setup required if xAI is already configured as a provider.
NovitaAI
NovitaAI is the AI-native cloud for builders and agents. Its three product lines are Model API for 200+ models, Agent Sandbox for building and running AI agents, and GPU Cloud for scalable compute, all available from one platform.config.yaml:
NOVITA_BASE_URL.
Ollama Cloud — Managed Ollama Models, OAuth + API Key
Ollama Cloud hosts the same open-weight catalog as local Ollama but without the GPU requirement. Pick it inmibyan model as Ollama Cloud, paste your API key from ollama.com/settings/keys, and Mibyan auto-discovers the available models.
config.yaml directly:
ollama.com/v1/models and cached for one hour. model:tag notation (e.g. qwen3-coder:480b-cloud) is preserved through normalization — don’t use dashes.
DeepInfra
DeepInfra (--provider deepinfra, DEEPINFRA_API_KEY) is discovered live from its catalog. Reasoning is controlled through DeepInfra’s top-level reasoning_effort field, so agent.reasoning_effort, /reasoning <level>, --reasoning and per-model agent.reasoning_overrides work in both directions: an effort turns thinking on for models that default off (DeepSeek-V4.x), /reasoning none turns it off for models that default on (GLM-4.6, Qwen3-Thinking). Leaving reasoning unset keeps DeepInfra’s per-model default; xhigh is native and ultra is sent as max.
AWS Bedrock
Anthropic Claude, Amazon Nova, DeepSeek v3.2, Meta Llama 4, and other models via AWS Bedrock. Uses the AWS SDK (boto3) credential chain — no API key, just standard AWS auth.
config.yaml:
AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY, AWS_PROFILE from ~/.aws/credentials, IAM role on EC2/ECS/Lambda, IMDS, or SSO. No env var is required if you’re already authenticated with the AWS CLI.
Bedrock uses the Converse API under the hood — requests are translated to Bedrock’s model-agnostic shape, so the same config works for Claude, Nova, DeepSeek, and Llama models. Set BEDROCK_BASE_URL only if you’re calling a non-default regional endpoint.
See the AWS Bedrock guide for a walkthrough of IAM setup, region selection, and cross-region inference.
Google Vertex AI
Gemini models on Google Cloud Vertex AI via Vertex’s OpenAI-compatible endpoint. Authentication is OAuth2 — a short-lived access token (~1 hour) minted from a service-account JSON or Application Default Credentials (ADC). There is no static API key; Mibyan mints and auto-refreshes the token for you, including re-minting on a mid-session401.
config.yaml (project/region are non-secret and live here; the credential path stays in .env):
VERTEX_PROJECT_ID / VERTEX_REGION env vars override the config.yaml values. Mibyan lazy-installs google-auth on first use; run mibyan setup if the managed install needs repair. See the Google Vertex AI guide for the full walkthrough, and the Google Gemini guide for the static-API-key AI Studio path instead.
Qwen Portal (OAuth)
Alibaba’s Qwen Portal with browser-based OAuth login. Pick Qwen OAuth (Portal) inmibyan model, sign in through the browser, and Mibyan persists the refresh token.
config.yaml:
mibyan_QWEN_BASE_URL only if the portal endpoint relocates (default: https://portal.qwen.ai/v1).
Alibaba Cloud (Coding Plan)
If you’re subscribed to Alibaba’s Coding Plan (a pricing SKU separate from standard DashScope API access), Mibyan exposes it as its own first-class provider:alibaba-coding-plan. Endpoint: https://coding-intl.dashscope.aliyuncs.com/v1. It’s OpenAI-compatible like the regular alibaba provider but with a different base URL and billing surface.
alibaba_coding uses the same DASHSCOPE_API_KEY your alibaba entry already uses — no separate key needed, just a different routing target. Before this provider was registered, users who set provider: alibaba_coding in config.yaml silently fell through to OpenRouter routing.
For the mainland-China endpoint (alibaba-coding-plan-cn, https://coding.dashscope.aliyuncs.com/v1) set ALIBABA_CODING_PLAN_CN_API_KEY. The CN provider still falls back to ALIBABA_CODING_PLAN_API_KEY / DASHSCOPE_API_KEY, but with only the shared key set the /model picker lists just the international row — set the CN key (or provider: alibaba-coding-plan-cn in config.yaml) to surface the CN one. The same applies to alibaba-token-plan-cn with ALIBABA_TOKEN_PLAN_CN_API_KEY.
MiniMax (OAuth)
MiniMax-M2.7 via browser OAuth login — no API key needed. Pick MiniMax (OAuth) inmibyan model, sign in through the browser, and Mibyan persists the access + refresh tokens. Uses the Anthropic Messages-compatible endpoint (/anthropic) under the hood.
config.yaml:
MiniMax-M2.7 (main) and MiniMax-M2.7-highspeed (wired as the default auxiliary model). The OAuth path ignores MINIMAX_API_KEY / MINIMAX_BASE_URL.
NVIDIA NIM
Nemotron and other open source models via build.nvidia.com (free API key) or a local NIM endpoint.config.yaml:
build.nvidia.com — no configuration needed. This routes consumption against the correct origin in NVIDIA’s billing dashboard.
GMI Cloud
Open and reasoning models via GMI Cloud — OpenAI-compatible API, API key authentication.config.yaml:
GMI_BASE_URL (default: https://api.gmi-serving.com/v1).
Actual Computer
Your own hardware as a private inference cluster via Actual Computer. Two serving modes, both using Chat Completions so reasoning and final content are returned together:- Hosted relay —
https://api.actual.inc, end-to-end encrypted, routes to your cluster. Authenticate with anac_inference key from actual.inc/user/keys. - Local daemon — on-device at
http://127.0.0.1:8080, fully offline. No API key needed: Mibyan detects the loopback base URL and authenticates with an internal placeholder automatically.
~/.mibyan/config.yaml; only the hosted API key belongs in .env:
- Model IDs come from your cluster’s
GET /v1/models— discover withmibyan modelorcurl -s https://api.actual.inc/v1/models -H "Authorization: Bearer $ACTUAL_API_KEY". - Bare hosts in
model.base_urlare normalized:http://127.0.0.1:8080becomeshttp://127.0.0.1:8080/v1automatically. The legacyACTUAL_BASE_URLenvironment variable is a fallback when no Actual URL is configured in YAML. - Actual uses
/v1/chat/completionsfor chat, compaction, title generation, and every other auxiliary task. This also applies to custom providers targetingapi.actual.inc, model switches, and fallbacks. Legacy Responses settings in the main model, custom provider, or auxiliary task configuration are overridden automatically. - Reasoning effort is clamped to Actual’s supported range (
none/low/medium/high/max) — a globalxhigh/ultrasetting will not 400 requests. - Small local models: Mibyan’ full default toolset plus the system prompt can exceed a 32k context window, producing an empty-stream error from llama.cpp-family servers. Restrict the toolset (
-t file,web) or load the model with a larger context. The optionalactual-setupskill (mibyan skills install official/devops/actual-setup) covers setup and troubleshooting in detail. - Aliases:
actual-computer,actualcomputer,aci.
StepFun
Step-series models via StepFun — OpenAI-compatible API, API key authentication.config.yaml:
STEPFUN_BASE_URL (default: https://api.stepfun.com/v1).
Hugging Face Inference Providers
Hugging Face Inference Providers routes to 20+ open models through a unified OpenAI-compatible endpoint (router.huggingface.co/v1). Requests are automatically routed to the fastest available backend (Groq, Together, SambaNova, etc.) with automatic failover.
config.yaml:
:fastest (default), :cheapest, or :provider_name to force a specific backend.
The base URL can be overridden with HF_BASE_URL.
Custom & Self-Hosted LLM Providers
Mibyan works with any OpenAI-compatible API endpoint. If a server implements/v1/chat/completions, you can point Mibyan at it. This means you can use local models, GPU inference servers, multi-provider routers, or any third-party API.
General Setup
Three ways to configure a custom endpoint: Interactive setup (recommended):config.yaml):
config.yaml, which is the source of truth for model, provider, and base URL.
Switching Models with /model
Once you have at least one custom endpoint configured, you can switch models mid-session:
Ollama — Local Models, Zero Config
Ollama runs open-weight models locally with one command. Best for: quick local experimentation, privacy-sensitive work, offline use. Supports tool calling via the OpenAI-compatible API.config.yaml directly:
vLLM — High-Performance GPU Inference
vLLM is the standard for production LLM serving. Best for: maximum throughput on GPU hardware, serving large models, continuous batching.max_position_embeddings by default. If that exceeds your GPU memory, it errors and asks you to set --max-model-len lower. You can also use --max-model-len auto to automatically find the maximum that fits. Set --gpu-memory-utilization 0.95 (default 0.9) to squeeze more context into VRAM.
Tool calling requires explicit flags:
Supported parsers:
hermes (Qwen 2.5, Hermes 2/3), llama3_json (Llama 3.x), mistral, deepseek_v3, deepseek_v31, xlam, pythonic. Without these flags, tool calls won’t work — the model will output tool calls as text.
Qwen reasoning parsers: Mibyan preserves structured reasoning metadata such as reasoning, reasoning_content, and streamed reasoning deltas when OpenAI-compatible servers return them. That metadata is treated as reasoning/thinking trace data, not as a replacement for the assistant’s visible answer. For Qwen reasoning models served by vLLM, make sure the final user-visible response still appears in content. If --reasoning-parser qwen3 leaves content empty in your deployment, either disable that parser or pass a server-supported request option such as chat_template_kwargs.enable_thinking: false through extra_body.
SGLang — Fast Serving with RadixAttention
SGLang is an alternative to vLLM with RadixAttention for KV cache reuse. Best for: multi-turn conversations (prefix caching), constrained decoding, structured output.--context-length to override. If you need to exceed the model’s declared maximum, set SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1.
Tool calling: Use --tool-call-parser with the appropriate parser for your model family: qwen (Qwen 2.5), llama3, llama4, deepseekv3, mistral, glm. Without this flag, tool calls come back as plain text.
llama.cpp / llama-server — CPU & Metal Inference
llama.cpp runs quantized models on CPU, Apple Silicon (Metal), and consumer GPUs. Best for: running models without a datacenter GPU, Mac users, edge deployment.-c): Recent builds default to 0 which reads the model’s training context from the GGUF metadata. For models with 128k+ training context, this can OOM trying to allocate the full KV cache. Set -c explicitly to at least 64,000 tokens for Mibyan. If using parallel slots (-np), the total context is divided among slots — with -c 64000 -np 4, each slot only gets 16k, which is below Mibyan’ minimum per active session.
Then configure Mibyan to point at it:
config.yaml so it persists across sessions.
LM Studio — Desktop App with Local Models
LM Studio is a desktop app for running local models with a GUI. Best for: users who prefer a visual interface, quick model testing, developers on macOS/Windows/Linux. Start the server from the LM Studio app (Developer tab → Start Server), or use the CLI:context_length unless you configured one in Mibyan, so LM Studio can apply its own model setting. Mibyan then uses only the context length LM Studio reports after loading.
To change context length in LM Studio:
- Click the gear icon next to the model picker
- Set “Context Length” to at least 64000 for a smooth experience
- Reload the model for the change to take effect
- If your machine cannot fit 64000, consider using a smaller model with larger context lengths.
lms load model-name --context-length 64000
You can use the CLI to estimate if the model will fit: lms load model-name --context-length 64000 --estimate-only
To set persistent per-model defaults: My Models tab → gear icon on the model → set context size.
:::
If you use LM Studio’s Just-In-Time loading / Auto-Evict feature and want LM Studio to manage model loading and eviction from normal chat requests, skip Mibyan’ explicit preload step:
WSL2 Networking (Windows Users)
Since Mibyan requires a Unix environment, Windows users run it inside WSL2. If your model server (Ollama, LM Studio, etc.) runs on the Windows host, you need to bridge the network gap — WSL2 uses a virtual network adapter with its own subnet, solocalhost inside WSL2 refers to the Linux VM, not the Windows host.
Option 1: Mirrored Networking Mode (Recommended)
Available on Windows 11 22H2+, mirrored mode makeslocalhost work bidirectionally between Windows and WSL2 — the simplest fix.
-
Create or edit
%USERPROFILE%\.wslconfig(e.g.,C:\Users\YourName\.wslconfig): -
Restart WSL from PowerShell:
-
Reopen your WSL2 terminal.
localhostnow reaches Windows services:
Hyper-V FirewallOn some Windows 11 builds, the Hyper-V firewall blocks mirrored connections by default. If
localhost still doesn’t work after enabling mirrored mode, run this in an Admin PowerShell:Option 2: Use the Windows Host IP (Windows 10 / older builds)
If you can’t use mirrored mode, find the Windows host IP from inside WSL2 and use that instead oflocalhost:
Server Bind Address (Required for NAT Mode)
If you’re using Option 2 (NAT mode with the host IP), the model server on Windows must accept connections from outside127.0.0.1. By default, most servers only listen on localhost — WSL2 connections in NAT mode come from a different virtual subnet and will be refused. In mirrored mode, localhost maps directly so the default 127.0.0.1 binding works fine.
Ollama on Windows (detailed): Ollama runs as a Windows service. To set
OLLAMA_HOST:
- Open System Properties → Environment Variables
- Add a new System variable:
OLLAMA_HOST=0.0.0.0 - Restart the Ollama service (or reboot)
Windows Firewall
Windows Firewall treats WSL2 as a separate network (in both NAT and mirrored mode). If connections still fail after the steps above, add a firewall rule for your model server’s port:11434, vLLM 8000, SGLang 30000, llama-server 8080, LM Studio 1234.
Quick Verification
From inside WSL2, test that you can reach your model server:base_url in your Mibyan config.
Troubleshooting Local Models
These issues affect all local inference servers when used with Mibyan.”Connection refused” from WSL2 to a Windows-hosted model server
If you’re running Mibyan inside WSL2 and your model server on the Windows host,http://localhost:<port> won’t work in WSL2’s default NAT networking mode. See WSL2 Networking above for the fix.
Tool calls appear as text instead of executing
The model outputs something like{"name": "web_search", "arguments": {...}} as a message instead of actually calling the tool.
Cause: Your server doesn’t have tool calling enabled, or the model doesn’t support it through the server’s tool calling implementation.
Model seems to forget context or give incoherent responses
Cause: Context window is too small. When the conversation exceeds the context limit, most servers silently drop older messages. Mibyan’s system prompt + tool schemas alone can use 4k–8k tokens. Diagnosis:127.0.0.1, LAN, Docker service names) says which window the server is serving and names the fix for any OpenAI-compatible server, not just Ollama: raise the server’s context (llama.cpp -c 64000, vLLM --max-model-len, Ollama OLLAMA_CONTEXT_LENGTH/Modelfile num_ctx) or set model.ollama_num_ctx in config.yaml to the window the server really serves (at least 64K). model.ollama_num_ctx is honoured on every local endpoint; only the automatic detection behind it uses Ollama’s /api/show.
”Context limit: 2048 tokens” at startup
Mibyan auto-detects context length from your server’s/v1/models endpoint. If the server reports a low value (or doesn’t report one at all), Mibyan uses the model’s declared limit which may be wrong.
Fix: Set it explicitly in config.yaml:
Responses get cut off mid-sentence
Possible causes:- Low output limit on the server — configure the server’s generation default (for example SGLang’s
--default-max-tokens). Mibyan does not expose an output-token cap setting. Response length is distinct from the conversation’s context window (context_length). - Context exhaustion — The model filled its context window. Increase
model.context_lengthor enable context compression in Mibyan.
LiteLLM Proxy — Multi-Provider Gateway
LiteLLM is an OpenAI-compatible proxy that unifies 100+ LLM providers behind a single API. Best for: switching between providers without config changes, load balancing, fallback chains, budget controls.mibyan model → Custom endpoint → http://localhost:4000/v1.
Example litellm_config.yaml with fallback:
ClawRouter — Cost-Optimized Routing
ClawRouter by BlockRunAI is a local routing proxy that auto-selects models based on query complexity. It classifies requests across 14 dimensions and routes to the cheapest model that can handle the task. Payment is via USDC cryptocurrency (no API keys).mibyan model → Custom endpoint → http://localhost:8402/v1 → model name blockrun/auto.
Routing profiles:
ClawRouter requires a USDC-funded wallet on Base or Solana for payment. All requests route through BlockRun’s backend API. Run
npx @blockrun/clawrouter doctor to check wallet status.Other Compatible Providers
Any service with an OpenAI-compatible API works. Some popular options:
Configure any of these with
mibyan model → Custom endpoint, or in config.yaml:
Context Length Detection
Context windows and output limits are different
context_length is the total context window — the combined budget for input and output tokens (e.g. 200,000 for Claude Opus 4.6). Mibyan uses this to decide when to compress history and to validate API requests.Output limits govern a single generated response, not the conversation history.
Mibyan no longer reads model.max_tokens, mibyan_MAX_TOKENS, provider output-cap
settings, or model_overrides.*.*.max_output_tokens. Remove these legacy settings.
Custom OpenAI-compatible endpoints receive no automatic catalog-sized output cap.
Their server defaults apply; these can be lower than the model maximum.
A reply that degenerates into a repetition loop is still stopped: within about 130,000
characters of the loop starting (visible or reasoning text), Mibyan closes the stream
and ends the turn with a “Repetition Detected” notice, so an uncapped endpoint cannot keep a
looping model running.Native Anthropic Messages (including the native Anthropic Bedrock path) requires
max_tokens, so Mibyan supplies an internal value. Bedrock Converse is a separate
protocol: its optional inferenceConfig.maxTokens is omitted by default, which
AWS documents as the model maximum.
Internal bounded tasks and provider-specific protocol requirements remain implementation
details. Omission does not universally select a model’s maximum output.Set context_length when auto-detection gets the window size wrong.- Config override —
model.context_lengthin config.yaml (highest priority). This is an explicit pin: it always wins over provider metadata, so Mibyan labels it(pinned)wherever the window is shown (welcome banner,/model,/usage, the status bar) and logs one warning at startup when the pin disagrees with the window the provider is known to advertise. The pin is dropped automatically when you switch model, provider or base URL. - Custom provider per-model —
providers.<name>.models.<id>.context_length - Persistent cache — previously discovered values (survives restarts)
- Endpoint
/models— queries your server’s API (local/custom endpoints) - Anthropic
/v1/models— queries Anthropic’s API formax_input_tokens(API-key users only) - OpenRouter API — live model metadata from OpenRouter
- Nous Portal — suffix-matches Nous model IDs against OpenRouter metadata
- models.dev — community-maintained registry with provider-specific context lengths for 3800+ models across 100+ providers
- Fallback defaults — broad model family patterns (128K default)
claude-opus-4.6 is 1M on Anthropic direct but 128K on GitHub Copilot).
To set the context length explicitly, add context_length to your model config:
mibyan model will prompt for context length when configuring a custom endpoint. Leave it blank for auto-detection.
Named Custom Providers
If you work with multiple custom endpoints (e.g., a local dev server and a remote GPU server), you can define them as named custom providers under theproviders: dict in config.yaml, keyed by provider name:
api (the endpoint base URL — base_url/url are accepted aliases), name (optional display name; defaults to the dict key), key_env or inline api_key or key_cmd (see below), transport (chat_completions / anthropic_messages / codex_responses), default_model, models, context_length, discover_models, extra_body, extra_headers, session_affinity_header (name of a header that carries the conversation id, for session-aware proxies; off unless set), ssl_ca_cert / ssl_verify, catalog_provider (see below), and enabled: false to hide an entry without deleting it.
Command-minted credentials (key_cmd)
Vision, thinking, and native local-model capability probes materialize the same
callable credential used by chat before building authentication headers. They
reuse the command token cache without replacing the chat client’s callable.
If a command cannot mint a string token, these best-effort probes send no bearer
rather than an object representation or a lower-priority configured credential.
Native local-model probes remove inherited Authorization on a failed explicit
callable while retaining unrelated configured headers. Chat retains its normal
error handling.
Enterprise gateways often issue short-lived bearer tokens (SSO/OIDC brokers, cloud IAM, internal auth proxies) rather than static API keys, so a token copied into .env goes stale mid-session and requests start returning 401. key_cmd names a command that prints a token; Mibyan runs it and caches the result until shortly before expiry, so long sessions keep working with no restart:
databricks auth token, gcloud auth print-access-token, az account get-access-token, vault read, or Claude Code-style apiKeyHelper scripts.
The command must print only the token on stdout: either bare, or as JSON with an access_token field (expires_in is honored; absolute expiry/expiresOn ISO timestamps too). Multi-line output is rejected rather than guessed at. If no expiry is advertised, the token is re-minted on a bounded window.
Precedence: an explicit --api-key flag still wins; otherwise key_cmd beats a static api_key/key_env on the same entry. The minted credential applies to the main agent turn and to auxiliary tasks (title generation, compression, vision, embedding) alike.
Model discovery also honors key_cmd for both providers: and legacy
custom_providers: entries, including mibyan model setup. Helpers run only when
an authenticated live catalog probe is needed: disabled discovery and warm catalog
cache reads do not mint tokens. Catalogs are scoped to the command identity, so
rotating a bearer does not invalidate the catalog. Probe helpers use their own
short-lived token source, not the inference client’s token cache; minted bearers
are never saved to config.yaml. If a helper fails, discovery falls back to the
configured model without exposing the helper’s output.
Not to be confused with secrets.command, which runs a helper once at startup to populate env vars process-wide. Use that for a vault/keychain helper handing back many secrets; use key_cmd when one provider’s credential must be re-minted during a session.
Legacy formatOlder configs used a top-level
custom_providers: list instead. It still works — Mibyan reads both — and mibyan update auto-migrates it to the providers: dict (config v12). Field names differ slightly in the dict format: legacy model is default_model, and legacy api_mode is transport.codex_responses proxies. A custom entry with transport: codex_responses (a local Codex proxy, for example) resolves the context window of Codex OAuth models (gpt-6-astra, gpt-5.6-sol/-terra/-luna, gpt-5.5, …) from the Codex OAuth table — 272K for most slugs — not from the 1.05M direct-API catalog, so compression fires before the Codex backend’s limit and its 272K billing tier. The decision follows the transport, not the hostname; the same holds for openai-codex behind mibyan_CODEX_BASE_URL or model.base_url. A per-model models.<id>.context_length, an entry-level context_length, or model.context_length still wins; the opt-in -900k picker variants keep their verified 900K.
Reasoning effort on custom endpoints. The configured reasoning_effort (/reasoning max, agent.reasoning_effort) reaches a custom endpoint unchanged on both the chat_completions and the codex_responses transport — up to max; only the Mibyan-internal ultra is clamped to max. Two exceptions follow the host rather than the entry: a custom entry pointed at api.openai.com keeps OpenAI’s per-model ladder (max is a gpt-5.6-only level there), and an entry pointed at a provider whose profile publishes a per-model vocabulary (Ramp Router) is clamped to that catalog. An endpoint that rejects the level answers with an HTTP 400 instead of Mibyan silently downgrading it. When no effort is configured at all, chat_completions requests carry reasoning_effort: medium — the same default the Nous Portal and OpenRouter routes apply — rather than leaving the endpoint’s own default in charge (kimi-k3’s is max, about 3x the reasoning tokens of medium); models marked supports_reasoning: false in the catalog or model_overrides, and Ollama models without the thinking capability, keep the field off.
Some OpenAI-compatible endpoints need provider-specific request body fields. Add an extra_body map to the matching custom provider and Mibyan will merge it into each chat-completions request for that endpoint:
enable_thinking under chat_template_kwargs instead of as a top-level extra_body field:
content empty:
extra_body follows the provider everywhere: it is merged at agent construction, survives every gateway turn (including turns where /fast layers service_tier/speed overrides on top — those merge over your extra_body rather than replacing it), and is re-derived on /model switches — switching to a named custom provider applies its extra_body, and switching away clears it so it never leaks to another provider.
The mibyan model → Custom Endpoint wizard now prompts for the API mode explicitly and persists your answer to config.yaml (as transport on the provider entry). URL-based auto-detection (e.g. /anthropic paths → anthropic_messages) still happens as a fallback when the field is left blank.
Native vision for custom-provider models. If your custom endpoint serves a vision-capable model that isn’t in models.dev, set model.supports_vision: true so Mibyan routes attached images natively (as image_url parts) instead of pre-processing them through vision_analyze. Single knob — no need to also set agent.image_input_mode: native.
providers.<name>.models.<id>.supports_vision) and accepts standard YAML booleans (true/false/yes/no/on/off/1/0).
A model_overrides entry that only corrects metadata (for example context_window) for a model the catalog does not know leaves vision and reasoning capability and the output-token limit unknown — vision_analyze, video_analyze and the reasoning-effort picker stay available. Only an explicit supports_vision: false / supports_reasoning: false in the override marks the model as text-only or non-reasoning.
Inheriting a catalogued vendor’s metadata (catalog_provider). When a named custom provider (a gateway, proxy or reseller) serves models that Mibyan already knows under a built-in provider, point the entry at that vendor and its models inherit the catalogued context window, output limit, vision and reasoning flags — no model_overrides needed:
catalog_provider accepts a Mibyan provider id (deepseek, anthropic, openai, …) or a models.dev id. It affects metadata lookups only — requests still go to your api URL with your credentials — and an explicit model_overrides entry for the same model still wins.
Switch between them mid-session with the triple syntax:
mibyan model menu.
Cookbook: Together AI, Groq, Perplexity
The cloud providers listed in Other Compatible Providers all speak OpenAI’s REST dialect, so they wire up the same way under theproviders: dict. Three worked recipes follow. Each drops into ~/.mibyan/config.yaml and the matching API key goes in ~/.mibyan/.env.
Together AI
Hosts open-weight models (Llama, MiniMax, Gemma, DeepSeek, Qwen) at prices significantly below first-party APIs. Good default for multi-model fleets./v1/models endpoint works, so mibyan model can auto-discover available models.
Groq
Ultra-fast inference (~500 tok/s on Llama-3.3-70B). Small catalog but strong for latency-sensitive interactive use.Perplexity
Useful when you want a model that does live web search and citation automatically. Strict about which models are available — check perplexity.ai/settings/api for the current list.api: https://api.perplexity.ai/v1 with api_mode: codex_responses) reserves the function names web_search, search_files, fetch_url, people_search and finance_search for its own built-in tools. Mibyan renames its client tools of the same name to mibyan_<name> on the wire and maps them back before dispatch, for the main agent loop and auxiliary calls (title generation, compression, MoA aggregation) alike — the same treatment OpenCode’s /v1/responses endpoints get.
Multiple providers in one config
The three recipes compose — use all of them together and switch per turn with/model custom:<name>:<model>:
Choosing the Right Setup
Optional API Keys
Self-Hosting Firecrawl
By default, Mibyan uses the Firecrawl cloud API for web search and scraping. If you prefer to run Firecrawl locally, you can point Mibyan at a self-hosted instance instead. See Firecrawl’s SELF_HOST.md for complete setup instructions. What you get: No API key required, no rate limits, no per-page costs, full data sovereignty. What you lose: The cloud version uses Firecrawl’s proprietary “Fire-engine” for advanced anti-bot bypassing (Cloudflare, CAPTCHAs, IP rotation). Self-hosted uses basic fetch + Playwright, so some protected sites may fail. Search uses DuckDuckGo instead of Google. Setup:-
Clone and start the Firecrawl Docker stack (5 containers: API, Playwright, Redis, RabbitMQ, PostgreSQL — requires ~4-8 GB RAM):
-
Point Mibyan at your instance (no API key needed):
FIRECRAWL_API_KEY and FIRECRAWL_API_URL if your self-hosted instance has authentication enabled.
OpenRouter Provider Routing
When using OpenRouter, you can control how requests are routed across providers. Add aprovider_routing section to ~/.mibyan/config.yaml:
:nitro to any model name for throughput sorting (e.g., anthropic/claude-sonnet-4:nitro), or :floor for price sorting. Per-model details: Provider Routing.
OpenRouter Pareto Code Router
OpenRouter ships an experimental coding-model router atopenrouter/pareto-code that auto-routes requests to the cheapest model meeting a coding-quality bar (ranked by Artificial Analysis). Pick this model and tune the min_coding_score knob in ~/.mibyan/config.yaml:
min_coding_scoreis only sent whenmodel.modelisopenrouter/pareto-code. On any other model the value is a no-op.- Set to empty string (or remove the line) to let OpenRouter pick the strongest available coder — its documented behavior when the plugins block is omitted.
- Selection is deterministic per score on a given day, but the actual model chosen can shift as the Pareto frontier moves (new models, benchmark updates).
- See OpenRouter’s Pareto Router docs for the full router behavior.
- To use the Pareto Code router for a specific auxiliary task (compression, vision, etc.) instead of the main agent, set
extra_body.pluginsunder that task — see Auxiliary Models → OpenRouter routing & Pareto Code for auxiliary tasks.
Fallback Providers
Configure a chain of backup providers Mibyan tries in order when the primary model fails (rate limits, server errors, auth failures). The canonical format is a top-levelfallback_providers: list:
providers.<name> block (provider: my-relay or provider: custom:my-relay) inherits that block’s transport / api_mode when the entry sets none, so a Responses-only or Anthropic-Messages relay keeps its declared wire on fallback. Set api_mode on the entry to override it.
The legacy single-pair fallback_model: dict is still accepted for back-compat:
openrouter, nous, novita, openai-codex, copilot, copilot-acp, anthropic, gemini, qwen-oauth, huggingface, zai, kimi-coding, kimi-coding-cn, minimax, minimax-cn, minimax-oauth, deepseek, nvidia, xai, xai-oauth, ollama-cloud, bedrock, ai-gateway, azure-foundry, opencode-zen, opencode-go, commandcode, commandcode-anthropic, kilocode, xiaomi, arcee, gmi, actual, stepfun, lmstudio, alibaba, alibaba-coding-plan, tencent-tokenhub, tencent-tokenplan, nebius-token-factory, router, custom.
See Also
- Configuration — General configuration (directory structure, config precedence, terminal backends, memory, compression, and more)
- Environment Variables — Complete reference of all environment variables

