> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mibyanai.com/llms.txt
> Use this file to discover all available pages before exploring further.

# LLM and Model Providers

This page covers setting up inference providers for Mibyan — from cloud APIs like OpenRouter and Anthropic, to self-hosted endpoints like Ollama and vLLM, to advanced routing and fallback configurations. You need at least one provider configured to use Mibyan.

## Inference Providers

You need at least one way to connect to an LLM. Use `mibyan model` to switch providers and models interactively, or configure directly:

| Provider | Setup |
| - | - |
| **Nous Portal** | `mibyan model` (OAuth, subscription-based) |
| **OpenAI Codex** | `mibyan model` → **ChatGPT or Codex Subscription** (ChatGPT OAuth, uses Codex models) |
| **GitHub Copilot** | `mibyan model` (OAuth device code flow, `COPILOT_GITHUB_TOKEN`, `GH_TOKEN`, or `gh auth token`) |
| **GitHub Copilot ACP** | `mibyan model` (spawns local `copilot --acp --stdio`) |
| **Anthropic** | `mibyan model` (Claude Max + extra usage credits via OAuth; also supports Anthropic API key or manual setup-token — see note below) |
| **OpenRouter** | `OPENROUTER_API_KEY` in `~/.mibyan/.env`, or `mibyan auth add openrouter --type oauth` (browser login via OpenRouter's PKCE flow; stores a key in the credential pool) |
| **Ramp Router** | `RAMP_ROUTER_API_KEY` in `~/.mibyan/.env` (provider: `router`; aliases: `ramp-router`, `ramp`, `router.com`; Responses-native gateway, live account-scoped catalog) |
| **Fireworks AI** | `FIREWORKS_API_KEY` in `~/.mibyan/.env` (provider: `fireworks`; aliases: `fireworks-ai`, `fw`) |
| **NovitaAI** | `NOVITA_API_KEY` in `~/.mibyan/.env` (provider: `novita`, 200+ models, Model API, Agent Sandbox, GPU Cloud) |
| **AI Gateway** | `AI_GATEWAY_API_KEY` in `~/.mibyan/.env` (provider: `ai-gateway`) |
| **z.ai / GLM** | `GLM_API_KEY` in `~/.mibyan/.env` (provider: `zai`) |
| **Kimi / Moonshot** | `KIMI_API_KEY` in `~/.mibyan/.env` (provider: `kimi-coding`) |
| **Kimi / Moonshot (China)** | `KIMI_CN_API_KEY` in `~/.mibyan/.env` (provider: `kimi-coding-cn`; aliases: `kimi-cn`, `moonshot-cn`) |
| **Arcee AI** | `ARCEEAI_API_KEY` in `~/.mibyan/.env` (provider: `arcee`; aliases: `arcee-ai`, `arceeai`) |
| **GMI Cloud** | `GMI_API_KEY` in `~/.mibyan/.env` (provider: `gmi`; aliases: `gmi-cloud`, `gmicloud`) |
| **Nebius Token Factory** | `NEBIUS_API_KEY` in `~/.mibyan/.env` (provider: `nebius-token-factory`; aliases: `nebius`, `nebius-tf`, `tokenfactory`) |
| **Actual Computer** | `ACTUAL_API_KEY` in `~/.mibyan/.env` for the hosted relay; set `model.base_url` in `config.yaml` for a local daemon (no key needed on loopback). Provider: `actual`; aliases: `actual-computer`, `actualcomputer`, `aci`. |
| **MiniMax** | `MINIMAX_API_KEY` in `~/.mibyan/.env` (provider: `minimax`) |
| **MiniMax China** | `MINIMAX_CN_API_KEY` in `~/.mibyan/.env` (provider: `minimax-cn`) |
| **xAI (Grok) — Responses API** | `XAI_API_KEY` in `~/.mibyan/.env` (provider: `xai`) |
| **xAI Grok OAuth (SuperGrok)** | `mibyan model` → "xAI Grok OAuth (SuperGrok / Premium+)" — browser login, no API key. See [guide](/desktop/guides/xai-grok-oauth) |
| **Qwen Cloud (Alibaba DashScope)** | `DASHSCOPE_API_KEY` in `~/.mibyan/.env` (provider: `alibaba`; mainland-China endpoint: `alibaba-cn`) |
| **Alibaba Cloud (Coding Plan)** | `ALIBABA_CODING_PLAN_API_KEY` (falls back to `DASHSCOPE_API_KEY`) (provider: `alibaba-coding-plan`, alias: `alibaba_coding`; mainland-China endpoint: `alibaba-coding-plan-cn` with `ALIBABA_CODING_PLAN_CN_API_KEY`, falling back to the shared keys) — separate billing SKU, different endpoint |
| **Alibaba Cloud (Token Plan)** | `ALIBABA_TOKEN_PLAN_API_KEY` in `~/.mibyan/.env` (provider: `alibaba-token-plan`; mainland-China endpoint: `alibaba-token-plan-cn` with `ALIBABA_TOKEN_PLAN_CN_API_KEY`, falling back to the shared key) — Model Studio flat-token tier |
| **Kilo Code** | `KILOCODE_API_KEY` in `~/.mibyan/.env` (provider: `kilocode`) |
| **Xiaomi MiMo** | `XIAOMI_API_KEY` in `~/.mibyan/.env` (provider: `xiaomi`, aliases: `mimo`, `xiaomi-mimo`) |
| **Tencent TokenHub** | `TOKENHUB_API_KEY` in `~/.mibyan/.env` (provider: `tencent-tokenhub`, aliases: `tencent`, `tokenhub`, `tencentmaas`) |
| **Tencent TokenPlan** | `TOKENPLAN_API_KEY` in `~/.mibyan/.env` (provider: `tencent-tokenplan`, aliases: `tokenplan`, `tencent-lkeap`; Anthropic Messages endpoint) |
| **OpenCode Zen** | `OPENCODE_ZEN_API_KEY` in `~/.mibyan/.env` (provider: `opencode-zen`) |
| **CommandCode** | `COMMANDCODE_API_KEY` in `~/.mibyan/.env` (provider: `commandcode`, alias: `commandcode-chat`; Claude models via `commandcode-anthropic`, alias: `commandcode-claude`). Works with GOAT/Pro/Max/Provider plans (not the \$1 Go plan — no API access). |
| **OpenCode Go** | `OPENCODE_GO_API_KEY` in `~/.mibyan/.env` (provider: `opencode-go`) |
| **DeepSeek** | `DEEPSEEK_API_KEY` in `~/.mibyan/.env` (provider: `deepseek`) |
| **Hugging Face** | `HF_TOKEN` in `~/.mibyan/.env` (provider: `huggingface`, aliases: `hf`) |
| **Google / Gemini** | `GOOGLE_API_KEY` (or `GEMINI_API_KEY`) in `~/.mibyan/.env` (provider: `gemini`) |
| **Google Vertex AI** | `mibyan model` → "Google Vertex AI" (provider: `vertex`; OAuth2 via service-account JSON or ADC, GCP billing) |
| **OpenAI API (direct)** | `OPENAI_API_KEY` in `~/.mibyan/.env` (provider: `openai-api`, optional `OPENAI_BASE_URL`) |
| **Azure AI Foundry** | `mibyan model` → "Azure AI Foundry" (provider: `azure-foundry`; uses Azure OpenAI / Foundry endpoint and key) |
| **AWS Bedrock** | `mibyan model` → "AWS Bedrock" (provider: `bedrock`; standard AWS credentials chain via boto3) |
| **NVIDIA Build** | `NVIDIA_API_KEY` in `~/.mibyan/.env` (provider: `nvidia`; NIM-hosted models on build.nvidia.com) |
| **Ollama Cloud** | `mibyan model` → "Ollama Cloud" (provider: `ollama-cloud`; cloud-hosted Ollama API) |
| **Qwen OAuth** | `mibyan model` → "Qwen OAuth" (provider: `qwen-oauth`; browser PKCE login) |
| **MiniMax OAuth** | `mibyan model` → "MiniMax (OAuth)" (provider: `minimax-oauth`; browser PKCE login) |
| **StepFun** | `STEPFUN_API_KEY` in `~/.mibyan/.env` (provider: `stepfun`) |
| **LM Studio** | `mibyan model` → "LM Studio" (provider: `lmstudio`, optional `LM_API_KEY`) |
| **Custom Endpoint** | `mibyan model` → choose "Custom endpoint" (saved in `config.yaml`) |

Both built-in OpenCode providers send an opaque, per-conversation `x-opencode-session` header on every request (main turns on every transport plus auxiliary calls such as compression, titles, approval checks, skills-hub lookups and `/btw` side questions — including the ones that run in the background after the turn has ended; headless Kanban `specify`/`decompose` and dashboard estimate calls use a per-task key; one-shots with no live session at all, such as Desktop commit-message generation from the review panel, send a fresh ephemeral key). OpenCode uses it to pin a conversation to one backend so its prompt cache stays warm; the value is derived from the Mibyan session id (or the Kanban task id) and carries no personal data.

The two built-in OpenCode providers each pin their own relay on `opencode.ai` (`opencode-zen` → `/zen/v1`, `opencode-go` → `/zen/go/v1`). A `model.base_url` left behind by the other relay is healed to the selected provider's relay, and the model you pick (`-m`, `/model`, a fallback entry or a channel override) decides which relay is used — so switching from a Zen model to a Go-only one never sends the request to Zen. The same per-model routing is re-applied when a session is resumed (`mibyan --resume`, `/resume`, the TUI and desktop resume paths): a wire format or relay URL recorded while the session ran a different OpenCode model never carries over to the model the session is reopened on. OpenCode models whose id carries a `-vision` marker (for example `deepseek-v4-flash-vision-exp`) are treated as vision-capable even before models.dev lists them, so `agent.image_input_mode: auto` attaches images natively without a `supports_vision` override. A custom provider you define under `providers:` whose name extends a family slug (for example `opencode-go-bridge`) still gets the family's per-model API-mode routing and `/v1` handling, but its `base_url` is taken as declared: name it after the relay it actually points at. Auxiliary tasks (`auxiliary.compression`, titles, vision, MoA) pointed at an OpenCode provider follow the same per-model table, so a Responses-only model such as `gpt-5.6-luna` or an Anthropic-wire one such as `minimax-m2.5` works there exactly as it does for the main conversation.

For the official API-key path, see the dedicated [Google Gemini guide](/desktop/guides/google-gemini).

<Tip>
  **Model key alias**

  In the `model:` config section, you can use either `default:` or `model:` as the key name for your model ID. Both `model: { default: my-model }` and `model: { model: my-model }` work identically.
</Tip>

### Nous Portal

[Nous Portal](https://portal.nousresearch.com) is Nous Research's unified subscription gateway and **the recommended way to run Mibyan**. One OAuth login covers 300+ frontier agentic models (Claude, GPT, Gemini, DeepSeek, Qwen, Kimi, GLM, MiniMax, Grok, ...) plus the Tool Gateway (web search, image generation, TTS, browser automation) — billed against your Nous subscription instead of separate per-provider accounts.

```bash theme={null}
mibyan setup --portal     # fresh install — OAuth + provider + gateway in one command
mibyan model              # existing install — pick "Nous Portal" from the list
mibyan portal info        # inspect login + routing at any time
```

Don't have a subscription yet? Get one at [portal.nousresearch.com/manage-subscription](https://portal.nousresearch.com/manage-subscription).

**For full details:** see the dedicated Nous Portal integration page (what's in the subscription, model catalog, troubleshooting) and the step-by-step Run Mibyan with Nous Portal guide.

**Client identification.** Every Portal request from Mibyan carries a `client=mibyan-client-v<version>` tag (e.g. `client=mibyan-client-v0.13.0`) auto-aligned to your installed release. This is sent on all Portal pathways — main chat loop, auxiliary calls, compression summarizer, web extraction — and lets Portal-side telemetry distinguish Mibyan traffic from other clients. No config required; the tag updates automatically when you `mibyan update`.

**JWT auth (automatic).** Mibyan prefers scoped `inference:invoke` JWTs for Portal requests with the legacy opaque session-key path as a fallback. No configuration is required — credentials are managed by the OAuth flow and rotate transparently. Revoked refresh tokens are quarantined to avoid replay loops.

<Info>
  **Codex Note**

  The OpenAI Codex provider authenticates via device code by default (open a URL, enter a code). Organizations that disable the device-code grant can opt in to the browser authorization-code + PKCE flow instead: `mibyan auth add openai-codex --browser` (one login) or `auth.codex_login_flow: browser` in `config.yaml` (every Codex login, including `mibyan model`). That flow listens on `http://localhost:1455/auth/callback` — the redirect URI registered for the Codex client, so the port is fixed; if it is already taken (a Codex CLI sign-in in progress) Mibyan says so and falls back to device code. Over SSH the listener needs a tunnel (`ssh -N -L 1455:127.0.0.1:1455 user@host`, see [OAuth over SSH](/desktop/guides/oauth-over-ssh)). Mibyan stores the resulting credentials in its own auth store under `~/.mibyan/auth.json` and can import existing Codex CLI credentials from `~/.codex/auth.json` when present. No Codex CLI installation is required. Automatic adoption of the Codex CLI login (when Mibyan' own refresh fails) is controlled by `auth.adopt_external_logins` — see [Borrowed CLI logins](/desktop/user-guide/security#borrowed-cli-logins).

  If a token refresh fails with a terminal error (HTTP 4xx, `invalid_grant`, revoked grant, etc.), Mibyan marks the refresh token as dead and stops replaying it so you don't see a flood of identical auth failures. The next request surfaces a typed re-auth message instead. Run `mibyan auth add openai-codex` (or `mibyan model` → **ChatGPT or Codex Subscription**) to start a fresh login (device code, or `--browser` for the loopback PKCE flow); the quarantine clears on the next successful exchange.

  Device login can fail with `[SSL: UNEXPECTED_EOF_WHILE_READING]` or a TLS handshake timeout on Python/OpenSSL 3.5+ when a middlebox rejects post-quantum groups such as X25519MLKEM768 (curl may still work). A one-off dropped connection is not fatal: while waiting for your browser approval Mibyan keeps polling through up to six consecutive transport errors (and retries the device-code request and token exchange twice) before giving up, so only a persistently broken network surfaces this error. Mibyan does not change default TLS policy. Point `OPENSSL_CONF` at a config that restricts `Groups` to classic curves before running `mibyan model`, or diagnose with TLS 1.2:

  ```ini theme={null}
  openssl_conf = openssl_init

  [openssl_init]
  ssl_conf = ssl_sect

  [ssl_sect]
  system_default = system_default_sect

  [system_default_sect]
  Groups = x25519:secp256r1:secp384r1:x448
  ```
</Info>

<Warning>
  Even when using Nous Portal, Codex, or a custom endpoint, some tools (vision, web summarization, MoA) use a separate "auxiliary" model. By default (`auxiliary.*.provider: "auto"`), Mibyan routes these tasks to your **main chat model** — the same model you picked in `mibyan model`. You can override each task individually to route it to a cheaper/faster model (e.g. Gemini Flash on OpenRouter) — see [Auxiliary Models](/desktop/user-guide/configuration#auxiliary-models).
</Warning>

<Tip>
  **Nous Tool Gateway**

  Paid Nous Portal subscribers also get access to the **Tool Gateway** — web search, image generation, TTS, and browser automation routed through your subscription. No extra API keys needed. On a fresh install, `mibyan setup --portal` logs you in, sets Nous as your provider, and turns the gateway on in one command. Existing users can enable it from `mibyan model` or per-tool from `mibyan tools`. Inspect routing at any time with `mibyan portal info`.
</Tip>

### Two Commands for Model Management

Mibyan has **two** model commands that serve different purposes:

| Command | Where to run | What it does |
| - | - | - |
| **`mibyan model`** | Your terminal (outside any session) | Full setup wizard — add providers, run OAuth, enter API keys, configure endpoints |
| **`/model`** | Inside a Mibyan chat session | Quick switch between **already-configured** providers and models |

If you're trying to switch to a provider you haven't set up yet (e.g. you only have OpenRouter configured and want to use Anthropic), you need `mibyan model`, not `/model`. Exit your session first (`Ctrl+C` or `/quit`), run `mibyan model`, complete the provider setup, then start a new session.

### Subscription plans: what your plan pays for

Several providers let you sign in to Mibyan with a **consumer subscription** (Claude Max, ChatGPT, SuperGrok / X Premium+, …) instead of an API key. What that subscription actually pays for — and what it doesn't — differs per provider, and it's the single most common source of billing surprises. The table below is the short version; each provider's own section has the details.

> Cells marked *not currently documented* mean exactly that: Mibyan docs do not yet specify the behavior. Don't assume — check your provider's billing dashboard, and treat these as open questions.

| Plan / path | Can Mibyan use it? | What gets consumed | What does NOT get consumed | Common surprise |
| - | - | - | - | - |
| **Anthropic — Claude Max + OAuth** | ✅ Yes — `mibyan model` → Anthropic OAuth. Requires Max **and** purchased extra usage credits | The **extra/overage credits** you've added on top of the Max plan | The **base Max plan allowance** (the usage included in Claude Code by default) | All Mibyan usage bills as "extra usage" even while your included Max allowance sits untouched |
| **Anthropic — Claude Pro** | ❌ No — Pro subscribers cannot use the OAuth path | Nothing (path unavailable) | Your Pro subscription | Pro looks like it should work; it doesn't. Use an `ANTHROPIC_API_KEY` instead (pay-per-token, independent of any Claude subscription) |
| **OpenAI Codex — ChatGPT plan OAuth** | ✅ Yes — `mibyan model` → **ChatGPT or Codex Subscription** (ChatGPT OAuth device-code login, uses Codex models) | *Not currently documented* | *Not currently documented* | Docs cover auth and token refresh only; plan-quota semantics are not yet documented |
| **xAI — SuperGrok / X Premium+ OAuth** | ✅ Yes — browser OAuth, no API key needed | Your **subscription quota** (documented explicitly for X Search: OAuth is preferred over an API key and "uses your subscription quota instead of API spend"). Inference quota semantics beyond that: *not currently documented* | `XAI_API_KEY` / pay-per-token API spend, when OAuth credentials are configured and preferred | `HTTP 403` after a successful login — xAI has restricted OAuth API access to specific SuperGrok tiers despite an active in-app subscription |
| **Google — Gemini consumer plan (Google AI Pro / Ultra)** | ❌ No documented path — the `gemini` provider is API-key only (`GOOGLE_API_KEY` / `GEMINI_API_KEY`); Vertex AI uses GCP billing | Your **API key's quota** (free tier or billing-enabled Google Cloud project) — *consumer-plan consumption not currently documented* | *Not currently documented* | Free-tier keys can be exhausted after a handful of agent turns, because Mibyan may make several model calls per user turn |

**Anthropic.** The OAuth path routes as Claude Code against your Anthropic account and **only works on a Claude Max plan with purchased extra usage credits** — the base Max allowance is never consumed by Mibyan, only the extra/overage credits on top. Claude Pro subscribers cannot use this path; the supported alternative is an `ANTHROPIC_API_KEY`, billed pay-per-token against that key's organization at standard API pricing. See [Anthropic (Native)](#anthropic-native) below.

**OpenAI Codex.** Mibyan authenticates via ChatGPT device-code OAuth, stores credentials in `~/.mibyan/auth.json`, and can import existing Codex CLI credentials from `~/.codex/auth.json`. Which ChatGPT plan tiers are eligible, and how Mibyan usage counts against your plan's Codex limits, are **not currently documented** — the Codex note under [Nous Portal](#nous-portal) covers authentication and token-refresh behavior only.

**xAI (SuperGrok / X Premium+).** Browser OAuth works with either an active SuperGrok subscription or an X Premium+ subscription on the linked X account, and the same bearer token is reused by direct-to-xAI tools (TTS, image gen, video gen, transcription, X Search). If inference returns `HTTP 403` after a successful login, that's a tier/entitlement restriction on xAI's side, not a stale token — the workaround is switching to an `XAI_API_KEY`. See [xAI (Grok)](#xai-grok--responses-api--prompt-caching) below and the [xAI Grok OAuth guide](/desktop/guides/xai-grok-oauth).

**Google Gemini.** There is currently no way to sign in to Mibyan with a consumer Gemini subscription — the `gemini` provider takes an API key, and [Google Vertex AI](#google-vertex-ai) bills to your GCP project. A billing-enabled Google Cloud project is recommended for agent use; free-tier quotas are too small for long-running agent sessions. See the [Google Gemini guide](/desktop/guides/google-gemini).

<Tip>
  **One subscription instead of five**

  If you'd rather not track per-provider plan semantics at all, [Nous Portal](#nous-portal) covers 300+ models under a single subscription with one OAuth login.
</Tip>

### Anthropic (Native)

Use Claude models directly through the Anthropic API — no OpenRouter proxy needed. Supports three auth methods:

When no explicit environment credential is selected, Mibyan-owned OAuth grants
in the credential pool take precedence over a borrowed Claude Code login. The
borrowed login remains the fallback when no owned OAuth grant is available —
unless `auth.adopt_external_logins: false` is set, in which case Mibyan never
reads or refreshes Claude Code's credentials (see
[Borrowed CLI logins](/desktop/user-guide/security#borrowed-cli-logins)).
Auxiliary authentication recovery refreshes the credential used by the failed
request, not an unrelated ambient login; rotating a borrowed login can otherwise
invalidate its owner's refresh token.

<Warning>
  **Requires Claude Max "extra usage" credits**

  When you authenticate via `mibyan model` → Anthropic OAuth (or via `mibyan auth add anthropic --type oauth`), Mibyan routes as Claude Code against your Anthropic account. **It only works if you're on a Claude Max plan and have purchased extra usage credits.** The base Max plan allowance (the usage included in Claude Code by default) is not consumed by Mibyan — only the extra/overage credits you've added on top are. Claude Pro subscribers cannot use this path.

  If you don't have Max + extra credits, use an `ANTHROPIC_API_KEY` instead — requests are billed pay-per-token against that key's organization (standard API pricing, independent of any Claude subscription).
</Warning>

```bash theme={null}
# With an API key (pay-per-token)
export ANTHROPIC_API_KEY=***
mibyan chat --provider anthropic --model claude-sonnet-4-6

# Preferred: authenticate through `mibyan model`
# Mibyan will use Claude Code's credential store directly when available
mibyan model

# Manual override with a setup-token (fallback / legacy)
export ANTHROPIC_TOKEN=***  # setup-token or manual OAuth token
mibyan chat --provider anthropic

# Auto-detect Claude Code credentials (if you already use Claude Code)
mibyan chat --provider anthropic  # reads Claude Code credential files automatically
```

When you choose Anthropic OAuth through `mibyan model`, Mibyan prefers Claude Code's own credential store over copying the token into `~/.mibyan/.env`. That keeps refreshable Claude credentials refreshable.

Or set it permanently:

```yaml theme={null}
model:
  provider: "anthropic"
  default: "claude-sonnet-4-6"
```

<Tip>
  **Aliases**

  `--provider claude` and `--provider claude-code` also work as shorthand for `--provider anthropic`.
</Tip>

### GitHub Copilot

Mibyan supports GitHub Copilot as a first-class provider with two modes:

**`copilot` — Direct Copilot API** (recommended). Uses your GitHub Copilot subscription to access GPT-5.x, Claude, Gemini, and other models through the Copilot API.

```bash theme={null}
mibyan chat --provider copilot --model gpt-5.4
```

**Authentication options** (checked in this order):

1. `COPILOT_GITHUB_TOKEN` environment variable
2. `GH_TOKEN` environment variable
3. `GITHUB_TOKEN` environment variable
4. `gh auth token` CLI fallback

If no token is found, `mibyan model` offers an **OAuth device code login** — the same flow used by the Copilot CLI and opencode.

<Warning>
  **Token types**

  The Copilot API does **not** support classic Personal Access Tokens (`ghp_*`). Supported token types:

  | Type | Prefix | How to get |
  | - | - | - |
  | OAuth token | `gho_` | `mibyan model` → GitHub Copilot → Login with GitHub |
  | Fine-grained PAT | `github_pat_` | GitHub Settings → Developer settings → Fine-grained tokens (needs **Copilot Requests** permission) |
  | GitHub App token | `ghu_` | Via GitHub App installation |

  If your `gh auth token` returns a `ghp_*` token, use `mibyan model` to authenticate via OAuth instead.
</Warning>

<Info>
  **Copilot auth behavior in Mibyan**

  Mibyan sends a supported GitHub token (`gho_*`, `github_pat_*`, or `ghu_*`) directly to `api.githubcopilot.com` and includes Copilot-specific headers (`Editor-Version`, `Copilot-Integration-Id`, `Openai-Intent`, `x-initiator`).

  On HTTP 401, Mibyan now performs a one-shot credential recovery before fallback:

  1. Re-resolve token via the normal priority chain (`COPILOT_GITHUB_TOKEN` → `GH_TOKEN` → `GITHUB_TOKEN` → `gh auth token`)
  2. Rebuild the shared OpenAI client with refreshed headers
  3. Retry the request once

  Some older community proxies use `api.github.com/copilot_internal/v2/token` exchange flows. That endpoint can be unavailable for some account types (returns 404). Mibyan therefore keeps direct-token auth as the primary path and relies on runtime credential refresh + retry for robustness.
</Info>

**API routing**: GPT-5+ models (except `gpt-5-mini`) automatically use the Responses API. All other models (GPT-4o, Claude, Gemini, etc.) use Chat Completions. Models are auto-detected from the live Copilot catalog.

**`copilot-acp` — Copilot ACP agent backend**. Spawns the local Copilot CLI as a subprocess:

```bash theme={null}
mibyan chat --provider copilot-acp --model copilot-acp
# Requires the GitHub Copilot CLI in PATH and an existing `copilot login` session
```

**Permanent config:**

```yaml theme={null}
model:
  provider: "copilot"
  default: "gpt-5.4"
```

| Environment variable | Description |
| - | - |
| `COPILOT_GITHUB_TOKEN` | GitHub token for Copilot API (first priority) |
| `mibyan_COPILOT_ACP_COMMAND` | Override the Copilot CLI binary path (default: `copilot`) |
| `mibyan_COPILOT_ACP_ARGS` | Override ACP args (default: `--acp --stdio`) |

### First-Class API-Key Providers

These providers have built-in support with dedicated provider IDs. Set the API key and use `--provider` to select:

```bash theme={null}
# Fireworks AI
mibyan chat --provider fireworks --model accounts/fireworks/models/kimi-k2p6
# Requires: FIREWORKS_API_KEY in ~/.mibyan/.env

# NovitaAI Model API
mibyan chat --provider novita --model moonshotai/kimi-k2.5
# Requires: NOVITA_API_KEY in ~/.mibyan/.env

# Ramp Router (model IDs come from your account's live catalog)
mibyan chat --provider router --model gpt-5.4-mini
# Requires: RAMP_ROUTER_API_KEY in ~/.mibyan/.env

# z.ai / ZhipuAI GLM
mibyan chat --provider zai --model glm-5
# Requires: GLM_API_KEY in ~/.mibyan/.env

# Kimi / Moonshot AI (international: api.moonshot.ai)
mibyan chat --provider kimi-coding --model kimi-for-coding
# Requires: KIMI_API_KEY in ~/.mibyan/.env

# Kimi / Moonshot AI (China: api.moonshot.cn)
mibyan chat --provider kimi-coding-cn --model kimi-k2.5
# Requires: KIMI_CN_API_KEY in ~/.mibyan/.env

# MiniMax (global endpoint)
mibyan chat --provider minimax --model MiniMax-M2.7
# Requires: MINIMAX_API_KEY in ~/.mibyan/.env

# MiniMax (China endpoint)
mibyan chat --provider minimax-cn --model MiniMax-M2.7
# Requires: MINIMAX_CN_API_KEY in ~/.mibyan/.env

# Qwen Cloud / DashScope (Qwen models)
mibyan chat --provider alibaba --model qwen3.5-plus
# Requires: DASHSCOPE_API_KEY in ~/.mibyan/.env

# Xiaomi MiMo
mibyan chat --provider xiaomi --model mimo-v2-pro
# Requires: XIAOMI_API_KEY in ~/.mibyan/.env

# Tencent TokenHub (Hy4 preview)
mibyan chat --provider tencent-tokenhub --model hy4-preview
# Requires: TOKENHUB_API_KEY in ~/.mibyan/.env

# Tencent TokenPlan (Hy4 preview via Anthropic Messages endpoint)
mibyan chat --provider tencent-tokenplan --model hy4-preview
# Requires: TOKENPLAN_API_KEY in ~/.mibyan/.env

# Arcee AI (Trinity models)
mibyan chat --provider arcee --model trinity-large-thinking
# Requires: ARCEEAI_API_KEY in ~/.mibyan/.env

# Meta Model API (Muse Spark family)
mibyan chat --provider meta-ai --model muse-spark-1.2
# Requires: MODEL_API_KEY in ~/.mibyan/.env

# GMI Cloud
# Use the exact model ID returned by GMI's /v1/models endpoint.
mibyan chat --provider gmi --model zai-org/GLM-5.1-FP8
# Requires: GMI_API_KEY in ~/.mibyan/.env

# Nebius Token Factory
mibyan chat --provider nebius --model deepseek-ai/DeepSeek-V4-Pro
# Requires: NEBIUS_API_KEY in ~/.mibyan/.env
```

Fireworks uses its native slash-form catalog IDs, such as `accounts/fireworks/models/kimi-k2p6`. Run `mibyan model`, choose **Fireworks AI**, and select from the live catalog or enter another Fireworks model ID. The default endpoint is `https://api.fireworks.ai/inference/v1`; configure a different endpoint through `model.base_url` in `config.yaml`, not `.env`.

Or set the provider permanently in `config.yaml`:

```yaml theme={null}
model:
  provider: "gmi"
  default: "zai-org/GLM-5.1-FP8"
```

Base URLs can be overridden with `NOVITA_BASE_URL`, `GLM_BASE_URL`, `KIMI_BASE_URL`, `MINIMAX_BASE_URL`, `MINIMAX_CN_BASE_URL`, `DASHSCOPE_BASE_URL`, `XIAOMI_BASE_URL`, `GMI_BASE_URL`, `META_BASE_URL`, or `TOKENHUB_BASE_URL` environment variables.

<Note>
  **Meta contributor tier**

  `muse-spark-1.2-contributor` and `muse-spark-1.3-contributor` are Meta's contributor tiers — Meta may train on your prompts and completions, so [interactive model selection asks for confirmation](/desktop/user-guide/configuring-models) before using either. For current pricing and rate limits, see [Meta Model API pricing and rate limits](https://dev.meta.ai/docs/pricing-rate-limits/). Use the standard `muse-spark-1.2` / `muse-spark-1.3` (no training) for confidential work.
</Note>

<Note>
  **Z.AI Endpoint Auto-Detection**

  When using the Z.AI / GLM provider, Mibyan automatically probes multiple endpoints (global, China, coding variants) to find one that accepts your API key. You don't need to set `GLM_BASE_URL` manually — the working endpoint is detected and cached automatically.
</Note>

### xAI (Grok) — Responses API + Prompt Caching

xAI is wired through the Responses API (`codex_responses` transport) for automatic reasoning support on Grok 4 models — no `reasoning_effort` parameter needed, the server reasons by default. Set `XAI_API_KEY` in `~/.mibyan/.env` and pick xAI in `mibyan model`, or drop `grok` as a shortcut into `/model grok-4-fast-reasoning`.

SuperGrok and X Premium+ subscribers can sign in with browser OAuth instead of using an API key — pick **xAI Grok OAuth (SuperGrok / Premium+)** in `mibyan model`, or run `mibyan auth add xai-oauth`. The same OAuth bearer token is automatically reused by direct-to-xAI tools (TTS, image gen, video gen, transcription). See the [xAI Grok OAuth guide](/desktop/guides/xai-grok-oauth) for the full flow — and if Mibyan runs on a remote host, also see [OAuth over SSH / Remote Hosts](/desktop/guides/oauth-over-ssh) for the required `ssh -L` tunnel.

When using xAI as a provider (any base URL containing `x.ai`), Mibyan automatically enables prompt caching by sending the `x-grok-conv-id` header with every API request. This routes requests to the same server within a conversation session, allowing xAI's infrastructure to reuse cached system prompts and conversation history.

No configuration is needed — caching activates automatically when an xAI endpoint is detected and a session ID is available. This reduces latency and cost for multi-turn conversations.

xAI also ships a dedicated TTS endpoint (`/v1/tts`). Select **xAI TTS** in `mibyan tools` → Voice & TTS, or see the [Voice & TTS](/desktop/user-guide/features/tts#text-to-speech) page for config.

**Retired xAI model migration (May 15, 2026):** xAI is retiring `grok-4*`, `grok-3`, `grok-code-fast-1`, and `grok-imagine-image-pro` on 2026-05-15. `mibyan doctor` and `mibyan chat` startup both detect any config still pointing at a retired ref and print the recommended replacement. Use `mibyan migrate xai` for a one-shot config rewrite — dry-run by default, add `--apply` to write changes (a timestamped copy of the previous config lands in `backups/config/` first).

```bash theme={null}
mibyan migrate xai          # preview replacements
mibyan migrate xai --apply  # rewrite ~/.mibyan/config.yaml in place
```

**xAI Web Search backend.** When the [Web Search](/desktop/user-guide/features/web-search) toolset is enabled, `web.backend: xai` routes search through xAI's hosted search endpoint using the same `XAI_API_KEY` / OAuth credentials. No additional setup required if xAI is already configured as a provider.

### NovitaAI

[NovitaAI](https://novita.ai) is the AI-native cloud for builders and agents. Its three product lines are Model API for 200+ models, Agent Sandbox for building and running AI agents, and GPU Cloud for scalable compute, all available from one platform.

```bash theme={null}
# Use any available model
mibyan chat --provider novita --model moonshotai/kimi-k2.5
# Requires: NOVITA_API_KEY in ~/.mibyan/.env

# Short alias
mibyan chat --provider novita-ai --model deepseek/deepseek-v3-0324
```

Or set it permanently in `config.yaml`:

```yaml theme={null}
model:
  provider: "novita"
  default: "moonshotai/kimi-k2.5"
  base_url: "https://api.novita.ai/openai/v1"
```

Get your API key at [novita.ai/settings/key-management](https://novita.ai/settings/key-management). The base URL can be overridden with `NOVITA_BASE_URL`.

### Ollama Cloud — Managed Ollama Models, OAuth + API Key

[Ollama Cloud](https://ollama.com/cloud) hosts the same open-weight catalog as local Ollama but without the GPU requirement. Pick it in `mibyan model` as **Ollama Cloud**, paste your API key from [ollama.com/settings/keys](https://ollama.com/settings/keys), and Mibyan auto-discovers the available models.

```bash theme={null}
mibyan model
# → pick "Ollama Cloud"
# → paste your OLLAMA_API_KEY
# → select from discovered models (gpt-oss:120b, glm-4.6:cloud, qwen3-coder:480b-cloud, etc.)
```

Or `config.yaml` directly:

```yaml theme={null}
model:
  provider: "ollama-cloud"
  default: "gpt-oss:120b"
```

The model catalog is fetched dynamically from `ollama.com/v1/models` and cached for one hour. `model:tag` notation (e.g. `qwen3-coder:480b-cloud`) is preserved through normalization — don't use dashes.

<Tip>
  **Ollama Cloud vs local Ollama**

  Both speak the same OpenAI-compatible API. Cloud is a first-class provider (`--provider ollama-cloud`, `OLLAMA_API_KEY`); local Ollama is reached via the Custom Endpoint flow (base URL `http://localhost:11434/v1`, no key). Use cloud for large models you can't run locally; use local for privacy or offline work.
</Tip>

### DeepInfra

DeepInfra (`--provider deepinfra`, `DEEPINFRA_API_KEY`) is discovered live from its catalog. Reasoning is controlled through DeepInfra's top-level `reasoning_effort` field, so `agent.reasoning_effort`, `/reasoning <level>`, `--reasoning` and per-model `agent.reasoning_overrides` work in **both directions**: an effort turns thinking on for models that default off (DeepSeek-V4.x), `/reasoning none` turns it off for models that default on (GLM-4.6, Qwen3-Thinking). Leaving reasoning unset keeps DeepInfra's per-model default; `xhigh` is native and `ultra` is sent as `max`.

### AWS Bedrock

Anthropic Claude, Amazon Nova, DeepSeek v3.2, Meta Llama 4, and other models via AWS Bedrock. Uses the AWS SDK (`boto3`) credential chain — no API key, just standard AWS auth.

```bash theme={null}
# Simplest — named profile in ~/.aws/credentials
mibyan chat --provider bedrock --model us.anthropic.claude-sonnet-4-6

# Or with explicit env vars
AWS_PROFILE=myprofile AWS_REGION=us-east-1 mibyan chat --provider bedrock --model us.anthropic.claude-sonnet-4-6
```

Or permanently in `config.yaml`:

```yaml theme={null}
model:
  provider: "bedrock"
  default: "us.anthropic.claude-sonnet-4-6"
bedrock:
  region: "us-east-1"          # or set AWS_REGION
  # profile: "myprofile"       # or set AWS_PROFILE
  # discovery: true            # auto-discover region from IAM
  # guardrail:                 # optional Bedrock Guardrails
  #   guardrail_identifier: "your-guardrail-id"
  #   guardrail_version: "DRAFT"
```

Authentication uses the standard boto3 chain: explicit `AWS_ACCESS_KEY_ID`/`AWS_SECRET_ACCESS_KEY`, `AWS_PROFILE` from `~/.aws/credentials`, IAM role on EC2/ECS/Lambda, IMDS, or SSO. No env var is required if you're already authenticated with the AWS CLI.

Bedrock uses the **Converse API** under the hood — requests are translated to Bedrock's model-agnostic shape, so the same config works for Claude, Nova, DeepSeek, and Llama models. Set `BEDROCK_BASE_URL` only if you're calling a non-default regional endpoint.

See the [AWS Bedrock guide](/desktop/guides/aws-bedrock) for a walkthrough of IAM setup, region selection, and cross-region inference.

### Google Vertex AI

Gemini models on Google Cloud Vertex AI via Vertex's OpenAI-compatible endpoint. Authentication is **OAuth2** — a short-lived access token (\~1 hour) minted from a service-account JSON or Application Default Credentials (ADC). There is **no static API key**; Mibyan mints and auto-refreshes the token for you, including re-minting on a mid-session `401`.

```bash theme={null}
# Service account JSON (recommended for servers / gateways)
echo "VERTEX_CREDENTIALS_PATH=/path/to/service-account.json" >> ~/.mibyan/.env
# or Application Default Credentials
gcloud auth application-default login

mibyan model   # → "Google Vertex AI" → project → region → model
```

Or in `config.yaml` (project/region are non-secret and live here; the credential path stays in `.env`):

```yaml theme={null}
model:
  provider: "vertex"
  default: "google/gemini-3-flash-preview"   # Vertex requires the google/ prefix
vertex:
  project_id: "my-gcp-project"   # blank → use the project embedded in the credentials
  region: "global"               # required for the Gemini 3.x previews
```

`VERTEX_PROJECT_ID` / `VERTEX_REGION` env vars override the `config.yaml` values. Mibyan lazy-installs `google-auth` on first use; run `mibyan setup` if the managed install needs repair. See the [Google Vertex AI guide](/desktop/guides/google-vertex) for the full walkthrough, and the [Google Gemini guide](/desktop/guides/google-gemini) for the static-API-key AI Studio path instead.

### Qwen Portal (OAuth)

Alibaba's Qwen Portal with browser-based OAuth login. Pick **Qwen OAuth (Portal)** in `mibyan model`, sign in through the browser, and Mibyan persists the refresh token.

```bash theme={null}
mibyan model
# → pick "Qwen OAuth (Portal)"
# → browser opens; sign in with your Alibaba account
# → confirm — credentials are saved to ~/.mibyan/auth.json

mibyan chat   # uses portal.qwen.ai/v1 endpoint
```

Or configure `config.yaml`:

```yaml theme={null}
model:
  provider: "qwen-oauth"
  default: "qwen3-coder-plus"
```

Set `mibyan_QWEN_BASE_URL` only if the portal endpoint relocates (default: `https://portal.qwen.ai/v1`).

<Tip>
  **Qwen OAuth vs Qwen Cloud (Alibaba DashScope)**

  `qwen-oauth` uses the consumer-facing Qwen Portal with OAuth login — ideal for individual users. The `alibaba` provider uses Qwen Cloud (Alibaba DashScope) with a `DASHSCOPE_API_KEY` — ideal for programmatic / production workloads. Both route to Qwen-family models but live at different endpoints.
</Tip>

### Alibaba Cloud (Coding Plan)

If you're subscribed to Alibaba's **Coding Plan** (a pricing SKU separate from standard DashScope API access), Mibyan exposes it as its own first-class provider: `alibaba-coding-plan`. Endpoint: `https://coding-intl.dashscope.aliyuncs.com/v1`. It's OpenAI-compatible like the regular `alibaba` provider but with a different base URL and billing surface.

```yaml theme={null}
model:
  provider: alibaba_coding     # alias for alibaba-coding-plan
  model: qwen3-coder-plus
```

Or from the CLI:

```bash theme={null}
mibyan chat --provider alibaba_coding --model qwen3-coder-plus
```

`alibaba_coding` uses the same `DASHSCOPE_API_KEY` your `alibaba` entry already uses — no separate key needed, just a different routing target. Before this provider was registered, users who set `provider: alibaba_coding` in `config.yaml` silently fell through to OpenRouter routing.

For the mainland-China endpoint (`alibaba-coding-plan-cn`, `https://coding.dashscope.aliyuncs.com/v1`) set `ALIBABA_CODING_PLAN_CN_API_KEY`. The CN provider still falls back to `ALIBABA_CODING_PLAN_API_KEY` / `DASHSCOPE_API_KEY`, but with only the shared key set the `/model` picker lists just the international row — set the CN key (or `provider: alibaba-coding-plan-cn` in `config.yaml`) to surface the CN one. The same applies to `alibaba-token-plan-cn` with `ALIBABA_TOKEN_PLAN_CN_API_KEY`.

### MiniMax (OAuth)

MiniMax-M2.7 via browser OAuth login — no API key needed. Pick **MiniMax (OAuth)** in `mibyan model`, sign in through the browser, and Mibyan persists the access + refresh tokens. Uses the Anthropic Messages-compatible endpoint (`/anthropic`) under the hood.

```bash theme={null}
mibyan model
# → pick "MiniMax (OAuth)"
# → browser opens; sign in with your MiniMax account (global or CN region)
# → confirm — credentials are saved to ~/.mibyan/auth.json

mibyan chat   # uses api.minimax.io/anthropic endpoint
```

Or configure `config.yaml`:

```yaml theme={null}
model:
  provider: "minimax-oauth"
  default: "MiniMax-M2.7"
```

Supported models: `MiniMax-M2.7` (main) and `MiniMax-M2.7-highspeed` (wired as the default auxiliary model). The OAuth path ignores `MINIMAX_API_KEY` / `MINIMAX_BASE_URL`.

<Tip>
  **MiniMax OAuth vs API key**

  `minimax-oauth` uses MiniMax's consumer-facing portal with OAuth login — no billing setup required. The `minimax` and `minimax-cn` providers use `MINIMAX_API_KEY` / `MINIMAX_CN_API_KEY` — for programmatic access. See the [MiniMax OAuth guide](/desktop/guides/minimax-oauth) for a full walkthrough.
</Tip>

### NVIDIA NIM

Nemotron and other open source models via [build.nvidia.com](https://build.nvidia.com) (free API key) or a local NIM endpoint.

```bash theme={null}
# Cloud (build.nvidia.com)
mibyan chat --provider nvidia --model nvidia/nemotron-3-super-120b-a12b
# Requires: NVIDIA_API_KEY in ~/.mibyan/.env

# Local NIM endpoint — override base URL
NVIDIA_BASE_URL=http://localhost:8000/v1 mibyan chat --provider nvidia --model nvidia/nemotron-3-super-120b-a12b
```

Or set it permanently in `config.yaml`:

```yaml theme={null}
model:
  provider: "nvidia"
  default: "nvidia/nemotron-3-super-120b-a12b"
```

<Tip>
  **Local NIM**

  For on-prem deployments (DGX Spark, local GPU), set `NVIDIA_BASE_URL=http://localhost:8000/v1`. NIM exposes the same OpenAI-compatible chat completions API as build.nvidia.com, so switching between cloud and local is a one-line env-var change.
</Tip>

Mibyan automatically attaches the NIM billing-origin header on every request to `build.nvidia.com` — no configuration needed. This routes consumption against the correct origin in NVIDIA's billing dashboard.

### GMI Cloud

Open and reasoning models via [GMI Cloud](https://www.gmicloud.ai/) — OpenAI-compatible API, API key authentication.

```bash theme={null}
# GMI Cloud
mibyan chat --provider gmi --model deepseek-ai/DeepSeek-V3.2
# Requires: GMI_API_KEY in ~/.mibyan/.env
```

Or set it permanently in `config.yaml`:

```yaml theme={null}
model:
  provider: "gmi"
  default: "deepseek-ai/DeepSeek-V3.2"
```

The base URL can be overridden with `GMI_BASE_URL` (default: `https://api.gmi-serving.com/v1`).

### Actual Computer

Your own hardware as a private inference cluster via [Actual Computer](https://actual.inc). Two serving modes, both using Chat Completions so reasoning and final content are returned together:

* **Hosted relay** — `https://api.actual.inc`, end-to-end encrypted, routes to *your* cluster. Authenticate with an `ac_` inference key from [actual.inc/user/keys](https://actual.inc/user/keys).
* **Local daemon** — on-device at `http://127.0.0.1:8080`, fully offline. No API key needed: Mibyan detects the loopback base URL and authenticates with an internal placeholder automatically.

```bash theme={null}
# Hosted relay (ACTUAL_API_KEY in ~/.mibyan/.env)
mibyan chat --provider actual --model <model-id-from-your-cluster>

# Local daemon (model.base_url in ~/.mibyan/config.yaml, no key)
mibyan chat --provider actual --model <installed-model-name>
```

Store provider settings in `~/.mibyan/config.yaml`; only the hosted API key belongs in `.env`:

```yaml theme={null}
model:
  provider: "actual"
  default: "<model-id>"
  base_url: "http://127.0.0.1:8080/v1" # Omit for the hosted relay.
```

Notes:

* Model IDs come from your cluster's `GET /v1/models` — discover with `mibyan model` or `curl -s https://api.actual.inc/v1/models -H "Authorization: Bearer $ACTUAL_API_KEY"`.
* Bare hosts in `model.base_url` are normalized: `http://127.0.0.1:8080` becomes `http://127.0.0.1:8080/v1` automatically. The legacy `ACTUAL_BASE_URL` environment variable is a fallback when no Actual URL is configured in YAML.
* Actual uses `/v1/chat/completions` for chat, compaction, title generation, and every other auxiliary task. This also applies to custom providers targeting `api.actual.inc`, model switches, and fallbacks. Legacy Responses settings in the main model, custom provider, or auxiliary task configuration are overridden automatically.
* Reasoning effort is clamped to Actual's supported range (`none/low/medium/high/max`) — a global `xhigh`/`ultra` setting will not 400 requests.
* Small local models: Mibyan' full default toolset plus the system prompt can exceed a 32k context window, producing an empty-stream error from llama.cpp-family servers. Restrict the toolset (`-t file,web`) or load the model with a larger context. The optional `actual-setup` skill (`mibyan skills install official/devops/actual-setup`) covers setup and troubleshooting in detail.
* Aliases: `actual-computer`, `actualcomputer`, `aci`.

### StepFun

Step-series models via [StepFun](https://platform.stepfun.com) — OpenAI-compatible API, API key authentication.

```bash theme={null}
# StepFun
mibyan chat --provider stepfun --model step-3.5-flash
# Requires: STEPFUN_API_KEY in ~/.mibyan/.env
```

Or set it permanently in `config.yaml`:

```yaml theme={null}
model:
  provider: "stepfun"
  default: "step-3.5-flash"
```

The base URL can be overridden with `STEPFUN_BASE_URL` (default: `https://api.stepfun.com/v1`).

### Hugging Face Inference Providers

[Hugging Face Inference Providers](https://huggingface.co/docs/inference-providers) routes to 20+ open models through a unified OpenAI-compatible endpoint (`router.huggingface.co/v1`). Requests are automatically routed to the fastest available backend (Groq, Together, SambaNova, etc.) with automatic failover.

```bash theme={null}
# Use any available model
mibyan chat --provider huggingface --model Qwen/Qwen3.5-397B-A17B
# Requires: HF_TOKEN in ~/.mibyan/.env

# Short alias
mibyan chat --provider hf --model deepseek-ai/DeepSeek-V3.2
```

Or set it permanently in `config.yaml`:

```yaml theme={null}
model:
  provider: "huggingface"
  default: "Qwen/Qwen3.5-397B-A17B"
```

Get your token at [huggingface.co/settings/tokens](https://huggingface.co/settings/tokens) — make sure to enable the "Make calls to Inference Providers" permission. Free tier included (\$0.10/month credit, no markup on provider rates).

You can append routing suffixes to model names: `:fastest` (default), `:cheapest`, or `:provider_name` to force a specific backend.

The base URL can be overridden with `HF_BASE_URL`.

## Custom & Self-Hosted LLM Providers

Mibyan works with **any OpenAI-compatible API endpoint**. If a server implements `/v1/chat/completions`, you can point Mibyan at it. This means you can use local models, GPU inference servers, multi-provider routers, or any third-party API.

### General Setup

Three ways to configure a custom endpoint:

**Interactive setup (recommended):**

```bash theme={null}
mibyan model
# Select "Custom endpoint (self-hosted / VLLM / etc.)"
# Enter: API base URL, API key, Model name
```

**Manual config (`config.yaml`):**

```yaml theme={null}
# In ~/.mibyan/config.yaml
model:
  default: your-model-name
  provider: custom
  base_url: http://localhost:8000/v1
  api_key: your-key-or-leave-empty-for-local
  # key_env: MY_PROVIDER_API_KEY  # env var holding the key (alternative to api_key)
```

<Warning>
  **Legacy env vars**

  `LLM_MODEL` in `.env` is **removed** — `config.yaml` is the single source of truth for model and endpoint configuration. `OPENAI_BASE_URL` is still honored, but **only** for the `openai-api` provider (it overrides the OpenAI endpoint for direct API-key access). For other providers and custom endpoints, use `mibyan model` or set `model.base_url` in `config.yaml` directly. If you have stale entries in your `.env`, they are automatically cleared on the next `mibyan setup` or config migration.
</Warning>

Both approaches persist to `config.yaml`, which is the source of truth for model, provider, and base URL.

### Switching Models with `/model`

<Warning>
  **mibyan model vs /model**

  **`mibyan model`** (run from your terminal, outside any chat session) is the **full provider setup wizard**. Use it to add new providers, run OAuth flows, enter API keys, and configure custom endpoints.

  **`/model`** (typed inside an active Mibyan chat session) can only **switch between providers and models you've already set up**. It cannot add new providers, run OAuth, or prompt for API keys. If you've only configured one provider (e.g. OpenRouter), `/model` will only show models for that provider.

  **To add a new provider:** Exit your session (`Ctrl+C` or `/quit`), run `mibyan model`, set up the new provider, then start a new session.
</Warning>

Once you have at least one custom endpoint configured, you can switch models mid-session:

```
/model custom:qwen-2.5          # Switch to a model on your custom endpoint
/model custom                    # Auto-detect the model from the endpoint
/model openrouter:claude-sonnet-4 # Switch back to a cloud provider
```

If you have **named custom providers** configured (see below), use the triple syntax:

```
/model custom:local:qwen-2.5    # Use the "local" custom provider with model qwen-2.5
/model custom:work:llama3       # Use the "work" custom provider with llama3
```

When switching providers, Mibyan persists the base URL and provider to config so the change survives restarts. When switching away from a custom endpoint to a built-in provider, the stale base URL is automatically cleared.

<Tip>
  `/model custom` (bare, no model name) queries your endpoint's `/models` API and auto-selects the model if exactly one is loaded. Useful for local servers running a single model.
</Tip>

Everything below follows this same pattern — just change the URL, key, and model name.

***

### Ollama — Local Models, Zero Config

[Ollama](https://ollama.com/) runs open-weight models locally with one command. Best for: quick local experimentation, privacy-sensitive work, offline use. Supports tool calling via the OpenAI-compatible API.

```bash theme={null}
# Install and run a model
ollama pull qwen2.5-coder:32b
ollama serve   # Starts on port 11434
```

Then configure Mibyan:

```bash theme={null}
mibyan model
# Select "Custom endpoint (self-hosted / VLLM / etc.)"
# Enter URL: http://localhost:11434/v1
# Skip API key (Ollama doesn't need one)
# Enter model name (e.g. qwen2.5-coder:32b)
```

Or configure `config.yaml` directly:

```yaml theme={null}
model:
  default: qwen2.5-coder:32b
  provider: custom
  base_url: http://localhost:11434/v1
  context_length: 64000   # See warning below
```

<Warning>
  **Ollama defaults to very low context lengths**

  Ollama does **not** use your model's full context window by default. Depending on your VRAM, the default is:

  | Available VRAM | Default context |
  | - | - |
  | Less than 24 GB | **4,096 tokens** |
  | 24–48 GB | 32,768 tokens |
  | 48+ GB | 256,000 tokens |

  Mibyan requires at least **64,000 tokens** of context for agent use with tools. Smaller windows are rejected at startup because the system prompt, tool schemas, and working conversation state need enough room for reliable multi-step workflows.

  **How to increase it** (pick one):

  ```bash theme={null}
  # Option 1: Set server-wide via environment variable (recommended)
  OLLAMA_CONTEXT_LENGTH=64000 ollama serve

  # Option 2: For systemd-managed Ollama
  sudo systemctl edit ollama.service
  # Add: Environment="OLLAMA_CONTEXT_LENGTH=64000"
  # Then: sudo systemctl daemon-reload && sudo systemctl restart ollama

  # Option 3: Bake it into a custom model (persistent per-model)
  echo -e "FROM qwen2.5-coder:32b\nPARAMETER num_ctx 64000" > Modelfile
  ollama create qwen2.5-coder-64k -f Modelfile
  ```

  **You cannot set context length through the OpenAI-compatible API** (`/v1/chat/completions`). It must be configured server-side or via a Modelfile. This is the #1 source of confusion when integrating Ollama with tools like Mibyan.
</Warning>

**Verify your context is set correctly:**

```bash theme={null}
ollama ps
# Look at the CONTEXT column — it should show your configured value
```

<Tip>
  List available models with `ollama list`. Pull any model from the [Ollama library](https://ollama.com/library) with `ollama pull <model>`. Ollama handles GPU offloading automatically — no configuration needed for most setups.
</Tip>

***

### vLLM — High-Performance GPU Inference

[vLLM](https://docs.vllm.ai/) is the standard for production LLM serving. Best for: maximum throughput on GPU hardware, serving large models, continuous batching.

```bash theme={null}
pip install vllm
vllm serve meta-llama/Llama-3.1-70B-Instruct \
  --port 8000 \
  --max-model-len 65536 \
  --tensor-parallel-size 2 \
  --enable-auto-tool-choice \
  --tool-call-parser mibyan
```

Then configure Mibyan:

```bash theme={null}
mibyan model
# Select "Custom endpoint (self-hosted / VLLM / etc.)"
# Enter URL: http://localhost:8000/v1
# Skip API key (or enter one if you configured vLLM with --api-key)
# Enter model name: meta-llama/Llama-3.1-70B-Instruct
```

**Context length:** vLLM reads the model's `max_position_embeddings` by default. If that exceeds your GPU memory, it errors and asks you to set `--max-model-len` lower. You can also use `--max-model-len auto` to automatically find the maximum that fits. Set `--gpu-memory-utilization 0.95` (default 0.9) to squeeze more context into VRAM.

**Tool calling requires explicit flags:**

| Flag | Purpose |
| - | - |
| `--enable-auto-tool-choice` | Required for `tool_choice: "auto"` (the default in Mibyan) |
| `--tool-call-parser <name>` | Parser for the model's tool call format |

Supported parsers: `hermes` (Qwen 2.5, Hermes 2/3), `llama3_json` (Llama 3.x), `mistral`, `deepseek_v3`, `deepseek_v31`, `xlam`, `pythonic`. Without these flags, tool calls won't work — the model will output tool calls as text.

**Qwen reasoning parsers:** Mibyan preserves structured reasoning metadata such as `reasoning`, `reasoning_content`, and streamed reasoning deltas when OpenAI-compatible servers return them. That metadata is treated as reasoning/thinking trace data, not as a replacement for the assistant's visible answer. For Qwen reasoning models served by vLLM, make sure the final user-visible response still appears in `content`. If `--reasoning-parser qwen3` leaves `content` empty in your deployment, either disable that parser or pass a server-supported request option such as `chat_template_kwargs.enable_thinking: false` through `extra_body`.

<Tip>
  vLLM supports human-readable sizes: `--max-model-len 64k` (lowercase k = 1000, uppercase K = 1024).
</Tip>

***

### SGLang — Fast Serving with RadixAttention

[SGLang](https://github.com/sgl-project/sglang) is an alternative to vLLM with RadixAttention for KV cache reuse. Best for: multi-turn conversations (prefix caching), constrained decoding, structured output.

```bash theme={null}
pip install "sglang[all]"
python -m sglang.launch_server \
  --model meta-llama/Llama-3.1-70B-Instruct \
  --port 30000 \
  --context-length 65536 \
  --tp 2 \
  --tool-call-parser qwen
```

Then configure Mibyan:

```bash theme={null}
mibyan model
# Select "Custom endpoint (self-hosted / VLLM / etc.)"
# Enter URL: http://localhost:30000/v1
# Enter model name: meta-llama/Llama-3.1-70B-Instruct
```

**Context length:** SGLang reads from the model's config by default. Use `--context-length` to override. If you need to exceed the model's declared maximum, set `SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1`.

**Tool calling:** Use `--tool-call-parser` with the appropriate parser for your model family: `qwen` (Qwen 2.5), `llama3`, `llama4`, `deepseekv3`, `mistral`, `glm`. Without this flag, tool calls come back as plain text.

<Warning>
  **SGLang defaults to 128 max output tokens**

  If responses seem truncated, check the server's generation default and configure it on the server (for example SGLang's `--default-max-tokens`). Mibyan does not expose an output-token cap setting.
</Warning>

***

### llama.cpp / llama-server — CPU & Metal Inference

[llama.cpp](https://github.com/ggml-org/llama.cpp) runs quantized models on CPU, Apple Silicon (Metal), and consumer GPUs. Best for: running models without a datacenter GPU, Mac users, edge deployment.

```bash theme={null}
# Build and start llama-server
cmake -B build && cmake --build build --config Release
./build/bin/llama-server \
  --jinja -fa \
  -c 64000 \
  -ngl 99 \
  -m models/qwen2.5-coder-32b-instruct-Q4_K_M.gguf \
  --port 8080 --host 0.0.0.0
```

**Context length (`-c`):** Recent builds default to `0` which reads the model's training context from the GGUF metadata. For models with 128k+ training context, this can OOM trying to allocate the full KV cache. Set `-c` explicitly to at least 64,000 tokens for Mibyan. If using parallel slots (`-np`), the total context is divided among slots — with `-c 64000 -np 4`, each slot only gets 16k, which is below Mibyan' minimum per active session.

Then configure Mibyan to point at it:

```bash theme={null}
mibyan model
# Select "Custom endpoint (self-hosted / VLLM / etc.)"
# Enter URL: http://localhost:8080/v1
# Skip API key (local servers don't need one)
# Enter model name — or leave blank to auto-detect if only one model is loaded
```

This saves the endpoint to `config.yaml` so it persists across sessions.

<Warning>
  **`--jinja` is required for tool calling**

  Without `--jinja`, llama-server ignores the `tools` parameter entirely. The model will try to call tools by writing JSON in its response text, but Mibyan won't recognize it as a tool call — you'll see raw JSON like `{"name": "web_search", ...}` printed as a message instead of an actual search.

  Native tool calling support (best performance): Llama 3.x, Qwen 2.5 (including Coder), Hermes 2/3, Mistral, DeepSeek, Functionary. All other models use a generic handler that works but may be less efficient. See the [llama.cpp function calling docs](https://github.com/ggml-org/llama.cpp/blob/master/docs/function-calling.md) for the full list.

  You can verify tool support is active by checking `http://localhost:8080/props` — the `chat_template` field should be present.
</Warning>

<Tip>
  Download GGUF models from [Hugging Face](https://huggingface.co/models?library=gguf). Q4\_K\_M quantization offers the best balance of quality vs. memory usage.
</Tip>

***

### LM Studio — Desktop App with Local Models

[LM Studio](https://lmstudio.ai/) is a desktop app for running local models with a GUI. Best for: users who prefer a visual interface, quick model testing, developers on macOS/Windows/Linux.

Start the server from the LM Studio app (Developer tab → Start Server), or use the CLI:

```bash theme={null}
lms server start                        # Starts on port 1234
lms load qwen2.5-coder --context-length 64000
```

Then configure Mibyan:

```bash theme={null}
mibyan model
# Select "LM Studio"
# Press Enter to use http://localhost:1234/v1
# Pick one of the discovered models
# If LM Studio server auth is enabled, enter LM_API_KEY when prompted
```

Mibyan preserves the context of an already-loaded LM Studio instance. For an unloaded model in the default explicit mode, Mibyan omits `context_length` unless you configured one in Mibyan, so LM Studio can apply its own model setting. Mibyan then uses only the context length LM Studio reports after loading.

To change context length in LM Studio:

1. Click the gear icon next to the model picker
2. Set "Context Length" to at least 64000 for a smooth experience
3. Reload the model for the change to take effect
4. If your machine cannot fit 64000, consider using a smaller model with larger context lengths.

Alternatively, use the CLI: `lms load model-name --context-length 64000`

You can use the CLI to estimate if the model will fit: `lms load model-name --context-length 64000 --estimate-only`

To set persistent per-model defaults: My Models tab → gear icon on the model → set context size.
:::

If you use LM Studio's Just-In-Time loading / Auto-Evict feature and want LM Studio to manage model loading and eviction from normal chat requests, skip Mibyan' explicit preload step:

```bash theme={null}
mibyan config set model.lmstudio_load_mode jit
```

Set it back to the default explicit preload behavior with:

```bash theme={null}
mibyan config set model.lmstudio_load_mode explicit
```

**Tool calling:** Supported since LM Studio 0.3.6. Models with native tool-calling training (Qwen 2.5, Llama 3.x, Mistral, Mibyan) are auto-detected and shown with a tool badge. Other models use a generic fallback that may be less reliable.

***

### WSL2 Networking (Windows Users)

Since Mibyan requires a Unix environment, Windows users run it inside WSL2. If your model server (Ollama, LM Studio, etc.) runs on the **Windows host**, you need to bridge the network gap — WSL2 uses a virtual network adapter with its own subnet, so `localhost` inside WSL2 refers to the Linux VM, **not** the Windows host.

<Tip>
  **Both in WSL2? No problem.**

  If your model server also runs inside WSL2 (common for vLLM, SGLang, and llama-server), `localhost` works as expected — they share the same network namespace. Skip this section.
</Tip>

#### Option 1: Mirrored Networking Mode (Recommended)

Available on **Windows 11 22H2+**, mirrored mode makes `localhost` work bidirectionally between Windows and WSL2 — the simplest fix.

1. Create or edit `%USERPROFILE%\.wslconfig` (e.g., `C:\Users\YourName\.wslconfig`):
   ```ini theme={null}
   [wsl2]
   networkingMode=mirrored
   ```

2. Restart WSL from PowerShell:
   ```powershell theme={null}
   wsl --shutdown
   ```

3. Reopen your WSL2 terminal. `localhost` now reaches Windows services:
   ```bash theme={null}
   curl http://localhost:11434/v1/models   # Ollama on Windows — works
   ```

<Note>
  **Hyper-V Firewall**

  On some Windows 11 builds, the Hyper-V firewall blocks mirrored connections by default. If `localhost` still doesn't work after enabling mirrored mode, run this in an **Admin PowerShell**:

  ```powershell theme={null}
  Set-NetFirewallHyperVVMSetting -Name '{40E0AC32-46A5-438A-A0B2-2B479E8F2E90}' -DefaultInboundAction Allow
  ```
</Note>

#### Option 2: Use the Windows Host IP (Windows 10 / older builds)

If you can't use mirrored mode, find the Windows host IP from inside WSL2 and use that instead of `localhost`:

```bash theme={null}
# Get the Windows host IP (the default gateway of WSL2's virtual network)
ip route show | grep -i default | awk '{ print $3 }'
# Example output: 172.29.192.1
```

Use that IP in your Mibyan config:

```yaml theme={null}
model:
  default: qwen2.5-coder:32b
  provider: custom
  base_url: http://172.29.192.1:11434/v1   # Windows host IP, not localhost
```

<Tip>
  **Dynamic helper**

  The host IP can change on WSL2 restart. You can grab it dynamically in your shell:

  ```bash theme={null}
  export WSL_HOST=$(ip route show | grep -i default | awk '{ print $3 }')
  echo "Windows host at: $WSL_HOST"
  curl http://$WSL_HOST:11434/v1/models   # Test Ollama
  ```

  Or use your machine's mDNS name (requires `libnss-mdns` in WSL2):

  ```bash theme={null}
  sudo apt install libnss-mdns
  curl http://$(hostname).local:11434/v1/models
  ```
</Tip>

#### Server Bind Address (Required for NAT Mode)

If you're using **Option 2** (NAT mode with the host IP), the model server on Windows must accept connections from outside `127.0.0.1`. By default, most servers only listen on localhost — WSL2 connections in NAT mode come from a different virtual subnet and will be refused. In mirrored mode, `localhost` maps directly so the default `127.0.0.1` binding works fine.

| Server | Default bind | How to fix |
| - | - | - |
| **Ollama** | `127.0.0.1` | Set `OLLAMA_HOST=0.0.0.0` environment variable before starting Ollama (System Settings → Environment Variables on Windows, or edit the Ollama service) |
| **LM Studio** | `127.0.0.1` | Enable **"Serve on Network"** in the Developer tab → Server settings |
| **llama-server** | `127.0.0.1` | Add `--host 0.0.0.0` to the startup command |
| **vLLM** | `0.0.0.0` | Already binds to all interfaces by default |
| **SGLang** | `127.0.0.1` | Add `--host 0.0.0.0` to the startup command |

**Ollama on Windows (detailed):** Ollama runs as a Windows service. To set `OLLAMA_HOST`:

1. Open **System Properties** → **Environment Variables**
2. Add a new **System variable**: `OLLAMA_HOST` = `0.0.0.0`
3. Restart the Ollama service (or reboot)

#### Windows Firewall

Windows Firewall treats WSL2 as a separate network (in both NAT and mirrored mode). If connections still fail after the steps above, add a firewall rule for your model server's port:

```powershell theme={null}
# Run in Admin PowerShell — replace PORT with your server's port
New-NetFirewallRule -DisplayName "Allow WSL2 to Model Server" -Direction Inbound -Action Allow -Protocol TCP -LocalPort 11434
```

Common ports: Ollama `11434`, vLLM `8000`, SGLang `30000`, llama-server `8080`, LM Studio `1234`.

#### Quick Verification

From inside WSL2, test that you can reach your model server:

```bash theme={null}
# Replace URL with your server's address and port
curl http://localhost:11434/v1/models          # Mirrored mode
curl http://172.29.192.1:11434/v1/models       # NAT mode (use your actual host IP)
```

If you get a JSON response listing your models, you're good. Use that same URL as the `base_url` in your Mibyan config.

***

### Troubleshooting Local Models

These issues affect **all** local inference servers when used with Mibyan.

#### "Connection refused" from WSL2 to a Windows-hosted model server

If you're running Mibyan inside WSL2 and your model server on the Windows host, `http://localhost:<port>` won't work in WSL2's default NAT networking mode. See [WSL2 Networking](#wsl2-networking-windows-users) above for the fix.

#### Tool calls appear as text instead of executing

The model outputs something like `{"name": "web_search", "arguments": {...}}` as a message instead of actually calling the tool.

**Cause:** Your server doesn't have tool calling enabled, or the model doesn't support it through the server's tool calling implementation.

| Server | Fix |
| - | - |
| **llama.cpp** | Add `--jinja` to the startup command |
| **vLLM** | Add `--enable-auto-tool-choice --tool-call-parser mibyan` |
| **SGLang** | Add `--tool-call-parser qwen` (or appropriate parser) |
| **Ollama** | Tool calling is enabled by default — make sure your model supports it (check with `ollama show model-name`) |
| **LM Studio** | Update to 0.3.6+ and use a model with native tool support |

#### Model seems to forget context or give incoherent responses

**Cause:** Context window is too small. When the conversation exceeds the context limit, most servers silently drop older messages. Mibyan's system prompt + tool schemas alone can use 4k–8k tokens.

**Diagnosis:**

```bash theme={null}
# Check what Mibyan thinks the context is
# Look at startup line: "Context limit: X tokens"

# Check your server's actual context
# Ollama: ollama ps (CONTEXT column)
# llama.cpp: curl http://localhost:8080/props | jq '.default_generation_settings.n_ctx'
# vLLM: check --max-model-len in startup args
```

**Fix:** Set context to at least **64,000 tokens** for agent use. See each server's section above for the specific flag.

The startup refusal for a local endpoint (`127.0.0.1`, LAN, Docker service names) says which window the server is serving and names the fix for any OpenAI-compatible server, not just Ollama: raise the server's context (llama.cpp `-c 64000`, vLLM `--max-model-len`, Ollama `OLLAMA_CONTEXT_LENGTH`/Modelfile `num_ctx`) or set `model.ollama_num_ctx` in `config.yaml` to the window the server really serves (at least 64K). `model.ollama_num_ctx` is honoured on every local endpoint; only the automatic detection behind it uses Ollama's `/api/show`.

#### "Context limit: 2048 tokens" at startup

Mibyan auto-detects context length from your server's `/v1/models` endpoint. If the server reports a low value (or doesn't report one at all), Mibyan uses the model's declared limit which may be wrong.

**Fix:** Set it explicitly in `config.yaml`:

```yaml theme={null}
model:
  default: your-model
  provider: custom
  base_url: http://localhost:11434/v1
  context_length: 64000
```

#### Responses get cut off mid-sentence

**Possible causes:**

1. **Low output limit on the server** — configure the server's generation default (for example SGLang's `--default-max-tokens`). Mibyan does not expose an output-token cap setting. Response length is distinct from the conversation's context window (`context_length`).
2. **Context exhaustion** — The model filled its context window. Increase `model.context_length` or enable [context compression](/desktop/user-guide/configuration#context-compression) in Mibyan.

***

### LiteLLM Proxy — Multi-Provider Gateway

[LiteLLM](https://docs.litellm.ai/) is an OpenAI-compatible proxy that unifies 100+ LLM providers behind a single API. Best for: switching between providers without config changes, load balancing, fallback chains, budget controls.

```bash theme={null}
# Install and start
pip install "litellm[proxy]"
litellm --model anthropic/claude-sonnet-4 --port 4000

# Or with a config file for multiple models:
litellm --config litellm_config.yaml --port 4000
```

Then configure Mibyan with `mibyan model` → Custom endpoint → `http://localhost:4000/v1`.

Example `litellm_config.yaml` with fallback:

```yaml theme={null}
model_list:
  - model_name: "best"
    litellm_params:
      model: anthropic/claude-sonnet-4
      api_key: sk-ant-...
  - model_name: "best"
    litellm_params:
      model: openai/gpt-4o
      api_key: sk-...
router_settings:
  routing_strategy: "latency-based-routing"
```

***

### ClawRouter — Cost-Optimized Routing

[ClawRouter](https://github.com/BlockRunAI/ClawRouter) by BlockRunAI is a local routing proxy that auto-selects models based on query complexity. It classifies requests across 14 dimensions and routes to the cheapest model that can handle the task. Payment is via USDC cryptocurrency (no API keys).

```bash theme={null}
# Install and start
npx @blockrun/clawrouter    # Starts on port 8402
```

Then configure Mibyan with `mibyan model` → Custom endpoint → `http://localhost:8402/v1` → model name `blockrun/auto`.

Routing profiles:

| Profile | Strategy | Savings |
| - | - | - |
| `blockrun/auto` | Balanced quality/cost | 74-100% |
| `blockrun/eco` | Cheapest possible | 95-100% |
| `blockrun/premium` | Best quality models | 0% |
| `blockrun/free` | Free models only | 100% |
| `blockrun/agentic` | Optimized for tool use | varies |

<Note>
  ClawRouter requires a USDC-funded wallet on Base or Solana for payment. All requests route through BlockRun's backend API. Run `npx @blockrun/clawrouter doctor` to check wallet status.
</Note>

***

### Other Compatible Providers

Any service with an OpenAI-compatible API works. Some popular options:

| Provider | Base URL | Notes |
| - | - | - |
| [Together AI](https://together.ai) | `https://api.together.xyz/v1` | Cloud-hosted open models |
| [Groq](https://groq.com) | `https://api.groq.com/openai/v1` | Ultra-fast inference |
| [DeepSeek](https://deepseek.com) | `https://api.deepseek.com/v1` | DeepSeek models |
| [Fireworks AI](https://fireworks.ai) | `https://api.fireworks.ai/inference/v1` | Fast open model hosting |
| [GMI Cloud](https://www.gmicloud.ai/) | `https://api.gmi-serving.com/v1` | Managed OpenAI-compatible inference |
| [Actual Computer](https://actual.inc) | `https://api.actual.inc/v1` | Private relay to your own cluster; local daemon at `http://127.0.0.1:8080/v1` |
| [Cerebras](https://cerebras.ai) | `https://api.cerebras.ai/v1` | Wafer-scale chip inference |
| [Mistral AI](https://mistral.ai) | `https://api.mistral.ai/v1` | Mistral models |
| [OpenAI](https://openai.com) | `https://api.openai.com/v1` | Direct OpenAI access |
| [Azure OpenAI](https://azure.microsoft.com) | `https://YOUR.openai.azure.com/` | Enterprise OpenAI |
| [LocalAI](https://localai.io) | `http://localhost:8080/v1` | Self-hosted, multi-model |
| [Jan](https://jan.ai) | `http://localhost:1337/v1` | Desktop app with local models |

Configure any of these with `mibyan model` → Custom endpoint, or in `config.yaml`:

```yaml theme={null}
model:
  default: meta-llama/Llama-3.1-70B-Instruct-Turbo
  provider: custom
  base_url: https://api.together.xyz/v1
  api_key: your-together-key
```

***

### Context Length Detection

<Note>
  **Context windows and output limits are different**

  **`context_length`** is the **total context window** — the combined budget for input *and* output tokens (e.g. 200,000 for Claude Opus 4.6). Mibyan uses this to decide when to compress history and to validate API requests.

  Output limits govern a single generated response, not the conversation history.
  Mibyan no longer reads `model.max_tokens`, `mibyan_MAX_TOKENS`, provider output-cap
  settings, or `model_overrides.*.*.max_output_tokens`. Remove these legacy settings.
  Custom OpenAI-compatible endpoints receive no automatic catalog-sized output cap.
  Their server defaults apply; these can be lower than the model maximum.
  A reply that degenerates into a repetition loop is still stopped: within about 130,000
  characters of the loop starting (visible or reasoning text), Mibyan closes the stream
  and ends the turn with a "Repetition Detected" notice, so an uncapped endpoint cannot keep a
  looping model running.

  Native Anthropic Messages (including the native Anthropic Bedrock path) requires
  `max_tokens`, so Mibyan supplies an internal value. Bedrock Converse is a separate
  protocol: its optional `inferenceConfig.maxTokens` is omitted by default, which
  [AWS documents as the model maximum](https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_InferenceConfiguration.html).
  Internal bounded tasks and provider-specific protocol requirements remain implementation
  details. Omission does not universally select a model's maximum output.

  Set `context_length` when auto-detection gets the window size wrong.
</Note>

Mibyan uses a multi-source resolution chain to detect the correct context window for your model and provider:

1. **Config override** — `model.context_length` in config.yaml (highest priority). This is an explicit **pin**: it always wins over provider metadata, so Mibyan labels it `(pinned)` wherever the window is shown (welcome banner, `/model`, `/usage`, the status bar) and logs one warning at startup when the pin disagrees with the window the provider is known to advertise. The pin is dropped automatically when you switch model, provider or base URL.
2. **Custom provider per-model** — `providers.<name>.models.<id>.context_length`
3. **Persistent cache** — previously discovered values (survives restarts)
4. **Endpoint `/models`** — queries your server's API (local/custom endpoints)
5. **Anthropic `/v1/models`** — queries Anthropic's API for `max_input_tokens` (API-key users only)
6. **OpenRouter API** — live model metadata from OpenRouter
7. **Nous Portal** — suffix-matches Nous model IDs against OpenRouter metadata
8. **[models.dev](https://models.dev)** — community-maintained registry with provider-specific context lengths for 3800+ models across 100+ providers
9. **Fallback defaults** — broad model family patterns (128K default)

For most setups this works out of the box. The system is provider-aware — the same model can have different context limits depending on who serves it (e.g., `claude-opus-4.6` is 1M on Anthropic direct but 128K on GitHub Copilot).

To set the context length explicitly, add `context_length` to your model config:

```yaml theme={null}
model:
  default: "qwen3.5:9b"
  base_url: "http://localhost:8080/v1"
  context_length: 131072  # tokens
```

For custom endpoints, you can also set context length per model:

```yaml theme={null}
providers:
  my-local-llm:
    api: "http://localhost:11434/v1"
    models:
      qwen3.5:27b:
        context_length: 64000
      deepseek-r1:70b:
        context_length: 65536
```

`mibyan model` will prompt for context length when configuring a custom endpoint. Leave it blank for auto-detection.

<Tip>
  **When to set this manually**

  * You're using Ollama with a custom `num_ctx` that's lower than the model's maximum
  * You want to limit context below the model's maximum (e.g., 8k on a 128k model to save VRAM)
  * You're running behind a proxy that doesn't expose `/v1/models`
</Tip>

***

### Named Custom Providers

If you work with multiple custom endpoints (e.g., a local dev server and a remote GPU server), you can define them as named custom providers under the `providers:` dict in `config.yaml`, keyed by provider name:

```yaml theme={null}
providers:
  local:
    api: http://localhost:8080/v1
    # api_key omitted — Mibyan uses "no-key-required" for keyless local servers
  work:
    api: https://gpu-server.internal.corp/v1
    key_env: CORP_API_KEY
    transport: chat_completions   # set explicitly by `mibyan model` → Custom Endpoint wizard; auto-detection still happens as a fallback
  anthropic-proxy:
    api: https://proxy.example.com/anthropic
    key_env: ANTHROPIC_PROXY_KEY
    transport: anthropic_messages  # for Anthropic-compatible proxies
```

Each entry accepts: `api` (the endpoint base URL — `base_url`/`url` are accepted aliases), `name` (optional display name; defaults to the dict key), `key_env` or inline `api_key` or `key_cmd` (see below), `transport` (`chat_completions` / `anthropic_messages` / `codex_responses`), `default_model`, `models`, `context_length`, `discover_models`, `extra_body`, `extra_headers`, `session_affinity_header` (name of a header that carries the conversation id, for session-aware proxies; off unless set), `ssl_ca_cert` / `ssl_verify`, `catalog_provider` (see below), and `enabled: false` to hide an entry without deleting it.

#### Command-minted credentials (`key_cmd`)

Vision, thinking, and native local-model capability probes materialize the same
callable credential used by chat before building authentication headers. They
reuse the command token cache without replacing the chat client's callable.
If a command cannot mint a string token, these best-effort probes send no bearer
rather than an object representation or a lower-priority configured credential.
Native local-model probes remove inherited Authorization on a failed explicit
callable while retaining unrelated configured headers. Chat retains its normal
error handling.

Enterprise gateways often issue short-lived bearer tokens (SSO/OIDC brokers, cloud IAM, internal auth proxies) rather than static API keys, so a token copied into `.env` goes stale mid-session and requests start returning 401. `key_cmd` names a command that *prints* a token; Mibyan runs it and caches the result until shortly before expiry, so long sessions keep working with no restart:

```yaml theme={null}
providers:
  my-gateway:
    base_url: "https://gateway.internal.example.com/v1"
    api_mode: chat_completions
    key_cmd: "my-auth-cli print-token --profile prod"
```

Works with any helper that prints a token — `databricks auth token`, `gcloud auth print-access-token`, `az account get-access-token`, `vault read`, or Claude Code-style `apiKeyHelper` scripts.

The command must print **only** the token on stdout: either bare, or as JSON with an `access_token` field (`expires_in` is honored; absolute `expiry`/`expiresOn` ISO timestamps too). Multi-line output is rejected rather than guessed at. If no expiry is advertised, the token is re-minted on a bounded window.

Precedence: an explicit `--api-key` flag still wins; otherwise `key_cmd` beats a static `api_key`/`key_env` on the same entry. The minted credential applies to the main agent turn and to auxiliary tasks (title generation, compression, vision, embedding) alike.

Model discovery also honors `key_cmd` for both `providers:` and legacy
`custom_providers:` entries, including `mibyan model` setup. Helpers run only when
an authenticated live catalog probe is needed: disabled discovery and warm catalog
cache reads do not mint tokens. Catalogs are scoped to the command identity, so
rotating a bearer does not invalidate the catalog. Probe helpers use their own
short-lived token source, not the inference client's token cache; minted bearers
are never saved to `config.yaml`. If a helper fails, discovery falls back to the
configured model without exposing the helper's output.

Not to be confused with `secrets.command`, which runs a helper **once at startup** to populate env vars process-wide. Use that for a vault/keychain helper handing back many secrets; use `key_cmd` when one provider's credential must be re-minted *during* a session.

<Note>
  **Legacy format**

  Older configs used a top-level `custom_providers:` list instead. It still works — Mibyan reads both — and `mibyan update` auto-migrates it to the `providers:` dict (config v12). Field names differ slightly in the dict format: legacy `model` is `default_model`, and legacy `api_mode` is `transport`.
</Note>

**Context window on `codex_responses` proxies.** A custom entry with `transport: codex_responses` (a local Codex proxy, for example) resolves the context window of Codex OAuth models (`gpt-6-astra`, `gpt-5.6-sol`/`-terra`/`-luna`, `gpt-5.5`, …) from the Codex OAuth table — 272K for most slugs — not from the 1.05M direct-API catalog, so compression fires before the Codex backend's limit and its 272K billing tier. The decision follows the transport, not the hostname; the same holds for `openai-codex` behind `mibyan_CODEX_BASE_URL` or `model.base_url`. A per-model `models.<id>.context_length`, an entry-level `context_length`, or `model.context_length` still wins; the opt-in `-900k` picker variants keep their verified 900K.

**Reasoning effort on custom endpoints.** The configured `reasoning_effort` (`/reasoning max`, `agent.reasoning_effort`) reaches a custom endpoint unchanged on both the `chat_completions` and the `codex_responses` transport — up to `max`; only the Mibyan-internal `ultra` is clamped to `max`. Two exceptions follow the host rather than the entry: a custom entry pointed at `api.openai.com` keeps OpenAI's per-model ladder (`max` is a gpt-5.6-only level there), and an entry pointed at a provider whose profile publishes a per-model vocabulary (Ramp Router) is clamped to that catalog. An endpoint that rejects the level answers with an HTTP 400 instead of Mibyan silently downgrading it. When no effort is configured at all, `chat_completions` requests carry `reasoning_effort: medium` — the same default the Nous Portal and OpenRouter routes apply — rather than leaving the endpoint's own default in charge (kimi-k3's is `max`, about 3x the reasoning tokens of `medium`); models marked `supports_reasoning: false` in the catalog or `model_overrides`, and Ollama models without the `thinking` capability, keep the field off.

Some OpenAI-compatible endpoints need provider-specific request body fields. Add an `extra_body` map to the matching custom provider and Mibyan will merge it into each chat-completions request for that endpoint:

```yaml theme={null}
providers:
  gemma-local:
    api: http://localhost:8080/v1
    default_model: google/gemma-4-31b-it
    extra_body:
      enable_thinking: true
      reasoning_effort: high
```

Use the shape your server documents. For example, vLLM Gemma deployments and some NVIDIA NIM endpoints expect `enable_thinking` under `chat_template_kwargs` instead of as a top-level `extra_body` field:

```yaml theme={null}
extra_body:
  chat_template_kwargs:
    enable_thinking: true
```

For Qwen reasoning models served by vLLM, this same shape can be used to disable thinking when a reasoning parser separates all generated text into reasoning fields and leaves the assistant `content` empty:

```yaml theme={null}
extra_body:
  chat_template_kwargs:
    enable_thinking: false
```

The configured `extra_body` follows the provider everywhere: it is merged at agent construction, **survives every gateway turn** (including turns where `/fast` layers `service_tier`/`speed` overrides on top — those merge over your `extra_body` rather than replacing it), and is **re-derived on `/model` switches** — switching to a named custom provider applies its `extra_body`, and switching away clears it so it never leaks to another provider.

The `mibyan model` → Custom Endpoint wizard now prompts for the API mode explicitly and persists your answer to `config.yaml` (as `transport` on the provider entry). URL-based auto-detection (e.g. `/anthropic` paths → `anthropic_messages`) still happens as a fallback when the field is left blank.

**Native vision for custom-provider models.** If your custom endpoint serves a vision-capable model that isn't in models.dev, set `model.supports_vision: true` so Mibyan routes attached images natively (as `image_url` parts) instead of pre-processing them through `vision_analyze`. Single knob — no need to also set `agent.image_input_mode: native`.

```yaml theme={null}
model:
  provider: custom
  base_url: http://localhost:8080/v1
  default: qwen3.6-35b-a3b
  supports_vision: true   # send images natively; otherwise vision_analyze pre-describes them
```

The same key is honored on per-named-provider models (`providers.<name>.models.<id>.supports_vision`) and accepts standard YAML booleans (`true/false/yes/no/on/off/1/0`).

A `model_overrides` entry that only corrects metadata (for example `context_window`) for a model the catalog does not know leaves vision and reasoning capability and the output-token limit **unknown** — `vision_analyze`, `video_analyze` and the reasoning-effort picker stay available. Only an explicit `supports_vision: false` / `supports_reasoning: false` in the override marks the model as text-only or non-reasoning.

**Inheriting a catalogued vendor's metadata (`catalog_provider`).** When a named custom provider (a gateway, proxy or reseller) serves models that Mibyan already knows under a built-in provider, point the entry at that vendor and its models inherit the catalogued context window, output limit, vision and reasoning flags — no `model_overrides` needed:

```yaml theme={null}
providers:
  my-gateway:
    api: https://gateway.example.com/v1
    key_env: GATEWAY_API_KEY
    catalog_provider: deepseek   # metadata lookups use DeepSeek's catalog entries
```

`catalog_provider` accepts a Mibyan provider id (`deepseek`, `anthropic`, `openai`, …) or a models.dev id. It affects metadata lookups only — requests still go to your `api` URL with your credentials — and an explicit `model_overrides` entry for the same model still wins.

Switch between them mid-session with the triple syntax:

```
/model custom:local:qwen-2.5       # Use the "local" endpoint with qwen-2.5
/model custom:work:llama3-70b      # Use the "work" endpoint with llama3-70b
/model custom:anthropic-proxy:claude-sonnet-4  # Use the proxy
```

You can also select named custom providers from the interactive `mibyan model` menu.

***

### Cookbook: Together AI, Groq, Perplexity

The cloud providers listed in [Other Compatible Providers](#other-compatible-providers) all speak OpenAI's REST dialect, so they wire up the same way under the `providers:` dict. Three worked recipes follow. Each drops into `~/.mibyan/config.yaml` and the matching API key goes in `~/.mibyan/.env`.

#### Together AI

Hosts open-weight models (Llama, MiniMax, Gemma, DeepSeek, Qwen) at prices significantly below first-party APIs. Good default for multi-model fleets.

```yaml theme={null}
# ~/.mibyan/config.yaml
providers:
  together:
    api: https://api.together.xyz/v1
    key_env: TOGETHER_API_KEY
    # transport: chat_completions  # default — no need to set

model:
  default: MiniMaxAI/MiniMax-M2.7   # or any model from together.ai/models
  provider: custom:together
```

```bash theme={null}
# ~/.mibyan/.env
TOGETHER_API_KEY=your-together-key
```

Switch models mid-session:

```
/model custom:together:meta-llama/Llama-3.3-70B-Instruct-Turbo
/model custom:together:google/gemma-4-31b-it
/model custom:together:deepseek-ai/DeepSeek-V3
```

Together's `/v1/models` endpoint works, so `mibyan model` can auto-discover available models.

#### Groq

Ultra-fast inference (\~500 tok/s on Llama-3.3-70B). Small catalog but strong for latency-sensitive interactive use.

```yaml theme={null}
# ~/.mibyan/config.yaml
providers:
  groq:
    api: https://api.groq.com/openai/v1
    key_env: GROQ_API_KEY

model:
  default: llama-3.3-70b-versatile
  provider: custom:groq
```

```bash theme={null}
# ~/.mibyan/.env
GROQ_API_KEY=your-groq-key
```

#### Perplexity

Useful when you want a model that does live web search and citation automatically. Strict about which models are available — check [perplexity.ai/settings/api](https://www.perplexity.ai/settings/api) for the current list.

```yaml theme={null}
# ~/.mibyan/config.yaml
providers:
  perplexity:
    api: https://api.perplexity.ai
    key_env: PERPLEXITY_API_KEY

model:
  default: sonar
  provider: custom:perplexity
```

```bash theme={null}
# ~/.mibyan/.env
PERPLEXITY_API_KEY=your-perplexity-key
```

Perplexity's Agent API (`api: https://api.perplexity.ai/v1` with `api_mode: codex_responses`) reserves the function names `web_search`, `search_files`, `fetch_url`, `people_search` and `finance_search` for its own built-in tools. Mibyan renames its client tools of the same name to `mibyan_<name>` on the wire and maps them back before dispatch, for the main agent loop and auxiliary calls (title generation, compression, MoA aggregation) alike — the same treatment OpenCode's `/v1/responses` endpoints get.

#### Multiple providers in one config

The three recipes compose — use all of them together and switch per turn with `/model custom:<name>:<model>`:

```yaml theme={null}
providers:
  together:
    api: https://api.together.xyz/v1
    key_env: TOGETHER_API_KEY
  groq:
    api: https://api.groq.com/openai/v1
    key_env: GROQ_API_KEY
  perplexity:
    api: https://api.perplexity.ai
    key_env: PERPLEXITY_API_KEY

model:
  default: MiniMaxAI/MiniMax-M2.7
  provider: custom:together      # boot to Together; switch freely after
```

<Tip>
  **Troubleshooting**

  * `mibyan doctor` should print no `Unknown provider` warnings for any of these names after the CLI validator fixes in #15083.
  * If a provider's `/v1/models` endpoint is unreachable (Perplexity is the common one), `mibyan model` will persist the model with a warning rather than hard-reject — see #15136.
  * To skip named providers entirely and use bare `provider: custom` with `CUSTOM_BASE_URL` env var, see #15103.
</Tip>

***

### Choosing the Right Setup

| Use Case | Recommended |
| - | - |
| **Just want it to work** | OpenRouter (default) or Nous Portal |
| **Local models, easy setup** | Ollama |
| **Production GPU serving** | vLLM or SGLang |
| **Mac / no GPU** | Ollama or llama.cpp |
| **Multi-provider routing** | LiteLLM Proxy or OpenRouter |
| **Cost optimization** | ClawRouter or OpenRouter with `sort: "price"` |
| **Maximum privacy** | Ollama, vLLM, or llama.cpp (fully local) |
| **Enterprise / Azure** | Azure OpenAI with custom endpoint |
| **Chinese AI models** | z.ai (GLM), Kimi/Moonshot (`kimi-coding` or `kimi-coding-cn`), MiniMax, Xiaomi MiMo, or Tencent TokenHub (first-class providers) |

<Tip>
  You can switch between providers at any time with `mibyan model` — no restart required. Your conversation history, memory, and skills carry over regardless of which provider you use.
</Tip>

## Optional API Keys

| Feature | Provider | Env Variable |
| - | - | - |
| Web scraping | [Firecrawl](https://firecrawl.dev/) | `FIRECRAWL_API_KEY`, `FIRECRAWL_API_URL` |
| Browser automation | [Browserbase](https://browserbase.com/) | `BROWSERBASE_API_KEY`, `BROWSERBASE_PROJECT_ID` |
| Image generation | [FAL](https://fal.ai/) | `FAL_KEY` |
| Premium TTS voices | [ElevenLabs](https://elevenlabs.io/) | `ELEVENLABS_API_KEY` |
| OpenAI TTS + voice transcription | [OpenAI](https://platform.openai.com/api-keys) | `VOICE_TOOLS_OPENAI_KEY` |
| Mistral TTS + voice transcription | [Mistral](https://console.mistral.ai/) | `MISTRAL_API_KEY` |
| Cross-session user modeling | [Honcho](https://honcho.dev/) | `HONCHO_API_KEY` |
| Semantic long-term memory | [Supermemory](https://supermemory.ai) | `SUPERMEMORY_API_KEY` |

### Self-Hosting Firecrawl

By default, Mibyan uses the [Firecrawl cloud API](https://firecrawl.dev/) for web search and scraping. If you prefer to run Firecrawl locally, you can point Mibyan at a self-hosted instance instead. See Firecrawl's [SELF\_HOST.md](https://github.com/firecrawl/firecrawl/blob/main/SELF_HOST.md) for complete setup instructions.

**What you get:** No API key required, no rate limits, no per-page costs, full data sovereignty.

**What you lose:** The cloud version uses Firecrawl's proprietary "Fire-engine" for advanced anti-bot bypassing (Cloudflare, CAPTCHAs, IP rotation). Self-hosted uses basic fetch + Playwright, so some protected sites may fail. Search uses DuckDuckGo instead of Google.

**Setup:**

1. Clone and start the Firecrawl Docker stack (5 containers: API, Playwright, Redis, RabbitMQ, PostgreSQL — requires \~4-8 GB RAM):
   ```bash theme={null}
   git clone https://github.com/firecrawl/firecrawl
   cd firecrawl
   # In .env, set: USE_DB_AUTHENTICATION=false, HOST=0.0.0.0, PORT=3002
   docker compose up -d
   ```

2. Point Mibyan at your instance (no API key needed):
   ```bash theme={null}
   mibyan config set FIRECRAWL_API_URL http://localhost:3002
   ```

You can also set both `FIRECRAWL_API_KEY` and `FIRECRAWL_API_URL` if your self-hosted instance has authentication enabled.

## OpenRouter Provider Routing

When using OpenRouter, you can control how requests are routed across providers. Add a `provider_routing` section to `~/.mibyan/config.yaml`:

```yaml theme={null}
provider_routing:
  sort: "throughput"          # "price" (default), "throughput", or "latency"
  # only: ["anthropic"]      # Only use these providers
  # ignore: ["deepinfra"]    # Skip these providers
  # order: ["anthropic", "google"]  # Try providers in this order
  # require_parameters: true  # Only use providers that support all request params
  # data_collection: "deny"   # Exclude providers that may store/train on data
  # models:                   # Per-model pins (same keys; unset keys fall through)
  #   "openai/gpt-6-astra": {only: ["openai"]}
  #   "anthropic/claude-fable-5.1": {only: ["anthropic"]}
```

**Shortcuts:** Append `:nitro` to any model name for throughput sorting (e.g., `anthropic/claude-sonnet-4:nitro`), or `:floor` for price sorting. Per-model details: [Provider Routing](/desktop/user-guide/features/provider-routing#per-model-overrides-models).

## OpenRouter Pareto Code Router

OpenRouter ships an experimental coding-model router at `openrouter/pareto-code` that auto-routes requests to the cheapest model meeting a coding-quality bar (ranked by [Artificial Analysis](https://artificialanalysis.ai/)). Pick this model and tune the `min_coding_score` knob in `~/.mibyan/config.yaml`:

```yaml theme={null}
model:
  provider: openrouter
  model: openrouter/pareto-code

openrouter:
  min_coding_score: 0.65   # 0.0–1.0; higher = stronger (more expensive) coders. Default 0.65.
```

Notes:

* `min_coding_score` is **only** sent when `model.model` is `openrouter/pareto-code`. On any other model the value is a no-op.
* Set to empty string (or remove the line) to let OpenRouter pick the strongest available coder — its documented behavior when the plugins block is omitted.
* Selection is deterministic per score on a given day, but the actual model chosen can shift as the Pareto frontier moves (new models, benchmark updates).
* See OpenRouter's [Pareto Router docs](https://openrouter.ai/docs/guides/routing/routers/pareto-router) for the full router behavior.
* To use the Pareto Code router for a specific **auxiliary task** (compression, vision, etc.) instead of the main agent, set `extra_body.plugins` under that task — see [Auxiliary Models → OpenRouter routing & Pareto Code for auxiliary tasks](/desktop/user-guide/configuration#openrouter-routing--pareto-code-for-auxiliary-tasks).

## Fallback Providers

Configure a chain of backup providers Mibyan tries in order when the primary model fails (rate limits, server errors, auth failures). The canonical format is a top-level `fallback_providers:` list:

```yaml theme={null}
fallback_providers:
  - provider: openrouter
    model: anthropic/claude-sonnet-4
  - provider: anthropic
    model: claude-sonnet-4
    # base_url: http://localhost:8000/v1    # optional, for custom endpoints
    # api_mode: chat_completions           # optional override (`transport:` is an accepted alias)
```

An entry that names a `providers.<name>` block (`provider: my-relay` or `provider: custom:my-relay`) inherits that block's `transport` / `api_mode` when the entry sets none, so a Responses-only or Anthropic-Messages relay keeps its declared wire on fallback. Set `api_mode` on the entry to override it.

The legacy single-pair `fallback_model:` dict is still accepted for back-compat:

```yaml theme={null}
fallback_model:
  provider: openrouter
  model: anthropic/claude-sonnet-4
```

When activated, the fallback swaps the model and provider mid-session without losing your conversation. The chain is tried entry-by-entry; activation is one-shot per session.

Supported providers: `openrouter`, `nous`, `novita`, `openai-codex`, `copilot`, `copilot-acp`, `anthropic`, `gemini`, `qwen-oauth`, `huggingface`, `zai`, `kimi-coding`, `kimi-coding-cn`, `minimax`, `minimax-cn`, `minimax-oauth`, `deepseek`, `nvidia`, `xai`, `xai-oauth`, `ollama-cloud`, `bedrock`, `ai-gateway`, `azure-foundry`, `opencode-zen`, `opencode-go`, `commandcode`, `commandcode-anthropic`, `kilocode`, `xiaomi`, `arcee`, `gmi`, `actual`, `stepfun`, `lmstudio`, `alibaba`, `alibaba-coding-plan`, `tencent-tokenhub`, `tencent-tokenplan`, `nebius-token-factory`, `router`, `custom`.

<Tip>
  Fallback is configured exclusively through `config.yaml` — or interactively via `mibyan fallback`. For full details on when it triggers, how the chain advances, and how it interacts with auxiliary tasks and delegation, see [Fallback Providers](/desktop/user-guide/features/fallback-providers).
</Tip>

***

## See Also

* [Configuration](/desktop/user-guide/configuration) — General configuration (directory structure, config precedence, terminal backends, memory, compression, and more)
* [Environment Variables](/desktop/reference/environment-variables) — Complete reference of all environment variables


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.