Skip to main content
Mibyan has a shared provider runtime resolver used across:
  • CLI
  • gateway
  • cron jobs
  • ACP
  • auxiliary model calls
Primary implementation:
  • mibyan_cli/runtime_provider.py — credential resolution, custom-endpoint runtime resolution
  • mibyan_cli/auth.py — provider registry, resolve_provider()
  • mibyan_cli/model_switch.py — shared /model switch pipeline (CLI + gateway)
  • agent/auxiliary_client.py — auxiliary model routing
  • providers/ — ABC + registry entry points (ProviderProfile, register_provider, get_provider_profile, list_providers)
  • plugins/model-providers/<name>/ — per-provider plugins (bundled) that declare api_mode, base_url, env_vars, fallback_models and register themselves into the registry on first access. User plugins at $mibyan_HOME/plugins/model-providers/<name>/ override bundled ones of the same name.
get_provider_profile() in providers/ returns a ProviderProfile for a given provider id. runtime_provider.py calls this at resolution time to get the canonical base_url, env_vars priority list, api_mode, and fallback_models without needing to duplicate that data in multiple files. Adding a new plugin under plugins/model-providers/<your-provider>/ (or $mibyan_HOME/plugins/model-providers/<your-provider>/) that calls register_provider() is enough for runtime_provider.py to pick it up — no branch needed in the resolver itself. If you are trying to add a new first-class inference provider, read Adding Providers and the Model Provider Plugin guide alongside this page.

Chat-completions reasoning shapes

OpenAI-compatible relays can return reasoning or reasoning_content as strings, text-part dictionaries, or lists of text parts and string fragments. Mibyan flattens these fields before string operations in the main stream, Relay recording, synchronous and asynchronous auxiliary streams, and completed-response reasoning extraction. Fragments retain their explicit whitespace; normalization adds no intra-field separator. Main-stream and Relay recording retain the existing paragraph breaks between complete bold reasoning headings. Reasoning stays separate from the visible answer. delta.reasoning_details is a list of opaque provider records. Both the main stream and Relay recording append these records in arrival order without flattening, merging, or rewriting their contents. Providers can emit complete records on the final delta or across multiple deltas; omit the field on chunks without new records. The collected records pass through response normalization and assistant-message storage into session replay, including nested signed payloads.

Resolution precedence

At a high level, provider resolution uses:
  1. explicit CLI/runtime request
  2. config.yaml model/provider config
  3. environment variables
  4. provider-specific defaults or auto resolution
That ordering matters because Mibyan treats the saved model/provider choice as the source of truth for normal runs. This prevents a stale shell export from silently overriding the endpoint a user last selected in mibyan model.

Providers

Current provider families include (see plugins/model-providers/ for the complete bundled set):
  • AI Gateway (Vercel)
  • OpenRouter
  • Nous Portal
  • OpenAI Codex
  • Copilot / Copilot ACP
  • Anthropic (native)
  • Google / Gemini (gemini)
  • Alibaba / DashScope (alibaba, alibaba-coding-plan)
  • DeepSeek
  • Z.AI
  • Kimi / Moonshot (kimi-coding, kimi-coding-cn)
  • MiniMax (minimax, minimax-cn, minimax-oauth)
  • Kilo Code
  • Hugging Face
  • OpenCode Zen / OpenCode Go
  • AWS Bedrock
  • Azure Foundry
  • NVIDIA NIM
  • xAI (Grok)
  • Arcee
  • GMI Cloud
  • StepFun
  • Qwen OAuth
  • Xiaomi
  • Ollama Cloud
  • LM Studio
  • Tencent TokenHub
  • Custom (provider: custom) — first-class provider for any OpenAI-compatible endpoint
  • Named custom providers (providers: dict in config.yaml; the legacy custom_providers list is still read for backward compatibility)

Output of runtime resolution

The runtime resolver returns data such as:
  • provider
  • api_mode
  • base_url
  • api_key
  • source
  • provider-specific metadata like expiry/refresh info

Why this matters

This resolver is the main reason Mibyan can share auth/runtime logic between:
  • mibyan chat
  • gateway message handling
  • cron jobs running in fresh sessions
  • ACP editor sessions
  • auxiliary model tasks

AI Gateway

Set AI_GATEWAY_API_KEY in ~/.mibyan/.env and run with --provider ai-gateway. Mibyan fetches available models from the gateway’s /models endpoint, filtering to language models with tool-use support.

OpenRouter, AI Gateway, and custom OpenAI-compatible base URLs

Mibyan contains logic to avoid leaking the wrong API key to a custom endpoint when multiple provider keys exist (e.g. OPENROUTER_API_KEY, AI_GATEWAY_API_KEY, and OPENAI_API_KEY). Each provider’s API key is scoped to its own base URL:
  • OPENROUTER_API_KEY is only sent to openrouter.ai endpoints
  • AI_GATEWAY_API_KEY is only sent to ai-gateway.vercel.sh endpoints
  • OPENAI_API_KEY is used for custom endpoints and as a fallback
Mibyan also distinguishes between:
  • a real custom endpoint selected by the user
  • the OpenRouter fallback path used when no custom endpoint is configured
That distinction is especially important for:
  • local model servers
  • non-OpenRouter/non-AI Gateway OpenAI-compatible APIs
  • switching providers without re-running setup
  • config-saved custom endpoints that should keep working even when OPENAI_BASE_URL is not exported in the current shell

Native Anthropic path

Anthropic is not just “via OpenRouter” anymore. When provider resolution selects anthropic, Mibyan uses:
  • api_mode = anthropic_messages
  • the native Anthropic Messages API
  • agent/anthropic_adapter.py for translation
Credential resolution for native Anthropic now prefers refreshable Claude Code credentials over copied env tokens when both are present. In practice that means:
  • Claude Code credential files are treated as the preferred source when they include refreshable auth
  • manual ANTHROPIC_TOKEN / CLAUDE_CODE_OAUTH_TOKEN values still work as explicit overrides
  • Mibyan preflights Anthropic credential refresh before native Messages API calls
  • Mibyan still retries once on a 401 after rebuilding the Anthropic client, as a fallback path

OpenAI Codex path

Codex uses a separate Responses API path:
  • api_mode = codex_responses
  • dedicated credential resolution and auth store support
  • a resumed session whose lingering Codex reasoning items (encrypted_content) are rejected — as a 400 invalid_encrypted_content or as a 401 token_expired — self-heals by stripping the cached items and replaying once, before any credential refresh or pool rotation. Reasoning minted after that strip keeps replaying; only a second rejection in the same session turns replay off

Auxiliary model routing

Auxiliary tasks such as:
  • vision
  • web extraction summarization
  • context compression summaries
  • skills hub operations
  • MCP helper operations
  • memory flushes
can use their own provider/model routing rather than the main conversational model. When an auxiliary task is configured with provider main, Mibyan resolves that through the same shared runtime path as normal chat. In practice that means:
  • env-driven custom endpoints still work
  • custom endpoints saved via mibyan model / config.yaml also work
  • auxiliary routing can tell the difference between a real saved custom endpoint and the OpenRouter fallback

Fallback models

Mibyan supports a configured fallback provider chain — a list of (provider, model) entries tried in order when the primary model encounters errors. The legacy single-pair fallback_model dict is still accepted for back-compat (and migrated on first write).

How it works internally

  1. Storage: AIAgent.__init__ stores the fallback_model dict and sets _fallback_activated = False.
  2. Trigger points: _try_activate_fallback() (forwarded to try_activate_fallback() in agent/chat_completion_helpers.py) is called from three places in the turn phases (agent/turn_api_error.py, agent/turn_response_check.py, agent/turn_recovery.py):
    • After max retries on invalid API responses (None choices, missing content)
    • On non-retryable client errors (HTTP 401, 403, 404)
    • After max retries on transient errors (HTTP 429, 500, 502, 503)
  3. Activation flow (_try_activate_fallback):
    • Returns False immediately if already activated or not configured
    • Calls resolve_provider_client() from auxiliary_client.py to build a new client with proper auth
    • Determines api_mode: codex_responses for openai-codex, anthropic_messages for anthropic, chat_completions for everything else
    • Swaps in-place: self.model, self.provider, self.base_url, self.api_mode, self.client, self._client_kwargs
    • For anthropic fallback: builds a native Anthropic client instead of OpenAI-compatible
    • Re-evaluates prompt caching (enabled for Claude models on OpenRouter)
    • Sets _fallback_activated = True — prevents firing again
    • Resets retry count to 0 and continues the loop
  4. Config flow:
    • CLI: reads the fallback chain via mibyan_cli/fallback_config.get_fallback_chain() → passes to AIAgent(fallback_model=...)
    • Gateway: gateway/run_config_loaders.py._load_fallback_model() reads config.yaml → passes to AIAgent
    • Validation: both provider and model keys must be non-empty, or fallback is disabled

What does NOT support fallback

  • Subagent delegation (tools/delegate_tool.py): subagents inherit the parent’s provider but not the fallback config
  • Auxiliary tasks: use their own independent provider auto-detection chain (see Auxiliary model routing above)
Unpinned cron jobs do support fallback: run_job() reads fallback_providers (or legacy fallback_model) from config.yaml and passes it to AIAgent(fallback_model=...), matching the gateway’s _load_fallback_model() pattern. A job with its own provider / model / base_url gets no chain, the same rule as a pinned delegation child. See Cron Internals.

Test coverage

Fallback behavior is exercised across several suites:
  • tests/agent/test_fallback_credential_isolation.py — credential isolation between primary and fallback
  • tests/mibyan_cli/test_fallback_cmd.py — the /fallback CLI command