- Credential pools — rotate across multiple API keys for the same provider (tried first)
- Primary model fallback — automatically switches to a different provider:model when your main model fails
- Auxiliary task fallback — independent provider resolution for side tasks like vision and compression
Primary Model Fallback
When your main LLM provider encounters errors — rate limits, server overload, auth failures, connection drops — Mibyan can automatically switch to a backup provider:model pair mid-session without losing your conversation.Configuration
The easiest path is the interactive manager:mibyan fallback reuses the provider picker from mibyan model — same provider list, same credential prompts, same validation. Use the subcommands add, list (alias ls), remove (alias rm), and clear to manage the chain. Changes persist under the top-level fallback_providers: list in config.yaml.
If you’d rather edit the YAML directly, add a top-level fallback_providers list to ~/.mibyan/config.yaml:
provider and model. Entries missing either field are ignored.
When a rate-limit response names its reset time, the primary is benched until exactly then (a provider that says nothing gets the exponential 60 s → 4 h backoff). Optionally, skip the switch when the primary reopens soon:
Gemini fallback entries accept
gemini, google, google-gemini, and
google-ai-studio. On Google’s native API endpoint, all use the native Gemini
client, including its generationConfig.thinkingConfig translation. A custom
OpenAI-compatible base URL continues to use the compatible client instead.
fallback_model vs fallback_providersfallback_providers (plural, list) is the current config shape and supports multiple fallbacks tried in order. fallback_model (singular) is the legacy single-fallback key — Mibyan still honors it for back-compat, but mibyan fallback writes the current fallback_providers key and migrates legacy config on write. When both are set, fallback_providers takes priority.Supported Providers
Custom Endpoint Fallback
For a custom OpenAI-compatible endpoint, addbase_url and optionally key_env:
When Fallback Triggers
The fallback activates automatically when the primary model fails with:- Rate limits (HTTP 429) — after exhausting retry attempts
- Server errors (HTTP 500, 502, 503) — after exhausting retry attempts
- Auth failures (HTTP 401, 403) — immediately (no point retrying)
- Not found (HTTP 404) — immediately
- Invalid responses — when the API returns malformed or empty responses repeatedly. An HTTP-200 body whose only assistant text is a router’s
Connect timeout, please try again later.with zero completion tokens counts as invalid too (streamed or not, in the main loop, the iteration-limit summary and auxiliary calls), so it is retried instead of shown as the answer. A streamed refusal (the model declining with an explanation on the refusal channel) is a terminalcontent_filterresult, not an empty response, so it is surfaced rather than retried. On the native Anthropic wire astop_reason: refusalarrives with an empty body; Mibyan reports the reason from the response’sstop_details(category and, when present, explanation) in the refusal message and in the log line (native_stop_reason=… stop_details=…).
- Resolves credentials for the fallback provider (including named custom providers using
key_cmd) - Builds a new API client, preserving a dynamic credential source across timeout and request-client rebuilds
- Swaps the model, provider, and client in-place
- Re-resolves the reasoning effort for the fallback model (its
agent.reasoning_overridesentry, else the globalagent.reasoning_effort) - Resets the retry counter and continues the conversation
mibyan chat --reasoning <level> is kept across that startup switch — it is your intent for the run.
Per-Turn, Not Per-SessionFallback is turn-scoped: each new user message starts with the primary model restored. If the primary fails mid-turn, fallback activates for that turn only. On the next message, Mibyan tries the primary again. Within a single turn, fallback activates at most once — if the fallback also fails, normal error handling takes over (retries, then error message). This prevents cascading failover loops within a turn while giving the primary model a fresh chance every turn.The per-turn retry is reset-aware: when the primary’s credentials report a rate-limit reset time that hasn’t elapsed yet (subscription windows like Claude Pro/Max’s 5-hour blocks or Codex weekly limits report these as hours or days), Mibyan skips the doomed retry and stays on the fallback until the reset passes — avoiding two pointless provider switches (and two prompt-cache invalidations) per turn. Expiry makes the primary eligible for a later retry; it does not schedule a retry or guarantee recovery. Transient 429s without a reset time use an exponential cooldown.When a switch arms that cooldown, the fallback notice includes its approximate remaining duration, for example:
Primary retry eligible in ~60 s; recovery is not guaranteed. Non-rate-limit switches and switches from an already-active cross-provider fallback do not announce a new primary cooldown.Examples
OpenRouter as fallback for Anthropic native:Where Fallback Works
Auxiliary Task Fallback
Mibyan uses separate lightweight models for side tasks. Each task has its own provider resolution chain that acts as a built-in fallback system.Tasks with Independent Provider Resolution
Auto-Detection Chain
When a task’s provider is set to"auto" (the default), Mibyan first tries the main provider + main model for that auxiliary task. If that route is unavailable or later fails with a capacity-style error, Mibyan follows your configured fallback policy and then stops:
custom. A healthy local endpoint with a different base URL remains eligible for fallback and subsequent auto routing. Aliases for the same custom endpoint share its health state. Built-in providers retain their shared-account health checks.
The task-specific chain is most precise and wins when present. The top-level fallback_providers chain is the same policy the main agent uses, so free-only or same-provider fallback rules apply to auxiliary tasks on auto as well.
Built-in text discovery chain (compression, web extract, title generation, etc.):
model.provider: auto or unset). Once you have picked a main provider, an unavailable main route with no fallback_chain / fallback_providers skips the auxiliary task with a warning instead of guessing another provider you happen to be logged into — an expired xAI or Codex session must never bill your Nous Portal or OpenRouter balance behind your back. Declare a fallback if you want one.
Configuring Auxiliary Providers
Each task can be configured independently inconfig.yaml:
fallback_chain; if omitted, provider: auto uses the top-level fallback_providers chain (the built-in discovery chain applies only when no main provider is selected).
Context compression is configured under auxiliary.compression:
provider to pick who handles the request, model to pick which model, and base_url to point at a custom endpoint (overrides provider).
Provider Options for Auxiliary Tasks
These options apply toauxiliary:, compression:, and fallback_providers: entries only — "main" is not a valid value for your top-level model.provider. For custom endpoints, use provider: custom in your model: section (see AI Providers).
Direct Endpoint Override
For any auxiliary task, settingbase_url bypasses provider resolution entirely and sends requests directly to that endpoint:
base_url takes precedence over provider. Mibyan uses the configured api_key for authentication, falling back to OPENAI_API_KEY if not set. It does not reuse OPENROUTER_API_KEY for custom endpoints.
Auxiliary Capacity-Error Fallback
When you set an explicit auxiliary provider (e.g.auxiliary.vision.provider: glm), Mibyan treats that as your preferred choice — but if the provider literally cannot serve the request because of a capacity error (HTTP 402 payment required, HTTP 429 daily-quota exhaustion, connection failure), Mibyan falls back through a layered chain instead of failing silently:
- Primary aux provider — the one you configured (tried first, always)
auxiliary.<task>.fallback_chain— your per-task override list, if you wrote one- Main agent provider + model — last-resort safety net (always tried, even if you didn’t write a chain)
- Warn + re-raise — if every layer fails, Mibyan logs
Auxiliary <task>: ... all fallbacks exhaustedat WARNING level and re-raises the original error
Retry-After: ...) are treated as request constraints, not capacity problems — they respect your explicit provider choice and do not trigger the fallback ladder. Only daily/monthly quota exhaustion, payment errors, and connection failures bypass the explicit-provider gate.
Auth errors (HTTP 401) on an explicit provider walk only step 2: if you wrote auxiliary.<task>.fallback_chain, its entries are tried in order (and a chain entry that dies mid-request hands off to the next one); the main agent model and the auto-detection chain are never consulted, because you did not opt that task into them. Without a chain the task fails on the auth error as before. Auth is credential-wide, so chain entries on the same provider label are skipped — point the spare at a different provider (a separate providers: entry counts).
For users on provider: auto (no explicit aux provider), the existing auto-detection chain runs in place of steps 2–3. Its first step is already the main agent model, so auto users get the same outcome with zero config.
Optional: per-task fallback chain
If you want a different fallback ordering than “main agent model first”, configurefallback_chain explicitly. Each entry needs at least provider; model, base_url, and api_key are optional.
fallback_chain to get fallback — the main-agent safety net runs regardless. Use it only when you specifically want a different order than the default.
Each fallback_chain entry may also declare its own timeout (seconds). Without it, a fallback candidate inherits the task-level timeout — which may be tuned for the primary provider. Declaring a per-entry timeout lets a slower-but-reliable fallback (e.g. a large-context summarizer) get the budget it actually needs instead of dying on the primary’s clock.
Provider quota errors that trigger fallback
Mibyan recognizes these as capacity-equivalent to 402 credit exhaustion (not transient rate limits):- Bedrock / LiteLLM:
Too many tokens per day,daily limit,tokens per day - Vertex AI / GCP:
quota exceeded,resource exhausted,RESOURCE_EXHAUSTED - Generic:
daily quota,quota_exceeded
Context Compression Fallback
Context compression uses theauxiliary.compression config block to control which model and provider handles summarization:
Legacy migrationOlder configs with
compression.summary_model / compression.summary_provider / compression.summary_base_url are automatically migrated to auxiliary.compression.* on first load (config version 17).Delegation Provider Override
Subagents spawned bydelegate_task inherit the parent agent’s primary fallback chain. You can still route subagents to a different primary provider:model pair for cost optimization:
Cron Job Providers
Unpinned cron jobs inherit your configuredfallback_providers chain (or legacy fallback_model), both when the primary’s credentials fail to resolve before the run and when the provider errors mid-run. A job pinned to its own provider, model or endpoint does not: if that route fails, the run fails (same-provider credential pool rotation still applies). This matches how a pinned delegation child behaves. Pin a cron job with provider and model overrides on the job itself:
cron.model / cron.model_provider instead. See Scheduled Tasks (Cron) for details.

