The API server exposes mibyan-agent as an OpenAI-compatible HTTP endpoint. Any frontend that speaks the OpenAI format — Open WebUI, LobeChat, LibreChat, NextChat, ChatBox, and hundreds more — can connect to mibyan-agent and use it as a backend.
Your agent handles requests with its full toolset (terminal, file operations, web search, memory, skills) and returns the final response. When streaming, tool progress indicators appear inline so frontends can show what the agent is doing.
One backend covers models + toolsMibyan itself needs a configured provider and tool backends for the API server to be useful. A Nous Portal subscription handles both — 300+ models plus web/image/TTS/browser via the Tool Gateway. Run mibyan setup --portal once before starting the API server and frontends like Open WebUI or LobeChat get a fully tool-equipped backend.
Quick Start
1. Enable the API server
Add to ~/.mibyan/.env:
2. Start the gateway
You’ll see:
3. Connect a frontend
Point any OpenAI-compatible client at http://localhost:8642/v1:
Or connect Open WebUI, LobeChat, or any other frontend — see the Open WebUI integration guide for step-by-step instructions.
Endpoints
POST /v1/chat/completions
Standard OpenAI Chat Completions format. Stateless — the full conversation is included in each request via the messages array.
Request:
Response:
Inline image input: user messages may send content as an array of text and image_url parts. Both remote http(s) URLs and data:image/... URLs are supported:
Uploaded files (file / input_file / file_id) and non-image data: URLs return 400 unsupported_content_type.
Streaming ("stream": true): Returns Server-Sent Events (SSE) with token-by-token response chunks. For Chat Completions, the stream uses standard chat.completion.chunk events plus Mibyan’ custom mibyan.tool.progress event for tool-start UX. For Responses, the stream uses OpenAI Responses event types such as response.created, response.output_text.delta, response.output_item.added, response.output_item.done, and response.completed.
All SSE streams (Chat Completions, Responses, /api/sessions/{id}/chat/stream, /v1/runs/{id}/events) emit a : keepalive comment line whenever no event has been sent for 10 seconds, so long tool calls do not trip client idle timeouts. Standard SSE clients ignore comment lines; custom parsers must skip lines that start with :.
Tool progress in streams:
- Chat Completions: Mibyan emits
event: mibyan.tool.progress for tool-start visibility without polluting persisted assistant text. Strict OpenAI clients that choke on named SSE events can turn these frames off with gateway.platforms.api_server.tool_progress_events: false (default true); content chunks are unaffected. The opt-out applies only to Chat Completions — /v1/runs/{id}/events always emits tool events, which is what the tool_progress_events feature in /v1/capabilities describes.
- Responses: Mibyan emits spec-native
function_call and function_call_output output items during the SSE stream, so clients can render structured tool UI in real time.
Model reasoning (emitted only when the model actually produces reasoning and the resolved reasoning config allows it; the input-side opt-out is model_options.reasoning.enabled: false):
- Chat Completions: reasoning deltas arrive as
choices[0].delta.reasoning_content chunks (the DeepSeek-style field Open WebUI, opencode and the Vercel AI SDK render as a thinking block); answer text stays in delta.content.
- Responses: each thinking burst is a spec-native
reasoning output item — response.output_item.added (item.type: "reasoning"), response.reasoning_summary_part.added, response.reasoning_summary_text.delta … response.reasoning_summary_text.done, response.reasoning_summary_part.done, response.output_item.done — closed before the next message or function_call item opens, and echoed in the response.completed output as {"id": "rs_…", "type": "reasoning", "status": "completed", "summary": [{"type": "summary_text", "text": "…"}]}. sequence_number stays monotonic across reasoning, text and tool events.
- Non-streaming:
/v1/chat/completions returns the turn’s reasoning on choices[0].message.reasoning_content; /v1/responses returns the same reasoning output item(s) ahead of the message (and of that step’s function_call items), also on GET /v1/responses/{id} replay.
- Echoing a prior response’s
output list back as the next input (what Responses SDK clients do) is fine: reasoning items are ignored on input rather than parsed as empty user turns.
- Support is advertised as
features.reasoning_streaming: true on GET /v1/capabilities.
POST /v1/responses
OpenAI Responses API format. Supports server-side conversation state via previous_response_id — the server stores full conversation history (including tool calls and results) so multi-turn context is preserved without the client managing it.
Request:
Response:
Tool calls in the output array were already executed server-side by the Mibyan agent — they are replayed with "status": "completed" for structured tool UI, never as pending calls for the client to execute.
With "stream": true, mid-turn assistant commentary (the openai-codex backend’s phase="commentary" progress preambles, or text a model writes alongside its tool calls) arrives as its own completed message output item carrying "phase": "commentary" (response.output_item.added + response.output_item.done, also listed in response.completed). It is never merged into the final answer item, so clients can render it as live progress and skip it when assembling the reply. Private reasoning never reaches this item. display.interim_assistant_messages: false (or the display.platforms.api_server override) suppresses it on every API-server surface.
Inline image input: input[].content can contain input_text and input_image parts. Both remote URLs and data:image/... URLs are supported:
Uploaded files (input_file / file_id) and non-image data: URLs return 400 unsupported_content_type.
Multi-turn with previous_response_id
Chain responses to maintain full context (including tool calls) across turns:
The server reconstructs the full conversation from the stored response chain — all previous tool calls and results are preserved. Chained requests also share the same session, so multi-turn conversations appear as a single entry in the dashboard and session history.
Each response’s output lists only that turn’s items (its function_call / function_call_output entries and final message), never earlier turns’ tool calls — including when Mibyan repaired the supplied history before the call (merged consecutive assistant or user items, dropped orphan tool results) or compacted it mid-chain. The stored chain is that repaired transcript, so the history does not grow by a second copy on every turn.
Named conversations
Use the conversation parameter instead of tracking response IDs:
The server automatically chains to the latest response in that conversation. Like the /title command for gateway sessions.
GET /v1/responses/{id}
Retrieve a previously stored response by ID.
DELETE /v1/responses/{id}
Delete a stored response.
GET /v1/models
Lists the agent as an available model. The advertised model name defaults to the profile name (or mibyan-agent for the default profile). Required by most frontends for model discovery.
/v1/models is intentionally the cheap OpenAI-compat surface. It does not
enumerate every authenticated provider/model combination Mibyan can route to,
and it does not do pricing or capability enrichment.
GET /api/model/options
Mibyan-aware clients can request the same curated provider/model inventory used
by the dashboard and TUI. This route uses the API server’s normal bearer
authentication and returns provider rows, model capability hints, and pricing
metadata that do not belong in the OpenAI-compatible /v1/models response:
That payload is the same substrate the dashboard Models page and the TUI
model.options RPC use. It returns authenticated providers, curated model
lists, per-model pricing, and model capability hints.
Normal opens are intentionally conservative for custom providers: Mibyan probes
only the currently selected custom endpoint so a stale or offline saved
endpoint does not block the picker. An explicit refresh flips to full probing
and busts the provider model cache:
Use /v1/models when an OpenAI-compatible client only needs a model name to
send back in chat/responses requests. Use /api/model/options when an
authenticated UI needs the richer Mibyan-specific picker metadata.
GET /v1/capabilities
Returns a machine-readable description of the API server’s stable surface for external UIs, orchestrators, and plugin bridges.
Use this endpoint when integrating dashboards, browser UIs, or control planes so they can discover whether the running Mibyan version supports runs, streaming, cancellation, and session continuity without depending on private Python internals.
Browser-extension control
Mibyan can route browser tools through an authenticated extension that controls
the browser session associated with the current Mibyan session. The feature is
disabled by default; set browser.extension_control.enabled to true to opt in:
The local API path also requires the API server bearer key. A controller may
register only for an existing server session. Mibyan derives the controller
principal from authenticated server state; a client-supplied principal_id is
ignored.
Discover the live contract through GET /v1/capabilities. The
browser_extension_control object reports whether the feature is enabled, the
protocol version, transport names, and the exact capability allowlist:
Requested capabilities outside that list are filtered out. Raw CDP, arbitrary
script evaluation, console access, uploads, image extraction, and vision are not
part of the controller protocol.
When a request has no bound controller identity, or when the feature is disabled,
Mibyan preserves the existing browser backend. Once the gateway binds a
controller principal and transport family to the request, that extension lane
is authoritative: missing, ambiguous, disconnected, or incapable controllers
fail closed instead of silently switching to a different local/cloud browser.
After an exact controller is selected, its result or error is authoritative and
Mibyan never retries the same action through another backend.
Local API registration
- Send an authenticated
POST /v1/browser-control/register with
protocol_version, session_id, controller_id, browser_profile_id, and
the requested capabilities.
- Mibyan returns a single-use ticket with a 30-second TTL and the filtered,
server-bound controller scope.
- Open
GET /v1/browser-control/ws with both WebSocket subprotocols:
mibyan-browser-control-v1 and
mibyan-browser-control-ticket.<ticket>.
The ticket is never accepted in the query string. Unknown, expired, reused, or
malformed tickets fail before WebSocket upgrade.
Controller frames
Mibyan sends browser.controller.command frames containing command_id,
action, immutable arguments, browser/controller ids, and the originating
tool_call_id. The controller replies with browser.controller.result, the
same command_id, an exact boolean ok, and either result or error.
Cancellation and timeout emit browser.controller.cancel; late results are
ignored.
An unexpected socket loss marks the controller offline and preserves work
already in flight until each command’s original deadline. A reconnect with the
same principal, profile, session, controller id, browser profile, and transport
identity refreshes the transport without admitting new work before any deferred
cancels are flushed. Negotiated capabilities may change on that reconnect; they
are not an identity field. A different controller id or browser profile in the
same authenticated session lane is a hard replacement: old pending work is
cancelled before the successor becomes routable. Send
browser.controller.detach on the authenticated controller transport for an
intentional hard detach — that immediately cancels pending work. Merely closing
the socket is treated as a recoverable disconnect.
The authenticated dashboard transport exposes the same registration, result,
heartbeat, capability, and ownership semantics over its Gateway RPC/event
channel. In both transports, selection requires one unambiguous exact match on
principal, profile, session, controller, browser profile, transport family, and
capability. Once selected, a controller failure is authoritative and is never
retried through a different browser backend.
Per-request model selection
Authenticated clients can override Mibyan’ default model selection per request
by sending:
model — the target model id for this turn
provider — the Mibyan provider slug to resolve credentials/runtime for this turn
model_options — request-scoped reasoning / service-tier controls
The same request fields are accepted on:
POST /v1/chat/completions
POST /v1/responses
POST /v1/runs
POST /api/sessions/{session_id}/chat
POST /api/sessions/{session_id}/chat/stream
Precedence is deterministic:
- Session
/model override, if that session already has one
- A static
gateway.platforms.api_server.model_routes mapping selected when
the request’s model is a configured route alias
- Direct request
model / provider when no route alias matches
- Global gateway config / environment defaults
model_options stays request-scoped regardless of which model/provider wins.
If a request sends a provider that conflicts with a configured model_routes
alias, Mibyan rejects the request with 400 instead of silently remixing route
credentials with another provider.
Bare model values on the OpenAI-compatible endpoints are opt-in. Generic
OpenAI clients routinely hardcode model names (gpt-4o, …), and existing
deployments rely on those falling back to the gateway default. On
POST /v1/chat/completions and POST /v1/responses, a model value sent
WITHOUT a provider is therefore ignored unless you enable:
Requests that include an explicit provider — and the Mibyan-native
/v1/runs and session-chat endpoints — always honor the requested model
regardless of this flag.
Example:
GET /health
Health check. Returns {"status": "ok"}. Also available at GET /v1/health for OpenAI-compatible clients that expect the /v1/ prefix.
GET /health/detailed
Authenticated readiness check for monitoring and control planes. It reports
bounded status for the active profile’s config, state database, configured
model, disk space, gateway/platform state, active API runs, pending process
completions, and active delegations. The response exposes status and counts,
not config values, credentials, paths, commands, queue payloads, or raw errors.
The public /health route remains a cheap liveness probe and does not run
readiness checks. A degraded readiness result still uses HTTP 200; inspect the
top-level status and readiness.checks fields.
Runs API (streaming-friendly alternative)
In addition to /v1/chat/completions and /v1/responses, the server exposes a runs API for long-form sessions where the client wants to subscribe to progress events instead of managing streaming themselves.
POST /v1/runs
Create a new agent run. Returns a run_id that can be used to subscribe to progress events.
Runs accept a simple input string and optional session_id, instructions, conversation_history, or previous_response_id. When session_id is provided, Mibyan surfaces it in the run status so external UIs can correlate runs with their own conversation IDs.
For safely retryable creation, send an Idempotency-Key header (1–255 visible ASCII characters). Mibyan durably reserves the key before starting work. An identical retry returns the original run_id with HTTP 202 and Idempotency-Replayed: true, including after a gateway restart and after the run has completed, failed, or been cancelled. Reusing the same key with a different JSON payload returns HTTP 409 with code idempotency_key_conflict. Keys are isolated by authenticated API profile/credential and retained for 24 hours after their last status update; clients should use unique, unguessable keys and must not reuse them for unrelated operations. Requests without the header retain the legacy behavior and always create a new run.
When session_id identifies an existing Mibyan session and no explicit
conversation_history or previous_response_id is supplied, the run loads
that session’s active transcript. Session turn leases serialize concurrent
writers and refresh the transcript after a contended wait.
GET /v1/runs/{run_id}
Poll the current run state. This is useful for dashboards that need status without holding an SSE connection open, or for UIs that reconnect after navigation.
model echoes what the request asked for. On a completed run, runtime is the provider/model pair that actually served the turn — after a fallback provider switch it names the fallback pair, so a cost-attribution poller books the run to the right provider. usage.cache_read_tokens / usage.cache_write_tokens are the session’s prompt-cache reads and writes, so cached input is not priced as full-price input. runtime has the same shape as on /v1/chat/completions and /v1/responses: route_source says how the runtime was chosen (global, raw_request, model_routes), and a request that named a model/provider also gets requested: {provider, model} so the asked-for and served pairs can be compared. The run.completed event on the events stream carries the same usage and runtime fields.
Statuses are retained briefly after terminal states (completed, failed, cancelled, or interrupted) for polling and UI reconciliation. When the gateway shuts down while a run is active, the run is persisted as interrupted (error Gateway shutdown interrupted the run., terminal event run.interrupted) before the agent is asked to stop, so a durable run never survives a restart as running; a late result from the interrupted turn cannot overwrite it.
While the gateway is still draining (a mibyan gateway stop/restart or SIGTERM with a turn in flight), every non-terminal run additionally carries shutdown_requested_at (Unix seconds) from the moment new turns are refused. status stays running because the turn is still being served; a poller that sees the field knows the process is on its way out and the run will end interrupted at the latest when the drain budget expires. Terminal runs never gain the field.
GET /v1/runs/{run_id}/events
Server-Sent Events stream of the run’s tool-call progress, token deltas, and lifecycle events. Designed for dashboards and thick clients that want to attach/detach without losing state.
Mid-turn assistant commentary — the openai-codex backend’s phase="commentary" progress
preambles, or text a model writes alongside its tool calls — arrives as message.interim
(text, already_streamed), the same contract the TUI gateway uses. already_streamed: true
means the text also went out as message.delta, so clients that render deltas can skip it.
The final answer still arrives only in run.completed; private reasoning never becomes
message.interim. Gate: display.interim_assistant_messages (default true).
Tool lifecycle events carry tool.started (tool, preview of the arguments) and
tool.completed (tool, duration in seconds, error, and a preview of the result). The
error flag reflects the tool’s own outcome — a non-zero terminal exit_code, a structured
{"error": ...} result, a denied approval — whether the result arrives as a JSON string or an
already-parsed object. The completion preview is the result text (structured results are
JSON-encoded), passed through forced secret redaction and then truncated to 500 characters, so a
client can tell an approval refusal (BLOCKED: ...) from an ordinary failure without receiving
the unbounded tool payload.
When the agent delegates work to background subagents, the stream also carries
subagent.start and subagent.complete lifecycle events, so clients can
observe delegation outcomes — including timeouts and failures — instead of the
run going silent while a child works. The subagent.complete payload carries
the child’s status, summary, duration, token/cost figures, a
child_session_id for correlation, and the delegation_id of the batch it
belongs to (so concurrent or nested fan-outs stay distinguishable); free-text fields pass forced secret
redaction before leaving the process. Per-tool child events
(subagent.tool, progress ticks) are intentionally not forwarded — they
are high-volume UI noise; use the per-child live transcript files for
play-by-play. These events are available while the parent stream is open; a
late detached completion does not reopen a finished run’s SSE stream or change
its terminal status.
Detached results and session history
Background delegation requires a continuation that reads server-side session
history: an explicit X-Mibyan-Session-Id on Chat Completions, a native
/api/sessions/{id}/chat request, or a Runs request using session history.
Header-less Chat Completions, Responses chains, and Runs requests with
previous_response_id or caller-supplied history instead execute delegation
synchronously, returning the result in the original turn. Merely deriving a
session ID from request content does not enable detached delivery.
For resumable requests, the completion is persisted once per delegation unit.
It is available through GET /api/sessions/{id}/messages and in the next real
client turn’s session history. Retries do not insert the same result again;
interim task-failure notices have separate identities. Delivery waits while a
client turn owns the session lease and follows compression continuations.
Chat Completions echoes the explicit session ID you supplied in both JSON and
streaming responses; keep sending that ID even after compression.
A completion never starts an unsolicited model turn or bypasses a pending
human confirmation. The client owns the next turn. Clients that continue using
their own history snapshots should use synchronous delegation rather than
expecting a server-side delivery row to be merged into those snapshots.
Unconsumed event buffers expire after five minutes so a detached client cannot
grow memory indefinitely. This expires transport state only: a run that is
still executing remains visible to status polling, approval, stop control, and
concurrency accounting until its executor work actually exits. A connected SSE
subscriber continues draining normally.
POST /v1/runs/{run_id}/stop
Interrupt a running agent turn. The endpoint returns immediately with {"status": "stopping"} while Mibyan asks the active agent to stop at the next safe interruption point.
The run stays tracked as stopping until the executor-backed work exits, then
settles as cancelled; requesting stop never hides a worker that is still
running.
POST /v1/runs/{run_id}/approval
Resolve a pending approval for a run that is waiting on a human decision (for example, a tool call gated behind an approval policy). The body carries the approval decision; the run resumes once the decision is recorded. This endpoint is advertised in /v1/capabilities as the run_approval feature so external UIs can detect support before surfacing an approval prompt.
MCP trust-gate consent — a write-capable tool on a server configured trust: untrusted — surfaces the same way: the run emits an approval.request event and parks in waiting_for_approval until this endpoint resolves it (once runs the tool, deny blocks it).
Jobs API (background scheduled work)
The server exposes a lightweight jobs CRUD surface for managing scheduled / background agent runs from a remote client. All endpoints are gated behind the same bearer auth.
GET /api/jobs
List all scheduled jobs.
POST /api/jobs
Create a new scheduled job. Body accepts the same shape as mibyan cron — prompt, schedule, skills, provider override, delivery target.
GET /api/jobs/{job_id}
Fetch a single job’s definition and last-run state.
PATCH /api/jobs/{job_id}
Update fields on an existing job (prompt, schedule, etc.). Partial updates are merged.
DELETE /api/jobs/{job_id}
Remove a job. Also cancels any in-flight run.
POST /api/jobs/{job_id}/pause
Pause a job without deleting it. Next-scheduled-run timestamps are suspended until resumed.
POST /api/jobs/{job_id}/resume
Resume a previously paused job.
POST /api/jobs/{job_id}/run
Trigger the job to run immediately, out of schedule.
Sessions API (session control over REST)
External UIs can manage Mibyan sessions over REST without standing up the dashboard. All endpoints are gated by API_SERVER_KEY and live under /api/sessions/*.
/v1/capabilities advertises the full surface via session_* feature flags and endpoints.session_* entries so external UIs can detect support and fall back safely. Inline images are supported in chat and chat/stream payloads (multimodal-aware path).
GET /v1/skills and GET /v1/toolsets let external clients enumerate the agent’s capabilities deterministically over REST instead of asking the model. Both are read-only and gated by API_SERVER_KEY.
/v1/skills returns the same metadata the skills hub uses internally. /v1/toolsets returns toolsets resolved for the api_server platform with the concrete tools list each one expands to. Both are advertised under endpoints.* in /v1/capabilities.
Long-term memory scoping (X-Mibyan-Session-Key)
Multi-user frontends like Open WebUI need a stable per-channel identifier for long-term memory (Honcho, etc.) that is independent of the transcript-scoped X-Mibyan-Session-Id (which rotates on /new). Pass X-Mibyan-Session-Key on /v1/chat/completions, /v1/responses, or /v1/runs and Mibyan threads it through to AIAgent(gateway_session_key=...), where the Honcho memory provider uses it to derive a stable scope.
Rules: max 256 chars, control characters (\r, \n, \x00) are rejected, and the value is echoed back on responses (JSON + SSE). /v1/capabilities advertises support via "session_key_header": "X-Mibyan-Session-Key". Without the key, Honcho’s per-session strategy produces a different scope per session_id — exactly the behavior Mibyan had before.
Automatic recall follows the transcript: the memory provider is initialised once per session and kept across requests, so a continued session (X-Mibyan-Session-Id, previous_response_id, or a declared X-Mibyan-Session-Key conversation) receives the recall the provider prepared after the previous turn, exactly like a Telegram or Discord chat does. A request without any continuation starts a fresh session and, like the first turn of any new CLI session, has nothing queued yet. Idle sessions release their provider after the same idle TTL as the gateway agent cache (agent.agent_cache.idle_ttl_secs, default one hour).
System Prompt Handling
When a frontend sends a system message (Chat Completions) or instructions field (Responses API), mibyan-agent layers it on top of its core system prompt. Your agent keeps all its tools, memory, and skills — the frontend’s system prompt adds extra instructions.
This means you can customize behavior per-frontend without losing capabilities:
- Open WebUI system prompt: “You are a Python expert. Always include type hints.”
- The agent still has terminal, file tools, web search, memory, etc.
Authentication
Bearer token auth via the Authorization header:
Configure the key via API_SERVER_KEY env var. If you need a browser to call Mibyan directly, also set API_SERVER_CORS_ORIGINS to an explicit allowlist.
Multi-profile routing (/p/<profile>/…)
When multi-profile gateway routing is
enabled (gateway.multiplex_profiles), the shared listener serves every
profile through a /p/<profile>/ URL prefix — and authentication is bound
to the routed profile:
- Requests to
/p/<profile>/v1/... must present that profile’s own
API_SERVER_KEY (from ~/.mibyan/profiles/<profile>/.env). The default
listener’s key is rejected on named-profile prefixes.
- Unprefixed routes and
/p/default/... keep using the default profile’s key.
- A named profile with no
API_SERVER_KEY of its own fails closed — its
prefix is unreachable until you set one.
- Runs are per-profile scoped:
/v1/runs/{run_id} and its events, stop,
steer, and approval routes only answer for the profile that created
the run (including runs started via /api/sessions/{id}/chat/stream);
another profile’s run id returns 404, never 403.
Breaking change (July 2026)Before this fix, a valid default-profile key was accepted on any
/p/<profile>/ prefix. If you relied on one shared key across profile
prefixes, set a distinct API_SERVER_KEY in each profile’s .env — reused
default keys on named prefixes now return 401.
SecurityThe API server gives full access to mibyan-agent’s toolset, including terminal commands. API_SERVER_KEY is required for every deployment, including the default loopback bind on 127.0.0.1. Keep API_SERVER_CORS_ORIGINS narrow to control browser access when you explicitly allow browser callers.
Configuration
Environment Variables
config.yaml
The same settings can live in ~/.mibyan/config.yaml under a nested gateway.api_server: section:
port, key, host, cors_origins, and model_name are automatically bridged into the platform’s extra settings, so they behave exactly like their API_SERVER_* environment-variable counterparts. Environment variables take precedence over config.yaml values. The block is also accepted under gateway.platforms.api_server: or a top-level platforms.api_server: section.
Concurrent-run cap
The API server limits how many agent runs may execute at once across the endpoints that start one directly: the OpenAI-compatible endpoints, the Runs endpoints, and the session-chat endpoints (POST /api/sessions/{id}/chat and its /stream variant, which carry cross-machine agent DMs). Cron-triggered runs (POST /api/jobs/{id}/run, POST /api/cron/fire) go through the cron scheduler and are governed by cron’s own limits, not this cap. The cap is read from gateway.api_server.max_concurrent_runs (default 10; 0 disables the limit, negative values clamp to 0). When the cap is reached, new run-starting requests are rejected with HTTP 429 Too many concurrent runs (max N) — clients should back off and retry.
Stored history size for previous_response_id chaining
Each stored /v1/responses snapshot embeds the full cumulative conversation history (that is what previous_response_id and conversation chaining replay), including every tool output verbatim. A conversation with a few large tool outputs can therefore make a single response_store.db write several hundred KB. Set gateway.api_server.history_tool_output_max_chars (default 0 = store verbatim) to cap each tool output and each string tool-call argument in the stored history at that many characters; anything longer is cut to the head plus a ...[N more chars] marker. User and assistant text is never touched, and the response.completed payload and incremental SSE events are unaffected. Because the stored history is what the model sees on the next chained turn, enabling the cap also trims what the model is replayed — leave it at 0 if your workflow needs complete tool outputs across turns.
All responses include security headers:
X-Content-Type-Options: nosniff — prevents MIME type sniffing
Referrer-Policy: no-referrer — prevents referrer leakage
CORS
The API server does not enable browser CORS by default.
For direct browser access, set an explicit allowlist:
When CORS is enabled:
- Preflight responses include
Access-Control-Max-Age: 600 (10 minute cache)
- SSE streaming responses include CORS headers so browser EventSource clients work correctly
X-Mibyan-Session-Id is an allowed request header, so browsers on an allowlisted origin can request session continuation.
Idempotency-Key is an allowed request header — clients can send it for deduplication (responses are cached by key for 5 minutes)
Most documented frontends such as Open WebUI connect server-to-server and do not need CORS at all.
Compatible Frontends
Any frontend that supports the OpenAI API format works. Tested/documented integrations:
Multi-User Setup with Profiles
To give multiple users their own isolated Mibyan instance (separate config, memory, skills), use profiles:
Each profile’s API server automatically advertises the profile name as the model ID:
http://localhost:8643/v1/models → model alice
http://localhost:8644/v1/models → model bob
In Open WebUI, add each as a separate connection. The model dropdown shows alice and bob as distinct models, each backed by a fully isolated Mibyan instance. See the Open WebUI guide for details.
Limitations
- Response storage — stored responses (for
previous_response_id) are persisted in SQLite and survive gateway restarts. Max 100 stored responses (LRU eviction).
- No file upload — inline images are supported on both
/v1/chat/completions and /v1/responses, but uploaded files (file, input_file, file_id) and non-image document inputs are not supported through the API.
- Simple OpenAI clients still see an alias —
/v1/models advertises the
stable Mibyan alias (mibyan-agent or the active profile name). Richer
clients can send explicit provider / model_options overrides on requests.
Proxy Mode
The API server also serves as the backend for gateway proxy mode. When another Mibyan gateway instance is configured with GATEWAY_PROXY_URL pointing at this API server, it forwards all messages here instead of running its own agent. This enables split deployments — for example, a Docker container handling Matrix E2EE that relays to a host-side agent.
See Matrix Proxy Mode for the full setup guide.