Commands, package names, and image names on this page come from the open-source project that Mibyan Desktop is built on, and can differ from the Mibyan Desktop installer. For the supported Mibyan install and update path, see Install and update.
ctx.llm is the supported way for a plugin to make an LLM call.
Chat completion, structured extraction, sync, async, with or without
images — same surface, same trust gate, same host-owned credentials.
Plugins reach for this when they need to do something that involves
the model but isn’t part of the agent’s conversation. A hook that
rewrites a tool error into something a non-engineer can read. A
gateway adapter that translates an inbound message before queuing
it. A slash command that summarises a long paste. A scheduled job
that scores yesterday’s activity and writes one line to a status
board. A pre-filter that decides whether a message is worth waking
the agent up for at all.
These are jobs the agent shouldn’t be in the loop on. They want one
LLM call, a typed answer, and to be done.
The smallest possible call
A more complete chat example
purpose is a free-form audit string — it shows up in agent.log
and in result.audit so operators can see which plugin made which
call. Optional but recommended for anything that fires often.
Structured output
When the plugin needs a typed answer, switch to the structured lane:jsonschema is
installed, and hands back a Python object on result.parsed. If the
model couldn’t produce valid JSON, result.parsed is None and
result.text carries the raw response.
What this lane gives you
- One call, four shapes.
complete()for chat,complete_structured()for typed JSON,acomplete()andacomplete_structured()for asyncio. Same arguments, same result objects. - Host-owned credentials. OAuth tokens, refresh flows, the
credential pool, per-task aux overrides — every credential
concept Mibyan already has applies. The plugin never sees a
token; the host attributes the call back through
result.audit. - Bounded. Single sync or async call. No streaming, no tool loops, no conversation state to manage. State the input, get the result, return.
- Fail-closed trust. A plugin you’ve never configured cannot
pick its own provider, model, agent, or stored credential. The
default posture is “use what the user is using.” Operators opt in
to specific overrides, per plugin, in
config.yaml.
Quick start
Two complete plugins below — one chat, one structured. Both ship inside a singleregister(ctx) function and need zero outside
configuration to run against whatever model the user has active.
Chat completion — /tldr
result.text is the model’s response; result.usage carries token
counts; result.provider and result.model carry attribution.
Structured extraction — /paste-to-tasks
mibyan-example-plugins
repo (companion repo for reference plugins — not bundled with
mibyan-agent itself). For the async surface (acomplete() /
acomplete_structured() with asyncio.gather()), see
plugin-llm-async-example
in the same repo.
When to use which
Everything else — provider selection, model resolution, auth, fallback,
timeout, vision routing — is the same across all four.
API surface
ctx.llm is an instance of agent.plugin_llm.PluginLlm.
complete()
messages is the standard OpenAI shape — a
list of {"role": "...", "content": "..."} dicts. Multi-turn
prompts (system + few-shot user/assistant pairs + final user) work
exactly as they would with the OpenAI SDK.
provider= and model= are independent and follow the same shape
as the host’s main config (model.provider + model.model). Set
just model= to use the user’s active provider with a different
model on it. Set both to switch providers entirely. Either argument
without operator opt-in raises PluginLlmTrustError.
complete_structured()
data: URL automatically). When json_schema or
json_mode=True is supplied, the host requests JSON output via
response_format, parses it locally as a fallback, and validates
against your schema if jsonschema is installed.
result.content_type == "json"—result.parsedis a Python object that matches your schema.result.content_type == "text"— parsing or validation failed; inspectresult.textfor the raw model response.
Async
Task-routed auxiliary calls
Passtask= to any of the four call shapes when a plugin needs its
own configured auxiliary route. Register that task during plugin setup;
plugin defaults apply until the operator overrides its provider and model
in auxiliary.<task>:
auxiliary.<task> overrides those
defaults and controls the deployment choice. A plugin can only use a task
it registered itself; unknown or foreign task names fail before provider
invocation. allow_task_override: true is an explicit operator grant for
using Mibyan built-in auxiliary tasks; it does not permit another plugin’s
tasks. Omit task= (or use "auto") to keep the active main provider/model.
Result attributes
usage carries input_tokens, output_tokens, total_tokens,
cache_read_tokens, cache_write_tokens, and cost_usd when the
provider returns those fields.
Trust gate
The default behaviour is fail-closed. With noplugins.entries
config block, a plugin can:
- run any of the four methods against the user’s active provider and model,
- set request-shaping arguments (
temperature,max_tokens,timeout,system_prompt,purpose,messages,instructions,input,json_schema),
provider=, model=, agent_id=, and profile=
arguments raise PluginLlmTrustError until the operator opts in.
Likewise, task= can use only the plugin’s registered auxiliary task
unless the operator grants allow_task_override for a built-in task.
Most plugins never need this section. A plugin that just calls
ctx.llm.complete(messages=...) with no overrides runs against
whatever the user has active and works zero-config. The block below
is only relevant when a plugin specifically wants to pin to a
different model or provider than the user.
name: field for flat plugins, or the
path-derived key for nested plugins (image_gen/openai,
memory/honcho, etc.).
What the gate enforces
Each override is independently gated. Granting
allow_model_override
does not also grant allow_provider_override — a plugin trusted
to pick a model is still pinned to the user’s active provider unless
it gets the provider gate as well.
What the gate does NOT need to enforce
- Request-shaping arguments —
temperature,max_tokens,timeout,system_prompt,purpose,messages,instructions,input,json_schema,schema_name,json_mode— are always allowed; they don’t pick credentials or routes. - The default deny posture means an unconfigured plugin can still do
useful work — it just runs against the active provider and model.
Operators only need to think about
plugins.entriesfor plugins that want finer routing.
What the host owns
A complete list of the thingsctx.llm does for the plugin so you
don’t have to:
- Provider resolution. Reads
model.provider+model.modelfrom the user’s config (or the explicit overrides when trusted). - Auth. Pulls API keys, OAuth tokens, or refresh tokens from
~/.mibyan/auth.json/ env, including the credential pool when one is configured. The plugin never sees them. - Vision routing. When image input is supplied and the user’s active text model is text-only, the host falls back to the configured vision model automatically.
- Fallback chain. If the user’s primary provider 5xxs or 429s, the request goes through Mibyan’ usual aggregator-aware fallback before it returns an error to the plugin.
- Timeout. Honours your
timeout=argument, falling back toauxiliary.<task>.timeoutconfig or the global aux default. - JSON shaping. Sends
response_formatto the provider when you ask for JSON, then re-parses locally from a code-fenced response if the provider returned one. - Schema validation. Validates against your
json_schemawhenjsonschemais installed; logs a debug line and skips strict validation otherwise. - Audit log. Each call writes one INFO line to
agent.logwith the plugin id, provider/model, purpose, and token totals.
What the plugin owns
- Request shape.
messagesfor chat,instructions+inputfor structured. The plugin builds the prompt; the host runs it. - Schema. Whatever shape you want back. The host doesn’t infer it for you.
- Error handling.
complete_structured()raisesValueErroron empty inputs and on schema-validation failure.PluginLlmTrustErrorfires when the trust gate denies an override. Anything else (provider 5xx, no credentials configured, timeout) raises whateverauxiliary_client.call_llm()raises. - Cost. Every call runs against the user’s paid provider. Don’t
loop on
complete()for every gateway message without thinking about token spend.
Where this fits in the plugin surface
Existingctx.* methods extend an existing Mibyan subsystem:
| ctx.register_tool | adds a tool the agent can call |
| ctx.register_platform | wires a new gateway adapter |
| ctx.register_image_gen_provider | replaces an image-gen backend |
| ctx.register_memory_provider | replaces the memory backend |
| ctx.register_context_engine | replaces the context compressor |
| ctx.register_hook | observes a lifecycle event |
ctx.llm is the first surface that lets a plugin run the same
model the user is talking to, out of band, without any of the
above. That’s its only job. If your plugin needs to register a
tool the agent invokes, use register_tool. If it needs to react
to a lifecycle event, use register_hook. If it needs to make its
own model call — for any reason, structured or not — ctx.llm.
Reference
- Implementation:
agent/plugin_llm.py - Tests:
tests/agent/test_plugin_llm.py - Reference plugins (companion repo):
plugin-llm-example— sync structured extraction with image inputplugin-llm-async-example— async withasyncio.gather()
- Auxiliary client (the engine under the hood): see Provider Runtime.

