Built-in tools and explicit deferralMibyan keeps its working-set core tools (
terminal, read_file, write_file,
patch, search_files, todo, memory, browser_*, web_search,
web_extract, clarify, execute_code, delegate_task, and the rest of
_mibyan_CORE_TOOLS) loaded directly by default. Cold, event-triggered built-ins
may be deferred when they are named in tools.tool_search.defer; the shipped
curated list covers tools such as computer_use, session_search, and selected
desktop helpers. MCP and non-core plugin tools remain eligible automatically.
An explicit defer list replaces the curated list, and defer: [] keeps every
tool eager.How it works
When Tool Search activates for a turn, the model sees three new tools in place of the deferred ones:calls takes one entry per invocation; a single local call is an array of
one. Only connectors__ names may be batched together; mixed and
multi-local batches are rejected.
A typical interaction looks like:
tool_search call is searched independently against the
same catalog (limit applies per query); the per-query groups carry tool
names only, while the shared tools map holds each matched tool’s
description and required parameter names once. Queries are stemmed, so
“issues” finds create_issue. Each query group that returns no matches
includes an available_sources summary of the connected servers so a lexical
miss is not mistaken for a missing capability.
tool_describe resolves every requested name in one call; unknown names
are reported in not_found without failing the rest of the batch.
When the model invokes tool_call, Mibyan unwraps the bridge and
dispatches the underlying tool exactly as if the model had called it
directly. Pre-tool-call hooks, guardrails, approval prompts, and
post-tool-call hooks all run against the real tool name — not against
tool_call. The activity feed in the CLI and gateway also unwraps so you
see the underlying tool, not the bridge.
When does it activate?
Tool Search uses tiered disclosure: the presence of any deferred tool (MCP/plugin or explicitly named built-in) activates the bridge; what scales with catalog size is how much of the catalog stays visible, not whether schemas defer.
The listing budget is
min(threshold_pct% of context, listing_max_tokens).
The decision is re-evaluated every time the tools array is built, so
adding or removing MCP servers mid-session moves the session between
tiers on the next assembly.
Configuration
defer list also includes the selected desktop GUI helpers listed
in mibyan_cli/config_defaults.py. It is the single source of truth for the
shipped curated set; the runtime fallback uses the same value.
Per-call array caps are internal safety bounds, not configuration. Over-cap
calls return an error so the model can retry with a smaller batch.
Why the listing exists
Without it, deferred capabilities are invisible — live benchmarking showed models substituting visible core tools (runninggh in the terminal instead
of searching for the deferred GitHub tool) or declaring a capability
nonexistent instead of calling tool_search. The listing applies the skills
pattern to tools: every capability stays discoverable by name at all times,
while full parameter schemas remain deferred. If the model sees the exact
tool name in the listing, it can skip tool_search and go straight to
tool_describe, saving a round trip.
You can also flip the legacy boolean shape:
Connectors (remote tools)
When you are signed in to the Nous Portal, the bridge additionally reaches connectors — remote tools served by the managed tool gateway. They are never registered locally:tool_search sends each query to the gateway, adds
the gateway’s hits to the local catalog as documents (tagged
source: "connectors", named connectors__<connector>__<tool>), and ranks
both with the same BM25 pass and the same rarest-token rule, so limit
caps the group as a whole and a connector tool that answers the query is
never pushed out by local tools that share one word with it. The gateway
call is bounded at 30 seconds; a slow or dark gateway degrades to local
results only. tool_describe fetches connector schemas from the gateway,
and tool_call sends each connector entry in a batch as its own gateway
request, in input order (a tool name the gateway does not know under its
conventional slug is retried once under the literal slug, so an entry can
cost two requests). If a connector ever shipped both GMAIL_X and a literal
X, both would compose to connectors__gmail__X, which runs GMAIL_X;
search keeps that twin, drops the other, and logs a warning. Results splice back into the batch’s original order
with recomputed counts.
CONNECTION_REQUIRED error. The manage_connections tool lists connectors and
their connection state and starts an authorization: in the desktop app the call
shows a card, blocks until each app is connected or skipped, and reports the
outcomes; elsewhere it returns a connect link per app for the user to open.
Disconnecting an account is done by the user in the Portal. The same tool also installs, enables and authorizes
local MCP servers from the catalog (targets with mcp: true), so it is
present whether or not you are signed in; only the managed-connector actions
need the sign-in.
The desktop backend’s account-list and disconnect APIs use the Portal’s
account-management service, including its organization membership checks and
disconnect audit. An unavailable Portal does not fall back to direct gateway
account management. Tool discovery, execution, and connection-status watching
continue through the gateway; the model tool cannot disconnect an account.
tool_call accepts a batch: calls is an array of {name, arguments}
entries (a single call is an array of one). Each connector entry in a batch
is dispatched as its own gateway request, one after another; local deferred
tools stay one entry per tool_call. A multi-entry batch that names a local
tool is rejected with a correction that restates the valid shape using the
caller’s own first entry, and a calls value emitted as a JSON string is
parsed like the array form. Approvals settle per entry before
dispatch, and a /stop between entries leaves the unstarted ones unsent
(their slots report INTERRUPTED).
When NOT to use it
Tool Search trades a fixed per-turn token cost (the three bridge tool schemas plus the catalog listing) and at least one extra round trip on cold tools (describe → call) for the savings on the deferred schemas. At tier 1 the listing keeps every capability visible, so the discovery round trip usually disappears — the model goes straight totool_describe. Live benchmarking showed the listing mode matching
eager loading’s task success while costing less than the bare bridge.
If you want the old always-eager behavior for a small toolset, set
enabled: off.
Trade-offs that don’t go away
These come from the prompt-cache integrity invariant — they are inherent to any progressive-disclosure design, not specific to this implementation:- One extra round trip on cold tools. The first time the model needs a deferred tool, it spends one or two extra model calls to find and load the schema. The token savings on the static side are real, but a portion is paid back at runtime.
- No cache benefit on deferred schemas. A loaded
tool_describeresult enters the conversation history (so it does get cached on subsequent turns) but it never benefits from the system-prompt cache prefix. - No provider-native validation for deferred schemas.
tool_describelets the model read a deferred tool’s schema, but the provider still sees only the generictool_call.argumentsobject. Mibyan therefore coerces and validates the underlying arguments locally before dispatch; the concrete tool or MCP server remains responsible for schemas Mibyan cannot safely validate, such as malformed schemas or external references. - Model-quality dependence. Tool Search assumes the model can write a reasonable search query for the tool it wants. Smaller models do this less well; the published Anthropic numbers (49% → 74% on Opus 4 with vs. without tool search) show the upside but also that ~26 points of accuracy is still retrieval failure.
- Toolset edits invalidate cache. Adding or removing a tool mid- session changes the bridge tools’ descriptions (which include the count of deferred tools) and the catalog, so the prompt cache is invalidated. This is the same trade-off as any toolset edit.
Implementation details
- Retrieval: BM25 over tokenized tool name, source name (the MCP
server or plugin toolset the tool belongs to, so searching
"linear"finds that server’s tools even when a tool’s own name doesn’t carry the service), description, and parameter names, with Snowball stemming (English) applied to both the index and the query so morphological variants match (“issues” findscreate_issue). A tool is a result only if it contains the query’s rarest token (the one in the fewest tool documents, so the word that names the intent:gmail,github,incident, notsendorcreate). A query whose rarest token appears in no tool returns an empty group with the connected sources and a retry hint, instead oflimittools that share one common word. - Relevance floor: a tool must match at least half of a query’s answerable terms (terms present anywhere in the catalog) before it is offered — sharing one incidental word with a long query is not a match. A hunt for a capability that doesn’t exist returns no results instead of a plausible-looking list the model rephrases against forever. The floor only engages from four answerable terms up, so short queries like “list issues” keep full recall, and an exact tool-name query always matches.
- Parallel execution unwraps the bridge. The batch planner decides
concurrency on the underlying tool of a
tool_call, not on the literal bridge name — so an MCP server opted in viasupports_parallel_tool_calls: truekeeps its concurrency when its tools are called through the bridge, andtool_search/tool_describelookups batch concurrently like any read-only tool. - Catalog is stateless across turns. It rebuilds from the current
tool-defs list every assembly — no session-keyed
Map. This avoids the class of bug where a stored catalog drifts out of sync with the live tool registry. - The catalog is scoped to the session’s toolsets.
tool_search,tool_describe, andtool_callonly ever see and invoke tools the session was actually granted. A subagent, kanban worker, or gateway session restricted to a subset of toolsets cannot use the bridge to discover or call a tool outside that subset — the deferred catalog is the deferrable slice of the session’s own enabled/disabled toolsets, not the whole process registry. - No JS sandbox. Mibyan uses the simpler “structured tools” mode (search / describe / call as plain functions). The JS-sandbox “code mode” some other implementations offer is a large surface area; we skip it.
See also
tools/tool_search.py— the implementationtests/tools/test_tool_search.py— the regression suite- The
openclaw-tool-search-reportPDF in the original implementation PR for the research that shaped the design

