> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mibyanai.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Tool Search

When you have many MCP servers or non-core plugin tools attached to a
session, their JSON schemas can consume a substantial fraction of the
context window on every turn — even when only a few of them are relevant
to what the user actually asked for.

**Tool Search** is Mibyan' opt-in progressive-disclosure layer for that
problem. When activated, MCP and plugin tools are replaced in the
model-visible tools array by three bridge tools, and the model loads each
specific tool's schema on demand.

<Info>
  **Built-in tools and explicit deferral**

  Mibyan keeps its working-set core tools (`terminal`, `read_file`, `write_file`,
  `patch`, `search_files`, `todo`, `memory`, `browser_*`, `web_search`,
  `web_extract`, `clarify`, `execute_code`, `delegate_task`, and the rest of
  `_mibyan_CORE_TOOLS`) loaded directly by default. Cold, event-triggered built-ins
  may be deferred when they are named in `tools.tool_search.defer`; the shipped
  curated list covers tools such as `computer_use`, `session_search`, and selected
  desktop helpers. MCP and non-core plugin tools remain eligible automatically.
  An explicit `defer` list replaces the curated list, and `defer: []` keeps every
  tool eager.
</Info>

## How it works

When Tool Search activates for a turn, the model sees three new tools in
place of the deferred ones:

```
tool_search(queries, limit?)   search the deferred-tool catalog (one or more queries)
tool_describe(names)           load the full schemas for one or more tools
tool_call(calls)               invoke deferred tools; `calls` is an array of {name, arguments}
```

`calls` takes one entry per invocation; a single local call is an array of
one. Only `connectors__` names may be batched together; mixed and
multi-local batches are rejected.

A typical interaction looks like:

```
Model: tool_search(["create a github issue", "send a slack message"])
  → { results: [ { query: "create a github issue",
                   matches: ["mcp_github_create_issue", ...] },
                 { query: "send a slack message",
                   matches: ["mcp_slack_post_message", ...] } ],
      tools: { mcp_github_create_issue: { description: "...",
                                          required: ["title"], ... },
               mcp_slack_post_message: { ... } } }
Model: tool_describe(["mcp_github_create_issue", "mcp_slack_post_message"])
  → { tools: { mcp_github_create_issue: { parameters: { ... } },
               mcp_slack_post_message: { parameters: { ... } } } }
Model: tool_call({ calls: [{ name: "mcp_github_create_issue",
                             arguments: { title: "...", body: "..." } }] })
  → { ok: true, issue_number: 42 }
```

Each query in a `tool_search` call is searched independently against the
same catalog (`limit` applies per query); the per-query groups carry tool
names only, while the shared `tools` map holds each matched tool's
description and required parameter names once. Queries are stemmed, so
"issues" finds `create_issue`. Each query group that returns no matches
includes an `available_sources` summary of the connected servers so a lexical
miss is not mistaken for a missing capability.
`tool_describe` resolves every requested name in one call; unknown names
are reported in `not_found` without failing the rest of the batch.

When the model invokes `tool_call`, Mibyan **unwraps the bridge** and
dispatches the underlying tool exactly as if the model had called it
directly. Pre-tool-call hooks, guardrails, approval prompts, and
post-tool-call hooks all run against the real tool name — not against
`tool_call`. The activity feed in the CLI and gateway also unwraps so you
see the underlying tool, not the bridge.

## When does it activate?

Tool Search uses **tiered disclosure**: the presence of *any* deferred
tool (MCP/plugin or explicitly named built-in) activates the bridge; what
scales with catalog size is how much of the catalog stays visible, not whether
schemas defer.

| Tier | Condition | What the model sees |
| - | - | - |
| **0** | No deferred tools | Every tool eager, no bridge. Pass-through. |
| **1** | Deferred catalog's listing fits the budget | Bridge + a skills-style manifest of every deferred tool (name + short description, degrading to names-only when over budget). Degradation is **per server**: when one oversized server (Cloudflare) is attached alongside small ones (Linear), the small servers keep their per-tool listings and only the oversized server collapses to a summary line. |
| **2** | Per-tool listing exceeds the budget even names-only for every server (e.g. Cloudflare's flat API surface alone: \~3,300 tools whose names are \~32K tokens) | Bare bridge + a one-line-per-server summary (server name + tool count), so the model knows which domains are reachable; individual tools are discoverable only through `tool_search`. |

The listing budget is `min(threshold_pct% of context, listing_max_tokens)`.
The decision is re-evaluated every time the tools array is built, so
adding or removing MCP servers mid-session moves the session between
tiers on the next assembly.

## Configuration

```yaml theme={null}
tools:
  tool_search:
    enabled: auto       # auto (default), on, or off
    threshold_pct: 5    # listing budget as a percentage of context
    search_default_limit: 5
    max_search_limit: 25
    listing: auto       # embed a grouped name+description catalog manifest
    listing_max_tokens: 4000
    defer:              # replace the curated default; [] keeps every tool eager
      - computer_use
      - session_search
      - image_generate
      - todo_list
      - process_manage
      - cronjob_manage
```

The default `defer` list also includes the selected desktop GUI helpers listed
in `mibyan_cli/config_defaults.py`. It is the single source of truth for the
shipped curated set; the runtime fallback uses the same value.

| Key | Default | Meaning |
| - | - | - |
| `enabled` | `auto` | `auto`/`on` activate whenever at least one deferred tool exists; `off` disables entirely (everything stays eager). `auto` is currently an alias of `on` — it is reserved for a future mode that inlines schemas when they fit the context and defers only when they don't. Pin `on` or `off` if you want today's behavior guaranteed across upgrades. |
| `threshold_pct` | `5` | Listing budget as a percentage of the active model's context length. Range 0–100. |
| `search_default_limit` | `5` | Hits returned per query when the model calls `tool_search` without a `limit`. |
| `max_search_limit` | `25` | Hard upper bound the model can request via `limit` (per query). Range 1–50. |
| `listing` | `auto` | Embed a skills-style manifest of every deferred tool (name + first sentence of its description, ≤60 chars, grouped by MCP server) in the `tool_search` bridge description. `auto` includes it when it fits the budget (falling back to names-only, then to the tier-2 server summary); `on`/`off` force either way. |
| `listing_max_tokens` | `4000` | Absolute cap on the embedded listing, regardless of context size. Range 200–60000. Large catalogs degrade to names-only or per-server summaries, keeping full schemas available through search. |
| `defer` | Curated list | Tool names replaced by the bridge by default. The list may include cold built-in tools as well as MCP/plugin tools; an explicit list replaces it, and `[]` disables deferral for every tool. |

Per-call array caps are internal safety bounds, not configuration. Over-cap
calls return an error so the model can retry with a smaller batch.

### Why the listing exists

Without it, deferred capabilities are *invisible* — live benchmarking showed
models substituting visible core tools (running `gh` in the terminal instead
of searching for the deferred GitHub tool) or declaring a capability
nonexistent instead of calling `tool_search`. The listing applies the skills
pattern to tools: every capability stays discoverable by name at all times,
while full parameter schemas remain deferred. If the model sees the exact
tool name in the listing, it can skip `tool_search` and go straight to
`tool_describe`, saving a round trip.

You can also flip the legacy boolean shape:

```yaml theme={null}
tools:
  tool_search: true   # equivalent to {enabled: auto}
```

## Connectors (remote tools)

When you are signed in to the Nous Portal, the bridge additionally reaches
**connectors** — remote tools served by the managed tool gateway. They are
never registered locally: `tool_search` sends each query to the gateway, adds
the gateway's hits to the local catalog as documents (tagged
`source: "connectors"`, named `connectors__<connector>__<tool>`), and ranks
both with the same BM25 pass and the same rarest-token rule, so `limit`
caps the group as a whole and a connector tool that answers the query is
never pushed out by local tools that share one word with it. The gateway
call is bounded at 30 seconds; a slow or dark gateway degrades to local
results only. `tool_describe` fetches connector schemas from the gateway,
and `tool_call` sends each connector entry in a batch as its own gateway
request, in input order (a tool name the gateway does not know under its
conventional slug is retried once under the literal slug, so an entry can
cost two requests). If a connector ever shipped both `GMAIL_X` and a literal
`X`, both would compose to `connectors__gmail__X`, which runs `GMAIL_X`;
search keeps that twin, drops the other, and logs a warning. Results splice back into the batch's original order
with recomputed counts.

```yaml theme={null}
tools:
  connectors:
    enabled: true   # false — never touch connector routes; the bridge
                    # behaves exactly as if the feature didn't exist
```

Signed out (or when the gateway does not serve connectors for your
account), everything above is invisible: local search behaves exactly as
described in the rest of this page, with no errors shown to the model.

A connector call that needs an account you haven't linked returns a
`CONNECTION_REQUIRED` error. The `manage_connections` tool lists connectors and
their connection state and starts an authorization: in the desktop app the call
shows a card, blocks until each app is connected or skipped, and reports the
outcomes; elsewhere it returns a connect link per app for the user to open.
Disconnecting an account is done by the user in the Portal. The same tool also installs, enables and authorizes
local MCP servers from the catalog (targets with `mcp: true`), so it is
present whether or not you are signed in; only the managed-connector actions
need the sign-in.

The desktop backend's account-list and disconnect APIs use the Portal's
account-management service, including its organization membership checks and
disconnect audit. An unavailable Portal does not fall back to direct gateway
account management. Tool discovery, execution, and connection-status watching
continue through the gateway; the model tool cannot disconnect an account.

`tool_call` accepts a batch: `calls` is an array of `{name, arguments}`
entries (a single call is an array of one). Each connector entry in a batch
is dispatched as its own gateway request, one after another; local deferred
tools stay one entry per `tool_call`. A multi-entry batch that names a local
tool is rejected with a correction that restates the valid shape using the
caller's own first entry, and a `calls` value emitted as a JSON string is
parsed like the array form. Approvals settle per entry before
dispatch, and a `/stop` between entries leaves the unstarted ones unsent
(their slots report `INTERRUPTED`).

## When NOT to use it

Tool Search trades a fixed per-turn token cost (the three bridge tool
schemas plus the catalog listing) and at least one extra round trip on
cold tools (describe → call) for the savings on the deferred schemas.
At tier 1 the listing keeps every capability visible, so the discovery
round trip usually disappears — the model goes straight to
`tool_describe`. Live benchmarking showed the listing mode matching
eager loading's task success while costing less than the bare bridge.

If you want the old always-eager behavior for a small toolset, set
`enabled: off`.

## Trade-offs that don't go away

These come from the prompt-cache integrity invariant — they are inherent
to any progressive-disclosure design, not specific to this implementation:

* **One extra round trip on cold tools.** The first time the model needs
  a deferred tool, it spends one or two extra model calls to find and
  load the schema. The token savings on the static side are real, but a
  portion is paid back at runtime.
* **No cache benefit on deferred schemas.** A loaded `tool_describe`
  result enters the conversation history (so it does get cached on
  subsequent turns) but it never benefits from the system-prompt cache
  prefix.
* **No provider-native validation for deferred schemas.** `tool_describe`
  lets the model read a deferred tool's schema, but the provider still sees
  only the generic `tool_call.arguments` object. Mibyan therefore coerces and
  validates the underlying arguments locally before dispatch; the concrete
  tool or MCP server remains responsible for schemas Mibyan cannot safely
  validate, such as malformed schemas or external references.
* **Model-quality dependence.** Tool Search assumes the model can write a
  reasonable search query for the tool it wants. Smaller models do this
  less well; the published Anthropic numbers (49% → 74% on Opus 4 with
  vs. without tool search) show the upside but also that \~26 points of
  accuracy is still retrieval failure.
* **Toolset edits invalidate cache.** Adding or removing a tool mid-
  session changes the bridge tools' descriptions (which include the
  count of deferred tools) and the catalog, so the prompt cache is
  invalidated. This is the same trade-off as any toolset edit.

## Implementation details

* **Retrieval:** BM25 over tokenized tool name, source name (the MCP
  server or plugin toolset the tool belongs to, so searching `"linear"`
  finds that server's tools even when a tool's own name doesn't carry
  the service), description, and parameter names, with Snowball
  stemming (English) applied to both the index and the query so
  morphological variants match ("issues" finds `create_issue`). A tool is
  a result only if it contains the query's rarest token (the one in the
  fewest tool documents, so the word that names the intent: `gmail`,
  `github`, `incident`, not `send` or `create`). A query whose rarest
  token appears in no tool returns an empty group with the connected
  sources and a retry hint, instead of `limit` tools that share one
  common word.
* **Relevance floor:** a tool must match at least half of a query's
  *answerable* terms (terms present anywhere in the catalog) before it
  is offered — sharing one incidental word with a long query is not a
  match. A hunt for a capability that doesn't exist returns no results
  instead of a plausible-looking list the model rephrases against
  forever. The floor only engages from four answerable terms up, so
  short queries like "list issues" keep full recall, and an exact
  tool-name query always matches.
* **Parallel execution unwraps the bridge.** The batch planner decides
  concurrency on the *underlying* tool of a `tool_call`, not on the
  literal bridge name — so an MCP server opted in via
  `supports_parallel_tool_calls: true` keeps its concurrency when its
  tools are called through the bridge, and `tool_search` /
  `tool_describe` lookups batch concurrently like any read-only tool.
* **Catalog is stateless across turns.** It rebuilds from the current
  tool-defs list every assembly — no session-keyed `Map`. This avoids
  the class of bug where a stored catalog drifts out of sync with the
  live tool registry.
* **The catalog is scoped to the session's toolsets.** `tool_search`,
  `tool_describe`, and `tool_call` only ever see and invoke tools the
  session was actually granted. A subagent, kanban worker, or gateway
  session restricted to a subset of toolsets cannot use the bridge to
  discover or call a tool outside that subset — the deferred catalog is
  the deferrable slice of the session's own enabled/disabled toolsets,
  not the whole process registry.
* **No JS sandbox.** Mibyan uses the simpler "structured tools" mode
  (search / describe / call as plain functions). The JS-sandbox "code
  mode" some other implementations offer is a large surface area; we
  skip it.

## See also

* `tools/tool_search.py` — the implementation
* `tests/tools/test_tool_search.py` — the regression suite
* The `openclaw-tool-search-report` PDF in the original implementation
  PR for the research that shaped the design


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.