Skip to main content
The execute_code tool lets the agent write Python scripts that call Mibyan tools programmatically, collapsing multi-step workflows into a single LLM turn. The script runs in a child process on the agent host, communicating with Mibyan over a Unix domain socket RPC.

How It Works

  1. The agent writes a Python script using from mibyan_tools import ...
  2. Mibyan generates a mibyan_tools.py stub module with RPC functions
  3. Mibyan opens a Unix domain socket and starts an RPC listener thread
  4. The script runs in a child process — tool calls travel over the socket back to Mibyan
  5. Only the script’s print() output is returned to the LLM; intermediate tool results never enter the context window
Available tools inside scripts: web_search, web_extract, read_file, write_file, search_files, patch, terminal (foreground only).

When the Agent Uses This

The agent uses execute_code when there are:
  • 3+ tool calls with processing logic between them
  • Bulk data filtering or conditional branching
  • Loops over results
The key benefit: intermediate tool results never enter the context window — only the final print() output comes back, dramatically reducing token usage.

Practical Examples

Data Processing Pipeline

Multi-Step Web Research

Bulk File Refactoring

Build and Test Pipeline

Execution Mode

execute_code has two execution modes controlled by code_execution.mode in ~/.mibyan/config.yaml: When to leave it on project: you want import pandas, from my_project import foo, or relative paths like open(".env") to work the same way they do in terminal(). This is almost always what you want. When to flip to strict: you need maximum reproducibility — you want the same interpreter every session regardless of which venv the user activated, and you want scripts quarantined from the project tree (no risk of accidentally reading project files through a relative path).
Fallback behavior in project mode: if VIRTUAL_ENV / CONDA_PREFIX is unset, broken, or points at a Python older than 3.8, the resolver falls back cleanly to sys.executable — it never leaves the agent without a working interpreter. Security-critical invariants are identical across both modes:
  • environment scrubbing (API keys, tokens, credentials stripped)
  • tool whitelist (scripts cannot call execute_code recursively, delegate_task, or MCP tools)
  • resource limits (timeout, stdout cap, tool-call cap)
Switching mode changes where scripts run and which interpreter runs them, not what credentials they can see or which tools they can call.

Persistent session kernel

Calls reuse a Python child for the same session, execution mode, interpreter, working directory, and tool set. Imports, variables, and loaded data can persist between cells. The child environment is fixed when the kernel starts. Pass reset: true to discard that kernel state. A timeout or interrupted kernel can also lose it. Do not assume that a later terminal environment change is already visible inside an existing kernel. The old code_execution.kernel_mode setting is no longer a separate switch.

Resource Limits

All limits are configurable via config.yaml:

State Between Calls (the session kernel)

On the local terminal backend, execute_code does not start a fresh interpreter for every call. Each session owns a persistent Python kernel, so variables, imports, and loaded data from one call are available in the next. The agent can load a dataset once and query it across several turns instead of re-reading it every time. Subagents get their own kernel; kernels are never shared across sessions. A subagent’s kernel lives exactly as long as the subagent: it is exempt from the live-kernel cap while the subagent runs (a wide fan-out no longer evicts a sibling’s kernel mid-task) and is disposed when the subagent finishes. What ends a kernel:
  • Timeout or interrupt. A cell that hits the timeout (or is interrupted) kills the kernel process and its state is lost on purpose; the result says so and the next call starts a fresh kernel.
  • reset=true. The agent can pass reset: true to discard the kernel’s state and start clean. This is also the way to pick up environment changes: a kernel’s environment is frozen when it spawns, so a newly allowlisted passthrough variable is invisible until the kernel is reset.
  • Idle timeout and eviction. Kernels die with the session, after code_execution.kernel_idle_timeout idle seconds (default 1800), or when more than code_execution.max_session_kernels (default 4) top-level sessions’ kernels are alive and the oldest is evicted (running subagents’ kernels do not count against the cap).
The security envelope is the same as a one-shot script: environment scrubbing, the tool whitelist, and the per-call tool budget all apply to every cell, and tool-call authority (approvals, session, allow-list) is rebound on each cell.
Remote backends (Docker, SSH, Modal) run a remote session kernel with the same contract. If the kernel cannot be spawned on the backend, Mibyan falls back to running each call as a standalone script and says so in the result. Large output. Stdout over 50 KB is shown head-and-tail inline, and the full text is saved under ~/.mibyan/cache/exec/ with the path included in the result, so the agent can page through it with read_file instead of re-running the script.

How Tool Calls Work Inside Scripts

When your script calls a function like web_search("query"):
  1. The call is serialized to JSON and sent over a Unix domain socket to the parent process
  2. The parent dispatches through the standard handle_function_call handler
  3. The result is sent back over the socket
  4. The function returns the parsed result
This means tool calls inside scripts behave identically to normal tool calls — same rate limits, same error handling, same capabilities. The only restriction is that terminal() is foreground-only (no background or pty parameters).

Error Handling

When a script fails, the agent receives structured error information:
  • Non-zero exit code: stderr is included in the output so the agent sees the full traceback
  • Timeout: Script is killed and the agent sees "Script timed out after 300s and was killed."
  • Interruption: If the user sends a new message during execution, the script is terminated and the agent sees [execution interrupted — user sent a new message]
  • Tool call limit: When the 50-call limit is hit, subsequent tool calls return an error message
The response always includes status (success/error/timeout/interrupted), output, tool_calls_made, and duration_seconds.

Security

Security ModelThe child process runs with a minimal environment. API keys, tokens, and credentials are stripped by default. The script accesses tools exclusively via the RPC channel — it cannot read secrets from environment variables unless explicitly allowed.
Environment variables containing KEY, TOKEN, SECRET, PASSWORD, CREDENTIAL, PASSWD, or AUTH in their names are excluded. Only safe system variables (PATH, HOME, LANG, SHELL, PYTHONPATH, VIRTUAL_ENV, etc.) are passed through.

Skill Environment Variable Passthrough

When a skill declares required_environment_variables in its frontmatter, those variables are automatically passed through to both execute_code and terminal child processes after the skill is loaded. This lets skills use their declared API keys without weakening the security posture for arbitrary code. For non-skill use cases, you can explicitly allowlist variables in config.yaml:
See the Security guide for full details.

mibyan_* variables in the child

The child process receives only a small, fixed set of operational mibyan_* variables by exact name:
  • mibyan_HOME
  • mibyan_PROFILE
  • mibyan_CONFIG
  • mibyan_ENV
(plus mibyan_RPC_DIR / mibyan_RPC_SOCKET / TZ / HOME, which Mibyan injects explicitly so the RPC channel works).
Behavior changeEarlier versions passed any variable whose name began with mibyan_ through to the child. That broad prefix was removed for security hardening: it could leak mibyan_*-named configuration that doesn’t match a secret substring (for example mibyan_BASE_URL, mibyan_KANBAN_DB, or a mibyan_*_WEBHOOK endpoint) into arbitrary sandboxed code.If an execute_code script — or a repo/plugin module it imports at import time — relied on a mibyan_* variable outside the four operational names above, it will now find that variable unset in the child. The drop is intentional, not a bug.
Workaround — opt the variable back in explicitly. Both routes pass the variable through execute_code and terminal children, and neither weakens the secret-stripping guarantee (Mibyan-managed provider credentials can never be re-allowed this way):
  1. Per-machine, in config.yaml — add the exact variable name to the passthrough allowlist:
  2. Per-skill, in the skill’s frontmatter — declare it so it is registered automatically whenever that skill is loaded:
Diagnosing it. When the child drops one or more non-allowlisted mibyan_* variables, Mibyan emits a one-line debug log naming them and pointing at the env_passthrough escape hatch. Run with debug logging (mibyan logs --level DEBUG, or check ~/.mibyan/logs/agent.log) and look for execute_code: dropped N non-allowlisted mibyan_* var(s) if a script behaves as though a mibyan_* variable is missing. Mibyan always writes the script and the auto-generated mibyan_tools.py RPC stub into a temp staging directory that is cleaned up after execution. In strict mode the script also runs there; in project mode it runs in the session’s working directory (the staging directory stays on PYTHONPATH so imports still resolve). The child process runs in its own process group so it can be cleanly killed on timeout or interruption.

execute_code vs terminal

Rule of thumb: Use execute_code when you need to call Mibyan tools programmatically with logic between calls. Use terminal for running shell commands, builds, and processes.

Platform Support

Code execution is available on Linux, macOS, and Windows. On Linux and macOS the RPC channel uses a Unix domain socket; on Windows, where AF_UNIX is unreliable, Mibyan automatically falls back to a loopback TCP socket for the sandbox RPC transport. Remote terminal backends (Docker/SSH/Modal/etc.) use a file-based RPC transport instead and additionally require Python 3 inside the backend.