How It Works
- Copy an image to your clipboard (screenshot, browser image, etc.)
- Attach it using one of the methods below
- Type your question and press Enter
- The image appears as a
[📎 Image #1]badge above the input - On submit, the image is sent to the model as a vision content block
Ctrl+C to clear all attached images.
Images are saved to ~/.mibyan/images/ as PNG files with timestamped filenames.
Paste Methods
How you attach an image depends on your terminal environment. Not all methods work everywhere — here’s the full breakdown:/paste Command
The most reliable explicit image-attach fallback.
/paste and press Enter. Mibyan checks your clipboard for an image and attaches it. This is the safest option when your terminal rewrites Cmd+V/Ctrl+V, or when you copied only an image and there is no bracketed-paste text payload to inspect.
Ctrl+V / Cmd+V
Mibyan now treats paste as a layered flow:- normal text paste first
- native clipboard / OSC52 text fallback if the terminal did not deliver text cleanly
- image attach when the clipboard or pasted payload resolves to an image or image path
file://... image URIs can attach immediately instead of sitting in the composer as raw text.
/terminal-setup for VS Code / Cursor / Windsurf
If you run the TUI inside a local VS Code-family integrated terminal on macOS, Mibyan can install the recommended workbench.action.terminal.sendSequence bindings for better multiline and undo/redo parity:
Cmd+Enter, Cmd+Z, or Shift+Cmd+Z are being intercepted by the IDE. Run it on the local machine only — not inside an SSH session.
Platform Compatibility
² See SSH & Remote Sessions below
³ The command writes local IDE keybindings and should not be run from the remote host
Platform-Specific Setup
macOS
No setup required. Mibyan usesosascript (built into macOS) to read the clipboard. For faster performance, optionally install pngpaste:
Linux (X11)
Installxclip:
Linux (Wayland)
Modern Linux desktops (Ubuntu 22.04+, Fedora 34+) often use Wayland by default. Installwl-clipboard:
WSL2
No extra setup required. Mibyan detects WSL2 automatically (via/proc/version) and uses powershell.exe to access the Windows clipboard through .NET’s System.Windows.Forms.Clipboard. This is built into WSL2’s Windows interop — powershell.exe is available by default.
The clipboard data is transferred as base64-encoded PNG over stdout, so no file path conversion or temp files are needed.
WSLg NoteIf you’re running WSLg (WSL2 with GUI support), Mibyan tries the PowerShell path first, then falls back to
wl-paste. WSLg’s clipboard bridge only supports BMP format for images — Mibyan auto-converts BMP to PNG using Pillow (if installed) or ImageMagick’s convert command.Verify WSL2 clipboard access
SSH & Remote Sessions
Clipboard image paste does not fully work over SSH. When you SSH into a remote machine, the Mibyan CLI runs on the remote host. Clipboard tools (xclip, wl-paste, powershell.exe, osascript) read the clipboard of the machine they run on — which is the remote server, not your local machine. Your local clipboard image is therefore inaccessible from the remote side.
Text can sometimes still bridge through terminal paste or OSC52, but image clipboard access and local screenshot temp paths remain tied to the machine running Mibyan.
Workarounds for SSH
-
Upload the image file — Save the image locally, upload it to the remote server via
scp, VSCode’s file explorer (drag-and-drop), or any file transfer method. Then reference it by path. (A/attach <filepath>command is planned for a future release.) -
Use a URL — If the image is accessible online, just paste the URL in your message. The agent can use
vision_analyzeto look at any image URL directly. -
X11 forwarding — Connect with
ssh -Xto forward X11. This letsxclipon the remote machine access your local X11 clipboard. Requires an X server running locally (XQuartz on macOS, built-in on Linux X11 desktops). Slow for large images. - Use a messaging platform — Send images to Mibyan via Telegram, Discord, Slack, or WhatsApp. These platforms handle image upload natively and are not affected by clipboard/terminal limitations.
Why Terminals Can’t Paste Images
This is a common source of confusion, so here’s the technical explanation: Terminals are text-based interfaces. When you press Ctrl+V (or Cmd+V), the terminal emulator:- Reads the clipboard for text content
- Wraps it in bracketed paste escape sequences
- Sends it to the application through the terminal’s text stream
osascript, powershell.exe, xclip, wl-paste) directly via subprocess to read the clipboard independently.
Supported Models
Image paste works with any vision-capable model. The image is sent as a base64-encoded data URL in the OpenAI vision content format:Image Routing (Vision-Capable vs Text-Only Models)
When a user attaches an image — from the CLI clipboard, the gateway (Telegram/Discord photo), or any other entry point — Mibyan routes it based on whether your current model actually supports vision:
You don’t configure this — Mibyan looks up your current model’s capability in the provider metadata and picks the right path automatically. The practical effect: you can switch between vision and non-vision models mid-session and image handling “just works” without changing your workflow. Text-only models get coherent context about the image rather than a broken multimodal payload they’d have to reject.
To override the automatic choice, set
agent.image_input_mode in config.yaml:
This is the knob to reach for when a backend accepts text but rejects native image input (for example an
openai-codex account whose backend answers image requests with server_error): keep your main model and point auxiliary.vision at a different vision-capable provider and model (with auxiliary.vision.provider: auto the describer would auto-detect the same main model again). That alone switches images to the description path in auto mode; agent.image_input_mode: text makes the same choice explicit.
Which auxiliary model handles the text-description path is configurable under auxiliary.vision — see Auxiliary Models.
vision_analyze has the same dual behavior
The vision_analyze tool itself follows the same routing. When the active main model is vision-capable and its provider supports image content inside tool results (currently the Anthropic, OpenAI, Azure-OpenAI, and Gemini 3.x stacks), vision_analyze short-circuits the auxiliary describer and returns the raw image pixels as a multimodal tool-result envelope. The main model sees the image natively on its next turn — no aux call, no text-summary information loss, no extra latency. One exception: an image that is already attached natively to the current user message is not re-embedded — vision_analyze on that same path returns a short text result saying the image is already in context (pass a region to zoom into part of it, which does embed the crop).
For text-only main models (or providers whose tool-result channel doesn’t carry images), vision_analyze falls back to the legacy path: it asks the configured auxiliary vision model to describe the image and returns the description as plain text. Either way the calling tool signature is the same — the tool decides which path to take at runtime based on the active model.
SVG and other non-raster images on Responses backends
Responses-style backends (for exampleopenai-codex) accept only inline JPEG, PNG, GIF and WebP; any other data:image/* part makes them reject the whole request, and because the part stays in history every later turn fails the same way. Mibyan handles this at the send layer: an inline SVG is rasterized to PNG when a rasterizer is installed (cairosvg, svglib+reportlab, rsvg-convert, or inkscape — the same soft dependencies vision_analyze uses), so the model still sees the drawing. Without a rasterizer, an SVG — and any other unsupported inline format such as BMP or TIFF — is replaced by a short text placeholder ([image omitted: image/svg+xml is not a supported image format]) while the valid images in the same message are still sent.
Native embeds ride the session: vision.embed_target_bytes and vision.max_calls_per_image
A native vision_analyze result bakes the image into the tool result, and that result is re-sent on every later API call of the session. Two config.yaml keys bound the recurring cost:
embed_target_bytes— images above the budget (or wider than 1568 px) are downscaled to a JPEG that fits. 256 KB keeps ordinary screenshots cheap; dense phone screenshots of tables can come out unreadable at that size, so raise it (say1048576) when the model keeps calling figures “unreadable”. Browser screenshots delivered natively use the same budget.max_calls_per_image— how often the same image (region crops of it included; local paths compare by resolved path) may be embedded per session. Once the cap is hit the tool returns"vision_analyze refused: this image has already been loaded into context N time(s) …"instead of another embed, so the model answers from what it already sees. Left unset, only delegateddelegate_tasksubagents are capped (at 3): they run unattended and cannot be steered mid-loop from the CLI, and a re-load loop there once burned 158 calls on five files. Set a number to cap every session, or0for unlimited everywhere.

