> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mibyanai.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Image Generation

> Generate images via FAL.ai — 11 models including FLUX 2, GPT Image (1.5 & 2), Nano Banana Pro, Ideogram, Recraft V4 Pro, Krea 2, and more, selectable via `mibyan tools`.

Mibyan generates images from text prompts via FAL.ai. Eleven models are supported out of the box, each with different speed, quality, and cost tradeoffs. The active model is user-configurable via `mibyan tools` and persists in `config.yaml`.

## Supported Models

| Model | Speed | Strengths | Price |
| - | - | - | - |
| `fal-ai/flux-2/klein/9b` *(default)* | `<1s` | Fast, crisp text | \$0.006/MP |
| `fal-ai/flux-2-pro` | \~6s | Studio photorealism | \$0.03/MP |
| `fal-ai/z-image/turbo` | \~2s | Bilingual EN/CN, 6B params | \$0.005/MP |
| `fal-ai/nano-banana-pro` | \~8s | Gemini 3 Pro, reasoning depth, text rendering | \$0.15/image (1K) |
| `fal-ai/gpt-image-1.5` | \~15s | Prompt adherence | \$0.034/image |
| `fal-ai/gpt-image-2` | \~20s | SOTA text rendering + CJK, world-aware photorealism | \$0.04–0.06/image |
| `fal-ai/ideogram/v3` | \~5s | Best typography | \$0.03–0.09/image |
| `fal-ai/recraft/v4/pro/text-to-image` | \~8s | Design, brand systems, production-ready | \$0.25/image |
| `fal-ai/qwen-image` | \~12s | LLM-based, complex text | \$0.02/MP |
| `fal-ai/krea/v2/medium/text-to-image` | \~15-25s | Illustration, anime, painting, expressive/artistic styles | \$0.030–0.035/image |
| `fal-ai/krea/v2/large/text-to-image` | \~25-60s | Photorealism, raw textured looks (motion blur, grain, film) | \$0.060–0.065/image |

Prices are FAL's pricing at time of writing; check [fal.ai](https://fal.ai/) for current numbers.

## Setup

<Tip>
  **Nous Subscribers**

  If you have a paid [Nous Portal](https://portal.nousresearch.com) subscription, you can use image generation through the **Tool Gateway** without a FAL API key. Your model selection persists across both paths. New installs can run `mibyan setup --portal` to log in and turn on every gateway tool at once; existing installs can pick **Nous Subscription** as the image-gen backend via `mibyan tools`.

  The **Nous Subscription** row is the only managed row. Its model picker spans every gateway the subscription runs — the FAL catalog above, native **Krea 2** (`krea-2-medium`, `krea-2-large`, `krea-2-medium-turbo`) and any Nous Portal image models — each model listed once, and the model you pick decides which gateway serves the request. Free tool-pool accounts see the FAL models only; Krea and Portal models are paid-subscription. With a Krea 2 model selected, `image_generate` also offers Krea's `creativity` setting and its `intensity`, `complexity` and `movement` sliders (-100 to 100).

  If the managed gateway returns `HTTP 4xx` for a specific model, that model isn't yet proxied on the portal side — the agent will tell you so, with remediation steps (switch to FAL.ai in `mibyan tools` with your own `FAL_KEY` for direct access, or pick a different model).
</Tip>

### Get a FAL API Key

1. Sign up at [fal.ai](https://fal.ai/)
2. Generate an API key from your dashboard

### Configure and Pick a Model

Run the tools command:

```bash theme={null}
mibyan tools
```

Navigate to **🎨 Image Generation**, pick your backend (Nous Subscription or FAL.ai), then the picker shows all supported models in a column-aligned table — arrow keys to navigate, Enter to select:

```
  Model                          Speed    Strengths                    Price
  fal-ai/flux-2/klein/9b         <1s      Fast, crisp text             $0.006/MP   ← currently in use
  fal-ai/flux-2-pro              ~6s      Studio photorealism          $0.03/MP
  fal-ai/z-image/turbo           ~2s      Bilingual EN/CN, 6B          $0.005/MP
  ...
```

Your selection is saved to `config.yaml`:

```yaml theme={null}
image_gen:
  provider: fal                 # `nous` if you picked Nous Subscription
  model: fal-ai/flux-2/klein/9b
  max_parallel_requests: 4      # concurrent images in one tool-call batch
```

`image_gen.provider` is the single selection key: `nous` routes through the managed Tool Gateway; a vendor name (`fal`, `openai`, `xai`, `krea`, ...) goes direct with your own key. The runtime always follows this stored selection — a `FAL_KEY` in `.env` is ignored while `provider: nous`, and `provider: fal` without `FAL_KEY` errors with `image_gen is configured to use fal (set via mibyan tools), but FAL_KEY is not set. Run 'mibyan tools' to change it.` rather than silently rerouting. Change providers via `mibyan tools`, not by adding/removing keys. (The old `use_gateway` boolean is legacy — still read as `nous` when `true`, but never written anymore.)

`max_parallel_requests` defaults to `4`. Mibyan clamps it to at least one and
to the global tool-worker limit, so image providers receive bounded parallel
requests without allowing an image batch to bypass the agent's concurrency cap.

### OpenRouter: the full Image API catalog

With `image_gen.provider: openrouter`, the model picker lists OpenRouter's
entire live image catalog — the dedicated
[Image API](https://openrouter.ai/docs/guides/overview/multimodal/image-generation)
models (Seedream, FLUX.2, Recraft, Qwen Image, MAI, Krea, Riverflow, Grok
Imagine, and more — 40+ ids) merged with the chat-completions image models.
The catalog is fetched live from `GET /images/models` and `GET /models`, so
new models appear in the picker as soon as OpenRouter serves them; no Mibyan
update needed. Generation routes each model to the surface that serves it
(dedicated `POST /images/generations` vs chat-completions) automatically.
Nous Portal proxies the chat-completions protocol only, so its picker offers
the chat-served models.

Optional per-request knobs for Image API models go under the scoped config
section (or `OPENROUTER_IMAGE_API_*` env vars):

```yaml theme={null}
image_gen:
  provider: openrouter
  model: bytedance-seed/seedream-4.5
  openrouter:
    resolution: 2K        # model-dependent: 1K / 2K / 4K
    quality: high         # gpt-image models
    output_format: png
```

### GPT-Image Quality

The `fal-ai/gpt-image-1.5` and `fal-ai/gpt-image-2` request quality is pinned to `medium` (\~$0.034–$0.06/image at 1024×1024). We don't expose the `low` / `high` tiers as a user-facing option so that Nous Portal billing stays predictable across all users — the cost spread between tiers is 3–22×. If you want a cheaper option, pick Klein 9B or Z-Image Turbo; if you want higher quality, use Nano Banana Pro or Recraft V4 Pro.

### Meta Model API: Muse Image

With `image_gen.provider: meta-ai`, images are generated through the
[Meta Model API](https://api.meta.ai) (`https://api.meta.ai/v1`), the same
OpenAI-compatible endpoint that serves the Muse Spark chat models. It is the
image-gen companion to the bundled `meta-ai` chat provider.

| Model | Speed | Strengths | Price |
| - | - | - | - |
| `muse-image-1.0` *(default)* | \~10s | Meta Model API image generation | \$0.01/image |

```yaml theme={null}
image_gen:
  provider: meta-ai
  model: muse-image-1.0
```

Auth reuses the same env vars as the Meta chat provider — `MODEL_API_KEY`
(Meta's documented name), with `META_API_KEY` / `META_MODEL_API_KEY` accepted
as aliases. Set `META_BASE_URL` to point at a proxy or alternate host. Text-to-image
only for now; responses are saved to `$mibyan_HOME/cache/images/`.

## FAL: GPT Image 2.5

Select **GPT Image 2.5 Flare** or **GPT Image 2.5 Sunburst** under
`mibyan tools` → Image Generation → FAL.ai. The model IDs are:

* `openai/gpt-image-2.5/flare/text-to-image`
* `openai/gpt-image-2.5/sunburst/text-to-image`

For example:

```bash theme={null}
mibyan config set image_gen.provider fal
mibyan config set image_gen.model openai/gpt-image-2.5/flare/text-to-image
```

Providing `image_url` or reference images automatically selects the corresponding
`openai/gpt-image-2.5/flare/edit` or `openai/gpt-image-2.5/sunburst/edit` endpoint.
Both accept up to 16 source images. Mibyan pins quality to `medium`, matching its
existing FAL GPT Image policy rather than FAL's higher-cost `high` default.
Landscape and portrait use 4:3 presets to satisfy the minimum pixel count;
square uses `square_hd`. Upscaling remains off unless requested.

FAL bills by tokens, not a fixed image price: $5/M text input, $1.25/M cached
text input, $10/M text output, $8/M image input, $2/M cached image input, and
$30/M image output, rounded up to \$0.0001 per request. See the
[Flare](https://fal.ai/models/openai/gpt-image-2.5/flare/text-to-image) and
[Sunburst](https://fal.ai/models/openai/gpt-image-2.5/sunburst/text-to-image)
pages. Direct FAL requires a funded `FAL_KEY`; managed-gateway availability
depends on that gateway's endpoint allowlist and is not implied by FAL availability.
Existing provider and model defaults are unchanged.

## OpenAI API: GPT Image 2.5

The **OpenAI** provider supports GPT Image 2.5 Flare (fast everyday creation)
and Sunburst (precision generation and editing), using `OPENAI_API_KEY`.
Select them through `mibyan tools` → Image Generation → OpenAI, or set:

```bash theme={null}
mibyan config set image_gen.provider openai
mibyan config set image_gen.openai.model gpt-image-2.5-flare
```

`gpt-image-2.5-flare` and `gpt-image-2.5-sunburst` use automatic quality.
Append `-low`, `-medium`, `-high`, `-xhigh`, or `-max` to select a fixed quality,
for example `gpt-image-2.5-sunburst-high`. Both support generation and editing
with up to 16 reference images. Existing GPT Image 2 selections and the
`gpt-image-2-medium` default are unchanged.

This is paid API usage, separate from a ChatGPT/Codex subscription. Both models
cost $5 per million text-input tokens, $8 per million image-input tokens, and
$30 per million image-output tokens (cached input rates are $1.25 and \$2,
respectively). Per-image cost varies with usage; the GPT Image 2 calculator
does not estimate 2.5 token consumption. See the official
[Flare](https://developers.openai.com/api/docs/models/gpt-image-2.5-flare) and
[Sunburst](https://developers.openai.com/api/docs/models/gpt-image-2.5-sunburst) docs.

The **OpenAI (Codex auth)** provider does not offer 2.5. The Codex backend
accepts any `model` value (including nonexistent ids) and generates with its
own server-managed engine, so a "selected" Flare or Sunburst tier would be a
label with no effect. Pick the direct OpenAI API provider or FAL for 2.5.

### Custom OpenAI-compatible image endpoint

The **OpenAI** provider can point at any OpenAI-compatible `/v1/images/generations`
endpoint (a local gateway, a task-scoped proxy, a third-party API gateway),
independently of the chat provider, and take its key from a variable of your choice:

```yaml theme={null}
image_gen:
  provider: openai
  openai:
    model: gpt-image-2-medium
    base_url: http://localhost:18081/v1   # → OPENAI_BASE_URL → api.openai.com
    key_env: IMAGE_GATEWAY_TOKEN          # → OPENAI_API_KEY
```

Only the variable *name* is stored in `config.yaml`; the secret stays in `.env`
or the process environment. Availability checks and
generation use the same resolution, so a configured `key_env` is enough — no
`OPENAI_API_KEY` is required. Requests go through Mibyan' own HTTP client, which
honours `HTTP(S)_PROXY`/`NO_PROXY` but ignores macOS system proxies (whose
exception list is invisible to Python), so `localhost` endpoints connect directly.
The `OpenAI-Project` header is sent blank on image requests: an `OPENAI_PROJECT_ID`
set for chat otherwise makes the image endpoint return 403 `model_not_found` on
projects with a model allow-list, while the key itself already carries the project.

**Gateway model names.** Catalog ids are mapped for OpenAI: `gpt-image-2-medium`
is sent as `model: gpt-image-2` + `quality: medium`. Any other value of
`image_gen.openai.model` (or `OPENAI_IMAGE_MODEL`) is sent verbatim as `model`
with **no** `quality` field, so a gateway that serves its own image model names
(`custom-image-model`, `grok-imagine-image`, ...) receives exactly that id and
never sees a quality enum it might reject. The shared top-level `image_gen.model`
is never passed through — it can hold another provider's id (a FAL path, for
instance) from an earlier selection.

**Reusing a named custom endpoint.** If the gateway is already declared under
`providers:` for chat, point the image provider at it by *name* instead of
repeating its URL and key:

```yaml theme={null}
providers:
  my-gateway:
    name: My Gateway
    api: https://gateway.example.com/v1
    key_env: MY_GATEWAY_KEY

image_gen:
  provider: openai
  openai:
    provider: my-gateway        # inherits api + key_env from providers.my-gateway
    model: grok-imagine-image   # sent verbatim, no quality
```

Resolution order is `image_gen.openai.base_url` → the named endpoint's URL →
`OPENAI_BASE_URL`, and the variable named by `image_gen.openai.key_env` → the
named endpoint's `api_key`/`key_env` → `OPENAI_API_KEY`; an explicit `base_url`
or `key_env` next to `provider` therefore overrides that part of the endpoint. A
name that matches no `providers:` entry is logged as a warning and ignored.

## Usage

The agent-facing schema is intentionally minimal — the model picks up whatever you've configured:

```
Generate an image of a serene mountain landscape with cherry blossoms
```

```
Create a square portrait of a wise old owl — use the typography model
```

```
Make me a futuristic cityscape, landscape orientation
```

## Image-to-Image / Editing

The same `image_generate` tool also **edits existing images** when the active
model supports it — pass a source image and the backend routes to its editing
endpoint automatically (mirrors how `video_generate` handles image-to-video).
Omit the source image and it's plain text-to-image.

```
Take this photo and make it a rainy Tokyo street at night → <image>
```

```
Blend these two product shots into one hero image → <image1> <image2>
```

Two inputs drive the edit:

* **`image_url`** — the primary source image to edit/transform (public URL or local path).
* **`reference_image_urls`** — additional style/composition references (capped per-model).

### Which backends support editing

| Backend | Image-to-image | Reference cap | How |
| - | - | - | - |
| **FAL.ai** (edit-capable models below) | ✓ | up to 16 (per model) | routes to the model's `/edit` endpoint |
| **OpenAI** (GPT Image 2 / 2.5 Flare / Sunburst) | ✓ | up to 16 | `images.edit()` |
| **xAI** (Grok Imagine) | ✓ | 1 | `/v1/images/edits` (`grok-imagine-image-quality`) |
| **Krea** (`Krea 2`) | ✓ | up to 10 | reference-guided generation (`image_style_references`); local files are embedded in the request, up to 3 MB in total |
| **OpenAI (Codex auth)** | ✓ | up to 16 | `POST /backend-api/codex/images/edits` with inline `images[]` data URLs (remote URLs are fetched client-side) |
| **OpenRouter** (Image API models) | ✓ | up to 14–16 (per model) | `input_references` on `POST /images/generations`; chat-served models use `image_url` content parts (up to 3) |

FAL models with an editing endpoint: `flux-2/klein/9b`, `flux-2-pro`,
`nano-banana-pro`, `gpt-image-1.5`, `gpt-image-2`, `ideogram/v3`, and
`qwen-image`, plus GPT Image 2.5 Flare and Sunburst above. Pure text-to-image FAL models (`z-image/turbo`, `recraft`,
`krea/*`) reject image inputs with a clear error pointing you at an
edit-capable model.

<Note>
  **OpenAI (Codex auth): the backend decides quality and size**

  Mibyan posts straight to the Codex backend's native
  `images/generations` / `images/edits` endpoints (the same route the official
  Codex client uses), so no chat model is involved and the call does not depend
  on which chat models your ChatGPT plan currently has. The backend, however,
  treats `model`, `quality` and `size` as advisory: it may return a different
  quality tier or geometry than requested (a portrait request can come back
  square). The result carries `reported_quality`, `reported_size` and
  `pixel_size` alongside what was requested, plus `imagegen_request_id` for
  OpenAI support. For exact control over quality and size, configure the
  **OpenAI** (API key), **FAL**, or **xAI** backend instead.
</Note>

The active model's editing capability is surfaced in the tool description at
runtime, so the agent knows whether `image_url` will be honored before it
calls the tool.

## Aspect Ratios

Every model accepts the same three aspect ratios from the agent's perspective. Internally, each model's native size spec is filled in automatically:

| Agent input | image\_size (flux/z-image/qwen/recraft/ideogram) | aspect\_ratio (nano-banana-pro) | image\_size (gpt-image-1.5) | image\_size (gpt-image-2) |
| - | - | - | - | - |
| `landscape` | `landscape_16_9` | `16:9` | `1536x1024` | `landscape_4_3` (1024×768) |
| `square` | `square_hd` | `1:1` | `1024x1024` | `square_hd` (1024×1024) |
| `portrait` | `portrait_16_9` | `9:16` | `1024x1536` | `portrait_4_3` (768×1024) |

GPT Image 2 maps to 4:3 presets rather than 16:9 because its minimum pixel count is 655,360 — the `landscape_16_9` preset (1024×576 = 589,824) would be rejected.

This translation happens in `_build_fal_payload()` — agent code never has to know about per-model schema differences.

## Upscaling

### Opt-in only

No model upscales by default. Modern image models emit their best quality
natively, and the available upscalers are *creative* enhancers (diffusion
passes) that can subtly redraw content — degrading rendered text, faces, and
fine detail. Upscaling only runs when the agent explicitly requests it.

### The `upscale` parameter (per-call opt-in)

* `upscale: true` — chain a high-resolution pass after generation:

| Backend | Upscaler |
| - | - |
| **FAL.ai** | Clarity Upscaler (2×, +\$0.03/MP) |
| **Krea** | Krea Enhance (2×, up to 8K ceiling) |
| Other backends | no upscaler; native resolution returned |

* `upscale: false` / omitted — native resolution (the default)

`video_generate` also accepts `upscale: true` on the FAL backend, chaining
ByteDance's **SeedVR2** video upscaler (2×, \$0.001/MP of output video) after
generation.

When the FAL image pass runs, it uses these settings:

| Setting | Value |
| - | - |
| Upscale factor | 2× |
| Creativity | 0.35 |
| Resemblance | 0.6 |
| Guidance scale | 4 |
| Inference steps | 18 |

If upscaling fails (network issue, rate limit), the original image is returned automatically. The response reports `upscaled: true/false` so the agent knows which resolution it got.

## How It Works Internally

1. **Model resolution** — `_resolve_fal_model()` reads `image_gen.model` from `config.yaml`, falls back to the `FAL_IMAGE_MODEL` env var, then to `fal-ai/flux-2/klein/9b`.
2. **Payload building** — `_build_fal_payload()` translates your `aspect_ratio` into the model's native format (preset enum, aspect-ratio enum, or GPT literal), merges the model's default params, applies any caller overrides, then filters to the model's `supports` whitelist so unsupported keys are never sent.
3. **Submission** — `_submit_fal_request()` routes via direct FAL credentials or the managed Nous gateway, according to the stored `image_gen.provider` selection.
4. **Upscaling** — runs only when the agent passed `upscale: true`; every model's catalog default is off.
5. **Delivery** — final image URL returned to the agent, which emits a `MEDIA:<url>` tag that platform adapters convert to native media.
6. **Usage accounting** — token-billed image models (OpenRouter chat-image and Image API models such as `google/gemini-3.1-flash-lite-image`, OpenAI `gpt-image`) return real token counts, so each call is recorded in `session_model_usage` as task `image_generation` under the billing provider and model, and shows up in `mibyan insights` and the dashboard's Usage analytics alongside other model calls. Per-image backends (FAL, xAI, Krea, ...) return no token usage and are not recorded there.

## Debugging

Enable debug logging:

```bash theme={null}
export IMAGE_TOOLS_DEBUG=true
```

Debug logs go to `./logs/image_tools_debug_<session_id>.json` with per-call details (model, parameters, timing, errors).

## Platform Delivery

| Platform | Delivery |
| - | - |
| **CLI** | Image URL printed as markdown \`\` — click to open |
| **Telegram** | Photo message with the prompt as caption |
| **Discord** | Embedded in a message |
| **Slack** | URL unfurled by Slack |
| **WhatsApp** | Media message |
| **Others** | URL in plain text |

## Limitations

* **Requires credentials** for the active backend (FAL `FAL_KEY` / Nous Subscription, `OPENAI_API_KEY`, xAI OAuth, `KREA_API_KEY`)
* **Editing is model-dependent** — image-to-image works only on edit-capable models (see the table above); text-to-image-only models reject image inputs with a clear error
* **Temporary URLs** — backends return hosted URLs that expire after hours/days; Mibyan materializes them to the local cache so delivery still works after expiry
* **Per-model constraints** — some models don't support `seed`, `num_inference_steps`, etc. The `supports` / `edit_supports` filter silently drops unsupported params; this is expected behavior


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.