Skip to main content
Commands, package names, and image names on this page come from the open-source project that Mibyan Desktop is built on, and can differ from the Mibyan Desktop installer. For the supported Mibyan install and update path, see Install and update.
Text-to-image generation, inpainting, and img2img.

Skill metadata

Reference: full SKILL.md

The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.

Stable Diffusion Image Generation

Guide to generating images with Stable Diffusion using the HuggingFace Diffusers library.

When to use Stable Diffusion

Use Stable Diffusion when:
  • Generating images from text descriptions
  • Performing image-to-image translation (style transfer, enhancement)
  • Inpainting (filling in masked regions)
  • Outpainting (extending images beyond boundaries)
  • Creating variations of existing images
  • Building custom image generation workflows
Key features:
  • Text-to-Image: Generate images from natural language prompts
  • Image-to-Image: Transform existing images with text guidance
  • Inpainting: Fill masked regions with context-aware content
  • ControlNet: Add spatial conditioning (edges, poses, depth)
  • LoRA Support: Efficient fine-tuning and style adaptation
  • Multiple Models: SD 1.5, SDXL, SD 3.0, Flux support
Use alternatives instead:
  • DALL-E 3: For API-based generation without GPU
  • Midjourney: For artistic, stylized outputs
  • Imagen: For Google Cloud integration
  • Leonardo.ai: For web-based creative workflows

Quick start

Installation

Basic text-to-image

Using SDXL (higher quality)

Architecture overview

Three-pillar design

Diffusers is built around three core components:

Pipeline inference flow

Core concepts

Pipelines

Pipelines orchestrate complete workflows:

Schedulers

Schedulers control the denoising process:

Swapping schedulers

Generation parameters

Key parameters

Reproducible generation

Negative prompts

Image-to-image

Transform existing images with text guidance:

Inpainting

Fill masked regions:

ControlNet

Add spatial conditioning for precise control:

Available ControlNets

LoRA adapters

Load fine-tuned style adapters:

Multiple LoRAs

Memory optimization

Enable CPU offloading

Attention slicing

xFormers memory-efficient attention

VAE slicing for large images

Model variants

Loading different precisions

Loading specific components

Batch generation

Generate multiple images efficiently:

Common workflows

Workflow 1: High-quality generation

Workflow 2: Fast prototyping

Common issues

CUDA out of memory:
Black/noise images:
Slow generation:

References

Resources