Skip to main content
Commands, package names, and image names on this page come from the open-source project that Mibyan Desktop is built on, and can differ from the Mibyan Desktop installer. For the supported Mibyan install and update path, see Install and update.
SAM: zero-shot image segmentation via points, boxes, masks.

Skill metadata

Reference: full SKILL.md

The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.

Segment Anything Model (SAM)

Guide to using Meta AI’s Segment Anything Model for zero-shot image segmentation.

When to use SAM

Use SAM when:
  • Need to segment any object in images without task-specific training
  • Building interactive annotation tools with point/box prompts
  • Generating training data for other vision models
  • Need zero-shot transfer to new image domains
  • Building object detection/segmentation pipelines
  • Processing medical, satellite, or domain-specific images
Key features:
  • Zero-shot segmentation: Works on any image domain without fine-tuning
  • Flexible prompts: Points, bounding boxes, or previous masks
  • Automatic segmentation: Generate all object masks automatically
  • High quality: Trained on 1.1 billion masks from 11 million images
  • Multiple model sizes: ViT-B (fastest), ViT-L, ViT-H (most accurate)
  • ONNX export: Deploy in browsers and edge devices
Use alternatives instead:
  • YOLO/Detectron2: For real-time object detection with classes
  • Mask2Former: For semantic/panoptic segmentation with categories
  • GroundingDINO + SAM: For text-prompted segmentation
  • SAM 2: For video segmentation tasks

Quick start

Installation

Download checkpoints

Basic usage with SamPredictor

HuggingFace Transformers

Core concepts

Model architecture

Model variants

Prompt types

Interactive segmentation

Point prompts

Box prompts

Combined prompts

Iterative refinement

Automatic mask generation

Basic automatic segmentation

Customized generation

Filtering masks

Batched inference

Multiple images

Multiple prompts per image

ONNX deployment

Export model

Use ONNX model

Common workflows

Workflow 1: Annotation tool

Workflow 2: Object extraction

Workflow 3: Medical image segmentation

Output format

Mask data structure

COCO RLE format

Performance optimization

GPU memory

Speed optimization

Common issues

References

Resources