Skip to main content
Commands, package names, and image names on this page come from the open-source project that Mibyan Desktop is built on, and can differ from the Mibyan Desktop installer. For the supported Mibyan install and update path, see Install and update.
Serverless GPU cloud for ML jobs and model APIs.

Skill metadata

Reference: full SKILL.md

The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.

Modal Serverless GPU

Guide to running ML workloads on Modal’s serverless GPU cloud platform.

When to use Modal

Use Modal when:
  • Running GPU-intensive ML workloads without managing infrastructure
  • Deploying ML models as auto-scaling APIs
  • Running batch processing jobs (training, inference, data processing)
  • Need pay-per-second GPU pricing without idle costs
  • Prototyping ML applications quickly
  • Running scheduled jobs (cron-like workloads)
Key features:
  • Serverless GPUs: T4, L4, A10G, L40S, A100, H100, H200, B200 on-demand
  • Python-native: Define infrastructure in Python code, no YAML
  • Auto-scaling: Scale to zero, scale to 100+ GPUs instantly
  • Sub-second cold starts: Rust-based infrastructure for fast container launches
  • Container caching: Image layers cached for rapid iteration
  • Web endpoints: Deploy functions as REST APIs with zero-downtime updates
Use alternatives instead:
  • RunPod: For longer-running pods with persistent state
  • Lambda Labs: For reserved GPU instances
  • SkyPilot: For multi-cloud orchestration and cost optimization
  • Kubernetes: For complex multi-service architectures

Quick start

Installation

Hello World with GPU

Run: modal run hello_gpu.py

Basic inference endpoint

Core concepts

Key components

Execution modes

GPU configuration

Available GPUs

GPU specification patterns

Container images

Persistent storage

Web endpoints

FastAPI endpoint decorator

Full ASGI app

Web endpoint types

Dynamic batching

Secrets management

Scheduling

Performance optimization

Cold start mitigation

Model loading best practices

Parallel processing

Common configuration

Modal 1.0 autoscaler renames (see the migration guide):
  • container_idle_timeout → scaledown_window
  • concurrency_limit → max_containers
  • keep_warm → min_containers
  • allow_concurrent_inputs=N → the @modal.concurrent(max_inputs=N) decorator

Debugging

Common issues

References

Resources