Skip to main content
Commands, package names, and image names on this page come from the open-source project that Mibyan Desktop is built on, and can differ from the Mibyan Desktop installer. For the supported Mibyan install and update path, see Install and update.
On-demand GPU cloud instances for ML training.

Skill metadata

Reference: full SKILL.md

The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.

Lambda Labs GPU Cloud

Guide to running ML workloads on Lambda Labs GPU cloud with on-demand instances and 1-Click Clusters.

When to use Lambda Labs

Use Lambda Labs when:
  • Need dedicated GPU instances with full SSH access
  • Running long training jobs (hours to days)
  • Want simple pricing with no egress fees
  • Need persistent storage across sessions
  • Require high-performance multi-node clusters (16-512 GPUs)
  • Want pre-installed ML stack (Lambda Stack with PyTorch, CUDA, NCCL)
Key features:
  • GPU variety: B200, H100, GH200, A100, A10, A6000, V100
  • Lambda Stack: Pre-installed PyTorch, TensorFlow, CUDA, cuDNN, NCCL
  • Persistent filesystems: Keep data across instance restarts
  • 1-Click Clusters: 16-512 GPU Slurm clusters with InfiniBand
  • Simple pricing: Pay-per-minute, no egress fees
  • Global regions: 12+ regions worldwide
Use alternatives instead:
  • Modal: For serverless, auto-scaling workloads
  • SkyPilot: For multi-cloud orchestration and cost optimization
  • RunPod: For cheaper spot instances and serverless endpoints
  • Vast.ai: For GPU marketplace with lowest prices

Quick start

Account setup

  1. Create account at https://lambda.ai
  2. Add payment method
  3. Generate API key from dashboard
  4. Add SSH key (required before launching instances)

Launch via console

  1. Go to https://cloud.lambda.ai/instances
  2. Click “Launch instance”
  3. Select GPU type and region
  4. Choose SSH key
  5. Optionally attach filesystem
  6. Launch and wait 3-15 minutes

Connect via SSH

GPU instances

Available GPUs

Instance configurations

Launch times

  • Single-GPU: 3-5 minutes
  • Multi-GPU: 10-15 minutes

Lambda Stack

All instances come with Lambda Stack pre-installed:

Verify installation

Python API

Installation

Authentication

List available instances

Launch instance

List running instances

Terminate instance

SSH key management

CLI with curl

List instance types

Launch instance

Terminate instance

Persistent storage

Filesystems

Filesystems persist data across instance restarts:

Create filesystem

  1. Go to Storage in Lambda console
  2. Click “Create filesystem”
  3. Select region (must match instance region)
  4. Name and create

Attach to instance

Filesystems must be attached at instance launch time:
  • Via console: Select filesystem when launching
  • Via API: Include file_system_names in launch request

Best practices

SSH configuration

Add SSH key

Multiple keys

Import from GitHub

SSH tunneling

JupyterLab

Launch from console

  1. Go to Instances page
  2. Click “Launch” in Cloud IDE column
  3. JupyterLab opens in browser

Manual access

Training workflows

Single-GPU training

Multi-GPU training (single node)

Checkpoint to filesystem

1-Click Clusters

Overview

High-performance Slurm clusters with:
  • 16-512 NVIDIA H100 or B200 GPUs
  • NVIDIA Quantum-2 400 Gb/s InfiniBand
  • GPUDirect RDMA at 3200 Gb/s
  • Pre-installed distributed ML stack

Included software

  • Ubuntu 22.04 LTS + Lambda Stack
  • NCCL, Open MPI
  • PyTorch with DDP and FSDP
  • TensorFlow
  • OFED drivers

Storage

  • 24 TB NVMe per compute node (ephemeral)
  • Lambda filesystems for persistent data

Multi-node training

Networking

Bandwidth

  • Inter-instance (same region): up to 200 Gbps
  • Internet outbound: 20 Gbps max

Firewall

  • Default: Only port 22 (SSH) open
  • Configure additional ports in Lambda console
  • ICMP traffic allowed by default

Private IPs

Common workflows

Workflow 1: Fine-tuning LLM

Workflow 2: Batch inference

Cost optimization

Choose right GPU

Reduce costs

  1. Use filesystems: Avoid re-downloading data
  2. Checkpoint frequently: Resume interrupted training
  3. Right-size: Don’t over-provision GPUs
  4. Terminate idle: No auto-stop, manually terminate

Monitor usage

  • Dashboard shows real-time GPU utilization
  • API for programmatic monitoring

Common issues

References

Resources