Skip to main content
Commands, package names, and image names on this page come from the open-source project that Mibyan Desktop is built on, and can differ from the Mibyan Desktop installer. For the supported Mibyan install and update path, see Install and update.
Run PyTorch training across GPUs with minimal changes.

Skill metadata

Reference: full SKILL.md

The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.

HuggingFace Accelerate - Unified Distributed Training

Quick start

Accelerate simplifies distributed training to 4 lines of code. Installation:
Convert PyTorch script (4 lines):
Run (single command):

Common workflows

Workflow 1: From single GPU to multi-GPU

Original script:
With Accelerate (4 lines added):
Configure (interactive):
Questions:
  • Which machine? (single/multi GPU/TPU/CPU)
  • How many machines? (1)
  • Mixed precision? (no/fp16/bf16/fp8)
  • DeepSpeed? (no/yes)
Launch (works on any setup):

Workflow 2: Mixed precision training

Enable FP16/BF16:

Workflow 3: DeepSpeed ZeRO integration

Enable DeepSpeed ZeRO-2 (pass a DeepSpeedPlugin, not a raw dict):
Or point at a full DeepSpeed JSON config via the plugin:
ds_config.json (a raw DeepSpeed config — passed via the plugin, NOT via --config_file):
Or via interactive config:
Launch (--config_file expects an accelerate YAML, not a raw DeepSpeed JSON):

Workflow 4: FSDP (Fully Sharded Data Parallel)

Enable FSDP:
Or via config:

Workflow 5: Gradient accumulation

Accumulate gradients:
Effective batch size: batch_size * num_gpus * gradient_accumulation_steps

When to use vs alternatives

Use Accelerate when:
  • Want simplest distributed training
  • Need single script for any hardware
  • Use HuggingFace ecosystem
  • Want flexibility (DDP/DeepSpeed/FSDP/Megatron)
  • Need quick prototyping
Key advantages:
  • 4 lines: Minimal code changes
  • Unified API: Same code for DDP, DeepSpeed, FSDP, Megatron
  • Automatic: Device placement, mixed precision, sharding
  • Interactive config: No manual launcher setup
  • Single launch: Works everywhere
Use alternatives instead:
  • PyTorch Lightning: Need callbacks, high-level abstractions
  • Ray Train: Multi-node orchestration, hyperparameter tuning
  • DeepSpeed: Direct API control, advanced features
  • Raw DDP: Maximum control, minimal abstraction

Common issues

Issue: Wrong device placement Don’t manually move to device:
Issue: Gradient accumulation not working Use context manager:
Issue: Checkpointing in distributed Use accelerator methods:
Issue: Different results with FSDP Ensure same random seed:

Advanced topics

Megatron integration: See references/megatron-integration.md for tensor parallelism, pipeline parallelism, and sequence parallelism setup. Custom plugins: See references/custom-plugins.md for creating custom distributed plugins and advanced configuration. Performance tuning: See references/performance.md for profiling, memory optimization, and best practices.

Hardware requirements

  • CPU: Works (slow)
  • Single GPU: Works
  • Multi-GPU: DDP (default), DeepSpeed, or FSDP
  • Multi-node: DDP, DeepSpeed, FSDP, Megatron
  • TPU: Supported
  • Apple MPS: Supported
Launcher requirements:
  • DDP: torch.distributed.run (built-in)
  • DeepSpeed: deepspeed (pip install deepspeed)
  • FSDP: PyTorch 1.12+ (built-in)
  • Megatron: Custom setup

Resources