Skip to main content
Commands, package names, and image names on this page come from the open-source project that Mibyan Desktop is built on, and can differ from the Mibyan Desktop installer. For the supported Mibyan install and update path, see Install and update.
Pretrain LLMs at scale with PyTorch 4D parallelism.

Skill metadata

Reference: full SKILL.md

The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.

TorchTitan - PyTorch Native Distributed LLM Pretraining

Quick start

TorchTitan is PyTorch’s official platform for large-scale LLM pretraining with composable 4D parallelism (FSDP2, TP, PP, CP), achieving 65%+ speedups over baselines on H100 GPUs. Installation:
Download tokenizer:
Start training on 8 GPUs:

Common workflows

Workflow 1: Pretrain Llama 3.1 8B on single node

Copy this checklist:
Step 1: Download tokenizer
Step 2: Configure training In torchtitan’s current layout, run configs are defined in a Python config registry (torchtitan/models/llama3/config_registry.py) and selected by name via CONFIG=<name> (or --config <name>). To customize, register your own config in the registry, or override individual fields on the command line (e.g. --optimizer.lr 3e-4 --training.steps 1000). The equivalent settings for an 8B run look like this (shown as fields; set them in the registry entry or as --section.key value overrides):
Step 3: Launch training
Step 4: Monitor and checkpoint TensorBoard logs are saved to ./outputs/tb/:

Workflow 2: Multi-node training with SLURM

Step 1: Configure parallelism for scale For 70B model on 256 GPUs (32 nodes):
Step 2: Set up SLURM script
Step 3: Submit job
Step 4: Resume from checkpoint Training auto-resumes if checkpoint exists in configured folder.

Workflow 3: Enable Float8 training for H100s

Float8 provides 30-50% speedup on H100 GPUs.
Step 1: Install torchao
Step 2: Configure Float8 In the current torchtitan, Float8 is applied at config time via the quantization parameter in your model_registry() call inside the config registry (not via a [quantize.linear.float8] TOML section). Add a Float8LinearConverter.Config:
Enable torch.compile in your run config too:
Step 3: Launch with compile

Workflow 4: 4D parallelism for 405B models

Step 1: Create seed checkpoint Required for consistent initialization across PP stages:
Step 2: Configure 4D parallelism
Step 3: Launch on 512 GPUs

When to use vs alternatives

Use TorchTitan when:
  • Pretraining LLMs from scratch (8B to 405B+)
  • Need PyTorch-native solution without third-party dependencies
  • Require composable 4D parallelism (FSDP2, TP, PP, CP)
  • Training on H100s with Float8 support
  • Want interoperable checkpoints with torchtune/HuggingFace
Use alternatives instead:
  • Megatron-LM: Maximum performance for NVIDIA-only deployments
  • DeepSpeed: Broader ZeRO optimization ecosystem, inference support
  • Axolotl/TRL: Fine-tuning rather than pretraining
  • LitGPT: Educational, smaller-scale training

Common issues

Issue: Out of memory on large models Enable activation checkpointing and reduce batch size:
Or use gradient accumulation:
Issue: TP causes high memory with async collectives Set environment variable:
Issue: Float8 training not faster Float8 only benefits large GEMMs. Filter small layers via the converter’s filter_fqns:
Issue: Checkpoint loading fails after parallelism change Use DCP’s resharding capability:
Issue: Pipeline parallelism initialization Create seed checkpoint first (see Workflow 4, Step 1).

Supported models

Performance benchmarks (H100)

Advanced topics

FSDP2 configuration: See references/fsdp.md for detailed FSDP2 vs FSDP1 comparison and ZeRO equivalents. Float8 training: See references/float8.md for tensorwise vs rowwise scaling recipes. Checkpointing: See references/checkpoint.md for HuggingFace conversion and async checkpointing. Adding custom models: See references/custom-models.md for TrainSpec protocol.

Resources