Commands, package names, and image names on this page come from the open-source project that Mibyan Desktop is built on, and can differ from the Mibyan Desktop installer. For the supported Mibyan install and update path, see Install and update.
Skill metadata
Reference: full SKILL.md
The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.
TorchTitan - PyTorch Native Distributed LLM Pretraining
Quick start
TorchTitan is PyTorch’s official platform for large-scale LLM pretraining with composable 4D parallelism (FSDP2, TP, PP, CP), achieving 65%+ speedups over baselines on H100 GPUs. Installation:Common workflows
Workflow 1: Pretrain Llama 3.1 8B on single node
Copy this checklist:torchtitan/models/llama3/config_registry.py) and selected by name via CONFIG=<name>
(or --config <name>). To customize, register your own config in the registry, or override
individual fields on the command line (e.g. --optimizer.lr 3e-4 --training.steps 1000).
The equivalent settings for an 8B run look like this (shown as fields; set them in the
registry entry or as --section.key value overrides):
./outputs/tb/:
Workflow 2: Multi-node training with SLURM
Workflow 3: Enable Float8 training for H100s
Float8 provides 30-50% speedup on H100 GPUs.quantization
parameter in your model_registry() call inside the config registry (not via a
[quantize.linear.float8] TOML section). Add a Float8LinearConverter.Config:
torch.compile in your run config too:
Workflow 4: 4D parallelism for 405B models
When to use vs alternatives
Use TorchTitan when:- Pretraining LLMs from scratch (8B to 405B+)
- Need PyTorch-native solution without third-party dependencies
- Require composable 4D parallelism (FSDP2, TP, PP, CP)
- Training on H100s with Float8 support
- Want interoperable checkpoints with torchtune/HuggingFace
- Megatron-LM: Maximum performance for NVIDIA-only deployments
- DeepSpeed: Broader ZeRO optimization ecosystem, inference support
- Axolotl/TRL: Fine-tuning rather than pretraining
- LitGPT: Educational, smaller-scale training
Common issues
Issue: Out of memory on large models Enable activation checkpointing and reduce batch size:filter_fqns:
Supported models
Performance benchmarks (H100)
Advanced topics
FSDP2 configuration: See references/fsdp.md for detailed FSDP2 vs FSDP1 comparison and ZeRO equivalents. Float8 training: See references/float8.md for tensorwise vs rowwise scaling recipes. Checkpointing: See references/checkpoint.md for HuggingFace conversion and async checkpointing. Adding custom models: See references/custom-models.md for TrainSpec protocol.Resources
- GitHub: https://github.com/pytorch/torchtitan
- Paper: https://arxiv.org/abs/2410.06511
- ICLR 2025: https://iclr.cc/virtual/2025/poster/29620
- PyTorch Forum: https://discuss.pytorch.org/c/distributed/torchtitan/44

