Commands, package names, and image names on this page come from the open-source project that Mibyan Desktop is built on, and can differ from the Mibyan Desktop installer. For the supported Mibyan install and update path, see Install and update.
Skill metadata
Reference: full SKILL.md
The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.
Flash Attention - Fast Memory-Efficient Attention
Quick start
Flash Attention provides 2-4x speedup and 10-20x memory reduction for transformer attention through IO-aware tiling and recomputation. PyTorch native (easiest, PyTorch 2.2+):Common workflows
Workflow 1: Enable in existing PyTorch model
Copy this checklist:torch.backends.cuda.sdp_kernel is deprecated; use
torch.nn.attention.sdpa_kernel with SDPBackend):
Workflow 2: Use flash-attn library for advanced features
For multi-query attention, sliding window, or H100 FP8. Copy this checklist:Workflow 3: H100 FP8 optimization (FlashAttention-3)
For maximum performance on Hopper GPUs (H100).Important: The pip packageflash-attn(2.8.x) ships FlashAttention-2 only — it does not contain FA3 or FP8 H100 kernels, andflash_attn_funcdoes not auto-use FP8. FlashAttention-3 is a separate beta build compiled from source from the repo’shopper/directory, exposed via theflash_attn_interfacemodule. FA3 supports FP16/BF16 forward+backward and FP8 forward only.
pip install flash-attn. Build it from the hopper/ subdirectory:
flash_attn_interface (distinct from the FA2 flash_attn).
FP8 is a forward-only path and expects float8_e4m3fn inputs:
When to use vs alternatives
Use Flash Attention when:- Training transformers with sequences >512 tokens
- Running inference with long context (>2K tokens)
- GPU memory constrained (OOM with standard attention)
- Need 2-4x speedup without accuracy loss
- Using PyTorch 2.2+ or can install flash-attn
- Standard attention: Sequences <256 tokens (overhead not worth it)
- xFormers: Need more attention variants (not just speed)
- Memory-efficient attention: CPU inference (Flash Attention needs GPU)
Common issues
Issue: ImportError: cannot import flash_attn Install with no-build-isolation flag:- <512 tokens: Minimal speedup (10-20%)
- 512-2K tokens: 2-3x speedup
-
2K tokens: 3-4x speedup
- Ampere (A100, A10): ✅ Full support
- Turing (T4): ✅ Supported
- Volta (V100): ❌ Not supported
Advanced topics
Integration with HuggingFace Transformers: See references/transformers-integration.md for enabling Flash Attention in BERT, GPT, Llama models. Performance benchmarks: See references/benchmarks.md for detailed speed and memory comparisons across GPUs and sequence lengths.Hardware requirements
- GPU: NVIDIA Ampere+ (A100, A10, A30) or AMD MI200+
- VRAM: Same as standard attention (Flash Attention doesn’t increase memory)
- CUDA: 12.0+ (11.8 minimum)
- PyTorch: 2.2+ for native support
Resources
- Paper: “FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness” (NeurIPS 2022)
- Paper: “FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning” (ICLR 2024)
- Blog: https://tridao.me/blog/2024/flash3/
- GitHub: https://github.com/Dao-AILab/flash-attention
- PyTorch docs: https://pytorch.org/docs/stable/generated/torch.nn.functional.scaled_dot_product_attention.html

