Skip to main content
Commands, package names, and image names on this page come from the open-source project that Mibyan Desktop is built on, and can differ from the Mibyan Desktop installer. For the supported Mibyan install and update path, see Install and update.
Fine-tune large LLMs with LoRA on limited GPU memory.

Skill metadata

Reference: full SKILL.md

The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.

PEFT (Parameter-Efficient Fine-Tuning)

Fine-tune LLMs by training <1% of parameters using LoRA, QLoRA, and 25+ adapter methods.

When to use PEFT

Use PEFT/LoRA when:
  • Fine-tuning 7B-70B models on consumer GPUs (RTX 4090, A100)
  • Need to train <1% parameters (6MB adapters vs 14GB full model)
  • Want fast iteration with multiple task-specific adapters
  • Deploying multiple fine-tuned variants from one base model
Use QLoRA (PEFT + quantization) when:
  • Fine-tuning 70B models on single 24GB GPU
  • Memory is the primary constraint
  • Can accept ~5% quality trade-off vs full fine-tuning
Use full fine-tuning instead when:
  • Training small models (<1B parameters)
  • Need maximum quality and have compute budget
  • Significant domain shift requires updating all weights

Quick start

Installation

LoRA fine-tuning (standard)

QLoRA fine-tuning (memory-efficient)

LoRA parameter selection

Rank (r) - capacity vs efficiency

Alpha (lora_alpha) - scaling factor

Target modules by architecture

Loading and merging adapters

Load trained adapter

Merge adapter into base model

Multi-adapter serving

PEFT methods comparison

IA3 (minimal parameters)

Prefix Tuning

Integration patterns

With TRL (SFTTrainer)

With Axolotl (YAML config)

With vLLM (inference)

Performance benchmarks

Memory usage (Llama 3.1 8B)

Training speed (A100 80GB)

Quality (MMLU benchmark)

Common issues

CUDA OOM during training

Adapter not applying

Quality degradation

Best practices

  1. Start with r=8-16, increase if quality insufficient
  2. Use alpha = 2 * rank as starting point
  3. Target attention + MLP layers for best quality/efficiency
  4. Enable gradient checkpointing for memory savings
  5. Save adapters frequently (small files, easy rollback)
  6. Evaluate on held-out data before merging
  7. Use QLoRA for 70B+ models on consumer hardware

References

Resources