Skip to main content
Commands, package names, and image names on this page come from the open-source project that Mibyan Desktop is built on, and can differ from the Mibyan Desktop installer. For the supported Mibyan install and update path, see Install and update.
TRL: SFT, DPO, GRPO, RLOO reward modeling for LLM RLHF.

Skill metadata

Reference: full SKILL.md

The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.

TRL - Transformer Reinforcement Learning

Quick start

TRL provides post-training methods for aligning language models with human preferences. Installation:
Supervised Fine-Tuning (instruction tuning):
DPO (align with preferences):

Common workflows

Workflow 1: Full RLHF pipeline (SFT → Reward Model → RLOO)

Complete pipeline from base model to human-aligned model.
Note (TRL 1.x): PPO has been removed from TRL — PPOTrainer, PPOConfig, and python -m trl.scripts.ppo no longer exist. Use an online-RL trainer TRL still ships: RLOO (RLOOTrainer / trl rloo) is the closest drop-in for a reward-model-driven RLHF pipeline, and GRPO (GRPOTrainer / trl grpo, see Workflow 3) is the memory-efficient alternative. The step below uses RLOO.
Copy this checklist:
Step 1: Supervised fine-tuning Train base model on instruction-following data:
Step 2: Train reward model Train model to predict human preferences:
Step 3: RLOO reinforcement learning Optimize policy using the reward model. PPO was removed in TRL 1.x; use the RLOO CLI (trl rloo) with the trained reward model passed via --reward_model_name_or_path:
Equivalent Python (RLOOTrainer / RLOOConfig):
Step 4: Evaluate

Workflow 2: Simple preference alignment with DPO

Align model with preferences without reward model. Copy this checklist:
Step 1: Prepare preference dataset Dataset format:
Load dataset:
Step 2: Configure DPO
Step 3: Train with DPOTrainer
CLI alternative:

Workflow 3: Memory-efficient online RL with GRPO

Train with reinforcement learning using minimal memory. For in-depth GRPO guidance — reward function design, critical training insights (loss behavior, mode collapse, tuning), and advanced multi-stage patterns — see references/grpo-training.md. A production-ready training script is in templates/basic_grpo_training.py. Copy this checklist:
Step 1: Define reward function
Or use a reward model:
Step 2: Configure GRPO
Step 3: Train with GRPOTrainer
CLI:

When to use vs alternatives

Use TRL when:
  • Need to align model with human preferences
  • Have preference data (chosen/rejected pairs)
  • Want to use reinforcement learning (RLOO, GRPO)
  • Need reward model training
  • Doing RLHF (full pipeline)
Method selection:
  • SFT: Have prompt-completion pairs, want basic instruction following
  • DPO: Have preferences, want simple alignment (no reward model needed)
  • RLOO: Have a reward model, want online RL (the reward-model-driven RLHF path; PPO was removed in TRL 1.x)
  • GRPO: Memory-constrained, want online RL with reward functions
  • Reward Model: Building RLHF pipeline, need to score generations
Use alternatives instead:
  • HuggingFace Trainer: Basic fine-tuning without RL
  • Axolotl: YAML-based training configuration
  • LitGPT: Educational, minimal fine-tuning
  • Unsloth: Fast LoRA training

Common issues

Issue: OOM during DPO training Reduce batch size and sequence length:
Or use gradient checkpointing:
Issue: Poor alignment quality Tune beta parameter:
Issue: Reward model not learning Check loss type and learning rate:
Ensure preference dataset has clear winners:
Issue: Online RL (RLOO/GRPO) training unstable Adjust the KL/beta regularization toward the reference policy:

Advanced topics

SFT training guide: See references/sft-training.md for dataset formats, chat templates, packing strategies, and multi-GPU training. DPO variants: See references/dpo-variants.md for IPO, cDPO, RPO, and other DPO loss functions with recommended hyperparameters. Reward modeling: See references/reward-modeling.md for outcome vs process rewards, Bradley-Terry loss, and reward model evaluation. Online RL methods: See references/online-rl.md for PPO, GRPO, RLOO, and OnlineDPO with detailed configurations. GRPO deep dive: See references/grpo-training.md for expert-level GRPO patterns — reward function design philosophy, training insights (why loss increases, mode collapse detection), hyperparameter tuning, multi-stage training, and troubleshooting. Production-ready template in templates/basic_grpo_training.py.

Hardware requirements

  • GPU: NVIDIA (CUDA required)
  • VRAM: Depends on model and method
    • SFT 7B: 16GB (with LoRA)
    • DPO 7B: 24GB (stores reference model)
    • RLOO 7B: 40GB (policy + reward model)
    • GRPO 7B: 24GB (more memory efficient)
  • Multi-GPU: Supported via accelerate
  • Mixed precision: BF16 recommended (A100/H100)
Memory optimization:
  • Use LoRA/QLoRA for all methods
  • Enable gradient checkpointing
  • Use smaller batch sizes with gradient accumulation

Resources