Commands, package names, and image names on this page come from the open-source project that Mibyan Desktop is built on, and can differ from the Mibyan Desktop installer. For the supported Mibyan install and update path, see Install and update.
Skill metadata
Reference: full SKILL.md
The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.
TRL - Transformer Reinforcement Learning
Quick start
TRL provides post-training methods for aligning language models with human preferences. Installation:Common workflows
Workflow 1: Full RLHF pipeline (SFT → Reward Model → RLOO)
Complete pipeline from base model to human-aligned model.Note (TRL 1.x): PPO has been removed from TRL —Copy this checklist:PPOTrainer,PPOConfig, andpython -m trl.scripts.ppono longer exist. Use an online-RL trainer TRL still ships: RLOO (RLOOTrainer/trl rloo) is the closest drop-in for a reward-model-driven RLHF pipeline, and GRPO (GRPOTrainer/trl grpo, see Workflow 3) is the memory-efficient alternative. The step below uses RLOO.
trl rloo) with the trained reward model passed via --reward_model_name_or_path:
RLOOTrainer / RLOOConfig):
Workflow 2: Simple preference alignment with DPO
Align model with preferences without reward model. Copy this checklist:Workflow 3: Memory-efficient online RL with GRPO
Train with reinforcement learning using minimal memory. For in-depth GRPO guidance — reward function design, critical training insights (loss behavior, mode collapse, tuning), and advanced multi-stage patterns — see references/grpo-training.md. A production-ready training script is in templates/basic_grpo_training.py. Copy this checklist:When to use vs alternatives
Use TRL when:- Need to align model with human preferences
- Have preference data (chosen/rejected pairs)
- Want to use reinforcement learning (RLOO, GRPO)
- Need reward model training
- Doing RLHF (full pipeline)
- SFT: Have prompt-completion pairs, want basic instruction following
- DPO: Have preferences, want simple alignment (no reward model needed)
- RLOO: Have a reward model, want online RL (the reward-model-driven RLHF path; PPO was removed in TRL 1.x)
- GRPO: Memory-constrained, want online RL with reward functions
- Reward Model: Building RLHF pipeline, need to score generations
- HuggingFace Trainer: Basic fine-tuning without RL
- Axolotl: YAML-based training configuration
- LitGPT: Educational, minimal fine-tuning
- Unsloth: Fast LoRA training
Common issues
Issue: OOM during DPO training Reduce batch size and sequence length:Advanced topics
SFT training guide: See references/sft-training.md for dataset formats, chat templates, packing strategies, and multi-GPU training. DPO variants: See references/dpo-variants.md for IPO, cDPO, RPO, and other DPO loss functions with recommended hyperparameters. Reward modeling: See references/reward-modeling.md for outcome vs process rewards, Bradley-Terry loss, and reward model evaluation. Online RL methods: See references/online-rl.md for PPO, GRPO, RLOO, and OnlineDPO with detailed configurations. GRPO deep dive: See references/grpo-training.md for expert-level GRPO patterns — reward function design philosophy, training insights (why loss increases, mode collapse detection), hyperparameter tuning, multi-stage training, and troubleshooting. Production-ready template in templates/basic_grpo_training.py.Hardware requirements
- GPU: NVIDIA (CUDA required)
- VRAM: Depends on model and method
- SFT 7B: 16GB (with LoRA)
- DPO 7B: 24GB (stores reference model)
- RLOO 7B: 40GB (policy + reward model)
- GRPO 7B: 24GB (more memory efficient)
- Multi-GPU: Supported via
accelerate - Mixed precision: BF16 recommended (A100/H100)
- Use LoRA/QLoRA for all methods
- Enable gradient checkpointing
- Use smaller batch sizes with gradient accumulation
Resources
- Docs: https://huggingface.co/docs/trl/
- GitHub: https://github.com/huggingface/trl
- Papers:
- “Training language models to follow instructions with human feedback” (InstructGPT, 2022)
- “Direct Preference Optimization: Your Language Model is Secretly a Reward Model” (DPO, 2023)
- “Group Relative Policy Optimization” (GRPO, 2024)
- Examples: https://github.com/huggingface/trl/tree/main/examples/scripts

