Skill metadata
Reference: full SKILL.md
The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.
slime: LLM Post-Training Framework for RL Scaling
slime is an LLM post-training framework from Tsinghua’s THUDM team, powering GLM-4.5, GLM-4.6, and GLM-4.7. It connects Megatron-LM for training with SGLang for high-throughput rollout generation.When to Use slime
Choose slime when you need:- Megatron-LM native training with SGLang inference
- Custom data generation workflows with flexible data buffers
- Training GLM, Qwen3, DeepSeek V3, or Llama 3 models
- Research-grade framework with production backing (Z.ai)
- You need enterprise-grade stability features → use miles
- You want flexible backend swapping → use verl
- You need PyTorch-native abstractions → use torchforge
Key Features
- Training: Megatron-LM with full parallelism support (TP, PP, DP, SP)
- Rollout: SGLang-based high-throughput generation with router
- Data Buffer: Flexible prompt management and sample storage
- Models: GLM-4.x, Qwen3, DeepSeek V3/R1, Llama 3
Architecture Overview
Installation
From Source
Quick Start: GRPO Training
Workflow 1: Standard GRPO Training
Use this workflow for training reasoning models with group-relative advantages.Prerequisites Checklist
- Docker environment or Megatron-LM + SGLang installed
- Model checkpoint (HuggingFace or Megatron format)
- Training data in JSONL format
Step 1: Prepare Data
Step 2: Configure Model
Choose a pre-configured model script:Step 3: Launch Training
Step 4: Monitor Training
- Check TensorBoard:
tensorboard --logdir outputs/ - Verify reward curves are increasing
- Monitor GPU utilization across nodes
Workflow 2: Asynchronous Training
Use async mode for higher throughput by overlapping rollout and training.When to Use Async
- Large models with long generation times
- High GPU idle time in synchronous mode
- Sufficient memory for buffering
Launch Async Training
Async-Specific Parameters
Workflow 3: Multi-Turn Agentic Training
Use this workflow for training agents with tool use or multi-step reasoning.Prerequisites
- Custom generate function for multi-turn logic
- Tool/environment interface
Step 1: Define Custom Generate Function
Step 2: Launch with Custom Function
examples/search-r1/ for a complete multi-turn search example.
Configuration Reference
Three Argument Categories
slime uses three types of arguments: 1. Megatron Arguments (passed directly):--sglang-):
Key Constraints
Data Buffer System
slime’s data buffer enables flexible data management:Basic Data Source
Buffered Data Source (Off-Policy)
Common Issues and Solutions
Issue: SGLang Engine Crash
Symptoms: Inference engine dies mid-training Solutions:Issue: Weight Sync Timeout
Symptoms: Training hangs after rollout Solutions:Issue: OOM During Training
Symptoms: CUDA OOM in backward pass Solutions:Issue: Slow Data Loading
Symptoms: GPU idle during data fetch Solutions:Supported Models
Each model has pre-configured scripts in
scripts/models/.
Advanced Topics
Co-location Mode
Share GPUs between training and inference to reduce memory:Custom Reward Model
Evaluation Multi-Task
Resources
- Documentation: https://thudm.github.io/slime/
- GitHub: https://github.com/THUDM/slime
- Blog: https://lmsys.org/blog/2025-07-09-slime/
- Examples: See
examples/directory for 14+ worked examples

