Skip to main content
RL post-training for LLMs with Megatron and SGLang.

Skill metadata

Reference: full SKILL.md

The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.

slime: LLM Post-Training Framework for RL Scaling

slime is an LLM post-training framework from Tsinghua’s THUDM team, powering GLM-4.5, GLM-4.6, and GLM-4.7. It connects Megatron-LM for training with SGLang for high-throughput rollout generation.

When to Use slime

Choose slime when you need:
  • Megatron-LM native training with SGLang inference
  • Custom data generation workflows with flexible data buffers
  • Training GLM, Qwen3, DeepSeek V3, or Llama 3 models
  • Research-grade framework with production backing (Z.ai)
Consider alternatives when:
  • You need enterprise-grade stability features → use miles
  • You want flexible backend swapping → use verl
  • You need PyTorch-native abstractions → use torchforge

Key Features

  • Training: Megatron-LM with full parallelism support (TP, PP, DP, SP)
  • Rollout: SGLang-based high-throughput generation with router
  • Data Buffer: Flexible prompt management and sample storage
  • Models: GLM-4.x, Qwen3, DeepSeek V3/R1, Llama 3

Architecture Overview

Installation

From Source

Quick Start: GRPO Training


Workflow 1: Standard GRPO Training

Use this workflow for training reasoning models with group-relative advantages.

Prerequisites Checklist

  • Docker environment or Megatron-LM + SGLang installed
  • Model checkpoint (HuggingFace or Megatron format)
  • Training data in JSONL format

Step 1: Prepare Data

Or with chat format:

Step 2: Configure Model

Choose a pre-configured model script:

Step 3: Launch Training

Step 4: Monitor Training

  • Check TensorBoard: tensorboard --logdir outputs/
  • Verify reward curves are increasing
  • Monitor GPU utilization across nodes

Workflow 2: Asynchronous Training

Use async mode for higher throughput by overlapping rollout and training.

When to Use Async

  • Large models with long generation times
  • High GPU idle time in synchronous mode
  • Sufficient memory for buffering

Launch Async Training

Async-Specific Parameters


Workflow 3: Multi-Turn Agentic Training

Use this workflow for training agents with tool use or multi-step reasoning.

Prerequisites

  • Custom generate function for multi-turn logic
  • Tool/environment interface

Step 1: Define Custom Generate Function

Step 2: Launch with Custom Function

See examples/search-r1/ for a complete multi-turn search example.

Configuration Reference

Three Argument Categories

slime uses three types of arguments: 1. Megatron Arguments (passed directly):
2. SGLang Arguments (prefixed with --sglang-):
3. slime Arguments:

Key Constraints

Example: 32 × 8 = 256 × 1

Data Buffer System

slime’s data buffer enables flexible data management:

Basic Data Source

Buffered Data Source (Off-Policy)


Common Issues and Solutions

Issue: SGLang Engine Crash

Symptoms: Inference engine dies mid-training Solutions:

Issue: Weight Sync Timeout

Symptoms: Training hangs after rollout Solutions:

Issue: OOM During Training

Symptoms: CUDA OOM in backward pass Solutions:

Issue: Slow Data Loading

Symptoms: GPU idle during data fetch Solutions:

Supported Models

Each model has pre-configured scripts in scripts/models/.

Advanced Topics

Co-location Mode

Share GPUs between training and inference to reduce memory:

Custom Reward Model

Evaluation Multi-Task


Resources