Skip to main content
Commands, package names, and image names on this page come from the open-source project that Mibyan Desktop is built on, and can differ from the Mibyan Desktop installer. For the supported Mibyan install and update path, see Install and update.
lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.).

Skill metadata

Reference: full SKILL.md

The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.

lm-evaluation-harness - LLM Benchmarking

What’s inside

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

Quick start

lm-evaluation-harness evaluates LLMs across 60+ academic benchmarks using standardized prompts and metrics. Installation:
Evaluate any HuggingFace model:
View available tasks:

Common workflows

Workflow 1: Standard benchmark evaluation

Evaluate model on core benchmarks (MMLU, GSM8K, HumanEval). Copy this checklist:
Step 1: Choose benchmark suite Core reasoning benchmarks:
  • MMLU (Massive Multitask Language Understanding) - 57 subjects, multiple choice
  • GSM8K - Grade school math word problems
  • HellaSwag - Common sense reasoning
  • TruthfulQA - Truthfulness and factuality
  • ARC (AI2 Reasoning Challenge) - Science questions
Code benchmarks:
  • HumanEval - Python code generation (164 problems)
  • MBPP (Mostly Basic Python Problems) - Python coding
Standard suite (recommended for model releases):
Step 2: Configure model HuggingFace model:
Quantized model (4-bit/8-bit):
Custom checkpoint:
Step 3: Run evaluation
Step 4: Analyze results Results saved to results/llama2-7b-eval.json:

Workflow 2: Track training progress

Evaluate checkpoints during training.
Step 1: Set up periodic evaluation Evaluate every N training steps:
Step 2: Choose quick benchmarks Fast benchmarks for frequent evaluation:
  • HellaSwag: ~10 minutes on 1 GPU
  • GSM8K: ~5 minutes
  • PIQA: ~2 minutes
Avoid for frequent eval (too slow):
  • MMLU: ~2 hours (57 subjects)
  • HumanEval: Requires code execution
Step 3: Automate evaluation Integrate with training script:
Or use PyTorch Lightning callbacks:
Step 4: Plot learning curves

Workflow 3: Compare multiple models

Benchmark suite for model comparison.
Step 1: Define model list
Step 2: Run evaluations
Step 3: Generate comparison table
Output:

Workflow 4: Evaluate with vLLM (faster inference)

Use vLLM backend for 5-10x faster evaluation.
Step 1: Install vLLM
Step 2: Configure vLLM backend
Step 3: Run evaluation vLLM is 5-10× faster than standard HuggingFace:

When to use vs alternatives

Use lm-evaluation-harness when:
  • Benchmarking models for academic papers
  • Comparing model quality across standard tasks
  • Tracking training progress
  • Reporting standardized metrics (everyone uses same prompts)
  • Need reproducible evaluation
Use alternatives instead:
  • HELM (Stanford): Broader evaluation (fairness, efficiency, calibration)
  • AlpacaEval: Instruction-following evaluation with LLM judges
  • MT-Bench: Conversational multi-turn evaluation
  • Custom scripts: Domain-specific evaluation

Common issues

Issue: Evaluation too slow Use vLLM backend:
Or reduce fewshot examples:
Or evaluate subset of MMLU:
Issue: Out of memory Reduce batch size:
Use quantization:
Enable CPU offloading:
Issue: Different results than reported Check fewshot count:
Check exact task name:
Verify model and tokenizer match:
Issue: HumanEval not executing code Code-executing tasks (HumanEval, MBPP, etc.) are gated behind an explicit confirmation flag — you must pass --confirm_run_unsafe_code to run them:
Without this flag lm-eval refuses to run the task rather than silently skipping code execution.

Advanced topics

Benchmark descriptions: See references/benchmark-guide.md for detailed description of all 60+ tasks, what they measure, and interpretation. Custom tasks: See references/custom-tasks.md for creating domain-specific evaluation tasks. API evaluation: See references/api-evaluation.md for evaluating OpenAI, Anthropic, and other API models. Multi-GPU strategies: See references/distributed-eval.md for data parallel and tensor parallel evaluation.

Hardware requirements

  • GPU: NVIDIA (CUDA 11.8+), works on CPU (very slow)
  • VRAM:
    • 7B model: 16GB (bf16) or 8GB (8-bit)
    • 13B model: 28GB (bf16) or 14GB (8-bit)
    • 70B model: Requires multi-GPU or quantization
  • Time (7B model, single A100):
    • HellaSwag: 10 minutes
    • GSM8K: 5 minutes
    • MMLU (full): 2 hours
    • HumanEval: 20 minutes

Resources