Skip to main content
Commands, package names, and image names on this page come from the open-source project that Mibyan Desktop is built on, and can differ from the Mibyan Desktop installer. For the supported Mibyan install and update path, see Install and update.
vLLM: high-throughput LLM serving, OpenAI API, quantization.

Skill metadata

Reference: full SKILL.md

The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.

vLLM - High-Performance LLM Serving

When to use

Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.

Quick start

vLLM achieves 24x higher throughput than standard transformers through PagedAttention (block-based KV cache) and continuous batching (mixing prefill/decode requests). Installation:
Basic offline inference:
OpenAI-compatible server:

Common workflows

Workflow 1: Production API deployment

Copy this checklist and track progress:
Step 1: Configure server settings Choose configuration based on your model size:
Step 2: Test with limited traffic Run load test before production:
Verify TTFT (time to first token) < 500ms and throughput > 100 req/sec. Step 3: Enable monitoring vLLM exposes Prometheus metrics at /metrics on the API port (default 8000):
Key metrics to monitor:
  • vllm:time_to_first_token_seconds - Latency
  • vllm:num_requests_running - Active requests
  • vllm:gpu_cache_usage_perc - KV cache utilization
Step 4: Deploy to production Use Docker for consistent deployment:
Step 5: Verify performance metrics Check that deployment meets targets:
  • TTFT < 500ms (for short prompts)
  • Throughput > target req/sec
  • GPU utilization > 80%
  • No OOM errors in logs

Workflow 2: Offline batch inference

For processing large datasets without server overhead. Copy this checklist:
Step 1: Prepare input data
Step 2: Configure LLM engine
Step 3: Run batch inference vLLM automatically batches requests for efficiency:
Step 4: Process results

Workflow 3: Quantized model serving

Fit large models in limited GPU memory.
Step 1: Choose quantization method
  • AWQ: Best for 70B models, minimal accuracy loss
  • GPTQ: Wide model support, good compression
  • FP8: Fastest on H100 GPUs
Step 2: Find or create quantized model Use pre-quantized models from HuggingFace:
Step 3: Launch with quantization flag
Step 4: Verify accuracy Test outputs match expected quality:

When to use vs alternatives

Use vLLM when:
  • Deploying production LLM APIs (100+ req/sec)
  • Serving OpenAI-compatible endpoints
  • Limited GPU memory but need large models
  • Multi-user applications (chatbots, assistants)
  • Need low latency with high throughput
Use alternatives instead:
  • llama.cpp: CPU/edge inference, single-user
  • HuggingFace transformers: Research, prototyping, one-off generation
  • TensorRT-LLM: NVIDIA-only, need absolute maximum performance
  • Text-Generation-Inference: Already in HuggingFace ecosystem

Common issues

Issue: Out of memory during model loading Reduce memory usage:
Or use quantization:
Issue: Slow first token (TTFT > 1 second) Enable prefix caching for repeated prompts:
For long prompts, enable chunked prefill:
Issue: Model not found error Use --trust-remote-code for custom models:
Issue: Low throughput (<50 req/sec) Increase concurrent sequences:
Check GPU utilization with nvidia-smi - should be >80%. Issue: Inference slower than expected Verify tensor parallelism uses power of 2 GPUs:
Enable speculative decoding for faster generation (pass config as JSON; --speculative-model was removed in favor of --speculative-config):

Advanced topics

Server deployment patterns: See references/server-deployment.md for Docker, Kubernetes, and load balancing configurations. Performance optimization: See references/optimization.md for PagedAttention tuning, continuous batching details, and benchmark results. Quantization guide: See references/quantization.md for AWQ/GPTQ/FP8 setup, model preparation, and accuracy comparisons. Troubleshooting: See references/troubleshooting.md for detailed error messages, debugging steps, and performance diagnostics.

Hardware requirements

  • Small models (7B-13B): 1x A10 (24GB) or A100 (40GB)
  • Medium models (30B-40B): 2x A100 (40GB) with tensor parallelism
  • Large models (70B+): 4x A100 (40GB) or 2x A100 (80GB), use AWQ/GPTQ
Supported platforms: NVIDIA (primary), AMD ROCm, Intel GPUs, TPUs

Resources