Commands, package names, and image names on this page come from the open-source project that Mibyan Desktop is built on, and can differ from the Mibyan Desktop installer. For the supported Mibyan install and update path, see Install and update.
Skill metadata
Reference: full SKILL.md
The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.
NeMo Curator - GPU-Accelerated Data Curation
NVIDIA’s toolkit for preparing high-quality training data for LLMs.When to use NeMo Curator
Use NeMo Curator when:- Preparing LLM training data from web scrapes (Common Crawl)
- Need fast deduplication (16× faster than CPU)
- Curating multi-modal datasets (text, images, video, audio)
- Filtering low-quality or toxic content
- Scaling data processing across GPU cluster
- 16× faster fuzzy deduplication (8TB RedPajama v2)
- 40% lower TCO vs CPU alternatives
- Near-linear scaling across GPU nodes
- datatrove: CPU-based, open-source data processing
- dolma: Allen AI’s data toolkit
- Ray Data: General ML data processing (no curation focus)
Quick start
Installation
Basic text curation pipeline
Major version rewrite (1.x): NeMo Curator was rewritten around a Ray-based pipeline/stage architecture. The oldShape of a 1.x pipeline (from the upstream quickstart):DocumentDataset+nemo_curator.modules.*/ScoreFilter/Modifycall-the-object-on-a-dataset API from 0.x is gone. In 1.x you composeProcessingStages into aPipelineand run it with an executor. The exact stage/import surface differs per modality — treat the examples in this skill below as conceptual (0.x-style) and follow the current quickstart and text guide for the exact 1.x APIs rather than copying imports verbatim.
Data curation pipeline
Stage 1: Quality filtering
Stage 2: Deduplication
Exact deduplication:Stage 3: PII redaction
Stage 4: Classifier filtering
GPU acceleration
GPU vs CPU performance
Multi-GPU scaling
Multi-modal curation
Image curation
Video curation
Audio curation
Common patterns
Web scrape curation (Common Crawl)
Distributed processing
Performance benchmarks
Fuzzy deduplication (8TB RedPajama v2)
- CPU (256 cores): 120 hours
- GPU (8× A100): 7.5 hours
- Speedup: 16×
Exact deduplication (1TB)
- CPU (64 cores): 8 hours
- GPU (4× A100): 0.5 hours
- Speedup: 16×
Quality filtering (100GB)
- CPU (32 cores): 2 hours
- GPU (2× A100): 0.2 hours
- Speedup: 10×
Cost comparison
CPU-based curation (AWS c5.18xlarge × 10):- Cost: 36/hour
- Time for 8TB: 120 hours
- Total: $4,320
- Cost: 65.54/hour
- Time for 8TB: 7.5 hours
- Total: $491.55
Supported data formats
- Input: Parquet, JSONL, CSV
- Output: Parquet (recommended), JSONL
- WebDataset: TAR archives for multi-modal
Use cases
Production deployments:- NVIDIA used NeMo Curator to prepare Nemotron-4 training data
- Open-source datasets curated: RedPajama v2, The Pile
References
- Filtering Guide - 30+ quality filters, heuristics
- Deduplication Guide - Exact, fuzzy, semantic methods
Resources
- GitHub: https://github.com/NVIDIA-NeMo/Curator
- Docs: https://docs.nvidia.com/nemo/curator/latest/
- Version: 1.2.0 (1.x is a Ray-based pipeline rewrite — see the quickstart before copying 0.x snippets)
- License: Apache 2.0

