Skip to main content
Commands, package names, and image names on this page come from the open-source project that Mibyan Desktop is built on, and can differ from the Mibyan Desktop installer. For the supported Mibyan install and update path, see Install and update.
Curate LLM training data: dedupe, filter, PII redaction.

Skill metadata

Reference: full SKILL.md

The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.

NeMo Curator - GPU-Accelerated Data Curation

NVIDIA’s toolkit for preparing high-quality training data for LLMs.

When to use NeMo Curator

Use NeMo Curator when:
  • Preparing LLM training data from web scrapes (Common Crawl)
  • Need fast deduplication (16× faster than CPU)
  • Curating multi-modal datasets (text, images, video, audio)
  • Filtering low-quality or toxic content
  • Scaling data processing across GPU cluster
Performance:
  • 16× faster fuzzy deduplication (8TB RedPajama v2)
  • 40% lower TCO vs CPU alternatives
  • Near-linear scaling across GPU nodes
Use alternatives instead:
  • datatrove: CPU-based, open-source data processing
  • dolma: Allen AI’s data toolkit
  • Ray Data: General ML data processing (no curation focus)

Quick start

Installation

Basic text curation pipeline

Major version rewrite (1.x): NeMo Curator was rewritten around a Ray-based pipeline/stage architecture. The old DocumentDataset + nemo_curator.modules.* / ScoreFilter / Modify call-the-object-on-a-dataset API from 0.x is gone. In 1.x you compose ProcessingStages into a Pipeline and run it with an executor. The exact stage/import surface differs per modality — treat the examples in this skill below as conceptual (0.x-style) and follow the current quickstart and text guide for the exact 1.x APIs rather than copying imports verbatim.
Shape of a 1.x pipeline (from the upstream quickstart):
The 0.x-style snippets in the sections that follow illustrate the concepts (quality filtering, exact/fuzzy/semantic dedup, PII redaction, classifier filtering). For runnable 1.x code, map each concept onto the corresponding stage from the modality guide.

Data curation pipeline

Stage 1: Quality filtering

Stage 2: Deduplication

Exact deduplication:
Fuzzy deduplication (16× faster on GPU):
Semantic deduplication:

Stage 3: PII redaction

Stage 4: Classifier filtering

GPU acceleration

GPU vs CPU performance

Multi-GPU scaling

Multi-modal curation

Image curation

Video curation

Audio curation

Common patterns

Web scrape curation (Common Crawl)

Distributed processing

Performance benchmarks

Fuzzy deduplication (8TB RedPajama v2)

  • CPU (256 cores): 120 hours
  • GPU (8× A100): 7.5 hours
  • Speedup: 16×

Exact deduplication (1TB)

  • CPU (64 cores): 8 hours
  • GPU (4× A100): 0.5 hours
  • Speedup: 16×

Quality filtering (100GB)

  • CPU (32 cores): 2 hours
  • GPU (2× A100): 0.2 hours
  • Speedup: 10×

Cost comparison

CPU-based curation (AWS c5.18xlarge × 10):
  • Cost: 3.60/hour×10=3.60/hour × 10 = 36/hour
  • Time for 8TB: 120 hours
  • Total: $4,320
GPU-based curation (AWS p4d.24xlarge × 2):
  • Cost: 32.77/hour×2=32.77/hour × 2 = 65.54/hour
  • Time for 8TB: 7.5 hours
  • Total: $491.55
Savings: 89% reduction ($3,828 saved)

Supported data formats

  • Input: Parquet, JSONL, CSV
  • Output: Parquet (recommended), JSONL
  • WebDataset: TAR archives for multi-modal

Use cases

Production deployments:
  • NVIDIA used NeMo Curator to prepare Nemotron-4 training data
  • Open-source datasets curated: RedPajama v2, The Pile

References

Resources