Skip to main content
Commands, package names, and image names on this page come from the open-source project that Mibyan Desktop is built on, and can differ from the Mibyan Desktop installer. For the supported Mibyan install and update path, see Install and update.
Fast BPE/WordPiece tokenization and custom vocab training.

Skill metadata

Reference: full SKILL.md

The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.

HuggingFace Tokenizers - Fast Tokenization for NLP

Fast, production-ready tokenizers with Rust performance and Python ease-of-use.

When to use HuggingFace Tokenizers

Use HuggingFace Tokenizers when:
  • Need extremely fast tokenization (<20s per GB of text)
  • Training custom tokenizers from scratch
  • Want alignment tracking (token → original text position)
  • Building production NLP pipelines
  • Need to tokenize large corpora efficiently
Performance:
  • Speed: <20 seconds to tokenize 1GB on CPU
  • Implementation: Rust core with Python/Node.js bindings
  • Efficiency: 10-100× faster than pure Python implementations
Use alternatives instead:
  • SentencePiece: Language-independent, used by T5/ALBERT
  • tiktoken: OpenAI’s BPE tokenizer for GPT models
  • transformers AutoTokenizer: Loading pretrained only (uses this library internally)

Quick start

Installation

Load pretrained tokenizer

Train custom BPE tokenizer

Training time: ~1-2 minutes for 100MB corpus, ~10-20 minutes for 1GB

Batch encoding with padding

Tokenization algorithms

BPE (Byte-Pair Encoding)

How it works:
  1. Start with character-level vocabulary
  2. Find most frequent character pair
  3. Merge into new token, add to vocabulary
  4. Repeat until vocabulary size reached
Used by: GPT-2, GPT-3, RoBERTa, BART, DeBERTa
Advantages:
  • Handles OOV words well (breaks into subwords)
  • Flexible vocabulary size
  • Good for morphologically rich languages
Trade-offs:
  • Tokenization depends on merge order
  • May split common words unexpectedly

WordPiece

How it works:
  1. Start with character vocabulary
  2. Score merge pairs: frequency(pair) / (frequency(first) × frequency(second))
  3. Merge highest scoring pair
  4. Repeat until vocabulary size reached
Used by: BERT, DistilBERT, MobileBERT
Advantages:
  • Prioritizes meaningful merges (high score = semantically related)
  • Used successfully in BERT (state-of-the-art results)
Trade-offs:
  • Unknown words become [UNK] if no subword match
  • Saves vocabulary, not merge rules (larger files)

Unigram

How it works:
  1. Start with large vocabulary (all substrings)
  2. Compute loss for corpus with current vocabulary
  3. Remove tokens with minimal impact on loss
  4. Repeat until vocabulary size reached
Used by: ALBERT, T5, mBART, XLNet (via SentencePiece)
Advantages:
  • Probabilistic (finds most likely tokenization)
  • Works well for languages without word boundaries
  • Handles diverse linguistic contexts
Trade-offs:
  • Computationally expensive to train
  • More hyperparameters to tune

Tokenization pipeline

Complete pipeline: Normalization → Pre-tokenization → Model → Post-processing

Normalization

Clean and standardize text:
Common normalizers:
  • NFD, NFC, NFKD, NFKC - Unicode normalization forms
  • Lowercase() - Convert to lowercase
  • StripAccents() - Remove accents (é → e)
  • Strip() - Remove whitespace
  • Replace(pattern, content) - Regex replacement

Pre-tokenization

Split text into word-like units:
Common pre-tokenizers:
  • Whitespace() - Split on spaces, tabs, newlines
  • ByteLevel() - GPT-2 style byte-level splitting
  • Punctuation() - Isolate punctuation
  • Digits(individual_digits=True) - Split digits individually
  • Metaspace() - Replace spaces with ▁ (SentencePiece style)

Post-processing

Add special tokens for model input:
Common patterns:

Alignment tracking

Track token positions in original text:
Use cases:
  • Named entity recognition (map predictions back to text)
  • Question answering (extract answer spans)
  • Token classification (align labels to original positions)

Integration with transformers

Load with AutoTokenizer

Convert custom tokenizer to transformers

Common patterns

Train from iterator (large datasets)

Performance: Processes 1GB in ~10-20 minutes

Enable truncation and padding

Multi-processing

Speedup: 5-8× with 8 cores

Performance benchmarks

Training speed

Hardware: 16-core CPU, tested on English Wikipedia

Tokenization speed

Test: English text, average sentence length 20 words

Memory usage

Supported models

Pre-trained tokenizers available via from_pretrained(): BERT family:
  • bert-base-uncased, bert-large-cased
  • distilbert-base-uncased
  • roberta-base, roberta-large
GPT family:
  • gpt2, gpt2-medium, gpt2-large
  • distilgpt2
T5 family:
  • t5-small, t5-base, t5-large
  • google/flan-t5-xxl
Other:
  • facebook/bart-base, facebook/mbart-large-cc25
  • albert-base-v2, albert-xlarge-v2
  • xlm-roberta-base, xlm-roberta-large
Browse all: https://huggingface.co/models?library=tokenizers

References

Resources