Skill metadata
Reference: full SKILL.md
The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.
Outlines: Structured Text Generation
When to Use This Skill
Use Outlines when you need to:- Guarantee valid JSON/XML/code structure during generation
- Use Pydantic models for type-safe outputs
- Support local models (Transformers, llama.cpp, vLLM)
- Maximize inference speed with zero-overhead structured generation
- Generate against JSON schemas automatically
- Control token sampling at the grammar level
API note (Outlines 1.x): This skill targets the current v1 API. The pre-1.0 helpers (outlines.models.transformers(...),outlines.generate.json/choice/regex/...) have been removed. In v1 you create a model withoutlines.from_transformers(...)(orfrom_vllm,from_llamacpp,from_openai) and then call the model directly with an output type:model(prompt, output_type). JSON/Pydantic outputs are returned as a JSON string — validate withYourModel.model_validate_json(result).
Installation
Quick Start
Basic Example: Classification
With Pydantic Models
Core Concepts
1. Constrained Token Sampling
Outlines constrains token generation at the logit level using a compiled automaton derived from your output type. How it works:- Convert the output type (JSON/Pydantic/regex/
Literal) to a schema/grammar - Compile the grammar into a token-level automaton
- Filter invalid tokens at each step during generation
- Fast-forward when only one valid token exists
- Zero overhead: Filtering happens at token level
- Speed improvement: Fast-forward through deterministic paths
- Guaranteed validity: Invalid outputs impossible
2. Output Types
In v1 you pass the desired output type directly as the second argument.Multiple choice (Literal)
JSON via Pydantic
Regex (pass a regex string)
Numeric types
3. Model Backends
Outlines supports multiple local and API-based backends viafrom_* factories.
Transformers (Hugging Face)
llama.cpp
vLLM (High Throughput)
OpenAI (server-side constrained JSON)
4. Pydantic Integration
Outlines has first-class Pydantic support with automatic schema translation. Generation returns a JSON string; callmodel_validate_json to get an instance.
Basic Models
Nested Models
Enums and Literals
Common Patterns
Pattern 1: Data Extraction
Pattern 2: Classification
Pattern 3: Structured Forms
Pattern 4: Multi-Entity Extraction
Pattern 5: Code Generation
Pattern 6: Batch Processing
Backend Configuration
Transformers
llama.cpp
vLLM (Production)
Best Practices
1. Use Specific Types
2. Add Constraints
3. Use Enums for Categories
4. Provide Context in Prompts
5. Handle Optional Fields
6. Always Validate JSON Output
Comparison to Alternatives
When to choose Outlines:
- Using local models (Transformers, llama.cpp, vLLM)
- Need maximum inference speed
- Want Pydantic model support
- Require zero-overhead structured generation
- Control token sampling process
- Instructor: Need API models with automatic retrying
- Guidance: Need token healing and complex workflows
- LMQL: Prefer declarative query syntax
Performance Characteristics
Speed:- Zero overhead: Structured generation as fast as unconstrained
- Fast-forward optimization: Skips deterministic tokens
- 1.2-2x faster than post-generation validation approaches
- Automaton compiled once per output type (cached)
- Minimal runtime overhead
- Efficient with vLLM for high throughput
- 100% valid outputs (guaranteed by the constrained automaton)
- No retry loops needed
- Deterministic token filtering
Resources
- Documentation: https://dottxt-ai.github.io/outlines/
- GitHub: https://github.com/dottxt-ai/outlines (12k+ stars)
- Discord: https://discord.gg/R9DSu34mGd
- Blog: https://blog.dottxt.co
See Also
references/json_generation.md- Comprehensive JSON and Pydantic patternsreferences/backends.md- Backend-specific configurationreferences/examples.md- Production-ready examples

