Skill metadata
Reference: full SKILL.md
The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.
LLaVA - Large Language and Vision Assistant
Open-source vision-language model for conversational image understanding.When to use LLaVA
Use when:- Building vision-language chatbots
- Visual question answering (VQA)
- Image description and captioning
- Multi-turn image conversations
- Visual instruction following
- Document understanding with images
- 23,000+ GitHub stars
- GPT-4V level capabilities (targeted)
- Apache 2.0 License
- Multiple model sizes (7B-34B params)
- GPT-4V: Highest quality, API-based
- CLIP: Simple zero-shot classification
- BLIP-2: Better for captioning only
- Flamingo: Research, not open-source
Quick start
Installation
Basic usage
Available models
CLI usage
Web UI (Gradio)
Multi-turn conversations
Common tasks
Image captioning
Visual question answering
Object detection (textual)
Scene understanding
Document understanding
Training custom model
Quantization (reduce VRAM)
Best practices
- Start with 7B model - Good quality, manageable VRAM
- Use 4-bit quantization - Reduces VRAM significantly
- GPU required - CPU inference extremely slow
- Clear prompts - Specific questions get better answers
- Multi-turn conversations - Maintain conversation context
- Temperature 0.2-0.7 - Balance creativity/consistency
- max_new_tokens 512-1024 - For detailed responses
- Batch processing - Process multiple images sequentially
Performance
On A100 GPU
Benchmarks
LLaVA achieves competitive scores on:- VQAv2: 78.5%
- GQA: 62.0%
- MM-Vet: 35.4%
- MMBench: 64.3%
Limitations
- Hallucinations - May describe things not in image
- Spatial reasoning - Struggles with precise locations
- Small text - Difficulty reading fine print
- Object counting - Imprecise for many objects
- VRAM requirements - Need powerful GPU
- Inference speed - Slower than CLIP
Integration with frameworks
LangChain
Gradio App
Resources
- GitHub: https://github.com/haotian-liu/LLaVA ⭐ 23,000+
- Paper: https://arxiv.org/abs/2304.08485
- Demo: https://llava.hliu.cc
- Models: https://huggingface.co/liuhaotian
- License: Apache 2.0

