Skill metadata
Reference: full SKILL.md
The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.
Scrapling
Scrapling is a web scraping framework with anti-bot bypass, stealth browser automation, and a spider framework. It provides three fetching strategies (HTTP, dynamic JS, stealth/Cloudflare) and a full CLI. This skill is for educational and research purposes only. Users must comply with local/international data scraping laws and respect website Terms of Service.When to Use
- Scraping static HTML pages (faster than browser tools)
- Scraping JS-rendered pages that need a real browser
- Bypassing Cloudflare Turnstile or bot detection
- Crawling multiple pages with a spider
- When the built-in
web_extracttool does not return the data you need
Installation
Quick Reference
CLI Usage
Extract Static Page
Extract JS-Rendered Page
Extract Cloudflare-Protected Page
POST Request
Output Formats
The output format is determined by the file extension:.html— raw HTML.md— converted to Markdown.txt— plain text.json/.jsonl— JSON
Python: HTTP Scraping
Single Request
Session (Persistent Cookies)
POST / PUT / DELETE
With Proxy
Python: Dynamic Pages (JS-Rendered)
For pages that require JavaScript execution (SPAs, lazy-loaded content):Wait for Specific Element
Disable Resources for Speed
Blocks fonts, images, media, stylesheets (~25% faster):Custom Page Automation
Python: Stealth Mode (Anti-Bot Bypass)
For Cloudflare-protected or heavily fingerprinted sites:Stealth Session
Element Selection
All fetchers return aSelector object with these methods:
CSS Selectors
XPath
Find Methods
Similar Elements
Find elements with similar structure (useful for product listings, etc.):Navigation
Python: Spider Framework
For multi-page crawling with link following:Multi-Session Spider
Route requests to different fetcher types:Pause/Resume Crawling
Pitfalls
- Browser install required: run
scrapling installafter pip install — without it,DynamicFetcherandStealthyFetcherwill fail - Timeouts: DynamicFetcher/StealthyFetcher timeout is in milliseconds (default 30000), Fetcher timeout is in seconds
- Cloudflare bypass:
solve_cloudflare=Trueadds 5-15 seconds to fetch time — only enable when needed - Resource usage: StealthyFetcher runs a real browser — limit concurrent usage
- Legal: always check robots.txt and website ToS before scraping. This library is for educational and research purposes
- Python version: requires Python 3.10+

