Skip to main content
Scrape sites with stealth browsing and Cloudflare bypass.

Skill metadata

Reference: full SKILL.md

The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.

Scrapling

Scrapling is a web scraping framework with anti-bot bypass, stealth browser automation, and a spider framework. It provides three fetching strategies (HTTP, dynamic JS, stealth/Cloudflare) and a full CLI. This skill is for educational and research purposes only. Users must comply with local/international data scraping laws and respect website Terms of Service.

When to Use

  • Scraping static HTML pages (faster than browser tools)
  • Scraping JS-rendered pages that need a real browser
  • Bypassing Cloudflare Turnstile or bot detection
  • Crawling multiple pages with a spider
  • When the built-in web_extract tool does not return the data you need

Installation

Minimal install (HTTP only, no browser):
With browser automation only:

Quick Reference

CLI Usage

Extract Static Page

With CSS selector and browser impersonation:

Extract JS-Rendered Page

Extract Cloudflare-Protected Page

POST Request

Output Formats

The output format is determined by the file extension:
  • .html — raw HTML
  • .md — converted to Markdown
  • .txt — plain text
  • .json / .jsonl — JSON

Python: HTTP Scraping

Single Request

Session (Persistent Cookies)

POST / PUT / DELETE

With Proxy

Python: Dynamic Pages (JS-Rendered)

For pages that require JavaScript execution (SPAs, lazy-loaded content):

Wait for Specific Element

Disable Resources for Speed

Blocks fonts, images, media, stylesheets (~25% faster):

Custom Page Automation

Python: Stealth Mode (Anti-Bot Bypass)

For Cloudflare-protected or heavily fingerprinted sites:

Stealth Session

Element Selection

All fetchers return a Selector object with these methods:

CSS Selectors

XPath

Find Methods

Similar Elements

Find elements with similar structure (useful for product listings, etc.):

Python: Spider Framework

For multi-page crawling with link following:

Multi-Session Spider

Route requests to different fetcher types:

Pause/Resume Crawling

Pitfalls

  • Browser install required: run scrapling install after pip install — without it, DynamicFetcher and StealthyFetcher will fail
  • Timeouts: DynamicFetcher/StealthyFetcher timeout is in milliseconds (default 30000), Fetcher timeout is in seconds
  • Cloudflare bypass: solve_cloudflare=True adds 5-15 seconds to fetch time — only enable when needed
  • Resource usage: StealthyFetcher runs a real browser — limit concurrent usage
  • Legal: always check robots.txt and website ToS before scraping. This library is for educational and research purposes
  • Python version: requires Python 3.10+