Companion repository for the book "Autonomous Web Scraping with AI: A Config-Driven Approach" - a comprehensive guide to modern web scraping, from static HTML extraction to AI-driven autonomous data collection.
Read the book online: https://heldernoid.github.io/scrapping
.
+-- book/ Book chapters (Markdown)
| +-- part-1-foundations/ Why scraping, rendering landscape, JSON config contract
| +-- part-2-extraction/ Selectors, field types, pagination patterns
| +-- part-3-rendering/ Static fetch, Playwright, auto-fallback engine
| +-- part-4-scale/ Scheduling, storage, alerts
| +-- part-5-ai-agents/ Config as language, autonomous agents, MCP
+-- notebooks/ Jupyter notebooks (runnable code demos)
+-- demo-sites/ 4 target sites for scraping practice
| +-- shopsphere-ssr/ Amazon-like marketplace (server-side rendered)
| +-- shopsphere-csr/ Amazon-like marketplace (client-side rendered)
| +-- jobhive-ssr/ Job board (server-side rendered)
| +-- jobhive-csr/ Job board (client-side rendered)
+-- configs/ Scraper configs for all 4 demo sites
+-- agents/ AI agent code (config generator, autonomous scraper, MCP server)
+-- examples/ Quick start scripts
A scraper config is a JSON object that completely specifies a scraping task:
{
"render_mode": "static",
"sources": [{"url_template": "https://jobs.example.com/listings?page={n}",
"pagination": {"start": 1, "step": 1, "max_pages": 50}}],
"listing": {"link_selector": "a.job-link", "link_prefix": "https://jobs.example.com"},
"fields": {
"title": {"selector": "h1.job-title", "retrieve": "plaintext"},
"company": {"selector": ".company-name", "retrieve": "plaintext"},
"salary": {"selector": ".salary", "retrieve": "regexp", "pattern": "\\$([\\d,]+)"},
"location": {"selector": ".job-location", "retrieve": "plaintext"},
"tags": {"selector": ".skill-tag", "retrieve": "plaintext", "multiple": true}
}
}This config:
- Generates pagination URLs automatically (
?page=1,?page=2, ...) - Follows links from listing pages to detail pages
- Extracts structured fields from each detail page
- Can be generated by an AI agent given a URL
| Site | URL | Rendering | Scraping approach |
|---|---|---|---|
| ShopSphere SSR | http://localhost:8001 | Server-side (Flask+Jinja2) | Static httpx fetch |
| ShopSphere CSR | http://localhost:8002 | Client-side (JS fetches API) | API direct or Playwright |
| JobHive SSR | http://localhost:8003 | Server-side (Flask+Jinja2) | Static httpx fetch |
| JobHive CSR | http://localhost:8004 | Client-side (JS fetches API) | API direct or Playwright |
SSR and CSR versions of each site have identical data and identical HTML structure in the rendered DOM. The only difference is whether JavaScript must execute first to build that DOM.
Using uv (recommended):
# Install uv (if not already installed)
curl -LsSf https://astral.sh/uv/install.sh | sh
cd /path/to/scrapping
uv sync
# Activate
source .venv/bin/activate # Linux/Mac
# or: .venv\Scripts\activate # Windows
# Or run scripts directly without activating
uv run examples/quick_start.py
uv run jupyter notebook notebooks/Using pip (alternative):
cd /path/to/scrapping
python -m venv .venv
source .venv/bin/activate # Linux/Mac
# or: .venv\Scripts\activate # Windows
pip install -r requirements.txtcp .env.example .env
# Edit .env and set your OPENROUTER_API_KEYcd demo-sites
docker compose up -dVerify sites are running:
curl http://localhost:8001/ # ShopSphere SSR
curl http://localhost:8002/ # ShopSphere CSR
curl http://localhost:8003/ # JobHive SSR
curl http://localhost:8004/ # JobHive CSR
# From the repo root
jupyter notebook notebooks/Work through the notebooks in order:
01-static-scraping-basics.ipynb- httpx + BeautifulSoup fundamentals02-css-selectors.ipynb- CSS selector reference03-pagination-patterns.ipynb- All pagination approaches04-playwright-dynamic-content.ipynb- Playwright for CSR sites05-json-config-approach.ipynb- Building and running configs06-ai-agent-autonomous-scraping.ipynb- AI-driven autonomous scraping
# Check demo site status
python agents/demo.py --demo status
# Scrape ShopSphere (server-side rendered, no Playwright needed)
python agents/demo.py --demo marketplace-ssr --max-items 5
# Scrape JobHive (server-side rendered)
python agents/demo.py --demo jobboard-ssr --max-items 5
# Compare what static scraper sees on SSR vs CSR
python agents/demo.py --demo compare
# Auto-generate a config from a URL and scrape it
python agents/autonomous_scraper.py --url http://localhost:8001 --max-items 10
# Generate a config only (no scrape)
python agents/config_generator.py http://localhost:8003python agents/mcp_server.pyAdd to Claude Code settings to use the scraper as MCP tools from within Claude Code.
The server runs Flask with Jinja2 templates. Every product and job is embedded directly in the HTML response. curl http://localhost:8001/products | grep "MacBook" will find the product.
The server returns an empty HTML shell:
<div id="app"><div class="loading">Loading products...</div></div>JavaScript in the browser calls /api/products?page=1 and builds the product cards from the JSON response. curl http://localhost:8002/products | grep "MacBook" finds nothing. Playwright is required.
The key demonstration: the same CSS selectors work on both SSR and CSR versions after rendering. Only render_mode in the config differs.
| Type | Returns | Use case |
|---|---|---|
plaintext |
element.get_text() |
Most text content |
attr |
element.get(attr) |
href, src, data-* attributes |
regexp |
First capture group | Extract numbers from text |
regexpall |
List of all matches | Multiple values in one element |
Add "multiple": true to any field to collect all matching elements as a list. Works with any retrieve type:
"gallery_images": {"selector": "img.gallery-image", "retrieve": "attr", "attr": "src", "multiple": true}
"tags": {"selector": "span.tag", "retrieve": "plaintext", "multiple": true}| Value | Stops when |
|---|---|
no_results |
Link selector returns 0 elements (default) |
last_page_text |
Specified text appears in page HTML |
| Value | Behavior |
|---|---|
static |
httpx fetch only, no JavaScript |
playwright |
Always use headless browser |
auto |
Try static, escalate to Playwright if probe selector fails |
The book is in book/ as Markdown files, organized by part:
- Part 1: Foundations - Why scraping, SSR vs CSR, the JSON config contract
- Part 2: Extraction - CSS selectors, field types and transforms, pagination
- Part 3: Rendering - Static fetching, Playwright, auto-fallback engine
- Part 4: Scale - Scheduling, dual-database storage, change detection
- Part 5: AI Agents - Config as language, autonomous agents, MCP integration
- Python 3.11+
- Docker (for demo sites)
- OpenRouter API key (for AI agent features)
- Playwright browser (for CSR scraping):
playwright install chromium
If you use this book or code in your work, please cite:
@misc{monteiro_2026_19513159,
author = {Monteiro, H{\'e}lder},
title = {Autonomous Web Scraping with {AI}: A Config-Driven Approach},
month = apr,
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.19513159},
url = {https://doi.org/10.5281/zenodo.19513159}
}Or in plain text:
Hélder Monteiro. Autonomous Web Scraping with AI: A Config-Driven Approach. Zenodo, 2026. https://doi.org/10.5281/zenodo.19513159
Code: MIT. Book content: Creative Commons Attribution 4.0.