Skip to content

Repository files navigation

Autonomous Web Scraping with AI: A Config-Driven Approach

Companion repository for the book "Autonomous Web Scraping with AI: A Config-Driven Approach" - a comprehensive guide to modern web scraping, from static HTML extraction to AI-driven autonomous data collection.

Read the book online: https://heldernoid.github.io/scrapping

What's in This Repository

.
+-- book/                    Book chapters (Markdown)
|   +-- part-1-foundations/  Why scraping, rendering landscape, JSON config contract
|   +-- part-2-extraction/   Selectors, field types, pagination patterns
|   +-- part-3-rendering/    Static fetch, Playwright, auto-fallback engine
|   +-- part-4-scale/        Scheduling, storage, alerts
|   +-- part-5-ai-agents/    Config as language, autonomous agents, MCP
+-- notebooks/               Jupyter notebooks (runnable code demos)
+-- demo-sites/              4 target sites for scraping practice
|   +-- shopsphere-ssr/      Amazon-like marketplace (server-side rendered)
|   +-- shopsphere-csr/      Amazon-like marketplace (client-side rendered)
|   +-- jobhive-ssr/         Job board (server-side rendered)
|   +-- jobhive-csr/         Job board (client-side rendered)
+-- configs/                 Scraper configs for all 4 demo sites
+-- agents/                  AI agent code (config generator, autonomous scraper, MCP server)
+-- examples/                Quick start scripts

The Core Idea

A scraper config is a JSON object that completely specifies a scraping task:

{
  "render_mode": "static",
  "sources": [{"url_template": "https://jobs.example.com/listings?page={n}",
               "pagination": {"start": 1, "step": 1, "max_pages": 50}}],
  "listing": {"link_selector": "a.job-link", "link_prefix": "https://jobs.example.com"},
  "fields": {
    "title":    {"selector": "h1.job-title",     "retrieve": "plaintext"},
    "company":  {"selector": ".company-name",     "retrieve": "plaintext"},
    "salary":   {"selector": ".salary",          "retrieve": "regexp", "pattern": "\\$([\\d,]+)"},
    "location": {"selector": ".job-location",    "retrieve": "plaintext"},
    "tags":     {"selector": ".skill-tag",       "retrieve": "plaintext", "multiple": true}
  }
}

This config:

  • Generates pagination URLs automatically (?page=1, ?page=2, ...)
  • Follows links from listing pages to detail pages
  • Extracts structured fields from each detail page
  • Can be generated by an AI agent given a URL

The Four Demo Sites

Site URL Rendering Scraping approach
ShopSphere SSR http://localhost:8001 Server-side (Flask+Jinja2) Static httpx fetch
ShopSphere CSR http://localhost:8002 Client-side (JS fetches API) API direct or Playwright
JobHive SSR http://localhost:8003 Server-side (Flask+Jinja2) Static httpx fetch
JobHive CSR http://localhost:8004 Client-side (JS fetches API) API direct or Playwright

SSR and CSR versions of each site have identical data and identical HTML structure in the rendered DOM. The only difference is whether JavaScript must execute first to build that DOM.

Quick Start

1. Set Up Python Environment

Using uv (recommended):

# Install uv (if not already installed)
curl -LsSf https://astral.sh/uv/install.sh | sh

cd /path/to/scrapping
uv sync

# Activate
source .venv/bin/activate  # Linux/Mac
# or: .venv\Scripts\activate  # Windows

# Or run scripts directly without activating
uv run examples/quick_start.py
uv run jupyter notebook notebooks/

Using pip (alternative):

cd /path/to/scrapping
python -m venv .venv
source .venv/bin/activate  # Linux/Mac
# or: .venv\Scripts\activate  # Windows
pip install -r requirements.txt

2. Configure API Key

cp .env.example .env
# Edit .env and set your OPENROUTER_API_KEY

3. Start Demo Sites

cd demo-sites
docker compose up -d

Verify sites are running:

curl http://localhost:8001/  # ShopSphere SSR
curl http://localhost:8002/  # ShopSphere CSR
curl http://localhost:8003/  # JobHive SSR
curl http://localhost:8004/  # JobHive CSR

4. Run the Jupyter Notebooks

# From the repo root
jupyter notebook notebooks/

Work through the notebooks in order:

  1. 01-static-scraping-basics.ipynb - httpx + BeautifulSoup fundamentals
  2. 02-css-selectors.ipynb - CSS selector reference
  3. 03-pagination-patterns.ipynb - All pagination approaches
  4. 04-playwright-dynamic-content.ipynb - Playwright for CSR sites
  5. 05-json-config-approach.ipynb - Building and running configs
  6. 06-ai-agent-autonomous-scraping.ipynb - AI-driven autonomous scraping

5. Run the Agent Demo

# Check demo site status
python agents/demo.py --demo status

# Scrape ShopSphere (server-side rendered, no Playwright needed)
python agents/demo.py --demo marketplace-ssr --max-items 5

# Scrape JobHive (server-side rendered)
python agents/demo.py --demo jobboard-ssr --max-items 5

# Compare what static scraper sees on SSR vs CSR
python agents/demo.py --demo compare

# Auto-generate a config from a URL and scrape it
python agents/autonomous_scraper.py --url http://localhost:8001 --max-items 10

# Generate a config only (no scrape)
python agents/config_generator.py http://localhost:8003

6. Start the MCP Server

python agents/mcp_server.py

Add to Claude Code settings to use the scraper as MCP tools from within Claude Code.

Demo Site Architecture

Server-Side Rendered (SSR) - ports 8001, 8003

The server runs Flask with Jinja2 templates. Every product and job is embedded directly in the HTML response. curl http://localhost:8001/products | grep "MacBook" will find the product.

Client-Side Rendered (CSR) - ports 8002, 8004

The server returns an empty HTML shell:

<div id="app"><div class="loading">Loading products...</div></div>

JavaScript in the browser calls /api/products?page=1 and builds the product cards from the JSON response. curl http://localhost:8002/products | grep "MacBook" finds nothing. Playwright is required.

The key demonstration: the same CSS selectors work on both SSR and CSR versions after rendering. Only render_mode in the config differs.

Scraper Config Reference

Field retrieve types

Type Returns Use case
plaintext element.get_text() Most text content
attr element.get(attr) href, src, data-* attributes
regexp First capture group Extract numbers from text
regexpall List of all matches Multiple values in one element

Add "multiple": true to any field to collect all matching elements as a list. Works with any retrieve type:

"gallery_images": {"selector": "img.gallery-image", "retrieve": "attr", "attr": "src", "multiple": true}
"tags":           {"selector": "span.tag",           "retrieve": "plaintext",           "multiple": true}

Pagination stop_condition values

Value Stops when
no_results Link selector returns 0 elements (default)
last_page_text Specified text appears in page HTML

render_mode values

Value Behavior
static httpx fetch only, no JavaScript
playwright Always use headless browser
auto Try static, escalate to Playwright if probe selector fails

Book Structure

The book is in book/ as Markdown files, organized by part:

  • Part 1: Foundations - Why scraping, SSR vs CSR, the JSON config contract
  • Part 2: Extraction - CSS selectors, field types and transforms, pagination
  • Part 3: Rendering - Static fetching, Playwright, auto-fallback engine
  • Part 4: Scale - Scheduling, dual-database storage, change detection
  • Part 5: AI Agents - Config as language, autonomous agents, MCP integration

Requirements

  • Python 3.11+
  • Docker (for demo sites)
  • OpenRouter API key (for AI agent features)
  • Playwright browser (for CSR scraping): playwright install chromium

Citation

If you use this book or code in your work, please cite:

@misc{monteiro_2026_19513159,
  author    = {Monteiro, H{\'e}lder},
  title     = {Autonomous Web Scraping with {AI}: A Config-Driven Approach},
  month     = apr,
  year      = {2026},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.19513159},
  url       = {https://doi.org/10.5281/zenodo.19513159}
}

Or in plain text:

Hélder Monteiro. Autonomous Web Scraping with AI: A Config-Driven Approach. Zenodo, 2026. https://doi.org/10.5281/zenodo.19513159

License

Code: MIT. Book content: Creative Commons Attribution 4.0.

About

Repo for the Autonomous Web Scraping with AI book.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages