Silkworm is an async-first Python web scraping framework built on wreq and scraper-rs. It combines a small, typed Spider/Request/Response API with middleware, output pipelines, and production crawl controls.
Documentation · Getting started · Examples · API reference
- Async crawling with concurrency, priorities, request deduplication, timeouts, and deadlock-free queue backpressure.
- Browser-impersonating HTTP through wreq, with redirects, proxies, cookies, retries, throttling, and robots.txt support.
- Push-style typed callbacks using
await self.emit(...)andawait response.follow(...). - Async CSS/XPath selection and optional declarative
Item,Text, andAttrextraction. - File, database, cloud, queue, and message-stream pipelines, with batch processing available on every built-in pipeline.
- Production controls for failure policies, stop limits, pause/resume, caching, graceful shutdown, metrics, and structured crawl statistics.
- Optional CDP, Servo, and OnionLink clients for rendered or onion-service pages.
Silkworm supports Python 3.13–3.15.
pip install silkworm-rsWith uv:
uv add silkworm-rsIntegrations are installed as optional extras. For example:
pip install "silkworm-rs[rsloop,polars]"See Getting Started for the complete extras and Python-version compatibility table.
from silkworm import HTMLResponse, Response, Spider, run_spider
from silkworm.pipelines import JsonLinesPipeline
class QuotesSpider(Spider):
name = "quotes"
start_urls = ("https://quotes.toscrape.com/",)
async def parse(self, response: Response) -> None:
if not isinstance(response, HTMLResponse):
return
for quote in await response.select(".quote"):
text = await quote.select_first(".text")
author = await quote.select_first(".author")
if text is not None and author is not None:
await self.emit({"text": text.text, "author": author.text})
next_link = await response.select_first("li.next > a")
if next_link is not None and (href := next_link.attr("href")):
await response.follow(href, callback=self.parse)
run_spider(
QuotesSpider,
item_pipelines=[JsonLinesPipeline("data/quotes.jl")],
)Callbacks return None; they report items and requests with emit and follow,
which apply backpressure while the callback is running.
Run a spider module or test its parser without writing a runner script:
silkworm crawl examples/quotes_spider.py -o data/quotes.jl -s max_items=100
silkworm parse https://quotes.toscrape.com/ --spider examples/quotes_spider.pySee the CLI reference for configuration, output formats, and exit codes.
| Topic | Guide |
|---|---|
| Installation, extras, and first spider | Getting Started |
| Spider, request, response, selectors, and callbacks | Core Concepts |
| Declarative item extraction | Declarative Extraction |
| Engine, HTTP, CDP, Servo, and OnionLink | Engine and HTTP Client |
| Request and response middleware | Middlewares |
| Output destinations and batch processing | Pipelines |
| asyncio, rsloop, uvloop, winloop, and Trio | Runners |
| Resumable and observable crawls | Production Crawling |
| Version changes | Migration Guide |
| Constraints and workarounds | Limitations |
uv venv --python python3.13
uv sync --group dev
just fmt && just lint && just typecheck && just testSee the development workflow and open an issue or pull request on GitHub.
Silkworm builds on wreq, scraper-rs, fast-h2m, turboxml, and optional integrations including OnionLink and Servo. Thank you to their maintainers and contributors.
MIT. See LICENSE.