Skip to content

Repository files navigation

silkworm-rs

PyPI - Version Tests Docs PyPI Downloads

Silkworm is an async-first Python web scraping framework built on wreq and scraper-rs. It combines a small, typed Spider/Request/Response API with middleware, output pipelines, and production crawl controls.

Documentation · Getting started · Examples · API reference

Highlights

  • Async crawling with concurrency, priorities, request deduplication, timeouts, and deadlock-free queue backpressure.
  • Browser-impersonating HTTP through wreq, with redirects, proxies, cookies, retries, throttling, and robots.txt support.
  • Push-style typed callbacks using await self.emit(...) and await response.follow(...).
  • Async CSS/XPath selection and optional declarative Item, Text, and Attr extraction.
  • File, database, cloud, queue, and message-stream pipelines, with batch processing available on every built-in pipeline.
  • Production controls for failure policies, stop limits, pause/resume, caching, graceful shutdown, metrics, and structured crawl statistics.
  • Optional CDP, Servo, and OnionLink clients for rendered or onion-service pages.

Install

Silkworm supports Python 3.13–3.15.

pip install silkworm-rs

With uv:

uv add silkworm-rs

Integrations are installed as optional extras. For example:

pip install "silkworm-rs[rsloop,polars]"

See Getting Started for the complete extras and Python-version compatibility table.

Quick start

from silkworm import HTMLResponse, Response, Spider, run_spider
from silkworm.pipelines import JsonLinesPipeline


class QuotesSpider(Spider):
    name = "quotes"
    start_urls = ("https://quotes.toscrape.com/",)

    async def parse(self, response: Response) -> None:
        if not isinstance(response, HTMLResponse):
            return

        for quote in await response.select(".quote"):
            text = await quote.select_first(".text")
            author = await quote.select_first(".author")
            if text is not None and author is not None:
                await self.emit({"text": text.text, "author": author.text})

        next_link = await response.select_first("li.next > a")
        if next_link is not None and (href := next_link.attr("href")):
            await response.follow(href, callback=self.parse)


run_spider(
    QuotesSpider,
    item_pipelines=[JsonLinesPipeline("data/quotes.jl")],
)

Callbacks return None; they report items and requests with emit and follow, which apply backpressure while the callback is running.

Command line

Run a spider module or test its parser without writing a runner script:

silkworm crawl examples/quotes_spider.py -o data/quotes.jl -s max_items=100
silkworm parse https://quotes.toscrape.com/ --spider examples/quotes_spider.py

See the CLI reference for configuration, output formats, and exit codes.

Documentation

Topic Guide
Installation, extras, and first spider Getting Started
Spider, request, response, selectors, and callbacks Core Concepts
Declarative item extraction Declarative Extraction
Engine, HTTP, CDP, Servo, and OnionLink Engine and HTTP Client
Request and response middleware Middlewares
Output destinations and batch processing Pipelines
asyncio, rsloop, uvloop, winloop, and Trio Runners
Resumable and observable crawls Production Crawling
Version changes Migration Guide
Constraints and workarounds Limitations

Development

uv venv --python python3.13
uv sync --group dev
just fmt && just lint && just typecheck && just test

See the development workflow and open an issue or pull request on GitHub.

Acknowledgements

Silkworm builds on wreq, scraper-rs, fast-h2m, turboxml, and optional integrations including OnionLink and Servo. Thank you to their maintainers and contributors.

License

MIT. See LICENSE.

About

Async web scraping framework on top of Rust. Works with Free-threaded Python (`PYTHON_GIL=0`).

Topics

Resources

Stars

76 stars

Watchers

0 watching

Forks

Releases

Used by

Contributors

Languages