The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Scrapewright listing page.
Give it a URL. It writes the scraper.
Most e-commerce catalog scraping splits into two worlds: sites on a known platform (Shopify, WooCommerce) that expose a clean JSON feed, and everything else — bespoke HTML where you hand-write a parser per site and re-write it every time the markup shifts. scrapewright collapses both into one call:
The LLM is a compiler, not a runtime. It runs once per site to produce a recipe of CSS selectors; every page after that is parsed by plain BeautifulSoup at zero marginal cost. That is the whole cost-control story — no per-page model calls, no token bill that scales with your crawl.
Everything normalizes to one Product shape, so downstream code never knows or
cares which path a record came from.
-o writes .csv (Excel-ready, UTF-8 BOM), .xlsx (pip install scrapewright[excel]),
or .jsonl; without it, products stream to stdout as JSONL.
detect answers the routing question before a job starts:
Twelve platforms are recognized: Shopify and WooCommerce publish a free
JSON catalog, so those route to catalog — deterministic, no LLM, no browser.
Magento, BigCommerce, Salesforce Commerce Cloud, Squarespace, Wix, Webflow,
PrestaShop, Shopware, Ecwid and OpenCart are recognized by fingerprint and
route to crawl, where the recipe path handles them like any custom site — the
point of naming them is knowing what you face, not writing twelve parsers.
Wix and Ecwid render client-side, so detection says crawl+js up front.
A site behind an anti-bot wall reports strategy: blocked with the HTTP status,
rather than pretending it found nothing.
Products are just the built-in default. Declare the fields you want and the same compile-once/replay-free loop works on any structured page — job posts, listings, registry records:
Field kinds are text (default), number, url, and list. Recipes are cached
per site and per schema, so one domain can be compiled against several field
sets without them overwriting each other.
scrapewright ships an MCP server, so an agent can call it as a tool instead of reading raw HTML itself:
Point any MCP client at that command and the agent gains five tools: detect_site,
scrape_catalog, extract_page, crawl_site, and list_learned_sites.
Drop this into your client's config — Claude Desktop, Cursor, or anything else that speaks MCP:
The key is only needed for sites on no known platform, where a recipe has to be written once. Shopify and WooCommerce stores work without it.
The economics are the point. An agent that reads pages itself pays model tokens per page, forever. These tools pay once per site — an agent crawling 500 pages spends one synthesis, not five hundred, and platform stores (Shopify, WooCommerce) cost nothing at all.
The same core behind an HTTP API, with keys, quotas, metering and background jobs:
| Endpoint | Purpose |
|---|---|
POST /v1/detect | platform + strategy (cheap) |
POST /v1/extract | one page -> structured record |
POST /v1/crawl | a whole site -> job id (crawls outlive a request) |
GET /v1/jobs/{id} | poll a crawl |
GET /v1/usage | what this key has consumed, against its plan |
One action costs real money: compiling a new site, a single LLM pass over a page, measured at $0.02 on a small product page and $0.15 on a heavy rendered one. Everything after that is BeautifulSoup — the ten-thousandth record from a compiled site is free to serve. So credits are priced off that one action, and everything else is denominated relative to it:
| Action | Credits |
|---|---|
| 1 record delivered | 1 |
| 1 browser render | 5 |
| 1 new site compiled | 300 |
page fetches, detect | free |
Margin is measured on compiling a site, because that is the only step that costs anything; a test fails if a price edit drops any pack below 60%. A free account can cost us at most $0.20 a month, even if every free credit goes to the most expensive action there is.
Credits are a ledger, not a counter — every grant and every charge is a row,
so a disputed bill can be reconstructed line by line, and a replayed payment
webhook cannot double-credit (grants take an idempotency key). Running out
returns 402 with the balance and what to do about it; a crawl is capped by the
credits on hand, so a job stops at what the caller can pay for instead of
overdrawing.
Stripe is wired in and turned on by environment, not by a code change:
| Endpoint | Purpose |
|---|---|
GET /v1/credits/packs | the price list — public, no key needed |
POST /v1/credits/checkout | start a purchase, returns a Stripe Checkout URL |
POST /v1/webhooks/stripe | payment notifications from Stripe |
The webhook endpoint takes no API key — Stripe is the caller, so the signature is the credential, and an unverified endpoint would be a free credit printer for anyone who guessed the URL. Three rules hold the integration up:
examples/stripe_smoke_test.py runs the whole path against Stripe's test mode
with the 4242 card. Any other provider plugs into the same two-method
BillingProvider protocol in scrapewright.service.billing; without one, the
service simply runs free, which is the right default for a demo or a self-hosted
instance.
Docker:
Add --js (or Scrapewright(js=True)) and pages that render their catalog in the
browser become extractable:
Rendering stays rare by construction: the static fetch runs first, and Chromium is
only started when the static HTML is an empty client-side shell or extraction on it
fails. A recipe learned from rendered HTML is tagged needs_js, so later runs on that
site skip the wasted static hop. The browser starts at most once per run and is reused
for every page.
Product shapeA record is usable when it carries a title, a price, and a URL. The
validator (scrapewright.coverage) reports the usable ratio across a batch —
the number a recipe is trusted on before it's cached.
| Module | Role |
|---|---|
detect | Platform registry: free-catalog probes, then fingerprints for 12 platforms; returns the strategy to use |
extract/shopify, extract/woocommerce | Deterministic catalog extractors |
extract/jsonld | schema.org/Product from <script type="application/ld+json"> — free, ~common |
extract/llm | Synthesizes a SelectorRecipe from HTML — the one-time compile step |
extract/selectors | Replays a recipe with BeautifulSoup — the deterministic runtime |
schema | Schema/Field — declare what to extract; PRODUCT_SCHEMA is the built-in default |
service/ | FastAPI app: API keys (stored hashed), record-based quotas, cost metering, background crawl jobs, pluggable billing |
service/credits | Credit prices, packs, and the free allowance |
service/stripe_billing | Stripe Checkout + signature-verified webhook |
service/pricing | Measured unit costs and the margin each pack clears |
mcp_server | Five MCP tools so AI agents can call scrapewright directly |
fetch | StaticFetcher (plain HTTP) and BrowserFetcher (headless Chromium), plus the shell heuristic that decides when a render is worth paying for |
crawl | Frontier: turns one listing URL into product URLs (pattern match + card-template fallback + pagination) — deterministic, no LLM |
cache | Persists recipes keyed by domain, so the compile happens once |
validate | Field-coverage scoring |
export | Batch → .csv / .xlsx / .jsonl |
pipeline | Orchestrates detect → extract → validate → cache → heal |
max_synth_per_run (default 3) — a site that resists synthesis cannot burn
one model call per page. The bill is bounded no matter how large the crawl.model and works with any
injected client; the default targets Anthropic's Claude via the official SDK.The deterministic paths are fully covered by offline fixtures — no network, no model calls — so CI is green without an API key:
v0.9 (alpha). Implemented and tested: an HTTP service with API keys, prepaid credits (priced off the one action that costs money, on an auditable ledger) and Stripe checkout with a signature-verified webhook, cost metering and background jobs; platform detection across 12 storefronts with a recommended strategy per site, catalog extraction (Shopify, WooCommerce), page extraction (JSON-LD, LLM-synthesized selectors), recipe caching, self-healing re-synthesis with a bounded per-run model budget, a crawl frontier (one listing URL → the whole site), JS rendering via an optional Playwright fetcher with automatic escalation, schema-agnostic extraction (bring your own fields), an MCP server for AI agents, coverage validation, and CSV / XLSX / JSONL export. 141 offline tests.
Known limit, stated plainly: it does not defeat anti-bot walls — deliberately out of scope. Sites behind Akamai/Fastly-style challenges return an honest miss.
Roadmap: pagination strategies for infinite-scroll listings, and a deployed instance of the service.
MIT — see LICENSE.