# scrapewright [![PyPI](https://img.shields.io/pypi/v/scrapewright)](https://pypi.org/project/scrapewright/) [![Python](https://img.shields.io/pypi/pyversions/scrapewright)](https://pypi.org/project/scrapewright/) [![License: MIT](https://img.shields.io/badge/license-MIT-green.svg)](LICENSE) **Give it a URL. It writes the scraper.** Most e-commerce catalog scraping splits into two worlds: sites on a known platform (Shopify, WooCommerce) that expose a clean JSON feed, and everything else — bespoke HTML where you hand-write a parser per site and re-write it every time the markup shifts. scrapewright collapses both into one call: 1. **Detect** the platform behind a URL. 2. For known platforms, **extract deterministically** from their public catalog API — free, stable, no LLM. 3. For custom HTML, **synthesize a reusable extractor once** with an LLM, cache it, and **replay it deterministically forever after**. The LLM is a *compiler*, not a runtime. It runs **once per site** to produce a recipe of CSS selectors; every page after that is parsed by plain BeautifulSoup at zero marginal cost. That is the whole cost-control story — no per-page model calls, no token bill that scales with your crawl. ``` ┌─────────────┐ store URL ───▶ │ detect │ └──────┬──────┘ ┌──────────────────┼──────────────────┐ ▼ ▼ ▼ shopify woocommerce generic HTML products.json wc/store/products (page mode) │ │ │ │ deterministic │ ▼ │ (free) │ cached recipe? ──yes──▶ replay (free) └────────┬─────────┘ │ no ▼ ▼ Product{} ◀───── selectors ── JSON-LD? ──yes──▶ Product{} (free) ▲ │ no │ ▼ └──────── replay ◀── LLM synthesizes recipe ONCE ──▶ cache ``` Everything normalizes to one `Product` shape, so downstream code never knows or cares which path a record came from. ## Install ```bash pip install scrapewright # deterministic paths (Shopify, Woo, JSON-LD) pip install "scrapewright[llm]" # + LLM recipe synthesis for custom HTML pip install "scrapewright[llm,js,excel,mcp]" # + JS rendering, XLSX, MCP server playwright install chromium # only needed for --js ``` ## Use it ```python from scrapewright import Scrapewright sw = Scrapewright() # Catalog mode — a whole Shopify/WooCommerce store, deterministically for product in sw.scrape_catalog("https://shop.example.com", max_items=200): print(product.brand, product.title, product.price, product.currency) # Page mode — one custom-HTML product page. # First call: tries JSON-LD (free); if absent, the LLM writes a recipe once. # Every later call on that domain: replayed from the cached recipe, no LLM. item = sw.scrape_page("https://boutique.example.com/products/wool-coat") print(item.model_dump(exclude={"raw"})) # Crawl mode — walk a WHOLE custom store from one listing/category URL. # The frontier discovers product pages (deterministic, free); the first page # pays the single synthesis cost, every other page replays the recipe. for product in sw.crawl("https://boutique.example.com/collection", max_items=100): print(product.title, product.price) ``` ### CLI ```bash scrapewright detect https://shop.example.com # platform + strategy scrapewright run https://shop.example.com --max 50 # scrape a catalog → JSONL scrapewright crawl https://boutique.example.com/collection -o products.xlsx scrapewright run https://shop.example.com -o products.csv # Excel-ready CSV scrapewright add https://boutique.example.com/products/coat # learn a site scrapewright run https://boutique.example.com/products/coat --no-llm scrapewright list # cached recipe domains ``` `-o` writes `.csv` (Excel-ready, UTF-8 BOM), `.xlsx` (`pip install scrapewright[excel]`), or `.jsonl`; without it, products stream to stdout as JSONL. ### Know what you are dealing with `detect` answers the routing question before a job starts: ``` $ scrapewright detect https://some-store.com https://some-store.com platform: bigcommerce catalog: - strategy: crawl note: BigCommerce (Stencil) markup ``` Twelve platforms are recognized: **Shopify** and **WooCommerce** publish a free JSON catalog, so those route to `catalog` — deterministic, no LLM, no browser. **Magento, BigCommerce, Salesforce Commerce Cloud, Squarespace, Wix, Webflow, PrestaShop, Shopware, Ecwid** and **OpenCart** are recognized by fingerprint and route to `crawl`, where the recipe path handles them like any custom site — the point of naming them is knowing what you face, not writing twelve parsers. Wix and Ecwid render client-side, so detection says `crawl+js` up front. A site behind an anti-bot wall reports `strategy: blocked` with the HTTP status, rather than pretending it found nothing. ### Bring your own schema Products are just the built-in default. Declare the fields you want and the same compile-once/replay-free loop works on any structured page — job posts, listings, registry records: ```bash scrapewright run https://jobs.example.com/p/123 -f title -f company -f salary:number -f tags:list --schema-name job ``` ```python from scrapewright import Scrapewright, Schema job = Schema.from_names(["title", "company", "salary:number", "tags:list"], name="job") record = Scrapewright().extract("https://jobs.example.com/p/123", job) print(record.data) # {'title': ..., 'company': ..., 'salary': ..., 'tags': [...]} ``` Field kinds are `text` (default), `number`, `url`, and `list`. Recipes are cached per site *and* per schema, so one domain can be compiled against several field sets without them overwriting each other. ### Use it from an AI agent (MCP) scrapewright ships an [MCP](https://modelcontextprotocol.io) server, so an agent can call it as a tool instead of reading raw HTML itself: ```bash pip install "scrapewright[mcp,llm]" scrapewright mcp ``` Point any MCP client at that command and the agent gains five tools: `detect_site`, `scrape_catalog`, `extract_page`, `crawl_site`, and `list_learned_sites`. Drop this into your client's config — Claude Desktop, Cursor, or anything else that speaks MCP: ```json { "mcpServers": { "scrapewright": { "command": "uvx", "args": ["--from", "scrapewright[mcp,llm]", "scrapewright", "mcp"], "env": { "ANTHROPIC_API_KEY": "sk-ant-..." } } } } ``` The key is only needed for sites on no known platform, where a recipe has to be written once. Shopify and WooCommerce stores work without it. The economics are the point. An agent that reads pages itself pays model tokens per page, forever. These tools pay **once per site** — an agent crawling 500 pages spends one synthesis, not five hundred, and platform stores (Shopify, WooCommerce) cost nothing at all. ### Run it as a service The same core behind an HTTP API, with keys, quotas, metering and background jobs: ```bash pip install "scrapewright[service,llm]" scrapewright keys create --label alice --plan free scrapewright serve --port 8000 ``` ```bash curl -X POST localhost:8000/v1/extract -H "X-API-Key: sw_..." -H "Content-Type: application/json" -d '{"url": "https://shop.example.com/products/coat"}' ``` | Endpoint | Purpose | |---|---| | `POST /v1/detect` | platform + strategy (cheap) | | `POST /v1/extract` | one page -> structured record | | `POST /v1/crawl` | a whole site -> job id (crawls outlive a request) | | `GET /v1/jobs/{id}` | poll a crawl | | `GET /v1/usage` | what this key has consumed, against its plan | #### Prepaid credits, no subscription One action costs real money: **compiling a new site**, a single LLM pass over a page, measured at $0.02 on a small product page and $0.15 on a heavy rendered one. Everything after that is BeautifulSoup — the ten-thousandth record from a compiled site is free to serve. So credits are priced off that one action, and everything else is denominated relative to it: | Action | Credits | |---|---| | 1 record delivered | 1 | | 1 browser render | 5 | | 1 new site compiled | 300 | | page fetches, `detect` | free | ``` $ scrapewright plans pack credits price $/credit margin starter 10,000 $10 0.00100 80.0% growth 50,000 $40 0.00080 75.0% scale 250,000 $150 0.00060 66.7% Free: 1,000 credits a month, resetting. ``` Margin is measured on compiling a site, because that is the only step that costs anything; a test fails if a price edit drops any pack below 60%. A free account can cost us at most $0.20 a month, even if every free credit goes to the most expensive action there is. Credits are a **ledger, not a counter** — every grant and every charge is a row, so a disputed bill can be reconstructed line by line, and a replayed payment webhook cannot double-credit (grants take an idempotency key). Running out returns `402` with the balance and what to do about it; a crawl is capped by the credits on hand, so a job stops at what the caller can pay for instead of overdrawing. ```bash scrapewright credits grant --pack starter --idempotency scrapewright credits balance ``` #### Taking payment Stripe is wired in and turned on by environment, not by a code change: ```bash pip install "scrapewright[service,stripe]" export STRIPE_SECRET_KEY=sk_test_... # absent -> nothing is for sale export STRIPE_WEBHOOK_SECRET=whsec_... # absent -> webhooks are refused scrapewright serve ``` | Endpoint | Purpose | |---|---| | `GET /v1/credits/packs` | the price list — public, no key needed | | `POST /v1/credits/checkout` | start a purchase, returns a Stripe Checkout URL | | `POST /v1/webhooks/stripe` | payment notifications from Stripe | The webhook endpoint takes **no API key** — Stripe is the caller, so the signature *is* the credential, and an unverified endpoint would be a free credit printer for anyone who guessed the URL. Three rules hold the integration up: * **Verify every signature.** No signing secret configured means webhooks are refused outright, rather than accepted unverified. * **Never trust an amount off the wire.** The event names a pack; how many credits that pack is worth is looked up from our own price list, so a tampered payload buys exactly what it paid for or nothing at all. * **Grant idempotently, keyed on the Checkout session.** Stripe retries deliveries, and one payment can produce several event types — the session id is what identifies the money that actually moved. `examples/stripe_smoke_test.py` runs the whole path against Stripe's test mode with the 4242 card. Any other provider plugs into the same two-method `BillingProvider` protocol in `scrapewright.service.billing`; without one, the service simply runs free, which is the right default for a demo or a self-hosted instance. Docker: ```bash docker build -t scrapewright . # static paths docker build -t scrapewright --build-arg WITH_JS=1 . # + headless Chromium docker run -p 8000:8000 -v sw-data:/data scrapewright ``` ### Client-side-rendered stores Add `--js` (or `Scrapewright(js=True)`) and pages that render their catalog in the browser become extractable: ```bash scrapewright run https://spa-store.example.com/products/x --page --js scrapewright crawl https://spa-store.example.com/shop --js -o products.xlsx ``` Rendering stays **rare by construction**: the static fetch runs first, and Chromium is only started when the static HTML is an empty client-side shell or extraction on it fails. A recipe learned from rendered HTML is tagged `needs_js`, so later runs on that site skip the wasted static hop. The browser starts at most once per run and is reused for every page. ## The `Product` shape ```python url: str # canonical product URL title: str brand: str | None price: Decimal | None # parsed from "1,250.00" / "1.250,00" / "€1290" alike currency: str | None available: bool | None images: list[str] # absolute URLs sizes: list[str] description: str | None sku: str | None source_platform: str # shopify | woocommerce | json-ld | selector ``` A record is **usable** when it carries a title, a price, and a URL. The validator (`scrapewright.coverage`) reports the usable ratio across a batch — the number a recipe is trusted on before it's cached. ## How the pieces fit | Module | Role | |---|---| | `detect` | Platform registry: free-catalog probes, then fingerprints for 12 platforms; returns the strategy to use | | `extract/shopify`, `extract/woocommerce` | Deterministic catalog extractors | | `extract/jsonld` | schema.org/Product from `