# ๐Ÿ”ฅ PyreCrawl โ€” Web Browsing Superpowers for Your AI Agent [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT) [![MCP](https://img.shields.io/badge/MCP-1.0-blue.svg)](https://modelcontextprotocol.io/) [![Python 3.10+](https://img.shields.io/badge/python-3.10+-blue.svg)](https://www.python.org/) [![PyPI](https://img.shields.io/pypi/v/pyrecrawl.svg)](https://pypi.org/project/pyrecrawl/) [![GitHub stars](https://img.shields.io/github/stars/SanggonBoy/PyreCrawl?logo=github)](https://github.com/SanggonBoy/PyreCrawl/stargazers) [![Downloads / 30d](https://img.shields.io/endpoint?url=https%3A%2F%2Fpyrecrawl-stats.fajarnugraha90543.workers.dev%2Fbadge%2Fdownloads)](https://pypistats.org/packages/pyrecrawl) [![Active users / 30d](https://img.shields.io/endpoint?url=https%3A%2F%2Fpyrecrawl-stats.fajarnugraha90543.workers.dev%2Fbadge%2Fusers)](https://github.com/SanggonBoy/PyreCrawl#-privacy--anonymous-usage-ping) **One command gives any AI agent the whole web.** Scrape, extract, crawl, map, and search โ€” self-hosted, no API keys, no rate limits, no subscription. PyreCrawl speaks **MCP** (Model Context Protocol), the standard tool interface for Claude, Cursor, VS Code, Codex, OpenCode, Hermes, and any MCP-compatible agent. A **smart auto-fallback ladder** always picks the cheapest method that succeeds: ``` fast HTTP โ”‚ (403/503/Cloudflare challenge or empty body) โ–ผ stealth browser (real Chromium + Cloudflare solver) โ”‚ (still blocked, or the page needs full JS rendering) โ–ผ deep processing (LLM-ready markdown, citations, structured extraction) ``` ## โšก Tools exposed | Tool | What it does | |---|---| | `scrape(url, prefer="auto")` | Single URL โ†’ LLM-ready markdown | | `extract(url, schema)` | Scrape + structured extraction (JsonCss schema) | | `map_site(root, include_pattern=None, limit=200)` | Enumerate all internal URLs | | `crawl(root, max_pages=5, prefer="auto", include_paths=None, exclude_paths=None, max_depth=0)` | Multi-page crawl with path filters + true BFS depth | | `document(url)` | PDF/DOCX/PPTX โ†’ markdown (no browser, optional `[docs]` extras) | | `search(query, limit=10)` | Web search via DuckDuckGo HTML (no API key) | | `search_papers(query, limit=8, source="arxiv", category=None)` | Academic search via arXiv + Crossref (no API key) โ€” feed `pdf_url` into `document` | | `batch_scrape(urls[], ...)` | Many URLs in ONE call โ€” parallel, deduped, cache-aware | | `deep_research(query, limit=5, scrape_top=3)` | Search โ†’ evidence pack with [n] citations (no LLM synthesis โ€” your agent does that) | | `monitor(url, action, css_selector=None)` | Change detection with persisted snapshots + unified diff | | `session(session, action, ...)` | Persistent browser session (cookies kept) โ€” login walls, multi-step flows, screenshots | | `cache(action)` | Inspect/clear/enable/disable the HTTP response cache | | `health()` | Versions + import sanity check | **MCP Resources** (read-only state without a tool call): `pyrecrawl://cache/stats` ยท `pyrecrawl://sessions` ยท `pyrecrawl://monitors` **MCP Prompts** (ready-made playbooks): `research(topic)` ยท `rag_ingest(site)` ยท `watch_page(url)` ### Env flags | Variable | Default | Effect | |---|---|---| | `PYRECRAWL_CACHE` | off | `1` = in-memory LRU (128 pages), or a directory path (reserved for disk mode) | | `PYRECRAWL_CACHE_TTL` | `900` | Cache entry lifetime in seconds | | `PYRECRAWL_MONITOR_DIR` | `~/.pyrecrawl/monitors` | Where monitor snapshots persist | | `PYRECRAWL_NO_TELEMETRY` | off | `1` = disable the anonymous startup ping (also honors `DO_NOT_TRACK=1`) | `prefer` options: `"auto"` (default ladder) ยท `"fast"` (HTTP only) ยท `"stealth"` (CF bypass) ยท `"llm"` (deep processing). --- ## ๐Ÿš€ Install & Use (one-liner) ### 1. Install #### [UV](https://docs.astral.sh/uv/) (recommended โ€” one command, zero Python setup) UV is a fast Python package manager that handles Python itself โ€” no need to install Python separately. Get it once: ```bash # macOS / Linux curl -LsSf https://astral.sh/uv/install.sh | sh # Windows (PowerShell) powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex" ``` [Learn more about UV โ†’](https://docs.astral.sh/uv/) Then run PyreCrawl directly โ€” no venv, no `pip install`, no Python download: ```bash uvx pyrecrawl@latest ``` #### Or via uv tool install (persistent, recommended for regular use) ```bash uv tool install pyrecrawl ``` #### Or via pipx (alternative) ```bash pipx install pyrecrawl ``` #### Or via pip into a venv ```bash pip install pyrecrawl ``` ### 2. One-time browser engines ```bash pyrecrawl setup ``` This installs Chromium + stealth browser engines (~2 min, one-time). ### 3. Register with your AI agent ```bash # Auto-detect installed agents and write their MCP configs pyrecrawl install # Or target specific agents pyrecrawl install claude-desktop cursor # Dry-run to preview what would change pyrecrawl install --dry-run ``` Supported agents: `claude-desktop`, `claude-code`, `cursor`, `vscode`, `codex`, `opencode`, `hermes`. ### 4. Start chatting After installing + registering, **restart your agent** (or start a new session). Then ask: > *"Scrape https://example.com and summarize it."* Available tools: | Tool | What it does | |------|-------------| | `scrape` | Fetch a single URL โ†’ markdown (auto-escalates past Cloudflare) | | `extract` | Scrape + structured extraction via CSS schema โ†’ JSON | | `map_site` | Enumerate all internal URLs from a root | | `crawl` | Multi-page crawl: discover + scrape in bulk | | `batch_scrape` | Fetch many URLs in one parallel call | | `search` | Web search via DuckDuckGo with anti-bot bypass | | `search_papers` | Academic paper search (arXiv / Crossref) | | `deep_research` | Search + scrape + citations in one call โ€” **primary research tool** | | `document` | Extract text from PDF/DOCX/PPTX URLs | | `monitor` | Track a URL for content changes over time | | `session` | Persistent browser session for login walls | | `cache` | Inspect or clear the response cache | | `health` | Verify engine availability + version | Plus 3 guided prompts: `research`, `rag_ingest`, `watch_page`. ### Quick examples Ask your agent naturally โ€” no special syntax needed: | You say | Agent uses | |---------|-----------| | *"Scrape https://example.com and summarize it"* | `scrape` โ†’ returns markdown โ†’ agent summarizes | | *"Research Rust memory safety vulnerabilities"* | `deep_research` โ†’ search + scrape + citations | | *"Deep research on AI regulation worldwide"* | `deep_research(iterations=3)` โ†’ multi-pass with refined queries | | *"Extract all product names and prices from this page"* | `extract` โ†’ CSS schema โ†’ structured JSON | | *"Crawl https://docs.example.com and give me an overview"* | `crawl` โ†’ multi-page โ†’ summary | | *"Monitor this page for price changes"* | `monitor` โ†’ baseline snapshot โ†’ periodic diff | | *"Find papers about transformer attention"* | `search_papers` โ†’ arXiv results | | *"What's the current cache hit rate?"* | `cache` โ†’ stats | --- ## ๐Ÿง  Skills โ€” Maximize Your Agent's Research Quality PyreCrawl tools give your agent **hands** (scrape, crawl, search). But the agent still needs a **brain** โ€” instructions on *when* to use which tool, *how* to chain research passes, and *what* anti-hallucination rules to follow. That's what **[PyreCrawl Skills](https://github.com/SanggonBoy/pyrecrawl-skills)** provides. | | MCP Tools (this repo) | Skills ([pyrecrawl-skills](https://github.com/SanggonBoy/pyrecrawl-skills)) | |---|---|---| | **Role** | Execute web operations | Tell the agent how to use them | | **Analogy** | Hands | Brain | | **Example** | `deep_research(query, iterations=3)` | "Run 3 passes, check gaps after each, cite everything" | | **Required?** | Yes (the engine) | Optional (but recommended for research quality) | **Quick setup:** ```bash # 1. Install the tools (you already have this) uvx pyrecrawl@latest # 2. Add the research skill to your project git clone https://github.com/SanggonBoy/pyrecrawl-skills.git /tmp/pyrecrawl-skills cp /tmp/pyrecrawl-skills/pyrecrawl-research/SKILL.md ./CLAUDE.md # or .cursorrules / AGENTS.md ``` > **Without skills:** Your agent has powerful tools but improvises usage. > **With skills:** Your agent follows a proven research protocol with anti-hallucination guardrails. --- ## ๐Ÿ“š Manual config (if `pyrecrawl install` doesn't match your setup) ### Claude Desktop **Config file** - Linux: `~/.config/Claude/claude_desktop_config.json` - macOS: `~/Library/Application Support/Claude/claude_desktop_config.json` - Windows: `%AppData%\Claude\claude_desktop_config.json` ```json { "mcpServers": { "pyrecrawl": { "command": "uvx", "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"] } } } ``` ### Claude Code **Config file**: project-scoped `.mcp.json` ```json { "mcpServers": { "pyrecrawl": { "command": "uvx", "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"] } } } ``` ### Cursor **Config file**: `~/.cursor/mcp.json` ```json { "mcpServers": { "pyrecrawl": { "command": "uvx", "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"] } } } ``` ### VS Code / Copilot **Config file**: `.vscode/mcp.json` (project-scoped) ```json { "servers": { "pyrecrawl": { "command": "uvx", "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"], "type": "stdio" } } } ``` ### Codex CLI **Config file**: `~/.codex/config.toml` ```toml [mcp_servers.pyrecrawl] command = "uvx" args = ["--from", "pyrecrawl", "pyrecrawl", "serve"] ``` ### OpenCode **Config file**: `~/.config/opencode/opencode.json` ```json { "mcp": { "pyrecrawl": { "type": "local", "command": ["uvx", "--from", "pyrecrawl", "pyrecrawl", "serve"], "enabled": true } } } ``` ### Hermes **Config file** - Linux/macOS: `~/.hermes/config.yaml` - Windows: `%LocalAppData%\hermes\config.yaml` ```yaml mcp_servers: pyrecrawl: command: uvx args: - --from - pyrecrawl - pyrecrawl - serve enabled: true ``` > **Windows note:** `uvx` must be on PATH. If not, use the full path to `uvx.exe` (e.g. `C:\Users\\AppData\Local\hermes\bin\uvx.exe`). --- ## ๐Ÿง  How the ladder chooses PyreCrawl runs each request through three tiers, stopping at the first one that returns a complete, LLM-ready result: | Concern | Fast tier | Stealth tier | Deep tier | |---|---|---|---| | Static HTML page | โœ… ~200ms | โ€” | โ€” | | Cloudflare-protected | โŒ | โœ… Turnstile solver | โ€” | | JS-heavy SPA | โŒ | โœ… real Chromium | โ€” | | Live DOM data (input `.value`, JS state) | โŒ | โœ… `js` param | โ€” | | LLM-ready markdown + citations | โ€” | โ€” | โœ… BM25, fit-markdown | | Structured extraction (CSS schema) | โ€” | โ€” | โœ… | | Deep crawl (BFS/DFS/BestFirst) | โ€” | โ€” | โœ… adaptive | The agent never has to pick. `prefer="auto"` does it every call. ### Live DOM data with `js` and `wait_for` Some sites keep the data you want in a DOM *property* (e.g. an ``'s `.value`) that JS writes after an XHR โ€” it never appears in the serialized HTML. The `scrape` tool accepts two stealth-tier params for exactly this: ```json { "url": "https://temp-mail.org/id", "prefer": "stealth", "wait_for": "document.getElementById('mail').value.includes('@')", "js": "document.getElementById('mail').value" } ``` - `wait_for` โ€” a JS **predicate expression** polled until truthy (bounded by `timeout`). Use it instead of guessing a sleep for anything that arrives asynchronously. - `js` โ€” a JS **expression** evaluated once the page settles; the value comes back in `meta.js_result`. Errors are captured in `meta.js_error` (the page result is still returned, never a crash). --- ## ๐Ÿ“Š Compared to Firecrawl (hosted) | | Firecrawl | PyreCrawl | |---|---|---| | Cost | Free 1k/mo, then $16โ€“333/mo | **Free, self-hosted** | | Local LLM support | โŒ | โœ… Ollama / any LLM | | Cloudflare bypass | โœ… (Fire-Engine, paid) | โœ… (free, built-in) | | Markdown + BM25 | โœ… | โœ… | | Self-host | โŒ | โœ… | | Academic paper search | โŒ | โœ… arXiv + Crossref (`search_papers`) | | Hosted search API | โœ… /search | โš ๏ธ DuckDuckGo HTML + arXiv/Crossref (no key) | --- ## ๐Ÿ”ง Development ```bash git clone https://github.com/SanggonBoy/PyreCrawl.git cd PyreCrawl uv venv --python 3.12 .venv source .venv/Scripts/activate # Windows; or .venv/bin/activate on macOS/Linux uv pip install -e ".[dev]" python -m playwright install chromium scrapling install ``` ### Run tests ```bash python scripts/selfcheck.py # real-network smoke test (13 tools + engines) python scripts/probe_stdio.py # stdio JSON-RPC probe python scripts/test_ladder_bug.py # SPA-shell ladder escalation regression python scripts/test_js_eval.py # stealth js/wait_for params regression python scripts/test_scope_selector.py # crawl css_selector/max_depth wiring python scripts/test_link_harvest.py # map/BFS link purity regression ``` --- ## ๐Ÿ“ฆ Publish Maintainers only: ```bash git tag vX.Y.Z git push origin vX.Y.Z ``` GitHub Actions builds + uploads to PyPI via [trusted publishing](https://docs.pypi.org/trusted-publishers/). --- ## ๐Ÿ”” Stay up to date PyreCrawl checks PyPI on every startup and reports the latest version โ€” your MCP agent sees this automatically via the `health()` tool response and can notify you inline. To check manually: ```bash pyrecrawl version ``` To upgrade: ```bash pyrecrawl update # runs: uv tool upgrade pyrecrawl ``` **Get notified of new releases:** click **Watch** โ†’ **Releases only** at the [GitHub repo](https://github.com/SanggonBoy/PyreCrawl) to receive email notifications when a new version is published. --- > [!NOTE] > PyreCrawl sends **one anonymous usage ping per 24 h** at server startup โ€” see > [Privacy](#-privacy--anonymous-usage-ping) for exactly what's sent and how to opt out. ## ๐Ÿ”’ Privacy โ€” anonymous usage ping PyreCrawl phones home **once per 24 h** with a tiny anonymous ping when the MCP server starts, so we can count real users (DAU/MAU) instead of raw downloads. | Sent (4 fields, ~100 bytes) | Never sent | |---|---| | Hashed machine id (SHA-256 of hostname+MAC โ€” not reversible) | Your IP (not stored) | | PyreCrawl version | Any URL you scrape | | Python version | Any page content or search queries | | OS family (`windows` / `linux` / `darwin`) | Anything else | Client code: [`src/pyrecrawl/telemetry.py`](src/pyrecrawl/telemetry.py) (~90 lines, stdlib only) ยท Collector: [`workers/telemetry/`](workers/telemetry/) โ€” a self-hostable Cloudflare Worker + D1, no third-party analytics service. Opt out any time: ```bash export PYRECRAWL_NO_TELEMETRY=1 # or the industry-standard DO_NOT_TRACK=1 ``` --- ## ๐Ÿ“œ Uninstall ```bash # Remove from all agent configs pyrecrawl uninstall # Remove the package uv tool uninstall pyrecrawl ``` --- ## ๐Ÿ›ก๏ธ License MIT โ€” see [LICENSE](LICENSE).