Scrape a single URL → LLM-ready markdown.
Use this when the user shares a URL and wants its content (read, analyze,
summarize, extract). Auto-escalates through fast→stealth→llm when blocked.
Args:
url: Target URL (http/https).
prefer: "auto" | "fast" | "stealth" | "llm".
auto = fast first, escalate to stealth on block/short page.
fast = cheap HTTP only (no JS).
stealth = real Chromium + Cloudflare solver.
llm = full Crawl4AI browser + BM25 fit-markdown.
timeout: per-attempt timeout in seconds.
include_html: include raw HTML in the response (large; off by default).
js: (stealth only) JS expression evaluated against the live page
after it settles. The value comes back in ``meta.js_result``.
Use for data that lives in DOM *properties* (e.g. an input's
``.value``) rather than in serialized HTML.
wait_for: (stealth only) JS predicate expression polled until truthy
(bounded by ``timeout``). Use to wait for content that arrives
asynchronously after ``network_idle``.
Returns:
{url, final_url, status, markdown, title, method, elapsed_ms, meta}
or {error, url, method} on failure.
Scrape + structured extraction using a CSS-based JSON schema.
Use when the user wants structured data (tables, lists, product info)
extracted from a page. Define a CSS schema to target specific elements.
The schema is a JsonCssExtractionStrategy schema:
{ "name": "PageItems", "baseSelector": "div.item",
"fields": [{"name": "title", "selector": "h2", "type": "text"}, ...] }
Returns parsed JSON in `data`.
Enumerate all internal URLs reachable from `root`.
Use when the user wants to map a site's structure or find all pages
before crawling. Often paired with `crawl` or `batch_scrape`.
Args:
root: Website root (e.g. "https://example.com/docs").
include_pattern: Optional regex; only URLs matching are returned.
limit: Hard cap on returned URLs.
Multi-page crawl: discover URLs on `root`, then scrape each.
Use this when the user wants to crawl an entire site section or docs,
or needs multiple pages scraped in bulk. For single pages use `scrape`;
for research questions use `deep_research`.
Args:
root: start URL.
max_pages: hard cap on pages scraped.
css_selector: scope each page's html/markdown to the matched element
(non-llm: lxml re-scope of the fetched HTML; llm: native crawl4ai
css_selector).
prefer: "auto" | "fast" | "stealth" | "llm" (llm = Crawl4AI BFS deep-crawl).
include_paths: regex — keep only URLs matching (matched against full URL).
exclude_paths: regex — drop URLs matching (e.g. `/tag/|/page/\d+`).
max_depth: 0 = flat harvest from the root page's links (default);
>0 = true BFS up to that link depth, honoring the filters.
Returns:
{root, pages: [{url, markdown, title, ...}], count, discovered, elapsed_ms}
or {error, root} on failure.
Extract text from a PDF/DOCX/PPTX URL → markdown (no browser).
Use when the user shares a link to a document (PDF, Word, PowerPoint)
and wants its text content. Also useful after `search_papers` to get
full text from a paper's pdf_url.
Content-type sniffed and routed to pypdf / python-docx / python-pptx.
Optional deps — install with `pip install 'pyrecrawl[docs]'`.
Web search via DuckDuckGo HTML (no API key required).
Use this for targeted searches where you need anti-bot bypass (Cloudflare
protection on DDG). For simple searches, the built-in web_search may suffice.
For research questions, prefer `deep_research` (search + scrape + citations).
Returns [{url, title, snippet}, ...]. The smart ladder bypasses
DDG's bot detection if needed.
Returns:
{query, results: [{url, title, snippet}], count}
or {error, query} on failure.
Scrape MANY URLs in ONE call (parallel, deduped, cache-aware).
Use when the user provides multiple URLs or you have a list of pages
to fetch. More efficient than calling `scrape` N times.
Args:
urls: Target URLs (deduped automatically; empties dropped).
prefer: "auto" | "fast" | "stealth" | "llm".
timeout: per-URL timeout in seconds.
max_concurrency: parallel workers (default 4).
include_html: include raw HTML per result (large; off by default).
Returns {requested, unique, succeeded, failed, results[]}.
Per-URL failures are isolated — other URLs still succeed.
Returns:
{requested, unique, succeeded, failed, results: [{url, markdown, ...}]}
Search the web, then pull the top sources as EVIDENCE (no LLM synthesis).
PRIMARY RESEARCH TOOL — use when the user asks to research, investigate,
deep-dive, fact-check, or learn about a topic. Returns a ``citations``
list with stable [n] numbers and an ``evidence`` list of per-source
markdown — the agent does the synthesis from evidence.
Multi-pass mode: set iterations=2-3 to auto-run additional searches with
refined queries (alternatives, criticism, latest developments) and append
deduplicated evidence. Each pass adds up to ``scrape_top`` new sources.
Args:
query: search string.
limit: how many search results to fetch.
scrape_top: how many of those to actually fetch content from.
prefer: "auto" | "fast" | "stealth" | "llm".
iterations: 1 (default, single pass), 2-3 (multi-pass with refined
queries targeting evidence gaps). Each pass searches from a
different angle and deduplicates by URL.
Returns:
{query, iterations_run, queries: [str], hits: [{url, title, snippet}],
citations: [{url, title}], evidence: [{url, title, markdown}],
scraped, used_engines, elapsed_ms}
or {error, query, hint} on failure.
Track a URL over time and report meaningful content changes.
Use when the user wants to watch a page for updates (price changes,
new blog posts, status updates). Ask me to 'set up monitoring for <url>'
for a guided setup playbook.
Args:
url: target URL.
action: "check" | "history" | "forget".
prefer: ladder preference, same as ``scrape``.
css_selector: scope the diff to one element (so banner /
nav changes don't trigger false positives).
Snapshots persist under ``PYRECRAWL_MONITOR_DIR`` (default
``~/.pyrecrawl/monitors/``). ``check`` returns ``status`` of
``new`` | ``unchanged`` | ``changed`` | ``error`` and a unified
diff when the page changed.
Returns:
{url, status: "new"|"unchanged"|"changed"|"error",
diff?: str, snapshot_chars?: int, elapsed_ms?: int}
Search academic papers via arXiv or Crossref — no API keys.
Args:
query: free-text search, e.g. "transformer attention scaling laws".
limit: max results (1-25 arXiv / 1-20 crossref).
source: "arxiv" (CS/physics/math preprints, default) or
"crossref" (all fields, DOI-backed).
category: optional arXiv category filter, e.g. "cs.LG", "cs.CV".
Returns papers with id/url/pdf_url/title/authors/summary/published.
Feed pdf_url into the `document` tool to extract full text.
Returns:
{query, source, papers: [{title, authors, abstract, url, pdf_url?, ...}]}
or {error, query} on failure.
Inspect the response cache: stats, clear, enable, or disable.
Use this to check cache hit rates before large batch jobs, or to
clear stale cached responses when a site's content has changed.
Args:
action: "stats" (default) | "clear" | "disable" | "enable".
Verify PyreCrawl is working: engine versions, dependencies, update status.
Use this at the start of a session or before a large scraping job to
confirm all engines are installed and up to date.
+1 more tools listed on main page