The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Crawlbrulee MCP listing page.
EU-native web scraping for AI agents & developers.
plug crawlbrulee into your agent. the official mcp server for crawlbrulee gives mcp-aware agents — Claude Code, Codex, Cursor, Claude Desktop — native tools to scrape pages, map sites, run background jobs, and check usage. one call turns any url into clean markdown, screenshots, metadata and links.
get a free api key → dashboard.crawlbrulee.com
the server:
npx-runnable — zero install.@crawlbrulee/sdk under the hood; this mcp is just a thin protocol adapter.this readme covers the mcp server itself — its tools and how to wire it into a host. for how the api behaves — endpoints, parameters, and error semantics — please see our api docs.
the same pattern works for Codex, Claude Desktop, and any other host that
accepts a stdio mcp launch command — set command: npx, args: ["-y", "@crawlbrulee/mcp"], and forward CRAWLBRULEE_API_KEY via the env block.
| env var | required | description |
|---|---|---|
CRAWLBRULEE_API_KEY | yes | api key sent as Authorization: Bearer …. get one at https://crawlbrulee.com. |
the mcp reads the env var on first tool invocation — not at startup — so a typo in your config surfaces as a clear tool-error message rather than the server failing to come up. see authentication for how the api consumes keys.
scrapefetch a single url and return the requested content (markdown, cleaned html, raw html, links, images, screenshot, page metadata).
input — only url is required; everything else has sane defaults.
output — full scrape result. page metadata (title, OG tags, etc.) is returned under metadata. extracted images are returned as absolute urls — query strings are preserved, and relative srcs are resolved against the page url. screenshots are returned as signed download urls the agent can fetch separately. in rare cases a screenshot can't be captured: when you requested other outputs too, the screenshot field is simply left out while the rest is still returned — but a screenshot-only call that can't deliver errors instead (unsupported_screenshot_output, HTTP 422, when the content type can't be screenshotted) and isn't billed. the result also carries a top-level response_meta.usage block:
alongside response_meta.usage, the result surfaces any non-fatal warnings — stable string codes an agent can switch on. an outsized page is truncated rather than refused, and the code names which part was cut:
| code | what it means for the payload |
|---|---|
screenshot_truncated | the page was taller than the scrolling-capture height cap; the screenshot covers the top of the page. |
links_truncated | the page had more than 30,000 links; the links array is cut at the cap and is incomplete. |
inline_images_truncated | the page had more than 10,000 inline images; the images array is cut at the cap and is incomplete. |
raw_html_truncated | the page body exceeded 10,000,000 characters; raw_html is cut at a tag boundary, never mid-tag. |
metadata_truncated | the page head exceeded 2,000,000 characters; metadata can be missing tags that sat past the cut. |
and if you requested an extract that doesn't apply to the content type (e.g. markdown of a pdf), the field name comes back in an unsupported_fields list — with the rest of the payload still returned.
every input field, its default, and its constraints are documented under the scrape endpoint — with extraction, screenshots, proxies & location, and caching covering the individual blocks.
scrape_asyncsubmit a scrape job to run asynchronously and get back a job_id immediately, instead of holding the connection open. use this for long-running scrapes (heavy js rendering, full-page screenshots of long pages); for a quick one-shot fetch prefer the synchronous scrape tool. then poll scrape_status until the job is done and fetch the page with scrape_result.
takes the same input as scrape plus an optional per-job completion webhook:
output — { "job_id": "..." }.
when a webhook is attached, we deliver a single signed scrape.complete POST to your endpoint once the job reaches a terminal state, with your metadata echoed under data.metadata and the job's usage under data.response_meta.usage — so you can react to completion (and track cost) without polling. verify the X-Cwbl-Signature header with the sdk's verifyWebhookSignature (configure the signing secret in the dashboard under account → webhooks).
the job lifecycle is documented under async scrape; the delivery contract and payload shape under webhooks, with the signature scheme in webhook verification.
scrape_statuslook up the current lifecycle status of an async job: pending, running, done, or failed (with an error message when failed). once the job is done the response also carries a response_meta.usage block (credits, billed engine, resolved proxy tier, screenshot_slices). a cache hit is represented by engine: "cache". poll until done, then call scrape_result.
scrape_resultfetch the extracted content of a completed async job — the same result shape as the synchronous scrape tool (including metadata and response_meta.usage). errors if the job is still pending/running, so check scrape_status first.
mapbuild (or fetch a cached) link-map for a website. combines sitemap discovery with homepage link extraction. use this to enumerate a site before scraping selected pages. each link is just { url }.
max_urls (default 5000, max 100000) is a discovery budget, not a trim at the end — discovery stops as soon as that many urls are found, so a smaller value is a faster, cheaper crawl. limit (default 5000, max 10000) only pages the answer.
returned urls are normalized the same way scrape normalizes its returned url, so map-then-scrape stays on one host. results are ordered with the most useful links first.
the response's response_meta carries pagination, truncation, and a usage block (credits, billed engine, resolved proxy tier). map responses do not include screenshot-slice accounting.
a map stopped by your own max_urls returns exactly that many links with response_capped: false — the signal that the site has more is truncation.discovery_cap_reason:
discovery_cap_reason is one of max_urls, time, file_budget, depth, file_size, unread_files, or null when nothing stopped discovery. only max_urls is a limit you can raise from the request. unread_files means a sitemap file the site publishes could not be read at all this time — often temporary, so asking again later can return more. time, file_budget, depth and file_size mean the site itself is big, slow or deep, and a retry will not help.
see the map endpoint for discovery rules and pagination semantics.
usagereturns the current billing-cycle snapshot: total / used / available credits, used quota percent, max concurrency, and cycle reset timestamp. takes no arguments. what a call costs, and how credits are counted, is documented under credits & pricing.
whoamireturns the organization name, token name, and truncated token preview for the configured api key. useful for confirming which account is in use before credit-consuming operations.
every tool returns an mcp error result (isError: true) when the api call fails. the error text follows a stable format:
agents can branch on the errorName code. the set comes from the sdk's ApiErrorName union plus two synthetic codes added by this mcp (missing_api_key, internal_error):
| code | meaning |
|---|---|
missing_api_key | CRAWLBRULEE_API_KEY is not set in the mcp host's env. |
invalid_credentials | server rejected the api key (revoked, wrong env, etc.). |
service_unavailable | temporary backend failure (HTTP 503). your key is fine — retry with backoff. |
too_many_requests | rate limit hit — back off and retry. |
usage_allocation_error | plan credit / concurrency cap exceeded. show usage to user. |
validation_error | input failed server validation. |
invalid_url | target url was rejected before fetching. |
blocked_url | target url is on the blocklist. |
antibot_blocked | origin's anti-bot defenses blocked the fetch. |
too_many_redirects | origin redirected the fetch in a loop (HTTP 422). the target's doing — don't retry blindly. |
page_too_large | the page's html was too large to process (HTTP 422). terminal — never retry it. |
scrape_error | origin returned an error during scraping. |
unsupported_screenshot_output | screenshot-only request on a content type that can't be screenshotted (HTTP 422). not billed. |
not_found | async job ID unknown (e.g. bad job_id to scrape_status / scrape_result). |
request_timeout | network / read timeout. safe to retry. |
client_closed_request | caller cancelled before completion. |
internal_server_error | unhandled server-side failure. |
crawlbrulee_error | sdk error without a typed name. |
internal_error | bug in this mcp — please open an issue. |
the api docs carry the canonical error reference — every error name, what causes it, and how to recover.
run the built mcp locally:
it will block waiting for an mcp client on stdio. combine with the MCP Inspector for interactive debugging.
this readme covers the mcp server itself — installing it, wiring it into a host, and the tools it exposes. for how the api behaves — endpoints, parameters, and error semantics — the api docs are canonical. the mcp guide covers host setup in more depth.
one api, many ways to call it:
@crawlbrulee/sdk (the sdk this mcp wraps)crawlbrulee on pypinpx crawlbrulee@crawlbrulee/mcp (this one)docs: crawlbrulee.com/docs · dashboard: dashboard.crawlbrulee.com