The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Crawlforge MCP Server listing page.
30 web scraping, crawling, deep-research & autonomous-extraction tools for Claude, Cursor & any MCP client.
Clean Markdown & structured JSON from any site. Get started with 1,000 free credits — no credit card required.
⭐ Star us on GitHub to follow along — it genuinely helps others discover the project.
agent, a unified multi-format scrape, document processing, stealth browsing, and more, callable directly from your AI assistant.extract_with_llm runs against a local Ollama model out of the box: no LLM API key, no per-token cost, and your data never leaves your machine. Cloud (OpenAI/Anthropic) is opt-in.agent — describe what you need in natural language; it plans, gathers, and shapes an answer under orchestrator-enforced hard stops (max steps/URLs/wall-clock) — no URLs required.| CrawlForge MCP | Firecrawl | Raw scraping API | |
|---|---|---|---|
| Native MCP server | ✅ 30 tools | ✅ | ❌ |
| Free tier | ✅ 1,000 credits, rollover | Limited | Varies |
| Self-hosted / local LLM extraction (Ollama) | ✅ default, $0/token | ❌ | ❌ |
| Autonomous agent (no URLs needed) | ✅ agent | ✅ | ❌ |
| Deep research with source verification | ✅ deep_research | Partial | ❌ |
| Browser automation / actions | ✅ scrape_with_actions | ✅ | Varies |
| Stealth / anti-detection engines | ✅ Chromium + Camoufox | ✅ | Add-on |
| Pre-built site templates | ✅ 10 sites | ❌ | ❌ |
| License | MIT | AGPL-3.0 | Proprietary |
Comparison reflects publicly documented capabilities at time of writing. CrawlForge is MIT-licensed and MCP-first — built to plug straight into AI coding assistants.
Every tool requires a CrawlForge API key — new accounts get 1,000 free trial credits to start. The recommended path signs you in through the browser, so the key is never pasted into a terminal (a coding agent can run this for you and relay the URL):
It prints an approval URL; open it, approve, and the key is stored in ~/.crawlforge/config.json. Then run crawlforge init to register the MCP server with your client. Or use the interactive wizard, which also configures your clients:
This will:
Don't have an API key? Get one free at https://www.crawlforge.dev/signup
One-step setup (v4.6.0+):
crawlforge initdetects your API key, installs the agent skill, and idempotently merges the MCP config stanza into Claude Code, Claude Desktop, and Cursor. Usecrawlforge init --all --yesto configure every detected client non-interactively.
Add to claude_desktop_config.json:
Location:
~/Library/Application Support/Claude/claude_desktop_config.json%APPDATA%/Claude/claude_desktop_config.json~/.config/Claude/claude_desktop_config.jsonRestart Claude Desktop to activate.
The setup wizard automatically configures Claude Code by adding to ~/.claude.json:
After setup, restart Claude Code to activate.
The setup wizard automatically configures Cursor by adding to ~/.cursor/mcp.json:
Restart Cursor to activate.
n8n's built-in MCP Client Tool node connects over Streamable HTTP (works on n8n Cloud and self-hosted). Run the server in HTTP mode:
Then point the MCP Client Tool node at http://<host>:10000/mcp with transport HTTP Streamable and a Bearer credential set to the same API key. On self-hosted n8n you can instead use the community n8n-nodes-mcp node over STDIO (npx -y crawlforge-mcp-server).
Full guide: docs/n8n-integration.md
Which launch command?
npx -y crawlforge-mcp-serverneeds no global install and always runs the published version (recommended for Claude Desktop). For a global install (npm i -g crawlforge-mcp-server), use the dedicatedcrawlforge-mcpbin — it resolves on yourPATH, so it survives Node/nvm version switches. The barecrawlforgecommand still launches the server when an MCP client spawns it over stdio (backward compatibility for configs created before v4.2.5); interactively it's the CLI — runcrawlforge mcpto start the server by hand.
CrawlForge requires a CrawlForge API key — every tool is metered and consumes credits. New accounts get 1,000 free trial credits to start. Get a key at crawlforge.dev/signup.
All Tools (API key required)
| Tool | Credits | What it does |
|---|---|---|
fetch_url | 1 | Fetch content from any URL |
extract_text | 1 | Extract clean text from web pages |
extract_links | 1 | Get all links from a page |
extract_metadata | 1 | Extract page metadata (title, OG tags, schema.org) |
scrape_template | 1 | Structured data from well-known sites (Amazon, GitHub, LinkedIn, YouTube, Reddit, Hacker News, npm, and more) without writing selectors |
list_ollama_models | 1 | List the Ollama models installed locally (helps you pick a model for extract_with_llm) |
get_batch_results | 1 | Retrieve paginated results for a batch_scrape job by batchId |
read_result | 1 | Search, slice, read lines or a JSON path from a result a tool returned with truncated: true and a result_handle (kept 1 hour on your own machine) — never fetch the page again |
scrape | 2 | Unified single-fetch, multi-format extraction. Pass a formats array (markdown/html/rawHtml/text/links/metadata/screenshot/json-schema, plus {type:"highlights",query} and {type:"question",question} for only the matching sentences, table rows and code blocks, verbatim with offsets into the markdown: +1 credit once per call, mode:"model" +3) plus onlyMainContent; one fetch serves every requested format with per-format partial-success warnings. escalate:true retries a blocked plain fetch once in the stealth browser and returns the page instead of the block (+5 as the projected ceiling; the charge drops back to 2 when the plain fetch succeeded and no escalation ran) |
scrape_structured | 2 | Extract structured data with CSS selectors |
extract_embedded_state | 2 | Read a page's embedded JavaScript state — __NEXT_DATA__, React Server Component payloads, Nuxt, Apollo, Redux, <script type="application/json"> — with a path to scope the result. No LLM in the extraction path |
extract_content | 2 | Enhanced content extraction |
map_site | 2 | Discover and map website structure (optional search= ranks the discovered URLs) |
process_document | 2 | Multi-format document processing |
localization | 2 | Multi-language and geo-location management |
track_changes | 3 | Monitor content changes over time |
analyze_content | 3 | Comprehensive content analysis |
extract_structured | 3 | LLM-powered schema-driven extraction (your own LLM key or local Ollama) |
extract_with_llm | 3 | Natural-language extraction. Defaults to a local Ollama model; pass provider: "openai" | "anthropic" with the matching key for cloud models (external LLM billed by your provider) |
summarize_content | 4 | Generate intelligent summaries |
crawl_deep | 4 | Deep crawl entire websites |
search_web | 5 | Search the web using Google Search API |
reddit_search | 5 | Search Reddit posts/comments or read a full thread — reddit.com blocks direct scraping, so this reads the Arctic Shift community archive (free, no Reddit credentials). A Reddit-wide search spends a web search to discover posts, so it is priced with search_web |
serp_rank | 5 | Check where a domain ranks in Google's real organic SERP for a keyword (the position search_web can't give). Powered by DataForSEO (DATAFORSEO_LOGIN/DATAFORSEO_PASSWORD, billed to your own DataForSEO account). Returns { configured:false } and charges 0 credits until configured |
batch_scrape | 5 | Process multiple URLs simultaneously |
scrape_with_actions | 5 | Browser automation chains |
generate_llms_txt | 5 | Generate AI interaction guidelines |
stealth_mode | 5 | Anti-detection browser management |
agent | 8 | Autonomous research/extraction from a natural-language prompt — no URLs required. Plans, gathers, and shapes an answer under hard safety stops (max steps/URLs/wall-clock enforced by the orchestrator, never the LLM) |
deep_research | 10 | Multi-stage research with source verification |
Ten tools (scrape, fetch_url, extract_content, crawl_deep, batch_scrape, stealth_mode, scrape_with_actions, process_document, deep_research, extract_embedded_state) accept max_inline_chars (default 40,000; env CRAWLFORGE_MAX_INLINE_CHARS): a result over it comes back as a preview plus a result_handle for read_result, with the full result kept for 1 hour under ~/.crawlforge/results/ on your own machine — nothing is uploaded.
For the full canonical capabilities reference (all tools, CLI commands, stealth engines, research workflow), see SKILL.md.
Every tool is metered and requires an API key. New accounts get 1,000 free trial credits — no credit card required to start.
| Plan | Credits | Best For |
|---|---|---|
| Free | 1,000 one-time | Testing & personal projects |
| Hobby ($19) | 5,000 / month | Small projects & development |
| Professional ($99) | 50,000 / month | Professional use & production |
| Business ($399) | 250,000 / month | Large scale operations |
All plans include:
CrawlForge tracks the current MCP spec (2025-06-18) plus select experimental extensions:
scrape, map_site, serp_rank, reddit_search, search_web, extract_structured, and crawl_deep return machine-parseable structuredContent alongside the usual text, validated against a published outputSchema; legacy clients keep working off the text.isError: true result the calling model can read and retry from, instead of a raw JSON-RPC protocol error.tools/list ordering (client prompt-cache friendly), and cacheable-result hints on read-only tools.crawl_deep, batch_scrape, deep_research, agent — for clients that support polling; synchronous results are still returned for clients that don't.See docs/mcp-spec-adoption.md for wire-level examples and client-compatibility notes.
extract_with_llm with Ollama)extract_with_llm defaults to a local Ollama model — no LLM-provider key, no per-token LLM costs, and no data leaving your machine (the CrawlForge credit cost still applies).
deep_research (Camoufox)deep_research automatically retries sources that block the normal fetch path (Reddit, Quora, forums, and Cloudflare/DataDome-protected pages return HTTP 403) through a real fingerprinted browser, then re-extracts from the rendered HTML. It's bounded (RESEARCH_MAX_STEALTH_RETRIES, default 8, plus a per-page timeout) and lazy — the browser stack only loads when a source is actually blocked.
Engine selection (RESEARCH_STEALTH_ENGINE):
auto (default) — prefer Camoufox (Firefox anti-detect), fall back to Chromium stealth, then plain fetch.camoufox — force Camoufox.chromium — force the Chromium stealth engine.Headless Chromium cannot clear modern challenges (Cloudflare Turnstile, DataDome) — Camoufox can. In testing it recovered Quora and Trustpilot pages that were otherwise fully blocked. To enable it, install the optional dependency and run its one-time binary fetch:
Without the Camoufox binary, deep_research silently falls back to Chromium stealth and then to plain fetch — no errors, just lower recovery on heavily-protected sites. Disable the whole fallback with RESEARCH_STEALTH_FALLBACK=false.
Note: Hard IP-reputation blocks (e.g. Reddit's edge
403) resist headless stealth from any IP and require residential/mobile proxies, which CrawlForge does not provide. See docs/stealth-engines.md for details.
Your configuration is stored at ~/.crawlforge/config.json:
Once configured, use these tools in your AI assistant:
~/.crawlforge/config.json{crawlforge.dev, www.crawlforge.dev, api.crawlforge.dev}, HTTPS required). Setting CRAWLFORGE_API_URL to an arbitrary host is blocked at parse time.scrape_with_actions accepts only 7 action types (wait, click, type, press, scroll, screenshot, executeJavaScript). No download, file-write, or arbitrary cross-page navigation primitives exist.executeJavaScript action throws by default. Set ALLOW_JAVASCRIPT_EXECUTION=true at deploy time to enable (not recommended in production).deep_research (>50 URLs), batch_scrape (sync mode, >25 URLs), crawl_deep (projected >500 pages), extract_structured (schema has >3 required fields with no LLM configured). Credit-low situations also elicit. Confirmation is best-effort: if the MCP client does not support elicitation the tool proceeds (fail-open).withAuth() and is metered — credits are checked and deducted before execution, and a valid API key is required for every tool (fail-closed since v3.0.18).See docs/sandboxing-and-approvals.md for the full reference.
v3.0.3 (2025-10-01): Removed authentication bypass vulnerability. All users must authenticate with valid API keys.
For the full security policy and how to report a vulnerability, see SECURITY.md.
MIT License - see LICENSE file for details.
Contributions are welcome! Please read our Contributing Guide first.
Built with ❤️ by the CrawlForge team