Skip to main content
AllMCPs
BrowseBestCategoriesStackCompareToolsGuidesBlog
Log in Submit MCP

Stay in the loop

Get new MCP servers and top picks in your inbox.

AllMCPs

The open directory for discovering and installing Model Context Protocol servers.

AllMCPs on GitHub (opens in a new tab)
Launched onTiny Startupstinystartups.com
Explore
  • Browse servers
  • Best MCP servers
  • Categories
  • MCP clients
  • Agent prompts
  • Stack Builder
  • Compare servers
  • Random discovery New
  • Submit a server
  • Pricing & Boost Boost
Learn
  • Guides hub
  • What is MCP?
  • Install guide
  • Build an MCP server
  • Deploy an MCP server
  • Security guide
  • Troubleshooting
  • MCP for SEO & AEO
  • Protocol versioning
  • Blog & updates
Tools
  • All developer tools
  • Config generator
  • Config validator
  • Config auditor
  • MCP playground
  • Token calculator
  • OpenAPI β†’ MCP
  • Badge generator
For agents
  • REST API docs
  • Trust & traffic Live
  • Remote MCP server SSE β†— (opens in a new tab)
  • llms.txt β†— (opens in a new tab)
  • Catalog JSON β†— (opens in a new tab)
Company
  • About
  • Advertise Sponsor
  • Contact
  • GitHub β†— (opens in a new tab)
  • Terms
  • Privacy
AllMCPs VerifiedAllMCPs VerifiedFeatured on Nick LaunchesFeatured on Nick LaunchesLaunch Llama NewsletterLaunch Llama NewsletterVerified DR - allmcps.comVerified DR - allmcps.comFeatured on SaaSGrowFeatured on SaaSGrowFeatured on Twelve ToolsFeatured on Twelve ToolsFeatured on Saaspa.geFeatured on Saaspa.geFeatured on Findly.toolsFeatured on Findly.toolsFeatured on Startup FameFeatured on Startup FameFeatured on LaunchKiwiFeatured on LaunchKiwiFeatured on ScrollLaunchFeatured on ScrollLaunchFeatured on DailyPingsFeatured on DailyPingsFazier badgeFazier badgeFeatured on NewTool.siteFeatured on NewTool.siteFeatured on saasfame.comFeatured on saasfame.comDR Checker - Domain RatingDR Checker - Domain RatingListed on Turbo0Listed on Turbo0Launched on LaunchBoard - Product Launch PlatformLaunched on LaunchBoard - Product Launch PlatformList on SimilarlabsList on Similarlabshttps://codetrendy.comhttps://codetrendy.comListed on DevTool.ioFeatured on BuildlistFeatured on BuildlistLaunched on Tiny StartupsFeatured on ShowMeBestAIFeatured on ShowMeBestAIFind us on LaunchZoneFind us on LaunchZoneAllMCPs VerifiedAllMCPs VerifiedFeatured on Nick LaunchesFeatured on Nick LaunchesLaunch Llama NewsletterLaunch Llama NewsletterVerified DR - allmcps.comVerified DR - allmcps.comFeatured on SaaSGrowFeatured on SaaSGrowFeatured on Twelve ToolsFeatured on Twelve ToolsFeatured on Saaspa.geFeatured on Saaspa.geFeatured on Findly.toolsFeatured on Findly.toolsFeatured on Startup FameFeatured on Startup FameFeatured on LaunchKiwiFeatured on LaunchKiwiFeatured on ScrollLaunchFeatured on ScrollLaunchFeatured on DailyPingsFeatured on DailyPingsFazier badgeFazier badgeFeatured on NewTool.siteFeatured on NewTool.siteFeatured on saasfame.comFeatured on saasfame.comDR Checker - Domain RatingDR Checker - Domain RatingListed on Turbo0Listed on Turbo0Launched on LaunchBoard - Product Launch PlatformLaunched on LaunchBoard - Product Launch PlatformList on SimilarlabsList on Similarlabshttps://codetrendy.comhttps://codetrendy.comListed on DevTool.ioFeatured on BuildlistFeatured on BuildlistLaunched on Tiny StartupsFeatured on ShowMeBestAIFeatured on ShowMeBestAIFind us on LaunchZoneFind us on LaunchZone
Β© 2026 Jackalope Digital LLC. All rights reserved.
  1. Home
  2. πŸ”Ž Search & Data Extraction
  3. CrawlEyes MCP Server
C
Health: Not checked yetWe have not completed a health check for this listing yet.No health check has run yet.

CrawlEyes MCP Server

User RatingsBe the first to rate and review this MCP server! Enrichment pendingWe haven’t run our AI enrichment pass on this listing yet, so the overview, use cases, and FAQ below may be sparse or missing. We work through the catalog over time β€” check back soon.
View Repository

Web scraping & search for AI agents: extract, SearXNG/Tavily search, rerank. No API key needed.

Quick Install

Automated & IDE Setup

Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β€” or use 1-click editor setup below.

Add to CursorAdd to VS Code
Manual Client & Custom JSON ConfigExpand JSON β–Ύ

Client Config & Setup

Choose your client or environment
Target File:~/Library/Application Support/Claude/claude_desktop_config.json
claude_desktop_config.json
{
  "mcpServers": {
    "crawleyes-mcp-server": {
      "command": "npx",
      "args": [
        "-y",
        "crawleyes-mcp-server"
      ]
    }
  }
}

πŸ’‘ Paste the JSON block into your client's configuration file under mcpServers, then restart the application.

Install Directory Badge Claim listing AlternativesπŸ”Ž More in Search & Data Extraction

Documentation Overview

CrawlEyes β€” Web Scraping & Search Toolkit for AI Agents

GitHub stars GitHub license Python MCP MIT

CrawlEyes gives AI agents reliable full-text extraction (web_extract) and robust search (web_search) backends β€” the "eyes" that let agents see and read the web. Built and tested against Hermes Agent.

Also ships as a standard MCP server, so any MCP client (Claude Desktop, Cursor, other agents) can reuse the same search + extraction capabilities.

CrawlEyes demo

Why CrawlEyes?

Most agent toolkits cover one slice of the pipeline. CrawlEyes is the rare all-in-one that you can actually run behind the Great Firewall without external accounts.

Typical agent toolkitCrawlEyes
πŸ” SearchAPI key required, often blocked in CNβœ… SearXNG (self-hosted) + Tavily keyless fallback β€” zero config, zero key
πŸ“„ ExtractionSeparate scraper, or Firecrawl SaaSβœ… Built-in Crawl4AI full-text extract, ~89% noise removal
🧠 Semantic rerankRarely includedβœ… Local fastembed rerank β€” no torch, ~50MB model
πŸ”Œ MCP serverOften missingβœ… Standard MCP tools (search + extract + deep_research + sitemap), any client
🌐 China-friendlyMostly English/GFW-blockedβœ… Tested on a real mainland China server (baidu + yandex)

Zero API keys. Zero external accounts. One command. CrawlEyes is the only toolkit in this space that combines search + extraction + semantic reranking + MCP in a single, China-friendly, self-hosted package.

Features

CapabilityWhereWhy it matters
Full-text extractionscripts/crawl4ai_cli.pyHeadless-browser scraping β†’ clean Markdown; handles ~80% of JS/dynamic/UA-blocked pages
Content denoising (P1)crawl4ai_cli.py --noise-filterPrunes nav/ads/comments via Crawl4AI's PruningContentFilter — measured 24.6k→2.8k chars (~89% noise removed) on a typical article
Retry with backoff (P3)crawl4ai_cli.py --retry NExponential backoff (1s/2s/4s) on transient failures
Browser session reuse (P4)crawl4ai_cli.py --session NAMEReuses the browser context across scrapes in one process β€” no cold-start per URL
Keyword-focused extractioncrawl4ai_cli.py --bm25 KEYWORDKeeps only paragraphs relevant to a keyword (experimental β€” BM25 is English-centric; works best on English docs)
Search (primary)SearXNG (self-hosted meta-search)Privacy-friendly search aggregator
Search (fallback)Tavily keyless APIZero-config, no-key fallback when SearXNG is down/empty
Search orchestrationplugins/searxng-tavily/Hermes plugin provider: SearXNG first β†’ auto-fallback to Tavily keyless; three-state circuit breaker (3 fails β†’ 60s cooldown β†’ half-open) + shared SQLite cache (TTL 3600s)
Semantic reranking (P2)scripts/crawl_search_standalone.pyLocal embedding rerank of search results with fastembed + BAAI/bge-small-zh-v1.5 (512-dim, no torch dependency, ~50MB, cached) β€” puts relevant results first. Measured: crawler-relevant items 0.817/0.732 float to top, irrelevant 0.302/0.139 sink
MCP server (P5)scripts/mcp_crawl_server.pyExposes search + extract + deep_research + sitemap as standard MCP tools (stdio default, or streamable-http for remote clients). Works in any MCP client, no Hermes dependency. Extracted content is sanitized against prompt-injection (strips invisible chars + prompt-hijack lines). Unified rate limiting + exponential backoff guard every tool (sliding window, per-tool cost) so concurrent agent calls can't hammer downstream services
Sitemap discovery (P1)crawleyes/sitemap.pysitemap(origin) β†’ parses sitemap.xml (plain / gzip / index-recursion) with robots.txt fallback, returns a deduped URL map. Zero-key way to discover a site's URL surface for whole-site fetch or deep-research seeding
Multi-format extract (P0)extract(..., format=)markdown (default) / fit (denoised) / raw (unfiltered) / markdown_with_citations β€” pick the level of cleanup you need
RAG-ready interfacescrawleyes/rag.pyOne-liners markdown(url) / search_markdown(query) β†’ clean, sanitized, LLM-ready Markdown for RAG corpora
Deep researchcrawleyes/deep_research.pydeep_research(topic) β†’ decomposes topic into sub-questions β†’ searches β†’ extracts β†’ synthesizes a cited Markdown report. Optional LLM (any OpenAI-compatible endpoint); degrades to evidence-aggregate mode without one
Verificationscripts/Clean subprocess scripts to verify each backend end-to-end per Hermes profile

Project layout

Code
plugins/searxng-tavily/   Hermes web-search provider plugin (SearXNG β†’ Tavily keyless fallback)
                          + three-state circuit breaker + shared SQLite cache
scripts/
  crawl4ai_cli.py          Universal scraping CLI (URL β†’ Markdown), with denoise/retry/session/BM25
  crawl_search_standalone.py  Standalone search (SearXNG β†’ Tavily) + optional semantic rerank.
                             No Hermes dependency β€” usable anywhere, powers the MCP server.
  mcp_crawl_server.py      Standard MCP server exposing search + extract + deep_research + sitemap (stdio)
  single_env_check.py      Verify crawl4ai provider registered+available+extracts (one profile)
  verify_searxng_tavily.py Verify searxng-tavily provider: normal path + forced fallback
  agent_link_check.py      Verify full agent tool chain: web_search_tool dispatch + logs

Quick start

1. Install Crawl4AI (China-friendly mirrors)

bash
python3 -m venv .venv
# Use Tsinghua PyPI mirror for speed (or any mirror you prefer)
.venv/bin/pip install -i https://pypi.tuna.tsinghua.edu.cn/simple crawl4ai
# Playwright browser kernel β€” use npmmirror binary mirror if cdn.playwright.dev is blocked
PLAYWRIGHT_DOWNLOAD_HOST=https://registry.npmmirror.com/-/binary/playwright \
  .venv/bin/python -m playwright install chromium
.venv/bin/crawl4ai-setup

2. Scrape a page

bash
.venv/bin/python scripts/crawl4ai_cli.py https://example.com          # stdout Markdown
.venv/bin/python scripts/crawl4ai_cli.py https://example.com -o out.md  # to file
.venv/bin/python scripts/crawl4ai_cli.py URL --text --max-words 5000   # plain text, truncated

# Multi-format extraction (markdown|fit|raw|markdown_with_citations)
.venv/bin/python scripts/crawl4ai_cli.py URL --format raw              # unfiltered source markdown
.venv/bin/python scripts/crawl4ai_cli.py URL --format markdown_with_citations  # + source URLs

# Respect robots.txt (opt-in, default off)
.venv/bin/python scripts/crawl4ai_cli.py URL --respect-robots

# Denoise nav/ads + retry 3x + reuse session across scrapes
.venv/bin/python scripts/crawl4ai_cli.py URL --noise-filter --retry 3 --session s1

3. Use the search + rerank (standalone, no Hermes)

server.ts
# Optional: local semantic rerank of results (fastembed + bge-small-zh, auto-downloaded)
.venv/bin/pip install -i https://pypi.tuna.tsinghua.edu.cn/simple fastembed

# SearXNG first, Tavily keyless fallback, then rerank
SEARXNG_URL=https://your-searxng .venv/bin/python -c "
import sys; sys.path.insert(0, 'scripts')
from crawl_search_standalone import CrawlSearch
r = CrawlSearch(rerank=True).search('your query')
print(r['data']['web'])"

China-network note: the embedding model downloads from HuggingFace, which is blocked on mainland networks. Set HF_ENDPOINT=https://hf-mirror.com and HF_HUB_DISABLE_XET=1 (hf-mirror doesn't support the xet protocol and returns 401 without this).

4. Run as an MCP server (any client)

bash
# Any MCP client can connect via stdio (default):
.venv/bin/python scripts/mcp_crawl_server.py
# Exposes tools:
#   search(query, limit)                 - SearXNG β†’ Tavily keyless, rerank, retry+rate-limit
#   extract(url, max_words, format)      - markdown|fit|raw|markdown_with_citations
#   deep_research(topic, num_questions)  - multi-round cited report
#   sitemap(origin, max_urls)            - URL map from sitemap.xml / robots.txt

# Or serve over HTTP (streamable-http) for remote clients:
.venv/bin/python -m crawleyes.mcp_crawl_server --transport http --port 8765 --host 127.0.0.1
#   β†’ clients connect to http://127.0.0.1:8765/mcp
#   (host/port configurable; default 127.0.0.1:8765)

For Hermes specifically, add to config.yaml:

yaml
mcp_servers:
  crawl:
    command: "/path/to/crawl/.venv/bin/python"
    args: ["/path/to/crawl/scripts/mcp_crawl_server.py"]
    timeout: 90
    connect_timeout: 60

4b. Firecrawl-compatible /scrape endpoint

Already using Firecrawl's Python SDK? Point it at CrawlEyes and keep your code:

bash
.venv/bin/python -m crawleyes.firecrawl_api --port 8899 --host 127.0.0.1
#   POST /v2/scrape  β†’  { success, data: { markdown, metadata } }
#   GET  /healthz    β†’  health check
server.ts
from firecrawl import Firecrawl
fc = Firecrawl(api_url="http://127.0.0.1:8899", api_key="ignored")
doc = fc.scrape(url="https://example.com")   # β†’ { markdown, metadata }

Read the full README β†’View source on GitHub β†’

Related MCP Servers

View all in Search & Data Extraction View all alternatives
  • Apexapi MCP logoApexapi MCP

    Call 120+ AI models and live web context (scrape/crawl/extract) from any MCP client with one key

    πŸ”Ž Search & Data Extraction0 views
    Compare vs Apexapi MCP β†’
  • R
    Riveter

    MCP server for Riveter's enrichment, scraping, and monitoring API

    πŸ”Ž Search & Data Extraction0 views
    Compare vs Riveter β†’
  • Tavily MCP Server logoTavily MCP Server

    MCP server for advanced web search using Tavily API.

    πŸ”Ž Search & Data Extraction0 views
    Compare vs Tavily MCP Server β†’
  • Apify MCP Server logoApify MCP Server

    Extract data from any website with thousands of scrapers, crawlers, and automations on Apify Store ⚑

    πŸ”Ž Search & Data Extraction0 views
    Compare vs Apify MCP Server β†’

Reviews

No reviews yet β€” be the first to share how this listing worked for you.

Frequently Asked Questions about CrawlEyes MCP Server

Add the following block to your claude_desktop_config.json under mcpServers: "mcpServers": { "crawleyes-mcp-server": { "command": "npx", "args": ["-y", "CrawlEyes MCP Server"] } }

AllMCPs Directory Badge

Full Badge Customizer

Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.

Badge Style:
Live Dynamic SVG PreviewCrawlEyes MCP Server AllMCPs Directory Badge
Markdown (GitHub README)
[![AllMCPs](https://allmcps.com/api/badge/crawleyes-mcp-server?style=directory)](https://allmcps.com/mcp/crawleyes-mcp-server)
HTML Embed
<a href="https://allmcps.com/mcp/crawleyes-mcp-server"><img src="https://allmcps.com/api/badge/crawleyes-mcp-server?style=directory" alt="CrawlEyes MCP Server on AllMCPs" /></a>

Technical Specs & Signals

CategoryπŸ”ŽSearch & Data Extraction
More technical detailsExpand β–Ύ
TransportSTDIO
RuntimeNode.js
Last updatedSep 7, 2026
Views0
Unique ViewsTotal visits recorded for this listing page on AllMCPs.
Installs0
Installs & Copy ActionsTotal times users copied install commands or configuration snippets for this server.
27Quality signal: Emerging Β· 27/100How this signal is calculated β–Ύ
Server availabilityNot measured

Not scored for repo-hosted servers β€” we can't reach the running server, only its GitHub page. Hosted MCP endpoints are health-checked live.

Verified ownership8/20
Documentation & tools11/30
Adoption & activity1/15
Community engagement0/10

A guidance signal from public completeness & health data β€” not a user rating. New listings start lower and rise as they add docs, get verified, and grow adoption. Signals we can't observe for a listing are skipped, not counted against it.

β˜… FeaturedMoxie Docs MCP logo

Moxie Docs MCP

MCP & Agent Skills for Automated Documentation, and codebase conventions + context

Explore Server β†’

Own this project?

This directory is pre-filled from public sources. Claim via GitHub README, site badge, or DNS TXT to unlock edit access and the Official badge and attach your website β€” proof is checked automatically, then reviewed by our team.

Free dofollow backlink: add your website and place the AllMCPs badge on it β€” no claim needed. We detect it automatically and keep it verified as long as the badge stays live.

Claim & get free dofollow

Share & Embed

Add our SVG badge (dark/light directory styles) or embeddable widget to your site.

Explore more

More in πŸ”Ž Search & Data Extraction β†’Best MCP servers for Web Search & Scraping β†’Alternatives to CrawlEyes MCP Server β†’Install in Claude DesktopInstall in CursorInstall in VS Code