Skip to main content
AllMCPs
BrowseBestCategoriesStackCompareToolsGuidesBlog
Log in Submit MCP

Stay in the loop

Get new MCP servers and top picks in your inbox.

AllMCPs

The open directory for discovering and installing Model Context Protocol servers.

AllMCPs on GitHub (opens in a new tab)
Launched onTiny Startupstinystartups.com
Explore
  • Browse servers
  • Best MCP servers
  • Categories
  • MCP clients
  • Agent prompts
  • Stack Builder
  • Compare servers
  • Random discovery New
  • Submit a server
  • Pricing & Boost Boost
Learn
  • Guides hub
  • What is MCP?
  • Install guide
  • Build an MCP server
  • Deploy an MCP server
  • Security guide
  • Troubleshooting
  • MCP for SEO & AEO
  • Protocol versioning
  • Blog & updates
Tools
  • All developer tools
  • Config generator
  • Config validator
  • Config auditor
  • MCP playground
  • Token calculator
  • OpenAPI β†’ MCP
  • Badge generator
For agents
  • REST API docs
  • Trust & traffic Live
  • Remote MCP server SSE β†— (opens in a new tab)
  • llms.txt β†— (opens in a new tab)
  • Catalog JSON β†— (opens in a new tab)
Company
  • About
  • Advertise Sponsor
  • Contact
  • GitHub β†— (opens in a new tab)
  • Terms
  • Privacy
AllMCPs VerifiedAllMCPs VerifiedFeatured on Nick LaunchesFeatured on Nick LaunchesLaunch Llama NewsletterLaunch Llama NewsletterVerified DR - allmcps.comVerified DR - allmcps.comFeatured on SaaSGrowFeatured on SaaSGrowFeatured on Twelve ToolsFeatured on Twelve ToolsFeatured on Saaspa.geFeatured on Saaspa.geFeatured on Findly.toolsFeatured on Findly.toolsFeatured on Startup FameFeatured on Startup FameFeatured on LaunchKiwiFeatured on LaunchKiwiFeatured on ScrollLaunchFeatured on ScrollLaunchFeatured on DailyPingsFeatured on DailyPingsFazier badgeFazier badgeFeatured on NewTool.siteFeatured on NewTool.siteFeatured on saasfame.comFeatured on saasfame.comDR Checker - Domain RatingDR Checker - Domain RatingListed on Turbo0Listed on Turbo0Launched on LaunchBoard - Product Launch PlatformLaunched on LaunchBoard - Product Launch PlatformList on SimilarlabsList on Similarlabshttps://codetrendy.comhttps://codetrendy.comListed on DevTool.ioFeatured on BuildlistFeatured on BuildlistLaunched on Tiny StartupsFeatured on ShowMeBestAIFeatured on ShowMeBestAIFind us on LaunchZoneFind us on LaunchZoneAllMCPs VerifiedAllMCPs VerifiedFeatured on Nick LaunchesFeatured on Nick LaunchesLaunch Llama NewsletterLaunch Llama NewsletterVerified DR - allmcps.comVerified DR - allmcps.comFeatured on SaaSGrowFeatured on SaaSGrowFeatured on Twelve ToolsFeatured on Twelve ToolsFeatured on Saaspa.geFeatured on Saaspa.geFeatured on Findly.toolsFeatured on Findly.toolsFeatured on Startup FameFeatured on Startup FameFeatured on LaunchKiwiFeatured on LaunchKiwiFeatured on ScrollLaunchFeatured on ScrollLaunchFeatured on DailyPingsFeatured on DailyPingsFazier badgeFazier badgeFeatured on NewTool.siteFeatured on NewTool.siteFeatured on saasfame.comFeatured on saasfame.comDR Checker - Domain RatingDR Checker - Domain RatingListed on Turbo0Listed on Turbo0Launched on LaunchBoard - Product Launch PlatformLaunched on LaunchBoard - Product Launch PlatformList on SimilarlabsList on Similarlabshttps://codetrendy.comhttps://codetrendy.comListed on DevTool.ioFeatured on BuildlistFeatured on BuildlistLaunched on Tiny StartupsFeatured on ShowMeBestAIFeatured on ShowMeBestAIFind us on LaunchZoneFind us on LaunchZone
Β© 2026 Jackalope Digital LLC. All rights reserved.
  1. Home
  2. πŸ”Ž Search & Data Extraction
  3. Lookacrawler
Lookacrawler logo
Health: Not checked yetWe have not completed a health check for this listing yet.No health check has run yet.

Lookacrawler

User RatingsBe the first to rate and review this MCP server! Enrichment pendingWe haven’t run our AI enrichment pass on this listing yet, so the overview, use cases, and FAQ below may be sparse or missing. We work through the catalog over time β€” check back soon.
View Repository

Free, open-source, token-efficient local alternative to Firecrawl with native MCP Server for LLMs (>73% token reduction, Playwright stealth anti-bot bypass).

Quick Install

Automated & IDE Setup

Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β€” or use 1-click editor setup below.

Add to CursorAdd to VS Code
Manual Client & Custom JSON ConfigExpand JSON β–Ύ

Client Config & Setup

Choose your client or environment
Target File:~/Library/Application Support/Claude/claude_desktop_config.json
claude_desktop_config.json
{
  "mcpServers": {
    "lucasmartins-ai-lookacrawler": {
      "command": "npx",
      "args": [
        "-y",
        "lucasmartins-ai-lookacrawler"
      ]
    }
  }
}

πŸ’‘ Paste the JSON block into your client's configuration file under mcpServers, then restart the application.

Install Directory Badge Claim listing AlternativesπŸ”Ž More in Search & Data Extraction

Documentation Overview

πŸ•·οΈ LookaCrawler

Free, open-source, token-efficient local alternative to Firecrawl with native Model Context Protocol (MCP) Server for LLMs.

License: MIT Bun TypeScript MCP Ready GitHub Stars Local-First


πŸ“Œ Why LookaCrawler?

Web crawling for Large Language Models (LLMs) is broken by default: modern web pages contain massive HTML bloat (scripts, tracking pixels, nested divs, navigation headers, stylesheets), costing thousands of wasted tokens per page.

LookaCrawler is an open-source, token-optimized local crawler that strips >73% to 90% of web bloat, extracts clean Markdown, bypasses anti-bot barriers with stealth Playwright drivers, and exposes a native Model Context Protocol (MCP) Server ready for Claude Desktop, Cursor, and Antigravity.


πŸ₯Š Comparison: LookaCrawler vs Alternatives

FeatureπŸ•·οΈ LookaCrawlerπŸ”₯ Firecrawl (Cloud)⚑ Jina Reader
Pricing / Cost$0.00 (100% Free Open Source)$16 to $99+/monthRate-limited API
Token Reduction>73% to 90% pruning + FootnotesStandard MarkdownBasic Markdown
Autonomous CrawlingNative map & crawl (BFS + Regex)Cloud CrawlerSingle-page only
Pre-Crawl ActionsNative Playwright (click, scroll, fill)Paid AddonNone
Link FormattingInline, References Footnotes, StripInline onlyInline only
Data Privacy100% Local (Zero Telemetry)Cloud ProviderCloud API
MCP IntegrationNative Tools + Resources + PromptsCommunity WrapperNone
Stealth & Anti-BotReal Chrome + Stealth FingerprintCloud ProxiesBasic Headers
Local SQLite CacheBuilt-in (24h TTL cache)Redis / Paid AddonNone
JS SPA SupportPlaywright + Chrome PoolCloud HeadlessHeadless

πŸš€ Key Features

  • Token Economy First: Automatically prunes scripts, styles, inline SVGs, tracking tags, navigations, footers, redundant forms, and boilerplate containers with high link density (>80%).
  • Advanced Link & Image Formatting:
    • link_format: Choose between inline (standard markdown), references (footnote citations [1], saving ~25% tokens on repetitive URLs), or strip (pure text).
    • image_mode: Choose between ignore (zero tokens), alt_only (preserves semantic context without URL bloat), or markdown (full ![alt](url)).
  • Autonomous Mapping & Recursive Crawling:
    • map_website: Inspects /robots.txt, sitemaps, and root anchors to discover all pages in a domain.
    • crawl_website: Breadth-first autonomous crawling with max depth, max pages, route regex filters, and real-time token accounting.
  • Pre-Crawl Browser Actions: Automate clicks, scrolls, typing, and waits in Playwright before extracting content (dismiss cookie banners, scroll for infinite loading, expand accordions).
  • Dual Hybrid Crawling Engine:
    • fast: Ultra-fast native HTTP GET with backoff. Auto-escalates to deep if an anti-bot challenge is encountered.
    • deep: Headless Playwright engine launching real Google Chrome with stealth patches (navigator.webdriver cleared, WebGL spoofed, CDP leaks stripped) to transparently crawl Cloudflare/Turnstile-protected pages.
  • Native MCP Ecosystem:
    • Tools: extract_web_content, crawl_website, map_website, batch_extract_web_content, extract_structured_data.
    • Resources: Live telemetry at crawler://metrics and cache analytics at crawler://cache/stats.
    • Prompts: Pre-engineered templates crawl-and-summarize and compare-pages.
  • Local SQLite Caching: Stores extracted Markdown in crawler_cache.sqlite to eliminate duplicate network calls.
  • Structured JSON & Metadata Extraction: Extracts Open Graph tags (og:title, og:description), publication dates, canonical URLs, and custom CSS selectors.

πŸ”Œ 1-Click MCP Setup (Claude Desktop & Cursor)

Add LookaCrawler to your claude_desktop_config.json or Cursor MCP settings:

config.json
{
  "mcpServers": {
    "lookacrawler": {
      "command": "bun",
      "args": ["run", "/absolute/path/to/lookacrawler/index.ts"]
    }
  }
}

Now you can prompt Claude or Cursor:

"Crawl https://example.com/docs and extract the API documentation using LookaCrawler."


πŸ“¦ Quick Start & CLI Usage

1. Installation

Requires Bun 1.1+ (high-performance runtime with native SQLite):

bash
# Clone the repository
git clone https://github.com/lucasmartins-ai/lookacrawler.git
cd lookacrawler

# Install dependencies
bun install

2. CLI Commands

bash
# Single URL fast Markdown extraction
bun run cli.ts extract https://news.ycombinator.com --mode fast

# Single URL fast Markdown extraction with reference footnotes
bun run cli.ts extract https://example.com --link-format references --output page.md

# Headless Playwright deep extraction with CSS selector target
bun run cli.ts extract https://example.com --mode deep --selector "main" --json

# Discover all website URLs and sitemaps
bun run cli.ts map https://example.com --max-urls 500

# Recursively crawl documentation with regex filtering and token accounting
bun run cli.ts crawl https://example.com/docs --max-depth 2 --max-pages 15 --link-format references

# Batch concurrent multi-URL crawling
bun run cli.ts batch https://site1.com https://site2.com --concurrency 4

# Structured JSON schema extraction
bun run cli.ts structured https://example.com --schema '{"title":"h1","links":"a"}'

# Start MCP Server via SSE on port 3000
bun run cli.ts serve --transport sse --port 3000

3. Docker Deployment

bash
# Build and run Docker container
docker build -t lookacrawler .
docker run -p 3000:3000 lookacrawler

πŸ§ͺ Architecture & Testing

text
Incoming URL ──► [Local SQLite Cache Check] ──(Hit)──► Return Cached Markdown
                       β”‚ (Miss)
                       β–Ό
            [Fast HTTP GET Request] ──(Blocked?)──► [Auto-Escalate to Deep Stealth]
                       β”‚                                      β”‚
                       β–Ό                                      β–Ό
            [HTML DOM Tree Parser] β—„β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                       β”‚
                       β–Ό
       [Aggressive Token Noise Pruner]
       (Strips SVG, Nav, Ads, Tracking, CSS, JS)
                       β”‚
                       β–Ό
         [Mozilla Readability Engine]
                       β”‚
                       β–Ό
         [Turndown Markdown Converter] ──► Return Clean LLM Markdown

Run test suite:

bash
bun test

⭐ Star & Support

If LookaCrawler saves you API fees and token costs:

  • ⭐ Star this repository to help other developers find it!
  • πŸ’‘ Open an Issue / PR for new stealth bypasses or crawler features.

Built by LookADev

lookacrawler is built and maintained by LookADev, an engineering studio specializing in AI agents, web architecture, and token optimization.

Start a project β†’ lookadev.com Β· Email: lucas@lookadev.com

πŸ“„ License

Open-source software licensed under the MIT License.

Read the full README β†’View source on GitHub β†’

Related MCP Servers

View all in Search & Data Extraction View all alternatives
  • Apify MCP Server logoApify MCP Server

    Extract data from any website with thousands of scrapers, crawlers, and automations on Apify Store ⚑

    πŸ”Ž Search & Data Extraction0 views
    Compare vs Apify MCP Server β†’
  • Firecrawl MCP Server logoFirecrawl MCP Server
    Verified

    Official Firecrawl server to search the web and scrape, crawl, map, and extract structured data from any site for LLMs. Handles JS-rendered pages, PDFs, and batch jobs; hosted remote MCP with OAuth or self-host.

    πŸ”Ž Search & Data Extraction5 views
    Compare vs Firecrawl MCP Server β†’
  • Open WebSearch logoOpen WebSearch

    Web search using free multi-engine search (NO API KEYS REQUIRED) β€” Supports Bing, Baidu, DuckDuckGo, Brave, Exa, and CSDN.

    πŸ”Ž Search & Data Extraction3 views
    Compare vs Open WebSearch β†’
  • ShadowCrawl logoShadowCrawl

    Rust MCP stealth scraper: anti-bot search/scrape with CDP fallback + HITL non-robot.

    πŸ”Ž Search & Data Extraction1 views
    Compare vs ShadowCrawl β†’

Reviews

No reviews yet β€” be the first to share how this listing worked for you.

Frequently Asked Questions about Lookacrawler

Add the following block to your claude_desktop_config.json under mcpServers: "mcpServers": { "lookacrawler": { "command": "npx", "args": ["-y", "lucasmartins-ai/lookacrawler"] } }

AllMCPs Directory Badge

Full Badge Customizer

Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.

Badge Style:
Live Dynamic SVG PreviewLookacrawler AllMCPs Directory Badge
Markdown (GitHub README)
[![AllMCPs](https://allmcps.com/api/badge/lucasmartins-ai-lookacrawler?style=directory)](https://allmcps.com/mcp/lucasmartins-ai-lookacrawler)
HTML Embed
<a href="https://allmcps.com/mcp/lucasmartins-ai-lookacrawler"><img src="https://allmcps.com/api/badge/lucasmartins-ai-lookacrawler?style=directory" alt="Lookacrawler on AllMCPs" /></a>

Technical Specs & Signals

CategoryπŸ”ŽSearch & Data Extraction
More technical detailsExpand β–Ύ
TransportSTDIO
RuntimeNode.js
Last updatedSep 7, 2026
Views0
Unique ViewsTotal visits recorded for this listing page on AllMCPs.
Installs0
Installs & Copy ActionsTotal times users copied install commands or configuration snippets for this server.
35Quality signal: Fair Β· 35/100How this signal is calculated β–Ύ
Server availabilityNot measured

Not scored for repo-hosted servers β€” we can't reach the running server, only its GitHub page. Hosted MCP endpoints are health-checked live.

Verified ownership8/20
Documentation & tools17/30
Adoption & activity1/15
Community engagement0/10

A guidance signal from public completeness & health data β€” not a user rating. New listings start lower and rise as they add docs, get verified, and grow adoption. Signals we can't observe for a listing are skipped, not counted against it.

β˜… Spotlight Slot

Feature Your MCP Server

Get maximum visibility for your server across our directory, search results, and detail pages.

Spotlight Your Server

Own this project?

This directory is pre-filled from public sources. Claim via GitHub README, site badge, or DNS TXT to unlock edit access and the Official badge and attach your website β€” proof is checked automatically, then reviewed by our team.

Free dofollow backlink: add your website and place the AllMCPs badge on it β€” no claim needed. We detect it automatically and keep it verified as long as the badge stays live.

Claim & get free dofollow

Share & Embed

Add our SVG badge (dark/light directory styles) or embeddable widget to your site.

Explore more

More in πŸ”Ž Search & Data Extraction β†’Best MCP servers for Web Search & Scraping β†’Alternatives to Lookacrawler β†’Install in Claude DesktopInstall in CursorInstall in VS Code