Analyse robots.txt content and report whether a path may be crawled.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
π‘ Paste the JSON block into your client's configuration file under mcpServers, then restart the application.
Load a URL in headless Chromium and see what will block your scraper β before you write it
Anti-bot vendors, rate limits, robots.txt rules, login walls, console errors, screenshots
Useful? A star or rating is how other developers find it β β GitHub Β· β Open VSX Β· β Marketplace
Run Scrape-LE: Check URL Scrapeability (Ctrl+Alt+S / Cmd+Alt+S), enter a URL, and the page loads in a real headless Chromium. The report lands in the output channel: HTTP status, page title, load time, console errors, a full-page screenshot, and four detections. Works in VS Code and VS Codeβbased editors like Cursor and VSCodium (installable from Open VSX).
One-time setup: run Scrape-LE: Setup Browser to install Chromium (~130MB, into Playwright's browser cache).
| Where | What you get | Install |
|---|---|---|
| VS Code | The same check, in your editor, on a keystroke | Marketplace |
| Cursor, VSCodium, Windsurf | The same extension | Open VSX |
| A terminal or a CI step | The same run over a whole tree, with exit codes | cargo install scrape-le Β· crates.io |
| Any MCP agent, via Node | analyze_robots_txt over stdio | npx scrape-le-mcp Β· npm |
| Zed | The MCP server as a context server | add it by hand (no listing yet) |
The same engine runs as an MCP server, so an agent can call it directly instead of you running a command.
| Editor | How |
|---|---|
| VS Code 1.101+ | Nothing to install β the extension registers analyze_robots_txt with agent mode |
| Zed | No listing yet β add the MCP server by hand |
| Claude Code | claude mcp add scrape-le -- npx -y scrape-le-mcp |
| Cursor, Windsurf, anything else | point it at npx scrape-le-mcp |
Given robots.txt contents and a path, reports whether the generic (User-agent: *) rules permit crawling it, plus the crawl delay, disallowed patterns and any sitemaps.
The server takes content and returns data β it reads no files and makes no network requests of its own. Published as scrape-le-mcp on npm and as io.github.nolindnaidoo/scrape-le in the MCP registry.
Most hosts read a JSON config. Add one entry:
-y skips the install prompt on first run. Pin a version if you would rather not track releases β scrape-le-mcp@2.2.6.
Prefer not to go through npx on every launch? Install it once and point at the binary instead:
It speaks MCP over stdio and needs no environment variables, no API key and no configuration of its own. To check it before wiring it into anything:
That prints the tool list and exits β if you see analyze_robots_txt, the server works.
The same check runs from a terminal or an agent loop: a Rust CLI in crate/ of this repository, sharing one signature corpus with the extension β crate/signatures/ and crate/fixtures/ β so CI fails if the two ever disagree about a URL.
The exit code is the answer: 0 clear Β· 1 a real no Β· 2 the question was malformed. ## Detections
| Detection | How it works |
|---|---|
| Anti-bot vendors | Response headers, script sources, DOM elements, and window globals fingerprint Cloudflare (incl. Turnstile challenges), reCAPTCHA, hCaptcha, DataDome, and PerimeterX |
| Rate limiting | X-RateLimit-* / RateLimit-* / Retry-After response headers, plus HTTP 429 |
| robots.txt | Fetches <origin>/robots.txt and evaluates the User-agent: * rules against your URL with RFC 9309 semantics β grouped agents, Allow/Disallow longest-match, * wildcards, $ anchors, crawl-delay, sitemaps |
| Authentication | HTTP 401/403, login forms (password + username fields), auth keywords in page text, auth path segments in the final URL |
Honest limitations: signatures are best-effort fingerprints of public integration patterns β a detected widget means the page can challenge you, not that it will, and a clean result is not proof a site allows scraping. Agent-specific robots.txt groups are ignored (only the * rules are reported). Pages get up to 5 seconds to go network-idle after load, so content rendered later than that can be missed by the page-level detections.
| Command | Description |
|---|---|
Scrape-LE: Check URL Scrapeability (Ctrl+Alt+S / Cmd+Alt+S) | Prompt for a URL and run the full check |
Scrape-LE: Check Selected URL | Run the check on the URL in the current selection (also in the right-click menu) |
Scrape-LE: Setup Browser | Install or verify the Chromium browser |
Scrape-LE: Open Settings | Open Scrape-LE settings |
Scrape-LE: Help & Troubleshooting | Built-in documentation |
| Setting | Default | Description |
|---|---|---|
scrape-le.browser.timeout | 30000 | Page-load timeout in ms (5000β120000) |
scrape-le.browser.viewport.width | 1280 | Viewport width |
scrape-le.browser.viewport.height | 720 | Viewport height |
scrape-le.browser.userAgent | "" | Custom User-Agent (empty = Chromium default) |
scrape-le.retry.userAgents | false | On a blocked or failed check, retry under common User-Agents and report which worked |
scrape-le.screenshot.enabled | true | Save a full-page screenshot per check |
scrape-le.screenshot.path | .vscode/scrape-le | Screenshot directory (workspace-relative or absolute) |
scrape-le.screenshot.format | png | png or jpeg |
scrape-le.screenshot.quality | 90 | JPEG quality 0β100 (ignored for png) |
scrape-le.checkConsoleErrors | true | Capture console and page errors while loading |
scrape-le.detections.antiBot | true | Anti-bot vendor detection |
scrape-le.detections.rateLimit | true | Rate-limit detection |
scrape-le.detections.robotsTxt | true | robots.txt fetch + evaluation |
scrape-le.detections.authentication | true | Authentication-wall detection |
scrape-le.notificationsLevel | important | all = every notification, important = warnings + errors, silent = errors only |
scrape-le.statusBar.enabled | true | Show the status bar item |
Twelve languages besides English:
German Β· Spanish Β· French Β· Indonesian Β· Italian Β· Japanese Β· Korean Β· Portuguese (Brazil) Β· Russian Β· Ukrainian Β· Vietnamese Β· Chinese (Simplified)
Both halves are covered β the manifest (command titles, setting names and descriptions) and everything shown while the extension runs (notifications, the status bar, quick-picks and prompts). The extension follows VS Code's display language, so it matches whatever the editor is already set to; no setting of its own.
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/scrape-le)<a href="https://allmcps.com/mcp/scrape-le"><img src="https://allmcps.com/api/badge/scrape-le?style=directory" alt="Scrape Le on AllMCPs" /></a>