The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Scrape Le listing page.
Load a URL in headless Chromium and see what will block your scraper — before you write it
Anti-bot vendors, rate limits, robots.txt rules, login walls, console errors, screenshots
Useful? A star or rating is how other developers find it — ★ GitHub · ★ Open VSX · ★ Marketplace
Run Scrape-LE: Check URL Scrapeability (Ctrl+Alt+S / Cmd+Alt+S), enter a URL, and the page loads in a real headless Chromium. The report lands in the output channel: HTTP status, page title, load time, console errors, a full-page screenshot, and four detections. Works in VS Code and VS Code–based editors like Cursor and VSCodium (installable from Open VSX).
One-time setup: run Scrape-LE: Setup Browser to install Chromium (~130MB, into Playwright's browser cache).
| Where | What you get | Install |
|---|---|---|
| VS Code | The same check, in your editor, on a keystroke | Marketplace |
| Cursor, VSCodium, Windsurf | The same extension | Open VSX |
| A terminal or a CI step | The same run over a whole tree, with exit codes | cargo install scrape-le · crates.io |
| Any MCP agent, via Node | analyze_robots_txt over stdio | npx scrape-le-mcp · npm |
| Zed | The MCP server as a context server | add it by hand (no listing yet) |
The same engine runs as an MCP server, so an agent can call it directly instead of you running a command.
| Editor | How |
|---|---|
| VS Code 1.101+ | Nothing to install — the extension registers analyze_robots_txt with agent mode |
| Zed | No listing yet — add the MCP server by hand |
| Claude Code | claude mcp add scrape-le -- npx -y scrape-le-mcp |
| Cursor, Windsurf, anything else | point it at npx scrape-le-mcp |
Given robots.txt contents and a path, reports whether the generic (User-agent: *) rules permit crawling it, plus the crawl delay, disallowed patterns and any sitemaps.
The server takes content and returns data — it reads no files and makes no network requests of its own. Published as scrape-le-mcp on npm and as io.github.nolindnaidoo/scrape-le in the MCP registry.
Most hosts read a JSON config. Add one entry:
-y skips the install prompt on first run. Pin a version if you would rather not track releases — scrape-le-mcp@2.2.6.
Prefer not to go through npx on every launch? Install it once and point at the binary instead:
It speaks MCP over stdio and needs no environment variables, no API key and no configuration of its own. To check it before wiring it into anything:
That prints the tool list and exits — if you see analyze_robots_txt, the server works.
The same check runs from a terminal or an agent loop: a Rust CLI in crate/ of this repository, sharing one signature corpus with the extension — crate/signatures/ and crate/fixtures/ — so CI fails if the two ever disagree about a URL.
The exit code is the answer: 0 clear · 1 a real no · 2 the question was malformed. ## Detections
| Detection | How it works |
|---|---|
| Anti-bot vendors | Response headers, script sources, DOM elements, and window globals fingerprint Cloudflare (incl. Turnstile challenges), reCAPTCHA, hCaptcha, DataDome, and PerimeterX |
| Rate limiting | X-RateLimit-* / RateLimit-* / Retry-After response headers, plus HTTP 429 |
| robots.txt | Fetches <origin>/robots.txt and evaluates the User-agent: * rules against your URL with RFC 9309 semantics — grouped agents, Allow/Disallow longest-match, * wildcards, $ anchors, crawl-delay, sitemaps |
| Authentication | HTTP 401/403, login forms (password + username fields), auth keywords in page text, auth path segments in the final URL |
Honest limitations: signatures are best-effort fingerprints of public integration patterns — a detected widget means the page can challenge you, not that it will, and a clean result is not proof a site allows scraping. Agent-specific robots.txt groups are ignored (only the * rules are reported). Pages get up to 5 seconds to go network-idle after load, so content rendered later than that can be missed by the page-level detections.
| Command | Description |
|---|---|
Scrape-LE: Check URL Scrapeability (Ctrl+Alt+S / Cmd+Alt+S) | Prompt for a URL and run the full check |
Scrape-LE: Check Selected URL | Run the check on the URL in the current selection (also in the right-click menu) |
Scrape-LE: Setup Browser | Install or verify the Chromium browser |
Scrape-LE: Open Settings | Open Scrape-LE settings |
Scrape-LE: Help & Troubleshooting | Built-in documentation |
| Setting | Default | Description |
|---|---|---|
scrape-le.browser.timeout | 30000 | Page-load timeout in ms (5000–120000) |
scrape-le.browser.viewport.width | 1280 | Viewport width |
scrape-le.browser.viewport.height | 720 | Viewport height |
scrape-le.browser.userAgent | "" | Custom User-Agent (empty = Chromium default) |
scrape-le.retry.userAgents | false | On a blocked or failed check, retry under common User-Agents and report which worked |
scrape-le.screenshot.enabled | true | Save a full-page screenshot per check |
scrape-le.screenshot.path | .vscode/scrape-le | Screenshot directory (workspace-relative or absolute) |
scrape-le.screenshot.format | png | png or jpeg |
scrape-le.screenshot.quality | 90 | JPEG quality 0–100 (ignored for png) |
scrape-le.checkConsoleErrors | true | Capture console and page errors while loading |
scrape-le.detections.antiBot | true | Anti-bot vendor detection |
scrape-le.detections.rateLimit | true | Rate-limit detection |
scrape-le.detections.robotsTxt | true | robots.txt fetch + evaluation |
scrape-le.detections.authentication | true | Authentication-wall detection |
scrape-le.notificationsLevel | important | all = every notification, important = warnings + errors, silent = errors only |
scrape-le.statusBar.enabled | true | Show the status bar item |
Twelve languages besides English:
German · Spanish · French · Indonesian · Italian · Japanese · Korean · Portuguese (Brazil) · Russian · Ukrainian · Vietnamese · Chinese (Simplified)
Both halves are covered — the manifest (command titles, setting names and descriptions) and everything shown while the extension runs (notifications, the status bar, quick-picks and prompts). The extension follows VS Code's display language, so it matches whatever the editor is already set to; no setting of its own.
/robots.txt. Nothing is sent anywhere else — no telemetry, no analytics.fetchRobotsTxt builds a URL from an arbitrary origin, which inside an agent loop is an SSRF primitive: the caller supplying the URL is the model, not you. The server analyses robots.txt content you already fetched, and a test asserts no tool accepts a url argument.| What | Where |
|---|---|
| What the tool is allowed to say — scope, output contract, refusals, non-goals | crate/SPEC.md |
| How the extension is built and held together — architecture, invariants, toolchain, release | AGENTS.md |
| How the CLI is built and held together | crate/AGENTS.md |
| What changed | CHANGELOG.md · crate/CHANGELOG.md |
| The tool's page, and the other fifteen | letools.dev/tools/scrape-le |
| Input | Size | Found | Time | Rate | Scan speed |
|---|---|---|---|---|---|
| Header signature scan | 2.83 MB | 20,000 | 5.19 ms | 3,852,946/sec | 544.6 MB/s |
| robots.txt path match | 3.32 MB | 60,000 | 9.64 ms | 6,223,689/sec | 344.3 MB/s |
Median of 7 runs after warmup, on Apple M5 Pro, 24 GB RAM, Node 24.3.0. Inputs are generated
by scripts/benchmark.ts rather than checked in, so the sizes above are
exactly what was measured. Reproduce with bun run benchmark.
These are machine-specific and are not asserted in CI — a benchmark that gates a build only tells you how busy the runner was.
| Metric | Coverage |
|---|---|
| Statements | 93.13% |
| Branches | 84.44% |
| Functions | 95.20% |
| Lines | 94.54% |
379 test cases across 29 files, plus an integration suite that runs
in a real VS Code extension host and an end-to-end test that installs the
built .vsix into a clean profile.
Generated from a real run — coverage/coverage-summary.json and
coverage/test-results.json — by scripts/coverage-readme.js; CI fails if
this section drifts. Reproduce with bun run test:coverage, and the case
count is the one vitest prints.
Sixteen single-purpose tools for the work in front of every model. Each ships a Rust CLI and an MCP server. One page: letools.dev
Get it out
Check it
Guard it
Each stands on its own: no shared crate, no published core. Where two of them agree, it is because the same answer was right twice.
Contact — nolindnaidoo.com · GitHub · LinkedIn
Rust — pixelcoords and pixelactions are one loop: pixelcoords answers where, pixelactions acts there. Their own tools, their own voice — not part of the LE family.
MIT © nolindnaidoo