The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Scrapiq listing page.
MCP server for Scrapiq — turn any URL into clean text, markdown, or JSON for LLM/RAG pipelines, directly from your MCP client.
Scrapiq is a lightweight open-source HTTP API that fetches a web page and returns clean content — boilerplate stripped. This server exposes it as a Model Context Protocol (MCP) tool so Claude Desktop, Cursor, and any MCP client can extract clean web content with one call.
Dependency-free: pure Python stdlib, no pip packages, no node_modules. Two transports: stdio for local clients, streamable HTTP for remote clients.
The server is listed in the official MCP Registry as io.scrapiq/scrapiq and runs at:
No key, no install — add that URL as a remote MCP server in any client that supports streamable HTTP:
Not on PyPI yet, so install straight from this repo:
Requires a running Scrapiq instance (see Scrapiq README — git clone, pip install -e ".[dev]", then scrapiq). Point the server at it:
Stateless: one POST /mcp per JSON-RPC message (or batch), replies with application/json. It issues no Mcp-Session-Id and offers no server→client SSE stream, so GET /mcp answers 405 by design. CORS is open, so browser-based clients (e.g. MCP Inspector) can call it directly.
Add to claude_desktop_config.json:
scrapiq_extractExtract a web page into clean structured content.
Arguments:
url (string, required) — the URL to extractformat (string, optional) — "markdown" (default) | "text" | "json"max_chars (integer, optional) — truncate content to N charsExample:
Returns title, content, links, and metadata — no ads, no nav, no scripts.
MIT