In-depth architectural comparison of the PDF Reader MCP and Web Hygiene MCP MCP servers. Compare execution transports, security boundaries, tool capabilities, quality scores, and ready-to-paste client installation snippets for Claude, Cursor, Windsurf, and VS Code.
At a Glance & Executive Verdict
PDF Reader MCP
Developer Tools · Local stdio
Quality: 69/100 (Great) | Auth: No auth required
Web Hygiene MCP
Developer Tools · Remote HTTP/SSE
Quality: 57/100 (Good) | Auth: API Key required
Verdict Summary: Choose PDF Reader MCP if you need specialized Developer Tools tools running via a local process. Choose Web Hygiene MCP if your workspace requires Developer Tools integration with remote web transport. Both servers can be configured concurrently in your client's mcpServers manifest.
Which MCP Server Should You Choose?
Choose PDF Reader MCP when:
You need dedicated capabilities in the Developer Tools domain.
You prefer local stdio subprocess transport architecture.
Your security boundary fits: No auth required (Free / Open Source).
You need dedicated capabilities in the Developer Tools domain.
You prefer remote streaming HTTP/SSE transport architecture.
Your security boundary fits: API Key required (BYOK (Pay Provider Direct)).
You have access to required keys: APIFY_API_TOKEN.
Primary tools included: Live HTTP URL and redirect checks, Robots.txt, llms.txt, sitemap, and feed inspection, Sitemap URL listing with filters and last-modified dates.
Document and media evidence when Markdown is not enough. video_timeline (bounded scene/cue timeline) and render_frame (actual decoded frames), and cite_check (quote/location support, not semantic truth), are anymd Pro; PDF operations: inspect (page facts, metadata), render_page (PNG images), extract_regions (crop bounding boxes), ocr_pages / analyze_regions (configured OCR or vision provider), structure (JSON with document map, elements, geometry; profile quality|research adds trust and accessibility reports), compare (page-level diff of sources[0] vs sources[1]).
outline
Return a document heading tree with stable node ids, title paths, page/slide/chapter ranges and Markdown byte ranges. PDF bookmarks or detected headings, Word headings, PowerPoint slides, EPUB chapters and headings, HTML and Markdown. format json (default) or tree. Pass a node id to read to fetch that section.
read
Read any document as clean Markdown: PDF, Word (DOCX), PowerPoint (PPTX), Excel (XLSX/XLS/ODS), CSV, EPUB, HTML or a web URL, Markdown/text, images (metadata + OCR), audio/video (metadata, chapters, subtitles). Pages/slides/sheets carry <!-- page N --> style markers for citation. Long documents stop at max_tokens (default 20000) and end with a cursor to continue; choose pages with pages: "1-5,8". A directory returns its readable files.
Ready-to-Paste Client Configurations
Paste either (or both) of these JSON server blocks into your client config file (e.g. claude_desktop_config.json or ~/.cursor/mcp.json).
PDF Reader MCP is categorized under Developer Tools and uses a local stdio subprocess. In contrast, Web Hygiene MCP belongs to Developer Tools using remote streaming HTTP/SSE transport. Select PDF Reader MCP when you need capabilities focused on developer tools and Web Hygiene MCP when you require tools for developer tools.
Search documents for text: one file, many files, whole directories (recursive, .gitignore aware), or URLs, across every format read supports. Returns each hit as file + page/slide/sheet + a snippet with the match in bold. mode auto (default) finds the exact phrase and falls back to BM25-ranked passages when there is none; literal or ranked force one. Narrow directories with glob, e.g. "*.pdf".
Web Hygiene MCP Tools (6)
Live HTTP URL and redirect checks
Robots.txt, llms.txt, sitemap, and feed inspection
Sitemap URL listing with filters and last-modified dates
Bulk broken-link and performance checks
RSS, Atom, and JSON Feed discovery and reading
Citation verification with optional Internet Archive lookup
Evidence-first PDF MCP. Agent Document Twin with citeable page+bbox evidence.
Live web checks for AI agents, from real HTTP requests at call time: site_overview (robots.txt, llms.txt, sitemaps, feeds), list_site_urls, check_url, check_links (404s, redirects, SSL, slow pages), discover_feeds, read_feed (RSS/Atom/JSON Feed), and verify_citations (does the cited page exist and really carry the quote, title, date? Internet Archive copy for dead links). Remote Streamable HTTP server hosted on Apify; pay per event via your own Apify account, nothing runs locally. Claude Code plugin with verify-sources and site-preflight skills available.