In-depth architectural comparison of the MinerU Ecosystem and Agent Search MCP MCP servers. Compare execution transports, security boundaries, tool capabilities, quality scores, and ready-to-paste client installation snippets for Claude, Cursor, Windsurf, and VS Code.
At a Glance & Executive Verdict
MinerU Ecosystem
Search & Data Extraction · Local stdio
Quality: 64/100 (Good) | Auth: API Key required
Agent Search MCP
Search & Data Extraction · Local stdio
Quality: 68/100 (Great) | Auth: No auth required
Verdict Summary: Choose MinerU Ecosystem if you need specialized Search & Data Extraction tools running via a local process. Choose Agent Search MCP if your workspace requires Search & Data Extraction integration with local subprocess execution. Both servers can be configured concurrently in your client's mcpServers manifest.
Which MCP Server Should You Choose?
Choose MinerU Ecosystem when:
You need dedicated capabilities in the Search & Data Extraction domain.
You prefer local stdio subprocess transport architecture.
Your security boundary fits: API Key required (Free / Open Source).
Official MinerU document parsing MCP (mineru-open-mcp on PyPI). Converts PDFs, doc/docx/ppt/pptx, images, and spreadsheets to Markdown via the MinerU API; free Flash mode without an API key (about 20 pages per file); optional MINERUAPITOKEN for higher limits.
Free multi-engine MCP search server — 8 free engines (DDG, Sogou, Bing, Baidu, Wikipedia, Startpage, Yandex, Mojeek), waterfall progressive search, multi-source verification, content enrichment, news search, language auto-detection, CLI. Zero API keys needed. npx agent-search-mcp
Category & Scope
Tools & Capabilities Breakdown
MinerU Ecosystem Tools (2)
parse_documents
Convert PDFs, Word (DOCX), PowerPoint (PPTX), Excel (XLSX in Flash mode), images, and public document URLs—or HTML page URLs—into Markdown using the MinerU cloud API.
**World effects:** Reads local paths you pass; fetches http(s) URLs. Uploads file or URL content to mineru.net for processing (do not use for data you must not send off-device). Does not modify or delete originals. May write Markdown (and sometimes images) under `output_dir` or the server default when results are saved or inline content is too large.
**Auth & limits:** Without `MINERU_API_TOKEN`, uses free Flash mode: Markdown-only output, service limits apply. With `MINERU_API_TOKEN`, higher per-file page limits and optional extra output formats per MinerU plans; token is read from env (or HTTP Bearer when using streamable HTTP).
**Use this when:** The user needs full-document extraction, tables/formulas as HTML/Latexa, batch conversion, or per-file PDF page ranges. **Do not use** for listing supported OCR script codes—call `get_ocr_languages` instead. **Not a substitute** for offline-only or strictly local parsers.
**Parameters (intent):** `file_sources` is a list of path/URL strings or `{"source": "…", "pages": "1-5"}` objects (PDF page ranges; Flash allows simple `N` or `N-M`). `enable_ocr` defaults to true. `language` is an OCR/script code (default `ch`); see `get_ocr_languages` for valid values. Set `model` to `"html"` only when every source is a web page URL; otherwise omit. `output_dir` overrides where large or batch results are written.
get_ocr_languages
Return the supported MinerU OCR / script language codes (e.g. ch, en, japan, latin). Read-only; no uploads. Use before setting `language` on `parse_documents` for scanned or multilingual documents. Do not use for converting files—call `parse_documents` instead.
Ready-to-Paste Client Configurations
Paste either (or both) of these JSON server blocks into your client config file (e.g. claude_desktop_config.json or ~/.cursor/mcp.json).
MinerU Ecosystem is categorized under Search & Data Extraction and uses a local stdio subprocess. In contrast, Agent Search MCP belongs to Search & Data Extraction using local stdio subprocess. Select MinerU Ecosystem when you need capabilities focused on search & data extraction and Agent Search MCP when you require tools for search & data extraction.
Quick public-web search for current facts and discovery.
Use free_search_advanced for verification or domain filters. Use free_extract for one selected URL.
Returns compact multi-source evidence with confidence, relevance, source_count, stop_reason, evidence_budget, and partialFailures. Provider families are counted once even when several adapters use the same upstream. Optional API providers run only when credentials and policy allow.
@readOnly true @idempotent true — makes outbound HTTP requests to configured search engines. Injection detection and SSRF protection active.
free_search_advanced
Verification-oriented search with domain filters and waterfall fallback.
Use for claim checking, Chinese-source search, or publisher restrictions. Start with free_search for quick discovery. The response exposes source_count, stop_reason, evidence_budget, and partialFailures.
@readOnly true @idempotent true — runs waterfall progressive search across policy-allowed engines. Makes outbound HTTP requests to search engines and optionally to Jina Reader for content enrichment.
free_extract
Extract one selected URL into clean Markdown.
Use after search when snippets are insufficient, or for a URL supplied by the user. It fetches one page; it does not search or bulk-extract URLs.
Behavior: Makes an outbound HTTP request to Jina Reader (r.jina.ai) which fetches and converts the page to markdown. Has SSRF protection: blocks private IPs, localhost, and metadata endpoints. 10s request timeout — pages exceeding this will fail with a timeout error. HTTP errors (4xx, 5xx) are returned as structured error responses.
fetch_github_readme
Fetch README content from a GitHub repository.
Best for: Getting project documentation quickly.
Not recommended for: Non-GitHub URLs — use free_extract instead.
@readOnly true @idempotent true — makes outbound HTTP requests to raw.githubusercontent.com.
fetch_csdn_article
Fetch content from a CSDN blog article.
Best for: Chinese developer blog content on CSDN.
Not recommended for: Other Chinese sites — use free_extract instead.
@readOnly true @idempotent true — makes outbound HTTP requests to the CSDN article URL.
fetch_juejin_article
Fetch content from a Juejin article.
Best for: Chinese developer articles on Juejin.
Not recommended for: Non-Juejin content — use free_extract instead.
@readOnly true @idempotent true — makes outbound HTTP requests to juejin.cn API.
search_with_synthesis
Opt-in deep search with waterfall verification and content enrichment. Returns structured results plus a prompt_hint; it does not call an external LLM.
Use when multi-source evidence needs a synthesis handoff. free_search is the default for discovery; this path uses more tokens and network time.
@readOnly true @idempotent true — runs waterfall search across free+paid engines with content enrichment.