Paper search, arXiv full-text reading, and citation graphs over arXiv, Semantic Scholar, OpenAlex.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent — or use 1-click editor setup below.
💡 Paste the JSON block into your client's configuration file under mcpServers, then restart the application.
Inspect callable tools, capabilities, and parameters exposed to AI agents by Paper Search (arXiv + Semantic Scholar + OpenAlex).
get_openalex_workCallable MCP tool function
get_openalex_citationsCallable MCP tool function
get_openalex_referencesCallable MCP tool function
search_openalex_worksCallable MCP tool function
search_openalex_authorsCallable MCP tool function
search_openalex_institutionsCallable MCP tool function
Remotely-callable MCP server for academic paper search, full-text retrieval & image→LaTeX, served at https://latex-tools.online/mcp.
Three corpora behind one normalized interface:
arxiv (default) — search, metadata, and full-text (HTML / markdown / LaTeX source)semanticscholar (alias s2) — the full S2 API surface: citation graph, authors, recommendations, full-text snippets, bulk datasetsopenalex (alias oa) — 316M all-field works: citation graph, authors with h-index, institutions, topics, influence metricsPlus a unified search_all that fuses all three corpora, image→LaTeX OCR, and LaTeX lint + PDF→text tooling.
| Tool | Purpose |
|---|---|
search_all(query, max_results=10, sources='arxiv,semanticscholar,openalex') | Unified search. Fans out to all three corpora concurrently, de-duplicates the same work (by DOI/title) and re-ranks with Reciprocal Rank Fusion. Each hit carries sources (who found it) + an ids map for follow-up calls. Prefer this for broad lookups. |
search_papers(query, source='arxiv', max_results=10, sort_by='relevance') | Single-corpus search. arXiv query accepts plain text or field syntax (ti: au: cat:cs.CL abs: + AND/OR). |
get_paper(paper_id, source='arxiv') | One paper's full record. S2 id accepts S2 id / DOI: / ARXIV: / CorpusId:. |
search_by_author(author, source='arxiv') | Papers by author, newest first. |
list_recent(category, source='arxiv') | Latest in a category (arXiv code or S2 field of study). |
list_categories(source='arxiv') | Common category codes. |
read_paper(paper_id, format='markdown') | FULL text (arXiv). markdown = body with formulas as $LaTeX$; html = raw LaTeXML page; latex = original manuscript .tex source. |
list_paper_sources() | Available corpora. |
read_paper fetch chain: arxiv.org/html/{id} → ar5iv fallback (markdown/html), or arxiv.org/e-print/{id} tarball main .tex (latex). Formulas are recovered from the LaTeXML alttext invariant.
| Tool | Purpose |
|---|---|
search_medical(query, study_types='rct,meta-analysis,systematic-review', year_from=0, max_results=10, fetch_fulltext=True) | Clinical literature search. Queries PubMed, filters by research type via Publication-Type tags and re-ranks by the evidence pyramid (meta-analysis / systematic review > RCT > cohort > ...), so real trials surface above high-cited reviews/guidelines that pure-citation ranking floats up. Open-access full text is attached from Europe PMC by PMID. If the type filter yields nothing it auto-relaxes (flagged filter_relaxed). query is English keyword/boolean text — do NL/multilingual query understanding upstream. Backed by NCBI E-utilities + Europe PMC (both free, no key required). |
Turn a formula or table image back into LaTeX (e.g. a figure cropped from a paper) without needing your own vision model. Backed by the co-located recognize service (PaddleOCR-VL / DeepSeek-OCR / texify).
| Tool | Purpose |
|---|---|
recognize_formula(image_url=... or image_base64=..., model='deepseek-ocr') | Formula image → LaTeX. image_url is downloaded server-side (with SSRF guards). Returns {latex, model, elapsed_ms}. |
recognize_table(image_url=... or image_base64=..., model='deepseek-ocr') | Table image → LaTeX tabular. |
list_ocr_models() | Available OCR models (deepseek-ocr, paddleocr-vl, texify). |
Companions to the LaTeX/PDF web tools at latex-tools.online — same backends, exposed over MCP.
| Tool | Purpose |
|---|---|
lint_latex(code) | Check a LaTeX snippet for errors and return an auto-fixed version. Returns {errors, fixed_code, summary_en, summary_zh, elapsed_ms}. |
extract_pdf(pdf_url=... or pdf_base64=..., formula=True, table=True) | PDF → clean Markdown/LaTeX text via MinerU (useful for papers with no open-access full text). pdf_url is downloaded server-side (SSRF-guarded). Content-addressed + cached: a recently-seen or small PDF returns content in one call; a fresh PDF (MinerU is GPU-heavy, minutes) returns status='running' + a task_id. |
extract_pdf_result(task_id) | Fetch an extract_pdf job by task_id. Returns content once status='done'; while 'running', content is null — call again shortly. |
get_openalex_work · get_openalex_citations · get_openalex_references · search_openalex_works (filters: year range, open-access, min-citations, institution)search_openalex_authors · search_openalex_institutionsget_openalex_trends · list_openalex_topicsget_paper_citations · get_paper_references · get_paper_authorsmatch_paper_title · autocomplete_paperssearch_papers_bulk (≤1000, sortable, token paging) · get_papers_batchsearch_authors · get_author · get_author_papers · get_authors_batchsearch_snippets (search inside paper body)recommend_papers_for_paper · recommend_papers_from_exampleslist_dataset_releases · get_dataset_release · get_dataset_download_links · get_dataset_diffs| Var | Default | Notes |
|---|---|---|
PAPER_MCP_HOST | 127.0.0.1 | |
PAPER_MCP_PORT | 9400 | |
PAPER_MCP_PATH | /mcp | |
SEMANTIC_SCHOLAR_API_KEY | — | optional; raises S2 rate limit. Set via /etc/paper-mcp.env in prod. |
MCP_MAX_PER_HOUR | 300 | Direct-client JSON-RPC POST budget per IP. |
MCP_WORKER_MAX_PER_HOUR | 300 | Trusted reverse-proxy Worker budget per HMAC-derived connection key. Raw keys are not retained. |
MCP_WORKER_SHARED_MAX_PER_HOUR | 2400 | Shared ceiling across all trusted Worker connections. |
MCP_RATE_COOLDOWN_SEC | 300 | Minimum fast-rejection cooldown after a bucket reaches its limit. |
paper-mcp.service on tencent-us (43.130.32.180), WorkingDirectory /opt/paper-mcp, loopback port 9400.https://latex-tools.online/mcp → 127.0.0.1:9400/mcp.X-MCP-Worker after validating the upstream platform. Never pass through a client-supplied value.$uri, not $request, for the MCP route./etc/paper-mcp.env (SEMANTIC_SCHOLAR_API_KEY).tencent-us operations backup, not by this source repository; never commit /etc/paper-mcp.env.This repo is the source of truth. The server runs an independent copy under /opt/paper-mcp (not auto-synced):
Production parity verified on 2026-07-23: main@07f6bbe8622aa063f56ee222a40d19c5d4264048 matches all 12 deployed Python source files byte-for-byte. The older copy embedded in latex-tools-deploy/paper-mcp/ is not a deployment source.
_USER_AGENT, backoff).read_paper covers ~80%+ of papers via official HTML; older scan-only papers may have no full text.docs repo on 2026-06-07; that copy is gone.MIT © MCPServings. See LICENSE.
Factual signals from GitHub, npm, and our automated checks — not a rating.
No reviews yet — be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/paper-search-arxiv-semantic-scholar-openalex)<a href="https://allmcps.com/mcp/paper-search-arxiv-semantic-scholar-openalex"><img src="https://allmcps.com/api/badge/paper-search-arxiv-semantic-scholar-openalex?style=directory" alt="Paper Search (arXiv + Semantic Scholar + OpenAlex) on AllMCPs" /></a>