The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Scholar MCP Server listing page.
A FastMCP server for the scholarly citation landscape (papers, patents, books, and standards), giving LLMs a unified way to search, cross-reference, and retrieve prior art across all four source types via Semantic Scholar, EPO Open Patent Services, Open Library, and standards bodies (NIST, IETF, W3C, ETSI), with OpenAlex enrichment and optional docling-serve PDF/full-text conversion.
Documentation | Config wizard | PyPI | Docker
externalIds are automatically enriched with publisher, edition, cover URL, and subject data from Open Library.sync-standards. ISO, IEC, IEEE have a live-fetch fallback for unsynced identifiers; CC and CEN have no live API and require a sync first. Citations matching standards patterns (RFC, ISO, NIST SP, IEEE, EN, CC) are automatically enriched with structured standard_metadata including identifier, title, body, status, and full-text URL when available (see docs/guides/standards.md)..deb and .rpm packages with systemd service and security hardening.Per-domain depth is uneven. Papers currently have the richest tool surface (citation graph, recommendations, cross-referencing to all three other domains); standards are the leanest. That reflects public data availability, not a value hierarchy: writing a paper typically needs all four source types for citations and prior art. Parity work is tracked in GitHub issues and milestones; the roadmap shows intent, not a completeness commitment.
With this server mounted in an MCP client (Claude, etc.), you can:
search_papers + get_citations + enrich_paper.find_bridge_papers + get_citation_graph.get_patent + batch_resolve + standards/book enrichment.generate_citations.resolve_standard_identifier + get_standard.If you add optional extras via the PROJECT-EXTRAS-START / PROJECT-EXTRAS-END sentinels in pyproject.toml, document them below:
Scholar-mcp ships two optional-dependency groups:
[mcp]: installs FastMCP; required to run scholar-mcp serve and expose tools over stdio/HTTP.[all]: currently identical to [mcp]; reserved for future optional backends.For MCP-server usage:
Installing the bare pvliesdonk-scholar-mcp package is enough for library use (from scholar_mcp import ...) but the scholar-mcp serve CLI requires [mcp].
To run the newest merged code instead of the newest release, use the rolling edge tag. It is rebuilt on every merge to main and carries no version identity. See Image tags for the full tag list.
A compose.yml ships at the repo root as a starting point. Copy .env.example to .env, edit, and docker compose up -d.
To attach a remote Python debugger (development only; the protocol is unauthenticated), see Remote debugging.
Download .deb or .rpm packages from the GitHub Releases page. Both install a hardened systemd unit; env configuration is sourced from /etc/scholar-mcp/env (copy from the shipped /etc/scholar-mcp/env.example).
Download the .mcpb bundle from the GitHub Releases page and double-click to install, or run:
Claude Desktop prompts for required env vars via a GUI wizard, with no manual JSON editing needed.
For manual Claude Desktop configuration and setup options, see Claude Desktop deployment.
Artifacts ship on three channels. Each row lists exactly what that channel publishes.
| Channel | Version identity | Artifacts |
|---|---|---|
edge (rolling) | None; the commit is the identity | Docker image :edge rebuilt on every merge to main; .mcpb bundle as the mcpb-bundle-edge workflow artifact; Claude Code plugin .zip as the plugin-zip-edge artifact; rolling unstable docs version. It leaves no git tag, GitHub release, or PyPI entry behind. |
| Pre-release | vX.Y.Z-rc.N, computed and reviewed in its release pull request | PyPI (as the pre-release X.Y.ZrcN); GitHub release with wheels, sdist, .deb/.rpm packages, .mcpb bundle, plugin .zip, and SBOM attached; Docker image under its immutable vX.Y.Z-rc.N tag plus the ordering-aware rolling rc tag. Skips the plugin marketplace, the MCP registry, and the docs deploy. |
| Stable | vX.Y.Z | Everything: PyPI, Docker (version tag plus ordering-aware latest / vX / vX.Y), .deb/.rpm, GitHub release assets (wheels, sdist, .mcpb bundle, plugin .zip, SBOM), plugin marketplace and MCP registry entries (when the release is the newest stable), versioned docs with an ordering-aware latest alias. |
Pre-releases reach PyPI so that a candidate's .mcpb bundle installs: the bundle points at PyPI rather than carrying the code. Ordinary installers never see them, because a PEP 440 resolver skips pre-releases unless the requirement pins one or you pass --pre. Ask for a candidate by name with pip install pvliesdonk-scholar-mcp==X.Y.ZrcN. PyPI spells it in the PEP 440 canonical form, while tags use SemVer. Rolling pointers are ordering-aware, so a patch release cut from an old release/X.Y branch never moves latest-style tags back to older content, and a candidate for an already-released version never moves rc. See Release process for the full model.
For library usage (embedding the domain logic without the MCP transport), import from the scholar_mcp package directly. Backend clients live under src/scholar_mcp/_s2_client.py, _epo_client.py, _openlibrary_client.py, and _standards_client.py.
The server registers a built-in get_server_info tool (via fastmcp_pvl_core.register_server_info_tool) so operators can confirm the deployed version with a single MCP call. The default response carries server_name, server_version, and core_version. Servers that talk to a remote upstream wire upstream version reporting inside the DOMAIN-UPSTREAM-START / DOMAIN-UPSTREAM-END sentinel in src/scholar_mcp/server.py; see CLAUDE.md for the wiring pattern.
Core environment variables shared across all fastmcp-pvl-core-based services:
| Variable | Default | Description |
|---|---|---|
SCHOLAR_MCP_KV_STORE_URL | file:///data/state | Persistent-state backend URL shared by every pvl-core subsystem that needs state. memory:// is in-process and lost on restart; file:///path persists on one server; redis://, dynamodb:// and mongodb:// each need their matching extra. When unset, defaults to file:///data/state (the volume family Docker images mount), or to memory://; with a warning; on a host where that directory is not usable. |
FASTMCP_LOG_LEVEL | INFO | Log level for FastMCP internals and app loggers (DEBUG / INFO / WARNING / ERROR / CRITICAL). The -v CLI flag overrides to DEBUG. |
FASTMCP_ENABLE_RICH_LOGGING | true | Set false for plain or structured JSON log output. |
Domain-specific variables go below under Domain configuration.
This server inherits opt-in per-subject authorization from fastmcp-pvl-core. The default posture is off: every authenticated caller can use every tool, resource, and prompt. Turn it on by pointing SCHOLAR_MCP_ACL_PATH at a TOML ACL file; the middleware is installed only when the path is set, and individual tools opt in by declaring meta={"required_scope": "<scope>"} at registration. A tool without required_scope is unrestricted regardless of caller.
Wire it in by uncommenting the acl_path field in src/scholar_mcp/config.py and the AuthorizationMiddleware stanza in src/scholar_mcp/server.py; both ship as commented stubs in the scaffold.
<kind>:<id> convention is documentation only; the library treats each subject as a literal string.* is the only library-treated special scope: it grants every required scope. Subject-side wildcards (* as an ACL key) are rejected at load time.read:project-foo or write:vault/personal; fastmcp-pvl-core treats every scope except * as opaque.The subject string used as a value in the bearer-tokens TOML (SCHOLAR_MCP_BEARER_TOKENS_FILE) is the same string used as a key in the ACL TOML. Same string, opposite roles, so keep the two files consistent when adding or removing a principal. See Mapped bearer tokens in the authentication guide for the bearer-tokens TOML schema.
In single-token mode (SCHOLAR_MCP_BEARER_TOKEN) every authenticated caller shares one subject, the library's default (currently "bearer-anon"); override it with SCHOLAR_MCP_BEARER_DEFAULT_SUBJECT; reference that string as the ACL key. When no auth is configured (no SCHOLAR_MCP_BEARER_TOKEN, SCHOLAR_MCP_BEARER_TOKENS_FILE, or OIDC env vars set, which is common in stdio dev rigs but also possible on HTTP), every request resolves to the literal subject "local". Reference that string as the ACL key for un-authenticated local sessions.
Callers authenticate via a bearer token or OIDC (mutually exclusive). See the Authentication guide for setup, mapped multi-subject tokens, OIDC, and troubleshooting.
After copier copy and gh repo create --push:
DOMAIN sentinel comment) in this README and in CLAUDE.md. The GENERATED-ENV-TABLE-* regions are not DOMAIN blocks; the config generator owns them and rewrites them on every run.uv sync --all-extras --all-groups.uv run pre-commit install.uv run pytest -x -q && uv run ruff check --fix . && uv run ruff format . && uv run mypy src/ tests/.CI workflows reference three repository secrets. Configure them via Settings → Secrets and variables → Actions or with gh secret set:
| Secret | Used by | How to generate |
|---|---|---|
RELEASE_TOKEN | release-prepare.yml, release.yml, release-notes.yml, copier-update.yml, renovate.yml, bootstrap.yml | Fine-grained PAT at https://github.com/settings/personal-access-tokens/new with contents: write, pull_requests: write, and administration: write (bootstrap applies the repository rulesets + auto-merge). Must belong to a repository admin: the shipped rulesets grant bypass to the admin role, and the release tag + GitHub release that knope creates after a release pull request merges rely on it (pull requests the token opens also need it so their CI runs). Scoped to this repo. |
CODECOV_TOKEN | ci.yml | https://codecov.io: sign in with GitHub and add the repo. The upload token is on its settings page. |
CLAUDE_CODE_OAUTH_TOKEN | claude.yml, claude-code-review.yml, release-notes.yml | Run claude setup-token locally and paste the result. |
Dependency updates are handled by Renovate (
renovate.yml), which reusesRELEASE_TOKEN. It maintainsuv.lockand auto-merges patch/minor bumps once theCI Successcheck is green;bootstrap.ymlenables auto-merge and applies the repository rulesets (.github/rulesets/) on first push. See Repository Protection for the per-branch posture and bypass model. GitHub Actions are updated in the copier template and arrive viacopier update, not per-repo.
GITHUB_TOKEN is auto-provided; no action needed.
The PR gate (matches CI):
Pre-commit runs a subset of the gate on each commit; see .pre-commit-config.yaml for details, or CLAUDE.md for the full Hard PR Acceptance Gates.
uv sync creates .venv/bin/* scripts with absolute shebangs pointing at the venv Python. If you move the repo after scaffolding (mv /old/path /new/path), uv run pytest fails with ModuleNotFoundError: No module named 'fastmcp' because the stale shebang resolves to a different interpreter than the venv's site-packages.
Fix:
uv run python -m pytest also works as a one-shot workaround (bypasses the stale entry-script shim).
uv.lock refresh after copier updateWhen copier update introduces new dependencies (such as a new extra added to pyproject.toml.jinja), the CI install step runs uv sync --locked, which fails against a stale lockfile. Run uv lock locally and commit the refreshed uv.lock alongside accepting the copier-update PR.
CI installs with --locked (and the review workflow with --frozen) so no job ever rewrites uv.lock in its own workspace: a job that re-locks hides the drift it just repaired, and a dirty workspace breaks any later git checkout in the same job. Lockfile drift then shows up as a red install step with a clear message, not as a silent mutation.
Domain environment variables use the SCHOLAR_MCP_ prefix:
| Variable | Default | Required | Description |
|---|---|---|---|
SCHOLAR_GITHUB_TOKEN | (none) | No | GitHub token used to raise rate limits when fetching standards documents from GitHub. Optional; unauthenticated requests work at reduced limits. |
SCHOLAR_MCP_READ_ONLY | true | No | When true, write-tagged tools (PDF download and conversion cache writes) are hidden. Set false to enable them. |
SCHOLAR_MCP_S2_API_KEY | (none) | No | Semantic Scholar API key. Optional but strongly recommended: unauthenticated requests are limited to ~1 req/s. Request one at https://www.semanticscholar.org/product/api#api-key-form. |
SCHOLAR_MCP_DOCLING_URL | (none) | No | Base URL of a running docling-serve instance for PDF conversion (such as http://localhost:5001). When unset, PDF conversion tools return an error. |
SCHOLAR_MCP_VLM_API_URL | (none) | No | OpenAI-compatible VLM endpoint for formula and figure enrichment during PDF conversion. |
SCHOLAR_MCP_VLM_API_KEY | (none) | No | API key for the VLM endpoint. |
SCHOLAR_MCP_VLM_MODEL | gpt-4o | No | Model name to use with the VLM endpoint. |
SCHOLAR_MCP_CACHE_DIR | /data/scholar-mcp | No | Directory for the SQLite cache database (cache.db) and downloaded PDFs (pdfs/, md/). |
SCHOLAR_MCP_CONTACT_EMAIL | (none) | No | Contact email for the OpenAlex polite pool (improves rate limits). Also enables Unpaywall lookups as a PDF fallback source. |
SCHOLAR_MCP_EPO_CONSUMER_KEY | (none) | No | EPO Open Patent Services consumer key. Optional; patent tools are hidden when unset. Register at https://developers.epo.org/user/register. |
SCHOLAR_MCP_EPO_CONSUMER_SECRET | (none) | No | EPO Open Patent Services consumer secret. Optional; patent tools are hidden when unset. |
SCHOLAR_MCP_GOOGLE_BOOKS_API_KEY | (none) | No | Google Books API key. Optional; book tools work unauthenticated at reduced rate limits. |
SCHOLAR_MCP_JOBS_SOFT_DEADLINE_S | 25.0 | No | Seconds a long-running tool call may run in the foreground before it is promoted to a background job and a job handle is returned instead. |
SCHOLAR_MCP_JOBS_RESULT_TTL_S | 3600.0 | No | Seconds a background-job record (working or finished) is retained for polling before it expires from the store. |
SCHOLAR_MCP_JOBS_MAX_PER_SUBJECT | 256 | No | Maximum live background jobs per calling subject; further promotions are rejected until older records expire. |
Domain-config fields are composed inside src/scholar_mcp/config.py between the CONFIG-FIELDS-START / CONFIG-FIELDS-END sentinels; env reads go through fastmcp_pvl_core.env(_ENV_PREFIX, "SUFFIX", default) so naming stays consistent, and field invariants go in __post_init__ between the CONFIG-VALIDATE-START / CONFIG-VALIDATE-END sentinels. Each field's metadata help and tags generate the table above directly, so keep them accurate and complete.
Scholar-mcp pings Semantic Scholar once on startup and every 7 days
thereafter to keep the configured key from being removed for inactivity
(Semantic Scholar may remove keys unused for 60+ days). If S2 starts
rejecting the key with 403 Forbidden, this shows up in the server logs
as s2_key_forbidden (on real tool calls) or s2_keepalive_key_forbidden
(from the background keepalive); grep for either to confirm a dead key
versus a transient upstream issue.
asyncio.to_thread(). Simpler client code, explicit offloading at the transport boundary.SCHOLAR_MCP_READ_ONLY=false. Safer default for first-run.fastmcp-pvl-core jobs layer, whether the slow part is a docling conversion, an EPO throttle being waited out, or a graph walk making one request per node. A call that beats SCHOLAR_MCP_JOBS_SOFT_DEADLINE_S returns its result directly; a slower one returns a handle to poll with get_job_result. No tool decides in advance whether to go background, so a cache hit needs no special case.scholar-mcp sync-standards, not live at runtime, which avoids paywalled-HTML scraping and keeps tool calls fast.API key optional but recommended: The server works without a Semantic Scholar API key, but unauthenticated requests are limited to ~1 req/s and will hit 429 throttles quickly during multi-step operations like citation graph traversal. Request a free key to get ~10 req/s.
Claude Desktop configuration (claude_desktop_config.json):
Tier 2 bodies (ISO, IEC, IEEE, CC, CEN) are populated from community-curated bulk dumps rather than live-scraped at MCP-server runtime. Run the sync on first install and periodically thereafter:
Schedule via cron, launchd, or a systemd timer. Weekly is sufficient; standards change slowly. First sync can take several minutes; subsequent runs that find no upstream changes exit within seconds.
29 tools, organised by scholarly source type.
| Tool | Description |
|---|---|
search_papers | Full-text search with year, venue, field-of-study, and citation-count filters. Returns up to 100 results with pagination. |
get_paper | Fetch full metadata for a single paper by DOI, S2 ID, arXiv ID, ACM ID, or PubMed ID. |
get_author | Fetch author profile with publications, or search by name. |
| Tool | Description |
|---|---|
get_citations | Forward citations (papers that cite a given paper) with optional filters. |
get_references | Backward references (papers cited by a given paper). |
get_citation_graph | BFS traversal from seed papers, returning nodes + edges up to configurable depth. |
find_bridge_papers | Shortest citation path between two papers. |
| Tool | Description |
|---|---|
recommend_papers | Paper recommendations from 1 to 5 positive examples and optional negative examples. |
generate_citations | Generate BibTeX, CSL-JSON, or RIS citations for up to 100 papers, with automatic entry type inference and optional OpenAlex venue enrichment. |
enrich_paper | Augment Semantic Scholar metadata with OpenAlex fields (affiliations, funders, OA status, concepts). |
| Tool | Description |
|---|---|
search_patents | Search patents across 100+ patent offices via EPO OPS with CPC / applicant / inventor / jurisdiction / date filters. |
get_patent | Fetch bibliographic / claims / description / family / legal / citations sections for a single patent by publication number. Citations include NPL-to-paper resolution via Semantic Scholar. |
get_citing_patents | Find patents that cite a given academic paper (best-effort; EPO OPS citation search coverage is incomplete). |
fetch_patent_pdf | Download a patent PDF via authenticated EPO OPS and optionally convert to Markdown. |
Patent tools are hidden when
SCHOLAR_MCP_EPO_CONSUMER_KEYandSCHOLAR_MCP_EPO_CONSUMER_SECRETare not set.fetch_patent_pdfis also write-tagged and hidden whenSCHOLAR_MCP_READ_ONLY=true.
| Tool | Description |
|---|---|
search_books | Search for books by title, author, ISBN, or keywords via Open Library. Returns up to 50 results. |
get_book | Fetch book metadata by ISBN-10, ISBN-13, Open Library work ID, or edition ID. Optionally download and cache the cover image locally. |
get_book_excerpt | Fetch a book excerpt and description from Google Books by ISBN. Shows preview availability and link. |
recommend_books | Recommend books for a subject via Open Library, sorted by popularity. |
Papers with an ISBN in their
externalIdsare automatically enriched withbook_metadata(publisher, edition, cover URL, subjects, and more) from Open Library when fetched viaget_paper,get_citations,get_references, orget_citation_graph. Book records also includeworldcat_url(when ISBN-13 is present),google_books_url, andsnippetfrom Google Books enrichment. Cover images can be downloaded and cached locally viaget_book.
| Tool | Description |
|---|---|
resolve_standard_identifier | Normalise a messy citation string such as "rfc9000" or "nist 800-53" to canonical form and body. |
search_standards | Search standards by identifier, title, or free text, optionally filtered to one body (NIST, IETF, W3C, ETSI). |
get_standard | Retrieve a standard by canonical or fuzzy identifier, optionally fetching and converting the full text via docling. |
Tier-1 bodies (NIST, IETF, W3C, ETSI) are supported with full metadata and optional full-text conversion. Tier-2 bodies (ISO, IEC, IEEE, CC, CEN/CENELEC) are populated locally via
scholar-mcp sync-standards.
| Tool | Description |
|---|---|
batch_resolve | Resolve up to 100 mixed identifiers (paper DOIs, patent numbers, ISBNs) to full metadata in one call, routing each to the right backend with OpenAlex fallback. |
| Tool | Description |
|---|---|
fetch_paper_pdf | Download PDF for a paper (S2 open-access, then ArXiv/PMC/Unpaywall fallback). |
convert_pdf_to_markdown | Convert a local PDF to Markdown via docling-serve. |
fetch_and_convert | Full pipeline: fetches the PDF with fallback sources, then converts it to Markdown and returns both. |
fetch_pdf_by_url | Download a PDF from any URL and optionally convert to Markdown. |
PDF tools are write-tagged and hidden when
SCHOLAR_MCP_READ_ONLY=true(the default).fetch_patent_pdf(above) and theget_standardfull-text mode cover the patent and standards equivalents.
| Tool | Description |
|---|---|
get_job_result | Retrieve the outcome of a background job by ID. |
Tools answer directly when the work is quick, including on a cache hit. A slower call returns
{"status": "working", "job_id": "...", "poll_with": "get_job_result"}; poll with the tool the handle names until the status is terminal.