The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the PyEuropePMC listing page.
PyEuropePMC is a Python client for Europe PMC. It searches the literature, downloads open-access full text, and parses JATS XML into metadata, plain text and structured sections. It also ships a command-line tool and an MCP server for AI agents.
QueryBuilder, pagination, and results as JSON, XML or Dublin Core. Search guidesemanticscholar extra and CORE an API key. Multi-source searchrdf_map.yml that ships with the package; pass RDFMapper(config_path=...) or PipelineConfig(rdf_config_path=...) to use your own. Data models and RDFPyEuropePMC supports Python 3.10 to 3.13. The core install covers search, full-text download, XML parsing, RDF, the command line and the MCP server. The extras add:
| Extra | Installs | Adds |
|---|---|---|
analytics | pandas, numpy | to_dataframe, citation_statistics, quality_metrics, remove_duplicates and CSV export |
visualization | matplotlib, seaborn, pandas, numpy | The plot_* functions and create_summary_dashboard |
export | pandas, tabulate, xlsxwriter | Excel and Markdown-table export (pyeuropepmc.utils.export) |
semanticscholar | semanticscholar | SemanticScholarClient and the semantic_scholar search source |
enrichment | semanticscholar | Semantic Scholar data in PaperEnricher |
bibliography | bibtexparser | BibTeX parsing, validation and RIS/CSL conversion, including the bib_* MCP tools |
zotero | pyzotero | The Zotero client |
agentic | openai, langchain, langchain-openai, langgraph, jinja2 | LLM analysis, the LLM MCP tools and pyeuropepmc claim |
ui | flask, tornado | The web UI (pyeuropepmc claim serve) |
signing | cryptography | Signed search logs for systematic reviews |
benchmark | huggingface-hub | Downloading the published benchmark datasets (pyeuropepmc benchmark download) |
rdf | nothing | Kept so existing installs keep working; the core rdflib writes JSON-LD itself |
standard | jupyterlab, notebook, ipykernel, ipython, ipywidgets, matplotlib, seaborn, pandas, numpy, tabulate, xlsxwriter, requests-cache, rich | Jupyter plus the analytics, plotting and export packages |
all | all of the above | Every optional feature |
Upgrading from 1.x? The migration guide lists what moved into extras.
section_type is front for the title and abstract, then body, back or appendix. Each block has a type such as paragraph, list, table or figure. Tables and figures carry label and caption, and tables also carry rows. The XML parsing guide covers the other extractors.
PyEuropePMC parses every XML document it reads with defusedxml: full-text articles, Europe PMC search results, arXiv and PubMed responses, and local files. A DOCTYPE declaration is accepted, but a document that declares entities, internal or external, is refused rather than expanded; FullTextXMLParser raises ParsingError. lxml is not used and does not need to be installed.
| Command | What it does |
|---|---|
pyeuropepmc unified_search | search several sources with deduplication, or compare-sources to see each source's results |
pyeuropepmc normalize | Turn JATS XML into clean text (text), sections (sections) or BioC JSON (bioc); classify a heading; batch a directory |
pyeuropepmc benchmark | Score and profile the XML parser on a local folder or a published dataset |
pyeuropepmc claim | Check the claims in a text against Europe PMC literature with LLM agents (needs the agentic extra and an API key) |
unified_search search queries Europe PMC, PubMed and arXiv unless you pass --source. Add --help to any command for its options.
pyeuropepmc-mcp serves 24 tools over the Model Context Protocol: multi-source search, paper details and citations, citation-graph walking, ClinicalTrials.gov search, a local full-text index, figure extraction, bibliography conversion and LLM-powered analysis. The server is part of the core install. pip install "pyeuropepmc[all]" enables every tool; otherwise the bib_* tools need bibliography, and the LLM tools need agentic plus an OpenAI-compatible API key (see Configuration below).
For Claude Desktop and other clients that start the server themselves:
Without an install, "command": "uvx" with "args": ["pyeuropepmc", "mcp"] does the same; this is what the MCP Registry entry tells clients to run.
The server has no authentication of its own, so keep the HTTP transport on 127.0.0.1 or put an authenticating proxy in front of it. The MCP server guide lists every tool. The server is listed in the MCP Registry as io.github.JonasHeinickeBio/pyeuropepmc.
Europe PMC needs no API key. Other services read these environment variables. A value passed in code, such as FullTextClient(email=...) or EnrichmentConfig(unpaywall_email=...), takes precedence.
| Variable | Used for |
|---|---|
UNPAYWALL_EMAIL, CROSSREF_EMAIL | Contact e-mail for Unpaywall and Crossref. FullTextClient needs one of them for its Unpaywall fallback. |
OPENALEX_EMAIL, DATACITE_EMAIL, ROR_EMAIL, ROR_CLIENT_ID | The other enrichment sources |
SEMANTIC_SCHOLAR_API_KEY | Semantic Scholar, with higher rate limits |
CORE_API_KEY | The core search source |
OPENAI_API_KEY, OPENAI_BASE_URL, OPENAI_MODEL | LLM features, with any OpenAI-compatible API; the model defaults to gpt-4o-mini |
ZOTERO_API_KEY, ZOTERO_LIBRARY_ID, ZOTERO_LOCAL | The Zotero client |
PYEUROPEPMC_MCP_TRANSPORT, PYEUROPEPMC_MCP_HOST, PYEUROPEPMC_MCP_PORT, PYEUROPEPMC_MCP_LOG_LEVEL | Defaults for the pyeuropepmc-mcp options |
The pyeuropepmc command also loads the first .env file it finds in the current directory, one of its parents, or ~/.config/pyeuropepmc/. Variables that are already set keep their values.
The guides are in docs/. Start with installation and the quick start, or go straight to the API reference, caching, the example scripts and the changelog.
Parser benchmark. On the 55 JATS articles in benchmark_xmls/xml, the parser's mean composite quality score is 0.998 (measured on 2026-09-15). The benchmarking guide explains the metrics. To reproduce the score from a source checkout, or to score a published dataset:
The weekly benchmark workflow times the API clients and opens a pull request that refreshes the section below.
Last updated: 2026-09-14
| Metric | Value |
|---|---|
| Benchmarked methods | 10 |
| Total requests | 224 |
| Mean call time | 0.871s |
| Success rate | 100.0% |
Generated: 2026-09-14 07:45:39
| Method | Mean Time | Std Dev | Mean Memory | Cache | Requests | Errors |
|---|---|---|---|---|---|---|
| get_article_details | 1.707s | 0.212s | 0.0MB | ❌ | 31 | 0 |
| p50: 1.620s · p95: 2.235s · ops: 0.59/s · runs: 30 | ||||||
| get_citations | 1.789s | 0.318s | 0.0MB | ❌ | 31 | 0 |
| p50: 1.666s · p95: 2.851s · ops: 0.56/s · runs: 30 |
| Method | Mean Time | Std Dev | Mean Memory | Cache | Requests | Errors |
|---|---|---|---|---|---|---|
| get_article_details | <1ms | <1ms | 0.0MB | ✅ | 1 | 0 |
| p50: 17µs · p95: 29µs · ops: 53390.19/s · runs: 30 | ||||||
| get_citations | <1ms | <1ms | 0.0MB | ✅ | 1 | 0 |
| p50: 19µs · p95: 30µs · ops: 48437.33/s · runs: 30 |
| Method | Mean Time | Std Dev | Mean Memory | Cache | Requests | Errors |
|---|---|---|---|---|---|---|
| search | 2.074s | 0.262s | 0.1MB | ❌ | 31 | 0 |
| p50: 1.964s · p95: 2.673s · ops: 0.48/s · runs: 30 | ||||||
| get_hit_count | 2.085s | 0.562s | 0.0MB | ❌ | 31 | 0 |
| p50: 1.933s · p95: 3.668s · ops: 0.48/s · runs: 30 |
| Method | Mean Time | Std Dev | Mean Memory | Cache | Requests | Errors |
|---|---|---|---|---|---|---|
| search | <1ms | <1ms | 0.0MB | ✅ | 1 | 0 |
| p50: 109µs · p95: 134µs · ops: 8795.90/s · runs: 30 | ||||||
| get_hit_count | <1ms | <1ms | 0.0MB | ✅ | 1 | 0 |
| p50: 111µs · p95: 137µs · ops: 8603.62/s · runs: 30 |
| Method | Mean Time | Std Dev | Mean Memory | Cache | Requests | Errors |
|---|---|---|---|---|---|---|
| check_fulltext_availability | 1.050s | 0.148s | 0.1MB | ❌ | 93 | 0 |
| p50: 0.998s · p95: 1.388s · ops: 0.95/s · runs: 30 |
| Method | Mean Time | Std Dev | Mean Memory | Cache | Requests | Errors |
|---|---|---|---|---|---|---|
| check_fulltext_availability | <1ms | <1ms | 0.0MB | ✅ | 3 | 0 |
| p50: 20µs · p95: 21µs · ops: 48496.45/s · runs: 30 |
| Method | No-Cache Mean | Cached Mean | Speedup (no/cache) |
|---|---|---|---|
| get_article_details | 1.707s | <1ms | >17074.4x |
| get_citations | 1.789s | <1ms | >17888.5x |
| Method | No-Cache Mean | Cached Mean | Speedup (no/cache) |
|---|---|---|---|
| check_fulltext_availability | 1.050s | <1ms | >10504.7x |
| Method | No-Cache Mean | Cached Mean | Speedup (no/cache) |
|---|---|---|---|
| get_hit_count | 2.085s | <1ms | 17941.89x |
| search | 2.074s | <1ms | 18245.92x |
Notes: Means are computed over measured iterations; '-' indicates missing data. Values like '<1ms' indicate very fast cached responses. Speedups shown as lower-bounds when cached times are too small to measure precisely.
Run the modular benchmark locally and regenerate these artifacts:
MODULAR_BENCHMARK_RESULTS.jsonContributions are welcome. The development guide covers setup, testing, code quality and the release process. Report bugs and suggest features in the issue tracker.
If PyEuropePMC supports your research, please cite it and give the version you used (pip show pyeuropepmc prints it):
The literature itself comes from Europe PMC; please acknowledge it as your data source.
PyEuropePMC is released under the MIT License. Articles you retrieve keep their own licences, which FullTextXMLParser.extract_license() reads from the XML.