The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Preprint Fulltext listing page.
English | 简体中文 | 繁體中文 | 한국어 | Deutsch | Español | Français | Italiano | 日本語
Retrieve the full text of bioRxiv / medRxiv / arXiv preprints as clean, structured, embedding-ready data — from a CLI, a Python library, or an MCP server.
preprint-fulltext turns a DOI (or a search) into structured sections
(abstract / introduction / methods / results / discussion), a single JSON/Markdown
document, or a chunked JSONL/Parquet corpus ready for embeddings and RAG. openRxiv
text-and-data-mining (TDM) compliance is enforced structurally, not left to the user.
"Embedding-ready" means the output is clean, section-aware, token-bounded chunks — ready to feed to your embedding model. Computing embeddings is an optional last step you own; this tool does not bundle an embedding model.
Preprint full text is scattered across incompatible channels: Europe PMC serves JATS
XML for the open-access subset, the openRxiv S3 buckets hold the authoritative
.meca corpus (requester-pays), OpenAlex is a catalog with n-gram-only full-text
search, and the bioRxiv/medRxiv websites render HTML. preprint-fulltext unifies
them behind one canonical data model and one shared JATS parser, so you get the
same structured output no matter where a document came from.
SKILL.md) that need to pull a preprint's
full text or search the literature mid-task.Language models and agents reason far more reliably over a paper's methods and results
than over its abstract alone — most scientific claims, protocols, quantities, and caveats
live in the body. preprint-fulltext gives Claude, Codex, and other agents that body as
clean, section-labeled, provenance- and license-tagged text, which is the substrate for
grounded scientific reasoning and deep research:
Because every Section/Chunk carries its kind (methods / results / …), source, and
license, an agent can cite precisely (which section of which paper/version) and stay
within-license while it reasons. Full text is retrieval, not memorization: the model
grounds its reasoning in the primary source instead of recalling a possibly-stale summary.
get <id> — one preprint's full text as structured JSON or Markdown. bioRxiv/
medRxiv route Europe PMC → S3 (opt-in HTML fallback); arXiv ids route to arXiv's
LaTeXML full text (native HTML → ar5iv). Latest version by default; --version selects one.search / discover — keyword, title, abstract, or author search across Europe PMC,
OpenAlex, and arXiv; topic/category/date discovery.ingest — resumable, incremental bulk ingestion from the openRxiv S3 buckets
into a chunked corpus (JSONL or Parquet) with a sidecar manifest.The MCP server is built in — no extra install and no third-party MCP framework. It's a
small, self-contained JSON-RPC 2.0 stdio server, so preprint-fulltext-mcp works out of the
box with only the core dependencies.
Set a contact email for the Europe PMC / OpenAlex polite pools (recommended), and an OpenAlex API key if you use OpenAlex (required by OpenAlex since 2026-02-13):
get emits a FullText document (JSON) or Markdown (--markdown). search /
discover stream one SearchHit per line (JSONL). ingest writes one Chunk per
line plus a <out>_manifest.jsonl audit/resume sidecar.
1. Read one paper's methods/results as text.
2. Build an embedding-ready corpus on a topic (free, no AWS).
3. Build the complete corpus for a month from S3 (requester-pays).
4. Find papers by author or title, then fetch.
5. Give a coding agent literature access — run preprint-fulltext-mcp and point your
agent at it (see skills/preprint-fulltext/SKILL.md).
Give a coding agent live preprint access. The server exposes four tools —
search_preprints, get_fulltext, get_metadata, resolve — over stdio. (Bulk ingest
is intentionally not a tool: it is long-running and incurs requester-pays cost.)
mcp-name: io.github.genecell/preprint-fulltext
It's a local stdio server, so it works in Claude Code / Cursor / VS Code / Windsurf / Zed / Codex / Cline — but not the claude.ai web app (there, use the Skill instead).
uvx (no install)uv runs the published package on demand — nothing to
pip install or keep on a PATH. Install uv once:
The launch command is uvx --from preprint-fulltext preprint-fulltext-mcp (the --from is
needed because the run command differs from the package name). First launch downloads the
package (~30 s); later launches are cached.
mcpServersOr edit ~/.claude.json (user) / project .mcp.json:
mcpServers (same shape)Cursor: ~/.cursor/mcp.json (global) or .cursor/mcp.json (project). Windsurf:
~/.codeium/windsurf/mcp_config.json. Cline: MCP Servers → Configure. Continue:
~/.continue/config.
servers + type.vscode/mcp.json (workspace) or user settings.json under "mcp":
Or one-shot: code --add-mcp '{"name":"preprint-fulltext","command":"uvx","args":["--from","preprint-fulltext","preprint-fulltext-mcp"]}'
context_servers (different shape)~/.config/zed/settings.json:
~/.codex/config.toml:
Or: codex mcp add preprint-fulltext -- uvx --from preprint-fulltext preprint-fulltext-mcp
If you already pip install preprint-fulltext, the server is on your PATH as
preprint-fulltext-mcp — use "command": "preprint-fulltext-mcp" (no args) in any config
above.
Env vars: set
CONTACT_EMAIL(Europe PMC / OpenAlex polite pools) andOPENALEX_API_KEY(only for OpenAlex search/discover) via the config'senvblock, or in your shell before launching the client. SeeSKILL.mdfor the full agent-facing tool reference.
| Verb | Default source | Notes |
|---|---|---|
get | auto (Europe PMC → S3, or arXiv) | bioRxiv/medRxiv: EPMC (CC/OA subset) → S3 (complete, needs AWS creds), --html opt-in fallback. arXiv ids → arXiv LaTeXML full text (native HTML → ar5iv). |
search | Europe PMC | Real relevance ranking; --source openalex|arxiv. |
discover | OpenAlex | 250M+ works, OA locations, topic/date; --source arxiv. |
ingest | S3 (or Europe PMC) | S3 = complete corpus; Europe PMC = free CC subset. arXiv bulk is out of scope (use arXiv's own S3 LaTeX bucket). |
Via environment variables (prefixed PREPRINT_FULLTEXT_ or the bare names below),
a .env file, or a preprint-fulltext.toml:
| Setting | Default | Purpose |
|---|---|---|
CONTACT_EMAIL | – | Polite-pool identity for Europe PMC / OpenAlex |
OPENALEX_API_KEY | – | Required by OpenAlex since 2026-02-13 |
AWS_REGION | us-east-1 | Region for the requester-pays openRxiv buckets |
PREPRINT_FULLTEXT_CACHE_DIR | ~/.cache/preprint-fulltext | Content-addressed cache |
PREPRINT_FULLTEXT_CHUNK_TOKENS | 512 | Max tokens per chunk |
PREPRINT_FULLTEXT_CHUNK_OVERLAP | 64 | Token overlap within a section |
Corpora are for the operator's own text/data mining under the openRxiv TDM terms.
preprint-fulltext does not re-host or redistribute preprint full text. Every
FullText/Chunk carries its license; the export gate has two modes:
--redistribution): works whose license permits redistribution
pass unchanged; all others are degraded to a link-back stub (metadata + URL,
no body text). Unknown/ambiguous licenses are treated as non-redistributable.Live tests are opt-in (they hit the real public APIs — Europe PMC, arXiv, and the bioRxiv/medRxiv JSON API):
The same live smoke runs in CI on demand (Actions → live-smoke) and weekly, to catch
upstream API drift; the default test workflow stays fully offline.
Agent docs (AGENTS.md, llms.txt, .cursor/rules/…, .github/copilot-instructions.md)
are generated from skills/preprint-fulltext/SKILL.md:
Min Dai — dai@broadinstitute.org (Gord Fishell Lab, Harvard Medical School / Broad Institute). Issues and pull requests welcome at https://github.com/genecell/preprint-fulltext.
BSD-3-Clause (see LICENSE). This covers the software only —
retrieved preprint content remains under its author-selected license.