Read PDF internals: text, structure, fonts, signatures and tags, with what was not read declared.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
One-click editor setup isnβt available for this listing yet β we donβt have a confirmed install command, and weβd rather show nothing than point your editor at the wrong package or host. Follow the projectβs own setup instructions, linked above.
English | ζ₯ζ¬θͺ
An MCP (Model Context Protocol) server specialized in deciphering PDF internal structures.
While typical PDF MCP servers are thin wrappers for text extraction, this project focuses on reading and analyzing the internal structure of PDF documents. Pair it with pdf-spec-mcp for specification-aware structural analysis and validation.
| Server | Role |
|---|---|
| pdf-spec-mcp | PDF specification knowledge (ISO 32000, PDF/A, PDF/UA) |
| pdf-reader-mcp (this) | Read and inspect PDF internal structure β what is in a PDF |
| pdf-verify-mcp | Authenticity verification β whether it is genuine: cryptographic signature verification, tamper detection, PAdES level, PDF/A validation, encrypted-PDF decryption |
pdf-reader-mcp inspects signature structure (inspect_signatures); for cryptographic signature verification, trust/revocation evaluation, and PDF/A conformance validation, use pdf-verify-mcp.
19 tools organized into three tiers:
| Tool | Description |
|---|---|
get_page_count | Lightweight page count retrieval |
get_metadata | Full metadata extraction (title, author, PDF version...) |
read_text | Text extraction with Y-coordinate reading order (opt-in split_columns: 2 | 3 for untagged multi-column PDFs, compact_whitespace for Japanese forms). Resolves /ActualText replacements (Β§14.9.4, both the structure-element and the Span marked-content path). Reports text extractability per page (Β§9.10.1) so an empty result is never mistaken for an empty page. For logical order in tagged PDFs, prefer extract_structured_text |
search_text | Full-text search with surrounding context. Searches the same text read_text returns, /ActualText included, so a hit means what a reader sees (a note names any page whose marked content could not be aligned) |
read_images | Embedded image XObjects as PNG or JPEG files, returned as MCP image content blocks so a vision model can read them. max_width / max_height downscale by area average; the response has a byte budget and names anything it leaves out |
read_url | Fetch a remote PDF and extract its text β nothing more. The bytes are not saved; to use the other 18 tools on a URL's PDF, download it first and pass the local path (see "read_url and the read-only boundary") |
render_page | Rasterise pages to PNG/JPEG via PDFium-WASM (optional dependency @hyzyla/pdfium). The next step when text extractability says no_text_layer / not_extractable β draws the whole page, vector art and forms included |
summarize | Quick overview report (metadata + text + image count + per-document text extractability) |
| Tool | Description |
|---|---|
inspect_structure | Object tree and catalog dictionary analysis |
inspect_tags | Tagged PDF structure tree visualization |
inspect_fonts | Font inventory (embedded/subset/type detection) |
inspect_annotations | Annotation listing (categorized by subtype) |
inspect_signatures | Digital signature field structure analysis |
extract_structured_text | Tagged PDF text in logical content order (ISO 32000-2 Β§14.8.2.5), each piece labelled with its structure type (H1 / P / Table β¦). Resolves /ActualText, separates /Alt and list labels, keeps page-spanning elements whole. include_bbox: true adds where each element is drawn β one rectangle per page, in the form add_annotation takes |
extract_tables | Tagged PDF <Table> subtree β Markdown table (preserves columns). A table continuing across a page break is ONE table (pages array) |
locate_objects | Object number β page and rectangle, in the coordinate form pdf-writer-mcp add_annotation takes. Bridges pdf-verify-mcp verify_integrity's "which objects changed" to "where they are". Each location names its basis: an annotation's own /Rect is exact, a content stream can only say "the whole page" |
| Tool | Description |
|---|---|
validate_tagged | Deprecated β PDF/UA pass/fail belongs to pdf-verify-mcp validate_conformance (flavour: "pdfua-1"). Kept until the next major |
validate_metadata | Deprecated β same migration path as above. Kept until the next major |
compare_structure | Structural diff between two PDFs (properties + fonts) |
read_url returns text, and only text. This is a decision, now stated rather than implied:
the fetched bytes are discarded after extraction, because saving them would make a reader
tool write to the file system, and every tool of this server is read-only
(readOnlyHint: true β all 19 of them).
To run search_text, inspect_structure, extract_tables, render_page or anything else
against a PDF that lives at a URL, download the file first β with whatever fetch capability
the calling environment has β and pass the local path. Fetching is the caller's
responsibility, deliberately: an agent environment always has a way to download a file, and a
reader that also writes files has stopped being a pure observer.
read_url remains the right tool for the one-shot question: what does the document at this
URL say?
summarize reporting hasText: false used to be a dead end: nothing in this server could
read the document any further. render_page closes that β it rasterises pages to PNG or JPEG
and returns them as MCP image content blocks, so a vision model can read a scan, a diagram, a
filled form, or handwriting.
pages is required: rendering is the most expensive operation here, and "all pages" of a
500-page scan should be a decision, not a default. The same 4 MB response budget as
read_images applies, with omissions named.
Rendering runs on PDFium compiled to WebAssembly (@hyzyla/pdfium, an optional
dependency). A WASM binary is the same bytes on every platform, so the published package still
behaves identically wherever npx runs it β the reason native addons are not used here.
Without the dependency installed, render_page reports what to install and every other tool
works normally. PDFium (BSD-3-Clause) is a different engine from the pdf.js this server reads
text with; the tool description says so, because a rendering difference between engines must
not be attributed to the file.
Measured before choosing this: pdf.js +
@napi-rs/canvas(1.0.7 and 0.1.80) segfaults the whole process on pages that draw images β exactly the pages this tool exists for β and renders blank pages whenstandardFontDataUrlis not configured.
read_images used to base64 imgData.data β pdfjs's decoded pixels. An 8Γ8 RGB image was
192 bytes with no PNG or JPEG signature anywhere in it, so the result could not be opened by
any viewer and could not be read by a vision model, which is the reason to extract an image in
the first place.
Images are now encoded (PNG by default, lossless; format: "jpeg" with quality when smaller
matters) and returned as MCP image content blocks, with the metadata alongside in a text
block. Both encoders are written out here β no native addon, no per-platform binary.
The response is bounded at 4 MB of encoded image data. A 200 dpi A4 scan is ~11.6 MB of pixels on its own, so images past the budget are named with the reason rather than dropped:
read_images returns the image XObjects a page draws. It is not a picture of the page β vector
drawings and text are not covered by it.
read_text used to answer with text or with nothing, and nothing meant three different things.
ISO 32000-2 Β§9.10.1 separates them, so this server does too. Every text-returning tool β
read_text, read_url, search_text, extract_structured_text, summarize β reports, per
page:
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/pdf-reader-pdf-agent-stack)<a href="https://allmcps.com/mcp/pdf-reader-pdf-agent-stack"><img src="https://allmcps.com/api/badge/pdf-reader-pdf-agent-stack?style=directory" alt="PDF Reader (PDF Agent Stack) on AllMCPs" /></a>