Document tooling for AI agents: PDF/text reading, chunking, schema-validated JSON output.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
π‘ Paste the JSON block into your client's configuration file under mcpServers, then restart the application.
An MCP (Model Context Protocol) server that gives an LLM deterministic document tooling for structured-data extraction workflows. The LLM does the reading and extraction reasoning; this server provides the parts that should never be left to a language model: reliable file access, parsing, chunking, JSON Schema validation, and guarded file output.
Built by Koray Nar as a portfolio project for the AI document-automation workflows he is building β the target use case is turning messy PDFs (purchase orders, invoices, reports) into schema-validated JSON. Published as part of a public portfolio. Pairs with Claude Code and Claude Desktop, and with any other MCP client.
An extraction agent fails in predictable places: it hallucinates file contents, loses track of long documents, silently produces JSON that almost matches the target schema, and writes output wherever it likes. This server removes those failure modes:
DOC_EXTRACT_ROOT).--- page N --- markers, so citations of
"page 3" mean page 3.| Tool | Arguments | What it does |
|---|---|---|
list_documents | directory, glob_pattern='*' | List files under a directory inside the allowed root, with size and modified time. Supports recursive globs like **/*.pdf. Patterns must be relative and free of ..; matches resolving outside the root are dropped. |
read_document | path, pages='' | Return a document's text. .pdf via pypdf with --- page N --- markers and optional 1-indexed page selection ('3', '1-5', '1-3,7'); .txt/.md/.json read directly; .csv rendered as an aligned text table. Clear error for unsupported types. |
document_info | path | Metadata without full content: type, size, modified time; page count and PDF metadata for PDFs; line count for text files. |
chunk_document | path, max_chars=4000, overlap=200 | Split a document into ordered overlapping chunks, each with an index, start offset, and (for PDFs) a page hint. |
validate_json | data, json_schema | Validate a JSON string against a JSON Schema (Draft 2020-12). Returns every validation error with a JSON Pointer path via Draft202012Validator.iter_errors. |
save_structured | path, data, format='json'|'csv' | Write extracted data inside the allowed root. CSV expects a JSON array of flat objects. Returns written path, row count, and byte count. |
All path arguments are resolved and refused if they escape the allowed root
(path traversal guard). The glob_pattern argument is confined the same way:
absolute patterns and patterns containing .. are rejected, and any match
that resolves outside the root (for example through a symlink) is silently
dropped from the listing. Guard failures are raised as MCP tool errors, so
the calling model sees the actual reason, not a masked generic error.
Requires Python 3.11+ and uv.
Run standalone (stdio transport):
Add to claude_desktop_config.json:
DOC_EXTRACT_ROOT defaults to the server's working directory if unset. Set it
to the folder your documents live in; nothing outside it can be read or
written.
list_documents(".", "*.pdf") β find the invoices.document_info("invoice.pdf") β check the page count.read_document("invoice.pdf", "1-3") or chunk_document(...) β get text.validate_json(data, json_schema) β fix every reported error, revalidate.save_structured("out/invoice.json", data, "json") β write the result..pdf, .txt, .md,
.csv, .json.The test suite builds a small two-page PDF fixture in-memory (a minimal hand-constructed PDF, no extra dependencies) and covers all six tools, the path-traversal guard, glob-pattern confinement (including symlink escapes), page-range errors, multi-error schema validation, a CSV round-trip, and tool registration plus error propagation through the MCP server object.
MIT Β© 2026 Koray Nar
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/doc-extract-mcp)<a href="https://allmcps.com/mcp/doc-extract-mcp"><img src="https://allmcps.com/api/badge/doc-extract-mcp?style=directory" alt="Doc Extract MCP on AllMCPs" /></a>