# doc-extract-mcp [Health: Active]

**Category:** 💻 Developer Tools  
**Repository:** https://github.com/koraynar/doc-extract-mcp  
**GitHub Stars:** 0  
**Views:** 0  
**Installs:** 0  
**Upvotes:** 0  
**Directory Page:** https://allmcps.com/mcp/doc-extract-mcp

## Description
Document tooling for AI agents: PDF/text reading, chunking, schema-validated JSON output.

## Claude Desktop Quick Installation
Remote MCP endpoint (confidence: high). Install path detected from listing signals. Add as a URL/SSE server in your client:

```json
"mcpServers": {
  "doc-extract-mcp": {
    "url": "https://docs.astral.sh/uv/"
  }
}
```

## Documentation & README

# doc-extract-mcp

An MCP (Model Context Protocol) server that gives an LLM deterministic document
tooling for structured-data extraction workflows. The LLM does the reading and
extraction reasoning; this server provides the parts that should never be left
to a language model: reliable file access, parsing, chunking, JSON Schema
validation, and guarded file output.

Built by Koray Nar as a portfolio project for the AI document-automation
workflows he is building — the target use case is turning messy PDFs (purchase
orders, invoices, reports) into schema-validated JSON. Published as part of a
public portfolio. Pairs with Claude Code and Claude Desktop, and with any
other MCP client.

## Why

An extraction agent fails in predictable places: it hallucinates file contents,
loses track of long documents, silently produces JSON that almost matches the
target schema, and writes output wherever it likes. This server removes those
failure modes:

- File access is confined to one allowed root (`DOC_EXTRACT_ROOT`).
- PDF text arrives with explicit `--- page N ---` markers, so citations of
  "page 3" mean page 3.
- Long documents are chunked deterministically with overlap and page hints.
- Extracted JSON is checked against a JSON Schema (Draft 2020-12) and **every**
  error is reported with a JSON Pointer path — not just the first — so the
  model can fix all mistakes in one pass.
- Output is written by the server (JSON or CSV), inside the same root, with a
  verifiable row/byte count.

## Tools

| Tool | Arguments | What it does |
|---|---|---|
| `list_documents` | `directory`, `glob_pattern='*'` | List files under a directory inside the allowed root, with size and modified time. Supports recursive globs like `**/*.pdf`. Patterns must be relative and free of `..`; matches resolving outside the root are dropped. |
| `read_document` | `path`, `pages=''` | Return a document's text. `.pdf` via pypdf with `--- page N ---` markers and optional 1-indexed page selection (`'3'`, `'1-5'`, `'1-3,7'`); `.txt`/`.md`/`.json` read directly; `.csv` rendered as an aligned text table. Clear error for unsupported types. |
| `document_info` | `path` | Metadata without full content: type, size, modified time; page count and PDF metadata for PDFs; line count for text files. |
| `chunk_document` | `path`, `max_chars=4000`, `overlap=200` | Split a document into ordered overlapping chunks, each with an index, start offset, and (for PDFs) a page hint. |
| `validate_json` | `data`, `json_schema` | Validate a JSON string against a JSON Schema (Draft 2020-12). Returns every validation error with a JSON Pointer path via `Draft202012Validator.iter_errors`. |
| `save_structured` | `path`, `data`, `format='json'\|'csv'` | Write extracted data inside the allowed root. CSV expects a JSON array of flat objects. Returns written path, row count, and byte count. |

All path arguments are resolved and refused if they escape the allowed root
(path traversal guard). The `glob_pattern` argument is confined the same way:
absolute patterns and patterns containing `..` are rejected, and any match
that resolves outside the root (for example through a symlink) is silently
dropped from the listing. Guard failures are raised as MCP tool errors, so
the calling model sees the actual reason, not a masked generic error.

## Quickstart

Requires Python 3.11+ and [uv](https://docs.astral.sh/uv/).

```bash
git clone https://github.com/koraynar/doc-extract-mcp.git
cd doc-extract-mcp
uv venv
uv pip install -e .
```

Run standalone (stdio transport):

```bash
DOC_EXTRACT_ROOT=/path/to/your/documents uv run doc-extract-mcp
```

### Claude Code

```bash
claude mcp add doc-extract --env DOC_EXTRACT_ROOT=/path/to/your/documents \
  -- uv run --directory /absolute/path/to/doc-extract-mcp doc-extract-mcp
```

### Claude Desktop

Add to `claude_desktop_config.json`:

```json
{
  "mcpServers": {
    "doc-extract": {
      "command": "uv",
      "args": [
        "run",
        "--directory",
        "/absolute/path/to/doc-extract-mcp",
        "doc-extract-mcp"
      ],
      "env": {
        "DOC_EXTRACT_ROOT": "/path/to/your/documents"
      }
    }
  }
}
```

`DOC_EXTRACT_ROOT` defaults to the server's working directory if unset. Set it
to the folder your documents live in; nothing outside it can be read or
written.

### Typical workflow

1. `list_documents(".", "*.pdf")` — find the invoices.
2. `document_info("invoice.pdf")` — check the page count.
3. `read_document("invoice.pdf", "1-3")` or `chunk_document(...)` — get text.
4. The LLM extracts fields into JSON.
5. `validate_json(data, json_schema)` — fix every reported error, revalidate.
6. `save_structured("out/invoice.json", data, "json")` — write the result.

## Limitations (honest ones)

- **Text-based PDFs only.** Extraction uses pypdf; scanned/image-only PDFs
  yield empty text. There is no OCR.
- **Extraction quality varies** with how the PDF was produced. Complex layouts
  (multi-column, heavy tables) may come out with imperfect reading order —
  that is a pypdf characteristic this server inherits.
- **No .docx / .xlsx support.** Supported types are `.pdf`, `.txt`, `.md`,
  `.csv`, `.json`.
- **The server does no extraction reasoning.** It will not find your invoice
  total; it makes sure the model that does is working from real text and that
  the result matches your schema.
- This is a working tool, built for the AI-automation work I'm building up and
  published as part of my portfolio — it is new and has no production mileage
  yet. It has tests and a path-confinement guard, but it has not been hardened
  beyond that — review before pointing it at sensitive directories.

## Development

```bash
uv venv
uv pip install -e '.[dev]'
uv run pytest
```

The test suite builds a small two-page PDF fixture in-memory (a minimal
hand-constructed PDF, no extra dependencies) and covers all six tools, the
path-traversal guard, glob-pattern confinement (including symlink escapes),
page-range errors, multi-error schema validation, a CSV round-trip, and tool
registration plus error propagation through the MCP server object.

## License

MIT © 2026 Koray Nar

