The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Techtenstein Pdf listing page.
MCP server that gives your Claude, Cline, or Cursor session the ability to extract text, tables, and metadata from any PDF URL — including scanned PDFs via OCR. Powered by the Techtenstein PDF Extract API.
pdf_extract_text(pdf_url, ocr=False) — Extract all text from a PDF as clean plain textpdf_extract_tables(pdf_url) — Extract all tables as structured row arrayspdf_metadata(pdf_url) — Get title, author, page count, creation date, encryption statusAdd to ~/Library/Application Support/Claude/claude_desktop_config.json:
Restart Claude Desktop. pdf_extract_text, pdf_extract_tables, and pdf_metadata will appear as available tools.
Cline auto-detects MCP servers from your Claude Desktop config. Same setup as above works.
Add to ~/.cursor/mcp.json:
Free tier (50 extractions/day, no card): https://apis.techtenstein.com
Paid tiers start at $5/month for 2,000 extractions.
Once installed, ask Claude:
"Extract the tables from this earnings report PDF: https://example.com/q4.pdf"
Claude will call pdf_extract_tables and return a clean structured view of every table on the page.
Or for scanned documents:
"This PDF is a scanned invoice. Extract the text: https://example.com/invoice.pdf"
Claude will call pdf_extract_text(pdf_url, ocr=True) and read the image-based text via OCR.
MIT