The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Docalyze listing page.
An MCP (Model Context Protocol) server that lets AI assistants read and visually analyze local documents — PDFs, Excel spreadsheets, CSV files, Word documents, PowerPoint presentations, and images.
No API keys required. The host AI (GitHub Copilot, Claude, etc.) does all the reasoning directly.
| Format | Extensions | Read | Visual |
|---|---|---|---|
.pdf | ✅ | ✅ | |
| Excel | .xlsx, .xls | ✅ | ✅ |
| CSV / TSV | .csv, .tsv | ✅ | — |
| JSON | .json | ✅ | — |
| Word | .docx | ✅ | ✅ |
| PowerPoint | .pptx | ✅ | ✅ |
| Plain text | .txt, .md | ✅ | — |
| Images | .png, .jpg, .jpeg, .gif, .bmp, .tiff, .webp | — | ✅ |
| Tool | Description |
|---|---|
list_documents | List files under a directory, filtered by glob pattern |
document_info | Get metadata (size, modified date, sheets) for a file |
read_document | Extract text content from a document with pagination |
visual_evaluate_document | Return page images inline so the AI can analyze charts, tables, and diagrams |
Search for docalyze in the MCP server gallery (Extensions sidebar → MCP tab) and click Install.
This requires uv or pipx installed — the npm wrapper calls uvx to run the Python package automatically.
Add to your VS Code mcp.json (or settings.json):
Or, if you installed via pip and want to use the entry point:
The base install handles PDF, Excel, CSV, JSON, and plain text. For additional formats:
The server reads documents from a configurable root directory. Set the DOCUMENTS_ROOT environment variable to change it:
If not set, it defaults to the directory containing the server script.
MIT