Upload and process PDF documents from URLs or local files. Supports OCR processing, hierarchical content extraction, and intelligent document analysis. Returns a unique doc_id for subsequent operations. Processing typically takes 0-3 minutes depending on document size (estimate: 2 seconds per page). Supports files up to 100MB.
Primary document retrieval tool. After orienting with get_folder_structure() (when available), use this for all document-related questions. The bare call returns root-level sub-folders and documents; pass folder_id to drill into a sub-folder level by level. Use sort="relevance" + query for semantic ranking. Do NOT jump to search_documents() first — it is an escalation path, only after browse_documents(sort="relevance") has failed.
ESCALATION tool — never the first step. Use only after `browse_documents(sort="relevance", query=...)` missed a document you strongly believe exists by precise term, acronym, or file-name fragment. Query must be keywords only — see the `query` schema describe. Each result has a `score` (6-10, higher is more relevant). On result: single best match → read it; several equally good → ask user; none → fall back to `browse_documents(sort="relevance")`.
Orientation step: show the folder hierarchy as a tree (like `tree -d`). **Call this before browse_documents()** to plan targeted retrieval — folder names reveal content domains (e.g. "Research", "Finance") and return folder IDs you pass to browse_documents(folder_id=…). Wait for this result before issuing browse_documents() or search_documents(). For folder_id="root" (default), returns the entire folder tree; for a specific folder_id, returns that subtree only.
Check a document's processing status and metadata. `status` is one of "pending", "queued", "processing", "completed", or "failed" — call this before `get_document_structure()` or `get_page_content()` to confirm the document is ready.
Extract a document's hierarchical outline (headers, sections, page references). REQUIRED for documents over 20 pages — call this first to locate relevant sections, then pass their page numbers to `get_page_content()`. Use the `part` parameter to iterate large outlines until `pagination.has_more` is false.
Extract page content from a processed document. Use tight, targeted page ranges — never the whole document at once. For documents over 20 pages, call `get_document_structure()` first to pick relevant sections. Embedded image paths in the response feed into `get_document_image()`.
Retrieve an image from a document — pass an `image_path` from a `get_page_content()` response: either an embedded image (from a content block) or a full rendered page (from the `page_images` array, useful when a page is scanned or its meaning depends on visual layout). Returns the image as an MCP image content block (base64 data + MIME type). Do NOT repeat the raw base64 value in your text response.
Permanently delete documents and all associated data. Only invoke when the user explicitly names the documents AND confirms deletion. Returns `results` — one entry per requested document: `{ doc_name, status: "deleted" | "not_found" | "failed", error? }`. Inspect each entry for per-document failures. This action is irreversible.