The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the MCP Local Rag listing page.
English | 简体中文 | Deutsch | Español | Português (Brasil) | Français
Search private documents from an MCP client or the terminal without sending them to an embedding API.
mcp-local-rag indexes PDF, DOCX, Markdown, and text files on your machine. Search combines semantic similarity with keyword matching, so queries can match both intent and exact technical terms such as API names, class names, and error codes.
No API key, Docker, Python, or external database is required.
Set BASE_DIR to that directory. It is also the security boundary for file operations. Replace
/absolute/path/to/your/documents below with the directory's absolute path.
mcp-local-rag uses the standard MCP protocol over a local stdio server, so it works with AI coding tools and other MCP hosts that support local MCP servers.
Use one of the examples below, or register npx -y mcp-local-rag and set BASE_DIR using your
client's MCP configuration format.
For Claude Code: Run this command:
For Codex: Add to ~/.codex/config.toml:
For OpenCode: Add to ~/.config/opencode/opencode.json (or opencode.jsonc):
For Cursor: Add to ~/.cursor/mcp.json:
Restart the client, then ask it to build the index:
The first sync downloads the default embedding model (about 90 MB) and may take 1–2 minutes before ingestion starts. Later runs use the local cache.
Once the sync completes:
To use the CLI without an MCP client:
The CLI uses the current directory as its document root by default. Run both commands from the
same directory so they use the same default index, or set BASE_DIR and DB_PATH explicitly.
Some document sets cannot be sent to a hosted embedding service because of confidentiality or organizational policy. Keeping the index local makes them searchable without adding a per-query API cost.
Semantic search alone can miss exact identifiers that matter in technical documentation. Keyword reranking keeps those terms visible without giving up natural-language retrieval.
| Input | How to ingest |
|---|---|
| PDF, DOCX, TXT, Markdown | File ingestion or directory sync |
| HTML already fetched by the client | ingest_data; cleaned with Readability and converted to Markdown |
| Plain text or Markdown held in memory | ingest_data with a stable source identifier |
HTML fetching is not built into the server. An MCP client can fetch a page and pass its HTML to
ingest_data.
Excel, PowerPoint, standalone images, and source-code file extensions are not supported by file ingestion. PDFs can optionally use a local vision model to describe figures, but this is not OCR or image search.
| Tool | Purpose |
|---|---|
sync_start | Reconcile the index with all configured roots or one path |
sync_status | Poll a running sync job |
ingest_file | Ingest or replace one file |
ingest_data | Ingest text, Markdown, or HTML already held by the client |
query_documents | Search with semantic matching and keyword boost |
read_chunk_neighbors | Read surrounding chunks from a search result |
list_files | Show supported files and their ingestion state |
delete_file | Delete an indexed file or an ingest_data item |
status | Show index and search status |
sync_start ingests new and changed files, skips byte-identical files, and removes index entries
for files that no longer exist:
The tool returns a jobId immediately. Clients should poll sync_status until its state becomes
succeeded or failed. A changed PDF keeps the visual profile it was indexed with; sync_start
cannot change it. Set STORE_IMAGES=true in the MCP server environment to store supported PDF and
DOCX images for new or changed files selected by sync; unchanged files remain skipped.
Only one sync job is retained by the server process. A newer job replaces a finished record, and restarting the server discards it.
ingest_file accepts PDF, DOCX, TXT, and Markdown. MCP file paths must be absolute and must stay
inside a configured document root:
Re-ingesting the same path replaces its existing chunks.
Results contain the text, source path, title, chunk index, relevance score, and any images stored
on that chunk. MCP returns each image as an image content block paired with its result identity;
CLI query includes an images array of { imageIndex, mimeType, data } on every result. Pass the
chunkIndex and either filePath or source from a result to read_chunk_neighbors when the
answer needs more context:
Both query_documents and list_files accept an optional absolute scope path prefix, or a
list of prefixes. A prefix matches the exact path and its descendants.
Use ingest_data after the MCP client fetches a page:
The server extracts the main article, converts it to Markdown, and stores it under the supplied source identifier. Reusing the same source updates the existing content.
Respect the source site's terms and copyright when indexing external content.
Visual mode adds a generated caption for figure-heavy PDF pages. It is opt-in and does not load a vision model during normal ingestion.
Image storage is independent of visual captions. Set STORE_IMAGES=true for the MCP server, or
pass --images to CLI ingestion and sync:
PDF storage uses detected figure/table regions. DOCX storage includes only PNG/JPEG images that
the existing Mammoth conversion emits as <img>; charts, SmartArt, and shapes are not separately
rendered. Stored images follow their surrounding text into the final semantic chunk and do not
alter ranking, scores, or result count.
visual / --visual | STORE_IMAGES / --images | PDF behavior |
|---|---|---|
| false | false | Text only; no visual captions or returned images. |
| true | false | Generated captions become searchable text; no images are stored or returned. |
| true | true | Generated captions become searchable text, and images from matched chunks are returned inline. |
| false | true | Images are attached to nearby retained PDF text and returned inline for matched chunks; the VLM is not imported, loaded, or run. |
| Profile | Model cache | Use case |
|---|---|---|
fast (default) | about 250 MB | Lightweight visual indexing |
quality | about 1.7 GB | Figures containing labels, annotations, or other in-image text |
Select the larger model with visualQuality: "quality" over MCP or
--visual-quality quality over CLI. Measured CPU inference was about three times as slow as
fast, though results depend on hardware and model updates.
quality CaptionsFrom 0.18.4 quality runs Qwen3.5-2B; earlier versions ran Qwen2.5-VL-3B. Captions already indexed
keep the wording the old model produced, and sync will not redo them, so re-ingest the files you
want refreshed:
Add --images if the file was ingested with it, because a run without it replaces the stored
images. The old model stays on disk. Once nothing else uses it, delete
onnx-community/Qwen2.5-VL-3B-Instruct-ONNX/ from the model cache directory — <cache-dir>, which
defaults to ./models/.
The profile a PDF was indexed with is recorded, and sync reuses it: a PDF indexed with fast or
quality is re-ingested with that same profile, and a PDF with no recorded profile is ingested as
text.
--visual overrides recorded profiles, so it also captions PDFs that were indexed as text.
Changing a profile re-ingests the PDF even when the file itself has not changed; running the same
command again does nothing and loads no model. Image settings are never recorded, so --images
and STORE_IMAGES never cause a re-ingest.
To turn captions off for a path, run ingest on it: a successful normal ingest clears the
recorded profile. To retry a page whose captioning failed, run ingest <path> --visual --visual-quality <profile> with the profile you want — a plain ingest clears it instead. If a
PDF's indexed rows disagree about the profile, sync stops before changing anything and names the
file; re-run it with --visual to settle the profile.
Captions are auxiliary text, not faithful transcriptions. Treat retrieved captions and document text as untrusted input rather than instructions.
At high limits, matched chunks and their attachments can approach the model/client context ceiling; choose the query limit with the calling model's available context in mind.
The CLI uses the same parser, embedder, and vector store without an MCP client:
Global options such as --db-path, --cache-dir, and --model-name go before the subcommand.
Subcommand options go after it:
Run npx mcp-local-rag --help for the complete command reference.
The CLI does not read MCP client configuration. Set the same environment variables or flags if
both interfaces should share an index. In particular, MODEL_NAME and the CLI --model-name
must match for a shared database.
Keyword boost is enabled by default. Relevance-gap grouping and the distance and file filters are optional controls for corpora that need tighter result selection.
| Variable | Default | Description |
|---|---|---|
RAG_HYBRID_WEIGHT | 0.6 | Keyword boost factor (0.0–1.0). 0 disables keyword reranking; 1 applies the maximum boost. |
RAG_GROUPING | (not set) | similar keeps the first relevance group; related keeps up to two, using significant vector-distance gaps as boundaries. |
RAG_MAX_DISTANCE | (not set) | Filter out low-relevance results (e.g., 0.5). |
RAG_MAX_FILES | (not set) | Limit results to top N files (e.g., 1 for single best file). |
For API specifications and other documents containing many identifiers, a stronger keyword weight can improve exact-term ranking:
0.7: slightly stronger exact-term reranking than the default1.0: maximum keyword boostDuring ingestion:
During search:
Agent Skills provide query and ingestion guidance for AI assistants:
Installed skills cover query formulation, result refinement, and HTML ingestion. Ask the assistant to use the mcp-local-rag skill explicitly if it does not activate automatically.
The MCP server reads environment variables. The CLI accepts the listed global environment
variables and flags; image storage on CLI ingestion and sync is enabled only with --images.
| Environment Variable | CLI Flag | Default | Description |
|---|---|---|---|
BASE_DIR | --base-dir | Current directory | One document root; the CLI flag is repeatable on ingest, list, and sync |
BASE_DIRS | N/A | (unset) | JSON array of document roots; takes precedence over BASE_DIR |
DB_PATH | --db-path | ./lancedb/ | Vector database location |
CACHE_DIR | --cache-dir | ./models/ | Model cache directory |
MODEL_NAME | --model-name | Xenova/all-MiniLM-L6-v2 | Hugging Face embedding model |
MAX_FILE_SIZE | --max-file-size | 104857600 (100MB) | Maximum file size in bytes |
CHUNK_MIN_LENGTH | --chunk-min-length | 50 | Minimum chunk length in characters (1–10000) |
STORE_IMAGES | N/A | false | MCP server only: store supported PDF/DOCX images and return them with matched chunks. CLI uses --images. |
RAG_DEVICE | N/A | cpu | ONNX Runtime execution device |
RAG_DTYPE | N/A | fp32 | Embedding dtype passed to the selected model |
BASE_DIR and BASE_DIRS)mcp-local-rag only allows file operations inside configured roots. For multiple roots,
BASE_DIRS must be a JSON array of non-empty paths:
Root configuration is resolved in this order:
--base-dir <path> flags (repeatable on ingest, list, and sync)BASE_DIRSBASE_DIREach source replaces the lower-priority source rather than merging with it. Invalid BASE_DIRS
configuration fails instead of falling back to BASE_DIR or the current directory. status
remains available in MCP so the client can report the configuration error.
DB_PATH and CACHE_DIR are relative to the process working directory by default. Set absolute
paths when the MCP client may start the server from different project directories.
Set MODEL_NAME or pass --model-name to choose a Hugging Face embedding model that fits the
language and domain of your documents.
mcp-local-rag generates embeddings with mean pooling and L2 normalization. When choosing a model, check whether these settings match its recommended inference setup, since the pooling method can affect retrieval quality.
Changing MODEL_NAME, RAG_DEVICE, or RAG_DTYPE can make existing vectors incompatible.
Use a new DB_PATH or delete the existing index and re-ingest after changing the embedding
configuration.
An example model for English documents is Xenova/bge-small-en-v1.5.
BASE_DIR, BASE_DIRS, or CLI --base-dir roots.DB_PATH. Read-only queries can run
while a sync is active.DB_PATH directory while no writer is active.Documents must be ingested first. Run "List all ingested files" to verify.
Check internet connection. If behind a proxy, configure network settings. The model can also be downloaded manually.
Default limit is 100MB. Split large files or increase MAX_FILE_SIZE.
Check chunk count with status. Large documents with many chunks may slow queries. Consider splitting very large files.
Ensure file paths are within one of the configured roots (BASE_DIR, any BASE_DIRS entry, or any CLI --base-dir). Use absolute paths.
BASE_DIRS accepts a JSON array of one or more non-empty path strings:
BASE_DIRS='["/Users/me/work","/Users/me/specs"]'BASE_DIRS=/a:/b (delimiter syntax not supported)BASE_DIRS='[]' (empty array)npx mcp-local-rag should run without errorsContributions welcome! See CONTRIBUTING.md for setup and guidelines.
MIT License. Free for personal and commercial use.
Built with Model Context Protocol by Anthropic, LanceDB, and Transformers.js.