Local Rust RAG server with MiniLM embeddings, hybrid search, incremental indexing, and MCP tools for document retrieval.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent — or use 1-click editor setup below.
One-click editor setup isn’t available for this listing yet — we don’t have a confirmed install command, and we’d rather show nothing than point your editor at the wrong package or host. Follow the project’s own setup instructions, linked above.
Inspect callable tools, capabilities, and parameters exposed to AI agents by Quillrag.
rag_indexIncrementally index a directory/file. Skips unchanged files, re-embeds only diffs. Pruning is scoped to the directory you pass: docs indexed from other roots are untouched. Symlinked files/dirs are followed.
rag_searchHybrid retrieval: dense MiniLM cosine + BM25 keyword, fused with Reciprocal Rank Fusion. Returns ranked chunks with source paths.
rag_statusDocument/chunk counts, bytes indexed, file-type breakdown.
rag_clearWipe everything.
The quillrag MCP server turns local files into a searchable retrieval-augmented generation store. It is distributed as a static Rust binary with the MiniLM embedding model included, so normal operation does not require Node, Python, pip, npm, or downloading a model. Processing, storage, and search remain on the local machine; the project states that it has no network code path after installation.
The default indexer handles common text and code formats, including Markdown, plain text, JSON, YAML, TOML, CSV, HTML, XML, logs, and many programming-language files. It skips dot-directories and several build or dependency directories such as node_modules, target, dist, venv, and __pycache__. Paragraphs are split into chunks with a 1000-character cap and 120-character overlap.
rag_index walks a file or directory and records document content, metadata, and embeddings. Content hashes let it skip unchanged files and re-embed only modified content. Pruning applies only to the root being indexed, so indexing one directory does not remove documents previously added from another root. Symlinked files and directories are followed, while cycles, broken links, and unreadable directories stop the walk before the store is changed.
rag_search combines dense retrieval from MiniLM embeddings with BM25 keyword retrieval. Reciprocal Rank Fusion merges the two rankings and returns ranked chunks with their source paths. Dense retrieval uses an exact, single-threaded scan of the stored vectors rather than an approximate nearest-neighbor index, so query time grows with corpus size.
The data is stored in one redb file, with a BM25 sidecar index rebuilt during indexing. The model loads lazily on the first indexing or search operation; the MCP handshake and rag_status do not load it. The project reports roughly 20 ms for startup and MCP initialization, about 25 ms for later searches on a small corpus, and higher latency as the number of vectors increases.
Download a release binary for Linux x86_64, macOS Apple Silicon, or Windows x86_64, or build it with cargo install --path .. Linux release binaries from version 0.1.6 require glibc 2.39 or newer; older Linux distributions should build from source with a local toolchain.
Start the MCP process with quillrag serve. A client configuration can pass the data directory through QUILLRAG_DATA:
The same engine is available from the CLI with index, search, status, and clear commands. File extensions can be extended with the CLI -e option or the corresponding extensions setting.
The quillrag MCP server exposes four tools:
rag_index incrementally indexes a file or directory, with optional behavior to avoid pruning or force re-embedding for the selected root.rag_search performs hybrid semantic and keyword retrieval and returns ranked chunks with source paths.rag_status reports document and chunk totals, indexed bytes, and file-type counts.rag_clear removes the complete store.Optional PDF and OCR extraction is available to applications embedding the library with those features. The README does not describe those extractors as part of the standard standalone binary setup.
Retrieval is an exact linear scan, not ANN search. The README estimates roughly 25 ms for 1,000 chunks, 250 ms for 10,000, and 2–5 seconds for 100,000; the latter figures are extrapolated. The binary is about 105 MB, with approximately 120 MB idle RAM usage and up to about 250 MB during batch embedding.
Older indexes affected by a historical chunk-path association issue should be rebuilt in a new data directory and verified before switching clients. Explicitly use rag_clear or the CLI clear when the whole store must be removed; indexing another root does not wipe unrelated roots.
Factual signals from GitHub, npm, and our automated checks — not a rating.
No reviews yet — be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/quillrag)<a href="https://allmcps.com/mcp/quillrag"><img src="https://allmcps.com/api/badge/quillrag?style=directory" alt="Quillrag on AllMCPs" /></a>