# lyonzin/knowledge-rag [Health: Active]

**Category:** 🧠 Knowledge & Memory  
**Repository:** https://github.com/lyonzin/knowledge-rag  
**GitHub Stars:** 277  
**Views:** 3  
**Installs:** 0  
**Upvotes:** 0  
**Directory Page:** https://allmcps.com/mcp/lyonzin-knowledge-rag

## Description
Local RAG system for Claude Code with hybrid search (BM25 + semantic), cross-encoder reranking, markdown-aware chunking, query expansion, and 28 MCP tools. Runs entirely offline with zero external servers.

## Tools
Capabilities this server exposes over MCP:

- **search_knowledge** — 
    Hybrid search combining semantic search + BM25 keyword search with cross-encoder reranking.

    Read-only. No side effects.

    Args:
        query: Search query text (1–3 keywords recommended; phrase queries also work)
        max_results: Maximum number of results (default: 5, max: 20)
        category: Optional category filter — one of: security, ctf, logscale, development, general,
            redteam, blueteam. Call list_categories() first to see available categories and counts.
        hybrid_alpha: Balance between semantic and keyword search. 0.0 = keyword-only (best for exact
            technical terms like CVE IDs or tool names), 0.3 = balanced default, 1.0 = semantic-only
            (best for conceptual or natural-language queries).
        min_score: Minimum normalized relevance score (0.0–1.0) to include a result. Results scoring
            below this threshold are discarded. Default 0.0 returns all results. Use 0.2–0.4 to cut
            low-relevance noise.
        snippet_mode: When true (default), truncates content to ~500 characters at a natural break
            point and adds a content_length field with the original size. Use get_document() to
            fetch full content when needed. Set to false to return full chunk content.
        search_method: Dispatch selector (v4.8.2+). One of ``"auto"`` (router picks FTS5 fast-path
            for lexical queries when enabled, hybrid otherwise), ``"hybrid"`` (force hybrid path —
            kill switch for suspected router misclassification), or ``"fts5"`` (force FTS5 fast-path
            — debug/testing; errors out when the feature is disabled or the index is not ready).
            Default ``"auto"`` preserves pre-v4.8.2 behavior byte-for-byte when the fast-path is
            disabled in config.

    Returns:
        JSON string with results including content chunks, source filepath, relevance score, and
        search method used. Returns chunks, not full document content.

    Usage: Primary search tool — use for any topic or keyword lookup. Prefer search_similar() when
    you already have a reference document and want more like it. Prefer get_document() when you
    already know the exact filepath and need the full content.
    
- **get_document** — 
    Get the full content of a specific document by filepath.

    Read-only. No side effects.

    Args:
        filepath: Relative path to the document within the documents directory
            (e.g., "security/technique.md"). Must be an indexed file — use
            list_documents() to browse available paths, or search_knowledge()
            to find the filepath by topic first.

    Returns:
        JSON string with full document content and metadata (filepath, category, size).

    Usage: Use when you need the complete text of a known file — search_knowledge()
    returns chunks, not full docs. Use search_knowledge() first to find the filepath
    if unknown. Use list_documents() to browse all available files by category.
    
- **reindex_documents** — Index or reindex all documents in the knowledge base (runs in background).

    ``force`` — smart reindex (detect changed files + rebuild BM25). Use after
    filesystem edits outside add_document/update_document.
    ``full_rebuild`` — nuclear rebuild (delete + re-embed). Use only after
    embedding-model change or index corruption. Mutually exclusive with resume.
    ``resume`` — pick up an interrupted smart reindex from
    ``data/reindex_checkpoint.json``. Falls back to a fresh smart run silently
    if the checkpoint is missing/corrupt/drifted (v4.8.0 Fase 4).

    Returns a JSON envelope. Poll ``get_reindex_status()`` until
    ``reindex.active`` becomes false. Add/update/URL tools already auto-index —
    use these flags only for the recovery/rebuild scenarios above.
    
- **get_reindex_status** — 
    Get the current status of a background reindex operation.

    Lightweight — does not compute full index statistics. Use this to poll progress
    after calling reindex_documents().

    Returns:
        JSON string with reindex status. When active: operation name, progress (processed/total),
        percent complete, indexed/skipped/errors counts, and start time. When inactive: active=false,
        plus last_result or last_error from the most recent completed reindex.

    Usage: Call repeatedly after reindex_documents() to monitor progress. When reindex.active
    becomes false, the operation is complete. Use get_index_stats() for full index health metrics.
    
- **list_categories** — 
    List all document categories with their document counts.

    Read-only. No side effects. Reflects the live index state.

    Returns:
        JSON string with category names, document counts per category, and total document count.

    Usage: Use before filtering search_knowledge() or list_documents() by category to see
    which categories exist and how many documents each contains. Use get_index_stats() instead
    for broader system health metrics (model name, cache hit rate, BM25 status).
    
- **list_documents** — 
    List all indexed documents, optionally filtered by category.

    Read-only. No side effects.

    Args:
        category: Optional category filter. Must be a valid category name — call
            list_categories() to see available options (e.g., security, ctf, logscale,
            development, general, redteam, blueteam).

    Returns:
        JSON string with list of document filepaths, categories, and metadata for each indexed file.

    Usage: Use to browse what's in the index or verify a specific file is indexed. Use
    list_categories() first to see valid category names. Use search_knowledge() when you
    want to find documents by topic rather than browsing the full list. Use get_document()
    to read a specific file once you have its filepath.
    
- **get_index_stats** — 
    Get statistics and health metrics for the knowledge base index.

    Read-only. No side effects.

    Returns:
        JSON string with system metrics: total documents, total chunks, embedding model name,
        BM25 status, query cache hit rate, and file watcher status.

    Usage: Use for system health checks — verifying the embedding model loaded, checking
    index population, or monitoring cache efficiency. Use list_categories() for per-category
    document counts instead. Use evaluate_retrieval() to measure actual search quality with
    test queries.
    
- **add_document** — 
    Add a new document to the knowledge base from raw text content.

    Mutating — writes a file to disk and indexes it immediately. No auth required.

    Args:
        content: Full text content of the document (markdown supported)
        filepath: Relative path within documents directory (e.g., "security/new-technique.md").
            The subdirectory should match the category.
        category: Document category — one of: security, ctf, logscale, development, general,
            redteam, blueteam (default: general)

    Returns:
        JSON string with indexing results (filepath, chunks created, status).

    Usage: Use to add new documents from text content. Use add_from_url() instead when
    the source is a web page. Use update_document() to replace content of an existing file.
    The document is immediately searchable after this call — no manual reindex needed.
    
- **update_document** — 
    Update the content of an existing document in the knowledge base.

    Mutating — overwrites the file on disk and re-indexes immediately. Old chunks are
    removed and replaced with new ones. Full content replacement, not a patch.

    Args:
        filepath: Full or relative path to the document file. Must be an already-indexed
            file — use list_documents() to find valid paths.
        content: New full-text content to replace the existing content entirely

    Returns:
        JSON string with update results (old chunk count, new chunk count, status).

    Usage: Use to replace a document's content completely. Use add_document() to create
    a new file instead. Use remove_document() to delete without replacing. Changes are
    immediately searchable — no manual reindex needed.
    
- **remove_document** — 
    Remove a document from the knowledge base index.

    Mutating — removes index entries. If delete_file=True, also permanently deletes
    the file from disk (irreversible, cannot be undone).

    Args:
        filepath: Path to the document file. Must be an indexed document — use
            list_documents() to find valid paths.
        delete_file: If True, permanently deletes the file from disk in addition to
            removing from the index (default: False).

    Returns:
        JSON string with removal results (filepath, status).

    Usage: Use to unindex a document while keeping the file on disk (default). Set
    delete_file=True only for permanent removal. Use update_document() to replace
    content instead of removing. Use reindex_documents(force=True) if you deleted
    the file manually on disk outside of this tool.
    
- **add_from_url** — 
    Fetch content from a URL, convert to markdown, and add to the knowledge base.

    Mutating — makes an outbound HTTP request (requires internet access), strips HTML,
    converts to markdown, saves to disk, and indexes immediately.

    Args:
        url: Full URL to fetch (https:// required). The page must be publicly accessible.
        category: Document category — one of: security, ctf, logscale, development, general,
            redteam, blueteam (default: general)
        title: Optional document title. Auto-detected from the page's <title> tag if omitted.

    Returns:
        JSON string with indexing results (detected title, filepath, chunks created, status).

    Usage: Use to ingest web content (writeups, blog posts, documentation pages) directly
    by URL. Use add_document() instead when you already have the text content. The document
    is immediately searchable after this call — no manual reindex needed.
    
- **search_similar** — 
    Find documents semantically similar to a given reference document.

    Read-only. No side effects. Uses the document's embedding for similarity comparison.

    Args:
        filepath: Path to the reference document (must already be indexed — use
            list_documents() to verify). E.g., "security/technique.md"
        max_results: Number of similar documents to return (default: 5, max: 20)

    Returns:
        JSON string with list of similar document filepaths and similarity scores (0.0–1.0).

    Usage: Use when you have a specific document and want to discover thematically related
    ones. Use search_knowledge() instead when you have a text query rather than a reference
    document. The reference document must be indexed — call list_documents() to confirm
    it exists before calling this tool.
    
- **evaluate_retrieval** — 
    Evaluate search quality by testing whether search_knowledge() retrieves expected documents.

    Read-only. Runs multiple search queries internally. No side effects on the index.

    Args:
        test_cases: JSON string array of test cases. Each item requires "query" (search string)
            and "expected_filepath" (path of the document that should appear in top-5 results).
            Example: [{"query": "suid exploit", "expected_filepath": "security/suid.md"}]

    Returns:
        JSON string with MRR@5 (Mean Reciprocal Rank), Recall@5, and per-query hit/miss breakdown.
        MRR@5 above 0.7 indicates good retrieval quality.

    Usage: Use to audit search quality after bulk document ingestion or after tuning
    hybrid_alpha. Use get_index_stats() for system health checks instead. Use
    search_knowledge() for actual document retrieval — this tool is for quality measurement only.
    

## Claude Desktop Quick Installation
Install path detected from listing signals. Uses `npx` (confidence: high):

```json
"mcpServers": {
  "knowledge-rag": {
    "command": "npx",
    "args": ["-y","skills"]
  }
}
```

## Documentation

## What lyonzin/knowledge-rag MCP server does

The lyonzin/knowledge-rag MCP server provides a local retrieval-augmented generation backend for MCP clients. It indexes documents stored in a configured knowledge-base directory and makes their contents available through search, browsing, retrieval, and maintenance tools. The repository describes support for Claude Code, Cursor, Windsurf, Cline, VS Code, Gemini CLI, and Zed, with Python 3.11+ and Windows, Linux, and macOS listed as supported environments.

The primary retrieval path searches text using both semantic embeddings and BM25 keyword matching, then applies cross-encoder reranking. Search results contain chunks, source paths, relevance scores, and the selected search method rather than complete documents. A separate tool retrieves the full text when the path is known.

## How it works

Documents are split into chunks and indexed locally. Semantic indexing uses an in-process FastEmbed ONNX model; the first query may download the model, after which the system can operate offline. Search can favor exact lexical matches, balance lexical and semantic signals, or use semantic similarity alone. The search tool also supports result limits, category filtering, minimum relevance scores, and abbreviated snippets.

The lyonzin/knowledge-rag MCP server supports seven document categories: security, ctf, logscale, development, general, redteam, and blueteam. Agents can inspect available categories and indexed file paths before issuing filtered searches or requesting a full file. For a known document, similarity search finds related indexed documents using its embedding.

Index maintenance is asynchronous. A smart reindex detects filesystem changes, while a full rebuild is intended for embedding-model changes or index corruption. A resume mode can continue an interrupted operation when a valid checkpoint exists. After starting a reindex, clients should poll the status tool until the operation is inactive.

## Setup and configuration

Install the Python package with `pip install knowledge-rag`, then run `knowledge-rag init` to create a configuration file and `documents/` directory. Add Markdown, PDFs, code, or other supported files to that directory and restart the MCP client. The first embedding-model download is approximately 200 MB according to the project documentation; later use can be offline.

The project also documents SSE and streamable HTTP transports for shared deployments. HTTP configuration can specify a host and port, bearer authentication, request limits, metrics, and JSON logging. These settings are configuration options rather than requirements for the local MCP workflow.

## Tools and capabilities

The MCP tool set includes capabilities for:

- Hybrid keyword and semantic search with reranking.
- Full-document retrieval by indexed filepath.
- Category and document listing.
- Adding, replacing, removing, and URL-importing documents.
- Similar-document discovery.
- Background reindexing and progress polling.
- Index health statistics and retrieval-quality evaluation using MRR@5 and Recall@5.

Document mutations write to disk and index immediately. Removing an item normally unindexes it while preserving the file; passing the deletion option permanently removes the file. URL ingestion requires internet access and only supports publicly accessible HTTPS pages.

## Limitations and notes

Search returns chunks by default, so agents needing complete context must call the document retrieval tool. Filepath-based operations require the document to be indexed, and category filters must use valid live categories. A full rebuild is more disruptive than a smart reindex and should be reserved for model changes or index problems.

Although the lyonzin/knowledge-rag MCP server is designed for local and offline use, URL import makes outbound HTTP requests. No external service or paid API key is identified as required by the provided material. GPU support is optional; the documentation states that CPU execution is available through FastEmbed ONNX.

_Full upstream README: https://allmcps.com/mcp/lyonzin-knowledge-rag/readme_

