The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Wso2 Docs MCP Server listing page.
"This is an unofficial community project. Not affiliated with or endorsed by WSO2."
A production-ready Model Context Protocol (MCP) server that provides AI assistants (Claude Desktop, Claude Code, Cursor, VS Code) with semantic search over WSO2 documentation via Retrieval-Augmented Generation (RAG).
Under the hood, it uses a blazing-fast dual-ingestion engine:
| Product | ID | URL |
|---|---|---|
| API Manager | apim | https://apim.docs.wso2.com |
| Micro Integrator | mi | https://mi.docs.wso2.com/en/4.4.0 |
| Ballerina Integrator | bi | https://bi.docs.wso2.com |
| Choreo | choreo | https://wso2.com/choreo/docs |
| Identity Server | is | https://is.docs.wso2.com/en/latest |
| Ballerina | ballerina | https://ballerina.io/learn |
| WSO2 Library | library | https://wso2.com/library |
Choose the setup path that fits your use case:
Install the package globally to get the wso2-docs-mcp-server, wso2-docs-crawl, and wso2-docs-migrate commands available system-wide:
Prefer no global install? You can use
npx wso2-docs-mcp-server,npx wso2-docs-crawl, andnpx wso2-docs-migratein every step below - just replace the bare command with itsnpxequivalent.
Download the docker-compose.yml and start the database:
Install Ollama and pull the default embedding model:
No Ollama? Skip this step. The server automatically falls back to HuggingFace ONNX - model downloads on first use with no extra setup.
Run migration again whenever you change
EMBEDDING_DIMENSIONS(i.e. switch embedding provider). The script detects and handles dimension changes automatically.
Available product IDs: apim, mi, bi, choreo, is, ballerina, library
The MCP server is launched on demand by your AI client - no background process needed.
Claude Desktop - edit ~/Library/Application Support/Claude/claude_desktop_config.json:
Claude Code - run once in your terminal:
Cursor - create .cursor/mcp.json in your project root:
VS Code - create .vscode/mcp.json:
Using
npxinstead of global install? Replace"command": "wso2-docs-mcp-server"with"command": "npx"and add"args": ["-y", "wso2-docs-mcp-server"].
Cloud embedding provider? Add the key to
env, e.g."EMBEDDING_PROVIDER": "openai", "OPENAI_API_KEY": "sk-...".
Install Ollama and start it:
No Ollama? Skip this step. The server detects Ollama is not running and automatically falls back to HuggingFace ONNX inference - the model downloads on first use with no extra setup.
Note: Run migration again whenever you change
EMBEDDING_DIMENSIONS(i.e. switch embedding provider). The script detects and handles dimension changes automatically.
For development (no build step):
Replace
/ABSOLUTE/PATH/TO/wso2-docs-mcp-serverwith your actual clone path.
Claude Desktop - edit ~/Library/Application Support/Claude/claude_desktop_config.json:
Claude Code:
See config-examples/claude_code.sh for a convenience script.
Cursor - create .cursor/mcp.json - see config-examples/cursor_mcp.json.
VS Code - create .vscode/mcp.json - see config-examples/vscode_mcp.json.
| Tool | Description |
|---|---|
search_wso2_docs | Semantic search across all products. Optional product and limit filters. |
get_wso2_guide | Search within a specific product (apim, mi, bi, choreo, is, ballerina, library). |
explain_wso2_concept | Broad concept search across all products, returns 8 top results. |
list_wso2_products | Returns all supported products with IDs and base URLs. |
The default EMBEDDING_PROVIDER=ollama runs entirely on your machine with no API key. The startup sequence is:
Both paths use nomic-embed-text / Xenova/nomic-embed-text-v1 by default and produce identical 768-dim vectors, so you can switch between them without re-indexing.
When Ollama is not available, the server auto-detects the best compute backend:
| Machine | Detection | ONNX dtype | Batch size | Throughput |
|---|---|---|---|---|
| Apple Silicon (M1/M2/M3/M4) | process.arch === 'arm64' | q8 INT8 | 32 | ~9 ms/chunk |
| NVIDIA GPU | nvidia-smi probe | fp32 | 64 | GPU-dependent |
| All others | fallback | q8 INT8 | 16 | ~10 ms/chunk |
Why q8 on Apple Silicon instead of CoreML/Metal?
CoreML compiles Metal shaders on first use (~20 min cold-start). For the typical chunk sizes produced by this server (6–20 chunks per page), the CPU↔GPU transfer overhead eliminates any inference gain. INT8 quantized inference on ARM NEON SIMD is consistently ~100× faster than fp32 CPU with zero cold-start cost.
Benchmark (Apple M-chip, Xenova/nomic-embed-text-v1):
Note: For small crawls (≤ 10 pages) total wall-clock time is dominated by network I/O (HTTPS fetches to docs sites), so the end-to-end improvement is modest. The embedding speedup becomes significant at scale - crawling 500+ pages where embedding previously accounted for hours of runtime. For best crawl performance, run Ollama (
ollama serve) which parallelises inference natively and has no per-chunk overhead.
| Variable | Default | Description |
|---|---|---|
DATABASE_URL | - | PostgreSQL connection string (required) |
EMBEDDING_PROVIDER | ollama | ollama | openai | gemini | voyage |
EMBEDDING_DIMENSIONS | 768 | Must match model output dimensions |
CRAWL_CONCURRENCY | 5 | Concurrent HTTP requests during crawl |
CHUNK_SIZE | 800 | Approximate tokens per chunk |
CHUNK_OVERLAP | 100 | Overlap tokens between chunks |
CACHE_TTL_SECONDS | 3600 | In-memory query cache TTL |
TOP_K_RESULTS | 10 | Default search result count |
| Variable | Default | Description |
|---|---|---|
OLLAMA_BASE_URL | http://localhost:11434 | Ollama server URL |
OLLAMA_EMBEDDING_MODEL | nomic-embed-text | Model pulled and used via Ollama |
HUGGINGFACE_EMBEDDING_MODEL | Xenova/nomic-embed-text-v1 | ONNX fallback when Ollama is not running |
| Variable | Default | Description |
|---|---|---|
OPENAI_API_KEY | - | Required if EMBEDDING_PROVIDER=openai |
OPENAI_EMBEDDING_MODEL | text-embedding-3-small | OpenAI model |
GEMINI_API_KEY | - | Required if EMBEDDING_PROVIDER=gemini |
GEMINI_EMBEDDING_MODEL | text-embedding-004 | Gemini model |
VOYAGE_API_KEY | - | Required if EMBEDDING_PROVIDER=voyage |
VOYAGE_EMBEDDING_MODEL | voyage-3 | Voyage model |
| Provider | Model | Dimensions |
|---|---|---|
| Ollama / HuggingFace | nomic-embed-text / Xenova/nomic-embed-text-v1 | 768 (default) |
| Ollama / HuggingFace | mxbai-embed-large / Xenova/mxbai-embed-large-v1 | 1024 |
| Ollama / HuggingFace | all-minilm / Xenova/all-MiniLM-L6-v2 | 384 |
| OpenAI | text-embedding-3-small | 1536 |
| OpenAI | text-embedding-3-large | 3072 |
| Gemini | text-embedding-004 | 768 |
| Voyage | voyage-3 | 1024 |
| Voyage | voyage-3-lite | 512 |