The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Pulse8 AI Cortex Knowledge Vault listing page.
The open-source knowledge layer for AI agents
Knowledge that compounds.
PULSE8.ai Cortex is the open-source knowledge layer for AI agents: Git-native memory, a typed knowledge graph, and MCP-powered retrieval on top of plain Markdown — so agents can build, evolve, and reuse persistent knowledge instead of re-deriving it on every query.
Under the hood it's a unified vault for AI agents and humans, backed by a typed knowledge graph, full-text + hybrid search, and a MarkItDown-powered file compiler. Drop files in (PDF, DOCX, PPTX, XLSX, HTML, images, and more), let agents read, write, search, link, and compile knowledge — no database required.
Inspired by Andrej Karpathy's LLM Wiki pattern — a persistent, compounding knowledge base maintained by LLMs instead of re-derived on every query. Search powered by Tobi Lütke's QMD.
Most AI agents can access tools, but they cannot accumulate knowledge.
Traditional RAG systems retrieve documents. PULSE8.ai Cortex builds a persistent, evolving knowledge layer that grows over time and becomes more valuable the more agents and humans interact with it.
With PULSE8.ai Cortex, agents can:
note, agent_def, session, daily, feedback)vault_context builds a ranked subgraph from any seed query| Aspect | Traditional RAG | PULSE8.ai Cortex |
|---|---|---|
| Focus | Documents | Knowledge |
| Memory | Session-based | Persistent |
| Structure | Chunks | Markdown + typed graph |
| Evolution | Static index | Continuous, file-watched |
| Versioning | None | Git-native |
| Agent collaboration | Limited | First-class (MCP) |
| Capability | PULSE8.ai Cortex | Traditional RAG | GraphRAG |
|---|---|---|---|
| Persistent knowledge | ✅ | ❌ | ⚠️ |
| Markdown-native storage | ✅ | ❌ | ❌ |
| MCP-compatible out of the box | ✅ | ❌ | ❌ |
| Knowledge graph | ✅ | ❌ | ✅ |
| Git versioning | ✅ | ❌ | ❌ |
| Agent memory layer | ✅ | ❌ | ⚠️ |
| Human + AI collaboration | ✅ | ❌ | ⚠️ |
| Continuous knowledge evolution | ✅ | ❌ | ⚠️ |
| Zero database required | ✅ | ❌ | ❌ |
PULSE8.ai Cortex speaks MCP — so it plugs into any AI client that does. The same vault is reachable over streamable HTTP or stdio, and mirrored 1:1 by a REST API at /api/v1/.
| Category | Compatible with |
|---|---|
| AI agents | Claude Desktop, Claude Code, OpenAI Agents, Gemini, custom agent frameworks |
| Development tools | Cursor, VS Code, JetBrains IDEs |
| Agent frameworks | LangGraph, LangChain, CrewAI, AutoGen |
| MCP ecosystem | MCP clients, MCP servers, MCP tool registries |
| Human tools | Obsidian, any Markdown editor, any Git client |
Because the vault is just files, humans and agents collaborate on the same knowledge — no proprietary format, no lock-in.
[!NOTE] PULSE8.ai Cortex requires Docker. An OpenRouter API key is optional — needed only for LLM-powered cross-referencing between wiki articles. File conversion works out of the box without any API key.
3. Connect your MCP client (e.g. Claude Desktop) to http://localhost:8420/mcp/.
To stop: ./scripts/stop.sh
Docker Desktop on macOS cannot expose the Metal GPU to containers, so containerized QMD embeds on CPU only — over an order of magnitude slower on non-trivial vaults. Run QMD natively instead; the qmd binary uses Metal automatically:
The QMD daemon's pid and log are kept in .qmd-native.pid / .qmd-native.log. To stop both: ./scripts/stop.sh --native-qmd (a plain ./scripts/stop.sh also cleans up a native QMD if one is running).
If you manage QMD yourself (already running elsewhere), start only the Cortex container:
To stop: ./scripts/stop.sh --cortex-only
For production deployments with NVIDIA GPU acceleration:
See docs/ec2-gpu-setup.md for a full guide on instance selection, NVIDIA toolkit installation, and cost estimates.
| Knowledge Graph | Typed graph engine (NetworkX) — wikilinks, tags, and custom edges, auto-maintained on every file change |
| Full-Text Search | QMD search with hybrid (BM25 + vector + re-ranking) by default; keyword and semantic modes selectable. Results cached with a configurable TTL. |
| File Compiler | Converts raw sources (PDF, DOCX, PPTX, XLSX, HTML, images, etc.) to Markdown via MarkItDown. LLM used only for cross-referencing. |
| MCP Server | Streamable HTTP + stdio transport — works with Claude Desktop, Cursor, and any MCP client |
| Feedback & Notifications | vault_feedback captures quality feedback as notes; optional Microsoft Teams webhook posts an adaptive card per submission |
| Daily Activity Log | Every write/ingest/compile is mirrored into daily/<date>.md as a greppable, wikilinked timeline |
| Bulk Ingest | Ingest dozens or hundreds of files at once from a local directory with SHA-256 dedup and bounded concurrency |
| REST API | FastAPI endpoints mirroring all MCP tools at /api/v1/, including multipart file upload and bulk ingest |
| Vault Watcher | Real-time filesystem monitoring — graph stays in sync automatically |
| Lineage & Audit | Every edge labeled extracted / inferred / manual; vault_trace answers "why does the vault say X" back to the source document |
| Graph Queries | vault_path (what connects X to Y), vault_impact (what's downstream of this note), vault_explain (entity summary with provenance) |
| Curation Report | Read counters + outcome feedback (useful / dead-end / corrected) surface stale, contradicted, and never-read notes at GET /api/v1/curation/report |
| Zero Database | Everything persists as Markdown + JSON on your filesystem |
45.0% overall accuracy on LongMemEval-S (500 questions, full haystacks, hybrid search) with 65.6% evidence recall@8 and zero judge errors — measured end-to-end through the public REST API: ingest → compile → graph → search → answer.
| Category | Accuracy | Recall@8 |
|---|---|---|
| single-session-assistant | 96.4% | 98.2% |
| single-session-user | 71.4% | 78.6% |
| knowledge-update | 60.3% | 75.6% |
| temporal-reasoning | 25.6% | 51.1% |
| multi-session | 24.8% | 54.1% |
| single-session-preference | 23.3% | 63.3% |
Every number is reproducible from a pinned config (dataset SHA-256, models, seed) with one command:
The harness (evals/) publishes per-question JSONL traces, separates the judge model from the answer model, uses the official LongMemEval per-type grading prompts, and includes blind human validation of the judge. Full methodology, caveats, and raw results: docs/benchmarks/.
Cortex is deterministic-first: ingestion (MarkItDown conversion), the knowledge graph (wikilinks, tags, derived_from edges), and QMD search all work with zero LLM calls. The LLM is an optional enrichment pass — cross-referencing, tagging, image captioning — not a dependency.
Pick a backend with LLM_BACKEND (env) / CORTEX_LLM_BACKEND (Python):
| Backend | What it covers |
|---|---|
openai-compatible (default) | OpenRouter, Azure OpenAI, Ollama, vLLM, LM Studio — anything speaking the OpenAI protocol. Point LLM_BASE_URL at your endpoint. |
bedrock | AWS Bedrock via the standard AWS credential chain (no API key). Requires boto3. |
none | Explicit zero-LLM mode. Guaranteed to construct no LLM client and make no model calls — suitable for air-gapped deployments. |
Air-gapped example with a local Ollama:
Or fully deterministic: LLM_BACKEND=none ./scripts/start.sh (no API key needed).
PULSE8.ai Cortex implements the resources-as-tool-inputs pattern recommended by the Microsoft Copilot Studio CAT team: token-heavy tool outputs (large search result sets, full context windows) can be kept server-side and passed between tools as lightweight handles, so the LLM context window stays small.
How it works. Pass as_resource: true to vault_search or vault_context (or ?as_resource=true on GET /api/v1/search). Instead of inlining the full payload, you get a handle:
Read it back through any of the three transports:
resources/read with the cortex://resource/{id} URI (Claude Desktop, Cursor, Copilot Studio MCP).vault_resource_read for clients that only expose tools to the planning layer (some Copilot Studio configurations).GET /api/v1/resources/{resource_id} (accepts the bare ID or the full URI).The store is in-memory, asyncio-safe, TTL-evicted, and LRU-bounded:
| Env var | Default | Purpose |
|---|---|---|
CORTEX_RESOURCE_TTL_SECONDS | 3600 | Max age before a stored resource is evicted lazily on read |
CORTEX_RESOURCE_MAX_ITEMS | 1000 | LRU cap before oldest entry is dropped |
The same ResourceStore is shared between MCP and REST — produce a handle via MCP, read it back via REST (or vice versa).
Microsoft Copilot Studio setup — agent instructions, tool selection, and the Custom Connector fallback — is documented in docs/copilot-studio.md. No Cortex code change required.
| Tool | Description |
|---|---|
vault_read | Read a note by path |
vault_write | Create or update a note |
vault_search | Search the vault (keyword / semantic / hybrid). Supports as_resource=true |
vault_link | Create, query, or delete graph edges |
vault_context | Build a context window: search → graph traversal → ranked subgraph. Supports as_resource=true |
vault_ingest | Ingest raw content or binary files (supports content_base64 for binary) |
vault_compile | Compile unprocessed raw sources into wiki Markdown via MarkItDown |
vault_feedback | Submit feedback on vault quality (status: OPEN; optional related_paths and outcome: useful / dead-end / corrected) |
vault_list_feedbacks | List feedback note metadata (paths, tags, status; not full body) |
vault_resource_read | Read a server-stored MCP resource by ID (fallback for clients without resources/read) |
vault_trace | Trace a note's lineage: provenance, raw sources, and edges labeled extracted / inferred / manual |
vault_path | Shortest paths between two notes — "what connects X to Y", every hop typed and origin-labeled |
vault_impact | Walk everything downstream of a note (change-impact analysis) |
vault_explain | Explain a note: summary, provenance, sources, links in/out, contradictions |
The vault is a plain directory of Markdown files organised by purpose. Cortex classifies each file into a typed node (NodeType) used by the graph engine and exposed in REST and MCP responses.
| Folder | NodeType | Purpose |
|---|---|---|
wiki/ | note | Compiled, interlinked knowledge articles |
raw/ | raw_source | Unprocessed sources (PDF, DOCX, TXT, …) the compiler reads from |
agents/ | agent_def | Agent definitions |
sessions/ | session | Per-session notes / conversation transcripts |
daily/ | daily | Daily notes (Obsidian Daily Notes convention) |
feedback/ | feedback | Feedback on vault quality (status, related_paths) |
.cortex/ | (skipped) | Cortex internals — graph.json, index.md, log.md, manifests |
Order of precedence (first match wins):
type:* — explicit override always wins (e.g. type: note in agents/foo.md resolves to NodeType.NOTE)raw/ agents/ sessions/ daily/ feedback/ inherit the folder's type with no filename suffix needed (e.g. daily/2026-06-10.md → daily).agent.md, .session.md, .memory.md are still honored anywhere (e.g. wiki/legacy.agent.md → agent_def)NodeType.NOTEIn practice this means you can drop YYYY-MM-DD.md straight into daily/, or an unsuffixed planner.md into agents/, and the graph and API will classify them correctly without any renaming.
Every vault_write, vault_ingest, and successful compile event (MCP and REST paths) is automatically mirrored into today's UTC daily note at daily/YYYY-MM-DD.md. The file is created on first event of the day and each subsequent event appends a ## [HH:MM] event | summary block plus a [[wiki-stem]] wikilink (so the watcher draws a LINKS_TO edge to the affected note). The format follows the Karpathy log.md greppable-prefix pattern — grep "^## \[" daily/2026-06-10.md gives a clean timeline of the day.
Writes targeting daily/, feedback/, or .cortex/ are deliberately not mirrored (would be self-referential noise). The hidden .cortex/log.md audit log is unaffected and continues to receive every operation.
For ingesting many files at once (dozens or hundreds of PDFs, papers, docs), use the one-click shell script instead of feeding them one at a time through MCP. It reads directly from a local directory (recursively, including subfolders) — no wire overhead, no running server required — deduplicates via SHA-256 hashing, compiles with bounded concurrency, and rebuilds the index once at the end. Subpaths are preserved under the vault raw folder (e.g. source/abcde/doc.html → raw/abcde/doc.html).
The script automatically loads your .env for the LLM key and vault path, prints a summary, then runs the full pipeline (copy, compile, reindex). No running Cortex server needed.
For programmatic use without MCP (requires running Cortex server):
The dedup manifest is stored at .cortex/ingest-manifest.json. Files are matched by content hash, not filename — renaming a file won't cause re-ingestion, and the same content under a different name will be skipped.
Copy the example and fill in your values:
| Variable | Required | Default | Description |
|---|---|---|---|
LLM_BACKEND | No | openai-compatible | LLM backend: openai-compatible, bedrock (AWS credential chain), or none (zero LLM calls) |
LLM_API_KEY | No | — | OpenRouter (or compatible) API key (for cross-referencing only) |
COMPILER_MODEL | No | anthropic/claude-sonnet-4 | Model for cross-reference detection |
LLM_BASE_URL | No | https://openrouter.ai/api/v1 | LLM API base URL |
VAULT_DIR | No | ./example_vault | Path to your vault directory |
INGEST_DIR | No | ./ingest | Path to bulk-ingest source directory (mounted as /ingest in Docker) |
QMD_REFRESH_INTERVAL_SECONDS | No | 900 | Periodic re-index interval (seconds; 0 to disable) |
QMD_SEARCH_MODE | No | hybrid | Default search mode when unspecified: hybrid (BM25 + vector + re-rank), semantic, or keyword |
QMD_CACHE_TTL_SECONDS | No | 30 | TTL for the search-result cache; raise it on read-heavy vaults to skip repeat QMD calls |
QMD_SEARCH_TIMEOUT_SECONDS | No | 120 | Per-request search timeout (increase for hybrid on CPU-only hosts) |
QMD_EMBED_TIMEOUT_MS | No | 600000 | Embed timeout in ms (increase for CPU-only deployments) |
QMD_URL | No | — | External QMD URL for cortex-only mode (e.g. http://host.docker.internal:3100) |
AUTH_METHOD | No | none | Authentication method: none, apikey, or oidc (see Authentication) |
API_KEY | No | — | Static API key for x-api-key header (used when AUTH_METHOD=apikey) |
OIDC_TENANT_ID | No | — | Microsoft Entra ID tenant ID (used when AUTH_METHOD=oidc) |
OIDC_CLIENT_ID | No | — | Microsoft Entra ID app (client) ID |
OIDC_CLIENT_SECRET | No | — | Microsoft Entra ID client secret |
OIDC_BASE_URL | No | http://localhost:8420 | Public base URL of the Cortex server (used for OAuth callbacks) |
TEAMS_WEBHOOK_URL | No | — | Incoming webhook / Power Automate URL; posts an adaptive card on each new feedback note |
TEAMS_APP_BASE_URL | No | — | Optional public Cortex base URL for a "View in Cortex" link on the Teams card |
OPENROUTER_API_KEY and CORTEX_LLM_API_KEY are accepted as aliases for LLM_API_KEY. Variables above are set in .env (Docker reads them via Compose) and map to the CORTEX_* settings used by the app.
Cortex supports two authentication methods that protect both the REST API (/api/v1/) and the MCP endpoint (/mcp/). Set AUTH_METHOD in .env to choose:
AUTH_METHOD | Description |
|---|---|
none | Default. All endpoints are open — no authentication required. |
apikey | Static API key. Clients pass x-api-key header. |
oidc | Microsoft Entra ID (Azure AD) with OAuth 2.0 + MFA support. |
AUTH_METHOD=apikey)The simplest option. Set the method and key in .env:
Clients pass it via the x-api-key header:
No OAuth discovery endpoints are served — no login popups. Requests without a valid key receive a 401.
AUTH_METHOD=oidc)For enterprise environments that require interactive login with MFA support:
This enables:
GET /api/v1/login. After login, pass the access token as Authorization: Bearer <token>. A valid x-api-key header is also accepted as a fallback when API_KEY is set.To use OIDC, register an app in the Azure Portal:
http://localhost:8420/api/v1/auth/callback (Web platform)openid, profile, and email (Microsoft Graph → Delegated).envAn example config is included at [claude_desktop_config.example.json](https://github.com/synpulse8-opensource/pulse8-ai-cortex-knowledge-vault/blob/HEAD/claude_desktop_config.example.json).
HTTP with API key (recommended) — PULSE8.ai Cortex runs as a persistent server:
HTTP without auth — when no authentication is configured:
Stdio — Claude Desktop launches the server on demand (no auth needed):
Add to your .cursor/mcp.json:
Watcher and Compiler are independent components:
.md file added, modified, or deleted triggers automatic node/edge updates.wiki/. Supported formats include PDF, DOCX, PPTX, XLSX, HTML, CSV, JSON, XML, images (EXIF/OCR), and plain text. The LLM is only used for optional cross-reference detection between articles.They connect indirectly: the compiler writes to wiki/, the watcher picks those up and updates the graph.
| Format | Extensions |
|---|---|
.pdf | |
| Microsoft Word | .docx |
| Microsoft PowerPoint | .pptx |
| Microsoft Excel | .xlsx, .xls |
| HTML | .html, .htm |
| Text-based | .csv, .json, .xml, .txt, .md |
| Images | .jpg, .png, etc. (EXIF metadata) |
Search uses a two-stage pipeline:
QMD answers "what's relevant?" — the graph answers "how are these results connected?"
| Script | Description |
|---|---|
scripts/serve.py | Dev server (HTTP or stdio based on CORTEX_MCP_TRANSPORT) |
scripts/compile.py | Batch-compile all raw sources |
scripts/reindex.py | Full reindex + graph rebuild |
scripts/bulk_ingest.sh | One-click bulk ingest from a local directory |
scripts/bulk_ingest.py | Python CLI for bulk ingest (called by bulk_ingest.sh) |
scripts/lint.py | Lint vault structure |
The vault directory is bind-mounted from your host into the containers. All data lives on your local disk and survives container restarts.
The QMD search index is stored in a Docker volume (qmd-cache). To force a full re-index:
Releases are automated through GitHub Actions. Publishing a GitHub Release triggers three workflows that build and publish everything:
| Workflow | Publishes to |
|---|---|
publish-pypi.yml | PyPI |
publish-docker.yml | GitHub Container Registry (ghcr.io) |
publish-mcp.yml | MCP Registry (GitHub OIDC auth) |
To cut a release:
Bump the version in pyproject.toml and server.json (keep them in sync), update CHANGELOG.md, and commit to main.
Create the GitHub Release — via the UI (Releases → "Draft a new release" → new tag vX.Y.Z) or the CLI:
That's it — the release event fires all three workflows. publish-mcp.yml waits for PyPI to serve the new version (so the mcp-name ownership marker in this README is verifiable), then publishes the server via GitHub OIDC under io.github.synpulse8-opensource/* (no token or local mcp-publisher needed).
Verify the registry entry once the workflow finishes:
[!IMPORTANT] PyPI versions are immutable — a version number can never be reused, even after deletion. Always increment to a new version; never re-release an existing one.
We believe AI agents need a dedicated knowledge layer — not another document store. If you share that vision:
Together we can build the knowledge layer for agentic AI.
We welcome contributions! Please open an issue to discuss your idea before submitting a pull request.
Use GitHub Issues to report bugs or request features.
PULSE8.ai Cortex builds on ideas and tools from the open-source community:
This project is licensed under the PULSE8.ai Cortex Open Source License (Apache License 2.0 with additional terms).