The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Kagura Memory Cloud listing page.
English · 日本語
Adaptive memory for AI agents and teams — self-hosted, beyond RAG.
An MCP server that gets smarter every time you search:
hybrid search + a neural memory graph that learns which memories belong together.
Works with Claude, ChatGPT, Gemini, and any MCP-compatible client.
Python SDK (KaguraClient, REST clients & FileIngestor)
Claude Code CLI recalling memories from Kagura over MCP — ▶ watch the demo
Your AI forgets everything after each conversation. Kagura fixes that — and gets smarter every time you search.
Most AI memory tools are just vector databases with a chat wrapper. Kagura is different — it implements the full LLM Knowledge Base pattern (Karpathy's LLM Wiki) at team scale:
| Approach | Storage | Compounding | Scale |
|---|---|---|---|
| Vector DB / RAG | Embedded chunks | None — retrieve-only | Any |
| Karpathy's LLM Wiki | Markdown files | LLM rewrites pages | Personal (~100 pages) |
| Kagura Memory Cloud | PostgreSQL + Qdrant + Neural graph | Hebbian + Sleep Maintenance | Team / org |
| Feature | Description |
|---|---|
| Adaptive Memory | Every search automatically strengthens connections between related memories. The more you use it, the better explore() discovers hidden relationships. |
| Hybrid Search | Semantic (OpenAI / self-hosted) + BM25 keyword — 96% top-1 accuracy |
| AI Reranking | Self-hosted (Ollama/vLLM — local, free), Voyage AI, or Cohere — cross-encoder reranking for precision |
| Neural Memory Graph | Hebbian learning builds a knowledge graph in the background. explore() traverses it for serendipitous discovery. |
| Agent Memory Substrate | Beyond a knowledge store: delivery modes (pinned / time-triggered), a server-stamped trust boundary, an agent state lane, and a retrieval-feedback signal — the primitives an autonomous agent loop needs. |
| Agent Control Plane (preview) | Workspace-scoped Agent Registry, subtractive context bindings, agent-bound member keys, lifecycle kill switches, and one-call session bootstrap. Introduced in v0.49.0. |
| 64 MCP Tools | Memory, Agent Substrate, Agent Control Plane, Neural edges, Contexts, Tags, Files (R2), Analyses (Memory Analysis), Resources, Secrets, Sleep Maintenance, Usage, API-Key Bindings |
| Multi-Provider | OpenAI or self-hosted (Ollama, vLLM — local, private, zero cost) for embeddings |
| Team Ready | Workspaces, RBAC, context isolation, shared memory |
| Web UI | Next.js dashboard — contexts, search settings, member management |
| 5-Minute Setup | ./setup.sh and you're done |
Karpathy's LLM Wiki pattern describes a 5-layer "living knowledge base" — beyond traditional RAG. Kagura implements all 5 layers at team scale:
| Layer | Kagura Implementation | Difference from Karpathy's pattern |
|---|---|---|
| Ingest | REST /api/v1/memory, MCP remember, R2 file storage, resource tokens | + binary blobs, + multi-tenant |
| Compile | MCP-as-compile-API — chat agent compiles via structured tool calls (remember(summary, content, type, tags)) + Sleep Maintenance for batch consolidation | Continuous micro-compile (not batch wiki rewrite) — schema-enforced output |
| Index | Triple index: BM25 (keyword) + Qdrant (semantic) + Hebbian graph (relational) — all auto-maintained | No manual index.md upkeep |
| Query | Hybrid Search + AI Reranker + explore graph traversal | Beyond markdown grep — supports semantic + relational queries |
| Enhance | Hebbian learning — every recall() strengthens edges between co-retrieved memories. Sleep Maintenance consolidates periodically. | Background graph evolution (zero LLM cost) vs LLM-driven page rewrites |
Compounding loop: Currently explicit (user/agent calls remember() after synthesizing answers). Auto-write-back of synthesized answers is intentionally opt-in to keep noise low.
Kagura separates precision search and discovery into two independent paths, each optimized for its purpose:
recall() — Precision search. Hybrid (semantic 60% + BM25 40%) with optional AI reranking. Returns the most relevant memories.explore() — Discovery. Traverses the Neural Memory graph to find related memories that keyword search would miss.recall() silently strengthens edges between co-retrieved memories. No explicit training needed — the graph grows organically as you use the system.This separation is intentional: mixing graph signals into recall degrades precision (validated via benchmarks). Instead, each path does what it's best at.
Data isolation: All data is filtered by workspace_id → context_id → user_id. Memories never leak across boundaries. Single Qdrant collection with payload filtering.
Tech stack: FastAPI (async) · PostgreSQL · Qdrant · Redis · Next.js 16 · OAuth2 · MCP over Streamable HTTP
Vector backend: Qdrant by default. A single-process self-hosted / CLI / edge deployment can instead run the embedded LanceDB backend — "Kagura Lite" (preview) with no separate Qdrant server (KAGURA_VECTOR_BACKEND=lance, cd backend && uv sync --locked --extra lite). Not for multi-worker / SaaS (LanceDB is single-writer). See Deployment → Embedded Vector Backend.
| Minimum | Recommended | |
|---|---|---|
| CPU | 2 cores | 4+ cores |
| RAM | 4 GB | 8+ GB |
| Disk | 10 GB free | 20+ GB free |
One-line setup:
With Claude Code:
Step-by-step setup:
.env.local settings (auto-configured by setup_env):
| Setting | Required | Description |
|---|---|---|
API_KEY_SECRET | Yes | Secret for API key encryption (auto-generated) |
JWT_SECRET | Yes | Secret for JWT tokens (auto-generated) |
OPENAI_API_KEY | Yes* | OpenAI API key for embeddings |
SELF_HOSTED_BASE_URL | No | Self-hosted backend URL (default: http://localhost:11434) |
EMBEDDING_PROVIDER | No | openai (default) or self_hosted |
GOOGLE_CLIENT_ID/SECRET | No | Google OAuth2 login (optional — password login available) |
GITHUB_CLIENT_ID/SECRET | No | GitHub OAuth2 login (optional) |
* Either OPENAI_API_KEY or a running self-hosted inference server (e.g. Ollama) is required for memory features.
| Command | Purpose |
|---|---|
python3 -m src.cli.setup_env | Generate secrets + configure .env.local (run before Docker) |
python3 -m src.cli.create_admin | Create admin + workspace + API key + .mcp.json + embedding setup |
python3 -m src.cli.reset_password | Reset password and/or MFA |
python3 -m src.cli.delete_admin | Delete admin (for re-creation) |
Run from
backend/directory. Docker API container must be running.
brew install python@3.11 nodesudo apt install docker.io docker-compose-v2 python3.11 nodejs npm.env.local (DATABASE_URL, QDRANT_URL, ENVIRONMENT=production, CORS_ORIGINS)frontend/.env.example to frontend/.env.local and set:
NEXT_PUBLIC_API_URL — backend URL (default: http://localhost:8080)NEXT_PUBLIC_APP_URL — frontend URL for metadataNEXT_PUBLIC_PLAN_FREE_DISPLAY_NAME / BASIC / PRO / PROMAX — plan display name customization (default: S/M/L/XL)Works with Claude Code, Claude Desktop, Claude Chat, ChatGPT, Gemini CLI, and any Streamable-HTTP MCP client.
Claude Code (3 steps):
http://localhost:3000/workspace/integrations/api-keys to create an API key.mcp.json.example to .mcp.json and fill in your workspace ID and API key:.mcp.json.example ships with the all-tools URL. Set "url" to one of:
http://localhost:8080/mcp/w/{workspace_id}http://localhost:8080/mcp/w/{workspace_id}?profile=corePick core when your client loads every tool schema at session start (it is about 65% smaller). It lists the 12 memory and context tools and leaves out Sleep, analyses, files, edges, secrets, resources and the agent control plane — those stay callable, they are just not listed; switch back to the default URL to see them. See Tool Profiles.
.mcp.jsonis in.gitignore— never commit it (contains API keys).
Full setup guide — every client, the memory-sync hook, the ready-to-use .claude/ templates, the kagura-memory Claude Code plugin (skills + tool-guardrail hooks), and the WSL2 networking note: MCP Client Setup
64 tools across 13 categories: Memory (remember / recall / explore …), Agent Substrate (pinned + time-triggered delivery, state, measurements, feedback), Agent Control Plane (preview), Neural Edges, Contexts, Tags, Files (R2), Analyses (Memory Analysis), Resources, Secrets (zero-knowledge), Sleep Maintenance, Usage, and API-Key Bindings — each with per-role access control.
Tool-by-tool reference with required roles: MCP Tools Reference
A client does not have to list all 64: the core URL above (?profile=core) lists 12, and ?tools=remember,recall lists exactly the tools you name — see Tool Profiles.
In addition to MCP tools, a full REST API is available:
/api/v1/memory/*)/api/v1/contexts/*)/api/v1/agents/*)/api/v1/files/*, up to 100 MiB); legacy /api/v1/attachments/* routes return 410 Gone/api/v1/analyses/*)/api/v1/resources/*)/api/v1/workspaces/*)/api/v1/admin/*)/api/v1/config/secrets/*)Full API documentation: http://localhost:8080/redoc
Two OAuth2 providers are supported:
GOOGLE_CLIENT_ID and GOOGLE_CLIENT_SECRETGITHUB_CLIENT_ID and GITHUB_CLIENT_SECRETUsers with the same email address across providers share a single account. Password + MFA login is available without any OAuth provider (see Quick Start). An existing account can also add a password from its profile and then sign in with its verified email address and that password; see Deployment → Email + password sign-in.
Plans control per-workspace resource limits (contexts / memories / MCP calls per day). Four tiers ship by default: S (free), M (basic), L (pro) and XL (promax). For self-hosted single-user setups, assign the XL (promax) plan to your workspace — it is the only tier that may create resources, connectors and public contexts (numeric limits are env-overridable; a tier's feature set is not). Defaults, environment-variable overrides, and optional Stripe self-service billing: Deployment → Plan Tiers
This project is designed to be developed with Claude Code and Kagura Memory Cloud itself — pre-configured slash commands, safety hooks, sub-agents, and rules load automatically from .claude/. Setup and the full tooling reference: Contributing → Development with Claude Code
API reference — two complementary entry points:
http://localhost:8080/redoc — auto-generated from FastAPI, always in sync with the running backendConcepts & guides:
KaguraClient (MCP) plus REST clients for resources, files, secrets, workspaces, and agent bootstrap, and a document FileIngestorSee CONTRIBUTING.md for development setup, code style, and PR workflow.