Local-first memory for Claude Code and any MCP client: hybrid search + knowledge graph, $0/token.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
π‘ Paste into ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows)
A memory server for Claude Code and any other MCP client. It gives your AI assistant a permanent, searchable memory that lives in one SQLite file on your machine. Store a decision today, ask about it next month, and the answer comes back. Everything runs locally: the embedding model, the search index, the knowledge graph. No cloud account, no API key, no per-token cost.
License: source-available and free for noncommercial use (PolyForm Noncommercial 1.0.0): personal projects, hobby, study, research, charity, education, and government. Commercial use requires a paid license (COMMERCIAL.md).
Who it's for: developers who want Claude (or Cursor, Codex, any MCP client) to remember decisions across sessions. Solo builders and hobbyists use it free. Teams share a knowledge base over git. And anyone who wants to replace a cloud memory service (mem0, Zep, Letta, Supermemory) with something that runs entirely on their own machine.
Run npx mcp-memory-graph serve and you get a local web dashboard for browsing and searching your memory outside Claude.

Search works by meaning, not keywords. The query below ("how do we handle payments") finds the Stripe, GDPR, and Postgres notes even though none of them contains that phrase β each result carries a confidence score and a match-type badge:

Browse and sort the whole store in one table β scope, type, tags, quality score, and how often each memory has been read:

mem0, Zep, Letta, and Supermemory are the usual names for AI memory, and several of them have open-source cores. This one is built around a different default: nothing leaves your machine and there's no infrastructure to run.
| MCP Memory Graph | Typical hosted memory service | |
|---|---|---|
| Where it runs | One SQLite file on your machine | A managed cloud service (some also self-host) |
| Embeddings | Local model in Node (MiniLM), no API key | Usually a cloud embedding API |
| Cost per token | $0 β nothing is metered | Usage-based, or a server you operate |
| Extra infrastructure | None | Often Postgres/pgvector, Redis, or a Python service |
| Claude Code integration | First-class: hooks auto-capture and recall | Manual wiring |
| Benchmarks | Committed corpus + runner, reproducible locally | Mostly self-reported |
The trade-off is honest: a single-process SQLite server tops out in the low hundreds of thousands of vectors (see Limitations), and a hosted service will scale past that without you thinking about it. If you're a solo developer or a small team who wants memory that's private, free, and zero-ops, that ceiling is rarely the thing you hit first.
AI assistants forget everything between sessions. Your decisions, your patterns, the bug you fixed last Tuesday: all gone when the conversation ends. This server fixes that.
claude -p) for learning extraction. You can turn that off with review_on_stop: false.You need Node.js 20 or newer and Claude Code installed.
1. Get the server. From npm (easiest):
Or from source:
2. Register the server with Claude Code (optional β init in step 3 does this for you at user scope):
3. Install the hooks (recommended):
This is the one command that wires everything up: it registers the MCP server (user scope), installs the auto-capture/recall hooks and the usage skill, writes config, and schedules a nightly cleanup. Answer the prompts, or pass --yes to accept the defaults. (Skip the auto-registration with --no-register if you manage claude mcp yourself.)
4. Try it. Open a Claude Code session and say:
Then, in a later session:
Claude searches its memory and answers with the stored decision. That's the whole loop.
5. Verify the install. Ask Claude:
It should list all 51 tools (45 memory_*, 3 vault_*, 3 core_memory_*).
The first time a memory tool runs, the embedding model (about 30 MB) downloads from HuggingFace and is cached at ~/.cache/huggingface/. Every start after that is instant.
To undo everything: npx mcp-memory-graph uninstall.
Upgrading the package updates the code that runs each session (hooks, tools, the server), so server-side fixes apply the next time a tool runs β nothing else needed for those.
But files that init wrote earlier are not rewritten by a package upgrade: the Claude Code hook registrations in settings.json and the macOS launchd plist at ~/Library/LaunchAgents/com.mcp-memory.consolidate.plist. If you installed before 2.6.3, that plist used a bare node that launchd (whose minimal PATH excludes nvm) could not run β so the nightly consolidation silently never fired. Re-run npx mcp-memory-graph init once after upgrading to regenerate it with an absolute node path and an output log. Verify it then runs:
To clear conflict noise that accumulated while the job wasn't running: npx mcp-memory-graph consolidate.
When you store a memory, the server turns the text into a vector (a list of 384 numbers that captures its meaning) using a small model that runs inside Node.js. It also indexes the text for keyword search. Both live in one SQLite file, by default at ~/.mcp-memory/memory.db.
When you search, the server runs both kinds of search at once, merges the rankings, and returns the best matches with a confidence label. A second model can then re-sort the top results for better precision (this is the reranker, on by default for MCP clients, and it costs about 200 ms).
On top of that sits a knowledge graph: memories link to entities and to each other, so the server can answer questions that need more than one hop, like "what does the payment service depend on?". A nightly "dream cycle" deduplicates, re-scores, prunes, and reports gaps.
Every number below was produced locally: real embedding model, real production handlers, no network. You can rerun all of them on your own machine.
A quick primer if benchmarks are new to you. A gold set is a list of questions where the right answer is known in advance. Precision@1 asks: was the top result the right one? Recall@5 asks: was the right answer anywhere in the top 5? MRR (mean reciprocal rank) rewards putting the right answer near the top. The reranker is a second model that re-sorts the top 50 results; it is slower but noticeably more accurate.
| precision@1 | precision@3 | MRR | search p95 | |
|---|---|---|---|---|
| Hybrid (RRF) | 0.563 | 0.750 | 0.704 | ~4 ms |
| + cross-encoder rerank (MCP default) | 0.813 | 0.875 | 0.867 | ~230 ms |
Reproduce with npm run bench. Full methodology, the gold set itself, and every miss are printed and documented in docs/BENCHMARKS.md.
With the real embedder and a file-backed SQLite database, retrieval p95 is 9.1 ms at 10,000 vectors and 30 ms at 50,000. The rerank pass adds a roughly constant 200 ms on top. Most memory products publish self-reported, cloud-hosted numbers; these are measured locally and reproducible from a committed corpus and runner.
Four public memory benchmarks, run untuned (stock MiniLM embedder, production handlers, zero benchmark-specific tweaks), matching or beating MemPalace on all four:
| Benchmark | Our result | Comparison |
|---|---|---|
| LongMemEval-S | R@5 = 97.8% | vs 96.6% published |
| ConvoMem | R@10 = 93.5% | vs 92.9% |
| LOCOMO | session R@10 = 82.2%, R@50 = 100% | vs 60.3% baseline |
| MemBench | hit@5 = 78.7% | vs their 80.3% tuned |
Run them yourself: npm run bench:longmemeval, bench:locomo, bench:convomem, bench:membench. The honest notes (where the reranker helps and where it hurts, the dedup floor on MemBench, gold-set size caveats) are in docs/BENCHMARKS.md.
rerank: true adds the cross-encoder pass. use_graph: true blends in HippoRAG Personalized PageRank multi-hop scores. as_of: <timestamp> searches the graph as it stood at a past moment.global, project, user, team, department.importance_score and confidence_score on every memory, from access frequency, recency, and content signals.claude -p reviews the transcript and stores zero to five curated learnings. (This replaces the older type: "agent" Stop hook, which is silently broken on macOS; see anthropics/claude-code#39184.)Five opt-in hooks, installed by init:
| Hook | When it fires | What it does |
|---|---|---|
| SessionStart | session begins | Status check (memory count, expired, stale docs) and surfaces the top memories for the project |
| UserPromptSubmit | each prompt that carries a task signal (a ticket/PR id or β₯2 keywords) | Keyword-searches the store and surfaces matching memories so you recall prior work before re-deriving it; stays silent on trivial prompts |
| PostToolUse | after a memory search | Tracks hits and misses to search-log.jsonl |
| PreCompact | before context compression | Optional learning extraction (off by default) |
| Stop | session ends | Spawns headless claude -p to review the session and store learnings |
The Stop hook detaches in about 30 ms and reviews in the background for 10 to 60 seconds. It needs the claude CLI on $PATH (or $CLAUDE_BIN), authenticated. Turn it off with review_on_stop: false in ~/.mcp-memory/config.json.
| Field | Purpose | Examples |
|---|---|---|
scope | Isolation level | global, project, user, team, department |
namespace | Sub-scope grouping | "my-project", "legal-team", "q4-audit" |
department | Organizational unit | legal, engineering, hr, sales, finance |
document_type | Content classification | contract, policy, code, incident, decision, report |
access_level | Data sensitivity | public, internal, confidential, restricted |
tags | Flexible categorization | ["renewal", "notice-period", "compliance"] |
language | Content language (ISO 639-1) | "en", "da", "de" |
source | Origin | file path, URL, system name |
author | Creator | person or system name |
metadata | Domain-specific JSON | {contract_type: "NDA", parties: ["A","B"]} |
expires_at | Auto-expiration date | ISO 8601 timestamp |
scope and namespace group content within one database. A shared-database MCP_API_NAMESPACE pin gives supported per-namespace multi-tenant isolation (schema v14); a separate database file per tenant is the strongest boundary. See docs/MULTI-TENANCY.md.
valid_from, valid_to) alongside transaction-time. Updates invalidate rather than delete: the prior fact gets a valid_to stamp instead of being overwritten, so history is never lost. Reads default to currently valid rows but accept as_of: <timestamp> for point-in-time recall. memory_history returns one memory's full timeline.memory_graph traverses entities and relationships up to 3 hops. memory_extract_entities stores LLM-extracted entities and relationships.use_graph: true on search runs Personalized PageRank over the entity and link graph for associative retrieval.memory_query answers a question with a tight subgraph. It seeds from hybrid search, walks the graph up to max_hops while avoiding hubs, and returns a token-budgeted context string instead of flooding the window.memory_communities finds densely connected entity clusters, for "what are the main themes in here?" questions.on_conflict), so new facts reconcile with existing ones instead of piling up duplicates.stability signal, so rarely reinforced knowledge slowly sinks in ranking, the way human memory fades.(scope, namespace) that the agent maintains itself (core_memory_get, core_memory_append, core_memory_replace). Appends that would overflow are refused, which forces deliberate compaction.memory_tiers reports a MemGPT-style hot / recall / archival distribution and lists the hot working set.memory_reflect gathers the most reflection-worthy memories and, in store mode, persists synthesized insights linked back to their sources.vault_sync reads a vault in. memory_export_vault writes memories out as .md files with YAML frontmatter that round-trips losslessly for every authored field (id, scope, namespace, tags, access_level, importance, timestamps). Two derived scores are not in the frontmatter and reset on re-import: confidence_score (to 0.6) and stability (to 1.0). Use memory_export (JSON) for a byte-perfect backup. One metadata key is reserved: metadata._vault holds internal sync bookkeeping and never appears in tool output or exported files.memory_canvas exports the graph as a JSON Canvas 1.0 .canvas file that opens as a spatial board in Obsidian.serve exposes /publish/:namespace (index, page, search, graph) as a read-only wiki. It is deliberately not behind bearer auth, but is hard-scoped to published access levels (MCP_PUBLISH_ACCESS_LEVELS, default public).memory_session_note appends to one "daily note" per session. memory_template returns structured note scaffolds per document type.memory init wizard: interactive setup (or --yes for defaults) that writes ~/.mcp-memory/config.json (or project-scoped config) plus the Claude Code wiring.memory export-graph writes a deterministic memory-graph.json you can commit and share. memory git-setup installs a .gitattributes entry and the memory-union merge driver so parallel commits merge instead of conflict.MCP_AGENT_ID (or pass agent_id per store) and memory_attribution reports how many valid memories each agent wrote.memory_questions surfaces what the graph is well placed to find: ambiguous links to confirm, frequently mentioned but under-documented entities, orphaned and stale memories.memory_forget soft-deletes by default (a tombstone via valid_to, recoverable, still visible via as_of). With hard: true it returns a portability export first, then permanently erases. memory_delete is unchanged.The server ships a browser dashboard for viewing and managing memories outside Claude. It runs on the same Express server as the MCP HTTP transport, so there is no separate process.
Six pages:
Tech: React 19, Vite, Tailwind CSS v4, shadcn/ui, Fuse.js, D3, Recharts.
Run it:
Docker: the image includes the built frontend. After docker compose up, the dashboard is at http://<host>:3200 alongside the MCP endpoint. Team members can browse the shared store from any browser, no Claude Code required.
The REST surface is for reading and managing. Creating memories goes through MCP (memory_store over POST /mcp); there is deliberately no POST /api/memories.
| Method | Path | Description |
|---|---|---|
GET | /api/stats | Memory counts and breakdowns |
GET | /api/search?q=... | Hybrid search with filters |
GET | /api/memories | List with pagination and sorting |
GET | /api/memories/:id | Single memory with metadata |
GET | /api/memories/:id/versions | Version history |
GET | /api/memories/:id/related | Semantically related memories |
PATCH | /api/memories/:id | Update content or metadata |
DELETE | /api/memories/:id | Delete a memory |
GET | /api/graph | Nodes and edges for graph visualization |
GET | /api/manifest | Integrity manifest (merkle root plus per-memory hashes) |
GET | /api/insights | Trends and themes summary |
GET | /api/health | Knowledge-gap report (recurring zero-result searches) |
GET | /api/webhooks | List webhook targets (gated by MCP_WEBHOOKS) |
POST | /api/webhooks | Register an SSRF-validated outbound target |
DELETE | /api/webhooks/:id | Remove a webhook target |
POST | /api/webhooks/dispatch | Drain the durable, HMAC-signed delivery queue |
The first nine are what the dashboard uses. All REST endpoints call the same handlers as the MCP tools; no business logic is duplicated.
The server tracks how knowledge is used, scores quality, learns from sessions, and consolidates itself over time.
Every memory gets an importance_score between 0 and 1:
Recency factor:
| Age | Factor |
|---|---|
| < 7 days | 1.0 |
| < 30 days | 0.7 |
| < 90 days | 0.4 |
| > 90 days | 0.1 |
Memories that are never accessed gradually lose importance. Auto-extracted memories start lower and get pruned if they never prove useful.
Note on access reinforcement. The formula above is the periodic recompute run by the consolidate Score stage. Each read (
memory_get,memory_search,memory_related) also applies a small immediate boost (importance_score += 0.03, capped at 1.0), and search uses importance as a mild rank multiplier (1 + importance * 0.5). A memory read 20 or more times approaches the ceiling from reads alone, and consolidate re-baselines it on the next run. This popularity weighting is intentional. If you want a fixed value that reads don't drift, set an explicitimportance_scoreonmemory_storeormemory_update.
When a search returns nothing, the query is logged. The dream cycle's gap stage surfaces these, so you can see what's missing from the store.
claude binary on $PATH (or $CLAUDE_BIN), authenticated without prompting. Optional; disable with review_on_stop: false.init doesUser scope writes hooks to ~/.claude/settings.json, so they fire in every Claude Code session. Project scope writes hooks to .claude/settings.json in the current directory and creates .mcp.json for automatic server discovery; collaborators who clone the project get the memory server registered automatically.
Init does seven things:
dist/hooks/.~/.mcp-memory/config.json (user scope) or <project>/.mcp-memory/config.json (project scope; the generated .mcp.json pins it via MCP_MEMORY_CONFIG_PATH)..claude/CLAUDE.md (project scope) or prints a snippet (user scope).claude mcp add -s user memory-server -- npx -y mcp-memory-graph for you (idempotent; best-effort β warns with the manual command if the claude CLI isn't on PATH; skip with --no-register). Project scope is registered via the committable .mcp.json instead. This makes step 2 of the Quick Start optional.mcp-memory-graph usage skill into ~/.claude/skills/ so Claude Code has inline guidance for all 51 tools, gotchas, and workflows. Skip with --no-skill.Under a non-interactive shell (agent/CI) the wizard is bypassed: defaults are applied and a report is printed showing what was set and how to change each value. Passing --yes applies the defaults silently (no report).
Key flags: --scope user|project, --schedule HH:MM[,HH:MM] (nightly consolidation time, default 03:00), --vault <path> (enable Obsidian vault round-trip), --no-review-on-stop (disable the end-of-session learning review), --no-skill (skip skill install), --no-register (skip the user-scope claude mcp add), --remote <url> (team server mode).
npx mcp-memory-graph uninstall reverses everything init did: removes hooks, the nightly schedule, the CLAUDE.md block, and the installed skill.
Every step is scriptable. There is no interactive-only path:
Claude Code gets the hooks; everyone else gets the same 51 tools, driven manually. The server is a standard MCP server, so any client works. A line in the client's rules file makes usage near-automatic.
Register the server. Example for Codex, in ~/.codex/config.toml (global) or .codex/config.toml (project, trusted only):
Or codex mcp add memory-graph -- node /abs/path/to/mcp-memory-graph/dist/index.js. Cursor, Windsurf, and other clients use their own MCP config format, but the server command (node .../dist/index.js) and the HTTP option are the same.
Then nudge the agent in its instructions file (Codex: AGENTS.md; Cursor: project rules):
Before answering questions about architecture, decisions, patterns, or past fixes, call
memory_searchon the memory-graph server first; store new decisions, patterns, and fixes withmemory_store.
The server runs three ways, from a single-user cache to a knowledge base shared across many machines. All three are local-first: nothing leaves the machines you choose to run it on.
npx mcp-memory-graph init registers a local stdio server plus the hooks. Memory lives in one SQLite file on your machine. Nothing else to run. Right choice for solo use.
Run one server that many clients connect to over HTTP. Everyone shares the same memory base, live.
Start the server (pick one):
Set MCP_AUTH_TOKEN whenever the server is reachable beyond loopback. It is a shared bearer token, one secret for all clients. The server refuses to start unauthenticated on a non-loopback bind unless you set MCP_AUTH_OPTIONAL=1. Terminate TLS at a reverse proxy or tunnel for anything off-host.
Connect a client, one command per machine:
For Claude Code this writes a project .mcp.json pointing at the shared server. The token is stored as an env-var reference ("Authorization": "Bearer ${MEMORY_MCP_TOKEN}"), so the committed .mcp.json never contains the secret. Non-Claude clients point at the same server through their own MCP config.
| Flag | Effect |
|---|---|
--token-env <NAME> | Reference this env var for the token (default MEMORY_MCP_TOKEN) |
--token <value> | Inline a literal token instead (avoid committing it) |
--no-auth | Omit the auth header (loopback or trusted network only) |
In remote mode the local capture and recall hooks are not installed. The memory lives on the server, not in a local file the hooks could read. The agent uses
memory_searchandmemory_storedirectly (the CLAUDE.md guidance is still written).
Prefer your knowledge base in git, reviewed through pull requests, with no server to run? Export memories to plain Markdown and share the folder as a git repo:
Each collaborator must run
vault-initonce in their own clone. The merge driver and post-merge rebuild hook live in local git config (.git/), not in the repo. A fresh clone withoutvault-initwill hit raw conflict markers in.memory/graph.jsonon its first concurrent pull. Re-runningvault-initis idempotent and does not clobber the committed sidecar.
Two recovery notes for team vaults:
memory rebuild can refuse with VaultIntegrityError because .memory/manifest.json is stale. Delete that file and re-run rebuild; it is derived state and regenerates..md while your database has newer state? Import first (vault_sync or rebuild), then export (memory sync). A full export from a stale database overwrites vault files, including your hand edit.MCP_AUTH_TOKEN is a single shared secret, fine for a trusted group; rotate it by restarting the server with a new value. For per-key RBAC (one server, N keys, each pinned to a namespace set and an access-level ceiling) use memory keys create|list|revoke (schema v16). The legacy shared token still works and is checked first. See docs/MULTI-TENANCY.md.--remote default keeps it in an env var by design.127.0.0.1 (the default) unless you front the server with a proxy that terminates TLS; then set MCP_BIND=0.0.0.0.Building an org-wide AI brain? One server, a key per employee, an org chart the AI can traverse (people, teams, SOPs, and tools as typed graph nodes), with enforced who-sees-what. The recipe, built on existing primitives, is in docs/ENTERPRISE-BRAIN.md.
| Variable | Default | Description |
|---|---|---|
MCP_MEMORY_DB_PATH | ~/.mcp-memory/memory.db | Database file location. The directory is created automatically. |
MCP_MEMORY_MODEL | Xenova/all-MiniLM-L6-v2 | HuggingFace embedding model name. Must be an ONNX model compatible with Transformers.js. |
MCP_MEMORY_DIMENSIONS | 384 | Embedding vector dimensions. Must match the model's output. |
MCP_MEMORY_CONFIG_PATH | ~/.mcp-memory/config.json | Override location for the configuration file. |
The full env reference (auth, rate limits, webhooks, vault, publish) is in docs/ENV.md.
Model identity is recorded and enforced. The database remembers which embedding model built it (
schema_meta.embedding_model). Starting the server with a differentMCP_MEMORY_MODELfails loudly instead of silently degrading every search (same dimension does not mean same vector space). To switch models: set the new model and runmemory rebuild(re-embeds from the vault), or export and re-import.
The config file controls self-improvement behavior, hook settings, and per-project overrides. Resolution order: MCP_MEMORY_CONFIG_PATH env, then <cwd>/.mcp-memory/config.json (project-scope init writes this), then ~/.mcp-memory/config.json. Created by npx mcp-memory-graph init, or write it by hand:
| Section | Key | Default | Description |
|---|---|---|---|
defaults | scope | "project" | Default scope for new memories |
defaults | namespace | "auto" | Default namespace ("auto" derives from project directory name) |
projects[] | path | Project root directory | |
projects[] | namespace | Namespace override for this project | |
projects[] | watch | Glob patterns for files to track for changes | |
consolidation | similarity_threshold | 0.85 | Cosine similarity threshold for deduplication (0.5-1.0) |
consolidation | prune_after_days | 30 | Days before pruning low-quality memories |
consolidation | min_importance_to_keep | 0.1 | Minimum importance score to survive pruning |
consolidation | max_operations | 100 | Max operations per consolidation run |
consolidation | schedule | [{ "hour": 3, "minute": 0 }] | One or more { hour, minute } entries (24-hour). Re-run init after changing to regenerate the launchd plist. |
hooks | extract_on_compact | false | Mine transcript before context compression (regex-based, off by default) |
hooks | extract_on_session_end | false | Extract learnings when session ends (regex-based, off by default) |
hooks | track_searches | true | Log search hits and misses to search-log.jsonl |
hooks | review_on_stop | true | Spawn headless claude -p at session end to review the transcript and store learnings. Set false to disable without removing the hook. |
extraction | categories | ["decision", "pattern", "error_fix", "convention"] | Learning categories to extract |
extraction | min_confidence | 0.4 | Minimum confidence for extracted learnings |
storage | db_path | scope-dependent | SQLite file location (~/.mcp-memory/memory.db for user scope, <project>/.mcp-memory/memory.db for project scope). MCP_MEMORY_DB_PATH overrides. |
vault | path | unset | Obsidian vault root used by vault_sync, memory_export_vault, and rebuild when no explicit path is passed. MCP_VAULT_PATH and --vault <path> override. |
vault | write_through | true | Mirror memory writes out to the vault as .md files when a vault is configured. MCP_VAULT_WRITE_THROUGH=0 overrides. |
| Command | Description |
|---|---|
npx mcp-memory-graph | Start the MCP server on stdio (default) |
npx mcp-memory-graph serve | Start the HTTP server: MCP transport, REST API, web dashboard |
npx mcp-memory-graph init | Interactive setup wizard: hooks, config, nightly schedule (user scope). Add --yes/-y for non-interactive |
npx mcp-memory-graph init --scope project | Setup for the current project only (creates .mcp.json and .claude/settings.json) |
npx mcp-memory-graph uninstall | Reverse init: remove hooks and schedule |
npx mcp-memory-graph consolidate | Run the dream cycle manually |
npx mcp-memory-graph export-graph [--out <path>] [--scope <s>] [--namespace <n>] | Write a committable, deterministic memory-graph.json for git sharing |
npx mcp-memory-graph git-setup | Install the .gitattributes entry and memory-union merge driver for conflict-free graph sharing |
npx mcp-memory-graph merge-graphs <ours> <theirs> <out> | Git union merge driver for memory-graph.json (invoked by git, not by hand) |
npx mcp-memory-graph vault-init [--vault <path>] | Make the vault a git repo: union merge driver, pull.rebase=false, post-merge and post-checkout rebuild hooks |
npx mcp-memory-graph sync | Export all valid memories plus the graph sidecar to the vault (.md files) |
npx mcp-memory-graph rebuild [--vault <path>] | Rebuild the SQLite index from the vault's .md files (collaborators run this after git pull) |
npx mcp-memory-graph migrate | Upgrade the database to the current schema version |
npx mcp-memory-graph backup [--out <path>] | WAL-safe online snapshot (retention: MCP_MEMORY_MAX_BACKUPS, default 10) |
npx mcp-memory-graph keys create|list|revoke | Per-key RBAC: mint, inspect, revoke API keys (namespace set plus access ceiling) |
memory_storeStore a new memory. The vector embedding is generated automatically.
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
content | string | Yes | The text content to store | |
title | string | No | Short title for the memory | |
scope | enum | No | globalΒΉ | global, project, user, team, department |
namespace | string | No | ΒΉ | Sub-scope (e.g., project name) |
importance_score | number | No | computed | 0-1 manual importance override |
agent_id | string | No | MCP_AGENT_ID env | Attribution for memory_attribution rollups |
on_conflict | enum | No | add | add, supersede, skip: write-gate behavior on near-duplicates |
document_type | string | No | contract, policy, code, incident, decision, etc. | |
source | string | No | Where this content came from | |
author | string | No | Who created it | |
department | string | No | legal, engineering, hr, sales, finance | |
tags | string[] | No | Tags for categorization | |
access_level | enum | No | internal | public, internal, confidential, restricted |
language | string | No | en | ISO 639-1 language code |
metadata | object | No | Domain-specific key-value pairs | |
expires_at | string | No | ISO 8601 expiration date |
ΒΉ When omitted, a loaded config file's defaults.scope and defaults.namespace ("auto" = project directory name) apply first; the hardcoded fallback is global with no namespace.
Example prompt:
memory_searchHybrid vector plus keyword search across stored memories.
How it works:
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
query | string | Yes | Natural language query or keywords | |
scope | enum | No | Filter by scope | |
namespace | string | No | Filter by namespace | |
department | string | No | Filter by department | |
document_type | string | No | Filter by document type | |
tags | string[] | No | Filter: must contain ALL specified tags | |
access_level | enum | No | Filter by access level | |
language | string | No | Filter by language | |
limit | number | No | 10 | Max results (1-100) |
offset | number | No | 0 | Pagination offset |
search_mode | enum | No | hybrid | hybrid, vector, or keyword |
temporal_decay | object | No | {type: "exponential", half_life_days: 30} or {type: "linear", max_age_days: 365} | |
date_from | string | No | Only memories after this date | |
date_to | string | No | Only memories before this date | |
min_confidence | number | No | Minimum confidence threshold (0-1) |
Example prompts:
Each result includes the memory content and metadata, the combined RRF score, a normalized confidence (0-1), a confidence_level label (high at 0.7 and above, medium at 0.4 and above, low below that), and a match_type (hybrid, vector, or keyword).
The default
detail_level: "summary"projection returnsconfidence_levelbut omits the numericconfidenceand the fullcontent, to save tokens. Passdetail_level: "full"when you need them.
memory_getRetrieve a specific memory by ID. For ingested documents, optionally include all child chunks.
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
id | string | Yes | Memory UUID | |
include_chunks | boolean | No | false | Include child chunks for ingested documents |
memory_updateUpdate an existing memory. If content changes, the embedding regenerates automatically. The previous version is saved to history.
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
id | string | Yes | Memory ID to update | |
content | string | No | New content (triggers re-embedding) | |
title | string | No | New title | |
metadata | object | No | Replacement metadata | |
tags | string[] | No | Replacement tags | |
expires_at | string/null | No | New expiry, or null to remove | |
changed_by | string | No | Who made this change |
memory_deleteDelete memories by ID or by filter. At least one of id or filter is required.
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | No | Delete a specific memory |
filter.scope | enum | No | Delete all in scope |
filter.namespace | string | No | Delete all in namespace |
filter.department | string | No | Delete all in department |
filter.before_date | string | No | Delete older than date |
filter.expired_only | boolean | No | Only delete expired memories |
memory_listBrowse memories with filtering, pagination, and sorting.
| Parameter | Type | Default | Description |
|---|---|---|---|
scope | enum | Filter by scope | |
namespace | string | Filter by namespace | |
department | string | Filter by department | |
document_type | string | Filter by type | |
limit | number | 20 | Max results (1-100) |
offset | number | 0 | Pagination offset |
sort_by | enum | created_at | created_at, updated_at, or title |
sort_order | enum | desc | asc or desc |
memory_ingestIngest a full document: it is chunked by content type, each chunk is embedded, and everything is stored with parent-child relationships. Use this for large documents.
| Parameter | Type | Default | Description |
|---|---|---|---|
content | string | Full document text (required) | |
title | string | Document title | |
content_type | enum | text | Chunking strategy: text, markdown, code, legal, structured |
chunk_size | number | 512 | Target chunk size in characters (~4 chars per token) |
chunk_overlap | number | 50 | Overlap between chunks for context |
source | string | Origin file or URL | |
document_type | string | Document classification | |
department | string | Department | |
author | string | Author | |
tags | string[] | Tags | |
metadata | object | Domain-specific metadata |
Chunking by content type:
| Type | Strategy | Splits on |
|---|---|---|
text | Paragraph | Double newlines (\n\n) |
markdown | Heading-aware | #, ##, ### headings |
code | Function-aware | function, class, const, interface boundaries |
legal | Sentence | Period, exclamation, question marks |
structured | Paragraph | Double newlines (same as text) |
memory_relatedFind memories semantically related to a given one. Uses vector similarity, so it finds connections keyword search misses.
| Parameter | Type | Default | Description |
|---|---|---|---|
id | string | Memory ID to find related for (required) | |
limit | number | 5 | Max results (1-50) |
min_similarity | number | Minimum similarity threshold (0-1) |
memory_versionsView a memory's version history. Every update creates a version record.
| Parameter | Type | Default | Description |
|---|---|---|---|
id | string | Memory ID (required) | |
limit | number | 10 | Max versions (1-50) |
memory_statsUsage statistics about stored memories.
| Parameter | Type | Description |
|---|---|---|
scope | enum | Filter stats by scope |
namespace | string | Filter stats by namespace |
department | string | Filter stats by department |
Returns totals for memories, documents, and chunks, breakdowns by scope, department, and type, storage size, and the expired count.
memory_exportExport current memory content as JSON for portability or migration. This is not a full backup: it serializes only currently live, top-level memories. It omits edit history, the knowledge graph, condense-undo originals, ingested child chunks, and soft-forgotten rows. For disaster recovery, copy the SQLite file (cp ~/.mcp-memory/memory.db ..., see the RUNBOOK); embeddings recompute deterministically on import.
| Parameter | Type | Default | Description |
|---|---|---|---|
scope | enum | Filter export | |
namespace | string | Filter export | |
department | string | Filter export |
Max 1000 records per export.
memory_importImport memories from JSON. Each item is embedded and stored.
| Parameter | Type | Default | Description |
|---|---|---|---|
data | array | Array of memory objects (required) | |
overwrite | boolean | false | Overwrite existing IDs |
vault_syncScan an Obsidian vault, parse the markdown, embed and store. See Obsidian Vault Integration below.
vault_statusSync status for a vault: files synced, pending, changed, and the last sync time.
vault_searchHybrid search scoped to one vault's memories.
By default this searches the namespace named after the vault's folder name. Memories exported from another namespace keep their original namespace in frontmatter. If a search over a freshly synced vault returns nothing, pass an explicit
namespace(and/orscope) override.
memory_consolidateThe dream cycle: deduplicate, score, prune, expire, and detect knowledge gaps.
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
scope | enum | No | Limit consolidation to a scope | |
namespace | string | No | Limit consolidation to a namespace | |
similarity_threshold | number | No | 0.85 | Cosine similarity for dedup (0.5-1.0) |
prune_expired | boolean | No | true | Remove expired memories |
prune_low_quality | boolean | No | false | Remove memories below min importance |
dry_run | boolean | No | false | Preview changes without applying |
max_operations | number | No | 100 | Cap on total operations per run |
Five stages run in order: Score (recalculate importance), Expire (enforce expires_at), Prune (drop low-quality when enabled), Dedup (merge near-duplicates), Gaps (surface zero-result searches). Returns a report with counts per stage.
Example prompts:
memory_extract_learningsMine a session transcript for decisions, patterns, error fixes, and conventions using heuristic pattern matching. No external LLM needed.
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
transcript | string | Yes | Session transcript text to mine | |
scope | enum | No | Scope for extracted memories | |
namespace | string | No | Namespace for extracted memories | |
department | string | No | Department for extracted memories | |
tags | string[] | No | Additional tags | |
source | string | No | Source attribution | |
categories | enum[] | No | all | Filter to decision, pattern, error_fix, convention |
auto_store | boolean | No | true | Automatically store extracted learnings |
Extraction looks for decision language ("we decided", "the fix was"), pattern language ("always use", "never do"), error fixes ("the problem was", "solved by"), and conventions ("our convention is", "standard practice"). Each hit is deduplicated against existing memories and stored with a lower initial confidence.
Parameters for the remaining tools are validated by Zod schemas in src/schemas/; each registration's full description lives in src/server.ts.
| # | Tool | Purpose |
|---|---|---|
| 18 | memory_tiers | MemGPT-style hot / recall / archival tier distribution plus the hot working set |
| 19 | memory_export_vault | Write memories out to an Obsidian vault as .md files with YAML frontmatter (reverse of vault_sync) |
| 20 | memory_canvas | Export the graph as a JSON Canvas 1.0 .canvas for Obsidian |
| 21 | memory_manifest | Lightweight content-free index (titles, types, tags, scores) to discover what exists |
| 22 | memory_graph | Query the knowledge graph: entities, relationships, linked memories, multi-hop traversal (depth 1-3) |
| 23 | memory_extract_entities | Store LLM-extracted entities and relationships for a memory |
| 24 | memory_condense | Apply agent-generated summaries to condense old memories (original preserved) |
| 25 | memory_restore | Restore a condensed memory to its original content and re-embed |
| 26 | memory_query | Answer a question with a tight, token-budgeted subgraph instead of flooding context |
| 27 | core_memory_get | Read the pinned, always-in-context core-memory block for a (scope, namespace) |
| 28 | core_memory_append | Append to the core-memory block (refused if it would overflow char_limit) |
| 29 | core_memory_replace | Replace text in the core-memory block (used to update or compact it) |
| 30 | memory_reflect | Generative-Agents-style reflection: gather material, or store a synthesized insight |
| 31 | memory_communities | GraphRAG community detection over the entity graph for corpus-level themes |
| 32 | memory_template | Fetch a structured note scaffold per document type |
| 33 | memory_session_note | Per-session "daily note" (appends to one memory per session_id) |
| 34 | memory_attribution | Roll up how many valid memories each agent_id wrote |
| 35 | memory_questions | "Questions to ask" digest: ambiguous links, under-documented entities, orphans |
| 36 | memory_forget | GDPR-grade forget: soft-delete (recoverable) by default, or hard erase-after-export |
| 37 | memory_history | Point-in-time bi-temporal timeline plus edit-version history for one memory |
| 38 | memory_unlinked_mentions | Entity names mentioned in memory text with no graph edge yet (suggested links) |
| 39 | memory_query_structured | Exact metadata filter query over top-level memories (no semantic ranking) |
| 40 | memory_version_diff | Line-level diff between two stored versions of a memory |
| 41 | memory_version_restore | Roll a memory back to a previous version (snapshots the current one first) |
| 42 | memory_verify | Verify the signed provenance envelope of memories (ed25519 over content_hash plus origin): per-memory ok/unsigned/content_mismatch/bad_signature/untrusted plus a summary. Opt-in signing via MCP_SIGN_MEMORIES; multi-machine allowlist via MCP_TRUSTED_PUBKEYS / trusted_pubkeys |
| # | Tool | Purpose |
|---|---|---|
| 43 | memory_webhook | Manage the event bus (gated by MCP_WEBHOOKS): register, list, delete SSRF-validated outbound targets, or dispatch the durable, HMAC-signed delivery queue (retry, circuit breaker, dead letter). Mutations emit created/updated/superseded/deleted/forgotten events |
| 44 | memory_insights | Advisor digest: unresolved conflicts, stale memories, most-contradicted facts, evidence-less decisions |
| 45 | memory_health | Store health roll-up: live/retired/stale counts, aging buckets, unresolved conflicts, webhook delivery health |
| 46 | memory_revalidate | Change propagation: list stale memories, preview a change's blast radius (dry run), or confirm a memory is current |
| 47 | memory_session_state | Resumable "where was I" session state, save and resume (versioned) |
| 48 | memory_expertise | Per-user expertise profile: observe a topic, get the profile |
| 49 | memory_export_dataset | Export learnings and reflections as JSONL training pairs (pairs/chatml/alpaca) for fine-tuning |
| 50 | memory_lesson | Capture a structured lesson or incident in one call: fills the matching section template (incident β Symptom/Root Cause/Fix/Prevention; lesson β What/Why it matters/How to apply) from your field values and stores it through the normal deduped write path |
The SQLite database is at schema version 18, with automatic forward migration from any earlier version. The core tables:
memories: all memory data, TEXT primary key (UUIDs), parent-child support for document chunks, plus access_count, last_accessed_at, importance_score, and confidence_score.memories_fts: FTS5 virtual table for keyword search with BM25 ranking, synced with the memories table.memories_vec: vec0 virtual table for vector search. 384-dimension float32 embeddings with scope and namespace metadata for pre-filtering.memory_versions: version history for every change.memory_access_log: every search, get, and related-memory access, with timestamps and query context.ingest_source_tracking: ingested files, for change detection on re-ingestion.Later schema versions add the knowledge-graph tables (entities, links, conflicts, communities), webhooks, session state, and the RBAC api_keys table. Every mutation keeps the three core tables in sync atomically inside a SQLite transaction; the repository.ts layer enforces this, and nothing else touches the tables directly.
Engineering:
Legal:
Finance:
HR:
Sales:
Point the server at a vault folder and every markdown file becomes a searchable memory, with frontmatter, tags, and wiki-links extracted. No Obsidian app needed; it reads the files straight from disk.
| Tool | Description |
|---|---|
vault_sync | Scan vault, parse files, embed and store. Incremental (mtime-based). |
vault_status | Sync status: files synced, pending, changed, last sync time. |
vault_search | Hybrid search scoped to a vault's memories. |
What gets extracted:
| Obsidian feature | Memory field |
|---|---|
YAML frontmatter title: | title |
YAML frontmatter tags: [...] | tags (merged with inline) |
YAML frontmatter author: | author |
| YAML frontmatter (all fields) | metadata.frontmatter |
Inline #tags in content | tags (merged with frontmatter) |
[[wiki-links]] | metadata.links array |
| File path relative to vault | source |
| Vault directory name | namespace |
Usage examples:
vault_sync parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
vault_path | string | Absolute path to vault directory (required) | |
chunk_size | number | 1024 | Target chunk size for large files |
chunk_overlap | number | 50 | Overlap between chunks |
force | boolean | false | Re-sync all files regardless of mtime |
include_patterns | string[] | Only sync matching globs (e.g., ["notes/**"]) | |
exclude_patterns | string[] | Skip matching globs (e.g., ["templates/**"]) |
How sync works: it scans the vault recursively for .md files (skipping .obsidian/, .trash/, .git/), compares modification times against the last sync, extracts frontmatter, wiki-links, and tags from new or changed files, embeds, and stores. Deleted files have their memories removed. Files larger than the chunk size are split with markdown-aware chunking. A second sync of an unchanged vault takes under a millisecond.
npx mcp-memory-graph init.npx mcp-memory-graph uninstall.access_level metadata (public, internal, confidential, restricted) for organizational awareness.Backup:
Reset:
When installed via npx mcp-memory-graph init, a nightly job runs all five dream-cycle stages plus access-log rotation (entries older than 90 days are dropped).
On macOS, a launchd plist is created at ~/Library/LaunchAgents/com.mcp-memory.consolidate.plist, scheduled for 3:00 AM. On Linux, init prints a cron suggestion:
Run it manually any time:
MCP_MEMORY_MODEL (with a rebuild), but the shipped benchmarks only validate the default.memory_extract_learnings uses pattern matching, not an LLM. It catches common phrasings and misses subtle ones. (The Stop hook's claude -p review is the LLM-quality path.)What's actually next, in rough order:
.md changes instead of manual rebuild.as_of content reconstruction. Point-in-time queries currently reconstruct validity (which facts were live); reconstructing the content of edited memories at that instant is the remaining half.| Component | Package | Purpose |
|---|---|---|
| MCP SDK | @modelcontextprotocol/sdk ^1.29 | Model Context Protocol server framework |
| Embeddings | @huggingface/transformers ^3.8 | Local ONNX model inference in Node.js |
| Database | better-sqlite3 ^12.10 | Synchronous SQLite with native bindings |
| Vector search | sqlite-vec 0.1.10-alpha.4 | vec0 virtual table for KNN search |
| Validation | zod ^3.25 | Schema validation for tool inputs |
| TypeScript | typescript ^5 | Strict mode, ES2022 target |
| Frontend | React 19, Vite, Tailwind CSS v4 | Web dashboard SPA |
| UI components | shadcn/ui | Accessible component primitives |
| Fuzzy search | fuse.js ^7 | Client-side autocomplete suggestions |
| Graph viz | d3-force, d3-zoom, d3-drag | Knowledge graph layout |
Source-available, not open source. Licensed under the PolyForm Noncommercial License 1.0.0: free for any noncommercial purpose (personal projects, hobby, study, research, charitable, educational, public-research, and government use). Commercial use requires a paid license; see COMMERCIAL.md.
If you're unsure whether your use counts as commercial, check the safe harbors in the license or just ask: yonasmougaard@gmail.com.
MCP memory server Β· Model Context Protocol Β· Claude Code memory Β· persistent AI memory Β· LLM long-term memory Β· AI agent memory Β· local-first memory Β· $0/token memory Β· hybrid vector + keyword search Β· semantic search Β· knowledge graph Β· bi-temporal memory Β· HippoRAG / Personalized PageRank Β· cross-encoder reranking Β· RAG memory Β· SQLite vector database Β· sqlite-vec Β· FTS5 / BM25 Β· local embeddings (all-MiniLM-L6-v2, Transformers.js) Β· Obsidian vault sync Β· JSON Canvas Β· GDPR forget Β· signed provenance Β· self-hosted memory.
Also searched as: a self-hosted, privacy-first alternative to mem0, Zep, Letta, Cognee, and Supermemory Β· long-term memory for Claude / Cursor / Codex Β· an Obsidian-backed knowledge base for AI agents Β· a local knowledge-graph memory that never leaves your machine.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/mcp-memory-graph-2)<a href="https://allmcps.com/mcp/mcp-memory-graph-2"><img src="https://allmcps.com/api/badge/mcp-memory-graph-2?style=directory" alt="Mcp Memory Graph on AllMCPs" /></a>