The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Memory for AI listing page.
An MCP server that turns a codebase into a persistent knowledge graph — functions, classes, call chains, HTTP routes, cross-service links — so an AI coding agent answers structural questions with graph queries instead of reading file after file.
Windows native x64 only; MSYS2 CLANG64 is the development toolchain. See support and verification. One self-contained native executable. 162 languages via vendored tree-sitter grammars, refined by embedded Hybrid-LSP type resolution. 23 MCP tools. No language runtime, no Docker, no API key, no telemetry — everything runs locally.
Research — design and evaluation are described in Codebase-Memory: Tree-Sitter-Based Knowledge Graphs for LLM Code Exploration via MCP (arXiv:2603.27277): across 31 real repositories, 10× fewer tokens and 2.1× fewer tool calls vs. file-by-file exploration, at 83% answer quality (92% for the file-by-file baseline).
30-second machine check first (details: docs/INSTALL.md — Preflight):
powershell -Command "$PSVersionTable.PSVersion" must print a version — search_code shells out to PowerShell at runtime; if the WindowsPowerShell\v1.0 directory is missing from PATH, add it before installing.git --version works (watcher freshness + detect_changes).memory-for-ai --version tells you — re-running the installer is the update, indexes survive.Windows (PowerShell):
Then restart your coding agent and say "Index this project". Done.
The installer downloads the verified release archive for your platform, verifies its SHA-256 against checksums.txt, installs the binary, and configures every coding agent it detects (Claude Code, Codex, Gemini CLI, Cursor, VS Code, Windsurf, and ~40 more — see Multi-agent support). Options: --skip-config (binary only), --dir=<path>, --clients=<list>, --project / --name=<name> (per-project mode). Full reference, including all package managers, manual MCP config, containers/CI, and uninstall: docs/INSTALL.md.
Antivirus note: Microsoft Defender may flag a release binary as
Trojan:Script/Wacatac.B!ml— a known false positive (typically 61 of ~62 engines clean; the same family flagsgh, llama.cpp, and Microsoft's own Go toolchain). Evidence and self-verification steps: Antivirus False Positives.
One binary can serve any number of repositories, but sometimes a project deserves its own fenced memory: an MCP server named after the repo, an index no other repo can see, and zero edits to global agent config. That is --project:
What it does: installs/refreshes the shared binary, writes a project-local .mcp.json entry named memory-for-ai-<repo-directory> pinned with --scope=<repo>, and indexes the repository immediately. An agent opened in that repo sees exactly one server serving exactly that graph; opening a different repo sees its own. Use the PowerShell installer for this supported platform. Details and guarantees: docs/INSTALL.md and docs/CONFIGURATION.md.
index_repository parses the whole tree (tree-sitter syntax pass + Hybrid-LSP type resolution), builds a graph of nodes (Function, Class, Route, Package, …) and edges (CALLS, IMPORTS, IMPLEMENTS, DATA_FLOWS, HTTP_CALLS, CROSS_*, …), and persists it to SQLite under ~/.cache/memory-for-ai/. A background watcher re-indexes on Git/filesystem changes. After that, the agent's questions become millisecond graph queries:
There is no built-in LLM: your MCP client is the intelligence layer; this tool is the structural memory. Typical wins:
trace_path, detect_changes follow indexed relationships across files. The counts are exact for stored edges, but event callbacks (such as JSX props) or partially parsed code may not be represented; verify coverage and source before relying on an empty trace.get_architecture: languages, packages, entry points, routes, hotspots, layers, community-detection clusters.query_graph (read-only Cypher subset).manage_adr persists architecture decisions beside it.Measured on Apple M3 Pro (see docs/MEASURING.md to reproduce on your own workload):
| Operation | Time | Notes |
|---|---|---|
| Linux kernel full index | 3 min | 28M LOC, 75K files → 4.81M nodes, 7.72M edges |
| Django full index | ~6 s | 49K nodes, 196K edges |
| Cypher query | <1 ms | Relationship traversal |
| Trace call path (depth 5) | <10 ms | BFS traversal |
| Dead-code detection | ~150 ms | Full graph scan |
Token efficiency (five structural queries on the same repo): ~3,400 tokens via the graph vs ~412,000 tokens via file-by-file exploration. A single-cause measurement recipe and the honest cost model (including the fixed per-session tool-list overhead) are documented in docs/MEASURING.md.
| Document | What it covers | Primary audience |
|---|---|---|
| docs/AGENT_GUIDE.md | Operating manual: mental model, all 23 tools, task→tool playbooks, correctness protocol, per-project tuning | AI coding agents (and their humans) |
| AGENTS.md / docs/DEVELOPMENT-STANDARD.md | Normative AI coding rules, warning policy, verification matrix, and documentation-sync checklist | Contributors and AI coding agents |
| docs/INSTALL.md | Machine preflight (platform, disk, PowerShell/git on PATH, old-version check), every install path: one-liners, per-project, package managers, containers/CI, update/uninstall, build from source, artifact verification | Whoever installs |
| docs/CONFIGURATION.md | Config files, config set keys, environment variables, scoped sessions | Operators, CI authors |
| docs/MEASURING.md | Measuring answer quality, latency/stability, and token/tool-call savings on your repo | Evaluators |
| docs/AB-RESULTS.md | Worked A/B measurement on this repository: 8 questions, graph vs file-by-file, with the freshness incident and honest limitations | Evaluators |
| docs/llms.txt | Machine-readable index of the above | AI agents |
| SECURITY.md | Reporting, release policy, antivirus false positives, supply chain | Everyone |
Every MCP tool also runs as a local one-shot command (no daemon, no standing process; stdout stays machine-clean):
cli <tool> --help prints the flags generated from that tool's input schema. Arguments can also be piped as JSON on stdin.
Tree-sitter gives a syntactic AST; it cannot tell that user.profile.display_name() resolves to Profile.display_name three modules away. memory-for-ai embeds a lightweight C implementation of language type-resolution algorithms — structurally inspired by tsserver/typescript-go, pyright, gopls, Roslyn, Eclipse JDT, and rust-analyzer — that refines call/usage edges on every parse. No language-server process, no per-project setup.
Full type-aware resolution for Python, TypeScript/JavaScript/JSX/TSX, PHP, C#, Go, C/C++, Java, Kotlin, Rust, Perl: imports, generics, inheritance, JSX dispatch, traits/late static binding (PHP), LINQ + records (C#), embedded structs (Go), templates/namespaces (C++), overload + lambda resolution (Java), extension + scope functions (Kotlin), trait methods + UFCS (Rust), MRO + Exporter (Perl). All other grammars fall back to tree-sitter-textual resolution, so every file still produces a graph.
162 languages, all parsed by vendored tree-sitter grammars compiled into the binary. Benchmarked against 64 real open-source repositories (78–49K nodes each):
| Tier | Score | Languages |
|---|---|---|
| Excellent (≥90%) | Lua, Kotlin, C++, Perl, Objective-C, Groovy, C, Bash, Zig, Swift, CSS, YAML, TOML, HTML, SCSS, HCL, Dockerfile | |
| Good (75–89%) | Python, TypeScript, TSX, Go, Rust, Java, R, Dart, JavaScript, Erlang, Elixir, Scala, Ruby, PHP, C#, SQL | |
| Functional (<75%) | OCaml, Haskell |
Also parsed (not yet benchmarked): Ada, Agda, Apex, ArkTS, Assembly, Astro, AWK, Beancount, BibTeX, Bicep, Bitbake, Blade, Cairo, Cap'n Proto, CFML, CFScript, Chialisp, Clojure, CMake, COBOL, Common Lisp, Crystal, CSV, CUDA, D, Devicetree, Diff, Dotenv, Elm, Emacs Lisp, F#, Fennel, Fish, Form, Fortran, Func, GDScript, Git Attributes, Gitignore, Gleam, GLSL, GN, Go Module, Go Template, GraphQL, Hare, HLSL, Hyprlang, INI, ISPC, Janet, Jinja2, JSDoc, JSON, JSON5, Jsonnet, Julia, Just, Kconfig, KDL, Lean 4, Linker Script, Liquid, LLVM IR, Luau, Magma, Makefile, Markdown, MATLAB, Mermaid, Meson, Mojo, Move, NASM, Nickel, Nix, ObjectScript Routine, ObjectScript UDL, Odin, Pascal, Pine Script, Pkl, PL/SQL, PO, Pony, PowerShell, Prisma, Properties, Protobuf, Puppet, PureScript, QML, Racket, Regex, Requirements, ReScript, RON, reStructuredText, Scheme, Slang, Smali, Smithy, Solidity, SOQL, SOSL, Squirrel, SSH config, Starlark, Svelte, Sway, SystemVerilog, TableGen, Tcl, Teal, Templ, Thrift, TLA+, Typst, Verilog, VHDL, Vim script, Vue, WGSL, WIT, Wolfram, XML, Zsh.
install auto-detects and configures 45 supported automatic/conditional client surfaces (39 automatic + 6 conditional/explicit) — Claude Code, Codex CLI, Gemini CLI, Zed, OpenCode, Cursor, VS Code, Windsurf, Kiro, Qwen Code, GitHub Copilot CLI, Junie, Factory Droid, Grok Build, Amp, Devin, and the rest of the matrix in docs/INSTALL.md. It writes only documented MCP entries plus durable instructions, skills, and lifecycle hooks where the client documents a safe contract; it never enables experimental flags, plugins, or permission bypasses. Custom-agent formats receive three tiered graph profiles — Scout (fast provisional discovery), Verify (default, evidence-checked), Auditor (bounded exhaustive verification) — each biased to prove graph evidence against source via check_index_coverage. Preview exactly what would be written on your machine with memory-for-ai install --dry-run.
Claude Code, Codex CLI, Gemini CLI, Zed, OpenCode, Antigravity, Aider, KiloCode, VS Code, Cursor, Windsurf, Augment / Auggie, OpenClaw, Kiro, Junie, Hermes, OpenHands, Cline, Warp, Qwen Code, GitHub Copilot CLI, Factory Droid, Crush, Goose, Mistral Vibe, Grok Build, Qoder CLI, Kimi Code CLI, GitLab Duo CLI, Rovo Dev CLI, Amp, Devin CLI / Local, Tabnine, Continue / cn (conditional), Visual Studio (conditional, Windows), TRAE (conditional), Roo Code (conditional), Amazon Q Developer IDE, CodeBuddy Code CLI, IBM Bob IDE (conditional), IBM Bob Shell, Pochi, Pi, Sourcegraph Cody (explicit opt-in), Oh My Pi (omp).
One session-coordination daemon is shared per account across clients: it owns watchers, shared indexing, and the optional graph UI (memory-for-ai --ui=true --port=9749, then open http://localhost:9749), so concurrent sessions never start duplicate services. All CBM processes must run the same exact build; native install/update/uninstall coordinate a safe account-wide activation window.
This tool reads your codebase and writes your agent configuration — that is its job. All processing is 100% local; there is no telemetry and nothing phones home (cbm makes no network request of its own accord). Every release is verified before publication: VirusTotal scan of all executable candidates (single documented Microsoft !ml tolerance), SLSA Level 3 build provenance (gh attestation verify …), Sigstore cosign keyless signatures, SHA-256 checksums.txt, and CodeQL SAST gating. Full policy and audit trail: SECURITY.md.
MIT — see LICENSE.