The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Coldstart listing page.
Self-maintaining codebase knowledge for AI coding agents.
Agent-written notes that stay anchored to your code — plus fast, deterministic navigation — so Claude Code, Codex, and Cursor stop rediscovering the repo every session.
Website · Docs · Blog · Philosophy · npm
Two layers, one tool:
coldstart kb) — durable, agent-written notes about how this codebase actually works: what a file is for, how a flow spans files, which invariants hold. Captured after real tasks, recalled when a later task matches, and kept honest by the index — every note is anchored to real files, and a note whose evidence drifted is flagged, not served as truth.coldstart find / coldstart gs) — a fast static index over file paths, symbol names, exports, and the import/call graph. It answers "which files are relevant to this task?" in milliseconds, with checkable evidence instead of a similarity score.No embeddings, no model to run, no service to babysit. Agents are already good at reading and reasoning about code; what they waste tokens on is finding the right file and re-deriving what the last session already figured out. coldstart does those two parts and gets out of the way.
Requires Node.js 18+.
A single coldstart init does everything — navigation and the notebook. It asks two things — the experience (cli, recommended, or mcp) and the client — then writes the agent-facing guidance into the client's own rules file (as an imported coldstart.md for Claude Code; inlined directly for Cursor and Codex, which don't resolve @file references), wires the client, and sets up the notebook (skeleton, git wiring, and — for Claude Code, Codex, and Cursor — the capture/recall hooks). Pass --experience / --client to skip the prompts. The client is never auto-detected; you always pick it.
coldstart.md and ensures CLAUDE.md imports it via @coldstart.md, and registers both the find/gs search hooks (a PostToolUse nudge + a PreToolUse find-dedup guard) and the notebook recall/capture hooks (UserPromptSubmit + Stop/SubagentStop) in .claude/settings.json — merged into any existing settings, never overwriting them. The mcp experience also writes .mcp.json.AGENTS.md (Codex has no @file include, so there's no separate coldstart.md), refreshed in place on re-run, and registers Codex-specific navigation plus notebook hooks in .codex/hooks.json. The capture hook understands Codex rollout and subagent transcripts. The mcp experience also writes [mcp_servers.coldstart] into .codex/config.toml..cursor/rules/coldstart.mdc — an always-applied rule that carries the full coldstart guidance inline (Cursor doesn't reliably resolve @file references in rules), rewritten on every init — and registers Cursor-specific navigation plus notebook hooks in .cursor/hooks.json (a preToolUse find-dedup guard, a postToolUse nudge, beforeSubmitPrompt recall, and stop/subagentStop capture — merged into any existing hooks). The capture hook parses Cursor's own conversation transcript. The mcp experience also writes .cursor/mcp.json.coldstart.md only, and prints the wiring directions (plus the MCP server entry for the mcp experience).On every client, the PostToolUse hook also delivers the "Edited together" signal without waiting to be asked. Once an agent has edited two different files, it names the files that git history says keep changing alongside them — the sibling implementation, the test in another language, the doc that drifts. Agents don't reliably run gs before editing, so the signal only reached them if they went looking; this puts it in front of them at the edit. It's advisory: the message says outright that the relation is a habit rather than a code dependency, and asks the agent to check rather than to change anything. Each file is named at most once per task, and the list resets whenever you send a new message.
init then warms the index in the background, so your first lookup is instant. Re-running init is safe — it never duplicates entries.
A version stamp in the keeper's lockfile makes the old background keeper shut down on the next lookup; a fresh one spawns from the new binary. No manual restart needed.
[!NOTE] Migrating from
coldstart-mcp: the package was renamedcoldstart-mcp→@cstart/coldstartat 2.0.0 (the CLI is now the primary surface).coldstart-mcpis deprecated but still installs; switch withnpm uninstall -g coldstart-mcp && npm install -g @cstart/coldstart && coldstart init. Thecoldstart-mcpbinary name is kept as an alias, so existing MCP configs keep working.
init writes per-repo wiring that a global npm uninstall can't reach (npm fires no reliable uninstall hook, and it has no record of which repos you init'd). So — like husky — coldstart ships an explicit reverse:
unwire removes only coldstart-owned markers from the files init touched — hook entries, the @coldstart.md import, the AGENTS.md block, the MCP server entry, and files coldstart fully owns (coldstart.md, .cursor/rules/coldstart.mdc) — never your own content in shared files. It sweeps all four clients, is idempotent (a second run reports everything already gone), and keeps the notebook by default since it's committed, shared data. Run it in each project first, then npm uninstall -g @cstart/coldstart to remove the package.
A repo-local knowledge base written and read by agents, in .coldstart/notebook/:
What a note is. Three shapes: a file note (what a file is for — a single summary, or per-symbol facets for hub files), a flow note (a cross-file story: ordered steps, invariants), and a lesson (a trap, rule, bug-cause, rationale, or confirmed absence). Every note carries anchors — concrete file paths and symbols its claims rest on.
Where notes reach the agent. Three surfaces, no new habits required:
Summary: lines on find results — a past agent's verified overview of a file, right where the file ranks. [fresh] means the file is byte-identical to when the summary was verified — the agent can rely on it without re-reading the file.kb search / kb lookup — a search engine over the notebook for mid-task vocabulary changes, and an exact-address lookup (path [symbol]) before editing a file.Why it can be trusted. This is the part that took the design work:
[evidence changed: <path>] and the guidance says re-verify — stale knowledge degrades into a labeled hypothesis instead of a confident lie..raw event log (commit it — merges are unions, so parallel branches of notes reconcile without conflicts). The Markdown notes are derived, regenerated mechanically, and gitignored.--into <id>) or declare it new (--new). Duplicates are gated at write time, not cleaned up later.Setup: the notebook comes with coldstart init — no separate step. It creates the notebook skeleton, sets union-merge for the logs, and (on Claude Code, Codex and Cursor) wires the two hooks — capture at session end, recall at prompt time. (coldstart kb init still exists as an alias if you want to (re-)wire just the notebook.) Other hosts can drive the notebook without the hooks: via the full kb CLI, or — for no-shell clients — the kb_search / kb_lookup / kb_write / kb_status / kb_repair / kb_repair_aliases MCP tools.
Language-agnostic. The notebook's freshness machinery is content-hash based, so it works on any codebase — including languages the navigation index doesn't parse. Where the index does parse, notes additionally get symbol-level freshness.
[!NOTE] The notebook is young. What's verified today: notes written by agents in real sessions checked out accurate against the code; the stale-note loop closes end-to-end (flag → re-read → correction); capture, recall, and concurrent writes hold up under stress. The bet — stated as a bet — is that a corpus like this compounds over a repo's lifetime: the second time any question comes up, the answer is one
Readaway instead of a re-derivation.
| What it answers | Replaces | |
|---|---|---|
find <terms> | "Which files are about this?" — ranks files by how many of your query terms they cover (filenames, path segments, exported symbols, plus a repo-wide name-reference pass). | a flurry of grep/glob while orienting |
gs <file> | "What is this file?" — top-level symbols with line ranges, who imports it, who calls each symbol, and name-related neighbors. | reading a whole file just to learn its shape and usage |
The intended flow: find a concept → pick the best path → gs that file for its shape and who uses it → Read only for the implementation inside a method body. Notebook summaries ride along on find results, so often the orientation step answers itself.
find — locate the files for a concept[!TIP] Pass every salient identifier from your task — the symbol, the domain noun, the rare token you half-remember — not one distilled keyword.
findranks files by how many of your terms each one covers and shows, per file, which terms it defines vs. imports and a preview of the lines where they cluster. Often that's enough to answer without opening anything.
Speed-wise, find competes with raw grep: its repo-wide reference pass runs on ripgrep — yours from PATH, the bundled copy, or an editor's (COLDSTART_RG overrides) — with git grep/grep fallbacks, and the ranked page comes from the pre-built index, not a scan.
Flags: --path GLOB (scope; comma-combine, ! excludes) · --tests (include test files) · --via (show name-reference relations) · --json
gs — drill into one fileReturns the file's symbols (with line ranges), its 1-hop internal imports, who imports it, and per-symbol cross-file callers — in one call. This is the answer to "who uses this file / who calls this symbol"; it is not a grep.
Flags: --symbol a,b (deliver named method bodies inline) · --match TERM (filter a god-file to one area; a|b = OR, /regex/ = regex) · --view symbols|imports|importers|callers · --json
graph — see the whole repo at onceFor humans, not agents. Writes one self-contained HTML file and opens it: every indexed file is a
point on a sphere, positioned by directory, and clicking one opens a 2D view of everything it is
connected to — with each relation named (imports, calls save(), load(),
edited together · 13 of 42 commits, same note · <title>). Click a neighbour to open its
connections; click one of those and the view slides along, so you always see two levels rather than
an ever-growing hairball.
No dependencies. The page is HTML, CSS, and a canvas script with your repo's data baked in — no server, no build step, and nothing leaves your machine. Mail it, drop it in a gist, open it on a plane. Try it on coldstart's own codebase.
Flags: --out PATH (default .coldstart/graph.html, gitignored for you) · --no-open · --json
coldstart ships as one binary with two front doors:
coldstart find … / coldstart gs … / coldstart kb …. For any shell-capable agent (Claude Code, Cursor, terminal use). This is the fast path.find and gs tools, plus the notebook as kb_search / kb_lookup / kb_write / kb_status / kb_repair / kb_repair_aliases, all byte-identical to the CLI. For clients like Claude Desktop that have no shell. (kb commit stays CLI/human-only — publishing notes to git is never an agent action.)Same engine, same index, same results. Pick whichever your agent can reach.
It works best with Claude Code, Codex, and Cursor: all three get platform-specific find/gs hooks and notebook recall/capture hooks from coldstart init. Any other client gets coldstart.md plus printed wiring directions.
coldstart has no embeddings, no generated summaries, no semantic layer computed at index time — on purpose. The semantic layer is the agent. Every consumer is already a frontier model; pre-computing meaning at index time only duplicates that, worse and stale. So the index keeps what's cheap to keep exact — paths, symbols, exports, the import/call graph — and returns why each file ranked.
The notebook is the same philosophy applied to memory: coldstart still computes no meaning of its own. It stores, anchors, and freshness-checks the meaning agents author — written at task time, by the reasoner that had the full context, about the question that actually mattered. The full argument is in PHILOSOPHY.md.
coldstart is one keeper, thin readers:
find/gs) and the MCP server are stateless readers over that cache. The first reader for a repo lazily spawns the keeper, so even uncommitted edits stay live.~/.coldstart/daemon/<root>.log and exits when its lockfile is removed.There is no cache TTL. The index is never discarded for being old — it's kept correct instead:
status shows), and a rotating fingerprint audit after each save catches watcher-missed events.The keeper also stamps the notebook's anchor freshness (a small sidecar, derived single-flight) — the notebook never loads the code index to answer a query.
restart is the right move whenever anything feels stale — a fresh keeper reconciles on start, so it comes back correct, not just alive. status answers "is my index fresh, and why?": liveness, cache age, the keeper's last reconcile/patch/rebuild/save stamps, and the tail of the repair log — no network probe.
Navigation index: TypeScript, JavaScript, JSX/TSX, Vue, Svelte, Astro, AngularJS 1.x, Java, Kotlin, Ruby (Rails-aware: has_many/belongs_to associations, routes.rb resources, controller↔view edges), Python (Django convention edges), Go, Rust, C#, PHP (Laravel convention edges), C++, Groovy (incl. Gradle DSL), GraphQL, YAML, TOML, XML, and .env files.
Not indexed: Swift, Dart — no extension mapping; these files are not walked or parsed.
The notebook works regardless — its freshness stamps are content-hash based, so notes on a Swift repo are as trustworthy as notes on a TypeScript one (they just lack symbol-level freshness detail).
gs gives you the shape.find says "no indexed file contains any of […]" → those identifiers aren't in the repo. Don't grep spelling variants.See PHILOSOPHY.md for why coldstart computes no semantics of its own, ARCHITECTURE.md for the index pipeline, process model, and notebook internals, and TROUBLESHOOTING.md for recovery procedures.
gs callers are one-hop and file-scoped. Member-expression calls (this.method(), api.method()) aren't cross-file resolved; named function/constant calls are. Chase further hops by calling gs on the caller files.import(variable)) and runtime-DSL references (polymorphic associations, gem/reflection-backed models) stay unresolved..raw logs are committed and union-merge across branches and machines.Two repositories, each run more than once, with the CLI. Each question is asked in two arms that differ only in whether coldstart is installed: the baseline is plain Claude Code using the search tools it ships with. Sonnet 5 on both sides, one fresh session per question.
| Repository | Language | Tokens vs. baseline | Recall | Questions |
|---|---|---|---|---|
| Arches | Python / Django | −64% | +2 pts | 27 |
| JMRI | Java | −31% | parity | 25 |
The questions come from closed issue reports, which are written before anyone knows where the fault lives, so they describe a symptom rather than naming a file. The correct answer is the set of source files in the commit that closed the issue, so ground truth comes from git history rather than from anyone's judgement.
Those two numbers are not averaged. Per-turn cost differs by language, and the saving comes from removing turns, so a question an agent can settle in two or three turns has little to give back. An earlier benchmark on Kafka, Django and Mastodon was discarded: the model has read all three, and that contaminated both arms. The full method, including what was thrown out and why, is at coldstartmcp.dev/benchmark. The harness itself is a separate project, coldbench, and works on any repository with git history, including private ones. Both question sets are published with their gold file lists and the issue number behind each question, at coldbench/examples.
Longer pieces on the problems behind this tool — what agent sessions actually cost, and what happened to the design when the measurements disagreed with the plan.
MIT — see LICENSE.