The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the CGIS Code Graph listing page.
Ask "what breaks if I change this?" and get the call chain, not a guess.
CGIS parses a repository with tree-sitter into a graph of fully qualified symbols and the calls, imports and containment between them, stores it in SQLite, and serves it to AI agents over MCP. An agent that would otherwise grep and read whole files asks the graph instead.
Real output — CGIS run on its own source.
That ships the MCP server, a skill that teaches the agent when to query the graph instead of reading files, and /cgis:ingest to build the graph on first use. The server is pulled from PyPI on demand via uvx, so there is nothing to clone or build.
Or install it for good: pip install codegraph-brain (Python 3.12+), then use cgis directly. The full command list is in CLI_USAGE.md.
CGIS runs on a working twelve-repository estate — four languages, 8,146 commits, shipping daily. On its FastAPI backend it classifies 82.7% of 87,845 edges definitively and prints the remaining 17.3% rather than inventing targets for them — including the part that is CGIS's own gap.
That share rose from 11.4% in #459, which stopped counting calls to missing symbols as resolved. About 4 points of what is left is a known resolver gap, not something undiscoverable: ingesting app/ strips the app. prefix its imports carry, and the import path does not yet reconcile the two — the same backend ingested at its package root reports 13.2%. The number is what the tool admits it cannot place today, and it is allowed to move the unflattering way.
Read the case study → — every figure measured and reproducible, including what CGIS doesn't cover.
Text retrieval hands an agent chunks that look related. It cannot say which of three functions named save a call reaches, or what sits five callers up. CGIS resolves every call site to a fully qualified name when the source allows it — and when it does not, the edge stays marked unresolved and is counted, never filled with a plausible guess.
| If you use… | CGIS adds |
|---|---|
| grep / file reads in the agent | Transitive callers and callees in one call, without spending context on whole files |
| LSP-backed symbol tools (e.g. Serena) | A persisted whole-repo graph for multi-hop impact, coupling, PageRank and drift |
| A repo map (e.g. aider) | Resolved edges you can traverse and audit, with the resolved/unresolved ratio reported |
The main tools:
| Tool | Answers |
|---|---|
cgis_ingest | Build or incrementally refresh the graph |
cgis_overview | Where to start: sizes and the largest packages, when you have no FQN yet |
cgis_find_symbol | Partial name → candidate FQNs |
cgis_analyze_impact | What breaks upstream if this changes? |
cgis_trace_flow | What does this call, transitively? |
cgis_get_structure | Class / module hierarchy |
cgis_context | A compact GraphRAG context package for one symbol |
cgis_metrics | Coupling, god classes, PageRank, package cohesion |
cgis_audit_reachability | Authz / IDOR coverage — does every handler reach its guard? |
cgis_drift | How far each domain has moved from its declared pattern |
cgis_validate | Graph integrity: resolved vs unresolved edges |
All 14 tools, with parameters: MCP_REFERENCE.md.
ResolverEngine maps raw calls to fully qualified names, or leaves them explicitly unresolved.The details — and a pipeline graph CGIS regenerates from its own source on every change — are in HOW_IT_WORKS.md.
Guardian is CGIS's built-in LLM reviewer — it reviews pull requests using the graph as context, not just the diff text. It runs in CI and posts inline comments anchored to the exact line.
qwen2.5-coder, llama3.1, granite-code, …) for free local inference, or at Mistral / Gemini in the cloud. You can even mix them — a strong cloud finder with a free local cross-model skeptic.No GPU on hand? Benchmark it on a notebook GPU → — free end to end, since the fixtures score without an LLM judge. Or point Guardian at a remote Ollama → — over an frp stcp tunnel, no public port, and a guard that refuses a review of a silently truncated prompt.
CGIS collects nothing: no telemetry, no analytics, no account. Your code and the graph built from it stay on your machine. The one exception is opt-in: Guardian, if you run it with a cloud model, sends the reviewed diff to the provider you chose. See PRIVACY.md.
Requires Python 3.12+ and uv.
See CONTRIBUTING.md for the standards: strict MyPy, linting, ontology compliance.
CGIS is free and you can run it yourself. If you would rather have the analysis than the tool, I run a fixed-price audit of your codebase's structure — authorisation coverage, blast radius, coupling, architectural drift — delivered in five working days, $2,400 fixed, with an explicit list of what the analysis cannot see. Read what's included →