Personal RAG over your GitHub history (commits, code, reviews), served to Claude Code over MCP.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent โ or use 1-click editor setup below.
We haven't yet run this listing's install command through our automated sandbox check. This isn't a red flag โ we're steadily working through the catalog.
๐ก Paste into ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows)
You reviewed a permission check six months ago. Claude doesn't remember it. github-twin does โ it indexes your commits and review comments, and surfaces them as retrieval hits whenever an agent writes or reviews new code in your style.
Try it now from Claude Code โ drop this into ~/.claude.json and reload:
Your code stays on your box. Embeddings are computed locally (Ollama or sentence-transformers); only the LLM seam (
gt summarize,gt distill,gt eval) optionally calls a hosted provider, and even that's swappable to local Ollama. Thegeminiembedder is the one exception โ opt-in only.
A personal RAG over your GitHub history, served to Claude Code (or any MCP client) as a stdio server. Two scopes, one codebase:
Retrieval is hybrid (BM25 + vector via RRF), AST-aware via tree-sitter for python/scala/javascript/typescript/go/rust, and contextually enriched at embed time with per-chunk headers + optional LLM-generated summaries.
The fastest path is uvx โ no virtualenv to manage, isolated per-tool:
If you prefer a project-local install:
gt and github-twin are the same Typer app โ use whichever fits your
muscle memory.
Pick whichever is least friction โ github-twin tries them in this order:
gh install needed):
Token persists in the OS keyring (macOS Keychain / Linux Secret
Service / Windows Credential Manager) or, when unavailable, a 0600
file under your data dir.gh CLI: if you've already run gh auth login,
gt picks up the token via gh auth token โ nothing to do.GITHUB_TOKEN env var: a classic PAT works too; useful for CI /
headless / docker. Required scopes: repo, read:org, user:email.The MCP server runs over stdio via github-twin serve (or gt serve).
Run uvx github-twin auth login once on the box that will host the
server.
Option A โ via the Claude Code plugin marketplace (lowest-friction):
This registers the MCP server entry automatically; set
GT_PATHS__DATA_DIR in your environment (or in ~/.claude.json's
env block for this server) to point at the DB directory.
Option B โ manual wiring: add an entry to ~/.claude.json (or
your mcp_servers.json):
If you'd rather not persist a token and instead supply it inline (CI,
ephemeral container), add "GITHUB_TOKEN": "ghp_..." to that env
block; it acts as the lowest-priority fallback.
Restart Claude Code; the find_*, predict_review_outcome,
summarize_review_patterns, and sync tools will be available.
Pick a directory to hold the SQLite DB, config, and ingested cache โ everything per-data-dir lives under this one root:
gt sync is incremental on subsequent runs.
config.toml lives next to the DB at <data_dir>/config.toml and is
created on the first gt init --embed-backend ... call. Default
<data_dir> is $XDG_DATA_HOME/github-twin (or ~/.local/share/github-twin)
when GT_PATHS__DATA_DIR is unset.
The retrieval surface (find_*, predict_review_outcome) always runs locally on the SQLite index โ no API call. LLM calls only happen in three places:
gt distill โ clusters review comments / commits into rules.gt summarize โ generates per-chunk NL summaries used by the embed-time
prefix.gt eval reviews / eval predictions โ held-out RAG-vs-baseline scoring.Each picks a backend by precedence Claude โ Gemini โ Ollama (whichever API key is set), or you can force one explicitly.
| Provider | Env var | What it covers |
|---|---|---|
| Anthropic (Claude) | ANTHROPIC_API_KEY | Distill / summarize / eval LLM. Best quality. |
| Google (Gemini, API key) | GEMINI_API_KEY or GOOGLE_API_KEY | Distill / summarize / eval LLM. Free tier is generous. |
| Google (Gemini, Vertex / ADC) | GT_GEMINI_PROJECT (+ optional GT_GEMINI_LOCATION, default us-central1) | Same backends, but auth via gcloud auth application-default login โ no key in your shell. API key wins if both are set. |
| Ollama (local) | OLLAMA_HOST (default http://127.0.0.1:11434) | Distill / summarize / eval LLM. Fully offline. |
The Vertex / ADC path needs the aiplatform.googleapis.com API enabled
on your project, and billing applies even for "free" Gemini models โ
the AI Studio free tier does not extend to Vertex. Project IDs are not
secrets; the credential itself lives at
~/.config/gcloud/application_default_credentials.json and is refreshed
by gcloud.
We keep the embedder backend separate from the LLM backend. Choose one:
nomic-embed-text, 768-dim, ~50ms/chunk).
Requires a running Ollama daemon. Zero cost, fully local.uv add 'github-twin[st]',
pulls torch). Useful when an Ollama daemon isn't available or you
want a specific HuggingFace model. Local.gemini-embedding-001 at 3072-dim by
default). Uses the google-genai dep that's already installed; auth
via GEMINI_API_KEY / GOOGLE_API_KEY, or via GT_GEMINI_PROJECT
gcloud auth application-default login) to route through
Vertex AI without managing a key. Remote โ this is the only
embedder that sends chunk text off-box. Pick it when you have Gemini
auth but no Ollama / [st] install, and your corpus is okay to
share with Google.The embedder is a per-DB commitment โ sqlite-vec bakes the vector
dimension into the table at first creation. Stamp the choice into
<data_dir>/config.toml at init time so every subsequent command
picks it up:
Re-running with the same values is a no-op; running with different
values against an existing config.toml fails loud rather than
silently changing the corpus. GT_EMBED__* env vars still work for
one-off overrides and CI.
A "cloud-LLM only" setup either needs an embedder process (Ollama /
[st]) or has to opt into the remote Gemini embedder.
When you gt init, the GH client needs:
repo โ private repos and PR comments on themuser:email โ verified email addresses for the user-mode identity sweepread:org โ org member listing and private org repo discoveryA fine-grained PAT works; classic tokens too.
Hybrid search by default: BM25 (SQLite FTS5) and vector similarity run in
parallel, then fuse via Reciprocal Rank Fusion (k=60). The vector leg
matches semantic intent; the BM25 leg catches exact identifiers
(getUserById, SQLITE_OPEN_READWRITE) that vector search routinely
misses. Design reference: Anthropic โ Contextual
Retrieval.
At embed time, each chunk gets a deterministic header prepended:
# path :: symbol_name (node_kind), plus the function's leading
docstring/comment when present, plus an optional LLM-generated summary
(see gt summarize). The header lets vector queries land on chunks
whose bodies only contain identifiers (e.g. natural-language queries
against a VaultSecretEq function).
BM25 query expansion is on by default (cfg.retrieval.query_expansion = "rule"), with rule-based code-shaped synonyms applied only to the
BM25 leg โ embeddings already capture synonymy, so expansion never
touches the vector query. Switch to "ollama" to add LLM-generated
alternates on top, cached on disk per-token.
predict_review_outcome stays on pure vector retrieval because its
inverse-distance vote weighting depends on calibrated L2 distance.
All retrieval tools accept optional repo= and author_login= filters.
| Tool | Returns |
|---|---|
find_review_comments(diff_hunk, language?, repo?, author_login?, k=5) | Past review comments on diffs similar to the input. |
find_style_examples(query, language?, repo?, author_login?, k=5) | Past code chunks matching a description. |
find_code(query, language?, repo?, path_glob?, node_kind?, k=5) | Source snippets from files at HEAD (org mode). |
find_applicable_rules(query, language?, repo?, author_login?, k=5) | Distilled code-pattern rules relevant to a coding task. |
predict_review_outcome(diff_or_summary, language?, repo?, author_login?, k=20) | Weighted prediction over nearest past PRs: {approved, changes_requested, commented}. |
summarize_review_patterns(language?, limit=20) | Distilled rules from clustered review comments (run gt distill first). |
sync(since?) | Incremental ingest + summarize + embed. |
Use github-twin <command> interchangeably with gt <command>.
| Surface | Env / config key | Default | Alt |
|---|---|---|---|
LLM (cfg.distill.backend, cfg.summarize.backend) | ANTHROPIC_API_KEY / GEMINI_API_KEY (or GT_GEMINI_PROJECT + ADC) / Ollama | auto (cloud > local) | force claude / gemini / ollama |
Embedder (cfg.embed.backend) | โ / GEMINI_API_KEY (or GT_GEMINI_PROJECT + ADC) | ollama (nomic-embed-text) | sentence_transformers via [st] extra, or gemini (gemini-embedding-001, remote) |
Vector store (cfg.vector_store.backend) | โ | sqlite-vec (brute-force KNN) | faiss via [faiss] extra |
BM25 query expansion (cfg.retrieval.query_expansion) | โ | rule (deterministic) | ollama (LLM, cached) or off |
All settings are layered: defaults โ <data_dir>/config.toml (or the
explicit --config PATH) โ env vars prefixed GT_ (nested via __,
e.g. GT_EMBED__BACKEND=sentence_transformers).
gt eval runs the same prompt with and without retrieval and measures
RAG's accuracy lift on real held-out data:
The harness pre-flights eligibility counts so typo'd --author or
--repo fail fast without burning LLM calls. The judge embedder
defaults to a different model than the retriever
(sentence-transformers BGE-small with the [st] extra installed) to
avoid measuring how well retrieval clusters its own outputs.
Spans for every MCP tool call, every embedder call, and every retrieval leg, exported via OTLP. Auto-detected โ nothing fires unless the environment is configured. Specifically:
Install the [otel] extra (carries the SDK + HTTP OTLP exporter):
Point at an OTLP HTTP collector via env vars:
OTEL_SDK_DISABLED=true forces it off even when an endpoint is set.
Without [otel] or without an endpoint env var, the code paths
still run but every span is a free no-op from opentelemetry-api's
built-in tracer. stdout is never used โ even with telemetry on โ
because MCP speaks JSON over stdin/stdout and a stray console exporter
would corrupt the channel. The OTLP HTTP exporter posts to your
collector; SDK warnings route through Python logging (stderr).
Wired into Claude Code:
Span names + key attributes you can pivot on:
| Span | Useful attributes |
|---|---|
mcp.tool.{find_review_comments,find_style_examples,find_code,find_applicable_rules,predict_review_outcome,summarize_review_patterns,sync} | gh_twin.tool.k, gh_twin.filter.*, gh_twin.result.count (or .prediction/.confidence for predict) |
embedder.embed | gh_twin.embed.input_chars, gh_twin.embed.model |
retrieval.hybrid_search | gh_twin.retrieval.{chunk_kind,k,expander,hits,top_distance} |
retrieval.vector_search (predict_review_outcome) | same shape, sans expander |
A broken or unreachable collector emits a single Failed to export span batch log line per flush attempt and never propagates into the
tool handler โ pinned by tests/test_observability.py.
gRPC users: install opentelemetry-exporter-otlp-proto-grpc alongside
the [otel] extra and the SDK picks it up automatically based on
OTEL_EXPORTER_OTLP_PROTOCOL=grpc.
One DB can hold many targets (user + N orgs + N repos); use a separate
GT_PATHS__DATA_DIR per DB when you want them isolated. The resolved
data dir is pure with respect to env:
GT_PATHS__DATA_DIR when set$XDG_DATA_HOME/github-twin/ when XDG_DATA_HOME is set~/.local/share/github-twin/The current working directory is never consulted.
Layout โ everything per-DB lives under one root:
If you ran an older release that wrote ./config.toml or ./data/ into
the current working directory, every CLI invocation logs a one-time WARN
with the exact mv command to migrate.
Versions come from git tags via hatch-vcs. Cutting a release is one command:
The push to a v* tag triggers .github/workflows/release.yml, which:
uv build across Python 3.12 / 3.13.release.yml, environment pypi).a/b/rc) are flagged so they don't replace "Latest"..claude-plugin/plugin.json on main to match the tag โ
sets version and pins the MCP server invocation to
uvx github-twin@X.Y.Z serve. The marketplace fetches the manifest
from HEAD, so this is what users get when their marketplace cache
refreshes.First-time setup, once per repo:
github-twin,
Workflow: release.yml, Environment: pypi.pypi. Add
yourself as a Required Reviewer for an extra approval step before
each publish (optional but recommended)..github/workflows/ci.yml runs on every PR and push to main โ the
release workflow re-runs the same checks before publishing, so a
broken main never produces a release.
embed.prefix.prefix_chunk): per-kind header
spliced before each chunk's text at embed time, never written back to
chunk.text. Bumps EMBED_TEXT_VERSION whenever the shape changes
so the next gt embed re-derives vectors.process.chunkers): tree-sitter walks emit chunks
per declarable unit; falls back to line-window for unsupported
languages or parse failures.store.query_expansion): BM25 leg only, vector leg always sees the
raw embedding โ pinned by test_hybrid_search.py.The original design plan lives in
getting_started.md along with the full
walkthrough.
MIT โ see LICENSE.
mcp-name: io.github.ChristopherDavenport/github-twin
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/github-twin)<a href="https://allmcps.com/mcp/github-twin"><img src="https://allmcps.com/api/badge/github-twin?style=directory" alt="Github Twin on AllMCPs" /></a>