Index your codebase. AI searches instead of re-reading files. 94% token savings.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
π‘ Paste into ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows)
Index your codebase. AI searches instead of re-reading files.
94% token savings, reproducibly benchmarked.
Website Β· Docs Β· Why CCE? Β· Benchmark Β· GitHub
Python 3.11+ Β· macOS Β· Linux Β· Windows
One command. Auto-detects your editor. Zero cloud, zero config.
Talk: We Cut 94% of Our AI Coding Tokens β AI Engineer World's Fair 2026
| Use case | How CCE helps | |
|---|---|---|
| π° | Reduce Claude Code costs | 94% fewer input tokens per session |
| π | Keep code private | Everything local, no cloud indexing |
| π | Multi-editor teams | One index across Claude Code, Cursor, VS Code, Gemini CLI |
| π§ | Cross-session memory | Decisions and context survive restarts |
| β‘ | Faster responses | Less context = faster Claude replies |
| π | Track actual savings | Dollar amounts, not estimates |
One command. 30 seconds.
Or if you prefer a persistent install:
Restart your editor. Done. Every question now hits the index instead of re-reading files.
Already have Ollama? Skip
[local]and useuv tool install code-context-engineinstead. CCE auto-detects Ollama at localhost:11434 and usesnomic-embed-text.
Python 3.11+ and a C compiler (for tree-sitter grammars).
| Platform | Setup |
|---|---|
| macOS | xcode-select --install |
| Ubuntu/Debian | sudo apt install build-essential cmake |
| Fedora/RHEL | sudo dnf install gcc gcc-c++ cmake |
| Windows | Visual Studio Build Tools (C++ workload) + CMake |
Tested on macOS, Linux, Windows with Python 3.11/3.12/3.13.
cce init auto-detects your editor and writes the right config. To target a
specific agent, use --agent claude, --agent codex, --agent copilot, --agent pi, or
--agent all.
| Editor | Config written | Instructions |
|---|---|---|
| Claude Code | .mcp.json | CLAUDE.md |
| VS Code / Copilot | .vscode/mcp.json | .github/copilot-instructions.md |
| Cursor | .cursor/mcp.json | .cursorrules |
| Gemini CLI | .gemini/settings.json | GEMINI.md |
| OpenAI Codex | ~/.codex/config.toml (user-global, per-project section) | AGENTS.md |
| OpenCode | opencode.json | |
| Tabnine | .tabnine/agent/settings.json | TABNINE.md |
| Pi | .mcp.json | AGENTS.md |
Multiple editors in the same project? All get configured in one command.
Codex note: Codex CLI reads MCP servers from ~/.codex/config.toml only β
it has no per-project config. cce init adds one [mcp_servers.cce-<project>-<hash>]
section per project so multiple projects coexist; cce uninstall removes only
the section for the current project.
Pi note: Pi does not support MCP natively. To use CCE with Pi, you need a
pi MCP adapter extension (e.g. pi-mcp-adapter)
that consumes the .mcp.json config and exposes CCE's tools to the Pi agent.
cce init sets up both .mcp.json and AGENTS.md (Pi loads the latter
automatically for startup instructions).
Supports Anthropic, OpenAI, and Google model pricing. Configure via pricing.model in ~/.cce/config.yaml.
Input tokens are 85-95% of your Claude Code bill. CCE cuts them by 94% (benchmarked on FastAPI).
| Without CCE | With CCE | |
|---|---|---|
| Session startup | Re-reads files every time | Queries the index |
| Finding a function | Read entire 800-line file | Get the 40-line function |
| Cross-session memory | None | Decisions + code areas persisted |
| Token cost (Sonnet, medium project) | ~$0.14/session | ~$0.04/session |
We benchmarked CCE against FastAPI (53 source files, 180K tokens) with 20 real coding questions. No cherry-picking, no synthetic queries.
Methodology: For each query, "without CCE" means reading the full content of every file the query touches. "With CCE" means the relevant chunks after compression.
Important baseline note: The 94% number is measured against full-file reads, not against what Claude Code actually does. In practice, Claude Code already uses grep, partial file reads, and targeted tools, so the real-world savings compared to normal Claude Code behavior will be lower than 94%. We use full-file as the baseline because it's reproducible and deterministic (no agent behavior variability). The benchmark measures CCE's retrieval efficiency, not a head-to-head comparison with Claude Code's built-in exploration.
| Metric | Result |
|---|---|
| Retrieval savings | 94% (83,681 β 4,927 tokens/query) |
| Compression (additional, on retrieved chunks) | 89% (4,927 β 523 tokens/query) |
| Recall@10 (found the right files) | 0.90 |
| Latency p50 | 0.4ms |
| Queries tested | 20 |
| Layer | What it does | Savings | Method |
|---|---|---|---|
| Retrieval | Full files β relevant code chunks | 94% | measured |
| Chunk Compression | Raw chunks β signatures + docstrings | 89% | measured |
| Grammar | Drops articles/fillers from memory text | 13% | measured |
Output compression (reducing Claude's reply length) provides additional savings (~65% estimated) but is not included in the headline number above.
| Repo | Language | Files | Retrieval savings | Recall@10 |
|---|---|---|---|---|
| FastAPI | Python | 53 | 94% | 0.90 |
| Django | Python (large) | 2,347 | 93% | 0.95 |
| Express | JavaScript | 6 | 94% | 1.00 |
| chi | Go | 94 | 76% | 0.67 |
| fiber | Go (monorepo) | 396 | 93% | 0.07 |
Django (2,347 files, 5.4M tokens) shows CCE scales to large codebases with 0.95 recall. Go's shorter files reduce the retrieval headroom (smaller baseline). Monorepos dilute recall at top-10 (fiber). Middleware queries with one-feature-per-file hit R=1.00 consistently.
Reproduce it yourself:
Full results in benchmarks/results/. Queries and methodology in benchmarks/.
11 MCP tools that Claude uses automatically:
| Tool | What it does |
|---|---|
context_search | Hybrid vector + BM25 search with graph expansion |
expand_chunk | Full source for a compressed result |
related_context | Find code via graph edges (calls, imports) |
session_recall | Recall decisions from past sessions |
session_timeline | Walk turn summaries for a session (drill into recall hits) |
session_event | Inspect raw tool input/output for a specific event |
record_decision | Save a decision for future sessions |
record_code_area | Record which files were worked in |
index_status | Check index freshness |
reindex | Re-index a file or the full project |
set_output_compression | Adjust response verbosity (off / lite / standard / max) |
Live dashboard with donut charts, file health, and session history:

Dollar estimates with multi-provider pricing (Anthropic, OpenAI, Google):
context_search. Hybrid vector + BM25 retrieval finds the right chunks. Code graph adds related files automatically.session_recall.cce savings shows exactly how much you saved.Re-indexing after edits takes under 1 second (96% embedding cache hit rate). Git hooks keep the index current automatically.
Output compression tools (like Caveman) save 20-75% on output tokens. Output is 5-15% of your bill. Net savings: ~11%.
CCE saves on input tokens (94% retrieval savings on FastAPI, reproducibly benchmarked). Input is 85-95% of your bill.
Not a text search. Tree-sitter AST parsing creates semantic chunks. Hybrid retrieval merges vector similarity with BM25 keyword matching via Reciprocal Rank Fusion. A confidence scorer blends similarity (50%), keyword match (30%), and recency (20%). Graph expansion walks CALLS/IMPORTS edges to pull in related code.
record_decision("use JWT for auth", reason="session tokens flagged by legal") is stored in SQLite and surfaces via session_recall in the next session. No re-explaining your architecture.
Not estimates. Actual tokens served vs full-file baseline, broken down by buckets (retrieval, compression, output, memory, grammar). Dollar costs fetched from Anthropic's pricing page. Savings summary shown at every session start.
Secret files (.env, *.pem, credentials.json) are never indexed. Content is scanned for AWS keys, GitHub tokens, Slack tokens, Stripe keys, JWTs, and generic credentials. PII (emails, IPs, SSNs, credit cards) is scrubbed from memory writes. All MCP file paths are validated against path traversal.
SHA-256 fingerprint per chunk, salted with model name. Re-index skips unchanged code. Binary float32 storage (10x smaller than JSON). Typical re-index: 96% cache hit, under 1 second.
Replaced LanceDB with sqlite-vec. Same cosine-distance quality, 99% smaller install. WAL mode + PRAGMA NORMAL for 80% write speedup. Vectors, FTS5, code graph, and compression cache all in three SQLite files.
Memory entries compressed without LLM calls. Drops articles, fillers, pronouns. Three levels (lite/full/ultra, 20-60% savings). Code, paths, URLs preserved byte-for-byte. Same input always yields same output.
5 Claude Code lifecycle hooks capture session context. Every hook runs curl ... || true, so a crashed server never blocks the user. SessionStart injects bootstrap context; others capture silently.
Dollar estimates in cce savings support 15+ models across Anthropic, OpenAI, and Google. Static pricing ships with CCE, live Anthropic pricing is fetched and cached 7 days. Configure pricing.model (e.g. gpt-4o, gemini-2.5-pro, sonnet) or override with pricing.input / pricing.output for custom rates.
Running dozens of cce serve processes (one per project per AI session) can exhaust system memory. The resource governor caps ONNX Runtime threads per process, uses advisory file locks so only one process indexes a given project at a time, backs off under Linux memory pressure (PSI), and auto-shuts down idle servers after 30 minutes. Configure via serve.idle_timeout_minutes and serve.max_ort_threads.
CCE's cross-session memory depends on the agent calling record_decision and record_code_area. Memory nudges make recording ambient: after N searches without a recording, context_search results include a short reminder. At session end, the Stop hook summarizes unrecorded activity. Nudges re-arm after the first recording so they stay useful without being noisy.
cce serve --http exposes a POST /search endpoint for custom agent integrations that speak HTTP instead of MCP stdio. Same hybrid retrieval pipeline, structured JSON response with confidence scores. Input validation clamps top_k (1..100) and confidence_threshold (0.0..1.0).
7 buckets track every token saved: retrieval, chunk compression, output compression, memory recall, grammar, turn summarization, progressive disclosure. Survives restarts. Powers CLI and dashboard analytics.
Run cce list for the full command reference.
Zero-config by default. Override what you need in ~/.cce/config.yaml or .context-engine.yaml:
Remote Ollama: If you run Ollama on another machine in your network, set compression.ollama_url (e.g. http://nas.local:11434) or export CCE_OLLAMA_URL (the env var wins). CCE probes the endpoint and falls back to truncation-only compression when it's unreachable, so a flaky link won't break indexing.
CCE also compresses Claude's responses (same concept as Caveman):
| Level | Style | Savings |
|---|---|---|
off | Full output | 0% |
lite | No filler or hedging | ~30% |
standard | Fragments, drop articles | ~65% |
max | Telegraphic | ~75% |
Tell Claude: "switch to max compression" or "turn off compression". Code blocks and commands are never compressed.
| Component | Size |
|---|---|
| Core install (Ollama backend) | ~17 MB |
With [local] extra (fastembed + ONNX) | ~189 MB |
| Embedding model (one-time download) | ~60 MB (fastembed) or managed by Ollama |
| Index per project (small/medium/large) | 5-60 MB |
No GPU required. With Ollama, embeddings are handled by the Ollama server. With the [local] extra, the embedding model runs on CPU via ONNX Runtime.
AST-aware chunking (tree-sitter parsed, 11 extensions):
| Language | Extensions |
|---|---|
| Python | .py |
| JavaScript | .js, .jsx |
| TypeScript | .ts, .tsx |
| PHP | .php |
| Go | .go |
| Rust | .rs |
| Java | .java |
| C# | .cs |
Language-aware fallback chunking (40+ extensions):
| Category | Languages |
|---|---|
| Web | HTML, CSS, SCSS, LESS, Vue, Svelte |
| Systems | C, C++, Zig, Nim |
| Mobile | Swift, Kotlin, Dart |
| Functional | Haskell, Scala, Clojure, Elixir, Erlang, F# |
| Scripting | Ruby, Perl, Lua, R, Bash/Zsh |
| Data/Config | JSON, YAML, TOML, XML, SQL, GraphQL, Protobuf |
| DevOps | Terraform, HCL, Dockerfile |
| Docs | Markdown |
All other text files are chunked by line range. Binary files are skipped.
| Page | Content |
|---|---|
| How Much Are You Spending on AI Coding Tokens? | The math on input vs output tokens |
| What is CCE? (Complete Guide) | Setup, tools, how it works, FAQ |
| How to Save Claude Code Tokens | Cost breakdown and savings guide |
| Benchmark Deep Dive | Full FastAPI benchmark methodology |
| Comparison with Alternatives | CCE vs Cursor, Aider, Continue, Greptile |
| Examples | Real conversations with Claude |
| How It Works | Full 9-stage pipeline |
| CLI Reference | Every command with output |
| Configuration | All config options |
No. Quality stays the same or slightly improves.
CCE replaces "dump the entire file" with "search for the relevant function." The model still gets the code it needs (0.90 Recall@10 in benchmarks). Less irrelevant context means less noise competing for attention, which can improve the model's focus on your actual question.
CCE writes output compression rules directly into your agent's instruction files (CLAUDE.md, AGENTS.md, .cursorrules, etc.) during cce init. These rules apply to the entire session, not just CCE tool responses, so every reply from the agent follows them.
Set the level in ~/.cce/config.yaml or .context-engine.yaml:
Then re-run cce init to update instruction files. Or change at runtime:
| Level | Savings | What it does |
|---|---|---|
off | 0% | No compression |
lite | ~25% | Removes filler/hedging/pleasantries + diff-only for code changes |
standard | ~70% | Drops articles, fragments, short synonyms + diff-only for code |
max | ~80% | Telegraphic style + diff-only for code |
Default is standard. All levels include code output rules that tell the model to show only changed lines (not full file rewrites), which is where most output tokens go in coding sessions. The max level produces very terse prose (similar to "caveman mode"). Code blocks, paths, and commands are never compressed regardless of level.
Most savings are input tokens (what goes into the model):
| Layer | Type | Typical savings |
|---|---|---|
| Retrieval | Input | 94% (full files β relevant chunks) |
| Chunk compression | Input | 89% (chunks β signatures) |
| Grammar compression | Input | 13% (article/filler removal) |
| Turn summarization | Input | varies (session history) |
| Progressive disclosure | Input | varies (tool payloads) |
| Output compression | Output | 25-80% (depends on level) |
Output tokens cost 5x more per token (e.g. Opus: $15/1M input vs $75/1M output), so even a small output reduction has outsized cost impact.
See CHANGELOG.md for shipped features.
Contributions welcome. See https://github.com/elara-labs/code-context-engine/blob/main/CONTRIBUTING.md for setup.
MIT. See LICENSE.
Claude Code Β· MCP Β· sqlite-vec Β· Tree-sitter Β· fastembed Β· Ollama
If CCE saves you tokens, give it a star.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/code-context-engine)<a href="https://allmcps.com/mcp/code-context-engine"><img src="https://allmcps.com/api/badge/code-context-engine?style=directory" alt="Code Context Engine on AllMCPs" /></a>