Save up to 80% tokens when AI reads code via AST-aware structural reading
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
We haven't yet run this listing's install command through our automated sandbox check. This isn't a red flag β we're steadily working through the catalog.
π‘ Paste the JSON block into your client's configuration file under mcpServers, then restart the application.
Token-efficient AI coding, enforced. Cuts context consumption in AI coding assistants by up to 90% without changing the way you work.
Why it matters more now: as frontier models move up in price, the tokens you don't spend reading code are worth more, not less. The savings are in tokens; the value is in tokens Γ price. Token Pilot keeps the expensive main thread lean so the premium model spends its budget on reasoning, not on re-reading files.
Three layers, each useful on its own, stronger together:
smart_read, read_symbol, read_for_edit, β¦). Ask for an outline or load one function by name instead of the whole file.Read on large files, recursive Grep, unbounded git diff) and redirect to token-efficient alternatives.tp-* subagents β Claude Code delegates with MCP-first behaviour and tight response budgets.Files under 200 lines are returned in full β zero overhead for small files.
Measured on public open-source repos. Files β₯50 lines only:
| Repo | Files | Raw Tokens | Outline Tokens | Savings |
|---|---|---|---|---|
| token-pilot (TS) | 55 | 102,086 | 8,992 | 91% |
| express (JS) | 6 | 14,421 | 193 | 99% |
| fastify (JS) | 23 | 50,000 | 3,161 | 94% |
| flask (Python) | 20 | 78,236 | 7,418 | 91% |
| Total | 104 | 244,743 | 19,764 | 92% |
smart_readoutline savings only. Real sessions additionally benefit from session cache,read_symbol, andread_for_edit. Reproduce:npx tsx scripts/benchmark.ts.
Creates (or merges into) .mcp.json with token-pilot + context-mode, then prompts to install tp-* subagents. Restart your AI assistant to activate.
Grep/Bash/Read calls; redirect to efficient alternatives β hooks & modestp-* subagents (Claude Code only) β MCP-first delegates with haiku/sonnet model tiers and budget enforcement β agents referencetools/list to save ~2 k tokens per session β profiles & config| Client | MCP tools | PreToolUse hooks | tp-* subagents |
|---|---|---|---|
| Claude Code | β | β | β |
| Codex CLI | β | β | β |
| Cursor | β | β | β |
| Gemini CLI | β | β | β |
| Cline (VS Code) | β | β | β |
| Antigravity | β | β | β |
Hooks are written for Claude Code (the plugin's own hooks/hooks.json, or
~/.claude/settings.json for an npm install) and for Codex CLI
(npx token-pilot install-hook --client=codex, then /hooks inside Codex to
trust them). The other clients get the MCP tools; their hook systems exist but
token-pilot does not write to them yet.
Manual config snippets for each client β installation guide
TOKEN_PILOT_MODE controls how aggressively Token Pilot redirects heavy native tool calls:
| Value | Behaviour |
|---|---|
advisory | Allow all β hooks pass through, advisory notes only |
deny (default) | Block heavy Grep/Bash patterns; intercept large Read calls |
strict | Deny + auto-cap MCP output (smart_read β€ 2 000 tokens, find_usages β list mode, smart_log β 20 commits) |
Token Pilot owns input tokens β the stuff Claude reads from files, git, search. The other half of a session (what Claude writes back, how it executes code, how it remembers state across days) is owned by separate tools. They compose cleanly:
| Tool | Owns | Typical savings |
|---|---|---|
| Token Pilot | code reads, git, search | 60-90% input |
| caveman | Claude's response prose (terse-speak skill) | ~75% output |
| ast-index | the structural indexer Token Pilot rides on | foundation |
| context-mode | sandboxed shell / python / js execution | 90%+ on big stdout |
A session that pairs token-pilot + caveman typically hits ~85-90% total reduction β each cuts a different half, no overlap. Install what you need; none of them assume the others are present.
Rules of thumb: read code β smart_read/read_symbol; execute code with big output β context-mode execute; bash-only agent β ast-index CLI. Never copy the whole stack into CLAUDE.md β Token Pilot's doctor warns when CLAUDE.md exceeds 60 lines.
TypeScript, JavaScript, Python, Go, Rust, Java, Kotlin, C#, C/C++, PHP, Ruby. Non-code (JSON/YAML/Markdown/TOML) gets structural summaries. Regex fallback handles most other languages.
Claude Code (plugin β recommended):
Other clients (Cursor, Codex, Cline, β¦):
The May 2026 Claude Code update changed a few things that affect how token-pilot is invoked. Nothing breaks on older versions β these are quality-of-life notes for the newer ones.
Run a tp-* agent directly without the plugin: prefix.
claude --agent tp-debugger "fix the stack trace" now works the same
as --agent token-pilot:tp-debugger. The Task tool dispatcher
resolves the short name automatically.
Cold ast-index calls β raise MCP_TOOL_TIMEOUT.
The first find_usages / outline / read_symbol on a large repo
triggers an index build. Default per-MCP-tool timeout (60 s) is
enough for ~50k-file repos; bigger ones benefit from
MCP_TOOL_TIMEOUT=120000 in ~/.claude/settings.json. Subsequent
calls hit the cache and return in ~50 ms.
Background sessions with --mcp-config.
Dispatching a worker via claude agents or --bg with
--mcp-config /path/to/other.json swaps the MCP set for that
session. If token-pilot is not in the override config, MCP tools
(smart_read, find_usages, β¦) are unavailable in that worker
even though the hooks (Read / Edit / Bash / Grep / Task) still
fire β hooks are project-level, MCP tools are session-level. Add
token-pilot to the override config or skip --mcp-config.
claude plugin details token-pilot.
Shows the projected per-turn token cost, the hook event names, and
the MCP server entry. The skill list, the agent list, and the LSP
list are all auto-discovered from the canonical sub-folders.
These fields come from reverse-engineering @anthropic-ai/claude-code@2.1.87
source (see the May 2026 Habr write-up). They work today but are
not in the official Claude Code docs, so use at your own risk.
memory: project)Every relevant tp-* agent (onboard, debugger, pr-reviewer,
history-explorer, audit-scanner) now ships with memory: project
in its frontmatter. Claude Code persists the agent's working notes
in the project so the agent gets faster on repeat invocations β
tp-onboard remembers your layout, tp-pr-reviewer remembers your
flagged patterns, etc. v0.35.0+.
requiredMcpServers)Every tp-* agent declares requiredMcpServers: ["token-pilot"].
Claude Code refuses to load the agent when the MCP server isn't
configured, so a stale install never produces a "tools not found"
loop. v0.35.0+.
once: true)The plugin ships a SessionStart hook flagged once: true β
Claude Code runs it once per project then auto-removes the entry.
It surfaces friendly hints when install-agents or
install-ast-index hasn't been run yet. v0.35.0+.
async: true)PostToolUse hooks (Bash, Task) are marked async: true so they
no longer add wall-clock to the hot path β telemetry writes fire
in the background.
If you want full auto-approval for safe commands, the YOLO classifier reads natural-language environment descriptions:
token-pilot's enforcement still runs on top (raw Read on large files is denied first, regardless of autoMode).
Factual signals from GitHub, npm, and our automated checks β not a rating.
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/token-pilot)<a href="https://allmcps.com/mcp/token-pilot"><img src="https://allmcps.com/api/badge/token-pilot?style=directory" alt="Token Pilot on AllMCPs" /></a>