The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Token Optimizer MCP listing page.
Spend less context, keep the conclusions, and audit every claim across 16 coding clients.
One local ledger for optimizer tools, live-graph substitutions, every agent, and the graph's own cost.
Your agent burns most of its context on work it already did: re-reading files that have not changed, dumping a whole file to see three lines, running unbounded searches, and re-deriving conclusions it reached last session and then forgot.
Token Optimizer attacks that on four fronts.
1. It makes the expensive call impossible. Install the plugin and a built-in
Read of a 200 KB file is denied, with the refusal naming the cached,
diffed replacement. Same for Grep, Glob, Edit, Write, and cat /
head / grep -r through the shell. Re-reading a file you already read this
session returns only a diff — usually the single biggest win, and one that
size-based rules structurally cannot catch. There is no setting to turn on.
2. It remembers what your agent worked out. A per-project knowledge graph accumulates findings, decisions and dead ends as a side effect of working, then feeds them back the moment the agent touches the relevant file. A finding costs ~150 tokens to carry. Re-deriving it costs 5k–50k.
3. It measures itself, in public, and tells you when it is losing. A materialized before/actual-return measurement for MCP progressive disclosure, with later expansions debited from the same net. Modeled graph substitutions and the randomized control arm for downstream graph effects remain separate. Every number is measured, visibly collecting, excluded, or absent.
4. It attributes the traffic. Returned context, optional cost equivalents, and net transport avoided are grouped by operation and MCP handshake identity. Codex, Claude Code, Gemini, and any other connected client get separate rows. Old records without identity remain explicitly unattributed instead of being assigned to whichever agent happens to be open now.
No account, no telemetry, no hosted service. MIT, so it is usable at work.
The screenshots in this README come from the shipped server reading persisted local data, not a design mockup. In the capture above it reports:
Collecting.The dashboard now reads native CLI usage receipts and prices uncached input, cache reads, cache writes, and output with the exact captured provider, model, route, request-time tier, and versioned official source. Ambiguous model ids stay Not priced instead of receiving a blended guess. API/list-price equivalents are kept separate from provider-reported charges and are never labeled as a subscription invoice. See the token accounting contract.
Claude Code — install the plugin, not the bare MCP server. The plugin is what enforces; adding the server alone just gives the model tools it can ignore.
That is the entire installation. All sixteen clients →
Then, whenever you want to know what to do next:
One ranked queue: what is costing the most per session, with an optional monthly cost equivalent only after you configure your own effective rate. Each line names how to fix it. Not a dashboard, not six reports — a queue.
Every agent session ends the same way: the reasoning evaporates. The next session re-derives it, at full price, forever.
This builds a living per-project graph — nodes for files, symbols, tasks and
findings; edges for derived_from, contains, supersedes, contradicts,
related — and it fills itself in from real work. No ingestion job, no
embedding model, no rebuild step, no query to formulate.
None of that is in your repository. It exists only because an agent once burned tokens finding it out — and every other tool throws it away at the end of the session.
What a default install actually produces. The structural graph — files,
symbols, tasks, and the edges between them — is captured from ordinary tool
traffic with no configuration at all. Findings are produced two ways. At session
end, derive reads evidence already on disk (command outcomes and exit codes,
red-to-green transitions, corrections, re-read churn) and writes findings from
it: no model call, no credential, nothing sent anywhere. And the active model
records durable conclusions itself through wiki_write.
The model-based semantic harvest is the third path, and the only one that
needs something you do not already have. It is not opt-in —
TOKEN_OPTIMIZER_HARVEST=0 turns it off — but its real gate is a credential:
with none it reports off:no-key, which is the state on CI, corporate machines,
and subscription-only logins. Point TOKEN_OPTIMIZER_HARVEST_ENDPOINT at a
local model and it runs free and private, with nothing leaving the machine.
npx token-optimizer-doctor states which of these is live.
| Classic RAG | This |
|---|---|
| Retrieves evidence; the model re-derives meaning each time | Retrieves verdicts — the reasoning already happened |
| Index built by a batch ingestion job | Accretes from real agent traffic — coverage follows attention |
| Similarity search | Traversal — this symbol and its callers |
| Model must formulate a query | Fires when the model reaches for a file |
| Staleness invisible; serves rotted chunks confidently | Staleness computed from content hashes, served with the invalidating diff |
| Returns only what is in the documents | Returns dead ends, which exist nowhere in your source tree |
Traversal plus lexical search: deterministic, instant, explainable, and it works offline.
A plain deny costs a full turn: the model calls Read, is refused, re-plans,
calls smart_read. But at refusal time we already hold the file and the
snapshot the graph stored — so the refusal carries the answer inside it.
Nothing to re-plan, no second call. Turn cost drops from one to zero.
And when the graph already holds the verdict a tool output would support, the output never enters context at all. Not compressed. Absent.
The overview answers the questions a token optimizer should answer first:
Not measured; no zero is invented.smart_read, smart_grep, smart_glob, smart_edit, or any other MCP
operation normally. Every successful result records returned context; tools
with a comparable materialized before-state also record a gross reduction;
later expand calls debit that reduction.wiki_write. Before a
new agent re-derives work, call wiki_read for the project or the files it is
about to touch. Native clients can also deliver matching knowledge
automatically.
Use wiki_query to read the graph directly — one finding by key, a ranked
BM25 search over claims, a node with its neighbours, or the graph's own audit
— which is how a subagent that never sees the SessionStart briefing reaches
what previous sessions established.http://localhost:3100. Use Overview for combined accounting and
What it knows for capture health, graph exploration, audits, and causal
evidence.npm run projects:discover -- /absolute/path/to/repos. This makes coverage
gaps explicit; it does not fabricate findings.For maintainers, this live smoke exercises the shipped stdio transport and creates separately attributed rows without seeding the analytics database:
The health panel is deliberately operational rather than a raw text dump. Each client has activity, runtime failures/timeouts, policy outcomes, and observed surface coverage. Diagnostics keep no prompts, commands, paths, or tool output.
Modeled substitution potential and causal graph effects are different claims.
The first is a full-file counterfactual and is never promoted to the verified
MCP headline. The second asks whether delivered knowledge prevented later reads;
it uses a control arm and remains Collecting until there are at least 20
treated file touches and 5 holdouts with valid downstream joins.
Drag to orbit, scroll to zoom, click a node for provenance, or switch to the bounded one-hop focus view. The default All known projects scope pools local graphs through opaque project IDs; filesystem paths never reach the browser.
Audit tab. Contradictions, stale findings, and low-confidence claims remain reviewable instead of silently becoming model truth.
Evidence console. Client/model/task cohorts, matched effects with 95% intervals, live outcome joins, harm feedback, and capability tiers for all 16 clients. Release and superiority claims fail closed while evidence is missing.
One-click Markdown export. The accumulated graph becomes documentation you can inspect, edit, and commit.
Server-side by design: the browser asks for a neighbourhood, a search result or a page. A mature graph holds thousands of nodes, and shipping it wholesale would make every page load a multi-megabyte download for a view that shows twenty things.
The default All known projects scope combines captured graphs through an opaque machine-local project registry; filesystem paths never reach the browser. Lifecycle hooks register repositories as they are used. To backfill existing local checkouts without reading their source files, run the bounded discovery command against one or more explicit roots:
Coverage distinguishes repositories with graph data from known repositories
whose capture has not started. The balance cards are backed by persisted events:
Memory deliveries counts graph context actually supplied to an agent,
Kept back for comparison counts randomized control touches, and Cost of
remembering combines delivered-context tokens with measured semantic-write
payload cost. Reading avoided stays Collecting or Not measured until at
least 20 treated file touches and 5 holdouts exist with a downstream join; the
dashboard does not manufacture a savings estimate from missing data.
See the causal evidence protocol, the cross-client capability contract, and the live evaluation suite.
Every native hook writes the same bounded JSONL lifecycle record, including Claude Code's custom router and compaction paths. Records carry the client and plugin versions, event, hashed session/turn correlation, latency, outcome, input/output byte counts, and response key shape. They deliberately retain no prompt, command, tool output, file content, or raw working-directory path. The fields include OpenTelemetry log severity and resource semantics so the local files can be collected without inventing a second schema.
Raw event rows are opt-in and capped at 1,000. The default report contains aggregate health and at most twenty recent failures/timeouts, keeping routine troubleshooting output small enough for CLI and model context windows.
The dashboard's Capture health panel shows runs, failures, timeouts and
p50/p95 latency by client. Logs rotate at 5 MiB, retain at most 40 files for 14
days, and live under .token-optimizer/logs when a state directory is set (or
~/.token-optimizer/logs otherwise). TOKEN_OPTIMIZER_LOG_DIR,
TOKEN_OPTIMIZER_LOG_MAX_BYTES, TOKEN_OPTIMIZER_LOG_MAX_FILES, and
TOKEN_OPTIMIZER_LOG_RETENTION_DAYS override those operational defaults.
Everyone else checkpoints and restores what you had — which spends the scarcest budget in the session replaying context you already paid for.
Selection here is derived, not a category list: cost-to-rederive × irrecoverability × reuse-probability, with dead ends and decisions on a floor,
because cheap-to-find is not the same as cheap-to-find-again. Restoration then
adapts to the situation — mid-problem, cold resume, or in-flow — within a
measured budget:
Where you were: does clock skew explain the 401s? ruled out: token signing, clock drift on the client untested: NTP skew on the server
That is resuming a thought. A summary describes one.
A large tool result becomes a preview chosen by the session's actual question, after parsing the output's shape (test report, diff, stack trace, log, JSON) — not the first 40 lines because they are first.
Every cut is named, because a model reasoning over a silent truncation
cannot know it is missing anything. expand serves from a content-addressed
store — it never re-runs your test suite — and expanding both teaches the next
preview and promotes what you needed into the graph, so the second expansion
never happens.
Provider caches are billable and provider-specific; a cache hit is not a free input token. For example, Anthropic publishes separate cache-read and cache-write multipliers, while OpenAI and Gemini expose their own cached-input usage and pricing rules. Token Optimizer reads native cache fields when the client supplies them and keeps reads, writes, uncached input, and output separate. It never applies one provider's cache multiplier to another client. The attribution view then does the part a hit rate cannot:
Attributed to a line, priced by what sits behind it. Keep-warm is decided by expected value from your observed gaps, per TTL tier — and when neither tier pays, it says so.
Everyone guesses from task shape and never checks. This reads which model ran each episode and what happened — retries, errors, turns — and prices both mistakes: what an overpowered model wastes, and what an underpowered one costs in retries. A tier that needs a retry in more than half its episodes is excluded at any price, because four cheap turns that fail are not cheap.
A report is read once and forgotten. Here a detection produces a durable, measured, reversible fix — a skip rule, a composite touch — plus a ~50-token session-start briefing so the waste never starts. Detectors are a shipped floor plus patterns derived from your project's own history, each carrying what it has actually saved:
Anything that touches your files is proposed as a diff and never applied.
fleet_audit ranks your whole machine by measured cost, and does something a
per-project scanner cannot: a fix proven in one project is offered to the others
containing the same file contents, carrying the evidence from where it was
measured. Matching is by content hash, never by filename.
It also runs the natural experiment nobody else can — enforcing clients versus directive ones — and reports it whichever way it falls, with the confound stated.
That is a bigger ask than a normal dependency makes, so:
Verify the release. Published from CI with npm provenance — npm audit signatures ties the artifact to the workflow run and the commit, without
trusting us. CHECKSUMS.sha256 ships alongside for offline checking.
If nothing seems to be happening, the lifecycle bundle is probably not
installed. An MCP server process cannot modify the host that launched it, and
npm 11 gates lifecycle scripts behind allow-scripts. Install the native
plugin/hook bundle listed for your client below; adding only the MCP server gives
the model tools but no pre-execution veto. For a legacy global Claude Code
installation, recovery is one line:
Check that it works — not that files exist.
It feeds a synthetic payload to the real hook binary and asserts a large read
is refused and a small one is not. A checklist would have passed on the exact
bug this project once shipped, where the plugin was connected, visible in
/mcp, and saving nothing. Every failure names its own fix.
Every refusal carries its own off switch. Enforcement that hides its disable is coercive, and the person who needs it is mid-refusal, not reading a README:
Removal is exact. The installer records every file it wrote, with hashes.
Uninstall removes only what still matches; anything you edited since is left in
place and named. Your own hooks are never touched, and we never rewrite your
settings.json — we merge into it.
Every tool in this space reports "tokens saved" computed from its own assumptions. That number cannot be wrong, because nothing checks it.
For optimizer results, this records the materialized before-state and the text actually returned to the client, then subtracts any linked expansion responses. Native graph-substitution counterfactuals stay labeled as modeled and outside the main headline.
The graph's broader claim—whether delivered knowledge prevented later re-reading—is causal, so it runs a randomized holdout. Delivery is silently withheld on a slice of file touches, stratified by file, and the effect is the difference in downstream reads between arms. That effect is not added to the combined net until the experiment can support it, and the page will tell you plainly:
the graph is NOT yet paying for itself
A tool that can only ever report good news is not reporting. The same discipline
runs throughout: an unmeasurable saving renders as unknown, never as zero,
and never as $0.00 — because "cannot tell yet" printed as "saved nothing" is a
silent false negative, and dollars get quoted to other people.
The tier is a protocol guarantee, not a preference:
The exact surfaces still differ. For example, a protocol that can replace a large read before it reaches the model provides a stronger token guarantee than one that can only observe it. The generated registry, adapters, dashboard, and certification report all read the same capability source so those claims cannot drift independently.
Every config shape is confirmed against that client's published documentation, with the source URL recorded in its README.
| Suite | Checks | What it proves |
|---|---|---|
test | 2,437 | Enforcement, staleness, injection, consolidation, disclosure, cache, routing, trust — driving the real hooks over stdin |
verify:clients | 253 | Every client config, lifecycle manifest, and enforcement surface matches its documented schema |
verify:harvest | 26 | Request shape, response parsing, and that no secret from a tool result crosses the wire |
verify:ui | 22 | Real headless Chromium: layout, label collisions, legibility |
doctor | 10 | The installed hooks actually refuse, and the server actually answers |
These are not decoration. They found six client configs that would have failed silently, an installer that destroyed user hooks, a Windows path bug that made enforcement blind to half its own refusals, and an uninstaller that printed a plan and deleted nothing.
| Token Optimizer | Typical alternatives | |
|---|---|---|
| License | MIT — commercial use fine | Often noncommercial-only; check before using at work |
| Default behaviour | Refuses the wasteful call | Suggests a better tool |
| Re-read of an unchanged file | Returns a diff | Returns the file again |
| Savings figure | Direct before/return ledger; causal graph effect separately held out | Computed from the tool's own assumptions |
| Cross-session memory | Findings, decisions, dead ends | Usually none |
| Compaction | Consolidation, ranked by cost-to-rederive | Checkpoint and replay |
| Cache economics | Measured from the transcript, attributed to a line | Rarely addressed |
| Model routing | Measured from episode outcomes | Guessed from task size |
| Cross-project | Fixes transfer by content hash | Per-project only |
| Clients | 16 | 3–6 typical |
| Telemetry | None | Varies |
Every client launches the same stdio server — npx -y @ooples/token-optimizer-mcp@latest —
but they differ in what they let a hook do, and that difference is the whole
product. A client with a pre-execution veto can have the wasteful call refused;
one without can only be told. Both are listed honestly below.
Ready-made configuration for all sixteen lives in
integrations/, generated from one source and validated by
npm run verify:clients. Full matrix: docs/CLIENT_SUPPORT.md.
Every MCP connection also receives capability-aware mandatory routing
instructions in its initialize response. That gives all clients a universal
always-on policy, but only the ten clients with native pre-tool surfaces can
hard-veto a wasteful built-in call; install their lifecycle bundle for actual
enforcement.
These ten clients expose a pre-execution hook. Their packaged lifecycle bundle
shares one capability-aware decision engine and defaults to assist: routing,
retrieval, capture and harvest all on, no refusals. Set
TOKEN_OPTIMIZER_MODE=enforce if you want expensive built-in calls vetoed, or
off to disable the hooks entirely.
| Client | Installable lifecycle surface |
|---|---|
| Claude Code | Native plugin: /plugin marketplace add ooples/token-optimizer-mcp, then /plugin install token-optimizer@token-optimizer |
| Codex | Native plugin in integrations/codex/plugin, or standalone hooks in integrations/codex |
| GitHub Copilot CLI | Project hooks in integrations/copilot |
| Gemini CLI | Gemini extension at the repository root, backed by integrations/gemini |
| Qwen Code | Extension bundle in integrations/qwen |
| Cursor | Project hooks and always-applied rule in integrations/cursor |
| Cline | Project hooks and rule in integrations/cline |
| OpenCode | In-process plugin and hooks in integrations/opencode |
| Kilo | In-process plugin and hooks in integrations/kilo |
| Windsurf | Project hooks and rule in integrations/windsurf |
Roo Code, Zed, Amp, Continue, Crush, and Droid do not expose a packaged pre-execution bridge that can safely veto built-in calls. Their generated, always-on rules make optimized routing mandatory whenever the exact MCP schema is visible, and fail open to a bounded native operation when it is not. The integration directories contain both the MCP config and the rules file at the paths documented by each host.
These paths are checked, not assumed. Verifying them against each client's published docs found six configs that would have installed cleanly and never loaded — including Kilo, whose schema shares nothing with the
mcpServersconvention the other clients use. A convention is not a schema.
For the best experience, install the Codex plugin. It bundles the MCP server, the token-optimization skill, session guidance, and a large-read hook:
Review and trust the bundled hooks with /hooks, then start a new conversation. If you prefer an MCP-only installation, use:
On Windows, if PowerShell blocks the codex.ps1 shim, use the command launcher directly:
This writes the server to ~/.codex/config.toml. Codex CLI, the Codex IDE extension, and the Codex app on the same host share that configuration.

Start a new Codex conversation after installation so the new tools are discovered. In an interactive CLI session, /mcp shows the tools available to the conversation.
The plugin supplies both automatically. For an MCP-only installation, add the guidance from integrations/AGENTS.md to a project or global AGENTS.md. A ready-made standalone hook is also available under integrations/codex/hooks; merge its hooks.json into ~/.codex/hooks.json, copy the script to ~/.codex/hooks/, and review it once with /hooks.
The Codex hook injects guidance at SessionStart. Under TOKEN_OPTIMIZER_MODE=enforce it blocks expensive native operations when the bundled MCP has an exact replacement, including a single unambiguous code-mode shell call such as cat or Get-Content; multi-operation orchestration remains advisory so unrelated work is not discarded, and a second attempt at the same target passes through if the MCP is unavailable. The default is assist, which keeps retrieval and capture but never vetoes; TOKEN_OPTIMIZER_MODE=advise adds the routing advisory without vetoes, and off disables the hooks. The AGENTS.md/skill guidance remains important.
If you prefer a smaller instruction block:
See the current Codex hooks documentation for hook trust, matching, and tool-coverage details.
If you prefer to edit ~/.codex/config.toml yourself:
Install the plugin, not the bare MCP server. The plugin is the only path that optimizes by default; adding the MCP server alone gives the model a set of tools it is free to never call.
That is the whole installation. There is nothing to configure and no flag to turn on.
From the first message of the next session, expensive built-in calls are refused and redirected to the optimized equivalent:
| You (or the model) do this | What happens |
|---|---|
Read a file over ~25 KB | Denied → smart_read (cached) |
Read any file already read this session | Denied → smart_read (returns only the diff) |
Grep file contents / Glob for files | Denied → smart_grep / smart_glob |
Edit a file over ~25 KB | Denied → smart_edit (returns a diff, not the file) |
cat/head/tail/Get-Content a large file | Denied → smart_read |
grep -r / rg across the tree | Denied → smart_grep |
| Context fills and compaction starts | optimize_session runs first |
The re-read case is usually the largest single win and the one most often missed: a 5 KB config read fifteen times across a session costs far more than one 200 KB file read once. Size-based rules never catch it.
Three properties, all tested:
git log | head,
and binary paths are never touched.One variable, no reinstall:
If you want the tools without the enforcement:
Verify with claude mcp get token-optimizer, or /mcp inside Claude Code. Then
add the recommendations from integrations/AGENTS.md
to your CLAUDE.md — but be aware that guidance in a context file is advisory,
and models routinely read past it.
The standalone global installer can also configure the Claude Code hooks and supported desktop clients:
On Windows, a restrictive PowerShell policy may need this user-scoped adjustment first:
Interactive global installs run the hook installer; CI and local dependency installs skip it. If automatic setup is skipped, use install-hooks.ps1 on Windows or install-hooks.sh on macOS/Linux. See the Claude Code MCP guide and this project's hook installation guide.
On current Copilot CLI releases:
If your Copilot CLI does not expose copilot mcp yet, save this as ~/.copilot/mcp-config.json:
A ready-made copy is available at integrations/copilot/mcp-config.json.
Inside an interactive Copilot session, /mcp show token-optimizer displays the connection status and available tools.
Keep integrations/AGENTS.md as the repository's AGENTS.md, or adapt the same guidance into .github/copilot-instructions.md.
For native lifecycle integration, copy the ready-made repository hooks into your project:
The hooks inject optimization guidance at sessionStart. Under TOKEN_OPTIMIZER_MODE=enforce they deny a large built-in view so Copilot retries with smart_read; the default assist leaves the call alone. Partial reads and files below 25 KB always pass through unchanged, and TOKEN_OPTIMIZER_MODE=advise gives recommendations without vetoes. Repository hooks work without overwriting user-level files; global hooks can instead be placed in ~/.copilot/hooks/ with their script paths adjusted for that directory.
Restart Copilot CLI after changing hook files. See GitHub's official MCP setup guide and hooks reference.
Add Token Optimizer directly at user scope:
Alternatively, install this repository as a Gemini extension so the MCP configuration and GEMINI.md guidance are packaged together:
Run /mcp inside Gemini CLI to inspect the connection. Restart Gemini CLI after installing or updating the extension.
Direct MCP users should copy GEMINI.md into the project or merge its rules into an existing GEMINI.md. Extension users receive that context file plus native hooks automatically.
The extension's SessionStart hook injects optimization guidance. Its AfterTool hook notices full-file read_file results over 25 KB and suggests smart_read. To make Gemini automatically replace those large results with a token-optimizer tail call, configure the extension setting Automatic large-read routing as true:
Automatic routing uses Gemini's native tailToolCallRequest: the smart_read result replaces the built-in read result before it reaches the model. Partial reads remain unchanged. Restart Gemini CLI after installing, updating, or reconfiguring the extension. See the official Gemini MCP guide, extension guide, and hooks reference.
Create or update opencode.json in your project:
For a global installation, merge the same mcp entry into ~/.config/opencode/opencode.json.
The output shows configured servers and their connection status.
Copy integrations/AGENTS.md to the project as AGENTS.md; the instructions entry above loads it. Then copy the ready-made local plugin:
The plugin preserves Token Optimizer usage state in OpenCode's compaction prompt. Under TOKEN_OPTIMIZER_MODE=enforce its tool.execute.before hook rejects full-file reads over 25 KB and steers the agent to smart_read; the default assist lets them through. Small and partial reads always pass normally, and TOKEN_OPTIMIZER_MODE=advise gives non-blocking guidance. Restart OpenCode after adding the plugin. See the official OpenCode MCP guide and plugin hook guide.
Any stdio-capable MCP client can launch Token Optimizer with:
Additional ready-made integration files are available for Claude Desktop, Codex, Gemini CLI, OpenCode, and GitHub Copilot.
You normally use Token Optimizer by asking your agent in plain language:
For clients that expose direct MCP tool calls, the core inputs are small JSON objects:
optimize_text returns measurements with every call. This example uses a deliberately repetitive payload to make every field easy to see; it is not a benchmark:
get_optimization_report applies the same versioned measurement contract as
the dashboard and aggregates qualifying operations into:
Two tools sound similar but serve different purposes:
| Tool | Use it for | Context-window effect |
|---|---|---|
optimize_text | Store bulky text under a key and return a compact reference | Reduces text kept in the active context |
compress_text | Produce Brotli/base64 data for storage or transport | May use more model tokens if pasted back into context |
If your goal is a smaller prompt, prefer optimize_text. Use compress_text only when you specifically need byte compression.
| Capability | Representative tools | What gets smaller or faster |
|---|---|---|
| Context and compression | optimize_text, get_cached, count_tokens, analyze_optimization, context_delta | Large payloads and repeated context |
| File and Git operations | smart_read, smart_write, smart_edit, smart_grep, smart_glob, smart_diff, smart_status | File contents, search results, and diffs |
| Caching | smart_cache, cache_warmup, cache_invalidation, cache_compression, predictive_cache | Repeated computation and retrieval |
| APIs and databases | smart_api_fetch, smart_sql, smart_graphql, smart_rest, smart_schema | Responses, schemas, and query analysis |
| Build and system tasks | smart_build, smart_test, smart_lint, smart_logs, smart_processes | Build logs and diagnostic output |
| Intelligence | smart-summarization, pattern-recognition, natural-language-query, recommendation-engine | Analysis and summaries |
| Analytics | get_optimization_report, get_action_analytics, get_hook_analytics, export_analytics | Token-savings visibility |
See docs/TOOLS.md for detailed tool inputs and examples.
Default local data locations include:
~/.token-optimizer-cache/~/.token-optimizer-mcp/analytics.db~/.token-optimizer/Set TOKEN_OPTIMIZER_CACHE_DIR to override the cache location.
This installs hooks that refuse your tool calls. That is a bigger ask than a normal dependency makes, so here is everything needed to check it and undo it.
Verify the release is genuine. The package is published from CI with npm provenance, which signs an attestation binding the artifact to the workflow run and the commit that built it:
That verifies without trusting us. A CHECKSUMS.sha256 is attached to each
GitHub release for offline checking (sha256sum -c CHECKSUMS.sha256) — useful
for mirrors, but weaker: it shares a trust root with the thing it hashes.
Check it actually works. Not that the files are in place — that it works:
This feeds a synthetic payload to the real hook binary and asserts a large read
is refused and a small one is not, that session-start emits the policy, that the
graph directory is writable, and that the MCP server starts and lists its tools.
Every failure names its own fix. There is also an install_doctor MCP tool.
Turn enforcement off, instantly. Every refusal says this, so you never have to come back here to find it:
Remove it. The installer records every file it wrote, with hashes, so removal is exact rather than best-effort:
It removes only files that still match what we wrote. Anything you have edited
since is left in place and named, because removing it would destroy your
work and removing it silently would be worse. Hooks you added yourself are not
in the manifest and are never touched. Config entries we added are listed for
you to remove — we do not rewrite your settings.json.
The detailed operational material below is intentionally retained for users who want to understand the complete tool surface, hooks pipeline, performance controls, analytics, and troubleshooting behavior.
The proposed category-level successor architecture, its 18 required
workstreams, and the evidence gates for claiming a universal cognitive runtime
are documented in
docs/UNIVERSAL_COGNITIVE_RUNTIME_PROGRAM.md.
Usage Example:
Optimized replacements for standard file tools with intelligent caching and diff-based updates:
Usage Example:
Intelligent caching and optimization for external data sources:
Usage Example:
Development workflow optimization with intelligent caching:
Usage Example:
Enterprise-grade caching strategies with 87-92% token reduction:
Usage Example:
Comprehensive monitoring with 88-92% token reduction through intelligent caching:
Usage Example:
System-level operations with smart caching:
Usage Example:
The MCP server is identical in every client, but lifecycle APIs are not. The repository ships client-native adapters instead of copying Claude event names into tools that would silently ignore them.
| Client | Native integration events | Refusal under MODE=enforce | Advisory escape hatch |
|---|---|---|---|
| Codex | SessionStart, PreToolUse | Deny replaceable large reads and single shell dumps | TOKEN_OPTIMIZER_MODE=advise |
| Claude Code | PreToolUse plus optional global pipeline | Deny replaceable large reads and noisy searches | TOKEN_OPTIMIZER_MODE=advise |
| GitHub Copilot CLI | sessionStart, preToolUse, postToolUse | Deny large view calls and steer to smart_read | TOKEN_OPTIMIZER_MODE=advise |
| Gemini CLI | SessionStart, BeforeTool, AfterTool | Deny replaceable large reads before they enter context | TOKEN_OPTIMIZER_MODE=advise |
| OpenCode | tool.execute.before, compaction hook | Reject large full-file reads and steer to smart_read | TOKEN_OPTIMIZER_MODE=advise |
The default is assist: routing, retrieval, capture and harvest are on and nothing is ever refused. That is the posture two independent harnesses measured as our best -- on THOL, assist scored 0.971 in 14.4 turns against control's 0.969 in 16.2, while enforce scored 0.960 in 20.3 turns and cost 1.471x control per task (median across 17 tasks; cheaper on only 2 of them). Set TOKEN_OPTIMIZER_MODE=enforce for the refusals described above, which use a 25,600-byte threshold; override it with TOKEN_OPTIMIZER_LARGE_READ_BYTES, use TOKEN_OPTIMIZER_MODE=advise for the routing advisory without vetoes, or TOKEN_OPTIMIZER_MODE=off to disable the lifecycle integration. Partial reads always pass through because they may already be more efficient than a full cached read, and under enforce one repeated attempt is allowed so a failed MCP server cannot permanently block work.
Granular token usage analytics for pinpointing optimization opportunities:
Usage Example:
Key Features:
This complete seven-phase pipeline applies to the optional Claude Code global-hook installation. Codex, Copilot, Gemini, and OpenCode use the smaller native adapters above because their event names, payloads, and result-replacement capabilities differ.
When global hooks are installed, token-optimizer-mcp runs automatically on every tool call:
Reduction targets in tool descriptions are design goals, not production
measurements. They never enter the verified ledger. Use the dashboard or
get_optimization_report for request-level measurements, and
npm run dashboard:audit-savings to inspect both qualifying and excluded rows.
No universal dollar value is claimed: configure an effective rate only when it
reflects your own provider, model, cache, route, plan, tier, and credits.
Token Optimizer works with stdio-capable MCP clients and includes first-party setup guidance for:
See Installation for the supported commands and configuration files.
The PowerShell hooks have been optimized to reduce overhead from 50-70ms to <10ms through:
Control hook behavior with these environment variables:
The MCP server exposes an 18-tool core catalog by default so tool schemas do not
consume a large share of the model context. Set
TOKEN_OPTIMIZER_TOOL_PROFILE=full before starting the server to expose all 103
specialized tools. TOKEN_OPTIMIZER_TOOL_PROFILE=core is the explicit form of
the default. Live graph capture uses
TOKEN_OPTIMIZER_TOOL_PROFILE=continuity, which exposes only capture and query.
The extended cognitive profile adds checkpoints, outcomes, and receipt
attestation. The native-token audit measures 480 startup tokens for continuity,
694 for extended cognitive, 4,815 for core, and 30,125 for full. Stateful
consumers normally receive zero MCP tools: host pre-action delivery adds only
the selected capsule through the client lifecycle channel. Other enabled MCP
servers add their own schemas independently.
The current hardened cross-CLI smoke does not qualify. In the final
Codex-to-Claude adversarial pair, both successors were correct and the runtime
capsule was delivered with exact provider model attestations and zero consumer
MCP tools, but runtime used 5.383% more tokens and 41.758% more latency. The
reciprocal runtime arm was blocked by Claude provider quota, and Antigravity
1.1.11 required account authentication before a Gemini model could run. These
signed outcomes keep the release verdict insufficient; they are evidence of
working provenance and fail-closed behavior, not powered effectiveness or
universal superiority.
The replacement-grade protocol is intentionally larger than that smoke:
54,054 all-family/all-arm trial envelopes, 113,022 provider
calls, three independent model families, direction-level non-inferiority, and
1,056 hard-negative opportunities per direction and arm so the Bonferroni-adjusted
95% family-wise false-delivery upper bound—not only the point estimate—must
remain below 1%. Run
npm run verify:ucr:study-design to validate the frozen metric coverage. The
full protocol and CLI driver contract are in
evals/ucr/FULL_STUDY_CONTRACT.md.
TOKEN_OPTIMIZER_USE_FILE_SESSION (default: false)
true to revert to file-based session tracking (legacy mode)$env:TOKEN_OPTIMIZER_USE_FILE_SESSION = "true"TOKEN_OPTIMIZER_SYNC_LOG_WRITES (default: false)
true to disable batched log writes$env:TOKEN_OPTIMIZER_SYNC_LOG_WRITES = "true"TOKEN_OPTIMIZER_DEBUG_LOGGING (default: true)
false to disable DEBUG-level logging$env:TOKEN_OPTIMIZER_DEBUG_LOGGING = "false"TOKEN_OPTIMIZER_DEV_PATH
~/source/repos/token-optimizer-mcp if not specified$env:TOKEN_OPTIMIZER_DEV_PATH = "C:\dev\token-optimizer-mcp"Performance Impact: Using in-memory mode (default) provides a 7x improvement in hook overhead:
To view your actual token SAVINGS, use the get_session_stats tool:
Output includes:
Example Output:
All operations are automatically tracked in session data files:
Location: ~/.claude-global/hooks/data/current-session.txt
Format:
New in v1.x: The savings object is now automatically updated every 10 operations, eliminating the need to manually call get_session_stats() for real-time monitoring. This provides instant visibility into token optimization performance.
How it works:
get_cache_stats() MCP toolNote: For detailed per-operation analysis, use get_session_stats(). The session file provides high-level aggregate metrics.
Analyze token usage across your entire project:
Provides:
Monitor cache hit rates and storage efficiency:
Metrics:
codex mcp get token-optimizer.codex mcp list./mcp in the interactive CLI and inspect the server status.codex is blocked on WindowsPowerShell may reject the codex.ps1 shim under a restrictive execution policy. Use codex.cmd for the installation and verification commands, or review your user-scoped PowerShell execution policy.
This removes the Codex registration; it does not delete your local cache or analytics database.
Symptom: Claude Code shows "Invalid Settings" error after running install-hooks
Cause: UTF-8 BOM (Byte Order Mark) was added to settings.json files
Solution: Upgrade to v3.0.2+ which fixes the BOM issue:
If you're already on v3.0.2+, manually remove the BOM:
Symptom: Token optimization not occurring automatically
Diagnosis:
Check if hooks are installed:
Verify dispatcher.ps1 exists:
Solution: Re-run the installer:
Symptom: Session stats show cache hit rate below 50%
Causes:
Solutions:
Warm up the cache before starting work:
Increase TTL for stable APIs:
Check cache size limits:
Symptom: Node.js process using excessive memory
Cause: Large cache in memory (L1/L2 tiers)
Solution: Configure cache limits:
Or clear the cache:
Symptom: Initial Read/Grep/Glob operations are slow
Cause: Cache is empty, building indexes
Solution: This is expected behavior. Subsequent operations will be 80-90% faster.
To pre-warm the cache:
Symptom: Cannot write to cache or log files
Cause: PowerShell execution policy or file permissions
Solution:
Set execution policy:
Check file permissions:
Re-run installer as Administrator if needed
Symptom: ~/.token-optimizer/cache.db is >1GB
Cause: Caching very large files or many API responses
Solution:
Clear old entries:
Reduce cache retention:
Manually delete cache (nuclear option):
If you encounter issues not covered here:
~/.claude-global/hooks/logs/dispatcher.log~/.claude-global/hooks/data/current-session.txtget_session_statsTo make Codex use your local build while developing:
MIT License - see LICENSE for details
Built for measurable token efficiency across supported coding agents by the ooples team.