ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Autonomous Agents β
β Claude Code β’ Cursor β’ Antigravity β’ Cline β’ Windsurf β
βββββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββ
β
βΌ (HTTP / SSE Outbound Prompts)
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β β‘ TOKENECTOMY AI GATEWAY (127.0.0.1:8080) β
β β
β βββββββββββββββββββββββββββ βββββββββββββββββββββββββββββ ββββββββββββββββββββββββββββββββββββββββββββββββ β
β β Zero-Cost Prompt β β Zero-Leak Credential β β Polyglot Trace Surgery β β
β β Cache β β Redaction β β (9 Languages) β β
β β <1ms Hit β’ 100% Free β β API Keys β’ JWTs β’ DB URIsβ β node_modules β’ site-packages β’ .cargo β β
β βββββββββββββ¬ββββββββββββββ βββββββββββββββ¬ββββββββββββββ ββββββββββββββββββββββββ¬ββββββββββββββββββββββββ β
β β β β β
β βββββββββββββββββββββββββββββββββΌβββββββββββββββββββββββββββββββββββββββββ β
β βΌ β
β Smart Upstream Auto-Router β
β (Anthropic β’ OpenAI β’ Groq β’ Ollama Local) β
ββββββββββββββββββββββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
ββββββββββββββββββββββββββββββββββΌβββββββββββββββββββββββββββββββββ
βΌ βΌ βΌ
api.anthropic.com api.openai.com localhost:11434
(Claude 3.5 / 3.7 Sonnet) (GPT-4o / o1 / o3-mini) (DeepSeek / Llama)
β‘ Quick Start (10s)
0. Zero-Config One-Liner Wrapper (wrap)
Run any coding agent or test suite directly behind Tokenectomy's privacy & caching gateway with zero manual environment variable setup:
# Wrap Claude Code CLI
razor wrap -- claude
# Wrap Aider with DeepSeek or Ollama (Auto-Transpiled on the fly!)
razor wrap -- aider --model deepseek/deepseek-chat
# Wrap test runs to scrub logs and prevent accidental credential leakage
razor wrap npm test
1. Launch the Standalone AI Gateway (Manual Mode)
npx -y tokenectomy-razor --proxy
π‘ Smart Multi-Provider Auto-Routing & Universal LLM Transpiler ON by default! Tokenectomy automatically analyzes request paths and headers. Need to run Claude Code against DeepSeek or Ollama? Tokenectomy transparently transpiles Anthropic /v1/messages format into OpenAI /v1/chat/completions on-the-fly and vice-versa.
2. Connect Your Favorite Agent
Point your coding agent or CLI to the local gateway on 127.0.0.1:8080:
| Agent / Environment | Setup Command or Configuration | Auto-Routed Upstream |
|---|
| Claude Code | export ANTHROPIC_BASE_URL="http://127.0.0.1:8080" | https://api.anthropic.com |
| Cursor | Models > OpenAI Base URL: http://127.0.0.1:8080/v1 | https://api.openai.com |
| Google Antigravity / Gemini | export OPENAI_BASE_URL="http://127.0.0.1:8080/v1" | https://api.openai.com |
| Cline / Roo Code | Provider: OpenAI Compatible β’ Base URL: http://127.0.0.1:8080/v1 | https://api.openai.com |
| Ollama / Local LLMs | Base URL: http://127.0.0.1:8080 | http://localhost:11434 |
3. Real-Time FinOps & Security Dashboard
Open http://127.0.0.1:8080/dashboard in your browser to observe live token reductions, prompt cache hits, dollar savings, and redacted credentials in real time.
4. Companion MCP Server Setup (Sub-Cortex)
For agents calling deep diagnostic tools (get_error_context, analyze_code, apply_code_patch):
Claude Code CLI
claude mcp add tokenectomy npx -y tokenectomy-razor --mcp
Google Antigravity / Gemini CLI
agy mcp add tokenectomy-razor -- npx -y tokenectomy-razor --mcp
Cursor Composer (.cursor/mcp.json)
{
"mcpServers": {
"tokenectomy": {
"command": "npx",
"args": ["-y", "tokenectomy-razor", "--mcp"]
}
}
}
Claude Desktop (claude_desktop_config.json)
{
"mcpServers": {
"tokenectomy": {
"command": "npx",
"args": ["-y", "tokenectomy-razor", "--mcp"]
}
}
}
5. Terminal Piping & CLI Scrubbing
npm test 2>&1 | npx tokenectomy-razor
π The Problem: Why Your AI Hits Rate Limits & Hallucinates
When your app crashes during development (Next.js, Express, FastAPI, Tokio), the runtime dumps hundreds of lines of third-party plumbing from node_modules or site-packages.
When you paste that raw crash dump into Cursor or Claude:
- Eats Your 5-Hour Rate Limit: A single Express/Prisma error can dump 5,000 to 45,000 tokens of third-party library code you never wrote. A few crash loops easily burn your session limit.
- Triggers AI Hallucinations: Claude gets lost in framework internals (
node_modules/express/lib/router/layer.js or starlette/routing.py) and tries to edit library files instead of your actual application code.
- Leaks Secrets & Credentials: Connection strings with raw database passwords, JWT bearer tokens, and cloud keys embedded in error traces get forwarded to external model servers.
Raw Terminal Crash (45,820 tokens + Leaked Secrets)
β
βΌ <0.2ms Local Rust DFA Engine
[Redact Passwords & Keys] βββΊ [Strip Third-Party Framework Frames] βββΊ [Isolate Root Cause]
β
βΌ
Clean Context (118 tokens β’ Zero Secrets β’ Sub-millisecond)
π Before & After Comparison
β Without Tokenectomy: AI Hallucinates & Burns Context
TypeError: Cannot read properties of undefined (reading 'token')
at loadComponents (/app/node_modules/next/dist/server/load-components.js:14:2)
at renderToHTML (/app/node_modules/next/dist/server/render.js:50:5)
at nextServer (/app/node_modules/next/dist/server/next-server.js:80:12)
at processTicksAndRejections (node:internal/process/task_queues:95:5)
at runMicrotasks (node:internal/process/task_queues:120:3)
at checkoutHandler (/app/pages/api/checkout.ts:42:15)
Database connection failed: postgresql://admin:super_secret_password@db.prod.internal:5432/primary
API key leaked: sk-ant-api03-abcdef1234567890abcdef1234567890
What Claude does: Analyzes load-components.js and next-server.js, speculates on Webpack / Next.js internals, and leaks connection credentials to remote inference logs.
β
With Tokenectomy: Clean Context & Instant Fix