Read large logs, JSON and docs in a fraction of the tokens. Offline, fact-preserving optimizer.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
One-click editor setup isnβt available for this listing yet β we donβt have a confirmed install command, and weβd rather show nothing than point your editor at the wrong package or host. Follow the projectβs own setup instructions, linked above.
A VL-JEPA-inspired pipeline that compresses images, text, conversations, and RAG documents locally via Ollama, then sends only compact payloads to any LLM API β every saving measured with a real tokenizer and checked for lost facts.
Use in Claude | Quick Start | Python API | REST API | Monitoring | Load Testing | Deployment | AI Tools | Benchmarks | Contributing
Every time you send an image or long prompt to GPT-4o / Claude / Gemini, you burn 1,000+ tokens on processing that could happen locally for free.
| Feature | Description |
|---|---|
| Local-First | Vision and text compression runs on Ollama (free, no API key needed) |
| Token Optimizer | Deterministic, fact-preserving compression: 66% fewer tokens on a realistic dev corpus in ~1ms (benchmark) |
| MCP Server | Works with Claude Desktop, Cursor, Cline, Continue, Zed |
| Selective Decoding | For video, only call API when scene changes (~2.85x fewer calls) with cosine similarity |
| Text Compression | Long prompts, conversations, RAG docs compressed locally |
| Speed Optimized | Connection pooling, model preloading, parallel processing |
| Multi-Provider | OpenAI, Anthropic, Google, Groq, DeepSeek, Together, Azure, AWS Bedrock, Ollama, or any OpenAI-compatible endpoint |
| REST API | FastAPI server for web application integration |
| Video Processing | Direct video file input with automatic frame extraction |
| Cost Tracking | Persistent cost tracking with SQLite analytics and exportable reports |
| Async Support | Non-blocking async methods for FastAPI, aiohttp, etc. |
| Streaming Responses | Stream responses from remote LLMs |
| Config Persistence | YAML/TOML config files with environment variable overrides |
| Structured Logging | JSON-formatted logging with rotation and correlation IDs |
| Docker Support | Dockerfile and docker-compose for easy deployment |
| Plugin System | Custom processors for domain-specific compression |
| Multi-Language | Support for 30+ languages with automatic detection |
LatentGate plugs into Claude as an MCP server. Its optimizer tools work offline, with no Ollama and no API key β Claude reads big logs, JSON dumps and docs through it and spends a fraction of the context. A 400-line error log goes from 14,400 to ~100 tokens.
Requires uv (uvx fetches LatentGate from PyPI on first run).
The first run downloads dependencies, which can exceed Claude's 30-second MCP startup limit
on a slow connection; run this once beforehand (later starts take ~2s):
Claude Code β plugin (MCP server + a skill that tells Claude when to use it):
Claude Code β MCP server only:
Claude Desktop β add to claude_desktop_config.json and restart:
Then just ask: "Read logs/app.log with latent-gate and tell me why orders fail."
| Tool | Needs Ollama | What it does |
|---|---|---|
read_file_optimized | no | Read a text file and return its optimized form (use for files you read, not files you edit) |
optimize_text | no | Optimize text you already have, optionally toward a question / max_tokens budget |
count_tokens | no | Count tokens (tiktoken o200k_base) |
compress_image | yes | Describe an image locally as a ~150-token scene payload |
compress_text / compress_conversation / compress_documents | yes | Optimizer + fact-checked local-LLM compression |
get_stats | yes | Session statistics |
This starts everything: Ollama β pulls models β API server β website. See scripts/quickstart.sh for options like --no-pull, --no-website, --port 9000.
For API deployments, restrict direct image-path reads to trusted directories:
Benchmark before releases so speed and savings are measured, not guessed:
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/latentgate)<a href="https://allmcps.com/mcp/latentgate"><img src="https://allmcps.com/api/badge/latentgate?style=directory" alt="LatentGate on AllMCPs" /></a>