The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Token Optimizer listing page.
Zero-dependency, fully-reversible context compression for AI agents.
slimctx compresses what your agent reads — tool outputs, logs, JSON, source files, prose — before it reaches the LLM. Same answers, fraction of the tokens. Pure Python stdlib: no ML models, no downloads, no network calls, ever. Auditable end to end in ~1,600 lines.
Live output of python3 benchmarks/demo.py — run it yourself, nothing is staged.
| Workload | Before | After | Savings | Key facts kept |
|---|---|---|---|---|
| Code search (100 results) | 5,557 | 916 | 84% | ✓ |
| SRE incident debugging | 61,699 | 298 | 100% | ✓ |
| GitHub issue triage | 12,836 | 975 | 92% | ✓ |
| Codebase exploration | 5,734 | 2,760 | 52% | ✓ |
Every run also verifies that each planted "needle" (the FIXME, the OOMKill,
the outlier) survives compression, and that every lossy transform is
byte-exact reversible. Reproduce with python3 benchmarks/bench.py.
[slimctx-ref <hash> ...] marker. The
model — or you — can always get the byte-exact original back.Headroom is the established project in this space and is more featureful today (provider proxy with SSE streaming, agent wrappers, cross-agent memory, an ML compression model). slimctx makes a different set of trade-offs, aimed at locked-down / client-site deployments:
| Headroom | slimctx | |
|---|---|---|
| Reversibility | JSON only (CCR); dropped text is gone | every lossy transform |
| Log handling | generic text scoring | template mining ([x1432] collapse) |
| Code handling | AST skeleton | AST skeleton + query-relevant bodies kept |
| Dependencies | Rust core, ONNX runtime, 261MB HF model | stdlib only |
| Network egress | HuggingFace pull on first run | none, ever |
| Store encryption | none (plaintext SQLite) | cipher hook (bring your own) |
| Determinism | cache-aligner component | by construction (pure functions + memo) |
| Audit surface | ~10s of KLOC across 3 languages | ~1,200 lines of Python |
If you need the proxy/wrap ecosystem, use Headroom. If you need something you can read in an afternoon, run air-gapped, and certify for a client environment, use slimctx.
As a library (any framework): call pipe.compress(messages) right
before your provider SDK call; expose pipe.retrieve as a tool named
retrieve so the model can pull originals.
As an MCP server (GitHub Copilot, Claude Code, Cursor, ...): ships built in, stdlib-only:
See USAGE.md for the GitHub Copilot (.vscode/mcp.json) setup
and a security deployment checklist.
Encrypted store:
Apache-2.0. Original implementation — no code derived from Headroom.