The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Distil listing page.
Your provider now edits the context window for you — clearing old tool results, summarizing history — by default, server-side, with no report of what changed.
Distil is the instrument that answers the only question that matters: did the agent still do the same thing?
|
We pointed it at the providers. Anthropic's default context-editing policy ( |
Distil also compresses — the tool output, logs, and history your agent re-sends every turn, reversibly.
It's the one context operation that ships with its own certificate, its own adversarial gate, and a number it is willing to refuse to print.
Compression that cannot be checked is a guess about your agent's behaviour.
Distil is built so every part of it is checkable, and so the checks are allowed to come back no.
1{A=B} − 1{A=A'} — a paired difference against the model's own self-agreement, with a bootstrap 95% CI, unclipped, so it is allowed to be negative. One reporting floor (50 A/B + 30 A/A) gates every surface; below it, every surface says below reporting floor instead of a number. The current live sample is below that floor, and the status line says so.distil_expand tool to recover the original mid-task. The gateway ships Tier-0 only rather than emit a stub it cannot restore.Edit(old_string=…) is a literal match. Reading exact-quote provenance from the shell command, not just the tool name, took byte-exact quote loss from 39.3% → 16.2% on real coding traffic — and it costs real savings, which we price rather than hide.distil validate --adversarial runs a COMA-class battery through the same path the proxy uses, and we publish the two cases that do not come back clean. Threat model →distil bench --curve traces savings against fact recall across the whole ladder, offline and free. The curve →distil wrap -- claude · codex · gemini · aider · opencode · qwen · goose · grok · openhands · copilot · kimi. Zero config, no code change.base_url client at it. Python, TypeScript, any language, any framework. Sync proxy, async proxy, and a standalone gateway, with the same provider coverage in each: Anthropic Messages, OpenAI Chat Completions and the Responses API, Azure OpenAI, and Gemini generateContent.from distil import compress_messages in your own agent loop.distil hook --install: Claude Code compresses its own tool output through
the documented PostToolUse extension point. No proxy, no credentials touched. distil quota shows
the rate-limit window it buys back. Details →Not sure which of those you want? Two questions pick your mode → — plain language, honest savings ranges, no jargon.
Will it save you money? On metered billing (an API key), yes — directly, off the bill. On a flat-rate Pro/Max subscription there is no per-token bill to cut, but there is a rate-limit window, and spending fewer tokens per turn leaves more of it for the next task.
distil quotashows that window live. Savings come from large, repetitive tool output: verbose JSON and duplicated log runs compress 25–99%, while prose and unique-line output compress ~0% — a short session that never reads a big file showing near 0% is the tool working correctly, not failing. Why →
◉ LIVE · measured from the opt-in census on a public git branch, never estimated
▶ Watch the counter tick live & audit every number →
| ⚡ Get the savings 2 min, no config pipx install distil-llmdistil onboard | 🔬 See the proof real harness benchmark ↓ · paper vs the others |
Use it · Library · Integrations · Install · Why trust it · Full Docs →
Building the agent yourself? Compress the message list where it lives — no proxy, no network hop:
Tool results get the reversible digest; user and system text get lossless transforms only; the model's own turns are never rewritten. Handles resolve across processes and restarts, so a digest made by the proxy expands here and vice versa. verbatim=True disables digests entirely.
Named compress_messages/expand_handle rather than compress/expand because distil.compress and distil.expand are modules — a top-level export sharing those names would resolve to the function or the module depending on unrelated import order.
TypeScript too — compress(messages) from the npm package, byte-identical to the Python engine. Full reference: Library API → · runnable examples: python_library.py · js_library.ts.
Maintain a framework? docs/INTEGRATING.md is the ~20 lines and the four rules — we would rather the integration live in your repo than ours.
Every other compressor asks you to trust it won't break your agent. Distil is the only one that proves it won't.
On 500 real coding tasks, compressed context matched full context within statistical noise: 42.0% vs 39.2% tasks solved. (SWE-bench Verified)
Honest scope: +2.8pp is a point estimate (CI −0.6..+6.2pp — non-inferiority certified, superiority not yet). Details, incl. what doesn't transfer →
| On a real 500-instance long-horizon agent (SWE-bench Verified, official harness) | task success | tied with full context? | reversible + certified? |
|---|---|---|---|
| Distil (gated + surprise digest, measured on v1.7) | 42.0% | ✅ tied (+2.8pp point est., CI −0.6..+6.2 — n.s.) | ✅ |
| Distil (relevance-gated, E8) | 36.8% | ✅ | ✅ |
| Headroom (lossy) | 32.6% | ❌ −6.6pp | ❌ |
| LLMLingua-2 (lossy — only 16/500 runs completed) | 2.4% | ❌ −36.8pp | ❌ |
| no compression (full) | 39.2% | — | — |
Headroom column read directly against the public headroom-ai 0.37.0 source, as of v0.37.0, 2026-09-04; every line there cites a file:line in that release. Facts, not adjectives — and where it is genuinely strong, we say so.
| Property | Distil | Headroom 0.37.0 (2026-09-04) |
|---|---|---|
| Per-request behavioural check | Paired A/A′/B replay, unclipped difference, bootstrap CI, one reporting floor | No shadow or dual-send path in the codebase; accuracy_guard="strict" is echoed on /healthz and /stats but nothing branches on it |
| Recovery of what was folded | Content-addressed store + agent-facing distil_expand, byte-exact, verified by a gate | A TTL cache (SQLite, 1800s, 1000-entry FIFO), no integrity or round-trip check |
| Lossy paths with no recovery | None — the gateway ships Tier-0 only rather than emit a stub it cannot restore | Four: OpenAI chat streaming, Responses under ChatGPT auth, Gemini streaming, Bedrock |
| Savings number | Counted, then calibrated against the provider's billed usage | Falls back to chars/3.5 |
| Exact-quote guarantee for coding agents | Provenance read from the shell command, not just the tool name; quote loss 39.3% → 16.2% | Not a property the tool has |
| Cache contract | Suffix-only, cache-monotonic, enforced as an invariant | Genuinely strong prompt-cache replay (overlay_cached_prefix) — real engineering |
| Adversarial gate | COMA-class battery in CI; the two cases that don't come back clean are published | None shipped |
| Degradation curve | Every ladder rung measured, offline and free | Point configuration only |
| Shipped default | Compresses | Mode cache — a full bypass on Bedrock, freeze-only on OpenAI |
On the same corpus, re-run 2026-09-04 with Headroom's model preloaded: distil 52.9% tokens / 58.7% $ / 100% decision-equivalent / PASS vs Headroom 1.7% / 2.0% / 81% / FAIL. On a read→edit→re-read coding workload Headroom reaches 35.6% tokens where distil's digest is 0.0% by design — that is the exact-quote guarantee being paid for, and both numbers are on one page with the raw output committed.
Distil is the only compressor statistically tied with full context — its v1.7 surprise-preserving digest reaches 42.0% vs 39.2% (paired non-inferiority certified; superiority not significant) while every lossy tool craters. And on the live head-to-head above (graded by claude-opus-4-8), it certifies 83.2% savings at a 0% decision-change rate, ~1,000× faster than the nearest tool (distil is pure-Python heuristics — no local ML model; competitors run transformer inference). Full breakdown ↓
One command sets you up and tells you what to do next:
It detects your environment (Claude Code · Codex · Gemini CLI; metered vs subscription) and hands you the exact commands. Or wrap your agent directly — no config, no code change:
Using Cursor, Cline, or Windsurf? They are IDE extensions — no argv to wrap and no documented env var, so
distil wrapcannot reach them. Run a proxy and point the editor's base-URL setting at it: docs/IDE-AGENTS.md. (GitHub Copilot is not redirectable at all, and that page says so rather than wasting your afternoon. The Continue CLI — as opposed to its VS Code extension — routes only through a config file, anddistil wrap -- cnmanages that file for you; see the same page.)
Each recognized agent (claude / codex / gemini / aider / opencode / qwen / goose) auto-selects the right env var and upstream — no --env-var or --upstream flag needed. Prints preset: <agent> detected → <VAR> on start. Explicit flags always win.
distil wrap againTired of typing distil wrap every time? Make it the default — once:
It detects your shell (zsh / bash / fish / PowerShell) and billing mode, writes the
right line to the rc file your shell actually reads, and tells you what it detected.
Want every SDK covered (not just the agent you type)? distil default --always-on
runs a persistent proxy service — powerful, but it pins ANTHROPIC_BASE_URL, so
every client on the machine goes through one local process.
That pin used to be a single point of failure: a proxy that was down for one
second meant sessions failing with ConnectionRefused, an error that names the
provider rather than distil. It no longer is. The service supervisor
(launchd/systemd) owns the listening socket, so a crash or a restart leaves
connections queued in the kernel backlog instead of refused — the client waits
about a second rather than dying. distil default --always-on also verifies the
service is genuinely registered and serving before it wires anything, and refuses
to wire at all if it isn't.
If you ever need out and distil is already uninstalled, sh ~/.distil/uninstall.sh
removes the pin, the service, and the shell block using nothing but sh.
Then watch genuine savings from your traffic — measured, not estimated:
Validate it on your traffic. --shadow runs a fraction of requests twice (compressed and full) and compares the agent's chosen next action:
Honest scope: that's next-action equivalence — a proxy, not task success (E7 shows it doesn't fully transfer under aggressive lossy compression). Distil fails safe to full context.
Will it save money? On metered billing (API key) — fewer tokens, fewer dollars, directly. On a flat-rate subscription there is no per-token bill, so the saving is rate-limit headroom: fewer tokens per turn means more turns before you hit the window (
distil quotashows it live). Coding agents: short sessions ~7%, big wins on long, many-turn sessions the model never re-reads.
You don't need byte-equivalence — you need decision-equivalence: your agent taking the same actions with compressed context. That's measurable and certifiable.
1{A=B} − 1{A=A'}, with a bootstrap 95% CI and no clamp at zero. The old ratio estimator printed exactly 100% whenever chance favoured it and could not express harm at all. One reporting floor now gates the status line, the proof ledger, shadow-stats, the census feed and the public dashboard alike — and prints below reporting floor rather than a flattering number.Edit still applies — an Edit(old_string=…) is a literal match against bytes the agent read earlier; digest that read and the edit silently does nothing while the agent reports success. Provenance is read from the shell command (cat, head, sed -n), not just the tool name — that is 33.6% of tool-result mass the name rule never covered. It costs savings, and the changelog prices it instead of hiding it.distil validate --adversarial runs seven COMA-class cases through the same public path the proxy uses. Trusted/untrusted budget isolation is structural: there is no keep budget shared between blocks anywhere, asserted as an equality in CI. Two results we publish rather than smooth over: dedup-baiting does fold the genuine error line (reversibility is what saves it), and decoy-verdict flooding is a real, unmitigated denial of savings — 0.0% on that block.distil bench --curve reports savings, fact recall, visible recall, facts lost and reversibility at every rung, offline and free.distil certify-trajectories bounds how many solvable tasks compression can cost (no other compressor certifies either level).distil_expand tool. Compress fearlessly.max_attempts = 5; ask "the connection timeout?" and it keeps deadline_ms. Two more layers grow from your own traffic, never from a shipped blob: associations distil learns from its content-free expand flywheel (hashed pairs, --expand sessions), and a learned relevance model that is promoted only after its held-out recall beats the lexical baseline on your labels — until promotion, the lexical + bridge layers are exactly what runs. An optional distributional-vector table can be supplied too (pure-Python cosine; none ships). Every layer is additive — it can only widen keeps, so reversibility and the certificate are untouched — and it needs no embeddings or model to work.distil dissect turns a wrap session into a report: savings by model/mechanism, the digest inventory, billed-usage calibration, latency by path, and a worth-your-attention anomaly list that catches silent failures automatically.Edit's old_string only has to exist byte-exact somewhere in the forwarded payload. → ADR 0010distil_expand call mid-stream, splicing the recovery in without buffering the turn (no TTFT tax on the reversible tier).Fidelity tiers: lossless (
--verbatim) · reversible (byte-recoverable on demand — default) · lossy (every other tool). Only Distil certifies the reversible tier (Headroom ships an uncertified retrieve; Distil's recovery is agent-facing — the model expands mid-task — and gated by the decision-equivalence certificate).
Don't take the table above on faith. distil bench re-certifies savings and decision-equivalence on a bundled 8-domain corpus, offline, in seconds — the same gate that runs in CI. How we evaluate — and why a compression ratio without a task-success delta is meaningless — is written up in docs/EVALUATION.md, including our own negative result:
Five gates, all in CI: bench (non-inferiority on the corpus), verify (byte-fidelity), retention (fact-level recall), fidelity (state probes, below), and validate — which drives the compressor against adversarial inputs (huge/unicode/nested/malformed/marker-injection/secret-looking) and asserts reversibility, reject-if-bigger, recency-exactness, fail-open, and content-free telemetry hold on every one. That last gate exists because a green unit suite kept coexisting with real-traffic bugs; validate is the adversarial layer that catches them.
Recall is not enough, and here's the case that proves it. A trajectory creates net/scratch_bench.py at turn 2 and deletes it at turn 4. Compress away turn 4 and every path token is still present — string recall reads 100% — while the agent now believes a file exists that doesn't, and will plan around it. distil fidelity folds tool calls into a file-state ledger and grades the final state, separating lost (path gone — the agent can see the gap) from stale (path present, state wrong — the agent acts confidently on a falsehood). On that case: string recall 100%, state fidelity 0%.
It reports three more things recall can't see: overclaim ("approximately 4200 ms" → "4200 ms" — the value survives, its uncertainty doesn't), continuation (does the agent still know what's left to do?), and error propagation (does a loss at turn k show up as a behaviour change at turn k+n?). The gate is on silent failures only — CI runs --max-silent 15 — because loud loss is already retention --max-lost's job, and gating one regression twice hides which property broke. The bound is the measured one, not zero: Tier-1 digests hedged spans behind restore handles and drops the qualifier on 9 of 171 claims, so gating at zero would assert a property the compressor does not have. On top of that, distil suite grades twelve public benchmarks whose answer keys were written by someone else — including BFCL, which compresses the tool schema and checks that every name the gold call needs — the function and each argument — survives. At matched savings (90.1% vs 89.3%) truncation keeps 0 of 70 names; distil keeps all 70 — though none of them visibly: the schema sits behind a restore handle, one distil_expand away. The suite prints that gap (visible → true support: bfcl 0%→100%) rather than the flattering number alone, because a reader who assumes the model can see a schema it must actually expand first has been misled by figures that are individually correct. Names are matched as identifiers — a quoted JSON token, escaping tolerated — not as prose: the generic matcher was crediting 11 of 85 golds by accident ('a' matching inside "tool-schemas"). Fifteen golds BFCL genuinely names a, b, c are excluded and counted, since a one-letter token can be neither credited nor failed honestly. Every row is labelled rich or thin payload, because a benchmark with nothing to compress is a control, not evidence — and a run that grades only controls exits 1. It needs no API key and no spend, so it is wired into make gate and the CI gate job rather than run before a launch. Full methodology, including what these probes found wrong with our own corpus, in docs/EVALUATION.md §6; how to run everything, in docs/RUNNING-EVALS.md.
Recall, and a number you can check yourself. The three gates above are graded on our corpus against our oracle — rigorous, but not checkable by you. distil retention --dataset hotpotqa grades against ground truth written by someone else (HotpotQA's gold supporting sentences, amid 8 distractor paragraphs), next to a truncation baseline tuned to distil's own savings on the same case:
| HotpotQA, n=100 | savings | answer recall | gold-sentence recall |
|---|---|---|---|
| distil (reversible) | 14.3% | 100.0% | 100.0% |
| truncation @ matched savings | 14.1% | 91.6% | 82.7% |
distil retention also splits recall into visible (in front of the model) and recoverable (one distil_expand away, verified against the handle's restore bytes). On the corpus that's 100% true recall with 0 lost, and being reversible instead of lossy is worth 21.4% recall — the mean across all 9 domains, each counted once. That's deliberately the macro average: the fact-weighted one reads 62.6%, but it's set by whichever domain carries the most probes, and one HTML fixture moved it from 9.8% to 62.6% without the compressor changing at all — the moat, as a measurement rather than an argument. distil retention --live reports the same on your own traffic; the meter stores counts only, never content.
And it found a real hole. The first thing the recall harness caught was not a regression but a missing capability: distil was compressing 0.0% of HTML tool results — minified markup is one long line, so line-folding had nothing to fold. Agents with a fetch or browser tool were paying full price for <script>, <style>, and nav chrome. Now:
| real page | before | after | saved | facts lost |
|---|---|---|---|---|
| Wikipedia article | 281,093 tok | 14,260 tok | 94.9% | 0 |
| Python docs page | 32,322 tok | 4,229 tok | 86.9% | 0 |
Reversible, which is the part a lossy extractor can't offer: the exact original stays behind the handle, so a bad heuristic call costs one distil_expand instead of the content.
To be precise about what each layer proves: the per-commit gates grade decision-equivalence with an offline deterministic oracle over the committed corpus (fast, free, runs on every push — but synthetic). A nightly live-cert job re-certifies the same trajectories against a real model (distil certify --runner anthropic), budget-capped with a hard --max-live-calls ceiling so an unattended run can never spend silently. The empirical results above (SWE-bench n=500, live head-to-head n=200) were graded by real models; the per-commit badge alone doesn't claim that.
Why trust the number? Token-savings numbers are easy to fake — measure quality at low compression, advertise savings at high compression. Distil refuses that: accuracy and compression are measured on the same trajectories, and a strategy that can't pass non-inferiority doesn't ship.
distil eval plots the certified compression frontier — a savings-vs-quality curve where every point carries its certification verdict, locating the cliff past which lossy compression drops decisions. The artifact no competitor publishes: benchmark.html.
Three results, all reproducible, all published with caveats:
llmlingua / headroom-ai (graded by claude-opus-4-8): 83.2% savings at 0% decision-change, ~1,000× faster (no ML model loaded vs. competitors' local transformer inference). The live proxy behavior is pinned to the certified strategy by tests/test_live_certified_equivalence.py; the one reviewed delta is a recency carve-out that keeps the freshest tool-result turns verbatim (an agent needs its freshest output byte-exact). Since 1.45 that carve-out applies only to content the provider has not cached — anchored to the client's cache_control breakpoint, and dropped entirely for providers that cache implicitly. A carve-out counted back from the end of the conversation slid forward as it grew, rewriting already-cached content one turn later and costing more in re-billed prefix than the digest saved. → benchmarkFull methodology, McNemar tests, per-instance data: docs/PAPER.md · PDF.
Measured on your traffic, never estimated, nothing leaves your machine:
x-distil-* response headers (tokens-saved, mode, compressible-tokens, expanded).distil leaderboard (--html for a page).distil proxy --shadow 0.05 reports the live decision-change rate — streaming-aware.distil proxy sidecar + set ANTHROPIC_BASE_URL once; every client routes through it.distil census on) shares your numbers-only totals — preview the exact payload with distil census show before consenting; TELEMETRY.md has the frozen schema. Default remains: nothing is sent.Dashboard, status-line plugin, federated leaderboard: Deploy & observability.
One proxy. Point any base_url-honoring client at it — Python, TypeScript, any language — and get cache-aware reversible compression with no code change.
| SDK / framework | Change | Example |
|---|---|---|
| Anthropic SDK (Py/TS) | base_url="http://127.0.0.1:8788" | examples/python_anthropic.py · examples/js_anthropic.ts |
Claude Agent SDK / claude -p (headless) | distil wrap -- <cmd> or ANTHROPIC_BASE_URL | examples/python_claude_agent_sdk.py |
| OpenAI SDK (Chat + Responses) | base_url="http://127.0.0.1:8788/v1" | examples/python_openai.py |
| Vercel AI SDK | createAnthropic({ baseURL: '…:8788' }) — or in-process: wrapLanguageModel({ model, middleware: distilMiddleware() }) | examples/js_vercel_ai_sdk.ts |
| LangChain (py/js) · LangGraph | anthropicApiUrl / base URL · pre_model_hook | examples/js_langchain.ts |
| LiteLLM | api_base="http://127.0.0.1:8788" | examples/python_litellm.py |
| Google Gemini | --upstream https://generativelanguage.googleapis.com | examples/python_gemini.py |
Codex · aider · Cursor-agent · any base_url client | distil wrap -- <agent> or OPENAI_BASE_URL | — |
Anything that speaks the Anthropic / OpenAI / Gemini wire format works — the proxy is framework-agnostic, so CrewAI, AutoGen, LlamaIndex, Agno, Strands, Bedrock, etc. route through it unchanged by pointing their client's base URL at distil.
Prefer in-process? Wrap the client directly — still no call-site change:
(OpenAI — Chat Completions and Responses API — and Gemini route through the proxy: distil wrap -- codex, or point OPENAI_BASE_URL at it. An in-process client wrap exists for the Anthropic SDK only.)
Framework hooks (no proxy, no network hop) — for agent frameworks that own the message list, compress it where it lives:
| Framework | Hook | Example |
|---|---|---|
| LiteLLM | distil.integrations.litellm.compress(kwargs) | examples/python_litellm.py |
| LangChain | distil.integrations.langchain.compress_messages(msgs) | — |
| LangGraph | pre_model_hook=pre_model_hook() (compresses graph state before the model node) | examples/python_langgraph.py |
| Agno | distil.integrations.agno.compressed_model(model) | — |
| Strands | distil.integrations.strands.compressing_hook() | — |
| LlamaIndex | DistilNodePostprocessor() (node postprocessor) · DistilLLM(llm) · compressing_tool(fn) | llamaindex.html |
langchain-distilListed in LangChain's own community middleware integrations. If you came from there, this is the package:
Tool and function messages get the reversible Tier-1 digest, human and system messages are Tier-0 lossless, and assistant messages are never rewritten — a model's own words are not distil's to edit. Every digest is byte-exact recoverable. Pass verbatim=True for Tier-0-only when no recovery tool is available.
It is a thin wrapper over the hooks in the table above, so it inherits the same certified compression path — nothing is re-implemented. distil-llm is a dependency; you do not install both by hand.
On a flat-rate Pro/Max plan there is no per-token bill to cut, so distil's dollar figures are notional. The rate-limit window is not notional: tokens spent on a 40 KB test log are quota unavailable for the next task.
The proxy can't help much here. Anthropic's consumer terms (§3, item 7) restrict automated access on
subscription credentials, so distil deliberately runs --lossless-only there and measures 0.27%.
Your account isn't worth a few percent.
A PostToolUse hook is a different mechanism — a documented, first-party extension point. Claude
Code compresses its own tool output, in its own process, before the model reads it:
Measured on a paired live A/B, both arms answering correctly: tool_result −38.6%,
cache_creation −67.4%, cost-weighted −68.3%, and decision-equivalence 5/5 across five
verifiable tasks. Critically cache_read did not collapse — a hook sees each result once and cannot
rewrite history, so compression is append-only by construction and the prompt cache survives.
Where it saves nothing. Tier-0 is JSON minification plus consecutive-run collapse, so savings are
shape-dependent: verbose JSON (npm/pip/kubectl/terraform) 28–33%, duplicated log runs up to
99%, and unique-line logs, prose, git log and git diff 0%. On distil's own eval corpus it
saves 0.00% — that corpus has no JSON and no consecutive duplicates. Published because quoting
only the favourable fixtures would be the overclaim we criticise in others.
Other agents: Gemini CLI's
AfterToolcan influence output indirectly (under evaluation); Codex CLI hooks are observe-only and reject output rewriting, so it's blocked upstream there.
Full page, with the method and the caveats →
Distil ships a Model Context Protocol server so an agent can compress its own tool output and get the exact bytes back later. Zero dependencies (stdlib JSON-RPC over stdio, no SDK), fully local — content never leaves the machine.
Add it in one line:
Haven't installed distil? Run it straight from PyPI — no install step:
Config lives in ~/Library/Application Support/Claude/claude_desktop_config.json (Claude Desktop,
macOS), .cursor/mcp.json (Cursor), or .vscode/mcp.json (VS Code). Restart the client after editing.
Verify it's up — no client needed:
| Tool | Does | Your agent reaches for it when |
|---|---|---|
distil_compress(text) | Returns a compact digest + an 8-hex handle; stores the original locally (encrypted, 0600) | A tool returned something huge and carrying it verbatim is wasteful |
distil_expand(handle) | Returns the exact original bytes — not a summary | The digest lost a detail it now needs: a line, a value, a stack frame |
distil_savings() | Cumulative tokens/dollars from the local ledger | You ask "how much has distil saved me?" |
Every tool is annotated (readOnlyHint, idempotentHint, openWorldHint: false), so a well-behaved
client knows distil_expand is a safe, repeatable, offline read without having to guess from prose.
This is the recall path, not the savings path. The MCP server doesn't compress your agent's traffic —
distil wrap -- <agent>does that, transparently, with no tool calls. What the MCP server adds is the other half: any agent, including one you didn't wrap, can calldistil_expandon a handle it sees in context and get the original back. Handles persist across sessions and processes, and age out afterDISTIL_RESTORE_TTL_DAYS(default 14).
New here? pipx install distil-llm, then distil onboard — it sets you up and guides you (see Use it now). Want to see it prove itself first instead? distil bench runs the certified gate in ~10s, no API key. The matrix below is for picking an install format — everything in it is an alternative, not a requirement.
⚠️ The one gotcha — the name. The PyPI package is
distil-llmbut the command isdistil(the bare name was taken). Sopipx install distil-llm→ rundistil ….pip install distilinstalls something else.
🔧 Seeing
Could not find a version that satisfies the requirement distil-llm (from versions: none)? The package is on PyPI — that error means yourpip/pipxis on a Python older than the package's floor, so pip filters every release out. Distil now supports Python 3.9+ (the version macOS ships), so a current install just works; if you still hit this on a very old Python, let uv provision one for you:uvx --python 3.12 --from distil-llm distil bench(oruv tool install --python 3.12 distil-llm). Check yours withpython3 --version.
🔧 Got an old version (e.g.
0.25.1) instead of the latest? Public PyPI always serves the newest (pip index versions distil-llmlists them). If you got an older one, yourpip/pipxis not resolving against public PyPI — almost always a stale internal mirror (Artifactory / CodeArtifact / Nexus that hasn't synced the latest yet — common right after a release) or a<1.0version pin in a constraints file /pip.conf. Diagnose and fix:
| Format | Command | Prereq |
|---|---|---|
| Zero install | uvx --from distil-llm distil bench | uv — auto-provisions Python 3.9+ |
| Isolated CLI | pipx install distil-llm → distil bench | Python 3.9+ (else pipx install --python python3.12 distil-llm) |
| Homebrew | brew install dshakes/tap/distil | Homebrew |
| Docker | docker run ghcr.io/dshakes/distil:latest bench (or docker build -t distil .) | Docker |
| Single file | make pyz → python dist/distil.pyz bench | Python 3.9+ |
| In a venv | pip install distil-llm (inside an active virtualenv) | Python 3.9+ |
| Node / JS / TS | npx distil-llm wrap -- <agent> · npm i distil-llm for baseURL helpers | Node 18+ (bridges to Python via uv/pipx) |
The import package and CLI are
distil; the PyPI distribution isdistil-llm(the bare name was taken — souvx/pipmust referencedistil-llm, notdistil). Distil is a CLI: install it isolated (pipx/uv/brew/Docker), because modern macOS/Linux block system-widepip install(PEP 668). Node / JS / TS:npx distil-llm wrap -- <agent>(the npm package bridges to the CLI), ornpm i distil-llmfordistilBaseURL()helpers to point any SDK at the proxy — or just setbase_urlyourself.
Basics are in Use it now and Works with every SDK. Beyond that:
| Goal | Command |
|---|---|
| Set up + a guided tour (start here) | distil onboard |
Make distil the default (no per-session wrap) | distil default · undo: distil default --undo |
| Remove distil's footprint (before uninstalling) | distil offboard · also clear data: distil offboard --purge |
| Diagnose your setup (ledger, shadow, proxy self-test, wiring) | distil doctor |
| Wire the savings status line into Claude Code | distil setup (compact segment: DISTIL_STATUSLINE=minimal) |
| Watch genuine savings accumulate | distil leaderboard · live TUI: distil dashboard |
| Session summary on exit (tokens, cost, shadow, restorability) | printed automatically by distil wrap — opt out with DISTIL_NO_LEDGER=1 |
| Deep-dive one session (savings, anomalies) | distil dissect (--html / --serve) |
| Live decision-equivalence on real traffic | distil wrap --shadow 0.1 -- claude → distil shadow-stats |
| Certify on your domain | distil ingest --input prod.jsonl --out ./mycorpus → distil conformal --corpus ./mycorpus |
| Recover digested detail from any agent (MCP) | distil mcp |
| Self-improving keep policy | distil learn / distil online |
Status line — one pattern in every state:
distil · <live> · total ▼<lifetime>.
state you see means saving distil · ⬢ digest · ▼12.0K · 40% smaller · $0.31 · total ▼27.0M · de 99%compressing (mode chip: ⬢ digest·◇ lossless·▪ verbatim;de= decision-equivalence)watching distil · ✓ on · waiting for a large read · total ▼27.0Mon, but no large content yet — savings come from big file/command output idle distil · ✓ on · total ▼27.0Mset up and on, no recent traffic not routed distil · off — session not routed · total ▼27.0Mthis session's requests go straight to the provider — start it with distil wrap(or the always-on env) to compressbypassing distil · ⚠ wrapped, agent bypassing proxy · total ▼27.0Mthe wrap is up but zero requests reached its proxy in 3+ minutes — the agent pinned its own endpoint. Fix: restart the wrap. Seen mostly with claude.ai-subscription (OAuth) sessions; routing those through a custom base URL is undocumented upstream, and a session occasionally ignores it. scripts/soak-report.shcaptures evidence if it persistsThe
desegment is live decision-equivalence evidence: a ✓/⚠/✗ rate once 50 A/B samples + 30 A/A samples accrue (A/B = compressed-vs-original; A/A = same request replayed against itself — the sampling-noise baseline),de n/50while collecting. Shadow sampling is on by default at 2% (--shadow 0disables;--shadow 1.0samples every request — proves equivalence in minutes at ~3× token cost, then drop back to the default 2%).Measured: The earlier number here (signature v3 / 1.13.0, 100% over 116 sampled requests, A/A 31/31) is withdrawn — the 1.51.1 changelog found that every replay carrying a prior
thinkingblock failed with a signature error, biasing the sample toward the minority of turns that had none. Honest current reading, live shadow build 1.51.1, lossless-only: 44 A/B and 11 A/A samples; raw agreement 81.8% [67.3, 91.8]; model self-agreement on identical input 84.8% [71.8, 92.4] over 46 byte-identical replays; statistically indistinguishable (p=0.63); 32 of the 44 A/B samples had no bytes changed by compression. Below the 50 A/B + 30 A/A reporting floor — not yet a verdict. The estimator (a ratio today, not a paired difference) is replaced in 1.52.0 by a paired design (three replays per sample, unclipped difference with a bootstrap CI, one reporting floor); the numbers above are from the pre-1.52.0 unpaired estimator and will be superseded once the paired sample clears the floor.
▼= tokens saved ·total= lifetime ·de= decision-equivalence (verdict once 50 A/B + 30 A/A shadow samples accrue). Sharing the line with git/cwd/model?DISTIL_STATUSLINE=minimal→distil ▼7.8K · 27M total. On a flat-rate subscription, dollars are notional and auto-hidden (DISTIL_SUBSCRIPTION=0/1).
You usually don't need to pick. distil onboard detects your billing and sets the right mode for you — it writes it into your setup so every session just works. Pass a flag to override for a specific session.
--safe) — The cautious setting: Distil only trims things it can rebuild perfectly (like extra blank space), and never summarizes. You save less, but there's zero chance of losing any detail. Picked automatically on a flat monthly subscription.For the technical breakdown:
| Mode | What it does | Savings | Safety | Auto-selected when |
|---|---|---|---|---|
--expand | Digest + injected expand tool so the model recovers content on demand | Most | Lossy-but-recoverable | Metered / API-key (PAYG) |
(default) digest | Tier-1 digest only — no tool injection | High | Reversible via RestoreStore | No flag passed |
--lossless-only / --safe | Lossless transforms only — no digests, no tool injection | Fewer | Zero unrecoverable content | Subscription / flat-rate |
--verbatim | Whitespace + JSON normalization only | Minimal | Most conservative | Debugging / auditing |
Subscription users should not force --expand; it crosses the lossless safety boundary. Coding re-reads? Add --session-delta either way.
Two techniques carry most of the win — they target where the money actually is in an agent loop, not where it looks like it is.
You re-send the growing context every step. With prompt caching a cache read is ~10× cheaper than fresh input, so the real cost is cache misses, not context size. Distil keeps the prefix byte-stable (schema canonicalization + lifting volatile fields like timestamps/UUIDs out of the prefix) and compresses only the volatile tail.
Naive recompression sends fewer tokens yet costs more than not compressing at all, because it rewrites the cached prefix every turn. Distil doesn't — that's the whole game most tools miss.
cache_control breakpoint advances, an SDK stamps index fields, a string becomes a text block. Measured offline, that used to cost the entire prefix, on every provider, on every turn (0% forwarded byte-identical). Distil now replays the bytes it forwarded last turn for the longest canonically-equal prefix — 100% on every rewrite shape that applies, at ~0.13 ms/request. It restores bytes, never decisions, so the exact-quote guarantee still wins over a cache hit. On by default; --no-prefix-replay opts out. → ADR 0011distil cache shows you whether it's working, and deliberately mixes two kinds of number: cache reads and writes come from the provider's own usage — ground truth about money — while prefix drift is distil's own diagnosis of why, from a content-free hash of the stable blocks it sent. A diagnosis with no measurement behind it is a guess, so it never prints one without the other. On a live three-turn session where the third turn prepends a session id to the system prompt, the two agree independently: the turn the hash flagged is the turn the provider re-billed 15,819 tokens to re-create. Turns that merely grew — a conversation doing what conversations do — are not drift, because a warning that fires on every healthy turn is one people switch off. With no proxied requests it exits non-zero rather than printing a reassuring zero. Full picture, including the one cache feature we deliberately don't ship and why, in docs/CACHE.md.
The eval isn't a ruler bolted on the side; it's a discovery engine. Remove a context block, replay, did any decision change? Blocks that never change a decision are provably free to drop.
The gate answers "is this strategy non-inferior on my corpus?". The Decision-Equivalence Risk Certificate answers the operational one: "for a risk budget I choose (say ≤5% decision-change), how hard can I compress with a guarantee that holds on my real traffic?"
Every certificate names the oracle that graded it. A certificate is evidence, and evidence that doesn't say what produced it isn't evidence — so Certificate.grader is stamped from the runner and printed in the guarantee. The default offline gate is graded by a deterministic synthetic oracle, not a model, and it says so verbatim: Graded by: deterministic (synthetic DECISION: oracle — NOT a model). Real-model evidence comes from distil certify --runner anthropic, and its certificates name that runner instead. You can always tell which layer a number came from, because the number carries it.
It's conformal risk control (Learn-Then-Test / CRC — distribution-free, finite-sample), not a heuristic threshold. The one load-bearing caveat: the guarantee requires exchangeability (calibration traffic ≈ live traffic) and is marginal over that distribution — recalibrate on drift. Full theory + citations: Concepts · docs/PAPER.md.
DERC certifies the step; this certifies the task. Our E7 experiment — and the 2024–26 agent-compression literature — shows per-step fidelity can pass while end-to-end success collapses, so distil also certifies the level users actually feel: run your eval suite twice (full context vs compressed), feed the matched outcomes in, and get a distribution-free bound on how many solvable tasks compression may cost you:
It refuses to certify on small samples, states its exchangeability assumptions in the certificate itself, and ships an anytime-valid drift monitor (trajectory_risk.drift_monitor) that tells you when live traffic has shifted enough that the certificate is stale. Matched failures also feed the outcome-guided policy (distil.compress.guideline): content classes that break tasks when digested get protected byte-exact, automatically.
40+ shipped capabilities, all real (no stubs): the cache-aware cost engine, causal pruning, the TOST gate + conformal certificate, the proxy + Anthropic/OpenAI/Gemini first-class adapters (Chat Completions, Responses API, and Gemini generateContent), an MCP server, LiteLLM/LangChain/LangGraph hooks, per-agent wrap presets, the Proof Ledger end-of-session printout, the multi-tenant gateway with issued keys and rate limits, encrypt-at-rest for the restore store, learned keep-models, output compression, and an optional Rust hot-path core (build-from-source via maturin; published wheels run the pure-Python engine, same API) — with zero runtime dependencies in the core.
Full module-by-module map: Architecture · Techniques · CLI reference.
Localhost-only by default — the proxy binds 127.0.0.1 and forwards only to the single configured upstream (no SSRF).
No secret/body logging — request bodies and credentials are never logged.
Auth-mode gating — a detected subscription/OAuth session auto-selects --lossless-only (Tier-0 verbatim: no Tier-1 digest stubs, no tool injection — provider-ToS-safe); distil wrap -- claude is safe by default, no flag needed. An explicit --expand opts into the recoverable digest even there (you authorized the recovery tool, so nothing is irreversibly lost — issue #28). Without an injected expand tool the agent cannot recover a stub, so --lossless-only folds directly into verbatim.
Encrypted at rest — digest originals in ~/.distil/restore/ are encrypted with HMAC-SHA256-CTR (encrypt-then-MAC, DSTL1 header, key at chmod 0600), protecting against backup/sync leakage and cross-user reads on shared filesystems. A same-UID attacker who can read both the data files and the key file is explicitly out of scope (see THREAT_MODEL.md). Legacy plaintext files load transparently. DISTIL_NO_ENCRYPT_AT_REST=1 opts out; handles age out after DISTIL_RESTORE_TTL_DAYS (default 14). No data is forwarded upstream.
Ops-ready — unauthenticated GET /distil/health liveness probe on every entry point (never touches the billed upstream); gateway accounting checkpoints to disk every 30 s (crash-safe, not just on graceful shutdown); DISTIL_DEBUG=1 surfaces everything the fail-open compression path swallows.
Upgrades apply to live sessions — distil wrap supervises its proxy as a subprocess on a wrap-owned socket; when a new version lands on disk (pipx/pip upgrade) the wrap hot-swaps in a fresh worker — same port, in-flight streams finish on the old one, the agent never restarts. Health-checked with automatic rollback: a broken upgrade keeps the old worker serving. POSIX; kill -USR1 <wrap pid> forces it, DISTIL_HOT_SWAP=0 opts out. On Windows the wrap keeps the historical in-thread proxy (no seamless swap) and warns on version skew instead — upgrades there apply on the next session.
OpenTelemetry GenAI spans (opt-in) — pip install 'distil-llm[otel]' and every proxied call emits a GenAI semantic-convention span (gen_ai.request.model, gen_ai.usage.input_tokens) plus distil's own story: distil.tokens.original vs distil.tokens.compressed, distil.compression.ratio, distil.shadow.sampled, and distil.session.id for per-session trace correlation — your existing OTel backend sees exactly what compression did to each request. Without the extra installed it's a single boolean check, zero overhead, and an OTel failure can never break the request path. The same numbers also export as OTel counters (distil.requests, distil.tokens.baseline/.sent/.saved), recorded at the same instrumentation point as the span attributes — so tracing and metrics can't disagree — and recorded before the span check, so they still work with tracing sampled off.
Prometheus endpoint (gateway) — GET /distil/metrics serves the standard text exposition format (distil_tokens_saved_total, distil_dollars_saved_total, distil_compression_ratio, …), written against the stdlib, so the scrape path adds no dependency and cannot fail to import. Series are labelled by tenant, so the endpoint sits behind exactly the same admin gate as /distil/stats: open on loopback for local use, and on any non-loopback bind it requires --admin-token and refuses without one. That gate is the point — an unauthenticated tenant-labelled /metrics is precisely the LiteLLM leak class of bug, and it is tested for directly (403 unauthenticated, 401 on a wrong token, plus label-injection and no-secrets-in-exposition tests). Full reference: docs/metrics.html.
Vision — repeated screenshots stop costing full price — a 1024×1024 image is ~1,400 input tokens, and an agent that screenshots a UI or polls a dashboard pays that on every turn the block stays in context. Distil elides only byte-identical repeats, replacing each with a recoverable reference: the first occurrence and every distinct image are untouched, nothing is re-encoded or downscaled, and distil_expand returns the original source byte-exact. URL sources are never treated as duplicates — two occurrences of one URL are not evidence of the same pixels. The prevailing alternative resizes, which is lossy by construction and unverifiable; this is certified at 100% decision-equivalence against a live vision model (A/A floor 100%, TOST p<0.0001), and the certificate ships in the package stating its own scope. Certify your own workload with distil certify --strategy vision --runner anthropic — your result outranks ours, including a failure. DISTIL_VISION=0 disables it. Full reference: Techniques § Vision.
Supply-chain hardening — releases carry PEP 740 Sigstore attestations (via PyPI trusted publishing), a CycloneDX SBOM on every GitHub release, and OpenSSF Scorecard weekly on main. The release job fails if PyPI does not report an attestation bundle for the version it just published, so this line cannot drift into being false. Verified for every release back to 1.19.0. Don't take our word for it: curl -s https://pypi.org/integrity/distil-llm/<version>/<filename>/provenance (note the integrity API — /pypi/<pkg>/<ver>/json does not carry an attestations field, and reading it there reports a false negative), or uvx pypi-attestations verify pypi --repository https://github.com/dshakes/distil pypi:distil_llm-<version>-py3-none-any.whl.
Kubernetes — a Helm chart for the multi-tenant gateway ships in packaging/helm/distil-gateway: auth-required, non-root and read-only-rootfs by default, PDB, HPA, NetworkPolicy restricting egress to DNS + 443, plus a ServiceMonitor and alert rules. A Grafana dashboard comes with it, and CI cross-checks every panel and alert against the metrics distil actually emits.
SSO / RBAC — the gateway accepts OIDC bearer tokens alongside its own dsk- keys, with three ordered roles (viewer < operator < admin). JWS verification is stdlib-only; RS256 needs the [oidc] extra and an RS256 token is refused when it is absent rather than accepted unverified.
Audit trail — every auth success, rejection, rate-limit and key issue/revoke is appended to $DISTIL_HOME/audit.jsonl (0600, JSONL, flock-guarded). Read it with distil gateway audit, or --json straight into a SIEM. Content-free like everything else distil writes: identifiers and outcomes, never prompt text, tool output, or the raw key.
Key lifetime — distil gateway keys issue --tenant acme --expires-in-days 90 gives a key a bounded life, enforced on every lookup; keys list shows active / expired / revoked separately so a sudden 401 doesn't send anyone hunting for a revocation that never happened. Keys without an expiry keep working forever, so nothing changes for existing deployments.
See Deploy & security for topologies (local sidecar, container sidecar, shared gateway), the security whitepaper for a review-ready data-handling and compliance summary, and SECURITY.md to report a vulnerability.
Self-calibrating token counts — the offline heuristic is directionally accurate; the compression ratio is exact regardless. distil is a proxy, so it sees the provider's real usage.* on every response — it learns the systematic correction from that (content-free, no network) and calibrates the absolute counts to your model + content mix automatically. The leaderboard shows "calibrated to your billed usage (N requests, ±X%)" once enough traffic has flowed; until then it's the raw heuristic (identity, so no skew). For per-string exactness there's still --tokenizer anthropic.
Default runner is a deterministic stand-in (offline gate with ground truth). Non-circular eval grades real agent traces with a real model — proof harness.
Credible grading, enforced: majority-vote (single samples let grader noise look like a decision change), a same-family grader, and grading the reversible tier with its distil_expand recovery loop.
No fabricated weights — the keep-model is a real logistic classifier (96.4% held-out accuracy, 0.98 F1; the committed metrics.json regenerates byte-identically from python -m distil.codec.learned, seed-pinned). The optional transformer codec ships no checkpoint in the package — a demo checkpoint is attached to the v0.1.0 release, and production means retraining on your own traces (distil train-transformer).
An outside benchmark caught a defect our own gates missed. 75 agent runs, graded from the API's own usage fields, found wrap --expand completing 6 of 15 coding tasks against bare Claude Code's 13 — seven runs wrote nothing to disk and reported success. Four causes, all fixed in 1.49.0, each pinned by a regression test verified to fail without its fix. The write-up, including what it does not establish, is here. A green test suite does not prove the work was done.
Distil is a compression engine with a correctness gate, not a context suite. We declined what can't go under the certificate:
| Adjacent feature | Our stance |
|---|---|
| Persistent memory / knowledge graph | Out of scope — a lossy store is the opposite of byte-reversible. |
| Hosted semantic cache | Out of scope — we make the provider's prompt cache pay off, not a second lossy one. |
| Editor/Copilot auth | Out of scope — Distil sits on the wire or in-process; never brokers credentials. |
What we did adopt (it survives the gate): a pluggable salience scorer to protect entities, cache-prefix observability, and framework hooks.
Distil compresses input/context (comprehensive) and output — generation-side verbosity shaping (PAYG, measured with distil output-savings) plus a reversible output-on-re-entry digest, so verbose past answers stop costing full price as history. Details: Output & I/O.
Every number reproduces from the bundled corpus (distil bench, no key). The non-circular proof harness grades real agent traces with a real model (τ-bench / SWE-bench): benchmarks/PROVE.md. Compiled paper, LaTeX source, and all committed results: docs/PAPER.md · docs/paper/ · paper PDF. Step-by-step: Reproduce the Numbers →
pipx install distil-llm && distil bench
certified savings across 9 domains in ~10 seconds — zero API key, zero runtime deps
Get started → · Wire it into your SDK · Read the proof · PyPI
A star is how the next engineer finds provable savings instead of a lossy guess — and
distil stats --badge gives you a shareable badge of your own measured number to
show alongside it. That badge + this repo are the whole marketing department.
PRs welcome — see CONTRIBUTING.md. The one rule that matters: a new compression strategy must pass make gate (non-inferior on every domain, byte-reversible). No green gate, no merge. That's the whole philosophy in one sentence.
Running the evals yourself — every gate is free and offline: see docs/RUNNING-EVALS.md.
pip install -U distil-llm (or uv tool upgrade distil-llm) tracks them.Apache-2.0 · “Same potency, less volume.”