The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Cap Shield listing page.
Context selection and compression for AI agents — with the recall measured, not claimed.
One file. No dependencies. No SDK required.
More context makes agents worse. ETH Zurich found context files LOWER task success versus giving the agent no repository context at all, while raising inference cost by over 20 %. Around two thirds of production agent failures trace to context problems, not to the model being incapable.
So the question is not how much you cut. It is whether what you kept was enough — and that is measured here, on a benchmark we did not choose: LongMemEval-S, 500 questions. Recall@10 of 93.8 % against a lexical baseline of 51.9 %.
Recall@10 is the strict measure: a question counts only when ALL gold sessions were found. Finding half the answer means the agent answers confidently on half a basis.
Two of the five tools need no account. Measure first, decide after.
Python 3.9+. Nothing else.
The optional SKILL.md tells an agent when to use these tools — and when not to.
The server is also reachable over HTTP:
Add it as a remote MCP server in any client that supports them, or open it
in an MCP inspector. measure_traffic and list_packages work with no
account and no key — you can measure your own traffic in a browser without
installing anything.
remember and assemble_context need a key and are not open over HTTP
yet; they answer with what to do instead.
A customer's dictionary is trained on their own traffic, in their own isolated store. A new version is adopted only if it measures better on held-out data neither version was trained on. A retraining that does not win is rejected and logged, and the old dictionary stays.
Old versions are never deleted, so packets compressed under any earlier version still unpack.
measure_traffic · no accountMeasure how much of your own agent traffic could be saved. NO ACCOUNT OR KEY NEEDED — use this first. Returns byte savings over the wire and, if a query is given, token savings from selective context retrieval. Nothing is stored: the text is compressed in memory and discarded. Rate limited to 20 calls per hour per IP.
list_packages · no accountList the available dictionaries with their MEASURED compression, including the ones that perform badly. Each entry says whether it works one message at a time or only batched, and how many messages came out LARGER. No key needed.
remember · requires a keyStore a memory entry for later retrieval. REQUIRES A KEY. This does not call any language model — it stores text in an isolated per-tenant archive. Use assemble_context to get relevant entries back.
assemble_context · requires a keyRetrieve the memory entries that answer a question, within a token budget. REQUIRES A KEY. Send the returned 'context' to your language model INSTEAD of the whole history. This does not call a model itself — it selects what to send. The budget is a ceiling, not a target: selection stops where relevance runs out, often well below it. The response says how many entries were left behind and why.
get_accountGet an account and an API key. Requires an email address. The key is returned ONCE and cannot be shown again — store it immediately. Beta quotas are low by design; they are hard stops, never overage billing.
The descriptions above are copied verbatim from the server. If they ever
differ from what tools/list returns, the server is right and this file
is stale.
An entry can be updated without storing it again in full. Every fifth version is complete; the ones between are stored as a delta against the previous version, with a SHA-256 checked on read. Lossless, and no version depends on more than four others.
Agent memory is mostly small edits to text that already exists. Storing each edit in full is what makes it expensive to keep.
Entries also pick their own compression strategy by size — and if compression does not pay off, the entry is stored raw and the response says so.
A saving you cannot check is a saving you have to take on trust. Every
hit carries method — vector or lexical — so you can see which mechanism
found it. Every assembly reports baseline_tokens (what the whole
history would have cost, in the same format), candidates_before_autocut,
autocut_removed, deduplicated and skipped_too_big.
Measuring the baseline in a different format once produced a 39-point error. The formats are identical for that reason.
Compression saves bytes over the wire. Selection saves tokens in the context. Two different mechanisms — adding them together produces a number that means nothing.
Compressed packets are decompressed before a model sees them, so this does not reduce inference cost. Saying otherwise is the easiest way to be wrong about this project.
Every figure is published live, including what has not been measured and which packages perform badly:
https://cap-shield-robin.fly.dev/.well-known/cap-shield.json
Fetch that rather than trusting this file. It goes stale; the document does not.
No account, nothing stored. The response includes the degraded share — how many of your messages came out larger.
Batching compresses several messages in the same context, which opens a CRIME/BREACH-style side channel: someone who can place chosen text in the same batch as a secret, and observe the batch size, learns something about the secret.
Only batch messages that already share a trust boundary. Optional padding closes the leak for under two bytes a message, and it is off by default — we say so rather than let you assume otherwise.
Individual packing does not have this problem at all.
Dictionary versions are never deleted, and the guarantee does not rest on us still being here: the archive export carries the dictionary binaries, and a standalone unpacker runs with no gateway, no network and no other part of the system.
https://cap-shield-robin.fly.dev/cap_unpack.py
It is served without a token, because whoever needs it most is whoever no longer has an account.
Beta. Server version 0.1.0.
Docs: https://cap-shield-robin.fly.dev/docs/quickstart Console: https://cap-shield-console.lovable.app
MIT — see LICENSE.