Scan prompts, tool definitions and model output for injection, and guard agent tool calls.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
One-click editor setup isnβt available for this listing yet β we donβt have a confirmed install command, and weβd rather show nothing than point your editor at the wrong package or host. Follow the projectβs own setup instructions, linked above.
A provenance-tracked tool-call guard for JavaScript agents. Tool results are tainted, destinations the user named are trusted, and every tool call the model proposes is checked before it runs: untrusted data does not reach a network, email, exec, file or payment sink. Fail-closed, in-process, zero runtime dependencies. Text detection, spotlighting and canaries ship as components, with their benchmark numbers published whether they flatter the library or not.
Prompt injection stopped being a text-classification problem the moment models started calling tools. The damage in an agent is a consequence, not a sentence: a URL lifted from an email ends up in an outbound request, an attendee address from a calendar entry becomes the recipient of send_email, a payload in a README lands in exec. Version 3 tracks where data came from and refuses to let untrusted data reach a dangerous sink. That is the capability model from Google DeepMind's CaMeL paper (arXiv 2503.18813), ported to a JavaScript tool-calling loop without the custom interpreter.
It is a policy layer inside your process. It is not an isolation boundary. It cannot see a flow that passes through the model's hidden state, a paraphrase that shares no identifiers with its source, or a source you never registered. The Limitations section lists what it misses, with the dataset rows that show it.
Or drive it by hand:
The default policies run in this order. plan-violation blocks any tool outside the plan() allow-list. injection-source-flow blocks when a source that itself scored as injection flows anywhere. untrusted-to-exfil-sink blocks untrusted data reaching network, email or messaging tools, and untrusted-to-exec does the same for exec and file writes. untrusted-to-payment asks for confirmation. injection-then-sink flags a sink call made in the same turn as an injection-scored source even when no flow was detected, because paraphrase is exactly the case the flow detectors miss. args-injection flags when the arguments themselves read as injection. Each of these is one of the patterns in Design Patterns for Securing LLM Agents (arXiv 2506.08837): Action-Selector via plan(), Plan-Then-Execute, Context-Minimisation via spotlighting.
Spotlighting (prompt-protection/spotlight, arXiv 2403.14720) marks untrusted spans by delimiting, datamarking or base64-encoding them, and gives you the system-prompt sentence that tells the model what the marker means. With wrapTools({ spotlight: 'datamark' }) the model sees marked text while the guard taints the original, and arguments are unmarked before flow detection so a copied span still matches.
Canaries (prompt-protection/canary) put a token in the system prompt and look for it in the output in exact, normalized, spaced, base64, hex, reversed and partial forms. There is also a shingle-similarity check between the output and the system prompt. Plain verbatim canaries were shown to fail against paraphrase (arXiv 2506.19109). Similarity closes part of that gap. Not all of it.
Text detection is a component, not the product: 106 input rules, 21 output rules and 9 tool-poisoning rules score text that has to be scored, and the Benchmark section says how well. An embedded 33 KB int8 n-gram classifier lives on prompt-protection/ml (preview tier, off by default, for reasons the benchmark explains). prompt-protection/lite is the rules-only entry at 20 KB gzipped.
npm run bench runs the shipped build against every set below and writes bench/results.json. The same run is a CI gate. Recall and false-positive rate are shown as regex / ml / hybrid; the shipped default is regex.
| Set | Licence | N (attack/benign) | Recall | False-positive rate |
|---|---|---|---|---|
| NotInject, over-defence benchmark (arXiv 2410.22770) | MIT | 339 (0/339) | n/a | 2.9% / 7.7% / 10.6% |
datasets/benign-hard + datasets/attacks, ours, written to evade proximity matching | CC-BY-4.0 | 285 (130/155) | 14.6% / 19.2% / 30.8% | 19.4% / 8.4% / 25.2% |
| in-the-wild jailbreaks, 900-row sample | MIT | 900 (300/600) | 46.3% / 48.7% / 66.3% | 20.5% / 26.3% / 37.2% |
| local held-out (never a test fixture) | MIT | 35 (20/15) | 75.0% / 45.0% / 80.0% | 6.7% / 6.7% / 13.3% |
| local tuning (doubles as test fixtures) | MIT | 134 (77/57) | 100% / 31.2% / 100% | 0.0% / 7.0% / 7.0% |
| Tool poisoning | MIT | 10 (5/5) | 100% | 0% |
| Output scan, canary variants, system-prompt similarity, credential/PII/relay rules | MIT | 18 (10/8) | 100% | 0% |
Agent flows, datasets/agent-flows.jsonl, 100 tool-call scenarios | CC-BY-4.0 | 100 (50/50) | agreement 100% on 99 scored rows, 1 documented miss Β· attack block-recall 82% Β· benign FPR 4% |
Some of these numbers are bad, and they are here on purpose. On NotInject the rules do well: 97.1% of short benign queries that merely contain "ignore" or "instruction" pass through. On the hard-negative set I wrote myself they false-positive on 19.4% of benign text. Questions about prompt injection, fiction, "grant admin access on Netflix" all trip them. They catch 14.6% of the attacks written to avoid canonical phrases. That is what pattern matching tops out at, and it is why provenance is the primary mechanism now. Both figures are CI gates at their current baseline; they can only go down from here.
The in-the-wild "regular" set is noisy. It includes SEO prompts that open with "Please ignore all previous instructions", so its false-positive column overstates. I kept it because it is external and unmodified.
The embedded model ships for transparency, not for use. It is trained on Apache and MIT datasets (deepset, gandalf, hackaprompt, SPML, plus about 17k mined benign rows) with a reproducible pipeline described in training/REPORT.md. In-distribution it looks great: 3-fold CV F1 0.98. Held out by dataset it does not: leave-one-dataset-out F1 0.53, in-the-wild AUROC 0.67. Adding hackaprompt in a second round lifted recall on unseen attacks from 4% to 19% on our set and lifted in-the-wild false positives from 17% to 24% with it. A bag of hashed n-grams does not transfer across jailbreak genres, so ml defaults to 'off'. If you want it anyway, analyzePrompt(text, { ml: 'escalate' }). Python and JS produce identical features and logits on 64 golden vectors under test, and the weights are 33 KB gzipped.
Latency: rule scan p99 about 0.1 ms, guard checkToolCall p99 about 3 ms with 64 registered sources, classifier about 0.15 ms. Bundle: core 65 KB gzipped with the weights included, lite 20 KB, guard 67 KB.
datasets/ is CC-BY-4.0 and disjoint from the test fixtures. It is also on the Hugging Face Hub as promptprotection/agent-security-datasets, with a card built from the benchmark results. attacks.jsonl has 130 rows across nine categories and 14 languages. benign-hard.jsonl has 155 benign prompts carrying trigger vocabulary, in NotInject's four categories plus developer jargon and security documentation. agent-flows.jsonl has 100 tool-call scenarios with the expected guard decision and the reason. node datasets/validate.mjs checks schema, uniqueness and disjointness from the fixtures.
Semantic paraphrase. Tainted prose rewritten so it shares no identifiers and no six-word shingles with its source is invisible to the guard. injection-then-sink covers the same-turn case only when the source itself scores as injection; af-037 in the agent-flows set is the documented miss.
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/prompt-protection)<a href="https://allmcps.com/mcp/prompt-protection"><img src="https://allmcps.com/api/badge/prompt-protection?style=directory" alt="Prompt Protection on AllMCPs" /></a>