Skip to main content
AllMCPs
BrowseBestCategoriesStackCompareToolsGuidesBlog
Log in Submit MCP

Stay in the loop

Get new MCP servers and top picks in your inbox.

AllMCPs

The open directory for discovering and installing Model Context Protocol servers.

AllMCPs on GitHub (opens in a new tab)
Launched onTiny Startupstinystartups.com
Explore
  • Browse servers
  • Best MCP servers
  • Categories
  • MCP clients
  • Agent prompts
  • Stack Builder
  • Compare servers
  • Random discovery New
  • Submit a server
  • Pricing & Boost Boost
Learn
  • Guides hub
  • What is MCP?
  • Install guide
  • Build an MCP server
  • Deploy an MCP server
  • Security guide
  • Troubleshooting
  • MCP for SEO & AEO
  • Protocol versioning
  • Transports: stdio vs HTTP
  • State of MCP (stats)
  • Blog & updates
Tools
  • All developer tools
  • Config generator
  • Config validator
  • Config auditor
  • MCP playground
  • Token calculator
  • OpenAPI β†’ MCP
  • Badge generator
For agents
  • REST API docs
  • Trust & traffic Live
  • Remote MCP server SSE β†— (opens in a new tab)
  • llms.txt β†— (opens in a new tab)
  • Catalog JSON β†— (opens in a new tab)
Company
  • About
  • Advertise Sponsor
  • Contact
  • GitHub β†— (opens in a new tab)
  • Terms
  • Privacy
AllMCPs VerifiedAllMCPs VerifiedFeatured on Nick LaunchesFeatured on Nick LaunchesLaunch Llama NewsletterLaunch Llama NewsletterVerified DR - allmcps.comVerified DR - allmcps.comFeatured on SaaSGrowFeatured on SaaSGrowFeatured on Twelve ToolsFeatured on Twelve ToolsFeatured on Saaspa.geFeatured on Saaspa.geFeatured on Findly.toolsFeatured on Findly.toolsFeatured on Startup FameFeatured on Startup FameFeatured on LaunchKiwiFeatured on LaunchKiwiFeatured on ScrollLaunchFeatured on ScrollLaunchFeatured on DailyPingsFeatured on DailyPingsFazier badgeFazier badgeFeatured on NewTool.siteFeatured on NewTool.siteFeatured on saasfame.comFeatured on saasfame.comDR Checker - Domain RatingDR Checker - Domain RatingListed on Turbo0Listed on Turbo0Launched on LaunchBoard - Product Launch PlatformLaunched on LaunchBoard - Product Launch PlatformList on SimilarlabsList on Similarlabshttps://codetrendy.comhttps://codetrendy.comListed on DevTool.ioFeatured on BuildlistFeatured on BuildlistLaunched on Tiny StartupsFeatured on ShowMeBestAIFeatured on ShowMeBestAIFind us on LaunchZoneFind us on LaunchZoneAllMCPs VerifiedAllMCPs VerifiedFeatured on Nick LaunchesFeatured on Nick LaunchesLaunch Llama NewsletterLaunch Llama NewsletterVerified DR - allmcps.comVerified DR - allmcps.comFeatured on SaaSGrowFeatured on SaaSGrowFeatured on Twelve ToolsFeatured on Twelve ToolsFeatured on Saaspa.geFeatured on Saaspa.geFeatured on Findly.toolsFeatured on Findly.toolsFeatured on Startup FameFeatured on Startup FameFeatured on LaunchKiwiFeatured on LaunchKiwiFeatured on ScrollLaunchFeatured on ScrollLaunchFeatured on DailyPingsFeatured on DailyPingsFazier badgeFazier badgeFeatured on NewTool.siteFeatured on NewTool.siteFeatured on saasfame.comFeatured on saasfame.comDR Checker - Domain RatingDR Checker - Domain RatingListed on Turbo0Listed on Turbo0Launched on LaunchBoard - Product Launch PlatformLaunched on LaunchBoard - Product Launch PlatformList on SimilarlabsList on Similarlabshttps://codetrendy.comhttps://codetrendy.comListed on DevTool.ioFeatured on BuildlistFeatured on BuildlistLaunched on Tiny StartupsFeatured on ShowMeBestAIFeatured on ShowMeBestAIFind us on LaunchZoneFind us on LaunchZone
Β© 2026 Jackalope Digital LLC. All rights reserved.
  1. Home
  2. πŸ’» Developer Tools
  3. Prompt Protection
P
Health: Not checked yetWe have not completed a health check for this listing yet.No health check has run yet.

Prompt Protection

User RatingsBe the first to rate and review this MCP server! Enrichment pendingWe haven’t run our AI enrichment pass on this listing yet, so the overview, use cases, and FAQ below may be sparse or missing. We work through the catalog over time β€” check back soon.
View RepositoryVisit Website

Scan prompts, tool definitions and model output for injection, and guard agent tool calls.

Quick Install

Automated & IDE Setup

Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β€” or use 1-click editor setup below.

One-click editor setup isn’t available for this listing yet β€” we don’t have a confirmed install command, and we’d rather show nothing than point your editor at the wrong package or host. Follow the project’s own setup instructions, linked above.

Manual Client & Custom JSON ConfigExpand JSON β–Ύ
No confirmed setup config for this listing yet. We only publish a config block when the install details come from the project itself β€” its README, its docs, or a verified owner. We haven’t found those for prompt-protection, and we’d rather show nothing than a guess you’d paste into your client. Follow the project’s own setup instructions for the current steps.
Install Directory Badge Claim listing AlternativesπŸ’» More in Developer Tools

Documentation Overview

prompt-protection

A provenance-tracked tool-call guard for JavaScript agents. Tool results are tainted, destinations the user named are trusted, and every tool call the model proposes is checked before it runs: untrusted data does not reach a network, email, exec, file or payment sink. Fail-closed, in-process, zero runtime dependencies. Text detection, spotlighting and canaries ship as components, with their benchmark numbers published whether they flatter the library or not.

CI npm npm downloads License: MIT TypeScript Zero dependencies

Live Demo β†’

What it is, and what it isn't

Prompt injection stopped being a text-classification problem the moment models started calling tools. The damage in an agent is a consequence, not a sentence: a URL lifted from an email ends up in an outbound request, an attendee address from a calendar entry becomes the recipient of send_email, a payload in a README lands in exec. Version 3 tracks where data came from and refuses to let untrusted data reach a dangerous sink. That is the capability model from Google DeepMind's CaMeL paper (arXiv 2503.18813), ported to a JavaScript tool-calling loop without the custom interpreter.

It is a policy layer inside your process. It is not an isolation boundary. It cannot see a flow that passes through the model's hidden state, a paraphrase that shares no identifiers with its source, or a source you never registered. The Limitations section lists what it misses, with the dataset rows that show it.

60-second agent quickstart

server.ts
import { generateText } from 'ai';
import { createGuard } from 'prompt-protection/guard';

const guard = createGuard({
  sinks: { http_post: 'network', send_email: 'email', run_shell: 'exec' }, // or rely on name heuristics
});

guard.analyzeUserTurn(userMessage);           // destinations the user names become trusted

const result = await generateText({
  model,
  tools: guard.wrapTools(tools),              // check β†’ execute β†’ taint result
  toolApproval: guard.vercelToolApproval(),   // block β†’ denied, confirm β†’ user-approval
  prompt: userMessage,
});

Or drive it by hand:

server.ts
guard.taint('read_email', emailBody);         // label a tool result as untrusted
const decision = guard.checkToolCall({ toolName: 'http_post', args: { url } });
// { action: 'block', policy: 'untrusted-to-exfil-sink', flows: [{ kind: 'identifier', path: 'args.url', … }] }

Architecture

Code
tool result ──taint──▢ provenance label ──▢ shingles + identifiers (URL, host, email, path, token)
                                                       β”‚
model proposes tool call ──▢ sink class ──▢ flow detection (exact / identifier / content) ──▢ policies ──▢ allow Β· flag Β· confirm Β· block
                                  β–²                                                             β–²
                       explicit map or name heuristics                          plan() allow-list Β· user-trusted destinations

The default policies run in this order. plan-violation blocks any tool outside the plan() allow-list. injection-source-flow blocks when a source that itself scored as injection flows anywhere. untrusted-to-exfil-sink blocks untrusted data reaching network, email or messaging tools, and untrusted-to-exec does the same for exec and file writes. untrusted-to-payment asks for confirmation. injection-then-sink flags a sink call made in the same turn as an injection-scored source even when no flow was detected, because paraphrase is exactly the case the flow detectors miss. args-injection flags when the arguments themselves read as injection. Each of these is one of the patterns in Design Patterns for Securing LLM Agents (arXiv 2506.08837): Action-Selector via plan(), Plan-Then-Execute, Context-Minimisation via spotlighting.

Spotlighting (prompt-protection/spotlight, arXiv 2403.14720) marks untrusted spans by delimiting, datamarking or base64-encoding them, and gives you the system-prompt sentence that tells the model what the marker means. With wrapTools({ spotlight: 'datamark' }) the model sees marked text while the guard taints the original, and arguments are unmarked before flow detection so a copied span still matches.

Canaries (prompt-protection/canary) put a token in the system prompt and look for it in the output in exact, normalized, spaced, base64, hex, reversed and partial forms. There is also a shingle-similarity check between the output and the system prompt. Plain verbatim canaries were shown to fail against paraphrase (arXiv 2506.19109). Similarity closes part of that gap. Not all of it.

Text detection is a component, not the product: 106 input rules, 21 output rules and 9 tool-poisoning rules score text that has to be scored, and the Benchmark section says how well. An embedded 33 KB int8 n-gram classifier lives on prompt-protection/ml (preview tier, off by default, for reasons the benchmark explains). prompt-protection/lite is the rules-only entry at 20 KB gzipped.

Benchmark

npm run bench runs the shipped build against every set below and writes bench/results.json. The same run is a CI gate. Recall and false-positive rate are shown as regex / ml / hybrid; the shipped default is regex.

SetLicenceN (attack/benign)RecallFalse-positive rate
NotInject, over-defence benchmark (arXiv 2410.22770)MIT339 (0/339)n/a2.9% / 7.7% / 10.6%
datasets/benign-hard + datasets/attacks, ours, written to evade proximity matchingCC-BY-4.0285 (130/155)14.6% / 19.2% / 30.8%19.4% / 8.4% / 25.2%
in-the-wild jailbreaks, 900-row sampleMIT900 (300/600)46.3% / 48.7% / 66.3%20.5% / 26.3% / 37.2%
local held-out (never a test fixture)MIT35 (20/15)75.0% / 45.0% / 80.0%6.7% / 6.7% / 13.3%
local tuning (doubles as test fixtures)MIT134 (77/57)100% / 31.2% / 100%0.0% / 7.0% / 7.0%
Tool poisoningMIT10 (5/5)100%0%
Output scan, canary variants, system-prompt similarity, credential/PII/relay rulesMIT18 (10/8)100%0%
Agent flows, datasets/agent-flows.jsonl, 100 tool-call scenariosCC-BY-4.0100 (50/50)agreement 100% on 99 scored rows, 1 documented miss Β· attack block-recall 82% Β· benign FPR 4%

Some of these numbers are bad, and they are here on purpose. On NotInject the rules do well: 97.1% of short benign queries that merely contain "ignore" or "instruction" pass through. On the hard-negative set I wrote myself they false-positive on 19.4% of benign text. Questions about prompt injection, fiction, "grant admin access on Netflix" all trip them. They catch 14.6% of the attacks written to avoid canonical phrases. That is what pattern matching tops out at, and it is why provenance is the primary mechanism now. Both figures are CI gates at their current baseline; they can only go down from here.

The in-the-wild "regular" set is noisy. It includes SEO prompts that open with "Please ignore all previous instructions", so its false-positive column overstates. I kept it because it is external and unmodified.

The embedded model ships for transparency, not for use. It is trained on Apache and MIT datasets (deepset, gandalf, hackaprompt, SPML, plus about 17k mined benign rows) with a reproducible pipeline described in training/REPORT.md. In-distribution it looks great: 3-fold CV F1 0.98. Held out by dataset it does not: leave-one-dataset-out F1 0.53, in-the-wild AUROC 0.67. Adding hackaprompt in a second round lifted recall on unseen attacks from 4% to 19% on our set and lifted in-the-wild false positives from 17% to 24% with it. A bag of hashed n-grams does not transfer across jailbreak genres, so ml defaults to 'off'. If you want it anyway, analyzePrompt(text, { ml: 'escalate' }). Python and JS produce identical features and logits on 64 golden vectors under test, and the weights are 33 KB gzipped.

Latency: rule scan p99 about 0.1 ms, guard checkToolCall p99 about 3 ms with 64 registered sources, classifier about 0.15 ms. Bundle: core 65 KB gzipped with the weights included, lite 20 KB, guard 67 KB.

Datasets

datasets/ is CC-BY-4.0 and disjoint from the test fixtures. It is also on the Hugging Face Hub as promptprotection/agent-security-datasets, with a card built from the benchmark results. attacks.jsonl has 130 rows across nine categories and 14 languages. benign-hard.jsonl has 155 benign prompts carrying trigger vocabulary, in NotInject's four categories plus developer jargon and security documentation. agent-flows.jsonl has 100 tool-call scenarios with the expected guard decision and the reason. node datasets/validate.mjs checks schema, uniqueness and disjointness from the fixtures.

Limitations

Semantic paraphrase. Tainted prose rewritten so it shares no identifiers and no six-word shingles with its source is invisible to the guard. injection-then-sink covers the same-turn case only when the source itself scores as injection; af-037 in the agent-flows set is the documented miss.

Read the full README β†’View source on GitHub β†’

Related MCP Servers

View all in Developer Tools View all alternatives
  • O
    Openapi MCP Server

    Connect any HTTP/REST API server using an Open API spec (v3)

    πŸ’» Developer Tools3 views
    Compare vs Openapi MCP Server β†’
  • C
    Claude Task Master

    AI-powered task management system for AI-driven development. Features PRD parsing, task expansion, multi-provider support (Claude, OpenAI, Gemini, Perplexity, xAI), and selective tool loading for optimized context usage.

    πŸ’» Developer Tools8 views
    Compare vs Claude Task Master β†’
  • M
    MCP Server Docker

    Integrate with Docker to manage containers, images, volumes, and networks.

    πŸ’» Developer Tools3 views
    Compare vs MCP Server Docker β†’
  • C
    Codemore

    The static analyzer your AI agent reads β€” fix-ready, machine-readable scan reports over MCP.

    πŸ’» Developer Tools2 views
    Compare vs Codemore β†’

Reviews

No reviews yet β€” be the first to share how this listing worked for you.

Frequently Asked Questions about Prompt Protection

We don't have a confirmed install command for prompt-protection yet, so we don't publish a generated one β€” a guessed package name would point at the wrong package or none at all. Follow the project's own README or setup instructions (https://github.com/mughalhere/prompt-protection) for the current steps.

AllMCPs Directory Badge

Full Badge Customizer

Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.

Badge Style:
Live Dynamic SVG PreviewPrompt Protection AllMCPs Directory Badge
Markdown (GitHub README)
[![AllMCPs](https://allmcps.com/api/badge/prompt-protection?style=directory)](https://allmcps.com/mcp/prompt-protection)
HTML Embed
<a href="https://allmcps.com/mcp/prompt-protection"><img src="https://allmcps.com/api/badge/prompt-protection?style=directory" alt="Prompt Protection on AllMCPs" /></a>

Technical Specs & Signals

CategoryπŸ’»Developer Tools
More technical detailsExpand β–Ύ
Last updatedSep 28, 2026
Views0
Unique ViewsTotal visits recorded for this listing page on AllMCPs.
Installs0
Installs & Copy ActionsTotal times users copied install commands or configuration snippets for this server.
27Quality signal: Emerging Β· 27/100How this signal is calculated β–Ύ
Server availabilityNot measured

Not scored for repo-hosted servers β€” we can't reach the running server, only its GitHub page. Hosted MCP endpoints are health-checked live.

Verified ownership8/20
Documentation & tools11/30
Adoption & activity1/15
Community engagement0/10

A guidance signal from public completeness & health data β€” not a user rating. New listings start lower and rise as they add docs, get verified, and grow adoption. Signals we can't observe for a listing are skipped, not counted against it.

β˜… Featured
M

Moxie Docs MCP

MCP & Agent Skills for Automated Documentation, and codebase conventions + context

Explore Server β†’

Own this project?

This directory is pre-filled from public sources. Claim via GitHub README, site badge, or DNS TXT to unlock edit access and the Official badge β€” proof is checked automatically, then reviewed by our team.

Free dofollow backlink: add your website and place the AllMCPs badge on it β€” no claim needed. We detect it automatically and keep it verified as long as the badge stays live.

Claim & get free dofollow

Share & Embed

Add our SVG badge (dark/light directory styles) or embeddable widget to your site.

Explore more

More in πŸ’» Developer Tools β†’Best MCP servers for Developers β†’Alternatives to Prompt Protection β†’Install in Claude DesktopInstall in CursorInstall in VS CodeSetup guides for all 13 MCP clients