The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Injectshield listing page.
Prompt-injection firewall for AI agents.
A drop-in REST API that detects and neutralizes injection attacks in any text — git commits, web pages, files, emails, user inputs — before they reach your AI agent's context window.
This repo is the open-source heuristic ruleset plus the source for the managed API at promptshield.pages.dev.
In May 2026 a viral HN thread demonstrated that a single git commit message could burn a Claude Code user's entire session quota via a schema-driven attack ("OpenClaw"). The pattern is general: any AI agent that ingests untrusted text — code review bots, documentation summarizers, RAG agents, support copilots — is exposed to prompt injection. Most teams ship without any input-side defense.
InjectShield is one layer of a defense-in-depth strategy. It's not a silver bullet. Use it alongside system-prompt hardening, tool sandboxing, and output filtering.
InjectShield ships a native MCP server at @injectshield/mcp. Once installed, your agent has three new tools — scan, scan_url, patterns — for input-side defense without writing any glue code.
For Cursor / Cline / other MCP clients, see packages/injectshield-mcp/README.md.
Or signup via the landing page: https://injectshield.dev — self-serve, email delivery.
Live:
https://api.injectshield.devOpen-source (this repo, MIT):
src/patterns.ts — the heuristic pattern library (~20 categorized rules).src/detect.ts — the detection engine (heuristic aggregation, sanitization).test/ — the test suite.server/, public/ — the full API + landing-page source.Managed only (paid tiers):
| Category | Examples |
|---|---|
instruction_injection | "ignore previous instructions", "new system prompt" |
system_override | system-prompt leak, role-tag forgery, ChatML/Llama special tokens |
role_hijack | "you are now…", DAN, Developer Mode |
exfiltration | data sent to attacker URLs, markdown image exfil |
schema_attack | OpenClaw-style schema references |
encoding_smuggle | base64-decoded directives |
invisible_text | zero-width / bidi / Unicode-Tag smuggling |
tool_abuse | synthetic tool-call directives in untrusted text |
jailbreak_classic | DAN, "no restrictions", etc. |
Found a novel attack? Open a PR adding a PatternRule to src/patterns.ts with:
id.category from the enum above.weight in [0, 1] — pick conservatively; the aggregation in detect.ts combines weights so every additional rule contributes meaningfully but isn't dominant.test/detect.test.ts covering both a positive and a likely-benign negative example.We auto-deploy merged patterns to the managed API. No-cost contributions get attribution in the changelog.
MIT. InjectShield reduces but does not eliminate prompt-injection risk.
Built on Cloudflare Pages (frontend) + Railway (API) + Postgres + Anthropic Claude (semantic layer). Pattern library informed by HackAPrompt, the PINT benchmark, and a long list of public attack examples.