The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the AgentShield listing page.
Stop prompt injections before they hit your LLM.
AgentShield is a fast, low-latency classifier that flags prompt-injection, jailbreak, and data-exfiltration attempts in ~50 ms — before they reach your LLM or agent.
benchmark/.Public API: https://api.agentshield.pro/v1/classify. Live site: agentshield.pro.
Async, retries, and middleware patterns: see packages/agentshield-sdk/README.md.
| Path | Purpose |
|---|---|
packages/agentshield-sdk/ | Official Python SDK (pip install agentshield-guard) — sync + async client, typed responses |
services/landing-page/ | FastAPI landing site, live demo proxy, self-serve signup, customer dashboard |
benchmark/ | Reproducible benchmark harness — datasets, runner, analysis, published report |
examples/ | Integration examples (LangChain, OpenAI SDK, FastAPI middleware) |
The core classification gateway is operated as a managed service; the SDK and benchmark give you everything you need to integrate and verify our numbers.
We publish our numbers and the exact code we used. To reproduce:
Results land in benchmark/results/. The published writeup is in benchmark/report/summary.md.
See agentshield.pro/blog for development updates.
Bug reports, dataset additions, and integration examples are welcome. Open an issue or a PR against main. For security issues, email security@agentshield.pro — please do not open public issues for vulnerabilities.
MIT — see LICENSE. Copyright © 2026 Eigenart Filmproduktion.
Third-party datasets in benchmark/datasets/ retain their original licenses (deepset/prompt-injections, PINT, jackhhao/jailbreak-classification, SPML Chatbot Prompt Injection). Pointers and attribution live in benchmark/datasets/ — please review each before redistributing.