AI security scanner - secrets, PII, prompt injection, and exfiltration detection.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
π‘ Paste into ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows)
Infrastructure security for AI agents. We attack what we defend.
Make agent behavior verifiable, auditable, and cryptographically provable across any harness, any platform. Built as a TypeScript monorepo with MCP integration, blockchain anchoring, and β as of Q4 β a published red-team engine that tests our own defenses.
The thesis: every security platform claims its defenses work. We're the first to publish the attacks that prove it. Same suite. Same scorecard. Same locked benchmark β applied to our own packages, every release.
The CLI detects your framework (LangChain, LlamaIndex, MCP, OpenAI, Anthropic, Microsoft, Google) and scaffolds the right security middleware for your stack. Or install everything at once:
The first agent security platform to publish its own offensive engine. Two new packages flipped the suite from purely defensive to defense + offense in the same monorepo, validated against each other:
| Package | Role | Version |
|---|---|---|
@weave_protocol/adversary | Offensive engine β 68 documented + novel attacks across 5 categories (IPI, tool-coercion, jailbreak, extraction, goal-corruption). Real Playwright browser target with 4 breach signal channels. Real-LLM demo mode via Anthropic API. | β v0.2.1 |
@weave_protocol/agentsecbench | Standardized benchmark β locked attack suites, tier grades AβF, paste-ready reports, side-by-side comparison | β v0.1.0 |
Trophy attacks β documented in-the-wild incidents reproduced in the corpus:
Why this matters: every model release, every WARD policy change, every adapter update can be re-benchmarked against the same locked suite. Did your score regress? agentsecbench compare will show you. Does your WARD policy actually defend anything? --measure-ward-delta will tell you. This is how a category gets defined.
See Adversary README β Β· See AgentSecBench README β Β· See METHODOLOGY.md β
Every enterprise agent question today is "what's my ceiling on this thing?" β measured in dollars, not just tool calls. @weave_protocol/witan@1.1.0 answers it. Per-window budgets (run / hour / day / week / month) that gate LLM calls and tool calls, with three actions: block, require approval/consensus, or notify.
Multi-provider LLM pricing built in β Anthropic, OpenAI, Google, and local (free). Per-tool amount caps (send_payment max $500/day). Interactive TTY approval prompt for human-in-loop terminals. Async callback for Slack/PagerDuty/custom UIs. Safe defaults β never silent approval in non-interactive contexts.
Backward compatible with the existing behavioral_limits.maxCostUSD. In-memory storage in v1.1 with a pluggable interface for the v1.2 Redis/SQLite backends. Programmatic API via import { SpendingTracker } from '@weave_protocol/witan/spending'.
See Witan spending caps README β
The thesis was that WARD.md could be a portable agent security standard β write it once, enforce it everywhere. As of today, that's shipped and live across the entire agent harness landscape:
| Runtime | Vendor | Enforcer | Status |
|---|---|---|---|
| MCP servers | Open standard | Hundredmen v1.1.0 | β Live on npm |
| Claude Code | Anthropic | adapter-claudecode v0.1.0 | β Live on npm |
Google Antigravity (desktop + agy CLI + SDK) | adapter-antigravity v0.1.0 | β Live on npm | |
| Microsoft Agent Framework | Microsoft | adapter-msaf v0.1.0 | β Live on npm |
| Browser agents | Open standard | browser v0.1.0 | β Live on npm |
The same WARD.md file in your project root is now read and enforced by Anthropic's, Google's, Microsoft's, MCP's, and the browser harness's runtimes β without any platform-specific edits.
@weave_protocol/browser adds runtime IPI (indirect prompt injection) scanning to browser-driving agents. 33 detection patterns cover the documented threat surface: hidden CSS payloads, role-hijack directives, tool-call mimicry, action-injection directives, payment-recipient proximity patterns (Atlan), copyright-DoS markers (Forcepoint), and more.
Pair with the Browser Guard extension for client-side visibility into what your agent sees vs. what you see.
Industry analysis of agent security trends, platform maturity, supply chain risks, and market gaps. Live at: tyox-all.github.io/Weave_Protocol/q3-2026.html
WardMiddleware class, one-line integration, Azure credential heuristic.agy CLI + SDK.weave init / audit / dashboard / doctor. One-command security setup.The suite is now organized into three layers β defense, offensive, and operations. All 17 packages live on npm under the @weave_protocol scope, plus one Python package on PyPI.
The packages that keep your agent within policy: declare it, enforce it across every harness, scan everything that enters, encrypt everything that exits.
| Package | Version | Description |
|---|---|---|
| π‘οΈ @weave_protocol/ward | 0.1.0 | WARD.md β agent security policy standard (parser, validator, runtime checks) |
| π‘οΈ @weave_protocol/adapter-claudecode | 0.1.0 | Claude Code adapter β enforces WARD.md via PreToolUse hooks |
| π‘οΈ @weave_protocol/adapter-antigravity | 0.1.0 | Google Antigravity adapter β enforces WARD.md across desktop, agy CLI, and SDK |
| π‘οΈ @weave_protocol/adapter-msaf | 0.1.0 | Microsoft Agent Framework adapter β middleware-based WARD enforcement |
| π @weave_protocol/browser | 0.1.0 | Browser agent security β runtime IPI scanner (33 patterns) for headless agents |
| π @weave_protocol/hundredmen | 1.1.0 | MCP proxy β intercept, scan, gate tool calls; enforces WARD.md as first gate |
| π‘οΈ @weave_protocol/mund | 0.2.2 | Scanner β secrets, PII, injection, MCP vetting, threat intel |
| ποΈ @weave_protocol/hord | 0.1.6 | Vault β encrypted storage with Yoxallismus dual-tumbler cipher |
| βοΈ @weave_protocol/domere | 1.3.4 | Judge β compliance (PCI-DSS, ISO27001, SOC2, HIPAA, GDPR, CCPA), blockchain anchoring |
| π₯ @weave_protocol/witan | 1.1.0 | Council β multi-agent consensus & governance, autonomous spending caps (Q4 v1.1) |
| π @weave_protocol/tollere | 0.2.2 | Customs β supply chain security (npm, PyPI, Docker, IDE extensions, sandwich detection) |
The red team. We attack what we defend.
| Package | Version | Description |
|---|---|---|
| βοΈ @weave_protocol/adversary | 0.2.1 | Offensive engine β 68 attacks Β· real Playwright browser target Β· real-LLM demo mode Β· WARD-aware attack selection |
| π― @weave_protocol/agentsecbench | 0.1.0 | Standardized benchmark β locked suites (ASB-Browser-v1), tier grading AβF, trophy attacks, WARD delta, paste-ready reports |
The front door, the dashboard, the bridges to other frameworks.
| Package | Version | Description |
|---|---|---|
| πΈοΈ @weave_protocol/cli | 0.1.0 | The weave CLI β init, audit, dashboard, doctor |
| π¦ @weave_protocol/full | 0.1.0 | Bundle β installs all packages in one command |
| π @weave_protocol/api | 1.1.1 | REST API + Operator Dashboard β npx @weave_protocol/api β http://localhost:3000/dashboard |
| π @weave_protocol/langchain | 1.0.2 | LangChain.js security callbacks & tool wrappers (0 audit vulnerabilities via npm overrides) |
| π weave-protocol-llamaindex | 0.1.0 | Python/LlamaIndex security callbacks & tools (on PyPI) |
Each package includes a SKILL.md file following the Claude Agent Skills specification. These teach AI agents how to use Weave Protocol tools effectively.
| Package | Skill Name | Triggers |
|---|---|---|
| πΈοΈ CLI | weave-cli | set up Weave, init project, scaffold security, audit, dashboard, doctor |
| π‘οΈ Ward | ward | WARD.md, agent security policy, guardrails, lock down agent |
| π‘οΈ adapter-claudecode | adapter-claudecode | secure Claude Code, install WARD hooks, block Claude Code actions |
| π‘οΈ adapter-antigravity | adapter-antigravity | secure Antigravity, agy hooks, block GCP credential reads |
| π‘οΈ adapter-msaf | adapter-msaf | secure MSAF agent, WardMiddleware, lock down Copilot SDK, Azure enforcement |
| π browser | browser-security | secure browser agent, IPI scanning, hidden CSS detection, page-context safety |
| π‘οΈ Mund | security-scanning | scan, detect secrets, check injection, vet MCP server, threat intel |
| ποΈ Hord | encrypting-data | encrypt, decrypt, vault, Yoxallismus, protect |
| βοΈ Domere | compliance-auditing | audit, checkpoint, SOC2, HIPAA, PCI-DSS, GDPR, CCPA, blockchain |
| π₯ Witan | consensus-governance | consensus, vote, approve, policy, escalate |
| π Hundredmen | security-inspection | intercept, drift, reputation, approve, block, live feed, enforce WARD |
| π Tollere | supply-chain-security | npm install, docker pull, install extension, typosquat, CVE, sandwich pattern |
| βοΈ Adversary | adversarial-testing | red-team agent, attack, penetration test, find vulnerabilities, IPI test, run attack corpus |
| π― AgentSecBench | security-benchmarking | benchmark agent, security score, tier grade, ASB-Browser, citable security report, compare runs |
| π Langchain | langchain-security | LangChain, callback, secure tool, RAG security, PII redaction |
| π API | weave-api-calling | REST API, HTTP endpoint, curl, fetch |
Installation:
The SKILL.md format is shared across Claude Code and Antigravity, so the same files work for both β only the install path differs.
For Microsoft Agent Framework, skills aren't used β MSAF is code-level. Use the WardMiddleware class from @weave_protocol/adapter-msaf instead.
Add to claude_desktop_config.json:
Drop a WARD.md in your project root. Any (or all) of the adapters will gate every tool call.
π Skill: weave-cli
WARD.md files declare what an agent is allowed to do, version-controlled alongside AGENTS.md and SKILL.md.
| Section | Controls |
|---|---|
| Filesystem | Read/write/execute/delete/list rules with glob patterns |
| Network | Outbound HTTP allowlist with optional method restrictions |
| Capabilities | Tools the agent may invoke (with optional approval gating) |
| Data Boundaries | Egress classifications (PII, PHI, credentials...) and redaction |
| Behavioral Limits | Iterations, runtime, cost, tokens, tool calls |
| Multi-Agent | Trust chain, isolation level, semantic drift threshold |
| Compliance | SOC2 / HIPAA / GDPR / CCPA / ISO27001 / PCI-DSS |
| Verification | Attestation backend (DΕmere), blockchain, frequency |
| Threat Model | In-scope / out-of-scope threats |
| Incident Response | Actions on violation (log / alert / terminate / attest) |
Enforced at runtime by five independent surfaces: Hundredmen (MCP), adapter-claudecode (Claude Code), adapter-antigravity (Antigravity), adapter-msaf (Microsoft Agent Framework), and browser (Browser agents).
π Skill: ward Β· π Spec: WARD.md SPEC β
All four enforcement surfaces share the same WARD.md file. Pick the adapter(s) for your harness:
WARD resolution (all adapters): $WEAVE_WARD_PATH β <cwd>/WARD.md β <cwd>/.weave/WARD.md β harness-specific user-global location.
Real-time security scanning for AI agents. Catches secrets (30+ patterns), PII, prompt injection, dangerous code, malicious MCP server descriptions. Threat intel auto-updates from community feeds.
π Skill: security-scanning
Encrypted storage with the Yoxallismus dual-tumbler cipher. AES-256-GCM, ChaCha20-Poly1305, Argon2id key derivation, secure memory handling.
π Skill: encrypting-data
Enterprise-grade verification, orchestration, compliance, and audit infrastructure. SOC2, HIPAA, PCI-DSS, ISO27001, GDPR, CCPA. Solana and Ethereum blockchain anchoring for immutable audit trails.
Blockchain Anchoring:
6g7raTAHU2h331VKtfVtkS5pmuvR8vMYwjGsZF1CUj2oBeCYVJYfbUu3k2TPGmh9VoGWeJwzm2hg2NdtnvbdBNCj0xAA8b52adD3CEce6269d14C6335a79df451543820π Skill: compliance-auditing
Multi-agent consensus and governance. Unanimous, majority, weighted, and quorum protocols. Rule enforcement, escalation, agent bus.
New in v1.1.0 β autonomous spending caps. Per-window budgets on LLM cost, tokens, tool calls, and per-tool spend limits. Gated by three actions: block, require_approval, notify. Multi-provider LLM pricing built in (Anthropic, OpenAI, Google, local). Interactive TTY prompt or async callback for approval workflows. Safe defaults for non-interactive contexts.
π Skill: consensus-governance
Real-time MCP security proxy. v1.1.0 enforces WARD.md as the first gate in the decision flow, ahead of reputation, drift, and approval checks.
π Skill: security-inspection
Supply chain security for AI-generated code. Catches malicious packages, Docker images, and IDE extensions before they reach node_modules/, your container, or your editor. npm, PyPI, Cargo, Go, Maven, Docker Hub, VS Code Marketplace, Open VSX, JetBrains.
π Skill: supply-chain-security
Where the other packages defend, Adversary attacks. 68 documented and novel attacks across 5 categories: IPI (33), tool-use coercion (15), jailbreak (10), extraction (5), goal corruption (5). Three targets: pattern-mock (CI smoke tests, no API), real-LLM (--real via Anthropic API, ~$0.02/full run), and real-browser (Playwright with four breach signal channels: network, form, DOM, console). WARD-aware attack selection prioritizes probes against capabilities your policy claims to enforce.
Locked scorecard schema v1.0 β consumed unchanged by AgentSecBench.
π Skill: adversarial-testing
The interpretation layer on top of Adversary. Locked, versioned attack suites that produce tier-graded reports. ASB-Browser-v1 (v1.0) is 40 curated attacks: all 33 IPI + 4 critical tool-coercion + 3 highest-impact extraction.
Tier grades AβF, four trophy attacks (Atlan, EchoLeak, Brave/Comet, Forcepoint), category gap analysis, optional WARD delta, plain-English interpretation prose. Reports are paste-ready Markdown β for blog posts, RFP responses, vendor audits, internal reviews.
π Skill: security-benchmarking Β· π Methodology: METHODOLOGY.md β
Live monitoring across all five enforcement surfaces in one view. The dashboard renders WARD.md at the top of a hierarchy diagram, fanning out to your configured enforcers (Hundredmen + the three vendor adapters + browser). Surfaces you're not using appear dimmed, so it's instantly clear what's protecting your agent versus what's available.
Includes a live activity feed (allows / denies / IPI detections / approvals), a WARD policy panel, and 24-hour aggregate stats. Auto-refreshes every 5 seconds. Monochrome design β built for ops rooms, not marketing decks.
π Skill: weave-api-calling
Security integration for LangChain.js applications. Drop-in callbacks, secured tool wrappers, RAG retriever scanning with PII redaction.
π Skill: langchain-security
The diagram shows the loop. Defense surfaces enforce the policy at runtime. The offensive engine attacks the agent. Scorecards feed back into WARD as new evidence β what attacks land, which need new policy domains, what regressed since the last release. The loop is what makes the moat.
Defense-in-depth across the entire AI agent lifecycle, validated continuously by an offensive engine that lives in the same monorepo:
adapter-claudecode for Claude Code (PreToolUse hooks)adapter-antigravity for Google Antigravity (PreToolUse hooks across desktop/CLI/SDK)adapter-msaf for Microsoft Agent Framework (middleware pipeline)browser for browser-driving agents (runtime IPI scanning)| CORS Layer | Weave Package | Function |
|---|---|---|
| Policy | π‘οΈ Ward | Declares allowed/denied actions, behavioral limits, attestation requirements |
| Policy Enforcement (Claude Code) | π‘οΈ adapter-claudecode | Reads WARD, gates Claude Code tool calls via hooks |
| Policy Enforcement (Antigravity) | π‘οΈ adapter-antigravity | Reads WARD, gates Antigravity calls across desktop/CLI/SDK |
| Policy Enforcement (MSAF) | π‘οΈ adapter-msaf | Reads WARD, gates Microsoft Agent Framework calls via middleware |
| Policy Enforcement (Browser) | π browser | Runtime IPI scanning for browser-driving agents |
| Policy Enforcement (MCP) | π Hundredmen | Reads WARD, gates tool calls at the MCP layer |
| Supply Chain | π Tollere | Vets dependencies, images, extensions before install |
| Origin Validation | π‘οΈ Mund | Validates input sources, detects injection |
| Context Integrity | ποΈ Hord | Protects data integrity through encryption |
| Deterministic Enforcement | βοΈ Domere | Ensures consistent policy application |
| Adversarial Validation | βοΈ Adversary + π― AgentSecBench | Continuously tests every layer above |
weave init)@weave_protocol/browser)@weave_protocol/adversary v0.2.1) β 68 documented + novel attacks, real Playwright browser target, real-LLM demo mode@weave_protocol/agentsecbench v0.1.0) β standardized benchmark, tier grades AβF@weave_protocol/witan v1.1.0) β per-window budgets on LLM cost + tokens + tool calls, gated by block / approval / notifyBug reports and feature requests welcome via GitHub Issues.
For security issues, please see SECURITY.md.
For all other inquiries: TYox-all@tutamail.com
See CONTRIBUTING.md for guidelines.
Apache 2.0 β See LICENSE
Built with β€οΈ for the AI agent ecosystem. We attack what we defend.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/mund)<a href="https://allmcps.com/mcp/mund"><img src="https://allmcps.com/api/badge/mund?style=directory" alt="Mund on AllMCPs" /></a>