Grade MCP servers AβF with the open behavioral litmus: reproducible, content-addressed evidence.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
We haven't yet run this listing's install command through our automated sandbox check. This isn't a red flag β we're steadily working through the catalog.
π‘ Paste the JSON block into your client's configuration file under mcpServers, then restart the application.
The open behavioral litmus harness for MCP servers β grade AβF, reproducible.
Point it at an npm ref, a pypi ref, a github/owner/repo ref (cloned, built, and run sandboxed β
Docker required; the grade pins the commit SHA), an https:// MCP endpoint, or a local entry
file. The harness connects
the way an agent would, fingerprints the exact tool surface, runs the four probe categories, and
prints the grade with the findings behind it β plus a deterministic evidence bundle on disk.
It runs the target's code (egress is Docker-sandboxed; without Docker, C-02 is skipped and the
grade caps at B), takes ~20β60s, and exits non-zero on D/F so it scripts anywhere. To dispute any
published grade, re-run this same command against the same server β open and deterministic means a
re-run reproduces the grade, or refutes it.
Looking up a grade someone already published takes under a second and runs nothing: the
polygraph.so index, or check_server from the MCP tools below.

For grade lookups, point any MCP client at polygraph's hosted endpoint, no install:
or the raw config:
This serves the lookup tools only (check_server, list_servers, request_grade); grading a
server yourself (run_litmus, run_skill_litmus) executes its code, so it needs the local
stdio install below.
The package also ships a stdio MCP server (polygraphso-litmus-mcp) with the full toolset, for
any MCP-capable client:
check_server β read a server's published grade in under a second (no execution); the
pre-flight check before recommending or installing a server.list_servers β servers with a published grade, A first; paged (default 25 per call, with grade/limit/offset filters and a full-corpus summary).request_grade β record a grade request with polygraph.so ($1 one-time fee; graded within 48h of payment β the response carries the payment link).run_litmus β grade a server now: the full harness, grade + evidence returned to the agent.run_skill_litmus β grade a Claude Code / Agent Skill (static scan, A/B/D/F).verify_attestation β read the onchain proof behind a published grade (EAS on Base).In Claude Code, the plugin wires the server plus two commands in one step:
then /polygraph:grade <server> and /polygraph:check <server>. Cursor and manual JSON setups
are on polygraph.so; full tool docs in
packages/litmus/README.md.
Fail a build when an MCP server or an Agent Skill it ships grades D/F under the open
behavioral litmus. For servers it is hybrid β a fast lookup of the published grade, then the harness
when ungraded; for skills it is a fast static scan. Un-gradeable targets warn unless strict.
It's on the GitHub Marketplace as
polygraphso/litmus@v1. For a security gate, pin to a commit SHA rather than the mutable @v1 tag:
Inputs: servers Β· skills Β· discover (default false) Β· min-grade Β· strict Β· working-directory Β· version Β· bearer. Outputs: result Β· failed Β· report.
Security. Grading a server runs its code (egress is Docker-sandboxed, but it still executes).
Trigger on pull_request, never pull_request_target. Keep discover off on public repos and name
targets explicitly β auto-discovered config is pull-request-controllable. bearer is sent as an
Authorization header to the target, so pass it only for an explicitly trusted, pinned remote β never
with discovery or on untrusted PRs, and keep it scoped and short-lived.
Not on GitHub? The gate is a plain command β npx @polygraphso/litmus@0.20.0 ci (pin the version) β
so it runs in any CI or as a pre-commit hook. A grade is a measurement, not a guarantee: re-run the
open harness to reproduce any result.
This is the source for @polygraphso/litmus,
the open behavioral litmus harness for MCP servers from polygraph.so.
The harness connects to an MCP server the way an agent would, fingerprints its exact tool surface, and runs four probe categories β C-01 tool-output injection (static, dynamic, and second-order β one tool's output weaponized as another's input), C-02 permission/egress (in a hardened default-deny Docker sandbox, matched host and port), C-03 sensitive-data handling (planted canaries), C-04 adversarial-input handling (malformed/oversized and jailbreak inputs) β then grades the server AβF. A passing grade is a measurement, not a guarantee; the methodology and its disclosed limits are at polygraph.so (the open source here is the ground truth).
Alongside the grade, an npm target's dependency tree is checked against the
osv.dev vulnerability database and any vulnerable dependencies are reported as
dependency advisories. This is a separate, point-in-time signal β it is advisory only: it
never affects the AβF grade and is not part of the reproducible evidence (vulnerability data changes
over time, so folding it into the grade would break re-run reproducibility). It applies to npm
targets only; other target kinds report it as skipped. Resolution runs
npm install --package-lock-only --ignore-scripts, which resolves the tree without downloading
tarballs or running any package code. Opt out with --no-deps-audit (or LITMUS_DEPS_AUDIT=0).
The same package also grades Claude Code / Agent Skills (a SKILL.md + bundle) under a
separate static litmus (litmus-skill-v3): a deterministic byte-scan β S-01 prompt
injection, S-03 data-exfiltration instructions, S-04 dangerous commands in the SKILL.md
body or bundled scripts (incl. base64-obfuscated curl | bash) β graded A/B/D/F and anchored
by a whole-directory content hash, plus a separate
advisory quality signal. It is static (no execution): an A is static-clean, not behavioral
proof. See packages/litmus/README.md.
The hosted, operator-run grading service is not in this repo β it lives in a separate private repo and consumes this package from npm like any other client.
This is a pnpm monorepo. Only @polygraphso/litmus is published; the
@polygraph/* packages are private building blocks that tsup bundles into it.
Factual signals from GitHub, npm, and our automated checks β not a rating.
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/litmus)<a href="https://allmcps.com/mcp/litmus"><img src="https://allmcps.com/api/badge/litmus?style=directory" alt="Litmus on AllMCPs" /></a>