The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Agentgate listing page.
English · 中文
A control plane for the tools agents run. It inventories what is in use, records the evidence behind every claim, states what a company refuses, and enforces that decision in CI and at runtime.
The whole project follows one rule:
cleanis emitted only when every check ran. Anything that could not be measured isunmeasured, and an artefact with an unmeasured part isincomplete— neverclean.
That rule is there because the usual failure of a security scanner is a green build for work nobody did. Here a check that crashes makes the result incomplete, so it cannot happen quietly.
| part | what it does | package |
|---|---|---|
| inventory | enumerate the registry, resolve packages, fetch repositories | packages/collect |
| evidence | join it into one record per server, with the bytes behind every claim | packages/collect |
| policy | scan configs, hooks, manifests and source for what a company would refuse | packages/guard |
| verification | check a claim against something outside the claim | packages/verify |
Server setup is in docs/operations/deployment-runbook.md. The first deployment, on 2026-09-16, is written up in docs/verification.md together with what was checked and what still is not.
docs/capabilities.md lists what this project can and cannot claim, one line each, every line carrying a command you can run. It exists because a consultant once wrote our capabilities down for us and included four we do not have.
The service runs at https://xn--5kvo87g.com/: landing page, pricing, the evidence index (rebuilt daily) and the API on the same host.
https://ciceroyang.github.io/agentgate/ is the landing page on GitHub Pages. The index is a single browsable page at https://ciceroyang.github.io/agentgate/evidence.html, rebuilt daily from the live registry — records are embedded, filtering happens locally, and there is nothing to sign up for. Pricing and a ten-minute walkthrough.
Upgrading to 0.6.0: read the upgrade guide before replacing an existing admission gate or watch installation. The 0.5.0 package can incorrectly pass a required-evidence gate when evidence is absent; do not use it for a new gate. Confirm the version you installed, and see the release record for public-package verification. A source checkout and the hosted service may run different versions.
Node 20 or newer, no dependencies. A clone already carries a sample index, so the service
answers immediately; refresh replaces it with a current one.
The package is on npm as @zhiliangtech/agentgate. Push a v* tag and CI publishes it with
provenance; publish-checklist.md has the setup and the
record of what was verified.
An unversioned npx command follows the latest dist-tag, not this checkout. Pin the accepted
version; editing a local version number does not change the public package.
With no policy file, check uses a built-in default that refuses nothing extra, and serve
answers from the snapshot the package shipped with. refresh writes to ./data next to you,
never into the installed package.
Docker works too, and runs the same command:
Run node bin/agentgate.mjs serve and open /inventory.html at the address it prints. Paste a
list of tool names or pick a text/JSON file, resolve ambiguous matches, fill in the version you
actually use, and download a standalone HTML report. The comparison happens in browser memory
against the embedded index snapshot: the list is not uploaded or stored, your machine is not
scanned, and no tool is executed. When a record carries a coverage block, the report also lists
which scanners ran and which did not, and why; a record whose own coverage block says a required
scanner did not finish will not be shown as matched, however complete the rest of its evidence looks.
The same thing without a browser:
An input is one name per line, a JSON array, or { "tools": [...] }. Each object may carry
name, server, package, registry and version — nothing else. Full client
configurations and credentials are rejected on purpose. The
inventory guide has the details.
Unmatched, ambiguous, missing-version, mismatched-version and incomplete-evidence items stay in the report. A version match is not proof of what is installed. The committed sample is historical and cannot produce a confirmed match; neither can old evidence without an exact content binding. Even a confirmed match is not a safety certification and not a new scan. Look at the scopes, the findings, the snapshot date and the gaps before you rely on it.
Exit code 0 means a report was produced, not that every tool passed. Malformed input or
unreadable data exits 2, and --out will not overwrite an existing file. When you want CI to
refuse something, use check, not inventory.
Nobody has this list by hand. discover reads the MCP configuration files already on the
machine and produces input that inventory --input accepts. It prints package coordinates as
lines when all exported identities are known; if any entry is alias-only, it uses JSON so an
alias such as tool@1.2.3 cannot be mistaken for a verified package and version:
It never prints an env value, a header or an argument, and a remote address is cut down to its
host, because paths and query strings carry tokens. It reads Codex's .codex/config.toml MCP
tables without starting the configured servers. Explicitly disabled entries remain visible in
--format json but are omitted from text and inventory exports. Unsupported MCP TOML shapes,
malformed files and unreadable files are listed with a reason and make the command exit 2.
Package names and versions are taken only from recognizable declared runner arguments; a custom
command or remote host is not treated as a verified package or runtime version.
One scan per directory, one verdict for the set. Any incomplete directory makes the audit incomplete, and a directory that does not exist counts as unmeasured rather than skipped.
Every run appends one line to a chained archive (prev is the previous line's hash) and stores
what it saw under snapshots/<sha256>.json. --verify recomputes the chain and every retained
snapshot, and exits 1 if anything does not match. Nothing is sent anywhere unless --webhook
names an address, and the archive is written before the push, so a chat service being down cannot
lose a capture.
For each AI-CAIQ item the mapping says what we can provide, where our coverage stops, and whether the answer is ours, the customer's, or an independent assessor's. It describes evidence. It is not a compliance conclusion and it does not reproduce the official text. All 58 items of the four domains a reviewer asks a vendor about are classified: 13 answers are ours, 41 are the customer's and 4 need an independent assessor.
The mapping says what we can provide. pack produces the thing itself: one directory a vendor
hands to the person reviewing them, where every answer we claim points at evidence in the same
directory and everything we could not measure is counted at the top.
It writes pack.json (machine readable), pack.html (for the reviewer), answers.aicaiq.md (all
58 items, each classified), manifest.txt (one sha256 per file) and manifest.sha256 (the seal on
the manifest). An answer whose evidence is missing reads unmeasured and the command exits 2, not
0. Example built from the live index:
docs/samples/evidence-pack-example — verifiable with
pack --verify. Contract: docs/spec/evidence-pack-v1.md.
Anything that speaks MCP can ask the index directly. Add this to claude_desktop_config.json, a
repo's .mcp.json, or whatever your client reads:
Four read-only tools: lookup_server (one record, with its coverage block), inventory_tools
(match the tools you actually use), coverage_report (how much of the index was measured) and
check_project (scan a local directory). It reads the local index, never writes, never uploads,
and never runs a scanned tool. An incomplete record is reported as incomplete, and a record that
the index does not have is reported as missing rather than safe. Details:
docs/spec/mcp-server-v1.md.
A policy states what a company refuses. It is data rather than code, and it has a spec: docs/spec/policy-v1.md.
The policy above requires indexed evidence. Supply a real index matching the local npm package's exact name and version: missing evidence exits 2; an explicit missing, malformed or sample index exits 3. A local source scan does not substitute for the required package evidence.
With no policy file and no --policy, the check still runs. It reports what the checks found and
says it used the built-in default, which refuses nothing extra; inventing obligations on your
behalf would make the result mean less, not more. A policy you name explicitly and that cannot
be read is an error, because that is a typo.
The same evaluation can go to a person instead of a terminal:
One static, printable file with no script in it. Anything that could not be measured gets its own section above the findings: a report that buries what it did not check reads as more complete than it is. This file is what the free checkup delivers.
There are three outcomes, and incomplete outranks findings. If a check failed to run, or an
evidence block the policy requires is unmeasured, the exit code is 2 however clean the
findings look. No threshold turns a partial answer into a pass.
| exit | meaning |
|---|---|
| 0 | clean |
| 1 | findings |
| 2 | incomplete |
A pull request that adds something the policy refuses will not merge, and the reason is posted on the pull request rather than left in a log nobody opens.
See examples/github-actions/policy.yml. The action runs the check, writes SARIF for code scanning, comments the report on the pull request, and exits with the check's own code — so an incomplete scan still fails the build at 2.
The same policy can apply to what has already shipped, if you put a gateway in front of the server instead of pointing your client at it:
A call the policy refuses is answered locally with a reason and never reaches the server. A forbidden tool is removed from the advertised list, so a client cannot ask for it at all. Every decision, allowed or refused, is appended to the log.
The index is kept, so two builds can be compared. The interesting column is the last one: changes that a release would have explained and did not.
A new finding on an unchanged version usually means a package was replaced without a release, a repository was edited in place, or the scan has started seeing something. That record cannot be back-filled. It only exists if someone was looking at the time.
And the scanner on a local project:
The tests are written by the same people who wrote the code. docs/verification.md records the checks that are not: a real MCP server through the gateway, and the list of what is still unverified.
To see whether the index's high and critical findings still match recorded human reviews, run
node scripts/review-criticals.mjs. A review has to bind the finding and its evidence to an
exact package version and to complete scanned-content provenance, including the SHA-256 digest
and the scope. A missing or changed binding needs another human review, and legacy approvals are
not upgraded automatically. --accept records a review that has already happened and refuses
incomplete provenance; it neither performs the review nor certifies third-party code.
--apply.This is an early open-source core. It covers collection, an evidence index, scanning, policy
checks in CI, a runtime gateway for MCP servers over stdio, historical diffs and a read-only
service. Deployment scripts and a runbook are in the tree, and the first deployment with its
checks is written up in docs/verification.md. That write-up says nothing
about the current health of the hosted service. The container path is not asserted but built: CI runs docker compose up --build and then a health check against the running container.
What the version identifiers promise, and which versions are supported, is written down in docs/spec/compatibility.md; SECURITY.md says how to report a vulnerability and what to expect. Neither is a substitute for gate 3 and gate 4 above — the enterprise surface is still missing and nobody outside this repository depends on it yet.
The enterprise features described in the pricing proposal — SSO/SAML, RBAC, multi-tenancy and signed audit export — are not implemented. The Team and Enterprise prices are unvalidated hypotheses; the free pilot is how we test whether anyone wants this. See the pilot scope and the licence.
One invariant is in the test suite: a crashed check can never produce clean. Run npm test
for the current numbers; this page does not repeat a test count.
AGPL-3.0-only. If you want to offer a modified agentgate as a closed service without publishing your changes — the case the AGPL does not permit — a commercial licence is available. See docs/product/licensing.md.