The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the 404.directory listing page.
Risk preflight for AI Agent actions.
Before an AI Agent uses a third-party tool, 404.directory checks the available evidence and returns an allow, review, or block decision.
For developers connecting Agents to external tools who want an explicit risk check before execution. This is an experimental preflight service; its decision depends on the available evidence.
Connect your Agent using Codex, Cursor, Claude Code, or the MCP SDK.
The tool workflow is:
Use evaluate_tool_risk for a registered third-party tool action. The separate evaluate_prediction_market workflow checks settlement wording, timing, public order-book liquidity, caller-observed eligibility, and execution mode. It does not predict the winner or place an order.
Official documentation search, deployment verification, and the curated read-only MCP gateway are supporting capabilities. See Product layers for the complete API map.
Trying it for the first time? Join the first-10 activation pilot and report your client, task category, and failure stage.
Agent-readable installation instructions: llms-install.md
Install the Agent Skill in Codex, Claude Code, Cursor, Cline, or another Agent Skills client:
The repository also conforms to Agent Plugins 1.0: compatible clients discover
the Agent Skill from skills/ and an identity-preserving bridge to the hosted
Streamable HTTP server from the root mcp.json. The bridge creates one random
ID in the client-managed PLUGIN_DATA directory. The raw ID stays local; the
service stores only an HMAC digest after a successful tool call.
Claude Code and Cowork use the native manifest in .claude-plugin/. It loads
the same Skill and identity-preserving bridge with Claude's persistent plugin
data directory, so updates keep the installation identity stable.
Install it directly in Claude Code while the official directory submission is under review:
| Layer | Purpose | Surface |
|---|---|---|
| First-party execution | Run first-party tools in this process | GET /tools, POST /understand, POST /verify/web, MCP tools |
| Curated remote execution | Search and call approved read-only remote MCP tools | MCP search_official_docs / inspect_tool_server / invoke_registered_tool |
| Ecosystem catalog + trust | Register / verify / trust / search third-party tools | /v1/*, MCP search_tools / get_tool / compare_tools / get_trust_score |
| Contextual risk preflight | Decide whether an Agent should proceed now | MCP evaluate_tool_risk / report_tool_outcome, REST /v1/evaluations/* |
| Prediction-market preflight | Check settlement and execution risk before action | MCP evaluate_prediction_market / report_prediction_market_outcome, REST /v1/prediction-markets/evaluations/* |
The current product is intentionally narrow: preflight one prediction-market decision or one registered third-party tool action, then capture a bounded outcome. Future identity, reputation, guarantee, and insurance layers remain hypotheses until real external Agent usage validates them.
| Tool | Endpoint | When to use |
|---|---|---|
understand_webpage | POST /understand | Understand an ordinary webpage (entities, state, actions) with no Agent-native API |
verify_web | POST /verify/web | Independently verify a public site after a deploy/update claim |
/v1)Requires a catalog backend (DATABASE_URL Postgres, or in-memory fallback when
CATALOG_MEMORY_FALLBACK=true).
Catalog keyword search uses catalog-lexical-v2 in both memory and PostgreSQL.
Try q=official%20documentation or q=OpenAI%20docs; words need not be adjacent.
All meaningful terms must match across the name, description, capabilities,
category or provider. Exact names rank first, then lexical relevance, existing
trust evidence and usage. Capability/protocol/category/trust filters remain
mandatory, and public search still excludes quarantined and suspended tools.
No matches returns count: 0, search.result_status: "no_matches", and a
recovery step pointing to MCP list_capabilities / REST /v1/capabilities.
An empty MCP search is a valid response, but is recorded as no_matches rather
than a successful Agent activation. It does not mean the task is impossible.
Search neither executes tools nor proves they are safe; preflight the exact
chosen slug. See search semantics and acceptance results.
Trust Profile dimensions (v1 algorithm, extensible factors JSON):
overall_scoreContextual preflight is available through POST /v1/evaluations; public
receipts are readable at GET /v1/evaluations/:id. One bounded outcome can be
attached through POST /v1/evaluations/:id/outcome using the one-time token
returned at evaluation time. Only the token hash is stored, and self-reported
outcomes never directly increase Trust. The older generic POST /v1/receipts
remains disabled because unbound anonymous submissions would poison Trust.
Copy-ready Agent trigger policy and examples:
docs/AGENT_RISK_PREFLIGHT.md
Privacy-safe product validation is public at
GET /v1/metrics/risk-evaluations: evaluation volume, decision distribution,
outcome-report rate, and behavior-change rate, without prompts or raw identity.
The prediction-market workflow is documented at
docs/PREDICTION_MARKET_PREFLIGHT.md.
Its privacy-safe aggregate metrics are available at
GET /v1/metrics/prediction-market-evaluations.
When the catalog is enabled, MCP also exposes:
evaluate_prediction_marketreport_prediction_market_outcomeevaluate_tool_riskreport_tool_outcomesearch_toolsget_toolcompare_toolsget_trust_scorerecommend_toolslist_capabilitiesget_capability_graphsearch_official_docsinspect_tool_serverinvoke_registered_toolalongside the existing executable tools.
evaluate_prediction_market is the primary first-use path: one call evaluates
an exact Polymarket market for settlement, liquidity, eligibility, and execution
risk without predicting or trading. evaluate_tool_risk is the second wedge,
used before an Agent installs or invokes an unfamiliar third-party tool.
search_official_docs remains a supporting path and returns bounded first-party
citations instead of raw provider indexes. Arbitrary URLs, authenticated
servers, non-active entries, unverified providers, and destructive tools are
rejected. Remote results are bounded and explicitly marked as untrusted data.
Clients that expose MCP Prompts also receive four task-oriented starting points:
preflight-prediction-market — turns an exact market and contemplated action
into an evaluate_prediction_market call;evaluate-agent-tool — finds a catalog candidate, calls the contextual risk
preflight, and requires an allow, review, or block result;research-official-docs — turns a real technical question into a
search_official_docs call;verify-public-deployment — turns a concrete public deployment claim into a
verify_web call.Rendering or opening a prompt never counts toward the 1,000-Agent target. Each
template explicitly requires a non-error tool result that materially answers
the user's task. The server records only aggregate prompts/list and
prompts/get activation stages, never prompt arguments or task text.
MCP prompt arguments are strings. evaluate-agent-tool requires an explicit
permissions argument, for example "public_network,credentials" or the JSON
array string '["public_network","credentials"]'. Use "[]" only when the
action requires no permissions. Missing, malformed, and unknown permissions
are rejected rather than silently treated as safe. This string encoding is
for prompts/get only; evaluate_tool_risk still accepts a JSON array.
For direct trading-Agent integration, see
docs/AGENT_INTEGRATION_QUICKSTART.md.
The privacy-safe first-10 cohort process is documented in
docs/FIRST_10_AGENT_PILOT.md.
Agents can explore shared-capability edges and get related-tool recommendations:
Similarity is Jaccard over capability sets, with small boosts for matching
protocol/category (cap_v1). This is the seed of the long-term Capability Graph.
Default: http://127.0.0.1:4040
With Postgres:
On boot, first-party tools are seeded into the catalog (SEED_FIRST_PARTY_TOOLS=true)
so GET /v1/tools/search?capability=web-verification returns verify_web.
The six operator-reviewed public MCP servers are also seeded as pending entries
when SEED_CURATED_MCP_SERVERS=true. The verification worker performs live MCP
admission before they become discoverable or executable.
VERIFICATION_WORKER_MODE=inline (loop inside HTTP process)Ownership Score ladder: first-party 1.0 → dns_txt 0.95 → github_bio 0.9 →
generic verified 0.8 → unverified 0.35.
The service inventory and the registered ecosystem catalog are distinct:
| Surface | Meaning |
|---|---|
MCP tools/list, GET /tools | The same enabled, callable 404 service tools (16 with the default native tools, catalog and gateway enabled) |
GET /tools/:name | The actual MCP argument schema, safety annotations, and explicit MCP / REST invocation routes |
GET /v1/tools/search | Registered target records, including seeded first-party and third-party tools; a match is not permission to execute |
GET /v1/capabilities | Capability labels for ecosystem records, not a list of callable 404 functions |
The homepage, installation guides, docs, server card, and discovery metadata
derive the enabled tool inventory from the real MCP registration at startup.
The three gateway tools (search_official_docs, inspect_tool_server,
invoke_registered_tool) are MCP-only: their metadata has invocation.rest: null.
For other tools, follow the declared REST path and parameter mapping instead
of assuming that /tools/:name executes a tool. HTTP contracts remain in
/openapi.json; MCP metadata schemas follow MCP's JSON Schema dialect, not
OpenAPI 3's schema dialect. Restart after changing registration/configuration.
See discovery consistency audit for validation, compatibility notes, and the local-only delivery boundary.
Homepage (GET /) is intentionally minimal: brand, tagline, tool names, and
links to Tools / MCP / OpenAPI / Docs / Health.
verify_web returns compact booleans in checks plus a structured evidence
object containing requested/final URLs, HTTP status comparison, expected-text
matching, TLS validation, the complete redirect chain, timestamp, and explicit
Claim → Evidence paths.
Tool execution is currently public and free. Rate limits use Vercel's trusted client-IP header (or the socket IP locally).
Point MCP clients at https://404.directory/mcp (or local http://127.0.0.1:4040/mcp).
The hosted endpoint can also be used directly as an OpenAI Responses API
remote MCP tool. A copy-ready payload with a privacy-safe installation token is
available in llms-install.md; see the
official OpenAI MCP guide.
To become eligible for verified counting, send a stable random, non-personal
identifier in X-404-Agent-ID. The server persists only an HMAC digest, never
the raw ID, prompts, arguments, or results. A successful call is necessary but
does not count by itself: independent-operator evidence must be admitted
separately. X-404-Source is an optional lowercase attribution label. Verified
public progress is available at GET /v1/metrics/verified-agents; unverified
installation diagnostics remain at GET /v1/metrics/agents. Complete client examples are at
https://404.directory/connect.
OpenAI Responses does not document arbitrary remote MCP request headers. Its
example instead uses the supported MCP authorization field with a generated
agent:<uuid>@<source> installation token. 404.directory accepts only that
strict non-personal shape as an Agent identity; unrelated OAuth bearer tokens
remain anonymous and are never treated as Agent IDs.
The privacy-safe activation funnel is available at
GET /v1/metrics/activation. It reports observed Connect views and installer
clicks plus de-duplicated external Agents that completed MCP initialize,
tools/list, prompts/list, prompts/get, attempted a tool call, failed a tool
call, or completed a successful tool execution. The per-source output separates
call rate, call success rate, prompt-to-success rate, and end-to-end activation
rate. Every funnel stage is diagnostic only. Successful execution is necessary
but does not count toward the 1,000-Agent target without a separate active
independent-operator evidence admission. Prompt names and arguments are not
stored in activation events. No raw Agent IDs, IPs, prompts, arguments, or
results are stored in the funnel.
GET /v1/metrics/verified-agents reports privacy-safe 7/30-day retention for
verified Agents. An Agent becomes eligible only after a complete observation
window and is retained only after another success on a later UTC day.
GET /v1/metrics/agents retains unverified installation diagnostics by safe
client label. GET /v1/metrics/reliability?days=30 aggregates external
execution evidence by tool, registered provider, client, and attribution source,
including sample size, success rate, P50/P95 latency, result count, and a finite
error taxonomy. Anonymous external executions can inform reliability but never
count toward the 1,000 verified-Agent target.
The official MCP Registry entry also declares X-404-Agent-ID as an install
input and defaults X-404-Source to official-registry, so compatible clients
can preserve a privacy-safe identity instead of silently creating anonymous
usage. The service remains usable without either header.
The dynamic install page also generates a one-click VS Code / GitHub Copilot Agent link with a unique non-personal ID already embedded:
https://404.directory/connect?source=github
Registry clients can display the same-domain, script-free service icon at
https://404.directory/icon.svg.
For clients or directories that accept only a stdio launch command, use the identity-preserving hosted bridge. It creates one random Agent ID per MCP client in the user's normal application-data directory and reuses it across restarts:
No account or API key is required. The bridge is dependency-free and forwards
only MCP JSON-RPC traffic to https://404.directory/mcp.
Codex supports MCP HTTP headers in ~/.codex/config.toml:
Do not add only the bare MCP URL if you want the Agent installation to retain a privacy-safe identity. Use the generated Codex configuration at https://404.directory/connect?source=github.
Tools are registered automatically from the Tool Registry — adding a tool does not require hand-writing separate MCP adapters.
ToolDefinition in src/tools/definitions/src/tools/create-registry.tsREST, OpenAPI, /tools/:name, and MCP pick it up from the registry. Keep
/tools compact so discovery cost does not grow with every schema.
Production runs on Google Cloud Run. Dockerfile uses Node slim and
installs only Chromium's headless shell so the retained Artifact Registry image
stays below the 0.5 GiB free storage allowance.
Catalog persistence is required in production. Without DATABASE_URL and
with CATALOG_MEMORY_FALLBACK=true (the local default), Registry / Trust /
telemetry evaporate on every cold start. Also: request-based Cloud Run does
not reliably run in-process setInterval workers — use an external worker.
Registry write APIs (POST /v1/tools, ownership challenge/verify, manual
verify) require Authorization: Bearer <REGISTRY_ADMIN_TOKEN|provider_api_key>.
New providers receive a one-time provider_api_key. Search defaults to
status=active only — pending tools stay quarantined.
Apply cloudrun.cleanup-policy.json to the source-deploy Artifact Registry
repository so superseded, untagged images do not accumulate storage charges.
Local Docker remains available:
Production hardening notes:
BROWSER_EGRESS_ALLOWED_PORTS narrow (default 80,443)RATE_LIMIT_* and verify/browser timeouts for your trafficTOOL_RATE_LIMIT_MAX than discoveryMCP_GATEWAY_MAX_RESULT_BYTES and MCP_GATEWAY_TIMEOUT_MSCATALOG_MEMORY_FALLBACK=false whenever DATABASE_URL is configuredVERIFICATION_WORKER_MODE=external on serverlessverify_web pins each connection to the exact public IP that passed DNS
validation, then re-resolves and re-validates every redirect hop. TLS SNI is
sent only for hostnames; IP-literal URLs (e.g. https://1.1.1.1) omit SNI and
validate the certificate against the IP insteadverify_web caps response bodies; 198.18.0.0/15 and other reserved ranges
are rejected in every environmentunderstand_webpage re-resolves and re-validates every browser request, but
Chromium request and routes it through a loopback-only forward proxy. That
proxy resolves the destination, rejects private/reserved addresses and
disallowed ports, then connects to the exact IP that passed validation.
Chromium's implicit loopback bypass and QUIC are disabled, while non-proxied
WebRTC UDP is blocked. Browser contexts also set serviceWorkers: "block".
A provider/network egress firewall is still recommended as an independent
second layer when the hosting platform supports oneServer-Timing, no-sniff/frame/referrer/
permissions/CSP headers; REST Tool results use Cache-Control: no-store,
while MCP streaming uses the SDK's no-cache, no-transform policy