Skip to main content
AllMCPs
BrowseBestCategoriesStackCompareToolsGuidesBlog
Log in Submit MCP

Stay in the loop

Get new MCP servers and top picks in your inbox.

AllMCPs

The open directory for discovering and installing Model Context Protocol servers.

Launched onTiny Startupstinystartups.com
Explore
  • Browse servers
  • Best MCP servers
  • Categories
  • MCP clients
  • Agent prompts
  • Stack Builder
  • Compare servers
  • Random discovery New
  • Submit a server
  • Pricing & Boost Boost
Learn
  • Guides hub
  • What is MCP?
  • Install guide
  • Build an MCP server
  • Deploy an MCP server
  • Security guide
  • Troubleshooting
  • MCP for SEO & AEO
  • Protocol versioning
  • Blog & updates
Tools
  • All developer tools
  • Config generator
  • Config validator
  • Config auditor
  • MCP playground
  • Token calculator
  • OpenAPI β†’ MCP
  • Badge generator
For agents
  • REST API docs
  • Trust & traffic Live
  • Remote MCP server SSE β†— (opens in a new tab)
  • llms.txt β†— (opens in a new tab)
  • Catalog JSON β†— (opens in a new tab)
Company
  • About
  • Advertise Sponsor
  • Contact
  • X (@AllMCPs) β†— (opens in a new tab)
  • GitHub β†— (opens in a new tab)
  • Terms
  • Privacy
AllMCPs VerifiedAllMCPs VerifiedFeatured on Nick LaunchesFeatured on Nick LaunchesLaunch Llama NewsletterLaunch Llama NewsletterVerified DR - allmcps.comVerified DR - allmcps.comFeatured on SaaSGrowFeatured on SaaSGrowFeatured on Twelve ToolsFeatured on Twelve ToolsFeatured on Saaspa.geFeatured on Saaspa.geFeatured on Findly.toolsFeatured on Findly.toolsFeatured on Startup FameFeatured on Startup FameFeatured on LaunchKiwiFeatured on LaunchKiwiFeatured on ScrollLaunchFeatured on ScrollLaunchFeatured on DailyPingsFeatured on DailyPingsFazier badgeFazier badgeFeatured on NewTool.siteFeatured on NewTool.siteFeatured on saasfame.comFeatured on saasfame.comDR Checker - Domain RatingDR Checker - Domain RatingListed on Turbo0Listed on Turbo0Launched on LaunchBoard - Product Launch PlatformLaunched on LaunchBoard - Product Launch PlatformList on SimilarlabsList on Similarlabshttps://codetrendy.comhttps://codetrendy.comListed on DevTool.ioFeatured on BuildlistFeatured on BuildlistLaunched on Tiny StartupsFeatured on ShowMeBestAIFeatured on ShowMeBestAIFind us on LaunchZoneFind us on LaunchZoneAllMCPs VerifiedAllMCPs VerifiedFeatured on Nick LaunchesFeatured on Nick LaunchesLaunch Llama NewsletterLaunch Llama NewsletterVerified DR - allmcps.comVerified DR - allmcps.comFeatured on SaaSGrowFeatured on SaaSGrowFeatured on Twelve ToolsFeatured on Twelve ToolsFeatured on Saaspa.geFeatured on Saaspa.geFeatured on Findly.toolsFeatured on Findly.toolsFeatured on Startup FameFeatured on Startup FameFeatured on LaunchKiwiFeatured on LaunchKiwiFeatured on ScrollLaunchFeatured on ScrollLaunchFeatured on DailyPingsFeatured on DailyPingsFazier badgeFazier badgeFeatured on NewTool.siteFeatured on NewTool.siteFeatured on saasfame.comFeatured on saasfame.comDR Checker - Domain RatingDR Checker - Domain RatingListed on Turbo0Listed on Turbo0Launched on LaunchBoard - Product Launch PlatformLaunched on LaunchBoard - Product Launch PlatformList on SimilarlabsList on Similarlabshttps://codetrendy.comhttps://codetrendy.comListed on DevTool.ioFeatured on BuildlistFeatured on BuildlistLaunched on Tiny StartupsFeatured on ShowMeBestAIFeatured on ShowMeBestAIFind us on LaunchZoneFind us on LaunchZone
Β© 2026 Jackalope Digital LLC. All rights reserved.
  1. Home
  2. Developer Tools
  3. Agent Reliability
A
Health: Not checked yetWe have not completed a health check for this listing yet.No health check has run yet.

Agent Reliability

User RatingsBe the first to rate and review this MCP server! Enrichment pendingWe haven’t run our AI enrichment pass on this listing yet, so the overview, use cases, and FAQ below may be sparse or missing. We work through the catalog over time β€” check back soon.
View Repository

Testing, benchmarking and auditing autonomous AI agents β€” methods, harnesses, evidence

Quick Install

Automated & IDE Setup

Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β€” or use 1-click editor setup below.

Add to CursorAdd to VS Code
Manual Client & Custom JSON ConfigExpand JSON β–Ύ

Client Config & Setup

Choose your client or environment
Target File:~/Library/Application Support/Claude/claude_desktop_config.json
claude_desktop_config.json
{
  "mcpServers": {
    "agent-reliability": {
      "command": "npx",
      "args": [
        "-y",
        "agent-reliability"
      ]
    }
  }
}

πŸ’‘ Paste the JSON block into your client's configuration file under mcpServers, then restart the application.

Install Directory Badge Claim listing AlternativesπŸ’» More in Developer Tools

Documentation Overview

Agent Reliability β€” MCP server

Testing, benchmarking and auditing autonomous AI agents β€” methods, harnesses, evidence

A remote MCP server over a curated knowledge graph. Every claim it returns is bound to a registered source: the tools hand back claims with their citations and a confidence value, so an agent can show its work instead of asserting.

Nothing to install. It is a hosted streamable-HTTP endpoint:

Code
https://agentreliability.dev/mcp

Add it to a client

Claude Code

Terminal
claude mcp add --transport http agent-reliability https://agentreliability.dev/mcp

Claude Desktop / any client reading mcpServers

config.json
{
  "mcpServers": {
    "agent-reliability": {
      "type": "streamable-http",
      "url": "https://agentreliability.dev/mcp"
    }
  }
}

No API key, no account, no auth. Read-only.

Check it answers, without any client at all:

Terminal
curl -s https://agentreliability.dev/mcp \
  -H 'Content-Type: application/json' \
  -H 'Accept: application/json, text/event-stream' \
  -H 'mcp-protocol-version: 2025-06-18' \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'

Tools

Eight, each with an outputSchema, each returning structuredContent.

toolargumentswhat it does
get_overviewβ€”Corpus overview: what this instance knows, counts by type, published tags, freshness. Start here when you land and do not yet know whether this corpus can answer your question.
searchquery, limit?Full-text search over the knowledge graph. Accent- and apostrophe-insensitive, so query in the user's own words; every hit carries its relevance score and the fields it matched.
answerquestionAnswer a question from the corpus. Returns the matched object's claims with sources and confidence β€” never an unsourced answer.
get_entityidFetch one knowledge object by id, with its claims and the sources each claim cites.
get_topictagList the knowledge objects carrying a tag (topics are content-backed tags).
get_relatedidGraph neighbours of an object: outgoing and incoming relations, each with its relation type.
get_sourcesobject_id?The whole source registry, or just the sources cited by one object. Use it to judge the corpus before trusting it.
get_latestlimit?Most recently verified knowledge objects β€” a freshness signal.

The intended path is get_overview β†’ search or answer β†’ get_entity β†’ get_related. get_overview exists because an agent that has just arrived needs to know whether this corpus can help before it spends a call guessing.

What is in the corpus

knowledge objects38
registered sources31
published topics93
typeobjects
entity25
guide9
comparison2
faq1
glossary1

Subject matter: evals and benchmarks (GAIA, AgentBench, Inspect), LLM-as-judge and its failure modes, Goodhart and benchmark contamination, fault injection and chaos testing, approval gates and autonomy levels, grounding and faithfulness.

Questions it is built to answer

  • How do I tell a real eval from a benchmark my agent has memorised?
  • What does calibration mean for an LLM judge, and how is it measured?
  • Which failure modes does fault injection actually catch?

What an answer actually looks like

A real call against the live endpoint β€” answer with "how do I tell a real eval from benchmark contamination" β€” returns this structuredContent, trimmed:

config.json
{
  "answered": true,
  "entity": {
    "id": "agent-reliability-glossary",
    "name": "Agent reliability glossary",
    "evidence_tier": "secondary",
    "confidence": 0.85,
    "last_verified": "2026-08-08",
    "canonical_url": "https://agentreliability.dev/k/agent-reliability-glossary"
  },
  "claims": [
    {
      "text": "An eval is a structured, repeatable test that measures an LLM or LLM-based system against a defined dimension; frameworks package evals as registries of reusable templates.",
      "sources": [{ "title": "openai/evals β€” framework for evaluating LLMs and LLM systems" }]
    }
  ]
}

Note what travels with the answer: the evidence tier, a confidence, the date it was last verified, and the source behind the claim β€” not as prose an agent has to parse, but as fields it can act on. An agent can decline to use a weak claim, or cite the primary source directly.

When the corpus cannot answer, answered is false. It does not improvise, and the miss is recorded so the gap can be filled.

Machine-readable surfaces

The MCP endpoint is one of several. The same corpus is served as plain files an agent can read directly:

surfacewhat it is
/llms.txtthe index, as text/plain
/llms-full.txtthe whole corpus in one file
/ai-index.jsonevery surface this instance publishes, with its content type
/api/index.jsonone JSON document per knowledge object
/api/sources.jsonthe source registry, in full
/.well-known/mcp/server.jsonthis server's manifest

Each knowledge object has a human page and a machine twin at the same id, with a canonical URL that agrees across all of them.

Behaviour worth knowing before you integrate

  • POST only. Every other method answers 405 with an Allow: POST, OPTIONS header.
  • Rate limit: 120 requests per minute per client, counted in a shared store, published on every response as RateLimit-Limit, RateLimit-Remaining and RateLimit-Reset (all three exposed via CORS). It fails open: if the store is unreachable the request is served.
  • Malformed input gets a spec-correct JSON-RPC error β€” -32700 for unparseable bodies, -32602 for an unknown tool β€” never an HTML error page.
  • Request bodies are capped and validated before transport.

Privacy

No accounts, no cookies, no ads. Usage is measured in aggregate with daily-rotating hashed identifiers and a 200-day retention; raw IPs are never stored. Full policy: PRIVACY.md.

Provenance and licence

Knowledge content is CC-BY-4.0: use it, cite it. The source registry is public precisely so a claim can be checked rather than trusted β€” get_sources returns what any given claim rests on.

Claims carry an evidence tier and a last_verified date. Where the evidence is weaker, the object says so rather than rounding up.

How it is built

Compiled and served by Citarium, a source-available framework for turning a knowledge graph into a website, an API, an MCP server and agent-readable files from a single source, under external evaluation.

The framework's code is licensed under the Business Source License 1.1 and its repository is not public. What is public β€” and what actually matters for trusting an answer β€” is this server, the corpus it serves, and the registered source behind every claim: get_sources returns what any given claim rests on, so it can be checked rather than trusted.

This repository is the server's public face: its manifest and its documentation. The corpus itself lives at agentreliability.dev.

Related MCP Servers

View all in Developer Tools View all alternatives
  • PraisonAI logoPraisonAI

    AI Agents Framework with Self Reflection and MCP support

    πŸ’» Developer Tools1 views
    Compare vs PraisonAI β†’
  • Agent Guild logoAgent Guild

    Attack-resistant reputation and trust layer for autonomous AI agents β€” discover, vet, and vouch.

    πŸ’» Developer Tools1 views
    Compare vs Agent Guild β†’
  • Flutter Skill logoFlutter Skill

    AI E2E testing bridge β€” give AI eyes and hands inside any app. 8 platforms, 40+ tools.

    πŸ’» Developer Tools1 views
    Compare vs Flutter Skill β†’
  • Labelhead Artist Momentum logoLabelhead Artist Momentum

    Trending hip-hop artist momentum scores across four cultural dimensions.

    πŸ’» Developer Tools0 views
    Compare vs Labelhead Artist Momentum β†’

Reviews

No reviews yet β€” be the first to share how this listing worked for you.

Frequently Asked Questions about Agent Reliability

Add the following block to your claude_desktop_config.json under mcpServers: "mcpServers": { "agent-reliability": { "command": "npx", "args": ["-y", "Agent Reliability"] } }

AllMCPs Directory Badge

Full Badge Customizer

Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.

Badge Style:
Live Dynamic SVG PreviewAgent Reliability AllMCPs Directory Badge
Markdown (GitHub README)
[![AllMCPs](https://allmcps.com/api/badge/agent-reliability?style=directory)](https://allmcps.com/mcp/agent-reliability)
HTML Embed
<a href="https://allmcps.com/mcp/agent-reliability"><img src="https://allmcps.com/api/badge/agent-reliability?style=directory" alt="Agent Reliability on AllMCPs" /></a>

Technical Specs & Signals

CategoryπŸ’»Developer Tools
More technical detailsExpand β–Ύ
TransportSTDIO
RuntimeNode.js
Views0
Unique ViewsTotal visits recorded for this listing page on AllMCPs.
Installs0
Installs & Copy ActionsTotal times users copied install commands or configuration snippets for this server.
27Quality signal: Emerging Β· 27/100How this signal is calculated β–Ύ
Server availabilityNot measured

Not scored for repo-hosted servers β€” we can't reach the running server, only its GitHub page. Hosted MCP endpoints are health-checked live.

Verified ownership8/20
Documentation & tools11/30
Adoption & activity1/15
Community engagement0/10

A guidance signal from public completeness & health data β€” not a user rating. New listings start lower and rise as they add docs, get verified, and grow adoption. Signals we can't observe for a listing are skipped, not counted against it.

β˜… FeaturedAllMCPs Server logo

AllMCPs Server

The official MCP server for AllMCPs.com - submit and manage tools directly from your AI. The open directory for MCP servers. Connect Claude, Cursor, Windsurf, and AI agents to databases, tools, files, and APIs. Explore 10,000+ servers. AllMCPs is the premier, open directory for discovering, evaluating, and installing Model Context Protocol (MCP) servers to equip AI agents and LLMs with real-world superpowers.

Explore Server β†’

Own this project?

This directory is pre-filled from public sources. Claim via GitHub README, site badge, or DNS TXT to unlock edit access and the Official badge and attach your website β€” proof is checked automatically, then reviewed by our team.

Free dofollow backlink: add your website and place the AllMCPs badge on it β€” no claim needed. We detect it automatically and keep it verified as long as the badge stays live.

Claim & get free dofollow

Share & Embed

Add our SVG badge (dark/light directory styles) or embeddable widget to your site.

Explore more

More in πŸ’» Developer Tools β†’Best MCP servers for Developers β†’Alternatives to Agent Reliability β†’Install in Claude DesktopInstall in CursorInstall in VS Code