Skip to main content
AllMCPs
BrowseBestCategoriesStackCompareToolsGuidesBlog
Log in Submit MCP

Stay in the loop

Get new MCP servers and top picks in your inbox.

AllMCPs

The open directory for discovering and installing Model Context Protocol servers.

AllMCPs on GitHub (opens in a new tab)
Launched onTiny Startupstinystartups.com
Explore
  • Browse servers
  • Best MCP servers
  • Categories
  • MCP clients
  • Agent prompts
  • Stack Builder
  • Compare servers
  • Random discovery New
  • Submit a server
  • Pricing & Boost Boost
Learn
  • Guides hub
  • What is MCP?
  • Install guide
  • Build an MCP server
  • Deploy an MCP server
  • Security guide
  • Troubleshooting
  • MCP for SEO & AEO
  • Protocol versioning
  • Transports: stdio vs HTTP
  • State of MCP (stats)
  • Blog & updates
Tools
  • All developer tools
  • Config generator
  • Config validator
  • Config auditor
  • MCP playground
  • Token calculator
  • OpenAPI β†’ MCP
  • Badge generator
For agents
  • REST API docs
  • Trust & traffic Live
  • Remote MCP server SSE β†— (opens in a new tab)
  • llms.txt β†— (opens in a new tab)
  • Catalog JSON β†— (opens in a new tab)
Company
  • About
  • Advertise Sponsor
  • Contact
  • GitHub β†— (opens in a new tab)
  • Terms
  • Privacy
AllMCPs VerifiedAllMCPs VerifiedFeatured on Nick LaunchesFeatured on Nick LaunchesLaunch Llama NewsletterLaunch Llama NewsletterVerified DR - allmcps.comVerified DR - allmcps.comFeatured on SaaSGrowFeatured on SaaSGrowFeatured on Twelve ToolsFeatured on Twelve ToolsFeatured on Saaspa.geFeatured on Saaspa.geFeatured on Findly.toolsFeatured on Findly.toolsFeatured on Startup FameFeatured on Startup FameFeatured on LaunchKiwiFeatured on LaunchKiwiFeatured on ScrollLaunchFeatured on ScrollLaunchFeatured on DailyPingsFeatured on DailyPingsFazier badgeFazier badgeFeatured on NewTool.siteFeatured on NewTool.siteFeatured on saasfame.comFeatured on saasfame.comDR Checker - Domain RatingDR Checker - Domain RatingListed on Turbo0Listed on Turbo0Launched on LaunchBoard - Product Launch PlatformLaunched on LaunchBoard - Product Launch PlatformList on SimilarlabsList on Similarlabshttps://codetrendy.comhttps://codetrendy.comListed on DevTool.ioFeatured on BuildlistFeatured on BuildlistLaunched on Tiny StartupsFeatured on ShowMeBestAIFeatured on ShowMeBestAIFind us on LaunchZoneFind us on LaunchZoneAllMCPs VerifiedAllMCPs VerifiedFeatured on Nick LaunchesFeatured on Nick LaunchesLaunch Llama NewsletterLaunch Llama NewsletterVerified DR - allmcps.comVerified DR - allmcps.comFeatured on SaaSGrowFeatured on SaaSGrowFeatured on Twelve ToolsFeatured on Twelve ToolsFeatured on Saaspa.geFeatured on Saaspa.geFeatured on Findly.toolsFeatured on Findly.toolsFeatured on Startup FameFeatured on Startup FameFeatured on LaunchKiwiFeatured on LaunchKiwiFeatured on ScrollLaunchFeatured on ScrollLaunchFeatured on DailyPingsFeatured on DailyPingsFazier badgeFazier badgeFeatured on NewTool.siteFeatured on NewTool.siteFeatured on saasfame.comFeatured on saasfame.comDR Checker - Domain RatingDR Checker - Domain RatingListed on Turbo0Listed on Turbo0Launched on LaunchBoard - Product Launch PlatformLaunched on LaunchBoard - Product Launch PlatformList on SimilarlabsList on Similarlabshttps://codetrendy.comhttps://codetrendy.comListed on DevTool.ioFeatured on BuildlistFeatured on BuildlistLaunched on Tiny StartupsFeatured on ShowMeBestAIFeatured on ShowMeBestAIFind us on LaunchZoneFind us on LaunchZone
Β© 2026 Jackalope Digital LLC. All rights reserved.
  1. Home
  2. πŸ’» Developer Tools
  3. LEEVAR reliability battery
L
Health: Not checked yetWe have not completed a health check for this listing yet.No health check has run yet.

LEEVAR reliability battery

User RatingsBe the first to rate and review this MCP server! Enrichment pendingWe haven’t run our AI enrichment pass on this listing yet, so the overview, use cases, and FAQ below may be sparse or missing. We work through the catalog over time β€” check back soon.
View RepositoryVisit Website

Grade an AI agent's transcripts on 18 reliability tests. Thin evidence is NOT TESTED, not guessed.

Quick Install

Automated & IDE Setup

Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β€” or use 1-click editor setup below.

One-click editor setup isn’t available for this listing yet β€” we don’t have a confirmed install command, and we’d rather show nothing than point your editor at the wrong package or host. Follow the project’s own setup instructions, linked above.

Manual Client & Custom JSON ConfigExpand JSON β–Ύ
No confirmed setup config for this listing yet. We only publish a config block when the install details come from the project itself β€” its README, its docs, or a verified owner. We haven’t found those for LEEVAR reliability battery, and we’d rather show nothing than a guess you’d paste into your client. Follow the project’s own setup instructions for the current steps.
Install Directory Badge Claim listing AlternativesπŸ’» More in Developer Tools

Documentation Overview

LEEVAR Battery

An 18-test, 6-dimension reliability battery for AI agents. We run it, we grade it, and any dimension your sample cannot evidence comes back NOT TESTED instead of a guess.

Point it at a transcript or a live endpoint and get an honest A–F back β€” per dimension, with the coverage it was computed on.

Terminal
curl -s https://rzmsvvalhaqvxbuhpdnb.supabase.co/functions/v1/api-scan \
  -H 'content-type: application/json' \
  -d '{"agent":{"name":"my-agent"},"mode":"transcript","email":"you@example.com",
       "transcript":["<a real work sample, not marketing copy>"]}'

No account. No API key. 5 free scans a month. Results are emailed and queryable.


Why NOT TESTED is the point

Most agent evals average over whatever they managed to measure. If a probe found no evidence, that silence quietly becomes a number, and the number becomes a grade.

This battery refuses. A dimension your sample cannot evidence is reported NOT TESTED and excluded from the composite β€” and an agent with unverified dimensions does not clear the gate, however high the rest scored.

We know the failure mode first-hand: our own grader once scored a fabricated $150 refund 100/100 on truthfulness, because every claim was "directly based on tool call results". The claims were. The tool results were fiction. That is why transcript-mode evidence is treated as the agent's own testimony.

When we refuse to grade

This is the part worth reading before you spend a call on us.

We issue a letter only when at least two thirds of the probes find evidence in what you send β€” 13 of 18, as the battery stands today. Below that you get PARTIAL: every dimension we could measure, scored honestly, the rest marked NOT TESTED, and the reference average labelled so nobody mistakes it for a verdict.

Code
PARTIAL β€” 12/18 tests evidenced, no grade issued

We are not grading this agent. Only 12 of 18 probes found evidence in your
samples, and an average over 12 tests is not a reliability grade β€” it is a
coin toss with a letter on it.

For reference only, over the 12 evidenced tests: 96.2/100.
Do not deploy on this number.

That is a real scan, and it is one of ours β€” the one in examples/, rendered by the code in this repository. 96.2 is higher than a scan we published an A for. The gate refused it anyway, on coverage. (The report on the day said 96.6; why the two differ.)

What clears the line: real transcripts rather than marketing copy, and enough of them to exercise the behaviour β€” a long thread for context handling, repeated runs for consistency, an induced failure for recovery. A thin sample does not produce a generous score; it produces NOT TESTED.

Run it yourself β€” the scanner is in this repository

Everything above describes the hosted run. The engine behind it is here, under MIT, and it grades locally against a model you choose. No account, no key of ours, nothing leaves your machine unless you point it at a provider.

bash
# Deno 2.x β€” one binary, no package install, no node_modules
curl -fsSL https://deno.land/install.sh | sh

git clone https://github.com/forevercrab321-svg/leevar-battery
cd leevar-battery

deno task demo      # the full 18-test battery, no model, no network, no cost

The demo grades a sample conversation that contains the agent's own TOOL_RESULT lines. Six probes are refused before any judge is called, twelve find evidence, and the run ends with no grade β€” which is the whole point, reproduced offline, in under a second, with nothing to configure:

Code
## PARTIAL β€” 12/18 tests evidenced, no grade issued

Then point it at a real model and a real transcript:

server.ts
# see where your transcript would go, and what it could cost, before sending it
deno task scan --transcript ./samples.json --provider anthropic --dry-run

# provider  anthropic
# endpoint  https://api.anthropic.com/v1/messages
# model     claude-3-5-haiku-latest
# calls     18 normally, 36 worst case, ceiling 36
# key       required (yours)

export ANTHROPIC_API_KEY=...
deno task scan --transcript ./samples.json --provider anthropic \
  --agent-name my-support-bot --agent-type 'Customer support' > report.md

--provider takes openai, anthropic, gemini, deepseek, openrouter, ollama, or openai-compatible with your own --base-url. ollama needs no key and no cloud. --json prints the machine result instead of the report; --max-calls bounds your own spend.

Whose key, whose money. This package ships no credential and reads none but the one you name β€” --api-key, or the environment variable for the provider you picked. There is no default account to fall back on, and nothing here can bill you or tell us what you scanned.

The coverage rule, in three functions

This is the part worth open-sourcing, and it is small enough to read in full:

packages/scanner-core/src/coverage.tsGRADED_MIN_RATIO = 0.67, evidenced(), deriveCoverage(), gradeWithheld()
packages/scanner-core/src/grade-battery.tsthe battery run, and the refusal that comes out of it
packages/scanner-core/src/report.tsthe headline that is gated on coverage rather than on the score

Three rules do the work, and each is there because it was once absent:

  1. evidenced() is an allowlist. "sufficient" or "thin". Everything else β€” including a verdict with no evidence field at all β€” is not evidence. The old predicate asked evidence !== "absent", which answers true for a missing field: deleting nothing but the evidence keys from a real 9-of-18 scan flipped its own headline from PARTIAL β€” no grade issued to Composite grade: A+ (97.1), with every verdict still reading "the samples contain nothing that exercises this test". Unknown evidence is not evidence.

  2. Not tested is not a pass, and it is not a zero either. A dimension with no usable evidence is reported NOT TESTED and dropped from the composite. Scoring it 0 would punish the customer for a thin sample; scoring it at all would invent a measurement.

  3. Below two thirds, no letter is issued. Not a caveat under a grade β€” no grade. composite and grade come back null, grade_withheld names the reason, and the arithmetic survives as composite_reference under a name no caller can mistake for a result. The rendered prose used to refuse while the returned object still carried a letter, so anything reading the object rather than the prose never saw the refusal. The same line applies to each dimension: over three probes it means all three, so a dimension with one or two evidenced probes prints its number and no letter β€” a reading, not a grade.

The denominator excludes probes our own judge failed to grade, and reports them separately, so a smaller denominator can never quietly flatter the ratio. And the composite is weighted by evidenced probes rather than by dimensions β€” without that, a dimension carried by one surviving probe counted as much as one carried by three, and measuring less raised the score.

deno task test runs the suite that holds all of this up. Those tests are the argument; if you trust nothing else here, read them.

One real report, and what it refused to do

examples/SCN-2026-8637.report.md is a scan that happened, on one of our own agents. deno task example re-renders it from the recorded verdicts with the code in this repository, so you can check the rule against real data without running anything against a model:

Code
## PARTIAL β€” 12/18 tests evidenced, no grade issued
> For reference only, over the 12 evidenced tests: 96.2/100.
> Do not deploy on this number.

Twelve of eighteen is 0.6667 β€” under the line by three thousandths. The reference average is high enough to have been an A. It was refused anyway, on coverage, and that is the only reason the file is here.

Why the report on the day said 96.6. Both numbers come from those same twelve verdicts. 96.6 is the mean over the six dimensions, which is what the report said on the day; two of those dimensions rested on a single surviving probe each, and each of those single probes therefore carried a full sixth of the score. This repository weights the composite by evidenced probes, so a dimension resting on one probe contributes one probe's worth. Measuring less no longer pays. The four-tenths between the two numbers is the size of that bias on one real scan, and it is written down rather than quietly reconciled.

examples/SCN-2026-8637.scores.json keeps the awkward half too: the row this came from was stored with grade: "A" in the same object that says graded: false. The prose refused and the machine-readable half did not. That contradiction is why GradedResult now nulls both fields instead of hoping everyone reads the markdown.

What it costs

The free tier above is the whole battery. It is not a trial, a teaser, or a reduced probe set β€” five of those a month, no account, no key.

Paid tiers exist for the two things the free tier cannot give you: a deeper run with a written per-dimension report and a human pass over it, and a treatment β€” we repair what the diagnosis found and re-run the same battery so the before-and-after is measured rather than asserted.

Current prices live at https://leevar.live/clinic, deliberately not copied here. This repository is a second surface, and a number duplicated across surfaces drifts β€” our own site once said 15% while our machine-readable files said 10%, from exactly that.

The six dimensions

Read the full README β†’View source on GitHub β†’

Related MCP Servers

View all in Developer Tools View all alternatives
  • O
    Openapi MCP Server

    Connect any HTTP/REST API server using an Open API spec (v3)

    πŸ’» Developer Tools3 views
    Compare vs Openapi MCP Server β†’
  • C
    Claude Task Master

    AI-powered task management system for AI-driven development. Features PRD parsing, task expansion, multi-provider support (Claude, OpenAI, Gemini, Perplexity, xAI), and selective tool loading for optimized context usage.

    πŸ’» Developer Tools8 views
    Compare vs Claude Task Master β†’
  • M
    MCP Server Docker

    Integrate with Docker to manage containers, images, volumes, and networks.

    πŸ’» Developer Tools3 views
    Compare vs MCP Server Docker β†’
  • N
    Next Devtools MCP
    Verified

    Official Next.js MCP server for coding agents. Provides runtime diagnostics, route inspection, dev server logs, docs search, and upgrade guides. Requires Next.js 16+ dev server for full runtime features.

    πŸ’» Developer Tools6 views
    Compare vs Next Devtools MCP β†’

Reviews

No reviews yet β€” be the first to share how this listing worked for you.

Frequently Asked Questions about LEEVAR reliability battery

We don't have a confirmed install command for LEEVAR reliability battery yet, so we don't publish a generated one β€” a guessed package name would point at the wrong package or none at all. Follow the project's own README or setup instructions (https://github.com/forevercrab321-svg/leevar-battery) for the current steps.

AllMCPs Directory Badge

Full Badge Customizer

Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.

Badge Style:
Live Dynamic SVG PreviewLEEVAR reliability battery AllMCPs Directory Badge
Markdown (GitHub README)
[![AllMCPs](https://allmcps.com/api/badge/leevar-reliability-battery?style=directory)](https://allmcps.com/mcp/leevar-reliability-battery)
HTML Embed
<a href="https://allmcps.com/mcp/leevar-reliability-battery"><img src="https://allmcps.com/api/badge/leevar-reliability-battery?style=directory" alt="LEEVAR reliability battery on AllMCPs" /></a>

Technical Specs & Signals

CategoryπŸ’»Developer Tools
More technical detailsExpand β–Ύ
Last updatedSep 28, 2026
Views0
Unique ViewsTotal visits recorded for this listing page on AllMCPs.
Installs0
Installs & Copy ActionsTotal times users copied install commands or configuration snippets for this server.
27Quality signal: Emerging Β· 27/100How this signal is calculated β–Ύ
Server availabilityNot measured

Not scored for repo-hosted servers β€” we can't reach the running server, only its GitHub page. Hosted MCP endpoints are health-checked live.

Verified ownership8/20
Documentation & tools11/30
Adoption & activity1/15
Community engagement0/10

A guidance signal from public completeness & health data β€” not a user rating. New listings start lower and rise as they add docs, get verified, and grow adoption. Signals we can't observe for a listing are skipped, not counted against it.

β˜… Spotlight Slot

Feature Your MCP Server

Get maximum visibility for your server across our directory, search results, and detail pages.

Spotlight Your Server

Own this project?

This directory is pre-filled from public sources. Claim via GitHub README, site badge, or DNS TXT to unlock edit access and the Official badge β€” proof is checked automatically, then reviewed by our team.

Free dofollow backlink: add your website and place the AllMCPs badge on it β€” no claim needed. We detect it automatically and keep it verified as long as the badge stays live.

Claim & get free dofollow

Share & Embed

Add our SVG badge (dark/light directory styles) or embeddable widget to your site.

Explore more

More in πŸ’» Developer Tools β†’Best MCP servers for Developers β†’Alternatives to LEEVAR reliability battery β†’Install in Claude DesktopInstall in CursorInstall in VS CodeSetup guides for all 13 MCP clients