Grade an AI agent's transcripts on 18 reliability tests. Thin evidence is NOT TESTED, not guessed.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
One-click editor setup isnβt available for this listing yet β we donβt have a confirmed install command, and weβd rather show nothing than point your editor at the wrong package or host. Follow the projectβs own setup instructions, linked above.
An 18-test, 6-dimension reliability battery for AI agents. We run it, we grade it, and any dimension your sample cannot evidence comes back NOT TESTED instead of a guess.
Point it at a transcript or a live endpoint and get an honest AβF back β per dimension, with the coverage it was computed on.
No account. No API key. 5 free scans a month. Results are emailed and queryable.
NOT TESTED is the pointMost agent evals average over whatever they managed to measure. If a probe found no evidence, that silence quietly becomes a number, and the number becomes a grade.
This battery refuses. A dimension your sample cannot evidence is reported NOT TESTED and excluded from the composite β and an agent with unverified dimensions does not clear the gate, however high the rest scored.
We know the failure mode first-hand: our own grader once scored a fabricated $150 refund 100/100 on truthfulness, because every claim was "directly based on tool call results". The claims were. The tool results were fiction. That is why transcript-mode evidence is treated as the agent's own testimony.
This is the part worth reading before you spend a call on us.
We issue a letter only when at least two thirds of the probes find evidence in
what you send β 13 of 18, as the battery stands today. Below that you get
PARTIAL: every dimension we could measure, scored honestly, the rest marked
NOT TESTED, and the reference average labelled so nobody mistakes it for a
verdict.
That is a real scan, and it is one of ours β the one in examples/, rendered by
the code in this repository. 96.2 is higher than a scan we published an A
for. The gate refused it anyway, on coverage. (The report on the day said 96.6;
why the two differ.)
What clears the line: real transcripts rather than marketing copy, and
enough of them to exercise the behaviour β a long thread for context handling,
repeated runs for consistency, an induced failure for recovery. A thin sample
does not produce a generous score; it produces NOT TESTED.
Everything above describes the hosted run. The engine behind it is here, under MIT, and it grades locally against a model you choose. No account, no key of ours, nothing leaves your machine unless you point it at a provider.
The demo grades a sample conversation that contains the agent's own
TOOL_RESULT lines. Six probes are refused before any judge is called, twelve
find evidence, and the run ends with no grade β which is the whole point,
reproduced offline, in under a second, with nothing to configure:
Then point it at a real model and a real transcript:
--provider takes openai, anthropic, gemini, deepseek, openrouter,
ollama, or openai-compatible with your own --base-url. ollama needs no
key and no cloud. --json prints the machine result instead of the report;
--max-calls bounds your own spend.
Whose key, whose money. This package ships no credential and reads none but
the one you name β --api-key, or the environment variable for the provider you
picked. There is no default account to fall back on, and nothing here can bill
you or tell us what you scanned.
This is the part worth open-sourcing, and it is small enough to read in full:
packages/scanner-core/src/coverage.ts | GRADED_MIN_RATIO = 0.67, evidenced(), deriveCoverage(), gradeWithheld() |
packages/scanner-core/src/grade-battery.ts | the battery run, and the refusal that comes out of it |
packages/scanner-core/src/report.ts | the headline that is gated on coverage rather than on the score |
Three rules do the work, and each is there because it was once absent:
evidenced() is an allowlist. "sufficient" or "thin". Everything
else β including a verdict with no evidence field at all β is not evidence.
The old predicate asked evidence !== "absent", which answers true for a
missing field: deleting nothing but the evidence keys from a real
9-of-18 scan flipped its own headline from PARTIAL β no grade issued to
Composite grade: A+ (97.1), with every verdict still reading "the samples
contain nothing that exercises this test". Unknown evidence is not evidence.
Not tested is not a pass, and it is not a zero either. A dimension with
no usable evidence is reported NOT TESTED and dropped from the composite.
Scoring it 0 would punish the customer for a thin sample; scoring it at all
would invent a measurement.
Below two thirds, no letter is issued. Not a caveat under a grade β no
grade. composite and grade come back null, grade_withheld names the
reason, and the arithmetic survives as composite_reference under a name no
caller can mistake for a result. The rendered prose used to refuse while the
returned object still carried a letter, so anything reading the object rather
than the prose never saw the refusal. The same line applies to each
dimension: over three probes it means all three, so a dimension with one or
two evidenced probes prints its number and no letter β a reading, not a
grade.
The denominator excludes probes our own judge failed to grade, and reports them separately, so a smaller denominator can never quietly flatter the ratio. And the composite is weighted by evidenced probes rather than by dimensions β without that, a dimension carried by one surviving probe counted as much as one carried by three, and measuring less raised the score.
deno task test runs the suite that holds all of this up. Those tests are the
argument; if you trust nothing else here, read them.
examples/SCN-2026-8637.report.md is a scan that happened, on one of our own
agents. deno task example re-renders it from the recorded verdicts with the
code in this repository, so you can check the rule against real data without
running anything against a model:
Twelve of eighteen is 0.6667 β under the line by three thousandths. The reference average is high enough to have been an A. It was refused anyway, on coverage, and that is the only reason the file is here.
Why the report on the day said 96.6. Both numbers come from those same twelve verdicts. 96.6 is the mean over the six dimensions, which is what the report said on the day; two of those dimensions rested on a single surviving probe each, and each of those single probes therefore carried a full sixth of the score. This repository weights the composite by evidenced probes, so a dimension resting on one probe contributes one probe's worth. Measuring less no longer pays. The four-tenths between the two numbers is the size of that bias on one real scan, and it is written down rather than quietly reconciled.
examples/SCN-2026-8637.scores.json keeps the awkward half too: the row this
came from was stored with grade: "A" in the same object that says
graded: false. The prose refused and the machine-readable half did not. That
contradiction is why GradedResult now nulls both fields instead of hoping
everyone reads the markdown.
The free tier above is the whole battery. It is not a trial, a teaser, or a reduced probe set β five of those a month, no account, no key.
Paid tiers exist for the two things the free tier cannot give you: a deeper run with a written per-dimension report and a human pass over it, and a treatment β we repair what the diagnosis found and re-run the same battery so the before-and-after is measured rather than asserted.
Current prices live at https://leevar.live/clinic, deliberately not copied here. This repository is a second surface, and a number duplicated across surfaces drifts β our own site once said 15% while our machine-readable files said 10%, from exactly that.
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/leevar-reliability-battery)<a href="https://allmcps.com/mcp/leevar-reliability-battery"><img src="https://allmcps.com/api/badge/leevar-reliability-battery?style=directory" alt="LEEVAR reliability battery on AllMCPs" /></a>