The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the RetriEVAL listing page.
LLM evaluation as an MCP server. Score your AI's outputs for faithfulness, relevancy, and hallucination from inside any MCP client — no pipeline, no test harness. Every result comes back with a link to a dashboard that keeps the history.
Try it live (no signup) · Watch the 2-minute demo · Dashboard
Five customer-support answers, scored on two metrics:
| metric | score | passing |
|---|---|---|
| answer_relevancy | 0.98 | 5/5 |
| faithfulness | 0.70 | 3/5 |
Every answer was on-topic and well-written. Two of them contradicted the policy they were supposedly grounded in — one promised free return shipping the policy doesn't offer, another invented a free overnight replacement. Reviewing by eye, you'd sign off on all five.
That gap is the point. Relevancy asks did it answer the question. Faithfulness asks is it actually in the source. You need both, and the second one catches the expensive failures.
Self-hosted. Clone it, point it at a judge, run it — your data never leaves your machine and there's no service to sign up for.
Then just ask:
Score these cases with faithfulness: [{"input": "...", "actual_output": "...", "retrieval_context": ["..."]}]
Pass your cases inline and nothing is stored — one call, no setup step. See Run locally (stdio) for client config, and Deploy as HTTP if you want your own always-on instance with a dashboard.
Want to try it before installing anything? There's a live sandbox at retrieval-mcp.com — no signup, runs on a free judge, nothing saved.
Keeping everything local: set RETRIEVAL_JUDGE_BACKEND=ollama and the judge
runs on your machine too, so no data leaves your network at any point. Useful if
you're evaluating anything you can't send to a third party.
faithfulness · answer_relevancy · contextual_precision ·
contextual_recall · contextual_relevancy · hallucination · bias ·
toxicity · summarization — plus authored G-Eval metrics you define in
plain language. All are normalized so higher = better (bias/toxicity report
the clean fraction), and each reasons before scoring.
load_golden_set accepts a file path (including uploaded files), an
http(s) URL, an inline JSON array, or JSONL text, in
JSON / JSONL / CSV / TSV. Field names are auto-normalized (question→input,
answer→actual_output, ground_truth→expected_output, contexts→context,
passages→retrieval_context, …), so most public benchmarks load as-is.
run_eval scores every case but returns only the 3 lowest-scoring by
default (tune with limit), with total_cases/shown and a pointer to
show_run_cases(run_id, offset, limit, metric) to page through the rest.
Judge spend is metered from real token usage and persisted. Set a hard cap:
Once cumulative spend hits the cap, further Anthropic calls stop and tools
return a clear budget_exceeded message. Check/clear with get_budget /
reset_budget. (Prices are approximate — override RETRIEVAL_PRICE_IN/OUT
$/1M tokens to match current pricing for your model.)
Run history uses a pluggable store, chosen by env:
~/.retrieval. Zero setup, local only.SUPABASE_URL + SUPABASE_SERVICE_KEY are set. Run
history lives in Postgres, shared by the local CLI, the deployed MCP, and the
website dashboard.Recommended hybrid flow:
supabase_schema.sql in Supabase (creates the runs table).SUPABASE_URL + SUPABASE_SERVICE_KEY on the MCP (local and/or Railway)
so every run is written centrally. Each run records its generator_model and
judge_model for cross-model comparison.web/ to Vercel (set the same Supabase env vars) and map it to
retrieval-mcp.com. The dashboard reads history via /api/runs (service key
stays server-side) and renders trend-by-model, a model leaderboard, and run
history. It shows sample data until Supabase is wired.Local stays your free sandbox (Ollama judge, file history); the website is the always-on window into the shared history.
Then: "Load examples/rag_golden.jsonl as 'space', run faithfulness, label it v1."
Deploy to Railway (or any host): the included Dockerfile / Procfile
work as-is. Set ANTHROPIC_API_KEY, RETRIEVAL_TOKEN, RETRIEVAL_BUDGET_USD
in the host env. Clients connect to https://<host>/mcp with header
Authorization: Bearer <token> — add it as a custom connector in
claude.ai / Claude Desktop, or point Agent Builder / CI at it. State (golden
sets, run history, spend) lives server-side, so it persists across machines.
demo/rag_demo.py builds a tiny end-to-end RAG over a small labeled dataset
(demo/labeled.json + demo/corpus.json): it retrieves with BM25, computes
recall@k against the gold passages (deterministic — the retriever's score),
generates an answer, then scores faithfulness (the generator's score). One answer
is deliberately hallucinated so you watch the two failure modes separate.
It also writes demo/generated_goldenset.jsonl — load that into the MCP
(load_golden_set → run_eval) for the judge-scored version. This is the bridge:
your pipeline emits predictions, the dataset supplies the labels, and RetriEval
scores retriever and generator independently.
query_rag(question) tool that calls your RAG endpoint or your
vector store (Chroma / Supabase pgvector), captures context + answer, and scores
in one shot.| Tool | Purpose |
|---|---|
list_metrics | built-in + authored metrics |
load_golden_set(name, source, fmt) | name a set for reuse (self-host only — shared and lost on restart) |
list_golden_sets | what's loaded |
author_metric(name, criteria, examples) | plain language → a scorer |
run_eval(metrics, cases, golden_set, threshold, outputs, label, limit) | score a set; pass cases inline (JSON/JSONL/CSV/TSV/path/URL) — nothing stored |
show_run_cases(run_id, offset, limit, metric) | page the rest |
evaluate_case(...) | one-off score |
ground_against_url(url, output, question) | check an output's consistency with a web page (no labels — consistency, not correctness) |
list_runs(golden_set, last_n) | saved runs |
plot_metric_trend / plot_run / compare_runs | inline charts |
get_budget / reset_budget | spend cap status / reset |
Apache License 2.0. Built by Hanns Carrillo.