The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Evals MCP Server listing page.
Author verifiable eval records through a draft → review → revise → submit loop with server-enforced graders; compile to JSONL/CSV/Inspect/lm-eval via MCP. STDIO or Streamable HTTP.
Nine tools for authoring eval records — the draft loop (create, revise, discard, submit), the standalone deterministic checker, and read/list/export:
| Tool | Description |
|---|---|
evals_describe_schema | Return the required and optional fields plus grader options for a task type. Call before drafting. |
evals_create_draft | Create a draft eval record carrying its own grader; returns the parsed record, a review protocol, and a verification subagent prompt. |
evals_get_record | Read a draft or submitted record by id; the id is stable across submit. |
evals_revise_draft | Apply a surgical set / append / unset patch to a draft by dotted path; re-runs the self-consistency check. |
evals_discard_draft | Delete a draft record by id. Draft-only. |
evals_run_check | Run a grader spec against candidate answers and get PASS/REJECT per candidate, decoupled from any saved record. |
evals_submit_draft | Finalize a draft through the committability gate, then freeze it. |
evals_list_records | Browse and filter records by status, domain, task type, or tag. Returns a compact summary per record. |
evals_export_records | Compile submitted records to JSONL, CSV, Inspect AI, or lm-evaluation-harness and write the artifact under exports/. |
evals_describe_schemaReturn what a record of a given task_type needs before you draft it.
task_type is one of numeric, exact_answer, set_answer, mcq, regex_answer, json_answer, free_responseevals_create_draftCreate and persist a draft eval record, then reflect it back as a review forcing function.
task_type discriminated union (per-type rules: mcq requires choices; free_response requires an llm_rubric grader)verification block and captures (EvalsIDs) when you already hold provenancedraft — passing self-consistency proves the grader discriminates, not that the gold is rightevals_revise_draftSurgically patch a draft so each change stays legible.
set (dotted-path → value), append (dotted-path → array items), and unset (dotted paths) operations — not full-record rewritestask_type cannot be patched (start a new draft to change the discriminant)evals_run_checkRun a grader against one or more candidates without touching a saved record.
candidates accepts strings, numbers, objects, or arrays — whatever the grader kind expectsgold for gold-relative kinds (exact_match); it is a no-op for target-embedding kinds like numeric and mcqllm_rubric cannot run on this server — submission relies on recorded independent verificationevals_submit_draftFinalize a draft through the committability gate, then freeze it.
captures from EVALS_CAPTURE_DIR, cross-checking the gold against the authoritative captured valuecontent_hash already submitted)submitted, stamps submitted_at and a checksum, and freezes it; otherwise refuses with a typed error and the record stays a draftfree_response llm_rubric is admitted on recorded independent verification and flagged server_verified: falseevals_export_recordsCompile submitted records to a downstream eval format.
jsonl (lossless, the lingua franca), csv (a flattened, lossy spreadsheet summary), inspect (UK AISI Inspect AI), lm-eval (EleutherAI lm-evaluation-harness)domain / task_type / tag filterexports/ and returns the file path, record count, byte size, and a short preview instead of dumping inline| Type | Name | Description |
|---|---|---|
| Resource | eval://record/{id} | A single draft or submitted record by id — the same payload evals_get_record returns, for resource-capable clients. |
All record data is also reachable through the tool surface — evals_get_record for a single record, evals_list_records to browse. The resource is a convenience mirror for clients that support resources, not the access path.
Built on @cyanheads/mcp-ts-core:
reason + recovery), surfaced to the agentnone, jwt, oauthEval authoring:
draft → review → surgical-revise → submit loop, with the server acting as both scribe (normalize, persist, compile) and adversarial checker (run the record's own grader, reject what doesn't hold up)discriminatedUnion keyed on task_type — numeric, exact_answer, set_answer, mcq, regex_answer, json_answer, free_responsenumeric via math.js, exact_match, set_match, regex, mcq, json_match) run server-side; llm_rubric relies on recorded independent verificationcaptures EvalsID field — link framework-written tool-call dumps, resolved from EVALS_CAPTURE_DIR and cross-checked against the gold (no server-to-server calls)EVALS_DATA_DIR — inspectable, diffable, version-controllable recordsAgent-friendly output:
evals_create_draft, evals_revise_draft) carry the loop's review mechanism in their responses — the parsed record parroted back, a per-field review protocol, and a ready subagent promptevals_list_records discloses truncation when the limit is hit, so a partial set is never mistaken for the whole corpusreason + recovery hint, so a rejected record tells the agent exactly what to fixAdd the following to your MCP client configuration file. Set EVALS_DATA_DIR to a writable folder — the server manages drafts/, submitted/, and exports/ under it.
Or with npx (no Bun required):
Or with Docker:
For Streamable HTTP, set the transport and start the server:
EVALS_DATA_DIR. No external API key is required.All server configuration is validated at startup via Zod schemas in src/config/server-config.ts.
| Variable | Description | Default |
|---|---|---|
EVALS_DATA_DIR | Root folder for record JSON; the store manages drafts/, submitted/, and exports/ under it. | ./evals-data |
EVALS_REQUIRE_CONFIRMATION | When true, evals_submit_draft requests human confirmation through multi-round input before finalizing. | false |
EVALS_DEFAULT_LICENSE | Default metadata.license applied when a draft omits one (e.g. CC-BY-4.0). | — |
EVALS_CAPTURE_DIR | Directory of framework-written tool-call captures; when set, captures EvalsIDs resolve to full dumps. | — |
MCP_TRANSPORT_TYPE | Transport: stdio or http. | stdio |
MCP_HTTP_PORT | Port for the HTTP server. | 3010 |
MCP_AUTH_MODE | Auth mode: none, jwt, or oauth. | none |
MCP_LOG_LEVEL | Log level (RFC 5424). | info |
OTEL_ENABLED | Enable OpenTelemetry instrumentation. | false |
See .env.example for the full list of optional overrides.
Build and run:
Run checks and tests:
The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/evals-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.
| Directory | Purpose |
|---|---|
src/index.ts | createApp() entry point — registers tools and the resource, inits the record-store and exporter services. |
src/config | Server-specific environment variable parsing and validation with Zod. |
src/mcp-server/tools | Tool definitions (*.tool.ts). |
src/mcp-server/resources | Resource definitions (*.resource.ts). |
src/services/eval-record | The record schema, draft builder, and submit gate. |
src/services/grader | Deterministic grader DSL execution and the committability check. |
src/services/record-store | On-disk JSON record CRUD, the draft→submitted move, and export writes. |
src/services/exporter | Compiling submitted records to JSONL/CSV/Inspect/lm-eval. |
tests/ | Unit and integration tests mirroring src/. |
See CLAUDE.md/AGENTS.md for development guidelines and architectural rules. The short version:
try/catch in tool logicctx.log for request-scoped logging; records persist to disk via the record-store service, not ctx.statecreateApp() arrays in src/index.tsIssues and pull requests are welcome. Run checks and tests before submitting:
Apache-2.0 — see LICENSE for details.