The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the MCP Eval Runner listing page.
npm mcp-eval-runner package
A standardized testing harness for MCP servers and agent workflows. Define test cases as YAML fixtures (steps → expected tool calls → expected outputs), run regression suites directly from your MCP client, and get pass/fail results with diffs — without leaving Claude Code or Cursor.
Tool reference | Configuration | Fixture format | Contributing | Troubleshooting | Design principles
expected_output without a server.output_contains, output_not_contains, output_equals, output_matches, schema_match, tool_called, and latency_under per step.{{steps.<step_id>.output}}.Add the following config to your MCP client:
By default, eval fixtures are loaded from ./evals/ in the current working directory. To use a different path:
Amp · Claude Code · Cline · Cursor · VS Code · Windsurf · Zed
Create a file at evals/smoke.yaml. Use live mode (recommended) by including a server block:
Then enter the following in your MCP client:
Your client should return a pass/fail result for the smoke test.
Fixtures are YAML (or JSON) files placed in the fixtures directory. Each file defines one test case.
| Field | Required | Description |
|---|---|---|
name | Yes | Unique name for the test case |
description | No | Human-readable description |
server | No | Server config — if present, runs in live mode; if absent, runs in simulation mode |
steps | Yes | Array of steps to execute |
server block (live mode)When server is present the eval runner spawns the server as a child process, connects via MCP stdio transport, and calls each step's tool against the live server.
steps arrayEach step has the following fields:
| Field | Required | Description |
|---|---|---|
id | Yes | Unique identifier within the fixture (used for output piping) |
tool | Yes | MCP tool name to call |
description | No | Human-readable step description |
input | No | Key-value map of arguments passed to the tool (default: {}) |
expected_output | No | Literal string used as output in simulation mode |
expect | No | Assertions evaluated against the step output |
Live mode — fixture has a server block:
Simulation mode — no server block:
expected_output (or empty string if absent).output_contains assertions will always fail if expected_output is not set.All assertions go inside a step's expect block:
Multiple assertions in one expect block are all evaluated; the step fails if any assertion fails.
Reference the output of a previous step in a downstream step's input using {{steps.<step_id>.output}}:
Piping works in both live mode and simulation mode.
create_test_caseFixtures created with the create_test_case tool do not include a server block. They always run in simulation mode. To use live mode, add a server block manually to the generated YAML file.
run_suite — execute all fixtures in the fixtures directory; returns a pass/fail summaryrun_case — run a single named fixture by namelist_cases — enumerate available fixtures with step counts and descriptionscreate_test_case — create a new YAML fixture file (simulation mode; no server block)scaffold_fixture — generate a boilerplate fixture with placeholder steps and pre-filled assertion commentsregression_report — compare the current fixture state to the last run; surfaces regressions and fixescompare_results — diff two specific runs by run IDgenerate_html_report — generate a single-file HTML report for a completed runevaluate_deployment_gate — CI gate; fails if recent pass rate drops below a configurable thresholddiscover_fixtures — discover fixture files across one or more directories (respects FIXTURE_LIBRARY_DIRS)--fixtures / --fixtures-dirDirectory to load YAML/JSON eval fixture files from.
Type: string
Default: ./evals
--db / --db-pathPath to the SQLite database file used to store run history.
Type: string
Default: ~/.mcp/evals.db
--timeoutMaximum time in milliseconds to wait for a single step before marking it as failed.
Type: number
Default: 30000
--watchWatch the fixtures directory and rerun the affected fixture automatically when files change.
Type: boolean
Default: false
--formatOutput format for eval results.
Type: string
Choices: console, json, html
Default: console
--concurrencyNumber of test cases to run in parallel.
Type: number
Default: 1
--http-portStart an HTTP server on this port instead of stdio transport.
Type: number
Default: disabled (uses stdio)
Pass flags via the args property in your JSON config:
Before publishing a new version, verify the server with MCP Inspector to confirm all tools are exposed correctly and the protocol handshake succeeds.
Interactive UI (opens browser):
CLI mode (scripted / CI-friendly):
Run before publishing to catch regressions in tool registration and runtime startup.
New assertion types go in src/assertions.ts — implement the Assertion interface and add a test. Integration tests live under tests/ as unit tests and under evals/ as eval fixtures.
This plugin is available on:
Search for mcp-eval-runner.