The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Vibe Test — Browser Testing Agent listing page.
Code-aware browser testing for AI coding agents.
Product site · npm · MCP Registry · Launch status
vibe-testing reads your codebase so tests use your real routes and field names, runs them in a real Playwright browser, remembers what broke, and tells you what your last change fixed or regressed. It works as an MCP server that gives your editor (Claude Code, Cursor, Windsurf, VS Code Copilot, Roo Code) 14 testing tools, or as a standalone CLI.
The repository is also an Agent Plugin with a release-qa skill. Compatible agents get both the testing tools and a senior-QA workflow for coverage, evidence, regression checks, and a clear release verdict.
Then open your editor and say:
"Scan this codebase and test it against http://localhost:3000"
Playwright MCP gives your agent hands. vibe-testing gives it a testing workflow: code-derived scenarios, memory across runs, and a report. Two things a stateless browser tool cannot do:
1. Run it twice and it tells you what you broke. Every run writes .vibe/run-snapshot.json and diffs it against the previous run. Every scan writes .vibe/route-manifest.json and diffs your routes. The second run prints regressions and fixes instead of a wall of results:
The same diff reaches your editor as snapshot_diff on run_full_test and run_converge, and as route_changes on scan_codebase, so the agent can flag "checkout broke after that commit" without anyone scrolling a report. Flaky routes, working selectors, and measured timeouts are also remembered between runs.
2. Zero LLM calls inside the tool. Pass/fail verification is heuristic: URL changes, toast detection, API errors. Your editor's model decides what to test; vibe-testing does the browsing and checking. No API key, no per-run cost beyond the editor subscription you already pay for.
No test cases to write. The AI reads your source code to understand real field names and routes, opens a browser, tests everything, and shows you what's broken.
init also:
~/.claude/settings.json, ~/.cursor/mcp.json, and so on) so the tools are available in every project, every session.env, vite.config, framework defaults)VIBE.md (edit with your test credentials) and vibe.config.jsonInstall this repository as a plugin when your agent supports the Agent Plugins standard. It includes the portable release-qa skill and starts the published npm MCP server with no API key.
Cursor: submit or install https://github.com/AishwaryShrivastav/vibe-testing as an Agent Plugin.
Claude Code:
Other compatible agents: load the repository root containing plugin.json, skills/, and mcp.json.
The direct MCP and CLI setup below remains available for editors without plugin support.
Claude Desktop and other MCPB hosts can install one local bundle. Build and validate it from this repository:
The artifact is written to artifacts/vibe-testing-<version>.mcpb. See the MCPB distribution guide for the bundle contents, validation steps, Smithery handoff, and the one-time Playwright Chromium prerequisite.
Detects and configures all installed editors. Done.
Add to ~/.claude/settings.json (global, works in every project):
Or add to .mcp.json in your project root (project-level only):
Add to ~/.cursor/mcp.json (global) or .cursor/mcp.json (project):
Add to ~/.codeium/windsurf/mcp_config.json:
Add to .vscode/mcp.json in your project:
Add to .roo/mcp.json:
14 tools available to your AI editor after setup:
| Tool | When to call | Returns |
|---|---|---|
configure | Start here on a new project. Detects framework, server, and authentication | Structured setup result and idempotent project files |
scan_codebase | Always first. Reads source code, finds routes/forms/tests/gaps | Routes, forms, coverage map, generated scenarios, route_changes since last scan |
get_context | Before writing test steps. Returns source files for a feature | Actual source code with real field names and selectors |
login | When app requires authentication | Post-login screenshot, token state, API calls observed |
scan_page_elements | To see all interactive elements on a page | Element list with selectors plus page screenshot |
explore_page | Broad "does everything work?" testing | Interaction results, API calls, errors, screenshot |
execute_scenario | Run specific test steps | Step-by-step logs plus screenshots |
get_coverage | View coverage map and untested routes | Coverage entries, gaps, available scenarios |
suggest_tests | Find coverage gaps after exploration | Prioritized, ready-to-run scenarios with steps |
take_screenshot | Quick visual verification | Screenshot of any URL |
generate_report | Build HTML report (auto-opens) | Report path plus summary |
run_full_test | One-shot: scan, execute, explore, report | Full results plus snapshot_diff vs last run |
run_converge | Iterative testing until thresholds | Summary across all rounds plus snapshot_diff vs last run |
cleanup | Close browsers, free resources | - |
scan_codebase
get_context
login
scan_page_elements / explore_page
execute_scenario
Step actions: navigate, fill, click, select, wait, assert, upload
CLI and MCP use the same scenario runner. An assert step evaluates every
supplied condition:
selector: the target must be visible. CSS, text=, label=, and
placeholder= are supported; the named locators match exactly.value: visible text must contain this string, on the selected element or on
the page body when no selector is supplied.url: the current URL must exactly match this absolute URL or path resolved
against the configured base URL, including query and fragment.For example:
These checks wait up to the step timeout (default 15000 ms per check). A failed
assertion stops the scenario, records a failed step and reason, and returns
status: "fail"; MCP also sets isError: true. Explicit conditions are used as
the verdict after the existing authentication and form API checks. The
description and expected_outcome prose are not executable assertions.
Without selector, value, or url, assert only checks page health: at least
10 characters of visible body text and no nonempty visible error indicators.
This compatibility check does not prove a described business outcome.
upload requires a file-input selector and a nonempty file path in value.
Relative paths resolve against the codebase root. Hidden file inputs are
supported. A missing file, directory, missing target, or target that is not a
file input returns status: "error". The action selects one file using
Playwright; add an assertion of the resulting UI to verify application-side
processing. It does not automate native file-picker dialogs.
Scenarios must contain at least one step. Unknown actions are errors.
take_screenshot
run_full_test
run_converge
Tell your AI editor:
The AI will:
scan_codebase, to understand routes, forms, existing testsget_context("login"), to read the actual login form source codelogin, to authenticate in a real browserexplore_page("/dashboard"), clicking everything and observing what breaksexplore_page("/settings"), samesuggest_tests, to find coverage gapsexecute_scenario x N, running targeted test flowsgenerate_report, HTML report opens automaticallycleanup, closing browsersThe AI will:
scan_codebase (if not already done)get_context("checkout"), reading CheckoutForm.tsx, api/orders/route.ts and so onloginexecute_scenario, filling the real form fields from source codegenerate_reportThe AI will:
login, testing the login flowtake_screenshot, visual confirmation of the post-login stateThe AI will run explore_page on every route, collecting API errors, broken elements, and failed interactions, then suggest_tests with the broken items marked as high priority.
What it creates:
| File | Where | Purpose |
|---|---|---|
.mcp.json | Project root | Claude Code MCP config (project-level) |
~/.claude/settings.json | Global | Claude Code MCP config (all projects) |
.cursor/mcp.json | Project root | Cursor MCP config |
~/.cursor/mcp.json | Global | Cursor MCP config (all projects) |
.cursor/rules/vibe-test.mdc | Project | Cursor rules, alwaysApply: true |
.windsurfrules | Project | Windsurf instructions |
~/.codeium/windsurf/mcp_config.json | Global | Windsurf MCP config (all projects) |
.vscode/mcp.json | Project | VS Code Copilot MCP config |
.github/copilot-instructions.md | Project | GitHub Copilot instructions |
.roo/mcp.json | Project | Roo Code MCP config |
CLAUDE.md | Project | Claude Code session instructions |
AGENTS.md | Project | Universal agent instructions (Codex, Devin, Zed) |
VIBE.md | Project | Test guidance, edit with your credentials |
vibe.config.json | Project | Config, URL auto-detected from your project |
Options:
After init, edit VIBE.md with your login URL and test credentials.
run options| Option | Default | Description |
|---|---|---|
--mode fast|deep | deep | fast: quick scan. deep: full feature extraction plus exploration |
--no-headed | - | Run browser headless (default: visible) |
--codebase <path> | cwd | Path to project root |
--scope <routes...> | all | Test only specific routes |
-c <path> | vibe.config.json | Config file path |
run and converge fail with a nonzero exit when no scenarios are discovered
or an execution batch is empty. MCP run_full_test and run_converge return
isError: true for these cases. Empty runs do not write a successful run
snapshot or replace the prior report.
converge options| Option | Default | Description |
|---|---|---|
--max-rounds <n> | 4 | Max follow-up rounds after baseline |
--target-pass-rate <r> | 0.92 | Stop when pass rate reaches this (0-1) |
--max-gaps <n> | 2 | Stop when critical plus important gaps fall to this |
Create VIBE.md in your project root. vibe-testing reads it automatically on every run.
See VIBE.example.md for the full template.
Created automatically by init with auto-detected URL. Edit as needed:
| Key | Description |
|---|---|
url | App URL, localhost or staging. Auto-detected by init. |
mode | fast (heuristic scan) or deep (full extraction plus exploration) |
auth.strategy | credentials (form login), basic (HTTP Basic Auth), or skip |
auth.login_url | Explicit login route for non-standard paths keyword matching would miss |
auth.credentials | Login credentials, used for login and for generated scenarios, persisted across runs |
never_interact | Text patterns or CSS selectors to skip during exploration |
scope.include | Route patterns to include; default /** includes / and all nested routes. * stays within one segment; a trailing /** includes the subtree root and its descendants. Other characters are literal. |
scope.exclude | Route patterns to exclude from testing |
scope.max_routes | Cap how many routes are tested per run |
scope.seed_routes | Concrete URLs for dynamic-segment routes the parser can't enumerate (e.g. /live/[slug] becomes /live/dev-mode-a-now). Each seeded route inherits requires_auth and the source file from its dynamic parent. |
browser.headed | true = visible browser. CLI default true, MCP server default false (headless) so editor sessions aren't disrupted by pop-up windows. |
browser.slowMo | Milliseconds between actions (useful for debugging) |
routes | auto (default) discovers routes from the codebase. config uses only routes explicitly listed in config. |
| Framework | Routes | API endpoints | Forms |
|---|---|---|---|
| Next.js App Router | yes | yes | yes |
| Next.js Pages Router | yes | yes | yes |
| Next.js (src/ variant) | yes | yes | yes |
| React SPA (react-router) | yes | - | yes |
| Vue + Vite (vue-router) | yes | - | yes |
| Nuxt | yes | yes | yes |
| SvelteKit | yes | yes | yes |
| Express / Fastify | - | yes | yes |
| Monorepos (Turborepo, pnpm, Lerna) | yes | yes | yes |
Existing test files are also read to build a coverage map: Jest, Vitest, Playwright, and Cypress suites are all parsed.
vibe-testing learns across runs and stores state in .vibe/:
[name='email'] worked on /login, uses it next run.vibe/route-manifest.json): every scan diffs against the previous one; new and removed routes surface as route_changes on scan_codebase results so the AI can cover them immediately.vibe/run-snapshot.json): every run captures per-route pass/fail and diffs against the prior run; snapshot_diff flags newly_passing (fixes), newly_failing (regressions), still_failing, plus added and removed routesrun_converge returns the same shape, so iterative runs in your editor highlight what you just broke.
Reset with npx vibe-testing@latest reset to start fresh.
When you ask your editor to "test the login flow", here is what it does:
Does vibe-testing use an LLM internally? No. It uses heuristic verification (URL changes, toast detection, API errors). Your editor's model is the brain: it sees screenshots and decides what to test next. Runs have no API cost.
What's the difference between explore_page and execute_scenario?
explore_page is broad: it clicks every button and input it finds and reports the results. execute_scenario is precise: you give it specific steps and it follows them exactly. Use explore_page to find what's on a page, then execute_scenario to test specific flows.
What's get_context for?
It returns the actual source code for a feature, so the AI knows [name='email'] instead of guessing #email-input. Always call it before writing test steps for a specific feature.
Does it handle SPAs with client-side routing? Yes. Playwright navigates the real browser, so client-side routing (React Router, Vue Router, and the rest) works naturally.
Does it handle login / authentication?
Yes. The login tool fills credentials in a real browser, captures auth tokens from localStorage/cookies, and keeps that session alive for authenticated tests. Credentials are persisted in .vibe/memory/ and reused automatically.
Will it click "Delete Account" or other destructive buttons?
No. Set never_interact in vibe.config.json or VIBE.md to blocklist dangerous actions. Any button whose text or selector matches is skipped during exploration.
Can I use it without an AI editor?
Yes. vibe-test run https://your-app.com runs standalone. It scans, generates scenarios, executes them, and produces an HTML report without needing an editor.
How do I test a staging environment?
Set url in vibe.config.json to your staging URL, or pass it as a CLI argument: npx vibe-testing@latest run https://staging.myapp.com.
Does it work with monorepos?
Yes. init detects Turborepo/pnpm/yarn workspaces and finds the frontend app automatically.
init installs the matching build. To install or repair it directly:
run and converge check for the browser before scanning your project and print the same recovery command when it is missing.
A Node 20 + Chromium image is included for environments that prefer container-based MCP servers:
See CHANGELOG.md for version history. Bug reports and feature requests: GitHub issues.
MIT, Aishwary Shrivastav
io.github.AishwaryShrivastav/vibe-testing at https://registry.modelcontextprotocol.io