Tests AI agents on your web app with Playwright, scoring task completion and blocking forbidden actions.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent โ or use 1-click editor setup below.
We haven't yet run this listing's install command through our automated sandbox check. This isn't a red flag โ we're steadily working through the catalog.
๐ก Paste the JSON block into your client's configuration file under mcpServers, then restart the application.
Inspect callable tools, capabilities, and parameters exposed to AI agents by Deskcert.
The deskcert MCP server evaluates whether an AI agent can perform approved tasks in a web application without attempting actions that the suite forbids. You author the test suite for the application under review rather than using a fixed benchmark environment. This makes it suitable for internal admin panels, dashboards, CRUD tools, and other browser-accessible systems.
Each task identifies the application through a target_url, describes the work the agent should perform, and lists named forbidden_actions. The runner checks task outcomes and applies the forbidden-action policy independently from the numeric score. An attempted forbidden action is intercepted before it reaches the page, recorded with its action name and step number, and causes the suite gate to fail.
The deskcert MCP server is scoped to web applications. It does not provide native desktop automation, VM snapshots, or operating-system-level control for Windows or macOS applications.
DeskCert drives the target site with Playwright. An agent adapter receives a screenshot and an accessibility-tree text dump, then returns the next action. The project includes a scripted adapter that requires no model or API key, while the adapter interface can be connected to an external or in-house agent loop.
The MCP mode is started with deskcert mcp and exposes run_suite over standard input and output. A deployment pipeline or coordinating agent can call that tool instead of launching the command-line runner directly. The same project also provides CLI commands for initializing suites, running them, and applying CI-specific exit codes.
For CI, exit code 0 indicates a pass, 1 indicates that the score is below the configured threshold, and 2 indicates at least one forbidden-action violation. This lets a pipeline distinguish an insufficient score from an agent attempting an explicitly blocked operation. The score reflects only the tasks and guardrails defined in the suite; it is not a general certification of the agent.
Install either package before using the deskcert MCP server:
npm install -g deskcert-clipip install deskcert-cliInstall the Chromium browser used by Playwright with npx playwright install chromium for the npm installation or playwright install chromium for the Python installation.
Run deskcert init to create an example suite and fixture application, or provide your own suite directory with --suite. A suite can target a local fixture, staging system, or internal application reachable through the relevant network. The schema accepts only http:// and https:// target URLs; file:// and javascript: URLs are rejected.
The Python package also includes deskcert serve-fixture, while the npm package can run the bundled fixture server directly with Node. Both package implementations expose the same init, run, ci, and mcp command surface and are expected to produce equivalent scoring results.
The deskcert MCP server provides the run_suite MCP tool over stdio. Its surrounding CLI supports these related capabilities:
init.run.ci.mcp.Suites are written in YAML and validated against the included JSON Schema. The browser runner can use the bundled scripted adapter for repeatable initial tests or CI self-tests, or an implementation of the two-method AgentAdapter interface for another agent framework.
Testing is limited to browser-driven web applications. The project does not claim to assess an agent across arbitrary desktop software or a complete operating-system environment. Results depend on the tasks, success checks, forbidden actions, and threshold selected in the suite.
A successful numeric score does not override a forbidden-action violation. Conversely, the suite score is not a broad safety rating: it describes the specific application and scenarios that were executed. The MCP transport documented here is stdio, so the calling client must be able to launch and communicate with a local MCP process.
Factual signals from GitHub, npm, and our automated checks โ not a rating.
No reviews yet โ be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/deskcert)<a href="https://allmcps.com/mcp/deskcert"><img src="https://allmcps.com/api/badge/deskcert?style=directory" alt="Deskcert on AllMCPs" /></a>