Scores AI-generated UI for genericness through an MCP tool backed by the slop-eval CLI and Anthropic's LLM judge.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
We haven't yet run this listing's install command through our automated sandbox check. This isn't a red flag β we're steadily working through the catalog.
π‘ Paste the JSON block into your client's configuration file under mcpServers, then restart the application.
Inspect callable tools, capabilities, and parameters exposed to AI agents by Slop Eval.
The slop-eval MCP server turns the slop-eval CLI into one generic MCP scoring tool. Its purpose is to evaluate whether an AI-generated interface looks generic, rather than to verify functional correctness or provide a formal design certification. The underlying scorer produces a composite result with rubric metadata, findings, a summary, and a disclaimer.
The rubric currently covers three areas: layout novelty, visual-identity distinctiveness, and component-pattern novelty. Each category receives a score from 0 to 10. Findings must include specific cited evidence; an uncited finding is treated as invalid by the scoring design.
For screenshot input, slop-eval sends the rendered image to an Anthropic LLM judge. The judge returns structured data through a forced tool-call schema rather than an unstructured conversational response. The resulting output includes individual findings and a composite score based on the versioned v1 rubric.
The CLI also supports a URL fallback. In that mode, the tool fetches raw HTML and text because the project does not bundle a headless browser. The judge therefore reasons about markup and copy instead of seeing the rendered layout. Screenshot and URL inputs are mutually exclusive.
Repeated evaluations of unchanged input can use content-hash caching. This skips another judge request when the input bytes have not changed. The project also defines a source interface for additional scoring sources; the current implemented scoring source is the LLM judge, while screenshot-diff support is documented as not scored.
Running the underlying package requires Node.js 18 or later for the npm distribution, or Python 3.9 or later for the PyPI distribution. The required environment variable is ANTHROPIC_API_KEY, supplied by the user. There is no shared or built-in key.
The CLI can be run with npx slop-eval-cli score --screenshot ./preview.png --json, or installed and built from the repository. JSON mode is intended for CI and agent workflows: it emits a parseable object on successful runs and error paths. The CLI uses exit code 0 for success, 1 when a configured threshold is not met, and 2 for usage or unrecoverable errors.
The slop-eval MCP server is described as a single generic MCP tool backed by the CLI. The underlying scoring workflow supports:
--fail-below option.The score is a heuristic signal from an LLM judge, not a certification. Results depend on the supplied input and the judge's assessment. URL mode is not equivalent to screenshot mode because it does not include a bundled browser or a rendered visual capture.
The tool requires access to Anthropic through the caller's API key, so it is not an offline evaluator. The repository describes one implemented scoring source and a screenshot-diff source that remains not_scored. The MCP material provided here does not identify a transport, MCP client compatibility list, or a separate tool name beyond the generic MCP tool description.
Factual signals from GitHub, npm, and our automated checks β not a rating.
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/slop-eval)<a href="https://allmcps.com/mcp/slop-eval"><img src="https://allmcps.com/api/badge/slop-eval?style=directory" alt="Slop Eval on AllMCPs" /></a>