# slop-eval [Health: Active]

**Category:** 💻 Developer Tools  
**Repository:** https://github.com/RudrenduPaul/slop-eval  
**GitHub Stars:** 0  
**npm Downloads (last month):** 188  
**Views:** 1  
**Installs:** 0  
**Upvotes:** 0  
**Directory Page:** https://allmcps.com/mcp/slop-eval

## Description
Wraps the slop-eval CLI as a single generic MCP tool for genericness scoring of AI UI output.

## Claude Desktop Quick Installation
Install path detected from listing signals. Uses `npx` (confidence: high):

```json
"mcpServers": {
  "slop-eval": {
    "command": "npx",
    "args": ["-y","slop-eval-cli"],
    "env": {
      "ANTHROPIC_API_KEY": ""
    }
  }
}
```

**Requires environment variables:** `ANTHROPIC_API_KEY` — the values above are empty placeholders; fill in real credentials before running (see the repository for what each one is for).

## Documentation

## What the slop-eval MCP server does

The slop-eval MCP server turns the slop-eval CLI into one generic MCP scoring tool. Its purpose is to evaluate whether an AI-generated interface looks generic, rather than to verify functional correctness or provide a formal design certification. The underlying scorer produces a composite result with rubric metadata, findings, a summary, and a disclaimer.

The rubric currently covers three areas: layout novelty, visual-identity distinctiveness, and component-pattern novelty. Each category receives a score from 0 to 10. Findings must include specific cited evidence; an uncited finding is treated as invalid by the scoring design.

## How it works

For screenshot input, slop-eval sends the rendered image to an Anthropic LLM judge. The judge returns structured data through a forced tool-call schema rather than an unstructured conversational response. The resulting output includes individual findings and a composite score based on the versioned v1 rubric.

The CLI also supports a URL fallback. In that mode, the tool fetches raw HTML and text because the project does not bundle a headless browser. The judge therefore reasons about markup and copy instead of seeing the rendered layout. Screenshot and URL inputs are mutually exclusive.

Repeated evaluations of unchanged input can use content-hash caching. This skips another judge request when the input bytes have not changed. The project also defines a source interface for additional scoring sources; the current implemented scoring source is the LLM judge, while screenshot-diff support is documented as not scored.

## Setup and configuration

Running the underlying package requires Node.js 18 or later for the npm distribution, or Python 3.9 or later for the PyPI distribution. The required environment variable is `ANTHROPIC_API_KEY`, supplied by the user. There is no shared or built-in key.

The CLI can be run with `npx slop-eval-cli score --screenshot ./preview.png --json`, or installed and built from the repository. JSON mode is intended for CI and agent workflows: it emits a parseable object on successful runs and error paths. The CLI uses exit code 0 for success, 1 when a configured threshold is not met, and 2 for usage or unrecoverable errors.

## Tools and capabilities

The slop-eval MCP server is described as a single generic MCP tool backed by the CLI. The underlying scoring workflow supports:

- Screenshot-based UI evaluation.
- URL-based evaluation using fetched HTML and text.
- Structured JSON output for automation.
- A versioned rubric with three scoring categories.
- Content-hash caching for unchanged inputs.
- Threshold-based failure through the CLI's `--fail-below` option.
- Programmatic library APIs in the TypeScript and Python distributions.

## Limitations and notes

The score is a heuristic signal from an LLM judge, not a certification. Results depend on the supplied input and the judge's assessment. URL mode is not equivalent to screenshot mode because it does not include a bundled browser or a rendered visual capture.

The tool requires access to Anthropic through the caller's API key, so it is not an offline evaluator. The repository describes one implemented scoring source and a screenshot-diff source that remains `not_scored`. The MCP material provided here does not identify a transport, MCP client compatibility list, or a separate tool name beyond the generic MCP tool description.

_Full upstream README: https://allmcps.com/mcp/slop-eval/readme_

