The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Dsh Verify listing page.
中文 | English
Witness — Check the behavior your browser spec describes. Browser assertions can expose some implementation gaps; they do not prove an app is complete or independently reviewed. (Witness is the product name;
dsh-verifyis the package name — same thing.)
If Witness catches something for you, ⭐ star the repo — it's how this project stays alive.
You asked an AI to build a web app. It said "done." Does it actually work?
dsh-verify opens a real browser and checks — so you never have to take the agent's word for it.
Requires Node.js 20+ and a one-time Chromium download. Install the current npm release with npm install --save-dev dsh-verify@0.9.6, then run npx dsh-verify init. Alternatively, run from a clone checkout:
Open either HTML report with its adjacent PNG file. The failing spec deliberately expects wrong text; inspect the assertion and screenshot. Existing missing-CSS failure report is another inspectable example.
init [directory] only writes a local fixture and two JSON specs. It never calls an LLM, launches a browser, or downloads anything. The destination must be new (its parent must exist); existing directories, files, and symlinks are refused without changes. No force/overwrite option. GitHub package and browser installation require internet; these starter checks use localhost only. gen and MCP generate_and_verify are optional, separate features that require an LLM key.
To check your own app, edit the JSON steps; keep serve for static files or use base for an existing URL. serve paths resolve from the command's current directory, not the spec file. A PASS applies only to the supplied checks in that run; it does not establish complete coverage, security, accessibility, or independent authorship. The app and spec may come from the same agent; review the spec separately when independence matters. Reports retain assertion details and explicitly requested screenshots, not a complete audit trail.


The quality gate for agent-built web apps. Works with any agent — DeepSeek Harness (dsh), Claude Code, Cursor, Copilot, Codex — and with any CI. You write what a human would check in a browser; a real browser executes it and returns a PASS/FAIL verdict with receipts (screenshots + diff images).
No LLM judges the outcome. The browser is the judge.

Same task. Same AI. Two builds. One missing CSS rule — the agent's self-review passed, a real browser caught it.
We ran a 4-agent web team (spec writer → frontend dev → QA → reviewer). Their own review said:
✅ "All requirements met. No issues found."
In a real browser, the dark-mode toggle did nothing — the .dark class was toggled, but the CSS rule was never written. Every agent self-test passed because there was nothing in the page for the agents to run. No one opened a real browser.
That's the gap: agents verify against what they believe they built, not against what a user actually experiences. These checks did not exercise the rendered CSS behavior; a browser assertion can cover it.
| Build | What the agents said | What a real browser says |
|---|---|---|
demo/buggy | "No issues found" | ❌ FAIL — background never changes |
demo/fixed | one CSS rule added | ✅ PASS — theme flips |
Same page. Same JS. One missing CSS rule. Two different verdicts.
| Entry point | What it's for | One-liner |
|---|---|---|
| MCP server | Your AI agent verifies its own deliverable, mid-session | claude mcp add dsh-verify -- npx -y -p dsh-verify dsh-verify-mcp |
| CLI | You or your CI verify a build/URL | npx dsh-verify init |
| GitHub Action | Every push runs real-browser checks | uses: 263311487-ux/dsh-verify/.github/actions/dsh-verify@main |
Then tell your agent, in plain words:
Verify http://localhost:3000 — click
#dark-toggle, then checkbodybackground-color changed. Screenshot it.
Tools exposed: verify_spec (run a spec JSON), verify_url (inline checks, no files), generate_and_verify (the AI drafts the checklist, real Chromium executes it), health.
The repo dogfoods it: the dogfood workflow asserts the fixed build passes and the buggy build fails on every push.
--json is available for CI. Exit 1 alone does not distinguish assertion failures from runtime errors; inspect the structured verdict and failed actions as the release checker does.expect_screenshot), refresh with --update-baselines.dsh-verify gen --url ... --prompt "..." learns the page in a real browser, has an LLM draft the checklist, then executes it deterministically. The AI drafts; it never judges.chromium | firefox | webkit per spec or --browser.Top-level fields: title, serve (static dir) or base (target URL), browser, steps. Run many at once with a glob; exit is 0 only if all pass.
An HTML report with adjacent image files — every step with a pass/fail badge, selector, and detail, plus screenshots:

Real-browser benchmark for agent-built web apps: same 3 tasks, same human checks, open entry. Run your model on the board in ~10 minutes:
Submit the result files for maintainer review before they appear on the leaderboard: agent-arena. Full rules in docs/ARENA.md.
The repo's own CI runs exactly that — engine self-tests, then asserts fixed passes and buggy fails — so the tool verifies itself on every push.
Historical snapshot, not a current model ranking. The committed records span 2026-08-17–18 UTC: 48 records, 44 PASS, 2 FAIL, 2 ERROR. Among the 46 records with a browser verdict, 44 passed and 2 failed; the two ERROR records are runner failures. The records cover two model labels × two strategies × three static tasks (four records per cell). Self-check allows up to two repair rounds, so compute budgets differ. Model labels are recorded identifiers, not independently verified identities.
Records do not pin historical verifier/dependency versions. Generated apps, model replies, and local report files are not committed, limiting exact reproduction. No new LLM runs were made for the 0.9.5 or 0.9.6 release. See methodology and limitations.
See docs/ARENA.md — methodology, the tasks, and how to run your own agent.
The following static image is illustrative only. It is not bound to a particular app, commit, spec, or successful run:
For a result tied to a specific build, publish the commit, spec, action run, verdict, and report together. See docs/verified-badge.md.
MIT