Windows computer-use MCP server: accessibility-first actions, text-tree UI, yields to the human
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
One-click editor setup isnβt available for this listing yet β we donβt have a confirmed install command, and weβd rather show nothing than point your editor at the wrong package or host. Follow the projectβs own setup instructions, linked above.
English Β· δΈζ
An MCP server + agent skill for computer use on Windows: accessibility element actions come first and screenshots are only the fallback. It ships a cross-session arbiter β when several agents share one machine it serializes them, and it yields while you are actually using the computer yourself.
This repository contains only the MCP server and the skill. It stays neutral toward any stdio MCP client (dsh / Claude Code / Codex / Cursor / Cline / ZCode β¦) β nothing here requires dsh.
Platform semantics (0.3.1 and later): the tools only work on Windows β they drive
user32/kernel32 and UI Automation. But the package also imports elsewhere and the server
starts there, answering tools/list as usual, so any client or directory crawler can
enumerate all 19 tools with their full schemas; actually calling a tool returns a clear
"requires Windows" error rather than the process failing to start at all. In 0.3.0 the import
itself raised, which made such crawlers unable to see the server at all
(tests/linux-handshake.py is the regression test for this property, and CI runs it on
ubuntu-latest).
A stdio MCP server exposing 19 tools. Every tool name is prefixed tool_, exactly as
tools/list returns it.
tool_skyshot (reads a window as a compact,
diffable text tree β three orders of magnitude smaller than a screenshot), tool_element_at_point,
tool_read_element, tool_find_elements, tool_capture_window (DPI-aware, cropped to the client area),
tool_list_windows / tool_find_window / tool_get_window_rect, tool_list_displays, tool_cursor_position,
tool_clipboard_read, tool_coexistence_statustool_element_action / tool_element_action_at β press / set_value / select / toggle / expand /
collapse / scroll_into_view / focus, delivered straight to the UIA element, so they
never steal focus and never care about z-ordertool_type_text (targeted PostMessage with effect_verified), tool_clipboard_write,
tool_open_application. They
synthesize no physical input, so they do not wait for you to stop working β and a clipboard
write still destroys whatever you last copied, so announce it when you do.tool_click_at's raw_event path and
tool_send_keys' global hotkeys. The gate waits for the machine to go input-quiet, then refuses
with user-active rather than fight you for the cursor. tool_click_at tries its element path
first (ax_press), which injects no physical input and therefore takes the mutex only; the
receipt's method field says which path actually ran."Read-only" here means it takes no mutating action and synthesizes no input, so it is safe to
call while someone is using the machine. Two of them have a side effect worth knowing:
tool_capture_window writes the screenshot to disk (save_path; a temp file when omitted), and
tool_skyshot updates the server-side diff baseline it diffs the next shot against.
A small tree is not proof of an empty window. tool_skyshot and tool_find_elements report
minimized, because a minimized window may be showing "nothing is displayed" rather than "nothing
is there" β and how much it hides depends on the application, which is why the receipt reports the
state instead of guessing at the cause. Measured here: an Edge window minimized before anything read
it exposes its browser chrome and no page at all (and include_offscreen does not recover it),
Explorer falls from 40 elements to 8 whenever it is minimized, and Notepad is unaffected. Restore the
window once and read it again if the result matters.
Every action returns a receipt rather than a self-reported success: action_sent /
effect_verified / foreground_changed / user-active / arbiter-busy. "The call was
accepted" and "the effect happened" are two different things, and the tool separates them for
the agent.
There are already several mature open-source Windows implementations. dsh-cua's differences are concentrated on one thing: sharing a machine with a human.
| dsh-cua | cua-driver | ahk-mcp | lean-computer-use-mcp | |
|---|---|---|---|---|
| Element actions delivered as UIA patterns (no focus steal, z-order irrelevant) | β | β (ax mode) | β reads via UIA, acts by coordinate click | via cua-driver |
| Recent human input β refuse | β
user-active | β | β | β |
| Cross-agent serialization (multi-process) | β named mutex | β | β | β |
| Per-action effect assertion | β
three-state effect_verified | reports a delivery tier | β | β state_changed heuristic only |
| Foreground-steal side effect measured | β
foreground_changed | β | β | β |
| Tool count | 19 | 59 | 15 | 6 |
The key distinction is two things that are routinely conflated:
PostMessage, so the cursor and keyboard focus are physically never touched. cua-driver has
it (ax mode). ahk-mcp does not, and the distinction is narrower than "no UIA": it reads
through UIA (ahk_uia_tree / ahk_uia_find / ahk_uia_url), but it has no UIA pattern
action β per its README it acts with coordinate clicks or synthetic keys, so an action does
move the real cursor.GetLastInputInfo, waits when it sees you using the machine, and on timeout
refuses (user-active) instead of barging in. As of 2026-09-25 a pattern search across
the other three codebases in that table found no equivalent β that is search evidence, not
proof, and it covers input-age detection only: cua-driver does have human-facing guards of a
different kind (a consent requirement, and foreground-steal detection with restore).effect_verified is likewise something the alternatives lack: it splits "the call was accepted"
from "the effect happened" and gives three states (true changed as expected / false accepted
but unchanged, downgraded to a failure / null no comparable state, i.e. unconfirmed). The
usual alternative is to re-observe once after the action and leave the judgement to the model.
What dsh-cua does not do (stated up front to avoid misunderstanding): no grounding of its
own β the server does not analyse pixels, so a text-only model cannot drive interfaces that
a tree cannot express (canvas, games, remote desktop). With a vision-capable model the pixel
path is supported end to end: capture_window returns the image together with a verified
imageβscreen mapping (bounds, scale, dpi_verified), and the model supplies the grounding.
Also not provided: record-and-replay, and an isolation sandbox. There are better-suited tools
for those.
Why not run the agent on a second desktop or a virtual display, so it never touches mine?
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/dsh-cua-windows-computer-use)<a href="https://allmcps.com/mcp/dsh-cua-windows-computer-use"><img src="https://allmcps.com/api/badge/dsh-cua-windows-computer-use?style=directory" alt="Dsh Cua Windows Computer Use on AllMCPs" /></a>