The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Dsh Cua Windows Computer Use listing page.
English · 中文
An MCP server + agent skill for computer use on Windows: accessibility element actions come first and screenshots are only the fallback. It ships a cross-session arbiter — when several agents share one machine it serializes them, and it yields while you are actually using the computer yourself.
This repository contains only the MCP server and the skill. It stays neutral toward any stdio MCP client (dsh / Claude Code / Codex / Cursor / Cline / ZCode …) — nothing here requires dsh.
Platform semantics (0.3.1 and later): the tools only work on Windows — they drive
user32/kernel32 and UI Automation. But the package also imports elsewhere and the server
starts there, answering tools/list as usual, so any client or directory crawler can
enumerate all 19 tools with their full schemas; actually calling a tool returns a clear
"requires Windows" error rather than the process failing to start at all. In 0.3.0 the import
itself raised, which made such crawlers unable to see the server at all
(tests/linux-handshake.py is the regression test for this property, and CI runs it on
ubuntu-latest).
A stdio MCP server exposing 19 tools. Every tool name is prefixed tool_, exactly as
tools/list returns it.
tool_skyshot (reads a window as a compact,
diffable text tree — three orders of magnitude smaller than a screenshot), tool_element_at_point,
tool_read_element, tool_find_elements, tool_capture_window (DPI-aware, cropped to the client area),
tool_list_windows / tool_find_window / tool_get_window_rect, tool_list_displays, tool_cursor_position,
tool_clipboard_read, tool_coexistence_statustool_element_action / tool_element_action_at — press / set_value / select / toggle / expand /
collapse / scroll_into_view / focus, delivered straight to the UIA element, so they
never steal focus and never care about z-ordertool_type_text (targeted PostMessage with effect_verified), tool_clipboard_write,
tool_open_application. They
synthesize no physical input, so they do not wait for you to stop working — and a clipboard
write still destroys whatever you last copied, so announce it when you do.tool_click_at's raw_event path and
tool_send_keys' global hotkeys. The gate waits for the machine to go input-quiet, then refuses
with user-active rather than fight you for the cursor. tool_click_at tries its element path
first (ax_press), which injects no physical input and therefore takes the mutex only; the
receipt's method field says which path actually ran."Read-only" here means it takes no mutating action and synthesizes no input, so it is safe to
call while someone is using the machine. Two of them have a side effect worth knowing:
tool_capture_window writes the screenshot to disk (save_path; a temp file when omitted), and
tool_skyshot updates the server-side diff baseline it diffs the next shot against.
A small tree is not proof of an empty window. tool_skyshot and tool_find_elements report
minimized, because a minimized window may be showing "nothing is displayed" rather than "nothing
is there" — and how much it hides depends on the application, which is why the receipt reports the
state instead of guessing at the cause. Measured here: an Edge window minimized before anything read
it exposes its browser chrome and no page at all (and include_offscreen does not recover it),
Explorer falls from 40 elements to 8 whenever it is minimized, and Notepad is unaffected. Restore the
window once and read it again if the result matters.
Every action returns a receipt rather than a self-reported success: action_sent /
effect_verified / foreground_changed / user-active / arbiter-busy. "The call was
accepted" and "the effect happened" are two different things, and the tool separates them for
the agent.
There are already several mature open-source Windows implementations. dsh-cua's differences are concentrated on one thing: sharing a machine with a human.
| dsh-cua | cua-driver | ahk-mcp | lean-computer-use-mcp | |
|---|---|---|---|---|
| Element actions delivered as UIA patterns (no focus steal, z-order irrelevant) | ✅ | ✅ (ax mode) | ❌ reads via UIA, acts by coordinate click | via cua-driver |
| Recent human input → refuse | ✅ user-active | ❌ | ❌ | ❌ |
| Cross-agent serialization (multi-process) | ✅ named mutex | ❌ | ❌ | ❌ |
| Per-action effect assertion | ✅ three-state effect_verified | reports a delivery tier | ❌ | ❌ state_changed heuristic only |
| Foreground-steal side effect measured | ✅ foreground_changed | ❌ | ❌ | ❌ |
| Tool count | 19 | 59 | 15 | 6 |
The key distinction is two things that are routinely conflated:
PostMessage, so the cursor and keyboard focus are physically never touched. cua-driver has
it (ax mode). ahk-mcp does not, and the distinction is narrower than "no UIA": it reads
through UIA (ahk_uia_tree / ahk_uia_find / ahk_uia_url), but it has no UIA pattern
action — per its README it acts with coordinate clicks or synthetic keys, so an action does
move the real cursor.GetLastInputInfo, waits when it sees you using the machine, and on timeout
refuses (user-active) instead of barging in. As of 2026-09-25 a pattern search across
the other three codebases in that table found no equivalent — that is search evidence, not
proof, and it covers input-age detection only: cua-driver does have human-facing guards of a
different kind (a consent requirement, and foreground-steal detection with restore).effect_verified is likewise something the alternatives lack: it splits "the call was accepted"
from "the effect happened" and gives three states (true changed as expected / false accepted
but unchanged, downgraded to a failure / null no comparable state, i.e. unconfirmed). The
usual alternative is to re-observe once after the action and leave the judgement to the model.
What dsh-cua does not do (stated up front to avoid misunderstanding): no grounding of its
own — the server does not analyse pixels, so a text-only model cannot drive interfaces that
a tree cannot express (canvas, games, remote desktop). With a vision-capable model the pixel
path is supported end to end: capture_window returns the image together with a verified
image→screen mapping (bounds, scale, dpi_verified), and the model supplies the grounding.
Also not provided: record-and-replay, and an isolation sandbox. There are better-suited tools
for those.
Why not run the agent on a second desktop or a virtual display, so it never touches mine?
Because Windows has nothing to build that on, and the mobile design that does work rests on
exactly the missing piece. On Android an app can create a VirtualDisplay and address input at
it — an input event carries a display id, so the agent's taps are routed to its own screen and
the human's touchscreen never notices. Windows routes input per desktop, not per display: a
desktop has one input queue and one cursor position.
That makes the obvious analogues dead ends:
| Idea | Why it does not isolate |
|---|---|
| Add a virtual monitor (an indirect display driver) | Another canvas, not another cursor — the pointer still has a single position across all monitors |
A Windows virtual desktop (Win+Ctrl+D) | A view switch inside the same session: same input queue, same cursor |
| A second session (RDP, or another user) | Isolation is real, but client Windows allows one interactive session per user at a time — connecting remotely locks the console, so the human loses their screen, which was the whole point |
A hidden Win32 desktop (CreateDesktop) | The agent would get its own input queue and cursor, but its windows are invisible, so it can only drive instances it launched — not the program you are looking at. (Inferred from the window-station/desktop model; not measured here.) |
That table is about input, and it is worth being exact about what a second display does
buy, because the answer is not "nothing". Measured on Edge with an isolated profile
(tests/verify-visible-vs-foreground.py): Chromium builds a page's accessibility tree for a
window that is merely shown. A window minimized before any query reported no page; the
same window reported a named page as soon as it was shown — while not the active window
(GetGUIThreadInfo(hwndActive) false) and fully covered (0% of its area on top). Every state
was read back from the OS rather than assumed, and a ForegroundWatch around each walk (1.4-2.7M
samples, foreground unchanged) shows the read takes no focus at all. A browser window parked on a
second display is therefore readable with skyshot while your focus never leaves your screen.
Input does not survive the move. A WM_CHAR posted to that same window is accepted by
PostMessage and changes nothing — covered or uncovered — while the identical post to the same
window when it is active inserts the character. So a second display buys the read and not
the act, and the surfaces that need acting on (Chromium internals, canvas, games) are exactly
the ones needing the activation it cannot provide. The constraint is input, and for input there
is still one cursor and one foreground window per desktop.
One qualification, because "my browser is on a second display and works fine" is a true report of
a different path. All of the above is about OS-level injection, and it does not generalise to
the browser's own protocol. Measured on the same machine, with a headed Edge window driven over CDP
(Input.dispatchMouseEvent, which is what Playwright's page.mouse.click sends): the click arrived
as a trusted event, the system cursor did not move by a single pixel (GetCursorPos identical
before and after), the foreground window did not change, and a further click still landed after
the window had been minimized (IsIconic true, foreground on another window). CDP injects into
the browser's own input pipeline, above the OS input queue, so the one-cursor/one-foreground rule
does not reach it at all — and neither does it reach anything else addressed through an
application's own API instead of through the screen. (The page's document.visibilityState still
read visible while minimized; that part is an artefact of the flags Playwright launches Chromium
with, not general Chromium behaviour. Delivering the click is a protocol property, independent of it.)
What decides the outcome is which layer the input enters at:
| Route | Cursor / foreground needed | Crosses a display or occlusion boundary |
|---|---|---|
| An application's own protocol (CDP, Playwright, a DOM/JS event, an app API) | no — nothing OS-level is injected | yes, and the window need not even be visible |
element_action (a UIA pattern delivered to an element) | no — but the window must exist and expose an element | yes |
PostMessage to a window (type_text) | measured on Chromium: yes — the target must be active for the text to appear | no |
Physical injection (SendInput, the raw-event path of click_at, send_keys) | yes — one cursor and one foreground window for the whole desktop | no |
So for a page, "put the browser on the virtual display" is a good answer, and this project does not compete with it: a page is already its own addressable, scriptable surface. dsh-cua is for the targets that have no such surface — Explorer, native dialogs, Office, canvas and games, applications with no scripting API — where OS input is the only route there is, and that route is exactly the one a second display does not help. Both statements are true.
So the conflict is not a gap in this implementation, it is an OS constraint: Windows has no second cursor. Given that, yielding is the only correct response — and most calls never get near the problem, because they inject no input at all:
element_action / element_action_at — a UIA pattern is delivered to the element: no physical
input, no cursor movement, no focus change. These are the paths that "can run while the user
types".type_text — a window-targeted PostMessage, not global keystrokes. It resolves the text
control inside the window first (a top-level window does not forward WM_CHAR to its child
edit) and reads that control back, so the receipt carries effect_verified rather than only
reporting that something was posted.skyshot, element_at_point, capture_window, …) — touches nothing.Exactly two paths inject physical input, and they exist because canvas-, game- and
Chromium-internal surfaces expose no element to address: the raw-event path of tool_click_at,
and the global hotkeys of tool_send_keys. Those two are what the arbiter guards.
You need Windows x64 + an interactive desktop session + Python ≥3.10 to actually drive a
desktop. (The package installs and starts on Linux/macOS too, tools/list answers normally,
and a tool call then reports "requires Windows" — see "platform semantics" above.)
However you install it, start the server with python -m dsh_cua:
Why the README does not say
dsh-cua-server: pip installs console scripts into the interpreter'sScriptsdirectory, and that directory is not necessarily on PATH — measured on a stock python.org 3.12 install, neither the User nor the Machine PATH contained it, sopip install dsh-cuasucceeded whiledsh-cua-serverreported command not found.python -mneeds no PATH entry at all. The console script is still shipped and works when PATH does contain it.Options 1 and 2 both work today: the package is published on PyPI (https://pypi.org/project/dsh-cua/). If
uvx/pipever 404s, use option 3 — it always works.
Any MCP client; name the server win32 (the skill's tool-name convention is
mcp__win32__*).
python -m (no PATH dependency, recommended):
uvx:
More shapes are in examples/: Claude Code / generic clients / a dsh
cordis.patch.yml fragment / the route modality declaration you need if you want the model to
read screenshots (tr-route-settings.yml).
skill/computer-use/SKILL.md is the companion doctrine for using
these tools: the observe → locate → act → verify loop, receipt semantics, retry safety, and the
discipline of coexisting with a human. The model can use the tools without it, but with it the
model picks the right path by itself — the measured difference is large.
The skill lives in this repository, not in the package — uvx and pip do not put a skill/
directory on your disk, so the cp below only works from a checkout. Without one, fetch the file:
From a clone, copy the directory instead:
| Tier | Operations | Gate |
|---|---|---|
| Read-only | the 12 observe tools | no gate, callable at any time |
| Soft | element actions, PostMessage typing, clipboard write, launching applications | cross-agent mutex (named mutex, multi-process, automatic serialization) |
| Hard | raw clicks, global hotkeys | mutex + GetLastInputInfo yielding: if the user typed recently it waits, and on timeout refuses with user-active instead of stealing the cursor |
Honest boundaries: yielding is a cooperation protocol, not a hard guarantee (the tight check
150 ms before injection narrows the window as much as possible); a few applications
self-activate even on set_value (the receipt reports foreground_changed truthfully); and
two operators on the same window has no technical solution — do not drive the same window the
agent is driving.
The tests need no human cooperation — "user input" is synthesized with one real 1-pixel cursor move, and the cursor is restored afterwards.
The tests need a real interactive desktop session (some checks create windows and address
them through UIA), so they cannot run on a GitHub-hosted runner. What CI does cover is the
part that needs no desktop: packaging and installation, module import, regressions for the diff
index and tree-line escaping, and the arbiter's decision logic — see
.github/workflows/ci.yml.
MIT