# dsh-cua - Windows Computer Use

**Category:** 💻 Developer Tools  
**Repository:** https://github.com/Hutusion/dsh-cua  
**Views:** 0  
**Installs:** 0  
**Upvotes:** 0  
**Directory Page:** https://allmcps.com/mcp/dsh-cua-windows-computer-use

## Description
Windows computer-use MCP server: accessibility-first actions, text-tree UI, yields to the human

## Claude Desktop Quick Installation
Heuristic fallback — verify the package name and runner against the repository README before running it. Uses `npx` (confidence: low):

```json
"mcpServers": {
  "dsh-cua-windows-computer-use": {
    "command": "npx",
    "args": ["-y","dsh-cua-windows-computer-use"]
  }
}
```

## Documentation & README

# dsh-cua

[![ci](https://github.com/Hutusion/dsh-cua/actions/workflows/ci.yml/badge.svg)](https://github.com/Hutusion/dsh-cua/actions/workflows/ci.yml)
[![PyPI](https://img.shields.io/pypi/v/dsh-cua)](https://pypi.org/project/dsh-cua/)
[![Hutusion/dsh-cua MCP server](https://glama.ai/mcp/servers/Hutusion/dsh-cua/badges/score.svg)](https://glama.ai/mcp/servers/Hutusion/dsh-cua)
[![LINUX DO](https://img.shields.io/badge/LINUX-DO-FFB003.svg?logo=data:image/svg%2bxml;base64,DQo8c3ZnIHhtbG5zPSJodHRwOi8vd3d3LnczLm9yZy8yMDAwL3N2ZyIgd2lkdGg9IjEwMCIgaGVpZ2h0PSIxMDAiPjxwYXRoIGQ9Ik00Ni44Mi0uMDU1aDYuMjVxMjMuOTY5IDIuMDYyIDM4IDIxLjQyNmM1LjI1OCA3LjY3NiA4LjIxNSAxNi4xNTYgOC44NzUgMjUuNDV2Ni4yNXEtMi4wNjQgMjMuOTY4LTIxLjQzIDM4LTExLjUxMiA3Ljg4NS0yNS40NDUgOC44NzRoLTYuMjVxLTIzLjk3LTIuMDY0LTM4LjAwNC0yMS40M1EuOTcxIDY3LjA1Ni0uMDU0IDUzLjE4di02LjQ3M0MxLjM2MiAzMC43ODEgOC41MDMgMTguMTQ4IDIxLjM3IDguODE3IDI5LjA0NyAzLjU2MiAzNy41MjcuNjA0IDQ2LjgyMS0uMDU2IiBzdHlsZT0ic3Ryb2tlOm5vbmU7ZmlsbC1ydWxlOmV2ZW5vZGQ7ZmlsbDojZWNlY2VjO2ZpbGwtb3BhY2l0eToxIi8+PHBhdGggZD0iTTQ3LjI2NiAyLjk1N3EyMi41My0uNjUgMzcuNzc3IDE1LjczOGE0OS43IDQ5LjcgMCAwIDEgNi44NjcgMTAuMTU3cS00MS45NjQuMjIyLTgzLjkzIDAgOS43NS0xOC42MTYgMzAuMDI0LTI0LjM4N2E2MSA2MSAwIDAgMSA5LjI2Mi0xLjUwOCIgc3R5bGU9InN0cm9rZTpub25lO2ZpbGwtcnVsZTpldmVub2RkO2ZpbGw6IzE5MTkxOSIgZmlsbC1vcGFjaXR5PSIxIi8+PHBhdGggZD0iTTcuOTggNzAuOTI2YzI3Ljk3Ny0uMDM1IDU1Ljk1NCAwIDgzLjkzLjExM1E4My40MjYgODcuNDczIDY2LjEzIDk0LjA4NnEtMTguODEgNi41NDQtMzYuODMyLTEuODk4LTE0LjIwMy03LjA5LTIxLjMxNy0yMS4yNjIiIHN0eWxlPSJzdHJva2U6bm9uZTtmaWxsLXJ1bGU6ZXZlbm9kZDtmaWxsOiNmOWFmMDA7ZmlsbC1vcGFjaXR5PSIxIi8+PC9zdmc+)](https://linux.do)

<!-- The LINUX DO badge is not decoration: linux.do's 开源推广 (open-source promotion) route
     requires the project to link back to the community. See
     https://linux.do/t/topic/1776670 -->

<!-- mcp-name: io.github.Hutusion/dsh-cua -->

**English** · [中文](https://github.com/Hutusion/dsh-cua/blob/main/README.zh-CN.md)

<!-- Links in this file are absolute on purpose: PyPI renders this same README at
     https://pypi.org/project/dsh-cua/, where a relative target (README.zh-CN.md,
     examples/, skill/...) resolves against pypi.org and 404s. GitHub needs no help with
     either form. Do not "simplify" these back to relative paths. -->

An MCP server + agent skill for computer use on Windows: **accessibility element actions come
first and screenshots are only the fallback**. It ships a cross-session arbiter — when several
agents share one machine it serializes them, and it yields while you are actually using the
computer yourself.

This repository contains only the MCP server and the skill. It stays **neutral toward any stdio
MCP client** (dsh / Claude Code / Codex / Cursor / Cline / ZCode …) — nothing here requires dsh.

**Platform semantics (0.3.1 and later)**: the tools only work on Windows — they drive
user32/kernel32 and UI Automation. But **the package also imports elsewhere and the server
starts there**, answering `tools/list` as usual, so any client or directory crawler can
enumerate all 19 tools with their full schemas; actually calling a tool returns a clear
"requires Windows" error rather than the process failing to start at all. In 0.3.0 the import
itself raised, which made such crawlers unable to see the server at all
(`tests/linux-handshake.py` is the regression test for this property, and CI runs it on
`ubuntu-latest`).

## What it is

A stdio MCP server exposing 19 tools. Every tool name is prefixed `tool_`, exactly as
`tools/list` returns it.

- **Observe** (read-only, callable at any time): `tool_skyshot` (reads a window as a compact,
  diffable text tree — three orders of magnitude smaller than a screenshot), `tool_element_at_point`,
  `tool_read_element`, `tool_find_elements`, `tool_capture_window` (DPI-aware, cropped to the client area),
  `tool_list_windows` / `tool_find_window` / `tool_get_window_rect`, `tool_list_displays`, `tool_cursor_position`,
  `tool_clipboard_read`, `tool_coexistence_status`
- **Element actions** (soft gate: serialized across agents, no physical input injected):
  `tool_element_action` / `tool_element_action_at` — press / set_value / select / toggle / expand /
  collapse / scroll_into_view / focus, delivered straight to the UIA element, so they
  **never steal focus and never care about z-order**
- **Other mutating calls** (soft gate too: the mutex, but **no** human-contention yield):
  `tool_type_text` (targeted PostMessage with `effect_verified`), `tool_clipboard_write`,
  `tool_open_application`. They
  synthesize no physical input, so they do **not** wait for you to stop working — and a clipboard
  write still destroys whatever you last copied, so announce it when you do.
- **Physical input** (hard gate: serialized across agents **and** yields to the human): exactly two
  things share your one cursor and one keyboard — `tool_click_at`'s **raw_event path** and
  `tool_send_keys`' global hotkeys. The gate waits for the machine to go input-quiet, then refuses
  with `user-active` rather than fight you for the cursor. `tool_click_at` tries its **element path
  first** (`ax_press`), which injects no physical input and therefore takes the mutex only; the
  receipt's `method` field says which path actually ran.

"Read-only" here means it takes no mutating action and synthesizes no input, so it is safe to
call while someone is using the machine. Two of them have a side effect worth knowing:
`tool_capture_window` writes the screenshot to disk (`save_path`; a temp file when omitted), and
`tool_skyshot` updates the server-side diff baseline it diffs the next shot against.

A small tree is not proof of an empty window. `tool_skyshot` and `tool_find_elements` report
`minimized`, because a minimized window may be showing "nothing is displayed" rather than "nothing
is there" — and how much it hides depends on the application, which is why the receipt reports the
state instead of guessing at the cause. Measured here: an Edge window minimized before anything read
it exposes its browser chrome and **no page at all** (and `include_offscreen` does not recover it),
Explorer falls from 40 elements to 8 whenever it is minimized, and Notepad is unaffected. Restore the
window once and read it again if the result matters.

Every action returns a **receipt** rather than a self-reported success: `action_sent` /
`effect_verified` / `foreground_changed` / `user-active` / `arbiter-busy`. "The call was
accepted" and "the effect happened" are two different things, and the tool separates them for
the agent.

## How it differs

There are already several mature open-source Windows implementations. dsh-cua's differences are
concentrated on one thing: **sharing a machine with a human.**

| | dsh-cua | [cua-driver](https://github.com/trycua/cua) | [ahk-mcp](https://github.com/anomalous3/ahk-mcp) | [lean-computer-use-mcp](https://github.com/Kvxw1105/lean-computer-use-mcp) |
|---|---|---|---|---|
| Element actions delivered as UIA patterns (no focus steal, z-order irrelevant) | ✅ | ✅ (ax mode) | ❌ reads via UIA, acts by coordinate click | via cua-driver |
| **Recent human input → refuse** | ✅ `user-active` | ❌ | ❌ | ❌ |
| Cross-agent serialization (multi-process) | ✅ named mutex | ❌ | ❌ | ❌ |
| Per-action effect assertion | ✅ three-state `effect_verified` | reports a delivery tier | ❌ | ❌ `state_changed` heuristic only |
| Foreground-steal side effect measured | ✅ `foreground_changed` | ❌ | ❌ | ❌ |
| Tool count | 19 | 59 | 15 | 6 |

**The key distinction is two things that are routinely conflated:**

- **"No focus steal" is a mechanism guarantee** — either a UIA pattern or a targeted
  `PostMessage`, so the cursor and keyboard focus are physically never touched. cua-driver has
  it (ax mode). **ahk-mcp does not**, and the distinction is narrower than "no UIA": it reads
  through UIA (`ahk_uia_tree` / `ahk_uia_find` / `ahk_uia_url`), but it has no UIA *pattern
  action* — per its README it acts with coordinate clicks or synthetic keys, so an action does
  move the real cursor.
- **"Yield the moment you move" is a timing guarantee** — it reads the age of the human's last
  input via `GetLastInputInfo`, waits when it sees you using the machine, and on timeout
  **refuses** (`user-active`) instead of barging in. As of 2026-09-25 a pattern search across
  the other three codebases in that table found no equivalent — that is search evidence, not
  proof, and it covers input-age detection only: cua-driver does have human-facing guards of a
  different kind (a consent requirement, and foreground-steal detection with restore).

`effect_verified` is likewise something the alternatives lack: it splits "the call was accepted"
from "the effect happened" and gives three states (`true` changed as expected / `false` accepted
but unchanged, downgraded to a failure / `null` no comparable state, i.e. unconfirmed). The
usual alternative is to re-observe once after the action and leave the judgement to the model.

**What dsh-cua does not do** (stated up front to avoid misunderstanding): no grounding of its
own — the server does not analyse pixels, so **a text-only model** cannot drive interfaces that
a tree cannot express (canvas, games, remote desktop). With a **vision-capable model** the pixel
path is supported end to end: `capture_window` returns the image together with a verified
image→screen mapping (`bounds`, `scale`, `dpi_verified`), and the model supplies the grounding.
Also not provided: record-and-replay, and an isolation sandbox. There are better-suited tools
for those.

## FAQ

**Why not run the agent on a second desktop or a virtual display, so it never touches mine?**

Because Windows has nothing to build that on, and the mobile design that does work rests on
exactly the missing piece. On Android an app can create a `VirtualDisplay` and address input at
it — an input event carries a display id, so the agent's taps are routed to its own screen and
the human's touchscreen never notices. **Windows routes input per desktop, not per display:** a
desktop has one input queue and one cursor position.

That makes the obvious analogues dead ends:

| Idea | Why it does not isolate |
|---|---|
| Add a virtual monitor (an indirect display driver) | Another canvas, not another cursor — the pointer still has a single position across all monitors |
| A Windows virtual desktop (`Win+Ctrl+D`) | A view switch inside the same session: same input queue, same cursor |
| A second session (RDP, or another user) | Isolation is real, but client Windows allows one interactive session per user at a time — connecting remotely locks the console, so the human loses their screen, which was the whole point |
| A hidden Win32 desktop (`CreateDesktop`) | The agent would get its own input queue and cursor, but its windows are invisible, so it can only drive instances it launched — not the program you are looking at. (Inferred from the window-station/desktop model; not measured here.) |

That table is about **input**, and it is worth being exact about what a second display *does*
buy, because the answer is not "nothing". Measured on Edge with an isolated profile
(`tests/verify-visible-vs-foreground.py`): Chromium builds a page's accessibility tree for a
window that is merely **shown**. A window minimized before any query reported **no** page; the
same window reported a named page as soon as it was shown — while **not the active window**
(`GetGUIThreadInfo(hwndActive)` false) and **fully covered** (0% of its area on top). Every state
was read back from the OS rather than assumed, and a `ForegroundWatch` around each walk (1.4-2.7M
samples, foreground unchanged) shows the read takes no focus at all. A browser window parked on a
second display is therefore readable with `skyshot` while your focus never leaves your screen.

Input does not survive the move. A `WM_CHAR` posted to that same window is **accepted** by
`PostMessage` and changes nothing — covered or uncovered — while the identical post to the same
window *when it is active* inserts the character. So a second display buys the **read** and not
the **act**, and the surfaces that need acting on (Chromium internals, canvas, games) are exactly
the ones needing the activation it cannot provide. The constraint is input, and for input there
is still one cursor and one foreground window per desktop.

One qualification, because "my browser is on a second display and works fine" is a true report of
a *different* path. All of the above is about **OS-level injection**, and it does not generalise to
the browser's own protocol. Measured on the same machine, with a headed Edge window driven over CDP
(`Input.dispatchMouseEvent`, which is what Playwright's `page.mouse.click` sends): the click arrived
as a **trusted** event, the system cursor did not move by a single pixel (`GetCursorPos` identical
before and after), the foreground window did not change, and a further click still landed **after
the window had been minimized** (`IsIconic` true, foreground on another window). CDP injects into
the browser's own input pipeline, above the OS input queue, so the one-cursor/one-foreground rule
does not reach it at all — and neither does it reach anything else addressed through an
application's own API instead of through the screen. (The page's `document.visibilityState` still
read `visible` while minimized; that part is an artefact of the flags Playwright launches Chromium
with, not general Chromium behaviour. Delivering the click is a protocol property, independent of it.)

What decides the outcome is **which layer the input enters at**:

| Route | Cursor / foreground needed | Crosses a display or occlusion boundary |
|---|---|---|
| An application's own protocol (CDP, Playwright, a DOM/JS event, an app API) | no — nothing OS-level is injected | yes, and the window need not even be visible |
| `element_action` (a UIA pattern delivered to an element) | no — but the window must exist and expose an element | yes |
| `PostMessage` to a window (`type_text`) | measured on Chromium: yes — the target must be **active** for the text to appear | no |
| Physical injection (`SendInput`, the raw-event path of `click_at`, `send_keys`) | yes — one cursor and one foreground window for the whole desktop | no |

So for a **page**, "put the browser on the virtual display" is a good answer, and this project does
not compete with it: a page is already its own addressable, scriptable surface. dsh-cua is for the
targets that have no such surface — Explorer, native dialogs, Office, canvas and games,
applications with no scripting API — where OS input is the only route there is, and that route is
exactly the one a second display does not help. Both statements are true.

So the conflict is not a gap in this implementation, it is an OS constraint: **Windows has no
second cursor.** Given that, yielding is the only correct response — and most calls never get
near the problem, because they inject no input at all:

- `element_action` / `element_action_at` — a UIA pattern is delivered to the element: no physical
  input, no cursor movement, no focus change. These are the paths that "can run while the user
  types".
- `type_text` — a window-targeted `PostMessage`, not global keystrokes. It resolves the text
  control inside the window first (a top-level window does not forward `WM_CHAR` to its child
  edit) and reads that control back, so the receipt carries `effect_verified` rather than only
  reporting that something was posted.
- Every read-only tool (`skyshot`, `element_at_point`, `capture_window`, …) — touches nothing.

Exactly two paths inject physical input, and they exist because canvas-, game- and
Chromium-internal surfaces expose no element to address: the raw-event path of `tool_click_at`,
and the global hotkeys of `tool_send_keys`. Those two are what the arbiter guards.

## Install

You need **Windows x64 + an interactive desktop session + Python ≥3.10** to actually drive a
desktop. (The package installs and starts on Linux/macOS too, `tools/list` answers normally,
and a tool call then reports "requires Windows" — see "platform semantics" above.)

```bash
# Option 1: uvx, zero install (recommended)
uvx dsh-cua                      # runs the stdio MCP server directly

# Option 2: pip
pip install dsh-cua

# Option 3: from source
pip install git+https://github.com/Hutusion/dsh-cua.git
```

However you install it, **start the server with `python -m dsh_cua`**:

```bash
python -m dsh_cua                # depends on no executable being on PATH
```

> **Why the README does not say `dsh-cua-server`**: pip installs console scripts into the
> interpreter's `Scripts` directory, and **that directory is not necessarily on PATH** —
> measured on a stock python.org 3.12 install, neither the User nor the Machine PATH contained
> it, so `pip install dsh-cua` succeeded while `dsh-cua-server` reported command not found.
> `python -m` needs no PATH entry at all. The console script is still shipped and works when
> PATH does contain it.
>
> Options 1 and 2 both work today: the package is published on PyPI
> (<https://pypi.org/project/dsh-cua/>). If `uvx`/`pip` ever 404s, use option 3 — it always
> works.

## Wiring it up

Any MCP client; name the server **`win32`** (the skill's tool-name convention is
`mcp__win32__*`).

**`python -m` (no PATH dependency, recommended)**:

```json
{ "mcpServers": { "win32": { "command": "python", "args": ["-m", "dsh_cua"] } } }
```

**`uvx`**:

```json
{ "mcpServers": { "win32": { "command": "uvx", "args": ["dsh-cua"] } } }
```

More shapes are in [`examples/`](https://github.com/Hutusion/dsh-cua/tree/main/examples): Claude Code / generic clients / a dsh
`cordis.patch.yml` fragment / the route modality declaration you need if you want the model to
read screenshots (`tr-route-settings.yml`).

## Skill (optional but strongly recommended)

[`skill/computer-use/SKILL.md`](https://github.com/Hutusion/dsh-cua/blob/main/skill/computer-use/SKILL.md) is the companion doctrine for using
these tools: the observe → locate → act → verify loop, receipt semantics, retry safety, and the
discipline of coexisting with a human. The model can use the tools without it, but with it the
model **picks the right path by itself** — the measured difference is large.

**The skill lives in this repository, not in the package** — `uvx` and `pip` do not put a `skill/`
directory on your disk, so the `cp` below only works from a checkout. Without one, fetch the file:

```bash
mkdir -p ~/.dsh/skills/computer-use
curl -fsSL https://raw.githubusercontent.com/Hutusion/dsh-cua/main/skill/computer-use/SKILL.md \
  -o ~/.dsh/skills/computer-use/SKILL.md      # use ~/.agents/skills/ for Claude Code
```

From a clone, copy the directory instead:

```bash
# Claude Code / generic agents
cp -r skill/computer-use ~/.agents/skills/
# dsh
cp -r skill/computer-use ~/.dsh/skills/
```

## Security model

| Tier | Operations | Gate |
|---|---|---|
| Read-only | the 12 observe tools | no gate, callable at any time |
| Soft | element actions, PostMessage typing, clipboard write, launching applications | cross-agent mutex (named mutex, multi-process, automatic serialization) |
| Hard | raw clicks, global hotkeys | mutex + `GetLastInputInfo` yielding: if the user typed recently it waits, and on timeout refuses with `user-active` instead of stealing the cursor |

Honest boundaries: yielding is a cooperation protocol, not a hard guarantee (the tight check
150 ms before injection narrows the window as much as possible); a few applications
self-activate even on `set_value` (the receipt reports `foreground_changed` truthfully); and
two operators on the same window has no technical solution — do not drive the same window the
agent is driving.

## Tests

```bash
python tests/verify-coexistence.py    # 25 checks: zero-input proof / cross-process mutex / synthetic human contention / kill switch
python tests/verify-p0-fixes.py       # the three P0s fixed in 0.2.0: each fails before the fix
```

The tests need no human cooperation — "user input" is synthesized with one real 1-pixel cursor
move, and the cursor is restored afterwards.

**The tests need a real interactive desktop session** (some checks create windows and address
them through UIA), so they **cannot run on a GitHub-hosted runner**. What CI does cover is the
part that needs no desktop: packaging and installation, module import, regressions for the diff
index and tree-line escaping, and the arbiter's decision logic — see
[`.github/workflows/ci.yml`](https://github.com/Hutusion/dsh-cua/blob/main/.github/workflows/ci.yml).

```bash
python tests/ci-desktop-free.py       # the local equivalent of the above, no desktop needed
```

## License

MIT

