Automates Windows desktop apps through MCP using screenshots, OCR, mouse, keyboard, windows, clipboard, and app controls.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
We haven't yet run this listing's install command through our automated sandbox check. This isn't a red flag β we're steadily working through the catalog.
π‘ Paste the JSON block into your client's configuration file under mcpServers, then restart the application.
Inspect callable tools, capabilities, and parameters exposed to AI agents by Oswright.
screenshot- Take a screenshot of the screen or a region. Returns the image as native MCP image content. Optionally saves to a file path.
get_screen_info- Get screen dimensions and monitor count.
find_text_on_screen- Find all occurrences of text on screen using OCR. Returns matches with coordinates and confidence.
read_screen_text- Read ALL visible text on the screen using OCR. Returns every detected text element with position.
find_image_on_screen- Find all occurrences of a template image on screen using OpenCV template matching.
mouse_click- Click the mouse at coordinates or current position. Returns screenshot.
The oswright MCP server exposes desktop automation tools to an MCP-compatible AI client. Agents can inspect the display, identify visible text with OCR, locate template images, move or click the mouse, enter text, press keys, and wait for interface changes. It can also list, focus, minimize, close, and capture individual windows.
The tool set supports both low-level coordinate actions and higher-level text-driven actions. For example, an agent can locate a label with OCR, fill a field, submit a form, and verify the resulting screen. Clipboard read/write tools provide another way to move text between the agent and desktop applications. launch_app starts an application directly and can wait for it to load.
The project describes support for Windows through Win32 APIs, with Linux and macOS support using pynput and platform-specific OCR fallbacks. The listed automation surface is especially relevant to applications that do not provide an API or that are difficult to drive through browser automation.
The server communicates over MCP using a standard stdio configuration. Each action can return a current screenshot, while screen analysis can use OCR, image matching, or Windows UI Automation where available. Text lookup tools return coordinates and confidence values, allowing an agent to act on visible labels rather than hard-coded positions.
Its incremental perception model keeps screen state between observations and rescans areas that changed. The README also describes cached OCR results, screen memory, speculative perception, adaptive waiting, and a resolution cascade that attempts cheaper lookup methods before more expensive ones. These mechanisms are intended to reduce repeated full-screen reads during multi-step tasks.
Coordinates are represented as physical pixels for DPI-aware interaction. The repository includes benchmark and test material for measuring latency, token usage, and task completion, but the reported measurements are machine- and workload-specific rather than universal guarantees.
Python 3.10 or newer is required. The documented installation uses uvx:
If uvx is unavailable, install the package with pip install oswright and configure the command as oswright, or run it with python -m oswright. The README provides setup instructions for Claude Desktop, Claude Code, VS Code, Cursor, Windsurf, and Cline. It also documents a custom stdio setup for Goose.
The oswright MCP server includes tools for:
Most interaction tools return a screenshot after acting, while screenshot and OCR tools provide visual or structured observations for the next step.
Desktop automation depends on the target application's visible state, screen content, window titles, and supported platform mechanisms. OCR can report confidence values and coordinates rather than semantic certainty, and image matching depends on a suitable template. The README notes that accessibility and pixel-based perception have different coverage: accessibility may miss web content or some Win32 controls, while pixel-based methods can miss elements that are easier to identify through accessibility.
The Windows implementation uses built-in Windows OCR and does not require the larger PyTorch installation described for some non-Windows OCR paths. A physical display is needed for desktop-driving tests; those tests skip when no display is available. The project does not provide a remote browser or application API: it operates the local desktop exposed to the process.
Factual signals from GitHub, npm, and our automated checks β not a rating.
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/oswright)<a href="https://allmcps.com/mcp/oswright"><img src="https://allmcps.com/api/badge/oswright?style=directory" alt="Oswright on AllMCPs" /></a>