# oswright [Health: Active]

**Category:** 💻 Developer Tools  
**Repository:** https://github.com/Ask-812/oswright  
**GitHub Stars:** 3  
**Views:** 1  
**Installs:** 0  
**Upvotes:** 0  
**Directory Page:** https://allmcps.com/mcp/oswright

## Description
Windows desktop automation that re-reads only what changed: 7x fewer tokens per task.

## Tools
Capabilities this server exposes over MCP:

- **screenshot** — - Take a screenshot of the screen or a region. Returns the image as native MCP image content. Optionally saves to a file path.
- **get_screen_info** — - Get screen dimensions and monitor count.
- **find_text_on_screen** — - Find all occurrences of text on screen using OCR. Returns matches with coordinates and confidence.
- **read_screen_text** — - Read ALL visible text on the screen using OCR. Returns every detected text element with position.
- **find_image_on_screen** — - Find all occurrences of a template image on screen using OpenCV template matching.
- **mouse_click** — - Click the mouse at coordinates or current position. Returns screenshot.
- **mouse_double_click** — - Double-click at coordinates or current position. Returns screenshot.
- **mouse_move** — - Move the mouse cursor to screen coordinates.
- **mouse_scroll** — - Scroll the mouse wheel. Returns screenshot.
- **mouse_drag** — - Drag from one point to another. Returns screenshot.
- **get_mouse_position** — - Get the current mouse cursor position.
- **type_text** — - Type text character by character. Returns screenshot.
- **press_key** — - Press a key or combo like `Enter`, `Ctrl+C`, `Alt+Tab`. Returns screenshot.
- **click_text** — - Find text via OCR and click on it. Auto-retries until found or timeout. Returns screenshot.
- **double_click_text** — - Find text via OCR and double-click on it. Returns screenshot.
- **right_click_text** — - Find text via OCR and right-click on it. Returns screenshot.
- **hover_text** — - Find text via OCR and hover over it. Returns screenshot.
- **fill_field** — - Find a label, click it, clear, and type a value. Returns screenshot.
- **fill_form** — - Fill multiple fields in one call. Reduces round-trips.
- **wait_for_text** — - Wait for text to appear on screen. Polls via OCR.
- **wait_for_text_gone** — - Wait for text to disappear from screen.
- **wait_for_time** — - Wait for a specified duration (capped at 30s), then screenshot.
- **list_windows** — - List all visible windows. Optionally filter by title substring.
- **focus_window** — - Bring a window to the foreground by title. Returns screenshot.
- **close_window** — - Close a window by title (sends WM_CLOSE). Returns screenshot.
- **minimize_window** — - Minimize a window by title. Returns screenshot.
- **screenshot_window** — - Capture a screenshot of just one window.
- **get_clipboard** — - Get the current text content of the system clipboard.
- **set_clipboard** — - Copy text to the system clipboard.
- **launch_app** — - Launch an application and optionally wait for it to load. Runs the program directly, never through a shell.
- **get_ocr_info** — - Get info about the active OCR backend and available backends.
- **observe** — - Report what changed on screen since the last observation. Rescans only the regions that moved. Prefer this over `screenshot` for tracking state.
- **find_element** — - Find on-screen text using the cheapest method that can answer. Reports which cascade rung responded.
- **click_element** — - Find text via the cascade and click it. The cheap alternative to `click_text`.
- **read_model_text** — - Read on-screen text from the incremental model without re-OCRing the display.
- **perception_stats** — - Report how much perception work the model has avoided.
- **remember_screen** — - Remember the current screen so future visits skip reading it. Persists across sessions.
- **atlas_stats** — - Report what the screen atlas has remembered and how often it helped.
- **get_ui_tree** — - Get the accessibility tree of the focused window. Returns all interactive elements with names, types, positions. Deterministic and instant.
- **click_ui_element** — - Click a UI element using the accessibility tree. More reliable than OCR.
- **fill_ui_element** — - Set the value of a UI element (e.g., text box). More reliable than OCR-based fill.
- **get_active_window** — - Get info about the currently focused window.
- **wait_for_change** — - Wait for the screen to visually change. Takes a baseline screenshot, polls until different.

## Claude Desktop Quick Installation
Install path detected from listing signals. Uses `uvx` (confidence: high):

```json
"mcpServers": {
  "oswright": {
    "command": "uvx",
    "args": ["oswright"]
  }
}
```

## Documentation

## What the oswright MCP server does

The oswright MCP server exposes desktop automation tools to an MCP-compatible AI client. Agents can inspect the display, identify visible text with OCR, locate template images, move or click the mouse, enter text, press keys, and wait for interface changes. It can also list, focus, minimize, close, and capture individual windows.

The tool set supports both low-level coordinate actions and higher-level text-driven actions. For example, an agent can locate a label with OCR, fill a field, submit a form, and verify the resulting screen. Clipboard read/write tools provide another way to move text between the agent and desktop applications. `launch_app` starts an application directly and can wait for it to load.

The project describes support for Windows through Win32 APIs, with Linux and macOS support using pynput and platform-specific OCR fallbacks. The listed automation surface is especially relevant to applications that do not provide an API or that are difficult to drive through browser automation.

## How the oswright MCP server works

The server communicates over MCP using a standard stdio configuration. Each action can return a current screenshot, while screen analysis can use OCR, image matching, or Windows UI Automation where available. Text lookup tools return coordinates and confidence values, allowing an agent to act on visible labels rather than hard-coded positions.

Its incremental perception model keeps screen state between observations and rescans areas that changed. The README also describes cached OCR results, screen memory, speculative perception, adaptive waiting, and a resolution cascade that attempts cheaper lookup methods before more expensive ones. These mechanisms are intended to reduce repeated full-screen reads during multi-step tasks.

Coordinates are represented as physical pixels for DPI-aware interaction. The repository includes benchmark and test material for measuring latency, token usage, and task completion, but the reported measurements are machine- and workload-specific rather than universal guarantees.

## Setup and configuration

Python 3.10 or newer is required. The documented installation uses `uvx`:

```json
{
  "mcpServers": {
    "oswright": {
      "command": "uvx",
      "args": ["oswright"]
    }
  }
}
```

If `uvx` is unavailable, install the package with `pip install oswright` and configure the command as `oswright`, or run it with `python -m oswright`. The README provides setup instructions for Claude Desktop, Claude Code, VS Code, Cursor, Windsurf, and Cline. It also documents a custom stdio setup for Goose.

## Tools and capabilities

The oswright MCP server includes tools for:

- Capturing the full screen, a region, or a specific window.
- Reading all visible text, finding text, and waiting for text to appear or disappear.
- Finding template images with OpenCV matching.
- Clicking, double-clicking, right-clicking, hovering, dragging, scrolling, and moving the mouse.
- Typing text, pressing keys or key combinations, and filling one or multiple fields.
- Reading the mouse position and screen dimensions, including monitor count.
- Listing, focusing, minimizing, and closing windows.
- Reading from and writing to the system clipboard.
- Launching applications and optionally waiting for startup.

Most interaction tools return a screenshot after acting, while screenshot and OCR tools provide visual or structured observations for the next step.

## Limitations and notes

Desktop automation depends on the target application's visible state, screen content, window titles, and supported platform mechanisms. OCR can report confidence values and coordinates rather than semantic certainty, and image matching depends on a suitable template. The README notes that accessibility and pixel-based perception have different coverage: accessibility may miss web content or some Win32 controls, while pixel-based methods can miss elements that are easier to identify through accessibility.

The Windows implementation uses built-in Windows OCR and does not require the larger PyTorch installation described for some non-Windows OCR paths. A physical display is needed for desktop-driving tests; those tests skip when no display is available. The project does not provide a remote browser or application API: it operates the local desktop exposed to the process.

_Full upstream README: https://allmcps.com/mcp/oswright/readme_

