# open-compute (computer-use) [Health: Active]

**Category:** 💻 Developer Tools  
**Repository:** https://github.com/ellmos-ai/open-compute-mcp  
**GitHub Stars:** 1  
**Views:** 0  
**Installs:** 0  
**Upvotes:** 0  
**Directory Page:** https://allmcps.com/mcp/open-compute-computer-use

## Description
Model-agnostic computer-use MCP: screenshot, safety-gated actions, Windows UIA targeting.

## Claude Desktop Quick Installation
Install path detected from listing signals. Uses `npx` (confidence: high):

```json
"mcpServers": {
  "open-compute-computer-use": {
    "command": "npx",
    "args": ["-y","open-compute-mcp"]
  }
}
```

## Documentation & README

<p align="center">
  <img src="https://raw.githubusercontent.com/ellmos-ai/open-compute-mcp/main/assets/wappen.jpg" alt="open-compute MCP server emblem" width="400">
</p>

# open-compute-mcp

**npm launcher for the [open-compute](https://github.com/ellmos-ai/open-compute) MCP server** —
model-agnostic **computer-use** tools exposed over the Model Context Protocol (MCP).

**EN** | [DE](https://github.com/ellmos-ai/open-compute-mcp/blob/HEAD/README_de.md)

[![CI](https://github.com/ellmos-ai/open-compute-mcp/actions/workflows/ci.yml/badge.svg)](https://github.com/ellmos-ai/open-compute-mcp/actions/workflows/ci.yml)
[![npm version](https://img.shields.io/npm/v/open-compute-mcp.svg)](https://www.npmjs.com/package/open-compute-mcp)
[![npm downloads](https://img.shields.io/npm/dt/open-compute-mcp.svg)](https://www.npmjs.com/package/open-compute-mcp)
[![GitHub Stars](https://img.shields.io/github/stars/ellmos-ai/open-compute-mcp.svg)](https://github.com/ellmos-ai/open-compute-mcp)
[![License: MIT](https://img.shields.io/github/license/ellmos-ai/open-compute-mcp.svg)](LICENSE)
[![Node.js](https://img.shields.io/badge/node-%3E%3D18-brightgreen.svg)](https://nodejs.org/)
[![Tests](https://img.shields.io/badge/tests-25%20passed-brightgreen.svg)](test)
[![MCP Enabled](https://img.shields.io/badge/MCP-server-blue.svg)](https://modelcontextprotocol.io)
[![Platform](https://img.shields.io/badge/platform-Windows%20%7C%20Linux%20%7C%20macOS-lightgrey.svg)](https://github.com/ellmos-ai/open-compute-mcp)
[![Privacy: Zero-Egress](https://img.shields.io/badge/privacy-100%25%20Offline%20%7C%20Zero--Egress-blue.svg)](SECURITY.md)
[![Security: Safety-Gated](https://img.shields.io/badge/security-Operator%20Ceiling%20%7C%20Safety--Gated-green.svg)](SECURITY.md)
[![Ecosystem: ellmos-ai](https://img.shields.io/badge/ecosystem-ellmos--ai-blueviolet.svg)](https://github.com/ellmos-ai)
[![Umbrella: open-bricks](https://img.shields.io/badge/umbrella-open--bricks-indigo.svg)](https://github.com/open-bricks)
[![Glama](https://glama.ai/mcp/servers/ellmos-ai/open-compute-mcp/badges/score.svg)](https://glama.ai/mcp/servers/ellmos-ai/open-compute-mcp)
[![LLM Ready](https://img.shields.io/badge/LLM-ready-success.svg)](https://github.com/ellmos-ai/open-compute-mcp/blob/main/llms.txt)

📦 **[View on npm →](https://www.npmjs.com/package/open-compute-mcp)** • 📋 **[Security Policy](https://github.com/ellmos-ai/open-compute-mcp/blob/HEAD/SECURITY.md)** • ⚖️ **[Licenses](https://github.com/ellmos-ai/open-compute-mcp/blob/HEAD/THIRD_PARTY_LICENSES.md)** • 🤖 **[LLM Context (llms.txt)](https://github.com/ellmos-ai/open-compute-mcp/blob/HEAD/llms.txt)**

---

### Quick Navigation

- [✨ Key Capabilities](#key-capabilities)
- [🏗️ Architecture](#architecture)
- [🛠️ Tools (16)](#tools)
- [🚀 Use with an MCP Client](#use-with-an-mcp-client)
- [🔄 Safe Interaction & Signal Lifecycle](#safe-interaction--signal-lifecycle)
- [⚙️ Configuration](#configuration-environment-variables)
- [🔒 Safety & Security](#safety)
- [🌐 ellmos-ai Ecosystem](#ellmos-ai-ecosystem)

---

> [!NOTE]
> **AI Assistant / Agent Integration**: This repository contains an [`llms.txt`](https://github.com/ellmos-ai/open-compute-mcp/blob/HEAD/llms.txt) file providing structured, machine-readable specifications of tools, safety modes (`OC_SAFETY_MODE`), and client configuration examples for RAG crawlers and autonomous agent frameworks.

The MCP **client is the reasoner** (no API key, model-agnostic): it calls `capture`
to see the screen, then acts with `do` / `click_name` / `invoke`. This is the keyless
Mode-A loop of open-compute, but as native tool-calls.

## Key Capabilities

1. **State-bound Perception & Window Targeting:** Captures/trees return one-shot observation IDs; window enumeration returns stable window/process IDs and issued tokens. WGC remains the GPU-window fallback.
2. **Fail-closed Action Execution:** Coordinates consume one observation and exact window binding; UIA names resolve exact-first; text is segmented with focus checks and character-count postconditions.
3. **Leased Signal Overlay & Abort Control:** The glowing border/cursor signal has owner/session metadata, a bounded TTL, turn-end cleanup, and immediate human abort.
4. **Multimodal Collaboration & Voice Notes:** Push-to-talk voice recording (`talk`), screen chat messaging (`chat`), directory monitoring (`watch_dir`), and macro replay (`rec_replay`).

## Architecture

```mermaid
graph TD
    A["AI Reasoner<br/>(Claude / Antigravity / Cursor)"] -- "MCP stdio (JSON-RPC)" --> B["npx open-compute-mcp<br/>(Node.js Launcher)"]
    B -- "Spawns via uvx" --> C["open-compute Python Engine<br/>(GitHub @ main)"]
    C -- "Screenshots / WGC" --> D["Windows Display"]
    C -- "UIA / Mouse / Keys" --> E["Windows Desktop Apps"]
    C -- "Glowing Border & Cursor" --> F["Signal Overlay UI"]

    subgraph Safety Gate
        C -. "OC_SAFETY_MODE<br/>(confirm / read_only / allow_all)" .-> C
        C -. "OC_DENY<br/>(hard action blacklist)" .-> C
    end
```

> This package is a **thin launcher**. It contains no server logic — it spawns the
> **Python** open-compute server (pulled from GitHub) and pipes MCP stdio through.
> Real screen capture and input require the **interactive Windows desktop session**.

## Requirements

- **Python 3.10+** and **[uv](https://docs.astral.sh/uv/)** on the host. The default
  launch uses `uvx` to fetch open-compute (with the `mcp` extra) **from GitHub** on
  first run — the `mcp` extra tracks the GitHub repo, so this works regardless of
  PyPI release timing.
- **Windows** for real capture/input (mss + UIA). Other platforms import the tools
  but cannot drive a desktop.

## Tools

| Tool | Purpose |
|---|---|
| `capture` | Return one-shot observation metadata plus an image (optionally one exact window). |
| `do` | Execute a safety-gated action; coordinates require `observation_id` + issued window descriptor/token. |
| `tree` | Return UIA elements and a one-shot observation ID for their coordinates. |
| `click_name` | Exact-first, ambiguity-safe click in a required issued window, with score/alternatives. |
| `invoke` | Exact-first, click-free UIA activation in a required issued window. |
| `list_windows` | List stable window/process IDs, exact titles, issued tokens, rects and centers. |
| `get_screen_size` | Virtual-desktop geometry + per-monitor breakdown (read-only). |
| `watch_dir` | Watch directories for file-system changes. |
| `push_status` | Feed-manager status (read-only). |
| `rec_replay` | Replay a `.clirec` macro (needs the optional `clirec` package). |
| `signal_show` | Show a configurable pre-action color/text countdown, then the mode-colored overlay, with owner/session lease and bounded TTL. |
| `signal_hide` | Hide the signal overlay. |
| `signal_status` | Owner/session/mode/visible/expires_at + pending abort message. |
| `signal_abort` | Ask the human for a short abort reason; the message is returned for the model. |
| `chat` | Human→model message about screen content, optionally with screenshot. |
| `talk` | Push-to-talk voice note → WAV path (hold key, speak, release; STT/TTS model-side). |

All coordinates are **normalized 0..1** relative to the virtual desktop. Tool
descriptions are localized in six languages (`de/en/es/ja/ru/zh`) via `OC_LANGUAGE`.

`do` also accepts the **hold primitives** `mouse_down` / `mouse_up` / `key_down` /
`key_up` for press-and-hold sequences (rubber-band selection, modifier-held
clicking, game input); anything still held is released when the server stops.
`capture(window=...)` falls back to Windows.Graphics.Capture when a plain grab of
a hardware-composited window (Roblox Studio, Blender, a GPU-accelerated browser)
comes back all-black — install the `wgc` extra for that.

## Safe Interaction & Signal Lifecycle

The v0.8 Python engine enforces observe → one action → automatic refresh. Keep
the full descriptor or `window_token` from `list_windows`, then pass it as
`expected_window` together with the latest `observation_id` from `capture` or
`tree`. `click_name`/`invoke` require that issued window too. Reuse, changed
state, focus mismatch, covered windows, and ambiguous UIA targets are rejected
before input. `type` returns requested/sent character
counts and complete/partial status without echoing the text. Signals have a
hard TTL and are removed at action turn end unless `keep_signal=true`.

An explicit `signal_show` starts the engine's configured pre-action grace
period. The static grace color is distinct from the mode color and the visible
text counts down `Start in N Sekunden` once per second. At zero, both phase and
color switch once to active. `signal_status` exposes the same phase, remaining
seconds, current color, and screenreader label. Duration, grace color, and text
template come from `OC_SIGNAL_GRACE_SECONDS` / `OC_SIGNAL_CONFIG`; `0` skips the
countdown. The design uses no flashing, pulsing, or motion animation.

```mermaid
sequenceDiagram
    autonumber
    actor Reasoner as AI Reasoner (Claude / AGY)
    participant Launcher as Node.js Launcher (open-compute-mcp)
    participant Engine as Python Engine (open-compute)
    participant UI as Windows Desktop / UIA
    actor Operator as Human Operator

    Note over Reasoner,Operator: Phase 1: Visual Perception & State Inspection
    Reasoner->>Launcher: capture(window?) / tree()
    Launcher->>Engine: Forward stdio JSON-RPC
    Engine->>UI: Grab Screen (mss/WGC) or Read UIA Tree
    UI-->>Engine: Frame Image / Semantic Element Tree
    Engine-->>Launcher: Observation ID + normalized response/image
    Launcher-->>Reasoner: State-bound visual observation

    Note over Reasoner,Operator: Phase 2: Signal Overlay Activation
    Reasoner->>Launcher: signal_show(mode="control")
    Launcher->>Engine: Invoke Signal Overlay
    Engine->>UI: Render static grace color + Start in N seconds
    UI-->>Operator: Text countdown + accessible window name
    Engine->>UI: At zero, switch once to the mode color

    Note over Reasoner,Operator: Phase 3: Action Request & Safety Gate
    Reasoner->>Launcher: do(one action, window token, observation_id) / click_name(target)
    Launcher->>Engine: Process Action Payload
    alt OC_SAFETY_MODE == "confirm" (Default)
        Engine-->>Launcher: Status "needs_confirmation" (Report Only)
        Launcher-->>Reasoner: Human confirmation needed
    else OC_SAFETY_MODE == "allow_all" (Isolated VM)
        Engine->>UI: Execute Mouse/Keyboard / Hold Primitives
        UI-->>Engine: Action Completed
        Engine-->>Launcher: Post-observation + window/modal/text postconditions
        Launcher-->>Reasoner: Action completed; old observation invalid
    end

    Note over Reasoner,Operator: Phase 4: Emergency Abort or Completion
    opt Operator Triggers Emergency Abort
        Operator->>Engine: Hotkey Pressed (Abort Signal)
        Engine->>UI: Auto-release all held keys/mouse buttons
        Engine-->>Reasoner: signal_abort message returned
    end
    Engine->>UI: Remove overlay on turn end/error/abort (unless keep_signal=true)
```

## Use with an MCP client

**Via this npm launcher (npx):**

```json
{
  "mcpServers": {
    "open-compute": {
      "command": "npx",
      "args": ["-y", "open-compute-mcp"]
    }
  }
}
```

**Directly via Python (uvx), no npm:**

```json
{
  "mcpServers": {
    "open-compute": {
      "command": "uvx",
      "args": ["--from", "open-compute[mcp,local,uia] @ git+https://github.com/ellmos-ai/open-compute.git", "open-compute-mcp"]
    }
  }
}
```

## Configuration (environment variables)

| Variable | Effect |
|---|---|
| `OPEN_COMPUTE_PYTHON` | Path to a `python.exe`; the launcher runs `-m open_compute.mcp_server` with it (use this if you installed open-compute into a specific environment). |
| `OPEN_COMPUTE_MCP_CMD` | Full command override (whitespace-split), e.g. `python -m open_compute.mcp_server`. |
| `OPEN_COMPUTE_GIT_REF` | Git ref (branch/tag/sha) to pin for the uvx launch (default: the repo's default branch). |
| `OPEN_COMPUTE_EXTRAS` | Extras for the default `uvx` launch (default `mcp,local,uia`). |
| `OC_LANGUAGE` | Language of the tool descriptions: `de`/`en`/`es`/`ja`/`ru`/`zh`. |
| `OC_SAFETY_MODE` | `confirm` (default) · `read_only` · `allow_all`. |
| `OC_DENY` | Comma-separated action types always denied (e.g. `type,launch_app`). |
| `OC_CAPTURE_SCALE` | Resize factor for every capture, `0.05`–`1.0`. **This launcher defaults to `0.5`** (see below); set `1.0` for full resolution. |
| `OC_CAPTURE_MAX_DIM` | Cap the longest edge in pixels (default off). Setting it suppresses the scale default, so the two never shrink twice. |
| `OC_CAPTURE_GRAYSCALE` | `1` drops colour. Shrinks the payload, **not** the token count — that follows pixel count alone. |
| `OC_SIGNAL_TTL` | Hard overlay lease limit in seconds (default 120). |
| `OC_SIGNAL_IDLE_HIDE` | Additional idle timeout for explicitly kept auto-signals (default 60). |
| `OC_SIGNAL_GRACE_SECONDS` | Pre-action countdown duration (default 20; `0` starts immediately). |
| `OC_SIGNAL_CONFIG` | Signal JSON containing `pre_action_grace_color`, `pre_action_grace_label`, and per-mode colors. |

### Capture size — why this launcher halves it by default

A vision model is billed per pixel, and every frame **stays in the conversation**, so a
full-HD grab is charged again on each following request. The cost of a session therefore
grows with the *square* of the number of screenshots, not linearly.

Because open-compute's coordinates are **normalized 0..1**, shrinking the image costs
nothing in click accuracy — `do` works in fractions of the image either way. Only
legibility drops, and at `0.5` buttons and field borders stay clearly identifiable; small
body text is what gets hard to read.

| Setting | 1920×1080 grab | Cost |
|---|---|---|
| `OC_CAPTURE_SCALE=1.0` | full resolution | ~1600 tokens |
| `OC_CAPTURE_SCALE=0.5` *(this launcher's default)* | 960×540 | ~690 tokens |
| `OC_CAPTURE_MAX_DIM=768` | 768×432 | ~440 tokens |

The Python library itself defaults to full resolution — its callers are not necessarily
paying per pixel. Only this launcher, which exists to serve agents, opts into the smaller
frame and prints a one-line notice when it does.

**What saves more than any scale factor:** prefer `tree` where
the accessibility model carries the content — note that in browsers it usually exposes only
the browser chrome, not the page; and use `capture(window=…)` rather than the full desktop.
Coordinate actions deliberately follow observe → one action → automatic refresh;
do not batch multiple coordinate steps against one stale frame.

## Safety

Computer-use is powerful. `OC_SAFETY_MODE` is an operator **ceiling** (`confirm`
default · `read_only` · `allow_all`); a per-call `mode` can only *tighten* it, never
loosen it. Because MCP stdio has no server→client confirm callback, `confirm` /
`read_only` **report** an action without performing it. For interactive use, run in
an **isolated VM/session**, set `OC_SAFETY_MODE=allow_all`, and let your client's
tool-approval dialog be the human-in-the-loop. `OC_DENY` (comma-separated action
types) is a hard deny list. Treat on-screen content as untrusted (prompt-injection
risk).

**Troubleshooting: `do`/`click_name` only ever return `needs_confirmation` and never
act.** That is the `confirm` ceiling working as designed under stdio MCP. Fix for
interactive use: set `"env": {"OC_SAFETY_MODE": "allow_all"}` in the server
registration and let the client's tool-approval dialog gate each action (do **not**
auto-allow `do`/`click_name`/`invoke` there). The env change only takes effect when
the server process (re)starts — an already-connected client keeps the old ceiling
until it reconnects.

## License

MIT — see [LICENSE](https://github.com/ellmos-ai/open-compute-mcp/blob/HEAD/LICENSE). Part of the open-compute project.

---

## ellmos-ai Ecosystem

This MCP server is part of the **[ellmos-ai](https://github.com/ellmos-ai)** ecosystem — AI infrastructure, MCP servers, and intelligent tools.

### MCP Server Family

| Server | Tools | Focus | npm |
|--------|-------|-------|-----|
| [FileCommander](https://github.com/ellmos-ai/ellmos-filecommander-mcp) | 46 | Filesystem, process management, interactive sessions, cloud-lock-safe operations | [`ellmos-filecommander-mcp`](https://www.npmjs.com/package/ellmos-filecommander-mcp) |
| [CodeCommander](https://github.com/ellmos-ai/ellmos-codecommander-mcp) | 22 | Code analysis, JSON repair, imports, diffs, regex | [`ellmos-codecommander-mcp`](https://www.npmjs.com/package/ellmos-codecommander-mcp) |
| [Clatcher](https://github.com/ellmos-ai/ellmos-clatcher-mcp) | 12 | File repair, format conversion, batch operations | [`ellmos-clatcher-mcp`](https://www.npmjs.com/package/ellmos-clatcher-mcp) |
| [n8n Manager](https://github.com/ellmos-ai/n8n-manager-mcp) | 18 | n8n workflow management via AI assistants | [`n8n-manager-mcp`](https://www.npmjs.com/package/n8n-manager-mcp) |
| [ControlCenter](https://github.com/ellmos-ai/ellmos-controlcenter-mcp) | 20 | MCP stack discovery, profile management, control plane | [`ellmos-controlcenter-mcp`](https://www.npmjs.com/package/ellmos-controlcenter-mcp) |
| [Homebase](https://github.com/ellmos-ai/ellmos-homebase-mcp) | 45 | Local-first LLM memory, knowledge, state, routing, swarm orchestration | [`ellmos-homebase-mcp`](https://www.npmjs.com/package/ellmos-homebase-mcp) (alpha) |
| [ServerCommander](https://github.com/ellmos-ai/ellmos-servercommander-mcp) | 8 | Server operations: health checks, log analysis, deploy dry-runs, mail diagnostics | [`ellmos-servercommander-mcp`](https://www.npmjs.com/package/ellmos-servercommander-mcp) (alpha) |
| [Blender Use](https://github.com/ellmos-ai/ellmos-blender-use-mcp) | 3 | Headless Blender asset QA and FBX reimport verification | [`ellmos-blender-use-mcp`](https://www.npmjs.com/package/ellmos-blender-use-mcp) (alpha) |
| **[Open Compute](https://github.com/ellmos-ai/open-compute-mcp)** | **16** | **Model-agnostic computer use: capture, safety-gated actions, Windows UIA, signal overlay & voice/chat** | **[`open-compute-mcp`](https://www.npmjs.com/package/open-compute-mcp)** (alpha) |

### AI Infrastructure & Sibling Tooling

| Project | Description |
|---|---|
| [BACH](https://github.com/ellmos-ai/bach) | Local-first text-based OS for LLM agents — 113+ handlers, 550+ tools, SQLite memory |
| [open-compute](https://github.com/ellmos-ai/open-compute) | Model-agnostic computer-use core powering Open Compute MCP |
| [clutch](https://github.com/ellmos-ai/clutch) | Provider-neutral LLM orchestration with auto-routing and budget tracking |
| [rinnsal](https://github.com/ellmos-ai/rinnsal) | Lightweight agent memory, connectors, and automation infrastructure |
| [ellmos-stack](https://github.com/ellmos-ai/ellmos-stack) | Self-hosted AI research stack (Ollama + n8n + Rinnsal + KnowledgeDigest) |
| [MarbleRun](https://github.com/ellmos-ai/MarbleRun) | Autonomous agent chain framework for Claude Code |
| [gardener](https://github.com/ellmos-ai/gardener) | Minimalist database-driven LLM OS prototype (4 functions, 1 table) |
| [ellmos-tests](https://github.com/ellmos-ai/ellmos-tests) | Testing framework for LLM operating systems (7 dimensions) |
| [sqlite-transit-sync](https://github.com/ellmos-ai/sqlite-transit-sync) | Safe, redacted, HMAC-verified SQLite snapshot synchronizer |
| [policy-registry](https://github.com/ellmos-ai/policy-registry) | Hierarchical policy & delegation authority engine |

### Open Bricks Umbrella

Our partner organization **[open-bricks](https://github.com/open-bricks)** bundles AI-native desktop applications — a modern, open-source software suite built for the age of AI. Sibling suites include [DevCenter](https://github.com/dev-bricks/DevCenter), [CodeBox](https://github.com/dev-bricks/CodeBox), [MethodenAnalyser](https://github.com/dev-bricks/MethodenAnalyser), [CleanMarkdown](https://github.com/doc-bricks/CleanMarkdown), and [PDFtoPDFocr](https://github.com/doc-bricks/PDFtoPDFocr).


