The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Screen Control listing page.
A local remote-control system for AI agents: watch your computer's screen live and send mouse/keyboard commands to it. Everything runs on your own machine — no data ever leaves it, no cloud middleman.
Every keystroke and every screenshot in this GIF went through the API — the agent never touched a physical keyboard.
AI agents today can write code and call APIs — but they can't see or touch your desktop. Screen Control gives any agent general-purpose computer use over a clean, safety-gated HTTP/MCP interface:
One process, zero configuration, works with any language that can speak HTTP — or natively through MCP in Claude Desktop, Cursor, VS Code and cloud agents.
Screen Control is the perception and actuation layer — the eyes and hands. The effective speed and capability of any agent using it are bounded by that agent itself and by the environment it runs in:
In practice this means: the same repo makes a fast reasoning model fast and capable, and makes a slow model slow — the toolchain is not the bottleneck. Real-time or action-heavy tasks need an agent with fast inference and tight tool-loop latency; slower agents should prefer deliberate, verification-heavy tasks.
| Feature | Description |
|---|---|
| 🖼️ Live screen feed | Continuously refreshing screenshot in the browser |
| 🖱️ Mouse control | Click, right-click, double-click, scroll, drag & drop via live screenshot |
| ⌨️ Keyboard control | Text typing (Unicode/Turkish included, layout-independent), keys and shortcuts (Ctrl+C, Alt+Tab…) |
| 👁️ OCR | Converts on-screen text to machine-readable format |
| 📷 Vision access | Raw-pixel paths for image-capable models: single frames, MJPEG stream, text-based motion detection |
| 🪟 Window management | List, focus, safe close (WM_CLOSE), kill (task-manager style) |
| 🖥️ Focus-free control | Read/write background windows via PostMessage without stealing focus |
| 🎮 Game mode | Camera look via relative mouse movement, hold-to-move keys |
| 🔐 Token auth | Every request requires X-Auth-Token (CSRF protection) |
| 🦺 Stuck-input watchdog | Auto-releases held keys after 30 s of inactivity |
| 🛟 Failsafe | Cursor to top-left corner aborts all commands (disabled in game mode) |
Per-Monitor DPI awareness. control.py calls
SetProcessDpiAwarenessContext(PER_MONITOR_AWARE_V2) at import time —
before the pyautogui import, because pyautogui touches coordinate APIs
during import and would otherwise lock the process to the interpreter
manifest's default (system-aware). With PMv2 active, every coordinate in the
system is a physical pixel end to end: mss capture, OCR bounding boxes,
pyautogui/SendInput clicks, ClipCursor. On High-DPI displays (125%/150%
scaling) nothing drifts between what OCR reports and where the mouse clicks.
Lock architecture. The server uses two independent locks instead of one global lock:
| Lock | Protects | Endpoints |
|---|---|---|
_input_lock | mouse, keyboard, game mode, window ops | /api/mouse, /api/key, /api/game, /api/window/post, ... |
_read_lock | capture, OCR, vision, enumeration | /api/screenshot, /api/ocr, /api/vision/*, /api/windows, ... |
A slow OCR (3–5 s on a busy screen) no longer freezes concurrent screenshot or vision reads — reads queue behind reads, inputs behind inputs.
This system is designed for a live perceive-act loop, not pre-written command chains:
This is enforced by the expect_hwnd guard: typing is refused (409) if
the foreground window doesn't match the target.
| Package | Purpose | Required? |
|---|---|---|
mss | Fast screen capture | ✅ Yes |
pyautogui | Mouse/keyboard control | ✅ Yes |
pyvda | Virtual desktop management | ✅ Yes |
flask | HTTP server | ✅ Yes |
Pillow | Image processing | ✅ Yes |
rapidocr-onnxruntime | OCR (screen text reading) | ⚠️ Optional |
Note: The OCR package is large and may take a while to install. If it fails, everything else still works — only the OCR feature is unavailable.
Returns the active backend name and what it can do. Agents should call this first (see ROADMAP.md for the multi-platform plan).
Feature values: true (supported), false (absent), null (unknown —
stub backend), "optional" (depends on an optional dependency).
| Capability | Windows | Linux X11 | Linux Wayland | macOS |
|---|---|---|---|---|
| Screen capture | Full | Full | Portal-dependent | Permission required |
| OCR | Full/optional | Full/optional | Full/optional | Full/optional |
| Mouse control | Full | Full | Restricted | Accessibility permission |
| Keyboard control | Full | Full | Restricted | Accessibility permission |
| Window enumeration | Full | WM-dependent | Limited | Accessibility/API-dependent |
| Background input | Strong | WM/app-dependent | Usually unavailable | Limited |
| Virtual desktops | Supported | DE/WM-dependent | DE/WM-dependent | Spaces-specific |
| Game mode | Supported | Experimental | Limited | Experimental |
Linux and macOS backends are currently fail-closed stubs: every operation returns
BACKEND_UNAVAILABLE(501) until implemented (ROADMAP Phases 5–7). Windows is the reference backend.
Every request must include the X-Auth-Token header. The token is
generated on each server start and written to .token.
| Code | Meaning |
|---|---|
| 401 | Missing or invalid token |
| 415 | POST without Content-Type: application/json |
Token bootstrap (for the bundled web UI):
The
/tokenendpoint is safe: Same-Origin Policy prevents foreign pages from reading it.
GET /api/screenshotReturns a JPEG screenshot.
| Parameter | Type | Default | Description |
|---|---|---|---|
monitor | int | 1 | Monitor index |
region | string | — | x,y,w,h sub-region |
GET /api/infoReturns screen dimensions and system state.
POST /api/mouse| action | Required params | Optional params | Description |
|---|---|---|---|
move | x, y | duration (default 0.15) | Move cursor to absolute position |
click | x, y | button (left/right), clicks (default 1) | Click at position |
scroll | clicks | x, y | Scroll wheel (positive=up) |
drag | x1, y1, x, y | duration, button | Drag between two points |
down | button (default "left") | — | Press and hold mouse button |
up | button (default "left") | — | Release held mouse button |
POST /api/key| action | Required params | Description |
|---|---|---|
press | key | Press and release a key |
down | key | Hold a key down (tracked for watchdog) |
up | key | Release a held key |
hotkey | keys (array) | Key combination (e.g. ["ctrl","c"]) |
type | text | Type text (Unicode, layout-independent) |
| Optional param | Default | Description |
|---|---|---|
expect_hwnd | — | Window handle to verify focus (409 if mismatch) |
interval | 0.03 | Delay between characters for type |
POST /api/ocrConverts on-screen text to machine-readable format.
| Param | Type | Default | Description |
|---|---|---|---|
region | array | — | [x, y, w, h] sub-region (faster) |
Three endpoints for models that can consume images:
| Endpoint | Description |
|---|---|
GET /api/vision/frame | Single JPEG frame (raw or base64) |
GET /api/stream | MJPEG live stream |
POST /api/vision/diff | Text-based motion detection (no vision needed) |
GET /api/vision/frame| Param | Default | Description |
|---|---|---|
scale | 1.0 | Downscale factor (0.5 = half size) |
gray | 0 | 1 for greyscale |
quality | 80 | JPEG quality (20-95) |
format | — | base64 for JSON response |
region | — | x,y,w,h sub-region |
GET /api/streamMJPEG live stream. Drop into <img src> or consume frame-by-frame.
| Param | Default | Description |
|---|---|---|
fps | 10 | Frames per second (1-30) |
quality | 70 | JPEG quality |
scale | 1.0 | Downscale factor |
region | — | x,y,w,h sub-region |
POST /api/vision/diffText-based motion detection — no vision model required.
| Body | Description |
|---|---|
{} | Compare against last stored frame |
{"grab":"gray"} | Store current frame for next comparison |
{"b64_prev":"..."} | Compare against provided previous frame |
GET /api/windowsList all visible windows.
POST /api/window| action | Required | Optional | Description |
|---|---|---|---|
focus | hwnd | — | Bring window to foreground |
close | hwnd | expect_title, expect_process | Safe close via WM_CLOSE |
kill | hwnd, pid | — | Force kill (task-manager style) |
topmost | hwnd | — | Set always-on-top |
untopmost | hwnd | — | Remove always-on-top |
maximize | hwnd | — | Maximise window |
Read and control windows without stealing focus — the user keeps working on their main desktop.
GET /api/window/captureCapture a window via PrintWindow (works even on another virtual desktop).
| Param | Description |
|---|---|
hwnd (required) | Window handle |
client | 1 = client area only |
ocr | 1 = return OCR text instead of image |
POST /api/window/postSend input to a window, choosing the delivery path automatically.
| action | Description |
|---|---|
type | Type text (Unicode-safe) |
key | Send a key press |
hotkey | Send a key combination |
click | Click at client coordinates |
scroll | Scroll the window |
drag | Drag inside the window |
Optional mode parameter controls routing:
| mode | Behavior |
|---|---|
auto (default) | Decided by input-mode probe (see below) |
background | Force PostMessage path (window keeps focus/z-order) |
focused | Force focus + SendInput path |
Routing rules (mode=auto):
postmessage — classic Win32 app: background PostMessage, no focus change.uia — WinUI/UWP/XAML surface (single DirectX canvas, no Win32 child
controls): posted messages are silently swallowed, so the window is focused
and the action is replayed through SendInput (client coords converted to
screen). This is the documented fallback for modern apps.focused — window is already foreground: focused SendInput path.invalid — HTTP 409; not a reachable top-level window.GET /api/window/input-modeClassify how a window receives input before posting to it. Returns one of
focused | postmessage | uia | invalid.
WinUI note: New Notepad (and other XAML-hosted apps) has no classic child Edit control to post to — the whole UI is one DirectX surface.
input-modereportsuiafor these;/api/window/postthen automatically uses the focused SendInput path./api/window/childrenremains useful for classic apps with real child controls.
GET /api/window/childrenList child controls of a window (class name + title + hwnd).
GET /api/desktopsList all virtual desktops.
POST /api/desktop| action | Params | Description |
|---|---|---|
switch | number | Switch to desktop N |
create | — | Create a new desktop |
| action | Params | Description |
|---|---|---|
start | sensitivity (default 12) | Lock cursor to center, enable game input |
move | dx, dy, sensitivity | Rotate camera (relative mouse) |
stop | — | Release cursor + all held input |
heartbeat | — | Keep-alive for long holds |
GET /api/heldReturns currently held keys/buttons and watchdog status.
POST /api/release_allEmergency: release everything (held keys, mouse buttons, game-mode cursor lock).
A dedicated, comprehensive guide for AI agents (LLMs, vision models, automation frameworks) is available in AGENT_GUIDE.md.
It covers:
expect_hwnd) to prevent wrong-window accidentsModel Context Protocol (MCP) turns this project into a plug-and-play toolbox for any MCP-capable agent: Claude Desktop, Claude Code, Cursor, VS Code Copilot Agent mode, custom cloud agents — no custom glue code, no curl scripts. The agent discovers and calls the tools natively.
mcp_server.py adds no new powers — every safety mechanism
(auth token, input/read locks, watchdog, Alt+F4 block, focus guard,
failsafe) stays enforced by server.py.
The MCP server auto-reads the token from .token (or the
SCREEN_CONTROL_TOKEN env var) — zero configuration.
Claude Desktop — claude_desktop_config.json:
Claude Code: claude mcp add screen-control -- python C:/path/to/screen-control/mcp_server.py
Cursor / VS Code: add the same entry to their MCP config files.
The HTTP transport is token-protected: every request must carry the
X-Auth-Token header (same token as the REST server). Query-string tokens
(?token=...) are rejected by design — URLs leak into proxy/tunnel
logs, browser history and shared links, and this token grants full desktop
control. Clients that cannot send custom headers should run a local stdio
mcp_server.py instead. Only GET /health is open, for liveness probes.
DNS-rebinding protection is disabled on this transport deliberately —
tunneled requests arrive with a foreign Host header, and the rebinding
threat is already covered by the token guard.
For a cloud agent, expose it through a tunnel:
Then configure the agent's MCP connection with <tunnel-url>/mcp plus the
token from .token as a header (X-Auth-Token).
For clients that cannot send custom headers (e.g. web connectors that only take an endpoint URL), create a scoped API key — a persistent, optionally time-limited credential — and embed it in the URL path:
Design guarantees (SC-06):
POST /api/keys
({"action":"revoke","name":"spark"}) — revocation takes effect
immediately on every endpoint.apikeys file stores only SHA-256 hashes, never raw keys⚠️ A tunnel exposes PC control to the internet. Keep the token secret, prefer short-lived tunnels and scoped keys for headerless connectors, and stop the server when not in use.
start-server.bat automates the whole cloud setup and prints everything
your cloud agent needs, ready to paste:
cloudflared.exe if missing (portable, no admin required)/token and /health probes)tunnel.log, and prints the summary:stop-server.bat stops all three (REST, MCP, tunnel) in one go.
| Category | Tools |
|---|---|
| Perception | get_info, ocr_screen, screenshot (real image block for vision models), motion_diff |
| Mouse / keyboard | mouse, keyboard (with expect_hwnd), get_held, release_all |
| Windows | list_windows, focus_window, window_children, window_input_mode, window_post, window_capture_ocr, close_window |
| Game mode | game (start / move / stop / heartbeat) |
| Consumer | Transport | Command |
|---|---|---|
| Claude Desktop / Cursor / VS Code (local) | stdio | python mcp_server.py |
| Claude Code | stdio | claude mcp add ... (above) |
| Cloud / remote agents | streamable-HTTP | start-server.bat (recommended) or python mcp_server.py --http --port 8751 + cloudflared tunnel --url http://127.0.0.1:8751 |
Note: This project targets MCP Python SDK 2.x (
MCPServerAPI). With SDK 1.x, replace the import withfrom mcp.server.fastmcp import FastMCP, ImageandMCPServerwithFastMCP.
Even bound to 127.0.0.1, a malicious page in the browser can trigger
non-preflighted requests (text/plain fetch, HTML form POST) to localhost.
The browser blocks the response but not the request — the server
would still execute the command.
Mitigation: Every request requires X-Auth-Token. A foreign page
cannot read this token (Same-Origin Policy), so it cannot authenticate.
Additional layers:
Content-Type: application/json (415 otherwise)Host header are refused with 421 — a
rebinding page that resolves its domain to 127.0.0.1 cannot read
/token or call the API/token and / responses carry Cache-Control: no-store so the
credential is never persisted by browsers or proxiesIn game mode, ClipCursor pins the cursor to a 2×2 box — the classic
pyautogui failsafe (cursor to top-left) does not work.
Mitigations:
Esc / Alt+Tab — real hardware input; this API cannot
block it, and it always worksPOST /api/release_all — instant release of everythingMitigations:
expect_hwnd guard on /api/key — if the foreground window doesn't
match, typing is refused with 409/api/window/post verifies the focus after the
focus switch and before any synthetic input (409 on mismatch) —
input is never replayed into whatever window happens to be foregroundfocus_window() raises on failure instead of silently returningMitigation: Blocked at the API level (403) on every delivery path —
the direct /api/key route, the background /api/window/post route
(PostMessage), and the focused fallback route share one safety policy
(control._assert_allowed):
Alt+F4 — the only banned Alt combo (Alt+Tab, Alt+menu are legitimate)Ctrl+Alt+Del — system security screenShift+Delete style — prevents permanent deletionMitigations:
winlogon.exe, csrss.exe, smss.exe, services.exe, lsass.exe,
svchost.exe, system, registry, dwm.exeexpect_process confirmation: a mismatch aborts the kill with
409 — protects against killing a newly-reused PIDA token-holding but misbehaving client should not be able to exhaust memory or starve the input lock.
Mitigations:
MAX_CONTENT_LENGTH = 1 MB — oversized request bodies are rejected (413)region width/height/area and scale are bounded (400 otherwise)text payloads are capped at 10,000 characters per input callThe server binds to 127.0.0.1 by default. To expose it to the network:
| Game Type | Suitable? | Notes |
|---|---|---|
| Minecraft (building) | ✅ Yes | Place blocks, walk, mine |
| Minecraft (PvP) | ❌ No | Too slow for fast combat |
| Turn-based games | ✅ Yes | Ample time for read→act→verify |
| RPG / adventure | ✅ Yes | Inventory, dialogue, exploration |
| Fast FPS | ❌ No | Reaction time insufficient |
| Puzzle games | ✅ Yes | Click-based, read-heavy |
If the consuming model can process images, use the vision endpoints directly:
This returns a single JPEG that the model can analyze for:
Use the diff endpoint for motion detection without vision:
The response tells you where things changed (tile coordinates) and how much (percentage), which is sufficient for:
| Approach | Payload | Use Case |
|---|---|---|
scale=1.0, gray=0 | ~500 KB | Full detail |
scale=0.5, gray=1 | ~50 KB | Good for most vision models |
scale=0.25, gray=1 | ~10 KB | Maximum compression |
diff (text) | ~1 KB | Text-only agents |
region=... | Variable | Focus on specific area |
The foreground window changed between the focus call and the type call.
Solution: always pass expect_hwnd and verify focus before typing.
The window may have been closed or may be a system window that
EnumWindows doesn't expose. Try:
Use POST /api/release_all or press Esc / Alt+Tab physically.
OCR on a full 1920×1080 screen can take from a few seconds up to ~30 s depending on your CPU and on-screen complexity. Use a region — small crops are typically 10× faster:
The system uses SendInput + KEYEVENTF_UNICODE which is layout-independent.
If characters still don't appear, the target app may not support Unicode
input — try POST /api/window/post with action: "type" instead.
The server must be running:
Tests authentication, blocked key combos, window management, safe close, critical process protection, and game mode — all non-destructive.
Expected output:
Launches a real application (mspaint or notepad), performs hold-to-draw game mechanics, verifies via pixel analysis, then safely closes with "Don't Save" dialog handling.
Note: This test launches a real application. It handles cleanup automatically (sends WM_CLOSE and clicks "Don't Save" if a dialog appears).
MIT License. See LICENSE for details.
Built with ❤️ for local automation and AI agent research.