AgentStack
MCP verified MIT Self-run

Desktop Touch Mcp

mcp-harusame64-desktop-touch-mcp · by Harusame64

Windows computer-use MCP server: drive any desktop app via semantic discover-then-act targeting (entities + leases, not pixel coordinates), with per-action perception guards, a native Rust UIA engine, and Chrome CDP. Works with Claude, Cursor, and any MCP client.

No reviews yet
0 installs
8 views
0.0% view→install

Install

$ agentstack add mcp-harusame64-desktop-touch-mcp

✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.13.1 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.13.1. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

Are you the author of Desktop Touch Mcp? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

desktop-touch-mcp

[](https://glama.ai/mcp/servers/Harusame64/desktop-touch-mcp)

[日本語](README.ja.md)

> Computer-use MCP server for Windows. Lets Claude, Cursor, or any MCP client see and operate your Windows 10/11 desktop — screenshots, UI Automation, Chrome CDP, keyboard / mouse, terminal — with semantic discover-then-act targeting that avoids pixel-coordinate guessing, and per-action perception guards that catch wrong-window typing before it happens.

npx -y @harusame64/desktop-touch-mcp

32 tools, native Rust engine (UIA in 2 ms), zero-config PowerShell fallback, full CJK support, MIT licensed. Add the snippet above to your Claude / Cursor / VS Code Copilot config and Claude can drive Notepad, Excel, Chrome, Windows Terminal, and any other app on your machine.

> Why this over pixel-clicking? Two ideas run through every tool: discover-then-actdesktop_discover returns interactive entities with short-lived leases instead of raw coordinates, so desktop_act operates on what you mean, not where it was — and per-action perception guards that verify the target window's identity and bounds before input lands, catching wrong-window typing and stale-coordinate clicks before they happen. > > Under the hood: an 82× average speedup from the Rust native engine (UIA focus queries in 2 ms, SSE2-accelerated image diffing at 13–15×), with a transparent PowerShell fallback when the engine is absent. The npm launcher fetches only the GitHub Release tag matching the installed version and verifies the Windows runtime zip before extraction.


Features

  • ⚡ High-performance Rust Native Core — The UIA bridge and image-diff engine are written in Rust (napi-rs + windows-rs) and loaded as a native .node addon. Direct COM calls from a dedicated MTA thread eliminate PowerShell process spawning — getFocusedElement completes in 2 ms (160× faster), and getUiElements returns full trees in ~100 ms with a batch BFS algorithm that minimizes cross-process RPC. Image-diff operations use SSE2 SIMD for 13–15× throughput. When the native engine is unavailable, every function transparently falls back to PowerShell — zero config required.
  • 🎯 Set-of-Marks (SoM) visual fallback — Games, RDP sessions, and non-accessible Electron apps return clickable elements even when UIA is completely blind. screenshot(detail="text") automatically detects UIA sparsity and activates a Hybrid Non-CDP pipeline: Rust-powered grayscale + bilinear upscale → Windows OCR → clustering → red bounding-box annotation with numbered badges ([1], [2]…). Two parallel representations returned: a visual PNG for spatial orientation and a semantic elements[] list with clickAt coords — no CDP required.
  • 🔁 One-call confirmation on visual-only targets — On UIA-blind targets (Electron, PWAs, games, custom canvases, RDP windows), desktop_act can fold the post-action confirmation into its own response: an optional roiCapture carrying a PNG crop of just the region that changed plus a lease-less preview of the controls now visible there. The agent confirms what its click did and finds the next target without a separate desktop_state + screenshot. On visual-only targets it is on by default for a visible change (returnCapture:"on-change"); pass returnCapture:"never" to suppress it, or "always" to force it. Never attached on structured targets (browser/CDP, UIA-rich native), where desktop_state is cheaper and exact — so those responses are unchanged.
  • LLM-native design — Built around how LLMs think, not how humans click. run_macro batches multiple operations into a single API call; diffMode sends only the windows that changed since the last frame. Minimal tokens, minimal round-trips.
  • Reactive Perception Graph — Register a lensId for a window or browser tab, pass it to action tools, and get guard-checked post.perception feedback after each action. It reduces repeated screenshot / desktop_state calls and prevents wrong-window typing or stale-coordinate clicks.
  • Full CJK support — Uses Win32 GetWindowTextW for window titles, avoiding nut-js garbling. IME bypass input supported for Japanese/Chinese/Korean environments.
  • 3-tier token reductiondetail="image" (~443 tok) / detail="text" (~100–300 tok) / diffMode=true (~160 tok). Send pixels only when you actually need to see them.
  • 1:1 coordinate modedotByDot=true captures at native resolution (WebP). Image pixel = screen coordinate — no scale math needed. With origin+scale passed to mouse_click, the server converts coords for you — eliminating off-by-one / scale bugs.
  • Browser capture data reductiongrayscale=true (~50% size), dotByDotMaxDimension=1280 (auto-scaled with coord preservation), and windowTitle + region sub-crops help exclude browser chrome and other irrelevant pixels. Typical reduction for heavy captures: 50–70%.
  • Chromium smart fallbackdetail="text" on Chrome/Edge/Brave auto-skips UIA (prohibitively slow there) and runs Windows OCR. hints.chromiumGuard + hints.ocrFallbackFired flag the path taken.
  • UIA element extractiondetail="text" returns button names and clickAt coords as JSON. Claude can click the right element without ever looking at a screenshot.
  • Auto-dock CLIwindow_dock(action='dock') snaps any window to a screen corner with always-on-top. Set DESKTOP_TOUCH_DOCK_TITLE='@parent' to auto-dock the terminal hosting Claude on MCP startup — the process-tree walker finds the right window regardless of title.
  • Emergency stop (Failsafe) — Move the mouse to the top-left corner (within 10px of 0,0) to immediately terminate the MCP server.

Requirements

| | | |---|---| | OS | Windows 10 / 11 (64-bit) | | Node.js | v20+ recommended (tested on v22+) | | PowerShell | 5.1+ (bundled with Windows) — used only as fallback when the Rust native engine is unavailable | | Claude CLI | claude command must be available |

> Note: nut-js native bindings require the Visual C++ Redistributable. > Download from Microsoft if not already installed.

> Note (Key Locker): The credential helper Key Locker uses is an unsigned executable, so on > some machines Windows SmartScreen or antivirus may show an "unknown publisher" warning the > first time it runs. This is expected — the helper ships with desktop-touch-mcp and runs locally > on your machine; you can allow it to proceed. Code signing is planned for a future release.


Installation

npx -y @harusame64/desktop-touch-mcp

The npm launcher resolves runtime strictly by npm package version. For package X.Y.Z, it fetches only GitHub Release tag vX.Y.Z, downloads desktop-touch-mcp-windows.zip, verifies its SHA256 digest, and only then expands it under %USERPROFILE%\.desktop-touch-mcp. Verified cached releases are reused on later runs.

Set DESKTOP_TOUCH_MCP_HOME to override the cache root directory.

> On a shared or CI network? The first run reads the GitHub Releases API to > locate the runtime zip. The anonymous limit is 60 requests/hour per IP, which a > shared public address (CI runners, office NAT) can exhaust before your download > even starts. Set GITHUB_TOKEN (or GH_TOKEN) in the environment and the > launcher authenticates the request, raising the limit to 5,000 requests/hour. > No token is needed on an ordinary home connection.

> Running the launcher from a source checkout? A source build's > bin/launcher.js carries a placeholder integrity hash (sha256: "PENDING") > instead of a finalized one. Rather than download and run an unverified runtime, > the launcher fails closed — this guard stops an accidentally published or > unfinalized launcher from silently starting unverified code. Published npm > releases always ship a real SHA256, so end users never see this. If you are > intentionally running the launcher from source, set > DESKTOP_TOUCH_MCP_ALLOW_UNVERIFIED=1 to skip integrity verification > (development only).

Register with Claude CLI

Add to ~/.claude.json under mcpServers:

{
  "mcpServers": {
    "desktop-touch": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@harusame64/desktop-touch-mcp"]
    }
  }
}

No system prompt needed. The command reference is automatically injected into Claude via the MCP initialize response's instructions field.

Register with other clients (HTTP mode)

Clients that require an HTTP endpoint (GPT Desktop, VS Code Copilot, Cursor, etc.) can use the built-in Streamable HTTP transport:

npx -y @harusame64/desktop-touch-mcp --http
# or with a custom port:
npx -y @harusame64/desktop-touch-mcp --http --port 8080

The server starts at http://127.0.0.1:23847/mcp (localhost only). Register the URL in your MCP client settings. A health check is available at http://127.0.0.1:/health.

In HTTP mode the system tray icon shows the active URL and provides quick-copy and open-in-browser shortcuts.

Development install

git clone https://github.com/Harusame64/desktop-touch-mcp.git
cd desktop-touch-mcp
npm install

Build after install:

npm run build

For a local checkout, register the built server directly:

{
  "mcpServers": {
    "desktop-touch": {
      "type": "stdio",
      "command": "node",
      "args": ["D:/path/to/desktop-touch-mcp/dist/index.js"]
    }
  }
}

> Note: Replace D:/path/to/desktop-touch-mcp with the actual path where you cloned this repository.


Tools (32 Optimized Tools)

> 📖 Full Reference: [docs/system-overview.md](docs/system-overview.md) — Exhaustive guide on parameters, return schemas, and coordinate math.

🌐 World-Graph V2 (Primary Path)

| Tool | Description | |---|---| | desktop_discover | Observe the desktop. Returns interactive entities with leases (UIA, CDP, Terminal, Visual SoM). | | desktop_act | Perform actions (click, type, drag, select) on entities via lease validation. Returns semantic diffs — plus an optional roiCapture (changed-region PNG + next-target preview) on visual-only targets. |

👁️ Observation & State

| Tool | Description | |---|---| | desktop_state | Lightweight check of focus, active window, cursor, and Auto-Perception attention signal. | | screenshot | Multi-mode capture: detail='text' (UIA/OCR), diffMode (P-frame), dotByDot (1:1), and background. Returns a cheap screenshot://by-ref/{id} link to the saved image instead of inlining pixels every time. | | screenshot_query / screenshot_gc | Inspect and prune the on-disk screenshot cache behind the by-ref links: screenshot_query lists saved captures without re-reading pixels; screenshot_gc reclaims space by retention policy (dry-run by default). | | workspace_snapshot | Instant session orientation: all window thumbnails + UI summaries in one call. | | server_status | Diagnostic check for native engine health and feature activation. |

⌨️ Input & Control

| Tool | Description | |---|---| | keyboard | Send keyboard input. Supports background input (WM_CHAR) and IME-safe clipboard bypass. | | mouse_click / mouse_drag | Precision coordinate-based interaction with homing and force-focus protection. | | scroll | Multi-strategy: raw (notches), to_element, smart (virtual lists), and capture (stitch). | | click_element | Legacy UIA-based click by name/ID (fallback when entities are unavailable). |

🌐 Browser CDP (Chrome/Edge/Brave)

| Tool | Description | |---|---| | browser_open / browser_navigate | Idempotent debug-mode launch and reliable navigation. | | browser_click / browser_fill / browser_form | High-level DOM interaction stable across repaints and framework re-renders. | | browser_eval | Deep inspection via js (scripting), dom (HTML), and appState (SPA data extraction). | | browser_overview / browser_search / browser_locate | Semantic discovery, grep-like DOM search, and pixel-accurate coordinate lookup. |

🛠️ Utilities & Workflow

| Tool | Description | |---|---| | terminal | Unified command execution: run (send + wait + read), read (OCR/UIA), and send. run completion modes: quiet, pattern, and exit (waits for the command to finish + returns its exit code — see [Terminal command completion](#terminal-command-completion-until)). | | wait_until | Efficient server-side polling for window, focus, text, or URL state changes. | | window_dock / focus_window | Window management: pin (always-on-top), unpin, dock (corner snap), and focus. | | workspace_launch | Launch apps and auto-detect new HWNDs (supports localized titles). | | run_macro | Batch up to 50 operations into a single round-trip for maximum efficiency. | | clipboard / notification_show | System-level text exchange and user alerts. | | key_locker | Manage credentials the terminal autofills for you (SSH key passphrases, sudo / login passwords). Secrets are entered once into the locker's own secure dialog and stored encrypted on this machine (Windows DPAPI); they are never shown to the assistant. action='launch_console' opens an autofill-capable console (returns a paneId to drive ssh/sudo into via terminal); save / list / forget / set_policy / status manage bindings. Autofill only fires in a console opened by launch_console. Disable with DESKTOP_TOUCH_DISABLE_KEY_LOCKER=1. |

📊 Office (Excel)

| Tool | Description | |---|---| | excel | Author and run Excel VBA macros via COM. action='run_vba' writes a macro into a managed Trusted Location and runs it; action='check_access_vbom' is a read-only preflight. Runs VBA where formula-only tools cannot. One-time setup: node scripts/enable-access-vbom.mjs. |


Standard workflow (v1.0.0)

The v2 World-Graph surface (desktop_discover / desktop_act) is the recommended dispatch path. The four-call shape works for native apps, browsers, and terminals identically.

desktop_state          → orient: focused window/element, modal, attention signal
desktop_discover       → find actionable entities (returns lease + windows[])
desktop_act(lease, …)  → act on entity (returns attention + post.perception)
desktop_state          → confirm the world changed as expected

Clicking — priority order:

browser_click(selector)               → Chrome / Edge (CDP, stable across repaints)
desktop_act(lease, action='click')    → native / dialog / visual (entity-based; use after desktop_discover)
click_element(name | automationId)    → native UIA fallback if desktop_act returns ok:false
mouse_click(x, y, origin?, scale?)    → pixel last resort; origin+scale from dotByDot screenshots only

Recovery hints — read response.attention after every observation and response.warnings[] on desktop_discover / desktop_act. Common reasons:

  • lease_expired / lease_generation_mismatch / lease_digest_mismatch / entity_not_found → re-call desktop_discover
  • modal_blockingresponse.blockingElement (when present) names the blocking modal; dismiss via click_element(name=blockingElement.name) then retry
  • entity_outside_viewportscroll(action='to_element' | 'raw'), then re-call desktop_discover
  • executor_failed → fall back to click_element / mouse_click / browser_click

Lease lifecycle:

  • Each desktop_discover response carries softExpiresAtMs (≈ 60 % of the TTL window). Past that timestamp the LLM should consider re-calling desktop_discover even though the lease is still technically valid — lease.expiresAtMs is the only correctness wall.
  • TTL adapts to view mode (action/explore/debug), entity count, and response payload size. Cap is 60 s.
  • Set DESKTOP_TOUCH_DISABLE_FUKUWARAI_V2=1 to fall back to the v1 tool surface (get_windows / get_ui_elements / `s

Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet — be the first.

Versions

  • v0.13.1 Imported from the upstream source.