AgentStack
MCP verified MIT Self-run

Ui Bridge Mcp

mcp-cholamb-ui-bridge-mcp · by cholamb

Computer Use for Windows - without the screenshots. Element-based MCP server: UIA + Win32 + CDP, 3-tier fallback ladder, ~ms per command.

No reviews yet
0 installs
6 views
0.0% view→install

Install

$ agentstack add mcp-cholamb-ui-bridge-mcp

✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

Are you the author of Ui Bridge Mcp? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

UiBridge MCP

Computer Use for Windows — without the screenshots.

[English](README.md) | [한국어](README.ko.md)

Your AI clicks buttons by name, not by guessing pixels on a screenshot — so it's fast (~ms per command), never misclicks when a window moves, and your screen never leaves your machine.

An MCP server that lets Claude Code (or any MCP client) control native Windows apps and web pages by element — not by pixel guessing. Interactions go through the Windows Accessibility API (UI Automation), Win32 messages, and the Chrome DevTools Protocol, so everything stays local on your machine.

Typical vision-based computer use takes a screenshot at every step and lets the model guess coordinates — slow, token-hungry, and fragile when the window moves or resizes. UiBridge does the opposite:

  • UI Automation (UIA) — click/type by element name or AutomationId
  • Chrome DevTools Protocol (CDP) — drive web page DOM with CSS selectors
  • Screen content is never sent to the model — only the text of the elements you ask for

The 3-tier fallback ladder

| Tier | Method | Covers | |---|---|---| | 1 | UIA / CDP (default) | Most apps that expose accessibility, all web pages | | 2 | Win32 messages (win32_find / win32_set_text / win32_click) | Apps that hide from UIA but still reuse standard Edit/Button child windows or named custom sub-views — no focus steal, works on background windows | | 3 | Annotated screenshot + coordinates (screenshot_window + click_at) | Purely custom-rendered surfaces — UIA element rectangles are numbered on the image (Set-of-Marks); the model picks a number, the code computes the click point |

Plus run_steps — batch multiple operations (with verify_window / wait_ms checkpoints) into a single call.

Install

python -m pip install -r requirements.txt
python install.py        :: registers the server in ~/.claude/settings.json
:: restart Claude Code — the ui-bridge tools appear

For web automation, a browser must be running in CDP debug mode. Easiest: just ask Claude to "launch the browser" — the launch_browser tool starts one automatically (isolated profile, so it won't clash with your normal Chrome). Or start it yourself:

start_browser_debug.bat

Verify the install with the self-test harnesses (uses Notepad):

python tests\uia_harness.py
python tests\stage3_harness.py
python tests\win32_harness.py

Requirements: Windows 10/11, Python 3.10+, Chrome or Edge (for web tools).

Usage examples

Once registered, just talk to Claude Code — it picks the right tools. Below, each prompt is followed by what actually happens under the hood.

1. Drive a desktop app

> "Open Notepad and type 'hello world' into it, then read it back."

inspect_tree(window_title="Notepad")            → finds the edit area (automation_id=15)
set_value(window_title="Notepad", automation_id="15", value="hello world")
get_value(window_title="Notepad", automation_id="15")   → "hello world"

No screenshot was taken, no coordinates were guessed. Works even if the window is resized, moved, or behind other windows.

2. Fill a web form

> "Log in to the admin page on the debug browser."

web_goto(url="https://example.com/login")        → waits for readyState=complete
web_fill(form_data={"#username": "admin", "#password": "..."})
web_click_element(selector="button[type=submit]")
web_wait(selector=".dashboard", timeout_ms=15000)

3. An app that hides from UIA (custom framework)

Some apps (e.g. chat apps built on custom renderers) expose almost nothing to UI Automation. Fall to tier 2:

win32_find(window_title="MyChatApp", class_name="Edit", visible_only=false)
   → [{hwnd: 132002, class: "Edit", size: [326, 28], ...}]
win32_set_text(hwnd=132002, text="search keyword")    → WM_SETTEXT, no focus steal
win32_find(window_title="MyChatApp", text_re="ChatRoomList")
   → even custom sub-views are Win32 child windows with names and rectangles

4. Purely custom-rendered surface (tier 3)

screenshot_window(window_title="MyChatApp", annotate=true)
   → {path: "...png", elements: [{n: 3, name: "Send", center: [1240, 890]}, ...]}
# the model looks at the numbered overlay, picks #3:
click_at(x=1240, y=890)

The model never guesses pixels — it picks an element number, the click point comes from the UIA rectangle.

5. Batch a known flow into one call

run_steps(steps=[
  {"tool": "set_value",     "args": {"window_title": "Notepad", "automation_id": "15", "value": "report done"}},
  {"tool": "wait_ms",       "args": {"ms": 200}},
  {"tool": "get_value",     "args": {"window_title": "Notepad", "automation_id": "15"}},
  {"tool": "verify_window", "args": {"window_title": "Notepad"}}
])

Stops at the first failing step and returns per-step results, so the agent can inspect and resume from exactly where it broke.

Performance

Connections are cached instead of re-created per command (auto-reconnect on drop):

  • Web element read: ~2.6 ms/call (vs seconds when handshaking per call)
  • Window connect: ~0.9 s first → ~5 ms cached
  • CDP port auto-discovery: if nothing listens on the default 9222, known

fallback ports are scanned (UIBRIDGE_CDP_FALLBACK_PORTS env var)

Tools (23)

  • Discovery: list_windows, inspect_tree
  • Interaction: click_element, type_text, get_value, set_value
  • Bookmarks / actions: bookmark_element, list_bookmarks, delete_bookmark,

list_actions, run_action — save frequently used elements by name, run parameterised multi-step sequences

  • Web (CDP): launch_browser (start a debug browser — call this first),

web_tabs, web_goto, web_page_info, web_click_element, web_type_text, web_read_text, web_find_elements, web_run_js, web_fill, web_wait

  • Win32 low-level (fallback tier 2): win32_find, win32_get_text,

win32_set_text, win32_click, win32_key

  • Screen & coordinates (fallback tier 3): screenshot_window, click_at,

drag, scroll_at, send_keys

  • Batch: run_steps

Config files

config/ ships with working examples (Calculator, Excel):

  • config/bookmarks.json — saved UI element bookmarks
  • config/apps.json — per-app locator hints
  • config/actions.json — parameterised multi-step action sequences

CLI

python cli.py --help

Runs the same operations without an MCP client — handy for scheduled tasks and deterministic scripts.

Troubleshooting

Web tools fail with "Cannot connect to browser CDP". The web_* tools need a browser running in CDP debug mode. Fix it either way:

  • Ask Claude to "launch the browser" (runs the launch_browser tool), or
  • Run start_browser_debug.bat in the ui-bridge folder.

I started Chrome with --remote-debugging-port but it still won't connect. This is the most common trap: if a normal Chrome window is already open, Chrome forwards the flag to that existing instance and no debug port opens. Use launch_browser / start_browser_debug.bat (they launch an isolated profile that sidesteps this), or close all Chrome windows first.

The debug browser isn't logged into my sites. It uses a separate profile (that's what avoids the clash above). Just log in once inside that window — the profile persists across runs.

Native (non-web) tools can't find my app. Run inspect_tree(window_title=...) to see what's exposed. If it's nearly empty, the app hides from UI Automation — use the win32_* tools (win32_find with visible_only=false), and screenshot_window as a last resort.

Security

  • No screenshots by default — tier 3 only saves to local temp and returns a path
  • Every interaction is a local API call; nothing leaves your machine
  • No hardcoded secrets, no telemetry

Contributing

Issues and PRs are welcome — especially reports of apps where the fallback ladder fails (attach inspect_tree / win32_find output if you can).

License

MIT © cholamb

Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet — be the first.

Versions

  • v0.1.0 Imported from the upstream source.