Install
$ agentstack add skill-sloemo01-hermes-skills-bundle-kimi-webbridge ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ● Network access Used
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Kimi WebBridge
Control the user's real browser (with their login sessions) via a local daemon at http://127.0.0.1:10086.
Tools
| Tool | Args | Returns | Note | |------|------|---------|------| | navigate | url, newTab(bool), group_title | {success, url, tabId} | First call opens a tab — see [Tabs](#tabs-and-the-current-tab). group_title sets the group's visible label | | find_tab | url, active(bool) | {success, url, tabId, borrowed} | Re-select a tab this session opened; active:true borrows the tab the user is viewing — see [Tabs](#tabs-and-the-current-tab) | | snapshot | — | {url, title, tree} with @e refs | Accessibility tree (text) — use this to read page content and locate elements | | click | selector (@e ref or CSS) | {success, tag, text} | Synthetic el.click() | | fill | selector, value | {success, tag, mode} | Works on `/ AND [contenteditable] (ProseMirror/Lexical/Slate). mode is "value" or "contenteditable" | | evaluate | code (supports async/await) | {type, value} | | | cdp | method, params | raw CDP response | Raw chrome.debugger passthrough — what evaluate is to JS, cdp is to CDP. Low-level escape hatch for cases the tools above don't cover | | screenshot | format(png\|jpeg), quality(0-100), optional selector (@e/CSS), optional path | {format, path, sizeBytes, mimeType} | Returns a file path, not base64 — see [Screenshots](#screenshots) | | network | cmd(start\|stop\|list\|detail), filter, requestId | request/response data | | | upload | selector, files(string[]) | {success, fileCount} | | | saveaspdf | paperformat, landscape, scale, printbackground, optional path | {path, sizeBytes, mimeType, pageTitle} | Render current page → PDF, returns a file path — see [Save as PDF](#save-the-current-page-as-pdf) | | listtabs | — | {success, tabs:[{tabId, url, title, active, groupTitle}]} | Inspect tabs in the current session | | closetab | — | {success, closed: bool} | Close the current tab in the session | | close_session | — | {success, closed: int} | Close all tabs in the session — closed` is the count. See [Sessions](#sessions) for when to call |
Tabs and the current tab
Single-tab tools (snapshot, click, fill, screenshot, save_as_pdf) act on the current tab — the one you most recently opened with navigate or selected with find_tab.
- Opening pages: use
newTab:truewhen pages should coexist (comparing, cross-referencing); omit it to send the current tab to a new URL. - Going back to an earlier tab: call
find_tabto make a tab you opened earlier in this session the current one again. Pass the tab's full URL — take it fromlist_tabsor the earliernavigateresult. A bare root domain (kimi.com) may miss awww.kimi.comtab, so prefer the exact URL. By defaultfind_tabsearches only this session's own tabs — it never reaches into the user's other tabs or windows. - Acting on a page the user already has open: pass
active:true("use my open X tab" / "the X page I'm viewing"). It borrows the tab the user is currently viewing (returnsborrowed:true); the borrowed tab is operated in place — it is not pulled into the session's tab group. - If
find_taberrors with "no tab matching … in this session", the page isn't open in this session —navigatewithnewTab:trueinstead.
curl -s -X POST http://127.0.0.1:10086/command \
-d '{"action":"find_tab","args":{"url":"https://www.kimi.com","active":true},"session":"k26-research"}'
Call Format
Every command carries a top-level session naming the current task — see [Sessions](#sessions) below. The examples in later sections omit it only for brevity; in real calls always include it. The command format depends on the user's OS.
macOS / Linux — inline JSON is fine:
curl -s -X POST http://127.0.0.1:10086/command \
-H 'Content-Type: application/json' \
-d '{"action":"navigate","args":{"url":"https://example.com","newTab":true,"group_title":"My task"},"session":"my-task"}'
Windows (PowerShell / cmd) — the shell corrupts non-ASCII characters (Chinese etc.) carried inline in command arguments or pipes; they reach the daemon as ? and the text is unrecoverable. Send every request as a file body instead:
- Write the JSON body to a uniquely-named temp file with your own file-write tool — never with shell
echo/heredoc, which corrupts non-ASCII the same way. Give every request its own filename with a random suffix (e.g.webbridge-req-.json) so concurrent requests never share a file and overwrite each other. - POST the file with
curl.exe— alwayscurl.exe, never barecurl, which Windows PowerShell aliases toInvoke-WebRequest:
curl.exe -s -X POST http://127.0.0.1:10086/command -H "Content-Type: application/json" --data-binary "@$env:TEMP\webbridge-req-.json"
- Delete the temp file as soon as the request returns — don't leave request bodies on disk.
Sessions
One task = one session = one tab group. A session collects every tab the task opens into one tab group, so the user sees a single group for "what the agent is doing right now". Pass it as a top-level field of the request body (not inside args).
- Pick one session name at the task's start, put it on every command, and never switch mid-task — even across different sites. Switching session names per site is the #1 cause of fragmented tab groups.
- Name it after the task, not the site (
camping-research,phone-compare). Use multiple sessions only for genuinely unrelated parallel tasks. group_titleis the human-readable group label — write it in the user's language, on the firstnavigateof the task.- When you create the group (the first
navigateof a task), tell the user once that this task's pages are collected under group «title», and that you'll close them whenever they ask.
# First tab: set session + a human label (in the user's language)
curl -s -X POST http://127.0.0.1:10086/command \
-d '{"action":"navigate","args":{"url":"https://www.kimi.com","newTab":true,"group_title":"K2.6 feature research"},"session":"k26-research"}'
# Another site, same task → same session → joins the same group automatically
curl -s -X POST http://127.0.0.1:10086/command \
-d '{"action":"navigate","args":{"url":"https://www.moonshot.cn","newTab":true},"session":"k26-research"}'
Closing is always user-initiated: call close_session only when the user explicitly asks ("close those", "clear the tabs"). It clears the whole group in one call.
Screenshots
The daemon writes the image to disk and returns {format, path, sizeBytes, mimeType} — never base64, since the model can't read raw image bytes. Take the .path and open it with the Read tool to actually see it.
# Default: PNG of the visible viewport, daemon picks a temp path
curl ... -d '{"action":"screenshot","args":{}}'
# Options (each independent): JPEG quality, element-only via @e/CSS selector, custom output path
curl ... -d '{"action":"screenshot","args":{"format":"jpeg","quality":60}}'
curl ... -d '{"action":"screenshot","args":{"selector":"@e123"}}'
A caller-supplied path is honored verbatim (parent dirs created, existing file overwritten) — use a unique name to avoid clobbering. save_as_pdf follows the same rule.
Prefer snapshot over CSS/JS selectors
snapshot returns interactive elements with @e refs based on semantic role/name. Use them directly with click/fill — they survive CSS class hash changes that break manually-written selectors.
Fall back to evaluate (JS) only when:
- The target has no
@eref in the snapshot - You need attributes not in the snapshot (e.g.,
href) - You need to dispatch complex event sequences, or scroll
Evaluate Tips
- Always use compact
JSON.stringify(data)— never addnull, 2formatting. Indentation and newlines can inflate the response several times over, causing truncation during transmission. evaluatecalls share the page's JS realm — re-declaring the sameconst/letacross two calls throwsSyntaxError. Wrap in an IIFE for a fresh scope:(() => { const x = ...; return x; })().
Text input — use fill
fill (selector = CSS or @e ref, plus the value) works on `/ (returns mode: "value") and on [contenteditable] rich editors — ProseMirror, TipTap, Lexical, Slate, Quill, etc. (returns mode: "contenteditable"`), firing the right input events so the page reacts.
fill is clear-and-insert: existing content is replaced. To append, read the current value via evaluate, concatenate, then fill with the result.
Form submit / special keys
There's no separate "press Enter" tool. To submit a form, click the submit button directly (click on the @e ref or selector). To dispatch a key event programmatically (e.g. Escape to close a modal):
{"action":"evaluate","args":{"code":"document.activeElement.dispatchEvent(new KeyboardEvent('keydown',{key:'Escape',bubbles:true}))"}}
Save the current page as PDF
save_as_pdf renders the current page to PDF and returns the file path. All args optional:
paper_format:letter(default) \|a4\|legal\|a3\|tabloidlandscape:false(default)scale:1.0(default), range[0.1, 2.0]print_background:true(default) — keep background colorspath: caller-supplied output path; if absent, daemon picks a default under OS temp dir using the page title as the filename
path semantics match screenshot: written verbatim, parent dirs auto-created, existing files overwritten.
Decoded PDF cap is 100 MB. Above that the daemon refuses; reduce scale or split the page.
Known limitations
- Sites that strictly check
event.isTrusted(some banking portals, captchas) ignoreclick/fillbecause those fire DOM-level synthetic events (isTrusted=false). For these, tell the user the page needs manual interaction. (Trusted input is possible at the protocol level via thecdpescape hatch, but treat that as advanced.) - Cross-origin iframes:
fill,click,evaluate, andsnapshotoperate on the top frame. If a target element lives in a same-page iframe from a different origin (e.g. embedded sandbox demos), navigate to the iframe's URL directly instead.
If a tool call fails (daemon or extension not ready)
If a tool call can't reach the daemon (connection refused), start it yourself — don't ask the user. This is safe to run anytime: it no-ops if the daemon is already up.
macOS / Linux:
~/.kimi-webbridge/bin/kimi-webbridge start
Windows (PowerShell):
& "$env:USERPROFILE\.kimi-webbridge\bin\kimi-webbridge.exe" start
Cross-platform (any shell):
${HOME:-$USERPROFILE}/.kimi-webbridge/bin/kimi-webbridge${EXE:-} start
Then retry the tool call. If it still fails — or the browser extension won't connect — point the user to the help page instead of deep-troubleshooting:
- English: https://www.kimi.com/features/webbridge
- 中文: https://www.kimi.com/zh-cn/features/webbridge
Never run stop / restart / uninstall automatically — those kill a running daemon. See references/operations.md for anything deeper.
Version mismatches
If a tool returns an error containing "Please update the Kimi WebBridge extension", the user's browser extension is older than this skill. Don't try to reconcile versions yourself — just tell the user, in their language, to update the extension and retry:
- English: https://www.kimi.com/features/webbridge
- 中文: https://www.kimi.com/zh-cn/features/webbridge
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: sloemo01
- Source: sloemo01/hermes-skills-bundle
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.