Install
$ agentstack add skill-eliasoulkadi-shokunin-agent-browser ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
agent-browser
Fast browser automation CLI for AI agents. Chrome/Chromium via CDP with accessibility-tree snapshots and compact @eN element refs.
Install: npm i -g agent-browser && agent-browser install
Workflow
Step 1: Navigate to a page
agent-browser goto "https://example.com"
agent-browser wait @e3 # wait for element to appear
agent-browser html # get current DOM snapshot
The goto command navigates to a URL and waits for the page to load. Use wait to ensure specific elements are present before interacting. Always snapshot with html after navigation to get fresh @eN refs.
Step 2: Interact with elements
agent-browser click @e5 # click element by accessibility ref
agent-browser type @e7 "hello world" # type into input
agent-browser select @e9 "option-value" # select dropdown option
agent-browser check @e12 # toggle checkbox/radio
agent-browser hover @e15 # hover over element
agent-browser press Enter # press keyboard key
agent-browser scroll down 500 # scroll page by pixels
agent-browser drag @e5 @e20 # drag element to target
agent-browser upload @e8 "C:\path\to\file.pdf" # file input upload
All interactions use accessibility-tree element references (@eN). These are stable within a single page snapshot but invalidate after navigation or DOM mutations.
Step 3: Extract data or capture evidence
agent-browser screenshot page.png # full-page screenshot
agent-browser screenshot --element @e5 el.png # element screenshot
agent-browser screenshot --full-page false vp.png # viewport-only screenshot
agent-browser text # visible text content
agent-browser html # full HTML source
agent-browser pdf output.pdf # generate PDF
agent-browser console # browser console logs
agent-browser network # network request log
agent-browser eval "document.querySelector('.price').innerText" # run JS
agent-browser performance # Core Web Vitals metrics
agent-browser accessibility # accessibility tree dump
Step 4: Fill and submit forms
agent-browser goto "https://example.com/login"
agent-browser type @e3 "user@example.com" # email field
agent-browser type @e5 "password123" # password field
agent-browser click @e7 # submit button
agent-browser wait @e10 # wait for post-login element
agent-browser screenshot logged-in.png # verify success
For complex multi-step forms, screenshot after each step for debugging. Use agent-browser html between steps to get updated element refs if the DOM changes.
Step 5: Handle authentication
agent-browser vault set example.com "user:pass" # store encrypted credentials
agent-browser vault get example.com # retrieve credentials
agent-browser auth login "https://example.com" # automated OAuth/login flow
agent-browser session save mysession # save cookies + localStorage
agent-browser session load mysession # restore session state
agent-browser session list # list saved sessions
Credentials are encrypted at rest. Sessions persist cookies, localStorage, and IndexedDB across runs. Use sessions to skip repeated login flows.
Step 6: Verify and debug
agent-browser console --errors # only error-level logs
agent-browser console --since 30s # recent console output
agent-browser network --status 4xx,5xx # failed requests
agent-browser network --host api.example.com # filter by host
agent-browser cookies # inspect cookies
agent-browser eval "document.title" # run JS in page context
agent-browser performance # Core Web Vitals metrics
Common Workflow Patterns
Navigate and interact
agent-browser goto
agent-browser click @e # click element by accessibility ref
agent-browser type @e "text" # type into element
agent-browser select @e "option" # select from dropdown
Extract data
agent-browser screenshot # full page screenshot
agent-browser html # get current page HTML
agent-browser text # get visible text content
agent-browser pdf # generate PDF
Browser state
agent-browser console # get browser console logs
agent-browser network # get network requests
agent-browser cookies # get/set cookies
agent-browser session save # save session state
agent-browser session load # restore session
Authentication
agent-browser vault set # store encrypted credentials
agent-browser vault get # retrieve and auto-fill
agent-browser auth login # automated login flow
Load specialized workflows
agent-browser skills get core # base workflow (always load first)
agent-browser skills get core --full # full command reference
agent-browser skills get electron # Electron desktop apps (VS Code, Slack, Discord, Figma)
agent-browser skills get slack # Slack workspace automation
agent-browser skills get dogfood # Exploratory testing / QA / bug hunts
agent-browser skills get vercel-sandbox # Inside Vercel Sandbox microVMs
agent-browser skills get agentcore # AWS Bedrock AgentCore cloud browsers
agent-browser skills list # all available workflows
Error Handling
| Error | Cause | Fix | |-------|-------|-----| | Element @eN not found | Element removed from DOM or re-rendered | Re-run agent-browser html to get fresh refs and retry | | Element @eN not interactable | Element hidden, covered, or disabled | Use agent-browser html to check visibility; scroll into view first | | Page load timeout | Network slow, infinite spinner, or redirect loop | Increase timeout with --timeout 30000; check agent-browser console for JS errors | | Navigation failed | Invalid URL, DNS failure, or certificate error | Verify URL format; check network with agent-browser network | | Authentication failed | Expired session, changed credentials, CAPTCHA | Re-login manually with agent-browser session save; update vault credentials | | Browser not found | Chrome/Chromium not installed | Run agent-browser install to download compatible Chromium | | Port already in use | Another instance of agent-browser running | Kill existing process or use --port to pick a different CDP port | | Screenshot black/empty | Page not fully rendered or GPU issue | Wait for visible element first; try --full-page false for viewport only | | Stale element reference | DOM mutated between snapshot and interaction | Always re-snapshot with agent-browser html before interacting | | Rate limiting / bot detection | Site detects automation | Use agent-browser session load with real browser fingerprint; add delays between actions |
Anti-Patterns
| Pattern | Problem | Fix | |---------|---------|-----| | Hardcoding @eN refs across pages | Refs change on every navigation | Always snapshot after navigation, never reuse stale refs | | Clicking without waiting | Element may not be ready | Use agent-browser wait @eN before interacting | | Blindly typing into any field | May type into wrong or readonly field | Verify element type from snapshot, use proper type/select/check | | Ignoring console errors | Hidden JS failures break interactions | Always check agent-browser console --errors after issues | | Forgetting to save session | Re-authentication needed on every run | agent-browser session save after manual login | | Running without headless mode check | Visual inspection needed for debugging | Use --headless false when debugging; screenshot after each step | | Bulk-extracting without rate limits | Triggering bot protection and IP bans | Add --delay 500 between actions; respect robots.txt | | Using XPath or CSS selectors directly | Fragile, breaks on minor DOM changes | Prefer accessibility-tree @eN refs; fall back to text content matching |
Troubleshooting Quick Reference
| Issue | First Command to Run | |-------|---------------------| | Element not found | agent-browser html | | Page not loading | agent-browser console | | Auth failing | agent-browser session load | | Weird layout | agent-browser screenshot debug.png | | Network errors | agent-browser network --status 4xx,5xx | | Performance issues | agent-browser performance | | Blank page after goto | agent-browser console --errors; agent-browser eval "document.readyState" | | Click does nothing | agent-browser html (check element type, visibility, disabled state) |
Why agent-browser
- Fast native Rust CLI, not a Node.js wrapper
- Works with any AI agent (Cursor, Claude Code, Codex, Continue, Windsurf, etc.)
- Chrome/Chromium via CDP with no Playwright or Puppeteer dependency
- Accessibility-tree snapshots with element refs for reliable interaction
- Sessions, authentication vault, state persistence, video recording
Checklist
- [ ] Page loaded and rendered without errors before interacting
- [ ] Selectors use stable attributes (data-testid, aria-label) not dynamic classes
- [ ] Screenshots taken at each meaningful step for debugging
- [ ] Timeout set per-action (default 30s) to avoid indefinite hangs
- [ ] Browser context closed in finally block to prevent orphan processes
Sources
- Accessibility Tree specification (w3.org/TR/accname-1.2)
- agent-browser CLI (npm: agent-browser)
- Vercel Sandbox microVMs (vercel.com/docs/sandbox)
- AWS Bedrock AgentCore (docs.aws.amazon.com/bedrock)
- Electron CDP support (electronjs.org/docs/latest/api/debugger)
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: EliasOulkadi
- Source: EliasOulkadi/shokunin
- License: MIT
- Homepage: https://eliasoulkadi.github.io/shokunin/
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.