Install
$ agentstack add mcp-phnx-labs-agents-cli ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ● Network access Used
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ● Environment & secrets Used
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
agents
The missing toolchain for CLI coding agents. Run any agent on your existing subscription. Spawn parallel teams in isolated terminals. Schedule routines, drive browsers and Electron apps, and store secrets behind Touch ID — all from one CLI.
Grok Droid
https://agents-cli.sh/demo.mp4
npm install -g @phnx-labs/agents-cli
# or
bun install -g @phnx-labs/agents-cli
Source: github.com/phnx-labs/agents-cli
Also available as ag -- all commands work with both agents and ag.
- [Pin versions per project](#pin-versions-per-project)
- [One config, every agent](#one-config-every-agent)
- [Run any agent](#run-any-agent)
- [Sessions across agents](#sessions-across-agents)
- [Run open models through Claude Code](#run-open-models-through-claude-code)
- [Teams](#teams)
- [Workflows](#workflows)
- [Browser](#browser)
- [Secrets](#secrets)
- [Routines](#routines)
- [PTY](#pty)
- [Portable setup](#portable-setup)
- [Private skills](#private-skills)
- [Security & Privacy](#security--privacy)
- [Compatibility](#compatibility)
- [FAQ](#faq)
Pin versions per project
# This project needs claude@2.0.65 -- newer versions changed tool calling.
agents use claude@2.0.65 -p
# The monorepo uses codex@0.116.0 across the team.
agents use codex@0.116.0 -p
This creates an agents.yaml at the project root:
# agents.yaml (commit this to your repo)
agents:
claude: "2.0.65"
codex: "0.116.0"
Think requirements.txt for CLI coding agents, on steroids. A shim reads agents.yaml from the project root and routes claude / codex / gemini / grok (and others) to the right version automatically. Each version gets its own isolated home -- switching backs up config and re-syncs resources.
agents add claude@2.0.65 # Install a specific version
agents add codex@latest # Install latest
agents add codex@oldest # Install the oldest published version
agents view # See everything installed
One config, every agent
# Set up the Notion MCP server once.
agents install mcp:com.notion/mcp
# It's now registered with Claude Code, Codex, Gemini CLI, and Cursor.
agents mcp list
Skills, slash commands, rules, hooks, and permissions work the same way -- install once in ~/.agents/, synced to every agent's native format automatically.
agents skills add gh:yourteam/python-expert # Knowledge pack -> all agents
agents commands add gh:yourteam/commands # Slash commands -> all agents
agents rules add gh:team/rules # AGENTS.md -> CLAUDE.md, GEMINI.md, .cursorrules
agents permissions add ./perms # Permissions -> auto-converted per agent
Write one AGENTS.md. It becomes CLAUDE.md for Claude Code, GEMINI.md for Gemini CLI, .cursorrules for Cursor.
Run any agent
agents run claude "Find all auth vulnerabilities in src/"
agents run codex "Fix the issues Claude found"
agents run gemini "Write tests for the fixed code"
Each resolves to the project-pinned version with skills, MCP servers, and permissions already synced. Single-typo names auto-correct across every command — agents view cladue resolves to claude, agents add codx@latest to codex.
Rate-limited? Keep working.
# Claude Code hits a rate limit -> Codex picks up automatically. Same project, same config.
agents run claude "refactor auth module" --mode edit --fallback codex,gemini
Multiple accounts? Spread the load.
# Picks the signed-in account you haven't used recently.
agents run claude "summarize recent commits" --strategy balanced
--strategy balanced spreads work across available versions of the same agent -- useful when you have multiple accounts and want to avoid burning through one.
Chain agents
agents run claude "Review PRs merged this week, summarize risks" \
| agents run codex "Write regression tests for the top 3 risks"
Supports plan (read-only) and edit modes, effort levels, JSON output for scripting, and timeout limits.
One protocol, every harness
# Typed event stream instead of raw stdout. Same command, any supported agent.
agents run claude "review this diff" --acp --json
--acp routes through the Agent Client Protocol so you get a unified event stream -- agent_message_chunk, tool_call, plan_update, stop_reason -- instead of writing a parser per CLI. File writes and shell commands flow through agents-cli, which means --mode plan becomes a real sandbox: the write RPC is denied, not just unused.
ACP adapters are documented for claude, codex, gemini, cursor, opencode, openclaw, and grok. Other harnesses keep running on the direct-exec path.
Sessions across agents
When you run multiple agents, conversations scatter across tools. Session search brings them together.
# Where was that auth conversation? Search Claude Code, Codex, Gemini CLI, OpenCode at once.
agents sessions "auth middleware"
# Filter by agent, project, or time window
agents sessions --agent codex --since 7d
agents sessions --project my-app
# Read a full conversation
agents sessions a1b2c3d4 --markdown
# Just the last 3 turns, user messages only
agents sessions a1b2c3d4 --last 3 --include user
Interactive picker when you're in a terminal. Structured output (--json, --markdown, filtered by role or turn count) when piped.
Backed by a SQLite + FTS5 index at ~/.agents/.history/sessions/sessions.db with incremental scanning -- warm reads in ~100ms. External tools can consume --json output as a programmatic observability layer; see [docs/05-sessions.md](docs/05-sessions.md) for the schema and [docs/06-observability.md](docs/06-observability.md) for the consumption patterns.
Run open models through Claude Code (experimental)
> Note: Profiles are experimental. Enable with agents beta profiles enable.
# Kimi K2.5 responding inside Claude Code's UI, tools, and skills.
# No proxy server. No LiteLLM. One OpenRouter key, stored in Keychain.
agents profiles add kimi
agents run kimi "refactor this file"
Built-in presets (all via OpenRouter, one shared key):
| Preset | Model | Notes | |---|---|---| | kimi | Kimi K2.5 | #1 HumanEval. Reasoning -- interactive only. | | minimax | MiniMax M2.5 | #1 SWE-bench Verified. Reasoning. | | glm | GLM 5 | #1 Chatbot Arena (open-weight). | | qwen | Qwen3 Coder Next | Latest coding Qwen. Print-safe. | | deepseek | DeepSeek Chat V3 | Latest non-reasoning. Print-safe. |
A profile swaps the model while keeping Claude Code as the agent runtime -- same UI, slash commands, skills, MCP tools. Under the hood: ANTHROPIC_BASE_URL + ANTHROPIC_MODEL, auth from Keychain at spawn time.
Custom endpoints (Ollama, vLLM) work too -- drop a YAML in ~/.agents/profiles/:
name: local-qwen
host: { agent: claude }
env:
ANTHROPIC_BASE_URL: https://ollama.example.com
ANTHROPIC_MODEL: qwen3.6:35b
auth:
envVar: ANTHROPIC_AUTH_TOKEN
keychainItem: agents-cli.ollama.token
Profile YAML has no secrets -- safe to agents repo push to a shared repo. agents profiles presets lists the full catalog.
Teams
agents teams create auth-feature
# Research first, then implement, then test.
agents teams add auth-feature claude "Research auth libraries" --name researcher
agents teams add auth-feature codex "Draft the migration" --name migrator --after researcher
agents teams add auth-feature claude "Write tests for the new code" --name tester --after migrator
agents teams start auth-feature # Fires teammates whose deps are done
agents teams status auth-feature # Who's working, what they changed, what they said
Teammates run detached -- close your terminal, they keep working. Check in with teams status, read full output with teams logs , clean up with teams disband.
Team state is observable via agents teams list --json / agents teams status --json (compact by default; add --verbose for the full per-teammate shape). External tools join it with sessions --json (teammates get isTeamOrigin: true) and cloud list --json (for --cloud teammates) to build a unified fleet view. See [docs/06-observability.md](docs/06-observability.md).
Workflows
Bundle an orchestrator prompt with optional subagents, skills, and plugins into a named, reusable pipeline. One bundle, one invocation.
# Use a workflow — workflow name goes in the agent slot
agents run code-review "review PR #42 on acme/api"
# List + inspect
agents workflows list
agents workflows view code-review
# Install from GitHub or local
agents workflows add gh:yourteam/code-review
agents workflows add ./my-workflow
A workflow is a directory:
~/.agents/workflows/code-review/
WORKFLOW.md # YAML frontmatter + orchestrator system prompt
subagents/ # optional: *.md files exposed to the orchestrator
security.md
style.md
skills/ # optional: knowledge packs scoped to this workflow
plugins/ # optional: plugin bundles
WORKFLOW.md's Markdown body is the orchestrator's system prompt. Files under subagents/ get copied to ~/.claude/agents/ at run time so the built-in Agent tool can dispatch to them by name — including in parallel. skills/ and plugins/ sync into the version home just for the run.
# WORKFLOW.md frontmatter
---
name: Code Review
description: Evidence-grounded PR review with file:line citations.
model: opus
tools:
- Read
- Grep
- Bash
- WebFetch
---
Workflows that need to write — post PR comments, edit files, send Slack — should run with --mode edit or --mode full. agents run defaults to --mode plan (read-only), which deadlocks at ExitPlanMode in headless runs.
Resolution is project > user > system: a /.agents/workflows// overrides a same-named workflow in ~/.agents/workflows/. Commit project workflows with your repo so teammates get the same pipeline.
Plugins
Bundle skills, commands, hooks, MCP servers, settings, and permissions under a single manifest. One source dir at ~/.agents/plugins//, mirrored into every installed Claude / OpenClaw version automatically.
# Install from a git URL or local path
agents plugins install hivemind@https://github.com/activeloopai/hivemind.git
agents plugins install ./my-plugin
# Apply to one agent (default version) or all supported
agents plugins sync rush-toolkit claude
agents plugins sync rush-toolkit
A plugin is a directory with a manifest:
~/.agents/plugins/my-plugin/
.claude-plugin/plugin.json # required: { name, version, description }
skills//SKILL.md # optional
commands/*.md # optional
hooks/hooks.json # optional — executable surface
.mcp.json # optional — executable surface
bin/, scripts/, settings.json # optional — executable surface
permissions/ # optional — executable surface
On sync, agents-cli copies the plugin into each version home's marketplace (/.claude/plugins/marketplaces/agents-cli/plugins//), registers the synthetic marketplace, and flips settings.json#enabledPlugins[@agents-cli] = true so Claude / OpenClaw load it.
Executable-surface gate
Plugins that ship hooks/, .mcp.json, bin/, scripts/, settings.json (non-permissions), or permissions/ can execute code on session events. agents-cli requires explicit consent before flipping enabledPlugins:
# Hooks-bearing plugins copy in but stay disabled by default
agents plugins install hivemind@https://github.com/activeloopai/hivemind.git \
--allow-exec-surfaces
# Same gate on re-sync (e.g., after upstream updates)
agents plugins sync hivemind claude --allow-exec-surfaces
Skills, commands, and subagents are declarative and never trip the gate. The gate is per-plugin, per-install: consenting to hivemind doesn't grant blanket exec-surface trust to anything else.
Version portability
Plugins live in the user repo (~/.agents/plugins/), not inside any single version home. Switching Claude via agents use claude@ re-syncs the plugin into the new version automatically — no re-install. New Claude versions added later pick it up on their first sync. Project-level /.agents/plugins// overrides a same-named user plugin (resolution is project > user > system, same as every other resource).
Browser
Give agents access to a real browser — no relay extension, no cloud service, no Playwright getting blocked.
# First run: omit --profile and we auto-pick the first installed Chromium-family
# browser. macOS prefers Chrome > Brave > Edge > Chromium > Comet; Linux prefers
# Chrome > Chromium > Brave > Edge; Windows prefers Edge (always preinstalled) >
# Chrome > Brave. The auto-picked profile is saved as "default" for later runs.
export AGENTS_BROWSER_TASK=$(agents browser start --url https://app.example.com)
# Or pin a named profile to a specific browser (chrome, comet, brave, chromium,
# edge, or custom) when you want isolation from "default".
agents browser profiles create work --browser chrome
# `start` writes the resolved name (e.g. `swift-crab-falcon-a3f92b1c`) to stdout
# and human-friendly commentary to stderr, so $(...) capture stays clean.
export AGENTS_BROWSER_TASK=$(agents browser start --profile work --url https://app.example.com)
agents browser refs # Get interactive element refs
agents browser click 42 # Click element ref 42
agents browser type 15 --text "hello" # Type into element ref 15
agents browser screenshot # Smart resizing, token-efficient
agents browser tabs # List tabs open for the current task
agents browser tab focus tab123 # Switch focus to another tab
agents browser done # Close task's tabs when finished
# Need to address a different task in the same shell? Override per call:
agents browser screenshot --task other-flow
Why this works where Playwright fails
Playwright and Puppeteer spin up fresh browser instances with automation flags. Sites like LinkedIn, Google, and most finance apps detect and block them immediately.
agents browser launches your existing residential Chrome (or Brave, Edge, Chromium) on your machine via CDP. Same browser fingerprint, same IP, same everything. Sites can't detect automation because you're using the same browser you'd use manually.
Token-efficient automation
The CLI handles the mechanical work so agents don't burn tokens on low-level browser commands. Screenshots are automatically resized without excessive compression — agents process smaller images while keeping the detail they need to make decisions.
Profile isolation
Multiple agents can run browser tasks simultaneously without stepping on each other. Each profile gets its own user data directory, cookies, and state. One agent logs into your work Slack, another into your personal email — no conflicts, no shared state.
agents browser profiles create work-slack --browser chrome
agents browser profiles create personal-gmail --browser chrome
# Two agents, two profiles, no interference
Safe credential access
Attach a [secrets bundle](#secrets) to a profile. The agent can log in without credentials in plaintext, and every secret access is recorded in the session log.
agents browser profiles create bank --browser chrome --secrets bank-creds
Electron apps
Control Electron apps (Slack, Discord, V
…
Source & license
This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: phnx-labs
- Source: phnx-labs/agents-cli
- License: MIT
- Homepage: https://agents-cli.sh
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.