Install
$ agentstack add mcp-mrgoonie-human-mcp ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ● Network access Used
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ● Environment & secrets Used
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Human MCP 👁️
> Bringing Human Capabilities to AI Agents
Human MCP v2.16.0 is a comprehensive Model Context Protocol server that provides AI coding agents with human-like capabilities including visual analysis, document processing, speech generation, content creation, image editing, browser automation, and advanced reasoning for debugging, understanding, and enhancing multimodal content.
"Human MCP" is a part of ClaudeKit
Features
🎯 Visual Analysis (Eyes) - ✅ Complete (4 tools)
- eyes_analyze: Analyze images, videos, and GIFs for UI bugs, errors, and accessibility
- eyes_compare: Compare two images to find visual differences
- eyesreaddocument: Extract text and data from PDF, DOCX, XLSX, PPTX, and more
- eyessummarizedocument: Generate summaries and insights from documents
✋ Content Generation & Image Editing (Hands) - ✅ Complete (18 tools)
- Image Generation (1 tool): geminigenimage - Generate images from text using Imagen API
- Video Generation (2 tools): geminigenvideo, geminiimageto_video - Create videos with Veo 3.0
- Music Generation (2 tools): minimaxgenmusic, elevenlabsgenmusic - Generate music with vocals
- Sound Effects (1 tool): elevenlabsgensfx - Generate sound effects from text descriptions
- AI Image Editing (5 tools): Gemini-powered editing with inpainting, outpainting, style transfer, object manipulation, composition
- Jimp Processing (4 tools): Local image manipulation - crop, resize, rotate, mask
- Background Removal (1 tool): rmbgremovebackground - AI-powered background removal
- Browser Automation (3 tools): playwrightscreenshotfullpage, playwrightscreenshotviewport, playwrightscreenshotelement - Automated web screenshots
🗣️ Speech Generation (Mouth) - ✅ Complete (4 tools)
- mouth_speak: Convert text to speech with 30+ voices and 24 languages
- mouth_narrate: Long-form content narration with chapter breaks
- mouth_explain: Generate spoken code explanations with technical analysis
- mouth_customize: Test and compare different voices and styles
🧠 Advanced Reasoning (Brain) - ✅ Complete (3 tools)
- mcp__reasoning__sequentialthinking: Native sequential thinking with thought revision
- brainanalyzesimple: Fast pattern-based analysis (problem solving, root cause, SWOT, etc.)
- brainpatternsinfo: List available reasoning patterns and frameworks
- brainreflectenhanced: AI-powered meta-cognitive reflection for complex analysis
Total: 29 MCP Tools Across 4 Human Capabilities
👁️ Eyes (4 tools) - Visual analysis and document processing ✋ Hands (18 tools) - Content generation, image editing, music/SFX, and browser automation 🗣️ Mouth (4 tools) - Speech generation and narration 🧠 Brain (3 tools) - Advanced reasoning and problem solving
Technology Stack
- Google Gemini 2.5 Flash - Vision, document, and reasoning AI
- Gemini Imagen API - High-quality image generation
- Gemini Veo 3.0 API - Professional video generation
- Gemini Speech API - Natural voice synthesis (30+ voices, 24 languages)
- Minimax API - Alternative speech (Speech 2.6), music (Music 2.5), video (Hailuo 2.3)
- ZhipuAI (Z.AI) API - Alternative vision (GLM-4.6V), image (GLM-Image), video (CogVideoX-3)
- ElevenLabs API - Text-to-speech (70+ languages), music generation, sound effects
- Playwright - Browser automation for web screenshots
- Jimp - Fast local image processing
- rmbg - AI-powered background removal (U2Net+, ModNet, BRIAI models)
Supported Providers & Models
Human MCP supports multiple AI providers per capability. Set via per-request provider parameter or environment variable defaults.
Vision (Eyes)
| Provider | Models | Features | Env Var | |----------|--------|----------|---------| | Google Gemini (default) | gemini-2.5-flash, gemini-2.5-pro | Image, video, GIF analysis; document processing | GOOGLE_GEMINI_API_KEY | | ZhipuAI | glm-4.6v, glm-4.6v-flash (FREE) | Image analysis via GLM-4.6V vision | ZHIPUAI_API_KEY |
Image Generation (Hands)
| Provider | Models | Features | Env Var | |----------|--------|----------|---------| | Google Gemini (default) | gemini-2.5-flash-image, gemini-3.1-flash-image-preview | Text-to-image, 14 aspect ratios, 5 styles | GOOGLE_GEMINI_API_KEY | | ZhipuAI | cogview-4-250304 | Text-to-image, CogView-4 | ZHIPUAI_API_KEY |
Video Generation (Hands)
| Provider | Models | Features | Env Var | |----------|--------|----------|---------| | Google Gemini (default) | veo-3.0-generate-001 | Text-to-video, image-to-video, 4-12s, camera controls | GOOGLE_GEMINI_API_KEY | | Minimax | MiniMax-Hailuo-2.3, MiniMax-Hailuo-2.3-Fast | Text-to-video, image-to-video, 768P/1080P | MINIMAX_API_KEY | | ZhipuAI | cogvideox-3 | Text-to-video, image-to-video, async polling | ZHIPUAI_API_KEY |
Speech / TTS (Mouth)
| Provider | Models | Features | Env Var | |----------|--------|----------|---------| | Google Gemini (default) | gemini-2.5-flash-preview-tts, gemini-2.5-pro-preview-tts | 31 voices, 24 languages, style prompts | GOOGLE_GEMINI_API_KEY | | Minimax | speech-2.6-hd, speech-2.6-turbo | Emotion control, speed adjustment | MINIMAX_API_KEY | | ElevenLabs | eleven_v3, eleven_multilingual_v2, eleven_flash_v2_5, eleven_turbo_v2_5 | 70+ languages, voice cloning, stability/style controls | ELEVENLABS_API_KEY |
Music Generation (Hands)
| Provider | Models | Features | Env Var | |----------|--------|----------|---------| | Minimax | music-2.5 | Vocals from lyrics, structure tags, MP3/WAV | MINIMAX_API_KEY | | ElevenLabs | Music API | Instrumental mode, 3s-10min, prompt-based | ELEVENLABS_API_KEY |
Sound Effects (Hands)
| Provider | Models | Features | Env Var | |----------|--------|----------|---------| | ElevenLabs | Sound Generation API | Text-to-SFX, 0.5-30s, looping, prompt influence | ELEVENLABS_API_KEY |
Provider Configuration
# Default provider per capability (optional, all default to "gemini")
SPEECH_PROVIDER=gemini # Options: gemini, minimax, elevenlabs
VIDEO_PROVIDER=gemini # Options: gemini, minimax, zhipuai
VISION_PROVIDER=gemini # Options: gemini, zhipuai
IMAGE_PROVIDER=gemini # Options: gemini, zhipuai
Or override per request:
{ "provider": "minimax", "prompt": "..." }
Google Gemini Documentation
- Gemini API
- Gemini Models
- Video Understanding
- Image Understanding
- Document Understanding
- Audio Understanding
- Speech Generation
- Image Generation
- Video Generation
Quick Start
Getting Your Google Gemini API Key
Before installation, you'll need a Google Gemini API key to enable visual analysis capabilities.
Step 1: Access Google AI Studio
- Visit Google AI Studio in your web browser
- Sign in with your Google account (create one if needed)
- Accept the terms of service when prompted
Step 2: Create an API Key
- In the Google AI Studio interface, look for the "Get API Key" button or navigate to the API keys section
- Click "Create API key" or "Generate API key"
- Choose "Create API key in new project" (recommended) or select an existing Google Cloud project
- Your API key will be generated and displayed
- Important: Copy the API key immediately as it may not be shown again
Step 3: Secure Your API Key
⚠️ Security Warning: Treat your API key like a password. Never share it publicly or commit it to version control.
Best Practices:
- Store the key in environment variables (not in code)
- Don't include it in screenshots or documentation
- Regenerate the key if accidentally exposed
- Set usage quotas and monitoring in Google Cloud Console
- Restrict API key usage to specific services if possible
Step 4: Set Up Environment Variable
Configure your API key using one of these methods:
Method 1: Shell Environment (Recommended)
# Add to your shell profile (.bashrc, .zshrc, .bash_profile)
export GOOGLE_GEMINI_API_KEY="your_api_key_here"
# Reload your shell configuration
source ~/.zshrc # or ~/.bashrc
Method 2: Project-specific .env File
# Create a .env file in your project directory
echo "GOOGLE_GEMINI_API_KEY=your_api_key_here" > .env
# Add .env to your .gitignore file
echo ".env" >> .gitignore
Method 3: MCP Client Configuration You can also provide the API key directly in your MCP client configuration (shown in setup examples below).
Step 5: Verify API Access
Test your API key works correctly:
# Test with curl (optional verification)
curl -H "Content-Type: application/json" \
-d '{"contents":[{"parts":[{"text":"Hello"}]}]}' \
-X POST "https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash:generateContent?key=YOUR_API_KEY"
Choosing Your Gemini Provider
Human MCP supports two ways to access Google's Gemini models:
Option 1: Google AI Studio (Default - Recommended for Getting Started)
Best for: Quick start, development, prototyping
Setup:
- Get your API key from Google AI Studio
- Set environment variable:
export GOOGLE_GEMINI_API_KEY=your_api_key
Pros:
- Simple setup (just one API key)
- No GCP account required
- Free tier available
- Perfect for development and testing
Option 2: Vertex AI (Recommended for Production)
Best for: Production deployments, enterprise use, GCP integration
Setup:
- Create a GCP project and enable Vertex AI API
- Set up authentication:
```bash # Option A: Application Default Credentials (for local dev) gcloud auth application-default login
# Option B: Service Account (for production) export GOOGLEAPPLICATIONCREDENTIALS=/path/to/service-account.json ```
- Set environment variables:
``bash export USE_VERTEX=1 export VERTEX_PROJECT_ID=your-gcp-project-id export VERTEX_LOCATION=us-central1 # optional, defaults to us-central1 ``
Pros:
- Enterprise-grade quotas and SLAs
- Better integration with GCP services
- Advanced IAM and security controls
- Usage tracking via Cloud Console
- Better for production workloads
Configuration Example (Claude Desktop):
{
"mcpServers": {
"human-mcp-vertex": {
"command": "npx",
"args": ["@goonnguyen/human-mcp"],
"env": {
"USE_VERTEX": "1",
"VERTEX_PROJECT_ID": "your-gcp-project-id",
"VERTEX_LOCATION": "us-central1"
}
}
}
}
Note: With Vertex AI, you don't need GOOGLE_GEMINI_API_KEY - authentication is handled by GCP credentials.
Vertex AI Authentication Methods
1. Application Default Credentials (ADC) - Best for local development
gcloud auth application-default login
2. Service Account - Best for production
# Download service account JSON from GCP Console
export GOOGLE_APPLICATION_CREDENTIALS=/path/to/service-account.json
3. Workload Identity - Best for GKE deployments
- Automatically configured when running on GKE
- No credentials file needed
- Recommended for Kubernetes deployments
Troubleshooting Vertex AI:
If you encounter authentication errors:
- Verify your GCP project ID:
gcloud config get-value project - Check ADC status:
gcloud auth application-default print-access-token - Ensure Vertex AI API is enabled: Visit Vertex AI Console
- Verify IAM permissions: Your account needs
Vertex AI Userrole
Cost Considerations:
Both Google AI Studio and Vertex AI use the same Gemini models and pricing, but:
- Google AI Studio: Generous free tier for testing
- Vertex AI: Production-grade quotas, better for high-volume usage
Alternative Methods for API Key
Using Google Cloud Console:
- Go to Google Cloud Console
- Create a new project or select existing one
- Enable the "Generative AI API"
- Go to "Credentials" > "Create Credentials" > "API Key"
- Optionally restrict the key to specific APIs and IPs
API Key Restrictions (Recommended):
- Restrict to "Generative AI API" only
- Set IP restrictions if using from specific locations
- Configure usage quotas to prevent unexpected charges
- Enable API key monitoring and alerts
Troubleshooting API Key Issues
Common Problems:
- Invalid API Key: Ensure you copied the complete key without extra spaces
- API Not Enabled: Make sure Generative AI API is enabled in your Google Cloud project
- Quota Exceeded: Check your usage limits in Google Cloud Console
- Authentication Errors: Verify the key hasn't expired or been revoked
Testing Your Setup:
# Verify environment variable is set
echo $GOOGLE_GEMINI_API_KEY
# Should output your API key (first few characters)
Prerequisites
- Node.js v22+ or Bun v1.2+
- Google Gemini API key (configured as shown above)
Development (For Contributors)
If you want to contribute to Human MCP development:
# Clone the repository
git clone https://github.com/human-mcp/human-mcp.git
cd human-mcp
# Install dependencies
bun install
# Copy environment template
cp .env.example .env
# Add your Gemini API key to .env
GOOGLE_GEMINI_API_KEY=your_api_key_here
# Start development server
bun run dev
# Build for production
bun run build
# Run tests
bun test
# Type checking
bun run typecheck
Usage with MCP Clients
Human MCP can be configured with various MCP clients for different development workflows. Follow the setup instructions for your preferred client below.
Claude Desktop
Claude Desktop is a desktop application that provides a user-friendly interface for interacting with MCP servers.
Configuration Location:
- macOS:
~/Library/Application Support/Claude/claude_desktop_config.json - Windows:
%APPDATA%\Claude\claude_desktop_config.json - Linux:
~/.config/Claude/claude_desktop_config.json
Configuration Example:
{
"mcpServers": {
"human-mcp": {
"command": "npx",
"args": ["@goonnguyen/human-mcp"],
"env": {
"GOOGLE_GEMINI_API_KEY": "your_gemini_api_key_here"
}
}
}
}
Claude Code (CLI)
Claude Code is the official CLI for Claude that supports MCP servers for enhanced coding workflows.
Prerequisites:
- Node.js v22+ or Bun v1.2+
- Google Gemini API key
- Claude Code CLI installed
Configuration Methods:
Claude Code offers multiple ways to configure MCP servers. Choose the method that best fits your workflow:
Method 1: Using Claude Code CLI (Recommended)
# Add Human MCP server with automatic configuration
claude mcp add --scope user human-mcp npx @goonnguyen/human-mcp --env GOOGLE_GEMINI_API_KEY=your_api_key_here
# Alternative: Add locally installed version
claude mcp add --scope project human-mcp npx @goonnguyen/human-mcp --env GOOGLE_GEMINI_API_KEY=your_api_key_here
# List configured MCP servers
claude mcp list
# Remove server if needed
claude mcp remove human-mcp
Method 2: Manual JSON Configuration
Configuration Location:
- All platforms:
~/.config/claude/config.json
Configuration Example:
{
"mcpServers": {
"human-mcp": {
"command":
…
## Source & license
This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.
- **Author:** [mrgoonie](https://github.com/mrgoonie)
- **Source:** [mrgoonie/human-mcp](https://github.com/mrgoonie/human-mcp)
- **License:** MIT
- **Homepage:** https://human.goclaw.sh
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.