AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
MCP verified MIT Self-run

Human Mcp

mcp-mrgoonie-human-mcp · by mrgoonie

Bringing Human Capabilities to AI Agents

No reviews yet
0 installs
37 views
0.0% view→install

Install

$ agentstack add mcp-mrgoonie-human-mcp

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets Used
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/mcp-mrgoonie-human-mcp)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
6mo ago

Declared compatibility

Claude CodeClaude DesktopCursorWindsurf

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Human Mcp? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Human MCP 👁️

> Bringing Human Capabilities to AI Agents

Human MCP v2.16.0 is a comprehensive Model Context Protocol server that provides AI coding agents with human-like capabilities including visual analysis, document processing, speech generation, content creation, image editing, browser automation, and advanced reasoning for debugging, understanding, and enhancing multimodal content.

"Human MCP" is a part of ClaudeKit

Features

🎯 Visual Analysis (Eyes) - ✅ Complete (4 tools)

  • eyes_analyze: Analyze images, videos, and GIFs for UI bugs, errors, and accessibility
  • eyes_compare: Compare two images to find visual differences
  • eyesreaddocument: Extract text and data from PDF, DOCX, XLSX, PPTX, and more
  • eyessummarizedocument: Generate summaries and insights from documents

Content Generation & Image Editing (Hands) - ✅ Complete (18 tools)

  • Image Generation (1 tool): geminigenimage - Generate images from text using Imagen API
  • Video Generation (2 tools): geminigenvideo, geminiimageto_video - Create videos with Veo 3.0
  • Music Generation (2 tools): minimaxgenmusic, elevenlabsgenmusic - Generate music with vocals
  • Sound Effects (1 tool): elevenlabsgensfx - Generate sound effects from text descriptions
  • AI Image Editing (5 tools): Gemini-powered editing with inpainting, outpainting, style transfer, object manipulation, composition
  • Jimp Processing (4 tools): Local image manipulation - crop, resize, rotate, mask
  • Background Removal (1 tool): rmbgremovebackground - AI-powered background removal
  • Browser Automation (3 tools): playwrightscreenshotfullpage, playwrightscreenshotviewport, playwrightscreenshotelement - Automated web screenshots

🗣️ Speech Generation (Mouth) - ✅ Complete (4 tools)

  • mouth_speak: Convert text to speech with 30+ voices and 24 languages
  • mouth_narrate: Long-form content narration with chapter breaks
  • mouth_explain: Generate spoken code explanations with technical analysis
  • mouth_customize: Test and compare different voices and styles

🧠 Advanced Reasoning (Brain) - ✅ Complete (3 tools)

  • mcp__reasoning__sequentialthinking: Native sequential thinking with thought revision
  • brainanalyzesimple: Fast pattern-based analysis (problem solving, root cause, SWOT, etc.)
  • brainpatternsinfo: List available reasoning patterns and frameworks
  • brainreflectenhanced: AI-powered meta-cognitive reflection for complex analysis

Total: 29 MCP Tools Across 4 Human Capabilities

👁️ Eyes (4 tools) - Visual analysis and document processing ✋ Hands (18 tools) - Content generation, image editing, music/SFX, and browser automation 🗣️ Mouth (4 tools) - Speech generation and narration 🧠 Brain (3 tools) - Advanced reasoning and problem solving

Technology Stack

  • Google Gemini 2.5 Flash - Vision, document, and reasoning AI
  • Gemini Imagen API - High-quality image generation
  • Gemini Veo 3.0 API - Professional video generation
  • Gemini Speech API - Natural voice synthesis (30+ voices, 24 languages)
  • Minimax API - Alternative speech (Speech 2.6), music (Music 2.5), video (Hailuo 2.3)
  • ZhipuAI (Z.AI) API - Alternative vision (GLM-4.6V), image (GLM-Image), video (CogVideoX-3)
  • ElevenLabs API - Text-to-speech (70+ languages), music generation, sound effects
  • Playwright - Browser automation for web screenshots
  • Jimp - Fast local image processing
  • rmbg - AI-powered background removal (U2Net+, ModNet, BRIAI models)

Supported Providers & Models

Human MCP supports multiple AI providers per capability. Set via per-request provider parameter or environment variable defaults.

Vision (Eyes)

| Provider | Models | Features | Env Var | |----------|--------|----------|---------| | Google Gemini (default) | gemini-2.5-flash, gemini-2.5-pro | Image, video, GIF analysis; document processing | GOOGLE_GEMINI_API_KEY | | ZhipuAI | glm-4.6v, glm-4.6v-flash (FREE) | Image analysis via GLM-4.6V vision | ZHIPUAI_API_KEY |

Image Generation (Hands)

| Provider | Models | Features | Env Var | |----------|--------|----------|---------| | Google Gemini (default) | gemini-2.5-flash-image, gemini-3.1-flash-image-preview | Text-to-image, 14 aspect ratios, 5 styles | GOOGLE_GEMINI_API_KEY | | ZhipuAI | cogview-4-250304 | Text-to-image, CogView-4 | ZHIPUAI_API_KEY |

Video Generation (Hands)

| Provider | Models | Features | Env Var | |----------|--------|----------|---------| | Google Gemini (default) | veo-3.0-generate-001 | Text-to-video, image-to-video, 4-12s, camera controls | GOOGLE_GEMINI_API_KEY | | Minimax | MiniMax-Hailuo-2.3, MiniMax-Hailuo-2.3-Fast | Text-to-video, image-to-video, 768P/1080P | MINIMAX_API_KEY | | ZhipuAI | cogvideox-3 | Text-to-video, image-to-video, async polling | ZHIPUAI_API_KEY |

Speech / TTS (Mouth)

| Provider | Models | Features | Env Var | |----------|--------|----------|---------| | Google Gemini (default) | gemini-2.5-flash-preview-tts, gemini-2.5-pro-preview-tts | 31 voices, 24 languages, style prompts | GOOGLE_GEMINI_API_KEY | | Minimax | speech-2.6-hd, speech-2.6-turbo | Emotion control, speed adjustment | MINIMAX_API_KEY | | ElevenLabs | eleven_v3, eleven_multilingual_v2, eleven_flash_v2_5, eleven_turbo_v2_5 | 70+ languages, voice cloning, stability/style controls | ELEVENLABS_API_KEY |

Music Generation (Hands)

| Provider | Models | Features | Env Var | |----------|--------|----------|---------| | Minimax | music-2.5 | Vocals from lyrics, structure tags, MP3/WAV | MINIMAX_API_KEY | | ElevenLabs | Music API | Instrumental mode, 3s-10min, prompt-based | ELEVENLABS_API_KEY |

Sound Effects (Hands)

| Provider | Models | Features | Env Var | |----------|--------|----------|---------| | ElevenLabs | Sound Generation API | Text-to-SFX, 0.5-30s, looping, prompt influence | ELEVENLABS_API_KEY |

Provider Configuration
# Default provider per capability (optional, all default to "gemini")
SPEECH_PROVIDER=gemini      # Options: gemini, minimax, elevenlabs
VIDEO_PROVIDER=gemini       # Options: gemini, minimax, zhipuai
VISION_PROVIDER=gemini      # Options: gemini, zhipuai
IMAGE_PROVIDER=gemini       # Options: gemini, zhipuai

Or override per request:

{ "provider": "minimax", "prompt": "..." }

Google Gemini Documentation

Quick Start

Getting Your Google Gemini API Key

Before installation, you'll need a Google Gemini API key to enable visual analysis capabilities.

Step 1: Access Google AI Studio
  1. Visit Google AI Studio in your web browser
  2. Sign in with your Google account (create one if needed)
  3. Accept the terms of service when prompted
Step 2: Create an API Key
  1. In the Google AI Studio interface, look for the "Get API Key" button or navigate to the API keys section
  2. Click "Create API key" or "Generate API key"
  3. Choose "Create API key in new project" (recommended) or select an existing Google Cloud project
  4. Your API key will be generated and displayed
  5. Important: Copy the API key immediately as it may not be shown again
Step 3: Secure Your API Key

⚠️ Security Warning: Treat your API key like a password. Never share it publicly or commit it to version control.

Best Practices:

  • Store the key in environment variables (not in code)
  • Don't include it in screenshots or documentation
  • Regenerate the key if accidentally exposed
  • Set usage quotas and monitoring in Google Cloud Console
  • Restrict API key usage to specific services if possible
Step 4: Set Up Environment Variable

Configure your API key using one of these methods:

Method 1: Shell Environment (Recommended)

# Add to your shell profile (.bashrc, .zshrc, .bash_profile)
export GOOGLE_GEMINI_API_KEY="your_api_key_here"

# Reload your shell configuration
source ~/.zshrc  # or ~/.bashrc

Method 2: Project-specific .env File

# Create a .env file in your project directory
echo "GOOGLE_GEMINI_API_KEY=your_api_key_here" > .env

# Add .env to your .gitignore file
echo ".env" >> .gitignore

Method 3: MCP Client Configuration You can also provide the API key directly in your MCP client configuration (shown in setup examples below).

Step 5: Verify API Access

Test your API key works correctly:

# Test with curl (optional verification)
curl -H "Content-Type: application/json" \
     -d '{"contents":[{"parts":[{"text":"Hello"}]}]}' \
     -X POST "https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash:generateContent?key=YOUR_API_KEY"

Choosing Your Gemini Provider

Human MCP supports two ways to access Google's Gemini models:

Option 1: Google AI Studio (Default - Recommended for Getting Started)

Best for: Quick start, development, prototyping

Setup:

  1. Get your API key from Google AI Studio
  2. Set environment variable: export GOOGLE_GEMINI_API_KEY=your_api_key

Pros:

  • Simple setup (just one API key)
  • No GCP account required
  • Free tier available
  • Perfect for development and testing
Option 2: Vertex AI (Recommended for Production)

Best for: Production deployments, enterprise use, GCP integration

Setup:

  1. Create a GCP project and enable Vertex AI API
  2. Set up authentication:

```bash # Option A: Application Default Credentials (for local dev) gcloud auth application-default login

# Option B: Service Account (for production) export GOOGLEAPPLICATIONCREDENTIALS=/path/to/service-account.json ```

  1. Set environment variables:

``bash export USE_VERTEX=1 export VERTEX_PROJECT_ID=your-gcp-project-id export VERTEX_LOCATION=us-central1 # optional, defaults to us-central1 ``

Pros:

  • Enterprise-grade quotas and SLAs
  • Better integration with GCP services
  • Advanced IAM and security controls
  • Usage tracking via Cloud Console
  • Better for production workloads

Configuration Example (Claude Desktop):

{
  "mcpServers": {
    "human-mcp-vertex": {
      "command": "npx",
      "args": ["@goonnguyen/human-mcp"],
      "env": {
        "USE_VERTEX": "1",
        "VERTEX_PROJECT_ID": "your-gcp-project-id",
        "VERTEX_LOCATION": "us-central1"
      }
    }
  }
}

Note: With Vertex AI, you don't need GOOGLE_GEMINI_API_KEY - authentication is handled by GCP credentials.

Vertex AI Authentication Methods

1. Application Default Credentials (ADC) - Best for local development

gcloud auth application-default login

2. Service Account - Best for production

# Download service account JSON from GCP Console
export GOOGLE_APPLICATION_CREDENTIALS=/path/to/service-account.json

3. Workload Identity - Best for GKE deployments

  • Automatically configured when running on GKE
  • No credentials file needed
  • Recommended for Kubernetes deployments

Troubleshooting Vertex AI:

If you encounter authentication errors:

  1. Verify your GCP project ID: gcloud config get-value project
  2. Check ADC status: gcloud auth application-default print-access-token
  3. Ensure Vertex AI API is enabled: Visit Vertex AI Console
  4. Verify IAM permissions: Your account needs Vertex AI User role

Cost Considerations:

Both Google AI Studio and Vertex AI use the same Gemini models and pricing, but:

  • Google AI Studio: Generous free tier for testing
  • Vertex AI: Production-grade quotas, better for high-volume usage
Alternative Methods for API Key

Using Google Cloud Console:

  1. Go to Google Cloud Console
  2. Create a new project or select existing one
  3. Enable the "Generative AI API"
  4. Go to "Credentials" > "Create Credentials" > "API Key"
  5. Optionally restrict the key to specific APIs and IPs

API Key Restrictions (Recommended):

  • Restrict to "Generative AI API" only
  • Set IP restrictions if using from specific locations
  • Configure usage quotas to prevent unexpected charges
  • Enable API key monitoring and alerts
Troubleshooting API Key Issues

Common Problems:

  • Invalid API Key: Ensure you copied the complete key without extra spaces
  • API Not Enabled: Make sure Generative AI API is enabled in your Google Cloud project
  • Quota Exceeded: Check your usage limits in Google Cloud Console
  • Authentication Errors: Verify the key hasn't expired or been revoked

Testing Your Setup:

# Verify environment variable is set
echo $GOOGLE_GEMINI_API_KEY

# Should output your API key (first few characters)

Prerequisites

Development (For Contributors)

If you want to contribute to Human MCP development:

# Clone the repository
git clone https://github.com/human-mcp/human-mcp.git
cd human-mcp

# Install dependencies  
bun install

# Copy environment template
cp .env.example .env

# Add your Gemini API key to .env
GOOGLE_GEMINI_API_KEY=your_api_key_here

# Start development server
bun run dev

# Build for production
bun run build

# Run tests
bun test

# Type checking
bun run typecheck

Usage with MCP Clients

Human MCP can be configured with various MCP clients for different development workflows. Follow the setup instructions for your preferred client below.

Claude Desktop

Claude Desktop is a desktop application that provides a user-friendly interface for interacting with MCP servers.

Configuration Location:

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
  • Windows: %APPDATA%\Claude\claude_desktop_config.json
  • Linux: ~/.config/Claude/claude_desktop_config.json

Configuration Example:

{
  "mcpServers": {
    "human-mcp": {
      "command": "npx",
      "args": ["@goonnguyen/human-mcp"],
      "env": {
        "GOOGLE_GEMINI_API_KEY": "your_gemini_api_key_here"
      }
    }
  }
}
Claude Code (CLI)

Claude Code is the official CLI for Claude that supports MCP servers for enhanced coding workflows.

Prerequisites:

  • Node.js v22+ or Bun v1.2+
  • Google Gemini API key
  • Claude Code CLI installed

Configuration Methods:

Claude Code offers multiple ways to configure MCP servers. Choose the method that best fits your workflow:

Method 1: Using Claude Code CLI (Recommended)

# Add Human MCP server with automatic configuration
claude mcp add --scope user human-mcp npx @goonnguyen/human-mcp --env GOOGLE_GEMINI_API_KEY=your_api_key_here

# Alternative: Add locally installed version
claude mcp add --scope project human-mcp npx @goonnguyen/human-mcp --env GOOGLE_GEMINI_API_KEY=your_api_key_here

# List configured MCP servers
claude mcp list

# Remove server if needed
claude mcp remove human-mcp

Method 2: Manual JSON Configuration

Configuration Location:

  • All platforms: ~/.config/claude/config.json

Configuration Example:

{
  "mcpServers": {
    "human-mcp": {
      "command":

…

## Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [mrgoonie](https://github.com/mrgoonie)
- **Source:** [mrgoonie/human-mcp](https://github.com/mrgoonie/human-mcp)
- **License:** MIT
- **Homepage:** https://human.goclaw.sh

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.