AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
MCP verified MIT Self-run

Vision Mcp Server

mcp-loveacup-vision-mcp-server · by Loveacup

MCP server providing multimodal vision capabilities via OpenAI-compatible API (Qwen3-VL, GPT-4o, etc.)

No reviews yet
0 installs
0 views
view→install

Install

$ agentstack add mcp-loveacup-vision-mcp-server

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets Used
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/mcp-loveacup-vision-mcp-server)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
6mo ago

Declared compatibility

Claude CodeClaude DesktopCursorWindsurf

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Vision Mcp Server? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

👁️ Vision MCP Server

[](LICENSE) [](https://nodejs.org/) [](https://modelcontextprotocol.io/)

Give your AI agent eyes. An MCP server providing multimodal vision capabilities — image analysis, OCR, image comparison, and video analysis — powered by any OpenAI-compatible vision model.

让你的 AI 代理拥有视觉能力。 通过任何 OpenAI 兼容的视觉模型,提供图像分析、OCR 文字识别、图像对比和视频分析。

[Features](#-features) · [Quick Start](#-quick-start) · [Tools](#️-tools-reference) · [Models](#-supported-models) · [中文说明](#-中文说明)


✨ Features

| Tool | Description | |------|-------------| | 🔍 analyze_image | Analyze images with natural language prompts | | 📝 ocr_image | Extract text from images (plain text / Markdown / JSON) | | 🔀 compare_images | Compare 2–4 images side by side | | 🎬 analyze_video | Analyze video content (requires video-capable model) |

Plus:

  • 🌐 OpenAI-compatible — Works with any vision model via standard API
  • 📁 Local files & URLs — Auto-converts local files to base64
  • ⚙️ Configurable — Environment variables, config files, or both

🚀 Quick Start

1. Install

git clone https://github.com/Loveacup/vision-mcp-server.git
cd vision-mcp-server
npm install && npm run build

2. Configure

Create a .env file in the project root:

VISION_BASE_URL=http://your-server:port/v1/chat/completions
VISION_MODEL=Qwen3-VL-32B
VISION_API_KEY=your-api-key    # optional for local models

📄 Or use config.json

{
  "baseUrl": "http://your-server:port/v1/chat/completions",
  "model": "Qwen3-VL-32B",
  "apiKey": "your-api-key",
  "maxTokens": 4096,
  "temperature": 0.7
}

3. Run

npm start

The server communicates over stdio, designed to be launched by an MCP client such as Claude Code.

🔌 Claude Code Integration

Add to your ~/.mcp.json:

{
  "mcpServers": {
    "vision": {
      "command": "node",
      "args": ["/path/to/vision-mcp-server/dist/index.js"],
      "env": {
        "VISION_BASE_URL": "http://your-server:port/v1/chat/completions",
        "VISION_MODEL": "Qwen3-VL-32B",
        "VISION_API_KEY": "your-api-key"
      }
    }
  }
}

> Replace /path/to/vision-mcp-server with the actual install path.

⚙️ Configuration Reference

Configuration priority: environment variables > config file > defaults

| Variable | Config Key | Default | Description | |---|---|---|---| | VISION_BASE_URL | baseUrl | (required) | OpenAI-compatible chat completions endpoint | | VISION_MODEL | model | Qwen3-VL-32B | Model name | | VISION_API_KEY | apiKey | (empty) | API key (optional for local models) | | VISION_MAX_TOKENS | maxTokens | 4096 | Max response tokens | | VISION_TEMPERATURE | temperature | 0.7 | Sampling temperature |

🛠️ Tools Reference

analyze_image

Analyze an image with a vision language model.

| Parameter | Type | Required | Default | Description | |---|---|---|---|---| | image | string | ✅ | — | Local file path or URL | | prompt | string | | "Describe this image in detail." | Analysis prompt | | detail | "low" \| "high" \| "auto" | | "auto" | Detail level |

ocr_image

Extract text from an image using OCR.

| Parameter | Type | Required | Default | Description | |---|---|---|---|---| | image | string | ✅ | — | Local file path or URL | | languages | string | | "" | Language hint, e.g. "zh,en" | | format | "plain" \| "markdown" \| "json" | | "plain" | Output format |

compare_images

Compare 2–4 images and describe differences/similarities.

| Parameter | Type | Required | Default | Description | |---|---|---|---|---| | images | string[] | ✅ | — | 2–4 image sources | | prompt | string | | "Compare these images..." | Comparison prompt |

analyze_video

Analyze video content. Requires a model with video support (e.g., Qwen3-VL).

| Parameter | Type | Required | Default | Description | |---|---|---|---|---| | video | string | ✅ | — | Local file path or URL | | prompt | string | | "Describe what happens in this video." | Analysis prompt |

🤖 Supported Models

| Model | Provider | Image | Video | Notes | |---|---|:---:|:---:|---| | Qwen3-VL | Self-hosted / API | ✅ | ✅ | Recommended. Full multimodal support | | GPT-4o | OpenAI | ✅ | ❌ | Strong image analysis | | LLaVA | Self-hosted | ✅ | ❌ | Open-source alternative | | InternVL | Self-hosted | ✅ | ⚠️ | Strong multilingual OCR |

Any model served via vLLM, Ollama, LMDeploy, or other OpenAI-compatible servers should work.

Supported formats: JPEG, PNG, GIF, WebP, BMP, SVG | MP4, AVI, MOV, MKV, WebM

📁 Project Structure

vision-mcp-server/
├── src/
│   ├── index.ts              # MCP server entry point
│   ├── config.ts             # Configuration loader
│   ├── types.ts              # TypeScript type definitions
│   ├── tools/
│   │   ├── analyze-image.ts
│   │   ├── ocr-image.ts
│   │   ├── compare-images.ts
│   │   └── analyze-video.ts
│   └── utils/
│       ├── api-client.ts     # OpenAI-compatible API client
│       └── file-handler.ts   # Local file → base64
├── package.json
├── tsconfig.json
├── .env.example
└── LICENSE

📄 License

[MIT](LICENSE)


🇨🇳 中文说明

功能

  • analyze_image — 使用视觉语言模型分析图像,支持自然语言提问
  • ocr_image — OCR 文字识别,支持纯文本、Markdown、JSON 输出
  • compare_images — 对比 2–4 张图像,识别差异和相似之处
  • analyze_video — 分析视频内容(需要 Qwen3-VL 等支持视频的模型)

快速开始

git clone https://github.com/Loveacup/vision-mcp-server.git
cd vision-mcp-server
npm install && npm run build

配置 .env

VISION_BASE_URL=http://your-server:port/v1/chat/completions
VISION_MODEL=Qwen3-VL-32B
VISION_API_KEY=your-api-key

在 Claude Code 的 ~/.mcp.json 中添加:

{
  "mcpServers": {
    "vision": {
      "command": "node",
      "args": ["/path/to/vision-mcp-server/dist/index.js"],
      "env": {
        "VISION_BASE_URL": "http://your-server:port/v1/chat/completions",
        "VISION_MODEL": "Qwen3-VL-32B",
        "VISION_API_KEY": "your-api-key"
      }
    }
  }
}

/path/to/vision-mcp-server 替换为实际安装路径。

Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.