AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
MCP verified MIT Self-run

Llm Benchmark Mcp Server

mcp-aiagentkarl-llm-benchmark-mcp-server · by AiAgentKarl

MCP Server for LLM comparison, benchmarks, and pricing — find the best model for any task

— No reviews yet
0 installs
32 views
0.0% view→install

Install

$ agentstack add mcp-aiagentkarl-llm-benchmark-mcp-server

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/mcp-aiagentkarl-llm-benchmark-mcp-server)

Reliability & compatibility

✓ Security review passed
0 installs to date
— no reviews yet
○ 5mo ago

Declared compatibility

Claude CodeClaude DesktopCursorWindsurf

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Llm Benchmark Mcp Server? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

LLM Benchmark MCP Server

MCP server that gives AI agents access to LLM benchmark data, pricing comparisons, and model recommendations.

Features

  • compare_models — Side-by-side benchmark comparison of LLMs (MMLU, HumanEval, MATH, GPQA, ARC, HellaSwag)
  • getmodeldetails — Detailed info about a specific model including strengths/weaknesses
  • recommend_model — Get the best model recommendation for your task and budget
  • listtopmodels — Top models ranked by category (coding, math, reasoning, chat)
  • get_pricing — Pricing comparison via OpenRouter API

Supported Models

GPT-4o, GPT-4o-mini, GPT-4 Turbo, o1, o3-mini, Claude 3.5 Sonnet, Claude 3.5 Haiku, Claude 3 Opus, Gemini 2.0 Flash, Gemini 2.0 Pro, Gemini 1.5 Pro, Llama 3.1 (8B/70B/405B), Llama 3.3 70B, Mistral Large, Mistral Small, Mixtral 8x22B, DeepSeek V3, DeepSeek R1, Qwen 2.5 72B

Installation

pip install llm-benchmark-mcp-server

Usage with Claude Desktop

Add to your claude_desktop_config.json:

{
  "mcpServers": {
    "llm-benchmark": {
      "command": "benchmark-server"
    }
  }
}

Or via uvx (no install needed):

{
  "mcpServers": {
    "llm-benchmark": {
      "command": "uvx",
      "args": ["llm-benchmark-mcp-server"]
    }
  }
}

Example Queries

  • "Compare GPT-4o vs Claude 3.5 Sonnet vs Gemini 2.0 Pro"
  • "Which model is best for coding on a low budget?"
  • "Show me the top 10 models for math"
  • "What does GPT-4o cost compared to Claude?"
  • "Give me details about DeepSeek R1"

Data Sources

  • Benchmarks: Hardcoded from official papers and public leaderboards (MMLU, HumanEval, MATH, GPQA, ARC-Challenge, HellaSwag)
  • Pricing: Live data from OpenRouter API
  • Arena Rankings: Chatbot Arena Leaderboard (when available)

More MCP Servers by AiAgentKarl

| Category | Servers | |----------|---------| | 🔗 Blockchain | Solana | | 🌍 Data | Weather · Germany · Agriculture · Space · Aviation · EU Companies | | 🔒 Security | Cybersecurity · Policy Gateway · Audit Trail | | 🤖 Agent Infra | Memory · Directory · Hub · Reputation | | 🔬 Research | Academic · LLM Benchmark · Legal |

→ Full catalog (40+ servers)

License

MIT

Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.