# vllm-project

> Open-source publisher. Listings imported from github.com/vllm-project — credited to the original author with their license.

- **Listings:** 7
- **Total installs:** 0
- **Profile:** https://agentstack.voostack.com/s/vllm-project
- **Website:** https://github.com/vllm-project

## Published listings

- [Semantic Router](https://agentstack.voostack.com/l/mcp-vllm-project-semantic-router) — MCP server · Free — `agentstack add mcp-vllm-project-semantic-router`
  System Level Intelligent Router for Mixture-of-Models at Cloud, Data Center and Edge
- [Vllm Deploy Simple](https://agentstack.voostack.com/l/skill-vllm-project-vllm-skills-vllm-deploy-simple) — Skill · Free · security-reviewed — `agentstack add skill-vllm-project-vllm-skills-vllm-deploy-simple`
  Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.
- [Vllm Deploy Docker](https://agentstack.voostack.com/l/skill-vllm-project-vllm-skills-vllm-deploy-docker) — Skill · Free · security-reviewed — `agentstack add skill-vllm-project-vllm-skills-vllm-deploy-docker`
  Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.
- [Vllm Deploy K8s](https://agentstack.voostack.com/l/skill-vllm-project-vllm-skills-vllm-deploy-k8s) — Skill · Free · security-reviewed — `agentstack add skill-vllm-project-vllm-skills-vllm-deploy-k8s`
  Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint. Use this skill whenever the user wants to deploy, run, or serve vLLM on a Kubernetes cluster, including creating deployments, services, checking existing deployments, or managing vLLM on K8s.
- [Vllm Bench Random Synthetic](https://agentstack.voostack.com/l/skill-vllm-project-vllm-skills-vllm-bench-random-synthetic) — Skill · Free · security-reviewed — `agentstack add skill-vllm-project-vllm-skills-vllm-bench-random-synthetic`
  Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics. Use when the user wants to quickly test vLLM serving performance without downloading external datasets.
- [Vllm Prefix Cache Bench](https://agentstack.voostack.com/l/skill-vllm-project-vllm-skills-vllm-prefix-cache-bench) — Skill · Free · security-reviewed — `agentstack add skill-vllm-project-vllm-skills-vllm-prefix-cache-bench`
  This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns. Use when the user asks to benchmark prefix caching hit rate, caching efficiency, or repeated-prompt performance in vLLM.
- [Vllm Bench Serve](https://agentstack.voostack.com/l/skill-vllm-project-vllm-skills-vllm-bench-serve) — Skill · Free · security-reviewed — `agentstack add skill-vllm-project-vllm-skills-vllm-bench-serve`
  Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve. Supports multiple datasets (random, sharegpt, sonnet, HF), backends (openai, openai-chat, vllm-pooling, embeddings), throughput/latency testing with request-rate control, and result saving. Use when benchmarking LLM serving performance, measuring TTFT/TPOT, or load testing inference APIs.

---
Seller on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Install any with `agentstack add <slug>`.
