AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified Apache-2.0 Self-run

Vllm Deploy K8s

skill-vllm-project-vllm-skills-vllm-deploy-k8s · by vllm-project

Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint. Use this skill whenever the user wants to deploy, run, or serve vLLM on a Kubernetes cluster, including creating deployments, services, checking existing deployments, or managing vLLM on K8s.

No reviews yet
0 installs
41 views
0.0% view→install

Install

$ agentstack add skill-vllm-project-vllm-skills-vllm-deploy-k8s

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-vllm-project-vllm-skills-vllm-deploy-k8s)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
5mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Vllm Deploy K8s? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

vLLM Kubernetes Deployment

A Claude skill for deploying vLLM to Kubernetes using YAML templates. Deploys a vLLM OpenAI-compatible server as a Kubernetes Deployment with a ClusterIP Service, GPU resources, and health probes.

What this skill does

  • Deploy vLLM as a Kubernetes Deployment + Service with NVIDIA GPU support
  • Check if a vLLM deployment already exists before deploying
  • Check if the Hugging Face token secret exists, and ask the user for their token if not
  • Use the vllm/vllm-openai:latest image by default (user can specify a different version)
  • Provide sensible default configuration that users can customize (model, replicas, GPU count, extra vLLM flags, etc.)

Prerequisites

  • kubectl configured with access to a Kubernetes cluster
  • NVIDIA GPU Operator or device plugin installed on cluster nodes
  • Hugging Face token (required for gated models like Llama, optional for public models)

Deployment Steps

Step 1: Check HF token secret

Before deploying, check if the hf-token Kubernetes secret exists in the target namespace:

kubectl get secret hf-token -n 
  • If the secret exists: proceed to Step 2.
  • If the secret does not exist: ask the user to provide their Hugging Face token, then create the secret:
kubectl create secret generic hf-token --from-literal=HF_TOKEN="" -n 

This is required for gated models (e.g., meta-llama/Meta-Llama-3.1-8B). For public models, the secret is optional but recommended to avoid rate limits.

Step 2: Check if deployment already exists

Before applying, check if a vLLM deployment already exists:

kubectl get deployment vllm -n 
  • If it exists: inform the user that the deployment already exists. Show the current image and status. Ask the user if they want to update it or skip.
  • If it does not exist: proceed to deploy.

Step 3: Deploy

Apply the template YAML files to deploy vLLM:

kubectl apply -f templates/vllm-service.yaml -n 
kubectl apply -f templates/vllm-deployment.yaml -n 

Step 4: Wait and verify

Wait for the deployment to roll out:

kubectl rollout status deployment/vllm -n  --timeout=600s

Verify the pod is running and ready:

kubectl get pods -n  -l app=vllm

Confirm the pod shows READY 1/1 and STATUS Running. If the pod is not ready yet, wait and check again. If it's in CrashLoopBackOff or Error, check the logs with kubectl logs -n -l app=vllm.

Step 5: Print deployment summary

Once the pod is ready, print a summary message to the user in this format (replace placeholders with actual values):

🎉 **vLLM Deployment Successful!**

| Resource | Name | Status |
|----------|------|--------|
| Deployment |  | / Ready |
| Service |  | ClusterIP: |
| Pod |  | Running |
| Image |  | |
| Model |  | |

 

**To test the API, run these two commands in your terminal:**

**1. Open a port-forward** (this connects your local port  to the vLLM service inside the cluster):

kubectl port-forward svc/vllm-svc : -n 

**2. In a separate terminal**, send a test request to the OpenAI-compatible API:

curl -s http://localhost:/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"","messages":[{"role":"user","content":"Hello!"}],"max_tokens":50}' | python3 -m json.tool

If everything is working, you'll get a JSON response with the model's reply.

Default Configuration

The templates use the following defaults:

| Parameter | Default Value | |-----------|---------------| | Image | vllm/vllm-openai:latest | | Model | Qwen/Qwen2.5-1.5B-Instruct | | Port | 8000 | | Replicas | 1 | | GPU count | 1 | | GPU memory utilization | 0.85 | | Tensor parallel size | 1 | | CPU request / limit | 12 / 128 | | Memory request / limit | 100Gi / 400Gi | | Shared memory (dshm) | 80Gi |

Customization

When the user requests changes, modify the template YAML files before applying. The following can be customized:

  • Image version: Change image: vllm/vllm-openai: in templates/vllm-deployment.yaml (default: latest). Use a specific version tag like v0.17.1 if the user requests it.
  • Model: Change the model name in the vllm serve command inside the Deployment args.
  • Extra vLLM flags: Append additional flags to the vllm serve command in the Deployment args (e.g., --max-model-len 4096, --kv-cache-dtype fp8, --enforce-eager, --generation-config vllm).
  • Replicas: Change replicas: in the Deployment spec.
  • GPU count: Change nvidia.com/gpu in both requests and limits under resources.
  • Tensor parallel size: Change --tensor-parallel-size flag to match the GPU count.
  • CPU/Memory resources: Change cpu and memory values under requests and limits.
  • Port: Change containerPort in the Deployment, port/targetPort in the Service, the port in all health probes (liveness, readiness, startup), AND add --port to the vllm serve command in args. All four must match.
  • Namespace: Apply to a specific namespace using -n .
  • Shared memory size: Change the sizeLimit of the dshm emptyDir volume.

Edit the template files using the Edit tool, then apply the modified templates.

Status Check

kubectl get deployment,svc,pods -n  -l app=vllm

Cleanup

When the user asks to clean up or delete the vLLM deployment, run the following steps:

  1. Delete the Deployment and Service:
kubectl delete -f templates/vllm-deployment.yaml -n 
kubectl delete -f templates/vllm-service.yaml -n 
  1. Ask the user if they also want to delete the HF token secret. If yes:
kubectl delete secret hf-token -n 
  1. Verify everything is cleaned up:
kubectl get deployment,svc,pods -n  -l app=vllm
  1. Print a summary message to the user:
vLLM deployment has been cleaned up from namespace .
Deleted: Deployment/vllm, Service/vllm-svc
HF token secret: 

Troubleshooting

  • Pod stuck in Pending: No GPU nodes available. Check kubectl describe pod for scheduling errors. Ensure NVIDIA GPU Operator or device plugin is installed.
  • Pod OOMKilled: Increase memory limits in the Deployment, or use a smaller model.
  • ImagePullBackOff: Check the image name and tag. Verify the node has access to Docker Hub / the container registry.
  • Startup probe failures (CrashLoopBackOff): Model download may be slow. Check logs with kubectl logs . Ensure hf-token secret exists for gated models. Increase failureThreshold on the startup probe if needed.
  • HF_TOKEN not working: Verify the secret exists: kubectl get secret hf-token -n . Check the token is valid.
  • GPU not detected in container: Ensure nvidia.com/gpu resource is requested and the NVIDIA device plugin is running on the node.

References

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.