Install
$ agentstack add mcp-mmontes11-k8s-ai β scanned Β· β verified, works with Claude Code, Cursor, and more.
Security review
β PassedNo issues found. Passed automated security review. Β· v0.1.0 How review works β
- β Prompt-injection patterns
- β Secret / credential exfiltration
- β Dangerous shell & filesystem operations
- β Untrusted network calls
- β Known-malicious package signatures
What it can access
- β Network access No
- β Filesystem access No
- β Shell / process execution No
- β Environment & secrets No
- β Dynamic code execution No
From automated source analysis of v0.1.0. βUsedβ means the capability is present in the source β more access means more to trust, not that itβs unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work βAbout
π§ k8s-ai
Tenant repository bootstrapped by k8s-infrastructure that contains the manifests for AI related applications
Overview
This repository manages AI workloads on Kubernetes using GitOps with Flux CD. It includes deployments for LLM inference services, web UIs, and model serving infrastructure.
Applications
Open WebUI
- Path:
./apps/open-webui - Type: HelmRelease (ollama-webui chart)
- Description: Web interface for interacting with LLMs
- Features:
- Persistent storage via PVC
- Integration with Ollama backend
- Model access control bypass enabled
ComfyUI
- Path:
./apps/comfyui - Type: Native Kubernetes resources
- Description: Graph-based interface for Stable Diffusion
- Image: mmontes11/docker-comfyui
- Features:
- Persistent volume for model caching
- Replication source/destination for data synchronization
- RESTic backup support
n8n
- Path:
./apps/n8n - Type: HelmRelease (n8n helm chart)
- Description: Workflow automation and integration platform
- Features:
- Persistent storage via PVC
- RESTic backup support
- Replication source/destination for data synchronization
opencode
- Path:
./apps/opencode - Type: Native Kubernetes resources
- Description: Coding agent and AI workspace for interactive development
- Image: mmontes11/docker-opencode
- Features:
- NVIDIA GPU support for accelerated model training and inference
- Persistent storage (100Gi PVC)
- Pre-configured development environment with tools
- RESTic backup support
- Replication source/destination for data synchronization
- Integration with GitHub, HuggingFace, and n8n via tokens
Infrastructure
Model Serving
Ollama
- Path:
./infrastructure/ollama - Description: Lightweight LLM inference server
- Features:
- Native GPU support
- Simple HTTP API
- Model caching
llama.cpp
- Path:
./infrastructure/llamacpp - Description: High-performance C/C++ inference engine optimized for CPU and GPU
- Features:
- Qwen3.6 MTP model support with 1.4-2.2x faster inference
- 256k context window for agentic AI workflows
- StatefulSet deployment with persistent storage
- Prometheus ServiceMonitor integration
- Ingress routing via HTTPRoute
vLLM
- Path:
./infrastructure/vllm - Description: High-throughput LLM serving with PagedAttention
- Use Case: Production workloads requiring high concurrency
KServe
- Path:
./infrastructure/kserve - Example:
./examples/llminferenceservice.yaml - Description: Kubernetes-native ML serving platform
- Features:
- LLMInferenceService CRD
- Custom model templates
MCP Servers
- MCP Kubernetes: Kubernetes model context protocol server
- MCP Grafana: Grafana monitoring integration
- MCP GitHub: GitHub API integration
- MCP Photoprism: Photo management (mmontes & xiaowen)
Architecture
βββ apps/ # Application deployments
β βββ comfyui/ # ComfyUI deployment
β βββ n8n/ # n8n workflow automation
β βββ opencode/ # opencode AI development workspace
β βββ open-webui/ # Open WebUI deployment
βββ clusters/ # Cluster-specific configurations
β βββ homelab/
β βββ apps.yaml # Application Kustomizations
β βββ infrastructure.yaml
β βββ namespaces.yaml
βββ examples/ # Example configurations
β βββ llminferenceservice.yaml
βββ infrastructure/ # Shared infrastructure
βββ kserve/ # KServe ML serving
βββ vllm/ # vLLM serving engine
βββ llamacpp/ # llama.cpp inference engine
βββ lws/ # LeaderWorkerSet
βββ ollama/ # Ollama LLM backend
βββ mcp-*/ # MCP server integrations
AI Benchmarks
LLM benchmarks using llama.cpp on Kubernetes: mmontes11/llm-bench
License
[MIT](./LICENSE)
Source & license
This open-source MCP server is cataloged on AgentStack and links to its original source β we do not rehost the code.
- Author: mmontes11
- Source: mmontes11/k8s-ai
- License: Apache-2.0
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.