AgentStack
MCP verified MIT Self-run

Self Hosted Ai Stack

mcp-hwdsl2-self-hosted-ai-stack · by hwdsl2

Deploy a complete self-hosted AI stack with Docker Compose: Ollama, LiteLLM, AnythingLLM, Whisper, WhisperLive, Kokoro, Embeddings, Docling and MCP Gateway. Local-first, private by default, with lightweight stacks, optional HTTPS and NVIDIA CUDA acceleration. Multi-arch: amd64, arm64.

No reviews yet
0 installs
14 views
0.0% view→install

Install

$ agentstack add mcp-hwdsl2-self-hosted-ai-stack

✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets Used
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

Are you the author of Self Hosted Ai Stack? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

[English](README.md) | [简体中文](README-zh.md) | [繁體中文](README-zh-Hant.md) | [Русский](README-ru.md)

Self-Hosted AI Stack

[](https://docs.docker.com/compose/)  [](https://hub.docker.com/u/hwdsl2)  [](https://opensource.org/licenses/MIT)

Includes Ollama, LiteLLM, AnythingLLM, Whisper, MCP Gateway, Embeddings, Docling, and Kokoro — fully configured and ready to run with Docker Compose.

  • Zero-config: all services auto-configure on first start
  • Secure by default: AnythingLLM password protection is enabled, and bundled API services auto-generate keys
  • HTTPS-ready: optional Caddy overlay provides automatic TLS and binds direct HTTP ports to localhost
  • Private: runs locally by default with optional external provider support via LiteLLM
  • Flexible: customize models, ports, providers, and API keys with simple env files
  • [Lightweight stacks](#lightweight-stacks) for lower memory requirements (as low as ~4.5 GB)
  • GPU acceleration via NVIDIA CUDA
  • Multi-arch: linux/amd64, linux/arm64

Community

  • 📬 Subscribe for project updates (1–2 emails/month) — get free AI and VPN deployment guides (PDF)
  • 💬 Join the r/selfhostedstack community for discussions and showcases
  • ⭐ Star the repository if you find it useful — it helps others discover it

Self-Hosted AI Stack is maintained by the author of Setup IPsec VPN (28k+ stars).

Included services

| Service | Role | Default port | |---|---|---| | Ollama (LLM) | Runs local LLM models (llama3, qwen, mistral, etc.) | 11434 | | AnythingLLM | Web-based chat UI — password-protected by default | 3001 | | LiteLLM | AI gateway with Admin UI — routes requests to Ollama and 100+ providers | 4000 | | Embeddings | Converts text to vectors for semantic search and RAG | 8000 | | Whisper (STT) | Transcribes spoken audio to text | 9000 | | WhisperLive (real-time STT) | Real-time speech-to-text transcription over WebSocket | 9090 | | Kokoro (TTS) | Converts text to natural-sounding speech | 8880 | | MCP Gateway | Provides MCP tools (filesystem, fetch, GitHub, search, databases) to AI clients | 3000 | | Docling | Converts documents (PDF, DOCX, etc.) to structured text/Markdown | 5001 |

Quick start

Requirements:

  • A Linux server (local or cloud) with Docker installed
  • At least 8 GB of RAM (with small models). For larger LLM models (8B+), 16 GB or more is recommended.
  • You can comment out services you don't need to reduce memory usage.

Start the full stack:

# Clone the repository to get the compose files
git clone https://github.com/hwdsl2/self-hosted-ai-stack
cd self-hosted-ai-stack
docker compose up -d

> Existing installs: If you cloned this project before it was renamed from docker-ai-stack, your existing checkout and deployment continue to work. GitHub redirects the old repository URL, and you do not need to rename your local directory, containers, volumes, or networks.

> PostgreSQL credentials: Fresh installs and existing default installs are handled automatically. If you previously set a custom database password, see [PostgreSQL credentials](#postgresql-credentials) before starting.

Pull a model (required before making LLM requests):

docker exec ollama ollama_manage --pull llama3.2:3b

Run the health check to verify all services are working:

./stack-check.sh

> Tip: On first start, services may take a few minutes to initialize. If any checks fail, wait and run ./stack-check.sh again. Use docker compose logs to check progress.

For detailed troubleshooting, see the [Troubleshooting](docs/troubleshooting.md) guide.

Get the LiteLLM master key (used to log into the Admin UI and for LLM requests):

docker exec litellm litellm_manage --showkey

Show core API keys (Ollama, LiteLLM, MCP Gateway)

docker exec ollama ollama_manage --showkey
docker exec litellm litellm_manage --showkey
docker exec mcp mcp_manage --showkey

Access AnythingLLM (Chat UI):

AnythingLLM is pre-configured to connect to your local LLM via LiteLLM. On first start, it may take a few minutes to become available (check progress with docker logs anythingllm).

Password-protected by default. A random admin password is auto-generated on first start, printed once to docker logs anythingllm, and saved to /app/server/storage/.initial_admin_password inside the anythingllm-data volume. The seeded password persists across container upgrades. Change it any time from Settings → Security; after you do, .initial_admin_password may no longer match the current login password.

Retrieve the auto-generated password:

# At any time from the data volume:
docker exec anythingllm cat /app/server/storage/.initial_admin_password

# Or from the live logs (only shown on first start):
docker compose logs anythingllm | grep -A4 "FIRST RUN"

Open http://:3001 in your browser and log in with the password above.

> Tip: When exposing AnythingLLM beyond localhost or a trusted LAN, use the included Caddy HTTPS overlay so the password is encrypted in transit and direct HTTP ports are bound to localhost. See [Internet-facing deployments](#internet-facing-deployments) below.

Access the LiteLLM Admin UI:

Open http://:4000/ui in your browser. Log in with username admin and your LiteLLM master key as the password. The UI provides virtual key management, spend tracking, and model configuration.

> Tip: In the Admin UI, click Playground in the left menu. Select a local model (e.g., ollama-chat/llama3.2:3b) from the dropdown and start chatting — a quick way to verify your local LLM is working end-to-end.

Stop the stack:

# Stop and remove all containers (data is preserved in Docker volumes)
docker compose down

GPU acceleration (NVIDIA CUDA)

For NVIDIA GPU acceleration, use the CUDA compose file:

docker compose -f docker-compose.cuda.yml up -d

> Tip: To avoid adding -f docker-compose.cuda.yml to every subsequent docker compose command (down, pull, logs, etc.), set it once for your shell session: > > ``bash > export COMPOSE_FILE=docker-compose.cuda.yml > ` > > Then run plain docker compose commands as usual. To make it persistent, add COMPOSEFILE=docker-compose.cuda.yml to a .env file in this directory. Run unset COMPOSEFILE` to switch back to the CPU configuration.

Requirements: NVIDIA GPU, NVIDIA driver 575.57.08+ (Linux) or 576.57+ (Windows), and the NVIDIA Container Toolkit installed on the host. CUDA images are linux/amd64 only.

> Podman users: The Compose deploy: GPU block is ignored by Podman. Use CDI instead — see [Using Podman](#using-podman).

Lightweight stacks

Don't need the full stack? Use a pre-configured subset from the stacks/ folder:

> Note: The lightweight stacks share default container names, ports, and Docker volume names. Run one stack variant at a time with the default compose files; stop the current variant before switching to another. To combine capabilities, use the full stack or customize Compose project names, container names, ports, and volumes.

| Stack | Services | Memory | Use case | |---|---|---|---| | [chat-ui](stacks/chat-ui/) | Ollama + LiteLLM + AnythingLLM | ~5 GB | Web-based ChatGPT-like chat interface | | [voice-pipeline](stacks/voice-pipeline/) | Whisper + Ollama + LiteLLM + Kokoro | ~6 GB | Speech-to-text → LLM → text-to-speech | | [voice-chat](stacks/voice-chat/) | Whisper + Ollama + LiteLLM + Kokoro + AnythingLLM | ~6.5 GB | Chat UI with voice input/output | | [rag-pipeline](stacks/rag-pipeline/) | Ollama + LiteLLM + Embeddings | ~5 GB | Semantic search + LLM Q&A | | [rag-pipeline-full](stacks/rag-pipeline-full/) | Ollama + LiteLLM + Embeddings + Docling | ~6 GB | Document parsing + semantic search + LLM Q&A | | [code-assistant](stacks/code-assistant/) | Ollama + LiteLLM + MCP Gateway + Embeddings | ~5 GB | AI coding with tools + semantic code search | | [ai-tools](stacks/ai-tools/) | Ollama + LiteLLM + MCP Gateway | ~5 GB | AI coding assistant with tool access | | [chat-only](stacks/chat-only/) | Ollama + LiteLLM | ~4.5 GB | Minimal local ChatGPT replacement |

git clone https://github.com/hwdsl2/self-hosted-ai-stack
cd self-hosted-ai-stack/stacks/chat-ui  # or voice-pipeline, voice-chat, rag-pipeline, rag-pipeline-full, code-assistant, ai-tools, chat-only
docker compose up -d

Architecture

graph LR
    A["🎤 Audio input"] -->|transcribe| W["Whisper(speech-to-text)"]
    D["📄 Documents"] -->|parse| DC["Docling(document → text)"]
    DC -->|embed| E["Embeddings(text → vectors)"]
    E -->|store| VDB["pgvector(in shared Postgres)"]
    W -->|query| E
    VDB -->|context| L["LiteLLM(AI gateway)"]
    W -->|text| L
    L -->|routes to| O["Ollama(local LLM)"]
    L -->|response| T["Kokoro TTS(text-to-speech)"]
    T --> B["🔊 Audio output"]
    C["🤖 AI client(Cline, Claude, etc.)"] -->|MCP tools| M["MCP Gateway(MCP endpoint)"]
    C -->|chat| L
    L -->|MCP protocol| M
    U["👤 User"] -->|chat| AN["AnythingLLM(chat UI)"]
    AN -->|LLM requests| L
    AN -->|MCP tools| M
    U -->|use| C
    U -->|speak| A
    U -->|upload| D

Notes:

  • Ollama's port (11434) and MCP Gateway's port (3000) are internal to the Docker network and not exposed to the host by default. Access your LLM through LiteLLM on port 4000.
  • Kokoro (TTS), Docling (document parsing), and WhisperLive (real-time STT) are disabled by default to reduce memory usage. Uncomment these services in docker-compose.yml to enable them.

Running without Docker Compose

If you prefer using docker run commands directly, first create a shared network so services can communicate:

docker network create ai-stack

Then generate a PostgreSQL password and start each service on the shared network:

> Note: With manual docker run, wait for each dependency to become ready before starting services that use it (for example, wait for PostgreSQL and any other dependencies, such as Ollama or MCP, before LiteLLM; if using AnythingLLM, wait for LiteLLM before starting it). The examples below generate one PostgreSQL password variable and reuse it for Postgres and LiteLLM.

LITELLM_POSTGRES_PASSWORD=$(LC_ALL=C tr -dc 'A-Za-z0-9'  **Note:** A shell `alias docker=podman` is **not** sufficient — aliases are not seen by scripts such as `stack-check.sh`. Use the `podman-docker` package (or a `docker` → `podman` symlink in your `PATH`) instead. Alternatively, `stack-check.sh` auto-detects Podman; you can also force it with `CONTAINER_ENGINE=podman ./stack-check.sh`.

**2. Install a Compose provider.** `podman compose` delegates to an external provider. Install either `podman-compose` or `docker-compose`:

```bash
# Fedora / RHEL / CentOS Stream
sudo dnf install -y podman-compose

# Debian / Ubuntu
sudo apt-get install -y podman-compose

3. Start the stack. With the shim installed, every command in this README works unchanged. Without it, substitute podman for docker:

git clone https://github.com/hwdsl2/self-hosted-ai-stack
cd self-hosted-ai-stack
podman compose up -d

Run the health check (auto-detects the engine):

./stack-check.sh

GPU acceleration (CDI). Podman does not read the Compose deploy.resources GPU block. Instead, use the Container Device Interface (CDI). After installing the NVIDIA Container Toolkit, generate a CDI spec:

sudo nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml

Then expose the GPU to the relevant services. For podman compose, replace the deploy: block in docker-compose.cuda.yml with a devices: entry for the ollama (and whisper) services:

    devices:
      - nvidia.com/gpu=all

For a plain podman run command, add --device nvidia.com/gpu=all.

SELinux. On SELinux-enabled hosts (Fedora, RHEL, CentOS Stream), bind-mounted files need a relabel suffix, or the container will be denied access. Add :z (shared) to the chat-ui-bootstrap.sh bind mount:

  • In docker-compose.yml: change ./chat-ui-bootstrap.sh:/usr/local/bin/chat-ui-bootstrap.sh:ro to ./chat-ui-bootstrap.sh:/usr/local/bin/chat-ui-bootstrap.sh:ro,z
  • In the podman run command above: change "$(pwd)/chat-ui-bootstrap.sh:/usr/local/bin/chat-ui-bootstrap.sh:ro" to "$(pwd)/chat-ui-bootstrap.sh:/usr/local/bin/chat-ui-bootstrap.sh:ro,z"

Named volumes do not need relabeling.

Next steps: Pull a model and access the services — follow the instructions in [Quick start](#quick-start) starting from "Pull a model." With the podman-docker shim installed, all commands work unchanged.

Connect MCP Gateway to LiteLLM

LiteLLM and MCP Gateway are automatically wired when using the compose files in this repository — no manual key setup is needed.

API keys are shared automatically between services via Docker shared volumes:

  • Ollama generates an API key on first start and copies it to a shared volume
  • MCP Gateway does the same
  • LiteLLM reads both keys from the shared volumes on startup

The LITELLM_MCP_URL=http://mcp:3000/mcp and LITELLM_OLLAMA_BASE_URL=http://ollama:11434 environment variables are pre-configured in the compose files, so all services are connected automatically with a single docker compose up -d.

Once connected, AI clients that call LiteLLM can use MCP tools (filesystem, fetch, GitHub, etc.) directly through the LiteLLM proxy.

Voice pipeline example

Transcribe a spoken question, get a local LLM response via Ollama, and convert it to speech:

Note: Kokoro (TTS) is disabled by default. To use this example, first uncomment the kokoro service in docker-compose.yml, then run docker compose up -d.

Tip: Need a sample audio file? Download this English speech sample (WAV, MIT License) from the Azure Samples repository:

curl -L -o sample_speech.wav \
    "https://github.com/Azure-Samples/cognitive-services-speech-sdk/raw/master/sampledata/audiofiles/katiesteve.wav"
LITELLM_KEY=$(docker exec litellm litellm_manage --getkey)
WHISPER_KEY=$(docker exec whisper whisper_manage --getkey)
KOKORO_KEY=$(docker exec kokoro kokoro_manage --getkey)

# Step 1: Transcribe audio to text (Whisper)
TEXT=$(curl -s http://localhost:9000/v1/audio/transcriptions \
    -H "Authorization: Bearer $WHISPER_KEY" \
    -F file=@sample_speech.wav -F model=whisper-1 | jq -r .text)

# Step 2: Send text to Ollama via LiteLLM and get a response
RESPONSE=$(curl -s http://localhost:4000/v1/chat/completions \
    -H "Authorization: Bearer $LITELLM_KEY" \
    -H "Content-Type: application/json" \
    -d "{\"model\":\"ollama/llama3.2:3b\",\"messages\":[{\"role\":\"user\",\"content\":\"$TEXT\"}]}" \
    | jq -r '.choices[0].message.content')

# Step 3: Convert the response to speech (Kokoro TTS)
curl -s http://localhost:8880/v1/audio/speech \
    -H "Content-Type: application/json" \
    -H "Authorization: Bearer $KOKORO_KEY" \
    -d "{\

…

## Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [hwdsl2](https://github.com/hwdsl2)
- **Source:** [hwdsl2/self-hosted-ai-stack](https://github.com/hwdsl2/self-hosted-ai-stack)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet — be the first.

Versions

  • v0.1.0 Imported from the upstream source.