# Serving Llms On Epyc

> >-

- **Type:** Skill
- **Install:** `agentstack add skill-amd-skills-serving-llms-on-epyc`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [amd](https://agentstack.voostack.com/s/amd)
- **Installs:** 0
- **Category:** [Agent Skills](https://agentstack.voostack.com/c/agent-skills)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [amd](https://github.com/amd)
- **Source:** https://github.com/amd/skills/tree/main/skills/serving-llms-on-epyc
- **Website:** https://developer.amd.com/

## Install

```sh
agentstack add skill-amd-skills-serving-llms-on-epyc
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# Serving LLMs on AMD EPYC (vLLM + zentorch, CPU)

Bring up a single vLLM OpenAI endpoint on an AMD EPYC host with the zentorch CPU
backend, sized to the hardware. Container-first (Docker or Podman); conda/host
is the fallback.

**This is single-socket serving:** one instance pinned to one socket and its memory
(vLLM scales poorly across sockets, so we do not span them). On a dual-socket host it
runs on a single socket; the multi-socket answer is **multiple instances (one per
socket)**, which is out of scope for this single-instance recipe.

Hard rule for this skill: **on any failure, report the cause + logs and STOP.
Do not retry, do not debug.** (Debugging is a separate workflow.)

**The agent does the serve flow itself** -- pull, configure, launch, poll --
using the runtime `validate.py` reports. Never hand the user per-serve commands.
Like serving-llms-on-instinct, an accessible container runtime is a one-time
**prerequisite**: if `validate.py` finds none, report its one-time fix (make
docker accessible / install podman / provide a conda env) and stop. Do not
attempt `sudo` or privilege escalation.

## Data file

Read `data/epyc.json` directly. It holds the container image, mandatory CPU run
flags, supported precision, the model-support policy, the default model, and the
verified throughput-flag gotcha. Do not hardcode the image tag from memory -- read it.

## Step 1: Detect the CPU

```bash
python3 scripts/detect.py            # add --host user@box for a remote host
```

Returns `cpu_model`, `is_amd_epyc`, `epyc_generation`
(Naples/Rome/Milan/Genoa/Bergamo/Siena/Turin), `zen_arch`, `avx512`,
`logical_cores`, `physical_cores`, `sockets`, `numa_nodes`, `memory_gb`. If
`is_amd_epyc` is `false`, stop: this skill targets AMD EPYC. (Other x86 may work
but is unsupported here.) Carry `epyc_generation` / `avx512` through the later
phases -- e.g. AVX-512 + bf16 land on Zen4+ (Genoa/Turin), and Turin packs up to
128 cores/socket, which the thread-binding in Step 5 sizes from.

## Step 2: Validate the runtime and environment

```bash
python3 scripts/validate.py --image 
```

Returns `ready`, `runtime` (`docker`, `podman`, or null), `runtime_detail`,
`conda_path_available`, `ram_gb`, and `errors/warnings/advisories`. Pick the path:
- `runtime` is `docker` or `podman` -> container path (Step 6), used verbatim.
- `runtime` null but `conda_path_available: true` -> conda/host path.
- `runtime` null and no conda -> `ready` is false. Report the one-time
  onboarding `fix` (make docker accessible / install podman / conda env) and stop.

Do not proceed if `ready` is `false`.

## Step 3: Resolve and validate the model

If the user named no model, use `default_model` from `data/epyc.json`
(`Qwen/Qwen3-0.6B` -- ungated, tiny, fast first success). Otherwise use theirs.

Check that vLLM actually supports the model (do **not** blanket-block multimodal):

```bash
python3 scripts/check_model.py --model-id  --vllm-version 
```

- Exit 0 = vLLM serves it as a generation endpoint (`kind` `text` or `multimodal`),
  or support is undeterminable (gated/offline) -- proceed; launch confirms.
- Exit 1 = positively unsupported: the architecture is not in vLLM's registry, or
  it is a `pooling`/embedding/reranker (not a chat/completion endpoint). Report the
  printed `message` and stop.
- A `multimodal` model is allowed; a vLLM-supported multimodal arch may still hit a
  GPU-only kernel on CPU, which surfaces at load (the no-retry rule then applies).

**Precision/dtype**: native CPU dtypes are `bf16` (default), `fp16`, `fp32`. Use
`bfloat16` unless the user asks otherwise.

For gated models (Llama, Gemma) `HF_TOKEN` must be set and the license accepted on
HuggingFace; if not, stop and say so.

## Step 4: Check it fits host RAM

RAM is the ceiling on CPU (weights + KV cache both live in RAM). Run on ONE line:

```bash
python3 scripts/estimate_memory.py --model-id  --ram-gb  --max-model-len  --num-prompts 
```

Exit 0 = fits, exit 1 = does not fit. If `fit.fits` is false: **do not launch.**
Tell the user `required_gb` vs `ram_gb` and the printed `fit.action` -- reduce
`--max-model-len` to `fit.suggested_max_model_len` and retry, or use a smaller
model. `--max-model-len` and `--num-prompts` are the two knobs that move KV.
Extra flag: `--weight-gb N` overrides weights if a model has no HF metadata
(rare). KV cache is bf16-only on zentorch CPU (no fp8 KV).

## Step 5: Size the CPU runtime from the hardware

```bash
eval "$(python3 scripts/cpu_tune.py)"      # or --format json to inspect
```

A single instance runs on **one socket, with its memory** (vLLM scales poorly across
sockets). `cpu_tune.py` exports `VLLM_CPU_OMP_THREADS_BIND` (the chosen socket's
physical cores) and `VLLM_CPU_KVCACHE_SPACE` (sized from that **socket's local RAM**,
not whole-system, so the KV pool stays on-socket). It does **not** set
`OMP_NUM_THREADS` (vLLM derives it) or `VLLM_CPU_NUM_OF_RESERVED_CPU` (vLLM's own default).

Socket choice on a dual-socket host (load-aware): it samples per-socket CPU busy%
(~0.5s) and prefers a free socket -- both free → socket 0; one free → that socket;
**both busy (≥ `--busy-threshold`, default 15%) → it `warning`s and proceeds on the
least-busy socket**. `--socket N` forces a choice. Single-socket hosts use socket 0.

For the chosen socket it also emits the memory-bound pin: `container_cpuset`
(`--cpuset-cpus= --cpuset-mems=`) for the container path, and
`conda_launch_prefix` (`numactl --cpunodebind/--membind`, falling back to `taskset`
CPU-only, or empty-with-note if neither tool exists) for conda. **Surface `warning`
to the user** if set. On NPS2/NPS4 a socket spans multiple NUMA nodes; memory is
bound across them and `nps_note` flags that finer binding could add performance.

## Step 6: Confirm the plan, then launch (container-first)

Before launching, present this summary and **wait for the user to confirm** -- do
not launch unprompted. This is the human gate before anything runs:

| Field | Value |
|---|---|
| Model / kind | `` -- `text` or `multimodal` (from `check_model.py`) |
| Path | container (``, image from `data/epyc.json`) or conda/host |
| Precision | `bfloat16` (or the user's choice) |
| Fit | required `` GB vs `` GB RAM |
| CPU sizing | socket `` (``), bind ``, KV `` GB (socket-local), mem bound to nodes `` |
| Hardware | EPYC `` (``), `` cores, AVX-512 `` |
| Port | `` |

If `cpu_tune.py` returned a `warning` (e.g. all sockets busy), include it here so the user sees it before confirming.

Proceed only on a clear "go". If the user declines or wants changes (model,
`--max-model-len`, port), stop and adjust -- do not launch.

Build the launch from `data/epyc.json`. The CLI is `vllm serve `.
**Do not pass `--device cpu`** on vLLM >= 0.20 -- the zentorch plugin
auto-selects the CPU platform and `vllm serve` rejects the flag. Only add it if
`vllm serve --help` lists it (older vLLM).

**Container path** (`runtime` from validate.py). The agent runs these itself,
including the pull. `RT` is the resolved runtime verbatim:
```bash
RT=""
$RT pull           # agent pulls; do not ask the user to
$RT run -d --name vllm-epyc \
              # --ipc=host --shm-size=16g --network=host
   \
               # --cpuset-cpus= --cpuset-mems=
  --env VLLM_CPU_OMP_THREADS_BIND="$VLLM_CPU_OMP_THREADS_BIND" \
  --env VLLM_CPU_KVCACHE_SPACE=$VLLM_CPU_KVCACHE_SPACE \
  --env HF_TOKEN=${HF_TOKEN} \
   \
  vllm serve  --dtype bfloat16 --port  --max-model-len 
```

**Conda/host path** (no container runtime, `conda_path_available` true). `eval`-ing
cpu_tune already exported the env vars; prefix the launch with `conda_launch_prefix`
from cpu_tune so memory is bound to the chosen socket (empty → unpinned, with a note):
```bash
 vllm serve  --dtype bfloat16 --port  --max-model-len  &
# e.g. numactl --cpunodebind=0 --membind=0 vllm serve ...
```

Optional throughput flags are **opt-in and must move together** (see Gotchas):
`TORCHINDUCTOR_FREEZING=1` + `VLLM_USE_AOT_COMPILE=0` (+ `ZENTORCH_WEIGHT_PREPACK=1`).
The base launch sets none of them.

## Step 7: Poll until up and responsive

A 503 while loading is normal. Poll until the server answers, then prove the
chat endpoint works. CPU first-token compile can take a minute or two.

```bash
# container alive (or process alive for conda) + /health
for i in $(seq 1 120); do
  # container path:
  $RT inspect -f '{{.State.Running}}' vllm-epyc 2>/dev/null | grep -q true || { echo "FAILED: container exited"; $RT logs --tail 50 vllm-epyc; break; }
  curl -sf http://localhost:/health >/dev/null 2>&1 && { echo "HEALTHY"; break; }
  sleep 3
done
```

Then validate the OpenAI endpoint is actually accessible:
```bash
curl -sf http://localhost:/v1/chat/completions -H 'Content-Type: application/json' \
  -d '{"model":"","messages":[{"role":"user","content":"hi"}],"max_tokens":8}'
```

Resource sanity (your validation list): `$RT stats --no-stream vllm-epyc`.

**If the server never becomes healthy or the endpoint does not respond: print
the container/process logs, state the failure, and STOP. Do not retry. Do not
start a debugging loop.**

## Step 8: On success, hand over the endpoint

Print a connection table (model, runtime, port, OMP threads, KV GB, max-model-len,
NUMA pinning) and a ready-to-run example:
```bash
curl -s http://localhost:/v1/chat/completions -H 'Content-Type: application/json' \
  -d '{"model":"","messages":[{"role":"user","content":"Hello"}]}'
```
To stop: `$RT rm -f vllm-epyc` (container) or `kill ` (conda).

## Offline (single-instance batch)

For a one-shot offline run instead of a server, replace Step 6-8 with a single
`vllm bench throughput` (or an offline `LLM.generate`) using the same sized env,
wait for completion, and report the metrics. Same no-retry / no-debug rule.

## Gotchas

See [reference.md](reference.md) for the full list. The load-bearing ones:

- **`--device cpu` was removed** from `vllm serve` in vLLM >= 0.20. The zentorch
  plugin auto-selects CPU. Passing it makes `vllm serve` error with
  "unrecognized arguments: --device cpu".
- **`TORCHINDUCTOR_FREEZING=1` alone crashes engine-core init** on vLLM 0.23 /
  zentorch 2.11 (`AssertionError: expected OutputCode, got function`). It only
  works with `VLLM_USE_AOT_COMPILE=0` set alongside it. Never set one without
  the other.
- **`--shm-size`**: vLLM needs a large `/dev/shm`; the container default (64MB)
  is too small. Use `--shm-size=16g` (in `data/epyc.json`).
- **NUMA / socket**: one instance is pinned to **one socket plus its memory** --
  CPU bind + `--cpuset-mems` (container) / `numactl --membind` (conda), with KV sized
  from that socket's local RAM. On a dual-socket host `cpu_tune.py` picks a free socket
  by load and `warning`s if both are busy. NPS2/NPS4 (multi-node socket) gets an
  `nps_note` that finer per-node binding could add more.

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [amd](https://github.com/amd)
- **Source:** [amd/skills](https://github.com/amd/skills)
- **License:** MIT
- **Homepage:** https://developer.amd.com/

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** yes
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-amd-skills-serving-llms-on-epyc
- Seller: https://agentstack.voostack.com/s/amd
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
