Install
$ agentstack add skill-amd-skills-serving-llms-on-epyc ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ● Network access Used
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
Serving LLMs on AMD EPYC (vLLM + zentorch, CPU)
Bring up a single vLLM OpenAI endpoint on an AMD EPYC host with the zentorch CPU backend, sized to the hardware. Container-first (Docker or Podman); conda/host is the fallback.
This is single-socket serving: one instance pinned to one socket and its memory (vLLM scales poorly across sockets, so we do not span them). On a dual-socket host it runs on a single socket; the multi-socket answer is multiple instances (one per socket), which is out of scope for this single-instance recipe.
Hard rule for this skill: on any failure, report the cause + logs and STOP. Do not retry, do not debug. (Debugging is a separate workflow.)
The agent does the serve flow itself -- pull, configure, launch, poll -- using the runtime validate.py reports. Never hand the user per-serve commands. Like serving-llms-on-instinct, an accessible container runtime is a one-time prerequisite: if validate.py finds none, report its one-time fix (make docker accessible / install podman / provide a conda env) and stop. Do not attempt sudo or privilege escalation.
Data file
Read data/epyc.json directly. It holds the container image, mandatory CPU run flags, supported precision, the model-support policy, the default model, and the verified throughput-flag gotcha. Do not hardcode the image tag from memory -- read it.
Step 1: Detect the CPU
python3 scripts/detect.py # add --host user@box for a remote host
Returns cpu_model, is_amd_epyc, epyc_generation (Naples/Rome/Milan/Genoa/Bergamo/Siena/Turin), zen_arch, avx512, logical_cores, physical_cores, sockets, numa_nodes, memory_gb. If is_amd_epyc is false, stop: this skill targets AMD EPYC. (Other x86 may work but is unsupported here.) Carry epyc_generation / avx512 through the later phases -- e.g. AVX-512 + bf16 land on Zen4+ (Genoa/Turin), and Turin packs up to 128 cores/socket, which the thread-binding in Step 5 sizes from.
Step 2: Validate the runtime and environment
python3 scripts/validate.py --image
Returns ready, runtime (docker, podman, or null), runtime_detail, conda_path_available, ram_gb, and errors/warnings/advisories. Pick the path:
runtimeisdockerorpodman-> container path (Step 6), used verbatim.runtimenull butconda_path_available: true-> conda/host path.runtimenull and no conda ->readyis false. Report the one-time
onboarding fix (make docker accessible / install podman / conda env) and stop.
Do not proceed if ready is false.
Step 3: Resolve and validate the model
If the user named no model, use default_model from data/epyc.json (Qwen/Qwen3-0.6B -- ungated, tiny, fast first success). Otherwise use theirs.
Check that vLLM actually supports the model (do not blanket-block multimodal):
python3 scripts/check_model.py --model-id --vllm-version
- Exit 0 = vLLM serves it as a generation endpoint (
kindtextormultimodal),
or support is undeterminable (gated/offline) -- proceed; launch confirms.
- Exit 1 = positively unsupported: the architecture is not in vLLM's registry, or
it is a pooling/embedding/reranker (not a chat/completion endpoint). Report the printed message and stop.
- A
multimodalmodel is allowed; a vLLM-supported multimodal arch may still hit a
GPU-only kernel on CPU, which surfaces at load (the no-retry rule then applies).
Precision/dtype: native CPU dtypes are bf16 (default), fp16, fp32. Use bfloat16 unless the user asks otherwise.
For gated models (Llama, Gemma) HF_TOKEN must be set and the license accepted on HuggingFace; if not, stop and say so.
Step 4: Check it fits host RAM
RAM is the ceiling on CPU (weights + KV cache both live in RAM). Run on ONE line:
python3 scripts/estimate_memory.py --model-id --ram-gb --max-model-len --num-prompts
Exit 0 = fits, exit 1 = does not fit. If fit.fits is false: do not launch. Tell the user required_gb vs ram_gb and the printed fit.action -- reduce --max-model-len to fit.suggested_max_model_len and retry, or use a smaller model. --max-model-len and --num-prompts are the two knobs that move KV. Extra flag: --weight-gb N overrides weights if a model has no HF metadata (rare). KV cache is bf16-only on zentorch CPU (no fp8 KV).
Step 5: Size the CPU runtime from the hardware
eval "$(python3 scripts/cpu_tune.py)" # or --format json to inspect
A single instance runs on one socket, with its memory (vLLM scales poorly across sockets). cpu_tune.py exports VLLM_CPU_OMP_THREADS_BIND (the chosen socket's physical cores) and VLLM_CPU_KVCACHE_SPACE (sized from that socket's local RAM, not whole-system, so the KV pool stays on-socket). It does not set OMP_NUM_THREADS (vLLM derives it) or VLLM_CPU_NUM_OF_RESERVED_CPU (vLLM's own default).
Socket choice on a dual-socket host (load-aware): it samples per-socket CPU busy% (~0.5s) and prefers a free socket -- both free → socket 0; one free → that socket; both busy (≥ --busy-threshold, default 15%) → it warnings and proceeds on the least-busy socket. --socket N forces a choice. Single-socket hosts use socket 0.
For the chosen socket it also emits the memory-bound pin: container_cpuset (--cpuset-cpus= --cpuset-mems=) for the container path, and conda_launch_prefix (numactl --cpunodebind/--membind, falling back to taskset CPU-only, or empty-with-note if neither tool exists) for conda. Surface warning to the user if set. On NPS2/NPS4 a socket spans multiple NUMA nodes; memory is bound across them and nps_note flags that finer binding could add performance.
Step 6: Confirm the plan, then launch (container-first)
Before launching, present this summary and wait for the user to confirm -- do not launch unprompted. This is the human gate before anything runs:
| Field | Value | |---|---| | Model / kind | ` -- text or multimodal (from check_model.py) | | Path | container (, image from data/epyc.json) or conda/host | | Precision | bfloat16 (or the user's choice) | | Fit | required GB vs GB RAM | | CPU sizing | socket (), bind , KV GB (socket-local), mem bound to nodes | | Hardware | EPYC (), cores, AVX-512 | | Port | ` |
If cpu_tune.py returned a warning (e.g. all sockets busy), include it here so the user sees it before confirming.
Proceed only on a clear "go". If the user declines or wants changes (model, --max-model-len, port), stop and adjust -- do not launch.
Build the launch from data/epyc.json. The CLI is vllm serve . Do not pass --device cpu on vLLM >= 0.20 -- the zentorch plugin auto-selects the CPU platform and vllm serve rejects the flag. Only add it if vllm serve --help lists it (older vLLM).
Container path (runtime from validate.py). The agent runs these itself, including the pull. RT is the resolved runtime verbatim:
RT=""
$RT pull # agent pulls; do not ask the user to
$RT run -d --name vllm-epyc \
# --ipc=host --shm-size=16g --network=host
\
# --cpuset-cpus= --cpuset-mems=
--env VLLM_CPU_OMP_THREADS_BIND="$VLLM_CPU_OMP_THREADS_BIND" \
--env VLLM_CPU_KVCACHE_SPACE=$VLLM_CPU_KVCACHE_SPACE \
--env HF_TOKEN=${HF_TOKEN} \
\
vllm serve --dtype bfloat16 --port --max-model-len
Conda/host path (no container runtime, conda_path_available true). eval-ing cputune already exported the env vars; prefix the launch with conda_launch_prefix from cputune so memory is bound to the chosen socket (empty → unpinned, with a note):
vllm serve --dtype bfloat16 --port --max-model-len &
# e.g. numactl --cpunodebind=0 --membind=0 vllm serve ...
Optional throughput flags are opt-in and must move together (see Gotchas): TORCHINDUCTOR_FREEZING=1 + VLLM_USE_AOT_COMPILE=0 (+ ZENTORCH_WEIGHT_PREPACK=1). The base launch sets none of them.
Step 7: Poll until up and responsive
A 503 while loading is normal. Poll until the server answers, then prove the chat endpoint works. CPU first-token compile can take a minute or two.
# container alive (or process alive for conda) + /health
for i in $(seq 1 120); do
# container path:
$RT inspect -f '{{.State.Running}}' vllm-epyc 2>/dev/null | grep -q true || { echo "FAILED: container exited"; $RT logs --tail 50 vllm-epyc; break; }
curl -sf http://localhost:/health >/dev/null 2>&1 && { echo "HEALTHY"; break; }
sleep 3
done
Then validate the OpenAI endpoint is actually accessible:
curl -sf http://localhost:/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"model":"","messages":[{"role":"user","content":"hi"}],"max_tokens":8}'
Resource sanity (your validation list): $RT stats --no-stream vllm-epyc.
If the server never becomes healthy or the endpoint does not respond: print the container/process logs, state the failure, and STOP. Do not retry. Do not start a debugging loop.
Step 8: On success, hand over the endpoint
Print a connection table (model, runtime, port, OMP threads, KV GB, max-model-len, NUMA pinning) and a ready-to-run example:
curl -s http://localhost:/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"model":"","messages":[{"role":"user","content":"Hello"}]}'
To stop: $RT rm -f vllm-epyc (container) or kill (conda).
Offline (single-instance batch)
For a one-shot offline run instead of a server, replace Step 6-8 with a single vllm bench throughput (or an offline LLM.generate) using the same sized env, wait for completion, and report the metrics. Same no-retry / no-debug rule.
Gotchas
See [reference.md](reference.md) for the full list. The load-bearing ones:
--device cpuwas removed fromvllm servein vLLM >= 0.20. The zentorch
plugin auto-selects CPU. Passing it makes vllm serve error with "unrecognized arguments: --device cpu".
TORCHINDUCTOR_FREEZING=1alone crashes engine-core init on vLLM 0.23 /
zentorch 2.11 (AssertionError: expected OutputCode, got function). It only works with VLLM_USE_AOT_COMPILE=0 set alongside it. Never set one without the other.
--shm-size: vLLM needs a large/dev/shm; the container default (64MB)
is too small. Use --shm-size=16g (in data/epyc.json).
- NUMA / socket: one instance is pinned to one socket plus its memory --
CPU bind + --cpuset-mems (container) / numactl --membind (conda), with KV sized from that socket's local RAM. On a dual-socket host cpu_tune.py picks a free socket by load and warnings if both are busy. NPS2/NPS4 (multi-node socket) gets an nps_note that finer per-node binding could add more.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: amd
- Source: amd/skills
- License: MIT
- Homepage: https://developer.amd.com/
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.