AgentStack
SKILL verified MIT Self-run

Gpu Host Tuning

skill-air-gapped-skills-gpu-host-tuning · by air-gapped

Audit AND tune Linux/GPU inference hosts — read-only host snapshot

No reviews yet
0 installs
11 views
0.0% view→install

Install

$ agentstack add skill-air-gapped-skills-gpu-host-tuning

✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

Are you the author of Gpu Host Tuning? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

gpu-host-tuning

Host-side tuning + audit for Linux GPU inference servers. Sits beneath any inference framework (vLLM, sglang, TensorRT-LLM, llama.cpp). Three modes:

  1. Audit — read-only snapshot
  2. Bench — ground-truth pinned-host↔GPU memcpy ceiling
  3. Tune — apply individual levers from the cheat-sheet

This file is a pointer map. The actual logic lives in scripts/ and the authoritative references in references/.

Quick start

# From the skill directory — typically ~/.claude/skills/gpu-host-tuning
# (personal) or .claude/skills/gpu-host-tuning (project install).

# Audit (read-only, ~60s)
./scripts/collect.sh

# Audit + pinned-memcpy bench (needs torch + CUDA, ~5 min)
./scripts/collect.sh --bench

The script prompts for the output parent dir on first interactive run and remembers the choice. Override via --out or HOST_AUDIT_DIR=. Default snapshot dirname is gpu-host-tuning--.

What the snapshot captures

One file per probe, numbered by section. See [references/probe-interpretation.md](references/probe-interpretation.md) for the full file-by-file decoder.

| Section | What | |---|---| | 00-09 meta | collector version, run timestamp, args | | 10-19 system + firmware | dmidecode (BIOS, CPU, memory DIMMs), lshw, /sys/class/dmi | | 20-29 CPU + power + C-states | governor, EPP, intelpstate / amdpstate, cpuidle states + disable mask, turbostat 5s residency, microcode, vulnerabilities, thermal zones | | 30-39 memory + NUMA | numactl -H, /proc/meminfo, THP, numabalancing, vm tunables, hugepages | | 40-49 kernel + limits | uname, /etc/os-release, /proc/cmdline, sysctl -a, ulimit, /sys/devices/system/cpu/vulnerabilities, dmesg, IRQ affinity, env vars in vllm processes | | 50-59 PCIe | lspci tree + verbose, AER counters, link width/speed for every NVIDIA device | | 60-69 GPU | nvidia-smi -q full, topo -m, nvlink --status, clocks/power/ECC, dmon 5s, dcgmi diag | | 70-79 network | NICs, IB (ibstat / ibvdevinfo), ethtool ring sizes, RDMA links | | 80-89 storage | lsblk, NVMe id-ctrl, smartctl, mount flags, io scheduler | | 90-99 container runtime | containerd version, CDI specs, cgroup v2, kubelet config, RKE2 config |

Three modes — what each maps to

| Mode | What | Reference | |---|---|---| | Audit | ./scripts/collect.sh writes the snapshot directory | references/probe-interpretation.md decodes each numbered file | | Bench | ./scripts/collect.sh --bench adds the pinned-memcpy CSV | references/session-findings.md lists baselines per chassis | | Tune | No script — apply individual levers from the cheat-sheet | references/recommended-tunings.md (lever-by-lever) and references/tuned-profiles.md (apply via tuned-adm) |

When to use which reference

| Goal | Read | |---|---| | Apply NVIDIA's stock DGX tunings | [references/tuned-profiles.md](references/tuned-profiles.md) | | See exactly what NVIDIA's settings packages flip (per-platform JSON, GRUB drop-ins, sysctl, units) | [references/nvidia-dgx-config-decoder.md](references/nvidia-dgx-config-decoder.md) | | Run a proper bring-up flow | [references/bringup-recipe.md](references/bringup-recipe.md) | | Find the lever the audit flagged | [references/recommended-tunings.md](references/recommended-tunings.md) | | Decode an audit output file | [references/probe-interpretation.md](references/probe-interpretation.md) | | Tune a Dell XE9680 (H100/H200, SPR/EMR) | [references/dell-xe9680.md](references/dell-xe9680.md) | | Tune a Dell XE9780 / XE9780L (B200/B300, Granite Rapids) | [references/dell-xe9780.md](references/dell-xe9780.md) | | Understand why cpufreq/cpuidle is empty inside a cloud VM | [references/virt-and-cloud-quirks.md](references/virt-and-cloud-quirks.md) | | See measured baselines from real boxes | [references/session-findings.md](references/session-findings.md) |

Comparing two snapshots

Two snapshots on the same host (e.g., pre-tune and post-tune) can be compared with diff -ruN snap_pre/ snap_post/. For a structured impact ranking, use references/probe-interpretation.md to interpret deltas.

Companion skills

  • [vllm-nvidia-hardware](../vllm-nvidia-hardware/) — per-SKU specs (HBM, TDP, NVLink, PCIe gen)
  • [vllm-deployment](../vllm-deployment/) — K8s manifest authoring, cache mounts, probes
  • [vllm-performance-tuning](../vllm-performance-tuning/) — vLLM-side knobs (above this skill's layer)

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet — be the first.

Versions

  • v0.1.0 Imported from the upstream source.