Install
$ agentstack add skill-togethercomputer-skills-together-gpu-clusters ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Together GPU Clusters
Overview
Use Together AI GPU clusters when the user needs infrastructure control instead of a managed inference product.
Typical fits:
- distributed training
- multi-node inference
- HPC or Slurm workloads
- custom Kubernetes jobs
- attached shared storage and cluster lifecycle management
When This Skill Wins
- Provision a cluster and manage it over time
- Choose between on-demand and reserved capacity
- Choose Kubernetes or Slurm as the orchestration layer
- Manage shared volumes and credentials
- Scale up, scale down, or troubleshoot node health
Hand Off To Another Skill
- Use
together-dedicated-endpointsfor managed single-model hosting - Use
together-dedicated-containersfor containerized inference without owning the full cluster - Use
together-sandboxesfor short-lived remote Python execution - Use
together-fine-tuningfor managed training jobs instead of raw cluster operations
Quick Routing
- Cluster creation, scaling, credentials, deletion
- Start with [scripts/managecluster.py](scripts/managecluster.py) or [scripts/managecluster.ts](scripts/managecluster.ts)
- Read [references/api-reference.md](references/api-reference.md)
- Shared storage lifecycle
- Use [scripts/managestorage.py](scripts/managestorage.py)
- Read [references/api-reference.md](references/api-reference.md)
- Kubernetes vs Slurm operations
- Read [references/cluster-management.md](references/cluster-management.md)
- Troubleshooting node health, PVCs, or scheduling
- Read [references/cluster-management.md](references/cluster-management.md)
- Together CLI workflows
- Read [references/cli.md](references/cli.md)
Workflow
- Decide whether the workload really needs cluster-level control.
- Choose on-demand vs reserved billing based on run duration and baseline utilization.
- Choose Kubernetes vs Slurm based on orchestration requirements and team tooling.
- Select region, GPU type, driver version, and shared storage plan.
- Provision first, then layer in access credentials, workload deployment, scaling, and health checks.
High-Signal Rules
- Python scripts require the Together v2 SDK (
together>=2.0.0). If the user is on an older version, they must upgrade first:uv pip install --upgrade "together>=2.0.0". - Prefer managed products unless the user explicitly needs raw infrastructure control.
- Treat storage lifecycle separately from cluster lifecycle; volumes can outlive clusters.
- When creating a cluster with new shared storage, prefer inline
shared_volumeover creating a volume separately and attaching viavolume_id. Separately created volumes may land in a different datacenter partition than the cluster, causing a "does not exist in the datacenter" error even when the volume shows as available. - GPU stock-outs (409 "Out of stock") are common. Always call
list_regions()first and be prepared to try multiple regions. - The API requires
cuda_versionandnvidia_driver_versionas separate fields in addition to the combineddriver_versionstring. Pass them viaextra_bodyin the Python SDK. - Credentials retrieval is part of provisioning. Do not stop at cluster creation if the user needs to run workloads immediately.
- Slurm and Kubernetes operational patterns differ materially; read the cluster-management reference before improvising.
- For repeated cluster operations, start from the scripts instead of rebuilding request shapes.
- Slurm startup scripts (worker/login init, worker/controller prolog and epilog, extra
slurm.conf) are Slinky v1.0 only. A non-zero exit from a worker prolog or epilog drains the node, and calling Slurm commands (squeue,scontrol,sacctmgr) inside any prolog/epilog can deadlock the scheduler.
Resource Map
- Cluster API reference: [references/api-reference.md](references/api-reference.md)
- Operational guide: [references/cluster-management.md](references/cluster-management.md)
- Operational troubleshooting: [references/cluster-management.md](references/cluster-management.md)
- CLI guide: [references/cli.md](references/cli.md)
- Python cluster management: [scripts/managecluster.py](scripts/managecluster.py)
- TypeScript cluster management: [scripts/managecluster.ts](scripts/managecluster.ts)
- Python storage management: [scripts/managestorage.py](scripts/managestorage.py)
Official Docs
- GPU Clusters Overview
- GPU Clusters Quickstart
- Clusters API
- Slurm Startup Scripts
- Instant GPU Clusters
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: togethercomputer
- Source: togethercomputer/skills
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.