Install
$ agentstack add skill-param087-agent-ml-skills-pytorch-training-loop Open-source listing — not yet scanned by AgentStack. Follow the source repository for install instructions.
Security review
⚠ Flagged1 finding(s); flagged for manual review. · v0.1.0 How review works →
- • Prompt-injection patterns
- • Secret / credential exfiltration
- • Dangerous shell & filesystem operations
- • Untrusted network calls
- • Known-malicious package signatures
- high Dangerous shell/eval execution.
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ● Dynamic code execution Used
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
PyTorch Training Loop
Overview
A correct PyTorch loop has a precise sequence of operations. Getting the order or the modes wrong produces silent bugs (no gradients, dropout active at eval, leaked compute graphs). This skill encodes the canonical, production-ready loop.
When to use
- Writing a training loop from scratch.
- Debugging a model that won't learn or OOMs.
- Reviewing PyTorch training code.
Canonical loop
import torch
from torch.amp import autocast, GradScaler
device = "cuda" if torch.cuda.is_available() else "cpu"
model.to(device)
optimizer = torch.optim.AdamW(model.parameters(), lr=3e-4, weight_decay=0.01)
scaler = GradScaler(enabled=(device == "cuda"))
best_val = float("inf")
for epoch in range(num_epochs):
# ---- TRAIN ----
model.train()
for x, y in train_loader:
x, y = x.to(device, non_blocking=True), y.to(device, non_blocking=True)
optimizer.zero_grad(set_to_none=True)
with autocast(device_type=device, enabled=(device == "cuda")):
out = model(x)
loss = criterion(out, y)
scaler.scale(loss).backward()
scaler.unscale_(optimizer)
torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)
scaler.step(optimizer)
scaler.update()
# ---- VALIDATE ----
model.eval()
val_loss = 0.0
with torch.no_grad():
for x, y in val_loader:
x, y = x.to(device), y.to(device)
val_loss += criterion(model(x), y).item() * x.size(0)
val_loss /= len(val_loader.dataset)
# ---- CHECKPOINT BEST ----
if val_loss < best_val:
best_val = val_loss
torch.save({"model": model.state_dict(),
"optimizer": optimizer.state_dict(),
"epoch": epoch}, "best.pt")
Non-negotiable rules
model.train()before training,model.eval()before validation/inference (toggles dropout & batchnorm).optimizer.zero_grad()every step — gradients accumulate otherwise.- Wrap validation/inference in
torch.no_grad()(orinference_mode()) to save memory. - Detach when logging:
loss.item(), notloss— keeping tensors leaks the graph and OOMs. - Clip gradients for RNNs/transformers to prevent explosions.
Reproducibility header
import torch, numpy as np, random
def seed_everything(seed=42):
random.seed(seed); np.random.seed(seed)
torch.manual_seed(seed); torch.cuda.manual_seed_all(seed)
torch.backends.cudnn.deterministic = True
torch.backends.cudnn.benchmark = False
Pitfalls
- Forgetting
zero_grad→ gradients pile up, training diverges. eval()never called → dropout/batchnorm corrupt validation metrics.- Accumulating
total_loss += loss(tensor, not.item()) → memory explosion. - DataLoader with
num_workers=0on large data → CPU-bound; raise workers +pin_memory=True. - LR too high → NaN loss; see the
ml-debuggingskill.
Hand-off
A checkpointed model + seed config that experiment-tracking logs and model-serving loads for inference.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: param087
- Source: param087/agent-ml-skills
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.