Investigate Terraform Drift
Determine why actual infrastructure differs from Terraform state and choose a safe reconciliation path.
Example Skill Name
One-line description of when this skill is used.
Triage Cloud Cost Spike
Identify the primary drivers of a sudden cloud cost increase and implement safe mitigations.
Recover Terraform State Lock
Safely assess and recover from a stuck Terraform state lock.
Triage Pending Pods
Diagnose pods stuck in Pending and identify scheduling constraints.
Triage Kubernetes Node Pressure
Diagnose node-level memory/disk/pid pressure and determine safe mitigations.
Triage Latency Regression
Identify what changed and where latency increased, using logs/metrics/traces.
Diagnose CrashLoopBackOff
Triage pods restarting repeatedly and identify the most likely root cause.
Sev1 First 15 Minutes
Execute the initial incident workflow to stabilize, communicate, and delegate.
Triage GCP Quota Exceeded
Diagnose GCP quota errors and identify the quota, scope, and fastest safe mitigation.
Triage Error Budget Burn
Investigate rapid SLO burn and identify whether the driver is errors, latency, or availability.
Triage Suspected Secret Exposure
Contain and respond to a suspected credential/secret exposure without increasing blast radius.
Triage Argo CD App OutOfSync
Identify why an Argo CD application is OutOfSync and resolve safely.
Triage AWS AccessDenied
Identify why an AWS API call is denied and what policy element blocks it.
Diagnose ImagePullBackOff
Determine why a pod cannot pull its container image and resolve safely.
Triage Kubernetes Service DNS
Diagnose in-cluster DNS resolution failures and isolate root causes.
Triage EKS Node NotReady
Diagnose EKS worker nodes in NotReady and determine safe remediation.