Triage Cloud Cost Spike
Identify the primary drivers of a sudden cloud cost increase and implement safe mitigations.
Recover Terraform State Lock
Safely assess and recover from a stuck Terraform state lock.
Triage Kubernetes Service DNS
Diagnose in-cluster DNS resolution failures and isolate root causes.
Sev1 First 15 Minutes
Execute the initial incident workflow to stabilize, communicate, and delegate.
Diagnose CrashLoopBackOff
Triage pods restarting repeatedly and identify the most likely root cause.
Triage Pending Pods
Diagnose pods stuck in Pending and identify scheduling constraints.
Triage Kubernetes Node Pressure
Diagnose node-level memory/disk/pid pressure and determine safe mitigations.
Triage Suspected Secret Exposure
Contain and respond to a suspected credential/secret exposure without increasing blast radius.
Triage Latency Regression
Identify what changed and where latency increased, using logs/metrics/traces.
Investigate Terraform Drift
Determine why actual infrastructure differs from Terraform state and choose a safe reconciliation path.
Triage Error Budget Burn
Investigate rapid SLO burn and identify whether the driver is errors, latency, or availability.
Triage GCP Quota Exceeded
Diagnose GCP quota errors and identify the quota, scope, and fastest safe mitigation.
Triage Argo CD App OutOfSync
Identify why an Argo CD application is OutOfSync and resolve safely.
Diagnose ImagePullBackOff
Determine why a pod cannot pull its container image and resolve safely.
Triage AWS AccessDenied
Identify why an AWS API call is denied and what policy element blocks it.
Example Skill Name
One-line description of when this skill is used.
Triage EKS Node NotReady
Diagnose EKS worker nodes in NotReady and determine safe remediation.