AgentStack
SKILL verified MIT Self-run

Incident Runbook

skill-byerlikaya-claude-starter-kit-incident-runbook · by byerlikaya

|

No reviews yet
0 installs
8 views
0.0% view→install

Install

$ agentstack add skill-byerlikaya-claude-starter-kit-incident-runbook

✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

Are you the author of Incident Runbook? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Incident Response & Runbook

Two modes: live incident (what to do right now) and aftermath (postmortem + runbook). Priority: stopping user impact > finding the root cause. No panic, one ordered step at a time.

Live incident — sequence

  1. Acknowledge & classify — what is the impact (who, how much), severity (SEV1 full outage … SEV3 minor).
  2. Mitigate the impact FIRST — rollback, turn off a feature flag, shift traffic, scale up. Without waiting on the root cause.
  3. Single coordinator — it is clear who decides; communication goes through one channel.
  4. Diagnose — last change? (deploy/migration/config) narrow it down with logs+metrics+traces (observability).
  5. Resolve — the smallest safe fix; then verify (health check).
  6. Close — confirm the impact is over; note the timeline (a postmortem input).

Mitigation reflexes

  • Last deploy suspect → rollback (vps-deploy revert).
  • Suspect feature → turn off the feature flag.
  • After a destructive migration → restore from backup (db-migration).
  • Dependency/service down → circuit breaker / graceful degradation.

Postmortem (blameless)

Once the incident is resolved, within 24-72 hours:

  • Timeline: detection → response → resolution (actual times).
  • Impact: who, for how long, what was lost.
  • Root cause: "5 whys"; the system/process is questioned, not the person (blameless).
  • Actions: concrete, owned, dated items that prevent a recurrence (no deferral).
  • If a lasting decision came out of it → adr.

Produce a runbook

For repeatable incidents, a step-by-step runbook: symptom → diagnostic commands → mitigation → verification → escalation. The runbook must be project-specific and executable (not generic); coordinate with docs-writer.

Invariant rules

  1. Stop the impact, then understand — the root cause does not hold up the resolution.
  2. Blameless culture — the postmortem questions the system, not the person.
  3. Actions are owned + dated — no "we'll look at it later".
  4. The runbook is executable — real commands/steps, not wishes.
  5. Make learning permanent — the lesson goes into an adr/runbook/monitoring, it does not get lost.

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet — be the first.

Versions

  • v0.1.0 Imported from the upstream source.