Install
$ agentstack add skill-byerlikaya-claude-starter-kit-incident-runbook ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
Incident Response & Runbook
Two modes: live incident (what to do right now) and aftermath (postmortem + runbook). Priority: stopping user impact > finding the root cause. No panic, one ordered step at a time.
Live incident — sequence
- Acknowledge & classify — what is the impact (who, how much), severity (SEV1 full outage … SEV3 minor).
- Mitigate the impact FIRST — rollback, turn off a feature flag, shift traffic, scale up. Without waiting on the root cause.
- Single coordinator — it is clear who decides; communication goes through one channel.
- Diagnose — last change? (deploy/migration/config) narrow it down with logs+metrics+traces (observability).
- Resolve — the smallest safe fix; then verify (health check).
- Close — confirm the impact is over; note the timeline (a postmortem input).
Mitigation reflexes
- Last deploy suspect → rollback (vps-deploy revert).
- Suspect feature → turn off the feature flag.
- After a destructive migration → restore from backup (db-migration).
- Dependency/service down → circuit breaker / graceful degradation.
Postmortem (blameless)
Once the incident is resolved, within 24-72 hours:
- Timeline: detection → response → resolution (actual times).
- Impact: who, for how long, what was lost.
- Root cause: "5 whys"; the system/process is questioned, not the person (blameless).
- Actions: concrete, owned, dated items that prevent a recurrence (no deferral).
- If a lasting decision came out of it →
adr.
Produce a runbook
For repeatable incidents, a step-by-step runbook: symptom → diagnostic commands → mitigation → verification → escalation. The runbook must be project-specific and executable (not generic); coordinate with docs-writer.
Invariant rules
- Stop the impact, then understand — the root cause does not hold up the resolution.
- Blameless culture — the postmortem questions the system, not the person.
- Actions are owned + dated — no "we'll look at it later".
- The runbook is executable — real commands/steps, not wishes.
- Make learning permanent — the lesson goes into an adr/runbook/monitoring, it does not get lost.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: byerlikaya
- Source: byerlikaya/claude-starter-kit
- License: MIT
- Homepage: https://www.npmjs.com/package/@byerlikaya/claude-starter-kit
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.