# Incident Slo Runbook

> Create or audit SLOs, SLIs, alert rules, incident response steps, escalation paths, postmortems, operational runbooks, and customer-impact communication. Use when defining production reliability, preparing launch readiness, responding to an outage, writing a runbook, tuning alerts, or closing the loop after an incident.

- **Type:** Skill
- **Install:** `agentstack add skill-majiayu000-spellbook-incident-slo-runbook`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [majiayu000](https://agentstack.voostack.com/s/majiayu000)
- **Installs:** 0
- **Category:** [Agent Skills](https://agentstack.voostack.com/c/agent-skills)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [majiayu000](https://github.com/majiayu000)
- **Source:** https://github.com/majiayu000/spellbook/tree/main/skills/incident-slo-runbook
- **Website:** https://github.com/majiayu000/spellbook#quick-start

## Install

```sh
agentstack add skill-majiayu000-spellbook-incident-slo-runbook
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# Incident SLO Runbook

## Purpose

Use this skill to connect observability to action. Metrics and logs are not enough; each critical user journey needs an SLO, alert, owner, response path, and post-incident learning loop.

## SLO Design

Define:

1. User journey or system capability.
2. SLI: request success, latency, freshness, durability, or job completion.
3. SLO target and measurement window.
4. Error budget and burn-rate alerts.
5. Exclusions with rationale.
6. Dashboard and data source.
7. Owner and escalation path.

Avoid vanity metrics. Prefer user-visible success and latency over internal counters unless internal counters are the only reliable proxy.

## Runbook Requirements

Each runbook should include:

- Symptom and alert name.
- Impacted users or systems.
- First 5-minute checks.
- Triage decision tree.
- Mitigation steps with commands.
- Rollback or failover path.
- Escalation owner.
- Customer/support communication note.
- Postmortem trigger.

Commands must be safe to run or explicitly labeled destructive.

## Incident Flow

1. Declare severity and incident commander.
2. Confirm impact from live evidence.
3. Stabilize with the lowest-risk mitigation.
4. Communicate status on a fixed cadence.
5. Preserve evidence before cleanup.
6. Write a blameless postmortem with action items and owners.

## Output Shape

```text
service_or_journey:
slo:
alerts:
dashboard_or_queries:
runbook:
escalation:
postmortem_template:
verification:
```

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [majiayu000](https://github.com/majiayu000)
- **Source:** [majiayu000/spellbook](https://github.com/majiayu000/spellbook)
- **License:** MIT
- **Homepage:** https://github.com/majiayu000/spellbook#quick-start

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-majiayu000-spellbook-incident-slo-runbook
- Seller: https://agentstack.voostack.com/s/majiayu000
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
