# Site Reliability Engineer

> SRE skills for reliability, monitoring, SLAs and incident management in production. Trigger this skill for availability questions, alerting, incident post-mortems, or defining reliability objectives (SLO/SLA) — complementary to DevOps, which focuses on deployment.

- **Type:** Skill
- **Install:** `agentstack add skill-benbasse-claude-skills-digital-solutions-sre`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [benbasse](https://agentstack.voostack.com/s/benbasse)
- **Installs:** 0
- **Category:** [Cloud & Infrastructure](https://agentstack.voostack.com/c/cloud-infrastructure)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [benbasse](https://github.com/benbasse)
- **Source:** https://github.com/benbasse/claude-skills-digital-solutions/tree/master/skills/dev/sre

## Install

```sh
agentstack add skill-benbasse-claude-skills-digital-solutions-sre
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# SRE (Site Reliability Engineer)

The SRE defines and protects the reliability users actually experience through clear service objectives, proactive monitoring and structured incident management.

## When to trigger this skill

- Defining SLOs/SLAs (e.g. availability of a critical API)
- Setting up alerting and system health dashboards
- Managing and running a post-mortem for a production incident
- Capacity planning ahead of a predictable high-load period

## Skills, responsibilities and best practices

- SLOs defined around real user experience, not just raw technical metrics
- Actionable alerting (every alert should map to a possible action)
- Blameless post-mortems focused on systemic causes
- Documented runbooks for recurring incidents
- Load testing ahead of predictable peaks (seasonal spikes, campaigns)

## Common pitfalls to avoid

- Piling on non-actionable alerts (alert fatigue, real incidents get ignored)
- No post-mortem after an incident -> the same failure repeats
- Not anticipating predictable load spikes

## Reference stack and tools

- Uptime monitoring (e.g. UptimeRobot, Better Stack)
- Centralized logs
- Application and infrastructure metric dashboards

## Typical deliverables

- Documented SLOs/SLAs
- Incident runbooks
- Post-mortem report

## Example prompts that should trigger this skill

- 'Define availability objectives for the order-tracking API'
- 'Run the post-mortem for last night's payment outage'

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [benbasse](https://github.com/benbasse)
- **Source:** [benbasse/claude-skills-digital-solutions](https://github.com/benbasse/claude-skills-digital-solutions)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-benbasse-claude-skills-digital-solutions-sre
- Seller: https://agentstack.voostack.com/s/benbasse
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
