Install
$ agentstack add skill-lifecycle-innovations-limited-claude-ops-ops-fires ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ● Network access Used
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
OPS ► FIRES
Runtime Context
Before executing, load available context:
- Daemon health: Read
${CLAUDE_PLUGIN_DATA_DIR:-$HOME/.claude/plugins/data/ops-ops-marketplace}/daemon-health.json
- Check
infra-monitorservice status — if not running, pre-gathered infra data may be stale - If
action_neededis not null → surface it immediately as a potential fire
- Secrets: AWS credentials are required for ECS/CloudWatch queries.
### Secret Resolution
- First: check
$AWS_ACCESS_KEY_ID/$AWS_PROFILEenv vars - Then:
doppler secrets get AWS_ACCESS_KEY_ID --plain(ifdopplerconfigured in prefs) - Then: use
password_manager_config.query_cmdfrom preferences - Sentry token:
$SENTRY_AUTH_TOKEN→ DopplerSENTRY_AUTH_TOKEN→ vault
- Preferences: Read
${CLAUDE_PLUGIN_DATA_DIR}/preferences.jsonforsecrets_managerconfig to know which vault to query.
CLI/API Reference
aws CLI
| Command | Usage | Output | | --------------------------------------------------------------------------------------------------------------------------------------------- | -------------- | ------------------------------------- | | aws ecs list-services --cluster --query 'serviceArns' | ECS services | ARN list | | aws ecs describe-services --cluster --services --query 'services[0].{status:status,running:runningCount,desired:desiredCount}' | Service health | JSON | | aws logs tail /ecs/ --since 1h --format short | ECS logs | Log lines (use with Monitor for live) |
gh CLI (GitHub)
| Command | Usage | Output | | --------------------------------------------------------------------------- | -------------- | ---------- | | gh run list --limit 20 --json status,conclusion,name,headBranch,createdAt | Recent CI runs | JSON array | | gh run view --repo --log-failed | Failed CI logs | Log output |
sentry-cli / Sentry API
| Command | Usage | Output | | -------------------------------------------------------------------------------------------------------------------------------- | --------------------------------- | ---------- | | sentry-cli issues list --project --status unresolved | Unresolved issues | Issue list | | curl -H "Authorization: Bearer $SENTRY_AUTH_TOKEN" "https://sentry.io/api/0/projects///issues/?query=is:unresolved" | API fallback when MCP unavailable | JSON array |
Agent Teams support
If CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1 is set, use Agent Teams when dispatching multiple fix agents simultaneously. This enables:
- Fix agents share findings (e.g., API agent discovers DB is the root cause → infra agent pivots to DB fix)
- You can prioritize: "CRITICAL ECS issue first, then CI failures"
- Real-time progress: agents report as they find root causes, you can merge fixes in optimal order
Team setup (only when flag is enabled, dispatch phase):
TeamCreate("fire-fixers")
Agent(team_name="fire-fixers", name="fix-[service]", ...)
If the flag is NOT set, use standard parallel subagents.
Pre-gathered infrastructure data
```! ${CLAUDEPLUGINROOT}/bin/ops-infra 2>/dev/null || echo '{"clusters":[],"error":"infra check failed"}'
### FinOps dashboard — open anomalies
Live anomaly feed from finops-dashboard (spend spikes, idle services,
expired credits, drift detections). High-severity items belong in the
FIRES table alongside infra outages. Falls open to `[]` if the dashboard
isn't configured.
```!
${CLAUDE_PLUGIN_ROOT}/scripts/finops-bridge.sh anomalies high 2>/dev/null || echo "[]"
CI failures (last 24h)
```! ${CLAUDEPLUGINROOT}/bin/ops-ci 2>/dev/null || echo '[]'
## External projects health
```!
${CLAUDE_PLUGIN_ROOT}/bin/ops-external 2>/dev/null || echo '[]'
Home automation (only if home_automation is configured in $PREFS_PATH)
```! if jq -e '.homeautomation' "${CLAUDEPLUGINDATADIR:-$HOME/.claude/plugins/data/ops-ops-marketplace}/preferences.json" >/dev/null 2>&1; then ${CLAUDEPLUGINROOT}/bin/ops-home snapshot 2>/dev/null || echo '{"configured":true,"error":"home probe failed"}' else echo '{"configured":false}' fi
## Your task
Analyze the pre-gathered data — including external projects. Then run parallel checks:
1. **ECS health** — parse infra data for unhealthy services, stopped tasks, failed deployments.
2. **Sentry** — if Sentry MCP is connected, query recent unresolved errors. Otherwise note it's unavailable.
3. **CI** — parse CI data for failing pipelines, broken main/dev branches.
4. **GitHub Actions** — `gh run list --limit 20 --json status,conclusion,name,headBranch,createdAt 2>/dev/null`
5. **External projects** — parse ops-external data. Flag `auth_expired` as HIGH (credential rotation needed), `unreachable`/`degraded` as MEDIUM, `not_configured` as LOW.
6. **Home automation** (only if home snapshot returned `configured:true`) — classify Homey incidents:
- Active critical alarm (smoke / water leak / security breach) → **P0 / CRITICAL** — cross-reference `/ops:ops-home alarm` for details.
- Major device offline (gateway, hub, primary thermostat) → **P1 / HIGH**.
- Energy spike > 3× 7-day baseline → **P2 / MEDIUM** — cross-reference `/ops:ops-home status`.
If snapshot returned `configured:false`, skip silently.
Classify each issue by severity:
| Severity | Criteria |
| -------- | ------------------------------------------------- |
| CRITICAL | Service down, DB unreachable, auth broken |
| HIGH | Elevated error rate, deploy stuck, CI main broken |
| MEDIUM | Non-critical service degraded, flaky tests |
| LOW | Warning-level, non-urgent |
---
## Output format
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ OPS ► FIRES DASHBOARD — [timestamp] ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
CRITICAL [service] — [issue] — [since]
HIGH [service] — [issue] — [since]
MEDIUM [service] — [issue] — [since]
ECS HEALTH [cluster] [service] [desired/running] [status]
CI STATUS [repo] [branch] [workflow] [status] [last run]
SENTRY (top errors, 24h) [error] [count] [first seen] [project]
EXTERNAL PROJECTS [alias] [source] [status] [details — e.g. auth_expired, unreachable]
HOME (only if home_automation is configured) [alarm/device] [type — smoke/water/security/offline/energy] [severity] [since] [If configured:false, omit this section entirely]
──────────────────────────────────────────────────────
Use **batched AskUserQuestion calls** (max 4 options each). Only show relevant actions (e.g., skip dispatch options if no issues found):
AskUserQuestion call 1:
[Dispatch fix agent for [top critical issue]] [Dispatch fix agent for [second issue]] [View logs for [service]] [More...]
AskUserQuestion call 2 (only if "More..."):
[Open Sentry dashboard] [Open GitHub Actions] [All clear — nothing to do]
If no fires: show "ALL SYSTEMS OPERATIONAL" with last-checked timestamps.
---
## Pre-dispatch staleness check (MANDATORY)
The pre-gathered CI data is cached and may be minutes-to-hours old. Before
dispatching ANY fix agent, verify the failure is still red on its branch HEAD.
This is defense-in-depth: even after the `bin/ops-ci` "current-state" filter
(which only emits workflows whose latest run on a tracked branch is failing),
a fix may have landed in the seconds since the cache was written. Dispatching
to a self-resolved fire wastes Sonnet quota — typically 50–150k tokens per
agent before it figures out there's nothing to fix.
For each fire the user selects:
```bash
gh run list --repo "$REPO" --workflow "$WORKFLOW" --branch "$BRANCH" --limit 1 \
--json conclusion,databaseId,createdAt --jq '.[0]'
- If
conclusion == "success"→ SKIP. Mark task completed with metadata{resolution: "self-resolved-pre-dispatch"}. Do NOT spawn agent. - If
conclusion == "failure"→ proceed to dispatch. - If
conclusion == null(in_progress) → wait 30s, recheck once, then proceed if still null.
For workflows scoped only to PRs (no main/dev runs), check the PR's combined CI status instead: gh pr checks --repo "$REPO" --json bucket,name.
Dispatch fix agent
When user selects to fix an issue, use AskUserQuestion to confirm the scope before dispatching:
Dispatch fix agent for: [issue title]
Severity: [CRITICAL/HIGH/MEDIUM]
Repo: [repo]
Error: [brief description]
The agent will:
- Investigate root cause in [repo]
- Create feature branch with fix
- Open PR for review
[Dispatch agent] [Show me the logs first] [Skip — I'll fix manually]
On confirmation, spawn an Agent with:
- The error details and logs
- Access to the relevant repo
- Instruction to create a feature branch, fix, and open a PR
- Report back when done or blocked
Use the agents/infra-monitor.md agent definition for infra issues.
If $ARGUMENTS contains a project alias, filter to that project's services only.
Native tool usage
Monitor — live service health
Use Monitor to stream ECS task logs or GitHub Actions runs when investigating fires:
Monitor(command: "aws logs tail /ecs/ --follow --since 5m")
Tasks — incident tracking
Use TaskCreate for each active fire. Update with TaskUpdate as fires are investigated/fixed/escalated.
WebFetch — status pages
When diagnosing fires, use WebFetch to check AWS status page (https://health.aws.amazon.com/health/status), Vercel status, or third-party API status pages.
WebSearch — known outage patterns
Use WebSearch to find if the error pattern matches a known AWS/infrastructure issue (e.g., "ECS task stopped CannotPullContainerError" → known ECR throttling).
Credential Expiry & Rate Limit Warnings (Phase 16)
The ops-daemon surfaces two additional fire categories in daemon-health.json:
credential_warnings— tokens/keys expiring within 7 days OR API keys older than 180 days. Fed by offline inspection ofpreferences.json(*_expires_at,*_created_atfields). No live API calls are made to validate credentials.rate_limit_warnings— integrations currently at ≥80% of their quota window. Fed by counters inrate-limits.json. Resets automatically when the window rolls over.
/ops:fires lists both alongside Sentry / infra / CI issues. Push notifications are dispatched by the daemon on the first crossing of the threshold — not re-sent until the next day (credentials) or window rollover (rate limits).
Ledger Integration
CLAIM_KEY: sentry:issue: (e.g. sentry:issue:MY-PROJECT-1A2B)
For non-Sentry fires (infra, CI, credential expiry), use:
- CI failure:
ci:run:: - Credential expiry:
credential:expiry:
Pre-flight skip-check
CLAIM_KEY="sentry:issue:"
ledger query --claim-key "$CLAIM_KEY" --since=-PT24H
If in_progress or done exists, skip the issue. If awaiting_sam exists, surface it as "fix already staged — needs your decision."
Claim + resolve
# Claim when beginning to investigate/fix
ledger write \
--claim-key "$CLAIM_KEY" \
--kind "fix" \
--status "in_progress" \
--title "Fire: " \
--ttl-sec 7200
# Resolve after fix is applied or escalated
ledger write \
--claim-key "$CLAIM_KEY" \
--kind "fix" \
--status "done" \
--title "Fire: " \
--context "fixed: | escalated: "
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: Lifecycle-Innovations-Limited
- Source: Lifecycle-Innovations-Limited/claude-ops
- License: MIT
- Homepage: https://github.com/Lifecycle-Innovations-Limited/claude-ops/wiki
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.