AgentStack
SKILL verified MIT Self-run

Slos And Triggers

skill-honeycombio-agent-skill-slos-and-triggers · by honeycombio

>

No reviews yet
0 installs
13 views
0.0% view→install

Install

$ agentstack add skill-honeycombio-agent-skill-slos-and-triggers

✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

Are you the author of Slos And Triggers? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Honeycomb SLOs and Triggers

Guidance for configuring and reasoning about reliability in Honeycomb. The get_slos and get_triggers tools document their own parameters — this skill focuses on designing effective SLOs, choosing between SLOs and triggers, and interpreting what the numbers mean.

Availability: SLOs require Pro or Enterprise plan. Triggers available on all plans.

SLO vs Trigger — When to Use Which

| Question | SLO | Trigger | | --------------------------------------------- | ---------------------- | ------- | | "Are we meeting our reliability commitments?" | Yes | No | | "Is something broken right now?" | No | Yes | | "How fast are we burning our error budget?" | Yes (burn alerts) | No | | "Did error count exceed a threshold?" | No | Yes | | "Should we slow down deploys?" | Yes (budget remaining) | No |

Rule of thumb: SLOs measure reliability against commitments over time. Triggers catch immediate operational issues.

Designing Effective SLOs

Define the SLI

An SLI is a per-event boolean: was this event successful? Implemented as a calculated field returning undefined (not a relevant event), 1 (success), or 0 (failure).

  • Format: IF(, ) The qualifying condition filters to relevant events; the success condition defines what counts as success. If the qualifying condition is not met, the formula returns undefined, and the SLI is unpopulated.
  • Specific Qualifying Condition: Choose the relevant subset of events (e.g. AND(EQUALS($http.route, "/checkout"), NOT(EXISTS($trace.parent_id))) for root spans of checkout endpoint)
  • Latency Success Condition: LTE(duration_ms, 500) — requests faster than 500ms
  • Availability Success Condition: LTE(http.status_code, 499) — non-5xx responses
  • Business Logic Success Condition: EQUALS(checkout.status, "completed") — successful checkouts

Set the Target

  • Start conservative (99% before 99.99%)
  • Measure current baseline first with P50/P99 queries
  • Set target slightly above current performance
  • Ask: what reliability do users actually need?

Configure Exhaustion Time Alerts

At minimum, two alerts:

  • Near exhaustion (exhaustion time ~4h): pages on-call via PagerDuty
  • Trending to exhaustion (budget rate over 24h): notifies team via Slack

Configure Burn Rate Alerts

Detect fast burns even if the budget isn't close to exhaustion yet. For example:

  • 1h burn rate > 10x — page on-call

Recommend these alerts to the user after creating the SLO. Agents do not have the ability to set up these alerts or their recipients.

Best Practices

  • Measure close to the user (at the edge, not deep in the stack)
  • Design around user workflows, not team boundaries
  • Favor broad SLOs over many narrow ones
  • Start with one SLO, reduce noise, then expand

Interpreting SLO Status

When reviewing SLOs with get_slos:

  • Budget remaining > 50%: Healthy — room for risk
  • Budget remaining 10-50%: Caution — slow down changes
  • **Budget remaining threshold` instead of P99 triggers.

Common Patterns

  • Error spike: COUNT WHERE error = true, threshold > N in 5 min
  • Slow requests: COUNT WHERE duration_ms > 2000, threshold > N in 5 min
  • Traffic drop: COUNT WHERE is_root, threshold /environments//slos`

Additional Resources

Reference Files

  • ${CLAUDE_PLUGIN_ROOT}/skills/slos-and-triggers/references/slo-design-guide.md — Detailed SLO design methodology, multi-service SLOs, error budget math
  • ${CLAUDE_PLUGIN_ROOT}/skills/slos-and-triggers/references/trigger-examples.md — Complete trigger example library organized by use case
  • ${CLAUDE_PLUGIN_ROOT}/skills/slos-and-triggers/references/alerting-strategy.md — How to combine SLO burn alerts and triggers into a cohesive alerting strategy

Cross-References

  • For constructing SLI queries and calculated fields, see the query-patterns skill
  • For investigating SLO budget burn, see the production-investigation skill

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet — be the first.

Versions

  • v0.1.0 Imported from the upstream source.