AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Observability

skill-xobotyi-cc-foundry-observability · by xobotyi

>-

No reviews yet
0 installs
50 views
0.0% view→install

Install

$ agentstack add skill-xobotyi-cc-foundry-observability

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-xobotyi-cc-foundry-observability)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
2mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Observability? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Observability

If you cannot ask arbitrary questions about your system's behavior from the outside, your system is not observable — it is merely monitored.


The Three Pillars

Logs — What Happened

Logs are timestamped, discrete event records. They capture what happened at a specific moment: an error thrown, a user action, a configuration loaded, a connection refused.

Use logs when you need:

  • Rich diagnostic context for a specific event
  • Debugging information with full error details and stack traces
  • Audit trails of who did what and when
  • Record of discrete state transitions

Logs are poor at:

  • Showing aggregate system health (use metrics)
  • Tracing request flow across services (use traces)
  • High-frequency numeric trends (too expensive at volume)

Metrics — How Is It Doing

Metrics are numeric measurements aggregated over time. They capture how the system is performing as quantitative time series: request rates, error percentages, latencies, queue depths, resource utilization.

Use metrics when you need:

  • Real-time health signals and alerting
  • Trend analysis over hours, days, weeks
  • Capacity planning and saturation monitoring
  • Pre-aggregated data that scales cheaply regardless of traffic

Metrics are poor at:

  • Explaining why something is broken (use logs)
  • Showing the path of a single request (use traces)
  • Storing per-event detail (cardinality explosion)

Traces — How Did It Flow

Traces record the causal chain of operations that make up a single request as it propagates through distributed components. A trace is a tree of spans, where each span represents one unit of work (an HTTP call, a database query, a queue publish).

Use traces when you need:

  • End-to-end latency breakdown across services
  • Dependency mapping and bottleneck identification
  • Understanding the path a failing request took
  • Correlating work across process and network boundaries

Traces are poor at:

  • Aggregate health monitoring (use metrics)
  • Detailed per-event diagnostics on a single node (use logs)
  • Cheap, long-term trend storage (traces are expensive at 100% sampling)

Choosing the Right Signal

  • "Is the system healthy right now?" — Metrics
  • "Why did this specific request fail?" — Traces + Logs
  • "What happened at 03:14 on node-7?" — Logs
  • "Where is the bottleneck in checkout flow?" — Traces
  • "Are error rates increasing over the last hour?" — Metrics
  • "What was the full stack trace of that exception?" — Logs
  • "Which downstream service is slow?" — Traces
  • "How much headroom does the database have?" — Metrics

Structured Logging

Always Structured

Emit logs as structured records (JSON or equivalent key-value format) with a consistent schema. Unstructured string logs are for local development only. Structured logs are machine-parseable, indexable, and filterable at scale.

Log Levels

Use levels consistently. Agree on what each level means across the team.

  • FATAL/CRITICAL — Process cannot continue; about to crash. Alerting: Page immediately
  • ERROR — Operation failed; requires investigation. Alerting: Alert / ticket
  • WARN — Unexpected condition; system compensated. Alerting: Monitor trend
  • INFO — Significant business or lifecycle event. Alerting: Dashboard
  • DEBUG — Diagnostic detail for developers. Alerting: Never in production by default
  • TRACE — Extremely verbose step-by-step flow. Alerting: Never in production

Rules:

  • Production defaults to INFO or above. DEBUG/TRACE are off unless explicitly enabled for a bounded investigation

window.

  • WARN is not a dumping ground. If it never leads to action, it is noise — downgrade to DEBUG or remove it.
  • ERROR means something is broken. Expected conditions (404 for missing resources, validation failures from bad input)

are not errors — log at INFO with a status field.

  • Log level must be configurable at runtime without restarts.

Structured Fields

Every log record should include these baseline fields:

  • timestamp: ISO 8601, UTC
  • level: Severity (ERROR, WARN, INFO, ...)
  • message: Human-readable summary of the event
  • service: Service name emitting the log
  • version: Service version / build / commit SHA
  • trace_id: Distributed trace ID (if in request context)
  • span_id: Current span ID (if in request context)

Add contextual fields relevant to the event:

  • user_id: User-initiated actions
  • request_id: Per-request correlation
  • duration_ms: Timed operations
  • error.type: Error class/name
  • error.message: Error description
  • error.stack: Stack trace (ERROR level only)
  • http.method, http.path, http.status: HTTP request/response
  • db.operation, db.duration_ms: Database calls

Sensitive Data

Never log:

  • Passwords, tokens, API keys, secrets
  • Full credit card numbers, SSNs, or equivalent PII
  • Session tokens or authentication cookies
  • Request/response bodies containing user-submitted personal data

When user identifiers are needed, log opaque IDs (user_id), not email addresses or names. If regulations (GDPR, HIPAA) apply, verify logged fields comply. When in doubt, omit the field.

Logging at Boundaries

At application startup:

  • INFO: service name, version, loaded configuration (without secrets), listen address
  • WARN: degraded mode (e.g., fallback to local cache because Redis is unreachable)
  • ERROR/FATAL: unrecoverable startup failures

Per incoming request:

  • INFO: method, path (scrubbed of PII), status code, duration, request dimensions (tenant, region)
  • WARN/ERROR: only for unexpected exceptions; catch at the top-level handler

Per outgoing dependency call:

  • INFO or DEBUG: target service, operation, status, duration
  • ERROR: failures in dependent services (Redis, database, queue, etc.)

Log Once, at the Right Level

Log a raised exception once. Do not catch-log-rethrow at every layer. Let exceptions propagate to the top-level handler, which logs with full context. Log and rethrow only when adding context that would otherwise be lost.


Metrics

Metric Types

  • Counter — Monotonically increasing; resets on restart. Use for totals: requests, errors, bytes sent
  • Gauge — Arbitrary value; goes up and down. Use for snapshots: queue depth, memory usage, connections
  • Histogram — Client-side aggregation into buckets. Use for distributions: request latency, payload size
  • Summary — Client-side quantile calculation. Use for pre-computed percentiles (less flexible than histogram)

Rules:

  • Use counters for events that accumulate. Derive rates with rate() / increase() — never store pre-computed rates.
  • Use gauges for current-state snapshots. Never rate() a gauge.
  • Use histograms for latency and size distributions. Histograms enable percentile calculation across instances;

summaries do not aggregate.

  • Export timestamps as Unix epoch seconds, not "time since" values.
  • Initialize all metrics with zero at startup to avoid missing-metric problems.

What to Measure

The Four Golden Signals (Google SRE)

For every user-facing service, measure these four:

  • Latency — Time to serve a request. Example: http_request_duration_seconds histogram
  • Traffic — Demand on the system. Example: http_requests_total counter by method/path
  • Errors — Rate of failed requests. Example: http_requests_total{status=~"5.."}
  • Saturation — How "full" the service is. Example: CPU usage, memory, queue depth, thread pool

Distinguish successful latency from error latency. A fast 500 is not good latency. Track both.

RED Method (Request-Centric)

For every microservice:

  • Rate — requests per second
  • Errors — failed requests per second
  • Duration — distribution of request latency

RED is a focused subset of the golden signals, optimized for request-driven services.

USE Method (Resource-Centric)

For every resource (CPU, memory, disk, network, thread pool):

  • Utilization — percentage of capacity in use
  • Saturation — backlog / queue depth
  • Errors — resource-level error count

RED tells you what is degraded from the user's perspective. USE tells you why at the infrastructure level. Use both together.

Service-Type Instrumentation
  • Online-serving (HTTP, gRPC) — Request rate, error rate, latency (p50/p90/p99), in-flight requests
  • Offline-processing (workers, pipelines) — Items in/out per stage, processing duration, last-processed timestamp,

queue depth

  • Batch jobs — Last successful completion time, job duration, records processed, exit status
  • Caches — Hit rate, miss rate, eviction count, latency to backend on miss
  • Thread/connection pools — Pool size, active count, queue length, wait time

Metric Naming

Metric names should be self-documenting. Follow these conventions:

  • Prefix with namespace. myapp_http_requests_total, not requests_total.
  • Use base units. Seconds (not milliseconds), bytes (not megabytes), ratio 0-1 (not percentage 0-100).
  • Suffix with unit. _seconds, _bytes, _total (for unit-less counters).
  • One metric, one unit, one quantity. Never mix request size with request duration in the same metric.
  • snake_case. http_request_duration_seconds, not httpRequestDurationSeconds.

| Good | Bad | | ------------------------------- | --------------------------------------- | | http_request_duration_seconds | request_latency (no unit, ambiguous) | | http_requests_total | http_responses_500_total (use labels) | | node_memory_usage_bytes | memory_mb (not base unit) | | process_cpu_seconds_total | cpu_percent (use ratio 0-1) |

Labels and Cardinality

Labels add dimensions to a metric. Every unique combination of label values creates a separate time series.

Good labels (bounded, low cardinality):

  • method (GET, POST, PUT, DELETE)
  • status_code (200, 404, 500 — or class: 2xx, 4xx, 5xx)
  • service, region, version

Dangerous labels (unbounded, high cardinality):

  • user_id, email, session_id
  • request_path with dynamic segments (/users/12345)
  • error_message (arbitrary strings)

Rules:

  • Keep label cardinality below 10 values per label for most metrics.
  • If a label can grow unbounded, it does not belong on a metric. Log it instead.
  • Use labels instead of encoding dimensions in the metric name. http_requests_total{method="GET"}, not

http_get_requests_total.

  • Ensure sum() or avg() across all label values is meaningful. If not, split into separate metrics.

Percentiles and Tail Latency

Averages hide outliers. A service with 100ms average latency may have 1% of requests taking 5 seconds. That 1% tail can dominate user experience when users hit multiple services per page load.

  • Always track p50, p90, p99 latency at minimum.
  • Use histograms with exponentially distributed bucket boundaries (e.g., 5ms, 10ms, 25ms, 50ms, 100ms, 250ms, 500ms, 1s,

2.5s, 5s, 10s).

  • Alert on p99, not mean. Mean latency alerts miss tail degradation.

Distributed Tracing

Core Concepts

  • Trace — End-to-end record of a single request across all services
  • Span — One unit of work within a trace (HTTP call, DB query, function)
  • Root span — First span in a trace; has no parent
  • Child span — Span nested under a parent; represents a sub-operation
  • Span context — Immutable bag of trace_id + span_id + flags, propagated across boundaries
  • Span attributes — Key-value metadata on a span (http.method, db.statement)
  • Span events — Timestamped annotations within a span's lifetime
  • Span links — Causal references between spans in different traces

Span Kinds

  • Client — Outgoing synchronous call. Example: HTTP request to another service
  • Server — Incoming synchronous call. Example: Handling an HTTP request
  • Producer — Creates async work. Example: Publishing to a message queue
  • Consumer — Processes async work. Example: Consuming from a message queue
  • Internal — No network boundary. Example: In-process function instrumentation

Context Propagation

Context propagation connects spans across process boundaries into a single trace. Without it, you get disconnected spans, not traces.

Rules:

  • Propagate context on every outgoing call. HTTP headers (W3C Trace Context or B3), message metadata, gRPC metadata

— every cross-process boundary must carry trace context.

  • Extract context on every incoming call. Extract trace context and create a child span under the propagated parent.
  • Use W3C Trace Context (traceparent/tracestate) as the default propagation format unless the ecosystem requires

otherwise (e.g., legacy B3).

  • Never generate a new trace ID when continuing an existing trace. A new trace ID means a broken trace.

What to Trace

Instrument at meaningful boundaries:

  • Incoming HTTP/gRPC requests — Always — auto-instrument
  • Outgoing HTTP/gRPC calls — Always — auto-instrument
  • Database queries — Always — auto-instrument or manual
  • Cache operations — Yes — hit/miss as attribute
  • Queue publish/consume — Yes — link producer and consumer spans
  • Significant business operations — Yes — manual spans for key logic
  • Tight loops / trivial functions — No — noise, performance cost

Span Attributes

Attach attributes that enable filtering and analysis:

  • http.method, http.route, http.status_code: HTTP spans
  • db.system, db.operation, db.statement: Database spans
  • messaging.system, messaging.operation: Queue spans
  • rpc.system, rpc.method: RPC spans
  • error (boolean), error.type, error.message: Error conditions
  • service.name, service.version: All spans (set on resource)

Use semantic conventions for attribute names rather than inventing custom ones. Consistent naming enables cross-service analysis.

Span Status

  • Unset — Completed without error (default). When: most successful operations
  • Error — Operation failed. When: server errors, exceptions
  • Ok — Explicitly marked successful. When: only when you need to override ambiguity

Leave status as Unset for normal success. Set Error only for actual failures. Do not set Error for client errors like 404 on a server span — the server operated correctly.

Sampling

At high traffic volumes, tracing 100% of requests is expensive. Sampling reduces cost while preserving signal. | Strategy | How It Works | Trade-off | | ------------------------ | ------------------------------------------------------ | ---------------------------------------------- | | Head-based | Decide at trace start whether to sample | Simple; may miss rare errors | | Tail-based | Decide after trace completes based on content | Catches errors; needs buffering infrastructure | | Always-on for errors | Sample 100% of error traces, probabilistic for success | Good default balance |

Rules:

  • Never drop error traces. If cost is a concern, sample successful traces at a lower rate but keep 100% of error and

high-latency traces.

  • Sample at the entry point (head) and propagate the decision. Each service deciding independently creates partial

traces.

  • Start with a low sampling rate (1-10%) and increase based on need, not the reverse.

Connecting the Pillars

The three pillars become powerful when correlated. An alert fires on a metric → you find the offending trace → the trace points to a span → the span's logs reveal the root cause.

Correlation Keys

  • trace_id — Links logs and spans to the same trace. Where: logs, span context
  • span_id — Links a log to the exact span that produced it. Where: logs, span context
  • request_id — Correlates all work for one inbound request. Whe

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.