# Observability

> >-

- **Type:** Skill
- **Install:** `agentstack add skill-xobotyi-cc-foundry-observability`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [xobotyi](https://agentstack.voostack.com/s/xobotyi)
- **Installs:** 0
- **Category:** [Agent Skills](https://agentstack.voostack.com/c/agent-skills)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [xobotyi](https://github.com/xobotyi)
- **Source:** https://github.com/xobotyi/cc-foundry/tree/master/plugins/backend/skills/observability

## Install

```sh
agentstack add skill-xobotyi-cc-foundry-observability
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# Observability

**If you cannot ask arbitrary questions about your system's behavior from the outside, your system is not observable —
it is merely monitored.**

---

## The Three Pillars

### Logs — What Happened

Logs are timestamped, discrete event records. They capture **what happened** at a specific moment: an error thrown, a
user action, a configuration loaded, a connection refused.

**Use logs when you need:**

- Rich diagnostic context for a specific event
- Debugging information with full error details and stack traces
- Audit trails of who did what and when
- Record of discrete state transitions

**Logs are poor at:**

- Showing aggregate system health (use metrics)
- Tracing request flow across services (use traces)
- High-frequency numeric trends (too expensive at volume)

### Metrics — How Is It Doing

Metrics are numeric measurements aggregated over time. They capture **how the system is performing** as quantitative
time series: request rates, error percentages, latencies, queue depths, resource utilization.

**Use metrics when you need:**

- Real-time health signals and alerting
- Trend analysis over hours, days, weeks
- Capacity planning and saturation monitoring
- Pre-aggregated data that scales cheaply regardless of traffic

**Metrics are poor at:**

- Explaining _why_ something is broken (use logs)
- Showing the path of a single request (use traces)
- Storing per-event detail (cardinality explosion)

### Traces — How Did It Flow

Traces record the causal chain of operations that make up a single request as it propagates through distributed
components. A trace is a tree of **spans**, where each span represents one unit of work (an HTTP call, a database query,
a queue publish).

**Use traces when you need:**

- End-to-end latency breakdown across services
- Dependency mapping and bottleneck identification
- Understanding the path a failing request took
- Correlating work across process and network boundaries

**Traces are poor at:**

- Aggregate health monitoring (use metrics)
- Detailed per-event diagnostics on a single node (use logs)
- Cheap, long-term trend storage (traces are expensive at 100% sampling)

### Choosing the Right Signal

- **"Is the system healthy right now?"** — Metrics
- **"Why did this specific request fail?"** — Traces + Logs
- **"What happened at 03:14 on node-7?"** — Logs
- **"Where is the bottleneck in checkout flow?"** — Traces
- **"Are error rates increasing over the last hour?"** — Metrics
- **"What was the full stack trace of that exception?"** — Logs
- **"Which downstream service is slow?"** — Traces
- **"How much headroom does the database have?"** — Metrics

---

## Structured Logging

### Always Structured

Emit logs as structured records (JSON or equivalent key-value format) with a consistent schema. Unstructured string logs
are for local development only. Structured logs are machine-parseable, indexable, and filterable at scale.

### Log Levels

Use levels consistently. Agree on what each level means across the team.

- **FATAL/CRITICAL** — Process cannot continue; about to crash. Alerting: Page immediately
- **ERROR** — Operation failed; requires investigation. Alerting: Alert / ticket
- **WARN** — Unexpected condition; system compensated. Alerting: Monitor trend
- **INFO** — Significant business or lifecycle event. Alerting: Dashboard
- **DEBUG** — Diagnostic detail for developers. Alerting: Never in production by default
- **TRACE** — Extremely verbose step-by-step flow. Alerting: Never in production

Rules:

- Production defaults to INFO or above. DEBUG/TRACE are off unless explicitly enabled for a bounded investigation
  window.
- WARN is not a dumping ground. If it never leads to action, it is noise — downgrade to DEBUG or remove it.
- ERROR means something is broken. Expected conditions (404 for missing resources, validation failures from bad input)
  are not errors — log at INFO with a status field.
- Log level must be configurable at runtime without restarts.

### Structured Fields

Every log record should include these baseline fields:

- `timestamp`: ISO 8601, UTC
- `level`: Severity (ERROR, WARN, INFO, ...)
- `message`: Human-readable summary of the event
- `service`: Service name emitting the log
- `version`: Service version / build / commit SHA
- `trace_id`: Distributed trace ID (if in request context)
- `span_id`: Current span ID (if in request context)

Add contextual fields relevant to the event:

- `user_id`: User-initiated actions
- `request_id`: Per-request correlation
- `duration_ms`: Timed operations
- `error.type`: Error class/name
- `error.message`: Error description
- `error.stack`: Stack trace (ERROR level only)
- `http.method`, `http.path`, `http.status`: HTTP request/response
- `db.operation`, `db.duration_ms`: Database calls

### Sensitive Data

Never log:

- Passwords, tokens, API keys, secrets
- Full credit card numbers, SSNs, or equivalent PII
- Session tokens or authentication cookies
- Request/response bodies containing user-submitted personal data

When user identifiers are needed, log opaque IDs (user_id), not email addresses or names. If regulations (GDPR, HIPAA)
apply, verify logged fields comply. When in doubt, omit the field.

### Logging at Boundaries

**At application startup:**

- INFO: service name, version, loaded configuration (without secrets), listen address
- WARN: degraded mode (e.g., fallback to local cache because Redis is unreachable)
- ERROR/FATAL: unrecoverable startup failures

**Per incoming request:**

- INFO: method, path (scrubbed of PII), status code, duration, request dimensions (tenant, region)
- WARN/ERROR: only for unexpected exceptions; catch at the top-level handler

**Per outgoing dependency call:**

- INFO or DEBUG: target service, operation, status, duration
- ERROR: failures in dependent services (Redis, database, queue, etc.)

### Log Once, at the Right Level

Log a raised exception **once**. Do not catch-log-rethrow at every layer. Let exceptions propagate to the top-level
handler, which logs with full context. Log and rethrow only when adding context that would otherwise be lost.

---

## Metrics

### Metric Types

- **Counter** — Monotonically increasing; resets on restart. Use for totals: requests, errors, bytes sent
- **Gauge** — Arbitrary value; goes up and down. Use for snapshots: queue depth, memory usage, connections
- **Histogram** — Client-side aggregation into buckets. Use for distributions: request latency, payload size
- **Summary** — Client-side quantile calculation. Use for pre-computed percentiles (less flexible than histogram)

Rules:

- Use counters for events that accumulate. Derive rates with `rate()` / `increase()` — never store pre-computed rates.
- Use gauges for current-state snapshots. Never `rate()` a gauge.
- Use histograms for latency and size distributions. Histograms enable percentile calculation across instances;
  summaries do not aggregate.
- Export timestamps as Unix epoch seconds, not "time since" values.
- Initialize all metrics with zero at startup to avoid missing-metric problems.

### What to Measure

#### The Four Golden Signals (Google SRE)

For every user-facing service, measure these four:

- **Latency** — Time to serve a request. Example: `http_request_duration_seconds` histogram
- **Traffic** — Demand on the system. Example: `http_requests_total` counter by method/path
- **Errors** — Rate of failed requests. Example: `http_requests_total{status=~"5.."}`
- **Saturation** — How "full" the service is. Example: CPU usage, memory, queue depth, thread pool

Distinguish **successful latency from error latency**. A fast 500 is not good latency. Track both.

#### RED Method (Request-Centric)

For every microservice:

- **R**ate — requests per second
- **E**rrors — failed requests per second
- **D**uration — distribution of request latency

RED is a focused subset of the golden signals, optimized for request-driven services.

#### USE Method (Resource-Centric)

For every resource (CPU, memory, disk, network, thread pool):

- **U**tilization — percentage of capacity in use
- **S**aturation — backlog / queue depth
- **E**rrors — resource-level error count

RED tells you _what_ is degraded from the user's perspective. USE tells you _why_ at the infrastructure level. Use both
together.

#### Service-Type Instrumentation

- **Online-serving** (HTTP, gRPC) — Request rate, error rate, latency (p50/p90/p99), in-flight requests
- **Offline-processing** (workers, pipelines) — Items in/out per stage, processing duration, last-processed timestamp,
  queue depth
- **Batch jobs** — Last successful completion time, job duration, records processed, exit status
- **Caches** — Hit rate, miss rate, eviction count, latency to backend on miss
- **Thread/connection pools** — Pool size, active count, queue length, wait time

### Metric Naming

Metric names should be self-documenting. Follow these conventions:

- **Prefix with namespace.** `myapp_http_requests_total`, not `requests_total`.
- **Use base units.** Seconds (not milliseconds), bytes (not megabytes), ratio 0-1 (not percentage 0-100).
- **Suffix with unit.** `_seconds`, `_bytes`, `_total` (for unit-less counters).
- **One metric, one unit, one quantity.** Never mix request size with request duration in the same metric.
- **snake_case.** `http_request_duration_seconds`, not `httpRequestDurationSeconds`.

| Good                            | Bad                                     |
| ------------------------------- | --------------------------------------- |
| `http_request_duration_seconds` | `request_latency` (no unit, ambiguous)  |
| `http_requests_total`           | `http_responses_500_total` (use labels) |
| `node_memory_usage_bytes`       | `memory_mb` (not base unit)             |
| `process_cpu_seconds_total`     | `cpu_percent` (use ratio 0-1)           |

### Labels and Cardinality

Labels add dimensions to a metric. Every unique combination of label values creates a separate time series.

**Good labels** (bounded, low cardinality):

- `method` (GET, POST, PUT, DELETE)
- `status_code` (200, 404, 500 — or class: 2xx, 4xx, 5xx)
- `service`, `region`, `version`

**Dangerous labels** (unbounded, high cardinality):

- `user_id`, `email`, `session_id`
- `request_path` with dynamic segments (`/users/12345`)
- `error_message` (arbitrary strings)

Rules:

- Keep label cardinality below 10 values per label for most metrics.
- If a label can grow unbounded, it does not belong on a metric. Log it instead.
- Use labels instead of encoding dimensions in the metric name. `http_requests_total{method="GET"}`, not
  `http_get_requests_total`.
- Ensure `sum()` or `avg()` across all label values is meaningful. If not, split into separate metrics.

### Percentiles and Tail Latency

Averages hide outliers. A service with 100ms average latency may have 1% of requests taking 5 seconds. That 1% tail can
dominate user experience when users hit multiple services per page load.

- Always track **p50, p90, p99** latency at minimum.
- Use histograms with exponentially distributed bucket boundaries (e.g., 5ms, 10ms, 25ms, 50ms, 100ms, 250ms, 500ms, 1s,
  2.5s, 5s, 10s).
- Alert on p99, not mean. Mean latency alerts miss tail degradation.

---

## Distributed Tracing

### Core Concepts

- **Trace** — End-to-end record of a single request across all services
- **Span** — One unit of work within a trace (HTTP call, DB query, function)
- **Root span** — First span in a trace; has no parent
- **Child span** — Span nested under a parent; represents a sub-operation
- **Span context** — Immutable bag of `trace_id` + `span_id` + flags, propagated across boundaries
- **Span attributes** — Key-value metadata on a span (http.method, db.statement)
- **Span events** — Timestamped annotations within a span's lifetime
- **Span links** — Causal references between spans in different traces

### Span Kinds

- **Client** — Outgoing synchronous call. Example: HTTP request to another service
- **Server** — Incoming synchronous call. Example: Handling an HTTP request
- **Producer** — Creates async work. Example: Publishing to a message queue
- **Consumer** — Processes async work. Example: Consuming from a message queue
- **Internal** — No network boundary. Example: In-process function instrumentation

### Context Propagation

Context propagation connects spans across process boundaries into a single trace. Without it, you get disconnected
spans, not traces.

Rules:

- **Propagate context on every outgoing call.** HTTP headers (W3C Trace Context or B3), message metadata, gRPC metadata
  — every cross-process boundary must carry trace context.
- **Extract context on every incoming call.** Extract trace context and create a child span under the propagated parent.
- **Use W3C Trace Context (`traceparent`/`tracestate`)** as the default propagation format unless the ecosystem requires
  otherwise (e.g., legacy B3).
- **Never generate a new trace ID** when continuing an existing trace. A new trace ID means a broken trace.

### What to Trace

Instrument at meaningful boundaries:

- **Incoming HTTP/gRPC requests** — Always — auto-instrument
- **Outgoing HTTP/gRPC calls** — Always — auto-instrument
- **Database queries** — Always — auto-instrument or manual
- **Cache operations** — Yes — hit/miss as attribute
- **Queue publish/consume** — Yes — link producer and consumer spans
- **Significant business operations** — Yes — manual spans for key logic
- **Tight loops / trivial functions** — No — noise, performance cost

### Span Attributes

Attach attributes that enable filtering and analysis:

- `http.method`, `http.route`, `http.status_code`: HTTP spans
- `db.system`, `db.operation`, `db.statement`: Database spans
- `messaging.system`, `messaging.operation`: Queue spans
- `rpc.system`, `rpc.method`: RPC spans
- `error` (boolean), `error.type`, `error.message`: Error conditions
- `service.name`, `service.version`: All spans (set on resource)

Use [semantic conventions](https://opentelemetry.io/docs/specs/semconv/) for attribute names rather than inventing
custom ones. Consistent naming enables cross-service analysis.

### Span Status

- `Unset` — Completed without error (default). When: most successful operations
- `Error` — Operation failed. When: server errors, exceptions
- `Ok` — Explicitly marked successful. When: only when you need to override ambiguity

Leave status as `Unset` for normal success. Set `Error` only for actual failures. Do not set `Error` for client errors
like 404 on a server span — the server operated correctly.

### Sampling

At high traffic volumes, tracing 100% of requests is expensive. Sampling reduces cost while preserving signal. |
Strategy | How It Works | Trade-off | | ------------------------ |
------------------------------------------------------ | ---------------------------------------------- | |
**Head-based** | Decide at trace start whether to sample | Simple; may miss rare errors | | **Tail-based** | Decide
after trace completes based on content | Catches errors; needs buffering infrastructure | | **Always-on for errors** |
Sample 100% of error traces, probabilistic for success | Good default balance |

Rules:

- Never drop error traces. If cost is a concern, sample successful traces at a lower rate but keep 100% of error and
  high-latency traces.
- Sample at the entry point (head) and propagate the decision. Each service deciding independently creates partial
  traces.
- Start with a low sampling rate (1-10%) and increase based on need, not the reverse.

---

## Connecting the Pillars

The three pillars become powerful when correlated. An alert fires on a metric → you find the offending trace → the trace
points to a span → the span's logs reveal the root cause.

### Correlation Keys

- `trace_id` — Links logs and spans to the same trace. Where: logs, span context
- `span_id` — Links a log to the exact span that produced it. Where: logs, span context
- `request_id` — Correlates all work for one inbound request. Whe

…

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [xobotyi](https://github.com/xobotyi)
- **Source:** [xobotyi/cc-foundry](https://github.com/xobotyi/cc-foundry)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-xobotyi-cc-foundry-observability
- Seller: https://agentstack.voostack.com/s/xobotyi
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
