Install
$ agentstack add skill-xobotyi-cc-foundry-observability ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Observability
If you cannot ask arbitrary questions about your system's behavior from the outside, your system is not observable — it is merely monitored.
The Three Pillars
Logs — What Happened
Logs are timestamped, discrete event records. They capture what happened at a specific moment: an error thrown, a user action, a configuration loaded, a connection refused.
Use logs when you need:
- Rich diagnostic context for a specific event
- Debugging information with full error details and stack traces
- Audit trails of who did what and when
- Record of discrete state transitions
Logs are poor at:
- Showing aggregate system health (use metrics)
- Tracing request flow across services (use traces)
- High-frequency numeric trends (too expensive at volume)
Metrics — How Is It Doing
Metrics are numeric measurements aggregated over time. They capture how the system is performing as quantitative time series: request rates, error percentages, latencies, queue depths, resource utilization.
Use metrics when you need:
- Real-time health signals and alerting
- Trend analysis over hours, days, weeks
- Capacity planning and saturation monitoring
- Pre-aggregated data that scales cheaply regardless of traffic
Metrics are poor at:
- Explaining why something is broken (use logs)
- Showing the path of a single request (use traces)
- Storing per-event detail (cardinality explosion)
Traces — How Did It Flow
Traces record the causal chain of operations that make up a single request as it propagates through distributed components. A trace is a tree of spans, where each span represents one unit of work (an HTTP call, a database query, a queue publish).
Use traces when you need:
- End-to-end latency breakdown across services
- Dependency mapping and bottleneck identification
- Understanding the path a failing request took
- Correlating work across process and network boundaries
Traces are poor at:
- Aggregate health monitoring (use metrics)
- Detailed per-event diagnostics on a single node (use logs)
- Cheap, long-term trend storage (traces are expensive at 100% sampling)
Choosing the Right Signal
- "Is the system healthy right now?" — Metrics
- "Why did this specific request fail?" — Traces + Logs
- "What happened at 03:14 on node-7?" — Logs
- "Where is the bottleneck in checkout flow?" — Traces
- "Are error rates increasing over the last hour?" — Metrics
- "What was the full stack trace of that exception?" — Logs
- "Which downstream service is slow?" — Traces
- "How much headroom does the database have?" — Metrics
Structured Logging
Always Structured
Emit logs as structured records (JSON or equivalent key-value format) with a consistent schema. Unstructured string logs are for local development only. Structured logs are machine-parseable, indexable, and filterable at scale.
Log Levels
Use levels consistently. Agree on what each level means across the team.
- FATAL/CRITICAL — Process cannot continue; about to crash. Alerting: Page immediately
- ERROR — Operation failed; requires investigation. Alerting: Alert / ticket
- WARN — Unexpected condition; system compensated. Alerting: Monitor trend
- INFO — Significant business or lifecycle event. Alerting: Dashboard
- DEBUG — Diagnostic detail for developers. Alerting: Never in production by default
- TRACE — Extremely verbose step-by-step flow. Alerting: Never in production
Rules:
- Production defaults to INFO or above. DEBUG/TRACE are off unless explicitly enabled for a bounded investigation
window.
- WARN is not a dumping ground. If it never leads to action, it is noise — downgrade to DEBUG or remove it.
- ERROR means something is broken. Expected conditions (404 for missing resources, validation failures from bad input)
are not errors — log at INFO with a status field.
- Log level must be configurable at runtime without restarts.
Structured Fields
Every log record should include these baseline fields:
timestamp: ISO 8601, UTClevel: Severity (ERROR, WARN, INFO, ...)message: Human-readable summary of the eventservice: Service name emitting the logversion: Service version / build / commit SHAtrace_id: Distributed trace ID (if in request context)span_id: Current span ID (if in request context)
Add contextual fields relevant to the event:
user_id: User-initiated actionsrequest_id: Per-request correlationduration_ms: Timed operationserror.type: Error class/nameerror.message: Error descriptionerror.stack: Stack trace (ERROR level only)http.method,http.path,http.status: HTTP request/responsedb.operation,db.duration_ms: Database calls
Sensitive Data
Never log:
- Passwords, tokens, API keys, secrets
- Full credit card numbers, SSNs, or equivalent PII
- Session tokens or authentication cookies
- Request/response bodies containing user-submitted personal data
When user identifiers are needed, log opaque IDs (user_id), not email addresses or names. If regulations (GDPR, HIPAA) apply, verify logged fields comply. When in doubt, omit the field.
Logging at Boundaries
At application startup:
- INFO: service name, version, loaded configuration (without secrets), listen address
- WARN: degraded mode (e.g., fallback to local cache because Redis is unreachable)
- ERROR/FATAL: unrecoverable startup failures
Per incoming request:
- INFO: method, path (scrubbed of PII), status code, duration, request dimensions (tenant, region)
- WARN/ERROR: only for unexpected exceptions; catch at the top-level handler
Per outgoing dependency call:
- INFO or DEBUG: target service, operation, status, duration
- ERROR: failures in dependent services (Redis, database, queue, etc.)
Log Once, at the Right Level
Log a raised exception once. Do not catch-log-rethrow at every layer. Let exceptions propagate to the top-level handler, which logs with full context. Log and rethrow only when adding context that would otherwise be lost.
Metrics
Metric Types
- Counter — Monotonically increasing; resets on restart. Use for totals: requests, errors, bytes sent
- Gauge — Arbitrary value; goes up and down. Use for snapshots: queue depth, memory usage, connections
- Histogram — Client-side aggregation into buckets. Use for distributions: request latency, payload size
- Summary — Client-side quantile calculation. Use for pre-computed percentiles (less flexible than histogram)
Rules:
- Use counters for events that accumulate. Derive rates with
rate()/increase()— never store pre-computed rates. - Use gauges for current-state snapshots. Never
rate()a gauge. - Use histograms for latency and size distributions. Histograms enable percentile calculation across instances;
summaries do not aggregate.
- Export timestamps as Unix epoch seconds, not "time since" values.
- Initialize all metrics with zero at startup to avoid missing-metric problems.
What to Measure
The Four Golden Signals (Google SRE)
For every user-facing service, measure these four:
- Latency — Time to serve a request. Example:
http_request_duration_secondshistogram - Traffic — Demand on the system. Example:
http_requests_totalcounter by method/path - Errors — Rate of failed requests. Example:
http_requests_total{status=~"5.."} - Saturation — How "full" the service is. Example: CPU usage, memory, queue depth, thread pool
Distinguish successful latency from error latency. A fast 500 is not good latency. Track both.
RED Method (Request-Centric)
For every microservice:
- Rate — requests per second
- Errors — failed requests per second
- Duration — distribution of request latency
RED is a focused subset of the golden signals, optimized for request-driven services.
USE Method (Resource-Centric)
For every resource (CPU, memory, disk, network, thread pool):
- Utilization — percentage of capacity in use
- Saturation — backlog / queue depth
- Errors — resource-level error count
RED tells you what is degraded from the user's perspective. USE tells you why at the infrastructure level. Use both together.
Service-Type Instrumentation
- Online-serving (HTTP, gRPC) — Request rate, error rate, latency (p50/p90/p99), in-flight requests
- Offline-processing (workers, pipelines) — Items in/out per stage, processing duration, last-processed timestamp,
queue depth
- Batch jobs — Last successful completion time, job duration, records processed, exit status
- Caches — Hit rate, miss rate, eviction count, latency to backend on miss
- Thread/connection pools — Pool size, active count, queue length, wait time
Metric Naming
Metric names should be self-documenting. Follow these conventions:
- Prefix with namespace.
myapp_http_requests_total, notrequests_total. - Use base units. Seconds (not milliseconds), bytes (not megabytes), ratio 0-1 (not percentage 0-100).
- Suffix with unit.
_seconds,_bytes,_total(for unit-less counters). - One metric, one unit, one quantity. Never mix request size with request duration in the same metric.
- snake_case.
http_request_duration_seconds, nothttpRequestDurationSeconds.
| Good | Bad | | ------------------------------- | --------------------------------------- | | http_request_duration_seconds | request_latency (no unit, ambiguous) | | http_requests_total | http_responses_500_total (use labels) | | node_memory_usage_bytes | memory_mb (not base unit) | | process_cpu_seconds_total | cpu_percent (use ratio 0-1) |
Labels and Cardinality
Labels add dimensions to a metric. Every unique combination of label values creates a separate time series.
Good labels (bounded, low cardinality):
method(GET, POST, PUT, DELETE)status_code(200, 404, 500 — or class: 2xx, 4xx, 5xx)service,region,version
Dangerous labels (unbounded, high cardinality):
user_id,email,session_idrequest_pathwith dynamic segments (/users/12345)error_message(arbitrary strings)
Rules:
- Keep label cardinality below 10 values per label for most metrics.
- If a label can grow unbounded, it does not belong on a metric. Log it instead.
- Use labels instead of encoding dimensions in the metric name.
http_requests_total{method="GET"}, not
http_get_requests_total.
- Ensure
sum()oravg()across all label values is meaningful. If not, split into separate metrics.
Percentiles and Tail Latency
Averages hide outliers. A service with 100ms average latency may have 1% of requests taking 5 seconds. That 1% tail can dominate user experience when users hit multiple services per page load.
- Always track p50, p90, p99 latency at minimum.
- Use histograms with exponentially distributed bucket boundaries (e.g., 5ms, 10ms, 25ms, 50ms, 100ms, 250ms, 500ms, 1s,
2.5s, 5s, 10s).
- Alert on p99, not mean. Mean latency alerts miss tail degradation.
Distributed Tracing
Core Concepts
- Trace — End-to-end record of a single request across all services
- Span — One unit of work within a trace (HTTP call, DB query, function)
- Root span — First span in a trace; has no parent
- Child span — Span nested under a parent; represents a sub-operation
- Span context — Immutable bag of
trace_id+span_id+ flags, propagated across boundaries - Span attributes — Key-value metadata on a span (http.method, db.statement)
- Span events — Timestamped annotations within a span's lifetime
- Span links — Causal references between spans in different traces
Span Kinds
- Client — Outgoing synchronous call. Example: HTTP request to another service
- Server — Incoming synchronous call. Example: Handling an HTTP request
- Producer — Creates async work. Example: Publishing to a message queue
- Consumer — Processes async work. Example: Consuming from a message queue
- Internal — No network boundary. Example: In-process function instrumentation
Context Propagation
Context propagation connects spans across process boundaries into a single trace. Without it, you get disconnected spans, not traces.
Rules:
- Propagate context on every outgoing call. HTTP headers (W3C Trace Context or B3), message metadata, gRPC metadata
— every cross-process boundary must carry trace context.
- Extract context on every incoming call. Extract trace context and create a child span under the propagated parent.
- Use W3C Trace Context (
traceparent/tracestate) as the default propagation format unless the ecosystem requires
otherwise (e.g., legacy B3).
- Never generate a new trace ID when continuing an existing trace. A new trace ID means a broken trace.
What to Trace
Instrument at meaningful boundaries:
- Incoming HTTP/gRPC requests — Always — auto-instrument
- Outgoing HTTP/gRPC calls — Always — auto-instrument
- Database queries — Always — auto-instrument or manual
- Cache operations — Yes — hit/miss as attribute
- Queue publish/consume — Yes — link producer and consumer spans
- Significant business operations — Yes — manual spans for key logic
- Tight loops / trivial functions — No — noise, performance cost
Span Attributes
Attach attributes that enable filtering and analysis:
http.method,http.route,http.status_code: HTTP spansdb.system,db.operation,db.statement: Database spansmessaging.system,messaging.operation: Queue spansrpc.system,rpc.method: RPC spanserror(boolean),error.type,error.message: Error conditionsservice.name,service.version: All spans (set on resource)
Use semantic conventions for attribute names rather than inventing custom ones. Consistent naming enables cross-service analysis.
Span Status
Unset— Completed without error (default). When: most successful operationsError— Operation failed. When: server errors, exceptionsOk— Explicitly marked successful. When: only when you need to override ambiguity
Leave status as Unset for normal success. Set Error only for actual failures. Do not set Error for client errors like 404 on a server span — the server operated correctly.
Sampling
At high traffic volumes, tracing 100% of requests is expensive. Sampling reduces cost while preserving signal. | Strategy | How It Works | Trade-off | | ------------------------ | ------------------------------------------------------ | ---------------------------------------------- | | Head-based | Decide at trace start whether to sample | Simple; may miss rare errors | | Tail-based | Decide after trace completes based on content | Catches errors; needs buffering infrastructure | | Always-on for errors | Sample 100% of error traces, probabilistic for success | Good default balance |
Rules:
- Never drop error traces. If cost is a concern, sample successful traces at a lower rate but keep 100% of error and
high-latency traces.
- Sample at the entry point (head) and propagate the decision. Each service deciding independently creates partial
traces.
- Start with a low sampling rate (1-10%) and increase based on need, not the reverse.
Connecting the Pillars
The three pillars become powerful when correlated. An alert fires on a metric → you find the offending trace → the trace points to a span → the span's logs reveal the root cause.
Correlation Keys
trace_id— Links logs and spans to the same trace. Where: logs, span contextspan_id— Links a log to the exact span that produced it. Where: logs, span contextrequest_id— Correlates all work for one inbound request. Whe
…
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: xobotyi
- Source: xobotyi/cc-foundry
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.