AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Building Dashboards

skill-pinchbench-skill-building-dashboards · by pinchbench

Designs and builds Axiom dashboards via API. Covers chart types, APL and metrics/MPL query patterns, SmartFilters, layout, and configuration options. Use when creating dashboards, migrating from Splunk, or configuring chart options.

— No reviews yet
0 installs
29 views
0.0% view→install

Install

$ agentstack add skill-pinchbench-skill-building-dashboards

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • ✓ Prompt-injection patterns
  • ✓ Secret / credential exfiltration
  • ✓ Dangerous shell & filesystem operations
  • ✓ Untrusted network calls
  • ✓ Known-malicious package signatures

What it can access

  • ● Network access Used
  • ✓ Filesystem access No
  • ✓ Shell / process execution No
  • ✓ Environment & secrets No
  • ✓ Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-pinchbench-skill-building-dashboards)

Reliability & compatibility

✓ Security review passed
0 installs to date
— no reviews yet
● 3mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Building Dashboards? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Building Dashboards

You design dashboards that help humans make decisions quickly. Dashboards are products: audience, questions, and actions matter more than chart count.

Philosophy

  1. Decisions first. Every panel answers a question that leads to an action.
  2. Overview → drilldown → evidence. Start broad, narrow on click/filter, end with raw logs.
  3. Rates and percentiles over averages. Averages hide problems; p95/p99 expose them.
  4. Simple beats dense. One question per panel. No chart junk.
  5. Validate with data. Never guess fields—discover schema first.

Entry Points

Choose your starting point:

| Starting from | Workflow | |---------------|----------| | Vague description | Intake → check dataset kind → design blueprint (APL or MPL) → queries per panel → deploy | | Template | Pick template → customize dataset/service/env → deploy | | Splunk dashboard | Extract SPL → translate via spl-to-apl → map to chart types → deploy | | Exploration | Use axiom-sre to discover schema/signals → productize into panels |


Intake: What to Ask First

Before designing, clarify:

  1. Audience & decision
  • Oncall triage? (fast refresh, error-focused)
  • Team health? (daily trends, SLO tracking)
  • Exec reporting? (weekly summaries, high-level)
  1. Scope
  • Service, environment, region, cluster, endpoint?
  • Single service or cross-service view?
  1. Dataset kind (mandatory first step)
  • Run scripts/metrics/datasets to identify each dataset's kind
  • If kind is otel:metrics:v1 → this is a metrics dataset. Follow the Metrics path below.
  • Otherwise → this is an events/logs dataset. Follow the APL path below.

> ⚠️ NEVER run getschema on a metrics dataset. APL queries against otel:metrics:v1 datasets return 0 rows without error — you will waste calls widening time ranges before realizing it's the wrong discovery method.

APL path (events/logs datasets):

  • Discover fields with getschema:

``apl ['dataset'] | where _time between (ago(1h) .. now()) | getschema ``

  • Continue to steps 4–5 below.

Metrics path (otel:metrics:v1 datasets):

  • Run scripts/metrics/metrics-spec — mandatory before composing any MPL query
  • Discover available metrics: scripts/metrics/metrics-info metrics
  • Discover tags: scripts/metrics/metrics-info tags
  • Explore tag values: scripts/metrics/metrics-info tags values
  • If discovery returns empty results, retry with --start set to 7 days ago — sparse metrics (sensors, batch jobs, crons) may not have data in the default 24h window
  • find-metrics searches tag values, not metric names — use it only when you know a specific entity name (service, host, device) to find which metrics are associated with it
  • Skip to the Metrics/MPL Blueprint below for panel design.
  1. Golden signals (APL path)
  • Traffic: requests/sec, events/min
  • Errors: error rate, 5xx count
  • Latency: p50, p95, p99 duration
  • Saturation: CPU, memory, queue depth, connections
  1. Drilldown dimensions (APL path)
  • What do users filter/group by? (service, route, status, pod, customer_id)

Dashboard Blueprint

Choose the blueprint that matches your dataset kind (identified in Intake step 3).

APL Blueprint (events/logs datasets)

1. At-a-Glance (Statistic panels)

Single numbers that answer "is it broken right now?"

  • Error rate (last 5m)
  • p95 latency (last 5m)
  • Request rate (last 5m)
  • Active alerts (if applicable)
2. Trends (TimeSeries panels)

Time-based patterns that answer "what changed?"

  • Traffic over time
  • Error rate over time
  • Latency percentiles over time
  • Stacked by status/service for comparison
3. Breakdowns (Table/Pie panels)

Top-N analysis that answers "where should I look?"

  • Top 10 failing routes
  • Top 10 error messages
  • Worst pods by error rate
  • Request distribution by status
4. Evidence (LogStream + SmartFilter)

Raw events that answer "what exactly happened?"

  • LogStream filtered to errors
  • SmartFilter for service/env/route
  • Key fields projected for readability

Metrics/MPL Blueprint (metrics datasets)

> Prerequisite: You MUST have run scripts/metrics/metrics-spec and scripts/metrics/metrics-info before designing panels. Never guess MPL syntax or metric/tag names.

> 🚨 ALIGNMENT RULE — non-negotiable for dashboard panels: Always align to the dashboard-supplied variable $__interval, not a fixed window. The dashboard runtime substitutes $__interval based on the active time range and panel width, so the same chart stays usable from a 5-minute to a 30-day view. Hard-coding align to 1m (or any constant) over-resolves long ranges and under-resolves short ones. > > ``mpl > | align to $__interval using avg ✅ dashboard panels > | align to 1m using avg ❌ fixed window — wrong granularity at most time ranges > ` > > **No param declaration needed in the chart query.apl** — the dashboard runtime injects param $__interval: Duration; automatically. (The Grafana datasource does the same via a preamble; the Axiom-native dashboard runtime behaves identically — verified against working production dashboards.) > > **Exceptions:** If you are pre-validating a query through scripts/metrics/metrics-query (which has no dashboard runtime), substitute a concrete duration for the test call only — do NOT commit that to the chart JSON. For genuinely sparse metrics where $__interval would round to an empty bucket (sensors, batch jobs, crons), a fixed wider window (e.g. 1h`) is acceptable; document why in the chart description.

1. At-a-Glance (Statistic panels)

Current values for key metrics — answer "what's the state right now?"

  • Latest value of primary metrics (e.g., current temperature, power draw)
  • Use group using avg or group using last depending on metric type (gauge vs counter)
2. Trends (TimeSeries panels)

Metric trends over time — answer "what changed?"

  • Primary metrics over time, grouped by key dimension
  • Use align to $__interval using avg|sum|last for proper time bucketing — $__interval is supplied by the dashboard runtime
  • Group by low-cardinality tags only (≤10 series per chart)
3. Breakdowns (TimeSeries or Table panels)

Per-entity detail — answer "where should I look?"

  • Metrics broken down by entity (room, host, pod, service)
  • Filter by tag values to keep series count manageable
  • Use separate panels per dimension rather than one overloaded chart
4. Entity State (TimeSeries or Table panels)

Boolean/state metrics — answer "what is on/off/active?"

  • Use align to $__interval using last for state metrics
  • Sparse metrics may need wider fixed align intervals (1h+) to show data — this is the documented exception to the $__interval rule

Layout Auto-Normalization

The console uses react-grid-layout which requires minH, minW, moved, and static on every layout entry. The dashboard-create and dashboard-update scripts auto-fill these if omitted, so layout entries only need i, x, y, w, h.


Required Chart Structure

Every chart MUST have a unique id field. Every layout entry's i field MUST reference a chart id. Missing or mismatched IDs will corrupt the dashboard in the UI (blank state, unable to save/revert).

{
  "charts": [
    {
      "id": "error-rate",
      "name": "Error Rate",
      "type": "Statistic",
      "query": { "apl": "..." }
    }
  ],
  "layout": [
    {"i": "error-rate", "x": 0, "y": 0, "w": 3, "h": 2}
  ]
}

Use descriptive kebab-case IDs (e.g. error-rate, p95-latency, traffic-rps). The dashboard-validate and deploy scripts enforce this automatically.


Metrics/MPL Chart Contract

Metrics-backed charts require both query.apl (the MPL pipeline string) and query.metricsDataset (the dataset name). The metricsDataset field is what tells the backend to interpret apl as MPL rather than APL — omitting it causes the chart to misbehave even if the pipeline string is well-formed.

> CRITICAL: Run scripts/metrics/metrics-spec before composing your first MPL query in a session. NEVER guess MPL syntax. > > API gotcha: Set query.metricsDataset to the dataset name (e.g. "otel-metrics"). The create API rejects query.mpl even though GET responses for existing metrics dashboards may include it — put the MPL string in query.apl instead.

{
  "type": "TimeSeries",
  "query": {
    "apl": "`otel-metrics`:`http.server.duration`\n| where `service.name` == \"api\"\n| align to $__interval using avg\n| group by `service.name` using avg",
    "metricsDataset": "otel-metrics"
  }
}

Validate queries with scripts/metrics/metrics-query before embedding in dashboard JSON.

See reference/metrics-mpl.md for the full contract and discovery scripts.


Chart Types

Note: Dashboard queries inherit time from the UI picker—no explicit _time filter needed.

Validation: TimeSeries, Statistic, Table, Pie, LogStream, Note, MonitorList are fully validated by dashboard-validate. Heatmap, ScatterPlot, SmartFilter work but may trigger warnings.

Statistic

When: Single KPI, current value, threshold comparison.

['logs']
| where service == "api"
| summarize 
    total = count(),
    errors = countif(status >= 500)
| extend error_rate = round(100.0 * errors / total, 2)
| project error_rate

Pitfalls: Don't use for time series; ensure query returns single row.

TimeSeries

When: Trends over time, before/after comparison, rate changes.

// Single metric - use bin_auto for automatic sizing
['logs']
| summarize ['req/min'] = count() by bin_auto(_time)

// Latency percentiles - use percentiles_array for proper overlay
['logs']
| summarize percentiles_array(duration_ms, 50, 95, 99) by bin_auto(_time)

Best practices:

  • Use bin_auto(_time) instead of fixed bin(_time, 1m) — auto-adjusts to time window
  • Use percentiles_array() instead of multiple percentile() calls — renders as one chart
  • Too many series = unreadable; use top N or filter

Table

When: Top-N lists, detailed breakdowns, exportable data.

['logs']
| where status >= 500
| summarize errors = count() by route, error_message
| top 10 by errors
| project route, error_message, errors

Pitfalls:

  • Always use top N to prevent unbounded results
  • Use project to control column order and names

Pie

When: Share-of-total for LOW cardinality dimensions (≤6 slices).

['logs']
| summarize count() by status_class = case(
    status 6 categories
- Always aggregate to reduce slices

### LogStream
**When:** Raw event inspection, debugging, evidence gathering.

```apl
['logs']
| where service == "api" and status >= 500
| project-keep _time, trace_id, route, status, error_message, duration_ms
| take 100

Pitfalls:

  • Always include take N (100-500 max)
  • Use project-keep to show relevant fields only
  • Filter aggressively—raw logs are expensive

Heatmap

When: Distribution visualization, latency patterns, density analysis.

['logs']
| summarize histogram(duration_ms, 15) by bin_auto(_time)

Best for: Latency distributions, response time patterns, identifying outliers.

Scatter Plot

When: Correlation between two metrics, identifying patterns.

['logs']
| summarize avg(duration_ms), avg(resp_size_bytes) by route

Best for: Response size vs latency correlation, resource usage patterns.

SmartFilter (Filter Bar)

When: Interactive filtering for the entire dashboard.

SmartFilter is a chart type that creates dropdown/search filters. Requires:

  1. A SmartFilter chart with filter definitions
  2. declare query_parameters in each panel query

Filter types:

  • selectType: "apl" — Dynamic dropdown from APL query
  • selectType: "list" — Static dropdown with predefined options
  • type: "search" — Free-text input

Panel query pattern:

declare query_parameters (country_filter:string = "");
['logs'] | where isempty(country_filter) or ['geo.country'] == country_filter

See reference/smartfilter.md for full JSON structure and cascading filter examples.

Monitor List

When: Display monitor status on operational dashboards.

No APL needed—select monitors from the UI. Shows:

  • Monitor status (normal/triggered/off)
  • Run history (green/red squares)
  • Dataset, type, notifiers

Note

When: Context, instructions, section headers.

Use GitHub Flavored Markdown for:

  • Dashboard purpose and audience
  • Runbook links
  • Section dividers
  • On-call instructions

Chart Configuration

Charts support JSON configuration options beyond the query. See reference/chart-config.md for full details.

Quick reference:

| Chart Type | Key Options | |------------|-------------| | Statistic | colorScheme, customUnits, unit, showChart (sparkline), errorThreshold/warningThreshold | | TimeSeries | aggChartOpts: variant (line/area/bars), scaleDistr (linear/log), displayNull | | LogStream/Table | tableSettings: columns, fontSize, highlightSeverity, wrapLines | | Pie | hideHeader | | Note | text (markdown), variant |

Common options (all charts):

  • overrideDashboardTimeRange: boolean
  • overrideDashboardCompareAgainst: boolean
  • hideHeader: boolean

APL Patterns

Time Filtering in Dashboards vs Ad-hoc Queries

Dashboard panel queries do NOT need explicit time filters. The dashboard UI time picker automatically scopes all queries to the selected time window.

// DASHBOARD QUERY — no time filter needed
['logs']
| where service == "api"
| summarize count() by bin_auto(_time)

Ad-hoc queries (Axiom Query tab, axiom-sre exploration) MUST have explicit time filters:

// AD-HOC QUERY — always include time filter
['logs']
| where _time between (ago(1h) .. now())
| where service == "api"
| summarize count() by bin_auto(_time)

Bin Size Selection

Prefer bin_auto(_time) — it automatically adjusts to the dashboard time window.

Manual bin sizes (only when auto doesn't fit your needs):

| Time window | Bin size | |-------------|----------| | 15m | 10s–30s | | 1h | 1m | | 6h | 5m | | 24h | 15m–1h | | 7d | 1h–6h |

Cardinality Guardrails

Prevent query explosion:

// GOOD: bounded
| summarize count() by route | top 10 by count_

// BAD: unbounded high-cardinality grouping
| summarize count() by user_id  // millions of rows

Field Escaping

Fields with dots need bracket notation:

| where ['kubernetes.pod.name'] == "frontend"

Fields with dots IN the name (not hierarchy) need escaping:

| where ['kubernetes.labels.app\\.kubernetes\\.io/name'] == "frontend"

Golden Signal Queries

Traffic:

| summarize requests = count() by bin_auto(_time)

Errors (as rate %):

| summarize total = count(), errors = countif(status >= 500) by bin_auto(_time)
| extend error_rate = iff(total > 0, round(100.0 * errors / total, 2), 0.0)
| project _time, error_rate

Latency (use percentiles_array for proper chart overlay):

| summarize percentiles_array(duration_ms, 50, 95, 99) by bin_auto(_time)

Layout Composition

Grid Principles

  • Dashboard width = 12 units
  • Typical panel: w=3 (quarter), w=4 (third), w=6 (half), w=12 (full)
  • Stats row: 4 panels × w=3, h=2
  • TimeSeries row: 2 panels × w=6, h=4
  • Tables: w=6 or w=12, h=4–6
  • LogStream: w=12, h=6–8

Section Layout Pattern

Row 0-1:  [Stat w=3] [Stat w=3] [Stat w=3] [Stat w=3]
Row 2-5:  [TimeSeries w=6, h=4] [TimeSeries w=6, h=4]
Row 6-9:  [Table w=6, h=4] [Pie w=6, h=4]
Row 10+:  [LogStream w=12, h=6]

Naming Conventions

  • Use question-style titles: "Error rate by route" not "Errors"
  • Prefix with context if multi-service: "[API] Error rate"
  • Include un

…

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.