Install
$ agentstack add skill-hoangsonww-claude-code-agent-monitor-benchmark ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Benchmark
Score a session against the rolling population average and report its percentile on cost, tokens, tool count, and complexity using Agent Monitor data.
Input
The user provides: $ARGUMENTS
This may be:
- A single session ID — benchmark that session
- "latest" — benchmark the most recent session
- "latest N" — benchmark the N most recent sessions, each vs the average
- empty — benchmark the most recent session (default)
Data Sources
| Endpoint | Returns | |----------|---------| | GET /api/sessions?limit=N | Population of sessions with cost, model, started_at, metadata (turncount, totalturndurationms) — builds the rolling baseline | | GET /api/pricing/cost/{sessionId} | { total_cost, breakdown:[{ input_tokens, output_tokens, cache_read_tokens, cache_write_tokens, cost }] } — the target session's cost and tokens | | GET /api/workflows/{sessionId} | complexity (score), stats (tool/event counts), toolFlow (distinct tools used) — the target session's tool count and complexity | | GET /api/analytics | avg_events_per_session, tool_usage, daily_sessions — corroborates population-level averages |
Report Sections
1. Build the Baseline
Fetch the population with GET /api/sessions?limit=200 (the rolling set). For each session gather cost (GET /api/pricing/cost/{id} or the list cost field), total tokens (sum of the 4 token types from the pricing breakdown), tool count and complexity (GET /api/workflows/{id}). Compute mean, median, and standard deviation for each metric across the population.
2. Measure the Target
For the requested session, pull the same four metrics:
- Cost —
total_costfromGET /api/pricing/cost/{id}. - Total tokens —
input + output + cache_read + cache_writesummed from the breakdown. - Tool count — distinct/total tools from
GET /api/workflows/{id}stats/toolFlow. - Complexity score —
complexity.scorefromGET /api/workflows/{id}.
3. Percentile and Deviation
For each metric report the target's percentile within the population (share of sessions at or below it) and its z-score (value − mean) / stddev. Label each: below average / typical / above average / outlier (|z| > 2).
4. Verdict
State whether the session was normal overall. If it is an outlier, name which metric drove it (e.g., complexity p96, cost p91 → an unusually heavy session).
Output
- A Markdown table: metric | session value | population mean | percentile | z-score | label.
- Currency in USD to 4 decimals; tokens and tool counts as integers; complexity to 2 decimals.
- Use ▲ for above-average and ▼ for below-average vs the mean.
- One-line verdict: "Normal session" or "Outlier — driven by (pNN)".
- When benchmarking multiple sessions, one row block per session plus a summary line.
- Read-only: percentiles come only from the fetched population; never fabricate the baseline.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: hoangsonww
- Source: hoangsonww/Claude-Code-Agent-Monitor
- License: MIT
- Homepage: https://hoangsonww.github.io/Claude-Code-Agent-Monitor/
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.