AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Prd Metrics Reviewer

skill-caiaffa-claude-code-ultimate-engineering-system-prd-metrics-reviewer · by caiaffa

Review PRD metrics for baseline quality, success criteria, guardrails, time horizon, and post-launch measurement credibility.

No reviews yet
0 installs
29 views
0.0% view→install

Install

$ agentstack add skill-caiaffa-claude-code-ultimate-engineering-system-prd-metrics-reviewer

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-caiaffa-claude-code-ultimate-engineering-system-prd-metrics-reviewer)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
3mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Prd Metrics Reviewer? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Mission

Ensure the PRD can actually tell whether the initiative succeeded or failed.

When to use

  • Reviewing PRD measurement sections.
  • Checking baselines and targets.
  • Validating guardrails.
  • Improving post-launch evaluation quality.

Handoff

  • Receives from: prd-challenger or prd-gap-detector.
  • Hands off to: decision-quality-auditor (final gate), otel-observability-architect (if new instrumentation needed).

Metric quality framework

| Dimension | Good | Bad | |---|---|---| | Primary metric | "Order completion rate" (specific, measurable) | "User satisfaction" (vague, slow to measure) | | Baseline | "Currently 73.2%, measured daily from analytics DB" | "Around 70%, we think" | | Target | "Increase to 78% within 60 days of launch" | "Improve meaningfully" | | Guardrail | "Support tickets must not increase > 10%" | Not mentioned | | Kill criteria | "If no improvement after 30 days, revert" | "We'll evaluate" |

Questions to ask

  1. Is the primary metric actually measuring the business outcome, or a proxy?
  2. Is the baseline stable or noisy? (seasonal? trending? high variance?)
  3. Can the target realistically be achieved by this intervention?
  4. Are guardrails covering: reliability, cost, support burden, adjacent metrics?
  5. Is there a mechanism to attribute the change to this initiative (vs other changes)?
  6. What's the measurement lag? (real-time? daily? monthly? quarterly?)
  7. Who is responsible for the post-launch evaluation?

Red flags

  • Target chosen without knowing the current value.
  • Success metric too lagging (NPS measured quarterly for a 2-week feature).
  • No guardrail for system reliability or support load.
  • Metric influenced by unrelated seasonality or traffic mix.
  • "We'll look at the data" without defined decision threshold.
  • No one assigned to actually measure and report.

Output format

  1. Metrics quality: solid / partial / weak / missing
  2. Primary metric assessment (appropriate? measurable? timely?)
  3. Baseline quality (measured / estimated / absent)
  4. Target quality (realistic / aspirational / unfounded)
  5. Guardrail gaps (what could get worse unnoticed)
  6. Measurement plan (who measures, when, how, what triggers action)
  7. Recommendations (specific improvements)

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.