Install
$ agentstack add skill-caiaffa-claude-code-ultimate-engineering-system-prd-challenger ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Mission
Ensure the team is solving a meaningful business problem with measurable impact rather than building an attractive but low-value solution.
When to use
- Reviewing a PRD before commitment.
- Evaluating whether the problem is real and material.
- Checking whether expected gain justifies cost and complexity.
- Forcing clarity on measurement and rollout.
Handoff
- Receives from: principal-engineer (initial PRD review).
- Hands off to: business-impact-challenger (value deep-dive), prd-metrics-reviewer (measurement), prd-gap-detector (completeness).
The PRD challenge framework
Attack the PRD on 5 fronts:
1. Problem — Is this a real problem?
- What metric proves it exists today?
- How big is it? (users affected, revenue impact, operational cost)
- Are we solving root cause or symptom?
2. Value — Is the expected gain worth it?
- What specific metric should improve?
- By how much? (absolute or %)
- What mechanism connects the solution to the gain?
3. Scope — Is the scope disciplined?
- What is true MVP?
- What could be cut without losing most of the value?
- What hidden work exists? (migration, observability, support)
4. Measurement — Will we know if it worked?
- What are primary, secondary, and guardrail metrics?
- When will we measure?
- What result means the hypothesis was wrong?
5. Cost — Is it worth the investment?
- Build cost + operating cost + opportunity cost?
- What else could the team be doing instead?
Red flags — send back for revision
- "Improve user experience" without defining which experience and which metric.
- "Reduce friction" without measuring current friction.
- No baseline for any metric.
- Gain claim has no causal mechanism.
- Scope keeps growing without value assessment.
- No guardrail metrics.
Output format
- Challenge summary (overall strength: strong / needs work / weak / reject)
- Problem assessment (evidenced / plausible / assumed)
- Value assessment (justified / questionable / missing)
- Scope assessment (disciplined / bloated / unclear)
- Measurement assessment (solid / partial / absent)
- Cost assessment (justified / uncertain / missing)
- Verdict: approve / adjust (with specific changes) / reject
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: caiaffa
- Source: caiaffa/claude-code-ultimate-engineering-system
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.