AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Shipkit Product Goals

skill-stefan-stepzero-shipkit-shipkit-product-goals · by stefan-stepzero

Derive measurable user-outcome criteria from the product blueprint. Writes goals/product.json (completion rates, UX quality, satisfaction). Evaluate mode compares actuals to targets.

No reviews yet
0 installs
28 views
0.0% view→install

Install

$ agentstack add skill-stefan-stepzero-shipkit-shipkit-product-goals

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-stefan-stepzero-shipkit-shipkit-product-goals)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
2mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Shipkit Product Goals? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

shipkit-product-goals — Product Success Criteria

Purpose: Derive measurable user-outcome criteria from the product blueprint. Each feature, UX pattern, and differentiator implies criteria for "how do we know this works?" — this skill makes those criteria explicit, measurable, and trackable.

What it does: Reads the product blueprint + discovery context + stage from strategic.json, proposes user-outcome criteria (P-), appends P- IDs to existing stage gates, then writes goals/product.json.

Protocol: This skill follows the canonical elicitation protocol defined in install/shared/references/elicitation-protocol.md (the mechanics — marker, state files, resume). The steps below are this skill's specific application of that protocol.

Calibration: Apply install/shared/references/ground-or-ask-calibration.md (the intelligence — propose vs ask). Ground first: This is a DERIVATION skill — the vast majority of P-* criterion fields are directly derivable from cited upstream artifacts (product-definition.json, product-discovery.json, goals/strategic.json). Propose every criterion you can ground, tagged with its source; flag low-leverage guesses as guessed. The only genuinely high-leverage ungrounded fields are the concrete numeric thresholds (e.g., "> 80% completion rate") when no signal from stage calibration, prior metrics, or product-definition implies a specific value. Ask those — and only those. Keep the question set tiny (often zero). Do not ask what a signal already answers; do not silently invent a high-leverage threshold.

Output: One JSON file:

  • goals/product.json — User-outcome criteria (PM owns)

> Strategic goals (stage, constraints, business metrics: S-) are handled by /shipkit-stage, owned by the Visionary agent. > Engineering goals (technical performance: E-) are handled by /shipkit-engineering-goals, owned by the Architect agent.


Modes

| Mode | Trigger | What It Does | |------|---------|-------------| | Define (default) | "Set goals", "success criteria", "product goals" | Derive P- criteria from product blueprint, write product.json, append P- to gates | | Evaluate | --evaluate, "check metrics" | Read metrics/latest.json, compare product criteria to targets, output gap report |


When to Invoke

User triggers:

  • "Define success criteria", "How do we measure success?"
  • "Set product goals", "What are our goals?"
  • "Evaluate goals", "Check metrics"

Workflow position:

  • After /shipkit-product-definition (reads the product blueprint)
  • After /shipkit-stage (reads stage and constraints from strategic.json)
  • Before /shipkit-engineering-goals (engineering reads product goals for alignment)
  • Before /shipkit-spec — criteria inform feature specifications
  • Before /shipkit-review-shipping — criteria become verification checks

Prerequisites

| File | Required? | Mode | Provides | If Missing | |------|-----------|------|----------|------------| | .shipkit/product-definition.json | Yes | Define | Features, patterns, differentiators | Route to /shipkit-product-definition | | .shipkit/product-discovery.json | Recommended | Define | Pain points for traceability | Proceed without traceability | | .shipkit/goals/strategic.json | Recommended | Define | Stage, constraints for calibration | Proceed with defaults; suggest /shipkit-stage | | .shipkit/goals/product.json | Yes | Evaluate | Current targets | Route to Define mode first | | .shipkit/metrics/latest.json | Yes | Evaluate | Current actuals | Report "no metrics available" |


Output File

| File | Owner Agent | Criteria Category | Contains | |------|------------|-------------------|----------| | goals/product.json | PM | user-outcome | Completion rates, UX quality, satisfaction, content quality |


Process — Define Mode

Completion Tracking (MANDATORY)

After loading context (Step 1), create tasks:

  1. TaskCreate: "Load context (product-definition + strategic.json)"
  2. TaskCreate: "Derive P-* criteria with rubrics from product blueprint"
  3. TaskCreate: "Classify checkability + verificationTool for each criterion"
  4. TaskCreate: "Map P-* criteria to existing gates in strategic.json"
  5. TaskCreate: "Archive existing artifact (if replacing)"
  6. TaskCreate: "Write goals/product.json"
  7. TaskCreate: "Update strategic.json gates with P-* IDs"
  8. TaskCreate: "Verify summary counts match actual array length"

Rules:

  • Writing product.json (task 6) is NOT done — strategic.json gates must also be updated (task 7)
  • TaskUpdate the gates task to completed only after reading strategic.json back and confirming P-* IDs appear in gate criteria arrays
  • Every criterion must have a rubric (3-5 levels) — bare thresholds fail the criteria derivation task
  • Do NOT present the final summary until ALL tasks show completed

Step 0: Check for Existing Files

  1. Check if .shipkit/goals/product.json exists
  2. If exists: archive current file to .shipkit/.archive/goals-product.YYYY-MM-DD.json and regenerate (fork context — no user prompt; let the reviewer catch over-eager rewrites)
  3. If legacy .shipkit/goals.json exists: migrate (see Migration section)
  4. If nothing exists: Skip to Step 1

Step 0.5: Check Elicitation State (Threshold Answers)

Read .shipkit/elicitation/product-goals/answers.md (or /elicitation/product-goals/answers.md when running under the orchestration engine — see install/shared/references/run-artifacts.md).

  • If real threshold answers are present → use those values when proposing thresholds in Step 3. Proceed with derivation as normal.
  • If absent or empty → proceed; the grounding pass in Step 3 will determine whether any thresholds remain ungrounded.

Step 1: Load Context

Read these files:

.shipkit/product-definition.json      → features, patterns, differentiators (REQUIRED)
.shipkit/product-discovery.json       → pain points for traceability (RECOMMENDED)
.shipkit/goals/strategic.json         → stage, constraints for calibration (RECOMMENDED)

If product-definition.json missing: Route to /shipkit-product-definition first. If strategic.json missing: Suggest running /shipkit-stage first. Proceed with defaults if user declines.


Step 2: Read Stage Context

Read stage and constraints from goals/strategic.json (set by /shipkit-stage):

| Stage | Product Criteria Depth | Focus | |-------|----------------------|-------| | POC | Basic: "core flow completable" (2-3 criteria) | Functional completeness | | Alpha | Light: core path usability (3-5 criteria) | Core usability | | MVP | Moderate: user outcome thresholds (5-10 criteria) | Usability + satisfaction | | Scale | Comprehensive: full UX quality (10-15 criteria) | Growth + user delight |

If goals/strategic.json is missing, default to MVP stage (fork context — no user prompt; dispatch /shipkit-stage first if stage needs to be set explicitly).


Step 3: Derive Criteria from Product Blueprint

For each section of product-definition.json, derive user-outcome criteria (P-*):

From UX patterns:

  • Usability criteria (can users complete the flow?)
  • Completion rate criteria (what % finish?)
  • Responsiveness criteria (how fast does it feel?)

From differentiators:

  • Validation criteria (does it actually differentiate?)
  • User perception criteria (do users notice/value it?)

From features:

  • Completeness gate (are all gate-scoped features functional?)
  • Integration gate (do features work together end-to-end?)

> Business metrics (S-* criteria) are defined by /shipkit-stage, not this skill.

Every threshold MUST include a rubric. A bare number like "> 80%" is meaningless without defining what each level looks like. For each criterion, generate a rubric with 3-5 level descriptors:

Example rubric for "Wizard completion rate":
  0-20%: Users abandon immediately — flow is broken or confusing
  20-50%: Users attempt but hit blockers — missing guidance or errors
  50-80%: Users complete with effort — friction points remain
  80-95%: Users complete smoothly — minor polish needed
  95-100%: Users complete effortlessly — flow is intuitive
  Target: > 80% (smooth completion)

The rubric makes the threshold defensible — reviewers can assess where the product actually falls and what gap to close.

See references/derivation-patterns.md for detailed derivation examples.


Step 3b: Classify Checkability + Verification Tool

For each derived criterion, assign checkability and verificationTool:

| verificationMethod | Context | checkability | verificationTool | |----------------------|---------|---------------|-------------------| | automated-test | Tests UI flow/rendering | verifiable | visual-qa | | automated-test | Tests code quality/coverage | verifiable | test or lint | | manual-check | Checks app behavior | verifiable | visual-qa or semantic-qa | | manual-check | Needs domain expert review | observable | none | | analytics | Any | observable | none | | user-feedback | Any | observable | none |

When uncertain: default to observable.

Key distinction: "E2E flow passes" = verifiable (Playwright can simulate). "80% of real users complete flow" = observable (needs real funnel data).


Step 4: Threshold Resolution — Ground or Ask

After completing Step 3 derivation, classify each proposed threshold:

| Class | When | Action | |-------|------|--------| | Grounded | Stage calibration table, product-definition patterns, or prior metrics imply the value | Propose, tagged with source. Proceed. | | Low-leverage ungrounded | Typical industry default (e.g., "> 80% MVP completion rate"), cheap to adjust later | Flag as guessed: true. Proceed. | | High-leverage ungrounded | Cannot be inferred from any signal AND shapes scope or sets a hard constraint (launch gate, partner SLA, satisfaction bar that is one-way-door) | Must be asked. |

If ANY high-leverage ungrounded thresholds exist — fork path:

  1. Write .shipkit/elicitation/product-goals/questions.md listing the specific thresholds that need values (one question per ungrounded threshold; include the derived criterion name and what the question is resolving).
  2. Write .shipkit/elicitation/product-goals/progress.json with status: in_progress.
  3. Do NOT write goals/product.json. Do NOT invent the threshold value.
  4. Emit the following as the final line of your output — nothing may follow it:

`` NEEDS_ELICITATION:shipkit-product-goals status=paused questions_file=.shipkit/elicitation/product-goals/questions.md reason=awaiting threshold values for high-leverage criteria ``

If all thresholds are grounded or low-leverage flagged:

Proceed directly to Step 5 (Map to Gates) and write goals/product.json in Step 6 as normal.

> Never silently invent an ungrounded high-leverage threshold. A guessed launch gate or SLA-bound success rate that turns out to be wrong is hard to discover late and expensive to reverse. If genuinely uncertain, emit the marker — even if the question count is just one or two.


Step 5: Map P-* to Existing Gates

Read gates from goals/strategic.json (defined by /shipkit-stage) and assign P-* criteria to them:

ADDING PRODUCT CRITERIA TO GATES:

MVP Launch Ready:
  Existing: S-001
  + P-001 (wizard completion time)
  + P-002 (wizard flow E2E)
  + P-003 (preview responsiveness)

Beta Ready:
  Existing: S-002, S-003
  + P-004 (teachers perceive speed advantage)
  + P-005 (teachers trust standards alignment)

If goals/strategic.json doesn't exist or has no gates, define gates locally in product.json and note they should be merged when /shipkit-stage runs.

Assign gates directly — no user prompt (fork context). The reviewer will flag misalignments in the loop's review cycle.


Step 5.5: Archive Existing Artifact

Artifact strategy: archive — Before writing, if the target file already exists, move it to .shipkit/archive/{filename}.{ISO-date}.json (create the archive/ directory if needed). Then write the new artifact fresh.


Step 6: Generate Product Goal File

After confirmation:

  1. Write .shipkit/goals/product.json — user-outcome criteria
  2. If goals/strategic.json exists, update its gates criteria arrays to include P-* IDs

See references/output-schema.md for full schema.


Step 7: Save and Suggest Next Steps

Product goals saved to .shipkit/goals/product.json

  Stage: {stage} (from strategic.json)
  Product criteria: {N}
  Gates updated: {N} (P-* criteria appended)

  Status: {N} not-measured

Next:
  1. /shipkit-engineering-goals — Add technical performance criteria (E-*)
  2. /shipkit-spec — Create specs for features (criteria inform acceptance tests)
  3. /shipkit-product-goals --evaluate — Check product criteria status after building

Process — Evaluate Mode

When invoked with --evaluate or "check metrics":

Step 1: Load Goal File

Read goals/product.json from .shipkit/goals/.

Step 2: Load Metrics

Read .shipkit/metrics/latest.json for current actuals.

If missing: Report "No metrics file found. Create .shipkit/metrics/latest.json with current measurements to enable evaluation."

Step 3: Compare Actuals to Targets

For each criterion in goals/product.json:

  • Match metric key from metrics/latest.json to criterion ID
  • Compare actual value to threshold
  • Update status: not-measured, below-threshold, at-threshold, exceeded

Step 5: Output Gap Report

Separate gaps into "Verifiable" (can be closed with tools) and "Observable" (needs real data):

PRODUCT GOAL EVALUATION — {date}

Stage: {stage}

═══ VERIFIABLE GAPS (run these tools to close) ═══

  ✗ P-003: Wizard flow E2E — not measured [visual-qa]
  ✓ P-004: Export works [visual-qa]

═══ OBSERVABLE GAPS (need real data) ═══

  ◌ P-001: Completion rate — needs real user funnel
  ◌ P-005: Teachers perceive speed advantage — needs user feedback

SUMMARY:
  Product criteria: {N}/{M} passing
  Verifiable: {V} passing, {V_gap} gaps
  Observable: {O} awaiting data

Note: Run /shipkit-stage --evaluate for full gate status across all goal files.
Note: Run /shipkit-engineering-goals --evaluate for technical criteria status.

Migration from Legacy Format

When .shipkit/goals.json (single file) exists or files have "source": "shipkit-goals" (old unified skill):

  1. Detect legacy format
  2. Offer migration
  3. Split criteria by category:
  • business-metric → strategic.json
  • user-outcome → product.json
  • technical-performance → leave for /shipkit-engineering-goals to handle
  1. Archive original: .shipkit/.archive/goals.YYYY-MM-DD.json

Goal File Schema

goals/product.json

{
  "$schema": "shipkit-artifact",
  "type": "goals-product",
  "version": "4.0",
  "lastUpdated": "YYYY-MM-DDTHH:MM:SSZ",
  "source": "shipkit-product-goals",

  "derivedFrom": {
    "productDefinition": ".shipkit/product-definition.json",
    "productDiscovery": ".shipkit/product-discovery.json"
  },

  "criteria": [
    {
      "id": "P-001",
      "name": "Feature completion rate",
      "category": "user-outcome",
      "metric": "% of users completing core flow",
      "threshold": "> 80%",
      "rubric": [
        { "range": "0-20%", "meaning": "Flow broken — users abandon immediately" },
        { "range": "20-50%", "meaning": "Major friction — users attempt but hit blockers" },
        { "range": "50-80%", "meaning": "Workable — most complete with effort" },
        { "range": "80-95%", "meaning": "Smooth — users complete without confusion" },
        { "range": "95-100%", "meaning": "Effortless — intuitive, no hesitation" }
      ],
      "currentValue": null,
      "verificationMethod": "analytics",
      "checkability": "observable

…

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [stefan-stepzero](https://github.com/stefan-stepzero)
- **Source:** [stefan-stepzero/shipkit](https://github.com/stefan-stepzero/shipkit)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.