AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Ab Test Analysis

skill-phuryn-pm-skills-ab-test-analysis · by phuryn

Analyze A/B test results with statistical significance, sample size validation, confidence intervals, and ship/extend/stop recommendations. Use when evaluating experiment results, checking if a test reached significance, interpreting split test data, or deciding whether to ship a variant.

No reviews yet
0 installs
1 views
0.0% view→install

Install

$ agentstack add skill-phuryn-pm-skills-ab-test-analysis

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-phuryn-pm-skills-ab-test-analysis)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
1mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Ab Test Analysis? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

A/B Test Analysis

Evaluate A/B test results with statistical rigor and translate findings into clear product decisions.

Context

You are analyzing A/B test results for $ARGUMENTS.

If the user provides data files (CSV, Excel, or analytics exports), read and analyze them directly. Generate Python scripts for statistical calculations when needed.

Instructions

  1. Understand the experiment:
  • What was the hypothesis?
  • What was changed (the variant)?
  • What is the primary metric? Any guardrail metrics?
  • How long did the test run?
  • What is the traffic split?
  1. Validate the test setup:
  • Sample size: Is the sample large enough for the expected effect size?
  • Use the formula: n = (Z²α/2 × 2 × p × (1-p)) / MDE²
  • Flag if the test is underpowered (<80% power)
  • Duration: Did the test run for at least 1-2 full business cycles?
  • Randomization: Any evidence of sample ratio mismatch (SRM)?
  • Novelty/primacy effects: Was there enough time to wash out initial behavior changes?
  1. Calculate statistical significance:
  • Conversion rate for control and variant
  • Relative lift: (variant - control) / control × 100
  • p-value: Using a two-tailed z-test or chi-squared test
  • Confidence interval: 95% CI for the difference
  • Statistical significance: Is p < 0.05?
  • Practical significance: Is the lift meaningful for the business?

If the user provides raw data, generate and run a Python script to calculate these.

  1. Check guardrail metrics:
  • Did any guardrail metrics (revenue, engagement, page load time) degrade?
  • A winning primary metric with degraded guardrails may not be a true win
  1. Interpret results:

| Outcome | Recommendation | |---|---| | Significant positive lift, no guardrail issues | Ship it — roll out to 100% | | Significant positive lift, guardrail concerns | Investigate — understand trade-offs before shipping | | Not significant, positive trend | Extend the test — need more data or larger effect | | Not significant, flat | Stop the test — no meaningful difference detected | | Significant negative lift | Don't ship — revert to control, analyze why |

  1. Provide the analysis summary:

``` ## A/B Test Results: [Test Name]

Hypothesis: [What we expected] Duration: [X days] | Sample: [N control / M variant]

| Metric | Control | Variant | Lift | p-value | Significant? | |---|---|---|---|---|---| | [Primary] | X% | Y% | +Z% | 0.0X | Yes/No | | [Guardrail] | ... | ... | ... | ... | ... |

Recommendation: [Ship / Extend / Stop / Investigate] Reasoning: [Why] Next steps: [What to do] ```

Think step by step. Save as markdown. Generate Python scripts for calculations if raw data is provided.


Further Reading

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

  • Author: phuryn
  • Source: phuryn/pm-skills
  • License: MIT
  • Homepage: https://www.productcompass.pm/p/pm-skills-2-red-team-ship

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.