Install
$ agentstack add skill-eigent-ai-agent-skills-ab-test-setup ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
A/B Test Setup
Overview
Use this skill to guide the full experiment lifecycle: hypothesis, design, sample size, implementation, analysis, and playbook documentation. Keep tests focused, measurable, and resistant to common errors like peeking early or testing too many changes at once.
Workflow
- Define the business goal, user segment, current baseline, target metric, and guardrail metrics.
- Write a specific hypothesis:
If we change X for audience Y, metric Z will improve because...
- Design the test:
- Control and variant.
- Primary metric.
- Secondary and guardrail metrics.
- Traffic split, eligibility, exclusions, and duration.
- Estimate sample size or minimum detectable effect when baseline traffic and conversion rates are available.
- Create an implementation checklist:
- Tracking, randomization, QA, exposure logging, analytics events, and rollback.
- Define decision rules before launch:
- Ship, revert, iterate, or continue testing.
- Analyze results after the test reaches the agreed sample size.
- Document what changed, what was learned, and follow-up experiments.
Test Backlog Pattern
When building a backlog, score each idea with ICE:
- Impact: expected business or user benefit.
- Confidence: evidence quality.
- Effort: complexity and implementation cost.
Prioritize tests that combine high impact, credible evidence, and low operational risk.
Example Prompts
I want to A/B test our signup CTA button. Current conversion rate is 3.2%, 8,000 visitors/month. Help me design the test, calculate the required sample size, and define what success looks like.Our A/B test just hit sample size. Here are the results [paste metrics]. Is this statistically significant? Should we ship the variant, revert, or keep testing?Build a prioritized A/B test backlog for our onboarding flow. Use ICE scoring. Sources to mine: our drop-off analytics, last month's support tickets, and these 3 heatmap observations.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: eigent-ai
- Source: eigent-ai/agent-skills
- License: Apache-2.0
- Homepage: https://www.eigent.ai/skills
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.