Install
$ agentstack add skill-vibbs-company-os-experiment-framework ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
Experiment Framework
Reference
- ID: S-QA-08
- Category: QA / Experimentation
- Inputs: PRD (success metrics), feature flag spec, analytics config, baseline conversion rates
- Outputs: experiment spec → artifacts/experiments/, experiment results → artifacts/experiments/
- Used by: QA & Release Agent
- Tool scripts: ./tools/artifact/validate.sh, ./tools/qa/experiment-report.sh
Purpose
Bring statistical rigor to product experiments. Most teams "ship and see" -- they launch a feature, watch a dashboard, and declare victory if a metric goes up. This leads to false conclusions, wasted engineering time, and product decisions driven by noise rather than signal. The experiment framework ensures every test has a clear hypothesis, proper sample size, controlled execution, and structured analysis.
Move beyond intuition-driven development to hypothesis-driven development. Every experiment answers a specific question with a quantified answer. When a feature "wins," you know the effect size, the confidence level, and whether it came at the cost of degrading something else. When a feature "loses," you learn something concrete rather than wondering if you just didn't run the test long enough.
When to Use
- Planning an A/B test for a new feature
- Validating whether a change actually improves metrics
- Feature flag strategy is
fullor feature uses experiment-type flags - Product Agent defines success metrics that need causal validation
- Activation experiments from activation-onboarding skill need statistical design
Experiment Framework Procedure
Step 0.5: Traffic Feasibility Check
Before designing any experiment:
- If
analytics.provideris configured, estimate current DAU from available data - Calculate minimum feasible MDE (Minimum Detectable Effect) given traffic and 90-day max duration
- If DAU 2 variants)
Sample Size Lookup Table (two-sided test, alpha=0.05, power=0.80):
| Baseline Rate | MDE (relative) | Sample Size per Variant | |--------------|-----------------|------------------------| | 1% | 20% | ~130,000 | | 5% | 10% | ~30,000 | | 10% | 10% | ~14,000 | | 20% | 10% | ~6,400 | | 50% | 5% | ~6,400 |
For custom calculations, use the full formula or reference an online sample size calculator. Always round UP to the nearest hundred.
Step 4: Estimate Duration
- Calculate daily eligible traffic (users who enter the experiment)
- Duration = (samplesizepervariant x numvariants) / daily_traffic
- Add buffer for weekday/weekend variation (multiply by 1.2)
- Minimum duration: 7 days (to capture weekly cycles)
- Maximum recommended: 90 days (if longer, the MDE is too small or traffic too low -- reconsider)
- If duration exceeds
experiments.max_concurrentcapacity, flag scheduling conflict
Step 5: Design Experiment Spec
Produce an experiment specification artifact:
---
id: EXP-XXX
type: experiment
status: draft
parent: PRD-XXX
depends_on: [RFC-XXX]
---
Include:
- Hypothesis (from Step 2)
- Primary metric + baseline + MDE
- Secondary metrics
- Guardrail metrics with thresholds
- Sample size per variant
- Estimated duration
- Variant descriptions (control = current, treatment = change)
- Feature flag name (links to feature-flags skill output)
- Assignment method (random, hash-based, segment-based)
- Exclusion criteria (new users only, specific segments, etc.)
Step 6: Define Experiment State Machine
Every experiment follows this lifecycle:
draft -> approved -> running -> analyzing -> concluded -> archived
| State | Entry Criteria | Exit Criteria | Who Owns | |-------|---------------|---------------|----------| | draft | Hypothesis + metrics defined | Spec reviewed by Product + Engineering | QA | | approved | Spec reviewed, sample size validated | Feature flag configured, tracking verified | QA | | running | Flag activated, users being assigned | Duration elapsed OR early termination triggered | QA + Engineering | | analyzing | Experiment stopped, data collected | Analysis complete, recommendation made | QA | | concluded | Decision made (ship/no-ship/iterate) | Flag cleaned up, code path removed or made permanent | Engineering | | archived | Learnings documented | -- | Product |
Step 7: Define Early Termination Rules
- Harm detected: If guardrail metric degrades >20% relative to control, stop immediately
- Clear winner: If primary metric shows >99% probability of being better (sequential testing), can stop early
- No-effect: If 80% of planned duration has passed and effect size is 5%, tracking failures >10%, stop and investigate
Step 8: Define Analysis Plan
- When to analyze: only after planned duration OR early termination trigger
- NEVER peek at results before planned duration (peeking inflates false positive rate)
- Statistical test:
- For proportions: two-proportion z-test or chi-squared test
- For continuous metrics: Welch's t-test or Mann-Whitney U
- For duration/time metrics: log-transform then t-test
- Report: effect size, confidence interval, p-value, practical significance
- Segmentation analysis: check if effect varies by user segment (new vs returning, plan tier, geography)
- Guardrail check: verify no guardrail metric degraded significantly
Step 9: Anti-Patterns
Document these to prevent common mistakes:
- Peeking: Checking results daily and stopping when "significant" -- inflates false positives to >20%
- Multiple testing: Testing 5 metrics without correction -- one will be "significant" by chance
- Survivorship bias: Only analyzing users who completed the flow, ignoring drop-offs
- Underpowered: Running experiments too short because "we need to move fast"
- History effects: External events (holidays, marketing campaigns) contaminating results
- Novelty effects: Short-term engagement bump from any change, not real improvement
Step 10: Validate and Save
- Save experiment spec to
artifacts/experiments/ - Run
./tools/artifact/validate.shto verify frontmatter - Link to parent PRD and RFC via
./tools/artifact/link.sh
Quality Checklist
- [ ] Hypothesis is clear and falsifiable
- [ ] Primary metric is defined with baseline rate
- [ ] MDE is specified and justified
- [ ] Sample size is calculated (not guessed)
- [ ] Duration is estimated with weekly cycle buffer
- [ ] Guardrail metrics are defined with degradation thresholds
- [ ] Early termination rules are documented
- [ ] Analysis plan specifies exact statistical test
- [ ] Anti-patterns section is included
- [ ] Experiment spec has valid artifact frontmatter
Cross-References
- feature-flags skill: Experiment-type flags from the feature-flags skill should have a corresponding experiment spec from this skill. The flag controls assignment; this skill designs the experiment.
- activation-onboarding skill: Activation experiments designed in the activation skill should use this framework for statistical rigor.
- instrumentation skill: Event tracking must be in place before an experiment can run. Verify tracking with the instrumentation skill's event taxonomy.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: vibbs
- Source: vibbs/company-os
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.