# Exp Lens Variance Stability

> Create a variance analysis profile assessing whether signals exceed noise and whether results are stable across random seeds. Stability lens answering "Is the signal larger than the noise?

- **Type:** Skill
- **Install:** `agentstack add skill-trecek-useful-claude-skills-exp-lens-variance-stability`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [Trecek](https://agentstack.voostack.com/s/trecek)
- **Installs:** 0
- **Category:** [Agent Skills](https://agentstack.voostack.com/c/agent-skills)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [Trecek](https://github.com/Trecek)
- **Source:** https://github.com/Trecek/useful-claude-skills/tree/main/.claude/skills/exp-lens-variance-stability

## Install

```sh
agentstack add skill-trecek-useful-claude-skills-exp-lens-variance-stability
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# Variance Stability Experimental Design Lens

**Philosophical Mode:** Stability
**Primary Question:** "Is the signal larger than the noise?"
**Focus:** Run-to-Run Variance, Seed Sensitivity, Nondeterminism Sources, Confidence Intervals, Noise Floor

## When to Use

- ML experiments with high variance across seeds
- Systems benchmarks with environmental noise
- Results where small differences are claimed as improvements
- User invokes `/exp-lens-variance-stability` or `/make-experiment-diag variance`

## Critical Constraints

**NEVER:**
- Modify any source code files
- Do not litter the codebase with useless comments, TODO markers, or explanatory annotations — the skill output and diagram speak for themselves

**ALWAYS:**
- Count the actual number of independent runs — single-run results must be flagged prominently
- Assess whether claimed improvements exceed the observed standard deviation
- Identify all sources of nondeterminism, not just random seeds
- Report when confidence intervals are absent — absence is a finding, not an omission
- BEFORE creating any diagram, LOAD the `/mermaid` skill using the Skill tool - this is MANDATORY

---

## Analysis Workflow

### Step 1: Launch Parallel Exploration Subagents

Spawn Explore subagents to investigate:

**Random Seed Management**
- Find how seeds are set and varied
- Look for: seed, random_state, torch.manual_seed, np.random, set_seed, PYTHONHASHSEED

**Nondeterminism Sources**
- Find sources of nondeterminism beyond seeds
- Look for: cudnn, benchmark, deterministic, parallel, async, thread, race, order, nondeterministic

**Multiple Run Protocol**
- Find how many runs are performed per condition
- Look for: n_runs, trials, repeat, replicate, mean, std, confidence, interval, aggregate

**Variance Reporting**
- Find how variance is reported (if at all)
- Look for: std, stderr, confidence, interval, range, median, quartile, bootstrap, error_bar

**Signal-to-Noise Assessment**
- Find whether claimed improvements exceed observed variance
- Look for: significant, difference, improvement, margin, effect_size, gap, overlap

### Step 2: Build Variance Profile

For each reported result:
1. How many independent runs?
2. What is the standard deviation across runs?
3. Does the claimed improvement exceed the noise floor?
4. Are confidence intervals reported?
5. What sources of nondeterminism exist beyond seeds?

Build the variance profile.

### Step 3: Analyze Signal vs Noise

**CRITICAL — Analyze Signal vs Noise:**
For every claimed improvement:
- Is the improvement magnitude larger than the run-to-run standard deviation?
- Could the ranking of methods change under reruns?
- Are the "best" results cherry-picked from multiple seeds?

### Step 4: Create the Diagram

Use the mermaid skill conventions to create a stochasticity diagram with:

**Direction:** `TB` (nondeterminism sources flow down through aggregation to reported results)

**Subgraphs:**
- NONDETERMINISM SOURCES
- VARIANCE AGGREGATION
- REPORTED RESULTS

**Node Styling:**
- `stateNode` class: Nondeterminism sources
- `handler` class: Aggregation methods
- `output` class: Reported results
- `gap` class: Unreported variance or single-run results
- `detector` class: Confidence intervals and statistical tests
- `cli` class: Seed management

### Step 5: Write Output

Write the diagram to: `temp/exp-lens-variance-stability/exp_diag_variance_stability_{YYYY-MM-DD_HHMMSS}.md`

---

## Output Template

```markdown
# Variance Stability Analysis: {Experiment Name}

**Lens:** Variance Stability (Stability)
**Question:** Is the signal larger than the noise?
**Date:** {YYYY-MM-DD}
**Scope:** {What was analyzed}

## Variance Profile

| Experiment | N Runs | Mean | Std | CI | Signal > Noise? |
|------------|--------|------|-----|----|-----------------|
| {experiment} | {n} | {mean} | {std} | {CI or "Not reported"} | {Yes/No/Unclear} |

## Nondeterminism Inventory

| Source | Type | Controlled? | Impact |
|--------|------|-------------|--------|
| {source} | {seed/hardware/async/etc} | {Yes/No/Partial} | {Low/Medium/High} |

## Stochasticity Diagram

```mermaid
%%{init: {'flowchart': {'nodeSpacing': 50, 'rankSpacing': 60, 'curve': 'basis'}}}%%
graph TB
    %% CLASS DEFINITIONS %%
    classDef cli fill:#1a237e,stroke:#7986cb,stroke-width:2px,color:#fff;
    classDef stateNode fill:#004d40,stroke:#4db6ac,stroke-width:2px,color:#fff;
    classDef handler fill:#e65100,stroke:#ffb74d,stroke-width:2px,color:#fff;
    classDef phase fill:#6a1b9a,stroke:#ba68c8,stroke-width:2px,color:#fff;
    classDef newComponent fill:#2e7d32,stroke:#81c784,stroke-width:2px,color:#fff;
    classDef output fill:#00695c,stroke:#4db6ac,stroke-width:2px,color:#fff;
    classDef detector fill:#b71c1c,stroke:#ef5350,stroke-width:2px,color:#fff;
    classDef gap fill:#ff6f00,stroke:#ffa726,stroke-width:2px,color:#000;
    classDef integration fill:#c62828,stroke:#ef9a9a,stroke-width:2px,color:#fff;

    subgraph NDSources ["NONDETERMINISM SOURCES"]
        direction TB
        SEED["Random Seed━━━━━━━━━━Controlled viaseed management"]
        HW["Hardware Variance━━━━━━━━━━GPU/CPU orderingdifferences"]
        ASYNC["Async Operations━━━━━━━━━━Thread/processrace conditions"]
    end

    subgraph Aggregation ["VARIANCE AGGREGATION"]
        direction TB
        MULTI["Multiple Runs━━━━━━━━━━N independentrepetitions"]
        CI["Confidence Interval━━━━━━━━━━Statistical boundson estimates"]
    end

    subgraph Results ["REPORTED RESULTS"]
        direction TB
        RESULT["Reported Result━━━━━━━━━━Mean ± stdwith CI"]
        SINGLE["Single-Run Result━━━━━━━━━━No variancereported"]
    end

    SEED --> MULTI
    HW --> MULTI
    ASYNC --> MULTI
    MULTI --> CI
    CI --> RESULT
    MULTI --> SINGLE

    %% CLASS ASSIGNMENTS %%
    class SEED cli;
    class HW,ASYNC stateNode;
    class MULTI handler;
    class CI detector;
    class RESULT output;
    class SINGLE gap;
```

## Seed Sensitivity Assessment

| Seed | Run Result | Rank Among Methods |
|------|-----------|-------------------|
| {seed} | {result} | {rank} |

## Reporting Completeness Checklist

- [ ] Number of runs reported per condition
- [ ] Standard deviation or standard error reported
- [ ] Confidence intervals reported
- [ ] All seeds or seed range disclosed
- [ ] Nondeterminism sources acknowledged

## Key Findings

- {Description of whether signals exceed noise and reporting completeness}
```

---

## Pre-Diagram Checklist

Before creating the diagram, verify:

- [ ] LOADED `/mermaid` skill using the Skill tool
- [ ] Using ONLY classDef styles from the mermaid skill (no invented colors)
- [ ] Diagram will include a color legend table

---

## Related Skills

- `/make-experiment-diag` - Parent skill for lens selection
- `/mermaid` - MUST BE LOADED before creating diagram
- `/exp-lens-reproducibility-artifacts` - For environment and artifact reproducibility
- `/exp-lens-error-budget` - For systematic error and bias analysis

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [Trecek](https://github.com/Trecek)
- **Source:** [Trecek/useful-claude-skills](https://github.com/Trecek/useful-claude-skills)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-trecek-useful-claude-skills-exp-lens-variance-stability
- Seller: https://agentstack.voostack.com/s/trecek
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
