# Exp Lens Sensitivity Robustness

> Create Sensitivity & Robustness experimental design analysis identifying load-bearing analytic choices and untested perturbations. Robustness lens answering "Which assumptions are load-bearing?

- **Type:** Skill
- **Install:** `agentstack add skill-trecek-useful-claude-skills-exp-lens-sensitivity-robustness`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [Trecek](https://agentstack.voostack.com/s/trecek)
- **Installs:** 0
- **Category:** [Agent Skills](https://agentstack.voostack.com/c/agent-skills)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [Trecek](https://github.com/Trecek)
- **Source:** https://github.com/Trecek/useful-claude-skills/tree/main/.claude/skills/exp-lens-sensitivity-robustness

## Install

```sh
agentstack add skill-trecek-useful-claude-skills-exp-lens-sensitivity-robustness
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# Sensitivity & Robustness Experimental Design Lens

**Philosophical Mode:** Robustness
**Primary Question:** "Which assumptions are load-bearing?"
**Focus:** Ablation Structure, Preprocessing Sensitivity, Metric Sensitivity, Hyperparameter Sensitivity, Distribution Shift

## When to Use

- Results may depend on specific preprocessing choices
- Need to verify robustness of conclusions across conditions
- Ablation study seems incomplete or cherry-picked
- User invokes `/exp-lens-sensitivity-robustness` or `/make-experiment-diag sensitivity`

## Critical Constraints

**NEVER:**
- Modify any source code or experiment files
- Do not litter the codebase with useless comments, TODO markers, or explanatory annotations — the skill output and diagram speak for themselves
- Treat "untested" as equivalent to "robust"

**ALWAYS:**
- Build a full sensitivity matrix (choices x perturbation types)
- Classify every analytic choice as load-bearing, minor, or untested
- Flag cases where the most impactful choices are the least tested
- Distinguish between ablations that were run and choices that were simply fixed
- BEFORE creating any diagram, LOAD the `/mermaid` skill using the Skill tool - this is MANDATORY

---

## Analysis Workflow

### Step 1: Launch Parallel Exploration Subagents

Spawn Explore subagents to investigate:

**Analytic Choices Made**
- Find all decision points in the analysis pipeline
- Look for: choice, option, default, parameter, threshold, method, alternative

**Ablation Coverage**
- Find which factors have been ablated
- Look for: ablation, without, remove, disable, vary, sweep, drop

**Preprocessing Variations**
- Find preprocessing steps that could be done differently
- Look for: normalize, tokenize, augment, crop, resize, filter, clean, impute

**Hyperparameter Sensitivity**
- Find which hyperparameters were tuned vs fixed
- Look for: learning_rate, batch_size, epochs, dropout, hidden_size, temperature, alpha, beta

**Distribution/Environment Variations**
- Find evidence of testing under different conditions
- Look for: shift, domain, transfer, cross, out_of_distribution, generalize, different

### Step 2: Build the Sensitivity Matrix

Rows = analytic choices. Columns = perturbation types (remove, change, stress).

For each cell: Does the conclusion survive the perturbation?

### Step 3: Classify Analytic Choices

**CRITICAL — Analyze Assumption Load:**
For every analytic choice:
- What happens if this choice were made differently?
- Is there evidence from ablations, sweeps, or prior literature?
- Are the most impactful choices the least tested?

Classify each choice as:
- **Load-bearing**: Conclusion changes if this choice changes
- **Minor**: Conclusion is robust to changes in this choice
- **Untested**: No evidence either way

### Step 4: Create the Optional Perturbation Diagram

If a diagram adds value, create a simplified flowchart. This is OPTIONAL for this hybrid lens — the tables are the primary output.

**Direction:** `TB` (choices flow down through perturbation to conclusion stability)

**Subgraphs:** "ANALYTIC CHOICES", "PERTURBATIONS TESTED", "CONCLUSION STABILITY"

**Node Styling:**
- `stateNode` class: analytic choice nodes
- `handler` class: perturbation type nodes
- `output` class: stable conclusion nodes
- `gap` class: load-bearing untested choice nodes
- `detector` class: sensitivity threshold nodes

### Step 5: Write Output

Write the analysis to: `temp/exp-lens-sensitivity-robustness/exp_diag_sensitivity_robustness_{YYYY-MM-DD_HHMMSS}.md`

---

## Output Template

```markdown
# Sensitivity & Robustness Analysis: {Experiment Name}

**Lens:** Sensitivity & Robustness (Robustness)
**Question:** Which assumptions are load-bearing?
**Date:** {YYYY-MM-DD}
**Scope:** {What was analyzed}

## Sensitivity Matrix

| Analytic Choice | Remove | Change Value | Stress Test | Overall Classification |
|----------------|--------|-------------|-------------|------------------------|
| {choice name} | Stable / Fragile / Untested | Stable / Fragile / Untested | Stable / Fragile / Untested | Load-bearing / Minor / Untested |

## Load-Bearing Assumptions

| Assumption | Evidence Type | Impact if Changed | Tested? |
|------------|--------------|-------------------|---------|
| {assumption} | Ablation / Sweep / Literature / None | High / Medium / Low | Yes / No |

## Ablation Coverage Assessment

| Factor | Ablated? | Result | Interpretation |
|--------|----------|--------|----------------|
| {factor name} | Yes / No | {delta metric if yes} | Conclusion holds / Fragile / Unknown |

## Robustness Profile

| Dimension | Status | Notes |
|-----------|--------|-------|
| Preprocessing choices | Robust / Fragile / Untested | {detail} |
| Hyperparameter choices | Robust / Fragile / Untested | {detail} |
| Metric choices | Robust / Fragile / Untested | {detail} |
| Distribution shift | Robust / Fragile / Untested | {detail} |

## Perturbation Diagram (Optional)

```mermaid
%%{init: {'flowchart': {'nodeSpacing': 50, 'rankSpacing': 60, 'curve': 'basis'}}}%%
flowchart TB
    %% CLASS DEFINITIONS %%
    classDef cli fill:#1a237e,stroke:#7986cb,stroke-width:2px,color:#fff;
    classDef stateNode fill:#004d40,stroke:#4db6ac,stroke-width:2px,color:#fff;
    classDef handler fill:#e65100,stroke:#ffb74d,stroke-width:2px,color:#fff;
    classDef phase fill:#6a1b9a,stroke:#ba68c8,stroke-width:2px,color:#fff;
    classDef newComponent fill:#2e7d32,stroke:#81c784,stroke-width:2px,color:#fff;
    classDef output fill:#00695c,stroke:#4db6ac,stroke-width:2px,color:#fff;
    classDef detector fill:#b71c1c,stroke:#ef5350,stroke-width:2px,color:#fff;
    classDef gap fill:#ff6f00,stroke:#ffa726,stroke-width:2px,color:#000;
    classDef integration fill:#c62828,stroke:#ef9a9a,stroke-width:2px,color:#fff;

    subgraph Choices ["ANALYTIC CHOICES"]
        CHOICE1["Preprocessing Choice━━━━━━━━━━{e.g., normalization}"]
        CHOICE2["Model Choice━━━━━━━━━━{e.g., architecture}"]
        LOAD_BEARING["Load-Bearing Choice━━━━━━━━━━{untested critical choice}"]
    end

    subgraph Perturbations ["PERTURBATIONS TESTED"]
        PERTURB1["Remove━━━━━━━━━━Ablation"]
        PERTURB2["Change Value━━━━━━━━━━Sensitivity sweep"]
        THRESHOLD["Sensitivity Threshold━━━━━━━━━━{delta that changes conclusion}"]
    end

    subgraph Stability ["CONCLUSION STABILITY"]
        STABLE["Stable Conclusion━━━━━━━━━━{holds under perturbation}"]
    end

    CHOICE1 --> PERTURB1
    CHOICE2 --> PERTURB2
    PERTURB1 --> THRESHOLD
    PERTURB2 --> THRESHOLD
    THRESHOLD --> STABLE
    LOAD_BEARING -.->|"untested risk"| THRESHOLD

    class CHOICE1,CHOICE2 stateNode;
    class LOAD_BEARING gap;
    class PERTURB1,PERTURB2 handler;
    class THRESHOLD detector;
    class STABLE output;
```

**Color Legend:**
| Color | Category | Description |
|-------|----------|-------------|
| Teal | Analytic Choices | Decision points in the pipeline |
| Yellow | Load-Bearing Untested | Critical choices with no perturbation evidence |
| Orange | Perturbations | Types of tests applied |
| Red | Sensitivity Thresholds | Points where conclusion may change |
| Dark Teal | Stable Conclusions | Results robust to perturbation |

## Recommendations

1. {Most urgent ablation to run — highest impact untested choice}
2. {Preprocessing sensitivity test needed}
3. {Distribution shift or domain generalization test needed}
```

---

## Pre-Diagram Checklist

Before creating the diagram, verify:

- [ ] LOADED `/mermaid` skill using the Skill tool
- [ ] Using ONLY classDef styles from the mermaid skill (no invented colors)
- [ ] Diagram will include a color legend table

---

## Related Skills

- `/make-experiment-diag` - Parent skill for lens selection
- `/mermaid` - MUST BE LOADED before creating diagram
- `/exp-lens-estimand-clarity` - For clarifying which conclusions are being stress-tested
- `/exp-lens-iterative-learning` - For tracking robustness improvements across experiment iterations

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [Trecek](https://github.com/Trecek)
- **Source:** [Trecek/useful-claude-skills](https://github.com/Trecek/useful-claude-skills)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-trecek-useful-claude-skills-exp-lens-sensitivity-robustness
- Seller: https://agentstack.voostack.com/s/trecek
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
