Install
$ agentstack add skill-trecek-useful-claude-skills-exp-lens-benchmark-representativeness ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
Benchmark Representativeness Experimental Design Lens
Philosophical Mode: Generalizability Primary Question: "Does this generalize beyond the test bed?" Focus: Task Distribution, Scenario Coverage, Missing Regions, Dataset Selection, Generalization Claims
When to Use
- Evaluating claims that extend beyond specific benchmarks
- Checking coverage of evaluation suite
- Assessing dataset diversity
- User invokes
/exp-lens-benchmark-representativenessor/make-experiment-diag benchmark
Critical Constraints
NEVER:
- Modify any source code files
- Do not litter the codebase with useless comments, TODO markers, or explanatory annotations — the skill output and diagram speak for themselves
ALWAYS:
- Focus on GENERALIZATION GAP between benchmark coverage and claimed scope
- Show which regions of the target space are untested
- Document the relationship between benchmark selection and generalization claims
- Include a coverage matrix mapping scenarios to metrics
- BEFORE creating any diagram, LOAD the
/mermaidskill using the Skill tool - this is MANDATORY
Analysis Workflow
Step 1: Launch Parallel Exploration Subagents
Spawn Explore subagents to investigate:
Benchmark & Dataset Inventory
- Find all datasets, benchmarks, test suites used
- Look for:
benchmark,dataset,test_suite,eval,corpus,split,GLUE,ImageNet
Task & Scenario Coverage
- Find what scenarios, conditions, and domains are tested
- Look for:
task,scenario,domain,category,difficulty,subset
Metric Coverage
- Find all evaluation metrics used
- Look for:
metric,accuracy,f1,bleu,rouge,latency,cost,fairness
Claimed Generalization Scope
- Find claims about generality in docs, papers, READMEs
- Look for:
generalize,real-world,production,deploy,robust,transfer,domain
Distribution Characteristics
- Find data distribution analysis, class balance, domain stats
- Look for:
distribution,balance,skew,size,demographics,diversity
Step 2: Build the Coverage Matrix
Build the coverage matrix: rows = scenarios/domains tested, columns = metrics measured. Identify which cells are populated and which are gaps. Compare the coverage to the stated generalization claims.
Step 3: CRITICAL — Analyze Generalization Gap
For every generalization claim:
- Target population: What is the full population the claim extends to?
- Benchmark representation: What subset of that population is represented in the benchmark?
- Untested regions: What regions of the space are untested?
- Coverage ratio: Is the coverage sufficient to support the claim?
Distinguish clearly:
- Strong claims (e.g., "production-ready"): require broad, diverse coverage
- Scoped claims (e.g., "best on GLUE"): only require benchmark-specific coverage
- Implicit claims: claims made in framing but not stated explicitly
Step 4: Create the Diagram
Use flowchart with:
Direction: TB (claims flow from benchmarks up to generalization)
Subgraphs:
BENCHMARKS USEDSCENARIOS TESTEDMETRICS MEASUREDGENERALIZATION CLAIMSUNTESTED REGIONS
Node Styling:
stateNodeclass: benchmarks/datasetshandlerclass: tested scenariosoutputclass: measured metricscliclass: generalization claimsgapclass: untested regions/missing coveragedetectorclass: validation of generalization
Step 5: Write Output
Write the diagram to: temp/exp-lens-benchmark-representativeness/exp_diag_benchmark_representativeness_{YYYY-MM-DD_HHMMSS}.md
Output Template
# Benchmark Representativeness Diagram: {System Name}
**Lens:** Benchmark Representativeness (Generalizability)
**Question:** Does this generalize beyond the test bed?
**Date:** {YYYY-MM-DD}
**Scope:** {What was analyzed}
## Coverage Matrix
| Scenario / Domain | {Metric A} | {Metric B} | {Metric C} | Coverage |
|-------------------|-----------|-----------|-----------|----------|
| {Scenario 1} | ✓ | ✓ | ✗ | Partial |
| {Scenario 2} | ✗ | ✗ | ✗ | None |
| {Scenario 3} | ✓ | ✓ | ✓ | Full |
## Benchmark Representativeness Diagram
```mermaid
%%{init: {'flowchart': {'nodeSpacing': 50, 'rankSpacing': 60, 'curve': 'basis'}}}%%
flowchart TB
%% CLASS DEFINITIONS %%
classDef cli fill:#1a237e,stroke:#7986cb,stroke-width:2px,color:#fff;
classDef stateNode fill:#004d40,stroke:#4db6ac,stroke-width:2px,color:#fff;
classDef handler fill:#e65100,stroke:#ffb74d,stroke-width:2px,color:#fff;
classDef phase fill:#6a1b9a,stroke:#ba68c8,stroke-width:2px,color:#fff;
classDef newComponent fill:#2e7d32,stroke:#81c784,stroke-width:2px,color:#fff;
classDef output fill:#00695c,stroke:#4db6ac,stroke-width:2px,color:#fff;
classDef detector fill:#b71c1c,stroke:#ef5350,stroke-width:2px,color:#fff;
classDef gap fill:#ff6f00,stroke:#ffa726,stroke-width:2px,color:#000;
classDef integration fill:#c62828,stroke:#ef9a9a,stroke-width:2px,color:#fff;
subgraph Benchmarks ["BENCHMARKS USED"]
direction TB
B1["{Benchmark 1}━━━━━━━━━━{size}, {domain}"]
B2["{Benchmark 2}━━━━━━━━━━{size}, {domain}"]
end
subgraph Scenarios ["SCENARIOS TESTED"]
direction TB
S1["{Scenario 1}━━━━━━━━━━{conditions}"]
S2["{Scenario 2}━━━━━━━━━━{conditions}"]
end
subgraph Metrics ["METRICS MEASURED"]
direction TB
M1["{Metric A}━━━━━━━━━━{what it captures}"]
M2["{Metric B}━━━━━━━━━━{what it captures}"]
end
subgraph Claims ["GENERALIZATION CLAIMS"]
direction TB
C1["{Claim 1}━━━━━━━━━━{source of claim}"]
C2["{Claim 2}━━━━━━━━━━{source of claim}"]
end
subgraph Gaps ["UNTESTED REGIONS"]
direction TB
G1["{Gap 1}━━━━━━━━━━{why it matters}"]
G2["{Gap 2}━━━━━━━━━━{why it matters}"]
end
VALIDATE["{Generalization Validity Check}━━━━━━━━━━Coverage vs. Claim scope"]
B1 --> S1
B2 --> S2
S1 --> M1
S2 --> M2
M1 --> C1
M2 --> C2
C1 --> VALIDATE
C2 --> VALIDATE
G1 -.->|missing| VALIDATE
G2 -.->|missing| VALIDATE
%% CLASS ASSIGNMENTS %%
class B1,B2 stateNode;
class S1,S2 handler;
class M1,M2 output;
class C1,C2 cli;
class G1,G2 gap;
class VALIDATE detector;
Color Legend: | Color | Category | Description | |-------|----------|-------------| | Dark Teal | Benchmarks | Datasets and test suites used | | Orange | Scenarios | Tested scenarios and conditions | | Teal | Metrics | Measured evaluation metrics | | Dark Blue | Claims | Generalization claims made | | Yellow/Amber | Gaps | Untested regions of target space | | Red | Validation | Generalization validity check |
Generalization Gap Analysis
| Claim | Evidence (Benchmarks) | Gap (Untested) | Risk | |-------|----------------------|----------------|------| | {Claim 1} | {What covers it} | {What is missing} | High/Med/Low | | {Claim 2} | {What covers it} | {What is missing} | High/Med/Low |
Representativeness Assessment
| Dimension | Current Coverage | Required for Claim | Verdict | |-----------|-----------------|-------------------|---------| | Domain diversity | {count} domains | {needed} | ✓/✗ | | Task variety | {count} tasks | {needed} | ✓/✗ | | Scale range | {min}–{max} | {needed} | ✓/✗ | | Distribution shift | {tested?} | {needed} | ✓/✗ |
---
## Pre-Diagram Checklist
Before creating the diagram, verify:
- [ ] LOADED `/mermaid` skill using the Skill tool
- [ ] Using ONLY classDef styles from the mermaid skill (no invented colors)
- [ ] Diagram will include a color legend table
---
## Related Skills
- `/make-experiment-diag` - Parent skill for lens selection
- `/mermaid` - MUST BE LOADED before creating diagram
- `/exp-lens-measurement-validity` - For metric quality analysis
- `/exp-lens-validity-threats` - For systematic threat inventory
## Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- **Author:** [Trecek](https://github.com/Trecek)
- **Source:** [Trecek/useful-claude-skills](https://github.com/Trecek/useful-claude-skills)
- **License:** MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.