Install
$ agentstack add skill-votee-ai-magic-data-agent-skills-magic-statistical-analysis ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ● Filesystem access Used
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
Natural Language Triggers
Activate this skill when the user says things like:
- "run a statistical test" / "compare these groups"
- "is this difference significant?" / "test the hypothesis"
- "compute correlations" / "what's the correlation?"
- "descriptive statistics" / "summarize the numbers"
These produce the SAME behavior as the statistical analysis workflow.
When to Use
- Need descriptive statistics with narrative interpretation
- Need hypothesis testing (group comparisons)
- Need correlation analysis with significance
- After magic-data-profiling or magic-data-cleaning, before reporting
- Results naturally feed into
magic-report-generationfor structured deliverables, ormagic-data-visualizationfor charts illustrating the findings
When NOT to Use: Use magic-data-profiling for initial exploration; use magic-data-exploration for pattern discovery.
Data Processing Expertise
Thinking
Before running any statistical test, ask:
- Am I exploring or confirming? — Exploratory analysis generates hypotheses; confirmatory analysis tests them. Using the same data for both inflates false discovery rates. If you explored the data to find a pattern and now want to test it, you need a held-out sample or an explicit acknowledgment that this is post-hoc.
- What is the minimum meaningful effect size? — A p-value 2 groups + normal = one-way ANOVA; >2 groups + non-normal = Kruskal-Wallis; both categorical = chi-square.
- Normality testing before parametric tests: Always check normality with the Shapiro-Wilk test before choosing a parametric test. If p dict:
cols = columns or df.select_dtypes(include=[np.number]).columns.tolist() stats = {} for col in cols: s = df[col].dropna() stats[col] = { "n": len(s), "mean": float(s.mean()), "std": float(s.std()), "median": float(s.median()), "skewness": float(s.skew()), "p25": float(s.quantile(0.25)), "p75": float(s.quantile(0.75)), } # Flag distribution shape stats[col]["shape"] = ( "symmetric" if abs(stats[col]["skewness"]) 0 else "left-skewed" ) return stats
### Hypothesis test selection and execution
```python
from scipy import stats as sp
import numpy as np
def auto_test(groups: list[np.ndarray], alpha: float = 0.05) -> dict:
# Check normality for each group
all_normal = all(sp.shapiro(g)[1] > 0.05 for g in groups if len(g) >= 3)
if len(groups) == 2:
g1, g2 = groups
if all_normal:
stat, p = sp.ttest_ind(g1, g2)
pooled_std = np.sqrt(((len(g1)-1)*g1.var(ddof=1) + (len(g2)-1)*g2.var(ddof=1)) / (len(g1)+len(g2)-2))
d = (g1.mean() - g2.mean()) / pooled_std if pooled_std else 0
return {"test": "t-test", "p": p, "effect_size": ("Cohen's d", d)}
else:
stat, p = sp.mannwhitneyu(g1, g2, alternative="two-sided")
r = 1 - (2 * stat) / (len(g1) * len(g2))
return {"test": "Mann-Whitney U", "p": p, "effect_size": ("rank-biserial r", r)}
else: # 3+ groups
if all_normal:
stat, p = sp.f_oneway(*groups)
grand = np.concatenate(groups)
ss_b = sum(len(g)*(g.mean()-grand.mean())**2 for g in groups)
eta2 = ss_b / ((grand - grand.mean())**2).sum()
return {"test": "ANOVA", "p": p, "effect_size": ("eta-squared", eta2)}
else:
stat, p = sp.kruskal(*groups)
return {"test": "Kruskal-Wallis", "p": p, "statistic": stat}
Correlation analysis with significance
from scipy import stats as sp
import pandas as pd, numpy as np
def correlate(df: pd.DataFrame, method: str = "auto", alpha: float = 0.05) -> list[dict]:
num_cols = df.select_dtypes(include=[np.number]).columns
all_normal = all(sp.shapiro(df[c].dropna())[1] > 0.05 for c in num_cols if len(df[c].dropna()) >= 3)
use = "pearson" if (method == "auto" and all_normal) or method == "pearson" else "spearman"
compute = sp.pearsonr if use == "pearson" else sp.spearmanr
results = []
for i, a in enumerate(num_cols):
for b in num_cols[i+1:]:
mask = df[a].notna() & df[b].notna()
if mask.sum() list[dict]:
adjusted_alpha = alpha / len(p_values)
return [{"p": p, "adjusted_alpha": adjusted_alpha,
"significant": p list[dict]:
n = len(p_values)
ranked = sorted(enumerate(p_values), key=lambda x: x[1])
results = [None] * n
for rank, (idx, p) in enumerate(ranked, 1):
threshold = (rank / n) * alpha
results[idx] = {"p": p, "rank": rank, "threshold": threshold,
"significant": p Scripts fall into three categories: **callable tools** (call directly via CLI),
> **scriptable tools** (call directly for standard use, or read + adapt for advanced use),
> and **reference implementations** (always read + adapt).
**Scriptable tools** -- call directly for standard use, read + adapt for advanced:
| Script | Tier | Standard CLI usage | When to customize |
|--------|------|-------------------|-------------------|
| `descriptive_stats.py` | A | `python3 descriptive_stats.py --input data.csv --output stats.json` | `--columns col1,col2` to restrict; `--explain` for verbose narrative |
| `correlation_analysis.py` | A | `python3 correlation_analysis.py --input data.csv --output corr.json` | `--method pearson\|spearman\|kendall` to override auto; `--columns` to restrict; `--min_significance` to adjust |
| `hypothesis_test.py` | B | `python3 hypothesis_test.py --input data.csv --output test.json --group_col region --value_col revenue` | `--group_col` and `--value_col` functionally required; `--test` to override auto; `--explain` for narrative |
**Reference implementations** -- read patterns, write custom code:
*(none — all statistics scripts are now scriptable)*
Scripts accept `--input` / `--output` named flags and `--input-format` (auto/csv/tsv/jsonl/json/parquet/excel).
## Checkpointing
Checkpointing is a judgment call — save intermediate results when the cost of re-running the preceding step justifies persisting the result.
**For statistical analysis specifically:**
| Signal | Checkpoint? | Reasoning |
|--------|-------------|-----------|
| After hypothesis testing | **Always** | Statistical results ARE the deliverable |
| After correlation analysis | **Yes** | Expensive on large datasets (pairwise computation) |
| After descriptive stats | **Optional** | Cheap to re-compute; save if feeding into a report |
| After regression/modeling | **Always** | Model fitting is expensive; coefficients are the output |
**Suggested names:** `stats_descriptive.json`, `stats_hypothesis.json`, `stats_correlation.json`, `stats_regression.json`
Include provenance metadata: input file, test parameters, sample sizes, alpha level, correction method.
**Save pattern:**
```python
import json, datetime
from pathlib import Path
def save_checkpoint(df, path: str, metadata: dict = None) -> str:
"""Save data + provenance metadata. Use at phase boundaries."""
p = Path(path)
p.parent.mkdir(parents=True, exist_ok=True)
if p.suffix == ".parquet":
df.to_parquet(p, index=False)
elif p.suffix == ".jsonl":
df.to_json(p, orient="records", lines=True, force_ascii=False)
else:
df.to_csv(p, index=False)
meta = {
"rows": len(df), "cols": list(df.columns),
"saved_at": datetime.datetime.utcnow().isoformat(),
**(metadata or {}),
}
p.with_suffix(".meta.json").write_text(
json.dumps(meta, indent=2, ensure_ascii=False)
)
return str(p)
Self-Healing
If analysis fails:
- Read error JSON — check
suggestionfield - Read script source to understand the failure mode
- Common fixes below
| Error | Likely Cause | Fix | |-------|-------------|-----| | Too few samples for test | Small group size | Use non-parametric test or note limitation | | Non-numeric group column | Wrong column specified | Check column types, use categorical column for grouping | | Constant column | Zero variance | Skip column, note in warnings | | Normality violated + small sample | Shapiro-Wilk p 0.9 between predictors | Consider dropping one of the correlated variables before modeling |
Reference Guides
| Topic | File | Load When | |-------|------|-----------| | Test selection | references/test_selection_guide.md | Choosing the right statistical test | | Interpretation | references/interpretation_guide.md | Writing statistical narratives | | Pitfalls | references/statistical_pitfalls.md | Results seem unexpected or contradictory |
Interactive Mode [Optional]
For agents working interactively with a user. Pipeline code generation skips this section.
PAUSE Gates
- If statistically significant findings are detected, present them with effect sizes, confidence intervals, and caveats. Include: test used, p-value, effect size interpretation, and sample size context. Ask the user which findings to highlight in reporting before proceeding.
Workspace Integration
- Update
workspace_state.mdwith analysis results. - Log statistical decisions in
logs/analysis_journal.md.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: Votee-AI
- Source: Votee-AI/magic-data-agent-skills
- License: Apache-2.0
- Homepage: https://docs.votee.ai/magic-agent-skills/data-agent/getting-started
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.