Install
$ agentstack add skill-ericwang915-data-scientist-skills-eda-profile ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
EDA Profile
Purpose
Generate a comprehensive exploratory data analysis profile in a single pass. Provides the complete picture of a dataset's structure, quality, and statistical properties.
How It Works
Step 1: Structure Overview
- Shape (rows × columns), memory usage
- Column names, data types, and inferred semantic types
- Identify: numeric, categorical, datetime, boolean, text, ID columns
- Sample rows (first, last, random)
Step 2: Univariate Analysis
For each column based on type:
- Numeric: mean, median, std, min, max, quartiles, skewness, kurtosis, histogram
- Categorical: unique count, top categories, frequency distribution, bar chart
- DateTime: range, gaps, frequency, time distribution
- Boolean: true/false ratio
- Text: length distribution, word count, common patterns
Step 3: Missing Data Profile
- Missing count and percentage per column
- Missing data patterns (which columns are missing together)
- Missingness mechanism hypothesis (MCAR/MAR/MNAR)
Step 4: Bivariate Analysis
- Correlation matrix (Pearson for numeric, Cramér's V for categorical)
- Top correlated pairs highlighted
- Potential multicollinearity flags (|r| > 0.8)
- Target variable relationships (if specified)
Step 5: Data Quality Flags
- Constant or near-constant columns
- High cardinality categoricals (potential ID columns)
- Potential data leakage indicators
- Suspicious distributions (uniform IDs that should be sequential, etc.)
Step 6: Key Insights Summary
- Top 5 findings that need attention
- Recommended next steps (cleaning, transformation, modeling)
Usage Examples
Example 1: New dataset
"Profile this dataset — I just received it and don't know what's in it"
Example 2: Pre-modeling
"Run EDA on this dataset before I build a prediction model.
The target variable is 'churn'."
Output Format
- Structure Summary: Table of columns with types, nulls, unique counts
- Statistical Profile: Descriptive statistics per column
- Visualizations: Histograms, bar charts, correlation heatmap
- Quality Flags: Issues ranked by severity
- Key Insights: Top findings and recommended actions
- Python Code: Reproducible profiling script
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: ericwang915
- Source: ericwang915/data-scientist-skills
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.