Install
$ agentstack add skill-ericwang915-data-scientist-skills-dimensionality-reduction ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Dimensionality Reduction
Purpose
Reduce the number of features while preserving important structure. Essential for visualization, denoising, and preprocessing.
How It Works
Method Selection
| Method | Preserves | Best For | Linear? | |--------|-----------|----------|---------| | PCA | Global variance | Feature compression, denoising | Yes | | t-SNE | Local structure | 2D/3D visualization | No | | UMAP | Local + global | Visualization, clustering prep | No | | SVD | Variance | Sparse data, NLP (LSA) | Yes | | LDA | Class separation | Supervised dimensionality reduction | Yes | | Autoencoder | Learned representation | Complex non-linear compression | No |
PCA Workflow
- Standardize features
- Compute covariance matrix and eigenvalues
- Choose components: explained variance ≥ 85-95%
- Transform and validate (scree plot, biplot)
t-SNE / UMAP Workflow
- Apply PCA first if >50 features (speed)
- Tune perplexity (t-SNE) or n_neighbors (UMAP)
- Generate 2D/3D embedding
- Color by labels or clusters for interpretation
Usage Examples
"Visualize this 50-feature customer dataset in 2D to see if
natural clusters exist"
"Reduce 200 features to the most important 20 using PCA
before training a model"
Output Format
- Method Choice: Rationale for selected approach
- Explained Variance: Scree plot, cumulative variance
- Visualization: 2D/3D scatter plot of reduced space
- Python Code: sklearn / umap-learn implementation
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: ericwang915
- Source: ericwang915/data-scientist-skills
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.