Install
$ agentstack add skill-ericwang915-data-scientist-skills-nlp-pipeline ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
NLP Pipeline
Purpose
Build end-to-end natural language processing pipelines, from raw text to model predictions.
How It Works
Step 1: Text Preprocessing
- Lowercasing, Unicode normalization
- Tokenization (word, subword, sentence)
- Stop word removal (task-dependent — keep for BERT)
- Lemmatization / stemming
- Special character and HTML tag removal
Step 2: Text Representation
| Method | When to Use | Library | |--------|-------------|---------| | Bag of Words | Simple baselines, small data | sklearn | | TF-IDF | Document classification, search | sklearn | | Word2Vec / GloVe | Word-level tasks, small models | gensim | | FastText | Morphologically rich languages | gensim | | BERT embeddings | State-of-the-art, contextual | transformers | | Sentence-transformers | Semantic similarity, search | sentence-transformers |
Step 3: NLP Tasks
| Task | Approach | |------|----------| | Sentiment Analysis | Fine-tuned BERT, or TF-IDF + logistic regression | | Text Classification | Fine-tuned transformer, or TF-IDF + SVM | | NER | spaCy, fine-tuned BERT for token classification | | Topic Modeling | LDA, BERTopic, NMF | | Summarization | T5, BART, or extractive methods | | Question Answering | RAG, fine-tuned QA models |
Step 4: Evaluate
- Classification: accuracy, F1, confusion matrix
- NER: entity-level precision, recall, F1
- Topic models: coherence score, human evaluation
- Embeddings: downstream task performance, nearest neighbor quality
Usage Examples
"Build a sentiment analysis pipeline for product reviews —
should I fine-tune BERT or use TF-IDF + logistic regression?"
"Extract company names and locations from these news articles"
Output Format
- Pipeline Design: Step-by-step architecture
- Model Selection: Chosen approach with rationale
- Python Code: Complete pipeline (spaCy, transformers, sklearn)
- Evaluation: Metrics on validation set
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: ericwang915
- Source: ericwang915/data-scientist-skills
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.