# Sklearn Pipelines

> Use when building scikit-learn models that must not leak preprocessing. Covers Pipeline, ColumnTransformer, custom transformers, and combining preprocessing with cross-validation correctly.

- **Type:** Skill
- **Install:** `agentstack add skill-param087-agent-ml-skills-sklearn-pipelines`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [param087](https://agentstack.voostack.com/s/param087)
- **Installs:** 0
- **Category:** [AI & ML](https://agentstack.voostack.com/c/ai-and-ml)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [param087](https://github.com/param087)
- **Source:** https://github.com/param087/agent-ml-skills/tree/main/skills/sklearn-pipelines

## Install

```sh
agentstack add skill-param087-agent-ml-skills-sklearn-pipelines
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# scikit-learn Pipelines

## Overview

A `Pipeline` chains preprocessing and the estimator into one object so that **every fit happens on training folds only**. This makes leakage structurally impossible and makes the model trivially serializable for serving. If you remember one thing from this pack: *wrap preprocessing in a Pipeline.*

## When to use

- Any sklearn model with preprocessing (scaling, encoding, imputing).
- You need cross-validation that includes preprocessing.
- You want one artifact to save and serve.

## Canonical pattern

```python
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.ensemble import HistGradientBoostingClassifier

num = ["age", "income", "tenure"]
cat = ["country", "plan"]

preprocess = ColumnTransformer([
    ("num", Pipeline([
        ("impute", SimpleImputer(strategy="median")),
        ("scale", StandardScaler()),
    ]), num),
    ("cat", Pipeline([
        ("impute", SimpleImputer(strategy="most_frequent")),
        ("ohe", OneHotEncoder(handle_unknown="ignore")),
    ]), cat),
])

model = Pipeline([
    ("prep", preprocess),
    ("clf", HistGradientBoostingClassifier(random_state=42)),
])

model.fit(X_train, y_train)        # all preprocessing fit on train only
preds = model.predict(X_test)      # preprocessing reused, no leakage
```

## Cross-validation the right way

```python
from sklearn.model_selection import cross_val_score, StratifiedKFold

cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
scores = cross_val_score(model, X, y, cv=cv, scoring="roc_auc")
# preprocessing is re-fit inside each fold automatically
```

Pair this with the `hyperparameter-tuning` skill — pass the whole pipeline to the search and tune with `clf__` / `prep__` prefixes.

## Custom transformer

```python
from sklearn.base import BaseEstimator, TransformerMixin

class LogTransform(BaseEstimator, TransformerMixin):
    def __init__(self, cols): self.cols = cols
    def fit(self, X, y=None): return self
    def transform(self, X):
        X = X.copy()
        X[self.cols] = np.log1p(X[self.cols])
        return X
```

## Pitfalls

- **`scaler.fit_transform(X)` before `train_test_split`** — the #1 leakage bug. Fit inside the pipeline instead.
- **`OneHotEncoder` without `handle_unknown="ignore"`** crashes on unseen test categories.
- **Imputing the target** — pipelines transform `X`, never `y`; impute/clean targets separately and deliberately.
- **Tuning preprocessing outside CV** — keep it in the pipeline so search respects fold boundaries.

## Hand-off

A single fitted `Pipeline` artifact that the model-evaluation, hyperparameter-tuning, and model-serving skills all consume directly (`joblib.dump(model, "model.joblib")`).

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [param087](https://github.com/param087)
- **Source:** [param087/agent-ml-skills](https://github.com/param087/agent-ml-skills)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-param087-agent-ml-skills-sklearn-pipelines
- Seller: https://agentstack.voostack.com/s/param087
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
