— No reviews yet
0 installs
13 views
0.0% view→install
Install
$ agentstack add skill-cheatthegod-biohermes-scrna-preprocessing-clustering ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Are you the author of Scrna Preprocessing Clustering? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claimAbout
scRNA Preprocessing And Clustering
Version Compatibility
Reference examples assume:
scanpy1.10+anndata0.10+pandas2.2+matplotlib3.8+
Before using code patterns, verify installed versions match the environment:
- Python:
python -c "import scanpy, anndata; print(scanpy.__version__, anndata.__version__)" - If signatures differ, inspect the installed API and adapt the pattern instead of retrying unchanged.
Overview
Use this skill to turn raw or minimally processed scRNA-seq data into an analysis-ready object with:
- QC-filtered cells and genes
- normalized expression values
- highly variable genes
- PCA and UMAP embeddings
- Leiden clusters
- saved
h5adartifact for annotation, DE, integration, or trajectory analysis
When To Use This Skill
- raw 10x matrices, filtered count matrices, or
h5adinputs need standard preprocessing - the user wants UMAP, clustering, or marker discovery
- downstream tasks depend on a stable single-cell object rather than ad hoc plots
Quick Route
- If the input is already a processed
h5ad, inspectadata.raw, embeddings, cluster columns, and QC columns before rerunning preprocessing. - If the input is raw counts, do QC first and only normalize after filtering obvious low-quality cells.
- If multiple batches are present, preprocess cleanly first, then consider integration instead of hiding batch effects with aggressive filtering.
Progressive Disclosure
- Read [technicalreference.md](technicalreference.md) for QC decision rules, assay caveats, and integration branching.
- Read [commandsandthresholds.md](commandsandthresholds.md) for concrete Scanpy code, default thresholds, and output conventions.
Default Rules
- Keep raw counts recoverable. Prefer
adata.raw = adata.copy()before regression or scaling. - Report thresholds explicitly. Do not silently drop cells or genes.
- Show QC distributions before applying hard filters.
- Use vector outputs such as
.pdfor.svgfor final figures when possible.
Expected Inputs
- 10x directory,
.h5,.h5ad, or count matrix - cell metadata if available
- species context for mitochondrial or ribosomal gene detection
Expected Outputs
results/processed.h5adqc/cell_qc_metrics.tsvqc/gene_qc_metrics.tsvfigures/qc_violin.pdffigures/pca_variance_ratio.pdffigures/umap_leiden.pdf
Preferred Tools
scanpyanndatapandasmatplotlibseaborn
Starter Pattern
import scanpy as sc
adata = sc.read_10x_mtx("counts/")
adata.var_names_make_unique()
adata.var["mt"] = adata.var_names.str.upper().str.startswith("MT-")
sc.pp.calculate_qc_metrics(adata, qc_vars=["mt"], inplace=True)
adata = adata[
(adata.obs["n_genes_by_counts"] >= 200)
& (adata.obs["n_genes_by_counts"] = 200`
- `max_genes = 3` for genes
### 4. Normalize, log-transform, and select HVGs
- normalize with `target_sum=1e4`
- `log1p`
- select `2000-4000` HVGs
- save raw counts before heavy transformations
### 5. Reduce dimensions and cluster
- PCA on HVGs
- neighbor graph using `10-30` PCs and `10-30` neighbors as a starting range
- UMAP for visualization
- Leiden across a small resolution grid such as `0.2`, `0.5`, `0.8`, `1.0`
### 6. Export analysis-ready artifacts
Always save:
- processed `h5ad`
- QC tables
- cluster assignments
- publication-ready QC and UMAP figures
## Output Artifacts
- `results/processed.h5ad`: main reusable AnnData object
- `results/cluster_assignments.tsv`: barcode plus cluster labels
- `qc/filter_summary.tsv`: counts before and after filtering
- `figures/umap_leiden.pdf`: main embedding figure
## Quality Review
- Median genes per cell should be plausible for the chemistry and tissue.
- Mitochondrial fraction should not dominate retained cells.
- PCA variance should decay smoothly rather than showing obvious technical axes only.
- UMAP should be reviewed together with QC metrics and batch labels, not alone.
- Cluster labels should not be finalized before marker inspection.
## Anti-Patterns
- reprocessing an already integrated object as if it were raw counts
- using a single universal mitochondrial threshold for every tissue
- interpreting UMAP separation as biology before checking batch and QC covariates
- discarding raw counts needed later for DE or pseudobulk
## Related Skills
- Cell Annotation
- Cell Communication
- Trajectory And Lineage
- Multiome And scATAC
## Optional Supplements
- `anndata`
- `scanpy`
## Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- **Author:** [cheatthegod](https://github.com/cheatthegod)
- **Source:** [cheatthegod/BioHermes](https://github.com/cheatthegod/BioHermes)
- **License:** MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.