Install
$ agentstack add skill-benchflow-ai-skillsbench-hierarchical-taxonomy-clustering ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Hierarchical Taxonomy Clustering
Create a unified multi-level taxonomy from hierarchical category paths by clustering similar paths and automatically generating meaningful category names.
Problem
Given category paths from multiple sources (e.g., "electronics -> computers -> laptops"), create a unified taxonomy that groups similar paths across sources, generates meaningful category names, and produces a clean N-level hierarchy (typically 5 levels). The unified category taxonomy could be used to do analysis or metric tracking on products from different platform.
Methodology
- Hierarchical Weighting: Convert paths to embeddings with exponentially decaying weights (Level i gets weight 0.6^(i-1)) to signify the importance of category granularity
- Recursive Clustering: Hierarchically cluster at each level (10-20 clusters at L1, 3-20 at L2-L5) using cosine distance
- Intelligent Naming: Generate category names via weighted word frequency + lemmatization + bundle word logic
- Quality Control: Exclude all ancestor words (parent, grandparent, etc.), avoid ancestor path duplicates, clean special characters
Output
DataFrame with added columns:
unified_level_1: Top-level category (e.g., "electronic | device")unified_level_2: Second-level category (e.g., "computer | laptop")unified_level_3throughunified_level_N: Deeper levels
Category names use | separator, max 5 words, covering 70%+ of records in each cluster.
Installation
pip install pandas numpy scipy sentence-transformers nltk tqdm
python -c "import nltk; nltk.download('wordnet'); nltk.download('omw-1.4')"
4-Step Pipeline
Step 1: Load, Standardize, Filter and Merge (step1_preprocessing_and_merge.py)
- Input: List of (DataFrame, source_name) tuples, each of the with
category_pathcolumn - Process: Per-source deduplication, text cleaning (remove &/,/'/-/quotes,'and' or "&", "," and so on, lemmatize words as nouns), normalize delimiter to
>, depth filtering, prefix removal, then merge all sources. source_level should reflect the processed version of the source level name - Output: Merged DataFrame with
category_path,source,depth,source_level_1throughsource_level_N
Step 2: Weighted Embeddings (step2_weighted_embedding_generation.py)
- Input: DataFrame from Step 1
- Output: Numpy embedding matrix (n_records × 384)
- Weights: L1=1.0, L2=0.6, L3=0.36, L4=0.216, L5=0.1296 (exponential decay 0.6^(n-1))
- Performance: For ~10,000 records, expect 2-5 minutes. Progress bar will show encoding status.
Step 3: Recursive Clustering (step3_recursive_clustering_naming.py)
- Input: DataFrame + embeddings from Step 2
- Output: Assignments dict {index → {level1: ..., level5: ...}}
- Average linkage + cosine distance, 10-20 clusters at L1, 3-20 at L2-L5
- Word-based naming: weighted frequency + lemmatization + coverage ≥70%
- Performance: For ~10,000 records, expect 1-3 minutes for hierarchical clustering and naming. Be patient - the system is working through recursive levels.
Step 4: Export Results (step4_result_assignments.py)
- Input: DataFrame + assignments from Step 3
- Output:
unified_taxonomy_full.csv- all records with unified categoriesunified_taxonomy_hierarchy.csv- unique taxonomy structure
Usage
Use scripts/pipeline.py to run the complete 4-step workflow.
See scripts/pipeline.py for:
- Complete implementation of all 4 steps
- Example code for processing multiple sources
- Command-line interface
- Individual step usage (for advanced control)
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: benchflow-ai
- Source: benchflow-ai/skillsbench
- License: Apache-2.0
- Homepage: https://www.skillsbench.ai
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.