Install
$ agentstack add mcp-starkbaknet-project-vectorizer ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ● Network access Used
- ● Filesystem access Used
- ✓ Shell / process execution No
- ● Environment & secrets Used
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Project Vectorizer
A powerful CLI tool that vectorizes codebases, stores them in a vector database, tracks changes, and serves them via MCP (Model Context Protocol) for AI agents like Claude, Codex, and others.
Latest Version: 0.1.4 | [Changelog](#changelog) | GitHub
📋 Table of Contents
- [Features](#features)
- [Installation](#installation)
- [Quick Start](#quick-start)
- [Performance Optimization](#performance-optimization)
- [CLI Commands](#cli-commands)
- [Configuration](#configuration)
- [Search Features](#search-features)
- [MCP Server](#mcp-server)
- [Advanced Usage](#advanced-usage)
- [Troubleshooting](#troubleshooting)
- [Changelog](#changelog)
- [Contributing](#contributing)
Features
🚀 Performance & Optimization
- Auto-Optimized Config - Auto-detect CPU cores and RAM for optimal settings (
--optimize) - Max Resources Mode - Use maximum system resources for fastest indexing (
--max-resources) - Smart Incremental - 60-70% faster indexing with intelligent change categorization
- Git-Aware Indexing - 80-90% faster by indexing only git-changed files
- Parallel Processing - Multi-threaded with auto-detected optimal worker count (up to 16 workers)
- Memory Monitoring - Real-time memory tracking with automatic garbage collection
- Batch Optimization - Memory-based batch size calculation for safe processing
🔍 Search & Indexing
- Code Vectorization - Parse and vectorize with sentence-transformers or OpenAI embeddings
- Multi-Level Chunking - Functions, classes, micro-chunks, and word-level chunks for precision
- Enhanced Single-Word Search - High-precision search for single keywords (0.8+ thresholds)
- Semantic + Exact Search - Combines semantic similarity with exact word matching
- Adaptive Thresholds - Automatically adjusts for optimal results
- Multiple Languages - 30+ languages (Python, JS, TS, Go, Rust, Java, C++, C, PHP, Ruby, Swift, Kotlin, and more)
🔄 Change Management
- Git Integration - Track changes via git commits with
index-gitcommand - Smart File Categorization - Detects New, Modified, and Deleted files
- Watch Mode - Real-time monitoring with configurable debouncing (0.5-10s)
- Incremental Updates - Only re-index changed content
- Hash-Based Detection - SHA256 file hashing for accurate change detection
🌐 AI Integration
- MCP Server - Model Context Protocol for AI agents (Claude, Codex, etc.)
- HTTP Fallback API - RESTful endpoints when MCP unavailable
- Semantic Search - Natural language queries for code discovery
- File Operations - Get content, list files, project statistics
🎨 User Experience
- Clean Progress Output - Single unified progress bar with timing information
- Suppressed Library Logs - No cluttered batch progress bars from dependencies
- Timing Information - Elapsed time for all operations (seconds or minutes+seconds)
- Verbose Mode - Optional detailed logging for debugging
- Professional UI - Rich terminal output with colors, panels, and formatting
- Real-time Updates - Live file names and status tags during indexing
💾 Database & Storage
- ChromaDB Backend - High-performance vector database
- Fast HNSW Indexing - Optimized similarity search algorithm
- Scalable - Handles 500K+ chunks efficiently
- Single Database - No external dependencies required
- Custom Paths - Configurable database location
Installation
From PyPI (Recommended)
# Install from PyPI
pip install project-vectorizer
# Verify installation
pv --version
From Source
# Clone repository
git clone https://github.com/starkbaknet/project-vectorizer.git
cd project-vectorizer
# Install
pip install -e .
# Or with development dependencies
pip install -e ".[dev]"
Quick Start
1. Initialize Your Project
# 🚀 Recommended: Auto-optimize based on your system (16 workers, 400 batch on 8-core/16GB RAM)
pv init /path/to/project --optimize
# Or with custom settings
pv init /path/to/project \
--name "My Project" \
--embedding-model "all-MiniLM-L6-v2" \
--chunk-size 256 \
--optimize
Output:
✓ Project initialized successfully!
Name: My Project
Path: /path/to/project
Model: all-MiniLM-L6-v2
Provider: sentence-transformers
Chunk Size: 256 tokens
Optimized Settings:
• Workers: 16
• Batch Size: 400
• Embedding Batch: 200
• Memory Monitoring: Enabled
• GC Interval: 100 files
2. Index Your Codebase
# 🚀 Recommended: First-time indexing with max resources (2-4x faster)
pv index /path/to/project --max-resources
# 🚀 Recommended: Smart incremental for updates (60-70% faster)
pv index /path/to/project --smart
# 🚀 Recommended: Git-aware for recent changes (80-90% faster)
pv index-git /path/to/project --since HEAD~5
# Standard full indexing
pv index /path/to/project
# Force re-index everything
pv index /path/to/project --force
# Combine for maximum performance
pv index /path/to/project --smart --max-resources
Output:
Using maximum system resources (optimized settings)...
• Workers: 16
• Batch Size: 400
• Embedding Batch: 200
Indexing examples/demo.py ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 100%
╭────────────────── Indexing Complete ──────────────────╮
│ ✓ Indexing complete! │
│ │
│ Files indexed: 48/49 │
│ Total chunks: 9222 │
│ Model: all-MiniLM-L6-v2 │
│ Time taken: 2m 16s │
│ │
│ You can now search with: pv search . "your query" │
╰───────────────────────────────────────────────────────╯
3. Search Your Code
# Natural language search
pv search /path/to/project "authentication logic"
# Single-word searches work great (high precision)
pv search /path/to/project "async" --threshold 0.8
pv search /path/to/project "test" --threshold 0.9
# Multi-word queries (semantic search)
pv search /path/to/project "user login validation" --threshold 0.5
# Find specific constructs
pv search /path/to/project "class" --limit 10
Output:
Search Results for: authentication logic
Found 5 result(s) with threshold >= 0.5
╭─────────────────────── Result 1 ───────────────────────╮
│ src/auth/login.py │
│ Lines 45-67 | Similarity: 0.892 │
│ │
│ def authenticate_user(username: str, password: str): │
│ """ │
│ Authenticate user credentials against database │
│ Returns user object if valid, None otherwise │
│ """ │
│ ... │
╰────────────────────────────────────────────────────────╯
4. Start MCP Server
# Start server (default: localhost:8000)
pv serve /path/to/project
# Custom host/port
pv serve /path/to/project --host 0.0.0.0 --port 8080
5. Monitor Changes in Real-Time
# Watch for file changes (default 2s debounce)
pv sync /path/to/project --watch
# Fast feedback (0.5s)
pv sync /path/to/project --watch --debounce 0.5
# Slower systems (5s)
pv sync /path/to/project --watch --debounce 5.0
Performance Optimization
Understanding the Optimization Flags
--optimize (Permanent)
Use when initializing a new project. Detects your system and saves optimal settings.
pv init /path/to/project --optimize
What it does:
- Detects CPU cores → sets
max_workers(e.g., 8 cores = 16 workers) - Calculates RAM → sets safe
batch_size(e.g., 16GB = 400 batch) - Sets memory thresholds based on total RAM
- Saves to config - All future operations use these settings
When to use:
- ✅ New projects
- ✅ Want permanent optimization
- ✅ Same machine for all operations
- ✅ "Set and forget" approach
--max-resources (Temporary)
Use when indexing to temporarily boost performance without changing config.
pv index /path/to/project --max-resources
pv index-git /path/to/project --since HEAD~1 --max-resources
What it does:
- Detects system resources (same as --optimize)
- Temporarily overrides config for this operation only
- Original config unchanged
When to use:
- ✅ Existing project without optimization
- ✅ One-time heavy indexing
- ✅ CI/CD with dedicated resources
- ✅ Don't want to modify config
Performance Benchmarks
System: 8-core CPU, 16GB RAM, SSD
| Mode | Files | Chunks | Time | Settings | | ------------------ | --------- | ------ | ------ | --------------------- | | Standard | 48 | 9222 | 4m 32s | 4 workers, 100 batch | | --max-resources | 48 | 9222 | 2m 16s | 16 workers, 400 batch | | Smart incremental | 5 changed | 412 | 24s | 16 workers, 400 batch | | Git-aware (HEAD~1) | 3 changed | 287 | 15s | 16 workers, 400 batch |
Key Findings:
--max-resources: 2x faster for full indexing- Smart incremental: 60-70% faster than full reindex
- Git-aware: 80-90% faster for recent changes
- Chunk size (128 vs 512): No performance difference (same ~2m 16s)
System Resource Detection
CPU Detection:
Detected: 8 cores
Optimal workers: min(8 * 2, 16) = 16 workers
Memory Detection:
Total RAM: 16GB
Available RAM: 8GB
Safe batch size: 8GB * 0.5 * 100 = 400
Embedding batch: 400 * 0.5 = 200
GC interval: 100 files
Memory Thresholds:
32GB+ RAM → threshold: 50000
16-32GB → threshold: 20000
8-16GB → threshold: 10000
/.vectorizer/config.json`
### Full Configuration Reference
```json
{
"chromadb_path": null,
"embedding_model": "all-MiniLM-L6-v2",
"embedding_provider": "sentence-transformers",
"openai_api_key": null,
"chunk_size": 128,
"chunk_overlap": 32,
"max_file_size_mb": 10,
"included_extensions": [
".py",
".js",
".ts",
".jsx",
".tsx",
".go",
".rs",
".java",
".cpp",
".c",
".h",
".hpp",
".cs",
".php",
".rb",
".swift",
".kt",
".scala",
".clj",
".sh",
".bash",
".zsh",
".fish",
".ps1",
".bat",
".cmd",
".md",
".txt",
".rst",
".json",
".yaml",
".yml",
".toml",
".xml",
".html",
".css",
".scss",
".sql",
".graphql",
".proto"
],
"excluded_patterns": [
"node_modules/**",
".git/**",
"__pycache__/**",
"*.pyc",
".pytest_cache/**",
"venv/**",
"env/**",
".env/**",
"build/**",
"dist/**",
"*.egg-info/**",
".DS_Store",
"*.min.js",
"*.min.css"
],
"mcp_host": "localhost",
"mcp_port": 8000,
"log_level": "INFO",
"log_file": null,
"max_workers": 4,
"batch_size": 100,
"embedding_batch_size": 100,
"parallel_file_processing": true,
"memory_monitoring_enabled": true,
"memory_efficient_search_threshold": 10000,
"gc_interval": 100
}
Key Settings Explained
Embedding Settings:
embedding_model: Model for embeddings (all-MiniLM-L6-v2, text-embedding-ada-002, etc.)embedding_provider: "sentence-transformers" (local) or "openai" (API)chunk_size: Tokens per chunk (128 for precision, 512 for context)chunk_overlap: Overlap between chunks (16-32 recommended)
Performance Settings:
max_workers: Parallel workers (auto-detected with --optimize)batch_size: Files per batch (auto-calculated with --optimize)embedding_batch_size: Embeddings per batchparallel_file_processing: Enable parallel processing (recommended: true)
Memory Settings:
memory_monitoring_enabled: Monitor RAM usage (recommended: true)memory_efficient_search_threshold: Switch to streaming for large resultsgc_interval: Garbage collection frequency (files between GC)
File Filtering:
included_extensions: File types to indexexcluded_patterns: Glob patterns to ignoremax_file_size_mb: Skip files larger than this
Server Settings:
mcp_host: MCP server hostmcp_port: MCP server portlog_level: INFO, DEBUG, WARNING, ERRORchromadb_path: Custom ChromaDB location (optional)
Environment Variables
Create .env file or export:
# OpenAI API Key (required for OpenAI embeddings)
export OPENAI_API_KEY="sk-..."
# Override config values
export EMBEDDING_PROVIDER="sentence-transformers"
export EMBEDDING_MODEL="all-MiniLM-L6-v2"
export CHUNK_SIZE="256"
export DEFAULT_SEARCH_THRESHOLD="0.3"
# Database
export CHROMADB_PATH="/custom/path/to/chromadb"
# Logging
export LOG_LEVEL="INFO"
export LOG_FILE="/var/log/vectorizer.log"
For complete list, see [docs/ENVIRONMENT.md](docs/ENVIRONMENT.md)
Editing Configuration
# View current config
cat /path/to/project/.vectorizer/config.json
# Edit manually
nano /path/to/project/.vectorizer/config.json
# Or regenerate with optimization
pv init /path/to/project --optimize
Search Features
Single-Word Search
Optimized for high-precision single-keyword searches.
# Programming keywords
pv search /path/to/project "async" --threshold 0.9
pv search /path/to/project "test" --threshold 0.8
pv search /path/to/project "class" --threshold 0.9
pv search /path/to/project "import" --threshold 0.85
# Works great for finding specific constructs
pv search /path/to/project "def" --threshold 0.9 # Python functions
pv search /path/to/project "function" --threshold 0.9 # JS functions
pv search /path/to/project "catch" --threshold 0.8 # Error handling
Features:
- Exact Word Matching: Prioritizes exact word boundaries
- Keyword Detection: Special handling for programming keywords
- Relevance Boosting: Huge boost for exact matches
- High Thresholds: Reliable results even at 0.8-0.9+
Multi-Word Search
Semantic search for phrases and concepts.
# Natural language
pv search /path/to/project "user authentication logic" --threshold 0.5
# Code patterns
pv search /path/to/project "error handling in database" --threshold 0.4
# Features
pv search /path/to/project "rate limiting middleware" --threshold 0.6
Search Result Ranking
Results ranked by:
- Exact word matches (highest priority)
- Content type (micro/word chunks get boost)
- Partial matches within larger words
- Semantic similarity from embeddings
Recommended Thresholds by Query Type
| Query Type | Threshold | Example | | -------------- | --------- | --------------------------------- | | Single keyword | 0.7-0.95 | "async", "test", "class" | | Two words | 0.5-0.8 | "error handling", "api routes" | | Short phrase | 0.4-0.7 | "user login validation" | | Complex query | 0.3-0.5 | "authentication with jwt tokens" | | Exploratory | 0.1-0.3 | "machine learning model training" |
MCP Server
Starting the Server
# Default (localhost:8000)
pv serve /path/to/project
# Custom settings
pv serve /path/to/project --host 0.0.0.0 --port 8080
Available MCP Tools
When running, AI agents can use these tools:
- search_code - Search vectorized codebase
``json { "query": "authentication logic", "limit": 10, "threshold": 0.5 } ``
- getfilecontent - Retrieve full file
``json { "file_path": "src/auth/login.py" } ``
- list_files - List all files
``json { "file_type": "py" // optional filter } ``
- getprojectstats - Get statistics
``json {} ``
HTTP Fallback API
If MCP unavaila
…
Source & license
This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: starkbaknet
- Source: starkbaknet/project-vectorizer
- License: MIT
- Homepage: https://pypi.org/project/project-vectorizer
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.