AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
MCP verified MIT Self-run

Project Vectorizer

mcp-starkbaknet-project-vectorizer · by starkbaknet

A CLI tool that vectorizes codebases, stores them in a database, tracks changes, and serves them via MCP (Model Context Protocol) for AI agents like Claude, Codex, and others.

No reviews yet
0 installs
11 views
0.0% view→install

Install

$ agentstack add mcp-starkbaknet-project-vectorizer

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access Used
  • Shell / process execution No
  • Environment & secrets Used
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/mcp-starkbaknet-project-vectorizer)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
10mo ago

Declared compatibility

Claude CodeClaude DesktopCursorWindsurf

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Project Vectorizer? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Project Vectorizer

A powerful CLI tool that vectorizes codebases, stores them in a vector database, tracks changes, and serves them via MCP (Model Context Protocol) for AI agents like Claude, Codex, and others.

Latest Version: 0.1.4 | [Changelog](#changelog) | GitHub


📋 Table of Contents

  • [Features](#features)
  • [Installation](#installation)
  • [Quick Start](#quick-start)
  • [Performance Optimization](#performance-optimization)
  • [CLI Commands](#cli-commands)
  • [Configuration](#configuration)
  • [Search Features](#search-features)
  • [MCP Server](#mcp-server)
  • [Advanced Usage](#advanced-usage)
  • [Troubleshooting](#troubleshooting)
  • [Changelog](#changelog)
  • [Contributing](#contributing)

Features

🚀 Performance & Optimization

  • Auto-Optimized Config - Auto-detect CPU cores and RAM for optimal settings (--optimize)
  • Max Resources Mode - Use maximum system resources for fastest indexing (--max-resources)
  • Smart Incremental - 60-70% faster indexing with intelligent change categorization
  • Git-Aware Indexing - 80-90% faster by indexing only git-changed files
  • Parallel Processing - Multi-threaded with auto-detected optimal worker count (up to 16 workers)
  • Memory Monitoring - Real-time memory tracking with automatic garbage collection
  • Batch Optimization - Memory-based batch size calculation for safe processing

🔍 Search & Indexing

  • Code Vectorization - Parse and vectorize with sentence-transformers or OpenAI embeddings
  • Multi-Level Chunking - Functions, classes, micro-chunks, and word-level chunks for precision
  • Enhanced Single-Word Search - High-precision search for single keywords (0.8+ thresholds)
  • Semantic + Exact Search - Combines semantic similarity with exact word matching
  • Adaptive Thresholds - Automatically adjusts for optimal results
  • Multiple Languages - 30+ languages (Python, JS, TS, Go, Rust, Java, C++, C, PHP, Ruby, Swift, Kotlin, and more)

🔄 Change Management

  • Git Integration - Track changes via git commits with index-git command
  • Smart File Categorization - Detects New, Modified, and Deleted files
  • Watch Mode - Real-time monitoring with configurable debouncing (0.5-10s)
  • Incremental Updates - Only re-index changed content
  • Hash-Based Detection - SHA256 file hashing for accurate change detection

🌐 AI Integration

  • MCP Server - Model Context Protocol for AI agents (Claude, Codex, etc.)
  • HTTP Fallback API - RESTful endpoints when MCP unavailable
  • Semantic Search - Natural language queries for code discovery
  • File Operations - Get content, list files, project statistics

🎨 User Experience

  • Clean Progress Output - Single unified progress bar with timing information
  • Suppressed Library Logs - No cluttered batch progress bars from dependencies
  • Timing Information - Elapsed time for all operations (seconds or minutes+seconds)
  • Verbose Mode - Optional detailed logging for debugging
  • Professional UI - Rich terminal output with colors, panels, and formatting
  • Real-time Updates - Live file names and status tags during indexing

💾 Database & Storage

  • ChromaDB Backend - High-performance vector database
  • Fast HNSW Indexing - Optimized similarity search algorithm
  • Scalable - Handles 500K+ chunks efficiently
  • Single Database - No external dependencies required
  • Custom Paths - Configurable database location

Installation

From PyPI (Recommended)

# Install from PyPI
pip install project-vectorizer

# Verify installation
pv --version

From Source

# Clone repository
git clone https://github.com/starkbaknet/project-vectorizer.git
cd project-vectorizer

# Install
pip install -e .

# Or with development dependencies
pip install -e ".[dev]"

Quick Start

1. Initialize Your Project

# 🚀 Recommended: Auto-optimize based on your system (16 workers, 400 batch on 8-core/16GB RAM)
pv init /path/to/project --optimize

# Or with custom settings
pv init /path/to/project \
  --name "My Project" \
  --embedding-model "all-MiniLM-L6-v2" \
  --chunk-size 256 \
  --optimize

Output:

✓ Project initialized successfully!

Name: My Project
Path: /path/to/project
Model: all-MiniLM-L6-v2
Provider: sentence-transformers
Chunk Size: 256 tokens

Optimized Settings:
  • Workers: 16
  • Batch Size: 400
  • Embedding Batch: 200
  • Memory Monitoring: Enabled
  • GC Interval: 100 files

2. Index Your Codebase

# 🚀 Recommended: First-time indexing with max resources (2-4x faster)
pv index /path/to/project --max-resources

# 🚀 Recommended: Smart incremental for updates (60-70% faster)
pv index /path/to/project --smart

# 🚀 Recommended: Git-aware for recent changes (80-90% faster)
pv index-git /path/to/project --since HEAD~5

# Standard full indexing
pv index /path/to/project

# Force re-index everything
pv index /path/to/project --force

# Combine for maximum performance
pv index /path/to/project --smart --max-resources

Output:

Using maximum system resources (optimized settings)...
  • Workers: 16
  • Batch Size: 400
  • Embedding Batch: 200

  Indexing examples/demo.py ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 100%

╭────────────────── Indexing Complete ──────────────────╮
│ ✓ Indexing complete!                                  │
│                                                       │
│ Files indexed: 48/49                                  │
│ Total chunks: 9222                                    │
│ Model: all-MiniLM-L6-v2                               │
│ Time taken: 2m 16s                                    │
│                                                       │
│ You can now search with: pv search . "your query"     │
╰───────────────────────────────────────────────────────╯

3. Search Your Code

# Natural language search
pv search /path/to/project "authentication logic"

# Single-word searches work great (high precision)
pv search /path/to/project "async" --threshold 0.8
pv search /path/to/project "test" --threshold 0.9

# Multi-word queries (semantic search)
pv search /path/to/project "user login validation" --threshold 0.5

# Find specific constructs
pv search /path/to/project "class" --limit 10

Output:

Search Results for: authentication logic

Found 5 result(s) with threshold >= 0.5

╭─────────────────────── Result 1 ───────────────────────╮
│ src/auth/login.py                                      │
│ Lines 45-67 | Similarity: 0.892                        │
│                                                        │
│ def authenticate_user(username: str, password: str):   │
│     """                                                │
│     Authenticate user credentials against database     │
│     Returns user object if valid, None otherwise       │
│     """                                                │
│     ...                                                │
╰────────────────────────────────────────────────────────╯

4. Start MCP Server

# Start server (default: localhost:8000)
pv serve /path/to/project

# Custom host/port
pv serve /path/to/project --host 0.0.0.0 --port 8080

5. Monitor Changes in Real-Time

# Watch for file changes (default 2s debounce)
pv sync /path/to/project --watch

# Fast feedback (0.5s)
pv sync /path/to/project --watch --debounce 0.5

# Slower systems (5s)
pv sync /path/to/project --watch --debounce 5.0

Performance Optimization

Understanding the Optimization Flags

--optimize (Permanent)

Use when initializing a new project. Detects your system and saves optimal settings.

pv init /path/to/project --optimize

What it does:

  • Detects CPU cores → sets max_workers (e.g., 8 cores = 16 workers)
  • Calculates RAM → sets safe batch_size (e.g., 16GB = 400 batch)
  • Sets memory thresholds based on total RAM
  • Saves to config - All future operations use these settings

When to use:

  • ✅ New projects
  • ✅ Want permanent optimization
  • ✅ Same machine for all operations
  • ✅ "Set and forget" approach
--max-resources (Temporary)

Use when indexing to temporarily boost performance without changing config.

pv index /path/to/project --max-resources
pv index-git /path/to/project --since HEAD~1 --max-resources

What it does:

  • Detects system resources (same as --optimize)
  • Temporarily overrides config for this operation only
  • Original config unchanged

When to use:

  • ✅ Existing project without optimization
  • ✅ One-time heavy indexing
  • ✅ CI/CD with dedicated resources
  • ✅ Don't want to modify config

Performance Benchmarks

System: 8-core CPU, 16GB RAM, SSD

| Mode | Files | Chunks | Time | Settings | | ------------------ | --------- | ------ | ------ | --------------------- | | Standard | 48 | 9222 | 4m 32s | 4 workers, 100 batch | | --max-resources | 48 | 9222 | 2m 16s | 16 workers, 400 batch | | Smart incremental | 5 changed | 412 | 24s | 16 workers, 400 batch | | Git-aware (HEAD~1) | 3 changed | 287 | 15s | 16 workers, 400 batch |

Key Findings:

  • --max-resources: 2x faster for full indexing
  • Smart incremental: 60-70% faster than full reindex
  • Git-aware: 80-90% faster for recent changes
  • Chunk size (128 vs 512): No performance difference (same ~2m 16s)

System Resource Detection

CPU Detection:

Detected: 8 cores
Optimal workers: min(8 * 2, 16) = 16 workers

Memory Detection:

Total RAM: 16GB
Available RAM: 8GB
Safe batch size: 8GB * 0.5 * 100 = 400
Embedding batch: 400 * 0.5 = 200
GC interval: 100 files

Memory Thresholds:

32GB+ RAM → threshold: 50000
16-32GB   → threshold: 20000
8-16GB    → threshold: 10000
/.vectorizer/config.json`

### Full Configuration Reference

```json
{
  "chromadb_path": null,
  "embedding_model": "all-MiniLM-L6-v2",
  "embedding_provider": "sentence-transformers",
  "openai_api_key": null,
  "chunk_size": 128,
  "chunk_overlap": 32,
  "max_file_size_mb": 10,
  "included_extensions": [
    ".py",
    ".js",
    ".ts",
    ".jsx",
    ".tsx",
    ".go",
    ".rs",
    ".java",
    ".cpp",
    ".c",
    ".h",
    ".hpp",
    ".cs",
    ".php",
    ".rb",
    ".swift",
    ".kt",
    ".scala",
    ".clj",
    ".sh",
    ".bash",
    ".zsh",
    ".fish",
    ".ps1",
    ".bat",
    ".cmd",
    ".md",
    ".txt",
    ".rst",
    ".json",
    ".yaml",
    ".yml",
    ".toml",
    ".xml",
    ".html",
    ".css",
    ".scss",
    ".sql",
    ".graphql",
    ".proto"
  ],
  "excluded_patterns": [
    "node_modules/**",
    ".git/**",
    "__pycache__/**",
    "*.pyc",
    ".pytest_cache/**",
    "venv/**",
    "env/**",
    ".env/**",
    "build/**",
    "dist/**",
    "*.egg-info/**",
    ".DS_Store",
    "*.min.js",
    "*.min.css"
  ],
  "mcp_host": "localhost",
  "mcp_port": 8000,
  "log_level": "INFO",
  "log_file": null,
  "max_workers": 4,
  "batch_size": 100,
  "embedding_batch_size": 100,
  "parallel_file_processing": true,
  "memory_monitoring_enabled": true,
  "memory_efficient_search_threshold": 10000,
  "gc_interval": 100
}

Key Settings Explained

Embedding Settings:

  • embedding_model: Model for embeddings (all-MiniLM-L6-v2, text-embedding-ada-002, etc.)
  • embedding_provider: "sentence-transformers" (local) or "openai" (API)
  • chunk_size: Tokens per chunk (128 for precision, 512 for context)
  • chunk_overlap: Overlap between chunks (16-32 recommended)

Performance Settings:

  • max_workers: Parallel workers (auto-detected with --optimize)
  • batch_size: Files per batch (auto-calculated with --optimize)
  • embedding_batch_size: Embeddings per batch
  • parallel_file_processing: Enable parallel processing (recommended: true)

Memory Settings:

  • memory_monitoring_enabled: Monitor RAM usage (recommended: true)
  • memory_efficient_search_threshold: Switch to streaming for large results
  • gc_interval: Garbage collection frequency (files between GC)

File Filtering:

  • included_extensions: File types to index
  • excluded_patterns: Glob patterns to ignore
  • max_file_size_mb: Skip files larger than this

Server Settings:

  • mcp_host: MCP server host
  • mcp_port: MCP server port
  • log_level: INFO, DEBUG, WARNING, ERROR
  • chromadb_path: Custom ChromaDB location (optional)

Environment Variables

Create .env file or export:

# OpenAI API Key (required for OpenAI embeddings)
export OPENAI_API_KEY="sk-..."

# Override config values
export EMBEDDING_PROVIDER="sentence-transformers"
export EMBEDDING_MODEL="all-MiniLM-L6-v2"
export CHUNK_SIZE="256"
export DEFAULT_SEARCH_THRESHOLD="0.3"

# Database
export CHROMADB_PATH="/custom/path/to/chromadb"

# Logging
export LOG_LEVEL="INFO"
export LOG_FILE="/var/log/vectorizer.log"

For complete list, see [docs/ENVIRONMENT.md](docs/ENVIRONMENT.md)

Editing Configuration

# View current config
cat /path/to/project/.vectorizer/config.json

# Edit manually
nano /path/to/project/.vectorizer/config.json

# Or regenerate with optimization
pv init /path/to/project --optimize

Search Features

Single-Word Search

Optimized for high-precision single-keyword searches.

# Programming keywords
pv search /path/to/project "async" --threshold 0.9
pv search /path/to/project "test" --threshold 0.8
pv search /path/to/project "class" --threshold 0.9
pv search /path/to/project "import" --threshold 0.85

# Works great for finding specific constructs
pv search /path/to/project "def" --threshold 0.9  # Python functions
pv search /path/to/project "function" --threshold 0.9  # JS functions
pv search /path/to/project "catch" --threshold 0.8  # Error handling

Features:

  • Exact Word Matching: Prioritizes exact word boundaries
  • Keyword Detection: Special handling for programming keywords
  • Relevance Boosting: Huge boost for exact matches
  • High Thresholds: Reliable results even at 0.8-0.9+

Multi-Word Search

Semantic search for phrases and concepts.

# Natural language
pv search /path/to/project "user authentication logic" --threshold 0.5

# Code patterns
pv search /path/to/project "error handling in database" --threshold 0.4

# Features
pv search /path/to/project "rate limiting middleware" --threshold 0.6

Search Result Ranking

Results ranked by:

  1. Exact word matches (highest priority)
  2. Content type (micro/word chunks get boost)
  3. Partial matches within larger words
  4. Semantic similarity from embeddings

Recommended Thresholds by Query Type

| Query Type | Threshold | Example | | -------------- | --------- | --------------------------------- | | Single keyword | 0.7-0.95 | "async", "test", "class" | | Two words | 0.5-0.8 | "error handling", "api routes" | | Short phrase | 0.4-0.7 | "user login validation" | | Complex query | 0.3-0.5 | "authentication with jwt tokens" | | Exploratory | 0.1-0.3 | "machine learning model training" |


MCP Server

Starting the Server

# Default (localhost:8000)
pv serve /path/to/project

# Custom settings
pv serve /path/to/project --host 0.0.0.0 --port 8080

Available MCP Tools

When running, AI agents can use these tools:

  1. search_code - Search vectorized codebase

``json { "query": "authentication logic", "limit": 10, "threshold": 0.5 } ``

  1. getfilecontent - Retrieve full file

``json { "file_path": "src/auth/login.py" } ``

  1. list_files - List all files

``json { "file_type": "py" // optional filter } ``

  1. getprojectstats - Get statistics

``json {} ``

HTTP Fallback API

If MCP unavaila

Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.