Install
$ agentstack add mcp-lyonzin-knowledge-rag Open-source listing, not yet scanned by AgentStack. Follow the source repository for install instructions.
Security review
⚠ Flagged1 finding(s); flagged for manual review. · v0.1.0 How review works →
- • Prompt-injection patterns
- • Secret / credential exfiltration
- • Dangerous shell & filesystem operations
- • Untrusted network calls
- • Known-malicious package signatures
- high Pipes remote content directly into a shell (remote code execution).
What it can access
- ● Network access Used
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Knowledge RAG
[](https://pypi.org/project/knowledge-rag/) [](https://www.npmjs.com/package/knowledge-rag) [](https://pepy.tech/projects/knowledge-rag)
[](https://github.com/lyonzin/knowledge-rag/actions/workflows/ci.yml) [](https://github.com/lyonzin/knowledge-rag/actions/workflows/security.yml) [](https://github.com/lyonzin/knowledge-rag/actions/workflows/quality-gate.yml) [](https://glama.ai/mcp/servers/lyonzin/knowledge-rag)
Your docs, your machine, zero cloud. Claude Code searches them natively.
Drop your PDFs, markdown, code, notebooks — 1800+ files, 39K chunks, indexed in under 3 minutes. Hybrid search (BM25 + semantic vectors + cross-encoder reranking) through 13 MCP tools. Everything runs locally via ONNX. No Docker, no Ollama, no API keys, no data leaves your machine.
pip install knowledge-rag → restart Claude Code → search_knowledge("your query")
13 MCP Tools | Hybrid Search + Reranking | 20 File Formats | Optional NVIDIA GPU | 100% Local
[What's New](#whats-new-in-v420) | [Supported Formats](#supported-formats) | [Installation](#installation) | [Configuration](#configuration) | [API Reference](#api-reference) | [Architecture](#architecture)
Star History
What's New in v4.2.0
Search Performance & Output Quality (v4.2.0)
128× faster BM25 search — replaced rank-bm25 full-corpus scan with a custom inverted-index implementation. Only documents containing query terms are scored, using numpy.argpartition for O(n) top-k selection. Adjacent chunk fetching now uses a single batched ChromaDB call instead of N round-trips, and an O(1) reverse lookup (_source_to_docid) eliminates linear scans.
Smarter output — two new parameters on search_knowledge:
snippet_mode(default:true) — truncates content to ~500 characters at natural break points, reducing token consumption by ~72%. Addscontent_lengthfield with original size; useget_document()for full content.min_score— filters results below a normalized relevance threshold (0.0–1.0). Eliminates low-quality noise from results. Response includesfiltered_by_scorecount for transparency.
Both parameters are fully backwards-compatible (existing callers see no change in behavior).
Enterprise Concurrent Access — SSE/HTTP Transport (v4.0.0)
The server now supports SSE and streamable-http transport modes. Instead of spawning a separate process per client (stdio), a single server process serves all clients with shared resources — 1 embedding model, 1 ChromaDB, 1 query cache.
# config.yaml
server:
transport: "sse" # "stdio" | "sse" | "streamable-http"
host: "127.0.0.1"
port: 8179
Or via CLI: knowledge-rag --transport sse
Optional enterprise features (all disabled by default):
- Rate limiting: Sliding-window counter, configurable RPM and burst
- Prometheus metrics:
/metricsendpoint on separate port - Bearer auth: Token validation for SSE/HTTP connections
All 13 MCP tools are instrumented with @rate_limited and @instrument decorators — zero overhead when features are disabled. Default transport remains stdio for full backwards compatibility.
> Migration: Existing users need zero changes. SSE mode is opt-in via server.transport: "sse" in config.yaml. See [Configuration](#configuration) for details.
Quality Gate — 7-Pillar PR Validation
Every PR (including dependabot bumps and one-line fixes) is now evaluated against 35+ automated checks spread across 7 pillars before any human review:
| Pillar | What it enforces | Tools | |---|---|---| | 1 Security | SAST, secrets, CVEs, supply chain | bandit, semgrep, gitleaks, pip-audit, dependency-review, Snyk, CodeQL, Socket | | 2 Stability | Flake detection, coverage trend, test count, deterministic runs | pytest-rerunfailures, codecov ±0.5pp, test-count guard | | 3 Memory Leak | RSS bounded under 1000-query load, no idle bloat | psutil-based baseline tests + nightly 50K-iteration soak | | 4 Versatility | 9 OS×Python combos, 14 format parsers, 4 config presets, locale tolerance, property-based fuzzing | matrix CI on Linux+Windows+macOS × 3.11+3.12+3.13, Hypothesis | | 5 Scalability | Performance regression > 10% blocks merge, public bench dashboard | pytest-benchmark, GH Pages chart | | 6 Versioning | Atomic version sync, API surface diff, conventional commits, CHANGELOG enforcement, backwards compat | griffe-style AST diff, custom guards | | 7 Quality | Type strictness, docstring coverage, complexity, dead code | mypy strict, interrogate ≥80%, radon, vulture |
Plus a nightly resilience workflow that runs chaos failure-injection (HF down, ChromaDB corruption, watchdog crash, ONNX zero-byte replay), determinism check (full suite × 3), and mutation testing on selected modules.
Read the full philosophy in [CONTRIBUTING.md](CONTRIBUTING.md). Report bugs via [SECURITY.md](SECURITY.md) or the [issue templates](.github/ISSUE_TEMPLATE/).
Critical Hotfix — No More Silent Zero-Vector Corruption (v3.8.1)
FastEmbedEmbeddings.__call__ no longer swallows exceptions and returns [[0.0]*dim, ...] when the ONNX model fails to load. That bug pre-existed in master but was silent: ChromaDB happily stored zero embeddings, count() reported normal numbers, smart-reindex skipped them as "already indexed", and queries returned garbage similarity with no error visible. Now raises EmbeddingModelLoadError / EmbeddingError loudly. All v3.8.0 users should upgrade. Full details in [Changelog](#v381-2026-05-10--hotfix).
Lazy-Loaded Embeddings — Cheaper Idle Processes (v3.8.0)
The FastEmbed ONNX model (~200MB resident) now loads on the first query, not at startup. Idle knowledge-rag processes are now genuinely cheap. Why this matters: MCP stdio is one-process-per-client by protocol — multiple Claude Code windows, Claude Desktop + IDE simultaneously, or review/approval flows that open extra connections all spawn their own processes. Before v3.8.0, every one of them paid the full embedding-model cost up front. Now only processes that actually serve queries load the model. Public API is unchanged.
Opt-In Single-Instance Guard (v3.8.0)
For users who measured their setup and want a hard cap of one server per data_dir:
export KNOWLEDGE_RAG_SINGLE_INSTANCE=1
A second instance exits immediately with code 75. OFF by default so multi-client MCP usage continues to work unchanged. Stale-PID recovery + SIGINT/SIGTERM cleanup wired correctly. Full guide in [docs/single-instance.md](docs/single-instance.md). Sample MCP config in [examples/mcp-config-single-instance.json](examples/mcp-config-single-instance.json).
5 Ways to Install
npx -y knowledge-rag # NPM — zero setup, auto-manages Python venv
pip install knowledge-rag # PyPI — classic Python install
curl -fsSL .../install.sh | bash # One-line installer (Linux/macOS/Windows)
docker pull ghcr.io/lyonzin/knowledge-rag # Docker — models pre-downloaded
git clone ... && pip install -r ... # From source
All methods produce the same MCP server. See [Installation](#installation) for full instructions.
Recent Highlights
- v4.0.0 — Enterprise concurrent access: SSE/HTTP transport (1 server → N clients), thread-safe shared state, optional rate limiting + Prometheus metrics, ChromaDB WAL mode,
--transportCLI - v3.9.0 — Quality Gate activated: 35+ automated PR checks across 7 pillars (Security, Stability, Memory Leak, Versatility, Scalability, Versioning, Quality) + nightly resilience suite (chaos, soak, determinism, mutation)
- v3.8.1 — Critical hotfix: loud-fail embeddings (no more silent zero-vector corruption); Windows CI flake erradicated (HFHUBOFFLINE + shell:bash + atexit wrapper)
- v3.8.0 — Lazy-load embeddings, opt-in single-instance guard, version sync across PyPI/NPM/Docker
- v3.6.0 — Multi-language code parsing (C/C++/JS/TS/XML), NPM wrapper, Docker image, automated release pipeline
- v3.5.2 — CUDA DLL auto-discovery from pip packages, graceful GPU→CPU fallback, explicit CPU provider (no CUDA noise when
gpu: false), BASE_DIR resolution fix for editable installs - v3.5.1 — Remove Python
**Tip:** The parser dispatch is extensible. Any format mapped inparserscan be enabled viasupportedformats` in config.yaml.
Features
| Feature | Description | |---------|-------------| | Hybrid Search | Semantic + BM25 keyword search with Reciprocal Rank Fusion | | Cross-Encoder Reranker | Xenova/ms-marco-MiniLM-L-6-v2 re-scores top candidates for precision | | GPU Acceleration | Optional ONNX CUDA support for 5-10x faster indexing | | YAML Configuration | Fully customizable via config.yaml with domain-specific presets | | Query Expansion | Configurable synonym mappings (69 security-term defaults) | | Markdown-Aware Chunking | .md files split by ##/### sections instead of fixed windows | | In-Process Embeddings | FastEmbed ONNX Runtime (BAAI/bge-small-en-v1.5, 384D) | | Keyword Routing | Word-boundary aware routing for domain-specific queries | | 20 Format Parsers | MD, TXT, PDF, PY, C, H, CPP, JS, JSX, TS, TSX, JSON, XML, CSV, DOCX, XLSX, PPTX, IPYNB + opt-in MQH/MQ4 | | Category Organization | Organize docs by folder, auto-tagged by path | | Incremental Indexing | Change detection via mtime/size — only re-indexes modified files | | Chunk Deduplication | SHA256 content hashing prevents duplicate chunks | | Query Cache | LRU cache with 5-min TTL for instant repeat queries | | Document CRUD | Add, update, remove documents via MCP tools | | URL Ingestion | Fetch URLs, strip HTML, convert to markdown, index | | Similarity Search | Find documents similar to a reference document | | Retrieval Evaluation | Built-in MRR@5 and Recall@5 metrics | | File Watcher | Auto-reindex on document changes via watchdog (5s debounce) | | Exclude Patterns | Glob-based file/directory exclusion during indexing | | MMR Diversification | Maximal Marginal Relevance reduces redundant results | | Persistent Model Cache | Embedding models cached in models_cache/ — survives reboots | | Auto-Migration | Detects embedding dimension mismatch and rebuilds automatically | | 13 MCP Tools | Full CRUD + search + evaluation via Claude Code |
Architecture
System Overview
flowchart TB
subgraph MCP["MCP SERVER (FastMCP)"]
direction TB
TOOLS["13 MCP Toolssearch | get | add | update | removereindex | reindex_status | list | stats | url | similar | evaluate"]
end
subgraph SEARCH["HYBRID SEARCH ENGINE"]
direction LR
ROUTER["Keyword Router(word boundaries)"]
SEMANTIC["Semantic Search(ChromaDB)"]
BM25["BM25 Keyword(inverted-index + expansion)"]
RRF["Reciprocal RankFusion (RRF)"]
RERANK["Cross-EncoderReranker"]
ROUTER --> SEMANTIC
ROUTER --> BM25
SEMANTIC --> RRF
BM25 --> RRF
RRF --> RERANK
end
subgraph STORAGE["STORAGE LAYER"]
direction LR
CHROMA[("ChromaDBVector Database")]
COLLECTIONS["Collectionssecurity | ctflogscale | development"]
CHROMA --- COLLECTIONS
end
subgraph EMBED["EMBEDDINGS (In-Process)"]
FASTEMBED["FastEmbed ONNXBAAI/bge-small-en-v1.5(384D, CPU or GPU)"]
CROSSENC["Cross-Encoderms-marco-MiniLM-L-6-v2"]
FASTEMBED --- CROSSENC
end
subgraph INGEST["DOCUMENT INGESTION"]
PARSERS["20 ParsersMD | PDF | TXT | PY | C | H | CPP | JS | JSX | TS | TSX | JSON | XML | CSVDOCX | XLSX | PPTX | IPYNB | MQH | MQ4"]
CHUNKER["ChunkingMD: section-awareOther: 1000 chars + 200 overlap"]
PARSERS --> CHUNKER
end
CLAUDE["Claude Code"] --> MCP
MCP --> SEARCH
SEARCH --> STORAGE
STORAGE --> EMBED
INGEST --> EMBED
EMBED --> STORAGE
Query Processing Flow
flowchart TB
QUERY["User Query'mimikatz credential dump'"] --> EXPAND
subgraph EXPANSION["Query Expansion"]
EXPAND["Synonym Expansionmimikatz -> mimikatz, sekurlsa, logonpasswords"]
end
EXPAND --> ROUTER
subgraph ROUTING["Keyword Routing"]
ROUTER["Keyword Router"]
MATCH{"Word BoundaryMatch?"}
CATEGORY["Filter: redteam"]
NOFILTER["No Filter"]
ROUTER --> MATCH
MATCH -->|Yes| CATEGORY
MATCH -->|No| NOFILTER
end
subgraph HYBRID["Hybrid Search"]
direction LR
SEMANTIC["Semantic Search(ChromaDB embeddings)Conceptual similarity"]
BM25["BM25 Inverted-Index(posting lists + numpy top-k)Exact term matching"]
end
subgraph FUSION["Result Fusion + Reranking"]
RRF["Reciprocal Rank Fusionscore = alpha * 1/(k+rank_sem)+ (1-alpha) * 1/(k+rank_bm25)"]
RERANK["Cross-Encoder RerankerRe-scores top 3x candidatesquery+doc pair scoring"]
SORT["Sort by Reranker ScoreNormalize to 0-1"]
ADJ["Adjacent Chunk Expansion(batch fetch ±1 chunk)"]
RRF --> RERANK --> SORT --> ADJ
end
subgraph OUTPUT["Output Processing"]
MINSCORE["min_score Filter(discard below threshold)"]
SNIPPET["snippet_mode Truncation(~500 chars at natural break)"]
MINSCORE --> SNIPPET
end
CATEGORY --> HYBRID
NOFILTER --> HYBRID
SEMANTIC --> RRF
BM25 --> RRF
ADJ --> MINSCORE
SNIPPET --> RESULTS["Resultssearch_method: hybrid|semantic|keywordscore + filtered_by_score + content_length"]
Document Ingestion Flow
flowchart LR
subgraph INPUT["Input"]
FILES["documents/├── security/├── development/├── ctf/└── general/"]
end
subgraph PARSE["Parse (20 formats)"]
MD["Markdown"]
PDF["PDF(PyMuPDF)"]
OFFICE["DOCX | XLSXPPTX | CSV"]
CODE["PY | C | H | CPP | JS | JSXTS | TSX | JSON | XML | IPYNB"]
end
subgraph CHUNK["Chunk"]
MDSPLIT["MD: Section-AwareSplit at ## headers"]
TXTSPLIT["Other: Fixed-Size1000 chars + 200 overlap"]
DEDUP["SHA256 DedupSkip duplicate content"]
end
subgraph EMBED["Embed"]
FASTEMBED["FastEmbed ONNXbge-small-en-v1.5(384D, CPU or GPU)"]
end
subgraph STORE["Store"]
CHROMADB[("ChromaDB")]
BM25IDX["BM25 Index"]
end
FILES --> MD & PDF & OFFICE & CODE
MD --> MDSPLIT
PDF & OFFICE & CODE --> TXTSPLIT
MDSPLIT --> DEDUP
TXTSPLIT --> DEDUP
DEDUP --> EMBED
EMBED --> STORE
hybrid_alpha Parameter Effect
flowchart LR
subgraph ALPHA["hybrid_alpha values"]
A0["0.0Pure BM25Instant"]
A3["0.3 (default)Keyword-heavyFast"]
A5["0.5Balanced"]
A7["0.7Semantic-heavy"]
A10["1.0Pure Semantic"]
end
subgraph USE["Best For"]
U0["CVEs, tool namesexact matches"]
U3["Technical queriesspecific terms"]
U5["General queries"]
U7["Conceptual queriesrelated topics"]
U10["'How to...' questionsconceptual search"]
end
A0 --- U0
A3 --- U3
A5 --- U5
A7 --- U7
A10 --- U10
Installation
Prerequisites
- Python 3.11+
- Claude Code CLI
- …or any other MCP client (Claude Desktop, Cursor, VS Code, Antigravity, opencode, Windsurf) — see [Use with other MCP clients](#use-with-other-mcp-clients)
- ~200MB disk for model cache (auto-downloaded on first run)
- Optional: NVIDIA GPU + CUDA 12 for accelerated embeddings (see [GPU Acceleration](#gpu-acceleration) below)
GPU Acceleration
GPU mode accelerates embedding generation during indexing and search. It requires an NVIDIA GPU with CUDA 12 support. No GPU? No problem — the server runs on CPU by default and GPU is entirely optional.
Requirements:
| Component | Minimum | How to check / get it | |------
…
Source & license
This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: lyonzin
- Source: lyonzin/knowledge-rag
- License: MIT
- Homepage: https://pypi.org/project/knowledge-rag/
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.