Install
$ agentstack add skill-stefan-jansen-claude-code-toolkit-rag-implementation ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
RAG Implementation Patterns
Comprehensive guide to implementing Retrieval-Augmented Generation (RAG) systems including vector database selection, chunking strategies, embedding models, retrieval optimization, and production deployment patterns.
Quick Reference
When to use this skill:
- Building RAG/semantic search systems
- Implementing document retrieval pipelines
- Optimizing vector database performance
- Debugging retrieval quality issues
- Choosing between vector database options
- Designing chunking strategies
- Implementing hybrid search
Technologies covered:
- Vector DBs: Qdrant, Pinecone, Chroma, Weaviate, Milvus
- Embeddings: OpenAI, Sentence Transformers, Cohere
- Frameworks: LangChain, LlamaIndex, Haystack
Part 1: Vector Database Selection
Database Comparison Matrix
| Database | Best For | Deployment | Performance | Cost | |----------|----------|------------|-------------|------| | Qdrant | Self-hosted, production | Docker/K8s | Excellent (Rust) | Free (self-host) | | Pinecone | Managed, rapid prototyping | Cloud | Excellent | Pay-per-use | | Chroma | Local development, embedded | In-process | Good (Python) | Free | | Weaviate | Complex schemas, GraphQL | Docker/Cloud | Excellent (Go) | Free + Cloud | | Milvus | Large-scale, distributed | K8s | Excellent (C++) | Free (self-host) |
Qdrant Setup (Recommended for Production)
from qdrant_client import QdrantClient
from qdrant_client.models import Distance, VectorParams, PointStruct
# Initialize client (local or cloud)
client = QdrantClient(url="http://localhost:6333") # or cloud URL
# Create collection
client.create_collection(
collection_name="documents",
vectors_config=VectorParams(
size=1536, # OpenAI text-embedding-3-small dimension
distance=Distance.COSINE # or DOT, EUCLID
)
)
# Insert vectors with payload
client.upsert(
collection_name="documents",
points=[
PointStruct(
id=1,
vector=[0.1, 0.2, ...], # 1536 dimensions
payload={
"text": "Document content",
"source": "doc.pdf",
"page": 1,
"metadata": {...}
}
)
]
)
# Search
results = client.search(
collection_name="documents",
query_vector=[0.1, 0.2, ...],
limit=5,
score_threshold=0.7 # Minimum similarity
)
Pinecone Setup (Managed Service)
from pinecone import Pinecone, ServerlessSpec
# Initialize
pc = Pinecone(api_key="your-key")
# Create index
pc.create_index(
name="documents",
dimension=1536,
metric="cosine",
spec=ServerlessSpec(cloud="aws", region="us-east-1")
)
# Get index
index = pc.Index("documents")
# Upsert vectors
index.upsert(vectors=[
("doc1", [0.1, 0.2, ...], {"text": "...", "source": "..."})
])
# Query
results = index.query(
vector=[0.1, 0.2, ...],
top_k=5,
include_metadata=True
)
Part 2: Chunking Strategies
Strategy 1: Fixed-Size Chunking (Simple, Fast)
def fixed_size_chunking(text: str, chunk_size: int = 512, overlap: int = 50) -> list[str]:
"""
Split text into fixed-size chunks with overlap.
Pros: Simple, predictable chunk sizes
Cons: May break mid-sentence, poor semantic boundaries
"""
words = text.split()
chunks = []
for i in range(0, len(words), chunk_size - overlap):
chunk = ' '.join(words[i:i + chunk_size])
chunks.append(chunk)
return chunks
# Usage
chunks = fixed_size_chunking(document, chunk_size=512, overlap=50)
When to use:
- Simple documents (logs, transcripts)
- Prototyping/MVP
- Consistent token budgets needed
Strategy 2: Semantic Chunking (Better Quality)
from langchain.text_splitter import RecursiveCharacterTextSplitter
def semantic_chunking(text: str, chunk_size: int = 1000, overlap: int = 200) -> list[str]:
"""
Split on semantic boundaries (paragraphs, sentences).
Pros: Preserves meaning, better retrieval quality
Cons: Variable chunk sizes, slower processing
"""
splitter = RecursiveCharacterTextSplitter(
chunk_size=chunk_size,
chunk_overlap=overlap,
separators=["\n\n", "\n", ". ", " ", ""], # Priority order
length_function=len
)
return splitter.split_text(text)
# Usage
chunks = semantic_chunking(document, chunk_size=1000, overlap=200)
When to use:
- Long-form documents (articles, books, reports)
- Quality > speed
- Natural language content
Strategy 3: Hierarchical Chunking (Best for Structured Docs)
def hierarchical_chunking(document: dict) -> list[dict]:
"""
Chunk based on document structure (sections, subsections).
Pros: Preserves hierarchy, enables parent-child retrieval
Cons: Requires structured input, more complex
"""
chunks = []
for section in document['sections']:
# Parent chunk (section summary)
chunks.append({
'text': section['title'] + '\n' + section['summary'],
'type': 'parent',
'section_id': section['id']
})
# Child chunks (paragraphs)
for para in section['paragraphs']:
chunks.append({
'text': para,
'type': 'child',
'parent_id': section['id']
})
return chunks
When to use:
- Technical documentation
- Books with TOC
- Legal documents
- Need to preserve context hierarchy
Strategy 4: Sliding Window (Maximum Context Preservation)
def sliding_window_chunking(text: str, window_size: int = 512, stride: int = 256) -> list[str]:
"""
Overlapping windows for maximum context.
Pros: No information loss at boundaries
Cons: Storage overhead (duplicate content)
"""
words = text.split()
chunks = []
for i in range(0, len(words) - window_size + 1, stride):
chunk = ' '.join(words[i:i + window_size])
chunks.append(chunk)
return chunks
When to use:
- Critical retrieval accuracy needed
- Short queries need broader context
- Storage cost not a concern
Part 3: Embedding Models
Model Selection Guide
| Model | Dimensions | Speed | Quality | Cost | Use Case | |-------|-----------|-------|---------|------|----------| | OpenAI text-embedding-3-small | 1536 | Fast | Excellent | $0.02/1M tokens | Production, general purpose | | OpenAI text-embedding-3-large | 3072 | Medium | Best | $0.13/1M tokens | High-quality retrieval | | all-MiniLM-L6-v2 | 384 | Very fast | Good | Free | Self-hosted, prototyping | | all-mpnet-base-v2 | 768 | Fast | Very good | Free | Self-hosted, quality | | Cohere embed-english-v3.0 | 1024 | Fast | Excellent | $0.10/1M tokens | Semantic search focus |
OpenAI Embeddings (Recommended)
from openai import OpenAI
client = OpenAI(api_key="your-key")
def get_embeddings(texts: list[str], model: str = "text-embedding-3-small") -> list[list[float]]:
"""
Generate embeddings using OpenAI.
Batch size: Up to 2048 inputs per request
Rate limits: Check tier limits
"""
response = client.embeddings.create(
model=model,
input=texts
)
return [item.embedding for item in response.data]
# Usage
chunks = ["chunk 1", "chunk 2", ...]
embeddings = get_embeddings(chunks)
Sentence Transformers (Self-Hosted)
from sentence_transformers import SentenceTransformer
# Load model (cached after first download)
model = SentenceTransformer('all-MiniLM-L6-v2')
def get_embeddings_local(texts: list[str]) -> list[list[float]]:
"""
Generate embeddings locally (no API costs).
GPU recommended for batches > 100
CPU acceptable for small batches
"""
return model.encode(texts, show_progress_bar=True).tolist()
# Usage
embeddings = get_embeddings_local(chunks)
Part 4: Retrieval Optimization
Technique 1: Hybrid Search (Dense + Sparse)
from qdrant_client.models import Filter, FieldCondition, MatchValue
def hybrid_search(query: str, query_vector: list[float], top_k: int = 10):
"""
Combine dense (vector) and sparse (keyword) search.
Dense: Semantic similarity
Sparse: Exact keyword matches
"""
# Dense search
dense_results = client.search(
collection_name="documents",
query_vector=query_vector,
limit=top_k * 2 # Get more candidates
)
# Sparse search (BM25 via metadata)
sparse_results = client.search(
collection_name="documents",
query_filter=Filter(
must=[
FieldCondition(
key="text",
match=MatchValue(value=query)
)
]
),
limit=top_k * 2
)
# Merge and re-rank
combined = merge_results(dense_results, sparse_results, weights=(0.7, 0.3))
return combined[:top_k]
Technique 2: Query Expansion
def expand_query(query: str) -> list[str]:
"""
Generate query variations for better recall.
Techniques:
- Synonym expansion
- Question reformulation
- Entity extraction
"""
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4",
messages=[{
"role": "system",
"content": "Generate 3 alternative phrasings of the user's query."
}, {
"role": "user",
"content": query
}]
)
expanded = [query] + response.choices[0].message.content.split('\n')
return expanded
# Usage
queries = expand_query("How to train neural networks?")
# → ["How to train neural networks?",
# "What are neural network training techniques?",
# "Neural network optimization methods",
# "Deep learning model training"]
Technique 3: Reranking
from sentence_transformers import CrossEncoder
# Load cross-encoder (better than bi-encoder for reranking)
reranker = CrossEncoder('cross-encoder/ms-marco-MiniLM-L-6-v2')
def rerank_results(query: str, results: list[dict], top_k: int = 5) -> list[dict]:
"""
Rerank initial results using cross-encoder.
More accurate but slower than initial retrieval
Use on top 20-50 candidates only
"""
# Score each query-document pair
pairs = [(query, result['text']) for result in results]
scores = reranker.predict(pairs)
# Combine scores with results
for result, score in zip(results, scores):
result['rerank_score'] = float(score)
# Sort and return top_k
reranked = sorted(results, key=lambda x: x['rerank_score'], reverse=True)
return reranked[:top_k]
Technique 4: Metadata Filtering
def filtered_search(
query_vector: list[float],
filters: dict,
top_k: int = 5
):
"""
Filter search by metadata (date, category, author, etc.)
Pre-filter: Faster but may miss results
Post-filter: More results but slower
"""
from qdrant_client.models import Filter, FieldCondition, Range
# Build filter conditions
conditions = []
if 'date_range' in filters:
conditions.append(
FieldCondition(
key="date",
range=Range(
gte=filters['date_range']['start'],
lte=filters['date_range']['end']
)
)
)
if 'category' in filters:
conditions.append(
FieldCondition(
key="category",
match=MatchValue(value=filters['category'])
)
)
# Search with filters
results = client.search(
collection_name="documents",
query_vector=query_vector,
query_filter=Filter(must=conditions) if conditions else None,
limit=top_k
)
return results
Part 5: Context Management
Pattern 1: Retrieved Context Optimization
def optimize_context(query: str, retrieved_docs: list[str], max_tokens: int = 4000) -> str:
"""
Optimize retrieved context to fit within LLM context window.
Strategies:
1. Relevance-based truncation
2. Extractive summarization
3. Overlap removal
"""
# Sort by relevance
sorted_docs = sorted(retrieved_docs, key=lambda d: d['score'], reverse=True)
# Build context within token budget
context_parts = []
total_tokens = 0
for doc in sorted_docs:
doc_tokens = estimate_tokens(doc['text'])
if total_tokens + doc_tokens dict:
"""
Generate answer with citation tracking.
Returns:
- answer: Generated text
- citations: List of source documents used
"""
from openai import OpenAI
client = OpenAI()
# Create source map
source_map = {i+1: source for i, source in enumerate(sources)}
numbered_context = "\n\n".join([
f"[{i+1}] {source['text']}"
for i, source in enumerate(sources)
])
response = client.chat.completions.create(
model="gpt-4",
messages=[{
"role": "system",
"content": "Answer using the provided sources. Cite sources as [1], [2], etc."
}, {
"role": "user",
"content": f"Context:\n{numbered_context}\n\nQuestion: {query}"
}]
)
answer = response.choices[0].message.content
# Extract citations from answer
import re
cited_nums = set(map(int, re.findall(r'\[(\d+)\]', answer)))
cited_sources = [source_map[num] for num in cited_nums if num in source_map]
return {
'answer': answer,
'citations': cited_sources,
'num_sources_used': len(cited_sources)
}
Part 6: Production Best Practices
Caching Strategy
from functools import lru_cache
import hashlib
class EmbeddingCache:
"""Cache embeddings to avoid recomputation."""
def __init__(self, cache_size: int = 10000):
self.cache = {}
self.max_size = cache_size
def get_or_compute(self, text: str, embed_fn) -> list[float]:
# Create cache key
key = hashlib.sha256(text.encode()).hexdigest()
if key in self.cache:
return self.cache[key]
# Compute and cache
embedding = embed_fn(text)
if len(self.cache) >= self.max_size:
# Evict oldest (FIFO)
self.cache.pop(next(iter(self.cache)))
self.cache[key] = embedding
return embedding
# Usage
cache = EmbeddingCache()
embedding = cache.get_or_compute(text, lambda t: get_embeddings([t])[0])
Async Processing
import asyncio
from typing import List
async def process_documents_async(documents: List[str], batch_size: int = 100):
"""
Process large document sets asynchronously.
Benefits:
- 10-50x faster for I/O-bound operations
- Better resource utilization
- Scalable to millions of documents
"""
async def process_batch(batch):
embeddings = await get_embeddings_async(batch)
await upsert_to_db_async(batch, embeddings)
# Split into batches
batches = [documents[i:i+batch_size] for i in range(0, len(documents), batch_size)]
# Process batches concurrently
await asyncio.gather(*[process_batch(batch) for batch in batches])
# Usage
asyncio.run(process_documents_async(documents))
Monitoring & Observability
import time
from dataclasses import dataclass
from datetime import datetime
@dataclass
class RAGMetrics:
"""Track RAG system performance."""
query_count: int = 0
avg_retrieval_time: float = 0.0
avg_generation_time: float = 0.0
cache_hit_rate: float = 0.0
avg_num_results: float = 0.0
class RAGMonitor:
def __init__(self):
self.metrics = RAGMetrics()
self.query_times = []
def log_query(self, retrieval_time: float, generation_time: float, num_results: int):
self.metrics.que
…
## Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- **Author:** [stefan-jansen](https://github.com/stefan-jansen)
- **Source:** [stefan-jansen/claude-code-toolkit](https://github.com/stefan-jansen/claude-code-toolkit)
- **License:** MIT
- **Homepage:** https://www.applied-ai.com/open-source/claude-code-toolkit/
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.