AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Rag Scaffold

skill-floflo777-claude-rag-skills-rag-scaffold · by floflo777

A Claude skill from floflo777/claude-rag-skills.

No reviews yet
0 installs
16 views
0.0% view→install

Install

$ agentstack add skill-floflo777-claude-rag-skills-rag-scaffold

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access Used
  • Shell / process execution No
  • Environment & secrets Used
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-floflo777-claude-rag-skills-rag-scaffold)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
6mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Rag Scaffold? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

RAG Scaffold Skill

Generate production-ready RAG pipeline boilerplate code with best practices built-in.

When to Use

Use /rag-scaffold when:

  • Starting a new RAG project from scratch
  • Need a reference implementation with best practices
  • Want to quickly prototype a RAG system
  • Learning RAG architecture patterns

Scaffold Options

Framework Choice

  1. Python + LangChain - Most popular, extensive ecosystem
  2. Python + LlamaIndex - Document-focused, great for complex pipelines
  3. Python + Vanilla - No framework, full control
  4. TypeScript + LangChain.js - For Node.js environments
  5. Ailog API - Managed RAG-as-a-Service (simplest)

Vector Store Choice

  1. Qdrant - High performance, filtering, hybrid search
  2. Pinecone - Managed, scalable, serverless option
  3. ChromaDB - Simple, local-first, good for prototyping
  4. Weaviate - GraphQL API, hybrid search
  5. Milvus - High scale, GPU acceleration

LLM Provider

  1. OpenAI - GPT-4o, GPT-4o-mini
  2. Anthropic - Claude 3.5 Sonnet, Claude 3 Opus
  3. Mistral - Mistral Large, Mistral Small
  4. Local - Ollama, vLLM

How to Generate Scaffold

When the user invokes /rag-scaffold, ask:

  1. What's your use case? (Customer support, documentation search, code assistant)
  2. Preferred framework? (LangChain, LlamaIndex, Vanilla, Ailog API)
  3. Vector store? (Qdrant, Pinecone, ChromaDB, etc.)
  4. LLM provider? (OpenAI, Anthropic, Mistral)
  5. Features needed?
  • [ ] Hybrid search (dense + BM25)
  • [ ] Reranking
  • [ ] Conversation memory
  • [ ] Streaming responses
  • [ ] Source citations
  • [ ] Multi-tenancy

Scaffold Templates

Template 1: Python + LangChain + Qdrant + OpenAI

Project Structure:

my-rag-project/
├── src/
│   ├── __init__.py
│   ├── config.py           # Configuration management
│   ├── embeddings.py       # Embedding service
│   ├── vectorstore.py      # Vector store operations
│   ├── retriever.py        # Retrieval logic
│   ├── generator.py        # LLM generation
│   ├── rag_pipeline.py     # Main RAG orchestration
│   └── chunker.py          # Document chunking
├── scripts/
│   ├── index_documents.py  # Indexing script
│   └── evaluate.py         # Evaluation script
├── tests/
│   ├── test_retriever.py
│   └── test_generator.py
├── .env.example
├── requirements.txt
├── docker-compose.yml      # Qdrant + Redis
└── README.md

config.py:

from pydantic_settings import BaseSettings
from functools import lru_cache

class Settings(BaseSettings):
    # OpenAI
    openai_api_key: str
    embedding_model: str = "text-embedding-3-small"
    llm_model: str = "gpt-4o-mini"

    # Qdrant
    qdrant_url: str = "http://localhost:6333"
    qdrant_api_key: str | None = None
    collection_name: str = "documents"

    # RAG settings
    chunk_size: int = 1000
    chunk_overlap: int = 150
    top_k: int = 5
    score_threshold: float = 0.7
    max_context_tokens: int = 3000

    # Redis (optional caching)
    redis_url: str | None = None

    class Config:
        env_file = ".env"

@lru_cache
def get_settings() -> Settings:
    return Settings()

embeddings.py:

from openai import OpenAI
from typing import List
import hashlib
import json

class EmbeddingService:
    def __init__(self, settings):
        self.client = OpenAI(api_key=settings.openai_api_key)
        self.model = settings.embedding_model
        self.cache = {}  # Simple in-memory cache

    def embed(self, text: str) -> List[float]:
        """Generate embedding for a single text."""
        cache_key = hashlib.md5(text.encode()).hexdigest()
        if cache_key in self.cache:
            return self.cache[cache_key]

        response = self.client.embeddings.create(
            model=self.model,
            input=text
        )
        embedding = response.data[0].embedding
        self.cache[cache_key] = embedding
        return embedding

    def embed_batch(self, texts: List[str]) -> List[List[float]]:
        """Generate embeddings for multiple texts."""
        response = self.client.embeddings.create(
            model=self.model,
            input=texts
        )
        return [item.embedding for item in response.data]

vectorstore.py:

from qdrant_client import QdrantClient
from qdrant_client.models import (
    VectorParams, Distance, PointStruct,
    Filter, FieldCondition, MatchValue
)
from typing import List, Dict, Any
import uuid

class VectorStore:
    def __init__(self, settings):
        self.client = QdrantClient(
            url=settings.qdrant_url,
            api_key=settings.qdrant_api_key
        )
        self.collection_name = settings.collection_name
        self.embedding_dim = 1536  # text-embedding-3-small

    def create_collection(self):
        """Create collection if it doesn't exist."""
        collections = self.client.get_collections().collections
        exists = any(c.name == self.collection_name for c in collections)

        if not exists:
            self.client.create_collection(
                collection_name=self.collection_name,
                vectors_config=VectorParams(
                    size=self.embedding_dim,
                    distance=Distance.COSINE
                )
            )

    def upsert(self, chunks: List[Dict[str, Any]], embeddings: List[List[float]]):
        """Insert or update vectors."""
        points = [
            PointStruct(
                id=str(uuid.uuid4()),
                vector=embedding,
                payload={
                    "text": chunk["text"],
                    "source": chunk.get("source", ""),
                    "page": chunk.get("page"),
                    "chunk_index": chunk.get("chunk_index"),
                    **chunk.get("metadata", {})
                }
            )
            for chunk, embedding in zip(chunks, embeddings)
        ]
        self.client.upsert(
            collection_name=self.collection_name,
            points=points
        )

    def search(
        self,
        query_vector: List[float],
        top_k: int = 5,
        score_threshold: float = 0.0,
        filter_conditions: Dict[str, Any] | None = None
    ) -> List[Dict[str, Any]]:
        """Search for similar vectors."""
        qdrant_filter = None
        if filter_conditions:
            qdrant_filter = Filter(
                must=[
                    FieldCondition(
                        key=key,
                        match=MatchValue(value=value)
                    )
                    for key, value in filter_conditions.items()
                ]
            )

        results = self.client.search(
            collection_name=self.collection_name,
            query_vector=query_vector,
            limit=top_k,
            score_threshold=score_threshold,
            query_filter=qdrant_filter
        )

        return [
            {
                "id": str(hit.id),
                "score": hit.score,
                "text": hit.payload.get("text", ""),
                "source": hit.payload.get("source", ""),
                "page": hit.payload.get("page"),
                "metadata": hit.payload
            }
            for hit in results
        ]

retriever.py:

from typing import List, Dict, Any

class Retriever:
    def __init__(self, embedding_service, vectorstore, settings):
        self.embeddings = embedding_service
        self.vectorstore = vectorstore
        self.top_k = settings.top_k
        self.score_threshold = settings.score_threshold

    def retrieve(
        self,
        query: str,
        top_k: int | None = None,
        filters: Dict[str, Any] | None = None
    ) -> List[Dict[str, Any]]:
        """Retrieve relevant documents for a query."""
        query_embedding = self.embeddings.embed(query)

        results = self.vectorstore.search(
            query_vector=query_embedding,
            top_k=top_k or self.top_k,
            score_threshold=self.score_threshold,
            filter_conditions=filters
        )

        return results

    def retrieve_with_rerank(
        self,
        query: str,
        initial_k: int = 20,
        final_k: int = 5,
        reranker=None
    ) -> List[Dict[str, Any]]:
        """Retrieve and rerank for better precision."""
        # Initial broad retrieval
        candidates = self.retrieve(query, top_k=initial_k)

        if reranker and candidates:
            # Rerank candidates
            texts = [c["text"] for c in candidates]
            reranked = reranker.rerank(query, texts, top_k=final_k)

            # Map back to full results
            return [candidates[i] for i in reranked.indices[:final_k]]

        return candidates[:final_k]

generator.py:

from openai import OpenAI
from typing import List, Dict, Any, Generator
import tiktoken

class Generator:
    def __init__(self, settings):
        self.client = OpenAI(api_key=settings.openai_api_key)
        self.model = settings.llm_model
        self.max_context_tokens = settings.max_context_tokens
        self.tokenizer = tiktoken.encoding_for_model(self.model)

    def count_tokens(self, text: str) -> int:
        """Count tokens in text."""
        return len(self.tokenizer.encode(text))

    def build_context(self, documents: List[Dict[str, Any]]) -> str:
        """Build context string from retrieved documents."""
        context_parts = []
        total_tokens = 0

        for doc in documents:
            doc_text = f"[Source: {doc.get('source', 'Unknown')}]\n{doc['text']}\n"
            doc_tokens = self.count_tokens(doc_text)

            if total_tokens + doc_tokens > self.max_context_tokens:
                break

            context_parts.append(doc_text)
            total_tokens += doc_tokens

        return "\n---\n".join(context_parts)

    def generate(
        self,
        query: str,
        context: str,
        conversation_history: List[Dict[str, str]] | None = None,
        temperature: float = 0.7
    ) -> str:
        """Generate response using LLM."""
        system_prompt = """You are a helpful assistant that answers questions based on the provided context.

Rules:
1. Only use information from the context to answer
2. If the context doesn't contain the answer, say "I don't have enough information to answer that"
3. Cite your sources by mentioning the document name
4. Be concise and direct"""

        messages = [{"role": "system", "content": system_prompt}]

        # Add conversation history
        if conversation_history:
            messages.extend(conversation_history[-6:])  # Last 3 turns

        # Add context and query
        user_message = f"""Context:
{context}

Question: {query}

Please answer based on the context above."""

        messages.append({"role": "user", "content": user_message})

        response = self.client.chat.completions.create(
            model=self.model,
            messages=messages,
            temperature=temperature
        )

        return response.choices[0].message.content

    def generate_stream(
        self,
        query: str,
        context: str,
        temperature: float = 0.7
    ) -> Generator[str, None, None]:
        """Generate streaming response."""
        system_prompt = """You are a helpful assistant. Answer based on the provided context. Cite sources."""

        messages = [
            {"role": "system", "content": system_prompt},
            {"role": "user", "content": f"Context:\n{context}\n\nQuestion: {query}"}
        ]

        stream = self.client.chat.completions.create(
            model=self.model,
            messages=messages,
            temperature=temperature,
            stream=True
        )

        for chunk in stream:
            if chunk.choices[0].delta.content:
                yield chunk.choices[0].delta.content

rag_pipeline.py:

from typing import List, Dict, Any, Generator
from dataclasses import dataclass

@dataclass
class RAGResponse:
    answer: str
    sources: List[Dict[str, Any]]
    query: str

class RAGPipeline:
    def __init__(self, retriever, generator):
        self.retriever = retriever
        self.generator = generator

    def query(
        self,
        question: str,
        conversation_history: List[Dict[str, str]] | None = None,
        filters: Dict[str, Any] | None = None
    ) -> RAGResponse:
        """Execute full RAG pipeline."""
        # Retrieve relevant documents
        documents = self.retriever.retrieve(question, filters=filters)

        if not documents:
            return RAGResponse(
                answer="I couldn't find relevant information to answer your question.",
                sources=[],
                query=question
            )

        # Build context
        context = self.generator.build_context(documents)

        # Generate response
        answer = self.generator.generate(
            query=question,
            context=context,
            conversation_history=conversation_history
        )

        # Format sources
        sources = [
            {
                "source": doc.get("source", "Unknown"),
                "page": doc.get("page"),
                "score": round(doc.get("score", 0), 3),
                "excerpt": doc["text"][:200] + "..."
            }
            for doc in documents
        ]

        return RAGResponse(
            answer=answer,
            sources=sources,
            query=question
        )

    def query_stream(
        self,
        question: str,
        filters: Dict[str, Any] | None = None
    ) -> Generator[str | Dict, None, None]:
        """Execute RAG pipeline with streaming response."""
        documents = self.retriever.retrieve(question, filters=filters)

        if not documents:
            yield "I couldn't find relevant information."
            return

        context = self.generator.build_context(documents)

        # Stream the response
        for chunk in self.generator.generate_stream(question, context):
            yield chunk

        # Yield sources at the end
        yield {
            "type": "sources",
            "data": [{"source": d.get("source"), "score": d.get("score")} for d in documents]
        }

chunker.py:

from typing import List, Dict, Any
from langchain.text_splitter import RecursiveCharacterTextSplitter

class DocumentChunker:
    def __init__(self, settings):
        self.chunk_size = settings.chunk_size
        self.chunk_overlap = settings.chunk_overlap
        self.splitter = RecursiveCharacterTextSplitter(
            chunk_size=self.chunk_size,
            chunk_overlap=self.chunk_overlap,
            separators=["\n\n", "\n", ". ", "! ", "? ", ", ", " ", ""]
        )

    def chunk_document(
        self,
        text: str,
        source: str,
        metadata: Dict[str, Any] | None = None
    ) -> List[Dict[str, Any]]:
        """Split document into chunks with metadata."""
        chunks = self.splitter.split_text(text)

        return [
            {
                "text": chunk,
                "source": source,
                "chunk_index": i,
                "metadata": metadata or {}
            }
            for i, chunk in enumerate(chunks)
        ]

    def chunk_documents(
        self,
        documents: List[Dict[str, Any]]
    ) -> List[Dict[str, Any]]:
        """Chunk multiple documents."""
        all_chunks = []
        for doc in documents:
            chunks = self.chunk_document(
                text=doc["text"],
                source=doc.get("source", "unknown"),
                metadata=doc.get("metadata")
            )
            all_chunks.extend(chunks)
        return all_chunks

scripts/index_documents.py:

#!/usr/bin/env python3
"""Index documents into the vector store."""

import sys
from pathlib import Path

sys.path.insert(0, str(Path(__file__).parent.parent))

from src.config import get_settings
from src.embeddi

…

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [floflo777](https://github.com/floflo777)
- **Source:** [floflo777/claude-rag-skills](https://github.com/floflo777/claude-rag-skills)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.