AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Ai Ml Integration Patterns

skill-mickeyyaya-refactoring-skills-ai-ml-integration-patterns · by mickeyyaya

Use when designing or reviewing AI/ML integration code — covers RAG pipeline design, prompt engineering, structured output with schema validation, tool use/function calling, LLM error handling, token budget management, hallucination mitigation, and AI/ML anti-patterns with examples in TypeScript and Python

No reviews yet
0 installs
17 views
0.0% view→install

Install

$ agentstack add skill-mickeyyaya-refactoring-skills-ai-ml-integration-patterns

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution Used

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-mickeyyaya-refactoring-skills-ai-ml-integration-patterns)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
5mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Ai Ml Integration Patterns? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

AI/ML Integration Patterns

Overview

Integrating LLMs into production systems introduces failure modes unique to probabilistic outputs: hallucinated facts, unbounded token costs, prompt injection attacks, and unvalidated structured responses. Use this guide when designing, building, or reviewing code that calls LLM APIs, builds RAG pipelines, or orchestrates AI agents.

When to use: Reviewing code that calls OpenAI, Anthropic, or other LLM APIs; evaluating RAG pipeline design; auditing prompt construction; any system that relies on LLM-generated structured output or tool execution loops.

Quick Reference

| Pattern | Core Idea | Primary Red Flag | |---------|-----------|-----------------| | RAG Pipeline | Ground LLM answers in retrieved documents | Retrieval without relevance filtering; missing context window budget | | Prompt Engineering | Structured prompts for reliable, reproducible outputs | Hardcoded prompts scattered in code; no version control | | Structured Output + Schema | Validate LLM JSON against a schema before using it | Trusting raw LLM output as typed data | | Tool Use / Function Calling | LLM selects and invokes registered tools; app executes | Executing tool calls without validating arguments | | LLM Error Handling | Retry rate limits, fall back on model failures, timeout on hangs | No retry on 429; no timeout on streaming calls | | Token Budget Management | Count, chunk, and truncate to stay within context limits | Unlimited context assembly; no chunk size cap | | Hallucination Mitigation | Source citation, confidence scoring, guardrails | LLM answers used directly with no grounding check | | Anti-Patterns | Common misuse patterns that cause prod failures | Prompt injection via user input; no output validation |


Patterns in Detail

1. RAG (Retrieval-Augmented Generation) Pipeline

RAG grounds LLM responses in real data by retrieving relevant documents at query time and injecting them into the context window.

Pipeline stages:

  1. Embedding — convert documents and queries to dense vectors
  2. Vector store — index and persist embeddings for ANN search
  3. Retrieval — query vector store with top-k similarity
  4. Reranking (optional) — re-score candidates with a cross-encoder
  5. Context window management — pack retrieved chunks within token budget
  6. Generation — LLM answers using grounded context

Red Flags:

  • Retrieval without relevance threshold — irrelevant chunks injected into context
  • No deduplication of retrieved chunks — redundant tokens waste context budget
  • Embedding model mismatch between indexing and query time
  • Context window assembly ignores token count — exceeds model limit at runtime
import { OpenAI } from "openai";
import { encode } from "gpt-tokenizer";

const openai = new OpenAI();

interface Chunk { id: string; text: string; score: number; source: string; }

async function buildRagContext(
  query: string,
  retrievedChunks: Chunk[],
  maxContextTokens = 3000
): Promise {
  // Filter by relevance threshold before packing
  const relevant = retrievedChunks.filter(c => c.score >= 0.75);

  const packed: string[] = [];
  const sources: string[] = [];
  let tokenCount = 0;

  for (const chunk of relevant) {
    const chunkTokens = encode(chunk.text).length;
    if (tokenCount + chunkTokens > maxContextTokens) break;
    packed.push(`[Source: ${chunk.source}]\n${chunk.text}`);
    sources.push(chunk.source);
    tokenCount += chunkTokens;
  }

  return { context: packed.join("\n\n---\n\n"), sources };
}

async function ragQuery(query: string, chunks: Chunk[]): Promise {
  const { context, sources } = await buildRagContext(query, chunks);

  const response = await openai.chat.completions.create({
    model: "gpt-4o",
    messages: [
      {
        role: "system",
        content: "Answer using ONLY the provided context. If the answer is not in the context, say 'I don't have enough information.'"
      },
      {
        role: "user",
        content: `Context:\n${context}\n\nQuestion: ${query}`
      }
    ],
    temperature: 0,
  });

  const answer = response.choices[0].message.content ?? "";
  return `${answer}\n\nSources: ${[...new Set(sources)].join(", ")}`;
}

Cross-reference: data-pipeline-patterns — chunking and embedding pipeline stages.


2. Prompt Engineering Patterns

Reliable LLM outputs require structured, versioned prompts — not ad-hoc string concatenation.

Patterns:

  • System prompt — define role, constraints, and output format once
  • Few-shot examples — show 2-5 input/output pairs to anchor behavior
  • Chain-of-thought — instruct the model to reason before answering ("Think step by step")
  • Structured output instruction — embed JSON schema or format spec in the prompt
  • Prompt versioning — store prompts in config or database with version identifiers

Red Flags:

  • Prompts built with string concatenation scattered across the codebase
  • User input injected directly into system prompts (prompt injection risk)
  • No few-shot examples for complex classification or extraction tasks
  • Chain-of-thought reasoning mixed with final output — parse errors at runtime
// BEFORE — ad-hoc, scattered, unversioned
const prompt = `You are a helpful assistant. User said: ${userMessage}`;

// AFTER — versioned prompt template with separation of concerns
interface PromptTemplate {
  version: string;
  systemPrompt: string;
  fewShotExamples: Array;
}

const SENTIMENT_PROMPT_V2: PromptTemplate = {
  version: "2.0",
  systemPrompt: [
    "You are a sentiment analysis engine.",
    "Classify the sentiment of the input as: positive, negative, or neutral.",
    "Respond ONLY with valid JSON: {\"sentiment\": \"\", \"confidence\": }",
    "Do not include any explanation or additional text."
  ].join("\n"),
  fewShotExamples: [
    { input: "I love this product!", output: '{"sentiment":"positive","confidence":0.98}' },
    { input: "This is terrible.", output: '{"sentiment":"negative","confidence":0.95}' },
    { input: "It works.", output: '{"sentiment":"neutral","confidence":0.80}' },
  ],
};

function buildMessages(template: PromptTemplate, userInput: string) {
  return [
    { role: "system" as const, content: template.systemPrompt },
    // few-shot examples as alternating user/assistant turns
    ...template.fewShotExamples.flatMap(ex => [
      { role: "user" as const, content: ex.input },
      { role: "assistant" as const, content: ex.output },
    ]),
    { role: "user" as const, content: userInput },
  ];
}

Python — chain-of-thought (XML tag parsing idiom differs from TS):

COT_SYSTEM_PROMPT = """You are a math reasoning assistant.
Think through each problem step by step, then provide your final answer.

Format your response as:

[your step-by-step thinking]

[final answer only]
"""

import re

def extract_cot_answer(response: str) -> tuple[str, str]:
    """Extract reasoning and final answer from chain-of-thought response."""
    reasoning_match = re.search(r"(.*?)", response, re.DOTALL)
    answer_match = re.search(r"(.*?)", response, re.DOTALL)
    reasoning = reasoning_match.group(1).strip() if reasoning_match else ""
    answer = answer_match.group(1).strip() if answer_match else response
    return reasoning, answer

3. Structured Output with Schema Validation

LLMs do not guarantee valid JSON or correct field types. Always validate structured output against a schema before using it in application logic.

Red Flags:

  • JSON.parse(llmResponse) with no schema validation — runtime crashes on malformed output
  • Optional fields accessed without null checks — implicit trust in LLM output shape
  • No retry on parse failure — single bad response causes permanent failure
  • Schema defined only in the prompt, not enforced in code
import { z } from "zod";
import { OpenAI } from "openai";

const openai = new OpenAI();

const ProductExtractionSchema = z.object({
  name: z.string().min(1),
  price: z.number().positive(),
  currency: z.enum(["USD", "EUR", "GBP"]),
  inStock: z.boolean(),
  tags: z.array(z.string()).default([]),
});

async function extractProduct(text: string) {
  for (let attempt = 1; attempt  {
  if (name === "get_weather") {
    const args = GetWeatherArgs.parse(rawArgs);  // validate before execution
    return JSON.stringify({ temperature: 22, condition: "sunny", location: args.location });
  }
  throw new Error(`Unknown tool: ${name}`);
}

async function runAgentLoop(userMessage: string, maxIterations = 10): Promise {
  const messages: OpenAI.Chat.ChatCompletionMessageParam[] = [{ role: "user", content: userMessage }];

  for (let i = 0; i  {
  const { maxAttempts = 3, timeoutMs = 30_000, useFallback = true } = options;
  const openai = new OpenAI({ timeout: timeoutMs });
  const PRIMARY_MODEL = "gpt-4o", FALLBACK_MODEL = "gpt-4o-mini";

  for (let attempt = 1; attempt  setTimeout(r, retryAfter * 1000));
      } else if (err instanceof APIConnectionTimeoutError) {
        if (attempt === maxAttempts) throw err;
        await new Promise(r => setTimeout(r, 1000 * attempt));
      } else if (err instanceof APIError && err.status >= 400 && err.status  setTimeout(r, 500 * 2 ** (attempt - 1)));
      }
    }
  }
  throw new Error("unreachable");
}

Cross-reference: error-handling-patterns — Retry with Exponential Backoff for generic retry utilities. Cross-reference: api-rate-limiting-throttling — rate limit detection and backoff strategies.


6. Token Budget Management

Exceeding the context window causes runtime errors. Unbounded context assembly silently inflates costs.

Strategies:

  • Counting — measure token usage before sending requests
  • Truncation — trim least-relevant content to fit the budget
  • Chunking — split large documents into overlapping windows for processing
  • Priority packing — allocate tokens by priority: system prompt > recent history > retrieved context

Red Flags:

  • No token count check before assembling the final prompt
  • Entire conversation history appended — grows unbounded over multi-turn sessions
  • Chunk size set in characters, not tokens — off by 3-4x for non-ASCII content
  • Overlap between chunks ignored — sentence boundaries split mid-thought
import { encode, decode } from "gpt-tokenizer";

const MODEL_TOKEN_LIMITS: Record = { "gpt-4o": 128_000, "gpt-4o-mini": 128_000 };
const countTokens = (text: string) => encode(text).length;
const truncateToTokenBudget = (text: string, max: number) => {
  const t = encode(text); return t.length ,
  userMessage: string,
  model = "gpt-4o"
): Array {
  let budget = (MODEL_TOKEN_LIMITS[model] ?? 8_000) - 2_048
    - countTokens(systemPrompt) - countTokens(userMessage);

  const includedHistory: typeof history = [];
  for (let i = history.length - 1; i >= 0 && budget > 0; i--) {
    const t = countTokens(history[i].content);
    if (t > budget) break;
    includedHistory.unshift(history[i]);
    budget -= t;
  }
  return [{ role: "system", content: systemPrompt }, ...includedHistory, { role: "user", content: userMessage }];
}

7. Hallucination Mitigation and Grounding

LLMs generate plausible-sounding but false information. Mitigation requires architectural controls — prompting alone is insufficient.

Techniques:

  • Source citation — require the model to cite which document each claim comes from
  • Confidence scoring — ask the model to rate certainty; threshold low-confidence answers
  • Guardrails — post-process outputs to detect and block unsafe or off-topic responses
  • Grounding checks — verify claims against the retrieved context programmatically
  • Abstain instruction — explicitly instruct the model to say "I don't know" rather than guess

Red Flags:

  • LLM answers without any retrieved context — no grounding possible
  • No instruction to abstain when uncertain — model invents answers
  • Guardrails applied only at the prompt level — no post-processing safety layer
  • Citations not verified against actual source content
import { z } from "zod";

interface GroundedAnswer {
  answer: string; citations: Array;
  confidence: "high" | "medium" | "low"; abstained: boolean;
}

const GROUNDED_SYSTEM_PROMPT = `You are a factual Q&A assistant. Answer ONLY using the provided context.
Cite sources as [source-id]. If not in context, set "abstained": true. Rate confidence high/medium/low.
Respond ONLY with JSON: {"answer":string,"citations":[{"sourceId":string,"excerpt":string}],"confidence":"high"|"medium"|"low","abstained":boolean}`;

const GroundedAnswerSchema = z.object({
  answer: z.string(),
  citations: z.array(z.object({ sourceId: z.string(), excerpt: z.string() })),
  confidence: z.enum(["high", "medium", "low"]),
  abstained: z.boolean(),
});

async function groundedQuery(
  question: string,
  context: string,
  openai: import("openai").OpenAI
): Promise {
  const response = await openai.chat.completions.create({
    model: "gpt-4o",
    response_format: { type: "json_object" },
    messages: [
      { role: "system", content: GROUNDED_SYSTEM_PROMPT },
      { role: "user", content: `Context:\n${context}\n\nQuestion: ${question}` },
    ],
    temperature: 0,
  });

  const raw = JSON.parse(response.choices[0].message.content ?? "{}");
  const result = GroundedAnswerSchema.parse(raw);

  // Post-hoc guardrail: flag low-confidence answers for review
  if (result.confidence === "low" && !result.abstained) {
    console.warn("Low-confidence answer returned without abstaining", { question });
  }
  return result;
}

Python — output guardrail post-processing (regex pattern matching idiom):

import re
from dataclasses import dataclass

@dataclass
class GuardrailResult:
    safe: bool
    reason: str | None

BLOCKED_PATTERNS = [
    re.compile(r"\b(password|api[_\s]key|secret[_\s]key|token)\b", re.IGNORECASE),
    re.compile(r"\b(ssn|social.security|credit.card)\b", re.IGNORECASE),
]

def apply_output_guardrails(text: str) -> GuardrailResult:
    """Post-process LLM output before returning to caller."""
    for pattern in BLOCKED_PATTERNS:
        if pattern.search(text):
            return GuardrailResult(safe=False, reason=f"Sensitive pattern detected: {pattern.pattern}")
    return GuardrailResult(safe=True, reason=None)

Cross-reference: security-patterns-code-review — input/output sanitization and injection prevention.


8. AI/ML Anti-Patterns

| Anti-Pattern | Description | Fix | |-------------|-------------|-----| | Prompt Injection | User input injected directly into system prompt — attacker controls model behavior | Separate user input from system instructions; sanitize or encode user content | | Unbounded Context | Full conversation history appended forever — token cost grows linearly | Implement rolling window or summarization; enforce token budget per request | | No Output Validation | Raw LLM JSON used as typed data without schema check | Always validate with Zod (TS) or Pydantic (Python) before use | | Hardcoded Prompts | Prompts inline in application code — no versioning, no A/B testing | Store prompts in config, database, or prompt management system | | Retry All Errors | Retrying 400 Bad Request or 404 Not Found — permanent errors waste quota | Classify errors: transient (429, 503) vs. permanent (400, 401, 404) | | No Timeout | LLM call with no timeout — hangs indefinitely on network failure | Always set request timeout; use streaming with read timeout | | Single Model Dependency | No fallback model — one provider outage causes full outage | Define primary + fallback model; implement model router | | Ignoring Token Costs | No token counting or budget — surprise bills at end of month | Count tokens before each request; set max_tokens on all calls | | Trusting LLM for Logic | Using LLM to make security, financial, or access control decisions | LLM o

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.