AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Agentsop Dspy

skill-agentsope-skillalchemy-agentsop-dspy · by agentsope

|

No reviews yet
0 installs
39 views
0.0% view→install

Install

$ agentstack add skill-agentsope-skillalchemy-agentsop-dspy

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution Used

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-agentsope-skillalchemy-agentsop-dspy)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
2mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Agentsop Dspy? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

DSPy SOP — Programming, Not Prompting

> "DSPy isn't a prompt-optimization agent framework. It's the LLM compiler for the shortest, cleanest code." > — Eito Miyamura [eito.substack.com/p/dspy-the-most-misunderstood-agent] > > "Prompts are effectively the weights of an LLM application." > — Core philosophy [arxiv.org/abs/2310.03714]


1. 何时激活 (When to activate)

Activate this skill when any of the following triggers are present in the user's intent or codebase:

| Trigger | Signal | |---|---| | Imports / mentions | import dspy, dspy.Signature, dspy.ChainOfThought, dspy.ReAct, Predict, MIPROv2, BootstrapFewShot, GEPA, teleprompter, compile( on an LM program | | Tasks | "auto-tune this prompt", "I want to swap GPT-4 for a smaller model without re-engineering prompts", "I have 50/200/1000 labeled examples — optimize this", "compile a pipeline for our metric", "distill GPT-4 into Llama-3-8B" | | Symptoms | Hand-written prompts grow past ~50 lines; brittleness on model swap; the team manually tunes few-shot examples; a metric exists but isn't being used to drive prompt design | | Cross-skill bridges | LangGraph node calls an LLM and needs better prompts → wrap the node body in a DSPy module. LlamaIndex retriever feeds a reranker → DSPy-compile the reranker against a labeled set |

Do NOT activate when:

  • The task is one-shot ("just answer this question once") — use raw client.messages.create.
  • No evaluation metric is possible and none is willing to be built — DSPy without a metric is just verbose prompting.
  • Prompts must remain human-authored verbatim for compliance, audit, or stylistic reasons.
  • The team is in rapid exploration mode where the task signature itself is changing daily — compile only after the signature stabilizes [dspy.ai/learn/optimization/overview/].

2. 核心心智模型 (Core mental model)

DSPy's full name is Declarative Self-improving Python. The three primitives form a PyTorch-like compile chain [arxiv.org/abs/2310.03714]:

┌─────────────┐    ┌──────────┐    ┌──────────────┐    ┌─────────┐
│  Signature  │ →  │  Module  │ →  │ Teleprompter │ →  │ Compile │
│ (what)      │    │ (how)    │    │ (optimizer)  │    │ (tune)  │
└─────────────┘    └──────────┘    └──────────────┘    └─────────┘
   I/O spec       Predict/CoT/      MIPROv2/GEPA/      Bake demos
   field names     ReAct/PoT        BootstrapFewShot   + instructions
   = semantic     = strategy        = search algorithm  into JSON

Three mental shifts the agent must internalize:

  1. Prompts are weights. The prompt string is not the artifact you ship — the compiled program (a JSON of demonstrations + instructions + structural choices) is. You ship program.json, not a .txt prompt [dspy.ai/tutorials/saving/].
  1. Signatures carry semantic load. question -> answer is not the same as query -> response. DSPy uses the field names as the only natural-language hint the optimizer has about intent before it sees data. Name them like you'd name function parameters in well-written code [dspy.ai/learn/programming/signatures/].
  1. Compile is a hyperparameter search, not a one-shot call. Compilation runs hundreds-to-thousands of LM calls (typically $2–$3 USD, 6–20 minutes, 3.2k API calls in the reference run). Costs scale with num_trials × |trainset| × |program LM calls| [dspy.ai/faqs/].

The PyTorch analogy is load-bearing. Signatures ≈ nn.Module.forward() shape contract. Modules ≈ nn.Linear / nn.Transformer. Teleprompters ≈ torch.optim.Adam. compile() ≈ training loop. save()/load() ≈ checkpoint.


3. SOP 工作流 (SOP workflow)

The DSPy team is explicit about a three-stage gate [dspy.ai/learn/]:

> "It's unproductive to launch optimization runs using a poorly designed program or a bad metric."

Do not skip stages. Each stage has an exit criterion.

Stage 1 — Programming (no optimizer yet)

  1. Pin the task as a Signature. Start inline ("question -> answer"); upgrade to a class-based dspy.Signature with InputField(desc=...) / OutputField(desc=...) when types matter or fields need disambiguation.
  2. Pick the lowest-power Module that works. Default to dspy.ChainOfThought. Use dspy.Predict for trivial classification, dspy.ReAct only when tools are needed, dspy.ProgramOfThought for arithmetic-heavy tasks [dspy.ai/learn/programming/modules/].
  3. Compose with plain Python control flow. Subclass dspy.Module, instantiate sub-modules in __init__, call them in forward(). No special DSL.
  4. Run zero-shot on 5–10 hand-picked examples. Look at outputs with dspy.inspect_history(n=3).

Exit criterion: the un-optimized program produces plausible outputs on 5+ examples. Not great — plausible.

Stage 2 — Evaluation (no optimizer yet)

  1. Build a dev set. Documented sweet spot: 30 examples = minimum useful, 300 = recommended, 200+ required for MIPROv2 to avoid overfitting [dspy.ai/learn/optimization/overview/].
  2. Write a metric: def metric(example, pred, trace=None) -> float|bool. Start with exact-match; only escalate to LLM-as-judge when the task demands it (open-ended generation, multi-criteria).
  3. Run dspy.Evaluate(devset=dev, metric=metric, num_threads=16) and record a baseline score.

Exit criterion: baseline score is stable across two runs (cache-free) AND the metric agrees with human judgment on 10 spot-checks.

Stage 3 — Optimization (compile)

  1. Pick optimizer by data + signal regime (Section 4 table). Decide which model optimizes vs. which model is the task model — they can differ.
  2. Use the unusual 20/80 split (20% train, 80% val) for prompt-based optimizers. GEPA uses standard ML splits (maximize train) [dspy.ai/learn/optimization/overview/].
  3. Start auto="light". Only escalate to "medium"/"heavy" if dev-set gains flatten and budget allows.
  4. Save the compiled program: compiled.save("v1.json") for state, or compiled.save("./v1/", save_program=True) for whole-program (preferred for production with metadata) [dspy.ai/tutorials/saving/].
  5. Deploy via FastAPI (dspy.asyncify) or MLflow (mlflow.dspy.log_model) [dspy.ai/tutorials/deployment/].

Exit criterion: compiled program beats baseline on a held-out test set (not the val set used in optimization) by ≥ task-relevant delta.

When to iterate back

Loop to Stage 1 if optimization plateaus. Per the docs: "Is your task well-defined? Do you need more data? Should your evaluation metric change?" — these are the questions to re-ask, not "should I try a different optimizer?" [dspy.ai/learn/optimization/overview/].


4. 操作模型 (Trigger / Action / Output / Evidence)

4.1 Choose the optimizer

| Trigger | Action | Output | Evidence | |---|---|---|---| | ≤10 labeled examples | BootstrapFewShot(metric=m, max_bootstrapped_demos=4, max_rounds=1) | Compiled program with self-generated demos | [dspy.ai/learn/optimization/optimizers/] | | 30–50 examples | BootstrapFewShotWithRandomSearch | Best-of-N candidate programs | [dspy.ai/learn/optimization/optimizers/] | | 200+ examples, willing to spend compute | MIPROv2(metric=m, auto="light") then escalate | Jointly-tuned instructions + few-shot demos via Bayesian optimization | [dspy.ai/api/optimizers/MIPROv2/] | | Need zero-shot prompts (no demos in final) | MIPROv2(..., max_bootstrapped_demos=0, max_labeled_demos=0) | Instruction-only optimization | [dspy.ai/learn/optimization/optimizers/] | | Have textual error feedback (test diffs, schema violations, judge rationales) | dspy.GEPA(metric=m_with_feedback) | Reflection-evolved prompts; sample-efficient | [dspy.ai/tutorials/gepaaiprogram/], [arxiv.org/abs/2507.19457] | | Already optimized with MIPROv2 / want to ship a smaller model | Chain into BootstrapFinetune(student=small_lm, teacher=optimized) | Finetuned weights (not just prompts) | [dspy.ai/api/optimizers/BootstrapFinetune/] | | Just want labeled demos in prompt (no search) | LabeledFewShot(k=8) | Trivial — fastest, cheapest, weakest | [dspy.ai/cheatsheet/] |

4.2 Module selection

| Trigger | Action | Why | |---|---|---| | Simple input → output | dspy.Predict(Sig) | Lowest overhead | | Reasoning helps | dspy.ChainOfThought(Sig) | Default choice per docs | | Math / counting / parsing | dspy.ProgramOfThought(Sig) | Code execution grounds the answer | | Tools (search, calc, API) | dspy.ReAct(Sig, tools=[...]) | Built-in tool loop | | Ensemble for hard cases | dspy.MultiChainComparison or dspy.majority | Vote across N CoT samples |

4.3 Metric design

| Trigger | Action | Caveat | |---|---|---| | Exact answer expected | lambda ex, pred: ex.answer.lower() == pred.answer.lower() | Cheap, deterministic | | Open-ended generation | LLM-as-judge with dspy.ChainOfThought(JudgeSig) | Watch for self-preference bias, recency bias, score-ID bias [arxiv.org/pdf/2509.26072] | | Multi-criteria (factuality + tone + length) | Sub-judge each dim, return bool during optimization (trace is not None) and float during evaluation | Documented pattern [dspy.ai/learn/evaluation/metrics/] | | Have rich error context | Return dspy.Prediction(score=..., feedback="missing field X") and use GEPA | Textual feedback is GEPA's superpower [dspy.ai/api/optimizers/GEPA/overview/] |

4.4 Cost guardrails

| Trigger | Action | Reference | |---|---|---| | Before any MIPROv2 call | Estimate: auto="light" ≈ a few $; auto="heavy" on 1000+ examples can hit tens of $ | [dspy.ai/faqs/] | | Budget tight | Use a cheap optimizer LM (e.g. gpt-4o-mini) to optimize prompts for a more expensive task LM — community-reported parity [github.com/stanfordnlp/dspy/issues/1596] | | Compile stuck mid-trial | Check issue #1970 pattern; reduce minibatch_size or kill and restart with smaller num_trials | | Need reproducibility | dspy.configure(track_usage=True) + log program.get_lm_usage() |


5. 困境决策案例 (Dilemma cases — ≥3)

Case A — "Optimizer cost vs gain: when is it worth compiling?"

困境 (Dilemma): User has a 3-stage RAG pipeline. Hand-tuned prompts already hit 72% on dev. MIPROv2 auto="heavy" would cost ~$40 and 4 hours. Worth it?

约束 (Constraints):

  • 250 labeled examples (above MIPROv2 200-example floor) [dspy.ai/learn/optimization/optimizers/].
  • Prompts already manually iterated — diminishing returns suspected.
  • Pipeline LM = GPT-4o ($-per-call adds up at trial scale).

决策步骤 (Decision steps):

  1. Check whether the prompts were ever validated against the metric, or just eye-balled. If eye-balled, even auto="light" (~$2) typically yields 10–30%+ on hand-tuned baselines per the paper's GPT-3.5/Llama2 results (25%/65% lift over standard few-shot) [arxiv.org/abs/2310.03714].
  2. Run auto="light" first as a cheap signal. The docs explicitly recommend "start with moderate values, observe behavior, and scale up only if you see clear gains" [github.com/stanfordnlp/dspy issue #1596].
  3. If light gives 20% of cases, fix the metric before compiling. Garbage metric → garbage compiled program.
  4. If sub-judge feedback is rich (e.g. "answer was verbose"), pipe textual feedback into dspy.GEPA instead of MIPROv2 — GEPA leverages text feedback for faster, more sample-efficient convergence [dspy.ai/api/optimizers/GEPA/overview/, arxiv.org/abs/2507.19457].
  5. Add a length penalty as a separate scalar in the metric — don't rely on the judge to penalize verbosity (judges over-prefer length).

结果: Multi-dimension metric with explicit length penalty + GEPA's textual feedback typically beats single-judge + MIPROv2 by 10–13% on AIME-style benchmarks [arxiv.org/abs/2507.19457] and is the empirically robust path.

可提取的操作: Never compile against a metric you haven't human-validated on ≥ 20 spot-checks. Decompose multi-criteria metrics. Prefer GEPA when you can express textual feedback.


Case D — "Compile-time hang / stuck trial — abort or wait?"

困境: MIPROv2 compile stuck mid-trial (no progress logs for 30 min). Reported pattern in issue #1970 [github.com/stanfordnlp/dspy/issues/1970]. Abort and restart, or wait?

约束:

  • $15 spent so far on the run.
  • Sunk cost vs. wasted further spend.
  • Possible causes: a single example triggers rate limits / context-length overflow / a tool call hangs.

决策步骤:

  1. Check dspy.inspect_history(n=3) — does the last LM call show truncation or rate-limit error?
  2. If context-length: reduce max_bootstrapped_demos and max_labeled_demos (default 4 each); the docs explicitly cite this as the #1 context-length fix [dspy.ai/faqs/].
  3. If rate-limit: lower num_threads in the underlying Evaluate; add retry/backoff in the LM client.
  4. If neither, abort. Restart with smaller minibatch_size (default 35; try 16) and smaller num_trials. Hanging is a known failure mode without graceful resume.
  5. Save partial progress: even mid-compile, student retains best demo candidates — check compiled._predictors state.

可提取的操作: **Compile is not atomic. Treat long hangs as failure. The cost of restart answer)` synthesizes. JetBlue's chatbot uses exactly this split — retrieval quality + answer quality as separate metrics, DSPy optimizes both [databricks.com/blog/optimizing-databricks-llm-pipelines-dspy].

vs Guidance / LMQL / Outlines

  • These control one LM call at the token level (grammars, regex, JSON schema).
  • DSPy controls multi-call programs at the optimization level.
  • Together: orthogonal. Use Outlines for "force valid JSON"; use DSPy for "make the JSON-emitting prompt good." DSPy's typed OutputField already pushes the LM toward structure but doesn't guarantee grammar conformance [dspy.ai/faqs/].

vs raw prompt engineering

  • Raw prompts win when: one-shot, signature unstable, no metric, audit constraints (Section 6).
  • DSPy wins when: pipeline ≥ 2 LM calls, you have ≥ 30 labeled examples, you'll swap models or scale, you have a metric (even an LLM-judge one — but harden it per Case C).

Production case study evidence

  • JetBlue + Databricks: RAG chatbot. Before DSPy = manual prompt tuning on retrieval/answer quality metrics. After = DSPy directly optimizes those metrics, faster development cycle [tastytechbytes.com/databricks-dspy-jetblue-ai-chatbot].
  • Haize Labs: automated LLM red-teaming.
  • In production at: Shopify, Databricks, Dropbox, JetBlue, Moody's, AWS, Sephora, VMware [dspy.ai].
  • Tobi Lütke (Shopify CEO): "Both DSPy and (especially) GEPA are currently severely under hyped in the AI context engineering world" [eito.substack.com].

Quick-reference appendix

Minimal end-to-end (verbatim from cheatsheet) [dspy.ai/cheatsheet/]

import dspy

# 1. Signature
class BasicQA(dspy.Signature):
    """Answer questions with short factoid answers."""
    question: str = dspy.InputField()
    answer: str = dspy.OutputField(desc="often between 1 and 5 words")

# 2. Module
qa = dspy.ChainOfThought(BasicQA)

# 3. Metric
def metric(ex, pred, trace=None):
    return ex.answer.lower() in pred.answer.lower()

# 4. Compile
from dspy.teleprompt import MIPROv2
optimizer = MIPROv2(metric=metric, auto="light")
compiled = optimizer.compile(qa, trainset=trainset)

# 5. Save / load
compiled.save("v1.json")

Decision tree (one screen)

Have a metric? ─── No ──► Stop. Build a metric first. (Or skip DSPy.)
   │
   Yes
   │
Have ≥ 30 examples? ─── No ──► Stop. Collect more data, or use LabeledFewShot(k=8) as floor.
   │
   Yes
   │
Have textual error feedback? ─── Yes ──► dspy.GEPA
   │
   No
   │
≤ 10 examples? ──► BootstrapFewShot
30–50?         ──► BootstrapFewShotWithRandomSearch
50–200?        ──► MIPROv2(auto="light", max_bootstrapped_demos=4)
200+?          ──► MIPROv2(auto="light" → "medium" if gains; "heavy" only if 300+ and budget)
Need to ship small model? ──► chain BootstrapFinetune after MIPROv2

Anatomy of a compiled `pr

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.