Install
$ agentstack add skill-agentsope-skillalchemy-agentsop-dspy ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ● Dynamic code execution Used
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
DSPy SOP — Programming, Not Prompting
> "DSPy isn't a prompt-optimization agent framework. It's the LLM compiler for the shortest, cleanest code." > — Eito Miyamura [eito.substack.com/p/dspy-the-most-misunderstood-agent] > > "Prompts are effectively the weights of an LLM application." > — Core philosophy [arxiv.org/abs/2310.03714]
1. 何时激活 (When to activate)
Activate this skill when any of the following triggers are present in the user's intent or codebase:
| Trigger | Signal | |---|---| | Imports / mentions | import dspy, dspy.Signature, dspy.ChainOfThought, dspy.ReAct, Predict, MIPROv2, BootstrapFewShot, GEPA, teleprompter, compile( on an LM program | | Tasks | "auto-tune this prompt", "I want to swap GPT-4 for a smaller model without re-engineering prompts", "I have 50/200/1000 labeled examples — optimize this", "compile a pipeline for our metric", "distill GPT-4 into Llama-3-8B" | | Symptoms | Hand-written prompts grow past ~50 lines; brittleness on model swap; the team manually tunes few-shot examples; a metric exists but isn't being used to drive prompt design | | Cross-skill bridges | LangGraph node calls an LLM and needs better prompts → wrap the node body in a DSPy module. LlamaIndex retriever feeds a reranker → DSPy-compile the reranker against a labeled set |
Do NOT activate when:
- The task is one-shot ("just answer this question once") — use raw
client.messages.create. - No evaluation metric is possible and none is willing to be built — DSPy without a metric is just verbose prompting.
- Prompts must remain human-authored verbatim for compliance, audit, or stylistic reasons.
- The team is in rapid exploration mode where the task signature itself is changing daily — compile only after the signature stabilizes [dspy.ai/learn/optimization/overview/].
2. 核心心智模型 (Core mental model)
DSPy's full name is Declarative Self-improving Python. The three primitives form a PyTorch-like compile chain [arxiv.org/abs/2310.03714]:
┌─────────────┐ ┌──────────┐ ┌──────────────┐ ┌─────────┐
│ Signature │ → │ Module │ → │ Teleprompter │ → │ Compile │
│ (what) │ │ (how) │ │ (optimizer) │ │ (tune) │
└─────────────┘ └──────────┘ └──────────────┘ └─────────┘
I/O spec Predict/CoT/ MIPROv2/GEPA/ Bake demos
field names ReAct/PoT BootstrapFewShot + instructions
= semantic = strategy = search algorithm into JSON
Three mental shifts the agent must internalize:
- Prompts are weights. The prompt string is not the artifact you ship — the compiled program (a JSON of demonstrations + instructions + structural choices) is. You ship
program.json, not a.txtprompt [dspy.ai/tutorials/saving/].
- Signatures carry semantic load.
question -> answeris not the same asquery -> response. DSPy uses the field names as the only natural-language hint the optimizer has about intent before it sees data. Name them like you'd name function parameters in well-written code [dspy.ai/learn/programming/signatures/].
- Compile is a hyperparameter search, not a one-shot call. Compilation runs hundreds-to-thousands of LM calls (typically $2–$3 USD, 6–20 minutes, 3.2k API calls in the reference run). Costs scale with
num_trials × |trainset| × |program LM calls|[dspy.ai/faqs/].
The PyTorch analogy is load-bearing. Signatures ≈ nn.Module.forward() shape contract. Modules ≈ nn.Linear / nn.Transformer. Teleprompters ≈ torch.optim.Adam. compile() ≈ training loop. save()/load() ≈ checkpoint.
3. SOP 工作流 (SOP workflow)
The DSPy team is explicit about a three-stage gate [dspy.ai/learn/]:
> "It's unproductive to launch optimization runs using a poorly designed program or a bad metric."
Do not skip stages. Each stage has an exit criterion.
Stage 1 — Programming (no optimizer yet)
- Pin the task as a Signature. Start inline (
"question -> answer"); upgrade to a class-baseddspy.SignaturewithInputField(desc=...)/OutputField(desc=...)when types matter or fields need disambiguation. - Pick the lowest-power Module that works. Default to
dspy.ChainOfThought. Usedspy.Predictfor trivial classification,dspy.ReActonly when tools are needed,dspy.ProgramOfThoughtfor arithmetic-heavy tasks [dspy.ai/learn/programming/modules/]. - Compose with plain Python control flow. Subclass
dspy.Module, instantiate sub-modules in__init__, call them inforward(). No special DSL. - Run zero-shot on 5–10 hand-picked examples. Look at outputs with
dspy.inspect_history(n=3).
Exit criterion: the un-optimized program produces plausible outputs on 5+ examples. Not great — plausible.
Stage 2 — Evaluation (no optimizer yet)
- Build a dev set. Documented sweet spot: 30 examples = minimum useful, 300 = recommended, 200+ required for MIPROv2 to avoid overfitting [dspy.ai/learn/optimization/overview/].
- Write a metric:
def metric(example, pred, trace=None) -> float|bool. Start with exact-match; only escalate to LLM-as-judge when the task demands it (open-ended generation, multi-criteria). - Run
dspy.Evaluate(devset=dev, metric=metric, num_threads=16)and record a baseline score.
Exit criterion: baseline score is stable across two runs (cache-free) AND the metric agrees with human judgment on 10 spot-checks.
Stage 3 — Optimization (compile)
- Pick optimizer by data + signal regime (Section 4 table). Decide which model optimizes vs. which model is the task model — they can differ.
- Use the unusual 20/80 split (20% train, 80% val) for prompt-based optimizers. GEPA uses standard ML splits (maximize train) [dspy.ai/learn/optimization/overview/].
- Start
auto="light". Only escalate to"medium"/"heavy"if dev-set gains flatten and budget allows. - Save the compiled program:
compiled.save("v1.json")for state, orcompiled.save("./v1/", save_program=True)for whole-program (preferred for production with metadata) [dspy.ai/tutorials/saving/]. - Deploy via FastAPI (
dspy.asyncify) or MLflow (mlflow.dspy.log_model) [dspy.ai/tutorials/deployment/].
Exit criterion: compiled program beats baseline on a held-out test set (not the val set used in optimization) by ≥ task-relevant delta.
When to iterate back
Loop to Stage 1 if optimization plateaus. Per the docs: "Is your task well-defined? Do you need more data? Should your evaluation metric change?" — these are the questions to re-ask, not "should I try a different optimizer?" [dspy.ai/learn/optimization/overview/].
4. 操作模型 (Trigger / Action / Output / Evidence)
4.1 Choose the optimizer
| Trigger | Action | Output | Evidence | |---|---|---|---| | ≤10 labeled examples | BootstrapFewShot(metric=m, max_bootstrapped_demos=4, max_rounds=1) | Compiled program with self-generated demos | [dspy.ai/learn/optimization/optimizers/] | | 30–50 examples | BootstrapFewShotWithRandomSearch | Best-of-N candidate programs | [dspy.ai/learn/optimization/optimizers/] | | 200+ examples, willing to spend compute | MIPROv2(metric=m, auto="light") then escalate | Jointly-tuned instructions + few-shot demos via Bayesian optimization | [dspy.ai/api/optimizers/MIPROv2/] | | Need zero-shot prompts (no demos in final) | MIPROv2(..., max_bootstrapped_demos=0, max_labeled_demos=0) | Instruction-only optimization | [dspy.ai/learn/optimization/optimizers/] | | Have textual error feedback (test diffs, schema violations, judge rationales) | dspy.GEPA(metric=m_with_feedback) | Reflection-evolved prompts; sample-efficient | [dspy.ai/tutorials/gepaaiprogram/], [arxiv.org/abs/2507.19457] | | Already optimized with MIPROv2 / want to ship a smaller model | Chain into BootstrapFinetune(student=small_lm, teacher=optimized) | Finetuned weights (not just prompts) | [dspy.ai/api/optimizers/BootstrapFinetune/] | | Just want labeled demos in prompt (no search) | LabeledFewShot(k=8) | Trivial — fastest, cheapest, weakest | [dspy.ai/cheatsheet/] |
4.2 Module selection
| Trigger | Action | Why | |---|---|---| | Simple input → output | dspy.Predict(Sig) | Lowest overhead | | Reasoning helps | dspy.ChainOfThought(Sig) | Default choice per docs | | Math / counting / parsing | dspy.ProgramOfThought(Sig) | Code execution grounds the answer | | Tools (search, calc, API) | dspy.ReAct(Sig, tools=[...]) | Built-in tool loop | | Ensemble for hard cases | dspy.MultiChainComparison or dspy.majority | Vote across N CoT samples |
4.3 Metric design
| Trigger | Action | Caveat | |---|---|---| | Exact answer expected | lambda ex, pred: ex.answer.lower() == pred.answer.lower() | Cheap, deterministic | | Open-ended generation | LLM-as-judge with dspy.ChainOfThought(JudgeSig) | Watch for self-preference bias, recency bias, score-ID bias [arxiv.org/pdf/2509.26072] | | Multi-criteria (factuality + tone + length) | Sub-judge each dim, return bool during optimization (trace is not None) and float during evaluation | Documented pattern [dspy.ai/learn/evaluation/metrics/] | | Have rich error context | Return dspy.Prediction(score=..., feedback="missing field X") and use GEPA | Textual feedback is GEPA's superpower [dspy.ai/api/optimizers/GEPA/overview/] |
4.4 Cost guardrails
| Trigger | Action | Reference | |---|---|---| | Before any MIPROv2 call | Estimate: auto="light" ≈ a few $; auto="heavy" on 1000+ examples can hit tens of $ | [dspy.ai/faqs/] | | Budget tight | Use a cheap optimizer LM (e.g. gpt-4o-mini) to optimize prompts for a more expensive task LM — community-reported parity [github.com/stanfordnlp/dspy/issues/1596] | | Compile stuck mid-trial | Check issue #1970 pattern; reduce minibatch_size or kill and restart with smaller num_trials | | Need reproducibility | dspy.configure(track_usage=True) + log program.get_lm_usage() |
5. 困境决策案例 (Dilemma cases — ≥3)
Case A — "Optimizer cost vs gain: when is it worth compiling?"
困境 (Dilemma): User has a 3-stage RAG pipeline. Hand-tuned prompts already hit 72% on dev. MIPROv2 auto="heavy" would cost ~$40 and 4 hours. Worth it?
约束 (Constraints):
- 250 labeled examples (above MIPROv2 200-example floor) [dspy.ai/learn/optimization/optimizers/].
- Prompts already manually iterated — diminishing returns suspected.
- Pipeline LM = GPT-4o ($-per-call adds up at trial scale).
决策步骤 (Decision steps):
- Check whether the prompts were ever validated against the metric, or just eye-balled. If eye-balled, even
auto="light"(~$2) typically yields 10–30%+ on hand-tuned baselines per the paper's GPT-3.5/Llama2 results (25%/65% lift over standard few-shot) [arxiv.org/abs/2310.03714]. - Run
auto="light"first as a cheap signal. The docs explicitly recommend "start with moderate values, observe behavior, and scale up only if you see clear gains" [github.com/stanfordnlp/dspy issue #1596]. - If
lightgives 20% of cases, fix the metric before compiling. Garbage metric → garbage compiled program. - If sub-judge feedback is rich (e.g. "answer was verbose"), pipe textual feedback into
dspy.GEPAinstead of MIPROv2 — GEPA leverages text feedback for faster, more sample-efficient convergence [dspy.ai/api/optimizers/GEPA/overview/, arxiv.org/abs/2507.19457]. - Add a length penalty as a separate scalar in the metric — don't rely on the judge to penalize verbosity (judges over-prefer length).
结果: Multi-dimension metric with explicit length penalty + GEPA's textual feedback typically beats single-judge + MIPROv2 by 10–13% on AIME-style benchmarks [arxiv.org/abs/2507.19457] and is the empirically robust path.
可提取的操作: Never compile against a metric you haven't human-validated on ≥ 20 spot-checks. Decompose multi-criteria metrics. Prefer GEPA when you can express textual feedback.
Case D — "Compile-time hang / stuck trial — abort or wait?"
困境: MIPROv2 compile stuck mid-trial (no progress logs for 30 min). Reported pattern in issue #1970 [github.com/stanfordnlp/dspy/issues/1970]. Abort and restart, or wait?
约束:
- $15 spent so far on the run.
- Sunk cost vs. wasted further spend.
- Possible causes: a single example triggers rate limits / context-length overflow / a tool call hangs.
决策步骤:
- Check
dspy.inspect_history(n=3)— does the last LM call show truncation or rate-limit error? - If context-length: reduce
max_bootstrapped_demosandmax_labeled_demos(default 4 each); the docs explicitly cite this as the #1 context-length fix [dspy.ai/faqs/]. - If rate-limit: lower
num_threadsin the underlying Evaluate; add retry/backoff in the LM client. - If neither, abort. Restart with smaller
minibatch_size(default 35; try 16) and smallernum_trials. Hanging is a known failure mode without graceful resume. - Save partial progress: even mid-compile,
studentretains best demo candidates — checkcompiled._predictorsstate.
可提取的操作: **Compile is not atomic. Treat long hangs as failure. The cost of restart answer)` synthesizes. JetBlue's chatbot uses exactly this split — retrieval quality + answer quality as separate metrics, DSPy optimizes both [databricks.com/blog/optimizing-databricks-llm-pipelines-dspy].
vs Guidance / LMQL / Outlines
- These control one LM call at the token level (grammars, regex, JSON schema).
- DSPy controls multi-call programs at the optimization level.
- Together: orthogonal. Use Outlines for "force valid JSON"; use DSPy for "make the JSON-emitting prompt good." DSPy's typed
OutputFieldalready pushes the LM toward structure but doesn't guarantee grammar conformance [dspy.ai/faqs/].
vs raw prompt engineering
- Raw prompts win when: one-shot, signature unstable, no metric, audit constraints (Section 6).
- DSPy wins when: pipeline ≥ 2 LM calls, you have ≥ 30 labeled examples, you'll swap models or scale, you have a metric (even an LLM-judge one — but harden it per Case C).
Production case study evidence
- JetBlue + Databricks: RAG chatbot. Before DSPy = manual prompt tuning on retrieval/answer quality metrics. After = DSPy directly optimizes those metrics, faster development cycle [tastytechbytes.com/databricks-dspy-jetblue-ai-chatbot].
- Haize Labs: automated LLM red-teaming.
- In production at: Shopify, Databricks, Dropbox, JetBlue, Moody's, AWS, Sephora, VMware [dspy.ai].
- Tobi Lütke (Shopify CEO): "Both DSPy and (especially) GEPA are currently severely under hyped in the AI context engineering world" [eito.substack.com].
Quick-reference appendix
Minimal end-to-end (verbatim from cheatsheet) [dspy.ai/cheatsheet/]
import dspy
# 1. Signature
class BasicQA(dspy.Signature):
"""Answer questions with short factoid answers."""
question: str = dspy.InputField()
answer: str = dspy.OutputField(desc="often between 1 and 5 words")
# 2. Module
qa = dspy.ChainOfThought(BasicQA)
# 3. Metric
def metric(ex, pred, trace=None):
return ex.answer.lower() in pred.answer.lower()
# 4. Compile
from dspy.teleprompt import MIPROv2
optimizer = MIPROv2(metric=metric, auto="light")
compiled = optimizer.compile(qa, trainset=trainset)
# 5. Save / load
compiled.save("v1.json")
Decision tree (one screen)
Have a metric? ─── No ──► Stop. Build a metric first. (Or skip DSPy.)
│
Yes
│
Have ≥ 30 examples? ─── No ──► Stop. Collect more data, or use LabeledFewShot(k=8) as floor.
│
Yes
│
Have textual error feedback? ─── Yes ──► dspy.GEPA
│
No
│
≤ 10 examples? ──► BootstrapFewShot
30–50? ──► BootstrapFewShotWithRandomSearch
50–200? ──► MIPROv2(auto="light", max_bootstrapped_demos=4)
200+? ──► MIPROv2(auto="light" → "medium" if gains; "heavy" only if 300+ and budget)
Need to ship small model? ──► chain BootstrapFinetune after MIPROv2
Anatomy of a compiled `pr
…
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: agentsope
- Source: agentsope/SkillAlchemy
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.