Install
$ agentstack add skill-omidzamani-dspy-skills-dspy-adapters-multimodal ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
DSPy Adapters and Multimodal I/O
Goal
Choose an adapter deliberately and model image, audio, and file inputs with DSPy's typed primitives.
Adapter Selection
| Adapter | Use it for | |---------|------------| | dspy.ChatAdapter() | Default, human-readable field markers, broad model compatibility | | dspy.JSONAdapter() | Structured JSON output and native function calling where supported | | dspy.XMLAdapter() | XML-tagged fields when XML is easier for the target LM to follow | | dspy.TwoStepAdapter() | A separate extraction pass when parsing needs extra help |
Configure globally or for a limited scope:
import dspy
dspy.configure(
lm=dspy.LM("openai/gpt-4o-mini"),
adapter=dspy.JSONAdapter(),
)
with dspy.context(adapter=dspy.XMLAdapter()):
result = dspy.Predict("question -> answer")(question="What is DSPy?")
Native Function Calling
JSONAdapter enables native function calling by default. ChatAdapter keeps text parsing by default. Override either behavior explicitly:
chat_native = dspy.ChatAdapter(use_native_function_calling=True)
json_manual = dspy.JSONAdapter(use_native_function_calling=False)
DSPy falls back to manual parsing when the configured LM does not support native function calling.
Image Inputs
class DescribeImage(dspy.Signature):
image: dspy.Image = dspy.InputField()
description: str = dspy.OutputField()
describe = dspy.Predict(DescribeImage)
result = describe(image=dspy.Image("./diagram.png"))
Pass a local path, HTTP URL, bytes, PIL image, or existing data URI directly to dspy.Image(...).
Audio and File Inputs
class SummarizeAudio(dspy.Signature):
audio: dspy.Audio = dspy.InputField()
summary: str = dspy.OutputField()
audio = dspy.Audio.from_file("./meeting.wav")
summary = dspy.Predict(SummarizeAudio)(audio=audio)
class SummarizeFile(dspy.Signature):
file: dspy.File = dspy.InputField()
summary: str = dspy.OutputField()
document = dspy.File.from_path("./research.pdf")
summary = dspy.Predict(SummarizeFile)(file=document)
Provider capabilities vary. Verify that the selected model accepts the media type before deployment.
Best Practices
- Start with
ChatAdapter; switch only for a measured reason. - Use typed signatures for structured output.
- Test adapter behavior against the exact production model.
- Avoid deprecated
Image.from_file()andImage.from_url()helpers; calldspy.Image(...). - Keep local file handling and uploaded file IDs within provider policy.
Related Skills
- Design signatures: [dspy-signature-designer](../dspy-signature-designer/SKILL.md)
- Build tool agents: [dspy-react-agent-builder](../dspy-react-agent-builder/SKILL.md)
Official Documentation
- Adapters guide: https://dspy.ai/learn/programming/adapters/
- Tools guide: https://dspy.ai/learn/programming/tools/
- XMLAdapter API: https://dspy.ai/api/adapters/XMLAdapter/
- Image API: https://dspy.ai/api/primitives/Image/
- Audio API: https://dspy.ai/api/primitives/Audio/
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: OmidZamani
- Source: OmidZamani/dspy-skills
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.