Install
$ agentstack add skill-black-forest-labs-skills-flux-3-audio-dialogue ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
FLUX 3 Audio and Dialogue
Name each layer separately: speech, voiceover, ambience, effects, music, or deliberate silence. One blurred description gives up control of all of them; every sound needs a physical source or narrative role.
Speech. Quote the exact line, name the visible speaker (or label the line voiceover/narration so it is not searching for a mouth to belong to), and add no on-screen text, no subtitles when text is unwanted:
A weather presenter on camera in front of a stylized storm map, speaking directly to
the lens: "Storm season is here, and this time, we're ready." Confident delivery,
clean studio lighting. No on-screen text, no subtitles.
Voice anchors: age range, accent when relevant, register, energy, recording distance. Reusing the same direction preserves a kind of voice, not the same performer across generations.
Speakability. Write for the clip's real duration: short sentences, one thought per line, room before and after the payoff; spell unusual names phonetically; shorten the line before speeding the delivery. A line that cannot finish comfortably needs a shorter script or a longer clip.
Effects are causal, not a detached list:
As the cup hits the tile, it cracks with one sharp ceramic snap.
Mix. Say what leads and what stays under it. Keep background voices out when one line matters:
Her line is foreground and fully intelligible. Café chatter and espresso hiss remain
low and diffuse. A restrained piano pulse enters beneath the final words without
masking them.
Silence and post. Set generate_audio: false for a deliberately silent source clip. Reserve for deterministic post: final loudness, EQ, ducking, and fades; guaranteed wording or speaker identity; frame-accurate sync; subtitles and captions; continuity across separately generated clips.
When a take misses, change one dimension at a time: speaker ownership, line length, delivery anchors, competing layers, action-to-effect causality, or generation versus post.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: black-forest-labs
- Source: black-forest-labs/skills
- License: MIT
- Homepage: https://docs.bfl.ai
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.