# Flux 3 Product Ads

> Use when building a finished product ad from FLUX 3 - shot design, voiceover, action-to-word sync, evidence-gated copy, deterministic assembly, and QC gates that catch clipped audio, floating products, off-model plates, and reports that claim a pass the build did not give.

- **Type:** Skill
- **Install:** `agentstack add skill-black-forest-labs-skills-flux-3-product-ads`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [black-forest-labs](https://agentstack.voostack.com/s/black-forest-labs)
- **Installs:** 0
- **Category:** [Content & Media](https://agentstack.voostack.com/c/content-and-media)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [black-forest-labs](https://github.com/black-forest-labs)
- **Source:** https://github.com/black-forest-labs/skills/tree/master/skills/flux-3-product-ads
- **Website:** https://docs.bfl.ai

## Install

```sh
agentstack add skill-black-forest-labs-skills-flux-3-product-ads
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# FLUX 3 product ads

A finished spot is two jobs, not one. FLUX 3 generates picture and voice. A
deterministic pass cuts, times, captions and masters them. Keep the boundary
sharp: the model handles what only a model can, and everything a computer can
compute exactly stays out of the model's hands.

Route facts this skill depends on live in `flux-3-generate`. Read it first for
`/v1/flux-3-video`, polling and download behaviour. Continuation behaviour, including
what a `v2v` link actually returns, lives in `flux-3-keyframes-continuation`.

Everything here has been run at 12, 30 and 60 seconds. Where a rule only holds at one
length, it says so. **Assume nothing written for a three-shot ten-second spot
generalises without being re-tested at length**: on the run that produced this
revision, two pieces of timing logic that were correct at 10 seconds turned out to
have a silent correctness bug and a complexity bug that only appear in a longer read.

## Shape of the work

1. Design shots that survive generation.
2. Generate picture plates and VO in the same round.
3. Screen VO by machine, then listen.
4. Derive timing from the audio.
5. Assemble deterministically.
6. Gate on measurements, including the failures that look like passes.

## 1. Shot design

Three cuts carry a 10-second spot: **reveal, proof, payoff**. Reveal establishes the
object, proof shows it doing the one thing the copy claims, payoff lands the brand.

Longer spots need more beats and more plates. A plate yields roughly 4 to 5 seconds
of usable middle once a dissolve is allowed for, so **shot count is length divided by
about 4.5, rounded up**, and a shortfall does not degrade gracefully: it fails the
build outright with no legal cut placement. Measured: 12s took 3 plates, 30s took 6,
60s took 11. A 60-second spot planned with 9 plates could not be cut at all.

At 30 seconds and beyond, reveal/proof/payoff no longer fills the time. What worked:
hook on the material, the object whole, a detail, the mechanism, a state change, what
the product does for you, brand. The structure that matters is that each beat earns
its own shot, not that it has three parts.

**Anchor motion at its endpoints.** A generated clip is trustworthy at its first and
last frame and inventive in between. Motion whose midpoint is implied by its endpoints
survives; motion that requires the model to invent geometry does not.

Works: a highlight travelling across a surface, a collar rotating through a short arc
and stopping, a handle rising and stopping, a lid parting slightly.

Two shots that *appear* to work and do not: **a hand drawing an ink line across
paper**, and **a phone descending into frame and landing on a product**. Both were
predicted to fail as "travel" and "an object entering the frame". Both arrived
looking correct, passed every signal gate, went into masters, and were rejected on
sight. The pen's ink ran ahead of the tip while the paper slid underneath; the
"phone" was a featureless slab as wide as the pad it landed on. Arrival is not
correctness, and this is the trap: the failure mode of a *nearly* achievable shot
is a plate that is wrong in the object rather than broken in the signal, which is
invisible to every measurement. See the semantic gate in section 6.

Fails: multiple revolutions, and rotation of two bodies at once. A crank asked for 2
to 3 full revolutions rotated to about 180 degrees, then deformed and vanished as the
arm occluded itself. A knurled collar asked to spin two revolutions *while* the head
it sits on tilted produced no motion at all: an inert clip where neither action
happened.

The predictor is not travel and not occlusion. It is **whether the model must invent
geometry it has never seen**. A hand crossing frame is a rigid body with a known
silhouette. A descending phone is a rigid body. A crank going around 720 degrees has
to render its own far side, and that is where it dies. Design against invented
geometry, not against movement.

**A quarter-turn is not a safe fallback for a rotation shot, it is an invisible one.**
The safe version of the failed collar shot technically succeeded and was still
unusable, because a small rotation of a knurled ring at macro reads as nothing
happening. When a rotation shot fails, the shot that replaces it should be **light
moving across a static object**, which is the most reliable motion in the system and
the one that consistently looks alive.

**When an object must arrive in frame, generate it leaving and reverse the clip.**
This is the highest-leverage trick in the section, because it converts an
invention problem into a translation problem. Three attempts at the same shot, one
variable:

| Attempt | Approach | Result |
|---|---|---|
| 1 | describe a descending phone; keyframe contains no phone | slab as wide as the pad |
| 2 | describe the phone in far more detail (two-thirds pad diameter, 8mm thick, dark glass front, metal rails, rounded corners, explicit no-warp instruction); same phone-less keyframe | still a slab, now tilted and overhanging the pad |
| 3 | keyframe already contains a correctly proportioned phone; generate the phone *lifting away*; `ffmpeg -vf reverse` in post | correct phone, correct scale, constant through the move |

Prompt detail did not fix invention. Removing the invention did. If the object is
in the keyframe, the model only has to move something it can already see, and
scale cannot drift because it was never chosen by the model. Cost: one filter.

Two things to watch. Author any lighting change in the direction that reads
correctly *after* reversing, and check the source still's camera height against its
neighbours, since a still shot for a different purpose may not cut with them.

**Reversing moves the action to the other end of the clip, and any timing you
already derived is now wrong.** This is the cost the trick hides, and it surfaces
as a sync failure in a spot that was passing before. The lift plate peaked at
5.79s of a 6.04s clip, a quarter-second before the end. Reversed, that same peak
sits at 0.17s. Nothing else changed: same duration, same frame count, same file
size class.

That matters because a plate can only be slipped *later* into its segment, never
earlier than its own first frame. So the reachable window for an action collapses
to roughly `[peak - segment_slack, peak]`, and an action at 0.17s can only ever
land in the first fraction of a second of its segment. The anchor word chosen for
the pre-reverse plate, several seconds in, became permanently unreachable, and the
build reported it correctly as a clamp:

```
slip : shot 3 in=0.00s -> action 14.32s vs word 15.62s (-1.30s)
       CLAMPED (reachable 13.00-14.32s, word at 15.62s)
```

The trap is that this looks like drift. The plates were byte-identical to the run
that passed, the VO was untouched, and the durations matched to the sample, so
every "did something move?" check comes back clean while the numbers disagree with
a note written a day earlier. The thing that moved was inside the file.

Two rules follow. **Re-derive anchors after any post step that changes where the
action sits in the clip**, reversing above all, and treat a reversal as a new plate
rather than an edit of the old one. And **when a reversed plate needs an early
action, anchor it to an early word**: a landing that peaks in the first frames
belongs on the first word of its line, not the word that named the motion when the
clip ran forwards.

**Test the prediction rather than trusting it.** Give a risky shot two prompts in the
same brief, a `hard_prompt` and a safe `prompt`, generate both, keep both. One extra
job per risky shot is what keeps this section honest as the model improves, and it is
how the two "impossible" shots above were found.

**Name the grounding in every prompt.** Say the product's contact shadow and its
reflection explicitly. A product can hold identity and motion perfectly and still look
pasted onto the frame because nothing defended its shadow. This passes identity checks
and reads as fake instantly to a human.

**Camera lock is advisory.** "Locked camera, no push-in, no pan, no zoom" still permits
parallax and drift. Design shots that tolerate a little movement rather than expecting
the prompt to forbid it.

## 1b. Reference stills come first, and the model gets a vote

Every plate inherits its still, because `i2v` treats the keyframe as literal. A still
that is wrong in a way you can live with poisons every clip generated from it, so the
reference pack has to be right before any video job runs.

**When the product is real, the pack is derived from the photograph, not written from
scratch.** This is the easier path and it skips the entire failure class below, because
no invariant paragraph has to describe the object well enough for a model to build it.
Take the supplied photo as the canonical still, then generate each remaining angle from
it with FLUX 2 `input_image` identity carry, seeded so the pack is reproducible. Write
the invariants anyway, by reading them off the photograph: they are what the identity
and semantic gates check against later. On a run built this way from one real product
photo, three generated angles held every named feature and no still needed regenerating.

Two things this does not buy you. The still is faithful and the *video* still drifts,
so the interior-frame and identity checks below apply unchanged. And a real photo is
usually a catalogue shot on seamless white, which gives you no set to cut to: every
plate looks like the same photo unless the shots differ in framing and scale, so design
the pack for genuinely different crops.

Expect the model to overrule the spec, and read it as information. On a three-product
run, two products came back with the model quietly substituting its own design: a
light channel specified as unlit rendered lit in every frame, and a lamp specified
with a two-segment elbow arm rendered as a single post with a yoke-mounted head in all
seven angles.

Both refusals were **coherent**: one alternative design, held consistently across every
angle. That is the signal. A model that disagrees at random gives you noise; a model
that disagrees identically seven times is telling you the invariant paragraph
describes something it cannot build. **Rewrite the spec to match what it reliably
builds, then regenerate.** Fighting a coherent refusal costs jobs and loses.

Regenerate a still when the error is one the video stage will amplify: stray text on a
prop, a wordmark in the wrong place, an object overhanging its base.

**Check identity across the pack, not one still at a time.** A pack can be clean
plate by plate and still be incoherent, because "is this a good photo of a lamp?"
is a different question from "is every one of these the same lamp?" On the lamp
above, the pack mixed two incompatible designs across its angles; each still looked
fine alone, and every plate generated from them inherited whichever design its
keyframe happened to carry. Pick one still as canonical, then compare each of the
others to it on **named, falsifiable features** rather than overall impression:

```
base:   shallow domed profile curving in one arc, vs a flat cylindrical puck
        with a vertical side wall
post:   smooth and unbroken from base to head, vs a collar, ring, knurling
        or joint partway up
head:   plain green cylinder roughly twice as long as wide, vs short/fat
        or a knurled metal barrel
ring:   exactly one knurled ring, at the FRONT of the head encircling the
        lens, vs any knurling on the post or base
```

Each feature names the correct form *and* the wrong one it gets confused with,
which is what makes a verdict checkable. Run the same comparison over the finished
plates too, at least twice per feature, and fold unanimously as with the semantic
gate in section 6. On a 14-plate run this returned 13 consistent and one off-model,
where a thin gold bezel had changed the head's proportions. The plate it flagged
had passed every signal gate. Its hard-motion twin, generated as the safe/risky
pair described above, was clean and already on disk, so the fix cost nothing.

This is worth doing before generation and again after, because the two runs answer
different questions: the first stops a poisoned pack, the second catches the plate
that drifted anyway.

**Never let a failed step's stale output become the next step's input.** When a
chained link failed, the runner picked up the previous day's file of the same name
and fed it to the following link, which then succeeded and produced a plausible
clip built on the wrong material. A failure that leaves the old file in place is
indistinguishable, to the next step, from a success. Check that a chain input is
newer than the run that is consuming it, or write links to run-scoped names so a
missing file is missing rather than stale.

## 2. Generate picture and VO together

Finished product ads carry voiceover. Budget it in the first generation round,
never bolt it on after picture lock: the VO determines the length of the spot,
so generating it last means re-cutting everything.

For voiceover generation, music beds, audio layering, speakability checks, and
deterministic audio finishing, use `flux-3-audio-dialogue`. This skill covers
ad-specific timing and assembly: how the VO sets spot length, how paragraphs
join at designed pauses, and how the mastered stem feeds the cut pipeline.

Get VO from audio-only jobs: a voice-booth scene whose picture is discarded and
whose audio is harvested. Generate at least two scripts against two speaker
profiles, because takes are not interchangeable and you want a real choice.

Do not use native picture audio in a finished spot. Generate picture silent.

**`duration` is a tempo control, not a length estimate.** The clip comes back at
exactly the length you ask for and the read stretches to fill it. The same 27-word
script, same speaker, asked for three lengths:

| asked | returned | words/sec |
| --- | --- | --- |
| 11s | 11.01s | 2.45 |
| 14s | 14.00s | 2.11 |
| 18s | 18.02s | 1.50 |

The script was delivered completely and correctly every time. Only the pace changed,
and at 18s it does not pad with silence, it drags. So pick the speaking rate you want
and derive the length: `duration = words / rate`, with roughly 2.1 to 2.5 words per
second reading unhurried and confident. Do not pad the estimate for "room to breathe";
that slows the delivery instead of adding a pause.

**One audio job caps at 20 seconds, about 45 words at a good pace.** A 30-second spot
needs around 75 words and a 60-second spot around 150, so any longer read is generated
**one paragraph per job** and joined in post at a designed pause. This is also better
writing, because ads pause. Trim each paragraph to its own speech before joining;
stacking booth room tone produces an audible seam.

**Generate a music bed the same way** whenever the spot runs past about 20 seconds.
Silence between paragraphs that is fine at 10 seconds sounds broken at 60. A bed is an
audio-only job with no speech in the prompt, looped with a crossfade to length, and
ducked under the voice with a sidechain compressor keyed off the VO. Attenuate the bed
first and let the duck shape an already-quiet bed: a bed loud enough to need heavy
ducking pumps audibly on every phrase.

## 3. Screen VO by machine, then listen

Transcribe each take and compute word error rate against the script. This
catches dropped and mangled lines cheaply.

**Invented brand names need phonetic matching, not exact matching.** ASR has
never seen your product name and will spell it plausibly wrong: "Ferrolane"
came back as "Feraline", "Feralaine" and "Fairlane". An exact-token check
rejected every correct take. Score brand names by phonetic similarity over a
sliding window and treat a close match as evidence the name was spoken.

Machine screening cannot approve a take. Pronunciation, cadence and synthetic
artefacts need a human ear. Mark every invented name unveri

…

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [black-forest-labs](https://github.com/black-forest-labs)
- **Source:** [black-forest-labs/skills](https://github.com/black-forest-labs/skills)
- **License:** MIT
- **Homepage:** https://docs.bfl.ai

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-black-forest-labs-skills-flux-3-product-ads
- Seller: https://agentstack.voostack.com/s/black-forest-labs
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
