AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Flux 3 Product Ads

skill-black-forest-labs-skills-flux-3-product-ads · by black-forest-labs

Use when building a finished product ad from FLUX 3 - shot design, voiceover, action-to-word sync, evidence-gated copy, deterministic assembly, and QC gates that catch clipped audio, floating products, off-model plates, and reports that claim a pass the build did not give.

No reviews yet
0 installs
22 views
0.0% view→install

Install

$ agentstack add skill-black-forest-labs-skills-flux-3-product-ads

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-black-forest-labs-skills-flux-3-product-ads)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
29d ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Flux 3 Product Ads? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

FLUX 3 product ads

A finished spot is two jobs, not one. FLUX 3 generates picture and voice. A deterministic pass cuts, times, captions and masters them. Keep the boundary sharp: the model handles what only a model can, and everything a computer can compute exactly stays out of the model's hands.

Route facts this skill depends on live in flux-3-generate. Read it first for /v1/flux-3-video, polling and download behaviour. Continuation behaviour, including what a v2v link actually returns, lives in flux-3-keyframes-continuation.

Everything here has been run at 12, 30 and 60 seconds. Where a rule only holds at one length, it says so. Assume nothing written for a three-shot ten-second spot generalises without being re-tested at length: on the run that produced this revision, two pieces of timing logic that were correct at 10 seconds turned out to have a silent correctness bug and a complexity bug that only appear in a longer read.

Shape of the work

  1. Design shots that survive generation.
  2. Generate picture plates and VO in the same round.
  3. Screen VO by machine, then listen.
  4. Derive timing from the audio.
  5. Assemble deterministically.
  6. Gate on measurements, including the failures that look like passes.

1. Shot design

Three cuts carry a 10-second spot: reveal, proof, payoff. Reveal establishes the object, proof shows it doing the one thing the copy claims, payoff lands the brand.

Longer spots need more beats and more plates. A plate yields roughly 4 to 5 seconds of usable middle once a dissolve is allowed for, so shot count is length divided by about 4.5, rounded up, and a shortfall does not degrade gracefully: it fails the build outright with no legal cut placement. Measured: 12s took 3 plates, 30s took 6, 60s took 11. A 60-second spot planned with 9 plates could not be cut at all.

At 30 seconds and beyond, reveal/proof/payoff no longer fills the time. What worked: hook on the material, the object whole, a detail, the mechanism, a state change, what the product does for you, brand. The structure that matters is that each beat earns its own shot, not that it has three parts.

Anchor motion at its endpoints. A generated clip is trustworthy at its first and last frame and inventive in between. Motion whose midpoint is implied by its endpoints survives; motion that requires the model to invent geometry does not.

Works: a highlight travelling across a surface, a collar rotating through a short arc and stopping, a handle rising and stopping, a lid parting slightly.

Two shots that appear to work and do not: a hand drawing an ink line across paper, and a phone descending into frame and landing on a product. Both were predicted to fail as "travel" and "an object entering the frame". Both arrived looking correct, passed every signal gate, went into masters, and were rejected on sight. The pen's ink ran ahead of the tip while the paper slid underneath; the "phone" was a featureless slab as wide as the pad it landed on. Arrival is not correctness, and this is the trap: the failure mode of a nearly achievable shot is a plate that is wrong in the object rather than broken in the signal, which is invisible to every measurement. See the semantic gate in section 6.

Fails: multiple revolutions, and rotation of two bodies at once. A crank asked for 2 to 3 full revolutions rotated to about 180 degrees, then deformed and vanished as the arm occluded itself. A knurled collar asked to spin two revolutions while the head it sits on tilted produced no motion at all: an inert clip where neither action happened.

The predictor is not travel and not occlusion. It is whether the model must invent geometry it has never seen. A hand crossing frame is a rigid body with a known silhouette. A descending phone is a rigid body. A crank going around 720 degrees has to render its own far side, and that is where it dies. Design against invented geometry, not against movement.

A quarter-turn is not a safe fallback for a rotation shot, it is an invisible one. The safe version of the failed collar shot technically succeeded and was still unusable, because a small rotation of a knurled ring at macro reads as nothing happening. When a rotation shot fails, the shot that replaces it should be light moving across a static object, which is the most reliable motion in the system and the one that consistently looks alive.

When an object must arrive in frame, generate it leaving and reverse the clip. This is the highest-leverage trick in the section, because it converts an invention problem into a translation problem. Three attempts at the same shot, one variable:

| Attempt | Approach | Result | |---|---|---| | 1 | describe a descending phone; keyframe contains no phone | slab as wide as the pad | | 2 | describe the phone in far more detail (two-thirds pad diameter, 8mm thick, dark glass front, metal rails, rounded corners, explicit no-warp instruction); same phone-less keyframe | still a slab, now tilted and overhanging the pad | | 3 | keyframe already contains a correctly proportioned phone; generate the phone lifting away; ffmpeg -vf reverse in post | correct phone, correct scale, constant through the move |

Prompt detail did not fix invention. Removing the invention did. If the object is in the keyframe, the model only has to move something it can already see, and scale cannot drift because it was never chosen by the model. Cost: one filter.

Two things to watch. Author any lighting change in the direction that reads correctly after reversing, and check the source still's camera height against its neighbours, since a still shot for a different purpose may not cut with them.

Reversing moves the action to the other end of the clip, and any timing you already derived is now wrong. This is the cost the trick hides, and it surfaces as a sync failure in a spot that was passing before. The lift plate peaked at 5.79s of a 6.04s clip, a quarter-second before the end. Reversed, that same peak sits at 0.17s. Nothing else changed: same duration, same frame count, same file size class.

That matters because a plate can only be slipped later into its segment, never earlier than its own first frame. So the reachable window for an action collapses to roughly [peak - segment_slack, peak], and an action at 0.17s can only ever land in the first fraction of a second of its segment. The anchor word chosen for the pre-reverse plate, several seconds in, became permanently unreachable, and the build reported it correctly as a clamp:

slip : shot 3 in=0.00s -> action 14.32s vs word 15.62s (-1.30s)
       CLAMPED (reachable 13.00-14.32s, word at 15.62s)

The trap is that this looks like drift. The plates were byte-identical to the run that passed, the VO was untouched, and the durations matched to the sample, so every "did something move?" check comes back clean while the numbers disagree with a note written a day earlier. The thing that moved was inside the file.

Two rules follow. Re-derive anchors after any post step that changes where the action sits in the clip, reversing above all, and treat a reversal as a new plate rather than an edit of the old one. And when a reversed plate needs an early action, anchor it to an early word: a landing that peaks in the first frames belongs on the first word of its line, not the word that named the motion when the clip ran forwards.

Test the prediction rather than trusting it. Give a risky shot two prompts in the same brief, a hard_prompt and a safe prompt, generate both, keep both. One extra job per risky shot is what keeps this section honest as the model improves, and it is how the two "impossible" shots above were found.

Name the grounding in every prompt. Say the product's contact shadow and its reflection explicitly. A product can hold identity and motion perfectly and still look pasted onto the frame because nothing defended its shadow. This passes identity checks and reads as fake instantly to a human.

Camera lock is advisory. "Locked camera, no push-in, no pan, no zoom" still permits parallax and drift. Design shots that tolerate a little movement rather than expecting the prompt to forbid it.

1b. Reference stills come first, and the model gets a vote

Every plate inherits its still, because i2v treats the keyframe as literal. A still that is wrong in a way you can live with poisons every clip generated from it, so the reference pack has to be right before any video job runs.

When the product is real, the pack is derived from the photograph, not written from scratch. This is the easier path and it skips the entire failure class below, because no invariant paragraph has to describe the object well enough for a model to build it. Take the supplied photo as the canonical still, then generate each remaining angle from it with FLUX 2 input_image identity carry, seeded so the pack is reproducible. Write the invariants anyway, by reading them off the photograph: they are what the identity and semantic gates check against later. On a run built this way from one real product photo, three generated angles held every named feature and no still needed regenerating.

Two things this does not buy you. The still is faithful and the video still drifts, so the interior-frame and identity checks below apply unchanged. And a real photo is usually a catalogue shot on seamless white, which gives you no set to cut to: every plate looks like the same photo unless the shots differ in framing and scale, so design the pack for genuinely different crops.

Expect the model to overrule the spec, and read it as information. On a three-product run, two products came back with the model quietly substituting its own design: a light channel specified as unlit rendered lit in every frame, and a lamp specified with a two-segment elbow arm rendered as a single post with a yoke-mounted head in all seven angles.

Both refusals were coherent: one alternative design, held consistently across every angle. That is the signal. A model that disagrees at random gives you noise; a model that disagrees identically seven times is telling you the invariant paragraph describes something it cannot build. Rewrite the spec to match what it reliably builds, then regenerate. Fighting a coherent refusal costs jobs and loses.

Regenerate a still when the error is one the video stage will amplify: stray text on a prop, a wordmark in the wrong place, an object overhanging its base.

Check identity across the pack, not one still at a time. A pack can be clean plate by plate and still be incoherent, because "is this a good photo of a lamp?" is a different question from "is every one of these the same lamp?" On the lamp above, the pack mixed two incompatible designs across its angles; each still looked fine alone, and every plate generated from them inherited whichever design its keyframe happened to carry. Pick one still as canonical, then compare each of the others to it on named, falsifiable features rather than overall impression:

base:   shallow domed profile curving in one arc, vs a flat cylindrical puck
        with a vertical side wall
post:   smooth and unbroken from base to head, vs a collar, ring, knurling
        or joint partway up
head:   plain green cylinder roughly twice as long as wide, vs short/fat
        or a knurled metal barrel
ring:   exactly one knurled ring, at the FRONT of the head encircling the
        lens, vs any knurling on the post or base

Each feature names the correct form and the wrong one it gets confused with, which is what makes a verdict checkable. Run the same comparison over the finished plates too, at least twice per feature, and fold unanimously as with the semantic gate in section 6. On a 14-plate run this returned 13 consistent and one off-model, where a thin gold bezel had changed the head's proportions. The plate it flagged had passed every signal gate. Its hard-motion twin, generated as the safe/risky pair described above, was clean and already on disk, so the fix cost nothing.

This is worth doing before generation and again after, because the two runs answer different questions: the first stops a poisoned pack, the second catches the plate that drifted anyway.

Never let a failed step's stale output become the next step's input. When a chained link failed, the runner picked up the previous day's file of the same name and fed it to the following link, which then succeeded and produced a plausible clip built on the wrong material. A failure that leaves the old file in place is indistinguishable, to the next step, from a success. Check that a chain input is newer than the run that is consuming it, or write links to run-scoped names so a missing file is missing rather than stale.

2. Generate picture and VO together

Finished product ads carry voiceover. Budget it in the first generation round, never bolt it on after picture lock: the VO determines the length of the spot, so generating it last means re-cutting everything.

For voiceover generation, music beds, audio layering, speakability checks, and deterministic audio finishing, use flux-3-audio-dialogue. This skill covers ad-specific timing and assembly: how the VO sets spot length, how paragraphs join at designed pauses, and how the mastered stem feeds the cut pipeline.

Get VO from audio-only jobs: a voice-booth scene whose picture is discarded and whose audio is harvested. Generate at least two scripts against two speaker profiles, because takes are not interchangeable and you want a real choice.

Do not use native picture audio in a finished spot. Generate picture silent.

duration is a tempo control, not a length estimate. The clip comes back at exactly the length you ask for and the read stretches to fill it. The same 27-word script, same speaker, asked for three lengths:

| asked | returned | words/sec | | --- | --- | --- | | 11s | 11.01s | 2.45 | | 14s | 14.00s | 2.11 | | 18s | 18.02s | 1.50 |

The script was delivered completely and correctly every time. Only the pace changed, and at 18s it does not pad with silence, it drags. So pick the speaking rate you want and derive the length: duration = words / rate, with roughly 2.1 to 2.5 words per second reading unhurried and confident. Do not pad the estimate for "room to breathe"; that slows the delivery instead of adding a pause.

One audio job caps at 20 seconds, about 45 words at a good pace. A 30-second spot needs around 75 words and a 60-second spot around 150, so any longer read is generated one paragraph per job and joined in post at a designed pause. This is also better writing, because ads pause. Trim each paragraph to its own speech before joining; stacking booth room tone produces an audible seam.

Generate a music bed the same way whenever the spot runs past about 20 seconds. Silence between paragraphs that is fine at 10 seconds sounds broken at 60. A bed is an audio-only job with no speech in the prompt, looped with a crossfade to length, and ducked under the voice with a sidechain compressor keyed off the VO. Attenuate the bed first and let the duck shape an already-quiet bed: a bed loud enough to need heavy ducking pumps audibly on every phrase.

3. Screen VO by machine, then listen

Transcribe each take and compute word error rate against the script. This catches dropped and mangled lines cheaply.

Invented brand names need phonetic matching, not exact matching. ASR has never seen your product name and will spell it plausibly wrong: "Ferrolane" came back as "Feraline", "Feralaine" and "Fairlane". An exact-token check rejected every correct take. Score brand names by phonetic similarity over a sliding window and treat a close match as evidence the name was spoken.

Machine screening cannot approve a take. Pronunciation, cadence and synthetic artefacts need a human ear. Mark every invented name unveri

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.