Install
$ agentstack add skill-gooseworks-ai-goose-video-animated-explainer-ad ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ● Environment & secrets Used
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
animated-explainer-ad
A guided workflow for producing animated explainer ads in the "villain-narrates-own-defeat" format. The skill is conversation-driven: the user gives a free-form prompt ("make an ad for Soteri Skin where Eczema is the villain"), and you walk them through gathering inputs, generating assets phase-by-phase, and composing the master. Each phase has a human review gate so the user catches drift before paying for downstream gen.
When to use this skill
Use when the user asks for any of:
- "Make an animated explainer ad" / "Pixar-style ad" / "cartoon ad" / "villain ad" / "absurdist explainer."
- "Like the Big Chill ad" / "like the Eczema ad" / "in that personified-problem style."
- A paid-social spot where the mechanism (pH/LOCK, adaptogen calming cortisol, dual-action cleanser) is the story.
- A 9:16 sound-off-safe ad with no on-camera talent.
Do NOT use when the user wants: photoreal product video, creator/UGC selfie cam, pure typography/data-viz motion graphics, before-after demos, or live-action footage. Reach for a different skill.
The format
The validated form, derived from two production runs (see references/case-studies.md):
- 9:16, 30–45 seconds, 10–13 scenes. Meta / TikTok / Reels feed-native.
- Bright Pixar/Disney 3D throughout — glossy, rounded, toy-like. Never photoreal.
- Single narrator voice — the personified villain. Same voice for the whole spot including the defeat. ElevenLabs
eleven_v3, ≤4 audio tags. - Roster ≤4 characters — villain + hero + 1 human victim + 1 set piece. More than 4 muddies the cause/effect chain.
- Recurring number/word motif appearing ≥3 times (teach beat, damage beat, payoff). "pH 4.9", "cortisol hormones."
- Climax beat is the product's "fix" moment, not the product appearance. Music resolves on the climax.
- Real product on the end card. PIL-composited from a real photo. Never an AI-rendered cartoon bottle.
- Brand text is never AI-rendered. Wordmark, claims, end-card copy, diegetic labels (
pH 4.9,MOISTURE BARRIER) are PIL or ffmpegdrawtextoverlays. - Burned word-by-word captions via libass.
- Whimsical Pixar/Disney instrumental score — sneaky → tense → uplift @ hero → resolved tail.
Pipeline overview
| Phase | What | Output | Approx cost | Gate? | |---|---|---|---|---| | 1 | Intake: gather brand + concept + script | project/brief.md + project/storyboard.html | $0 | ✅ user approves | | 2 | Character anchors via nano-banana | project/anchors/.png (×4) | ~$0.32 | ✅ user approves grid | | 3 | Voice casting + VO render | project/audio/vo/vo-.mp3 (×N) | ~$0.50 | ✅ user picks voice | | 4 | Scene keyframes via nano-banana with chained anchor refs | project/keyframes/scene-.png | ~$0.66 | ✅ user approves contact sheet | | 5 | i2v scene clips via Seedance Pro | project/clips/scene-.mp4 | ~$15.00 | ✅ user approves preview, then remaining | | 6 | End card: PIL composite real product + typeset brand | project/clips/scene-.mp4 | $0 | — | | 7 | Music bed via ElevenLabs music API | project/audio/music.mp3 | ~$0.20 | — | | 8 | Compose: retime, concat, mix, burn captions | project/finals/master.mp4 | $0 | ✅ user ships it |
Total: ~$15–25 per cut. Wall time ~2 hours including reviews.
Setup (run once per machine)
Before phase 1, confirm the user has:
# 1. Python deps
pip install -r requirements.txt
# 2. API keys in .env (copy from .env.example)
cp .env.example .env
# Edit .env to set FAL_KEY and ELEVENLABS_API_KEY
# 3. ffmpeg installed
ffmpeg -version
If anything's missing, surface it before starting phase 1 — don't proceed with broken tooling.
Phase 1 — Intake (gather brand + concept + script)
Goal: turn a free-form user prompt into a locked brief + storyboard, with a 10–13 scene VO script the user has approved.
Step 1: extract from the user's opening prompt
Parse what they've given you. Typical prompt: "Make an animated explainer ad for [brand] where the villain is [problem] and the hero is [product]. The mechanism is [X]."
Extract whatever's there:
- brand — name (e.g. "Soteri Skin")
- product — exact product + variant (e.g. "Baby Eczema Relief Cream")
- problem — one word ideally (e.g. "Eczema", "Stress", "Dullness")
- mechanism — the ownable claim (e.g. "pH/LOCK at 4.9", "Rhodiola Rosea calms cortisol")
- audience — who's watching (e.g. "parents of babies with eczema")
Step 2: ask for what's missing (one batch of questions)
Use a single batched question to fill gaps. Do not drag the user through a 10-question interview. The minimum you need before drafting the script:
- brand, product, problem, mechanism, audience (above)
- product hero image path — absolute path to a real product photo (≥1000×1000, clean background). Non-negotiable: see Decision Rule "real product on end card."
- brand palette — 2–4 hex colours (primary + accent + CTA). Ask if the user has brand guidelines; otherwise infer reasonable defaults from the product photo and confirm.
- recurring motif — the number or word the script will anchor on (must appear ≥3 times). Often equals the mechanism's key value ("pH 4.9", "30-day reset").
- target duration — default 38s. Range 30–45.
Step 3: propose the character roster
Based on the brand + problem + mechanism, propose 3–4 characters and confirm with the user. Strict cap: ≤4. A graphic prop (a pH meter, a clock face) is NOT a character — don't count it.
The 4-role template:
- Villain = the personified problem. Single narrator. Visual descriptor: ~2 lines (silhouette + skin + face + vibe).
- Hero = the personified product, visually opposite the villain. (Eczema is spiky-red; Soteri is smooth-cream-green.)
- Human victim = the person the user watching identifies with. The baby, the frazzled woman, the stressed founder.
- Set piece = what the villain attacks. The moisture barrier brick wall, the adrenal beans, the collagen scaffold.
If the brand has minions (subordinate creatures spawned by the villain), make them visually smaller copies of the villain — same army, same palette. Big Chill v3/v4 made them visually distinct and the cause/effect chain muddied; v5 collapsed them and it snapped into focus.
Step 4: draft the 12-scene spine
Use this canonical structure (adapt scene counts if the user asks for a shorter cut):
| # | Beat | VO contains | Caption | Diegetic label | |---|---|---|---|---| | 01 | Villain intro | "I'm [villain]. I live [where]." | "I'm [villain]." | — | | 02 | Teach the set piece | Names the mechanism's home (the barrier, the scaffold, the glands) | "Meet the [set piece]." | [SET PIECE NAME] | | 03 | Damage primer | "So I break it" / "I make it work overtime" | "So I break it." | — | | 04 | Reveal the secret | "Here's my little secret — it's really a [mechanism] problem" (optional — skip if mechanism is obvious from teach) | "It's really a [X] problem." | — | | 05 | Teach the number motif | "Healthy [X] sits at [MOTIF]" | "Healthy = [MOTIF]" | [MOTIF] | | 06 | Damage: mechanism breaks | "I push [X] higher and the wall cracks" | "[X] goes wrong → bad thing." | — | | 07 | Damage list item | "And then [bad thing 1]" | "[bad thing 1]." | — | | 08 | Damage peak | "And nobody [sleeps / can focus / feels good] anymore" | "[peak damage]." | — | | 09 | Hero arrives | "[nervous] Until [PRODUCT] shows up" | "Until [PRODUCT] shows up." | — | | 10 | CLIMAX: the fix | "Its [mechanism] snaps it right back to [MOTIF]" | "[MECHANISM] → [MOTIF]" | [MECHANISM] · [MOTIF] | | 11 | Defeat + payoff | "[Calm victim]. And me? Nowhere left to live." | "[Calm state]." | — | | 12 | End card | "That's [PRODUCT]. [Claim 1]. [Claim 2]." | end card type | end card |
The damage list (scenes 6–8) scales 1–4 items based on how many things the villain does. Most brands have 1–3; a 4-item damage list is rare.
Step 5: write the per-scene script
For each scene write:
vo— the spoken line (≤15 words usually; villain delivery is slow). Audio tags[whispers] [menacing] [nervous] [sighs]— ≤4 across the whole script. eleven_v3 inflates pauses around tags by ~400ms each.caption— the burned caption (shorter than the VO; e.g. VO "Healthy baby skin sits slightly acidic — right around 4.9." → caption "Healthy skin = pH 4.9").visual— one sentence describing what's on screen, in present tense, focused on the action, not the static composition.characters— which roster members appear (used to thread the right anchor refs in phase 4).diegetic_label(optional) — composited type that names the mechanism. Brand text NEVER AI-rendered (Decision Rule 7).duration_target_sec— your estimate of how long the scene needs. Refined in phase 3 once VO is rendered.
Step 6: lock the brief
Write everything to project/brief.md — frontmatter + character descriptors + scene table. Generate project/storyboard.html — a single HTML page that renders the scene table with placeholders where keyframes/clips will land. Use the minimal storyboard template provided below (or load the user's references/case-studies.md examples).
Hand the storyboard to the user. Wait for explicit approval of the script + scene table + character descriptors. This is the cheapest gate — script changes here are free; after phase 4 they cost re-rolled keyframes; after phase 5 they cost re-rolled clips.
Once approved, commit project/brief.md and move to phase 2.
Phase 2 — Character anchors
Goal: lock the visual identity of each character before any scene generation. Catches drift early.
For each character in the brief, call:
python scripts/render_anchor.py \
--project project/ \
--name \
--descriptor "" \
--negative "" \
--out project/anchors/.png
The script:
- Wraps the descriptor in the validated Pixar-3D prompt prefix:
"character portrait, neutral background, full body, glossy Pixar/Disney 3D, soft global illumination, shallow depth of field, toy-like, rounded forms". - Hits FAL nano-banana at 1024×1024 portrait.
- Normalises to 1080×1920 with padding.
- Writes a
.meta.jsonnext to the PNG with the exact prompt used.
After all anchors render, assemble a contact sheet (script auto-emits project/anchors/_grid.png) and hand to the user.
Wait for explicit approval. Catch drift here. Common issues:
- Villain reads as creepy/uncanny instead of comedic. Fix: re-render with "comedic, smug, pesky, never scary" in the descriptor.
- Hero looks generic. Fix: add the emblem/cape/distinguishing accent.
- Character has an unintended feature (Soteri's Eczema initially grew bat-wings — a "gremlin" interpretation drift). Fix: add to
--negative.
Re-rolls of a single anchor cost ~$0.08. Don't be precious about re-rendering until the user is happy.
Write project/anchors/LOCKED.json once approved:
{
"characters": {
"eczema": {
"role": "villain",
"anchor": "eczema.png",
"descriptor": "..."
}
},
"scene_threads": {
"scene-01": ["eczema"],
"scene-02": ["eczema", "moisture_barrier"]
}
}
scene_threads is derived from the brief's per-scene characters field. Phase 4 reads this to know which anchor PNGs to pass as chained refs for each scene.
Phase 3 — Voice casting + VO render
Goal: pick a single villain narrator voice the user is happy with, then render all VO lines.
Step 1: voice casting A/B
Shortlist 3–6 candidate ElevenLabs voices. Default shortlist for cartoon-villain:
| Voice | ID | Vibe | |---|---|---| | Austin | Bj9UqZbhQsanLzgalpEG | Texan, raspy, authentic — validated on Soteri | | Dylo | JjsQrIrIBD6TZ656NQfi | Young, fierce — validated on Big Chill | | Tom | mdzEgLpu0FjTwYs5oot0 | Cartoon-villain character voice (rejected on Soteri) | | Adam | pNInz6obpgDQGcFmaJgB | Deep narrative — for slower, more deliberate villains | | Bella | EXAVITQu4vr4xnSDxMaL | Female villain option |
For each candidate, render scene 01:
python scripts/render_vo.py \
--voice-id \
--text "" \
--out project/audio/casting/-scene-01.mp3 \
--speed 1.12
Play all candidates back to the user (just list the file paths — user opens in QuickTime). Wait for them to pick one.
Both reference runs rejected their first voice pick. Don't skip this step.
Step 2: render all VO lines
In the chosen voice, render every scene's VO line:
for scene in :
python scripts/render_vo.py \
--voice-id \
--text "" \
--out project/audio/vo/vo--.mp3 \
--speed 1.12
The script applies silence-trim (silenceremove start=0.05s end=0.12s) per line so they cut cleanly in the compose phase.
Step 3: measure VO durations + update scene timing
After all VO renders:
python scripts/measure_vo.py --project project/
This writes project/scene_timing.json with the actual VO duration per scene. The compose script reads this. If total VO duration > target_duration_sec × 1.15, warn the user and offer two fixes:
- Re-render at speed 1.18–1.20 (fast — single re-render).
- Trim the script (better — chronic speed-up reads as rushed).
The compose-stage atempo=1.3 is available as a last resort but flag it loudly; it's a smell.
Phase 4 — Scene keyframes
Goal: generate one keyframe PNG per scene (excluding the end card) using nano-banana with the locked character anchors as chained refs. This is what catches character drift before you spend on Seedance clip gen.
For each scene 01..N-1:
python scripts/render_keyframe.py \
--project project/ \
--scene \
--visual "" \
--negative "no text, no signage, no labels, no brand names" \
--refs project/anchors/.png project/anchors/.png \
--out project/keyframes/scene-.png
The script:
- Builds a nano-banana prompt:
". Pixar 3D, glossy, rounded, soft global illumination, shallow DOF, toy-like."+ the negative constraints. - Passes the character anchor PNGs as nano-banana
mediasrefs (chained-ref strategy). - Generates at the highest supported portrait resolution.
- Normalises raw output to 1080×1920 (nano sometimes returns 768×1344; the script auto-crops + scales).
- Saves raw to
project/keyframes/_raw/scene-.png, normalised to the main path, meta to.meta.json.
Once all keyframes render, the script emits project/keyframes/_grid.png — a contact sheet of all keyframes at thumbnail size. Hand to user.
Wait for explicit approval of the contact sheet. Common issues:
- Character drift (villain morphs across scenes). Fix: re-roll with a tighter negative constraint.
- Wrong scale (character too small / too large in frame). Fix: add "low-angle shot of " or "wide framing showing small" to the visual.
- Anachronistic background (modern device when the world is cartoon). Fix: add to negative constraints.
Re-rolls cost ~$0.08 per keyframe. Budget 2–4 re-rolls in a 12-scene cut.
Brand text contamination guard: never include the brand name in any keyframe prompt. Nano-banana will hallucinate "SOTERY SKINS" signage in backgrounds. Brand text lives only on the PIL-composited end card (phase 6) and the burned captions (phase 8).
Phase 5 — Scene clips (i2v with preview gate)
Goal: animate each keyframe into a 3–7s scene clip via Seedance Pro i2v. Use a preview gate to catch motion issues before committing the remaining budget.
Step 1: preview gate
Pick 3–4 scenes that cover all characters + the climax (e.g. for a 12-scene cut: scenes 01, 06, 10, 11 — covers all 4 characters and the climax). Generate:
python scripts/render_clip.py \
--project project/ \
--scene \
--keyframe project/keyframes/scene-.png \
--motion-prompt "" \
--duration \
--out project/clips/scene-.mp4
The script:
- Wraps the motion prompt in the validated anti-shake suffix: `"Pixar 3D animation style preserved. NO shake, NO wobble, NO earthquake. Smooth steady camera. Full-frame v
…
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: gooseworks-ai
- Source: gooseworks-ai/goose-video
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.