# Animated Explainer Ad

> Make a 9:16 ~30–45s animated explainer ad in bright Pixar/Disney 3D where a personified villain (the problem — eczema, stress, dullness, fees) narrates the spot in a single character voice, including their own defeat by the product. Use when the user wants a paid-social ad that teaches a mechanism through cartoon personification, has one ownable claim to anchor on, and has a real product photo fo…

- **Type:** Skill
- **Install:** `agentstack add skill-gooseworks-ai-goose-video-animated-explainer-ad`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [gooseworks-ai](https://agentstack.voostack.com/s/gooseworks-ai)
- **Installs:** 0
- **Category:** [Agent Skills](https://agentstack.voostack.com/c/agent-skills)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [gooseworks-ai](https://github.com/gooseworks-ai)
- **Source:** https://github.com/gooseworks-ai/goose-video/tree/main/skills/templates/animated-explainer-ad

## Install

```sh
agentstack add skill-gooseworks-ai-goose-video-animated-explainer-ad
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# animated-explainer-ad

A guided workflow for producing animated explainer ads in the "villain-narrates-own-defeat" format. The skill is conversation-driven: the user gives a free-form prompt ("make an ad for Soteri Skin where Eczema is the villain"), and you walk them through gathering inputs, generating assets phase-by-phase, and composing the master. Each phase has a human review gate so the user catches drift before paying for downstream gen.

## When to use this skill

Use when the user asks for any of:

- "Make an animated explainer ad" / "Pixar-style ad" / "cartoon ad" / "villain ad" / "absurdist explainer."
- "Like the Big Chill ad" / "like the Eczema ad" / "in that personified-problem style."
- A paid-social spot where the *mechanism* (pH/LOCK, adaptogen calming cortisol, dual-action cleanser) is the story.
- A 9:16 sound-off-safe ad with no on-camera talent.

Do NOT use when the user wants: photoreal product video, creator/UGC selfie cam, pure typography/data-viz motion graphics, before-after demos, or live-action footage. Reach for a different skill.

## The format

The validated form, derived from two production runs (see `references/case-studies.md`):

- **9:16, 30–45 seconds, 10–13 scenes.** Meta / TikTok / Reels feed-native.
- **Bright Pixar/Disney 3D throughout** — glossy, rounded, toy-like. Never photoreal.
- **Single narrator voice** — the personified villain. Same voice for the whole spot including the defeat. ElevenLabs `eleven_v3`, ≤4 audio tags.
- **Roster ≤4 characters** — villain + hero + 1 human victim + 1 set piece. More than 4 muddies the cause/effect chain.
- **Recurring number/word motif** appearing ≥3 times (teach beat, damage beat, payoff). "pH 4.9", "cortisol hormones."
- **Climax beat is the product's "fix" moment**, not the product appearance. Music resolves on the climax.
- **Real product on the end card.** PIL-composited from a real photo. Never an AI-rendered cartoon bottle.
- **Brand text is never AI-rendered.** Wordmark, claims, end-card copy, diegetic labels (`pH 4.9`, `MOISTURE BARRIER`) are PIL or ffmpeg `drawtext` overlays.
- **Burned word-by-word captions** via libass.
- **Whimsical Pixar/Disney instrumental score** — sneaky → tense → uplift @ hero → resolved tail.

## Pipeline overview

| Phase | What | Output | Approx cost | Gate? |
|---|---|---|---|---|
| 1 | Intake: gather brand + concept + script | `project/brief.md` + `project/storyboard.html` | $0 | ✅ user approves |
| 2 | Character anchors via nano-banana | `project/anchors/.png` (×4) | ~$0.32 | ✅ user approves grid |
| 3 | Voice casting + VO render | `project/audio/vo/vo-.mp3` (×N) | ~$0.50 | ✅ user picks voice |
| 4 | Scene keyframes via nano-banana with chained anchor refs | `project/keyframes/scene-.png` | ~$0.66 | ✅ user approves contact sheet |
| 5 | i2v scene clips via Seedance Pro | `project/clips/scene-.mp4` | ~$15.00 | ✅ user approves preview, then remaining |
| 6 | End card: PIL composite real product + typeset brand | `project/clips/scene-.mp4` | $0 | — |
| 7 | Music bed via ElevenLabs music API | `project/audio/music.mp3` | ~$0.20 | — |
| 8 | Compose: retime, concat, mix, burn captions | `project/finals/master.mp4` | $0 | ✅ user ships it |

**Total:** ~$15–25 per cut. Wall time ~2 hours including reviews.

## Setup (run once per machine)

Before phase 1, confirm the user has:

```bash
# 1. Python deps
pip install -r requirements.txt

# 2. API keys in .env (copy from .env.example)
cp .env.example .env
# Edit .env to set FAL_KEY and ELEVENLABS_API_KEY

# 3. ffmpeg installed
ffmpeg -version
```

If anything's missing, surface it before starting phase 1 — don't proceed with broken tooling.

---

## Phase 1 — Intake (gather brand + concept + script)

**Goal:** turn a free-form user prompt into a locked brief + storyboard, with a 10–13 scene VO script the user has approved.

### Step 1: extract from the user's opening prompt

Parse what they've given you. Typical prompt: *"Make an animated explainer ad for [brand] where the villain is [problem] and the hero is [product]. The mechanism is [X]."*

Extract whatever's there:

- **brand** — name (e.g. "Soteri Skin")
- **product** — exact product + variant (e.g. "Baby Eczema Relief Cream")
- **problem** — one word ideally (e.g. "Eczema", "Stress", "Dullness")
- **mechanism** — the ownable claim (e.g. "pH/LOCK at 4.9", "Rhodiola Rosea calms cortisol")
- **audience** — who's watching (e.g. "parents of babies with eczema")

### Step 2: ask for what's missing (one batch of questions)

Use a single batched question to fill gaps. Do **not** drag the user through a 10-question interview. The minimum you need before drafting the script:

- brand, product, problem, mechanism, audience (above)
- **product hero image path** — absolute path to a real product photo (≥1000×1000, clean background). Non-negotiable: see Decision Rule "real product on end card."
- **brand palette** — 2–4 hex colours (primary + accent + CTA). Ask if the user has brand guidelines; otherwise infer reasonable defaults from the product photo and confirm.
- **recurring motif** — the number or word the script will anchor on (must appear ≥3 times). Often equals the mechanism's key value ("pH 4.9", "30-day reset").
- **target duration** — default 38s. Range 30–45.

### Step 3: propose the character roster

Based on the brand + problem + mechanism, propose 3–4 characters and confirm with the user. **Strict cap: ≤4.** A graphic prop (a pH meter, a clock face) is NOT a character — don't count it.

The 4-role template:

1. **Villain** = the personified problem. Single narrator. Visual descriptor: ~2 lines (silhouette + skin + face + vibe).
2. **Hero** = the personified product, visually *opposite* the villain. (Eczema is spiky-red; Soteri is smooth-cream-green.)
3. **Human victim** = the person the user watching identifies with. The baby, the frazzled woman, the stressed founder.
4. **Set piece** = what the villain attacks. The moisture barrier brick wall, the adrenal beans, the collagen scaffold.

If the brand has minions (subordinate creatures spawned by the villain), make them visually *smaller copies of the villain* — same army, same palette. Big Chill v3/v4 made them visually distinct and the cause/effect chain muddied; v5 collapsed them and it snapped into focus.

### Step 4: draft the 12-scene spine

Use this canonical structure (adapt scene counts if the user asks for a shorter cut):

| # | Beat | VO contains | Caption | Diegetic label |
|---|---|---|---|---|
| 01 | Villain intro | "I'm [villain]. I live [where]." | "I'm [villain]." | — |
| 02 | Teach the set piece | Names the mechanism's home (the barrier, the scaffold, the glands) | "Meet the [set piece]." | [SET PIECE NAME] |
| 03 | Damage primer | "So I break it" / "I make it work overtime" | "So I break it." | — |
| 04 | Reveal the secret | "Here's my little secret — it's really a [mechanism] problem" (optional — skip if mechanism is obvious from teach) | "It's really a [X] problem." | — |
| 05 | Teach the number motif | "Healthy [X] sits at [MOTIF]" | "Healthy = [MOTIF]" | [MOTIF] |
| 06 | Damage: mechanism breaks | "I push [X] higher and the wall cracks" | "[X] goes wrong → bad thing." | — |
| 07 | Damage list item | "And then [bad thing 1]" | "[bad thing 1]." | — |
| 08 | Damage peak | "And nobody [sleeps / can focus / feels good] anymore" | "[peak damage]." | — |
| 09 | Hero arrives | "[nervous] Until [PRODUCT] shows up" | "Until [PRODUCT] shows up." | — |
| 10 | **CLIMAX: the fix** | "Its [mechanism] snaps it right back to [MOTIF]" | "[MECHANISM] → [MOTIF]" | [MECHANISM] · [MOTIF] |
| 11 | Defeat + payoff | "[Calm victim]. And me? Nowhere left to live." | "[Calm state]." | — |
| 12 | End card | "That's [PRODUCT]. [Claim 1]. [Claim 2]." | end card type | end card |

The damage list (scenes 6–8) scales 1–4 items based on how many things the villain does. Most brands have 1–3; a 4-item damage list is rare.

### Step 5: write the per-scene script

For each scene write:

- `vo` — the spoken line (≤15 words usually; villain delivery is slow). Audio tags `[whispers] [menacing] [nervous] [sighs]` — **≤4 across the whole script.** eleven_v3 inflates pauses around tags by ~400ms each.
- `caption` — the burned caption (shorter than the VO; e.g. VO "Healthy baby skin sits slightly acidic — right around 4.9." → caption "Healthy skin = pH 4.9").
- `visual` — one sentence describing what's on screen, in present tense, focused on the *action*, not the static composition.
- `characters` — which roster members appear (used to thread the right anchor refs in phase 4).
- `diegetic_label` (optional) — composited type that names the mechanism. Brand text NEVER AI-rendered (Decision Rule 7).
- `duration_target_sec` — your estimate of how long the scene needs. Refined in phase 3 once VO is rendered.

### Step 6: lock the brief

Write everything to `project/brief.md` — frontmatter + character descriptors + scene table. Generate `project/storyboard.html` — a single HTML page that renders the scene table with placeholders where keyframes/clips will land. Use the minimal storyboard template provided below (or load the user's `references/case-studies.md` examples).

**Hand the storyboard to the user. Wait for explicit approval** of the script + scene table + character descriptors. This is the cheapest gate — script changes here are free; after phase 4 they cost re-rolled keyframes; after phase 5 they cost re-rolled clips.

Once approved, commit `project/brief.md` and move to phase 2.

---

## Phase 2 — Character anchors

**Goal:** lock the visual identity of each character before any scene generation. Catches drift early.

For each character in the brief, call:

```bash
python scripts/render_anchor.py \
  --project project/ \
  --name  \
  --descriptor "" \
  --negative "" \
  --out project/anchors/.png
```

The script:
- Wraps the descriptor in the validated Pixar-3D prompt prefix: `"character portrait, neutral background, full body, glossy Pixar/Disney 3D, soft global illumination, shallow depth of field, toy-like, rounded forms"`.
- Hits FAL nano-banana at 1024×1024 portrait.
- Normalises to 1080×1920 with padding.
- Writes a `.meta.json` next to the PNG with the exact prompt used.

After all anchors render, assemble a contact sheet (script auto-emits `project/anchors/_grid.png`) and hand to the user.

**Wait for explicit approval.** Catch drift here. Common issues:
- Villain reads as creepy/uncanny instead of comedic. Fix: re-render with "comedic, smug, pesky, never scary" in the descriptor.
- Hero looks generic. Fix: add the emblem/cape/distinguishing accent.
- Character has an unintended feature (Soteri's Eczema initially grew bat-wings — a "gremlin" interpretation drift). Fix: add to `--negative`.

Re-rolls of a single anchor cost ~$0.08. Don't be precious about re-rendering until the user is happy.

Write `project/anchors/LOCKED.json` once approved:

```json
{
  "characters": {
    "eczema": {
      "role": "villain",
      "anchor": "eczema.png",
      "descriptor": "..."
    }
  },
  "scene_threads": {
    "scene-01": ["eczema"],
    "scene-02": ["eczema", "moisture_barrier"]
  }
}
```

`scene_threads` is derived from the brief's per-scene `characters` field. Phase 4 reads this to know which anchor PNGs to pass as chained refs for each scene.

---

## Phase 3 — Voice casting + VO render

**Goal:** pick a single villain narrator voice the user is happy with, then render all VO lines.

### Step 1: voice casting A/B

Shortlist 3–6 candidate ElevenLabs voices. Default shortlist for cartoon-villain:

| Voice | ID | Vibe |
|---|---|---|
| Austin | `Bj9UqZbhQsanLzgalpEG` | Texan, raspy, authentic — validated on Soteri |
| Dylo | `JjsQrIrIBD6TZ656NQfi` | Young, fierce — validated on Big Chill |
| Tom | `mdzEgLpu0FjTwYs5oot0` | Cartoon-villain character voice (rejected on Soteri) |
| Adam | `pNInz6obpgDQGcFmaJgB` | Deep narrative — for slower, more deliberate villains |
| Bella | `EXAVITQu4vr4xnSDxMaL` | Female villain option |

For each candidate, render scene 01:

```bash
python scripts/render_vo.py \
  --voice-id  \
  --text "" \
  --out project/audio/casting/-scene-01.mp3 \
  --speed 1.12
```

Play all candidates back to the user (just list the file paths — user opens in QuickTime). **Wait for them to pick one.**

Both reference runs rejected their first voice pick. Don't skip this step.

### Step 2: render all VO lines

In the chosen voice, render every scene's VO line:

```bash
for scene in :
  python scripts/render_vo.py \
    --voice-id  \
    --text "" \
    --out project/audio/vo/vo--.mp3 \
    --speed 1.12
```

The script applies silence-trim (`silenceremove start=0.05s end=0.12s`) per line so they cut cleanly in the compose phase.

### Step 3: measure VO durations + update scene timing

After all VO renders:

```bash
python scripts/measure_vo.py --project project/
```

This writes `project/scene_timing.json` with the actual VO duration per scene. The compose script reads this. If total VO duration > `target_duration_sec × 1.15`, warn the user and offer two fixes:

1. **Re-render at speed 1.18–1.20** (fast — single re-render).
2. **Trim the script** (better — chronic speed-up reads as rushed).

The compose-stage `atempo=1.3` is available as a last resort but flag it loudly; it's a smell.

---

## Phase 4 — Scene keyframes

**Goal:** generate one keyframe PNG per scene (excluding the end card) using nano-banana with the locked character anchors as chained refs. This is what catches character drift before you spend on Seedance clip gen.

For each scene 01..N-1:

```bash
python scripts/render_keyframe.py \
  --project project/ \
  --scene  \
  --visual "" \
  --negative "no text, no signage, no labels, no brand names" \
  --refs project/anchors/.png project/anchors/.png \
  --out project/keyframes/scene-.png
```

The script:
- Builds a nano-banana prompt: `". Pixar 3D, glossy, rounded, soft global illumination, shallow DOF, toy-like."` + the negative constraints.
- Passes the character anchor PNGs as nano-banana `medias` refs (chained-ref strategy).
- Generates at the highest supported portrait resolution.
- Normalises raw output to 1080×1920 (nano sometimes returns 768×1344; the script auto-crops + scales).
- Saves raw to `project/keyframes/_raw/scene-.png`, normalised to the main path, meta to `.meta.json`.

Once all keyframes render, the script emits `project/keyframes/_grid.png` — a contact sheet of all keyframes at thumbnail size. Hand to user.

**Wait for explicit approval** of the contact sheet. Common issues:
- Character drift (villain morphs across scenes). Fix: re-roll with a tighter negative constraint.
- Wrong scale (character too small / too large in frame). Fix: add "low-angle shot of " or "wide framing showing  small" to the visual.
- Anachronistic background (modern device when the world is cartoon). Fix: add to negative constraints.

Re-rolls cost ~$0.08 per keyframe. Budget 2–4 re-rolls in a 12-scene cut.

**Brand text contamination guard:** never include the brand name in any keyframe prompt. Nano-banana will hallucinate "SOTERY SKINS" signage in backgrounds. Brand text lives only on the PIL-composited end card (phase 6) and the burned captions (phase 8).

---

## Phase 5 — Scene clips (i2v with preview gate)

**Goal:** animate each keyframe into a 3–7s scene clip via Seedance Pro i2v. Use a preview gate to catch motion issues before committing the remaining budget.

### Step 1: preview gate

Pick 3–4 scenes that cover all characters + the climax (e.g. for a 12-scene cut: scenes 01, 06, 10, 11 — covers all 4 characters and the climax). Generate:

```bash
python scripts/render_clip.py \
  --project project/ \
  --scene  \
  --keyframe project/keyframes/scene-.png \
  --motion-prompt "" \
  --duration  \
  --out project/clips/scene-.mp4
```

The script:
- Wraps the motion prompt in the validated anti-shake suffix: `"Pixar 3D animation style preserved. NO shake, NO wobble, NO earthquake. Smooth steady camera. Full-frame v

…

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [gooseworks-ai](https://github.com/gooseworks-ai)
- **Source:** [gooseworks-ai/goose-video](https://github.com/gooseworks-ai/goose-video)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** yes
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-gooseworks-ai-goose-video-animated-explainer-ad
- Seller: https://agentstack.voostack.com/s/gooseworks-ai
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
